Strategy acquisition method and device, equipment, storage medium and program product
By introducing a combination of noise detection model and reinforcement learning model in reinforcement learning, detecting and considering the impact of noise, the problem of neglecting strategy learning by noise is solved, and the accuracy and adaptability of the optimal strategy are improved.
Patent Information
- Application Number
- CN202510316498.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
AI Technical Summary
In reinforcement learning, when an agent performs actions in the environment to receive rewards, the prior art tends to ignore the impact of noise on strategy learning, resulting in insufficient accuracy in obtaining optimal strategy.
By introducing a combination of the noise detection model and the reinforcement learning model, the noise in the environmental state is detected and supervised learning is performed, the optimal strategy under noise supervision is output, and the optimal strategy under noise-free supervision is combined to determine the final target action strategy.
Improves the accuracy and adaptability of policy acquisition to ensure that the optimal policy can be effectively executed in the presence of noisy environments.
Smart Images

Figure CN120258171A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of reinforcement learning, and particularly relates to a method, device, equipment, storage medium and program product for obtaining a policy. Background Art
[0002] In reinforcement learning, an agent obtains rewards by performing actions in an environment and adjusts its behavior based on these rewards to maximize the long-term cumulative rewards. The agent learns an optimal policy through this process, that is, the best actions to take in different states. It can be seen that after the environmental modeling and reward model are confirmed in reinforcement learning, the overall process is the internal policy convergence action, which easily ignores other factors that affect policy learning, thus affecting the accuracy of obtaining the optimal policy. Summary of the Invention
[0003] Embodiments of this application provide a method, device, equipment, storage medium and program product for obtaining a policy, which can improve the accuracy of obtaining a policy.
[0004] In a first aspect, an embodiment of this application provides a method for obtaining a policy. The method includes:
[0005] Input a first environmental state into a converged target model to obtain a first action policy corresponding to the first environmental state; wherein, the target model includes a noise detection model and a first reinforcement learning model; the noise detection model is used to detect noise in the first environmental state; the first reinforcement learning model updates the action policy based on the first environmental state and the noise detection result of the noise detection model;
[0006] Input the first environmental state into a converged second reinforcement learning model to obtain a second action policy corresponding to the first environmental state;
[0007] Determine a target action policy corresponding to the first environmental state according to the first action policy and the second action policy.
[0008] In a second aspect, an embodiment of this application provides a device for obtaining a policy. The device includes:
[0009] A first acquisition module, configured to input a first environmental state into a converged target model to obtain a first action policy corresponding to the first environmental state; wherein, the target model includes a noise detection model and a first reinforcement learning model; the noise detection model is used to detect noise in the first environmental state; the first reinforcement learning model updates the action policy based on the first environmental state and the noise detection result of the noise detection model;
[0010] A second acquisition module, configured to input the first environmental state into a converged second reinforcement learning model to obtain a second action policy corresponding to the first environmental state;
[0011] A determination module, configured to determine a target action policy corresponding to the first environmental state according to the first action policy and the second action policy.
[0012] In a third aspect, an embodiment of the present application provides a policy acquisition device, including: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the policy acquisition method described in the first aspect is implemented.
[0013] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the policy acquisition method described in the first aspect is implemented.
[0014] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the policy acquisition method described in the first aspect.
[0015] In an embodiment of the present application, on the one hand, the first environmental state can be input into a converged target model to obtain a first action policy corresponding to the first environmental state, where the target model includes a noise detection model and a first reinforcement learning model, and the first reinforcement learning model updates the action policy based on the first environmental state and the noise detection result of the noise detection model; on the other hand, the first environmental state can be input into a converged second reinforcement learning model to obtain a second action policy corresponding to the first environmental state. Compared with the second reinforcement learning model, the target model further integrates the noise detection model and uses the noise detection result of the noise detection model to perform supervised learning on the first reinforcement learning model. Therefore, the first action policy is an action policy learned under noise supervision, considering the influence of noise on the update of the action policy. In this way, the adaptability of the first action policy to the first environmental state can be improved. Thus, finally, according to the first action policy and the second action policy, the target action policy corresponding to the first environmental state is determined, so that the target action policy can be adapted to the first environmental state, thereby improving the acquisition accuracy of the action policy. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0017] Figure 1It is a schematic diagram of route screening provided by an embodiment of the present application;
[0018] Figure 2 It is a flowchart of a strategy acquisition method provided by an embodiment of the present application;
[0019] Figure 3 It is a schematic diagram of the working principle of a target model provided by an embodiment of the present application;
[0020] Figure 4a It is a schematic diagram of the working principle of a first reinforcement model provided by an embodiment of the present application;
[0021] Figure 4b It is a schematic diagram of the structure of a noise detection model provided by an embodiment of the present application;
[0022] Figure 5 It is a structural diagram of a strategy acquisition device provided by an embodiment of the present application;
[0023] Figure 6 It is a structural diagram of a strategy acquisition device provided by an embodiment of the present application. Detailed implementation manners
[0024] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than limiting the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.
[0025] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, elements defined by the statement "comprising..." do not exclude the presence of additional identical elements in the process, method, article or device comprising the said elements.
[0026] In the embodiments of the present application, considering that a policy is used to define the actions that an agent should take in a given environmental state, therefore, the policy can also be referred to as an action policy, that is, in the embodiments of the present application, the policy and the action policy can be mutually replaced.
[0027] In reinforcement learning, noise can affect policy learning. Exemplarily, taking the graph line screening shown as an example, assuming that S16 is also an impassable area, then it is impossible to reach S24 starting from S1. Then, if the noise just acts on S16, then the noise will affect the reward output by the reward model and can directly act on the update of the policy. Figure 1 As shown, taking the graph line screening as an example, assuming that S16 is also an area that cannot be crossed, then it is impossible to reach S24 starting from S1. Then, if the noise just acts on S16, then the noise will affect the reward output by the reward model and can directly act on the update of the policy.
[0028] Based on this, the embodiments of the present application provide a policy acquisition method, which takes into account the influence of noise on policy learning, so that the obtained optimal policy can be adapted to the environmental state, thereby improving the acquisition accuracy of the optimal policy.
[0029] The policy acquisition method of the embodiments of the present application can be applied to terminals, servers, service platforms, the cloud, distributed systems, the Internet of Things, vehicle networking systems, etc. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.
[0030] In addition, the policy acquisition method of the embodiments of the present application can be but is not limited to being applied to the following fields:
[0031] The field of autonomous driving, such as optimizing the driving behavior of vehicles through policy learning.
[0032] The field of robot control, such as enabling robots to autonomously learn how to complete complex tasks, such as grasping objects or navigation.
[0033] The field of games, such as training agents to formulate winning strategies in complex game environments.
[0034] The field of finance, such as optimizing trading strategies to maximize profits by learning historical data and market environments.
[0035] Next, in conjunction with the accompanying drawings, the policy acquisition method provided by the embodiments of the present application will be described in detail through some embodiments and their application scenarios.
[0036] See Figure 2 , Figure 2 which is a flowchart of the policy acquisition method provided by the embodiments of the present application. As Figure 2 shown, the policy acquisition method may include the following steps:
[0037] Step 201: Input the first environmental state into the converged target model to obtain the first action policy corresponding to the first environmental state. The target model includes a noise detection model and a first reinforcement learning model. The noise detection model is used to detect the noise in the first environmental state. The first reinforcement learning model updates the action policy based on the first environmental state and the noise detection result of the noise detection model.
[0038] Step 202: Input the first environmental state into the converged second reinforcement learning model to obtain the second action policy corresponding to the first environmental state.
[0039] Step 203: Determine the target action policy corresponding to the first environmental state according to the first action policy and the second action policy.
[0040] In the embodiment of the present application, the policy acquisition device pre-stores two models, namely: the converged target model; the converged second reinforcement learning model. The converged model can be understood as a trained model or a well-trained model, that is, the model meets the training convergence condition.
[0041] Both the target model and the second reinforcement learning model can be used for policy learning to obtain the optimal policy, but there are differences in the optimal policies obtained by the two, which are specifically described as follows:
[0042] In the embodiment of the present application, for the convenience of distinction, the optimal policy output by the target model for policy learning is called the first action policy, and the optimal policy output by the second reinforcement learning model for policy learning is called the second action policy.
[0043] The target model includes a noise detection model and a first reinforcement learning model. It can first detect the noise in the first environmental state through the noise detection model and output the noise detection result of the first environmental state. The noise detection result output by the noise detection model can specifically be manifested as the detected noise signal (noise_signal, which can also be called the anomaly detection signal or noise label (noise_flag)). The noise signal can indicate the type of detected noise, such as: 2-second human figure noise, that is, the noise is a temporary human figure that appears for 2 seconds and will disappear automatically later. In some embodiments, the noise can be manifested as an obstacle, such as: a suddenly appearing ghost car, ghost person, ghost obstacle, etc. In other embodiments, the noise can also be manifested as an impassable area. In some embodiments, the noise can include but is not limited to at least one of perception noise, noise after perception fusion, noise generated by the modeling environment, and noise that appears during prediction, etc.
[0044] The model structures of the first reinforcement learning model and the second reinforcement learning model in the target model can be the same. However, compared with the second reinforcement learning model, the input of the first reinforcement learning model further includes the noise detection result output by the model detection model, and the noise detection result will act on the policy learning of the first reinforcement learning model. In the target model, the noise detection model can be regarded as the supervision model of the first reinforcement learning model, and the first reinforcement learning model is supervised by noise for supervised learning, so that the first reinforcement learning model considers noise in policy learning, while the second reinforcement learning model has no noise detection network supervision.
[0045] That is to say, in the embodiment of this application, the policy learning of the first reinforcement learning model is policy learning under noisy supervision, and the policy learning of the second reinforcement learning model is policy learning under noiseless supervision. It can be considered that the first reinforcement learning model processes noisy data, and the second reinforcement learning model processes noiseless data.
[0046] Therefore, the optimal policy output by the first reinforcement learning model can be regarded as the optimal policy learned under noisy supervision, considering the influence of noise on policy update; the optimal policy output by the second reinforcement learning model can be regarded as the optimal policy learned under noiseless supervision, without considering the influence of noise on policy update. Furthermore, it can be understood that in the presence of noise, compared with the optimal policy output by the second reinforcement learning model, the optimal policy output by the first reinforcement learning model will be more adaptable to the environment.
[0047] Based on this, in the embodiment of this application, the first environmental state can be input into the target model through step 201, so that the target model outputs the optimal policy (i.e., the first action policy) considering the influence of noise. By executing step 202, the first environmental state is input into the second reinforcement learning model, so that the second reinforcement learning model outputs the optimal policy (i.e., the second action policy) without considering the influence of noise. Then, by executing step 203, the first action policy and the second action policy can be compared to determine the influence of noise on the policy, and further determine the target action policy. The target action policy is regarded as the optimal policy corresponding to the first environmental state, so that the target action policy can be adapted to the first environmental state, thereby improving the acquisition accuracy of the action policy.
[0048] The first environmental state can be understood as the state of the environment at a certain moment (such as the current moment), which can be determined by the data collected by the intelligent perception layer. In some embodiments, the first environmental state can be obtained in the following ways, but this does not limit the acquisition method of the first environmental state:
[0049] First, data collection can be carried out: collect data from sensors such as radars, cameras, and maps. After that, these data can be preprocessed, and the preprocessing can include but is not limited to:
[0050] Time synchronization: Ensure that the data from different sensors are aligned in time;
[0051] Data cleaning: Filter out invalid or missing data;
[0052] Data format conversion: Convert the sensor data into sequence data suitable for processing by a Transformer, which is used to generate an environmental state suitable for processing by a noise detection model and a reinforcement learning model.
[0053] After preprocessing the data, feature extraction can be performed. The embodiments of the present application do not limit the feature extraction methods for the data of each sensor. In some embodiments, as Figure 3 shown, the radar features can be extracted by processing the radar point cloud data through a PointNet network; the camera features can be extracted by processing the camera data using C3D (3nets) feature extraction + iDT classification + linearSVM training; the map features can be extracted by processing the map coordinates + road data using a GNN network.
[0054] After completing the feature extraction, the extracted features can be first fused, and then, as Figure 3 shown, after processing the fused data such as denoising, filtering, coordinate system conversion, etc., vectors and sequences are obtained. Finally, the processed vectors and sequences are input into the Transformer so that the Transformer generates an environmental state suitable for processing by a noise detection model and a reinforcement learning model. In the embodiments of the present application, for the environmental states input into different models, their forms of representation can be different. In one example, as Figure 3 shown, the environmental state input into the reinforcement learning model can be represented as a feature vector, and the environmental state input into the noise detection model can be the data processed by the radar, camera, map, etc.
[0055] For step 201, as Figure 3 shown, the data processed by the radar, camera, map, etc. output by the Transformer can be input into the noise detection model in the target network to detect the noise in these data through the noise detection model, and the noise detection result is input into the first reinforcement learning model in the target model. The feature vector output by the Transformer is input into the first reinforcement learning model, and the first reinforcement learning model updates the action policy based on the noise detection result output by the noise detection model for the feature vector output by the Transformer, and outputs the first action policy.
[0056] In some embodiments, the model can also be referred to as a network. The noise detection model can be referred to as an a network, and the reinforcement learning model can be referred to as a b network.
[0057] It should be noted that the embodiments of the present application do not limit the model structures of the noise detection model and the reinforcement learning model.
[0058] In some embodiments, the noise detection model can be a micro-analysis network model constructed using a Transformer as shown in Figure 4a However, it is not limited thereto. In these embodiments, the noise detection model is a Transformer model constructed with an encoder, a decoder, and a multi-head attention mechanism module.
[0059] The reinforcement learning model can use Q-learning, Deep Q-Network (DQN), Deep Deterministic Policy Gradient (DDPG), or Proximal Policy Optimization (PPO) reinforcement learning algorithms to train the agent, and iteratively back out the optimal policy through the Bellman formula. In some embodiments, the reinforcement learning model can be a network of traditional reinforcement learning (perception-fusion-planning-control network) as shown in Figure 4b It should be noted that Figure 4b the reinforcement learning model shown in Figure 4b is the first reinforcement learning model. In
[0060] For the policy acquisition method of the embodiments of the present application, on the one hand, the first environmental state can be input into the converged target model to obtain the first action policy corresponding to the first environmental state, where the target model includes a noise detection model and a first reinforcement learning model, and the first reinforcement learning model updates the action policy based on the first environmental state and the noise detection result of the noise detection model; on the other hand, the first environmental state can be input into the converged second reinforcement learning model to obtain the second action policy corresponding to the first environmental state. Compared with the second reinforcement learning model, the target model further integrates the noise detection model and uses the noise detection result of the noise detection model to perform supervised learning on the first reinforcement learning model. Therefore, the first action policy is an action policy learned under noise supervision, considering the influence of noise on policy update. In this way, the adaptability of the first action policy to the first environmental state can be improved. Thus, finally, according to the first action policy and the second action policy, the target action policy corresponding to the first environmental state is determined, which can make the target action policy adaptable to the first environmental state, thereby improving the acquisition accuracy of the action policy.
[0061] In some embodiments, inputting the first environmental state into a target model for convergence to obtain a first action policy corresponding to the first environmental state includes:
[0062] Input the first environmental state into a noise detection network, and detect the noise in the first environmental state through the noise detection network to obtain the noise signal of the first environmental state;
[0063] Input the first environmental state and the noise signal into a first reinforcement learning model, and perform a target operation through the first reinforcement learning model to obtain a first action policy:
[0064] Wherein, the target operation includes:
[0065] Select a target action according to the first environmental state and the current action policy;
[0066] Perform the target action in an environment with the added noise signal to obtain the immediate reward, future reward corresponding to the first environmental state, and a second environmental state; the second environmental state is the next environmental state of the first environmental state;
[0067] Determine the first action policy according to the immediate reward, future reward, and the probability of the first environmental state transitioning to the second environmental state.
[0068] It can be understood that in these embodiments, the input of the noise detection network is the environmental state, and the output is the noise signal of the environmental state.
[0069] The main difference between the input of the first reinforcement learning model and the input of the second reinforcement learning model is that: compared with the second reinforcement learning model, the input of the first reinforcement learning model further includes the noise signal of the environmental state, that is, the output of the noise detection network will be used as an input of the first reinforcement learning model. Based on this, the input of the first reinforcement learning model can at least include the environmental state and its noise signal.
[0070] In some embodiments, the input of the first reinforcement learning model can include the environment, the agent, the environmental state, and the noise signal of the environmental state, and the output is the action policy corresponding to the environmental state; in other embodiments, the first reinforcement learning model can be pre-integrated with the environment and the agent, and in these embodiments, the input of the first reinforcement learning model can be the environmental state and the noise signal of the environmental state, which can be specifically determined according to the actual situation, and the embodiments of the present application do not limit this.
[0071] In specific implementation, first input the first environmental state into the noise detection network to obtain the noise signal of the first environmental state; then, the first environmental state and its noise signal can be input into the first reinforcement learning model to enable the agent to perform policy learning to obtain the optimal policy of the first environmental state under the influence of the noise signal of the first environmental state, that is, the first action policy.
[0072] The agent can perform policy learning by executing target operations. First, it can select a target action according to the first environmental state and the current policy, and then execute the target action in an environment with a noise signal added, so as to obtain the immediate reward, future reward, and the next environmental state affected by the noise. Then, the policy can be updated according to the immediate reward and future reward affected by the noise, as well as the state transition probability, to determine the optimal policy.
[0073] In these embodiments, the first environmental state is detected for noise by a noise detection model, and a noise signal of the first environmental state is given. Then, according to the immediate reward and future reward affected by the noise, as well as the state transition probability, the first action policy of the first environmental state is determined. In this way, the first action policy can reflect the influence of the noise on the policy, thereby improving the adaptability of the first action policy to the first environmental state.
[0074] In some embodiments, before inputting the first environmental state into the converged target model to obtain the first action policy corresponding to the first environmental state, the method may further include:
[0075] Using the first training sample set to train the non-converged noise detection model to obtain a converged noise detection model; the first training sample set includes multiple first training samples, and each first training sample includes an environmental state and a noise label of the environmental state;
[0076] Using the second training sample set to train the non-converged first reinforcement learning model to obtain a converged first reinforcement learning model; the second training sample set includes multiple second training samples, and each second training sample includes an environmental state and an action policy label of the environmental state;
[0077] Combining the converged noise detection model and the converged first reinforcement learning model to obtain a non-converged target model;
[0078] Using the third training sample set to train the non-converged target model to obtain a converged target model; the third training sample set includes multiple third training samples, and each third training sample includes an environmental state and an action policy label of the environmental state.
[0079] In these embodiments, to obtain a converged target model, the noise detection model and the first reinforcement learning model in the target model can be separately trained first. After the two are separately converged, they are combined, and then combined training is carried out. In this way, the reliability of policy acquisition of the target model can be improved. Of course, in other embodiments, the two models in the target model can also be directly combined and trained to obtain a converged target model to simplify the training steps of the target model.
[0080] For the above training steps, specifically, for each training sample, the following steps can be executed:
[0081] Input the training sample into the unconverged model to obtain the prediction result corresponding to the training sample;
[0082] Determine the loss function value of the unconverged model according to the prediction result and the corresponding label;
[0083] In the case where the loss function value of the unconverged model does not meet the training convergence condition, adjust the model parameters of the unconverged model to obtain the updated unconverged model, and return to input the training sample into the unconverged model to obtain the prediction result corresponding to the training sample until the training convergence condition is met to obtain the converged model.
[0084] The training of the above models will be specifically described below:
[0085] 1. Training of the noise detection model.
[0086] A loss function can be defined to measure the difference between the model output and the true label. During noise monitoring, the true label can be the noise area or period manually marked during monitoring.
[0087] Data with noise labels can be used to train the noise detection model. Through supervised learning, the model will learn to map the input features to the probability of the presence of noise.
[0088] After training is completed, the noise detection model can be used to infer new collected data. The model will infer the probability of the presence of noise at each time step or in each area. Specifically, a threshold can be set to determine whether noise exists. If the probability output by the model exceeds this threshold, it is considered that noise exists at this time step or area. After that, the noise signal can be fed back to the first reinforcement learning model for further analysis.
[0089] 2. Training of the first reinforcement learning model.
[0090] The first reinforcement learning model can use Q-Learning or DQN or DDPG or PPO reinforcement learning algorithms to train the agent, and iteratively derive the optimal planning strategy backward through the Bellman formula.
[0091] The training steps of the first reinforcement learning model can include:
[0092] 1) Goal: Obtain the state value function V(s) and action value function q(s,a) under the optimal policy.
[0093] Derivation of the Bellman formula
[0094] Then:
[0095] Among them, in the stochastic policy, π(a│s) represents the probability of selecting action a in state s, while in the deterministic policy, π(a│s) directly gives the action a that the agent will take in state s. p(r|a,s) is the probability of obtaining reward r by taking action a in state s, p(s′|s,a) is the probability of the state s transitioning to state s′ by taking action a, and γ is the discount factor, representing the current value of future rewards.
[0096] Future rewards can be obtained through Monte Carlo sampling, which can be specifically implemented through the following steps:
[0097] Sampling: {x1, x2, x3, …, x N}.
[0098] Estimation:
[0099] Result: As N increases, the expected E will gradually approach the actual value.
[0100] Law of large numbers: When N approaches infinity,
[0101] The unbiased estimate that approaches approaches 0, and when the variance approaches 0, It approaches a constant.
[0102] 2) Using the policy iteration algorithm:
[0103] 1. Initialization: Given an initial policy π0.
[0104] 2. Policy Evaluation (PE): According to the current policy πk and the Bellman formula, calculate the state value function vπk(s). It can be implemented through an iterative method, such as vk+1 = rπk + γPπkvk, until vk converges to vπk. In some embodiments, Policy Evaluation can also be expressed as:
[0105]
[0106] 3. Policy Improvement (PI): According to the current state value function vπk, update the policy πk+1 such that for all states s and actions a, πk+1(a|s) = 1 if and only if a = argmaxaqπk(s,a), that is, select the action a that maximizes qπk(s,a) as the new policy. In some embodiments, policy improvement can be implemented through the following formula:
[0107]
[0108]
[0109] 4. Iteration: Repeat steps 2 and 3 until the policy πk+1 no longer changes, at which point the optimal policy π* is found.
[0110] 5. According to the contraction mapping theorem, the iteration process will eventually converge to a unique fixed point. At this time, the optimal policy is found, and the optimal state value function v*(s) and the optimal policy π* are obtained.
[0111] In the above formula, Vπ(s) is the value function of the current policy; π(a|s) is the probability distribution of the policy output action given a state; R(s,a) is the reward of the action in the current state; P(s'|s,a) is the state transition probability; γ is the discount factor.
[0112] III. For the combined training of the noise detection model and the first reinforcement learning model.
[0113] 1. Initialization:
[0114] ● Define the state space s, action space a, reward function R(s,a), policy π(a|s), and parameters required for the learning algorithm (such as learning rate, discount factor, etc.).
[0115] ● Set the noise detection flag noise_flag to record whether a noise signal is detected.
[0116] ● Initialize the experience storage (such as the experience replay buffer).
[0117] 2. Training loop:
[0118] For each training episode:
[0119] a. Initialize the state s as the starting state of the environment.
[0120] b. Loop until the episode ends (reaching the termination condition):
[0121] ● Select an action a according to the current policy π(a|s).
[0122] ● Execute the action a and observe the results (the next state s', immediate reward r, whether it ends done, and the noise signal noise_signal passed by the external model).
[0123] ● If the noise_signal indicates the existence of a 2-second human figure noise, set noise_flag = True.
[0124] (Assume the noise is a temporary human figure that appears for 2 seconds and then disappears automatically.)
[0125] ● Store the experience tuple (s, a, r, s', noise_flag) in the experience storage.
[0126] ● Analyze the change of the immediate reward r and the future reward according to the value of noise_flag. Specifically, if noise_flag = True, record the immediate reward and the future reward under the influence of noise, and calculate the difference from the case without noise.
[0127] ● Update the policy π(a|s) using a learning algorithm (such as Q-learning, deep Q-network, policy gradient, etc.), considering the influence of noise on the reward (for example, by adjusting the learning rate, adding a regularization term, or using a robust optimization method).
[0128] ● Update the state s = s', and prepare for the next iteration.
[0129] It can be understood that the training of the second reinforcement learning model is similar to that of the first reinforcement learning model. For details, please refer to the training of the first reinforcement learning model, which will not be elaborated here.
[0130] The following specifically describes the acquisition of the target action policy.
[0131] In some embodiments, according to the first action policy and the second action policy, determining the target action policy corresponding to the first environmental state includes:
[0132] Determine the degree of interference of the noise on the action policy according to the first action policy and the second action policy;
[0133] In the case where the noise interferes with the action policy, determine the target action policy corresponding to the first environmental state according to the comparison result of the deviation value and the preset threshold; wherein, the deviation value is the deviation value between the first action policy and the second action policy;
[0134] In the case where the noise does not interfere with the action policy, determine the first action policy or the second action policy as the target action policy.
[0135] In these embodiments, the first action policy and the second action policy can be compared first to determine the influence of the noise on the action policy, that is, to determine the degree of interference of the noise on the action policy, and then determine the target action policy based on the determination result.
[0136] Specifically, if the first action policy and the second action policy are the same, it can be determined that the noise does not interfere with / affect the action policy. In this case, the first action policy or the second action policy can be directly used.
[0137] If the first action strategy and the second action strategy are different, it can be determined that noise interferes with / affects the action strategy. In this case, the deviation value between the two action strategies can be further compared with a preset threshold to determine the target action strategy, where the preset threshold is the lowest threshold representing the safe execution of the strategy under the influence of noise and can be specifically set according to actual requirements, which is not limited in the embodiments of the present application.
[0138] Further, determining the target action strategy corresponding to the first environmental state according to the comparison result between the deviation value and the preset threshold may include:
[0139] When the deviation value is less than or equal to the preset threshold, the second action strategy is determined as the target action strategy;
[0140] When the deviation value is greater than the preset threshold, the first action strategy is determined as the target action strategy.
[0141] If the deviation value between the two action strategies is less than or equal to the preset threshold, it indicates that the influence of noise on the strategy does not affect the safe execution of the strategy, and the execution of the second action strategy is within the safe range. The noise in the environmental state can be ignored, and the second action strategy is selected as the target action strategy. If the deviation value between the two action strategies is greater than the preset threshold, it indicates that the influence of noise on the strategy will affect the safe execution of the strategy, and the execution of the second action strategy is not within the safe range. The noise in the environmental state cannot be ignored, and the first action strategy is selected as the target action strategy.
[0142] By determining the target strategy as described above, the safe execution of the target strategy can be ensured, thereby improving the reliability of determining the target strategy.
[0143] In some other embodiments, after obtaining the first action strategy and the second action strategy, the deviation value between the two action strategies can be directly compared with the preset threshold to determine the target strategy. In this way, the determination of the target strategy can be simplified and the determination efficiency of the target strategy can be improved.
[0144] The embodiments of the present application do not limit the calculation method of the deviation value between the two action strategies. In some embodiments, after inputting the first environmental state into the converged second reinforcement learning model to obtain the second action strategy corresponding to the first environmental state, before determining the target action strategy corresponding to the first environmental state according to the comparison result between the deviation value and the preset threshold, the method further includes:
[0145] Determine the deviation value according to the objective functions in the first action strategy and the second action strategy;
[0146] Wherein, the objective function includes at least one of the following: optimal state value function; optimal action value function; reward function.
[0147] In the embodiments of the present application, the optimal strategy may include: an optimal state value function and an optimal action value function, and / or, a reward function, where the reward function can represent the maximum reward that can be obtained by accumulating all actions under the strategy.
[0148] Based on this, in these embodiments, the deviation value between two action strategies can be calculated according to the values of the optimal state value function, the optimal action value function, and / or the reward function. When specifically implementing, when calculating the deviation value according to one of the above functions, the difference between the function values of the two action strategies can be directly used as the deviation value between the two action strategies. When calculating the deviation value according to multiple of the above functions, the difference between the values of the same function in the two action strategies can be calculated first, and then these differences are weighted and summed to obtain the deviation value between the two action strategies. The weight values corresponding to each function can be equal or unequal, and can be specifically set based on actual needs.
[0149] Since the above functions are all related to rewards, and noise will affect rewards, therefore, by determining the deviation value between two action strategies through the above functions, the influence degree of noise on the strategy can be accurately reflected. In this way, by determining the target action strategy based on the deviation value between two action strategies determined in this manner, the reliability of determining the target action strategy can be improved.
[0150] It should be noted that the various optional embodiments or implementation manners introduced in the embodiments of the present application can be combined with each other or implemented separately without conflict, and the embodiments of the present application do not make any limitations in this regard.
[0151] Based on the strategy acquisition method provided in the above embodiments, correspondingly, the present application also provides a specific implementation manner of a strategy acquisition device. Please refer to the following embodiments.
[0152] See Figure 5 , the strategy acquisition device provided in the embodiments of the present application may include:
[0153] A first acquisition module 501, configured to input a first environmental state into a converged target model to obtain a first action strategy corresponding to the first environmental state; wherein, the target model includes a noise detection model and a first reinforcement learning model; the noise detection model is used to detect the noise in the first environmental state; the first reinforcement learning model updates the action strategy based on the first environmental state and the noise detection result of the noise detection model;
[0154] A second acquisition module 502, configured to input the first environmental state into a converged second reinforcement learning model to obtain a second action strategy corresponding to the first environmental state;
[0155] A determination module 503, configured to determine a target action strategy corresponding to the first environmental state according to the first action strategy and the second action strategy.
[0156] The policy acquisition device provided by the embodiment of the present application can implement each process in the method embodiment. To avoid repetition, it will not be elaborated here.
[0157] Figure 6 The figure shows a schematic hardware structure diagram of policy acquisition provided by the embodiment of the present application.
[0158] In the policy acquisition device, it may include a processor 601 and a memory 602 storing computer program instructions.
[0159] Specifically, the above-mentioned processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0160] The memory 602 may include a mass storage for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 602 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 602 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 602 is a non-volatile solid state memory.
[0161] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0162] The processor 601 reads and executes the computer program instructions stored in the memory 602 to implement any one of the policy acquisition methods in the above embodiments.
[0163] In one example, the policy acquisition device may further include a communication interface 606 and a bus 610. Among them, as Figure 6As shown, a processor 601, a memory 602, and a communication interface 606 are connected via a bus 610 and communicate with each other.
[0164] The communication interface 606 is mainly used to implement communication between various modules, devices, units, and / or equipment in the embodiments of the present application. The bus 610 includes hardware, software, or both, and couples the components of the policy acquisition device to each other. By way of example and not limitation,
[0165] the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 610 may include one or more buses. Although the embodiments of the present application describe and illustrate a specific bus, the present application contemplates any suitable bus or interconnect.
[0166] In addition, in combination with the insulation resistance detection method in the above embodiments, the embodiments of the present application may be implemented by providing a computer storage medium. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the insulation resistance detection methods in the above embodiments is implemented.
[0167] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0168] The functional blocks shown in the above block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs, floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.
[0169] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, can be different from the order in the embodiments, or several steps can be executed simultaneously.
[0170] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It should also be understood that each block in the block diagrams and / or flowcharts, and the combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0171] As described above, the above is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present application.
Claims
1. A strategy acquisition method, characterized in that, Including: Input the first environmental state into a converged target model to obtain a first action policy corresponding to the first environmental state; wherein, the target model includes a noise detection model and a first reinforcement learning model; the noise detection model is used to detect noise in the first environmental state; the first reinforcement learning model updates the action policy based on the first environmental state and the noise detection result of the noise detection model; Input the first environmental state into a converged second reinforcement learning model to obtain a second action policy corresponding to the first environmental state; Determine a target action policy corresponding to the first environmental state according to the first action policy and the second action policy.
2. The method according to claim 1, wherein The step of inputting the first environmental state into a converged target model to obtain a first action policy corresponding to the first environmental state includes: Input the first environmental state into the noise detection network, and detect the noise in the first environmental state through the noise detection network to obtain a noise signal of the first environmental state; Input the first environmental state and the noise signal into the first reinforcement learning model, and execute a target operation through the first reinforcement learning model to obtain the first action policy: Wherein, the target operation includes: Select a target action according to the first environmental state and the current action policy; Execute the target action in an environment with the added noise signal to obtain an immediate reward, a future reward, and a second environmental state corresponding to the first environmental state; the second environmental state is the next environmental state of the first environmental state; Determine the first action policy according to the immediate reward, the future reward, and the probability of the first environmental state transitioning to the second environmental state.
3. The method according to claim 1, characterized in that Before the step of inputting the first environmental state into a converged target model to obtain a first action policy corresponding to the first environmental state, the method further includes: Train an unconverged noise detection model using a first training sample set to obtain a converged noise detection model; the first training sample set includes multiple first training samples, and each first training sample includes an environmental state and a noise label of the environmental state; Train an unconverged first reinforcement learning model using a second training sample set to obtain a converged first reinforcement learning model; the second training sample set includes multiple second training samples, and each second training sample includes an environmental state and an action policy label of the environmental state; Combine the converged noise detection model and the converged first reinforcement learning model to obtain an initial model; Train the initial model using a third training sample set to obtain the converged target model; the third training sample set includes multiple third training samples, and each third training sample includes an environmental state and an action policy label of the environmental state.
4. The method according to claim 1, characterized in that The step of determining a target action policy corresponding to the first environmental state according to the first action policy and the second action policy includes: Determine the interference of the noise on the action policy according to the first action policy and the second action policy; In the case where the noise interferes with the action policy, determine the target action policy corresponding to the first environmental state according to the comparison result between the deviation value and the preset threshold; wherein, the deviation value is the deviation value between the first action policy and the second action policy; In the case where the noise does not interfere with the action policy, determine the first action policy or the second action policy as the target action policy.
5. The method according to claim 4, characterized in that, The determining the target action policy corresponding to the first environmental state according to the comparison result between the deviation value and the preset threshold includes: In the case where the deviation value is less than or equal to the preset threshold, determine the second action policy as the target action policy; In the case where the deviation value is greater than the preset threshold, determine the first action policy as the target action policy.
6. The method according to claim 4, wherein After obtaining the second action policy corresponding to the first environmental state by inputting the first environmental state into the converged second reinforcement learning model, and before determining the target action policy corresponding to the first environmental state according to the comparison result between the deviation value and the preset threshold, the method further includes: Determine the deviation value according to the objective function in the first action policy and the second action policy; Wherein, the objective function includes at least one of the following: optimal state value function; optimal action value function; reward function.
7. A strategy acquisition device, characterized in that, The apparatus includes: A first acquisition module, configured to input a first environmental state into a converged target model to obtain a first action policy corresponding to the first environmental state; wherein, the target model includes a noise detection model and a first reinforcement learning model; the noise detection model is configured to detect noise in the first environmental state; the first reinforcement learning model updates the action policy based on the first environmental state and the noise detection result of the noise detection model; A second acquisition module, configured to input the first environmental state into a converged second reinforcement learning model to obtain a second action policy corresponding to the first environmental state; A determination module, configured to determine the target action policy corresponding to the first environmental state according to the first action policy and the second action policy.
8. A strategy acquisition device, characterized in that The device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the policy acquisition method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the policy acquisition method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that, When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the policy acquisition method according to any one of claims 1 to 6.