Multi-aircraft cooperative non-cooperative target capturing method based on distributed reinforcement learning algorithm

By using a distributed reinforcement learning algorithm, a policy network is set up for each UAV. Combined with priority experience replay and temporal difference method, the problems of high interception control accuracy and high training cost in traditional methods are solved, and efficient and stable multi-UAV cooperative interception is achieved.

CN121028845APending Publication Date: 2025-11-28BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511058718.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Traditional multi-target cooperative interception methods struggle to achieve precise guidance and control under complex and highly dynamic conditions. Furthermore, existing learning algorithms suffer from high training costs and difficulty in obtaining samples, resulting in poor interception stability.

Method used

By employing a distributed reinforcement learning algorithm, a policy network is set up for each UAV, and combined with priority experience replay and temporal difference method, the cooperative non-cooperative target acquisition of multiple UAVs is achieved by utilizing relative position and velocity information.

Benefits of technology

It improves the efficiency and accuracy of drone interception, reduces training time, and enhances interception stability and data utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121028845A_ABST
    Figure CN121028845A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-aircraft cooperative non-cooperative target capturing method based on a distributed reinforcement learning algorithm, and the method comprises the following steps: setting a simulation environment in which an unmanned plane group intercepts a target, obtaining flight information, and storing the flight information in an experience pool; data are sampled from the experience pool through a strategy network, a maneuvering strategy is generated, and the unmanned aerial vehicle performs maneuvering according to the maneuvering strategy; evaluating the value of the maneuvering strategy through an evaluation network, and storing the value into an experience pool; sampling data from the experience pool, and updating parameters of the strategy network; the strategy model is obtained through multiple times of updating, the strategy model is arranged on the unmanned aerial vehicle, the unmanned aerial vehicle flies according to a maneuvering strategy output by the strategy model, and target interception is achieved. According to the multi-aircraft cooperative non-cooperative target capturing method based on the distributed reinforcement learning algorithm provided by the invention, the training efficiency is improved, and the obtained maneuvering strategy is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a multi-aircraft cooperative non-cooperative target capturing method based on a distributed reinforcement learning algorithm and belongs to the technical field of guidance. BACKGROUND

[0002] The existing multi-target cooperative interception method mainly includes target assignment and cooperative guidance interception. The target assignment divides the cluster interception task into multiple one-to-one interception problems, mainly using hidden enumeration method, dynamic programming method, genetic algorithm, ant colony algorithm and particle swarm optimization algorithm. The cooperative guidance interception mainly adopts four types of methods, namely formation cooperative attack method, attack time cooperative method, angle constraint cooperative method and intelligent cooperative interception method. The formation cooperative method is divided into directed communication topology and undirected communication topology according to the communication mode, the attack time cooperative method is mainly divided into specified time attack and multi-aircraft simultaneous arrival, and the angle constraint cooperative method mainly includes landing constraint and line of sight angle convergence constraint.

[0003] However, the traditional cooperative guidance interception part has complex calculation dimension, and it is difficult to obtain accurate guidance control under the conditions of complexity, diversity and high dynamic characteristics, and the traditional cooperative guidance interception is mainly for stationary or slow-moving inanimate targets, and is not suitable for cluster penetration.

[0004] There is also a method for realizing cooperative interception based on a learning algorithm in the prior art, however, these methods require a large number of learning samples, have high training cost, and it is difficult to obtain samples, which is difficult to guarantee the sample coverage, resulting in long training time and poor interception stability after training.

[0005] Therefore, it is necessary to further study the multi-target cooperative interception method to solve the above problems. SUMMARY

[0006] In order to overcome the above problems, the present application provides a multi-aircraft cooperative non-cooperative target capturing method based on a distributed reinforcement learning algorithm, which comprises the following steps:

[0007] S1, setting a simulation environment for the unmanned aerial vehicle group to intercept the target, obtaining the flight information of each unmanned aerial vehicle in the unmanned aerial vehicle group, the flight information of other unmanned aerial vehicles observed by the unmanned aerial vehicle in the cluster, and the target flight information observed by the unmanned aerial vehicle, and storing the flight information into an experience pool;

[0008] S2, sampling data from the experience pool through a policy network to generate a maneuvering strategy, and the unmanned aerial vehicle performs a maneuvering action according to the maneuvering strategy, and obtains the flight information of the unmanned aerial vehicle and the target flight information observed by the unmanned aerial vehicle in the maneuvering action process, and stores the flight information and the maneuvering action value into the experience pool;

[0009] S3. Evaluate the value of the maneuver strategy through the evaluation network and store the value in the experience pool;

[0010] S4. Sample data from the experience pool and update the parameters of the policy network;

[0011] S5. Repeat steps S2 to S4 to obtain the strategy model. Set the strategy model on the UAV. The UAV flies according to the maneuver strategy output by the strategy model to intercept the target.

[0012] In a preferred embodiment, the flight information includes position, speed, and line-of-sight angular velocity.

[0013] In a preferred embodiment, a separate policy network is set up for each UAV, and multiple policy networks update parameters simultaneously.

[0014] In a preferred embodiment, the maneuver actions corresponding to the maneuver policies generated by multiple policy networks are stored in the same experience pool.

[0015] In a preferred embodiment, during the policy network update process in S4, a priority experience replay method is used.

[0016] In a preferred embodiment, in S4, the evaluation network is also updated.

[0017] In a preferred embodiment, the optimal value of the evaluation network is obtained using the time difference method during the evaluation network update process.

[0018] In the time difference method, the TD error is obtained by minimizing the distribution of the value function of the observed state and the next state.

[0019] In a preferred embodiment, the value in the evaluation network is set as follows:

[0020]

[0021] Among them, Y t Let γ represent the reinforcement learning value function, n represent the number of samplings, N represent the number of simulation steps, γ represent the decay factor, and r represent the reward value. t+n Let Q represent the reward value at time t+n. θ′ s represents the action value of the UAV when the parameter is θ′. t+n μ represents the state of the drone at time t+n. θ′ This represents the strategy function adopted by the UAV when the parameter is θ′.

[0022] In a preferred embodiment, a priority experience replay area is set in the experience pool. In S2 to S4, when sampling data from the experience pool, the sampling probability of data in the priority experience replay area is different from the sampling probability of data in other areas.

[0023] In a preferred embodiment, in S4, the parameter update of the policy network is expressed as:

[0024] ω←ω+β t δ ω

[0025]

[0026] Where ω is the parameter of the policy network, β t M represents the learning rate of the policy network, M represents the data transfer batch size, and i represents the sampling step number. This represents the gradient of the loss function with respect to the policy network parameters. Y represents the importance sampling weight at step i. i Let represent the reinforcement learning value function at step i.

[0027] The beneficial effects of this invention include:

[0028] (1) Applying distributed technology to reinforcement learning solves the problem of slow training convergence of traditional reinforcement learning algorithms and greatly improves efficiency.

[0029] (2) By adopting the N-step return technique, the sum of N-step rewards is combined with the reward calculation, which can more accurately predict future rewards and improve the training efficiency of the value network.

[0030] (3) Integrating the priority experience replay technology into reinforcement learning improves the efficiency of data utilization in the reinforcement learning exploration and optimization process, which is beneficial to the training of UAV policy networks and increases the training speed. Attached Figure Description

[0031] Figure 1 A flowchart of a multi-aircraft cooperative non-cooperative target acquisition method based on a distributed reinforcement learning algorithm according to a preferred embodiment of the present invention is shown.

[0032] Figure 2 The position curves during the UAV's interception of the target aircraft in Example 1 are shown;

[0033] Figure 3 The normal overload curve of the UAV in Example 1 is shown;

[0034] Figure 4 The axial overload curve of the UAV in Example 1 is shown;

[0035] Figure 5The flight speed curve of the UAV in Example 1 is shown;

[0036] Figure 6 The heading angle curve of the UAV in Example 1 is shown. Detailed Implementation

[0037] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Through these descriptions, the features and advantages of the present invention will become clearer and more apparent.

[0038] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments. Although various aspects of embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless specifically indicated otherwise.

[0039] A multi-aircraft cooperative non-cooperative target acquisition method based on a distributed reinforcement learning algorithm, provided by the present invention, includes the following steps:

[0040] S1. Set up a simulation environment for the drone swarm to intercept the target, obtain the flight information of each drone in the drone swarm, the flight information of other drones in the swarm observed by the drone, and the target flight information observed by the drone, and store the flight information in the experience pool.

[0041] S2. Data is sampled from the experience pool through the policy network to generate a maneuvering policy. The UAV performs maneuvering actions according to the maneuvering policy, and the flight information of the UAV and the target flight information observed by the UAV during the maneuvering action are obtained. The flight information and maneuvering action values ​​are stored in the experience pool.

[0042] S3. Evaluate the value of the maneuver strategy through the evaluation network and store the value in the experience pool;

[0043] S4. Sample data from the experience pool and update the parameters of the policy network;

[0044] S5. Repeat steps S2 to S4 to obtain the strategy model. Set the strategy model on the UAV. The UAV flies according to the maneuver strategy output by the strategy model to intercept the target.

[0045] Preferably, in the simulation environment, the motion model of the UAV is set as follows:

[0046]

[0047] Where x and y are the plane coordinates of the aircraft, v is the speed of the aircraft, and γ is the angle between the direction of the aircraft's velocity and the x-axis, i.e., the trajectory deflection angle.

[0048] Furthermore, the flight information includes position, speed, and line-of-sight angular velocity.

[0049] Preferably, after acquiring flight information, the relative position and relative speed between the observed target and the observing UAV, as well as the relative position and relative speed between other UAVs and the observing UAV, are also obtained and stored in the experience pool as a single data entry along with the flight information.

[0050] In this invention, by adding relative position and relative velocity information as prediction conditions for the policy network, the accuracy of the output maneuver strategy of the policy network is improved.

[0051] Policy networks and evaluation networks are commonly used networks in reinforcement learning. Their specific structures are not particularly limited in this invention, and those skilled in the art can freely choose according to actual needs.

[0052] Furthermore, similar to traditional reinforcement learning, the parameters of the policy network are updated using a gradient algorithm.

[0053] Unlike traditional reinforcement learning, this invention sets up a separate policy network for each UAV, and multiple policy networks update parameters simultaneously.

[0054] Traditional reinforcement learning uses only one policy network, and updates are performed on this single policy network. The conventional approach involves setting up a policy network and an evaluation network for a single drone, learning evaluation and rewards based on the drone's maneuvers to determine the final policy network, and then distributing this final policy network across different drones. In this invention, a separate policy network is set up for each drone, and multiple policy networks simultaneously perform step S2. This distributes the data sampling process across multiple parallel policy networks, thereby achieving efficient data sampling and utilization.

[0055] Furthermore, in order to enable multiple policy networks to execute in parallel, the maneuver actions corresponding to the maneuver policies generated by multiple policy networks are stored in the same experience pool, and the same evaluation network is used to evaluate the multiple policy networks. Based on the evaluation values, the parameters of the multiple policy networks are updated simultaneously.

[0056] In a preferred embodiment, the initial values ​​of the parameters and the initial value of the learning rate are the same for different policy networks.

[0057] In a preferred embodiment, during the policy network update process in S4, a priority experience replay method is used, so that high-quality data can be repeatedly used to update network parameters.

[0058] Prioritized experience replay is a commonly used method in reinforcement learning. It can give higher sampling weights to samples with high learning efficiency, thereby improving learning efficiency. For details, please refer to the paper Prioritized Experience Replay by David Silver's group at ICLR 2016. It will not be elaborated in this invention.

[0059] Furthermore, according to the present invention, the evaluation network is also updated in S4.

[0060] Preferably, during the evaluation network update process, the optimal value of the evaluation network is obtained using the temporal difference method, which is a commonly used method in reinforcement learning.

[0061] Unlike the traditional time difference method, in this invention, the TD error of the time difference method is obtained by minimizing the distribution of the value function of the observed state and the next state. Compared with the traditional time difference method, this improvement generates more stable input data for evaluating network updates.

[0062] More preferably, in the time-difference method, the value of the maneuver strategy evaluated by the evaluation network is expressed as:

[0063] Q w (x,a)=E(Z w (x,a))

[0064] The value function maps the state-action information in the experience pool to a distribution on the weight space, where w represents the weight space, s represents the current state of the UAV, a represents the UAV's maneuver, and Q represents the UAV's action. w (s,a) represents the action-value function, E() represents the expected operation, and Z... w (s,a) represents the distribution of the value function.

[0065] The function that minimizes the observed state and the next state is expressed as:

[0066]

[0067] Where L(w) represents the loss function for the value network weights, and d() represents the distance metric between distributions. Z represents the distributed Bellman operator. w′ (s,a) represents the value function corresponding to the target neural network.

[0068] In a preferred embodiment, the value in the evaluation network is set as follows:

[0069]

[0070] Among them, Y t Let γ represent the reinforcement learning value function, n represent the number of samplings, N represent the total number of simulation steps, γ represent the decay factor, and r represent the reward value. t+n Let Q represent the reward value at time t+n. θ′ s represents the action value of the UAV when the parameter is θ′. t+n μ represents the state of the drone at time t+n. θ′ This represents the strategy function adopted by the UAV when the parameter is θ′.

[0071] In this invention, by combining TD update with n-step regression technology through the above settings, n-step regression can generate a reward for the next n steps when calculating the error, thereby more accurately estimating the state action.

[0072] In traditional reinforcement learning training, parameters are typically updated by uniformly sampling data from an experience pool. This invention improves upon this traditional method, specifically...

[0073] A priority experience replay area is set up in the experience pool to store the data generated by the training of the policy network. Furthermore, in S2 to S4, when sampling data from the experience pool, the sampling probability of data in the priority experience replay area is different from that of data in other areas. That is, the data stored in the priority experience replay area is weighted by importance weight and sampled using non-uniform probability.

[0074] The above improvements allow for adjusting the sampling probability weights based on the TD error value, meaning that samples with high TD errors are more likely to be sampled than other samples.

[0075] More preferably, the sampled data is represented as (x i:i+N ,a i:i+N ,r i:i+N-1 )

[0076] Where i represents the sampling step number, x i:i+N Let a represent the set of UAV state information at all times from step i to step i+N. i:i+N Let r represent the set of drone action information at all times from step i to step i+N. i:i+N-1 This represents the set of drone reward information for all times from step i to step ii+N-1.

[0077] In a preferred embodiment, in S4, the parameter update of the policy network is expressed as:

[0078] ω←ω+β t δ ω

[0079]

[0080] Where ω is the parameter of the policy network, β t M represents the learning rate of the policy network, and M represents the data transfer batch size. This represents the gradient of the loss function with respect to the policy network parameters when the model parameters are ω. Y represents the importance sampling weight at step i. i Let represent the reinforcement learning value function at step i.

[0081] In a preferred embodiment, the parameter update of the evaluation network is expressed as:

[0082] θ←θ+α t δ θ

[0083]

[0084] Where θ represents the parameters used to evaluate the network, and α t This represents the learning rate used to evaluate the network. This represents the gradient of the loss function at step i with respect to the parameters of the value network. Z represents the gradient of the loss function with respect to the policy network. ω (x i ,α) represents the distribution of the value function at the i-th step, π θ (x i Let represent the policy network at step i.

[0085] Example

[0086] Example 1

[0087] A simulation experiment was conducted to demonstrate the cooperative interception of multiple drones based on a distributed reinforcement learning algorithm, including the following steps:

[0088] S1. Set up a simulation environment for the drone swarm to intercept the target, obtain the flight information of each drone in the drone swarm, the flight information of other drones in the swarm observed by the drone, and the target flight information observed by the drone, and store the flight information in the experience pool.

[0089] S2. Data is sampled from the experience pool through the policy network to generate a maneuvering policy. The UAV performs maneuvering actions according to the maneuvering policy, and the flight information of the UAV and the target flight information observed by the UAV during the maneuvering action are obtained. The flight information and maneuvering action values ​​are stored in the experience pool.

[0090] S3. Evaluate the value of the maneuver strategy through the evaluation network and store the value in the experience pool;

[0091] S4. Sample data from the experience pool and update the parameters of the policy network;

[0092] S5. Repeat steps S2 to S4 to obtain the strategy model. Set the strategy model on the UAV. The UAV flies according to the maneuver strategy output by the strategy model to intercept the target.

[0093] In S1, a protected area is set in the simulation environment with coordinates (-62.5, 50.9). Four target aircraft are set to attack the protected area, denoted as A-UAV1, A-UAV2, A-UAV3, and A-UAV4. The initial positions of the four target aircraft are (1354.2, -188.5), (-1132.8, -907), (1381.2, -604.1), and (768.4, 1338.2), respectively, and their initial velocities are 21.4 m / s, 22.5 m / s, 17.3 m / s, and 15.6 m / s, respectively.

[0094] Five drones were set up in the drone swarm to intercept the target aircraft, designated D-UAV1, D-UAV2, D-UAV3, D-UAV4, and D-UAV5. The initial positions of the five drones were (-19.4, 83.5), (-71.4, 102.6), (-98.7, 98.5), (-135.9, 40.5), and (-71.7, 60.5), respectively, with initial velocities of 20.1 m / s, 16.3 m / s, 17.2 m / s, 13.6 m / s, and 11.2 m / s, respectively. The drones could obtain flight information of other drones and the target aircraft through their onboard equipment. After obtaining the flight information, they also obtained the relative position and relative velocity between the observed target and the observing drone, as well as the relative position and relative velocity between other drones and the observing drone. This data, along with the flight information, was stored as a single data entry in the experience pool.

[0095] A separate policy network is set up for each drone, and multiple policy networks update parameters simultaneously.

[0096] The maneuvers corresponding to the maneuver policies generated by multiple policy networks are stored in the same experience pool.

[0097] In S4, the policy network is updated using a priority experience replay method, and the evaluation network is also updated.

[0098] In the evaluation of network updates, the time difference method is used to obtain the optimal value of the evaluation network.

[0099] In the time-difference method, the TD error is obtained by minimizing the distribution of the value function of the observed state and the next state. The value function of minimizing the observed state and the next state is expressed as:

[0100]

[0101] In evaluating the network, the value is set as follows:

[0102]

[0103] A priority experience replay area is set in the experience pool. In S2 to S4, when sampling data from the experience pool, the sampling probability of data in the priority experience replay area is different from that of data in other areas.

[0104] In S4, the parameter update of the policy network is represented as:

[0105] ω←ω+β t δ ω

[0106]

[0107] The parameter update of the evaluation network is represented as:

[0108] θ←θ+α t δ θ

[0109]

[0110] Figure 2 The figure shows the position curve of the drone intercepting the target aircraft. As can be seen from the figure, the drone's trajectory is smooth and the strike accuracy is high.

[0111] Figure 3 The normal overload curve of the UAV is shown. Figure 4 The axial overload curve of the UAV is shown. It can be seen that both the normal overload curve and the axial overload curve are smooth and meet the overload constraint requirements, thus ensuring the flight stability of the UAV.

[0112] Figure 5 The drone's flight speed curve is shown, indicating that the drone maintains a constant speed during the interception process and can autonomously adjust its speed appropriately, thus improving interception efficiency.

[0113] Figure 6 The heading angle curve of the drone is shown. It can be seen that the heading angle change curve of the drone is smooth during the interception process, which helps to improve the interception accuracy and prevent the drone's actuators from being damaged.

[0114] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," "outer," "front," and "rear," etc., indicate the orientation or positional relationship based on the orientation or positional relationship in the working state of this invention, and are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. Furthermore, the terms "first," "second," "third," and "fourth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0115] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0116] The present invention has been described above with reference to preferred embodiments; however, these embodiments are merely exemplary and illustrative. Various substitutions and modifications can be made to the present invention based on these embodiments, all of which fall within the scope of protection of the present invention.

Claims

1. A multi-aircraft cooperative non-cooperative target acquisition method based on distributed reinforcement learning algorithm, characterized in that, Includes the following steps: S1. Set up a simulation environment for the drone swarm to intercept the target, obtain the flight information of each drone in the drone swarm, the flight information of other drones in the swarm observed by the drone, and the target flight information observed by the drone, and store the flight information in the experience pool. S2. Data is sampled from the experience pool through the policy network to generate a maneuvering policy. The UAV performs maneuvering actions according to the maneuvering policy, and the flight information of the UAV and the target flight information observed by the UAV during the maneuvering action are obtained. The flight information and maneuvering action values ​​are stored in the experience pool. S3. Evaluate the value of the maneuver strategy through the evaluation network and store the value in the experience pool; S4. Sample data from the experience pool and update the parameters of the policy network; S5. Repeat steps S2 to S4 to obtain the strategy model. Set the strategy model on the UAV. The UAV flies according to the maneuver strategy output by the strategy model to intercept the target.

2. The multi-aircraft cooperative non-cooperative target acquisition method based on distributed reinforcement learning algorithm according to claim 1, characterized in that, The flight information includes position, speed, and line-of-sight angular velocity.

3. The multi-aircraft cooperative non-cooperative target acquisition method based on distributed reinforcement learning algorithm according to claim 1, characterized in that, A separate policy network is set up for each drone, and multiple policy networks update parameters simultaneously.

4. The multi-aircraft cooperative non-cooperative target acquisition method based on distributed reinforcement learning algorithm according to claim 3, characterized in that, The maneuvers corresponding to the maneuver policies generated by multiple policy networks are stored in the same experience pool.

5. The multi-aircraft cooperative non-cooperative target acquisition method based on distributed reinforcement learning algorithm according to claim 1, characterized in that, In S4, the policy network update process uses a priority experience replay method.

6. The multi-aircraft cooperative non-cooperative target acquisition method based on distributed reinforcement learning algorithm according to claim 1, characterized in that, In S4, the evaluation network was also updated.

7. The multi-aircraft cooperative non-cooperative target acquisition method based on distributed reinforcement learning algorithm according to claim 6, characterized in that, In the evaluation of network updates, the time difference method is used to obtain the optimal value of the evaluation network. In the time difference method, the TD error is obtained by minimizing the distribution of the value function of the observed state and the next state.

8. A multi-aircraft cooperative non-cooperative target acquisition method based on a distributed reinforcement learning algorithm according to claim 6, characterized in that, In evaluating the network, the value is set as follows: Among them, Y t Let γ represent the reinforcement learning value function, n represent the number of samplings, N represent the number of simulation steps, γ represent the decay factor, and r represent the reward value. t+n Let Q represent the reward value at time t+n. θ′ s represents the action value of the UAV when the parameter is θ′. t+n μ represents the state of the drone at time t+n. θ′ This represents the strategy function adopted by the UAV when the parameter is θ′.

9. A multi-aircraft cooperative non-cooperative target acquisition method based on a distributed reinforcement learning algorithm according to claim 1, characterized in that, A priority experience replay area is set in the experience pool. In S2 to S4, when sampling data from the experience pool, the sampling probability of data in the priority experience replay area is different from that of data in other areas.

10. A multi-aircraft cooperative non-cooperative target acquisition method based on a distributed reinforcement learning algorithm according to claim 1, characterized in that, In S4, the parameter update of the policy network is represented as: ω←ω+β t d ω Where ω is the parameter of the policy network, β t M represents the learning rate of the policy network, M represents the data transfer batch size, and i represents the sampling step number. Rp represents the gradient of the loss function with respect to the policy network parameters. i Y represents the importance sampling weight at step i. i Let represent the reinforcement learning value function at step i.