Full-size helicopter target hovering method based on reinforcement learning
Through a reinforcement learning-based method, combined with dynamics model, reward mechanism, curiosity exploration mechanism and self-attention mechanism, the problem of difficulty in controlling a full-size helicopter at hover or low speed is solved, and the efficient completion and stability improvement of the target point hover task is achieved.
Patent Information
- Application Number
- CN202510234203.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
The control of full-size helicopters at hover or low speed is more difficult, and the prior art is difficult to effectively achieve target point hover tasks, resulting in less practicality of the tasks.
A full-size helicopter dynamics model is adopted based on reinforcement learning, a reward mechanism for hover tasks is designed, a curiosity exploration mechanism for random network distribution is introduced, and a self-attention mechanism is introduced into the Actor network to improve the exploration ability and decision-making efficiency of strategies.
This significantly shortens the training time, improves the final performance in the later stage of training, ensures position stability at the target point and superior attitude control, avoids strategy jitter, and enhances the stability of the model in different situations.
Smart Images

Figure CN120162883A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of full-size helicopter target hovering methods, and specifically to a full-size helicopter target hovering method based on reinforcement learning. Background Art
[0002] A helicopter is an aircraft that relies on rotors to generate lift and can take off and land vertically and hover in the air for a long time. It is widely used in civil and military scenarios, including low-altitude operations, transportation, agriculture, and rescue. Due to its complex structure and rotor system, the dynamics of a helicopter are more complex than those of traditional layout aircraft such as fixed-wing aircraft, which makes the control of a helicopter more difficult.
[0003] Helicopter control has been widely studied, and various control methods such as PID, LQR, H∞, feedback linearization, and sliding mode control have been applied to helicopters. An important factor that makes helicopter control difficult is the strong nonlinearity of the dynamics. In recent years, with the breakthrough of machine learning, deep neural networks can fit nonlinear functions with arbitrary precision. Deep reinforcement learning (DRL) combines the nonlinear fitting ability of deep learning with the sequential decision-making ability of reinforcement learning. Through the interactive training of the controller, DRL has been successfully applied to the control of complex systems, such as Go and robot motion. There have been successful cases of applying reinforcement learning to helicopter control in the prior art. Ng et al. obtained the dynamic model of the Yamaha R-50 helicopter through locally weighted regression and learned a hovering controller through the PEGASUS algorithm. The results were very remarkable, and the trained controller made the flight stability of the helicopter exceed that of human pilots. In addition, the researchers improved the controller so that it could accurately perform competition actions. Based on the same method, the research team also learned the controllers for helicopter inverted hovering, aerobatic flight, and autorotation descent, promoting the technical development of autonomous helicopter aerobatic flight control. However, all of these used small unmanned helicopters, which are much smaller in size and weight than full-size helicopters. Due to their lighter structure, relatively simple power system, and higher maneuverability and response speed, there are significant differences in control strategies compared to full-size helicopters.
[0004] In the literature "Robust Adaptive Control for a Small Unmanned Helicopter Using Reinforcement Learning", the authors proposed a robust adaptive control method based on the combination of reinforcement learning and sliding mode control to solve the attitude control problem of small unmanned helicopters under dynamic uncertainties and external disturbances. However, this method is currently only applicable to the attitude control of small unmanned helicopters and does not involve translational motion control.
[0005] In the literature "Reinforcement Learning Control for a 2-DOF Helicopter With State Constraints Theory and Experiments", the authors proposed a control strategy based on reinforcement learning to achieve precise trajectory tracking control of a nonlinear two-degree-of-freedom helicopter system under the conditions of uncertainty and state constraints. Its main drawback is that the convergence speed of the reinforcement learning process is relatively slow, resulting in insufficient real-time performance of the control system.
[0006] Although deep reinforcement learning has been applied in helicopter control, it rarely involves full-scale helicopters. Due to the flexibility of large-scale rotor blades, the dynamic nonlinearity of full-scale helicopters is more severe. At the same time, helicopters are very unstable at low speeds, which makes the control of full-scale helicopters more difficult during hovering or at low speeds. Currently, the general hovering tasks are usually hovering at the current point, and the practicality of the tasks is relatively low. The target-point hovering task can be applied to trajectory control, landing and other tasks. To solve the above problems, this application proposes a target hovering method for full-scale helicopters based on reinforcement learning. Summary of the Invention
[0007] The purpose of the present invention is to provide a target hovering method for full-scale helicopters based on reinforcement learning to solve the problems raised in the above background technology.
[0008] To achieve the above purpose, the present invention provides the following technical solution: A target hovering method for full-scale helicopters based on reinforcement learning, including the following steps:
[0009] Step S1, establish a dynamic model of a full-scale helicopter;
[0010] Step S2, design a reward mechanism for the hovering task, and design the reward mechanism for the hovering task in two stages based on the distance of the helicopter from the target point;
[0011] Step S3, introduce a curiosity exploration mechanism with a random network distribution;
[0012] Step S4, introduce a self-attention mechanism into the Actor network of reinforcement learning, and introduce the self-attention mechanism into the second layer of the Actor network;
[0013] Step S5, train the agent, test the training results of the agent and display them.
[0014] Preferably, in the modeling process of step S1, the flight action characteristics of the helicopter are reproduced by a simulator, and the dynamic state of the helicopter is described by the following observed values: power, longitudinal airspeed, lateral airspeed, downward airspeed, northward speed, eastward speed, descent rate, roll angle, pitch angle, yaw angle, roll rate, pitch rate, yaw rate, X-axis position, Y-axis position, altitude and ground height.
[0015] Preferably, in step S1, the actual operation behavior of a human pilot is simulated through the following control operations, specifically, collective pitch control, longitudinal cyclic, lateral cyclic and rudder.
[0016] Preferably, the reward mechanism for the hovering task in step S2 is divided into two stages. When the helicopter is far from the target point and when the helicopter approaches the target point, the specific design of the reward function is as follows:
[0017]
[0018] Where d threshold is a set distance threshold for distinguishing different reward strategies, and r υ = v|| -||υ ⊥ ||, and υ || is the projection of the velocity in the target direction, and ||υ ⊥ || is the magnitude of the velocity component perpendicular to the target direction, representing the penalty for the vertical velocity. r xyz represents the negative value of the distance between the current position and the target point:
[0019] r xyz = -||x current -x target ||
[0020] Where, x current and x target represent the position vectors of the current position of the helicopter and the target point respectively. r pqr represents the sum of the differences between the current position and the target attitude:
[0021] r pqr = -(|p current -p target | + |q current -q target | + |r current -r target |)
[0022] p, q, r respectively represent the three angles of the attitude, namely, roll angle, pitch angle and yaw angle.
[0023] Preferably, in step S3, an RND intrinsic reward mechanism is introduced into SAC. When the agent encounters a new state, the error of the prediction network will increase, thereby generating a higher intrinsic reward to encourage the agent to explore these new areas. Specifically, the intrinsic reward r i is calculated by the following formula: r i =||f target (s t+1 )-f predict (s t+1 )|| 2 , and the intrinsic advantage is estimated by the intrinsic reward calculated by the RND model, and its definition is:
[0024]
[0025] where represents the intrinsic reward, is the Q value output by the intrinsic critic network, and the intrinsic advantage function and the original extrinsic advantage are combined as:
[0026]
[0027] where α and β respectively represent the weight factors of the extrinsic and intrinsic rewards.
[0028] Preferably, step S4 is specifically as follows: given an input state sequence {s1, s2,..., sn}, the self-attention mechanism first performs three linear transformations on each state vector to obtain corresponding Query (query), Key (key), and Value (value) vectors respectively. These vectors can be regarded as representations of the state sequence in different spaces and are used to capture the associations between states. The Query vector represents the information that the network currently wants to query, the Key vector represents the content that can be matched with the Query, and the Value stores the specific information related to the Key. Next, the Softmax function is used to normalize these attention scores and convert them into a probability distribution. Then, these weights are used to perform a weighted sum on the corresponding Value vectors to generate the final output of the self-attention mechanism.
[0029] Preferably, the specific process of step S5 is to iteratively train the agent. In each iteration loop, first obtain the current state from the environment, use the Actor network to generate the current action, then interact with the environment to obtain the next state and reward, and then calculate the reward and advantage using the reward function and the advantage function. Finally, update the Actor network and the Critic network.
[0030] The beneficial effects of the present invention compared with the prior art are:
[0031] 1) The improved algorithm of the present invention is close to the optimal in fewer time steps, which greatly shortens the training time. At the same time, the final performance of the method of the present invention in the later stage of training is also superior to SAC and TD3. In terms of the volatility of rewards, the algorithm of the present invention and SAC are relatively stable, while TD3 shows greater volatility;
[0032] 2) The algorithm of the present invention reaches the maximum value in 600 steps, and the convergence speed is significantly faster than SAC and TD3;
[0033] 3) In terms of the degree of oscillation, the curve of the method of the present invention is smoother as a whole, with less fluctuation, which means that the algorithm performs more stably during the training process, and the improvement of the strategy shows continuity and consistency, avoiding large-scale strategy jitter. This stability is particularly important for practical applications because it enables the model to maintain good performance in different scenarios;
[0034] 4) Compared with the traditional hovering task, the design of the reward function greatly broadens the exploration space of the helicopter;
[0035] 5) The self-attention mechanism is introduced. This improvement enables the model to more effectively understand and process the relationship between different states, thereby improving the accuracy and stability of task completion. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is the overall framework diagram of the present invention;
[0037] Figure 2 This is a framework diagram of the curiosity exploration mechanism based on random network distribution;
[0038] Figure 3 This is the framework diagram of the SAC algorithm that introduces the self-attention mechanism;
[0039] Figure 4 is the reward curve of this method and other reinforcement learning models during 1200 steps of training;
[0040] Figure 5 It is the position error comparison of different methods when the helicopter is flying to the target point (50,50);
[0041] Figure 6 The comparison of attitude control effects of different methods when the helicopter flies to the target point (50,50) is shown. DETAILED DESCRIPTION
[0042] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0043] Embodiment 1
[0044] Please refer to Figures 1-6 , a full-scale helicopter target hovering method based on reinforcement learning in the illustration, includes the following steps:
[0045] Step S1, establish a full-scale helicopter dynamics model;
[0046] Step S2, design a reward mechanism for the hovering task, and design the reward mechanism for the hovering task in two stages based on the distance between the helicopter and the target point;
[0047] Step S3, introduce a curiosity exploration mechanism of random network distribution;
[0048] Step S4, introduce a self-attention mechanism into the Actor network of reinforcement learning, and introduce a self-attention mechanism into the second layer of the Actor network;
[0049] Step S5, train the agent, test the training results of the agent and display them.
[0050] Based on the Soft Actor-Critic (SAC) algorithm, the present invention combines the intrinsic reward mechanism of random network distribution (RND) and the self-attention mechanism to improve the exploration ability and decision-making efficiency of the policy.
[0051] Specifically, Heli-gym is an open-source reinforcement learning environment built on the OpenAI Gym framework and designed for 6-degree-of-freedom helicopter flight. The simulator is based on the Minimum Complexity Helicopter Model, and aerodynamic dynamics is added to the model, which can simulate the performance of the helicopter under various flight conditions. This environment specifically refers to the aerodynamic characteristics of the AW109 helicopter, a light twin-engine helicopter known for its efficient aerodynamic design and flexible maneuverability. By using the dynamic parameters of the AW109, the simulator can more realistically reproduce the flight action characteristics of the helicopter. This precise modeling enables the reinforcement learning algorithm to be trained in a virtual environment approaching real flight, thereby effectively improving the performance and robustness of the autonomous flight controller.
[0052] The state information provided by this helicopter simulation environment includes the following items. These observations describe the dynamic state of the helicopter, covering information such as speed, angle, position, and altitude:
[0053] Power: Measured in horsepower (hp), ranging from 0 to infinity.
[0054] Longitudinal Air Speed: Measured in feet per second (ft / s), ranging from negative infinity to positive infinity.
[0055] Lateral Air Speed: Measured in feet per second (ft / s), ranging from negative infinity to positive infinity.
[0056] Downward Air Speed: Measured in feet per second (ft / s), ranging from negative infinity to positive infinity.
[0057] North Velocity: Measured in feet per second (ft / s), ranging from negative infinity to positive infinity.
[0058] East Velocity: Measured in feet per second (ft / s), ranging from negative infinity to positive infinity.
[0059] Descend Rate: Measured in feet per second (ft / s), ranging from negative infinity to positive infinity.
[0060] Roll Angle: Measured in radians (rad), ranging from -π to π.
[0061] Pitch Angle: Measured in radians (rad), ranging from -π to π.
[0062] Yaw Angle: Measured in radians (rad), ranging from -π to π.
[0063] Roll Rate, Body Frame: Measured in radians per second (rad / s), ranging from negative infinity to positive infinity.
[0064] Pitch Rate, Body Frame: Measured in radians per second (rad / s), ranging from negative infinity to positive infinity.
[0065] Yaw Rate, Body Frame: Measured in radians per second (rad / s), ranging from negative infinity to positive infinity.
[0066] X - axis position (X Loc, Earth): in feet, ranging from negative infinity to positive infinity.
[0067] Y - axis position (Y Loc, Earth): in feet, ranging from negative infinity to positive infinity.
[0068] Altitude above sea level (Sea Altitude): in feet, ranging from 0 to infinity.
[0069] Altitude above ground (Ground Altitude): in feet, ranging from 0 to infinity.
[0070] In this environment, the control of the helicopter simulates the actual operation behavior of a human pilot, mainly achieved through the following key control mechanisms:
[0071] Collective pitch control (Collective)
[0072] Longitudinal cyclic control (Longitudinal Cyclic)
[0073] Lateral cyclic control (Lateral Cyclic)
[0074] Rudder control (Pedal).
[0075] In this embodiment, step S2 improves the reward mechanism for the hovering task, which is divided into two stages: when the helicopter is far from the target point, the reward function is mainly based on speed to encourage the helicopter to move towards the target point as fast as possible, thereby reducing the constraints on attitude; when the helicopter approaches the target point, the reward function focuses on position error and attitude stability to ensure that the helicopter can hover precisely and stably at the target position. This design significantly broadens the exploration space of the helicopter compared to traditional hovering tasks. The specific design of the reward function is as follows:
[0076]
[0077] where d threshold is a set distance threshold used to distinguish different reward strategies, r υ = υ || - ||υ ⊥ ||, and υ || is the projection of the velocity in the target direction, ||υ ⊥ || is the magnitude of the velocity component perpendicular to the target direction, representing the penalty for vertical velocity, r xyz represents the negative value of the distance between the current position and the target point:
[0078] r xyz = - |||x current - xtarget ||
[0079] Among them, x current and x target respectively represent the position vectors of the current position of the helicopter and the target point, and r pqr represents the sum of the differences between the current position and the target attitude:
[0080] r pqr = -(|p current - p target | + |q current - q target | + |r current - r target |)
[0081] p, q, and r respectively represent the three angles of the attitude, namely the roll angle, the pitch angle, and the yaw angle.
[0082] In step S3, in order to encourage the agent to explore more environmental states, the RND intrinsic reward mechanism is introduced in SAC. RND consists of a fixed target network and a trainable prediction network. The target network outputs a stable feature representation, while the prediction network learns to imitate the output of the target network. When the agent encounters a new state, the error of the prediction network will increase, thereby generating a higher intrinsic reward to encourage the agent to explore these new areas. Specifically, the intrinsic reward r i is calculated by the following formula:
[0083] r i = ||f target (s t+1 ) - f predict (s t+1 )|| 2 ,
[0084] This mechanism is combined with the extrinsic reward r e of SAC, enabling the agent to consider both the direct feedback of the environment and the exploration of new potential states.
[0085] During the training process, the agent calculates the advantages based on the intrinsic reward and the extrinsic reward respectively. The intrinsic advantage is estimated by the intrinsic reward calculated by the RND model, and its definition is:
[0086]
[0087] Among them, represents the intrinsic reward, is the Q value output by the intrinsic critic network. The intrinsic advantage function and the original extrinsic advantage are combined as:
[0088]
[0089] Among them, α and β respectively represent the weight factors of extrinsic and intrinsic rewards. By adjusting these parameters, exploration and exploitation can be balanced, and the advantages of the combination can be used to update the policy gradient of the Actor network, so as to continuously optimize its action selection strategy while ensuring that the agent effectively explores the environment.
[0090] The self-attention mechanism is a technique that can dynamically capture the dependencies between input elements. It was first applied in the Transformer model. It adjusts the weights of different elements by calculating the similarities of Query, Key, and Value vectors, thereby generating a more globally informative representation. The self-attention mechanism can capture the dependencies between distant elements and can be efficiently computed in parallel.
[0091] Specifically, given an input state sequence {s1, s2,..., sn}, the self-attention mechanism first performs three linear transformations on each state vector to obtain the corresponding Query (query), Key (key), and Value (value) vectors. These vectors can be regarded as representations of the state sequence in different spaces and are used to capture the associations between states. The Query vector represents the information that the network currently wants to query, the Key vector represents the content that can be matched with the Query, and the Value stores the specific information related to the Key. By performing the dot product operation of Query and Key for each state, the similarity or correlation between different states, that is, the attention scores, can be obtained. Next, the Softmax function is used to normalize these attention scores and convert them into a probability distribution to represent the degree of attention of the network to different states. Then, these weights are used to perform a weighted sum of the corresponding Value vectors to generate the final output of the self-attention mechanism. This output not only contains the information of the current state but also comprehensively considers the dependencies between all other states, ensuring that the network can capture the global associations between different states in the input sequence.
[0092] The present invention introduces the self-attention mechanism into the second layer of the Actor network of SAC. When the traditional Actor network processes multi-dimensional input states, it is difficult to dynamically capture the complex relationships between various features. The self-attention mechanism can automatically extract the correlations between the state input features by calculating Query, Key, and Value vectors, thereby helping the network to generate more accurate actions based on these relationships. The input state is first encoded into Query, Key, and Value vectors, and then a new representation is obtained through the calculation of attention weights. This design can not only capture the mutual dependencies between input features but also more flexibly adjust the control strategy of the agent in policy learning.
[0093] Specifically, in the Actor network, after the input state passes through the self-attention mechanism, it can better extract and express important information in the state, so as to select appropriate actions. This shows good effects in practical applications, especially in complex scenarios where the environmental state is constantly changing. Through this improvement, SAC can converge faster in the learning stage and execute the policy more stably in the testing stage.
[0094] Next, the present invention iteratively trains the agent. In each iteration loop: First, obtain the current state from the environment, use the Actor network to generate the current action, then interact with the environment to obtain the next state and reward. After that, calculate the reward and advantage using the reward function and the advantage function. Finally, update the Actor network and the Critic network.
[0095] Figure 4 Shows the reward curves of the method of the present invention and other reinforcement learning models during the 1200-step training process, where each step contains 1000 training updates, totaling 1.2 million updates. This reflects the complexity and high difficulty of the helicopter control task from the side. The improved algorithm of the present invention approaches the optimum in fewer time steps, significantly shortening the training time. At the same time, the final performance of the improved algorithm in the later stage of training is also better than that of SAC and TD3. In terms of reward volatility, the method of the present invention is relatively stable compared with SAC, while TD3 shows greater fluctuations. The experimental results verify the effectiveness of the method of the present invention.
[0096] Figure 5 Shows the comparison of the position errors of different methods during the process of the helicopter flying towards the target point (50, 50). It can be seen from the figure that the method of the present invention is comparable to TD3 in terms of the speed of reaching the target point and is significantly better than SAC. At the same time, in terms of the position stability after reaching the target point, the method of the present invention performs more excellently, being better than the other two methods. Specifically, the statistical average position error after 400 steps shows that the error of the method of the present invention is 1.43, while the errors of SAC and TD3 are 1.83 and 2.01 respectively. This result further proves the significant advantages of the method of the present invention in position maintenance and stability.
[0097] Figure 6It shows the comparison of the attitude control effects of different methods when the helicopter flies towards the target point (50, 50). It can be seen from the figure that after reaching the target point, the method of the present invention performs more excellently in the control of the roll angle and the pitch angle. For the attitude angle (YawAngle), the absolute value of the yaw angle of the method of the present invention is slightly larger than that of other methods. This is because during the training process, the present invention does not constrain the absolute value of the attitude angle, but pays more attention to the minimization of the angular velocity. Since the goal is to make the fluctuation of the system as small as possible, the smaller the variance, the more stable the system. The calculation results show that after 400 steps, the variances of the yaw angles of each method are SAC: 0.015, TD3: 0.051, and OURS: 0.017 respectively. It can be seen that the control of the yaw angle by the method of the present invention is comparable to that of SAC and much lower than that of TD3, showing better stability. Generally speaking, the method of the present invention performs outstandingly in the control of the roll angle and the pitch angle, and at the same time maintains a level comparable to that of SAC in the control of the yaw angle, being much better than TD3.
[0098] As can be seen from the above, the improved algorithm is superior to the original SAC method in terms of convergence speed, training stability, and final performance, demonstrating its potential and advantages in reinforcement learning tasks.
[0099] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0100] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A full-size helicopter target hovering method based on reinforcement learning, characterized in that: The steps include: Step S1, establishing a full-scale helicopter dynamics model; Step S2, designing a reward mechanism for the hovering task. Based on the distance between the helicopter and the target point, the reward mechanism for the hovering task is designed in two stages. Step S3, introducing a curiosity exploration mechanism with random network distribution; Step S4, introducing a self-attention mechanism into the Actor network of reinforcement learning, and introducing a self-attention mechanism into the second layer of the Actor network; Step S5: train the intelligent agent, test the training results of the intelligent agent and display them.
2. The method for hovering a full-size helicopter target based on reinforcement learning according to claim 1, characterized in that: The modeling process in step S1 reproduces the flight action characteristics of the helicopter through a simulator, and uses the following observations to describe the dynamic state of the helicopter: power, longitudinal air speed, lateral air speed, downward air speed, north speed, east speed, descent rate, roll angle, pitch angle, yaw angle, roll rate, pitch rate, yaw rate, X-axis position, Y-axis position, altitude and ground height.
3. The method for hovering a full-size helicopter target based on reinforcement learning according to claim 2, characterized in that: The step S1 simulates the actual operation behavior of the human pilot through the following manipulation controls, specifically collective pitch control, longitudinal cyclic stick, lateral cyclic stick and rudder.
4. The method for hovering a full-size helicopter target based on reinforcement learning according to claim 3, characterized in that: The reward mechanism of the hovering task in step S2 is divided into two stages, when the helicopter is far away from the target point and when the helicopter is close to the target point, the specific design of the reward function is as follows: where d threshold is a distance threshold set to distinguish different reward strategies, r v =v || -||v ⊥ ||, and v || is the projection of the velocity in the target direction, ||v ⊥ || is the magnitude of the velocity component perpendicular to the target direction, indicating the penalty for vertical velocity, r xyz Indicates the negative value of the distance between the current position and the target point: r xyz =-||x current -x target || Among them, x current and x target Represent the current position of the helicopter and the position vector of the target point, r pqr Represents the sum of the differences between the current position and the target posture: r pqr =-(|p current -p target |+|q current -q target |+|r current -r target |) p, q, and r represent the three angles of attitude, namely the roll angle, pitch angle, and yaw angle.
5. The method for hovering a full-size helicopter target based on reinforcement learning according to claim 4, characterized in that: The step S3 introduces the RND intrinsic reward mechanism in SAC. When the agent encounters a new state, the error of the prediction network will increase, thereby generating a higher intrinsic reward to encourage the agent to explore these new areas. Specifically, the intrinsic reward r i Calculated by the following formula: r i =||f target (s t+1 )-f predict (s t+1 )|| 2 , inherent advantages It is estimated by the intrinsic reward calculated by the RND model, which is defined as: in, Intrinsic rewards, is the Q value output by the intrinsic critic network, combining the intrinsic advantage function with the original extrinsic advantage: Among them, α and β represent the weight factors of extrinsic and intrinsic rewards, respectively.
6. The method for hovering a full-size helicopter target based on reinforcement learning according to claim 5, characterized in that: The step S4 is specifically given the input state sequence {s1, s2, ..., sn}. The self-attention mechanism first performs three linear transformations on each state vector to obtain corresponding Query, Key and Value vectors respectively. These vectors can be regarded as representations of the state sequence in different spaces and are used to capture the association between states. The Query vector represents the information that the network currently wants to query, the Key vector represents the content that can be matched with the Query, and the Value stores specific information related to the Key. Next, the Softmax function is used to normalize these attention scores and convert them into a probability distribution. Then, these weights are used to perform weighted summation on the corresponding Value vectors to generate the final output of the self-attention mechanism.
7. The method for hovering a full-size helicopter target based on reinforcement learning according to claim 6, characterized in that: The specific process of step S5 is to iteratively train the intelligent agent. In each iterative cycle, first obtain the current state from the environment, use the Actor network to generate the current action, then interact with the environment to obtain the next state and reward, then use the reward function and advantage function to calculate the reward and advantage, and finally update the Actor network and Critic network.