Civil aircraft collision release strategy optimization method

By applying reinforcement learning and optimization of collision liberation strategies in civil aircraft collision liberation systems, the problems of low computing efficiency and long response time of existing systems are solved, and higher flexibility and security are achieved.

CN120048162APending Publication Date: 2025-05-27EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510229865.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-27

Smart Images

  • Figure CN120048162A_ABST
    Figure CN120048162A_ABST
Patent Text Reader

Abstract

The invention discloses a collision release strategy optimization method for a civil aircraft, which is characterized in that related parameters for determining collision release are optimized by adopting a reinforcement learning method, an optimized collision release strategy is generated, and the method specifically comprises the following steps: 1) acquiring information of a local aircraft and an intrusion aircraft; 2) constructing a collision model between the local machine and the intrusion machine; 3) determining a collision release strategy based on the prediction model; 4) constructing a strategy joint optimization direction; and 5) optimizing a collision release strategy, generating an optimized collision release strategy and the like. Compared with the prior art, the method has the advantages that the flexibility of the local aircraft to deal with various flight conditions in a collision release strategy is improved, the vertical acceleration of the local aircraft is limited and optimized, the flight stability of the local aircraft is effectively improved, the parameter of estimated response time of a pilot is optimized, and the safety threshold value of collision release is effectively expanded; in addition, the vertical interval is restrained and maximized, and the strongest safety performance is guaranteed on the basis of flight stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of airborne collision avoidance systems, and more specifically to an optimization method for collision release of civil aircraft. Background Art

[0002] The Traffic Alert and Collision Avoidance System (TCAS) is an important part of modern aviation safety. It aims to continuously monitor the positions, speeds, and flight paths of aircraft in the surrounding airspace through on-board equipment, and provide early warnings and avoidance instructions to pilots when potential collision threats are detected. TCAS was developed by the Federal Aviation Administration of the United States, CAA (Civil Aviation Authority) centers of other countries, and the aviation industry over many years. It is a complete solution that can effectively reduce the risk of mid-air collisions between aircraft. After years of technical improvements and developments, the advanced TCAS II has been obtained. Compared with the first-generation system, TCAS II provides more accurate and complex functions, especially in terms of collision avoidance strategies. It can not only provide Traffic Advisories (TA), but also generate Resolution Advisories (RA), that is, specific avoidance instructions (such as climbing or descending) to ensure a safe separation between aircraft.

[0003] TCAS II continuously observes the airspace around the aircraft itself. With the help of the ATC transponder, it searches for responses from other aircraft in the nearby airspace. If other aircraft enter the monitoring area of the aircraft itself, based on their responses to the aircraft's own inquiries, it obtains their relative distance, azimuth, and altitude information, and conducts threat detection and assessment on the encountering aircraft according to the horizontal miss distance (HMD) and vertical miss distance (VMD) at the closest point between the aircraft itself and the intruding aircraft. It predicts the time τ for the two aircraft to reach the closest point of approach (CPA), as well as the HMD and VMD. Through the time τ and the CAS collision avoidance logic, it continuously tracks and evaluates the threat level of the intruding aircraft to the aircraft itself. The warnings are divided into the above two levels of TA and RA. Then, according to the collision threat assessment information, a collision release strategy is given. TCAS II uses a four-beam directional antenna to determine the azimuth angle of the intruding aircraft, and its accuracy is 8-10°. Since the horizontal flight speed of the aircraft is much greater than the vertical flight speed, such azimuth angle accuracy is too low when the aircraft is flying at high speed. Therefore, the aircraft uses a one-dimensional vertical maneuver method for collision release.

[0004] The current collision avoidance strategy adopted by TCAS II is relatively fixed, and it does not obtain the optimal vertical acceleration of the local aircraft during collision avoidance and the expected reaction time of the pilot to ensure the maximum safety threshold. Therefore, the effectiveness of the collision avoidance strategy is not entirely determined by the detection accuracy of the system itself, but also involves the tuning of corresponding parameters. These parameters control how the system flexibly adjusts the behavior guided by the collision avoidance strategy according to different flight environments, aircraft dynamics, and airspace density.

[0005] In summary, the existing multi-aircraft collision warning systems have problems such as low computing efficiency and long response time when dealing with the collision risks between multiple devices, seriously affecting the safety of air traffic management. Summary of the Invention

[0006] The object of the present invention is to provide an optimization method for collision avoidance of civil aircraft in view of the deficiencies of the prior art. By using the method of reinforcement learning, the relevant parameters for determining collision avoidance are optimized to generate an optimized collision avoidance strategy. This method effectively improves the flight stability of the local aircraft by restricting and tuning the vertical acceleration of the local aircraft, and optimizes the expected reaction time of the pilot, effectively expanding the safety threshold of collision avoidance. In addition, the vertical separation is also constrained and maximized, ensuring the strongest safety performance on the basis of flight stability. Combining the above optimization directions, a corresponding reward function is set, and reinforcement learning is used for training and optimization again. While ensuring the basic safety performance, the flexibility and safety of the local aircraft in dealing with various flight situations in the collision avoidance strategy are effectively improved, with fast response speed and high accuracy, effectively solving the problems of low computing efficiency and long response time of the airborne collision avoidance system, and having good application prospects.

[0007] The specific technical solution for achieving the object of the present invention is: an optimization method for the collision avoidance strategy of civil aircraft, characterized in that the method specifically includes the following steps:

[0008] Step 1: Obtain the information of the local aircraft and the intruding aircraft

[0009] The obtained information includes: the initial vertical height and initial vertical height rate of the local aircraft; the sensitivity level, initial vertical height, and initial vertical height rate of the intruding aircraft; the horizontal relative distance between the intruding aircraft and the local aircraft, and the horizontal distance change rate of the intruding aircraft relative to the local aircraft. The specific steps are as follows:

[0010] 1-1: The local aircraft sends an inquiry every 1 s and monitors the received information content, collects and classifies the monitored data information, and stores it in a specific input data structure.

[0011] 1-2: Coarsely filter the data, obtain the priority of the data according to the timestamp of the data and the judgment conditions, and process the information of the intruding aircraft with higher priority first. The judgment conditions include: 1) The smaller the distance between the intruding aircraft and the local aircraft, the higher the priority; 2) The priority is higher when the address information of the intruding aircraft and the local aircraft are in the same system; 3) The priority is higher when the monitored data information is complete and the verification is accurate.

[0012] 1-3: Obtain the corresponding local and intruding aircraft information according to the content of the local inquiry information sent and the intruding aircraft response information and the time difference between the two, and store the corresponding information in a computer text document.

[0013] Step 2: Construct a collision model between the local aircraft and the intruding aircraft

[0014] 2-1: Construct a position relationship model between the local aircraft and the intruding aircraft;

[0015] 2-2: According to the horizontal relative distance R between the local aircraft and the intruding aircraft and the horizontal distance change rate of the intruding aircraft relative to the local aircraft Calculate the moment τ when the relative distance between the two aircraft is the smallest. The calculation formula is as follows:

[0016]

[0017] To ensure that the value range of τ is between 15 and 35 seconds, correct τ. The corrected Formula is calculated by the following formula:

[0018]

[0019] In the formula, D m Is a correction parameter introduced according to the sensitivity level of the intruding aircraft.

[0020] 2-3: Input the relevant information of the local aircraft and the intruding aircraft into the above collision model, and select the minimum value of τ and As the moment when the relative distance position between the two aircraft is the smallest, which is the time to reach the closest point of approach (Closest Point of Approach, abbreviated as CPA) between the local aircraft and the intruding aircraft.

[0021] Step 3: The prediction model determines the collision avoidance strategy

[0022] 3-1: According to the current civil aircraft performance, select the parameters adopted by the vertical maneuver method during coarse prediction under different collision avoidance states. The parameters mainly include: acceleration Accel and pilot reaction time Delay.

[0023] 3-2: According to the vertical flight models of the own aircraft and the intruder aircraft established during the rough prediction process, calculate the altitude values H OWN and H INT of the own aircraft and the intruder aircraft at the closest point of approach (CPA), respectively, so as to obtain the closest vertical separation distance ΔH between the two. Among them, the flight model of the intruder aircraft is a uniformly accelerated linear motion in the vertical direction, and the own aircraft model is divided into three parts: delayed uniform motion, uniformly accelerated motion, and target uniform motion. The corresponding calculation formulas are as follows:

[0024]

[0025] ΔH = |H OWN - H INT |.

[0026] Among them, H O and H I represent the initial vertical altitudes of the own aircraft and the intruder aircraft respectively, and T A represents the duration of the uniformly accelerated part of the own aircraft, which is determined by the initial vertical altitude rate of the own aircraft and the initial vertical altitude rate of the intruder aircraft .

[0027] 3-3: Substitute the parameters in different collision avoidance states into the prediction model, calculate the obtained closest vertical separation distance ΔH, compare it with the safety threshold value H LIMIT of the vertical separation distance, obtain a preliminary collision avoidance strategy, and determine the target vertical speed of the own aircraft to be achieved under this strategy. The preliminary collision avoidance strategy mainly includes: nominal climb, increased climb, nominal descent, and increased descent, and each corresponds to different target vertical speeds

[0028] Step 4: Construct the joint optimization direction of the strategy

[0029] To maximize the safety threshold, the optimization directions in the strategy include: minimizing the vertical acceleration Accel of the own aircraft to ensure the stability of the own aircraft's flight; maximizing the pilot's expected reaction time Delay to ensure sufficient reaction time is reserved; maximizing the difference between the closest vertical separation distance ΔH and the safety threshold value H LIMIT of the vertical separation distance to obtain the safest collision avoidance strategy parameters under the current strategy.

[0030] Step 5: Optimize the collision avoidance strategy by using the method of reinforcement learning to generate an optimized collision avoidance strategy

[0031] 5-1: Select the input state s in reinforcement learning as the own aircraft information, intruder aircraft information, interaction information between the two, and the target vertical speed of the own aircraft obtained from the rough prediction The relevant information of the host and the intruder aircraft specifically includes: the initial vertical height H of the host O and the initial vertical height rate The initial vertical height H of the intruder aircraft I , the initial vertical height rate and the sensitivity level Sensitivity_levels, the horizontal relative distance R between the intruder aircraft and the host, the rate of change of the horizontal distance of the intruder aircraft relative to the host

[0032] Step 5-2: Determine the output action a in reinforcement learning, including: the vertical acceleration Accel taken by the host and the expected reaction time of the pilot

[0033] Step 5-3: Use the Actor-Critic architecture for reinforcement learning training, where the goal of the Actor network is to adjust the θ π parameter to learn a policy π(a t |s t ; θ π ), and the Critic is used to evaluate the value of the state-action pair, that is, the state-action value function Q(s t ,a t ; θ q )

[0034] Step 5-4: Determine the content of the reward function in reinforcement learning according to the policy joint optimization direction, including: the vertical acceleration part R of the host Accel , the time planning part R Duration and the vertical interval threshold comparison part R △H , and the specific expression of the design of the reward function is as follows

[0035] R Accel =-0.05·e 0.45(|Accel|-4) ;

[0036]

[0037] R △H =0.3·(log(1 + △H)-log(1 + H LIMIT ));

[0038]

[0039] R Total =R Accel +R Duration +R △H +R Penalty .

[0040] Among them, H LIMITThe value range of Penalty is 300 - 700 ft, which will be selected according to the altitude levels of different aircraft. In addition, a penalty part R

[0041] is added to ensure that the selected parameters will not collide and the state is below the vertical separation threshold. Repeat this process continuously until the aircraft reaches the target vertical speed and achieves collision avoidance.

[0042] Step 5 - 6: Update the policy of the Actor using policy gradients. The specific gradient update formula is as follows:

[0043]

[0044] where A t is the advantage function.

[0045] The advantage of the current action relative to other actions is measured by subtracting the baseline (such as the state value function where R t represents the expected return), and can be calculated by the Critic network designed above. The specific calculation formula is:

[0046] A t = Q(s t , a t ; θ q ) - V(s t ), Q(s t , a t ; θ q ) ≈ Ε[R t |s t , a t .

[0047] Meanwhile, the gap between the currently estimated state value and the true return based on the current state and subsequent states is quantified by the Temporal Difference (TD) Error, and then the state - action pair value function Q(s t , a t ) in the Critic is updated. The TD Error and the update are specifically expressed as:

[0048] δ t = R t + γQ(s t+1 , a t+1 ) - Q(s t , a t );

[0049] Q(s t ,a t ) ← Q(s t ,a t ) + α·δ t 。

[0050] Wherein, δ t is the TD Error, γ is the weight factor used to control the weight of future rewards, 0 ≤ γ ≤ 1, Q(s t ,a t ) and Q(s t+1 ,a t+1 ) are the current state - action value function and the state - action value function at the next moment respectively.

[0051] During training, the loss function of the Actor network is The loss function of the Critic network is L Critic = Ε[(δ t ) 2 , and the two cooperate with each other for training to minimize their loss functions, so as to obtain a collision release strategy with higher safety and stronger flexibility.

[0052] Compared with the prior art, the present invention has the following beneficial technical effects and remarkable technical progress:

[0053] 1) By restricting and optimizing the vertical acceleration of the aircraft, the stability of the aircraft flight is effectively improved;

[0054] 2) By optimizing the parameter of the pilot's expected reaction time, it helps to expand the safety threshold of collision release;

[0055] 3) By constraining and maximizing the vertical separation, the strongest safety performance is ensured on the basis of flight stability;

[0056] 4) Through reinforcement learning training and optimization, the flexibility of the aircraft in dealing with various flight situations in the collision release strategy is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is the flow chart of the present invention;

[0058] Figure 2 is the schematic diagram of the two - aircraft collision detection model;

[0059] Figure 3 is the schematic diagram of the prediction strategy;

[0060] Figure 4 is the structural diagram of the policy optimization network. DETAILED DESCRIPTION OF THE INVENTION

[0061] Refer to Figure 1 , the present invention optimizes the relevant parameters for determining collision release by using the reinforcement learning method, specifically including the following steps:

[0062] Step 1: Information collection

[0063] 1-1: Obtain the initial vertical height and initial vertical height rate of the own aircraft;

[0064] 1-2: Obtain the sensitivity level, initial vertical height and initial vertical height rate of the intruding aircraft;

[0065] 1-3: Obtain the horizontal relative distance between the intruding aircraft and the own aircraft and the horizontal distance change rate of the intruding aircraft relative to the own aircraft.

[0066] Step 2: Build a collision model

[0067] 2-1: Build a collision model between the own aircraft and the intruding aircraft by using the position relationship between the own aircraft and the intruding aircraft;

[0068] 2-2: According to the horizontal relative distance R between the own aircraft and the intruding aircraft, and the horizontal distance change rate of the intruding aircraft relative to the own aircraft Calculate the moment τ when the relative distance between the two aircraft is the smallest, and use the sensitivity level of the intruding aircraft as the correction parameter D m Correct τ to obtain the corrected Make the value range of τ between 15 and 35 seconds;

[0069] 2-3: Input the obtained information of the own aircraft and the intruding aircraft into the collision model, and select the value of τ and The value that makes the vertical separation smaller as the moment when the relative distance between the two aircraft is the smallest, that is, the time to reach the closest point between the own aircraft and the intruding aircraft.

[0070] Step 3: Determine the collision release strategy

[0071] 3-1: Establish a motion prediction model for the own aircraft and the intruding aircraft;

[0072] 3-2: Use the motion prediction model to calculate its collision release strategy.

[0073] Step 4: Build the strategy joint optimization direction

[0074] 4-1: Minimize the vertical acceleration of the own aircraft to ensure the flight stability of the own aircraft;

[0075] 4-2: Maximize the pilot's expected reaction time to ensure sufficient reaction time is reserved;

[0076] 4-3: Maximize the closest vertical separation distance to obtain the safest collision release strategy parameters under the current strategy.

[0077] Step 5: Optimize the collision release strategy by using the method of reinforcement learning to generate an optimized collision release strategy.

[0078] In order to more clearly and comprehensively illustrate the technical means, technical improvements and beneficial effects of the present invention, the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0079] Embodiment 1

[0080] In this embodiment, the sensitivity level of the intruder aircraft is taken as 6, that is, the nominal flight altitude is in the range of FL100 - FL200. At this time, the correction coefficient D m = 0.8, and the safety threshold value H of the vertical separation distance LIMIT = 600 ft.

[0081] Refer to Figure 1 , a method for optimizing the collision release strategy of a civil aircraft, comprising the following specific steps:

[0082] Step 1: Obtain the information of the own aircraft and the intruder aircraft

[0083] The information of the own aircraft and the intruder aircraft includes: the initial vertical altitude and the initial vertical altitude rate of the own aircraft; the sensitivity level, the initial vertical altitude and the initial vertical altitude rate of the intruder aircraft; the horizontal relative distance between the intruder aircraft and the own aircraft, and the horizontal distance change rate of the intruder aircraft relative to the own aircraft. The specific steps are as follows:

[0084] 1-1: The own aircraft sends an inquiry every 1 s and monitors the received information content, and collects and classifies the monitored data information and stores it in a specific input data structure.

[0085] 1-2: Roughly filter the data, obtain the priority of the data according to the time stamp of the data and the judgment conditions, and process the information of the intruder aircraft with higher priority first. The judgment conditions include: 1) The smaller the distance between the intruder aircraft and the own aircraft, the higher the priority; 2) The priority of the address information of the intruder aircraft and the own aircraft in the same system is higher; 3) The priority of the monitored data information that is complete and accurately verified is higher.

[0086] 1-3: Obtain the corresponding information of the own aircraft and the intruder aircraft according to the content of the inquiry information sent by the own aircraft and the response information of the intruder aircraft and the time difference between the two, and store the corresponding information in a computer text document.

[0087] Step 2: Construct a collision model between the own aircraft and the intruder aircraft

[0088] Refer to Figure 2 , the steps to build the collision model between the host aircraft and the intruder aircraft are as follows:

[0089] 2-1: Build the position relationship model between the host aircraft and the intruder aircraft.

[0090] 2-2: According to the horizontal relative distance R between the host aircraft and the intruder aircraft and the horizontal distance change rate of the intruder aircraft relative to the host aircraft calculate the moment τ when the relative distance between the two aircraft is the smallest, and its calculation formula is as follows:

[0091]

[0092] To ensure that the value range of τ is between 15 and 35 seconds, correct τ, and its corrected is calculated by the following formula:

[0093]

[0094] In the formula, D m is the correction parameter introduced according to the sensitivity level of the intruder aircraft.

[0095] 2-3: Input the relevant information of the host aircraft and the intruder aircraft into the above collision model, and select the minimum value of τ and as the moment when the relative distance between the two aircraft is the smallest, which is the time to reach the closest point of approach (Closest Point of Approach, abbreviated as CPA) between the host aircraft and the intruder aircraft.

[0096] Step 3: The prediction model determines the collision avoidance strategy

[0097] Refer to Figure 3 , the prediction steps of the collision avoidance strategy are as follows:

[0098] 3-1: Select the parameters adopted by the vertical maneuver method in different collision avoidance states during the rough prediction. The parameters mainly include: acceleration Accel and pilot reaction time Delay. The detailed values of the parameters are mainly that the absolute value of acceleration Accel is 8 ft / s at the nominal avoidance state 2 , and the pilot reaction time Delay is 2.5 s; the absolute value of acceleration Accel is 11 ft / s at the increased avoidance state 2 , and the pilot reaction time Delay is 5 s.

[0099] 3-2: According to the vertical flight model of the host aircraft and the intruder aircraft established during the rough prediction as shown in Figure 2 , calculate the altitude values H OWN and HINT , so as to obtain the closest vertical interval distance ΔH. Among them, the flight model of the intruding aircraft is a uniformly accelerated linear motion in the vertical direction, and the model of the own aircraft is divided into three parts: delayed uniform motion, uniformly accelerated motion, and target uniform motion. The corresponding calculation formulas are as follows:

[0100]

[0101] ΔH = |H OWN - H INT |.

[0102] Among them, H O and H I respectively represent the initial vertical heights of the own aircraft and the intruding aircraft, and T A represents the duration of the uniformly accelerated part of the own aircraft, which is determined by the initial vertical height rate of the own aircraft and the initial vertical height rate of the intruding aircraft .

[0103] 3 - 3: Substitute the parameters in different collision avoidance states into the Figure 3 prediction model shown in the figure, calculate the closest vertical interval distance ΔH, and compare it with the safety threshold value H LIMIT of the vertical interval distance to obtain a preliminary collision avoidance strategy, and determine the target vertical speed of the own aircraft that needs to be achieved under this strategy. The preliminary collision avoidance strategies mainly include four types: nominal climb, increased climb, nominal descent, and increased descent, and each corresponds to different target vertical speeds

[0104] Step 4: Construct the combined optimization direction of the strategy.

[0105] To maximize the safety threshold, the optimization directions in the strategy should include: minimizing the vertical acceleration Accel of the own aircraft to ensure the stability of the own aircraft's flight; maximizing the pilot's expected reaction time Delay to ensure sufficient reaction time; maximizing the difference between the closest vertical interval distance ΔH and the safety threshold value H LIMIT of the vertical interval distance to obtain the safest collision avoidance strategy parameters under the current strategy.

[0106] Step 5: Optimize the collision avoidance strategy using the method of reinforcement learning to generate an optimized collision avoidance strategy

[0107] Refer to Figure 4 , and establish a reinforcement learning framework to optimize the collision avoidance strategy according to the combined optimization direction. The specific steps include:

[0108] 5-1: Select the input state s in reinforcement learning as the local information, the information of the intruding aircraft, the interaction information between the two, and the predicted local target vertical speed obtained by rough prediction. The relevant information of the local aircraft and the intruding aircraft specifically includes: the initial vertical height H of the local aircraft O and the initial vertical height rate The initial vertical height H of the intruding aircraft I , the initial vertical height rate and the sensitivity level Sensitivity_levels, the horizontal relative distance R between the intruding aircraft and the local aircraft, the rate of change of the horizontal distance of the intruding aircraft relative to the local aircraft

[0109] 5-2: Determine the output action a in reinforcement learning and the value range of its various parameters, including: the vertical acceleration Accel taken by the local aircraft, and its value range is an integer from 4 to 11 ft / min; and the expected reaction time Delay of the pilot. To give the pilot enough reaction time, the value range of τ is between 15 and 35 seconds, so its value range should not be greater than τ and is an integer from 5 to 15 s.

[0110] 5-3: Use the Actor-Critic architecture for reinforcement learning training. The goal of the Actor network is to adjust the θ π parameter to learn a policy π(a t |s t ; θ π ), and the Critic is used to evaluate the value of the state-action pair, that is, the state-action value function Q(s t , a t ; θ q ).

[0111] 5-4: Determine the content of the reward function in reinforcement learning according to the policy joint optimization direction, including: the local vertical acceleration part R Accel , the time planning part R Duration and the vertical interval threshold comparison part R △H , and the design of the reward function is specifically expressed as follows:

[0112] R Accel =-0.05·e 0.45(|Accel|-4) ;

[0113]

[0114] R △H =0.3·(log(1 + △H)-log(1 + H LIMIT ));

[0115]

[0116] R Total = R Accel + R Duration + R △H + R Penalty 。

[0117] Among them, the value range of H LIMIT is 300 - 700 ft, which will be selected according to the altitude levels of different aircraft flights. In addition, a penalty part R Penalty is added to ensure that the selected parameters will not collide and the state is below the vertical separation threshold.

[0118] 5 - 5: During the training process, the state s is input into the policy generated by the Actor network to obtain the corresponding action a, and the obtained action a is input into the world model, that is, the above-mentioned collision model of the own aircraft and the intruder aircraft, to obtain the next state. This process loops continuously until the own aircraft reaches the target vertical speed and achieves collision avoidance.

[0119] 5 - 6: Use policy gradient to update the policy of the Actor. The specific gradient update formula is as follows:

[0120]

[0121] Among them, A t is the advantage function, which measures the advantage of the current action relative to other actions by subtracting the baseline (such as the state value function where R t represents the expected return), and can be calculated by the Critic network designed above. The specific calculation is shown in the following formula:

[0122] A t = Q(s t , a t ; θ q ) - V(s t ), Q(s t , a t ; θ q ) ≈ Ε[R t |s t , a t .

[0123] At the same time, the gap between the currently estimated state value and the true return based on the current state and subsequent states is quantified by the Temporal Difference (TD) Error, and then the state-action pair value function Q(s t , a t ) in the Critic is updated. The specific expressions of the TD Error and the update are shown in the following formula:

[0124] δ t = R t + γQ(s t+1 , a t+1 ) - Q(s t , a t );

[0125] Q(s t , a t ) ← Q(s t , a t ) + α·δ t .

[0126] Where δ t is the TD Error, γ is the weight factor used to control the weight of future rewards, 0 ≤ γ ≤ 1, Q(s t , a t ) and Q(s t+1 , a t+1 ) are the current state - action value function and the state - action value function at the next moment, respectively.

[0127] During training, the loss function of the Actor network is The loss function of the Critic network is L Critic = Ε[(δ t ) 2 , and the two cooperate with each other during training to minimize their loss functions, so as to obtain a collision avoidance strategy with higher safety and stronger flexibility.

[0128] Based on the above embodiments, a civil aircraft collision avoidance strategy optimization simulation experiment is carried out in the present invention.

[0129] Assume that in this experiment, the initial vertical height H O = 10030 ft and the initial vertical height rate The initial vertical height H I = 12380 ft, the initial vertical height rate The horizontal relative distance R between the intruder aircraft and the own aircraft = 30380 ft, the horizontal distance change rate of the intruder aircraft relative to the own aircraft

[0130] Based on the prediction model, the strategy taken is nominal descent, and the corresponding target vertical speed of the own aircraft is - 1500 ft / min. Substituting it into the collision avoidance strategy optimization network, the vertical acceleration Accel = - 7 ft / s 2, the pilot's expected reaction time Delay = 7s, and the vertical separation distance △H = 1718.58ft. Although there is a loss in the vertical separation distance compared to the original strategy, it still ensures a value far greater than the threshold. And based on the original Accel = -8ft / s 2 , it effectively improves flexibility, stability, and safety on the basis of Delay = 2.5s.

[0131] In summary, the present invention establishes a collision model between the own aircraft and the intruding aircraft, gives an alarm and a rough prediction of the collision avoidance strategy, and obtains the target vertical speed that the own aircraft needs to reach when avoiding the collision. However, by restricting and optimizing the vertical acceleration of the own aircraft, the stability of the own aircraft's flight is effectively improved; the optimization of the parameter of the pilot's expected reaction time helps to expand the safety threshold for collision avoidance; in addition, the vertical separation is also restricted and maximized, ensuring the strongest safety performance on the basis of flight stability; a corresponding reward function is set according to the above optimization directions, and then the reinforcement learning with the Actor-Critic structure is used for training and optimization. Through the training and optimization of reinforcement learning, the present invention effectively improves the flexibility of the own aircraft in dealing with various flight situations in the collision avoidance strategy. The present invention provides an effective way for the optimization of the civil aircraft collision avoidance strategy.

[0132] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for optimizing collision escape strategy of a civil aircraft, characterized in that: The parameters related to collision resolution are optimized by using a reinforcement learning method to generate an optimized collision resolution strategy, which specifically includes the following steps: Step 1: Information Collection 1-1: Get the initial vertical height and initial vertical height rate of the aircraft; 1-2: Obtain the sensitivity level, initial vertical height and initial vertical height rate of the intruder aircraft; 1-3: Obtain the horizontal relative distance between the intruder and the own aircraft and the rate of change of the horizontal distance between the intruder and the own aircraft; Step 2: Build the collision model 2-1: Use the position relationship between the own aircraft and the intruder aircraft to build a collision model between the own aircraft and the intruder aircraft; 2-2: Based on the horizontal relative distance R between the own aircraft and the intruder aircraft, and the rate of change of the horizontal distance of the intruder aircraft relative to the own aircraft Calculate the time τ when the relative distance between the two aircraft is the smallest, and use the sensitivity level of the intruder as the correction parameter D m Correct τ and get the corrected The value range of τ is between 15 and 35 seconds; 2-3: Input the acquired information of the host and the intruder into the collision model, select τ and The value that makes the vertical separation smaller is taken as the moment when the relative distance between the two aircraft is the smallest, that is, the time when the aircraft and the intruder aircraft are closest to each other; Step 3: Determine the collision resolution strategy 3-1: Establish the motion prediction model of the host aircraft and the intruder aircraft; 3-2: Use the motion prediction model to calculate its collision escape strategy; Step 4: Construct strategy joint optimization direction 4-1: Minimize the vertical acceleration of the aircraft to ensure the stability of the aircraft's flight; 4-2: Maximize the pilot's expected reaction time to ensure that sufficient reaction time is reserved; 4-3: Maximize the closest vertical separation distance to obtain the safest collision escape strategy parameters under the current strategy; Step 5: Use reinforcement learning method to optimize the collision escape strategy and generate an optimized collision escape strategy.

2. The method for optimizing collision escape strategies for civil aircraft according to claim 1, characterized in that: The steps of step 1 to obtain the information of the local machine and the intruder machine are as follows: Step 1: The machine sends out inquiries every 1 second and monitors the received information, collects and classifies the monitored data and stores it in the data input structure; Step 2: Filter the input data, obtain the data priority according to the data timestamp and judgment conditions, and process the intruder information with high priority first. The judgment conditions include: 1) the smaller the distance between the intruder and the host, the higher the priority; 2) the address information of the intruder and the host are in the same system, the higher the priority; 3) the data information obtained by monitoring is complete and the verification is accurate, the higher the priority; Step 3: Based on the query information sent by the aircraft, the response information of the intruder aircraft and the time difference between the two, the corresponding information of altitude, altitude change rate, horizontal relative distance and horizontal relative distance change rate are obtained.

3. The method for optimizing collision escape strategies for civil aircraft according to claim 1, characterized in that: The time τ when the relative distance between the two aircraft in step 2 is the smallest is calculated by the following formula: The modified Calculated by the following formula: Where D m Correction parameter introduced for the sensitivity level of the intruder.

4. The method for optimizing collision escape strategies for civil aircraft according to claim 1, characterized in that: The step 3 specifically includes: Step 1: Select the acceleration and pilot reaction time of the vertical maneuver method under different collision release states during prediction; Step 2: Based on the vertical flight model of the aircraft and the intruder established in the prediction process, calculate the altitude values ​​of the aircraft and the intruder at the closest point respectively, and obtain the shortest vertical separation distance between the two. Step 3: Substitute the parameters under different collision escape states into the prediction model, calculate the minimum vertical separation distance, and obtain a preliminary collision escape strategy based on the comparison between the minimum vertical separation distance and the safety threshold value of the vertical separation distance, and determine the target vertical speed of the aircraft that needs to be achieved under this strategy.

5. The method for optimizing collision escape strategies for civil aircraft according to claim 1, characterized in that: The step 5 specifically includes: Step 1: Determine the input state s in reinforcement learning and its related information between the local machine and the intruder machine; Step 2: Determine the value range of the output action a and its parameters in the reinforcement learning, wherein the parameters include: the acceleration taken by the aircraft and the expected reaction time of the pilot; Step 3: Use the Actor-Critic architecture for reinforcement learning training, where the goal of the Actor network is to adjust θ π Parameters learn a policy π(a t |s t θ π ), Critic is used to evaluate the value of the state-action pair, namely the state-action value function Q(s t ,a t θ q ); Step 4: Determine the content of the reward function in the reinforcement learning according to the strategy joint optimization direction, wherein the content of the reward function includes: the vertical acceleration part of the aircraft, the time planning part, and the vertical interval threshold comparison part; Step 5: During the training process, the state s is input into the Actor network, the generated strategy obtains the corresponding action a, and the obtained action a is input into the world model, that is, the collision model between the local machine and the intruder to obtain the next state This cycle continues until the aircraft reaches the target vertical speed and is free from collision; Step 6: Use policy gradient to update the Actor's policy, and use Temporal Difference (TD) Error to update the Critic's state-action value function, and work together to obtain an optimized collision resolution strategy.

6. The method for optimizing collision escape strategies for civil aircraft according to claim 5, characterized in that: The state s is the information of the own aircraft, the intruder aircraft, the interaction information between the two, and the predicted vertical speed of the own aircraft target.

7. The method for optimizing collision escape strategies for civil aircraft according to claim 5, characterized in that: The relevant information of the host and the intruder specifically includes: the initial vertical height H of the host O and initial vertical height rate Initial vertical height of the intruder H I , initial vertical height rate and sensitivity level Sensitivity_levels; the horizontal relative distance R between the intruder and the own aircraft, the rate of change of the horizontal distance of the intruder relative to the own aircraft