Automatic landing control method and device based on hierarchical reinforcement learning and medium

By using a hierarchical reinforcement learning network to make landing decisions and control the flight of the aircraft, the adaptability and robustness problems of modern automatic landing technology in complex environments are solved, and safe and accurate landing is achieved under changing conditions.

CN120631017APending Publication Date: 2025-09-12BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510760383.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Modern automatic landing technology has poor adaptability and robustness in complex and changing weather and environmental conditions and relies on precise sensors and complex control algorithms.

Method used

A hierarchical reinforcement learning network is adopted, and attention weights are assigned to environmental parameters and aircraft state parameters through expert experience empowerment. Upper and lower reinforcement learning networks are constructed to perform landing decisions and flight control respectively, reducing the accuracy requirements for guidance equipment and improving adaptability.

Benefits of technology

Achieve safe and accurate landing of aircraft in complex environments, improve the adaptability and robustness of the automatic landing process, reduce the accuracy requirements of guidance equipment, and enhance adaptive control capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631017A_ABST
    Figure CN120631017A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic landing control method and device based on hierarchical reinforcement learning and a medium, and relates to the technical field of unmanned aerial vehicles, and the method comprises the steps: employing an expert experience weighting device to add attention weights to environmental parameters and aircraft state parameters, and obtaining the weighted environmental parameters and aircraft state parameters; according to the empowered environment parameters, the empowered aircraft state parameters and the decision memory, using an upper reinforcement learning network to make a landing decision to obtain a decision result, and updating the decision memory according to the decision result; and according to a decision result, performing flight control on the aircraft through a lower-layer reinforcement learning network by using the weighted environment parameters and the weighted aircraft state parameters to complete landing. The aircraft landing decision is made through the upper reinforcement learning network, the aircraft is adaptively controlled through the lower reinforcement learning network, the aircraft is guided to land automatically, and the adaptability and robustness of the aircraft in the automatic landing process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of drone technology, and in particular to an automatic landing control method, device, and medium based on hierarchical reinforcement learning. Background Art

[0002] Modern autoland technology has made significant progress and is widely used in the aviation sector, particularly on large civil airliners and drones. Autoland technology, which fully controls the aircraft's landing flight through an onboard automated flight system, significantly improves landing accuracy. As a key technology in the aviation sector, the safety and efficiency of autoland technology are crucial for ensuring aviation safety and improving air transport efficiency.

[0003] However, modern automatic landing technology still has some limitations and shortcomings. Traditional automatic landing systems rely on precise sensor data and complex control algorithms, such as the accuracy requirements of guidance equipment, dependence on external equipment and signals, etc. Under complex and changeable weather and environmental conditions, the adaptability and robustness of automatic landing technology are poor. Summary of the Invention

[0004] The purpose of this application is to provide an automatic landing control method, device and medium based on hierarchical reinforcement learning, which can improve the adaptability and robustness of the automatic landing process of the aircraft.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In the first aspect, the present application provides an automatic landing control method based on hierarchical reinforcement learning, the automatic landing control method based on hierarchical reinforcement learning comprising: obtaining environmental parameters and aircraft state parameters; constructing a hierarchical reinforcement learning network; the hierarchical reinforcement learning network comprising: an upper reinforcement learning network and a lower reinforcement learning network; the upper reinforcement learning network is used for landing decision-making; the lower reinforcement learning network is used for flight control; using an expert experience weighter to assign attention weights to the environmental parameters and the aircraft state parameters to obtain weighted environmental parameters and weighted aircraft state parameters; the expert experience weighter The verification weighter is a three-layer fully connected neural network model; based on the weighted environmental parameters, the weighted aircraft state parameters and the decision memory, the upper-layer reinforcement learning network is used to make a landing decision to obtain a decision result, and the decision memory is updated according to the decision result; the decision memory is used to store the decision result; the decision result is: landing or go-around; based on the decision result, the weighted environmental parameters and the weighted aircraft state parameters are used to control the flight of the aircraft through the lower-layer reinforcement learning network; it is determined whether the aircraft has landed, and if so, the landing task is completed, otherwise the environmental parameters and aircraft state parameters are re-acquired.

[0007] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described automatic landing control method based on hierarchical reinforcement learning.

[0008] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned automatic landing control method based on hierarchical reinforcement learning.

[0009] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0010] This application improves the computational efficiency of the hierarchical reinforcement learning network by assigning attention weights to environmental parameters and aircraft state parameters through an expert empowerer; makes landing decisions for the aircraft through an upper-layer reinforcement learning network, and controls the flight of the aircraft through a lower-layer reinforcement learning network, guiding the aircraft to land automatically, thereby reducing the accuracy requirements for the guidance equipment, and can adaptively control the aircraft under complex and changeable weather and environmental conditions, achieving safe and accurate landing of the aircraft, and improving the adaptability and robustness of the aircraft's automatic landing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0012] Figure 1 A schematic diagram of the process of an automatic landing control method based on hierarchical reinforcement learning provided in an embodiment of the present application Figure 1 .

[0013] Figure 2 A schematic diagram of the process of an automatic landing control method based on hierarchical reinforcement learning provided in an embodiment of the present application Figure 2 .

[0014] Figure 3 Schematic diagram of the training process of the upper-layer reinforcement learning network and the lower-layer reinforcement learning network provided in the embodiment of the present application.

[0015] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0017] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0018] Example 1, as Figure 1-Figure 2 As shown, this embodiment provides an automatic landing control method based on hierarchical reinforcement learning, including:

[0019] S1. Obtain environmental parameters and aircraft status parameters.

[0020] Furthermore, the environmental parameters include at least one of: landing runway direction, runway length, runway position, wind speed and direction, rain and snow conditions, and route occupancy.

[0021] Furthermore, the aircraft status parameter includes at least any one of: altitude, speed, heading, position, attitude angle, angle of attack, aircraft status, and fuel status.

[0022] S2. Construct a hierarchical reinforcement learning network. The hierarchical reinforcement learning network includes an upper-layer reinforcement learning network and a lower-layer reinforcement learning network. The upper-layer reinforcement learning network is used for landing decision making, while the lower-layer reinforcement learning network is used for flight control.

[0023] S3. Use the expert experience weighter to assign attention weights to the environmental parameters and aircraft state parameters to obtain weighted environmental parameters and weighted aircraft state parameters; the expert experience weighter is a three-layer fully connected neural network model.

[0024] S4. Based on the weighted environmental parameters, weighted aircraft state parameters, and decision memory, the upper-level reinforcement learning network is used to make a landing decision, obtain a decision result, and update the decision memory based on the decision result; the decision memory is used to store the decision result; the decision result is: landing or go-around.

[0025] S5. Based on the decision result, the weighted environmental parameters and the weighted aircraft state parameters are used to control the flight of the aircraft through the lower-level reinforcement learning network.

[0026] Step S5 specifically includes:

[0027] S51. When the decision result is landing, based on the weighted environmental parameters, the weighted aircraft state parameters and the decision result, the lower-level reinforcement learning network is used to calculate the control amount of the aircraft at the current moment using the first network parameters and the first reward function.

[0028] S52. When the decision result is a go-around, the control amount of the aircraft at the current moment is calculated using the lower-level reinforcement learning network with the second network parameters and the second reward function based on the weighted environmental parameters, the weighted aircraft state parameters and the decision result.

[0029] S53. Perform flight control on the aircraft based on the current control value of the aircraft.

[0030] In actual application, when the upper-layer reinforcement learning network determines that landing is possible, the landing process continues, and the lower-layer reinforcement learning network calculates the control amount of the aircraft at the current moment according to the first network parameters and the first reward function; otherwise, it switches to the go-around task, and the network parameters of the lower-layer reinforcement learning network are switched to the second network parameters and the second reward function to calculate the control amount of the aircraft at the current moment.

[0031] S6. Determine whether the aircraft has landed. If so, complete the landing mission. Otherwise, reacquire environmental parameters and aircraft status parameters.

[0032] In actual application, after starting the automatic landing mission, the environmental parameters and aircraft state parameters are first obtained. The environmental parameters and flight parameters are input into the expert experience empowerment device. The attention weights of each parameter are annotated according to the stored expert experience. The weighted parameters (weighted environmental parameters and weighted aircraft state parameters) are obtained. Together with the decision memory, they are input into the upper-layer reinforcement learning network to make the landing decision, and the decision result memory is passed to the lower-layer network. When the upper-layer reinforcement learning network determines that the landing is normal (the decision result is landing by landing), the landing process continues. Otherwise, it switches to the missed approach mission and switches the network parameters and reward function of the lower-layer reinforcement learning network. The lower-layer reinforcement learning network receives the weighted environmental parameters and weighted aircraft state parameters as well as the decision result of the upper-layer reinforcement learning network, completes the calculation according to the specific task, outputs the control amount of the aircraft at the current moment, and passes it to the flight control system. After completing the calculation at this moment, it is determined whether the aircraft has landed. If the landing is not completed, it enters the next cycle, re-acquires the environmental parameters and aircraft state parameters, and performs hierarchical reinforcement learning network landing decision and flight control. Otherwise, the program is terminated and the automatic landing mission is completed.

[0033] Furthermore, if Figure 3 As shown in Figure 2, the training process of the upper reinforcement learning network is as follows:

[0034] 1) Obtain empirical environmental parameters, empirical aircraft state parameters and empirical decision results from the expert experience database.

[0035] 2) Use the expert experience weighter to assign attention weights to the experience environment parameters and the experience aircraft state parameters to obtain the weighted experience environment parameters and the weighted experience aircraft state parameters.

[0036] 3) Taking the weighted empirical environment parameters and the weighted empirical aircraft state parameters as input, the empirical decision results as output, and the decision memory as the memory pool, the upper-layer reinforcement learning network is trained with the goal of ensuring that the number of iteration steps is greater than the preset number of steps and the success rate of the decision results is greater than the threshold, a trained upper-layer reinforcement learning network is obtained, and the decision memory is updated at the same time.

[0037] In practical applications, when the upper-layer reinforcement learning network is trained, an expert experience database is first imported to train the upper-layer decision network. A set of environmental parameters and aircraft state parameters are extracted from the expert experience database as state parameters and fed into the expert experience weighter. This weighter is a three-layer fully connected neural network designed to initially extract state parameter features and output weights for each parameter. The weighted empirical environmental parameters and weighted empirical aircraft state parameters are then fed into the upper-layer reinforcement learning network, which then outputs a decision result. This result is then compared with the expert's empirical decision results. Backpropagation is then used to update the parameters of the expert experience weighter and the upper-layer reinforcement learning network. The network then determines whether the number of iterations exceeds the target number and whether the success rate of the decision results exceeds a threshold. If any of these conditions are not met, the network returns to the expert experience database for training using a new set of data (the empirical environmental parameters and empirical aircraft state parameters). If all conditions are met, training proceeds to the lower-layer network.

[0038] Furthermore, if Figure 3 As shown in Figure 2, the training process of the lower reinforcement learning network is as follows:

[0039] 1) Use the flight simulation platform to obtain the simulation environment parameters, simulated aircraft state parameters and the control parameters of the simulated aircraft at the current moment.

[0040] 2) Use the expert experience weighter to assign attention weights to the simulated environment parameters and simulated aircraft state parameters to obtain the weighted simulated environment parameters and the weighted simulated aircraft state parameters.

[0041] 3) The weighted simulation environment parameters and the weighted simulation aircraft state parameters are used as state parameters, the control amount of the simulated aircraft at the current moment is used as the action parameter, and the benefit corresponding to the action parameter is used as the reward function. The goal is to minimize the loss function, and the lower-layer reinforcement learning network is trained to obtain a trained lower-layer reinforcement learning network.

[0042] Furthermore, the loss function is a minimum variance loss function.

[0043] Furthermore, the loss function is calculated as follows:

[0044]

[0045] In the formula, Loss is the loss function, n is the number of samples, y i The control reward value of the aircraft at the current moment i calculated by the flight simulation platform, is the predicted reward value of the control amount of the aircraft at the current moment i.

[0046] In the actual application process, first connect to the flight simulation platform, obtain the simulation environment parameters and the simulated aircraft state parameters, send them to the lower-level reinforcement learning network for training, and calculate the expected benefit function R'(S,A). Among them, S is a vector composed of state parameters, A is the control amount of the aircraft, and R'(S,A) represents the expected reward function value corresponding to executing action A under the S state. Select the action with the largest reward function value to execute, send the control amount of the aircraft to the simulator, obtain the actual feedback of the simulator (the control amount of the simulated aircraft at the current moment), calculate the actual reward function R(S,A), and convert the corresponding<S,A,R,S’> The tuple (S is the current state, A is the action, R is the reward value, and S' is the next state) is stored in the experience pool (decision memory). A loss function is used to calculate the difference between the current Q network output and the target network output. Backpropagation is then used to update the network parameters using the loss function. The network parameters are then checked to see if the number of iterations exceeds the target number and if the success rate of the decision result exceeds a threshold. If any of the conditions are not met, a new set of training is performed. If all conditions are met, the network is saved and a check is performed to see if both the landing and go-around tasks have been trained. If either task is incomplete, the state of that task is switched, the simulation platform scenario and reward function are changed, and a new round of training is performed. When both tasks have been trained and both networks have been saved, the training task is complete.

[0047] The technical effects of this application are as follows:

[0048] This application improves the computational efficiency of the hierarchical reinforcement learning network by assigning attention weights to environmental parameters and aircraft state parameters through an expert empowerer; uses the hierarchical reinforcement learning network to decompose complex tasks into two levels of subtasks: go-around decision-making and flight control. For each subtask, the corresponding reinforcement learning algorithm and strategy are trained with the help of an expert database for various situations such as crosswinds, runway icing, route occupancy, and fuel shortage during the automatic landing process, and reinforcement learning is performed independently, effectively reducing the complexity and difficulty of learning. When the hierarchical reinforcement learning network is used to guide the aircraft to perform automatic landing control, the upper-level reinforcement learning network makes landing decisions for the aircraft, and the lower-level reinforcement learning network performs flight control of the aircraft to guide the aircraft to land automatically, reducing the accuracy requirements for the guidance equipment. It can adaptively control the aircraft under complex and changeable weather and environmental conditions, achieve safe and accurate landing of the aircraft, improve the adaptability and robustness of the aircraft's automatic landing process, and bring stronger adaptability, optimization performance, and explainability to the automatic landing system.

[0049] Example 2: This application also provides a computer device, which can be a server or a terminal, and its internal structure diagram can be as follows: Figure 4 As shown. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, memory and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store processing data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an automatic landing control method based on hierarchical reinforcement learning is implemented.

[0050] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0051] In embodiment 3, the present application further provides a computer-readable storage medium storing a computer program, which implements the above methods when executed by a processor.

[0052] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0053] All actions of acquiring signals, information or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0054] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0055] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. An automatic landing control method based on hierarchical reinforcement learning, characterized in that: The automatic landing control method based on hierarchical reinforcement learning includes: Obtain environmental parameters and aircraft status parameters; Constructing a hierarchical reinforcement learning network; the hierarchical reinforcement learning network includes: an upper reinforcement learning network and a lower reinforcement learning network; the upper reinforcement learning network is used for landing decision-making; the lower reinforcement learning network is used for flight control; Using an expert experience weighter to assign attention weights to environmental parameters and aircraft state parameters to obtain weighted environmental parameters and weighted aircraft state parameters; the expert experience weighter is a three-layer fully connected neural network model; Based on the weighted environmental parameters, the weighted aircraft state parameters, and the decision memory, a landing decision is made using the upper-layer reinforcement learning network to obtain a decision result, and the decision memory is updated according to the decision result; the decision memory is used to store the decision result; the decision result is: landing or go-around; According to the decision results, the weighted environmental parameters and weighted aircraft state parameters are used to control the flight of the aircraft through the lower-level reinforcement learning network; Determine whether the aircraft has landed. If so, complete the landing mission. Otherwise, reacquire the environmental parameters and aircraft status parameters.

2. The automatic landing control method based on hierarchical reinforcement learning according to claim 1, characterized in that: Based on the decision results, the weighted environmental parameters and weighted aircraft state parameters are used to control the aircraft flight through the lower-level reinforcement learning network, specifically including: When the decision result is landing, the control amount of the aircraft at the current moment is calculated using the lower-layer reinforcement learning network with the first network parameters and the first reward function based on the weighted environmental parameters, the weighted aircraft state parameters, and the decision result; When the decision result is a go-around, the control amount of the aircraft at the current moment is calculated using the lower-layer reinforcement learning network with the second network parameters and the second reward function based on the weighted environmental parameters, the weighted aircraft state parameters and the decision result; The aircraft is controlled based on its current control parameters.

3. The automatic landing control method based on hierarchical reinforcement learning according to claim 1, characterized in that: The training process of the upper reinforcement learning network is as follows: Obtaining empirical environmental parameters, empirical aircraft state parameters and empirical decision results from the expert experience database; The expert experience weighter is used to assign attention weights to the experience environment parameters and the experience aircraft state parameters to obtain the weighted experience environment parameters and the weighted experience aircraft state parameters; Taking the weighted empirical environment parameters and the weighted empirical aircraft state parameters as input, the empirical decision results as output, and the decision memory as the memory pool, the upper-layer reinforcement learning network is trained with the goal of having more iteration steps than the preset number of steps and a success rate of the decision results greater than the threshold, a trained upper-layer reinforcement learning network is obtained, and the decision memory is updated at the same time.

4. The automatic landing control method based on hierarchical reinforcement learning according to claim 1, characterized in that: The training process of the lower layer reinforcement learning network is as follows: Using the flight simulation platform to obtain simulation environment parameters, simulation aircraft state parameters and the current control parameters of the simulation aircraft; Using the expert experience weighter to assign attention weights to the simulated environment parameters and simulated aircraft state parameters, the weighted simulated environment parameters and weighted simulated aircraft state parameters are obtained; The weighted simulation environment parameters and the weighted simulation aircraft state parameters are used as state parameters, the control amount of the simulated aircraft at the current moment is used as action parameters, the benefits corresponding to the action parameters are used as reward functions, and the loss function is minimized as the goal. The lower-layer reinforcement learning network is trained to obtain a trained lower-layer reinforcement learning network.

5. The automatic landing control method based on hierarchical reinforcement learning according to claim 4, characterized in that: The loss function is a minimum variance loss function.

6. The automatic landing control method based on hierarchical reinforcement learning according to claim 4, characterized in that: The calculation formula of the loss function is as follows: In the formula, Loss is the loss function, n is the number of samples, y i The control reward value of the aircraft at the current moment i calculated by the flight simulation platform, is the predicted reward value of the control amount of the aircraft at the current moment i.

7. The automatic landing control method based on hierarchical reinforcement learning according to claim 1, characterized in that: The environmental parameters include at least one of: landing runway direction, runway length, runway position, wind speed and direction, rain and snow conditions, and route occupancy.

8. The automatic landing control method based on hierarchical reinforcement learning according to claim 1, characterized in that The aircraft status parameters include at least any one of: altitude, speed, heading, position, attitude angle, angle of attack, aircraft status and fuel status.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the automatic landing control method based on hierarchical reinforcement learning according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the automatic landing control method based on hierarchical reinforcement learning described in any one of claims 1 to 8 is implemented.