A flight scheduling model pre-training method based on multimodal data

Through multimodal data pre-training methods, combined with the screen recording and voice data of the control automation system, an air traffic control reinforcement learning evaluation model was constructed, which solved the problems of slow and unstable training of the deep reinforcement learning flight allocation model and achieved more efficient and stable flight allocation model training.

CN119132112BActive Publication Date: 2025-10-03THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411011909.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2025-10-03
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

The deep reinforcement learning flight scheduling model has slow training speed, low data utilization and instability, resulting in unstable flight scheduling model training results, which makes it difficult to meet the requirements of aircraft safety flight in complex environments.

Method used

Through multimodal data pre-training methods, combined with the screen recording data and voice data of the control automation system, deep neural networks and natural language processing technologies are used to extract aircraft position, speed, heading and control intentions, build an air traffic control reinforcement learning evaluation model, and conduct supervised pre-training to improve data utilization and model stability.

Benefits of technology

The training speed of the flight scheduling model has been accelerated, the data utilization rate has been improved, the training results of the flight scheduling model have been made closer to the normal controller behavior, and the stability of the model and the training efficiency have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119132112B_ABST
    Figure CN119132112B_ABST
Patent Text Reader

Abstract

The present invention provides a flight scheduling model pre-training method based on multimodal data, comprising: step 1, obtaining screen recording data from a control automation system to obtain aircraft position information; obtaining the aircraft's speed, heading, altitude, and flight number; step 2, obtaining voice data corresponding to the screen recording data from the control automation system, converting the voice data into text, and extracting the control intent; step 3, matching the aircraft's position, speed, heading, altitude, and control intent by flight number to form training data, scoring the data, and training an air traffic control reinforcement learning evaluation model; and step 4, training an air traffic control deep reinforcement learning flight scheduling model. The present invention can pre-train a flight scheduling model by extracting aircraft position information and flight number from the control screen recording data, control intent from the control voice data, and performing manual evaluation, thereby accelerating the training speed of the flight scheduling model and optimizing the training effect of the flight scheduling model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of civil aviation air traffic control, and in particular to a flight dispatch model pre-training method based on multimodal data. Background Art

[0002] With the rapid development of the aviation industry, demand for daily travel and cargo transportation is rapidly increasing. Busy airports are experiencing over 1,000 flights per day. To ensure safe flight, certain efficiency sacrifices are necessary, leading to other impacts such as increased average flight times, deviations from standard arrival and departure routes, and increased pressure on air traffic controllers. Reinforcement learning, a key branch of machine learning, essentially describes and addresses the problem of how intelligent agents learn strategies to maximize rewards or achieve specific goals through their interactions with their environment. Therefore, we are exploring the use of deep reinforcement learning flight scheduling models for autonomous control.

[0003] Deep reinforcement learning methods suffer from numerous drawbacks, including slow training speeds, low data utilization, and unstable training results. Deep reinforcement learning typically requires extensive interaction and trial-and-error to learn effective strategies, especially when faced with complex environments and tasks. This can result in significantly higher data sample sizes than other machine learning methods. Deep reinforcement learning experiences instabilities during model training, such as large fluctuations in the learning curve, slow convergence, or even no convergence, due to factors such as the dynamic balance between exploration and exploitation, sparse or delayed rewards, and environmental non-stationarity. This is particularly true for flight scheduling model training, which has a high trial-and-error cost. Therefore, pre-training the air traffic control deep reinforcement learning flight scheduling model with multi-modal data to allow it to absorb domain expert knowledge is of practical significance for improving training speed and enhancing training stability. Summary of the Invention

[0004] Purpose of the invention: The purpose of this method is to improve the training speed of the air traffic control deep reinforcement learning flight allocation model, improve the data utilization rate of the air traffic control deep reinforcement learning flight allocation model, and enhance the stability of the air traffic control deep reinforcement model.

[0005] This method is used in the training process of the deep reinforcement learning flight allocation model for air traffic control. It extracts historical control experience through historical data in multiple modal forms, trains an evaluation model, and pre-trains the deep reinforcement learning flight allocation model for air traffic control. Combined with the evaluation model, it improves the training speed, improves data utilization, and enhances model stability.

[0006] To achieve the above objectives, the present invention discloses a flight dispatch model pre-training method based on multimodal data, comprising the following steps:

[0007] Step 1: Obtain screen recording data from the control automation system, perform coordinate transformation operations through the deep neural network target detection model (reference: YOLOv7, Wang CY, Bochkovskiy A, Liao HY M. YOLOv7: Trainable bag-of-freebiessets new state-of-the-art for real-time object detectors[C] / / Proceedings ofthe IEEE / CVF conference on computer vision and pattern recognition.2023:7464-7475.) to obtain the position information of the aircraft on the screen; obtain the aircraft's speed, heading, altitude, flight number, and part of the control intention through the text recognition detection model (reference: PGNet, Wang P, Zhang C, Qi F, et al. Pgnet: Real-time arbitrarily-shaped textspotting with point gathering network[C] / / Proceedings of the AAAI Conference on Artificial Intelligence.2021,35(4):2782-2790.) or plan data matching;

[0008] Step 2: Obtain the speech data corresponding to the screen recording data of the control automation system. Convert the speech data into text using a deep neural network speech recognition model (reference: Whisper, Radford A, Kim JW, Xu T, et al. Robust speech recognition via large-scale weak supervision [C] / / International conference on machine learning. PMLR, 2023: 28492-28518.). Then extract the control intent using an intent extraction model (reference: CII-BERT, Yu Yaoyao, Wang Xuan, Chen Si, et al. Joint extraction of air traffic control information and intent based on natural language processing [J]. Journal of Xihua University (Natural Science Edition), 2024.) and combine it with the control intent in the screen recording data.

[0009] Step 3: The aircraft's position, speed, heading, altitude, and control intent are matched by flight number to form training data. The data is scored (which can be manually scored) to train an air traffic control reinforcement learning evaluation model. The air traffic control reinforcement learning evaluation model includes a five-layer fully connected neural network. The input is the aircraft's position, speed, heading, altitude, and control intent, and the output is the data score.

[0010] Step 4: Use the training data and evaluation model to train the air traffic control deep reinforcement learning flight allocation model (reference: Liu Z, Xu Q, Shi Y, et al. Generation method of control strategy for aircraftsbased on hierarchical reinforcement learning[C] / / Artificial intelligence in China: proceedings of the 3rd international conference on artificial intelligence in China. Singapore: Springer Singapore, 2022: 109-116.).

[0011] In step 1, the control automation system screen recording data includes flight number FN, altitude ALT, speed SPD, heading HDG, position POS and altitude and heading adjustment intentions.

[0012] In step 1, the formula for the coordinate transformation operation is:

[0013] P xy = T(P uv ) (1)

[0014] Among them, P xy Represents the coordinate information converted to longitude and latitude, T represents the plane coordinate conversion function, P uv Represents the target pixel center point output by the deep neural network target detection model used in step 1.

[0015] In step 1, the plan data matching is performed by calculating the track position, traversing the plan data, and searching for the plan of the nearest route. The formula is:

[0016] P true =Min(Dis(P xy ,P)) (2)

[0017] Among them, P trueRefers to the matching plan data, Min refers to the minimum function, Dis refers to the distance calculation function, and P refers to all plan data.

[0018] In step 2, the speech data is converted into text through the deep neural network speech recognition model. The formula is:

[0019] W=ASR(A) (3)

[0020] Where W is the converted text information, ASR represents the deep neural network speech recognition model, and A represents the original input speech.

[0021] In step 2, the voice data corresponds to the screen recording data time. The data alignment method uses the voice data time as the benchmark and the flight number as the matching condition to perform threshold judgment on the altitude, speed and heading values. The formula is:

[0022]

[0023] Match represents whether the time of voice and data is aligned. 1 represents alignment and 0 represents misalignment. new Represents the current altitude, speed or heading value, V old represents the altitude, speed or heading value at the last sampling moment, and θ represents the threshold.

[0024] In step 2, the control intent includes the flight number, altitude adjustment, heading adjustment, and speed adjustment intent. The control intent in the screen recording data and the control intent in the voice data are combined into the formula:

[0025]

[0026] A represents the final intention, A Voice Represents the intent extracted from the voice data, A Video Represents the intent extracted from the screen recording data. Represents empty intent.

[0027] In step 3, the altitude ALT, speed SPD, heading HDG, and position POS obtained in step 1 are normalized to their maximum and minimum values ​​using the following formula:

[0028]

[0029] Where X′ represents the normalized data, X represents the data before normalization, and X min Represents the minimum value of all data in this category, X max Represents the maximum value among all data in this category.

[0030] In step 3, the data except the flight number in the control intent in step 2 is parsed into a one-hot encoding. The format of the one-hot encoding is:

[0031] [o1,o2,o3,…,o n ] (7)

[0032] Among them, if o i =1, then the rest o j =0, i and j are 1 to n and j≠i, n represents the number of categories of control intention data; n The one-hot encoding of the nth category of data in the regulatory intent;

[0033] The one-hot encoded data and the normalized data are concatenated into a matrix as the input data for the air traffic control reinforcement learning evaluation model. The concatenation order is: [LON, LAT, ALT, SPD, HDG, O], where LON and LAT are the longitude and latitude in the position information, respectively, and O is the one-hot encoded control intent data. The normalized data in step 1 and the one-hot encoded data in step 2 are preliminarily evaluated and scored by the automatic evaluation model. The scoring formula of the automatic evaluation model is:

[0034] S=c+W1T reduce -W2D route -W3C conflict (8)

[0035] Where c represents a constant value in the range of [0-100]. When a serious conflict occurs, the value of c is small, and in other cases the value of c is 90. W1, W2, and W3 represent weight coefficients, T reduce Represents the time that the flight time is reduced compared to the planned data, D route Represents the degree of deviation of the flight path from the original path, C conflict Represents the number of normal conflicts. Normal and severe conflicts are defined by different safety intervals. S is the score generated by the evaluation model. The air traffic controllers adjust the scoring results and normalize them to serve as label data for the air traffic control reinforcement learning evaluation model. This constitutes the training dataset, which is then used to train the evaluation model.

[0036] In step 4, based on the matching data obtained in step 3, the data ranked by the highest score is filtered, and the normalized [LON, LAT, ALT, SPD, HDG] in the matching data is used as input. The control intent data O after one-hot encoding in the matching data is used as output to perform supervised pre-training on deep reinforcement learning. The loss function uses the cross entropy loss function Loss, and the formula is:

[0037]

[0038] Where M represents the number of categories of regulatory intent; y ic is a sign function that takes 1 when the true value of sample i is c, otherwise it takes 0; ic The probability that the model output sample i is category c; when the accuracy reaches the threshold, the training process is stopped, the loss function is modified to the air traffic control reinforcement learning evaluation model, the network weights of the air traffic control deep reinforcement learning flight allocation model and the air traffic control reinforcement learning evaluation model are fixed, and exploration is carried out in a simulated training environment. When a certain amount of exploration data is accumulated, the weights of the air traffic control reinforcement learning evaluation model are continued to be fixed, and the weights of the air traffic control deep reinforcement learning flight allocation model are started to be trained. The goal is to maximize the output of the air traffic control reinforcement learning evaluation model.

[0039] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the method.

[0040] Beneficial effects: The present invention can effectively accelerate the exploration and training speed of the flight dispatch model, improve the data utilization rate of the model in the exploration and training stages, and make the flight dispatch model training results closer to normal controller behavior. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more apparent.

[0042] Figure 1 It is a schematic diagram of the workflow of the present invention.

[0043] Figure 2 This is a schematic diagram of the screen recording data of the control automation system. DETAILED DESCRIPTION

[0044] like Figure 1 As shown, the present invention provides a flight dispatch model pre-training method based on multimodal data, comprising the following steps:

[0045] Step 1: Get the screen recording data of the control automation system, such as Figure 2 As shown in Figure 1, this type of data contains rich flight information, such as flight number FN, altitude ALT, speed SPD, heading HDG and position POS and other key parameters. Figure 2For example, the flight in the dashed box has CXA8077 as its flight number, 0886 as its current altitude, 0978 as the target altitude assigned by the controller, 080 as its speed, and P215 as the next waypoint to be passed. The target altitude assigned by the controller is the intended destination. This information is extracted using the Pytesseract image and text recognition model. The deep neural network target tracking model performs coordinate transformation on the aircraft's horizontal position using the following formula:

[0046] P xy =T(P uv ) (1)

[0047] Among them, P xy Represents the coordinate information converted to longitude and latitude, T represents the plane coordinate conversion function, P uv Represents the target pixel center point output by the target tracking model. The aircraft's position on the screen is obtained through coordinate conversion. Using the current position and the latitude and longitude of the next point, the aircraft's current heading can be determined. When relevant information is missing from the sign, flight information can be obtained using the planned data matching method. The planned data matching method calculates the track position, traverses the planned data, and searches for the nearest route plan for matching. The formula is as follows:

[0048] P true =Min(Dis(P xy ,P)) (2)

[0049] Among them, P true Refers to the matching plan data, Min refers to the minimum value function, Dis refers to the distance calculation function, P cy Refers to the latitude and longitude coordinates after coordinate conversion, and P refers to all planned data. This information will serve as the input to the air traffic control reinforcement learning evaluation model and the air traffic control deep reinforcement learning flight allocation model;

[0050] Step 2: Simultaneously collect voice data corresponding to the time of the control automation system screen recording data, and use a deep neural network speech recognition algorithm, such as the Whisper model, to process the voice data and convert it into text data. The formula is as follows:

[0051] W=ASR(A) (3)

[0052] Where W represents the converted text, ASR stands for the neural network speech recognition model, and A represents the original input speech. After obtaining the text, template matching and deep neural network natural language processing algorithms are used to extract key command information from the text data, including the flight number FN, altitude adjustment value ALTA, heading adjustment value HDGA, speed adjustment value SPDA, and route adjustment value RA. For example, in the following command: "Egret eight holes turn turn up to nine turns hold," "Egret" can be converted to "CXA," "eight holes turn turn" can be converted to "8077," and the flight number is "CXA8077." The command type is "altitude adjustment," and "nine turns" can be converted to "9700" according to air traffic control standards, indicating an altitude adjustment to 9,700 meters.

[0053] Step 3: Match the data from Step 1 and Step 2 based on the flight numbers extracted from Step 1 and Step 2. Since the flight number is unique within the same time period, it can be used as a matching basis. At the same time, time alignment is performed based on the speed change time in the screen recording data using the following formula:

[0054]

[0055] Match represents whether the time of voice and data is aligned, 1 represents alignment, and 0 represents misalignment. new Represents the current altitude / speed / heading value, V old represents the altitude / speed / heading value at the last sampling moment, and θ represents the threshold value, which is set to 50 in this method. At the same time, the control intent is combined according to the following formula:

[0056]

[0057] A represents the final intention, A Voice Represents the intent extracted from the voice data, A Video Represents the intent extracted from the screen recording data. Represents the empty intention. Because the control instruction already has a clear intention value of adjusting the altitude to 9700 meters, the intention extracted from the control instruction is used as the final intention.

[0058] Normalize the data in step 1, including altitude ALT, speed SPD, heading HDG, and position POS, to the maximum and minimum values, and adjust the data range to the range of 0 to 1. The formula is as follows:

[0059]

[0060] Where X′ represents the normalized data, X represents the data before normalization, and X min Represents the minimum value among all data in this category.

[0061] Parse the control instruction category and instruction value data in step 2 into one-hot encoding. The format of one-hot encoding is:

[0062] [o1,o2,o3,…,o n ] (7)

[0063] If o i =1, then the rest o j =0,(j≠i). The normalized data and the one-hot encoded data are concatenated into a matrix as the input data for the air traffic control reinforcement learning evaluation model. The concatenation order is: [LON, LAT, ALT, SPD, HDG, O], where LON and LAT are the longitude and latitude in the POS information, and O is the one-hot encoded information. The original video data and voice data obtained in steps 1 and 2 are preliminarily evaluated and scored by the automatic evaluation model. The scoring formula of the automatic evaluation model is:

[0064] S=c+W1T reduce -W2D route -W3C conflict (8)

[0065] Where c represents a constant value in the range [0-100]. When a serious conflict occurs, c is set to 20, and in other cases the value of c is 90. W1, W2, and W3 represent weight coefficients, which are initially set to 1, 0.2, and 10. reduce Represents the time that the flight time is reduced compared to the planned data, D route Represents the degree of deviation of the flight path from the original path, C conflict Represents the number of normal conflicts. Normal and severe conflicts are defined by different safety intervals. A severe conflict is defined when the horizontal separation between aircraft is less than 3 nautical miles, and a normal conflict is defined when the horizontal separation between aircraft is greater than 3 nautical miles but less than 5 nautical miles. The scores are corrected by trainers and normalized to their minimum and maximum values, serving as label data for the air traffic control reinforcement learning evaluation model. This data is used to construct the evaluation model training dataset. This data will be used to train the evaluation model, enabling it to accurately assess the rationality and effectiveness of air traffic control decisions. The air traffic control reinforcement learning evaluation model is constructed using a five-layer fully connected neural network. Multiple rounds of training are performed using the evaluation model until the loss decreases and stabilizes, and the training accuracy exceeds 98%.

[0066] Step 4: Based on the matching data obtained in step 3, filter the top 30% of the data ranked by score, use the normalized aircraft information data [LON, LAT, ALT, SPD, HDG] in the matching data as input, and use the one-hot encoded control intent data O in the matching data as output to perform supervised pre-training on deep reinforcement learning. The loss function uses the cross entropy loss function, and the formula is:

[0067]

[0068] Where M represents the number of categories of regulatory intent, y ic is a sign function that takes 1 when the true value of sample i is c, otherwise it takes 0. ic The probability that the model output sample i is class c is represented by the probability that the sample i is class c. When the accuracy rate reaches above 70%, training is terminated and the loss function is modified to the ATC RL evaluation model. The network weights of the ATC deep RL flight scheduling model and the ATC RL evaluation model are fixed, and exploration is conducted in a simulated training environment. Once a sufficient amount of exploration data has been accumulated, the weights of the ATC RL evaluation model are fixed and training of the ATC deep RL flight scheduling model begins, with the goal of maximizing the output of the ATC RL evaluation model. Training is terminated when the score of the ATC deep RL flight scheduling model increases and stabilizes.

[0069] After using this method, compared with the ordinary air traffic control deep reinforcement learning flight allocation model training method based on autonomous exploration, the time required for the autonomous exploration stage of the flight allocation model in the training process is greatly shortened. Artificial experience is used to improve the convergence speed of the flight allocation model, reduce the unreasonable output actions of the flight allocation model, and make the output actions of the flight allocation model more in line with the controller's operating habits.

[0070] This invention provides a method for pre-training a flight scheduling model based on multimodal data. There are numerous methods and approaches for implementing this technical solution. The above is merely a preferred embodiment of the invention. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the invention, and such improvements and modifications are also within the scope of protection of the invention. Any components not specified in this embodiment may be implemented using existing technologies.

Claims

1. A flight dispatch model pre-training method based on multimodal data, characterized in that: The following steps are involved: Step 1: Obtain screen recording data from the control automation system and perform coordinate conversion using a deep neural network target detection model to obtain the position information of the aircraft on the screen. Obtain aircraft speed, heading, altitude, flight number, and part of the control intent through text recognition detection model or plan data matching; Step 2: Acquire the voice data corresponding to the screen recording data of the control automation system, convert the voice data into text using a deep neural network speech recognition model, and extract the control intent using an intent extraction model. The result is combined with the control intent in the screen recording data. Step 3: The aircraft's position, speed, heading, altitude, and control intent are matched by flight number to form training data, which is then scored to train an air traffic control reinforcement learning evaluation model. The model comprises a five-layer fully connected neural network, which takes the aircraft's position, speed, heading, altitude, and control intent as input and outputs the data score. Step 4: Use the training data and evaluation model to train the ATC deep reinforcement learning flight dispatch model; In step 1, the control automation system screen recording data includes flight number FN, altitude ALT, speed SPD, heading HDG, position POS and altitude and heading adjustment intentions; In step 1, the formula for the coordinate transformation operation is: P xy = T(P uv ) (1) Among them, P xy Represents the coordinate information converted to longitude and latitude, T represents the plane coordinate conversion function, P uv represents the target pixel center point output by the deep neural network target detection model used in step 1; In step 1, the plan data matching is performed by calculating the track position, traversing the plan data, and searching for the plan of the nearest route. The formula is: P true =Min(Dis(P xy ,P)) (2) Among them, P true Refers to the matched plan data, Min refers to the minimum value function, Dis refers to the distance calculation function, and P refers to all plan data; In step 2, the speech data is converted into text through the deep neural network speech recognition model. The formula is: W=ASR(A) (3) Where W is the converted text information, ASR represents the deep neural network speech recognition model, and A represents the original input speech; In step 2, the voice data corresponds to the screen recording data time. The data alignment method uses the voice data time as the benchmark and the flight number as the matching condition to perform threshold judgment on the altitude, speed and heading values. The formula is: Match represents whether the time of voice and data is aligned, 1 represents alignment, and 0 represents misalignment; V new Represents the current altitude, speed or heading value, V old represents the altitude, speed or heading value at the last sampling moment, and θ represents the threshold; In step 2, the control intent includes the flight number, altitude adjustment, heading adjustment, and speed adjustment intent. The control intent in the screen recording data and the control intent in the voice data are combined into the following formula: A represents the final intention, A Voice Represents the intent extracted from the voice data, A Video Represents the intent extracted from the screen recording data. represents empty intention; In step 3, the altitude ALT, speed SPD, heading HDG, and position POS obtained in step 1 are normalized to their maximum and minimum values ​​using the following formula: where X ′ represents the data after normalization, X represents the data before normalization, and X min Represents the minimum value of all data in this category, X max Represents the maximum value of all data in this category; In step 3, the data except the flight number in the control intent in step 2 is parsed into a one-hot encoding. The format of the one-hot encoding is: [o1,o2,o3,…,o n ](7) Among them, if o i =1, then the rest o j =0, i and j are 1 to n and j≠i, n represents the number of categories of control intention data; n The one-hot encoding of the nth category of data in the regulatory intent; The one-hot encoded data and the normalized data are concatenated into a matrix as the input data for the air traffic control reinforcement learning evaluation model. The concatenation order is: [LON, LAT, ALT, SPD, HDG, O], where LON and LAT are the longitude and latitude in the position information, respectively, and O is the one-hot encoded control intent data. The normalized data in step 1 and the one-hot encoded data in step 2 are preliminarily evaluated and scored by the automatic evaluation model. The scoring formula of the automatic evaluation model is: S=c+W1T reduce -W2D route -W3C conflict (8) Where c represents a constant value in the range of [0-100]. When a serious conflict occurs, the value of c is small, and in other cases the value of c is 90; W1, W2, and W3 represent weight coefficients; T reduce Represents the time that the flight time is reduced compared to the planned data, D route Represents the degree of deviation of the flight path from the original path, C conflict represents the number of ordinary conflicts; ordinary conflicts and serious conflicts are defined by different safety intervals; S is the score obtained by the evaluation model; The scoring results are corrected and normalized, and used as label data for the air traffic control reinforcement learning evaluation model to construct a training data set, and the evaluation model is trained using the training data set.

2. The method according to claim 1, characterized in that In step 4, based on the matching data obtained in step 3, the data ranked by the highest score is filtered, and the normalized [LON, LAT, ALT, SPD, HDG] in the matching data is used as input. The control intent data O after one-hot encoding in the matching data is used as output to perform supervised pre-training on deep reinforcement learning. The loss function uses the cross entropy loss function Loss, and the formula is: Where M represents the number of categories of regulatory intent; y ic is a sign function that takes 1 when the true value of sample i is c, otherwise it takes 0; ic The probability that the model output sample i is category c; when the accuracy reaches the threshold, the training process is stopped, the loss function is modified to the air traffic control reinforcement learning evaluation model, the network weights of the air traffic control deep reinforcement learning flight allocation model and the air traffic control reinforcement learning evaluation model are fixed, and exploration is carried out in a simulated training environment. When a certain amount of exploration data is accumulated, the weights of the air traffic control reinforcement learning evaluation model are continued to be fixed, and the weights of the air traffic control deep reinforcement learning flight allocation model are started to be trained. The goal is to maximize the output of the air traffic control reinforcement learning evaluation model.

Citation Information

Patent Citations

  • Air traffic control simulation training evaluation method

    CN106709642A

  • Flight release rate prediction system

    CN112132366A