Airplane motion trajectory determination method and device, electronic equipment and storage medium

By using a pre-trained neural network model for reinforcement learning and combining environmental and aircraft state data to generate accurate aircraft motion trajectories, the problem of insufficient flexibility of fixed-wing aircraft dynamics models under environmental uncertainty is solved, achieving higher prediction accuracy and adaptability.

CN119249899BActive Publication Date: 2025-10-24POLIXIR TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411368402.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-10-24
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

In the existing technology, the fixed-wing aircraft dynamics model has poor flexibility in predicting the aircraft's flight trajectory and cannot meet the needs of various practical applications, especially the lack of accuracy under the influence of environmental weather uncertainties.

Method used

By determining the environmental data of the target aircraft's predicted flight area, combining the initial state data and the preset aircraft action plan, and using a pre-trained neural network model for reinforcement learning, the predicted motion trajectory for different predicted time periods is generated.

Benefits of technology

The accuracy and flexibility of the predicted motion trajectory are improved, which can adapt to the aircraft motion prediction needs under different environmental conditions and meet the users' various prediction needs in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119249899B_ABST
    Figure CN119249899B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an aircraft motion trajectory determination method and device, electronic equipment and a storage medium. The method comprises: determining target environment data of a target aircraft corresponding to a predicted flight area in a to-be-predicted time period; wherein the to-be-predicted time period comprises at least one to-be-predicted time; determining predicted state data of the target aircraft corresponding to each to-be-predicted time in the to-be-predicted time period based on the target environment data, initial state data of the target aircraft, a preset aircraft action scheme and a pre-trained trajectory determination model; wherein the trajectory determination model is obtained by pre-strengthening learning of a neural network model; and generating a predicted motion trajectory of the target aircraft in the to-be-predicted time period based on the predicted state data corresponding to each to-be-predicted time. The technical solution of the embodiments of the present application can improve the accuracy of the predicted motion trajectory, has strong flexibility and is convenient for meeting various prediction requirements of users in actual applications.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of aircraft control, and in particular to a method and device for determining a flight trajectory of an aircraft, an electronic device and a storage medium. BACKGROUND

[0002] At present, a fixed-wing aircraft dynamics model is usually built through simulation to predict the flight trajectory of an aircraft in a flight process through the fixed-wing aircraft dynamics model.

[0003] However, in the process of implementing the present application, it is found that the prior art at least has the following technical problems: since the flight state and trajectory of an aircraft are affected by the environment and weather, there is uncertainty. The fixed-wing aircraft dynamics model can only predict the flight state and trajectory under a fixed environment, and the flexibility is poor, which cannot meet the various prediction requirements of users in actual applications. SUMMARY

[0004] Embodiments of the present application provide a method and device for determining a flight trajectory of an aircraft, an electronic device and a storage medium to achieve the purpose of improving the accuracy of predicting a flight trajectory, and generate a corresponding predicted flight trajectory for different to-be-predicted time periods, which has strong flexibility and is convenient for meeting various prediction requirements of users in actual applications.

[0005] According to an aspect of the present application, a method for determining a flight trajectory of an aircraft is provided, comprising:

[0006] determining target environment data of a target aircraft corresponding to a to-be-predicted flight area in a to-be-predicted time period; wherein the to-be-predicted time period comprises at least one to-be-predicted time point;

[0007] based on the target environment data, initial state data of the target aircraft, a preset aircraft action scheme and a pre-trained trajectory determination model, determining predicted state data of the target aircraft corresponding to each to-be-predicted time point in the to-be-predicted time period; wherein the trajectory determination model is obtained by pre-strengthening learning of a neural network model;

[0008] based on the predicted state data corresponding to each to-be-predicted time point, generating a predicted flight trajectory of the target aircraft in the to-be-predicted time period.

[0009] According to another aspect of the present application, a device for determining a flight trajectory of an aircraft is provided, comprising:

[0010] a target environment data determination module configured to determine target environment data of a target aircraft corresponding to a to-be-predicted flight area in a to-be-predicted time period; wherein the to-be-predicted time period comprises at least one to-be-predicted time point;

[0011] The prediction state data determination module is configured to determine, based on the target environment data, initial state data of the target aircraft, a preset aircraft action scheme, and a pre-trained trajectory determination model, prediction state data corresponding to each of the to-be-predicted time instants in the to-be-predicted time period.

[0012] The prediction running trajectory generation module is configured to generate, based on the prediction state data corresponding to each of the to-be-predicted time instants, a prediction running trajectory of the target aircraft in the to-be-predicted time period.

[0013] According to another aspect of the present application, an electronic device is provided, which comprises:

[0014] at least one processor; and

[0015] a memory in communication with the at least one processor; wherein

[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the aircraft running trajectory determination method according to any one of the embodiments of the present application.

[0017] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to perform the aircraft running trajectory determination method according to any one of the embodiments of the present application when executed.

[0018] The technical scheme of the embodiments of the present application determines the target environment data of the to-be-predicted flight region corresponding to the target aircraft in the to-be-predicted time period, wherein the to-be-predicted time period comprises at least one to-be-predicted time instant; and based on the target environment data, initial state data of the target aircraft, a preset aircraft action scheme, and a pre-trained trajectory determination model, prediction state data corresponding to each of the to-be-predicted time instants in the to-be-predicted time period is determined, wherein the trajectory determination model is obtained by pre-strengthening learning of a neural network model; thus, the prediction state data is determined while considering the target environment data of the to-be-predicted flight region, which is beneficial to improving the accuracy of the determined prediction state data; finally, based on the prediction state data corresponding to each of the to-be-predicted time instants, a prediction running trajectory of the target aircraft in the to-be-predicted time period is generated, which further improves the accuracy of the prediction running trajectory, and a corresponding prediction running trajectory is generated for different to-be-predicted time periods, which is flexible and convenient for meeting various prediction requirements of users in actual applications.

[0019] It is to be understood that the embodiments described herein are merely exemplary of the application and that a myriad of modifications, both as to the nature and number of elements within the execution of the application and as to the modes of execution thereof, can be made by those skilled in the art, without expressly quantifying the application and without departing from the scope of the application. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiments description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0021] Figure 1 is a flow chart of a method for determining an aircraft motion trajectory according to an embodiment of the present application;

[0022] Figure 2 is a flow chart of another method for determining an aircraft motion trajectory according to an embodiment of the present application;

[0023] Figure 3 is a structural schematic diagram of a device for determining an aircraft motion trajectory according to an embodiment of the present application;

[0024] Figure 4 is a structural schematic diagram of an electronic device for implementing the method for determining an aircraft motion trajectory according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to make the technical personnel in the art better understand the present application, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of the present application.

[0026] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "comprise" and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] It should be noted that in the technical solutions of the present disclosure, the collection, collection, updating, analysis, processing, use, transmission, storage, etc. of user personal information are in line with relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain user personal information security and network security.

[0028] Figure 1 A flowchart of a method for determining an aircraft motion trajectory is provided according to an embodiment of the present application. The present embodiment can be applicable to the case of predicting the flight trajectory of an aircraft in a to-be-predicted time period through the state data of the aircraft. The method can be executed by an aircraft motion trajectory determination device, which can be realized in the form of hardware and / or software.

[0029] As shown in Figure 1 , the method of the present embodiment can specifically include:

[0030] S110, determining target environment data of a target aircraft corresponding to a to-be-predicted flight area in a to-be-predicted time period; wherein the to-be-predicted time period includes at least one to-be-predicted time point.

[0031] The target aircraft is an aircraft whose motion trajectory needs to be predicted, and the to-be-predicted time period can be a future time period in which the target aircraft is expected to fly. The to-be-predicted time period can be divided into at least one to-be-predicted time point. For example, the to-be-predicted time period can be divided into time intervals, so that each division point corresponds to a to-be-predicted time point, and the length of the time interval is less than the length of the to-be-predicted time period. By determining the predicted state data of the target aircraft at each to-be-predicted time point, the predicted motion trajectory of the target aircraft in the to-be-predicted time period is formed. The target environment data includes at least one of wind speed data, wind direction data and temperature data in the to-be-predicted flight area in the to-be-predicted time period. The to-be-predicted flight area is a pre-determined aviation area that the target aircraft needs to fly in the to-be-predicted time period.

[0032] In the present embodiment, weather forecast data of the to-be-predicted flight area in the to-be-predicted time period can be obtained, and the target environment data can be determined based on the weather forecast data. Specifically, weather data corresponding to the to-be-predicted time point in the weather forecast data can be extracted as the target environment data.

[0033] S120, determining the predicted state data of the target aircraft at each to-be-predicted time point in the to-be-predicted time period based on the target environment data, the initial state data of the target aircraft, the preset aircraft action scheme and the pre-trained trajectory determination model.

[0034] The trajectory determination model is obtained by performing reinforcement learning on a neural network model in advance. The preset aircraft action scheme is composed of preset action data of a control stick of the target aircraft corresponding to each to-be-predicted time. The initial state data can be state data corresponding to a starting position of the target aircraft entering a to-be-predicted flight region. The predicted state data is state data obtained by predicting the target aircraft at the to-be-predicted time. The state data includes at least one of position, attitude, and speed of the target aircraft.

[0035] Optionally, the neural network model is a network model constituted by connecting a residual network model as a specific network framework in front of and behind each other through a plurality of residual network models. For example, the neural network model can be constituted by connecting a 5-layer residual network model in front of and behind each other.

[0036] In this embodiment, based on the target environment data, the initial state data of the target aircraft, the preset aircraft action scheme, and the pre-trained trajectory determination model, the implementation of the predicted state data of the target aircraft corresponding to each to-be-predicted time in the to-be-predicted time period includes: sequentially traversing each to-be-predicted time in the to-be-predicted time period in the order of time from front to back of the to-be-predicted time; for the currently traversed current to-be-predicted time, based on the target environment data, the initial state data of the target aircraft, and the preset aircraft action scheme, determining the current input data corresponding to the current to-be-predicted time, inputting the current input data into the pre-trained trajectory determination model, and obtaining the predicted state data corresponding to the current to-be-predicted time.

[0037] Specifically, each to-be-predicted time in the to-be-predicted time period is sequentially traversed in the order of time from front to back of the to-be-predicted time, and the predicted state data corresponding to the currently traversed current to-be-predicted time is determined. Optionally, if the current to-be-predicted time is the first to-be-predicted time in the to-be-predicted time period, the initial state data, the time environment data corresponding to the first to-be-predicted time, and the time action data are input into the pre-trained trajectory determination model as the current input data, and the predicted state data of the target aircraft corresponding to the first to-be-predicted time is output. If the to-be-predicted time for which the predicted state data is to be determined is a to-be-predicted time other than the first to-be-predicted time in the to-be-predicted time period, the predicted state data of the target aircraft corresponding to the previous to-be-predicted time, the time action data and the time environment data of the current to-be-predicted time are input into the pre-trained trajectory determination model as the current input data, and the predicted state data of the target aircraft corresponding to the current to-be-predicted time is output.

[0038] In this embodiment, the prediction state data corresponding to the previous to-be-predicted time is taken as the input data corresponding to the current to-be-predicted time, and the time action data and the time environment data of the current to-be-predicted time are combined, so that the environment information of the current to-be-predicted time is considered, which is beneficial to improve the accuracy of the prediction state data determined for each to-be-predicted time.

[0039] In S130, the prediction motion trajectory of the target aircraft in the to-be-predicted time period is generated based on the prediction state data corresponding to each to-be-predicted time.

[0040] Specifically, the prediction state data corresponding to each to-be-predicted time and the time sequence of the to-be-predicted time are combined, and the prediction motion trajectory is obtained based on the combined data. For example, the combined data is fitted, and the fitted trajectory is determined as the prediction motion trajectory.

[0041] The technical scheme of the embodiment of the application determines the target environment data of the to-be-predicted flight area corresponding to the target aircraft in the to-be-predicted time period, wherein the to-be-predicted time period includes at least one to-be-predicted time; and based on the target environment data, the initial state data of the target aircraft, the preset aircraft action scheme and the pre-trained trajectory determination model, the prediction state data corresponding to each to-be-predicted time of the target aircraft in the to-be-predicted time period is determined, wherein the trajectory determination model is obtained by strengthening the learning of the neural network model; thus, the prediction state data is determined while considering the target environment data of the to-be-predicted flight area, which is beneficial to improve the accuracy of the determined prediction state data; finally, the prediction motion trajectory of the target aircraft in the to-be-predicted time period is generated based on the prediction state data corresponding to each to-be-predicted time, which further improves the accuracy of the prediction motion trajectory, generates the corresponding prediction motion trajectory for different to-be-predicted time periods, has strong flexibility and is convenient for meeting various prediction requirements of users in actual application.

[0042] Figure 2 is a flowchart of another aircraft motion trajectory determination method provided by the embodiment of the application. Based on the above-mentioned embodiment, optionally, the trajectory determination model includes a state transition network; the reinforcement learning process of the trajectory determination model includes: obtaining a historical running data set of the target aircraft; wherein the historical running data set is composed of a plurality of initial real state data, a plurality of action data groups and a plurality of environment data groups; the action data group includes at least one historical action data, and the environment data group includes at least one historical environment data; based on the historical running data set, the to-be-trained state transition network is trained, and the state transition network after the training is completed is taken as the trajectory determination model. Wherein, the explanations of the same or corresponding terms as in the above-mentioned embodiments are not repeated here. As shown in the figure, the method includes: Figure 2

[0043] ​S210, acquire a historical running data set of the target aircraft; wherein the historical running data set is composed of a plurality of initial real state data, a plurality of action data groups and a plurality of environment data groups; the action data group contains at least one historical action data, and the environment data group contains at least one historical environment data.

[0044] Optionally, the trajectory determination model comprises a state transition network. The state transition network is a neural network. The historical running data set can be actual running data of the target aircraft in a historical flight time period. The historical environment data can be at least one of real wind direction, real wind speed and real temperature in the historical flight time period.

[0045] In a specific implementation, real state data, real action data and environment data of the target aircraft corresponding to a historical flight time period closest to the current time can be acquired to form the historical running data set. In order to facilitate the subsequent model training process, the data in the acquired historical running data set can be preprocessed. Specifically, before training the state transition network to be trained based on the historical running data set, the method further comprises: performing a data cleaning operation on the data in the historical running data set, and updating the historical running data set based on the data obtained after the data cleaning operation.

[0046] The data cleaning operation comprises at least one of missing value processing, abnormal value processing, error data processing and repeated data processing.

[0047] In a specific implementation, the specific way of missing value processing can be to fill the missing values in the historical running data set by using the median, mean or mode of the data in the historical running data set; or the average of two adjacent values adjacent to the missing value can also be determined, and the average is filled as the missing value. For example, if there is a missing value in the state data group, the median, mean or mode of the state data group can be determined, and the missing state data is filled with the median, mean or mode. The specific way of abnormal value processing can be to delete the abnormal value or repair the abnormal value based on the values adjacent to the abnormal value. The way of repeated data processing can be to delete the repeated data. The way of error data processing can be to delete the error data.

[0048] The embodiment performs the data cleaning operation on the data in the historical running data set, thereby facilitating the effective completion of the model training process, and is beneficial to improving the accuracy of the model training result and reducing errors.

[0049] S220, train the state transition network to be trained based on the historical running data set, and acquire the state transition network after the training as the trajectory determination model.

[0050] In the embodiment, in order to accurately and effectively train the state transition network, the data in the historical running data set can be subjected to a data alignment operation. Specifically, the data in the action data group and the environment data group are aligned based on the data generation time to establish a data correspondence relationship between the historical action data and the historical environment data at different data generation times. The starting real state data is the state data generated at the initial data generation time of the historical flight time period.

[0051] In a specific implementation, the implementation manner of training the state transition network based on the historical running data set specifically includes: traversing each starting real state data; for each traversed starting real state data, determining input data based on the starting real state data, the action data group and the environment data group corresponding to the starting real state data, inputting the input data into the state transition network to be trained, outputting historical predicted state data corresponding to the input data, determining a current reward value corresponding to the input data based on the discriminator network to be trained, updating the input data based on the historical predicted state data, inputting the updated input data into the state transition network to be trained, outputting historical predicted state data corresponding to the updated input data, repeatedly performing the operations of updating the input data based on the historical predicted state data, outputting the historical predicted state data and determining the current reward value until a preset ending condition is met to stop the operations, updating the network parameters of the state transition network to be trained based on each current reward value obtained in the current traversal, and updating the network parameters of the discriminator network to be trained based on each historical predicted state data obtained in the current traversal, the state data group corresponding to the starting real state data, the action data group and the environment data group, so that the discriminator network determines a next reward value based on the updated network parameters in the next traversal operation until a preset training condition is met to end the training, and the traversal of the starting real state data is stopped; wherein the state data group is a data group in which the data generated at the initial data generation time is the starting real state data; and the trajectory determination model is composed of the state transition network after the training.

[0052] It should be noted that the data generation times of the historical action data in each action data group and the data generation times of the historical environment data in each environment data group are continuous time points with equal time intervals. For example, the action data group includes 10 historical real action data, and the data generation times corresponding to each historical real action data are distributed with equal time intervals.

[0053] In this embodiment, each starting real state data can be traversed, for the first starting real state data traversed, in the action data group and the environment data group corresponding to the starting real state data, the historical action data and the historical environment data corresponding to the data generation time of the starting real state data are determined. For example, the starting real state data is the data generated at the starting time of the historical flight time period, and the historical action data and the historical environment data generated at the starting time in the action data group and the environment data group corresponding to the starting real state data can be determined. The starting real state data, the historical action data and the historical environment data corresponding to the starting real state data are input as input data into the state transition network to be trained, and the historical prediction state data corresponding to the input data is output. The historical prediction state data obtained based on the starting real state data is the prediction data generated at the next time of the data generation time of the starting real state data of the target aircraft. Based on the historical prediction state data and the pre-constructed discriminator network to be trained, the current reward value corresponding to the input data is determined.

[0054] The historical prediction state data, the historical action data and the historical environment data corresponding to the historical prediction state data are updated as input data, and the updated input data is input into the state transition network to be trained, and the historical prediction state data corresponding to the updated input data is output. The operation of updating the input data based on the historical prediction state data, outputting the historical prediction state data and determining the current reward value is repeatedly performed until the preset ending condition is met, then the operation is stopped, and the current traversal is ended. For example, the preset ending condition is that the number of repeated operations reaches a preset number threshold.

[0055] The adjustment parameters of the state transition network can be determined through the PPO (Proximal Policy Optimization) algorithm and each current reward value obtained in the current traversal, and the network parameters of the state transition network are adjusted based on the adjustment parameters to maximize the cumulative reward.

[0056] In addition, based on each historical prediction state data obtained in the current traversal, the state data group, the action data group and the environment data group corresponding to the starting real state data, the network parameters of the discriminator network to be trained are updated, so that the next reward value is determined based on the discriminator network with updated network parameters in the next traversal operation, and the training is ended when the preset training condition is met, and the traversal of the starting real state data is stopped. The preset training condition includes that the loss function corresponding to the discriminator network converges.

[0057] In a specific implementation, the network parameters of the discriminator network to be trained are updated based on the current iteration of each historical predicted state data, the state data group corresponding to the starting real state data, the action data group and the environment data group, including: updating the network parameters of the discriminator network to be trained based on the generative network adversarial algorithm, the current iteration of each historical predicted state data, the state data group corresponding to the starting real state data, the action data group and the environment data group.

[0058] Specifically, each current iteration of the historical predicted state data is performed, and the current iteration of the historical predicted state data, the historical real state data in the state data group corresponding to the historical predicted state data, the historical action data and the historical environment data are input to the discriminator network to be trained, the loss function corresponding to the discriminator network is determined based on the output result of the discriminator network to be trained, and the network parameters of the discriminator network are adjusted based on the Adam optimization algorithm; and the next reward value corresponding to the state transition network at the next iteration is determined by the discriminator network after adjusting the network parameters.

[0059] In a specific implementation, the network parameters of the discriminator network are adjusted based on each iteration of the obtained historical predicted state data, until the preset training condition is met, the iteration of the state data group is stopped, and the trajectory determination model is composed of the state transition network after training. The preset training condition includes that the loss function corresponding to the discriminator network converges. The loss function can be a cross-entropy loss function.

[0060] The embodiment trains the discriminator network through the generative network adversarial algorithm, so that the discriminator network can more accurately distinguish between predicted state data and real state data, which is conducive to improving the accuracy of the trajectory determination model.

[0061] S230, determining target environment data of a target flight area corresponding to a target aircraft in a to-be-predicted time period; wherein the to-be-predicted time period includes at least one to-be-predicted time.

[0062] S240, determining predicted state data of the target aircraft at each to-be-predicted time in the to-be-predicted time period based on the target environment data, initial state data of the target aircraft, a preset aircraft action scheme and a pre-trained trajectory determination model; wherein the trajectory determination model is obtained by pre-strengthening learning of a neural network model.

[0063] S250, generating a predicted motion trajectory of the target aircraft in the to-be-predicted time period based on the predicted state data corresponding to each to-be-predicted time.

[0064] In this embodiment, the trajectory determination model is trained in a reinforcement learning manner, so that the trajectory determination model can process an environment with high uncertainty and dynamics, so as to adapt to a complex environment in the real world and accurately determine the aircraft motion trajectory.

[0065] Figure 3 Fig. 1 is a structural schematic diagram of an aircraft motion trajectory determination device according to an embodiment of the present application. The device is used to execute the aircraft motion trajectory determination method provided in any of the above embodiments. The device and the aircraft motion trajectory determination method in each of the above embodiments belong to the same inventive concept. Details not described in the embodiment of the aircraft motion trajectory determination device can be referred to the above embodiments of the aircraft motion trajectory determination method. As shown in the figure, the device comprises: Figure 3

[0066] A target environment data determination module 10 is configured to determine target environment data of a target flight area corresponding to a target aircraft in a to-be-predicted time period. The to-be-predicted time period comprises at least one to-be-predicted time point.

[0067] A predicted state data determination module 11 is configured to determine predicted state data of the target aircraft at each to-be-predicted time point in the to-be-predicted time period based on the target environment data, initial state data of the target aircraft, a preset aircraft action scheme, and a pre-trained trajectory determination model. The trajectory determination model is obtained by performing reinforcement learning on a neural network model in advance.

[0068] A predicted running trajectory generation module 12 is configured to generate a predicted motion trajectory of the target aircraft in the to-be-predicted time period based on the predicted state data corresponding to each to-be-predicted time point.

[0069] Optionally, the predicted state data determination module 11 comprises:

[0070] A to-be-predicted time point traversal sub-module is configured to traverse each to-be-predicted time point in the to-be-predicted time period in a time sequence from front to back.

[0071] A current input data determination sub-module is configured to determine, for a currently-traversed to-be-predicted time point, current input data corresponding to the currently-traversed to-be-predicted time point based on the target environment data, the initial state data of the target aircraft, and the preset aircraft action scheme, input the current input data into the pre-trained trajectory determination model, and obtain the predicted state data corresponding to the to-be-predicted time point.

[0072] ​Based on any optional technical solution in the embodiments of the present invention, optionally, the target environment data includes at least one of wind speed data, wind direction data and temperature data in the flight area to be predicted within the predicted time period; the preset aircraft action plan is composed of preset action data of the joystick of the target aircraft corresponding to each predicted moment.

[0073] Based on any optional technical solution in the embodiments of the present invention, optionally, the trajectory determination model includes a state transition network; and the device further includes:

[0074] A historical operation data set acquisition module is used to acquire a historical operation data set of the target aircraft; wherein the historical operation data set is composed of a plurality of initial real state data, a plurality of action data groups, and a plurality of environment data groups; the action data group includes at least one historical action data, and the environment data group includes at least one historical environment data;

[0075] The trajectory determination model composition module is used to train the state transition network to be trained based on the historical operation data set, and obtain the state transition network after training as the trajectory determination model.

[0076] Based on any optional technical solution in the embodiments of the present invention, optionally, the trajectory determination model component module includes:

[0077] The current reward value determination submodule is used to traverse each starting real state data; for each traversal of the starting real state data, the input data is determined based on the starting real state data, the action data group corresponding to the starting real state data, and the environment data group, the input data is input to the state transfer network to be trained, and the historical predicted state data corresponding to the input data is output. Based on the discriminator network to be trained, the current reward value corresponding to the input data is determined, the input data is updated based on the historical predicted state data, and the updated input data value is input to the state transfer network to be trained, and the historical predicted state data corresponding to the updated input data is output, and the updating of the input data based on the historical predicted state data, the output of the historical predicted state data, and the determination of the current reward value are repeated. The operation of the previous reward value is stopped until the preset end condition is met, and the network parameters of the state transition network to be trained are updated based on the current reward values ​​obtained in the current traversal, and the network parameters of the discriminator network to be trained are updated based on the historical predicted state data obtained in the current traversal, the state data group corresponding to the initial real state data, the action data group and the environment data group, so as to determine the next reward value based on the discriminator network after the updated network parameters in the next traversal operation, until the training is terminated when the preset training condition is met, and the traversal of the initial real state data is stopped; wherein the state data group is a data group whose data generated at the moment of initial data generation is the initial real state data; the trajectory determination model is composed of the state transition network after the training is completed.

[0078] In any optional technical solution in the embodiments of the application, optionally, the current reward value determination sub-module comprises:

[0079] The discriminator network training unit is configured to update the network parameters of the discriminator network to be trained based on the generative network adversarial algorithm, the historical prediction state data obtained in the current iteration, the state data group corresponding to the starting real state data, the action data group, and the environment data group.

[0080] In any optional technical solution in the embodiments of the application, optionally, the method further comprises:

[0081] The data cleaning module is configured to perform a data cleaning operation on the data in the historical operation data set before training the state transition network to be trained based on the historical operation data set, and update the historical operation data set based on the data obtained after the data cleaning operation.

[0082] The data cleaning operation comprises at least one of missing value processing, abnormal value processing, error data processing, and duplicate data processing.

[0083] The technical solution of the embodiments of the application determines the target environment data of the target aircraft corresponding to the to-be-predicted flight region in the to-be-predicted time period, wherein the to-be-predicted time period comprises at least one to-be-predicted time point, and the prediction state data of the target aircraft at each to-be-predicted time point in the to-be-predicted time period is determined based on the target environment data, the initial state data of the target aircraft, the preset aircraft action scheme, and the pre-trained trajectory determination model, wherein the trajectory determination model is obtained by performing reinforcement learning on a neural network model. Thus, the prediction state data is determined while considering the target environment data of the to-be-predicted flight region, which is beneficial to improving the accuracy of the determined prediction state data. Finally, the prediction motion trajectory of the target aircraft in the to-be-predicted time period is generated based on the prediction state data corresponding to each to-be-predicted time point, which further improves the accuracy of the prediction motion trajectory. The corresponding prediction motion trajectory is generated for different to-be-predicted time periods, which is flexible and convenient for meeting various prediction requirements of users in actual applications.

[0084] It should be noted that in the embodiments of the above aircraft motion trajectory determination device, each unit and module included is only divided according to functional logic, but is not limited to the above division, as long as the corresponding functions can be implemented. In addition, the specific names of each functional unit are only for easy mutual differentiation, and do not limit the protection scope of the application.

[0085] Figure 4is a schematic diagram of the structure of an electronic device that implements the aircraft motion trajectory determination method according to an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0086] like Figure 4 As shown, the electronic device 20 includes at least one processor 21, and a memory connected to the at least one processor 21, such as a read-only memory (ROM) 22, a random access memory (RAM) 23, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 21 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 22 or the computer program loaded from the storage unit 28 to the random access memory (RAM) 23. Various programs and data required for the operation of the electronic device 20 can also be stored in the RAM 23. The processor 21, ROM 22 and RAM 23 are connected to each other via a bus 24. An input / output (I / O) interface 25 is also connected to the bus 24.

[0087] Multiple components in the electronic device 20 are connected to the I / O interface 25, including an input unit 26, such as a keyboard, a mouse, etc.; an output unit 27, such as various types of displays, speakers, etc.; a storage unit 28, such as a magnetic disk, an optical disk, etc.; and a communication unit 29, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 29 allows the electronic device 20 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0088] Processor 21 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 21 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any other suitable processor, controller, microcontroller, etc. Processor 21 executes the various methods and processes described above, such as the aircraft motion trajectory determination method.

[0089] In some embodiments, the aircraft motion trajectory determination method can be implemented as a computer program tangibly embodied in a computer readable storage medium, e.g., storage unit 28. In some embodiments, parts or all of the computer program can be loaded and / or installed onto electronic device 20 via, e.g., ROM 22 and / or communication unit 29. When the computer program is loaded onto RAM 23 and executed by processor 21, one or more steps of the aircraft motion trajectory determination method described above can be performed. Alternatively, in other embodiments, processor 21 can be configured to perform the aircraft motion trajectory determination method by other means, e.g., with the aid of firmware.

[0090] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0091] Computer programs used to implement the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor of the machine, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0092] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0093] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0094] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0095] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are a host product in the cloud computing service system to solve the defects of great management difficulty and weak business scalability in traditional physical hosts and VPS services.

[0096] The embodiment also provides a computer program product comprising a computer program which, when executed by a processor, implements the method for determining an aircraft motion trajectory as provided in any embodiment of the present application.

[0097] The computer program product, in implementation, can be written in one or more programming languages or combinations of languages to implement the computer program code for performing the operations of the present application, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, through the Internet using an Internet service provider).

[0098] It should be understood that the various forms of flow shown above can be reordered, added to, or deleted from, with steps. For example, the steps described in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, and the present application is not limited herein.

[0099] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method of determining a trajectory of an aircraft, characterized in that, The method comprises: determining target environment data of a target flight area corresponding to a target aircraft in a to-be-predicted time period; wherein the to-be-predicted time period comprises at least one to-be-predicted time point; based on the target environment data, the initial state data of the target aircraft, a preset aircraft action scheme and a pre-trained trajectory determination model, determining the predicted state data of the target aircraft at each to-be-predicted time point in the to-be-predicted time period; wherein the trajectory determination model is obtained by pre-strengthening learning of a neural network model; based on the predicted state data corresponding to each to-be-predicted time point, generating a predicted motion trajectory of the target aircraft in the to-be-predicted time period; the trajectory determination model comprises a state transition network; the reinforcement learning process of the trajectory determination model comprises: obtaining a historical running data set of the target aircraft; wherein the historical running data set comprises a plurality of initial real state data, a plurality of action data groups and a plurality of environment data groups; at least one historical action data is included in the action data group, and at least one historical environment data is included in the environment data group; traversing each initial real state data; for each traversed initial real state data, based on the initial real state data, the action data group and the environment data group corresponding to the initial real state data, determining input data, inputting the input data into the to-be-trained state transition network, outputting historical predicted state data corresponding to the input data, based on the to-be-trained discriminator network, determining a current reward value corresponding to the input data, updating the input data based on the historical predicted state data, and inputting the updated input data into the to-be-trained state transition network, outputting historical predicted state data corresponding to the updated input data, repeating the operations of updating the input data based on the historical predicted state data, outputting the historical predicted state data and determining the current reward value until a preset ending condition is met, stopping the operation, updating the network parameters of the to-be-trained state transition network based on each current reward value obtained in the current traversal, and updating the network parameters of the to-be-trained discriminator network based on each historical predicted state data obtained in the current traversal, the state data group, the action data group and the environment data group corresponding to the initial real state data, so that the discriminator network determines the next reward value based on the updated network parameters in the next traversal operation, until the preset training condition is met and the training is ended, and the traversal of the initial real state data is stopped; wherein the state data group is a data group generated at the initial data generation time point, and the data is the initial real state data; the preset training condition comprises that the loss function corresponding to the discriminator network converges; the trajectory determination model is composed of the state transition network after training.

2. The method of claim 1, wherein, the trajectory determination model is composed of the state transition network after training. traverse each of the to-be-predicted time points in the to-be-predicted time period in chronological order of the to-be-predicted time points; for the currently traversed current to-be-predicted time point, based on the target environment data, initial state data of the target aircraft, and a preset aircraft action scheme, determine current input data corresponding to the currently traversed current to-be-predicted time point, input the current input data into a pre-trained trajectory determination model, and obtain predicted state data corresponding to the current to-be-predicted time point.

3. The method of claim 1, wherein, The target environment data includes at least one of wind speed data, wind direction data, and temperature data of the to-be-predicted flight region in the to-be-predicted time period; and the preset aircraft action scheme is composed of preset action data of a control stick of the target aircraft corresponding to each of the to-be-predicted time points.

4. The method of claim 1, wherein, The network parameters of the to-be-trained discriminator network are updated based on the generated network adversarial algorithm, the each of the historical predicted state data obtained through the current traversal, the state data group corresponding to the initial real state data, the action data group, and the environment data group. The network parameters of the to-be-trained discriminator network are updated based on the generated network adversarial algorithm, the each of the historical predicted state data obtained through the current traversal, the state data group corresponding to the initial real state data, the action data group, and the environment data group.

5. The method of claim 1, wherein, Before the state transfer network to be trained is trained based on the historical running data set, the method further includes: performing a data cleaning operation on the data in the historical running data set, and updating the historical running data set based on the data obtained after the data cleaning operation; The data cleaning operation includes at least one of missing value processing, abnormal value processing, error data processing, and repeated data processing.

6. An aircraft motion trajectory determination apparatus, characterized in that, includes: a target environment data determination module configured to determine target environment data of a target aircraft corresponding to a to-be-predicted flight region in a to-be-predicted time period; the to-be-predicted time period includes at least one to-be-predicted time point; a predicted state data determination module configured to determine predicted state data of the target aircraft corresponding to each of the to-be-predicted time points in the to-be-predicted time period based on the target environment data, initial state data of the target aircraft, a preset aircraft action scheme, and a pre-trained trajectory determination model; the trajectory determination model is obtained by performing reinforcement learning on a neural network model in advance; a predicted running trajectory generation module configured to generate a predicted motion trajectory of the target aircraft in the to-be-predicted time period based on the predicted state data corresponding to each of the to-be-predicted time points; The device further includes: a historical running data set acquisition module configured to acquire a historical running data set of the target aircraft; the historical running data set is composed of a plurality of initial real state data, a plurality of action data groups, and a plurality of environment data groups; the action data group includes at least one historical action data, and the environment data group includes at least one historical environment data; a trajectory determination model composition module configured to train a state transfer network to be trained based on the historical running data set, and obtain the state transfer network after training as the trajectory determination model; the trajectory determination model composition module includes: The current reward value determination submodule is configured to traverse each starting real state data; for each traversed starting real state data, determine input data based on the starting real state data, an action data group corresponding to the starting real state data, and an environment data group; input the input data into the state transition network to be trained, output historical predicted state data corresponding to the input data, determine a current reward value corresponding to the input data based on the discriminator network to be trained, update the input data based on the historical predicted state data, and input the updated input data into the state transition network to be trained, output historical predicted state data corresponding to the updated input data, repeat the operations of updating the input data based on the historical predicted state data, outputting the historical predicted state data, and determining the current reward value until a preset ending condition is met, stop the operations, update network parameters of the state transition network to be trained based on each current reward value obtained in the current traversal, and update network parameters of the discriminator network to be trained based on each historical predicted state data obtained in the current traversal, the state data group corresponding to the starting real state data, the action data group, and the environment data group, so that the discriminator network determines a next reward value based on the updated network parameters in the next traversal operation until a preset training condition is met, the training is ended, and the traversal of the starting real state data is stopped; the state data group is a data group generated at an initial data generation moment, and the data in the data group is the starting real state data; the preset training condition includes convergence of a loss function corresponding to the discriminator network; and the trajectory determination model is composed of the state transition network after the training is ended.

7. An electronic device, comprising: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the aircraft motion trajectory determination method in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the aircraft motion trajectory determination method in any one of claims 1-5 when executed.

Citation Information

Patent Citations

  • Hybrid target track prediction method and system

    CN114819068A

  • A system and method for optimising flight efficiency

    WO2023242433A1