Methods, devices and equipment for aircraft trajectory planning and communication resource allocation

CN122575191APending Publication Date: 2026-08-14BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]本申请提供一种飞行器航迹规划和通信资源分配方法、装置及设备,用以解决现有技术缺乏对轨迹规划与通信资源的联合规划,导致了资源浪费的问题

Benefits of technology

[0048]本申请提供的一种飞行器航迹规划和通信资源分配方法、装置及设备,包括:获取飞行器组的状态数据以及位置数据;根据飞行器组的位置数据,确定无人机机组的时序特征;根据状态数据与时序特征,确定主飞行器的执行参数,执行参数包括轨迹控制参数及通信接入参数,其中,轨迹控制参数用于控制主飞行器的飞行轨迹,通信接入参数用于调整主飞行器通信接入的无人机;根据执行参数控制飞行器组。相较于现有技术缺乏对轨迹规划与通信资源的联合规划,导致了资源浪费。本申请通过点云时空特征提取与状态数据的联合规划,生成包括轨迹控制参数及通信接入参数的执行参数,实现对飞行轨迹与通信资源分配的协同优化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575191A_ABST
    Figure CN122575191A_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, and device for aircraft trajectory planning and communication resource allocation, relating to the field of aircraft trajectory planning. Applied to an aircraft group, the aircraft group includes a main aircraft and a drone group, with the drone group including at least one drone. The method includes: acquiring the status data and position data of the aircraft group; determining the temporal characteristics of the drone group based on the position data of the aircraft group; determining the execution parameters of the main aircraft based on the status data and temporal characteristics, the execution parameters including trajectory control parameters and communication access parameters, wherein the trajectory control parameters are used to control the flight trajectory of the main aircraft, and the communication access parameters are used to adjust the communication access of the drones of the main aircraft; and controlling the aircraft group according to the execution parameters. This application solves the technical problem of the prior art lacking joint planning of trajectory planning and communication resources, leading to resource waste.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of aircraft trajectory planning, and in particular to a method, apparatus and equipment for aircraft trajectory planning and communication resource allocation. Background Technology

[0002] With the acceleration of urbanization and the exacerbation of ground traffic congestion, urban air transportation is becoming an important component of the next generation of transportation systems. Electric vertical takeoff and landing (EVTOL) aircraft and drones are widely used in low-altitude airspace for logistics delivery, urban passenger transport, and emergency rescue.

[0003] Current aircraft mainly rely on pre-given environmental prior information to complete trajectory planning, and then use a single indicator, such as communication rate or sensing accuracy, to plan the flight trajectory and communication resources separately.

[0004] However, existing technologies lack joint planning of trajectory planning and communication resources, leading to resource waste. Summary of the Invention

[0005] This application provides a method, apparatus, and equipment for aircraft trajectory planning and communication resource allocation, in order to solve the problem of resource waste caused by the lack of joint planning of trajectory planning and communication resources in the prior art.

[0006] In a first aspect, this application provides a method for aircraft trajectory planning and communication resource allocation, applied to an aircraft group, the aircraft group including a main aircraft and an unmanned aerial vehicle (UAV) group, wherein the UAV group includes at least one UAV, the method comprising:

[0007] Acquire the status and position data of the aircraft group. The status data includes the mission data of the aircraft group, the communication data between the main aircraft and the UAV group, and the energy consumption data of the aircraft group.

[0008] Based on the position data of the aircraft group, the temporal characteristics of the UAV group are determined, where the temporal characteristics represent the movement trend of the UAV group over time;

[0009] Based on the status data and timing characteristics, the execution parameters of the main aircraft are determined. The execution parameters include trajectory control parameters and communication access parameters. The trajectory control parameters are used to control the flight trajectory of the main aircraft, and the communication access parameters are used to adjust the communication access of the main aircraft to the UAV.

[0010] Control the aircraft group according to the execution parameters.

[0011] In one possible design, the execution parameters of the main aircraft are determined based on state data and timing characteristics, including:

[0012] The state data and time series features are integrated into the state features of the aircraft group;

[0013] The state features are input into the first policy network to generate execution parameters. The first policy network is a multi-layer nonlinear transformation algorithm.

[0014] In one possible design, the location data is point cloud data acquired at each time point;

[0015] Based on the position data of the aircraft group, determine the temporal characteristics of the UAV group, including:

[0016] Feature extraction is performed on the point cloud data at each time point to obtain the four-dimensional features of each UAV; the four-dimensional features include the distance features of the UAV relative to the host aircraft, radial velocity features, azimuth features, and pitch angle features.

[0017] The four-dimensional features are aggregated into a global feature vector using the max pooling algorithm;

[0018] The global feature vector is input into multiple fully connected layers, and features are extracted layer by layer through activation functions in each of the multiple fully connected layers to obtain the final global feature vector.

[0019] The long short-term memory algorithm is used to extract features from the final global feature vector to obtain temporal features.

[0020] In one possible design, the trajectory control parameters include the repulsive force parameters, tangential guidance strength parameters, and tangential guidance direction parameters of the artificial potential field model, which is used to simulate the trajectory operation of the aircraft group.

[0021] Controlling the aircraft group according to execution parameters includes:

[0022] Input the repulsive force parameter, tangential guidance intensity parameter, and tangential guidance direction parameter into the artificial potential field model to generate trajectory adjustment parameters;

[0023] Adjust the flight trajectory of the main aircraft according to the trajectory adjustment parameters;

[0024] Adjust the communication access parameters of the main aircraft to the corresponding UAV.

[0025] In one possible design, after controlling the aircraft group according to the execution parameters, the following is also included:

[0026] Acquire the status update data of the aircraft group. The status update data includes the update task data of the aircraft group after executing the execution parameters, the update communication data between the main aircraft and the UAV group, and the update energy consumption data of the aircraft group.

[0027] Based on the preset reward function, state data, and updated state data, calculate the reward benefits of the main aircraft after executing the execution parameters;

[0028] Periodically acquire multiple pieces of experience data, including: status data, execution parameters, updated status data, and reward benefits;

[0029] A predetermined number of empirical data points are selected as training samples from multiple empirical data sets.

[0030] The first policy network is optimized using training samples.

[0031] In one possible design, the first policy network is optimized using training samples, including:

[0032] Obtain each updated state data from the empirical data in the training samples;

[0033] Each updated state data is input into the second policy network to generate update execution parameters corresponding to each experience data. The policy parameter values ​​of the second policy network are different from those of the first policy network.

[0034] Each update execution parameter is input into the first value network to generate a target value score for each piece of experience data. The target value score is used to represent the future return corresponding to the updated state data.

[0035] Each piece of empirical data from the training samples is input into the second value network to generate a predicted value score corresponding to each piece of empirical data. The predicted value score is used to represent the future return corresponding to the state data. The value parameter values ​​of the second value network are different from those of the first value network.

[0036] Based on each target value score and each predicted value score, the value parameter values ​​corresponding to the second value network are adjusted to obtain the optimized second value network.

[0037] Based on the policy gradient algorithm and the predicted value score, the policy parameter values ​​corresponding to the first policy network are adjusted to obtain the optimized first policy network.

[0038] Secondly, this application provides an apparatus for aircraft trajectory planning and communication resource allocation, comprising:

[0039] The acquisition module is used to acquire the status data and position data of the aircraft group. The status data includes the mission data of the aircraft group, the communication data between the main aircraft and the UAV group, and the energy consumption data of the aircraft group.

[0040] The first determining module is used to determine the temporal characteristics of the UAV group based on the position data of the aircraft group, wherein the temporal characteristics represent the motion trend of the UAV group over time;

[0041] The second determining module is used to determine the execution parameters of the main aircraft based on the status data and timing characteristics. The execution parameters include trajectory control parameters and communication access parameters. The trajectory control parameters are used to control the flight trajectory of the main aircraft, and the communication access parameters are used to adjust the communication access of the main aircraft to the UAV.

[0042] The control module is used to control the aircraft group according to the execution parameters.

[0043] Thirdly, this application provides an aircraft trajectory planning and communication resource allocation device, including: a memory and a processor;

[0044] The memory stores the instructions that the computer executes;

[0045] The processor executes computer execution instructions stored in memory, causing the processor to execute the aircraft trajectory planning and communication resource allocation method as described in the first aspect of the invention.

[0046] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the aircraft trajectory planning and communication resource allocation method as described in the first aspect of the invention.

[0047] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the aircraft trajectory planning and communication resource allocation method described in the first aspect of the invention.

[0048] This application provides a method, apparatus, and device for aircraft trajectory planning and communication resource allocation, comprising: acquiring state data and position data of an aircraft group; determining the temporal characteristics of the UAV group based on the position data of the aircraft group; determining the execution parameters of the main aircraft based on the state data and temporal characteristics, the execution parameters including trajectory control parameters and communication access parameters, wherein the trajectory control parameters are used to control the flight trajectory of the main aircraft, and the communication access parameters are used to adjust the communication access of the UAVs of the main aircraft; and controlling the aircraft group according to the execution parameters. Compared with the prior art, which lacks joint planning of trajectory planning and communication resources, resulting in resource waste, this application generates execution parameters including trajectory control parameters and communication access parameters through joint planning of point cloud spatiotemporal feature extraction and state data, thereby achieving collaborative optimization of flight trajectory and communication resource allocation. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 A schematic diagram of a system architecture for an aircraft trajectory planning and communication resource allocation method provided in this application embodiment;

[0051] Figure 2 A flowchart illustrating an aircraft trajectory planning and communication resource allocation method provided in this application embodiment. Figure 1 ;

[0052] Figure 3 This is a schematic diagram of the eVTOL-UAV-BS channel architecture provided in the embodiments of this application;

[0053] Figure 4a A schematic diagram of the azimuth angle between the eVTOL and the UAV provided in the embodiments of this application;

[0054] Figure 4b A schematic diagram showing the elevation angle of the eVTOL and the UAV provided in the embodiments of this application;

[0055] Figure 5 This is a schematic diagram of the aircraft group architecture in a sensor-integrated scenario provided in the embodiments of this application;

[0056] Figure 6 A flowchart illustrating an aircraft trajectory planning and communication resource allocation method provided in this application embodiment. Figure 2 ;

[0057] Figure 7a A schematic diagram of a deep reinforcement learning architecture based on temporal point clouds provided in an embodiment of this application;

[0058] Figure 7b A schematic diagram of a deep reinforcement learning process based on temporal point clouds provided for an embodiment of this application;

[0059] Figure 8 A schematic diagram of the structure of the aircraft trajectory planning and communication resource allocation device provided in the embodiments of this application;

[0060] Figure 9 This is a schematic diagram of the structure of an aircraft trajectory planning and communication resource allocation device provided in an embodiment of this application. Detailed Implementation

[0061] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0062] In the embodiments of this application, the terms "first" and "second" are used to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, nor do they necessarily imply difference. It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner. In the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more.

[0063] It should be noted that the phrase "at...time" in the embodiments of this application can refer to the instant at which a certain situation occurs, or to a period of time after the occurrence of a certain situation; the embodiments of this application do not specifically limit this. Furthermore, the aircraft trajectory planning and communication resource allocation method provided in the embodiments of this application is merely an example; the aircraft trajectory planning and communication resource allocation method may also include more or less content.

[0064] With the acceleration of global urbanization and the increasing severity of ground traffic congestion, urban air mobility is becoming a key component of the next generation of urban transportation systems. It provides a new dimension for alleviating urban traffic pressure by expanding the perspective of transportation from the traditional two-dimensional ground to the three-dimensional low-altitude airspace. Among these technologies, unmanned aerial vehicles (UAVs) and electric vertical take-off and landing (eVTOL) aircraft, with their advantages of high maneuverability, low noise, and low emissions, have become the core carriers for low-altitude passenger and freight transport within and between cities, significantly improving urban transportation efficiency. Therefore, they are widely considered enabling technologies that will reshape the three-dimensional urban transportation landscape and have the potential to spawn large-scale emerging markets in the future.

[0065] With the rapid development of urban air traffic and the low-altitude economy, the demand for autonomous flight, dynamic obstacle avoidance, and collaborative resource optimization of low-altitude aircraft in complex low-altitude environments is becoming increasingly prominent.

[0066] Unlike traditional drones primarily used for cargo transport and inspection, eVTOLs typically undergo five phases in urban aerial missions: takeoff, climb, cruise, approach, and landing. The approach phase, due to its complex operating environment and demanding flight maneuvers, represents the most concentrated area of ​​safety risks. The challenges faced by eVTOLs in this phase are multi-dimensional: at the aircraft level, it requires high-precision trajectory control and complex energy management; at the environmental level, it must respond in real-time to static obstacles such as buildings and power lines, as well as the uncertainties introduced by other dynamic flying targets. The presence of these dynamic targets severely challenges the traditional perception-decision closed loop based on cooperative communication and prior maps, revealing the limitations of directly transferring existing integrated communication and perception models from existing drones to eVTOL passenger transport scenarios. Therefore, developing a new airborne perception paradigm capable of real-time, robust perception and response to dynamic targets has become an urgent need to ensure the safe and reliable operation of UAMs (Urban Aerial Vehicles).

[0067] As a key component in realizing the aforementioned perception-decision closed loop, trajectory optimization technology for aircraft has been extensively studied, aiming to avoid collisions between eVTOLs or drones in dense airspace.

[0068] Optionally, existing technologies propose a hierarchical trajectory planning framework that combines an improved hybrid A* algorithm with local corridor planning based on Marden's theorem to ensure trajectory safety and smoothness.

[0069] Optionally, existing technologies also propose an adaptive evolutionary multi-objective distribution estimation algorithm, which optimizes high-dimensional objectives through mechanisms such as adaptive selection rate and covariance matrix evolution.

[0070] Optionally, existing technologies also propose an iterative algorithm based on the extended Kalman filter framework to solve the problem of optimizing UAV trajectories.

[0071] However, most of the aforementioned studies rely on pre-given environmental information and lack real-time perception capabilities for dynamic obstacles, leading to trajectory failures in dynamic scenarios. Therefore, using real-time point clouds from airborne millimeter-wave radar to drive trajectory optimization is more closely aligned with real-world urban scenarios and has significant practical implications. However, the aforementioned trajectory optimization studies generally design perception, communication, and trajectory optimization modules in a fragmented manner, resulting in low spectrum utilization and high information latency. To address the frequency contention problem between perception and communication, integrated communication and perception technology has been proposed. Current research focuses on optimizing UAV trajectories while simultaneously improving communication fairness and system energy efficiency.

[0072] Optionally, existing technologies propose an overall optimization problem, namely, optimizing user association and drone power allocation while ensuring the accuracy of multi-drone trajectories.

[0073] Optionally, existing technologies have designed an integrated UAV-communication and sensing system that jointly optimizes UAV trajectory, user association, target perception selection, and transmission beamforming, while maximizing the system's achievable rate under the constraints of the perception frequency and beam pattern gain of a given target.

[0074] However, these studies are all limited to two-dimensional coordinates, treating UAVs as point targets without using real multi-dimensional radar point clouds, and failing to fully consider the impact of dynamic obstacles in trajectory optimization. Furthermore, these studies are mostly constrained by fixed trajectories or single performance metric optimization (such as communication rate or sensing accuracy), failing to systematically coordinate trajectory flexibility, energy efficiency improvement, and the integrated sensing-communication requirements. Therefore, research on jointly optimizing trajectories and resource allocation based on millimeter-wave radar sensing results is of significant value.

[0075] In summary, the application of integrated communication and sensing technology in UAVs is relatively mature. Numerous studies in recent years have demonstrated that this technology enables UAVs to be applied in various fields such as cargo transportation and disaster relief. However, in the low-altitude flight systems of future smart cities, the energy consumption function and safety threshold of eVTOL differ significantly from those of UAVs. While existing eVTOL research has made some progress, research on the integrated communication-sensing-control design of eVTOLs in low-altitude airspace remains scarce. Most studies still separate environmental perception, trajectory planning, and resource optimization, and typically rely on static maps or prior location information to simplify the modeling of dynamic obstacles. This approach fails to meet the demands of dynamic UAV interference, obstacle uncertainty, and multi-objective coupled optimization in real urban airspace.

[0076] The existing technical solutions have the following limitations:

[0077] Optionally, dynamic obstacle modeling is coarse: existing technologies simplify dynamic drones into static point targets or ideal geometric objects without considering their actual motion trajectory and spatial structure characteristics, resulting in insufficient obstacle avoidance safety.

[0078] Optionally, the perception-communication-trajectory modules are separated: existing methods design environmental perception, trajectory planning and communication resource allocation independently, lacking joint utilization of real-time perception information of dynamic obstacles, resulting in low spectrum utilization and large information delay.

[0079] To address the aforementioned problems, the inventors, during their research on the low efficiency of aircraft trajectory planning and communication resource allocation, discovered that existing technologies lack joint planning for trajectory planning and communication resources, leading to resource waste. Therefore, the inventors considered using point cloud spatiotemporal feature extraction and state data joint planning to generate execution parameters, including trajectory control parameters and communication access parameters, to achieve coordinated optimization of flight trajectory and communication resource allocation. Based on this, embodiments of this application provide an aircraft trajectory planning and communication resource allocation method, apparatus, and device, applicable to the field of aircraft trajectory planning, aiming to solve the problem of low efficiency in existing aircraft trajectory planning and communication resource allocation technologies.

[0080] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0081] Figure 1 This is a schematic diagram of a system architecture for an aircraft trajectory planning and communication resource allocation method provided in an embodiment of this application. The aircraft trajectory planning and communication resource allocation system is a computer device. Figure 1 In the above architecture, at least one of data acquisition device 101, processing device 102 and display device 103 is included.

[0082] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the processing system architecture of the aircraft trajectory planning and communication resource allocation method. In other feasible embodiments of this application, the above architecture may include more or fewer components than illustrated, or combine some components, or divide some components, or arrange different components, which can be determined according to the actual application scenario and is not limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of both.

[0083] In the specific implementation process, the data acquisition device 101 may include an input / output interface or a communication interface. The data acquisition device 101 can be connected to the processing device through the input / output interface or the communication interface to obtain the status data and position data of the aircraft group.

[0084] The processing device 102 can obtain the execution parameters of the main aircraft based on the status data and position data of the aircraft group.

[0085] The display device 103 can also be a touch screen or the screen of a terminal device, used to receive user commands while displaying the execution parameters of the main aircraft, so as to realize interaction with the user.

[0086] It should be understood that the aforementioned processing device can be implemented by a processor reading instructions from memory and executing those instructions, or it can be implemented by a chip circuit.

[0087] Furthermore, the network architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0088] The technical solution of this application will be described in detail below with reference to specific embodiments:

[0089] Figure 2 A flowchart illustrating an aircraft trajectory planning and communication resource allocation method provided in this application embodiment. Figure 1 ,like Figure 2 As shown, this method is applied to an aircraft group, which includes a main aircraft and a drone group, wherein the drone group includes at least one drone.

[0090] S201. Acquire the status data and position data of the aircraft group.

[0091] The status data includes mission data of the aircraft group, communication data between the main aircraft and the UAV group, and energy consumption data of the aircraft group.

[0092] In this embodiment, during the approach of the main aircraft, the onboard millimeter-wave radar of the main aircraft actively transmits signals and receives echo signals from the UAVs at each moment. The echo signals are then converted into range and velocity dimension spectra using a Fast Fourier Transform, thereby obtaining the position data of the main aircraft and surrounding UAVs.

[0093] The approach process is a crucial and complex phase in the flight mission, marking the period from the start to the end of the flight mission.

[0094] In this embodiment, during the approach of the main aircraft, the perception portion of the dynamic environment—state data—is continuously acquired. This state data includes mission data of the aircraft group, communication data between the main aircraft and the UAV group, and energy consumption data of the aircraft group.

[0095] Optionally, mission data includes the status information of the start and end points of the flight mission, communication data includes the communication capacity and Age of Information (AoI) between the main aircraft and the UAV crew, and energy consumption data of the aircraft crew includes energy consumption and energy efficiency.

[0096] Optionally, the status information of the start and end points of the flight mission includes the location information, time information, speed information, flight attitude information, and system information of the start and end points of the flight mission.

[0097] Optionally, the communication capacity is obtained from the communication model. The communication model consists of the following:

[0098] For example, since the time scale of trajectory planning is much larger than the small-scale fading variation period of the wireless channel, a channel model between the main aircraft, the UAV crew, and the base station (BS) was established to focus on the large-scale channel characteristics.

[0099] In one possible embodiment, Figure 3 This is a schematic diagram of the eVTOL-UAV-BS channel architecture provided in the embodiments of this application, as shown below. Figure 3 As shown, the channel state between eVTOL and the UAV is affected by the low-altitude environment (such as other UAVs), and may be in line-of-sight (LoS) or non-line-of-sight (NLoS) conditions. Therefore, we introduce a probabilistic model to characterize the probability of LoS / NLoS links occurring.

[0100] Optionally, the line-of-sight channel probability between eVTOL and the UAV is:

[0101]

[0102]

[0103] in, : Represents the probability value of the line-of-sight channel between eVTOL and UAV at time t. Its value ranges from 0 to 1. The closer the value is to 1, the greater the probability of the line-of-sight channel existing. The closer the value is to 0, the smaller the probability of the line-of-sight channel existing. exp: Natural exponential function. : is a function related to the height difference. At time t, it reflects the impact of the relative positional relationship between the eVTOL and the UAV in the vertical direction on the line-of-sight channel probability. : Represents the spatial distance (norm form, usually understood as the straight-line distance between two points) between UAV u and eVTOLe at time t. This distance affects the probability of the existence of a line-of-sight channel. Generally speaking, the greater the distance, the lower the probability of the existence of a line-of-sight channel; arcsin: arcsine function; At time t, the coordinates of eVTOL in the vertical direction (usually understood as the height direction); The coordinates of the UAV u in the vertical direction at time t; These represent parameters related to the environment.

[0104] Furthermore, the probabilistic non-line-of-sight channel gain of eVTOL and UAVs can be expressed as:

[0105]

[0106] in, Indicates the path loss index; The gain, representing the probabilistic non-line-of-sight reference distance of 1 meter, can be calculated using the following formula:

[0107]

[0108] in, This indicates additional loss beyond line-of-sight. Represents the probability of non-line-of-sight distance; This is the gain at a reference distance of 1 meter; where and These are the antenna gains for the UAV transmitter and the eVTOL receiver, respectively. It is the wavelength.

[0109] Furthermore, the line-of-sight channel gain of eVTOL and UAVs It can be represented as:

[0110]

[0111] in, This represents the path loss index for line-of-sight channels.

[0112] Furthermore, when eVTOL communicates with the uth UAV, we introduce a decision variable. To represent the line-of-sight situation between eVTOL and the UAV, and the channel gain of eVTOL and the UAV. and The correspondence can be expressed as:

[0113]

[0114] Among them, if =1, then it is determined to be pure line-of-sight distance. If the value is 0, it is determined to be probabilistic non-line-of-sight. Therefore, the Signal to Interference plus Noise Ratio (SINR) can be expressed as:

[0115]

[0116] in, It is the power of additive white Gaussian noise in the eVTOL and UAV channels. This indicates the transmission power between the drone and eVTOL. This indicates the interference power of drones other than the ones that have accessed the channel.

[0117] Furthermore, the jamming drones are line-of-sight; therefore, eVTOL affects the channel capacity of the drones. It can be represented as:

[0118]

[0119] in, Indicates bandwidth.

[0120] Optionally, communication between eVTOL and the base station is affected by surrounding buildings; therefore, the communication model is constructed as a probabilistic non-line-of-sight channel. The channel gain of the eVTOL-base station channel is also considered. It can be represented as:

[0121]

[0122] in, b represents the path loss index; b represents the spatial coordinates of the base station. This represents the spatial distance (norm form, usually understood as the straight-line distance between two points) between base station b and eVTOLe at time t. This distance affects the channel gain; generally, the greater the distance, the smaller the channel gain may be. The calculation method and Similarly, it is determined by the probabilistic non-line-of-sight model.

[0123] Furthermore, since the drone is relatively far from the eVTOL when transmitting data to the base station, we assume that the communication between the eVTOL and the base station is not interfered with by the drone. Therefore, the signal-to-noise ratio (SNR) of the eVTOL-base station channel is... for:

[0124]

[0125] in, This represents the power of eVTOL and additive white noise in the base station channel.

[0126] Furthermore, the channel capacity of eVTOL and base stations. It can be represented as:

[0127]

[0128] in, Indicates bandwidth.

[0129] Furthermore, to ensure that all communication data can be uploaded from the U-shaped drone to the base station, the constraints can be expressed as follows:

[0130]

[0131] in, As decision variables, The time indicates that there is an access between eVTOL and UAV u in time slot t.

[0132] Optionally, communication freshness is obtained from an information age model, the composition of which is as follows:

[0133] For example, communication freshness is a key indicator for measuring the freshness of information in a wireless communication system. It is defined as the difference between the timestamp of the moment the receiver receives the latest data packet and the current time. The definition of communication freshness helps quantify the real-time performance and efficiency of information transmission in a wireless communication system.

[0134] Furthermore, the entire service cycle T is discretized into... There are several equal-length time intervals, each interval having a length of [missing information]. The time index is Within any time interval n, assuming the positions of eVTOL and the UAV remain unchanged, AoI state updates only occur at the interval boundaries.

[0135] Furthermore, let's assume The AoI (Aspect of Interest) representing mission data from drones is calculated using the following formula:

[0136]

[0137] in, It is the timestamp when the drone was activated.

[0138] Furthermore, we can derive the UAV's state information from the t-th time interval, which can be represented as:

[0139]

[0140] Within time slot n, If eVTOL successfully connects and receives data from the drone, the drone's AoI will be reset at the start of the next time slot. ; If a drone is not served within time slot t, its information age will accumulate to the length of one time slot. .

[0141] Furthermore, the average AoI for all drone state updates can be expressed as:

[0142]

[0143] Specifically, the overall communication freshness of the system in any time slot n is measured by the average AoI of all UAVs.

[0144] Optionally, energy consumption and energy efficiency are obtained from an energy consumption model, the composition of which is as follows:

[0145] For example, eVTOL is equipped with Each propeller consumes a total energy of propulsion energy. and communication energy consumption composition.

[0146] Furthermore, within a continuous time interval t, the approach trajectory of the eVTOL can be decomposed into a superposition of horizontal and vertical directions. Therefore, the propulsion power of the eVTOL consists of the power in the horizontal direction. and vertical power express. It can be represented as:

[0147]

[0148] in, It is the thrust of a single propeller. This is the total weight of the eVTOL. Indicates the area of ​​the propeller disk. This represents the propeller's correction factor, used to correct the gap between the ideal theoretical efficiency and the actual efficiency; its value is 0.9. This indicates the air density at the current altitude of the eVTOL.

[0149] Optional, The calculation formula is as follows:

[0150]

[0151] in, This indicates the number of propellers.

[0152] Optional, The calculation formula is as follows:

[0153]

[0154] in, This indicates the radius of the propeller disk.

[0155] Optional, The calculation formula is as follows:

[0156]

[0157] in, The density of air at sea level. Indicates altitude Temperature at the location , The value represents the sea level temperature, and g is the acceleration due to gravity. This represents the vertical temperature lapse rate, where M is the air gas constant. Altitude It represents the difference between the local ground level and the sea level.

[0158] Furthermore, the horizontal propulsion power The specific calculation formula is as follows;

[0159]

[0160] in, This represents the total air resistance experienced by the eVTOL. Indicates horizontal velocity. Total air resistance. The specific calculation formula is as follows:

[0161]

[0162] in, This represents the air density at height h. This represents the surface area of ​​the eVTOL in contact with air. This is the drag coefficient, representing the starting efficiency of the eVTOL. The specific calculation formula is as follows:

[0163]

[0164] in, Indicates the zero-lift drag coefficient. Indicates the lift coefficient. Indicates the aspect ratio of the wing. The Oswald efficiency factor is represented by the following formula:

[0165]

[0166] Furthermore, the propulsion energy consumption of eVTOL The specific calculation formula is as follows:

[0167]

[0168] in, This represents the power used by the eVTOL for hovering at time t. Hovering is an important flight state for the eVTOL, during which the propulsion system needs to generate sufficient upward thrust to balance the eVTOL's gravity and keep it at a specific altitude. It is the power required to maintain this hovering state, and it varies with factors such as time.

[0169] in, This represents the power required by the eVTOL at time t for propulsion-related actions other than hovering (such as attitude adjustment and horizontal movement). During flight, the eVTOL needs to perform various actions besides hovering to meet the requirements of the flight mission. The power used for these actions will also vary with time and other factors.

[0170] Furthermore, the energy consumption associated with eVTOL communication includes the power consumption of circuit elements when receiving and transmitting mission data signals, treating the communication-related power as a constant. eVTOL communication power consumption It can be represented as:

[0171]

[0172] in, It is a binary variable representing the communication link access status of eVTOL.

[0173] In summary, the overall energy consumption of eVTOL can be expressed as:

[0174]

[0175] S202. Determine the timing characteristics of the UAV group based on the position data of the aircraft group.

[0176] Among them, the temporal characteristics represent the movement trend of the drone unit over time.

[0177] The location data consists of point cloud data acquired at each time point.

[0178] In this embodiment, at each moment, based on the position data of the aircraft group, the four-dimensional features of each UAV (distance features, radial velocity features, azimuth features, and pitch features of the UAV relative to the main aircraft) are obtained, and the temporal features reflecting the dynamic changes of the environment (the changes of the UAV group over time) are extracted based on the four-dimensional features.

[0179] Among them, the four-dimensional features are generated by the radar perception model, and the composition formula of the radar perception model is as follows:

[0180] For example, the entire service cycle is time T. The positions of eVTOL and UAV are considered to remain constant within the same time interval t, but they can change in consecutive time intervals.

[0181] Specifically, a three-dimensional Cartesian coordinate system is used, and eVTOL is used in conjunction with each UAV ( The horizontal position changing over time is represented as follows:

[0182]

[0183] in, The eVTOL position coordinates at time t are the position state vector, which is a 3×1 column vector; It is a three-dimensional real number space; The coordinates of each UAV are specified; each coordinate is subject to the following constraints: , and ; These represent the x-axis, y-axis, and z-axis coordinates of the topology, respectively.

[0184] Specifically, in the radar perception model, UAVs can be detected by... It is represented by an ellipsoidal radar point cloud (RPC) composed of points.

[0185] in, , It is the maximum number of points in the sample.

[0186] Specifically, the Cartesian coordinates of the i-th point in the UAV's RPC can be represented as:

[0187]

[0188] in, ; This represents the Cartesian coordinate vector of the i-th point in the UAV's RPC at time t relative to a certain reference coordinate system (usually an inertial coordinate system or a local horizontal coordinate system, etc.). It is a 3×13×1 column vector, that is, it has three dimensions.

[0189] Specifically, since the eVTOL has sensing capabilities, the distance between the eVTOL and the (i, u)th point of the UAV u is... It can be obtained through the following formula:

[0190]

[0191] Where T represents the transpose of a vector.

[0192] The relative position vector between eVTOL and UAVs can be expressed as:

[0193]

[0194] In our system, both eVTOLs and UAVs are mobile. The speed of the eVTOL can be expressed as... The horizontal and vertical velocities of the drone u at the i-th point can be obtained through... Calculate the relative velocity between the eVTOL and the drone u at the (i, u)th point. The relative velocity can be expressed as:

[0195]

[0196] Optional, Figure 4a A schematic diagram of the azimuth angle between the eVTOL and the UAV provided in the embodiments of this application; Figure 4b This is a schematic diagram showing the elevation angle of the eVTOL and the UAV provided in an embodiment of this application. Figure 4a , Figure 4b As shown, in addition to distance (R) and velocity (V), the present invention also selected azimuth angle ( ) and elevation angle ( () as a four-dimensional feature.

[0197] Specifically, the azimuth and elevation angles between eVTOL and the (i, u)th point of the u-th UAV are respectively and It can be obtained through calculation.

[0198]

[0199]

[0200] in, : The coordinate value of eVTOL in the y-axis direction at time t; : The coordinates of the (i,u)th point of the u-th UAV in the y-axis direction at time t; : The coordinate value of eVTOL in the x-axis direction at time t; : The coordinates of the (i,u)th point of the u-th UAV in the x-axis direction at time t; : The coordinate value of eVTOL in the z-axis direction at time t; : The coordinates of the (i,u)th point of the u-th UAV in the z-axis direction at time t.

[0201] More specifically, by modeling with radar perception models, four-dimensional features (R, V, ...) of the UAV can be obtained. , ).

[0202] Specifically, the implementation steps of S202 include:

[0203] First, feature extraction is performed on the point cloud data at each time point to obtain the four-dimensional features of each UAV.

[0204] For example, point cloud data at various times is input into a radar perception model to generate four-dimensional features (R, V, ...) for each UAV. , ).

[0205] Among them, the four-dimensional features include the distance features of the UAV relative to the host aircraft, the radial velocity features, the azimuth features, and the pitch features.

[0206] For example, after generating four-dimensional features, there are also processing steps: multiple four-dimensional features are mapped to high dimensions through a multilayer perceptron (MLP) model. A nonlinear transformation algorithm is introduced into the multilayer perceptron model to extract more complex feature representations of the point cloud—point-by-point high-dimensional features.

[0207] Specifically, this MLP model employs a weight-sharing mechanism, which not only significantly reduces the number of model parameters but also enhances the model's generalization ability and translation invariance across different points.

[0208] Furthermore, in the MLP model, batch normalization (BN) is performed after each layer to accelerate training convergence and improve the stability of the MLP model. At time t, the output of the MLP model can be expressed as:

[0209]

[0210] in, Represented as point-by-point high-dimensional features, It is represented as the ReLU activation function, BN[] represents the normalization operation, and MLP() represents the convolution operation.

[0211] Secondly, the four-dimensional features are aggregated into a global feature vector using the max pooling algorithm.

[0212] For example, in order to handle the independence of point order and variability of point quantity while preserving key structural information of the point cloud, max-pooling is used. The pointwise high-dimensional features output by the MLP model are aggregated to obtain the global feature vector.

[0213] Specifically, for each point-by-point high-dimensional feature, the maximum value among all points is independently taken, thus obtaining a global feature vector whose point arrangement order remains unchanged and whose dimension is fixed. , can be represented as:

[0214]

[0215] Next, the global feature vector is input into multiple fully connected layers, and features are extracted layer by layer through activation functions in the multiple fully connected layers to obtain the final global feature vector.

[0216] For example, to further enhance the expressive power of features and achieve the mapping from the global feature vector of the point cloud to high-level semantic features, this invention will aggregate the obtained global feature vector. The data is then fed into a feature enhancement model consisting of multiple fully connected layers for further transformation.

[0217] Specifically, this model extracts more discriminative and high-level semantic feature representations through multi-layer nonlinear transformations, providing richer feature inputs for subsequent time-series modeling. The specific calculation formula is as follows:

[0218]

[0219] Here, FC() represents a fully connected operation. Let represent the activation function. For consistency, the final global feature vector output by the FC layer is denoted as . .

[0220] For example, if there are three fully connected layers, the global feature vector is input into the first layer, and the first layer outputs... ,Bundle As the input to the second layer, the output of the second layer... ,Bundle As input to the third layer, the third layer finally outputs the final global feature vector.

[0221] Finally, the long short-term memory algorithm is used to extract features from the final global feature vector to obtain temporal features.

[0222] For example, in order to fully explore the dynamic evolution patterns in UAV time-series point cloud data, the final global feature vectors at T consecutive time points are... The data is fed into a Long Short-Term Memory (LSTM) network to extract time-dependent temporal features.

[0223] Specifically, at time step t, the LSTM model shows the cell state at the previous time step. Hidden state and the current input As input, the cell state is the core memory carrier that is transmitted throughout the entire sequence along the time axis and is the physical basis for realizing long-term dependency modeling; the hidden state is the instantaneous representation output by the network at each time step, carrying the sequence summary information that is most effective for completing the current task up to the current moment, namely trajectory information and communication access information.

[0224] Furthermore, the updated cell state is output after calculation via the gating mechanism. With hidden state The gating mechanism inside LSTM can be represented as:

[0225]

[0226] in, This can be represented as vector concatenation. This represents the sigmoid activation function. For element-wise multiplication, This represents the hyperbolic tangent activation function. , , , and , , , These are the learnable parameters of the LSTM model. Through the aforementioned gating mechanism, LSTM can effectively capture the dynamic evolution pattern of UAV point cloud features over time, and its final output is the hidden state. This serves as a temporal feature describing the dynamics of the environment, which is then input into the subsequent reinforcement learning decision model.

[0227] S203. Determine the execution parameters of the main aircraft based on the status data and timing characteristics.

[0228] The execution parameters include trajectory control parameters and communication access parameters.

[0229] Specifically, the implementation steps of S203 include:

[0230] First, the state data and time series features are integrated into the state features of the aircraft group.

[0231] For example, the temporal features output by the LSTM model The movement history and trends of dynamic obstacles (drones) were captured, forming part of the perception of the dynamic environment. Complete state observation is essential for completing the navigation task. It also needs to include target information, which can be represented as:

[0232]

[0233] in, Indicates the target-related state. Indicates the communication capacity-related status. Indicates the AoI status of each UAV. This represents the system's energy consumption and energy efficiency status. For a unified representation, this invention uses the time-series hidden features output by the LSTM. With task decision state Together, they constitute the final decision input, i.e., the state features:

[0234]

[0235] Secondly, the state features are input into the first policy network to generate execution parameters. The first policy network is a multi-layer nonlinear transformation algorithm.

[0236] In this embodiment, the state features are input into the first policy network, which maps the state features into execution parameters through a multi-layer nonlinear transformation algorithm.

[0237] Among them, the trajectory control parameters are used to control the flight trajectory of the main aircraft, and the communication access parameters are used to adjust the communication access of the main aircraft to the UAV.

[0238] For example, the Actor-Critic reinforcement learning model can be used, which includes a deep neural network called the first policy network.

[0239] Specifically, the first strategy network , will state Mapped to execution parameters .

[0240] in, These are the policy parameters of the first policy network.

[0241] S204. Control the aircraft group according to the execution parameters.

[0242] The trajectory control parameters include the repulsive force parameters, tangential guidance intensity parameters, and tangential guidance direction parameters of the artificial potential field model. The artificial potential field model is used to simulate the trajectory operation of the aircraft group.

[0243] Specifically, the implementation steps of S204 include:

[0244] First, the repulsive force parameters, tangential guidance strength parameters, and tangential guidance direction parameters are input into the artificial potential field model to generate trajectory adjustment parameters.

[0245] In this embodiment, the repulsive force parameter, tangential guidance intensity parameter, and tangential guidance direction parameter are obtained from the execution parameters at the current moment and input into the artificial potential field model to generate the trajectory adjustment parameters at the current moment.

[0246] For example, in time slot t, the execution parameters are defined as follows:

[0247]

[0248] in, For the repulsive force parameters of the artificial potential field model, For the tangential guiding intensity parameters of the artificial potential field model, The tangential guidance direction parameter is the parameter of the artificial potential field model. The trajectory control parameter affects the next motion direction of eVTOL by adjusting the strength and direction of the repulsive term and the tangential guidance term in the artificial potential field model. These are communication access parameters used to represent the target UAV accessed by the current eVTOL in the current time slot.

[0249] The artificial potential field model abstracts the motion environment of each UAV into a virtual potential energy field. In this field, the target point is set as the position with the lowest global potential energy, generating a "gravitational field" that pulls the main aircraft closer to the target; while each UAV, acting as an obstacle, is given high potential energy, forming a "repulsive field" that pushes the main aircraft away. These two potential fields are superimposed in space to form a composite artificial potential field. The main aircraft only needs to move along the direction of the fastest decrease in potential energy in this field (i.e., the negative gradient direction) to naturally approach the target point while avoiding obstacles, and finally stop at the endpoint where the global potential energy is lowest.

[0250] Secondly, the flight trajectory of the main aircraft is adjusted according to the trajectory adjustment parameters.

[0251] Finally, adjust the communication access parameters of the main aircraft to the corresponding UAV.

[0252] In this embodiment, the drone access target at the current moment is determined according to the communication access parameters, and the connection between the host aircraft and the drone is adjusted. Each communication access adjustment can connect to at most one drone.

[0253] This embodiment provides a method for aircraft trajectory planning and communication resource allocation, including: acquiring state data and position data of an aircraft group; determining the temporal characteristics of the UAV group based on the position data of the aircraft group; determining the execution parameters of the main aircraft based on the state data and temporal characteristics, the execution parameters including trajectory control parameters and communication access parameters, wherein the trajectory control parameters are used to control the flight trajectory of the main aircraft, and the communication access parameters are used to adjust the communication access of the UAVs of the main aircraft; and controlling the aircraft group according to the execution parameters. Compared with the prior art, which lacks joint planning of trajectory planning and communication resources, resulting in resource waste, this application generates execution parameters including trajectory control parameters and communication access parameters through joint planning of point cloud spatiotemporal feature extraction and state data, realizing the collaborative optimization of flight trajectory and communication resource allocation.

[0254] In one possible embodiment, Figure 5 This is a schematic diagram of the aircraft group architecture in the scenario of sensor integration provided in the embodiments of this application, as shown below. Figure 5 As shown, the aircraft group architecture consists of a mission-performing eVTOL and a set of u dynamic unmanned aerial vehicles (where u={1,...,u,...,U}).

[0255] Before the mission begins, the aircraft group architecture is first established: the eVTOL mission start point, target end point, and the initial position and trajectory parameters of the dynamic UAVs are set. Preferably, the number of dynamic UAVs is 3.

[0256] Within the aircraft cluster architecture, the eVTOL is equipped with a millimeter-wave radar with sensing capabilities, enabling it to acquire multi-dimensional millimeter-wave radar point clouds of the drones in real time. These drones act as communication relays, providing services to ground users. Once the mission begins, the eVTOL continuously senses surrounding drones, establishes communication connections with them, and assists the drones in transmitting mission data to the base station.

[0257] In this embodiment, the entire aircraft architecture, under the constraints of real-time perception data, communication capacity, and obstacle avoidance requirements, jointly optimizes the trajectory and communication resource allocation of eVTOL, balancing communication service fairness and system energy efficiency.

[0258] Figure 6 A flowchart illustrating an aircraft trajectory planning and communication resource allocation method provided in this application embodiment. Figure 2 ,like Figure 6 As shown, after step S203 above, the following is also specifically included:

[0259] S601, Obtain status update data for the aircraft group.

[0260] The status update data includes update task data of the aircraft group after executing the execution parameters, update communication data between the main aircraft and the UAV group, and update energy consumption data of the aircraft group.

[0261] In this embodiment, after executing the execution parameters at the current time t, the state update data at time t+1 is obtained.

[0262] S602. Calculate the reward benefits of the main aircraft after executing the execution parameters based on the preset reward function, state data, and updated state data.

[0263] In this embodiment, the change value of the state data is obtained based on the state data and the updated state data. The change value is then substituted into a preset reward function to calculate the reward.

[0264] For example, the design of the reward function is crucial. Various constraints are incorporated into the reward function: trajectory completion, communication fairness, system energy efficiency, obstacle avoidance constraints, and AoI constraints. The specific reward function can be expressed as follows:

[0265]

[0266] Where R(t) is the reward for the main aircraft after executing the current execution parameters at time t; , , These represent the weighting coefficients for basic rewards, communication fairness, and system energy efficiency, respectively. , This represents the penalty for violating the safe distance and the penalty coefficient for not complying with AoI constraints.

[0267] in, Basic reward items; For fair rewards in communications; This is an energy efficiency award. This is a safety penalty item; This is an AoI constraint term.

[0268] Specifically, the basic reward is used to guide the low-altitude aircraft towards the target direction; the communication fairness reward is used to increase the minimum service capacity in multi-UAV scenarios; the energy efficiency reward is used to increase the communication benefits per unit of energy consumption; the safety penalty is used to constrain the low-altitude aircraft to maintain a safe distance from dynamic UAVs; and the AoI constraint is used to ensure the timeliness of information updates. Through the weighted combination of the above multiple rewards, the trajectory planning and resource allocation processes simultaneously take into account task completion, communication fairness, system energy efficiency, flight safety, and information freshness, achieving multi-objective collaborative optimization.

[0269] For example, obstacle avoidance constraints are achieved through an obstacle avoidance model, which is constructed using location data.

[0270] Specifically, to integrate discrete point cloud data (position data) into the trajectory optimization framework, a compact geometric representation needs to be constructed for each UAV target. Within a continuous time t, for each sensed UAV, based on its radar point cloud (position data), the following ellipsoidal model is constructed:

[0271] If the geometric center of the point cloud is considered as its centroid, then its centroid... It can be represented as:

[0272]

[0273] in, : Represents the position vector of the geometric center (centroid) of the point cloud of UAV u at time t. It is an important basis for subsequent operations such as building an ellipsoidal model. The centroid position obtained by calculation can represent the approximate center position of UAV in space, which can be used to further analyze the motion state of UAV, its relative positional relationship with other objects, etc. It is a normalization coefficient, where np represents the total number of points in the UAV's radar point cloud. This coefficient is used to average the positions of all points in the point cloud, ensuring that the calculation results reflect the overall center position of the point cloud, rather than being biased towards a localized area due to the number of points. This is the summation symbol, indicating a summation operation performed on np points in the UAV's radar point cloud. This summation process aggregates the positional information of all points in the point cloud for subsequent calculation of the average position. This represents the position vector of the i-th point in the UAV's radar point cloud. Each point has its coordinate information in three-dimensional space. By adding the position vectors of these points and averaging them, the geometric center of the entire point cloud can be obtained. It is the coordinate representation of the centroid position vector in three-dimensional space, representing the coordinate values ​​of the centroid in the x, y, and z coordinate axes at time t. The superscript T indicates that it is a column vector.

[0274] Furthermore, the covariance between each point in the point cloud and the centroid. It can be represented as:

[0275]

[0276] Furthermore, the covariance matrix M u Eigenvalue decomposition of (t) can be expressed as:

[0277]

[0278] Among them, Qu Λ(t) = [q1(t), q2(t), q3(t)] is the eigenvector matrix, defining the three principal axes of the ellipsoid. u (t)=diag(λ1(t),λ2(t),λ3(t)) is an eigenvalue diagonal matrix that satisfies λ1≥λ2≥λ3≥0.

[0279] Furthermore, the length of the ellipsoidal semi-axis can be obtained from the following formula:

[0280]

[0281] in, This corresponds to a 95% confidence level, ensuring that the ellipsoid contains most of the point cloud data. In the global coordinate system, a point on the ellipsoid surface is defined as p = [x, y, z]. T The implicit equation of the ellipsoid can be expressed as:

[0282]

[0283] in, , .

[0284] Furthermore, in the reward function, for the eVTOL position p e(t) Based on the drone's ellipsoidal parameters, collision risk is determined using the following inequality:

[0285]

[0286] in, This is a safety boundary factor used to compensate for perception errors and system latency. If this inequality holds, it indicates that the eVTOL has entered the drone's hazardous area, posing a collision risk.

[0287] S603: Periodically acquire multiple pieces of experience data.

[0288] The experience data includes: status data, execution parameters, updated status data, and reward benefits.

[0289] In this embodiment, the current state data, execution parameters, updated state data, and reward are taken as a set of experience data and stored in the experience cache. Then, the execution parameters for each time moment are retrieved and executed again to obtain the experience data corresponding to each time moment, and stored in the experience cache.

[0290] For example, the empirical data at the current moment is (s(t), a(t), r(t), s(t+1)).

[0291] S604. Select a preset number of empirical data points from multiple empirical data points as training samples.

[0292] In this embodiment, a preset number of experience data points are periodically selected from multiple experience data points in the experience cache as training samples for optimization processing.

[0293] S605. Optimize the first policy network using training samples.

[0294] In this embodiment, the policy parameter values ​​of the first policy network are optimized and trained multiple times using each piece of empirical data in the training samples.

[0295] Specifically, the implementation steps of S605 include:

[0296] First, obtain the updated state data from the empirical data in the training samples.

[0297] For example, the updated state data in the empirical data at the current moment is s(t+1).

[0298] Secondly, the updated status data is input into the second policy network to generate update execution parameters corresponding to each empirical data.

[0299] In this embodiment, each update status data is input into the second policy network, and the update execution parameters corresponding to each update status data are generated through the second policy network.

[0300] For example, input s(t+1) into the second policy network Generate the update execution parameter a corresponding to s(t+1). ' (t+1).

[0301] Among them, the policy parameter values ​​of the second policy network and the first policy network different.

[0302] Next, the update execution parameters are input into the first value network to generate the target value score for each empirical data.

[0303] The target value score is used to represent the future reward corresponding to the updated state data.

[0304] In this embodiment, each update execution parameter is input into the first value network, and the target value score of each empirical data is calculated through the first value network.

[0305] For example, s(t+1) and the predicted update execution parameter a ' (t+1) are input together into the first value network. Let the first value network evaluate the target value score of this "state-action pair". The specific calculation formula is as follows:

[0306]

[0307] in, For instant rewards, This is a discount factor used to balance the importance of current and future rewards.

[0308] Next, the empirical data from the training samples are input into the second value network to generate the predicted value score corresponding to each empirical data.

[0309] The predictive value score is used to represent the future return corresponding to the state data.

[0310] The value parameter values ​​of the second value network are different from those of the first value network.

[0311] In this embodiment, the state data and execution parameters from each experience data are input into the second value network, and the predicted value score corresponding to each experience data is generated through the second value network.

[0312] For example, through a second value network Assessment in status Next action The expected cumulative return – the predicted value score, of which These are the value parameters of the second value network. The goal of the second value network is to accurately estimate the expected cumulative return under the current strategy—the predicted value. The specific calculation formula is as follows:

[0313]

[0314] in, It is the expectation operator of the second value network, which represents the expected value of the subsequent cumulative return when making action choices according to this specific strategy; This represents the immediate reward obtained at time t+k. In reinforcement learning tasks, the immediate reward is a scalar value fed back by the environment based on the agent's state and actions at a specific time. It reflects the quality of the action at that time.

[0315] Then, based on each target value score and each predicted value score, the value parameter values ​​corresponding to the second value network are adjusted to obtain the optimized second value network.

[0316] In this embodiment, each target value score and each predicted value score are input into a loss function. The mean square error between each target value score and each predicted value score is calculated using the loss function to obtain the optimization parameters of the second value network. The optimized second value network is then obtained based on the optimization parameters. The specific calculation formula is as follows:

[0317]

[0318] in, This represents the loss function value of the second value network. In reinforcement learning, minimizing this loss function optimizes the parameters of the second value network, enabling it to more accurately estimate the value of state-action pairs. To calculate the mean square error between the target value score and the predicted value score, the second value network can continuously adjust its value parameters by minimizing this mean square error, so that the predicted value score is as close as possible to the target value score, thereby improving the accuracy of value estimation.

[0319] Finally, based on the policy gradient algorithm and the predicted value score, the policy parameter values ​​corresponding to the first policy network are adjusted to obtain the optimized first policy network.

[0320] In this embodiment, the policy gradient algorithm maximizes the predicted value score of the first policy network through gradient ascent, obtains the policy parameter value of the first policy network when the predicted value score is maximized, and optimizes the first policy network based on the policy parameter value.

[0321] In addition, the second policy network and the first value network are periodically updated slowly using a soft update algorithm, so that the policy parameter values ​​and value parameter values ​​of the second policy network and the first value network gradually approach the policy parameter values ​​and value parameter values ​​of the first policy network and the second value network.

[0322] Specifically, the calculation formula for the soft update algorithm is as follows:

[0323]

[0324] in, This is the soft update coefficient. and These represent the policy parameter values ​​and value parameter values ​​for the second policy network and the first value network, respectively. and These are the policy parameter values ​​and value parameter values ​​for the first policy network and the second value network, respectively.

[0325] In this embodiment, by acquiring the status update data of the aircraft group and designing a multi-weighted reward function that includes trajectory completion, communication fairness, system energy efficiency, obstacle avoidance constraints, and AoI constraints, combined with periodic experience data accumulation and a dual-network (policy network + value network) optimization architecture, a soft update algorithm is used to achieve progressive synchronization of policy parameters and value parameters. Finally, multi-objective collaborative optimization of task completion, communication fairness, system energy efficiency, flight safety, and information freshness is achieved in trajectory planning and resource allocation.

[0326] This application also provides a possible embodiment in which the reward function is constructed based on the system's joint optimization problem. The joint optimization problem can achieve an effective trade-off between communication resource fairness and energy efficiency in the system, specifically expressed as follows:

[0327]

[0328] in, Weighting factors representing fairness and energy efficiency in communication.

[0329] Furthermore, joint optimization problems also include various constraints, which need to be satisfied during the operation of the system.

[0330] Optionally, the first constraint constrains the flight boundary of the eVTOL:

[0331]

[0332] Optionally, the second constraint constrains the task boundary of the drone's motion:

[0333]

[0334] in, Let q represent the elliptical flight trajectory of the u-th UAV. u,1 ,q u,2 ,q u,3 These are the three principal axis vectors of the ellipse.

[0335] Optionally, the third constraint ensures that the distance between the eVTOL and the drone is greater than the safe distance d. min :

[0336]

[0337] Optionally, the fourth constraint constrains the velocity boundary:

[0338]

[0339] Optionally, the fifth constraint constrains the binary relationship between eVTOL and the drone:

[0340]

[0341] Optionally, the sixth constraint restricts eVTOL to accessing only one drone at a time within time t:

[0342]

[0343] Optionally, the seventh constraint ensures that, under the AoI constraint, each drone will connect to eVTOL at least once:

[0344]

[0345] Optionally, the eighth constraint determines the line-of-sight and probabilistic line-of-sight determination between the eVTOL and UAV channels:

[0346]

[0347] Optionally, the ninth constraint ensures that all communication data can be uploaded from the U-shaped drone to the base station:

[0348]

[0349] Optionally, the tenth constraint imposes a communication capacity constraint between each UAV and eVTOL.

[0350] This indicates the minimum communication capacity required for each drone to meet mission requirements:

[0351]

[0352] In one possible embodiment, Figure 7a A schematic diagram of a deep reinforcement learning architecture based on temporal point clouds provided in an embodiment of this application; Figure 7b A schematic diagram of a deep reinforcement learning process based on temporal point clouds is provided for an embodiment of this application, as shown below. Figure 7a and Figure 7b As shown:

[0353] Step 1: Establish a three-dimensional dynamic airspace scene, and set the eVTOL mission start point, target end point, and the initial position and trajectory parameters of the dynamic UAVs. Preferably, the number of dynamic UAVs is 3.

[0354] Step 2: Deploy a millimeter-wave radar on the eVTOL airframe and set the radar's duty cycle. At each time step, the millimeter-wave radar actively transmits signals and receives environmental echoes, generating a four-dimensional radar point cloud feature, which includes range, radial velocity, azimuth angle, and elevation angle.

[0355] Step 3: Organize and cache the radar point cloud data of the current and historical multiple time steps to form a point cloud input sequence of continuous time steps.

[0356] Step 4: Input the four-dimensional radar point cloud of each time step into the point cloud feature extraction network. After point-by-point feature mapping and max pooling operations, obtain the global spatial features of a single frame.

[0357] Step 5: Input the global spatial features of continuous time steps into the LSTM network to extract temporal features that reflect the dynamic changes of the environment.

[0358] Step 6: Combine the temporal features with the target location, communication capacity, AoI state, and energy consumption state to form a reinforcement learning state vector.

[0359] Step 7: Input the state vector into the policy network and output the joint action for the current time step. The joint action includes trajectory control parameters and UAV access selection parameters. Preferably, the action vector has a dimension of 4, where the first 3 dimensions are trajectory control parameters and the last dimension is access selection parameters.

[0360] Step 8: Update the eVTOL flight status based on the policy network output, determine the drone access target for the current time step, and calculate communication capacity, safe distance, system power consumption, and AoI metrics. At most one drone can access the network per time step.

[0361] Step 9: Calculate the instant reward based on trajectory advancement, communication fairness, system energy efficiency, safe distance constraints, and AoI constraints.

[0362] Step 10: Store the state transition samples in the experience replay buffer, randomly select a small batch of samples to update the policy network and value network, and perform a soft update on the target network.

[0363] Step 11: Repeat steps 2 to 10 until the model converges; deploy the converged strategy in online flight decision-making to achieve real-time trajectory planning and joint optimization of communication resources for eVTOL.

[0364] Figure 8 A schematic diagram of the structure of the aircraft trajectory planning and communication resource allocation device provided in the embodiments of this application is shown below. Figure 8 As shown, the device includes: an acquisition module 81, a first determination module 82, a second determination module 83, and a control module 84.

[0365] The acquisition module 81 is used to acquire the status data and position data of the aircraft group. The status data includes the mission data of the aircraft group, the communication data between the main aircraft and the UAV group, and the energy consumption data of the aircraft group.

[0366] The first determining module 82 is used to determine the temporal characteristics of the UAV group based on the position data of the aircraft group, wherein the temporal characteristics represent the motion trend of the UAV group over time.

[0367] The second determining module 83 is used to determine the execution parameters of the main aircraft based on the status data and timing characteristics. The execution parameters include trajectory control parameters and communication access parameters. The trajectory control parameters are used to control the flight trajectory of the main aircraft, and the communication access parameters are used to adjust the communication access of the main aircraft to the UAV.

[0368] Control module 84 is used to control the aircraft group according to the execution parameters.

[0369] In one possible design, the execution parameters of the main aircraft are determined based on state data and timing characteristics, including:

[0370] The second determining module 83 is also used to fuse the state data and time series features into the state features of the aircraft group;

[0371] The state features are input into the first policy network to generate execution parameters. The first policy network is a multi-layer nonlinear transformation algorithm.

[0372] In one possible design, the location data is point cloud data acquired at each time point;

[0373] Based on the position data of the aircraft group, determine the temporal characteristics of the UAV group, including:

[0374] The first determining module 82 is also used to extract features from the point cloud data at each time point to obtain the four-dimensional features of each UAV; wherein, the four-dimensional features include the distance features, radial velocity features, azimuth features and pitch features of the UAV relative to the host aircraft.

[0375] The four-dimensional features are aggregated into a global feature vector using the max pooling algorithm;

[0376] The global feature vector is input into multiple fully connected layers, and features are extracted layer by layer through activation functions in each of the multiple fully connected layers to obtain the final global feature vector.

[0377] The long short-term memory algorithm is used to extract features from the final global feature vector to obtain temporal features.

[0378] In one possible design, the trajectory control parameters include the repulsive force parameters, tangential guidance strength parameters, and tangential guidance direction parameters of the artificial potential field model, which is used to simulate the trajectory operation of the aircraft group.

[0379] Controlling the aircraft group according to execution parameters includes:

[0380] The control module 84 is also used to input the repulsive force parameters, tangential guidance intensity parameters and tangential guidance direction parameters into the artificial potential field model to generate trajectory adjustment parameters;

[0381] Adjust the flight trajectory of the main aircraft according to the trajectory adjustment parameters;

[0382] Adjust the communication access parameters of the main aircraft to the corresponding UAV.

[0383] In one possible design, after controlling the aircraft group according to the execution parameters, the following is also included:

[0384] Acquire the status update data of the aircraft group. The status update data includes the update task data of the aircraft group after executing the execution parameters, the update communication data between the main aircraft and the UAV group, and the update energy consumption data of the aircraft group.

[0385] Based on the preset reward function, state data, and updated state data, calculate the reward benefits of the main aircraft after executing the execution parameters;

[0386] Periodically acquire multiple pieces of experience data, including: status data, execution parameters, updated status data, and reward benefits;

[0387] A predetermined number of empirical data points are selected as training samples from multiple empirical data sets.

[0388] The first policy network is optimized using training samples.

[0389] In one possible design, the first policy network is optimized using training samples, including:

[0390] Obtain each updated state data from the empirical data in the training samples;

[0391] Each updated state data is input into the second policy network to generate update execution parameters corresponding to each experience data. The policy parameter values ​​of the second policy network are different from those of the first policy network.

[0392] Each update execution parameter is input into the first value network to generate a target value score for each piece of experience data. The target value score is used to represent the future return corresponding to the updated state data.

[0393] Each piece of empirical data from the training samples is input into the second value network to generate a predicted value score corresponding to each piece of empirical data. The predicted value score is used to represent the future return corresponding to the state data. The value parameter values ​​of the second value network are different from those of the first value network.

[0394] Based on each target value score and each predicted value score, the value parameter values ​​corresponding to the second value network are adjusted to obtain the optimized second value network.

[0395] Based on the policy gradient algorithm and the predicted value score, the policy parameter values ​​corresponding to the first policy network are adjusted to obtain the optimized first policy network.

[0396] This embodiment provides an aircraft trajectory planning and communication resource allocation device, which can execute an aircraft trajectory planning and communication resource allocation method of the above embodiment. Its implementation principle and technical effect are similar, and will not be described again here.

[0397] In the specific implementation of the aforementioned method for aircraft trajectory planning and communication resource allocation, each module can be implemented as a processor. The processor can execute computer execution instructions stored in the memory, thereby enabling the processor to execute the aforementioned method for aircraft trajectory planning and communication resource allocation.

[0398] Figure 9 This is a schematic diagram of the structure of an aircraft trajectory planning and communication resource allocation device provided in an embodiment of this application. Figure 9 As shown, the aircraft trajectory planning and communication resource allocation device 90 includes at least one processor 91 and a memory 92. The device also includes a communication component 93. The processor 91, memory 92, and communication component 93 are connected via a bus 94.

[0399] In the specific implementation process, at least one processor 91 executes computer execution instructions stored in memory 92, causing at least one processor 91 to execute a method in the field of aircraft trajectory planning as executed by the above-mentioned aircraft trajectory planning and communication resource allocation equipment.

[0400] The specific implementation process of processor 91 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0401] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0402] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage.

[0403] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0404] The above description of the functions implemented by the aircraft trajectory planning and communication resource allocation equipment and the main control equipment has introduced the solutions provided by the embodiments of the present invention. It is understood that, in order to achieve the above functions, the aircraft trajectory planning and communication resource allocation equipment or the main control equipment includes hardware structures and / or software modules corresponding to the execution of each function. By combining the units and algorithm steps of the various examples described in the embodiments of the present invention, the embodiments of the present invention can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of the embodiments of the present invention.

[0405] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the above-described method in the field of aircraft trajectory planning.

[0406] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0407] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in aircraft trajectory planning and communication resource allocation equipment or master control equipment.

[0408] This application also provides a computer program product, comprising: a computer program stored in a readable storage medium, wherein at least one processor of the aircraft trajectory planning and communication resource allocation device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the aircraft trajectory planning and communication resource allocation device to perform the scheme provided in any of the above embodiments.

[0409] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disk, or optical disk.

[0410] The technical solutions of this application have been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of this application is obviously not limited to these specific embodiments. The above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for aircraft trajectory planning and communication resource allocation, characterized in that, Applied to an aircraft group, the aircraft group comprising a main aircraft and a drone group, wherein the drone group includes at least one drone, the method includes: Acquire the status data and position data of the aircraft group, wherein the status data includes the mission data of the aircraft group, the communication data between the main aircraft and the UAV group, and the energy consumption data of the aircraft group; Based on the position data of the aircraft group, the temporal characteristics of the UAV group are determined, wherein the temporal characteristics represent the motion trend of the UAV group over time; Based on the state data and the timing characteristics, the execution parameters of the main aircraft are determined. The execution parameters include trajectory control parameters and communication access parameters. The trajectory control parameters are used to control the flight trajectory of the main aircraft, and the communication access parameters are used to adjust the communication access of the main aircraft to the UAV. The aircraft group is controlled according to the execution parameters.

2. The method according to claim 1, characterized in that, The step of determining the execution parameters of the main aircraft based on the state data and the timing characteristics includes: The state data and the time series features are fused together to form the state features of the aircraft group; The state features are input into a first policy network to generate the execution parameters. The first policy network is a multi-layer nonlinear transformation algorithm.

3. The method according to claim 2, characterized in that, The location data is point cloud data acquired at each time point; Determining the temporal characteristics of the UAV group based on the position data of the aircraft group includes: Feature extraction is performed on the point cloud data at each time point to obtain the four-dimensional features of each UAV; wherein, the four-dimensional features include the distance feature, radial velocity feature, azimuth feature, and pitch feature of the UAV relative to the host aircraft; The four-dimensional features are aggregated into a global feature vector using the max pooling algorithm; The global feature vector is input into multiple fully connected layers, and features are extracted layer by layer through activation functions in the multiple fully connected layers to obtain the final global feature vector. The temporal features are obtained by extracting features from the final global feature vector using the Long Short-Term Memory (LSTM) algorithm.

4. The method according to claim 1, characterized in that, The trajectory control parameters include the repulsive force parameters, tangential guidance intensity parameters, and tangential guidance direction parameters of the artificial potential field model. The artificial potential field model is used to simulate the trajectory operation of the spacecraft group. The step of controlling the aircraft group according to the execution parameters includes: The repulsive force parameter, the tangential guidance intensity parameter, and the tangential guidance direction parameter are input into the artificial potential field model to generate trajectory adjustment parameters; The flight trajectory of the main aircraft is adjusted according to the trajectory adjustment parameters; The communication access parameters are used to adjust the communication access of the UAV to the main aircraft.

5. The method according to claim 2, characterized in that, After controlling the aircraft group according to the execution parameters, the method further includes: The update status data of the aircraft group is obtained. The update status data includes the update task data of the aircraft group after executing the execution parameters, the update communication data between the main aircraft and the UAV group, and the update energy consumption data of the aircraft group. Based on the preset reward function, the state data, and the updated state data, calculate the reward benefit of the main aircraft after executing the execution parameters; Multiple pieces of experience data are periodically acquired, including: the status data, the execution parameters, the updated status data, and the reward income; A predetermined number of the aforementioned empirical data are selected as training samples from a plurality of the aforementioned empirical data; The first policy network is optimized using the training samples.

6. The method according to claim 5, characterized in that, Using the training samples, the first policy network is optimized, including: Each of the updated state data is obtained from each of the empirical data in the training samples; Each of the updated state data is input into the second policy network to generate update execution parameters corresponding to each of the experience data, wherein the policy parameter values ​​of the second policy network are different from those of the first policy network; Each of the update execution parameters is input into the first value network to generate a target value score for each of the experience data, wherein the target value score is used to represent the future return corresponding to the update state data; Each piece of experience data in the training samples is input into the second value network to generate a predicted value score corresponding to each piece of experience data. The predicted value score is used to represent the future return corresponding to the state data. The value parameter values ​​of the second value network are different from those of the first value network. Based on the target value scores and the predicted value scores, the value parameter values ​​corresponding to the second value network are adjusted to obtain an optimized second value network. Based on the policy gradient algorithm and the predicted value score, the policy parameter values ​​corresponding to the first policy network are adjusted to obtain an optimized first policy network.

7. A trajectory and communication planning device for an aircraft group, characterized in that, include: The acquisition module is used to acquire the status data and position data of the aircraft group. The status data includes the mission data of the aircraft group, the communication data between the main aircraft and the UAV group, and the energy consumption data of the aircraft group. The first determining module is used to determine the temporal characteristics of the UAV group based on the position data of the aircraft group, wherein the temporal characteristics represent the motion trend of the UAV group over time; The second determining module is used to determine the execution parameters of the main aircraft based on the state data and the timing characteristics. The execution parameters include trajectory control parameters and communication access parameters. The trajectory control parameters are used to control the flight trajectory of the main aircraft, and the communication access parameters are used to adjust the communication access of the main aircraft to the UAV. A control module is used to control the aircraft group according to the execution parameters.

8. A trajectory and communication planning device for an aircraft group, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.

10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-6.