Unmanned aerial vehicle take-off and landing point intelligent site selection method and system for low-altitude logistics
Through the combination of edge computing and reinforcement learning decision engines, multi-source data is collected and processed in real time, efficient location selection of drone take-off and landing points is achieved, and problems of long response time and poor adaptability in the existing technology are solved, and the efficiency and safety of low-altitude logistics are improved.
Patent Information
- Application Number
- CN202510724798.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The existing technology has a long response time in the selection of drone take-off and landing points, which cannot quickly adapt to dynamic logistics needs and environmental changes, resulting in low efficiency of drone logistics and difficult to ensure flight safety and efficiency.
Through edge computing terminals, multi-source data is collected in real time, five-dimensional real-time state vectors are built, combined with reinforcement learning decision engine and multi-objective reward function, action space is designed and lightweight training models are adopted to achieve efficient site selection of drone take-off and landing points.
It improves the real-time response capabilities of low-altitude logistics, enhances the adaptability of complex terrain, optimizes multi-objective decision-making, ensures the robustness and compliance of the system, and provides a reliable and intelligent management solution.
Smart Images

Figure CN120235367A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of logistics management, and specifically provides an intelligent siting method and system for UAV take-off and landing points for low-altitude logistics. Background Art
[0002] With the rapid development of the low-altitude logistics industry, UAVs, with their characteristics of high efficiency and flexibility, are increasingly widely used in fields such as urban express delivery and medical supplies transportation. However, there are still obvious shortcomings in the key link of siting UAV take-off and landing points.
[0003] When dealing with sudden demands in low-altitude logistics (such as a surge in medical supplies or an explosion of e-commerce promotion orders), it mostly relies on static clustering analysis methods of historical data (such as DBSCAN). This method leads to a serious lag in the adjustment of take-off and landing points, with a response time exceeding 4 hours, making it difficult to meet the demand characteristics of "small batches, high frequency, and immediacy" of UAV logistics. At the same time, the existing technology lacks effective processing of real-time dynamic data and cannot quickly perceive and adapt to dynamic factors such as changes in logistics demand, airspace status adjustment, and environmental condition fluctuations, making the siting results difficult to ensure UAV flight safety and logistics efficiency, and restricting the further development and large-scale application of the low-altitude logistics industry. Summary of the Invention
[0004] The purpose of the present invention is to provide an intelligent siting method and system for UAV take-off and landing points for low-altitude logistics to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: An intelligent siting method for UAV take-off and landing points for low-altitude logistics, including the following steps: Collect logistics order data, airspace status data, and environmental perception data in real time through an edge computing terminal, and perform normalization processing on the collected data; Construct a five-dimensional real-time state vector including a demand intensity factor, UAV cluster status, airspace compliance index, terrain resistance coefficient, and cost constraint vector based on the processed data; Construct a reinforcement learning decision engine, design an action space including node operations, task allocation, and resource allocation, and filter the action space based on airspace rules and physical constraints; Construct a multi-objective reward function including timeliness reward, cost reward, safety reward, and compliance reward, and achieve balanced optimization of multiple objectives through dynamic weight adjustment; Train a deep reinforcement learning model using the proximal policy optimization algorithm, and deploy the model in combination with edge computing and federated learning mechanisms; Generate a siting plan for UAV take-off and landing points through the reinforcement learning decision engine according to the real-time state vector; Output the site selection plan to the UAV scheduling system and the airspace management system, and optimize the model parameters through a closed-loop feedback mechanism.
[0006] Preferably, constructing a five-dimensional real-time state vector including a demand intensity factor, a UAV cluster state, an airspace compliance index, a terrain resistance coefficient, and a cost constraint vector based on the processed data, includes: Calculate the demand intensity factor based on exponential weighted moving average, and determine the emergency order weighting factor in combination with the order type; Calculate the UAV cluster state vector, including the remaining power of the normalized UAVs, the load rate of the UAV cluster, and the number of faulty UAVs; Calculate the airspace compliance index of the candidate area according to the dynamic no-fly zone data provided by the airspace management system; Calculate the terrain slope between the takeoff and landing points based on digital elevation data, and determine the terrain resistance coefficient in combination with the flight period; Obtain the cost constraint vector through Bayesian network training, including the normalized construction cost, maintenance cost, and noise complaint risk value.
[0007] Preferably, constructing the reinforcement learning decision engine, designing an action space including node operation, task assignment, and resource allocation, and filtering the action space based on airspace rules and physical constraints, includes: Define three types of atomic action combinations of node operation, task assignment, and resource allocation; Generate temporary takeoff and landing points in the compliant area based on the order heat map, and perturb the coordinates of the temporary takeoff and landing points; Use the Hungarian algorithm to solve the matching problem between tasks and takeoff and landing points, and the objective function includes flight cost and takeoff and landing point startup cost; Real-time shield invalid actions that violate airspace rules or physical constraints.
[0008] Preferably, constructing a multi-objective reward function including timeliness reward, cost reward, safety reward, and compliance reward, and achieving balanced optimization of multiple objectives through dynamic weight adjustment, includes: Construct the timeliness reward based on the decision time and order type; Construct the cost reward based on the flight cost, takeoff and landing point startup cost, and dynamic cost threshold; Construct the safety reward based on the obstacle density of the space system; Construct the compliance reward based on the approval result of the airspace management system; Adjust the weights of the multi-objective reward function according to the preset period and cluster load status.
[0009] Preferably, training the deep reinforcement learning model using the proximal policy optimization algorithm, and deploying the model in combination with edge computing and federated learning mechanisms, specifically including: Construct a neural network structure including an input layer, a hidden layer, and an output layer, where the hidden layer uses an activation function and sets an overfitting prevention mechanism; Perform online learning on the edge computing terminal and update the model parameters in combination with the federated learning mechanism; Participate in model training after performing privacy protection processing on the locally collected order data.
[0010] Preferably, it further includes the following steps: When detecting an abnormal change in the order volume, trigger an emergency response strategy to adjust the action space and the calculation logic of the reward function; When the model output is abnormal, generate a fallback site selection plan through the historical case matching mechanism.
[0011] Preferably, the edge computing terminal integrates a meteorological sensor and a high-precision positioning module for collecting wind speed, wind direction, and takeoff and landing point position data; The airspace management system provides dynamic no-fly zone data; The normalization processing of the data includes format conversion and standardization processing of different protocol data.
[0012] The present invention also provides an intelligent site selection system for drone takeoff and landing points for low-altitude logistics, including: A data collection module, which is used to collect logistics order data, airspace status data, and environmental perception data in real time through an edge computing terminal, and perform normalization processing on the collected data; A decision engine module, which is used to construct a five-dimensional real-time state vector including a demand intensity factor, a drone cluster state, an airspace compliance index, a terrain resistance coefficient, and a cost constraint vector; The decision engine module is also used to design an action space including node operations, task allocation, and resource allocation, and filter the action space based on airspace rules and physical constraints; The decision engine module is also used to construct a multi-objective reward function including timeliness reward, cost reward, safety reward, and compliance reward, and achieve balanced optimization of multiple objectives through dynamic weight adjustment; The decision engine module is also used to train a deep reinforcement learning model using the proximal policy optimization algorithm and deploy the model in combination with the edge computing and federated learning mechanisms; The decision engine module is also used to generate a site selection plan for drone takeoff and landing points according to the real-time state vector; An execution feedback module, which is used to output the site selection plan to the drone scheduling system and the airspace management system, and optimize the model parameters through a closed-loop feedback mechanism.
[0013] The present invention also provides an electronic device, which is a physical device, and the electronic device includes: A processor and a memory, the memory being communicatively connected to the processor; The memory is used to store at least one executable instruction executed by the processor, and the processor is used to execute the executable instruction to implement the intelligent site selection method for the take-off and landing points of drones for low-altitude logistics as described above.
[0014] The present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the intelligent site selection method for the take-off and landing points of drones for low-altitude logistics as described above is implemented.
[0015] Compared with the prior art, the beneficial effects of the present invention are: By constructing a dynamic data acquisition layer to obtain multi-source information in real time, and combining with a reinforcement learning decision engine for intelligent analysis, efficient site selection for drone take-off and landing points is achieved. By defining a state space including multi-dimensions such as demand intensity and drone cluster status, the complex scenarios of low-altitude logistics are accurately characterized; an efficient action space is designed and combined with an action mask mechanism to improve the accuracy and efficiency of decision-making; a multi-objective reward function is constructed and the weights are dynamically adjusted to balance multiple factors such as timeliness, cost, and safety; at the same time, a lightweight training and execution process is adopted to ensure the stable operation of the system in different scenarios; the present invention significantly improves the real-time response ability of low-altitude logistics, enhances the adaptability to complex terrains, optimizes multi-objective decision-making, and ensures the robustness and compliance of the system, providing a reliable solution for the intelligent management of low-altitude logistics. Description of the Drawings
[0016] Figure 1 It is the main flowchart of an intelligent site selection method for the take-off and landing points of drones for low-altitude logistics provided by an embodiment of the present invention; Figure 2 It is the structural schematic diagram of an intelligent site selection system for the take-off and landing points of drones for low-altitude logistics provided by an embodiment of the present invention; Figure 3 It is the structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] The execution entity of the method in this embodiment is a terminal, which can be a device such as a mobile phone, a tablet computer, a personal digital assistant (PDA) of a palm computer, a notebook, or a desktop computer. Of course, it can also be other devices with similar functions, and this embodiment does not impose any restrictions.
[0019] Please refer to Figure 1 , the present invention provides an intelligent siting method for drone takeoff and landing points for low-altitude logistics, including: Step 100: Real-time collect logistics order data, airspace status data, and environmental perception data through an edge computing terminal, and perform normalization processing on the collected data.
[0020] In this embodiment, the edge computing terminal integrates a meteorological sensor and a high-precision positioning module for collecting wind speed, wind direction, and takeoff and landing point position data; The airspace management system provides dynamic no-fly zone data; The normalization processing of the data includes format conversion and standardization processing of data with different protocols.
[0021] Specifically, in the low-altitude logistics system, high-performance edge computing terminals such as Huawei Atlas 500 are deployed. This terminal has a powerful real-time data collection ability and can comprehensively collect detailed logistics order information, including the order quantity, destination, weight, volume, etc.; accurately obtain airspace status data, such as the available airspace range at different altitude levels, flight restricted areas, the distribution of other aircraft within the current airspace, etc.; and through various sensors, accurately perceive the environment, covering meteorological data such as wind speed, wind direction, temperature, humidity, visibility, as well as geographical information data such as geographical terrain and building distribution; the multi-source data collected is transmitted and normalized through an efficient Apache Kafka data bus. With its excellent performance, Apache Kafka can ensure that the data throughput is stably maintained at ≥100,000 messages / second, ensuring the fast and orderly flow of data and providing a solid data foundation for subsequent data analysis and decision-making. As a data bus, Apache Kafka can achieve data normalization through its ecosystem components and extension solutions. For example, by integrating with Apache Flink or SparkStreaming, and using its richer transformation operators (such as window aggregation, regular expression parsing) to implement complex normalization logic.
[0022] Step 200: Based on the processed data, construct a five-dimensional real-time state vector S(t)=[S1, S2, S3, S4, S5] including a demand intensity factor, a drone cluster state, an airspace compliance index, a terrain resistance coefficient, and a cost constraint vector.
[0023] Specifically, the step 200 includes: Step 210, Demand Intensity Factor (S1): Calculate the demand intensity factor based on the exponentially weighted moving average , and determine the weighted factor for urgent orders in combination with the order type; Among them, is the order volume in the current 5 minutes, and this data is obtained in real time through the API interface provided by the logistics platform. To ensure the timeliness and accuracy of the data, the acquisition frequency is set to ≥100 times / second. For example, during the peak period of logistics business, the order data is updated frequently every second, enabling the system to grasp the latest order dynamics in real time; is the historical mean calculated based on the exponentially weighted moving average (EWMA), where the weighting coefficient takes the value of 0.3. This calculation method can highlight the influence weight of recent data more prominently. When analyzing the trend of order volume over time, the system can more sensitively capture the recent order volume fluctuations, rather than simply averaging the historical data; is the weighted factor for urgent orders, which is assigned differently according to the types of materials involved in the orders. For orders of medical supplies, due to their urgent needs related to life rescue, etc., the assignment is 1.5; for orders of fresh food cold chain, considering the requirements of commodity freshness and timeliness, the assignment is 1.2; for ordinary orders without special urgent time limit requirements, the assignment is 1.
[0024] Step 220, Drone Cluster Status (S2): Calculate the drone cluster status vector , including the normalized remaining battery power of the drones, the load rate of the drone cluster, and the number of faulty drones; Among them, is the normalized remaining battery power of the drones, is the load rate of the drone cluster, is the number of faulty drones; Battery normalization value is used to reflect the relative position of the current battery power of the drones within their battery power range, and its calculation formula is , where, represents the critical battery power of 20%. When the battery power of the drones drops to this value, immediate consideration should be given to landing for charging to ensure safe flight, corresponds to the fully charged state of 100%; Load ratio is used to measure the overall load of the drone cluster, and its calculation formula is , in the formula, represents the real-time load of the i th drone, is the maximum load of the i th drone,n represents the number of the entire UAV cluster. By accumulating and averaging the load ratios of each UAV, it can intuitively reflect whether the cluster load is approaching the saturation state; In addition, represents the number of faulty UAVs. This value is used to evaluate the health status of the UAV cluster. The number of faulty UAVs directly affects the stability and reliability of the cluster task execution.
[0025] Step 230, Airspace Compliance Index (S3): Calculate the airspace compliance index of the candidate area according to the dynamic no - fly zone data provided by the airspace management system ; Among them, represents the effective area of the candidate area, that is, the area that complies with specific airspace usage rules and can be safely used for UAV take - off, landing and other operations, is the total area of the candidate area, covering the entire range of the evaluated area, is the raster mask matrix obtained by UTM system conversion. This matrix is constructed based on various airspace restriction conditions (such as no - fly zones, height - restricted zones, etc.). The value at the corresponding airspace restriction position is 0, and the value at the position where the airspace can be normally used is 1, which is used to screen out the qualified areas, is the area identification matrix, which is used to clarify the target area to which each raster belongs. By multiplying and accumulating the corresponding elements of the two, the proportion of the effective area can be accurately obtained, quantifying the airspace compliance degree.
[0026] Step 240, Terrain Resistance Coefficient (S4): Calculate the terrain slope between the take - off and landing points based on the digital elevation data, and determine the terrain resistance coefficient in combination with the flight time period; Specifically, let the two take - off and landing points be and , and the formula for calculating the average slope is Among them, and respectively represent and 's altitude, is the distance between the two points on the horizontal plane; Considering the particularity of the night flight environment, compared with the daytime, night flight will increase a certain amount of additional resistance. Through actual tests and research and analysis, in this embodiment, when in the night flight state (i.e., = 1), an additional 10% resistance will be added to more accurately simulate and consider factors such as flight energy consumption. The corresponding resistance correction coefficient 's calculation formula is ; Through the above formula, the impact of terrain slope on the flight of the drone can be comprehensively considered, and reasonable corrections can be made for the additional resistance during night flight, thereby providing more accurate data support for the flight path planning and energy consumption estimation of the drone, ensuring that the correlation and differences between takeoff and landing points can be scientifically evaluated under different terrain and time conditions, and laying a solid foundation for subsequent decision-making.
[0027] Step 250, cost constraint vector (S5): The cost constraint vector is obtained through Bayesian network training, including the normalized construction cost, maintenance cost, and noise complaint risk value; During the training process, population density data is taken into consideration. This data reflects the density of personnel distribution in the area and is of great significance for the potential impact around the drone takeoff and landing points. At the same time, land rent data is input, which is related to the land cost investment required for the construction of the takeoff and landing points. After complex and accurate calculation and processing by the Bayesian network, the normalized construction cost, maintenance cost, and noise complaint risk value are finally output. These values are all limited within the range of [0,1] for subsequent comprehensive analysis and application.
[0028] Step 300, construct a reinforcement learning decision engine, design an action space including node operation, task allocation, and resource allocation, and filter the action space based on airspace rules and physical constraints.
[0029] Specifically, step 300 includes: Step 310, define three types of atomic action combinations of node operation, task allocation, and resource allocation, support 32 composite strategies, and filter invalid actions in real time through an action mask mechanism: Node operation covers the enabling and disabling of existing takeoff and landing points and the generation of new temporary takeoff and landing points; task allocation aims to reasonably allocate orders to each takeoff and landing point to achieve efficient distribution; resource allocation involves the reasonable arrangement of drone cluster resources such as power and load. These 3 types of atomic actions form 32 composite strategies through different combinations, which can flexibly meet the diverse needs of low-altitude logistics scenarios. The action mask mechanism is like an intelligent filter, which quickly screens out invalid actions that do not conform to the actual situation during the real-time decision-making process, greatly improving the decision-making efficiency and accuracy; Step 320, generate temporary takeoff and landing points in the compliance area based on the order heat map, and perform perturbation processing on the coordinates of the temporary takeoff and landing points; Among them, to more accurately reflect the actual demand distribution, based on the Delaunay triangulation algorithm, operations are carried out near the peak points where the temperature on the order heat map exceeds 40°C. Considering the rationality and flexibility of site selection, the range is limited to the area within a radius of 500m around the peak points, and this area must meet the compliance indicators This compliance area can ensure that the selected node has a good basis for real-world applications. A virtual node is generated within this compliance area, and its coordinates are calculated by the peak point coordinates. , that is, the virtual node coordinates are ,in and The values are randomly selected within the range of ±200m. The introduction of this random disturbance avoids the problem of site selection aggregation that may arise from being completely based on peak points, helps to cover the demand area more evenly, and improves the efficiency and rationality of subsequent logistics scheduling.
[0030] Step 330, the Hungarian algorithm is used to solve the matching problem between the mission and the take-off and landing point. The objective function includes the flight cost and the take-off and landing point startup cost. The objective function is: , in, Represents the distance from the mission starting point to the take-off and landing point j The comprehensive flight cost is composed of several key factors. The specific formula is: , It represents the straight-line flight distance between the mission starting point and the take-off and landing point on the horizontal plane, in kilometers. The coefficient 0.02 reflects the weight of the impact of flight distance on cost, that is, the cost increase caused by each kilometer of flight distance; It is the altitude difference between the mission start point and the take-off and landing point, in meters. The coefficient 0.1 reflects the effect of altitude difference on cost, which means the cost change associated with each meter of altitude difference. It represents the load rate of the drone when performing the mission, which is a proportional value between 0 and 1. The coefficient 0.05 indicates the impact of the load rate on the cost, that is, the cost change caused by a unit change in the load rate; It is used to measure the take-off and landing points. j The single startup cost covers the one-time investment costs such as startup energy consumption and equipment loss of the equipment related to the take-off and landing points. As a 0-1 decision variable, when the task i Assigned to take-off and landing points j hour, The value is 1, otherwise it is 0, so as to accurately control the matching relationship between the mission and the take-off and landing points.
[0031] Step 340, shielding invalid actions that violate airspace rules or physical constraints in real time.
[0032] It is understandable that, according to established airspace rules, such as no-fly zones, height-limited zones, etc., and physical constraints, including but not limited to terrain and landforms (obstacles such as mountains and water areas), obstacles (high-rise buildings, communication towers, etc.), the action planning of the drone is screened in real time. Once a rule-violating action is detected, it is immediately identified as an invalid action and blocked to ensure that the flight path of the drone is always legal and safe.
[0033] Step 400, construct a multi-objective reward function including timeliness reward, cost reward, safety reward, and compliance reward , and achieve balanced optimization of multiple objectives through dynamic weight adjustment.
[0034] Specifically, step 400 includes: Step 410, construct a timeliness reward based on the decision time and order type ; , wherein, represents the decision time, which is accurate to milliseconds and is the time interval from the decision start moment to the final determination of the plan, is the urgency identifier. When the task is determined to be urgent, takes the value of 1, otherwise 0. This formula means that the shorter the decision time and the more urgent the task, the higher the timeliness reward.
[0035] Step 420, construct a cost reward based on the flight cost, takeoff and landing point startup cost, and dynamic cost threshold ; , In the formula, represents the total cost during the flight of the drone, covering expenses directly related to flight such as fuel consumption and equipment wear and tear; is the setup cost of the takeoff and landing point, including one-time expenditures such as site rental and equipment deployment; is the dynamic cost threshold, which is calculated by the formula , where is the average value of past flight and setup costs, is the standard deviation. This formula reflects the relationship between the actual cost and the cost threshold, and the lower the cost, the higher the cost reward.
[0036] Step 430, construct a safety reward based on the obstacle density of the spatial voxel ; , This formula is calculated based on a 50m³ voxel, It represents the obstacle density within each 50 m³ voxel, that is, the proportion of the space occupied by obstacles within the voxel. By performing a product operation on all relevant voxels, the lower the obstacle density within the voxel, the higher the safety reward, thus measuring the safety level of the environment around the takeoff and landing points.
[0037] Step 440, construct a compliance reward based on the approval result of the airspace management system ; When the flight plan successfully passes the UTM (Unmanned Aircraft System Traffic Management) approval, the compliance reward takes a value of 1; if it fails to pass the approval, it takes a value of -0.5. The UTM approval mainly considers whether the flight plan complies with a series of regulatory requirements such as airspace management rules and safety standards.
[0038] Step 450, adjust the weights of the multi-objective reward function according to the preset time period and cluster load status; The system can automatically switch the weights according to different time periods (morning rush hour, off-peak, night) and load rates. For example, during the morning rush hour, due to the extremely high demand for logistics timeliness, set , at this time, the weight of the timeliness reward is the largest, highlighting the principle of timeliness priority. During the off-peak period, each factor is relatively balanced, and the weights may be adjusted to , , , , during the night period, considering flight safety and cost control, the weights will be different again, such as , , , , the load rate will also affect the weights. When the load rate is high, in order to ensure the smooth completion of the task, the weights of the cost reward and safety reward may be appropriately increased.
[0039] Step 500, train the deep reinforcement learning model using the proximal policy optimization algorithm and deploy the model in combination with edge computing and federated learning mechanisms.
[0040] Specifically, step 500 includes: Step 510, construct a neural network structure including an input layer, a hidden layer, and an output layer, and the hidden layer uses an activation function and sets an overfitting prevention mechanism; Using the PPO algorithm, the neural network includes a 512-dimensional input layer, two 256-neuron hidden layers (ReLU activation + dropout = 0.2), and a policy / value network output layer.
[0041] Step 520: Conduct online learning on the edge computing terminal and update the model parameters in combination with the federated learning mechanism. Conduct online learning on the edge computing terminal, set the computing power of the edge node ≥ 2 TOPS, and make the decision latency < 200 ms.
[0042] Among them, in the intelligent site selection process, this embodiment uses the Proximal Policy Optimization (PPO) algorithm to drive the decision-making process. The input layer of the constructed neural network is set to 512 dimensions, which can fully absorb various complex data information. It contains 2 hidden layers with 256 neurons each in the middle. Each layer uses the ReLU activation function to enhance the non-linear expression ability of the model, and is paired with a dropout ratio of 0.2 to effectively prevent overfitting. The end of the network is equipped with a policy / value network output layer, which is responsible for outputting the final decision result. To ensure the efficient operation of the neural network, clear requirements are put forward for the computing power of the edge node, which needs to reach a level of ≥ 2 TOPS to ensure quick response when processing large-scale data, and the decision latency is strictly controlled within < 200 ms to achieve near-real-time decision feedback.
[0043] Step 530: After performing privacy protection processing on the locally collected order data, participate in model training. The local node performs differential privacy processing (ε = 0.5) on the order data, and the central server aggregates the model through the FedAvg algorithm, triggering a parameter update every 500 decisions.
[0044] It can be understood that to ensure data privacy and security, this embodiment uses the federated learning mechanism. When the local node processes the order data, it introduces differential privacy technology and sets the privacy budget parameter ε to 0.5, minimizing the risk of data leakage while ensuring data availability. The central server then aggregates the model parameters uploaded by each local node through the FedAvg algorithm. After every 500 decisions are accumulated, a parameter update operation is triggered, prompting the model to continuously optimize within the entire federated system and improving the accuracy and generalization ability of the global model.
[0045] Step 540: When an abnormal change in the order volume is detected, trigger an emergency response strategy to adjust the action space and the calculation logic of the reward function. When the order volume change rate ΔQ > 300%, trigger a level-three emergency response, skip complex calculations, and increase the decision speed by another 50%. Step 550: When the model output is abnormal, generate a fallback site selection plan through the historical case matching mechanism.
[0046] When the model fails, retrieve the historical optimal solution through the cosine similarity Sim > 0.8 as the fallback action.
[0047] Specifically, when the order demand change rate ΔQ > 300%, it means that an extreme order peak occurs. At this time, the core area supply guarantee mechanism is immediately triggered. This mechanism will skip the regular complex calculation process and instead adopt a more efficient decision-making logic, enabling the decision-making speed to increase by 50% in this extreme situation and fully ensuring the material supply in the core area. In addition, to cope with the scenario of model failure, this embodiment also sets up safety guarantee measures. Through the cosine similarity algorithm, the decision-making schemes in the historical data are retrieved. When the cosine similarity Sim of the retrieval result > 0.8, it is determined that the historical scheme is highly similar to the current scenario, and the historical optimal solution is directly called to maintain the basic operation of the system and avoid business stagnation caused by model anomalies.
[0048] Step 600, according to the real-time state vector, generate a siting plan for the takeoff and landing points of the unmanned aerial vehicle through the reinforcement learning decision-making engine.
[0049] Step 700, output the siting plan to the unmanned aerial vehicle scheduling system and the airspace management system, and optimize the model parameters through a closed-loop feedback mechanism.
[0050] Output the siting plan to the unmanned aerial vehicle scheduling system and the airspace management system to ensure that the unmanned aerial vehicle scheduling system can efficiently plan the flight routes and task assignments of the unmanned aerial vehicles according to this plan, and the airspace management system can reasonably control the airspace resources for low-altitude logistics accordingly. At the same time, through the closed-loop feedback mechanism, continuously collect the actual operation data of the unmanned aerial vehicles and the system feedback information, and dynamically adjust and optimize the model parameters to improve the adaptability and accuracy of the model to different environments and task requirements.
[0051] In this embodiment, by constructing a dynamic data acquisition layer to obtain multi-source information in real time, combined with a reinforcement learning decision-making engine for intelligent analysis, the efficient siting of the takeoff and landing points of the unmanned aerial vehicle is realized. By defining a state space with multiple dimensions including demand intensity, unmanned aerial vehicle cluster status, etc., the complex scenario of low-altitude logistics is accurately characterized; designing an efficient action space and combining an action mask mechanism improves the accuracy and efficiency of decision-making; constructing a multi-objective reward function and dynamically adjusting the weights balance multiple factors such as timeliness, cost, and safety; at the same time, adopting a lightweight training and execution process ensures the stable operation of the system in different scenarios; the present invention significantly improves the real-time response ability of low-altitude logistics, enhances the adaptability to complex terrains, optimizes multi-objective decision-making, and ensures the robustness and compliance of the system, providing a reliable solution for the intelligent management of low-altitude logistics.
[0052] Based on the above embodiments, as Figure 2As shown in the figure, the present invention also provides an intelligent site selection system for UAV takeoff and landing points for low-altitude logistics, which is used to support the intelligent site selection method for UAV takeoff and landing points for low-altitude logistics in the above embodiments. The intelligent site selection system for UAV takeoff and landing points for low-altitude logistics includes: A data acquisition module 11, which is used to collect logistics order data, airspace status data, and environmental perception data in real time through an edge computing terminal, and perform normalization processing on the collected data; A decision engine module 12, which is used to construct a five-dimensional real-time state vector including a demand intensity factor, UAV cluster status, airspace compliance index, terrain resistance coefficient, and cost constraint vector; The decision engine module 12 is also used to design an action space including node operations, task allocation, and resource allocation, and filter the action space based on airspace rules and physical constraints; The decision engine module 12 is also used to construct a multi-objective reward function including timeliness reward, cost reward, safety reward, and compliance reward, and achieve balanced optimization of multiple objectives through dynamic weight adjustment; The decision engine module 12 is also used to train a deep reinforcement learning model using the proximal policy optimization algorithm, and deploy the model in combination with the edge computing and federated learning mechanisms; The decision engine module 12 is also used to generate a site selection plan for UAV takeoff and landing points according to the real-time state vector; An execution feedback module 13, which is used to output the site selection plan to the UAV scheduling system and the airspace management system, and optimize the model parameters through a closed-loop feedback mechanism.
[0053] In an optional embodiment, the decision engine module 12 is also used to: Calculate the demand intensity factor based on the exponentially weighted moving average, and determine the emergency order weighting factor in combination with the order type; calculate the UAV cluster state vector, including the normalized remaining battery power of the UAV, the UAV cluster load rate, and the number of faulty UAVs; calculate the airspace compliance index of the candidate area according to the dynamic no-fly zone data provided by the airspace management system; calculate the terrain slope between takeoff and landing points based on digital elevation data, and determine the terrain resistance coefficient in combination with the flight period; obtain the cost constraint vector through Bayesian network training, including the normalized construction cost, maintenance cost, and noise complaint risk value.
[0054] In an optional embodiment, the decision engine module 12 is also used to: Define three types of atomic action combinations: node operations, task allocation, and resource allocation; generate temporary takeoff and landing points in the compliant area based on the order heat map, and perform perturbation processing on the coordinates of the temporary takeoff and landing points; use the Hungarian algorithm to solve the matching problem between tasks and takeoff and landing points, and the objective function includes flight cost and takeoff and landing point startup cost; real-time shield invalid actions that violate airspace rules or physical constraints.
[0055] In an alternative embodiment, the decision engine module 12 is further configured to: Construct a timeliness reward based on the decision time and order type; construct a cost reward based on the flight cost, takeoff and landing point startup cost, and dynamic cost threshold; construct a safety reward based on the obstacle density of the spatial voxel; construct a compliance reward based on the approval result of the airspace management system; adjust the weights of the multi-objective reward function according to the preset time period and cluster load status.
[0056] In an alternative embodiment, the decision engine module 12 is further configured to: Construct a neural network structure including an input layer, a hidden layer, and an output layer, where the hidden layer uses an activation function and sets an overfitting prevention mechanism; perform online learning on the edge computing terminal, and update the model parameters in combination with the federated learning mechanism; participate in model training after performing privacy protection processing on the locally collected order data.
[0057] In an alternative embodiment, the decision engine module 12 is further configured to: When an abnormal change in the order volume is detected, trigger an emergency response policy to adjust the action space and the reward function calculation logic; when the model output is abnormal, generate a guaranteed site selection plan through a historical case matching mechanism.
[0058] In this embodiment, by constructing a dynamic data acquisition layer to obtain multi-source information in real time, and combining with a reinforcement learning decision engine for intelligent analysis, efficient site selection of the UAV takeoff and landing point is achieved. By defining a state space including multiple dimensions such as demand intensity and UAV cluster status, the complex scenario of low-altitude logistics is accurately characterized; an efficient action space is designed and combined with an action mask mechanism to improve the accuracy and efficiency of decision-making; a multi-objective reward function is constructed and the weights are dynamically adjusted to balance multiple factors such as timeliness, cost, and safety; at the same time, a lightweight training and execution process is adopted to ensure the stable operation of the system in different scenarios. The present invention achieves the effects of significantly improving the real-time response ability of low-altitude logistics, enhancing the adaptability to complex terrains, optimizing multi-objective decision-making, and ensuring the robustness and compliance of the system, providing a reliable solution for the intelligent management of low-altitude logistics.
[0059] Furthermore, the intelligent UAV takeoff and landing point site selection system for low-altitude logistics can run the above-mentioned intelligent UAV takeoff and landing point site selection method and system for low-altitude logistics. For specific implementation, reference can be made to the method embodiment, which will not be elaborated here.
[0060] Based on the above embodiments, as Figure 3 shown, the present invention further provides an electronic device, where the electronic device includes: At least one processor 22, at least one memory 21, a communication interface 23, and a communication bus 24, and the processor 22 is communicatively connected to the memory 21; In this embodiment, the memory 21 can be implemented in any suitable manner. For example, the memory 21 can be a read-only memory, a mechanical hard disk, a solid-state drive, a USB flash drive, etc.; the memory 21 is used to store at least one executable instruction executed by the processor. In this embodiment, the processor 22 can be implemented in any suitable manner. For example, the processor 22 can take the form of, for example, a microprocessor or a processor, and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc.; the processor is used to execute the executable instruction to implement the intelligent location selection method and system for the takeoff and landing points of low-altitude logistics drones as described above.
[0061] Based on the above embodiments, the present invention further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the intelligent location selection method and system for the takeoff and landing points of low-altitude logistics drones as described above.
[0062] Those of ordinary skill in the art can realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0063] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices, equipment, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0064] In several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or equipment can be in an electrical, mechanical, or other forms.
[0065] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0066] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0067] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage media include: U disk, mobile hard disk, read-only storage server, random access storage server, disk or optical disk, and other media that can store program instructions.
[0068] In addition, it should be noted that the combination of the various technical features in this case is not limited to the combination described in the claims of this case or the combination described in the specific embodiments. All technical features described in this case can be freely combined or combined in any way unless there is a contradiction between them.
[0069] It should be noted that the above examples are only specific embodiments of the present invention, and the present invention is obviously not limited to the above examples, and there are many similar variations. All variations directly derived or associated from the contents disclosed by the technicians in this field should fall within the protection scope of the present invention.
[0070] The above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An intelligent site selection method for UAV takeoff and landing points for low-altitude logistics, characterized in that Including the following steps: Collect logistics order data, airspace status data, and environmental perception data in real time through an edge computing terminal, and perform normalization processing on the collected data; Construct a five-dimensional real-time state vector based on the processed data, including a demand intensity factor, an unmanned aerial vehicle (UAV) cluster status, an airspace compliance index, a terrain resistance coefficient, and a cost constraint vector; Construct a reinforcement learning decision engine, design an action space including node operations, task allocation, and resource allocation, and filter the action space based on airspace rules and physical constraints; Construct a multi-objective reward function including timeliness rewards, cost rewards, safety rewards, and compliance rewards, and achieve balanced optimization of multiple objectives through dynamic weight adjustment; Use the proximal policy optimization algorithm to train a deep reinforcement learning model, and combine edge computing and federated learning mechanisms for model deployment; According to the real-time state vector, generate a site selection plan for UAV takeoff and landing points through the reinforcement learning decision engine; Output the site selection plan to the UAV scheduling system and the airspace management system, and optimize the model parameters through a closed-loop feedback mechanism.
2. The method according to claim 1, characterized in that The construction of the five-dimensional real-time state vector based on the processed data, including a demand intensity factor, an unmanned aerial vehicle (UAV) cluster status, an airspace compliance index, a terrain resistance coefficient, and a cost constraint vector, includes: Calculate the demand intensity factor based on exponential weighted moving average, and determine the emergency order weighting factor in combination with the order type; Calculate the UAV cluster status vector, including the normalized remaining battery power of UAVs, the UAV cluster load rate, and the number of faulty UAVs; Calculate the airspace compliance index of the candidate area according to the dynamic no-fly zone data provided by the airspace management system; Calculate the terrain slope between takeoff and landing points based on digital elevation data, and determine the terrain resistance coefficient in combination with the flight period; Obtain the cost constraint vector through Bayesian network training, including the normalized construction cost, maintenance cost, and noise complaint risk value.
3. The method according to claim 1, wherein The construction of the reinforcement learning decision engine, design an action space including node operations, task allocation, and resource allocation, and filter the action space based on airspace rules and physical constraints, includes: Define three types of atomic action combinations: node operations, task allocation, and resource allocation; Generate temporary takeoff and landing points in the compliance area based on the order heat map, and perform perturbation processing on the coordinates of the temporary takeoff and landing points; Use the Hungarian algorithm to solve the matching problem between tasks and takeoff and landing points, and the objective function includes flight cost and takeoff and landing point startup cost; Real-time block invalid actions that violate airspace rules or physical constraints.
4. The method according to claim 1, wherein The construction of the multi-objective reward function including timeliness rewards, cost rewards, safety rewards, and compliance rewards, and achieve balanced optimization of multiple objectives through dynamic weight adjustment, includes: Construct timeliness rewards based on decision-making time and order type; Construct cost rewards based on flight cost, takeoff and landing point startup cost, and dynamic cost threshold; Construct safety rewards based on the obstacle density of the space system; Construct compliance rewards based on the approval results of the airspace management system; Adjust the weights of the multi-objective reward function according to the preset period and cluster load status.
5. The method according to claim 1, characterized in that, The use of the proximal policy optimization algorithm to train a deep reinforcement learning model, and combine edge computing and federated learning mechanisms for model deployment, specifically includes: Construct a neural network structure including an input layer, a hidden layer, and an output layer, where the hidden layer adopts an activation function and sets an overfitting prevention mechanism; Perform online learning on the edge computing terminal and update the model parameters in combination with the federated learning mechanism; Participate in model training after performing privacy protection processing on the locally collected order data.
6. The method according to claim 1, wherein It also includes the following steps: When detecting an abnormal change in the order volume, trigger an emergency response strategy to adjust the action space and the calculation logic of the reward function; When the model output is abnormal, generate a guaranteed site selection plan through the historical case matching mechanism.
7. The method according to claim 1, wherein The edge computing terminal integrates a meteorological sensor and a high-precision positioning module for collecting wind speed, wind direction, and takeoff and landing point position data; The airspace management system provides dynamic no-fly zone data; The normalization processing of the data includes format conversion and standardization processing of data with different protocols.
8. An intelligent siting system for UAV takeoff and landing points for low-altitude logistics, characterized in that, It includes: A data collection module, which is used to collect logistics order data, airspace status data, and environmental perception data in real time through the edge computing terminal, and perform normalization processing on the collected data; A decision engine module, which is used to construct a five-dimensional real-time state vector including a demand intensity factor, an unmanned aerial vehicle (UAV) cluster status, an airspace compliance index, a terrain resistance coefficient, and a cost constraint vector; The decision engine module is also used to design an action space including node operations, task allocation, and resource allocation, and filter the action space based on airspace rules and physical constraints; The decision engine module is also used to construct a multi-objective reward function including timeliness reward, cost reward, safety reward, and compliance reward, and achieve balanced optimization of multiple objectives through dynamic weight adjustment; The decision engine module is also used to train a deep reinforcement learning model using the proximal policy optimization algorithm and deploy the model in combination with the edge computing and federated learning mechanisms; The decision engine module is also used to generate a site selection plan for the takeoff and landing points of UAVs according to the real-time state vector; An execution feedback module, which is used to output the site selection plan to the UAV scheduling system and the airspace management system, and optimize the model parameters through a closed-loop feedback mechanism.
9. An electronic device, characterized in that, The electronic device includes: A processor and a memory, and the memory is communicatively connected to the processor; The memory is used to store at least one executable instruction executed by the processor, and the processor is used to execute the executable instruction to implement the intelligent site selection method for UAV takeoff and landing points for low-altitude logistics according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, it implements the intelligent site selection method for UAV takeoff and landing points for low-altitude logistics according to any one of claims 1 to 7.
Citation Information
Patent Citations
Intelligent order management system and method
CN118505356A
Logistics unmanned aerial vehicle airport site selection method based on multi-source data driving
CN119026767A
Dynamic and static combined urban logistics unmanned aerial vehicle task planning method
CN119323323A
Unmanned aerial vehicle track and resource optimization method based on ISAC system and federated learning
CN119342440A
Logistics unmanned aerial vehicle take-off and landing point selection method, device, equipment, medium and product
CN119558584A
Cited By
Urban end logistics-oriented unmanned aerial vehicle take-off and landing site selection method
CN120525580A
Unmanned aerial vehicle resource scheduling method and system based on unmanned aerial vehicle information and multi-modal data
CN121212747A
Site selection method, system and equipment for vertical take-off and landing field in complex terrain area and medium
CN121352409A
Local node dynamic selection method for unmanned aerial vehicle low-altitude operation system
CN122176968A
Local node dynamic selection method for unmanned aerial vehicle low-altitude operation system
CN122176968B