Space-time adaptive matching system and method based on multi-target reinforcement learning
By adopting a space-time adaptive matching system with multi-objective reinforcement learning in the flight scheduling system, the existing system's lack of dynamic adaptability and neglect of multi-dimensional optimization goals is solved, and the adaptive scheduling and path planning of the drone in three-dimensional space is realized, and the system's adaptability and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202510166031.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-14
AI Technical Summary
Existing flight scheduling systems lack dynamic adaptability and are difficult to use real-time data for efficient scheduling. The path planning algorithms are mostly based on a single optimization goal, ignoring multi-dimensional resource efficiency and task priority.
Adopting a space-time adaptive matching system based on multi-objective reinforcement learning is adopted to realize the adaptive scheduling and path planning of drones in three-dimensional space through reinforcement learning and multi-objective optimization strategies. The system includes a data fusion cleaning module, an intelligent path planning module, a multi-objective scheduling optimization module, an adaptive task optimization module and a reinforcement learning drive module.
It realizes adaptive path adjustment and dynamic task priority adjustment of drones in complex dynamic environments, improves the system's adaptability and real-time response capabilities, and ensures efficient task execution and maximum resource utilization.
Smart Images

Figure CN120010515A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of aircraft path planning, specifically a spatiotemporal adaptive matching system based on multi-objective reinforcement learning. Background Art
[0002] The existing flight scheduling system is based on static mission planning, lacks dynamic adaptability, and is difficult to use real-time data for efficient scheduling. Existing path planning algorithms are mostly based on a single optimization goal, ignoring multi-dimensional resource efficiency and mission priority, and the A* algorithm cannot adapt to dynamic low-altitude environments and cannot cope with complex safety requirements and environmental changes. Summary of the invention
[0003] In view of the above-mentioned deficiencies in the prior art, the present invention proposes a spatiotemporal adaptive matching system based on multi-objective reinforcement learning. Through reinforcement learning and multi-objective optimization strategies, adaptive scheduling and path planning of UAVs in three-dimensional space are realized, intelligent aircraft mission management is provided, and operational safety in complex airspace is ensured. While meeting the multiple requirements of the low-altitude industry for safety and efficiency, it reduces dependence on static rules and expert experience, can dynamically adjust strategies according to real-time feedback and environmental changes, is highly adaptable, and meets the core technical requirements of future low-altitude equipment in operation services, supervision, and the safety standardization system of the entire industrial chain.
[0004] The present invention is achieved through the following technical solutions:
[0005] The present invention relates to a spatiotemporal adaptive matching system based on multi-objective reinforcement learning, comprising: a data fusion cleaning module, an intelligent path planning module, a multi-objective scheduling optimization module, an adaptive task optimization module and a reinforcement learning driving module, wherein: the data fusion cleaning module receives the three-dimensional coordinates, device status and task priority data of an aircraft from multi-source sensors, user equipment and a real-time monitoring system in the data receiving stage, and then uses a spatiotemporal convolutional network (STGCN) to clean and fuse the data, standardizes the data from different sources and removes abnormal data to obtain consistent spatiotemporal data, then extracts spatiotemporal features and outputs them to the three-dimensional space intelligent path planning module and the reinforcement learning driving module respectively; the intelligent path planning module performs three-dimensional space modeling of the flight area through a grid division technology, generates an initial path through an improved A* algorithm, and then generates a task priority data through a reinforcement learning driving module. Deep reinforcement learning is used to further optimize the path, so that the aircraft can avoid obstacles and adjust the path adaptively in a complex dynamic environment in real time. The multi-objective scheduling optimization module adopts the NSGA-II multi-objective optimization algorithm, which comprehensively considers flight time, task priority, and resource consumption to form a scheduling optimization plan, generates a global optimal solution set through Pareto frontier analysis, and continuously optimizes the scheduling efficiency in combination with deep reinforcement learning. The adaptive task optimization module monitors the aircraft status and task progress in real time during the task execution phase, and dynamically adjusts the task priority and scheduling strategy for path planning and resource allocation strategies through reinforcement learning technology based on feedback information. The reinforcement learning driving module receives the data and parameters fed back by the data fusion cleaning module and the adaptive task optimization module, and provides algorithm support for the reinforcement learning part in the three-dimensional space intelligent path planning module and the multi-objective scheduling optimization module.
[0006] The consistent spatiotemporal data is obtained by parsing the data packet content, converting and standardizing the data from different sources, automatically detecting and removing outliers using STGCN, and using the task model and resource model as labels to identify the unique data source, thereby ensuring that each time series data point can correspond to the exact source; based on the source information of the process instance, the data is grouped and sorted in timestamp order to obtain the time series data sequence of a single process instance.
[0007] The spatio-temporal convolutional network is implemented using the technology described in "Spatio-Temporal Graph Convolutional Networks: A Deep Learning Framework for Traffic Forecasting", IJCAI 2018) and performs spatio-temporal feature extraction through the following steps:
[0008] Step 1: Model the multi-source sensor data (such as aircraft coordinates and equipment status) as a graph structure, where nodes represent data sources and edges represent spatiotemporal correlations. Use the spatial graph convolution layer in the spatiotemporal convolutional network (GCN) to perform graph convolution operations to aggregate the features of adjacent nodes and capture the dependencies in the spatial dimension.
[0009] Step 2: Use the temporal convolution layer in the spatiotemporal convolutional network to perform dilated convolution to slide along the time axis to extract temporal dynamic features, such as the continuous changes in the trajectory of the aircraft or the temporal fluctuations of the task priority.
[0010] Step 3: Use the skip connection in the spatiotemporal convolutional network to fuse spatial and temporal features and enhance the model's ability to learn complex spatiotemporal patterns.
[0011] The improved A* algorithm specifically includes:
[0012] Step 1: Based on the three-dimensional space division grid model, the airspace is divided into three-dimensional grids with equal spacing, and the grid unit is defined as Grid (x, y, z), where x, y, and z are the three-dimensional coordinates of the center point of the grid. The grid represents the space and obstacle distribution of the aircraft in the airspace, and the path planning problem is transformed into a discretized path search problem in three-dimensional space;
[0013] The three-dimensional space division grid model includes: an airspace density perception unit, a grid resolution adjustment unit and a dynamic mapping unit, wherein: the airspace density perception unit calculates the complexity of the local airspace based on real-time sensor data (such as obstacle distribution, aircraft density, and meteorological conditions); the grid resolution adjustment unit dynamically adjusts the grid resolution according to the complexity of the local airspace; the dynamic mapping unit updates the obstacle status in real time through a spatiotemporal convolutional network (STGCN), and the predicted trajectory of the dynamic obstacle is modeled through a Kalman filter and fed back to the grid model.
[0014] The complexity of the local airspace Where: N obs is the number of obstacles, N uav is the number of aircraft, V celL is the current grid cell volume.
[0015] The dynamic adjustment means that a smaller grid (1m×1m×1m) is used in high-density areas (such as urban airspace) to improve obstacle avoidance accuracy; a larger grid (10m×10m×10m) is used in low-density areas (such as open airspace) to reduce the amount of calculation.
[0016] Step 2: Construct a path search goal to find a path P = {p1, p2, ..., p n}, so that the total cost function f(p i ) is the smallest, specifically: f(p i )=g(p i )+h(p i ), where: g(p i ) is the starting point to the current node p i The actual cost, h(p i ) is the heuristic estimated cost from the current node to the target node;
[0017] Step 3: Combine reinforcement learning to h(p i ) for dynamic optimization, specifically: h ′ (p i )=α·h(p i )+β·Q(p i ,a), where: α and β are adjustment coefficients, satisfying α+β=1, which can be adaptively adjusted according to the task priority to balance the impact of the task priority of the path node and the path cost. Q(p i ,a) is the Q value in reinforcement learning, indicating that at node p i The expected cost of the path after taking action a;
[0018] Step 4: During the path update process, Q-learning adjusts the path based on real-time feedback. Under the environmental feedback at time t, the Q value of each node n is updated as follows: Where: η is the learning rate, γ is the discount factor, and r is the i The immediate reward at the point, a is the current action, a′ is the next action, p i+1 To perform action a ′ The next node to be reached after
[0019] Step 5: By dynamically calculating the task priority and incorporating it into the weight adjustment of path planning, ensure that the aircraft can prioritize key tasks while taking into account the rationality of resource allocation and the efficiency of path planning. The dynamic adjustment of task priority is based on the analytic hierarchy process (AHP) and Bayesian optimization, and the priority score is determined by comprehensively evaluating the importance of the task, the current status of the aircraft (such as battery power, load capacity), and environmental changes (such as airspace congestion).
[0020] Priority of tasks described Where: u i is the priority score of the ith task, which is obtained by weighted summation of the following factors: T i Task urgency (such as task time limit), weighted by α T ; S i Aircraft status (such as remaining power, load capacity), weighted by αS ; E i Environmental complexity (such as high-density traffic areas), with a weight of α E , the expression of priority score is: i =α T T i +α S S i +α E e i . Where: T i +S i +E i =1, each weight can be dynamically adjusted through Bayesian optimization to adapt to changes in task scenarios. The calculated task priority w i It will be mapped to the weight of each node in the path planning process to guide the heuristic search of the A* algorithm and the reward distribution of reinforcement learning. Weight adjustment enables the system to balance multi-task requirements in real time, improve the flexibility of path planning and the overall efficiency of task completion.
[0021] The NSGA-II multi-objective optimization algorithm specifically includes:
[0022] Step A: In the multi-objective optimization process, the following objective function is used for optimization:
[0023] a. Minimize task completion time: The goal is to reduce the execution time of each task as much as possible, specifically: Where: T i represents the execution time of the i-th task, and n is the total number of tasks;
[0024] b. Maximizing resource utilization: Aims to improve the efficiency of resource utilization in the system, specifically: Where: R j is the available quantity of the jth resource, U j is the utilization rate of the jth resource, and m is the total number of resources.
[0025] c. Optimal path and energy consumption minimization: Considering the energy consumption optimization in drone path planning, the goal is to select the path with the lowest energy consumption, specifically: Where: E k represents the energy consumption of the kth path, and p is the number of all possible paths.
[0026] Step B: NSGA-II combines Pareto frontier solution and realizes multi-objective optimization. The Pareto frontier analysis includes:
[0027] a. Initial population generation: The system first generates an initial population of N individuals and randomly initializes the solution of each individual.
[0028] b. Non-dominated sorting: Perform non-dominated sorting on the individuals in the population and calculate the dominance of each individual. An individual A dominates another individual B if and only if: and This method is used to calculate the Pareto frontier solution set so that each individual is as close to the optimal solution as possible.
[0029] c. Crowding distance calculation: Calculate the crowding distance d for each individual i , to assess its relative distribution to other individuals, specifically: in: and They respectively represent the function values of adjacent individuals in the objective function f.
[0030] d. Selection and crossover mutation: Use tournament selection to select individuals with higher fitness and lower crowding from the parent generation for crossover and mutation to generate the next generation population.
[0031] Step C: By introducing deep reinforcement learning (DRL), the scheduling strategy can be adaptively adjusted to cope with complex and dynamic task requirements, including:
[0032] a. State representation: Use indicators such as task completion time, resource usage, and path energy consumption as the representation of the state space s.
[0033] b. Reward function: Define the reward function Where: w i Represents the weight of the i-th goal. The system continuously updates the weight through reinforcement learning to achieve a dynamic balance between different goals.
[0034] Step D: The final scheduling optimization solution is a diversified solution set P generated by the Pareto frontier solution technology, each of which contains the optimization results of task scheduling, resource allocation and path selection. For the actual application of system scheduling, the system can select the optimal solution from the Pareto solution set according to specific application requirements (such as the urgency of the task or the current usage of resources) to meet the current scheduling goals.
[0035] The reinforcement learning technology specifically includes:
[0036] Step i: Define the system state space S and action space A, and collect the current state P and system feedback data during each task progress. Select action a based on the state and feedback data. t , thereby updating the scheduling strategy;
[0037] Step ii: Update the Q value according to the Q-learning algorithm defined in the intelligent path planning module, and replace r with the value at node p. i The instantaneous reward at is specifically set to the value R(s) of the reward function R at time t t ,a t ), adjust the strategy according to the Q value of the current state and action in each feedback;
[0038] Step iii: Perform weighted summation of historical states and feedback data to update the strategy weights of task priority and path selection. Suppose the historical state set is {s t-n ,…,s t}, the corresponding feedback weight is {w t-n ,…,w t}, then the priority adjustment value of a task at the current time t Where: w i By normalizing the distance from the current moment, The closer the feedback data is to the current moment, the greater its weight;
[0039] Step iv: By updating the Q value, the system continuously optimizes the scheduling parameters so that it can achieve adaptive optimization in the face of environmental changes and adjustments to task requirements. Technical Effects
[0040] The present invention is based on the path planning and reinforcement learning collaborative optimization mechanism of the dynamic adaptive three-dimensional grid model: an adaptive modeling method for dynamically adjusting the three-dimensional space grid resolution is proposed, and the grid size (1m in high-density area) is adjusted in real time according to the airspace density (obstacle distribution, number of aircraft, meteorological conditions) 3 , low density area 10m 3 ), and based on this model, the heuristic function of the A* algorithm is improved. The Q value weights (h ′ (p i )=α·h(p i )+β·Q(p i ,a)) , deeply integrates the task priority (hierarchy analysis method AHP score) and the path cost to achieve real-time obstacle avoidance and path optimization in a three-dimensional dynamic environment. At the same time, by combining the NSGA-II multi-objective optimization algorithm with deep reinforcement learning (DRL), the dynamic weight adjustment Update the Pareto frontier solution set in real time. Use the feedback mechanism of reinforcement learning (the state space includes the position of the aircraft, the priority of the task, and the distribution of obstacles) to drive the objective function weight (w i) Adaptive optimization is used to generate a global optimal scheduling solution that adapts to environmental changes. The present invention applies STGCN to the cleaning of multi-source data of drones and the dynamic mapping of task priorities. The spatiotemporal dependency of multi-source data (coordinates, states, tasks) is extracted through spatiotemporal graph convolution (aggregating adjacent node features in spatial dimensions) and temporal hole convolution (capturing time series fluctuations), and the task priority weights (u are dynamically adjusted based on Bayesian optimization) are dynamically adjusted based on Bayesian optimization. i =α T T i +α S S i +α E E i ), realize the real-time linkage between task priority and path planning, and build an adaptive closed-loop optimization system with reinforcement learning as the core. By real-time monitoring of aircraft status (power, position, load), environmental changes (airspace congestion, obstacle movement) and task progress, the path planning and resource allocation strategies are dynamically adjusted. Introduce a weighted fusion mechanism for historical feedback data Combined with Q-learning online update strategy, self-learning and self-optimization of task scheduling are achieved.
[0041] Compared with the prior art, the present invention does not rely on fixed rules or static models, but optimizes the scheduling plan according to real-time feedback and task progress through an adaptive adjustment mechanism, significantly improving the adaptability and real-time response capabilities of the system. Through the multi-objective optimization algorithm, the system can process multiple scheduling goals at the same time and generate multiple optimization plans to ensure the high efficiency of task execution and the maximization of resource utilization. Through intelligent data processing and anomaly detection technology, the accuracy and reliability of scheduling decisions are further improved. Overall, the method of the present invention not only improves the efficiency and accuracy of task scheduling, but also has strong applicability and flexibility. It can be widely used in various complex low-altitude economic mission scenarios to meet changing needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is the main flow chart of the method of the present invention;
[0043] Figure 2 It is the main diagram topology module structure diagram of the system of the present invention;
[0044] Figure 3 A schematic diagram showing a time comparison of path planning in an embodiment of the present invention;
[0045] Figure 4 A schematic diagram of a three-dimensional distribution of a Pareto front solution set according to an embodiment of the present invention;
[0046] Figure 5 Schematic diagram of performance comparison between STGCN and Kalman filtering according to an embodiment of the present invention. DETAILED DESCRIPTION
[0047] like Figure 1 As shown, this embodiment involves a spatiotemporal adaptive matching system based on multi-objective reinforcement learning, including: a data fusion and cleaning module, an intelligent path planning module, a multi-objective scheduling optimization module, an adaptive task optimization module and a reinforcement learning driving module.
[0048] The data fusion and cleaning module includes: a multi-source data receiving unit, a spatiotemporal standardization unit, an anomaly detection and cleaning unit, and a multimodal fusion unit, wherein: the multi-source data receiving unit performs data collection and preliminary classification processing according to the input information of multi-source sensors, user equipment and real-time monitoring systems to obtain the original spatiotemporal data stream (including heterogeneous data such as aircraft three-dimensional coordinates, equipment status, and task priority), the spatiotemporal standardization unit performs format parsing and unified standardization processing (such as coordinate system conversion and timestamp alignment) according to the data format difference information, and obtains the intermediate data with consistent spatiotemporal dimensions, the anomaly detection and cleaning unit performs outlier detection and noise removal processing (such as outlier filtering and missing value interpolation) based on the spatiotemporal correlation characteristics of the spatiotemporal convolutional network (STGCN), and obtains the cleaned high-quality spatiotemporal data set, the multimodal fusion unit performs multi-source data fusion (including spatial graph convolution to aggregate adjacent node features and temporal hole convolution to extract temporal dependencies) according to the label identification information of the task model and the resource model, outputs the fused consistent spatiotemporal data, and distributes it to the intelligent path planning module and the reinforcement learning driving module.
[0049] The intelligent path planning module includes: a three-dimensional grid modeling unit, an initial path generation unit, a reinforcement learning optimization unit and a path dynamic adjustment unit, wherein: the three-dimensional grid modeling unit performs dynamic adaptive grid division processing (high-density area 1m 3 、Low-density area 10m 3 ), a three-dimensional space grid model is obtained, and the initial path generation unit uses the improved A* algorithm (fusion task priority weight h(p i )) performs heuristic search to generate an initial feasible path. The reinforcement learning optimization unit dynamically adjusts the Q value of the path node through Q-learning based on real-time environmental feedback (dynamic obstacle trajectory, aircraft status). The optimized path with real-time obstacle avoidance and optimal energy consumption is obtained. The path dynamic adjustment unit redistributes the path weight according to the changes in task priority (hierarchy analysis method AHP score) and resource allocation constraints, and outputs the final adaptively adjusted three-dimensional path plan.
[0050] The multi-objective scheduling optimization module includes: a multi-objective definition unit, a NSGA-II solution unit, a deep reinforcement learning collaboration unit and a solution selection unit, wherein: the multi-objective definition unit performs objective function modeling processing according to the optimization requirements of task completion time f1, resource utilization f2, and path energy consumption f3 to obtain a multi-objective optimization problem description, the NSGA-II solution unit performs Pareto frontier solution set generation processing according to the initial population (randomly generated solution set) and non-dominated sorting rules to obtain a global non-dominated solution set, and the deep reinforcement learning collaboration unit performs reward function generation processing according to the real-time state space S (aircraft position, task progress, environmental changes). Dynamically adjust the target weight w i Drive NSGA-II to iteratively optimize and obtain the dynamically updated Pareto optimal solution set. The solution selection unit screens the solution set according to the current task urgency and resource constraints through Bayesian optimization and hierarchical analysis method (AHP), and outputs the final scheduling solution (such as the shortest time solution or the lowest energy consumption solution).
[0051] The adaptive task optimization module includes: a real-time state monitoring unit, a priority dynamic evaluation unit, a strategy dynamic adjustment unit and a closed-loop feedback unit, wherein: the real-time state monitoring unit performs state acquisition and feature extraction processing according to the aircraft sensor data (power, position, load) and environmental monitoring information (obstacle movement, airspace congestion) to obtain a real-time state vector, and the priority dynamic evaluation unit adjusts the strategy according to the task urgency T i , Aircraft status S i 、Environmental complexity E i The weighted score (u i =α T T i +α S S i +α E E i ), dynamically calculate the task priority and obtain the real-time priority weight w i The strategy dynamic adjustment unit adjusts the priority weight and historical feedback data (V t =Σw i ·Q(s i ,a)), perform online optimization processing of path planning and resource allocation strategies, and output dynamically adjusted scheduling instructions. The closed-loop feedback unit evaluates the strategy effect and corrects parameters based on the task execution results and environmental change feedback, forming a closed-loop optimization mechanism to ensure continuous system adaptation.
[0052] The reinforcement learning driving module includes: a data integration unit, an algorithm configuration unit, a model training unit and a parameter real-time updating unit, wherein: the data integration unit performs state-action pair integration processing according to the spatiotemporal data output by the data fusion cleaning module and the feedback information of the adaptive task optimization module to obtain a reinforcement learning training data set; the algorithm configuration unit performs initialization configuration of Q-learning parameters (learning rate η, discount factor γ) and deep Q network (DQN) structure according to the task scenario requirements (such as obstacle avoidance sensitivity, energy consumption limit) to obtain a customized reinforcement learning model; the model training unit performs model training processing through experience replay and policy gradient descent according to historical data and real-time interaction data to obtain an optimized Q value function and policy network; the parameter real-time updating unit performs online policy update and model parameter synchronization processing according to dynamic changes in the environment (such as sudden obstacles, sudden changes in task priority), to ensure that the reinforcement learning algorithm of each module is always in the optimal state.
[0053] After specific practical experiments, Figure 2 In the application scenario shown in the figure, the learning rate η of the reinforcement learning driving module is set to 0.01, the discount factor γ is set to 0.9; the NSGA-II population size N is set to 200, the number of iterations is 100; the STGCN anomaly detection threshold is set to 3σ (standard deviation) and the experimental data obtained by running the above method are: the improved A* algorithm is effective in the dynamic grid model (high density area 1m 3 ) is 0.8 seconds for generating the initial path, which is 68% higher than the static A* algorithm (2.5 seconds); in the sudden obstacle scene, the reinforcement learning optimization unit shortens the path dynamic adjustment response time to 0.2 seconds, and the obstacle avoidance success rate reaches 95.6% (the traditional method is 78.4%); in the simulation of 72 hours of continuous operation, the system task completion rate reaches 99.2%, which is 11.7 percentage points higher than the existing static scheduling system (87.5%); the average resource utilization rate reaches 88.7% (the traditional system is 72.4%), and the redundant resource consumption is reduced by 23%. It includes: Web application layer, business processing layer and data layer.
[0054] The interaction between the user and the system at the Web application layer provides functions such as task ordering, status viewing, path display and result feedback. The user enters the basic information of the flight mission through the Web interface and views the system's path planning results and task execution progress in real time. A dynamic responsive user interface is built using HTML5, CSS3, and React.js to support the input of flight mission information (such as flight objectives, task priorities, resource requirements, etc.). Users can upload files (JSON) or directly enter task parameters through a form, and these inputs are submitted to the service layer through a RESTful API. In addition, by visually displaying the path and task allocation of the aircraft, users can clearly see the execution of the task. The path planning results, task execution status, resource utilization and other information of the flight mission will be visualized through Matplotlib and D3.js. The path information is displayed in a graphical way, showing the position, status and resource allocation of the aircraft in real time.
[0055] The business processing layer is the core of the method, which implements the core algorithms of data processing from input to output, task scheduling and path planning, and completes the adaptive matching and optimized scheduling of flight missions.
[0056] The multiple modules of this layer include: data fusion and cleaning module, intelligent path planning module, multi-objective scheduling optimization module, adaptive task optimization module and reinforcement learning driving module, among which: the data fusion and cleaning module receives data from multiple sources of the Web application layer, and uses PyTorch to implement the spatiotemporal convolutional network (STGCN), while convolving the features of the time series and space dimensions to extract spatiotemporal dependencies. De-noise the data from different sources, remove abnormal data, and then use NumPy to perform Z-Score normalization on the data, mapping the data to a standard normal distribution with a mean of 0 and a variance of 1. Output the cleaned and fused spatiotemporal data; the intelligent path planning module uses grid division to model the three-dimensional space, receives the spatiotemporal data containing the starting position and target position of the aircraft, obstacle information in the flight area, and the priority of the task after fusion and cleaning, and divides the three-dimensional space of the flight area into small units of the same size, and each grid represents a feasible area. Then use the heuristic function to evaluate the cost of the path, and implement the A* algorithm based on the SciPy library to find the path with the lowest cost from the starting point to the target in a known grid environment, and then generate a preliminary shortest path. Then, optimize based on the generated preliminary path by calling the reinforcement learning driver module. In each state, the aircraft selects actions according to the current environment and learns according to the reward function. During the flight, the aircraft updates its state through real-time sensor data (such as position, speed, battery status, etc.) and re-evaluates the path according to changes in the environment. The reinforcement learning module performs real-time optimization of the path based on the latest state information. The optimized path is adjusted through continuous state feedback, so that the aircraft can adapt to environmental changes in real time;
[0057] The multi-objective scheduling optimization module uses NSGA-II implemented based on the DEAP library as the core algorithm. The generation of the initial solution is constructed by receiving the preliminary path generated by the intelligent path planning module. The task scheduling plan is optimized by selecting, crossing and mutating each generation of population. Each optimization will gradually approach the Pareto frontier according to the goals of the aircraft's resource consumption, task completion time, priority, etc., and finally generate multiple scheduling plans that meet different needs; the adaptive task optimization module receives the aircraft's continuous feedback of status information, path progress, resource consumption and other real-time data, and feeds back the task priority, path planning and resource allocation strategy parameters to the multi-objective scheduling optimization module and the reinforcement learning driver module based on the task execution progress, aircraft status, environmental changes and other factors, so that it can continuously adjust the objective function and update the Q value function to achieve dynamic optimization and adjustment of the plan; the reinforcement learning driver module PyTorch implements the deep Q network (DQN), approximates the Q value function through the deep neural network (CNN), and defines the state space S including the current position, speed, task priority, distribution of environmental obstacles, etc. of the aircraft: S = {(x, y, z), ν x ,ν y ,ν z ,t,o}, where: (x,y,z) is the position of the aircraft, ν x ,v y ,v z is the velocity component of the aircraft in three-dimensional space, t is the remaining time of the task, and o is the position and state of the obstacle. The action space defines that each action of the aircraft at each moment in the grid environment can be left, right, up, down, forward, or backward: A = {left, right, up, down, forward, backward}. The reward function design takes into account aspects such as path length, obstacle avoidance, resource consumption, and task priority: R = -α·path_length+β·avoided_obstacle-γ·resource_consumption+δ·task_priority, where: α, β, γ, δ are weight coefficients that control the impact of each factor on the final reward. After the model is initialized, it interacts with the environment, takes actions and observes the results, and collects data on the state, action, reward, and next state. Based on the collected experience, the Q value is updated or the policy network is optimized to gradually approach the optimal strategy. It also collaborates and interacts with other modules to continuously learn and adapt to task requirements and environmental changes, making flight mission scheduling more intelligent and efficient.
[0058] The data layer stores and manages various types of data generated in the system, which not only ensures data persistence and efficient access, but also supports real-time data query, update and storage of historical data. According to the characteristics of different data, the data layer adopts different types of database storage solutions. InfluxDB is used as the storage system for spatiotemporal data streams. The spatiotemporal data streams are collected through the Web application layer or external sensor module and submitted to the back-end service through the RESTful API. The back-end service writes the data into InfluxDB through the InfluxDB client library; PostgreSQL is used as a relational database to store system status data and historical task records. The required data is obtained from the Web application through the RESTful API interface, and the back-end service stores the obtained structured data in the PostgreSQL database; MongoDB is used to store external system data. After the external system data (such as real-time airspace data) is transmitted to the system through the Open API interface, it is stored in MongoDB. Through the RESTful API, the service layer interacts with the MongoDB database, requests real-time airspace data, etc., to perform task scheduling and path optimization.
[0059] Table 1 Comparison of technical characteristics
[0060] Compared with the existing technology, this method has shown significant advantages in multiple key technical characteristics, especially in terms of reliability, efficiency, adaptability and effectiveness. The improvement of these technical characteristics is due to the comprehensive method adopted by the present invention, which not only introduces deep reinforcement learning and multi-objective optimization algorithms into the existing path planning and task scheduling, but also makes innovations in multiple dimensions such as data processing, model optimization and real-time decision-making.
[0061] The present invention greatly improves reliability. Existing task scheduling and path planning systems often rely on simple data cleaning and receiving mechanisms, which makes the system prone to data loss, inconsistency and delay when processing multi-source heterogeneous data, thereby affecting the accuracy of scheduling decisions and the overall reliability of the system. The present invention eliminates noise and outliers in multi-source data during the data fusion and cleaning stage by introducing the spatiotemporal convolutional network (STGCN), ensuring the consistency and accuracy of the data in time and space dimensions. Figure 5As shown in the figure, the anomaly detection accuracy of STGCN is 98.5%, while that of the traditional method (Kalman filter) is 89.2%; the false alarm rate of STGCN is 1.3%, while that of the traditional method is 8.7%. STGCN can efficiently process large-scale spatiotemporal data, minimize the impact of data loss and abnormal interference, provide reliable input data for subsequent path planning and task scheduling, and ensure the high reliability of the entire system;
[0062] The present invention has shown a significant improvement in efficiency. Existing path planning methods, especially the A* algorithm, are usually only applicable to two-dimensional static environments, and the path planning efficiency is greatly reduced when encountering dynamic obstacles or complex environments. The present invention combines the A* algorithm with reinforcement learning, and optimizes the path planning results in real time through deep reinforcement learning, so that the system can quickly calculate the optimal path in a complex and dynamic environment. This method has significantly improved the calculation speed and path optimization efficiency compared to existing algorithms. Figure 3 As shown in the figure, when the high-density area uses 1m 3 When the grid is used, the obstacle avoidance accuracy is improved by 40%, while the overall computational effort only increases by 15% (the traditional fixed grid computational effort increases by 50%).
[0063] In addition, the reinforcement learning mechanism in three-dimensional grid modeling and path planning works together to improve the computational efficiency of the system, ensuring that the aircraft can respond quickly and make optimal decisions in a changing environment, with strong real-time performance; adaptability is another key feature of the present invention. Existing path planning and task scheduling systems usually rely on preset rules or manual adjustment strategies, and lack the ability to respond to environmental changes and task requirements in real time. In contrast, the present invention introduces an adaptive reinforcement learning mechanism, which enables the system to dynamically adjust task priorities, resource allocation, and path planning strategies based on environmental changes, aircraft status, and task progress. The system can not only adjust the task scheduling of the aircraft in real time, but also flexibly and adaptively adjust the optimal solution according to changes in task requirements and fluctuations in the external environment. This dynamic adjustment capability significantly enhances the system's adaptability in dealing with emergencies and complex tasks, allowing the aircraft to better respond to different scenarios and task requirements, ensuring the smooth completion of the task;
[0064] In terms of effectiveness, the present invention uses the NSGA-II multi-objective optimization algorithm and the Pareto frontier solution to effectively solve the problem that existing optimization methods are difficult to balance multiple objectives. Existing multi-objective optimization algorithms often only provide a single scheduling solution, which is difficult to meet the multiple optimization requirements of time, resources and task priorities in complex tasks. By introducing NSGA-II, the present invention can not only handle multiple optimization objectives, but also generate multiple feasible optimal solutions through the Pareto frontier, providing users with more choices. Figure 4As shown in the figure, the present invention generates 15 non-dominated solutions, while the traditional method NSGA-II only generates 8; and the task completion time f1 is shortened from 20 minutes to 15 minutes, and the resource utilization f2 is increased from 75% to 92%. This multi-scheme selection and flexibility enables the system to dynamically generate qualified scheduling schemes according to the different requirements of the tasks, ensuring that the scheduling efficiency and resource utilization of the system in various application scenarios are maximized, fully meeting the needs of users.
[0065] In summary, the present invention improves the performance of the system in terms of reliability, efficiency, adaptability and effectiveness by innovatively combining reinforcement learning, spatiotemporal convolutional networks and multi-objective optimization algorithms. Unlike the solutions in the prior art that rely on preset rules and simple algorithms, the system of the present invention can perform real-time optimization and adaptive scheduling in a dynamic and complex environment, providing a more flexible and efficient task scheduling and path planning solution. Through the improvement of these technical characteristics, the present invention can provide higher quality intelligent services in multiple fields such as aircraft task scheduling, path planning and resource allocation, meeting the needs of different tasks and environments.
[0066] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principle and purpose of the present invention. The protection scope of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. Each implementation scheme within its scope shall be subject to the constraints of the present invention.
Claims
1. A spatiotemporal adaptive matching system based on multi-objective reinforcement learning, characterized in that: include: The data fusion cleaning module, the intelligent path planning module, the multi-objective scheduling optimization module, the adaptive task optimization module and the reinforcement learning driving module are as follows: the data fusion cleaning module receives the three-dimensional coordinates, equipment status and task priority data of the aircraft from multi-source sensors, user equipment and real-time monitoring systems during the data receiving stage, and then uses the spatiotemporal convolutional network (STGCN) to clean and fuse the data, standardizes the data from different sources and removes abnormal data to obtain consistent spatiotemporal data, then extracts spatiotemporal features and outputs them to the three-dimensional space intelligent path planning module and the reinforcement learning driving module respectively; the intelligent path planning module performs three-dimensional space modeling of the flight area through grid division technology, generates the initial path through the improved A* algorithm, and then further optimizes the path through reinforcement learning to make the aircraft It can avoid obstacles and adaptively adjust paths in real time in complex dynamic environments; the multi-objective scheduling optimization module adopts the NSGA-II multi-objective optimization algorithm, comprehensively considers flight time, task priority, and resource consumption to form a scheduling optimization plan, generates a global optimal solution set through Pareto frontier analysis, and continuously optimizes scheduling efficiency in combination with deep reinforcement learning; the adaptive task optimization module monitors the aircraft status and task progress in real time during the task execution phase, and dynamically adjusts the task priority and scheduling strategy for path planning and resource allocation strategies through reinforcement learning technology based on feedback information; the reinforcement learning drive module receives data and parameters fed back by the data fusion cleaning module and the adaptive task optimization module, and provides algorithm support for the reinforcement learning part in the three-dimensional space intelligent path planning module and the multi-objective scheduling optimization module.
2. The spatiotemporal adaptive matching system based on multi-objective reinforcement learning according to claim 1, characterized in that: The consistent spatiotemporal data is analyzed by parsing the data packet content, converting and standardizing the data from different sources, using STGCN to automatically detect and remove outliers, and using the task model and resource model as labels to identify the unique data source, thereby ensuring that each time series data point can correspond to the correct source; For the source information based on the process instance, the data is grouped and sorted in timestamp order to obtain the time series data sequence of a single process instance; The spatiotemporal convolutional network extracts spatiotemporal features through the following steps: Step 1: Model the multi-source sensor data as a graph structure, where nodes represent data sources and edges represent spatiotemporal correlations. Use the spatial graph convolution layer in the spatiotemporal convolutional network (GCN) to perform graph convolution operations to aggregate the features of adjacent nodes and capture the dependencies in the spatial dimension. Step 2: Use the temporal convolution layer in the spatiotemporal convolutional network to perform dilated convolution along the time axis to extract temporal dynamic features, such as the continuous changes in the trajectory of the aircraft or the temporal fluctuations of the task priority. Step 3: Use the skip connection in the spatiotemporal convolutional network to fuse spatial and temporal features and enhance the model's ability to learn complex spatiotemporal patterns.
3. The spatiotemporal adaptive matching system based on multi-objective reinforcement learning according to claim 1, characterized in that: The improved A* algorithm specifically includes: Step 1: Based on the three-dimensional space division grid model, the airspace is divided into three-dimensional grids with equal spacing, and the grid unit is defined as Grid (x, y, z), where x, y, and z are the three-dimensional coordinates of the center point of the grid, respectively. The grid represents the space and obstacle distribution of the aircraft in the airspace, and the path planning problem is converted into a discretized path search problem in three-dimensional space. Step 2: Construct a path. The search goal is to find a path P = {p1, p2, .., p n }, so that the total cost function f(p i ) is the smallest, specifically: f(p i )=g(p i )+h(p i ), where: g(p i ) is the starting point to the current node p i The actual cost, h(p i ) is the heuristic estimated cost from the current node to the target node; Step 3: Combine reinforcement learning to h(p i ) is dynamically optimized, specifically: h′(p i )=α·h(p i )+β·Q(p i , a), where: α and β are adjustment coefficients, satisfying α+β=1, which can be adaptively adjusted according to the task priority to balance the impact of the task priority of the path node and the path cost, Q(p i , a) is the Q value in reinforcement learning, indicating that at node p i The expected cost of the path after taking action a; Step 4: During the path update process, Q-learning adjusts the path based on real-time feedback and updates the Q value of each node n under the environmental feedback at time t, specifically: Where: η is the learning rate, γ is the discount factor, r is the immediate reward at node pi, a is the current action, a′ is the next action, P i+1 is the next node reached after executing action a′; Step 5: By dynamically calculating the task priority and incorporating it into the weight adjustment of path planning, we ensure that the aircraft can prioritize the execution of key tasks while taking into account the rationality of resource allocation and the efficiency of path planning. The dynamic adjustment of task priority is based on the analytic hierarchy process (AHP) and Bayesian optimization. The priority score is determined by comprehensively evaluating the importance of the task, the current status of the aircraft (such as battery power, load capacity), and environmental changes (such as airspace congestion).
4. The spatiotemporal adaptive matching system based on multi-objective reinforcement learning according to claim 1, characterized in that: The NSGA-II multi-objective optimization algorithm specifically includes: Step A: In the multi-objective optimization process, the following objective function is used for optimization: a1. Minimize task completion time: The goal is to reduce the execution time of each task as much as possible, specifically: Where: T i represents the execution time of the i-th task, and n is the total number of tasks; a2. Maximizing resource utilization: Aims to improve the efficiency of resource utilization in the system, specifically: Where: R j is the available quantity of the jth resource, U j is the utilization rate of the jth resource, and m is the total number of resources; a3. Optimal path and energy consumption minimization: Considering the energy consumption optimization in UAV path planning, the goal is to select the path with the lowest energy consumption, specifically: Where: E k represents the energy consumption of the kth path, and p is the number of all possible paths; Step B: NSG-II combines Pareto frontier solution and realizes multi-objective optimization. The Pareto frontier analysis includes: b1. Initial population generation: The system first generates an initial population of N individuals and randomly initializes the solution of each individual; b2. Non-dominated sorting: Perform non-dominated sorting on the individuals in the population and calculate the dominance of each individual. An individual A dominates another individual B if and only if: f i (A)≤f i (B) and f j (A)<fj(B), this method is used to calculate the Pareto frontier solution set so that each individual is as close to the optimal solution as possible; b3. Crowding distance calculation: Calculate the crowding distance d for each individual i , to assess its relative distribution to other individuals, specifically: in: and Respectively represent the function values of adjacent individuals in the objective function f; b4. Selection and crossover mutation: Use tournament selection to select individuals with higher fitness and lower crowding from the parent generation for crossover and mutation to generate the next generation population; Step c: By introducing deep reinforcement learning (DRL), the scheduling strategy can be adaptively adjusted to cope with complex and dynamic task requirements, including: c1. State representation: Use indicators such as task completion time, resource usage, and path energy consumption as the representation of the state space s; c2. Reward function: Define the reward function Where: w i represents the weight of the i-th goal. The system continuously updates the weight through reinforcement learning to achieve a dynamic balance between different goals. Step D: The final scheduling optimization solution is a diversified solution set P generated by the Pareto frontier solving technology. Each solution contains the optimization results of task scheduling, resource allocation and path selection. For the actual application of system scheduling, the system can select the optimal solution from the Pareto solution set according to specific application requirements to meet the current scheduling goals.
5. The spatiotemporal adaptive matching system based on multi-objective reinforcement learning according to claim 1, characterized in that: The reinforcement learning technology specifically includes: Step i: Define the system state space S and action space A, and collect the current state P and system feedback data during each task progress. Select action a based on the state and feedback data. t , thereby updating the scheduling strategy; Step ii: Update the Q value according to the Q-learning algorithm defined in the intelligent path planning module, and replace r with the value at node p. i The instantaneous reward at is specifically set to the value R(s) of the reward function R at time t t ,a t ), adjust the strategy according to the Q value of the current state and action in each feedback; Step iii: Perform weighted summation of historical states and feedback data to update the strategy weights of task priority and path selection. Suppose the historical state set is {s t-n , ..., s t }, the corresponding feedback weight is {w t-n , ..., w t }, then the priority adjustment value of a task at the current time t Where: w i By normalizing the distance from the current moment, The closer the feedback data is to the current moment, the greater its weight; Step iv: By updating the Q value, the system continuously optimizes the scheduling parameters so that it can achieve adaptive optimization in the face of environmental changes and adjustments to task requirements.
6. The spatiotemporal adaptive matching system based on multi-objective reinforcement learning according to claim 1, characterized in that: The data fusion and cleaning module includes: a multi-source data receiving unit, a spatiotemporal standardization unit, an anomaly detection and cleaning unit, and a multimodal fusion unit, wherein: the multi-source data receiving unit performs data collection and preliminary classification processing according to the input information of multi-source sensors, user equipment and real-time monitoring systems to obtain the original spatiotemporal data stream; the spatiotemporal standardization unit performs format analysis and unified standardization processing according to the data format difference information to obtain intermediate data with consistent spatiotemporal dimensions; the anomaly detection and cleaning unit performs outlier detection and noise removal processing based on the spatiotemporal correlation characteristics of the spatiotemporal convolutional network (STGCN) to obtain a cleaned high-quality spatiotemporal data set; the multimodal fusion unit performs multi-source data fusion according to the label identification information of the task model and the resource model, outputs the fused consistent spatiotemporal data, and distributes it to the intelligent path planning module and the reinforcement learning driving module.
7. The spatiotemporal adaptive matching system based on multi-objective reinforcement learning according to claim 1, characterized in that: The intelligent path planning module includes: a three-dimensional grid modeling unit, an initial path generation unit, a reinforcement learning optimization unit and a path dynamic adjustment unit, wherein: the three-dimensional grid modeling unit performs dynamic adaptive grid division processing according to the airspace density index and task requirement information to obtain a three-dimensional space rasterization model; the initial path generation unit performs heuristic search through an improved A* algorithm according to the starting point, target point and grid model information to generate an initial feasible path; the reinforcement learning optimization unit dynamically adjusts the Q value of the path node through Q-learning according to real-time environmental feedback to obtain an optimized path with real-time obstacle avoidance and optimal energy consumption; the path dynamic adjustment unit performs path weight redistribution processing according to task priority changes and resource allocation constraints, and outputs the final adaptively adjusted three-dimensional path plan.
8. The spatiotemporal adaptive matching system based on multi-objective reinforcement learning according to claim 1, characterized in that: The multi-objective scheduling optimization module includes: a multi-objective definition unit, an NSGA-II solution unit, a deep reinforcement learning collaboration unit and a solution selection unit, wherein: the multi-objective definition unit performs objective function modeling processing according to the optimization requirements of task completion time f1, resource utilization f2, and path energy consumption f3 to obtain a multi-objective optimization problem description; the NSGA-II solution unit performs Pareto frontier solution set generation processing according to the initial population and non-dominated sorting rules to obtain a global non-dominated solution set; the deep reinforcement learning collaboration unit performs a reward function based on the real-time state space S. Dynamically adjust the target weight w i Drive NSGA-II to iteratively optimize and obtain the dynamically updated Pareto optimal solution set. The solution selection unit screens the solution set according to the current task urgency and resource constraints through Bayesian optimization and hierarchical analysis method (AHP) and outputs the final scheduling solution.
9. The spatiotemporal adaptive matching system based on multi-objective reinforcement learning according to claim 1, characterized in that: The adaptive task optimization module includes: a real-time state monitoring unit, a priority dynamic evaluation unit, a strategy dynamic adjustment unit and a closed-loop feedback unit, wherein: the real-time state monitoring unit performs state acquisition and feature extraction processing according to the aircraft sensor data and environmental monitoring information to obtain a real-time state vector, and the priority dynamic evaluation unit performs a dynamic strategy adjustment unit according to the task urgency T i , Aircraft status S i 、Environmental complexity E i The weighted score is used to dynamically calculate the task priority and obtain the real-time priority weight w i The strategy dynamic adjustment unit performs online optimization processing of path planning and resource allocation strategies according to priority weights and historical feedback data, and outputs dynamically adjusted scheduling instructions. The closed-loop feedback unit evaluates strategy effects and corrects parameters based on task execution results and environmental change feedback, forming a closed-loop optimization mechanism to ensure continuous system adaptation.
10. The spatiotemporal adaptive matching system based on multi-objective reinforcement learning according to claim 1, characterized in that: The reinforcement learning driving module includes: a data integration unit, an algorithm configuration unit, a model training unit and a parameter real-time updating unit, wherein: the data integration unit performs state-action pair integration processing according to the spatiotemporal data output by the data fusion cleaning module and the feedback information of the adaptive task optimization module to obtain a reinforcement learning training data set; the algorithm configuration unit performs initialization configuration of Q-learning parameters and deep Q network (DQN) structure according to the task scenario requirements to obtain a customized reinforcement learning model; the model training unit performs model training processing according to historical data and real-time interaction data through experience replay and policy gradient descent to obtain an optimized Q value function and policy network; the parameter real-time updating unit performs online policy update and model parameter synchronization processing according to dynamic changes in the environment to ensure that the reinforcement learning algorithm of each module is always in the optimal state.
Citation Information
Patent Citations
Dynamic test flight task planning method based on deep reinforcement learning
CN115983566A
Unmanned aerial vehicle rapid path planning method and system based on deep learning
CN116449860A
Low-altitude airspace management method and system based on communication and sensing integration
CN118280168A
Service-oriented manufacturing resource optimization scheduling system based on adaptive learning algorithm
CN118586643A
Aircraft maintenance decision support system based on artificial intelligence
CN118863861A
Cited By
Console light switching control method based on artificial intelligence
CN120358652A
Multi-target energy consumption analysis and dynamic optimization system based on deep reinforcement learning
CN120469244A
Intelligent path navigation system for photovoltaic installation robot
CN120558240A
Biological feature recognition method driven by PPG big data
CN120611260A
Photovoltaic cleaning robot path planning method fusing group cascade power generation analysis
CN120628144A