Multi-target adaptive unmanned aerial vehicle path planning method and system for mobile crowd sensing

Through multi-objective utility functions, deterministic mapping and PEP-DQN algorithm, the real-time and energy optimization problems of drone path planning in dynamic mobile crowd perception scenarios are solved, efficient energy balance and real-time response of multi-drone collaborative operations are achieved, and task coverage and computing efficiency are improved.

CN120803054APending Publication Date: 2025-10-17JIANGSU VOCATIONAL COLLEGE OF BUSINESS +2
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511090453.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In dynamic mobile crowd sensing scenarios, existing drone path planning methods are difficult to achieve real-time performance, energy consumption optimization, and coverage balance in multi-drone collaborative operations. Especially in large-scale tasks and complex environments, the computational complexity is high, and traditional methods cannot quickly respond to changes in task allocation.

Method used

By adopting multi-objective utility function design, deterministic mapping strategy, PEP-DQN algorithm architecture and priority experience replay mechanism, and constructing a path planning method in a hybrid action space, combined with UAV energy consumption and mission benefits, an efficient and low-complexity decision-making mechanism is realized, which alleviates the cold start problem in deep reinforcement learning and improves strategy stability and sample utilization.

Benefits of technology

In a dynamic environment, efficient energy balance and collaborative task execution of UAV path planning are achieved, which significantly reduces energy consumption, improves task completion efficiency and strategy robustness, and meets the real-time response requirements of complex MCS scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803054A_ABST
    Figure CN120803054A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target adaptive unmanned aerial vehicle path planning method and system for mobile crowd sensing, the system comprises a platform layer, an unmanned aerial vehicle layer and a task layer, and the method comprises the following steps: constructing a multi-target utility function, and balancing task income and energy consumption; designing a task point-flight parameter deterministic mapping strategy in the hybrid action space; a PEP-DQN algorithm architecture is adopted, so that the decision-making efficiency and the strategy stability are improved; and a priority experience playback mechanism and an expert sample guiding mechanism are introduced. According to the method and the system, real-time performance and optimal energy consumption coordination of multi-unmanned aerial vehicle mixed action space path planning in a dynamic mobile crowd sensing scene are realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a mobile group cognitive sensing unmanned aerial vehicle path planning method and system, in particular to a multi-target adaptive unmanned aerial vehicle path planning method and system for mobile group cognitive sensing. BACKGROUND

[0002] Mobile group cognitive sensing (MCS) realizes dynamic data collection and processing of multi-scene environment by constructing large-scale distributed sensing network with intelligent terminals, which has broad application prospects in intelligent transportation, public safety and environmental monitoring fields. MCS system needs to complete task allocation and path planning in real time and efficiently between a large number of dispersed sensing tasks and mobile users, in order to balance coverage, timeliness and cost control. However, in special scenarios such as industrial parks, high-rise building dense areas or disaster sites, ground equipment often causes data blind area due to limited accessibility, and unmanned aerial vehicles (UAVs) can break through the space limit and undertake high-risk and repetitive tasks due to their three-dimensional maneuvering ability and rapid deployment advantage, which can significantly improve the data integrity and quality of the sensing network.

[0003] Although UAVs have unique advantages in MCS, they are limited by on-board battery capacity, sensitive to energy consumption, and the flight environment is often full of dynamic obstacles and unknown risks. Traditional path planning methods based on A*, RRT* and their various improved algorithms can generate feasible flight paths in static or semi-static environments, but they rely on prior maps and offline calculations, and are difficult to quickly respond to dynamic changes in task allocation, which can easily produce path redundancy and cause energy consumption to surge. In addition, when multiple UAVs perform collaborative tasks at the same time, traditional methods often need to perform brute force search or iterative optimization in high-dimensional parameter space, and the computational complexity increases exponentially with the task size and environmental complexity, which is difficult to meet the real-time requirements.

[0004] In the face of the above challenges, deep reinforcement learning (DRL) has become a powerful tool for solving dynamic path planning and energy consumption optimization, as it can adaptively learn decision-making strategies through interaction with the environment. Among them, the parameterized deep Q network (P-DQN) framework proposed by Xiong et al. realizes seamless integration of discrete task point selection and continuous flight parameters, and avoids exhaustive search of continuous parameters through deterministic policy network, thereby achieving preliminary success in hybrid action space. Fan et al. further introduced a hierarchical policy and priority experience replay mechanism on P-DQN, which optimized sample utilization efficiency and improved training convergence speed and policy stability. However, the existing P-DQN method still has problems such as slow cold start, insufficient policy robustness and limited energy consumption control ability when facing large-scale tasks and multi-UAV collaborative scenarios.

[0005] In the context of UAV-assisted MCS, the closest early work to the present invention is the research by Zhou et al. published in IEEE Transactions on Communications in 2018. They addressed the MCS system assisted by fixed-wing UAVs, transforming the task partitioning and path planning into a two-stage two-sided matching problem from the perspective of energy efficiency: the first stage generates UAV routes through dynamic programming or genetic algorithms, and the second stage completes the task allocation of UAVs and sub-regions using the Gale-Shapley algorithm. This method has been theoretically and simulated to significantly reduce energy consumption and improve data coverage. However, the method of Zhou et al. still requires offline calculation of the matching process and cannot quickly adjust the flight path when the task and environment change in real-time.

[0006] Following this, Gong, Zhang, and Li et al. published a research on online task allocation and path planning in IEEE Transactions on Vehicular Technology in 2018. They proposed four Greedy algorithms (QPA, TDA, DBA, B-DBA) to dynamically allocate tasks based on the ratio of task quality gain to user travel cost and adjust user paths online to ensure high task quality and low travel consumption in dynamic environments where users and tasks arrive. This method emphasizes the simplicity and efficiency of online greedy allocation, but still lacks in multi-UAV collaboration and energy consumption balancing.

[0007] In 2023, Wang et al. proposed a joint task allocation and path planning model for truck and UAV collaboration in Peer-To-Peer Networking and Applications. They used variable neighborhood search to simultaneously allocate tasks and optimize paths in a mixed fleet of logistics vehicles and UAVs. Experiments showed that this method had significant advantages in comprehensive task completion rate and travel cost. Although this solution alleviates the endurance pressure of UAVs by introducing ground vehicles, its ability to dynamically avoid obstacles and optimize three-dimensional energy consumption of UAVs still relies on traditional heuristic search, making it difficult to fully utilize the advantages of autonomous collaboration of UAVs.

[0008] In the field of deep reinforcement learning, in addition to single-machine P-DQN, academia has explored multi-agent P-DQN (DeepMAPQN) and parameterized actor-critic frameworks (HP-DQN, MAHHQN, etc.), trying to introduce coordination mechanisms into multi-UAV decision-making. However, due to factors such as computational complexity and non-stationarity, these methods have not been widely applied to large-scale MCS scenarios.

[0009] In summary, existing solutions, whether based on dynamic programming methods, greedy online allocation algorithms, or hybrid path planning combined with ground vehicles, have achieved varying degrees of success in task coverage or local efficiency, but it is difficult to simultaneously achieve high efficiency and low energy consumption, multi-UAV collaboration, and real-time online decision-making in dynamic MCS scenarios. SUMMARY

[0010] The technical problem to be solved by the present application is to realize the real-time and energy consumption optimal collaboration of multi-unmanned aerial vehicle hybrid action space path planning in the dynamic mobile crowd sensing scene. Specifically, it relates to: 1. How to construct an efficient and low complexity decision mechanism between discrete task points and continuous flight parameters; 2. How to alleviate the cold start and low sample utilization efficiency in deep reinforcement learning, and improve the policy convergence speed and stability; 3. How to realize the multi-objective balance optimization of coverage rate and energy consumption in multi-unmanned aerial vehicle cooperative operation.

[0011] In order to solve the above technical problems, the multi-objective adaptive unmanned aerial vehicle path planning method for mobile crowd sensing of the present application comprises the following steps: constructing a multi-objective utility function: considering the energy consumption and task benefit of the unmanned aerial vehicle, maximizing the system utility to guide the path decision, the optimization target of the multi-objective utility function is to maximize the cumulative reward, and the constraint conditions are that the unmanned aerial vehicles do not collide and the energy consumption of each task movement does not exceed the energy consumption limit of the unmanned aerial vehicle; designing a deterministic mapping strategy of task points and flight parameters in the hybrid action space: introducing a deterministic mapping function to connect the discrete task point selection and the continuous flight parameter adjustment, in the hybrid action space decision, the state includes task information and unmanned aerial vehicle usage, position information of the task and the unmanned aerial vehicle and unmanned aerial vehicle individual attributes, the action includes the next target task point of the discrete variable and the adjusted flight speed and heading angle of the continuous variable, the user equipment executes the action in the state to obtain the reward, and the reward is used to evaluate the advantages and disadvantages of the agent unloading decision; adopting PEP-DQN algorithm architecture: introducing the state and the deterministic mapping function of the discrete action to the continuous parameter, through Bellman equation transformation, target value definition, loss function setting and network parameter update formula, the decision efficiency and the policy stability are improved; introducing the priority experience replay mechanism and the expert sample guidance mechanism: in the experience replay mechanism, the priority replay strategy based on TD error is fused, the expert sample guidance strategy is matched, the utilization rate of key samples is improved, and the initial policy quality is optimized.

[0012] In the above method, a decay factor is introduced in the multi-objective utility function to measure the execution efficiency of the task within a given time, and the decay factor is

[0013]

[0014] Where, t K represents the deadline of the task, t iK represents the actual completion time of the task, and ∈ is a parameter for adjusting the influence of task completion time on benefit.

[0015] In the above method, the task movement energy consumption is

[0016]

[0017] Among them, X I and Y I and Z I Represents the X, Y and Z coordinates of the drone, and and Represents the X coordinate, Y coordinate and Z coordinate of the perception task T, L I Represents the energy consumption required for the drone to move, V I Indicates the moving speed of the drone.

[0018] In the above method, the Bellman equation is transformed into

[0019]

[0020] Among them, k t The discrete action chosen at a specific time t, x k is the corresponding continuous action; the target value is

[0021]

[0022] Among them, ω t and θ t is the parameter of the network at time t; the loss function is

[0023]

[0024] The network parameter update formula is:

[0025]

[0026] Among them, α t and β t are the update steps of the evaluation network Q(ω) and the target network x(θ) at time t, and are the stochastic gradients of the evaluation network Q(ω) and the target network x(θ) at time t, respectively.

[0027] In the above method, the priority experience replay mechanism dynamically adjusts the sampling probability according to the learning value of the sample, giving priority to learning high-value samples; the expert sample guidance mechanism generates high-quality samples through a greedy algorithm in the early stage of training, and gives priority to selecting the server configuration with the best computing resources as the initial strategy.

[0028] The mobile crowd sensing oriented multi-objective adaptive UAV path planning system comprises a platform layer, a UAV layer and a task layer, wherein: the platform layer performs multi-dimensional feature analysis on the demand submitted by a data requester through an intelligent task analysis module, generates an optimal task allocation sequence based on a task attribute feature matching algorithm; the UAV layer considers the endurance capability, load limit and other UAV performance parameters based on a dynamic programming algorithm, and performs real-time path optimization on the allocated task sequence; the planning result is mapped to the task layer in real time through a bidirectional feedback mechanism, forming a closed-loop optimization from task allocation to path planning to effect evaluation.

[0029] In the above system, an online re-planning module for dynamic obstacle environment is further included for realizing real-time response and energy efficiency optimization in the event of sudden changes.

[0030] In the above system, the platform layer further comprises a multi-objective utility function module for comprehensively considering the energy consumption and task benefits of the UAV to guide path decision.

[0031] In the above system, the UAV layer further comprises a deterministic mapping module for realizing deterministic mapping of task points and flight parameters in a hybrid action space to improve computational efficiency; the UAV layer further comprises a PEP-DQN algorithm module for improving decision efficiency and policy stability to realize organic unification of discrete and continuous action spaces; the UAV layer further comprises a priority experience replay and expert sample guidance module for improving the utilization rate of key samples and optimizing the initial policy quality, accelerating training convergence and relieving the cold start problem.

[0032] The method and system of the present application propose a path planning strategy based on an improved parameterized deep Q network (PEP-DQN) to realize efficient energy consumption balance and collaborative task execution in a dynamic environment, aiming at the contradiction between real-time response and energy consumption optimization in UAV path planning in a mobile crowd sensing system. The method first comprehensively considers the energy consumption and task benefits of the UAV in the utility function, maximizes the system utility to guide path decision, thereby effectively controlling energy consumption while meeting the sensing coverage rate. Secondly, by introducing a deterministic mapping function, the connection between discrete task point selection and continuous flight parameter adjustment is established, avoiding the exhaustive search overhead of traditional P-DQN on continuous parameters, and significantly improving the computational efficiency. In order to accelerate training convergence and relieve the cold start problem, a priority replay strategy based on TD error and an expert sample guidance strategy are integrated into the experience replay mechanism to improve the utilization rate of key samples and optimize the initial policy quality. Finally, through multi-scenario simulation experiments, the algorithm proposed in the present application shows significant advantages in energy consumption reduction, task completion efficiency and policy robustness, verifying its feasibility and practicality in complex MCS scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 Multi-objective adaptive UAV path planning system for mobile crowd sensing;

[0034] Figure 2 Multi-objective adaptive UAV path planning method flow chart for mobile crowd sensing;

[0035] Figure 3 PEP-DQN reinforcement learning model. DETAILED DESCRIPTION

[0036] Reference Figure 1 The multi-objective adaptive UAV path planning system for mobile crowd sensing adopts a hierarchical architecture design, and the system is composed of a platform layer, a UAV layer and a task layer to form a three-level linkage mechanism.

[0037] Specifically, in the platform layer, the MCS platform analyzes the demand submitted by the data requester through the intelligent task analysis module, generates the optimal task allocation sequence based on the task attribute feature matching algorithm; in the UAV layer, each agent considers the endurance capability, load limit and other UAV performance parameters based on the dynamic programming algorithm, and optimizes the allocated task sequence in real time; finally, the planning result is mapped to the task layer in real time through the bidirectional feedback mechanism, forming a closed-loop optimization system of "task allocation-path planning-effect evaluation". This hierarchical collaborative mechanism effectively improves the task execution efficiency and resource utilization of the system in complex environments.

[0038] Reference Figure 2 The multi-objective adaptive UAV path planning method for mobile crowd sensing solves the above problems by adopting the following technical steps: 1. Constructing a multi-objective utility function to balance task benefits and energy consumption; 2. Designing a deterministic mapping strategy for task points-flight parameters in a mixed action space; 3. Using PEP-DQN algorithm architecture to improve decision efficiency and strategy stability; 4. Introducing priority experience replay mechanism and expert sample guidance mechanism to accelerate convergence and alleviate cold start. The overall technical framework of this paper is shown in Figure 2 .

[0039] Step 1: Construct a multi-objective utility function to balance task benefits and energy consumption:

[0040] The UAV model can be described as , where I represents the UAV number.

[0041] The UAV position: X I and Y I and Z I represent the X coordinate and Y coordinate and Z coordinate of the UAV, which are used to determine the current position of the UAV in the task scene.

[0042] The UAV speed V I : VI represents the moving speed of the UAV. It represents the flight speed of the UAV, usually in units of meters per second (m / s).

[0043] heading angle represents the horizontal direction angle (yaw angle), which represents the orientation of the UAV in the horizontal plane, usually in the range of [-π, π]. I represents the pitch angle, which represents the orientation of the UAV in the vertical plane, usually in the range of [-π / 2, π / 2].

[0044] remaining battery power E I : E I represents the remaining battery power of the UAV, usually in percentage or joules (J).

[0045] UAV energy consumption L I : L I represents the energy consumption required for the UAV to move, usually in percentage or joules (J).

[0046] Task T can be described by a set , which is defined as follows:

[0047] Position of the task: and and represent the x-coordinate and y-coordinate and z-coordinate of the perception task T, respectively, which are used to determine the position of the task in the scene.

[0048] Task priority: represents the priority of the task T K . Each task is assigned a priority parameter, and the higher the value, the higher the priority. The priority of the task is directly related to the perception benefit of the task, and also affects the perception benefit of the mobile user and the UAV when executing the task. High priority tasks often have a greater impact on perception quality. If the task is not completed within the current time slot, its priority will gradually decrease.

[0049] Task completion speed: represents the completion speed of the task T K , which measures the execution efficiency of the task within a given time. The task T K is provided with a deadline t K , and if the task is completed closer to the deadline, the perception quality benefit will be significantly reduced. In order to quantify the impact of time on perception quality, we introduce a decay factor, which is defined by the following formula and multiplied by the task perception benefit:

[0050]

[0051] where tK denotes the deadline of the task, t iK denotes the actual time of completing the task, ∈ is a parameter to adjust the impact of the time of completing the task on the reward.

[0052] the basic perceived reward of the task denotes the basic reward perceived at the time of completing the task T in an ideal case. This reward reflects the basic return of completing the task without the influence of time decay and other factors.

[0053] the task assignment target denotes that the task is assigned to the UAV corresponding to the serial number.

[0054] Through the definition of these parameters, the model can more finely describe the spatial, temporal and reward characteristics of the task, which helps to optimize the path planning strategy and execution effect.

[0055] reward function R I is used to evaluate the performance of the UAV I after performing the action at time t, and is designed in the form of multi-objective optimization:

[0056] R I = W1*R task -W2*R energy -W3*R collision #(2)

[0057] wherein:

[0058] task completion reward R task is the reward obtained by the UAV after completing the task, which is calculated as follows:

[0059]

[0060] the energy consumption penalty of the UAV is the energy consumption consumed by the UAV in performing the task movement, and the calculation formula is as follows:

[0061]

[0062] According to the UAV path planning model defined above, the optimization target of the UAV path planning system is usually to maximize the cumulative reward G t :

[0063]

[0064] C2:R energy ≤ E I #(7)

[0065] Among them, constraint condition C1 ensures that the drone will not collide during movement. Constraint condition C2 ensures that the drone does not exceed its own energy consumption limit each time it moves. The specific optimization objectives include first maximizing task coverage, second covering as many high-priority task points as possible within a limited time, and finally minimizing energy consumption, optimizing flight path and speed, and prolonging the endurance of the drone.

[0066] Through the above multi-objective utility function design, the agent can consider the perception coverage and energy consumption control simultaneously during the training process, ensuring that the finally learned path planning strategy is both efficient and energy-saving in dynamic environments.

[0067] Step two designs a deterministic mapping strategy of task points-flight parameters in the hybrid action space:

[0068] In the scenario of mobile crowd sensing (MCS) assisted by drones, appropriate path planning strategies need to be developed for different task attributes in the scenario and drones to ensure that each task can be successfully completed. Therefore, the mobile crowd sensing platform can be regarded as a reinforcement learning agent, whose task is to design the optimal path planning strategy according to the task decision sequence in the environment and the drone.

[0069] Under normal circumstances, if the precise state transition probability matrix P can be obtained, the task offloading problem can be perfectly solved through the four-tuple (S, A, R, P) relying on the theory of Markov decision process. However, due to the continuous change of the state of the task and the drone over time, it is extremely difficult to obtain accurate transition probabilities. To solve this problem, this patent considers using a model-free reinforcement learning method based on the three-tuple (S, A, R).

[0070] 1. State state: The goal of reinforcement learning is to gradually approach the omniscient perspective by continuously learning strategies from historical data. Therefore, a comprehensive and accurate definition of the state is of great importance to improving decision-making efficiency. This patent considers the task information and the usage of the drone, the location information of the task and the drone, and the individual attributes of the drone, and defines the state s t as follows:

[0071] s t = {P(t), C(t), O(t), I(t)},#(8)

[0072] where P(t) = {P1(t), P2(t), P3(t), …, P k (t)} represents the basic attributes of k tasks at time t, C(t) = {C1(t), …, C k (t)} represents which drone the k tasks are assigned to. O(t) = {O1(t), …, O m (t), …, Ok+m (t)} represents the position information of m UAVs and k tasks within time slot t. Finally, i(t) = {I1(t),…,I m (t)} represents the basic attributes of m UAVs within time slot t.

[0073] 2. Action a: At time slot t, the MCS platform makes action decision based on the environment state s t . This decision includes the next target task point of discrete variables and the adjusted flight speed and heading angle of continuous variables. Therefore, the action a t at time slot t is defined as follows:

[0074] a t = {λ(t), θ(t)}, # (9)

[0075] where λ(t) = {λ1(t), λ2(t), λ3(t),…, λ k (t)} represents the sequence of tasks performed by the UAV within time slot t, This decision represents the flight speed and heading angle of the UAV within time slot t. Due to the sequence of tasks performed, the flight speed and heading angle of the UAV are continuous variables, so the decision space of the platform is composed of both discrete variables and continuous variables. In addition, the decision-making process follows a certain order: first, the sequence of tasks performed is determined, and then the flight speed and heading angle of the UAV are determined.

[0076] 3. Reward r: At time slot t, the user equipment performs action a t in state s t and obtains immediate reward r t , which is used to evaluate the quality of the agent's offloading decision. In the reinforcement learning framework, the core goal of the agent is to choose the optimal action to maximize the reward given the environment. The optimization goal of this patent is to maximize the system utility, which is given in formulas 4 and 5. Therefore, the reward function is defined as follows:

[0077]

[0078] where μ represents the penalty term. When the task allocation decision of the MCS platform does not meet the constraints listed in formulas 6 and 7, it will be subject to the corresponding penalty. This mechanism aims to provide feedback signals to the agent, indicating that the action it has chosen has suboptimal properties, thereby guiding the agent to optimize its strategy to meet the constraint conditions and improve the quality of decision-making.

[0079] Traditional reinforcement learning algorithms usually assume that the action space is completely continuous or completely discrete. Algorithms that focus on continuous action spaces, such as EDDPG (Ensemble Deep Deterministic Policy Gradient) and DQN (Deep Q-Network) and its improved algorithm Double DQN, face significant challenges when dealing with discrete action space problems. Although quantizing the action space can alleviate this problem to a certain extent, this approach may lead to dimensionality disaster or information loss. To address the mixed action space problem in human-machine collaborative task allocation, this paper proposes an improved algorithm PEP-DQN (Progressive Enhancement P-DQN), which is an enhanced version based on P-DQN (Parameterized Deep Q-Network)

[16] . The P-DQN algorithm can effectively handle discrete-continuous mixed action space problems and extend the traditional DQN algorithm to a hybrid structure. PEP-DQN optimizes the continuous action selection process by introducing a deterministic mapping function from states and discrete actions to continuous parameters, thereby significantly improving training efficiency and algorithm stability. The action-value function (AVF) of this algorithm maps system states and mixed actions into real values, which can determine the optimal discrete action without exhaustive continuous parameter search. Specifically, the action space can be formally defined as:

[0080] A={(k,x k )|x k ∈X k ,k∈[K]},#(11)

[0081] In this case, [K] represents a high-level discrete action set, where k is a discrete action selected from [K]. k represents the set of low-level continuous actions associated with discrete action k, x k It means from X k The action selection in P-DQN follows a greedy strategy.

[0082] Through the above strategy, PEP-DQN does not need to perform brute force search in the continuous parameter space, but directly generates high-quality parameters through deterministic mapping, significantly reducing the amount of computation and maintaining real-time decision-making capabilities in dynamic environments.

[0083] Adopting the PEP-DQN algorithm architecture to improve decision-making efficiency and strategy stability:

[0084] After completing the multi-objective utility function design and hybrid action space decision-making strategy, the next step is to build and optimize the PEP-DQN reinforcement learning framework so that it can converge quickly and run stably in dynamic MCS scenarios.

[0085] The PEP-DQN algorithm proposed in this patent innovatively realizes the organic unification of discrete and continuous action spaces, showing significant advantages compared to traditional algorithms such as DDQN (Double Deep Q-Network) and EDDPG (Ensemble Deep Deterministic Policy Gradient). The core innovation of this algorithm lies in its unique mixed action space processing mechanism: the agent can make discrete decisions in certain dimensions while performing continuous action selection in other dimensions. This multi-dimensional mixed decision architecture not only enhances the algorithm's expressive ability but also significantly improves its adaptability to complex tasks, enabling it to effectively handle real-world scenarios that contain both discrete and continuous decision elements. This design breaks through the limitations of traditional algorithms in single action space types, providing a new solution for reinforcement learning applications in mixed action space problems.

[0086] In PEP-DQN, consider an MDP with an action space defined in equation 9, for a∈A, there is Q(s,a)=Q(s,k,x k ). Set the discrete action selected at a certain t time as k t , and the corresponding continuous action as x k , then the Bellman equation becomes:

[0087]

[0088] Like the idea of DQN, use the deep neural network Q(s,k,x k ;ω) to approximate Q(s,k,x k ), when ω is fixed, we hope to find a set of θ that satisfies:

[0089]

[0090] Let ω t and θ t be the network parameters at t time, define the target value y t of the nth step as:

[0091]

[0092] In the training phase, update the parameters by minimizing the size of the loss function, the loss function of this algorithm is expressed as follows:

[0093]

[0094] The network parameter update formula is:

[0095]

[0096] where, a t and b t are the update steps of the evaluation network Q(ω) and the target network x(θ) at time t, respectively, and are the stochastic gradients of the evaluation network Q(ω) and the target network x(θ) at time t, respectively.

[0097] The network structure diagram of the PEP-DQN algorithm is shown in Figure 3 .

[0098] Through the above series of improvements, compared with the traditional P-DQN algorithm, the PEP-DQN can realize a more stable and faster training process than the original P-DQN. The time complexity of the algorithm is affected by the training period episode and the time step T, so the time complexity of the algorithm is O(n T n episode ), wherein n T represents the number of time steps in each episode, and n episode represents the number of episodes.

[0099] The priority experience replay mechanism and the expert sample guiding mechanism are introduced to accelerate convergence and alleviate cold start.

[0100] The PEP-DQN algorithm framework is proposed, which significantly improves the performance of the P-DQN algorithm by introducing deterministic functions, action-value function (AVF) optimization, priority experience replay mechanism, and expert sample guidance. Specifically, PEP-DQN establishes a mapping relationship between state-discrete action and continuous parameters through deterministic functions, effectively avoiding the time-consuming exhaustive search process in traditional methods, improving computational efficiency while enhancing algorithm stability. At the optimization level, the algorithm maps state-action pairs to real value space through action-value function, realizing efficient optimization of mixed action space. To further improve sample utilization efficiency, the algorithm uses a priority experience replay mechanism based on TD error, dynamically adjusts the sampling probability according to the learning value of the sample, thereby improving the utilization rate of training data. To address the cold start problem, the study proposes a guiding strategy based on expert samples: in the early stage of training, high-quality samples are generated through a greedy algorithm, and the server configuration with the optimal computing resources is preferentially selected as the initial strategy. This design not only effectively accelerates the convergence speed of the training process, but also reduces the risk of the algorithm falling into local optimum by introducing prior knowledge, especially the robustness in complex environments is significantly improved.

[0101] By incorporating the balance optimization goal of energy consumption and task benefit into the utility function, introducing a deterministic mapping function to realize efficient integration of discrete task points and continuous flight parameters, and combining the priority experience replay based on TD error and the expert sample guidance mechanism, the convergence speed and stability of the strategy under the mixed action space are effectively improved.

Claims

1. A multi-target adaptive UAV path planning method for mobile crowd intelligence perception, characterized by: The following steps are involved: Construct a multi-objective utility function: Considering the drone's energy consumption and mission benefits, the path decision is guided by maximizing system utility. The optimization goal of the multi-objective utility function is to maximize the cumulative reward while satisfying the constraints of non-collision between drones and the constraints that the energy consumption of each mission movement does not exceed the drone's own energy consumption limit. Design a deterministic mapping strategy between task points and flight parameters in a hybrid action space: Introduce a deterministic mapping function to connect discrete task point selection with continuous flight parameter adjustment. In hybrid action space decision-making, the state includes task information and drone usage, task and drone location information, and individual drone attributes. Actions include the next target task point (discrete variables) and the adjustment of flight speed and heading angle (continuous variables). User devices receive rewards for executing actions within the state, and the rewards are used to evaluate the quality of the agent's offloading decisions. Adopt the PEP-DQN algorithm architecture: Introduce a deterministic mapping function from state and discrete actions to continuous parameters. Through Bellman equation transformation, target value definition, loss function setting, and network parameter update formula, decision-making efficiency and strategy stability are improved. Introducing a priority experience replay mechanism and an expert sample guidance mechanism: Integrating a priority replay strategy based on TD error into the experience replay mechanism and matching it with an expert sample guidance strategy to improve the utilization of key samples and optimize the quality of the initial strategy.

2. The multi-target adaptive UAV path planning method for mobile crowd intelligence perception according to claim 1 is characterized by: The attenuation factor is introduced into the multi-objective utility function to measure the execution efficiency of the task within a given time. The attenuation factor is Among them, t K represents the deadline of the task, t iK represents the actual time to complete the task, and ∈ is a parameter that adjusts the impact of task completion time on revenue.

3. The multi-target adaptive UAV path planning method for mobile crowd intelligence perception according to claim 1 is characterized in that : The task movement energy consumption is Among them, X I and Y I and Z I Represents the X, Y and Z coordinates of the drone, and and Represents the X coordinate, Y coordinate and Z coordinate of the perception task T, L I Represents the energy consumption required for the drone to move, V I Indicates the moving speed of the drone.

4. The multi-target adaptive UAV path planning method for mobile crowd intelligence perception according to claim 1 is characterized in that : The Bellman equation is transformed into Among them, r t is the immediate reward, k t The discrete action chosen at a specific time t, x k is the corresponding continuous action; the target value is Among them, ω t and θ t is the parameter of the network at time t; the loss function is The network parameter update formula is: Among them, α t and β t are the update steps of the evaluation network Q(ω) and the target network x(θ) at time t, and are the stochastic gradients of the evaluation network Q(ω) and the target network x(θ) at time t, respectively.

5. The multi-target adaptive UAV path planning method for mobile crowd intelligence perception according to claim 1 is characterized in that: The priority experience replay mechanism dynamically adjusts the sampling probability according to the learning value of the sample, giving priority to learning high-value samples; the expert sample guidance mechanism generates high-quality samples through a greedy algorithm in the early stage of training, and gives priority to selecting the server configuration with the best computing resources as the initial strategy.

6. A multi-target adaptive UAV path planning system for mobile crowd intelligence perception, characterized by: It includes platform layer, UAV layer and task layer, among which: the platform layer performs multi-dimensional feature analysis on the requirements submitted by the data requester through the intelligent task parsing module, and generates the optimal task allocation sequence based on the task attribute feature matching algorithm; each intelligent agent in the UAV layer performs real-time path optimization of the assigned task sequence based on the dynamic planning algorithm, taking into account UAV performance parameters such as endurance and load limit; the planning results are mapped to the task layer in real time through a two-way feedback mechanism, forming a closed-loop optimization from task allocation to path planning to effect evaluation.

7. The multi-target adaptive UAV path planning system for mobile crowd intelligence perception according to claim 6 is characterized by: It also includes an online replanning module for dynamic obstacle environments, which is used to achieve real-time response and energy efficiency optimization in the event of sudden changes.

8. The system according to claim 6, wherein: The platform layer also includes a multi-objective utility function module, which is used to comprehensively consider the energy consumption and mission benefits of the drone and guide path decision-making.

9. The system according to claim 6, wherein: The drone layer also includes a deterministic mapping module for realizing deterministic mapping of task points and flight parameters in the hybrid action space, thereby improving computational efficiency. The drone layer also includes a PEP-DQN algorithm module for improving decision-making efficiency and strategy stability, thereby realizing the organic unity of discrete and continuous action spaces. The drone layer also includes a prioritized experience replay and expert sample guidance module for improving the utilization rate of key samples and optimizing the quality of initial strategies, thereby accelerating training convergence and alleviating cold start problems.

Citation Information

Patent Citations

  • Routing optimization architecture and method based on deep reinforcement learning under SDN architecture

    CN113395207A

  • Item delivery unmanned aerial vehicle cluster flight path planning method based on improved DQN algorithm

    CN116414146A

  • Local path planning method based on deep reinforcement learning

    CN116578080A

  • Unmanned aerial vehicle agricultural bird repelling method and system based on topological sorting reward mechanism

    CN117441701A

  • Unmanned aerial vehicle task scheduling and flight path design method

    CN118468695A