Unmanned aerial vehicle data acquisition method based on prediction enhanced deep reinforcement learning network
By introducing a prediction-enhanced deep reinforcement learning network into the UAV system, the predictor and exploration adapter are used to enhance the foresight and collaborative efficiency of the strategy, which solves the problems of insufficient environmental prediction and low collaborative efficiency in multi-UAV systems, and achieves more efficient trajectory planning and task completion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIDIAN UNIV
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-08
AI Technical Summary
Traditional multi-agent deep reinforcement learning methods lack the ability to predict the future state of the environment in UAV systems, resulting in a lack of foresight in action strategies, insufficient efficiency in multi-UAV collaboration, insufficient exploration capabilities, and low training efficiency due to sparse rewards.
We introduce a prediction-enhanced deep reinforcement learning network to enhance the foresight of the policy through predictors and exploration adapters, and design an adaptive intrinsic reward mechanism to enhance state representation and collaborative efficiency by using predictors and exploration adapters to guide UAVs in location exploration and regional cooperation.
Achieving stable and effective trajectory planning in complex and dynamic environments avoids redundant visits and regional congestion, enhances the foresight, exploratory, and collaborative capabilities of multiple UAVs, and significantly improves the quality of trajectory planning.
Smart Images

Figure CN121995936A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multi-UAV intelligent control technology, specifically relating to a UAV data acquisition method based on a predictive-enhanced deep reinforcement learning network. Background Technology
[0002] With the development of UAV applications in scenarios such as inspection, surveying, and emergency support, multi-UAV systems are gradually becoming mainstream in large-scale continuous missions. In order to achieve efficient cooperative trajectory planning in dynamic environments, multi-agent deep reinforcement learning (MADRL) has become an important solution.
[0003] Traditional MADRL solutions such as MATD3 and MADDPG, while possessing a certain degree of intelligence, still have the following shortcomings: 1) Lack of ability to predict the future state of the environment: This leads to a lack of foresight in action strategies, making it difficult to maintain stable performance over long time scales; 2) Insufficient collaboration efficiency among multiple drones: There are often problems such as different drones concentrating in similar areas, repeated visits, and wasted resources; 3) Insufficient or unstable exploration ability: Without additional guidance and rewards, the agent is prone to getting stuck in local optima in the early stages; 4) Sparse rewards lead to low training efficiency: In multi-drone missions, external rewards are often sparse, making it difficult for the model to obtain effective training signals.
[0004] Therefore, a multi-agent reinforcement learning method is needed that can simultaneously address the issues of "lack of foresight" and "insufficient collaborative efficiency". Summary of the Invention
[0005] To address the aforementioned problems in the existing technology, this invention provides a method for UAV data acquisition based on a prediction-enhanced deep reinforcement learning network.
[0006] The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a method for UAV data acquisition based on a predictive-enhanced deep reinforcement learning network, characterized in that it is applied to each UAV in a UAV system, and the method includes: The drone acquires its current state and inputs it into its own pre-trained policy network, which then outputs the current action. The drone performs the action at the current moment to collect data from the corresponding ground data acquisition node; The trained policy network is obtained by using reinforcement learning to jointly train the Actor network, Critic network, predictor, and explorer adapter deployed on each UAV in the UAV system. The predictor and the explorer adapter assist in the training of the Actor network and the Critic network. The outputs of the predictor and the explorer adapter are used to calculate the adaptive intrinsic reward of the UAV during the training process. The adaptive intrinsic reward is used to guide the UAV to perform location exploration and regional cooperation.
[0007] In some embodiments, the current state includes: the state information of the ground data acquisition nodes within the current sensing range, and the state information of the UAV itself at the current moment. The state information of the ground data acquisition nodes within the current sensing range includes: the position of all ground data acquisition nodes within the current sensing range of the UAV, the remaining data volume, and the data generation time of each remaining data packet. The state information of the UAV itself at the current moment includes: the position of the UAV at the current moment and the remaining energy. The actions at the current moment include: the angle used to control the direction of movement of the drone at the current moment, and the flight distance of the drone at the current moment.
[0008] In some embodiments, the predictor is a neural network, and the predictor is used to predict based on the input of any drone. The state information of the ground data acquisition nodes within the sensing range of the time step outputs a prediction probability, where the prediction probability represents the... The probability that a ground data acquisition node within the perception range of a time step belongs to any one of the UAVs.
[0009] In some embodiments, the loss function of the predictor is expressed as follows: ; in, Indicates the loss value. Represents the cross-entropy loss function. Represents the heat function. Indicates the first in the unmanned aerial vehicle system The first drone, the "each drone" represents any one of the aforementioned drones. The value ranges from 1 to , This indicates the total number of drones in the unmanned aerial vehicle system. Indicates the first A drone Status information of ground data acquisition nodes within the sensing range of the time step. express The ground data acquisition node it represents belongs to the first The probability of a drone.
[0010] In some embodiments, the exploration adapter is configured to, based on input from any one of the drones in the unmanned aerial vehicle system... The state of the time step generates the state of any one of the drones in The degree of spatial exploration at each time step, wherein any one of the drones is in The expression for the degree of spatial exploration at a time step is as follows: ; in, Indicates the first in the unmanned aerial vehicle system A drone The degree of spatial exploration in time step The value ranges from 1 to , This indicates the total number of drones in the unmanned aerial vehicle (UAV) system. Indicates the first A drone The state of the time step, This is the running average. This indicates an embedded network with fixed network parameters in the exploration adapter. This indicates the embedded network in the exploration adapter that has network parameters that need to be trained.
[0011] In some embodiments, the loss function of the exploration adapter is expressed as follows: ; in, Indicates the loss value. This indicates the operation of seeking the expected value.
[0012] In some embodiments, during training, any one of the drones in the unmanned aerial vehicle system performs... After the time step action, the result is The total reward for each time step includes: Environmental rewards for time steps and The adaptive intrinsic reward of the time step, wherein, for any one of the drones The expression for the environmental reward at each time step is as follows: ; ; in, Indicates the first A drone Environmental rewards for each time step, the first "One drone" refers to any one drone in the drone system. The value ranges from 1 to , This indicates the total number of drones in the unmanned aerial vehicle system. This represents the total number of ground data acquisition nodes on the ground. , This represents the total number of steps in each training session. This indicates the time taken at each time step. , Representing drones from the first The moment when data packets are collected at each ground data acquisition node , Indicates the first The generation time of the earliest data packet generated among the ground data acquisition nodes. The value ranges from 1 to , Indicates the first The weight of each ground data acquisition node, This represents a preset constant term. This indicates the preset AOI threshold. Indicates the first A drone The time step takes into account energy consumption or penalties for hitting obstacles.
[0013] In some embodiments, any one of the drones The expression for the adaptive intrinsic reward at each time step is as follows: ; ; ; in, Represents any one of the drones Adaptive intrinsic reward at time step This indicates that any one of the drones is in The degree of spatial exploration in time step This indicates that any one of the drones is in Time step location diversity reward This indicates that any one of the drones is in Regional cooperation rewards for time steps Represents any one of the drones The state of the time step, express Included The state information of the time step itself. express Included Status information of ground data acquisition nodes within the sensing range of the time step. This indicates that any one of the drones is in The set of positions of all time steps preceding the current time step. Represents a set Any position in Expressing the request Position and location in The Euclidean distance between them This indicates finding the absolute value. express The ground data acquisition node it represents belongs to the first The probability of a drone. This represents the network parameters of the predictor. In some embodiments, any one of the drones is performing After the time step action, the result is The expression for the total reward at each time step is as follows: ; in, This indicates that any one of the drones is performing... After the time step action, the result is Total reward for time step This represents the weighting coefficient.
[0014] In some embodiments, the application scenarios of the unmanned aerial vehicle system include: inspection scenarios, emergency response scenarios, and mobile sensing scenarios.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention, based on the traditional Multi-Agent Response Technology (MATD3) framework, introduces a predictor and an explorer adapter to enhance the foresight of the policy and designs an adaptive intrinsic reward mechanism. The predictor and explorer adapter enhance state representation, while the adaptive intrinsic reward mechanism improves the collaborative efficiency among UAVs, enabling multiple UAVs to achieve stable and effective trajectory planning in unknown environments, avoiding repeated visits, collisions, or regional congestion. This gives multiple UAVs stronger foresight, exploratory capabilities, and collaborative abilities in complex dynamic environments, significantly improving trajectory planning quality and enabling wide application in collaborative UAV tasks such as inspection, emergency response, and mobile sensing.
[0016] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the UAV data acquisition method based on a prediction-enhanced deep reinforcement learning network provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the training framework provided in an embodiment of the present invention; Figure 3 These are heatmaps of simulated drone trajectories from RC-MATD3 and MATD3; Figure 4 This is a comparison chart of the simulation benefits of the two algorithms. Detailed Implementation
[0018] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0019] Figure 1 This is a flowchart illustrating a UAV data acquisition method based on a prediction-enhanced deep reinforcement learning network, provided by an embodiment of the present invention. This method is applied to each UAV in the UAV system, such as... Figure 1 As shown, the method includes: S101. The UAV acquires the current state and inputs it into its own pre-trained policy network. The pre-trained policy network then outputs the current action.
[0020] In this invention, the current state includes: the state information of the ground data acquisition nodes within the current sensing range, and the current state information of the UAV itself. Specifically, the state information of the ground data acquisition nodes (PoIs) within the current sensing range includes: the positions of all ground data acquisition nodes within the UAV's sensing range, the remaining data volume of each ground data acquisition node, and the data generation time of each remaining data packet in each ground data acquisition node. The current state information of the UAV itself includes: the UAV's current position and remaining energy. The current actions include: the angle used to control the UAV's direction of movement at the current moment, and the UAV's current flight distance.
[0021] S102. The UAV executes the action at the current moment to collect data from the corresponding ground data acquisition node; wherein, the trained policy network is a trained Actor network obtained by jointly training the Actor network, Critic network, predictor and explorer adapter deployed on each UAV in the UAV system using reinforcement learning method. The predictor and explorer adapter are used to assist the training of the Actor network and Critic network, and the output of the predictor and the output of the explorer adapter are used to calculate the UAV's adaptive intrinsic reward during the training process. The adaptive intrinsic reward is used to guide the UAV to conduct location exploration and area cooperation.
[0022] This invention does not limit the number of Actor networks and Critic networks deployed on each drone. For example, Figure 2 This is a schematic diagram of the training framework provided by the present invention. For example... Figure 2 As shown, the Agent represents a drone. Each drone can be deployed with one Actor network, one Target Actor network, one Critic1 network, one Target Critic1 network, one Critic2 network, and one Target Critic2 network. As training progresses, the output of the Actor network gradually fits the Target Actor network; the same applies to the Critic1 and Critic2 networks. Furthermore, Figure 2 In this context, O represents the observed state during training, and R represents the total reward obtained.
[0023] In this invention, the number of layers and neurons in the Actor network and Critic network are set according to the needs of the actual scenario, and this invention does not limit this. In this invention, during each training session, the parameters of the Actor network, Critic network, predictor, and explorer adapter are updated, and the training method adopts existing training methods in reinforcement learning. It should be noted that the predictor is a neural network; for example, the predictor adopts a DNN network architecture. The predictor is used to determine the location of any input drone... The state information of the ground data acquisition nodes within the sensing range of the time step outputs a prediction probability, which represents the probability of the time step. The predictor calculates the probability that a ground data acquisition node within the perception range of a given time step belongs to any given UAV. The predictor is trained using a self-supervised method, taking the state information of the ground data acquisition node as input and outputting the probability distribution of data acquisition at that node being handled by different UAVs. At each time step, based on the UAV that actually completes data acquisition at the ground data acquisition node, the predictor automatically generates the corresponding UAV identity as training reference information, and calculates the error of the predictor's output based on this reference information, thereby updating the predictor's network parameters.
[0024] For example, the application scenarios of the above-mentioned drone system include drone collaborative scenarios such as inspection scenarios, emergency response scenarios, and mobile sensing scenarios, to perform drone collaborative tasks in these scenarios.
[0025] In this invention, during each training session, all ground data acquisition nodes have the following three performance metrics: a) All ground data acquisition nodes share a single overall AoI (Age of Information), using... It means that, among them, The expression is as shown in formula (1): (1); in, This represents the total number of ground data acquisition nodes on the ground. , This represents the total number of steps in each training session (i.e., the number of time steps in each training session is as follows). ), This indicates the time taken at each time step. , Representing drones from the first The moment when data packets are collected at each ground data acquisition node , Indicates the first The generation time of the earliest data packet generated among the ground data acquisition nodes. The value ranges from 1 to , Indicates the first The weight of each ground data acquisition node, Indicates the first The waiting time before the earliest generated data packet in each ground data acquisition node is collected by the drone.
[0026] b) All data packets in all ground data acquisition nodes have an AoI threshold, using... express, The expressions are shown in formulas (2) to (3): (2); (3); in, This represents a preset constant term. This indicates the preset AOI threshold.
[0027] c) Each ground data acquisition node corresponds to a data acquisition rate, using... express, The expression is as shown in formula (4): (4); in, Indicates the first time steps (i.e., time steps) or (time step) Indicates to Find the gradient. Indicates the first Each ground data acquisition node at time step The amount of data.
[0028] In this invention, the ultimate goal of each training iteration is to jointly minimize... And maximize the collection ratio, and the energy consumption does not exceed the maximum energy consumption. Treating the penalty function as a constant penalty, as shown in formula (3) above, the overall objective can then be expressed as formula (5), with constraints as formula (6): (5); (6); in, Indicates the first A drone The time step takes into account energy consumption or penalties for hitting obstacles. The value ranges from 1 to , This represents the total number of drones in the unmanned aerial vehicle (UAV) system, the first... Each drone represents any one drone in a drone system. Indicates the first A drone Energy consumption per time step.
[0029] To address the aforementioned issues, this invention requires constructing a partially observable Markov decision process model for the UAV before training. The UAV Markov model comprises the following three aspects: 1) Observation space (i.e., state): Observation space The first in the unmanned aerial vehicle system The state of any drone (i.e., any drone) at time step is represented as follows: Furthermore, referring to the above Figure 2 , include and Two parts, Indicates the first A drone Status information of ground data acquisition nodes within the sensing range of the time step. Indicates the first A drone The location of each ground data acquisition node within the sensing range of the time step, the remaining amount of data, and the data generation time of each remaining data packet; Indicates the first A drone The state information of the time step itself, that is, the state of the first time step. A drone The position of the time step and the remaining energy.
[0030] 2) Motion space: Motion space For the first For a single unmanned aerial vehicle (UAV), It is a 2-tuple: ,in, Indicates the use of controlling the first A drone The angle of the motion direction of the time step. Indicates the first A drone The flight distance at the time step, which is the maximum distance As a boundary.
[0031] 3) Reward function: The first A drone The expression for the environmental reward at each time step is given by formula (7): (7).
[0032] In this invention, thorough exploration is crucial for AoI threshold-aware UAV data acquisition tasks because the long decision-making cycle leads to an exponentially increasing search space. Therefore, this invention introduces a composite intrinsic reward calculated from spatial diversity to help UAVs more effectively cooperate in exploring the environment. This invention extracts two types of spatial diversity as intrinsic rewards. Intuitively, UAVs need to visit different locations to prevent the AoI from occasionally exceeding the threshold at a particular PoI. Furthermore, the distance between UAVs should be avoided to prevent wasting data acquisition resources. Since UAVs can only access a portion of their own observations (i.e., their own state), this invention utilizes… Encouraging location diversity through spatial k-nearest neighbors (KNN) and employing This allows us to construct several recent PoI states to predict the likelihood of a drone collecting data at the current PoI. In this way, we can fully leverage regional diversity and encourage a more direct division of workload among drones. Therefore, the... A drone The inherent reward of the time step includes the drone's performance. Time step location diversity reward and the drone in Regional cooperation rewards for time steps .
[0033] Specifically, location diversity is encouraged by designing an intrinsic reward based on KNN. A context buffer can be used to store the state information of each drone at each time step, thus enabling the optimization of the KNN-based reward system for each time step. For individual drones, up to At time step, the state information of the drone itself stored in the scenario buffer is { , , ..., },in, , , ..., Then it is the first A drone The historical state information of the time step itself, and, Representing the A drone The position and remaining energy of the time step are considered similarly, and so on. Therefore, the first... A drone Time step location diversity reward The expression is: (8); in, Indicates the first A drone The set of positions of all time steps preceding the current time step. Represents a set Any position in Expressing the request Position and location in The Euclidean distance between them This indicates the absolute value. Therefore, the location diversity reward... Encourage drones to visit places they haven't been to recently, as these places may contain PoIs with AoI exceeding the threshold.
[0034] Specifically, effective division of labor is a direct and efficient way for drones to cooperate in space. Each drone should be responsible for a different area, maximizing collection capabilities while minimizing redundant movement. To leverage this regional diversity, this invention uses a predictor to output a predictor for the area being covered. State belongs to the first The probability of a drone And take this probability as Part of the intrinsic reward of the time step, therefore, the first A drone Regional cooperation rewards for time steps The expression is: (9); in, express The ground data acquisition node it represents belongs to the first The probability of a drone. This represents the network parameters of the predictor. If the drones are in historical... The differences in state are only slight, so the probability This will result in low entropy, leading to low rewards for regional cooperation. To maximize expected future rewards, drones need to learn to achieve spatial division of labor when regional rewards provide positive feedback. To obtain reasonable probabilities, an information-theoretic objective is used to maximize the mutual information between the PoI state and the drone's identity, specifically as follows: (10), where, Indicates the amount of information. Indicates the first A drone With the Mutual information exists between the identities of the individual drones. Maximizing this mutual information encourages drones to discover new trajectory patterns because if drone cooperation is easier to identify, it becomes easier to infer which drones are responsible for a given observation. Therefore, the loss function of the predictor is expressed as follows: (11); in, Indicates the loss value. Represents the cross-entropy loss function. This represents the heat function.
[0035] In this invention, to control the degree of spatial exploration during training, an exploration adapter is added. This adapter controls the extent of spatial exploration. At the beginning of the training phase, the invention aims for a sufficiently large intrinsic reward to help the UAV discover interesting locations; however, as training progresses, the underlying DNN becomes familiar with the environment, thus reducing the potential impact of the intrinsic reward. In this invention, the exploration adapter includes two embedding networks. One embedding network has fixed parameters and serves as the target embedding network (called the Target Network). The other embedding network's parameters need to be updated during training (called the Predictor Network) to ultimately fit its own network parameters to those of the target embedding network. Specifically, the Target Network is a neural network with a fixed structure and weights frozen after initialization. It takes the state as input and outputs a fixed-dimensional feature vector (e.g., 64-dimensional). The Predictor Network has a similar structure to the Target Network, but its parameters are trainable, and its goal is to fit the output of the Target Network by minimizing the mean squared error (MSE). When the drone encounters a state it has never seen before, the Predictor Network cannot accurately predict the output of the Target Network, resulting in a large prediction error and thus generating a high intrinsic reward, incentivizing the drone to explore. Conversely, for common states, the Predictor Network has learned to predict the output of the Target Network better, with a smaller error, and the intrinsic reward is correspondingly lower, naturally guiding the agent to prioritize exploring unknown or information-rich areas. For example, each embedded network also adopts a DNN network architecture.
[0036] In this invention, the exploration adapter is used to determine the input based on the first... A drone The state at time step generates the first... A drone The degree of spatial exploration in time step ,in, The expression is as follows: (12); in, The running average, i.e. yes The average value, This refers to an embedded network with fixed network parameters, i.e., a target network. Indicates the output of the Target Network. This represents an embedding network with network parameters that need to be trained, i.e., a Predictor Network. This represents the output of PredictorNetwork.
[0037] In this invention, the expression for the loss function of the exploration adapter is as follows: (13); in, Indicates the loss value. This indicates the operation of seeking the expected value.
[0038] Therefore, in this invention, during each training process, the first... A drone The expression for the total reward at each time step is: (14); in, Indicates the first A drone is executing After the time step action, the result is Total reward for time step Indicates the weighting coefficient. The value can be set according to actual needs, and this invention does not limit it.
[0039] Refer to the above Figure 2 In this invention, the principle of jointly training the Actor network, Critic network, predictor, and explorer adapter deployed on each UAV in the UAV system is as follows: 1) Initialize the network parameters of the policy network, the network parameters of the value network, the network parameters of the two embedded networks, and the network parameters of the predictor for each UAV. At the same time, establish an experience replay buffer pool to store the interaction data between the UAV and the environment (i.e., experience data).
[0040] 2) During training, the process is executed cyclically, using time steps as the unit. For example... Figure 2 As shown, within each time slot (i.e., each time step), each UAV, based on the current environmental observation information (i.e., ), through local policy network Select consecutive actions (i.e., obtain) ), and interact with the environment (i.e., execute) ), receive instant rewards (i.e., get ) and the observed state at the next moment (i.e., obtaining Subsequently, combining the probability output by the predictor and the spatial exploration degree output by the exploration adapter, the immediate reward is calculated according to the above formula (14). (exist Figure 2The Chinese character is represented as The adjustments are made to obtain a comprehensive reward (i.e., the reward is...). ,exist Figure 2 middle, Represented as , Represented as , Represented as , Represented as ), and will include the current state, action, reward, and next state (i.e. , , , These, etc., are stored as interactive experiences in the experience replay buffer pool.
[0041] 3) When the number of samples (i.e., empirical data) in the experience buffer pool reaches the update condition, a small batch of samples is randomly sampled from the buffer pool for network parameter updates. Specifically, the loss function of the value network (i.e., Critic network) and the loss function of the policy network (i.e., Actor network) corresponding to each UAV are calculated, and all losses are weighted and summed. The parameters of the policy network and value network are updated using the gradient descent method, thereby realizing centralized training and distributed execution of multiple UAVs. It should be noted that the loss function of the value network and the loss function of the policy network adopt existing loss functions, and this invention does not limit this. At the same time, the network parameters of the predictor are updated according to the above formulas (11) and (13). The algorithm also updates the network parameters of the Predictor Network in the adapter, gradually bringing it closer to the output of the Target Network. This dynamically adjusts the intrinsic reward intensity, guiding the UAV to continuously explore the state space that has not yet been fully explored. This process iterates until the algorithm converges or reaches the preset number of training rounds, ultimately yielding an optimized multi-UAV cooperative trajectory control strategy.
[0042] For example, the flow of the training method of the present invention is shown in Table 1 below: Table 1
[0043] To demonstrate the technical effects of this invention, the following shows its optimization effect on drone crowdsourcing scenarios, and compares it with the traditional MATD3 effect. Specifically, Figure 3Figures (a) and (b) in the figure are schematic diagrams of the trajectories of the simulated UAVs of RC-MATD3 (representing the present invention) and MATD3, respectively. Calculations based on these figures show that the acquisition coverage of RC-MATD3 is 89.34%, while that of MATD3 is 72.95%, representing an improvement of approximately 23%. Figure 4 This is a comparison chart of the simulation benefits of the two algorithms, from... Figure 4 It can be seen that the efficiency of RC-MATD3 is improved by approximately 58%. Clearly, the experiments demonstrate that the method proposed in this invention significantly improves coverage and efficiency compared to MATD3.
[0044] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0045] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0046] In this specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. While different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce a good effect.
[0047] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A method for UAV data acquisition based on a prediction-enhanced deep reinforcement learning network, characterized in that, The method, applied to each drone in an unmanned aerial vehicle (UAV) system, includes: The drone acquires its current state and inputs it into its own pre-trained policy network, which then outputs the current action. The drone performs the action at the current moment to collect data from the corresponding ground data acquisition node; The trained policy network is obtained by using reinforcement learning to jointly train the Actor network, Critic network, predictor, and explorer adapter deployed on each UAV in the UAV system. The predictor and the explorer adapter assist in the training of the Actor network and the Critic network. The outputs of the predictor and the explorer adapter are used to calculate the adaptive intrinsic reward of the UAV during the training process. The adaptive intrinsic reward is used to guide the UAV to perform location exploration and regional cooperation.
2. The UAV data acquisition method based on prediction-enhanced deep reinforcement learning network according to claim 1, characterized in that, The current state includes: the state information of the ground data acquisition nodes within the current perception range, and the state information of the UAV itself at the current moment. The state information of the ground data acquisition nodes within the current perception range includes: the position of all ground data acquisition nodes within the current perception range of the UAV, the remaining amount of data, and the data generation time of each remaining data packet. The state information of the UAV itself at the current moment includes: the position of the UAV at the current moment and the remaining energy. The actions at the current moment include: the angle used to control the direction of movement of the drone at the current moment, and the flight distance of the drone at the current moment.
3. The UAV data acquisition method based on prediction-enhanced deep reinforcement learning network according to claim 2, characterized in that, The predictor is a neural network, and the predictor is used to predict based on the input of any drone. The state information of the ground data acquisition nodes within the sensing range of the time step outputs a prediction probability, where the prediction probability represents the... The probability that a ground data acquisition node within the perception range of a time step belongs to any one of the UAVs.
4. The UAV data acquisition method based on prediction-enhanced deep reinforcement learning network according to claim 3, characterized in that, The expression for the loss function of the predictor is as follows: ; in, Indicates the loss value. Represents the cross-entropy loss function. Represents the heat function. Indicates the first in the unmanned aerial vehicle system The drone, the first "each drone" represents any one of the aforementioned drones. The value ranges from 1 to , This indicates the total number of drones in the unmanned aerial vehicle system. Indicates the first A drone Status information of ground data acquisition nodes within the sensing range of the time step. express The ground data acquisition node it represents belongs to the first The probability of a drone.
5. The UAV data acquisition method based on prediction-enhanced deep reinforcement learning network according to claim 2, characterized in that, The exploration adapter is used to detect any one of the drones in the unmanned aerial vehicle system based on the input. The state of the time step generates the state of any one of the drones in The degree of spatial exploration at each time step, wherein any one of the drones is in The expression for the degree of spatial exploration at a time step is as follows: ; in, Indicates the first in the unmanned aerial vehicle system A drone The degree of spatial exploration in time step The value ranges from 1 to , This indicates the total number of drones in the unmanned aerial vehicle (UAV) system. Indicates the first A drone The state of the time step, This is the running average. This indicates an embedded network with fixed network parameters in the exploration adapter. This indicates the embedded network in the exploration adapter that has network parameters that need to be trained.
6. The UAV data acquisition method based on prediction-enhanced deep reinforcement learning network according to claim 5, characterized in that, The expression for the loss function of the exploration adapter is as follows: ; in, Indicates the loss value. This indicates the operation of seeking the expected value.
7. The UAV data acquisition method based on prediction-enhanced deep reinforcement learning network according to claim 2, characterized in that, During the training process, any one of the drones in the unmanned aerial vehicle system is performing... After the time step action, the result is The total reward for a time step includes: Environmental rewards for time steps and The adaptive intrinsic reward of the time step, wherein, for any one of the drones The expression for the environmental reward at each time step is as follows: ; ; in, Indicates the first A drone Environmental rewards for each time step, the first "One drone" refers to any one drone in the drone system. The value ranges from 1 to , This indicates the total number of drones in the unmanned aerial vehicle system. This represents the total number of ground data acquisition nodes on the ground. , This represents the total number of steps in each training session. This indicates the time taken at each time step. , Representing drones from the first The moment when data packets are collected at each ground data acquisition node , Indicates the first The generation time of the earliest data packet generated among the ground data acquisition nodes. The value ranges from 1 to , Indicates the first The weight of each ground data acquisition node, This represents a preset constant term. This indicates the preset AOI threshold. Indicates the first A drone The time step takes into account energy consumption or penalties for hitting obstacles.
8. The UAV data acquisition method based on prediction-enhanced deep reinforcement learning network according to claim 7, characterized in that, any one of the drones The expression for the adaptive intrinsic reward at each time step is as follows: ; ; ; in, Represents any one of the drones Adaptive intrinsic reward at time step This indicates that any one of the drones is in The degree of spatial exploration in time step This indicates that any one of the drones is in Time step location diversity reward This indicates that any one of the drones is in Regional cooperation rewards for time steps Represents any one of the drones The state of the time step, express Included The state information of the time step itself. express Included Status information of ground data acquisition nodes within the sensing range of the time step. This indicates that any one of the drones is in The set of positions of all time steps preceding the current time step. Represents a set Any position in Expressing the request Position and location in The Euclidean distance between them This indicates finding the absolute value. express The ground data acquisition node it represents belongs to the first The probability of a drone. This represents the network parameters of the predictor.
9. The UAV data acquisition method based on prediction-enhanced deep reinforcement learning network according to claim 8, characterized in that, Any one of the drones is executing After the time step action, the result is The expression for the total reward of a time step is as follows: ; in, This indicates that any one of the drones is performing... After the time step action, the result is Total reward for time step This represents the weighting coefficient.
10. The UAV data acquisition method based on prediction-enhanced deep reinforcement learning network according to claim 1, characterized in that, The application scenarios of the unmanned aerial vehicle (UAV) system include: inspection scenarios, emergency response scenarios, and mobile sensing scenarios.