Safety cooperative control method and system for realizing mobile crowd perception by using unmanned aerial vehicle
By constructing a multi-UAV mobile crowd sensing model and an improved MADDPG algorithm, combined with the ERNIE and Stackelberg game models, the problem of UAV scheduling deviation is solved, and stable and robust control of the UAV system in complex environments is achieved.
Patent Information
- Application Number
- CN202411914160.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-12-24
AI Technical Summary
Existing drone scheduling algorithms have difficulty achieving accurate scheduling in the face of environmental factors and malicious attacks, resulting in scheduling deviations and affecting the safety and robustness of drones.
A multi-UAV mobile crowd perception model is constructed. Through the improved multi-agent deep deterministic policy gradient algorithm (MADDPG), combined with the knowledge fusion enhanced language representation ERNIE model and the Stackelberg game model, gradient descent and policy update are performed to enhance the stability and robustness of UAVs in complex environments.
In the presence of scheduling deviations, the stability and robustness of the UAV system are ensured, the impact of scheduling deviations is reduced, and the control effect of the UAV in uncertain environments is improved.
Smart Images

Figure CN119759058B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of unmanned aerial vehicle scheduling, and particularly relates to a safe cooperative control method and system for realizing mobile crowd sensing by unmanned aerial vehicles. BACKGROUND
[0002] Mobile crowd sensing (MCS) based on intelligent devices has become an extremely attractive paradigm. With the development of 5G and above technologies, real-time applications of unmanned aerial vehicles (UAVs) become possible, through the use of UAVs as base stations (BSs) in the air to move and collect data from multiple users. Existing UAV scheduling algorithms all achieve global optimal solutions on the basis of precise scheduling of UAVs. In actual situations, due to the influence of various factors of the environment, it is often difficult to achieve precise scheduling of UAVs, and even individual UAVs may have a large deviation under malicious attacks. Research shows that attackers will modify or inject false data to destroy the accuracy of scheduling decisions, causing the scheduling of UAVs to deviate, thereby affecting the work of the next stage of UAVs. In this case, a robust control method needs to be designed to alleviate the impact of scheduling deviation.
[0003] The existing deep reinforcement learning-based solutions do not consider the bias of unmanned aerial vehicle control and attacks on the trajectory of the unmanned aerial vehicle, which will lead to a decline in algorithm efficiency and even threaten flight safety. For example, the patent application with the publication number CN110806756B proposes a DDPG (Deep Deterministic Policy Gradient) algorithm-based autonomous guidance control of unmanned aerial vehicles, focusing on improving the autonomy and task execution efficiency of unmanned aerial vehicles using deep learning. As another example, the patent application with the publication number CN115729258A proposes to model the unmanned aerial vehicle scheduling problem as a constrained cooperative Markov game, design state, action, reward and loss functions, and solve the optimal scheduling strategy to maximize the perception revenue of the unmanned aerial vehicle under the constraint of limited resources such as battery charging budget. However, these methods do not consider that attackers will modify or inject false data to damage the accuracy of scheduling decisions, thereby causing deviations in unmanned aerial vehicle scheduling. The patent application with the publication number CN116931543A proposes several key improvements to MADDPG (Multi-Agent Deep Deterministic Policy Gradient) to improve the safety and robustness of multi-unmanned aerial vehicle data collection, especially for physical perception inconsistency (PSI) and position attack protection. However, this method does not have sufficient protection measures when facing complex environmental changes or malicious attacks. SUMMARY
[0004] The present application aims to solve the problems in the prior art and provide a safe cooperative control method and system for unmanned aerial vehicle implementation of mobile crowd perception, which alleviates the safety and robustness problems caused by cooperative control bias of unmanned aerial vehicles.
[0005] To achieve the above-mentioned purpose, the present application has the following technical solutions:
[0006] In a first aspect, a safe cooperative control method for unmanned aerial vehicle implementation of mobile crowd perception is provided, comprising:
[0007] Constructing a multi-unmanned aerial vehicle mobile crowd perception model;
[0008] Initializing the state of the multi-unmanned aerial vehicle, setting a random action exploration process, selecting the initial action of the multi-unmanned aerial vehicle, and initializing the critic network and the policy network;
[0009] Calculate rewards based on the multi-UAV mobile crowd perception model, build an experience pool, and update the status of multiple UAVs;
[0010] Draw samples from the experience pool, calculate the expected return of multi-drone data collection, and update the critic network based on the expected return;
[0011] Construct the ERNIE model and Stackelberg game model based on knowledge fusion enhancement, perform gradient descent based on the gradient calculated by the Stackelberg game model, and solve the problem of minimizing the gradient descent result;
[0012] Update target policy network;
[0013] After the target policy network is updated, different perturbations are added to observe the fluctuations in the reward value under different perturbations. If the convergence conditions are met, the current policy is saved; otherwise, samples are drawn from the experience pool again.
[0014] As a preferred solution, the multi-UAV mobile crowd perception model includes a system model and a threat model. The system model includes a multi-UAV flight model, a communication model and an energy consumption model. The multi-UAV flight model defines two attributes of multi-UAV flight: angle and speed. The communication model includes the distance from multiple UAVs to ground users, channel gain, signal-to-noise ratio and unloading rate. The energy consumption model includes communication power, parasitic power, induction power and delay. The threat model is a modeling of attackers.
[0015] As a preferred solution, the construction process of the multi-UAV flight model includes:
[0016] Let U{U|U=1,2,…,U}, M{M|M=1,2,…,M} represent the UAV and mobile user in the 3D target area respectively, and set their coordinates as In addition, there is a height h b high-rise building B{B|B=1,2,…,B}, when h b ≥h u When the drone is in the same building as the building, it needs to avoid the corresponding building;
[0017] In each time slot [t, t+1), each UAV u moves at a speed In a certain direction Movement τ time, v max is the maximum speed of the UAV; each mobile user m starts from Move to And collect data separately, given the smart device equipped with the expected sampling frequency is
[0018] The construction process of the communication model includes:
[0019] The path loss, PL, model for Line of Sight, LoS, and Non-Line of Sight, NLoS, links is as follows:
[0020]
[0021] where a LoS , b LoS , a NLoS , b NLoS are environmental parameters on the floating intercept and slope, is the three-dimensional distance between the mobile user m and the UAV u;
[0022] When a user uploads data to a UAV, the surrounding buildings and other users can block the LoS transmission path. The LoS probability for the time period [t, t+1) is calculated as follows:
[0023]
[0024] where, is the Euclidean distance between the user m and the UAV u, where the UAV u is located at height h u , and each user is modeled as a cylinder with average height h uscr and average diameter g user ; h device is the height at which the user carrying the smart device is located; and l is the density of blockers. It is assumed that the density of mobile users follows a Poisson distribution with parameter l and the mobile users carrying the smart device are located at height h device , the average path loss is:
[0025]
[0026] Let the maximum coupling loss, MCL, be the maximum loss that the system can tolerate at the conducted power level and still operate;
[0027] The construction process of the energy consumption model includes:
[0028] The construction of the user residual data and the Aoi model includes the following steps:
[0029] At the time point [t, t+1), the user m attempts to upload all residual data to the nearest UAV u, and if the PL t (u, m) is tolerable, it means that the upload is successful, and the residual data of the user m is updated by the following expression:
[0030]
[0031] where G Tx and G RxThe gain of the Tx and Rx antennas respectively, and finally the data acquisition amount of each user is:
[0032]
[0033] The update expression of user Aoi is:
[0034]
[0035] The energy consumption of each time slot of the UAV is:
[0036]
[0037] In the formula, c1, c2, and c3 are constants, which depend on the weight of the UAV, the rotor, the blade, and the air density; v tip and The tip speed and average speed of the rotor respectively; is the flight speed of the UAV u at time t; and τ is the moving time;
[0038] When constructing the threat model, the attacker destroys the accuracy of the scheduling decision by modifying or injecting false data.
[0039] As a preferred scheme, in the steps of initializing the state of the multiple UAVs, setting a random action exploration process, selecting the initial action of the multiple UAVs, and initializing the critic network and the policy network, according to a = φ (o) + X k selecting the action of the UAV, wherein a is the action of the UAV, φ is the deterministic policy in the multi-agent deep deterministic policy gradient MADDPG algorithm, and X k is the random action exploration process of the UAV at time slot k.
[0040] As a preferred scheme, in the steps of calculating the reward based on the multiple UAV mobile crowd perception model, establishing the experience pool, and updating the state of the multiple UAVs, by calculating the reward, wherein:
[0041] represents the punishment when the UAV u collides with an obstacle or runs out of energy; C collect and C AoI are constants; and M is the number of users.
[0042] The experience pool contains the old state, action, reward, and new state, and (s, r, a, s') is written into the experience pool F. The experience pool F contains (s', s, a1,..., a u ,..., a U ), and records the experience of all UAVs.
[0043] The old state s is replaced by a new state s' of the next time slot.
[0044] As a preferred solution, the step of drawing a sample from the experience pool, calculating the expected return of multi-UAV data collection, and updating the critic network according to the expected return comprises:
[0045] The expected return gradient of multi-UAV data collection is calculated according to the following formula:
[0046]
[0047] In the formula, o is the observation state of the UAV, w contains all the observation results of the UAV, {θ1,..., θ U} is the strategy set of the multi-UAV, φ u is the strategy of the UAV u under the deterministic strategy, A φ (w, a1,..., a u ) is the centralized action value function under the deterministic strategy φ, containing all the UAV actions and state information; a is the action given by the current strategy, and F is the experience pool.
[0048] The specific way of updating the critic network is:
[0049] L(θ) = E F [(A φ (w, a1,..., a U )-b 2 )]
[0050] b = r + ζA φ′ (w', a'1,..., a' U )| a′=φ′ (o)
[0051] In the formula, ζ represents a discount factor, 0 < ζ < 1; L(θ) is a loss function, b is a target value of an advantage function, r is a reward value; w' is the next state reached after performing an action in the current state; a' = φ'(o) represents the optimal action output by the strategy network in the next state according to the strategy.
[0052] As a preferred solution, the ERNIE model based on knowledge fusion enhancement is constructed in the following manner:
[0053]
[0054] The Stackelberg game model is constructed in the following manner:
[0055]
[0056] U θis an update operation based on gradient ascent, which is used to gradually maximize the divergence of the perturbation and the original observation value, where denotes the operator combination Specifically as follows:
[0057]
[0058] The Stackelberg game model updates the gradient in the following manner:
[0059]
[0060] According to Update the target network;
[0061] In the formula, o k is the observation value of the kth agent; θ represents the parameters of the policy; δ K (o, θ) is the perturbation value generated by θ after K-step update; is the policy of the kth agent; D is a distance function used to measure the difference between the outputs of two policies; α, η are learning rates.
[0062] As a preferred scheme, when different perturbations are added after the target policy network is updated, the perturbation values are set to 0, 0.1, 0.2 and 0.3 respectively; the reward value fluctuation under different action perturbations is observed, and if the reward value fluctuation is less than 0.1, the convergence is satisfied.
[0063] In the second aspect, a safety cooperative control system for realizing mobile crowd perception of a UAV is provided, comprising:
[0064] A perception model construction module is configured to construct a multi-UAV mobile crowd perception model;
[0065] An initialization module is configured to initialize the state of the multi-UAV, set a random action exploration process, select the initial action of the multi-UAV, and initialize the critic network and the policy network;
[0066] An experience pool establishment module is configured to calculate the reward based on the multi-UAV mobile crowd perception model, establish an experience pool, and update the state of the multi-UAV;
[0067] An expected return calculation module is configured to extract samples from the experience pool, calculate the expected return of data collection of the multi-UAV, and update the critic network according to the expected return;
[0068] A gradient descent solving module is configured to construct an ERNIE model based on knowledge fusion enhancement and a Stackelberg game model, perform gradient descent through the gradient calculated by the Stackelberg game model, and solve the result of minimizing the gradient descent;
[0069] A target policy network updating module is configured to update the target policy network.
[0070] A policy screening module is configured to add different disturbances after the target policy network is updated, observe fluctuations in the reward value under different disturbances, save the current policy if a convergence condition is met, and return to extract samples from the experience pool again if the convergence condition is not met.
[0071] In a third aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the safety cooperative control method for mobile crowd perception by a UAV.
[0072] Compared with the prior art, the present application has at least the following beneficial effects:
[0073] In the crowd sensing technology, when the mobile crowd perception problem is achieved by deploying a UAV, the existing strategy function learned based on a deep reinforcement learning algorithm relies heavily on the precise scheduling of the UAV in the calculation process. In the real environment, the scheduling of the UAV inevitably has various deviations, which destroys the reinforcement learning algorithm, and further affects the scheduling control of the entire system, and even threatens the flight safety. The present application sets a random action exploration process, improves the deterministic policy in the multi-agent deep deterministic policy gradient (MADDPG) algorithm, constructs an enhanced language representation ERNIE model based on knowledge fusion, and performs gradient descent through the gradient calculated by the Stackelberg game model, so as to ensure that a relatively smooth strategy function can be obtained in the presence of scheduling deviations, thereby reducing the influence of the scheduling deviations, and enabling the control method to remain stable and robust in an uncertain environment. BRIEF DESCRIPTION OF DRAWINGS
[0074] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0075] Figure 1 The safety cooperative control method for mobile crowd perception by a UAV according to the embodiments of the present application is shown in the flowchart.
[0076] Figure 2 The application scenario of the safety cooperative control method for mobile crowd perception by a UAV according to the embodiments of the present application is shown in the schematic diagram. DETAILED DESCRIPTION
[0077] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, other embodiments can be obtained by those skilled in the art without creative effort.
[0078] Please refer to Figure 1 、 2 The embodiment of the present application proposes a safety cooperative control method for realizing mobile crowd perception by a UAV, which is used to alleviate the safety and robustness problems caused by cooperative control deviation of the UAV, and mainly includes the following steps:
[0079] S1, constructing a multi-UAV mobile crowd perception model;
[0080] S2, initializing the state of the multi-UAV, setting a random action exploration process, selecting the initial action of the multi-UAV, initializing the critic network and the policy network;
[0081] S3, calculating the reward based on the multi-UAV mobile crowd perception model, establishing an experience pool, and updating the state of the multi-UAV;
[0082] S4, extracting samples from the experience pool, calculating the expected return of the multi-UAV data collection, and updating the critic network according to the expected return;
[0083] S5, constructing a language representation ERNIE model based on knowledge fusion enhancement and a Stackelberg game model, performing gradient descent through the gradient calculated by the Stackelberg game model, and solving the result of minimizing the gradient descent;
[0084] S6, updating the target policy network;
[0085] S7, after updating the target policy network, adding different disturbances, observing the fluctuation of the reward value under different disturbances, if the convergence condition is met, saving the current policy, otherwise returning to extract samples from the experience pool.
[0086] In one possible implementation, the multi-UAV mobile crowd perception model of step S1 includes a system model and a threat model. The system model includes a multi-UAV flight model, a communication model and an energy consumption model, and includes two roles: a UAV cluster and a ground user. The multi-UAV flight model defines two attributes of the multi-UAV flight: angle and speed. The communication model includes the distance from the multi-UAV to the ground user, the channel gain, the signal-to-noise ratio and the offloading rate. The energy consumption model includes the communication power, the parasitic power, the induction power and the time delay. The threat model is a modeling of an attacker.
[0087] Further, the construction process of the multi-UAV flight model includes:
[0088] Let U{U|U=1,2,…,U}, M{M|M=1,2,…,M} represent the UAV and mobile user in the 3D target area respectively, and set their coordinates as In addition, there is a height h b high-rise building B{B|B=1,2,…,B}, when h b ≥h u When the drone is in the same building as the building, it needs to avoid the corresponding building;
[0089] In each time slot [t, t+1), each UAV u moves at a speed In a certain direction Movement τ time, v max is the maximum speed of the UAV; each mobile user m starts from Move to And collect data separately, given the smart device equipped with the expected sampling frequency is
[0090] The construction process of the communication model includes:
[0091] The path loss PL model for line-of-sight LoS and non-line-of-sight NLoS links is as follows:
[0092]
[0093] Where α LoS , β LoS , α NLoS , β NLoS are the environmental parameters on the floating intercept and slope, is the three-dimensional distance between mobile user m and drone u;
[0094] When a user uploads data to a drone, surrounding buildings and other users may block the LoS transmission path. The LoS probability for the period [t, t+1) is calculated as follows:
[0095]
[0096] Where, is the Euclidean distance between user m and drone u, where drone u is at height h u , and each user is modeled as having an average height h uscr and average diameter g user Cylinder; h device is the height of the user carrying the smart device; λ is the density of the blocker; assuming that the density of mobile users follows the Poisson distribution with parameter λ and the mobile user carrying the smart device is located at h device Height, so the average path loss is:
[0097]
[0098] Let the maximum coupling loss MCL be the maximum loss that the system can tolerate at the conducted power level and still operate.
[0099] The construction process of the energy consumption model includes:
[0100] The construction of the user residual data and the Aoi model includes the following steps:
[0101] At the time point [t, t+1), the user m attempts to upload all residual data to the nearest UAV u, if PL t (u, m) is tolerable, it means that the upload is successful, and the residual data of the user m is updated by the following expression:
[0102]
[0103] In the formula, G Tx and G Rx are the gains of the Tx and Rx antennas, and finally the data acquisition amount of each user is:
[0104]
[0105] The update expression of the user Aoi is:
[0106]
[0107] The energy consumption of the UAV in each time slot is:
[0108]
[0109] In the formula, c1, c2, and c3 are constants, which depend on the weight, rotor, blade, and air density of the UAV; v tip and are the tip speed and average speed of the rotor, respectively; is the flight speed of the UAV u at time t; and τ is the moving time.
[0110] When constructing the threat model, the attacker destroys the accuracy of the scheduling decision by modifying or injecting false data.
[0111] In a possible implementation, step S2 selects the UAV action according to a = φ (o) + X k , in which a is the UAV action, φ is the deterministic policy in the multi-agent deep deterministic policy gradient MADDPG algorithm, and X k is the random action exploration process of the UAV at time slot k.
[0112] In a possible implementation, step S3 selects the UAV action by Compute the reward, where: represents the penalty when the drone u hits an obstacle or runs out of energy; C collect with C AoI is a constant; M is the number of users.
[0113] The experience pool contains the old state, action, reward, and new state, and writes (s, r, a, s') into the experience pool F, which contains (s', s, a1,..., a u ,..., a U ), recording the experience of all drones;
[0114] The old state s is replaced with the new state s' of the next time slot.
[0115] In one possible implementation, step S4 calculates the expected return gradient of multi-drone data collection according to the following formula:
[0116]
[0117] In the formula, o is the observation state of the drone, w contains all the observation results of the drones, {θ1,..., θ U} is the strategy set of the multi-drone, φ u is the strategy of the drone u under the deterministic strategy, A φ (w, a1,..., a u ) is the centralized action value function under the deterministic strategy φ, containing all the drone action and state information; a is the action given by the current strategy, and F is the experience pool.
[0118] The specific way to update the critic network is:
[0119] L(θ) = E F [A φ (w, a1,..., a U ) - b 2 ]
[0120] b = r + ζA φ′ (w', a'1,..., a' U )| a′=φ′ (o)
[0121] In the formula, ζ represents the discount factor, 0 < ζ < 1; L(θ) is the loss function, b is the target value of the advantage function, r is the reward value; w' is the next state reached after performing the action in the current state; a' = φ'(o) represents the optimal action output by the strategy network in the next state according to the strategy network.
[0122] In a possible implementation, the language representation ERNIE model based on knowledge fusion enhancement in step S5 is constructed as follows:
[0123]
[0124] The Stackelberg game model is constructed as follows:
[0125]
[0126] U θ is an update operation based on gradient ascent, which is used to gradually maximize the divergence between the perturbation and the original observation, where Represents operator combination The details are as follows:
[0127]
[0128] The Stackelberg game model updates the gradient as follows:
[0129]
[0130] according to Update target network;
[0131] Where o k is the observation value of the kth agent; θ represents the parameters of the strategy; δ K (o, θ) is the perturbation value generated by θ after K steps of updating; is the strategy of the kth agent; D is a distance function used to measure the difference between the outputs of two strategies; α, η are learning rates.
[0132] In a possible implementation, step S6 is based on Update the target network.
[0133] In one possible implementation, in step S7, when different disturbances are added after the target policy network is updated, the disturbance values are set to 0, 0.1, 0.2, and 0.3, respectively; the reward value fluctuations under different action disturbances are observed. If the reward value fluctuation is less than 0.1, convergence is satisfied and the current policy is saved. Otherwise, return to step S4 to re-extract samples from the experience pool.
[0134] The safety cooperative control method for realizing mobile crowd perception of the unmanned aerial vehicle according to the embodiment of the application first considers the scheduling bias problem of multiple unmanned aerial vehicles under mobile crowd perception based on deep reinforcement learning, constructs an unmanned aerial vehicle flight model and a communication control model, improves the strategy update of the MADDPG algorithm, so that a smooth strategy function can be obtained, and the influence of the scheduling bias is reduced. By applying the method to the existing unmanned aerial vehicle system, environment-related parameters are designed in the unmanned aerial vehicle flight model and the communication control model, including three-dimensional coordinates of the unmanned aerial vehicle, flight speed, flight angle, link loss and some related evaluation indexes for calculating the reward value. In the process of improving the MADDPG algorithm, a regularizer is added in the strategy search By using the improved strategy update mechanism, it is ensured that the algorithm can still obtain a relatively smooth strategy function in the presence of scheduling bias, thereby reducing the influence of the scheduling bias, so that the algorithm can remain stable and robust in an uncertain environment. Through actual application verification, it is proved that the method is applicable and robust in the existing unmanned aerial vehicle system, and can effectively operate in various complex environments.
[0135] Another embodiment of the application also provides a safety cooperative control system for realizing mobile crowd perception of an unmanned aerial vehicle, comprising:
[0136] A perception model construction module is configured to construct a multiple unmanned aerial vehicle mobile crowd perception model.
[0137] An initialization module is configured to initialize the state of the multiple unmanned aerial vehicles, set a random action exploration process, select the initial action of the multiple unmanned aerial vehicles, initialize the critic network and the strategy network.
[0138] An experience pool establishment module is configured to calculate the reward based on the multiple unmanned aerial vehicle mobile crowd perception model, establish an experience pool, and update the state of the multiple unmanned aerial vehicles.
[0139] An expected return calculation module is configured to extract samples from the experience pool, calculate the expected return of data collection of the multiple unmanned aerial vehicles, and update the critic network according to the expected return.
[0140] A gradient descent solving module is configured to construct an ERNIE model based on knowledge fusion enhancement and a Stackelberg game model, perform gradient descent through the gradient calculated by the Stackelberg game model, and solve the result of minimizing the gradient descent.
[0141] A target strategy network updating module is configured to update the target strategy network.
[0142] A strategy screening module is configured to increase different disturbances after updating the target strategy network, observe the fluctuation of the reward value under different disturbances, save the current strategy if the convergence condition is met, and return to extract samples from the experience pool again if the convergence condition is not met.
[0143] The present application aims at the problem of UAV scheduling deviation, and based on the existing MADDPG algorithm, the problem caused by the UAV scheduling deviation is alleviated by adding an adversarial regularizer in the original algorithm strategy search and minimizing the adversarial regularizer, and the safety and robustness of the multi-UAV cooperative control are effectively improved.
[0144] Another embodiment of the present application also provides an electronic device, comprising:
[0145] The memory stores at least one instruction, and the processor executes the instruction stored in the memory to realize the safe cooperative control method of the UAV for mobile crowd perception.
[0146] Another embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the safe cooperative control method of the UAV for mobile crowd perception.
[0147] For example, the instruction stored in the memory can be divided into one or more modules / units, which are stored in the computer readable storage medium and executed by the processor to complete the safe cooperative control method of the UAV for mobile crowd perception. The one or more modules / units can be a series of computer readable instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the server.
[0148] The electronic device can be a smart phone, a notebook, a palm computer, a cloud server and other computing devices. The electronic device can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the electronic device can further include more or less components, or combine certain components, or different components, for example, the electronic device can further include an input / output device, a network access device, a bus, etc.
[0149] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0150] The memory can be an internal storage unit of the server, such as a hard disk or a memory of the server. The memory can also be an external storage device of the server, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory can also include both the internal storage unit and the external storage device of the server. The memory is used to store the computer readable instructions and other programs and data required by the server. The memory can also be used to temporarily store data that has been output or will be output.
[0151] It should be noted that the information interaction and execution process between the above module units are based on the same concept as the method embodiments, and the specific functions and technical effects brought about can be referred to the method embodiments part, which will not be repeated here.
[0152] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of functional units and modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit or module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific name of each functional unit or module is only for convenient distinction, and does not limit the protection scope of the present application. The specific working process of the unit or module in the above system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0153] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium at least includes any entity or device capable of carrying the computer program code to the photographing device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc.
[0154] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in a certain embodiment can be referred to the relevant description of other embodiments.
[0155] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A safe collaborative control method for UAVs to achieve mobile crowd perception, characterized by: include: Construct a multi-UAV mobile crowd perception model; Initialize the states of multiple drones, set up a random action exploration process, select the initial actions of multiple drones, and initialize the critic network and policy network; Calculate rewards based on the multi-UAV mobile crowd perception model, build an experience pool, and update the status of multiple UAVs; Draw samples from the experience pool, calculate the expected return of multi-drone data collection, and update the critic network based on the expected return; Construct the ERNIE model and Stackelberg game model based on knowledge fusion enhancement, perform gradient descent based on the gradient calculated by the Stackelberg game model, and solve the problem of minimizing the gradient descent result; Update target policy network; After the target policy network is updated, different perturbations are added to observe the fluctuations in the reward value under different perturbations. If the convergence conditions are met, the current policy is saved; otherwise, samples are drawn from the experience pool again. The multi-UAV mobile crowd perception model includes a system model and a threat model. The system model includes a multi-UAV flight model, a communication model, and an energy consumption model. The multi-UAV flight model defines two attributes: the angle and speed of multi-UAV flight. The communication model includes the distance from multiple UAVs to ground users, channel gain, signal-to-noise ratio, and offloading rate. The energy consumption model includes communication power, parasitic power, inductive power, and delay. The threat model is a model for modeling attackers. The language representation ERNIE model based on knowledge fusion enhancement is constructed as follows: The Stackelberg game model is constructed as follows: U θ is an update operation based on gradient ascent, which is used to gradually maximize the divergence between the perturbation and the original observation, where Represents operator combination The details are as follows: The Stackelberg game model updates the gradient as follows: according to Update target network; Where o k is the observation value of the kth agent; θ represents the parameters of the strategy; δ K (o, θ) is the perturbation value generated by θ after K steps of updating; is the strategy of the kth agent; D is a distance function used to measure the difference between the outputs of two strategies; α, η are learning rates.
2. The method for safe collaborative control of drones to achieve mobile crowd perception according to claim 1 is characterized in that: The construction process of the multi-UAV flight model includes: Let U{U|U=1,2,…,U}, M{M|M=1,2,…,M} represent the UAV and mobile user in the 3D target area respectively, and set their coordinates as In addition, there is a height h b high-rise building B{B|B=1,2,…,B}, when h b ≥h u When the drone is in the same building as the building, it needs to avoid the corresponding building; In each time slot [t, t+1), each UAV u moves at a speed In a certain direction Movement τ time, v max is the maximum speed of the UAV; each mobile user m starts from Move to And collect data separately, given the smart device equipped with the expected sampling frequency is The construction process of the communication model includes: The path loss PL model for line-of-sight LoS and non-line-of-sight NLoS links is as follows: Where α LoS , β LoS , α NLoS , β NLoS are the environmental parameters on the floating intercept and slope, is the three-dimensional distance between mobile user m and drone u; When a user uploads data to a drone, surrounding buildings and other users may block the LoS transmission path. The LoS probability for the period [t, t+1) is calculated as follows: Where, is the Euclidean distance between user m and drone u, where drone u is at height h u , and each user is modeled as having an average height h user and average diameter g user Cylinder; h device is the height of the user carrying the smart device; λ is the density of the blocker; assuming that the density of mobile users follows the Poisson distribution with parameter λ and the mobile user carrying the smart device is located at h device Height, the average path loss is: Let the maximum coupling loss MCL be the maximum loss that the system can tolerate and still operate at the conducted power level; The construction process of the energy consumption model includes: Constructing user residual data and Aoi model includes the following steps: At the time point [t, t+1), user m tries to upload all remaining data to the nearest drone u. If PL t (u,m) is tolerable, which means the upload is successful, and the remaining data of user m Update with the following expression: Where G Tx and G Rx are the gains of the Tx and Rx antennas respectively. The final data collection amount for each user is: The update expression for user Aoi is: The energy consumption of the drone in each time slot is: Where c1, c2, c3 are constants that depend on the weight of the drone, rotor, blades, and air density; v tip and are the tip speed and average speed of the rotor respectively; is the flight speed of UAV u at time t; τ is the moving time; When building threat models, attackers can disrupt the accuracy of scheduling decisions by modifying or injecting erroneous data.
3. The method for safe collaborative control of drones to achieve mobile crowd perception according to claim 1 is characterized in that: In the steps of initializing the states of multiple drones, setting a random action exploration process, selecting the initial actions of multiple drones, and initializing the critic network and the policy network, according to a=φ(o)+X k Select the drone action, where a is the drone action, φ is the deterministic policy in the multi-agent deep deterministic policy gradient MADDPG algorithm, and X k is the random action exploration process of the UAV in time slot k.
4. The method for safe collaborative control of drones to achieve mobile crowd perception according to claim 1 is characterized in that: In the steps of calculating rewards based on the multi-UAV mobile crowd perception model, establishing an experience pool, and updating the multi-UAV status, Calculate the reward, where: Indicates the penalty when the drone u hits an obstacle or runs out of energy; C collect with C AoI is a constant; M is the number of users; The experience pool contains the old state, action, reward and new state. Write (s, r, a, s′) into the experience pool F. The experience pool F contains (s′, s, a1, ..., a u ,...,a U ), all drone experiences were recorded; The new state s′ for the next time slot replaces the old state s.
5. The method for safe collaborative control of drones to achieve mobile crowd perception according to claim 1 is characterized in that: The steps of extracting samples from the experience pool, calculating the expected return of multi-drone data collection, and updating the critic network based on the expected return include: The expected benefit gradient of multi-UAV data collection is calculated as follows: Where o is the observation state of the UAV, w contains the observation results of all UAVs, {θ1,...,θ U } is the strategy set of multiple drones, φ u is the strategy of drone u under the deterministic strategy, A φ (w,a1,...,a u ) is the centralized action value function under the deterministic strategy φ, which contains all drone actions and state information; a is the action given by the current strategy, and F is the experience pool; The specific way to update the critic network is: L(θ)=E F [(A φ (w, a1,..., a U )-b 2 )] b=r+ζA φ′ (w′,a′1,...,a′ U )| a′=φ′(o) Where ζ represents the discount factor, 0<ζ<1; L(θ) is the loss function, b is the target value of the advantage function, and r is the reward value; w′ is the next state after executing the action in the current state; a′=φ′(o) represents the optimal action output by the policy network in the next state.
6. The method for safe collaborative control of drones to achieve mobile crowd perception according to claim 1 is characterized in that: When adding different disturbances after the target policy network is updated, the disturbance values are set to 0, 0.1, 0.2, and 0.3 respectively; the reward value fluctuations under different action disturbances are observed. If the reward value fluctuation is less than 0.1, convergence is satisfied.
7. A safety collaborative control system for UAVs to achieve mobile crowd perception, characterized by: A method for implementing a safe collaborative control of a drone for mobile crowd perception as claimed in any one of claims 1 to 6, comprising: Perception model construction module, used to construct a multi-UAV mobile crowd perception model; The initialization module is used to initialize the states of multiple drones, set a random action exploration process, select the initial actions of multiple drones, and initialize the critic network and policy network; The experience pool establishment module is used to calculate rewards based on the multi-UAV mobile crowd perception model, establish the experience pool, and update the status of multiple UAVs; An expected return calculation module, which is used to draw samples from the experience pool, calculate the expected return of multi-drone data collection, and update the critic network based on the expected return; The gradient descent solution module is used to build the language representation ERNIE model and Stackelberg game model based on knowledge fusion enhancement, perform gradient descent based on the gradient calculated by the Stackelberg game model, and solve the result of minimizing the gradient descent; A target policy network update module, used to update the target policy network; The strategy screening module adds different perturbations after the target strategy network is updated, observes the fluctuation of the reward value under different perturbations, and saves the current strategy if the convergence conditions are met. Otherwise, it returns to re-sample from the experience pool.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for safe collaborative control of a drone for mobile crowd perception as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
DDPG-based autonomous guidance and control method for unmanned aerial vehicles
CN110806756B
Unmanned aerial vehicle crowd sensing scheduling method, system, equipment and medium
CN115729258A
Safe flight control method, system and equipment for multi-unmanned aerial vehicle data collection and medium
CN116931543A
Unmanned aerial vehicle data acquisition trajectory and user association joint optimization method based on reinforcement learning in wireless network
CN115616906A
Multi-unmanned aerial vehicle base station collaborative coverage path planning method based on deep reinforcement learning
CN116227767A