A reinforcement learning unmanned aerial vehicle network construction system and method based on absolute positioning and relative positioning
By combining reinforcement learning methods of absolute and relative positioning, deployment and movement decisions for UAVs are formulated, solving the dynamic adaptation problem of UAV networking schemes, improving the efficiency and ranging accuracy of UAV networking, and realizing efficient networking of UAVs in dynamic scenarios.
Patent Information
- Application Number
- CN202310998339.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-08
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2043-08-08
AI Technical Summary
Existing drone networking solutions are mostly for static scenarios and lack dynamic adaptive deployment solutions. Furthermore, the relative positioning and ranging technology for drones has insufficient measurement accuracy, which cannot meet the needs of actual dynamic scenarios.
A reinforcement learning approach based on absolute and relative positioning is adopted, combined with genetic and greedy algorithms to make deployment and movement decisions for UAVs. An adaptive extended Kalman filter is used to improve ranging accuracy, enabling autonomous flight and dynamic networking of UAVs.
A dynamic UAV networking scheme is provided, which improves networking efficiency and ranging accuracy, reduces system control load, and has universality and real-time performance.
Smart Images

Figure CN116980911B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a UAV network construction system and method, in particular to a UAV network construction system based on absolute positioning and relative positioning, and belongs to the technical field of UAV network construction. BACKGROUND
[0002] At present, mainstream wireless communication technologies often need to rely on perfect network infrastructure. Because of the mobile flexibility and other characteristics, UAVs can be used as relay nodes for communication services in extreme scenarios, and the development of networking schemes is a current research hotspot. Existing research on UAV networking schemes is mostly based on static scenarios of ground nodes, that is, the ground nodes are static, which does not conform to the actual situation. Therefore, there is a lack of a UAV networking scheme that can adaptively deploy according to the dynamic change of the distribution characteristics of ground nodes; most of the existing research on the mobile behavior of UAV clusters is based on the absolute positioning information of UAVs, which may not meet the requirements in some special scenarios (such as war scenarios), so it is also necessary to study the use of relative positioning information of UAVs to make mobile decisions for UAV clusters; most of the existing research is in a centralized solution mode, and a distributed architecture can be introduced to reduce the total control load of the system; in addition, it is also a promising research direction to give UAVs the ability of autonomous flight, and there is currently a lack of a UAV relative positioning networking scheme that is not limited by the topology of UAV networking and is universal; finally, the measurement accuracy of the UAV relative positioning ranging technology also has room for further improvement.
[0003] In the prior art, application No. 202210494754.1 discloses a maximum coverage deployment method and system for capsule airports and UAV applications, which comprises the following steps: in the given area W, the positions of m capsule airports determined from q candidate positions, n UAVs carried by each capsule airport, and the air-to-ground channel model of the UAV and the user, the optimal fixed height H of the UAV and the maximum coverage radius R of the UAV are solved; each UAV is deployed at the optimal plane position to ensure that each UAV can provide the maximum network service capacity under the condition of covering the most users, and then the height of the UAV is optimized to obtain a more optimal height, thereby reducing power loss. The flowchart is shown in Figure 1
[0004] The method used in the above scheme can maximize the number of users covered by the UAVs, but has the following disadvantages:
[0005] (1) The scenario of the document is a static scenario, which does not conform to the actual situation, and the dynamic movement of UAVs needs to be considered;
[0006] (2) the document is a centralized solution, can consider introducing distributed architecture, thereby reducing the load of the system, therefore, the urgent need for a new solution to solve the above problems. SUMMARY
[0007] The present application is just for the problems existing in the prior art, provide a kind of based on absolute positioning and relative positioning's reinforcement learning unmanned aerial vehicle network construction system and method, in the field of unmanned aerial vehicle networking provides a new unmanned aerial vehicle networking scheme.Method uses the absolute information of internet of things node in network, uses genetic algorithm to solve the deployment decision of unmanned aerial vehicle and completes the coverage of network.Due to the mobile characteristics of internet of things node, every time interval, system will make the deployment decision of unmanned aerial vehicle again.Unmanned aerial vehicle also needs to move according to the latest distribution characteristics of internet of things node, method uses greedy algorithm to solve the movement decision of unmanned aerial vehicle.Therefore, the present application provides a dynamic unmanned aerial vehicle networking scheme.
[0008] In order to achieve the above object, the technical scheme of the present application is as follows, a kind of based on absolute positioning and relative positioning's reinforcement learning unmanned aerial vehicle network construction system, the system includes center side and edge side two parts, center side includes unmanned aerial vehicle deployment decision making module, unmanned aerial vehicle movement decision making module and information transceiver module;Edge side includes unmanned aerial vehicle flight decision making module, unmanned aerial vehicle communication module and ground node communication module.
[0009] In the above scheme, in the edge side, wherein ground node communication module is deployed on ground node, unmanned aerial vehicle communication module and unmanned aerial vehicle flight decision making module are deployed on each unmanned aerial vehicle for edge coverage network, ground node communication module uploads the absolute position information of ground node in edge side, unmanned aerial vehicle communication module receives the unmanned aerial vehicle deployment decision and movement decision instruction transmitted from center side;Unmanned aerial vehicle flight decision making module selects reference point and determines reward expression according to the deployment decision and movement decision of unmanned aerial vehicle, combines self-adapting extended Kalman filter ranging accuracy improvement strategy, uses reinforcement learning method to make flight decision of unmanned aerial vehicle, completes the adaptive flight of unmanned aerial vehicle.
[0010] Three modules in center side are all deployed in cloud computing center, information transceiver module completes the communication between each communication module in edge side, receives the absolute position information of ground node in edge side, returns unmanned aerial vehicle deployment decision and unmanned aerial vehicle movement decision,
[0011] Unmanned aerial vehicle deployment decision making module uses genetic algorithm, according to the distribution characteristics of ground node and the preset number of unmanned aerial vehicles, with the optimization target of maximizing network coverage rate, solves the best deployment position of unmanned aerial vehicle,
[0012] The unmanned aerial vehicle movement decision making module uses a greedy algorithm to solve the movement decision of the unmanned aerial vehicle according to the last deployment position of the unmanned aerial vehicle, with the optimization target of minimizing the total distance of the movement of the unmanned aerial vehicle cluster.
[0013] A reinforcement learning unmanned aerial vehicle network construction method based on absolute positioning and relative positioning, the method comprising the following steps:
[0014] Step 1: The ground node communication module uploads the absolute position information of the ground node to the information transceiver module on the center side,
[0015] Step 2: The center side solves the deployment decision and movement decision of the unmanned aerial vehicle based on the absolute position information of the ground node;
[0016] Step 3: The information transceiver module on the center side returns the deployment and movement decisions of the unmanned aerial vehicle to the unmanned aerial vehicle communication module on the edge side,
[0017] Step 4: The unmanned aerial vehicle flight decision making module uses a reinforcement learning method to make the flight decision of the unmanned aerial vehicle in combination with the adaptive extended Kalman filter ranging accuracy improvement strategy.
[0018] In step 2, the following steps are taken:
[0019] Step 2.1: The information transceiver module on the center side acquires the absolute position information of the ground node,
[0020] Step 2.2: The unmanned aerial vehicle deployment decision making module uses a genetic algorithm to solve the deployment decision of the unmanned aerial vehicle, and the specific steps are as follows,
[0021] Step 2.2.1: Determine the coding rule to adapt the unmanned aerial vehicle network coverage problem to the terminology of the genetic algorithm: an individual / chromosome represents the deployment decision of the unmanned aerial vehicle, a population represents a set of deployment decisions of the unmanned aerial vehicle, and an individual I is represented as: I={<x1,y1>,<x2,y2>,,<x k ,y k >}, where x k represents the absolute horizontal coordinate of the unmanned aerial vehicle k, y k represents the absolute vertical coordinate of the unmanned aerial vehicle k, and k is the preset number of unmanned aerial vehicles,
[0022] Step 2.2.2: Determine the fitness of the individual, as shown in equation (1).
[0023]
[0024] Where N covered represents the number of ground nodes covered by the unmanned aerial vehicle deployment scheme represented by the individual, and N represents the total number of ground nodes,
[0025] Step 2.2.3: Generate an initial population using a random strategy,
[0026] Step 2.2.4: Calculate the fitness of each individual in the population, perform crossover, mutation, and selection operations on the individual, exchange some dimensions of the individual, regenerate values for some dimensions of the individual using a random strategy, and finally, replace some individuals in the evolution population with better fitness than the original population,
[0027] Step 2.2.5: Determine whether the algorithm termination condition is met, if yes, output the optimal individual and enter step 2.3, otherwise, enter step 2.2.4,
[0028] Step 2.3: The UAV movement decision module uses a greedy algorithm to solve the UAV movement decision, the specific steps are as follows,
[0029] Step 2.3.1: Calculate the distance between each UAV in the last deployment decision and each UAV in the current deployment decision, generate a distance pair and sort them from low to high according to the distance, where t represents the time, represents the deployment position of UAV a at t-1, represents the deployment position of UAV b at t, d a,b represents the distance between UAV a at t-1 and UAV b at t,
[0030] Step 2.3.2: Select the element with the smallest distance in the distance pair set, which represents the movement of the last deployed UAV to the current deployment position (i.e., the movement decision of the UAV), then update the distance pair set and delete elements with the same first item as the selected distance pair,
[0031] Step 2.3.3: Determine whether the distance pair set is empty, if yes, output the movement decision of the UAV and enter step 3, otherwise, enter step 2.3.2.
[0032] Step 4: Specifically as follows:
[0033] Step 4.1: Combine the adaptive extended Kalman filter ranging accuracy improvement strategy to train the DQN network on the local UAV;
[0034] Step 4.2: The UAV cluster determines the movement priority according to the movement distance, moves in order according to the movement priority, selects the action with the maximum Q value on the trained DQN network to complete the movement, and finally reaches the specified position to complete the deployment.
[0035] Compared with the prior art, the present application has the following advantages: 1. The purpose of the present application provides a reinforcement learning unmanned aerial vehicle network construction system and method based on absolute positioning and relative positioning, which provides a new unmanned aerial vehicle networking scheme in the field of unmanned aerial vehicle networking. The method uses the absolute information of the Internet of Things nodes in the network to solve the deployment decision of the unmanned aerial vehicle using a genetic algorithm to complete the coverage of the network. Due to the mobile characteristics of the Internet of Things nodes, the system will re-determine the deployment decision of the unmanned aerial vehicle every period of time. Therefore, the unmanned aerial vehicle also needs to move according to the latest distribution characteristics of the Internet of Things nodes. The method uses a greedy algorithm to solve the movement decision of the unmanned aerial vehicle. Therefore, the present application provides a dynamic unmanned aerial vehicle networking scheme;
[0036] 2. The scheme comprehensively utilizes absolute positioning and relative positioning technology, uses absolute positioning technology in the stage of making deployment decisions and movement decisions of the unmanned aerial vehicle, and can guarantee the quality of the unmanned aerial vehicle network construction scheme; uses relative positioning technology in the stage of making flight decisions of the unmanned aerial vehicle, and can greatly reduce the load of the system main control module. The scheme combines the advantages of absolute positioning and relative positioning technology, and has the advantages of strong real-time performance and high networking efficiency;
[0037] 3. In the process of moving the unmanned aerial vehicle, each unmanned aerial vehicle uses a reinforcement learning method to make flight decisions of the unmanned aerial vehicle. The method provides a reinforcement learning method for making flight decisions of the unmanned aerial vehicle without considering the topology structure of the unmanned aerial vehicle network, using relative positioning information of the unmanned aerial vehicle, and has universality;
[0038] 4. Finally, in the process of relative positioning and distance measurement of the unmanned aerial vehicle, the method proposes an adaptive extended Kalman filter distance measurement precision improvement strategy. The strategy uses AEKF to calculate the relative distance between the unmanned aerial vehicles, and can effectively improve the distance measurement precision. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 The flow chart of the prior art maximum coverage deployment method and system for capsule airports and unmanned aerial vehicles;
[0040] Figure 2 The system framework diagram of the reinforcement learning unmanned aerial vehicle network construction system based on absolute positioning and relative positioning;
[0041] Figure 3 The flow chart of the reinforcement learning unmanned aerial vehicle network construction system and method based on absolute positioning and relative positioning;
[0042] Figure 4 Flight decision making scene diagram;
[0043] Figure 5 Average reward curve diagram;
[0044] Figure 6Trajectory test chart
[0045] Figure 7 Scenario diagram of adaptive extended Kalman filter ranging accuracy improvement strategy
[0046] Figure 8 Experimental result diagram A of adaptive extended Kalman filter ranging accuracy improvement strategy
[0047] Figure 9 Experimental result diagram B of adaptive extended Kalman filter ranging accuracy improvement strategy. DETAILED DESCRIPTION
[0048] In order to deepen the understanding of the present application, the present embodiment will be described in detail below with reference to the accompanying drawings.
[0049] Example 1: see Figure 2 A reinforcement learning unmanned aerial vehicle network construction system based on absolute positioning and relative positioning, the system includes two parts of center side and edge side, the center side contains unmanned aerial vehicle deployment decision making module, unmanned aerial vehicle movement decision making module and information transceiver module; the edge side contains unmanned aerial vehicle flight decision making module, unmanned aerial vehicle communication module and ground node communication module. In this scheme, among the edge side, the ground node communication module is deployed on the ground node, the unmanned aerial vehicle communication module and the unmanned aerial vehicle flight decision making module are deployed on each unmanned aerial vehicle used for edge coverage network, the ground node communication module uploads the absolute position information of the ground node on the edge side, the unmanned aerial vehicle communication module receives the unmanned aerial vehicle deployment decision and movement decision instructions transmitted from the center side; the unmanned aerial vehicle flight decision making module selects reference points and determines reward expression according to the deployment decision and movement decision of the unmanned aerial vehicle, uses reinforcement learning method to make flight decision of the unmanned aerial vehicle, and completes adaptive flight of the unmanned aerial vehicle. The three modules in the center side are all deployed in the cloud computing center, the information transceiver module completes communication between each communication module in the edge side, receives the absolute position information of the ground node in the edge side, returns the unmanned aerial vehicle deployment decision and the unmanned aerial vehicle movement decision, the unmanned aerial vehicle deployment decision making module uses genetic algorithm, according to the distribution characteristics of the ground node and the preset number of unmanned aerial vehicles, takes maximizing network coverage rate as the optimization target, solves the best deployment position of the unmanned aerial vehicle, and the unmanned aerial vehicle movement decision making module uses greedy algorithm, according to the last deployment position of the unmanned aerial vehicle, takes minimizing total distance of unmanned aerial vehicle cluster movement as the optimization target, and solves the movement decision of the unmanned aerial vehicle.
[0050] Example 2: see Figures 3-9 A reinforcement learning unmanned aerial vehicle network construction method based on absolute positioning and relative positioning, the method comprises the following steps:
[0051] Step 1: The ground node communication module uploads the absolute position information of the ground node to the information transceiver module on the center side.
[0052] Step 2: The center solves the UAV's deployment and movement decisions based on the absolute position information of the ground nodes;
[0053] Step 3: The information transceiver module on the center side returns the drone deployment and movement decisions to the drone communication module on the edge side.
[0054] Step 4: The UAV flight decision-making module combines the adaptive extended Kalman filter ranging accuracy improvement strategy with the reinforcement learning method to make the UAV flight decision.
[0055] Step 2 is specifically as follows: Step 2.1: The information transceiver module on the center side obtains the absolute position information of the ground node.
[0056] Step 2.2: The UAV deployment decision making module uses genetic algorithm to solve the UAV deployment decision. The specific steps are as follows:
[0057] Step 2.2.1: Determine the encoding rules and adapt the UAV network coverage problem to the genetic algorithm terminology: Individuals / chromosomes represent the deployment decisions of UAVs, and the population represents the set of deployment decisions of UAVs. Individual I is represented as: I = {<x1, y1>, <x2, y2>, <x k ,y k >}, where x k Indicates the absolute horizontal coordinate of UAV k, y k Indicates the absolute vertical coordinate of drone k, where k is the preset number of drones.
[0058] Step 2.2.2: Determine the individual's fitness, as shown in formula (1).
[0059]
[0060] where N covered Indicates the number of ground nodes that the UAV deployment plan represented by this individual can cover, and N represents the total number of ground nodes.
[0061] Step 2.2.3: Generate the initial population using a random strategy.
[0062] Step 2.2.4: Calculate the fitness of each individual in the population and perform crossover, mutation, and selection operations on the individuals. Select two individuals from the original population, swap some of their dimensions, and complete the crossover operation. Then, use a random strategy to regenerate values for some of the dimensions in the individuals, completing the mutation operation. Finally, select some individuals with better fitness in the evolved population to replace the individuals with poor fitness in the original population, completing the selection operation.
[0063] Step 2.2.5: Determine whether the algorithm termination condition is reached, yes, output the optimal individual, enter step 2.3, otherwise enter step 2.2.4,
[0064] Step 2.3: The UAV movement decision module uses the greedy algorithm to solve the movement decision of the UAV, and the specific steps are as follows,
[0065] Step 2.3.1: Calculate the distance between each UAV in the last UAV deployment decision and each UAV in the current UAV deployment decision, and generate a distance pair And sort them from low to high according to the distance size, where t represents the time, Represents the deployment position of UAV a at t-1, Represents the deployment position of UAV b at t, d a,b Represents the distance between UAV a at t-1 and UAV b at t,
[0066] Step 2.3.2: Select the element with the smallest distance in the distance pair set, which represents the movement of the last deployed UAV to the position of the current deployment (i.e. the movement decision of the UAV), and then update the distance pair set by deleting the elements in the distance pair set that are the same as the first element in the selected distance pair,
[0067] Step 2.3.3: Determine whether the distance pair set is empty, empty, output the movement decision of the UAV, enter step 3, otherwise enter step 2.3.2,
[0068] Step 3: The information receiving and sending module on the center side returns the UAV deployment and movement decision to the UAV communication module on the edge side,
[0069] Step 4: The UAV flight decision making module combines the adaptive extended Kalman filter ranging accuracy improvement strategy and uses reinforcement learning method to make flight decision of the UAV, and the specific steps are as follows:
[0070] Step 4.1: Combine the adaptive extended Kalman filter ranging accuracy improvement strategy to train the DQN network locally, and the training process is as follows:
[0071] Step 4.1.1: Define state, as shown in Table 1.
[0072] Table 1 Composition of UAV State
[0073]
[0074] d1, d2, d3 respectively represent the distance between the UAV and the selected three relative positioning reference points, and boundry represents the side length of the square region where the UAV is located. The reference points are selected by random strategy.
[0075] Step 4.1.2: Define action, as shown in Table 2.
[0076] Table 2 Composition of UAV action
[0077]
[0078] Where v refers to the speed of the UAV movement, and Δt refers to the time of each movement of the UAV.
[0079] Step 4.1.3: Randomly select three UAVs as reference points, and use the adaptive extended Kalman filter distance measurement accuracy improvement strategy to measure the relative distance between the UAV and the reference points, as shown in steps 4.1.3.1 to 4.1.3.3.
[0080] Step 4.1.3.1: Control process modeling. As shown in Figure 7 , taking the distance measurement process between a UAV A and a reference point B as an example. , respectively, the true distance between A and B at t-2, t-1, t time, and the cosine value of the angle α in Figure 7 is calculated by means of the cosine law, as shown in equation (2).
[0081]
[0082] The turning angle of the UAV at t-1 time is known, then Figure 7 The angle β can be calculated by equation (3).
[0083]
[0084] Step 4.1.3.2: Model . According to the cosine law, the can be calculated as shown in equation (4), where w<t> is the process noise at t time.
[0085]
[0086] Nonlinear processing is performed on equation (4), as shown in equation (5).
[0087]
[0088] Where d<t-2>, d<t-1>, d<t> respectively refer to the estimated distance between A and B at t-2, t-1, t time.
[0089] Step 4.1.3.3: Use AEKF to calculate d<t>, as shown in Table 3.
[0090] Table 3 AEKF calculation of d<t> process
[0091]
[0092]
[0093] Step 4.1.4: Determine the reward expression. Define four Boolean variables flag, flag1, flag2, flag3, whose meanings are shown in equations (6) to (9) respectively, i.e. flag represents whether the UAV has reached the destination node, flag1 represents whether the distance between the UAV and the reference node 1 at time t is closer to the distance between the destination node and the reference node 1 than the distance at the previous time, flag2 represents whether the distance between the UAV and the reference node 2 at time t is closer to the distance between the destination node and the reference node 2 than the distance at the previous time, and flag3 represents whether the distance between the UAV and the reference node 3 at time t is closer to the distance between the destination node and the reference node 3 than the distance at the previous time.
[0094]
[0095]
[0096]
[0097]
[0098] where d1 <t>, d2 <t>, d3 <t> represent the distances between the UAV and the selected three relative positioning reference nodes at time t, and d'1, d'2, d'3 represent the distances between the destination of the UAV and the three reference nodes, respectively, and the destination of the UAV is obtained by the UAV movement decision.
[0099] The reward is defined according to equations (10) and (11).
[0100]
[0101]
[0102] The definition of the reward in equation (10) will make the UAV eventually tend to select flight actions that make d1 <t>, d2 <t>, d3 <t> closer to d'1, d'2, d'3, and the distances between the destination of the UAV and the three reference nodes are d'1, d'2, d'3, so that the UAV can be trained to fly to the destination. Equation (11) represents whether the UAV has reached the destination at the end of the current episode, and if not, a certain penalty is given, and if so, a reward is given.
[0103] Step 4.1.5: Train the DQN network according to the algorithm pseudo code shown in Table 4.
[0104] Table 4 Training process of DQN network
[0105]
[0106] Step 4.2: The UAV cluster determines the moving priority according to the moving distance, moves in order according to the moving priority, selects the action with the maximum Q value on the trained DQN network to complete the movement, and finally reaches the designated position to complete the deployment.
[0107] The flight decision-making scene diagram of the system is shown in Figure 4 The average reward curve of the UAV training is shown in Figure 5
[0108] After 4000 iterations of the reinforcement learning model, the success rate of the UAV reaching the destination has reached 100%. After 50 tests of the model, the test trajectory is shown in Figure 6
[0109] Figure 8 The filtering effect under the condition of fixed Q=50^2 and R=50^2 is shown. The horizontal axis represents time, and the vertical axis represents distance. The solid line represents the theoretical data calculated according to the coordinates, i.e. the accurate distance value between the two UAVs. The blue scatter points represent the distance value measured directly by the sensor, the yellow scatter points represent the estimated value after EKF processing, and the green scatter points represent the estimated value after AEKF processing. It can be seen that the accuracy of AEKF is greater than that of EKF and the distance value measured directly by the sensor.
[0110] As shown in Figure 9 , DiffEKF represents the mean square error between the estimated data of the contrast EKF and the true data, and DiffAEKF represents the mean square error between the estimated data of the AEKF proposed by the application and the true data. The left vertical coordinate represents the corresponding mean square error, and the right vertical coordinate represents the ratio of DiffAEKF to DiffEKF. It can be seen from Figure 9 that the AEKF proposed by the application can effectively improve the distance estimation accuracy.
[0111] It should be noted that the above embodiments are not intended to limit the scope of protection of the present application, and any equivalent transformation or substitution made on the basis of the above technical solutions falls within the scope of protection of the claims of the present application.
Claims
1. An absolute positioning and relative positioning based reinforcement learning unmanned aerial vehicle network construction system, characterized in that, The system comprises a center side and an edge side, the center side comprising a UAV deployment decision-making module, a UAV movement decision-making module and an information transceiver module; the edge side comprising a UAV flight decision-making module, a UAV communication module and a ground node communication module; The absolute information of the Internet of Things nodes in the network is used to solve the deployment decision of the UAVs by using a genetic algorithm to complete the coverage of the network. Based on the movement characteristics of the Internet of Things nodes, the UAVs need to move according to the latest distribution characteristics of the Internet of Things nodes. A greedy algorithm is used to solve the movement decision of the UAVs. An absolute positioning technology is used in the stages of making the deployment decision and the movement decision of the UAVs, and a relative positioning technology is used in the stage of making the flight decision of the UAVs. In the process of relative positioning and distance measurement of the UAVs, an adaptive extended Kalman filter distance measurement accuracy improvement strategy is adopted. The strategy uses AEKF to calculate the relative distance between the UAVs. The three modules in the center side are deployed in a cloud computing center. The information transceiver module completes the communication between each communication module in the edge side, receives the absolute position information of the ground nodes in the edge side, and returns the UAV deployment decision and the UAV movement decision. The UAV deployment decision-making module uses a genetic algorithm to solve the optimal deployment position of the UAVs according to the distribution characteristics of the ground nodes and the preset number of UAVs, with the optimization target of maximizing the network coverage rate. The UAV movement decision-making module uses a greedy algorithm to solve the movement decision of the UAVs according to the last deployment position of the UAVs, with the optimization target of minimizing the total distance of the UAV cluster movement. 2.The absolute positioning and relative positioning based reinforcement learning UAV network construction system of claim 1, wherein, In the edge side, the ground node communication module is deployed in the ground nodes, the UAV communication module and the UAV flight decision-making module are deployed in each UAV for edge coverage network, the ground node communication module uploads the absolute position information of the ground nodes in the edge side, the UAV communication module receives the UAV deployment decision and movement decision instructions from the center side; the UAV flight decision-making module selects reference points and determines the reward expression according to the UAV deployment decision and movement decision, combines the adaptive extended Kalman filter distance measurement accuracy improvement strategy, and uses a reinforcement learning method to make the flight decision of the UAVs to complete the adaptive flight of the UAVs. 3.A method for constructing a reinforcement learning unmanned aerial vehicle network based on absolute positioning and relative positioning, characterized in that, The system in claim 1 is adopted, and the method comprises the following steps: Step 1: The ground node communication module uploads the absolute position information of the ground nodes to the information transceiver module in the center side, Step 2: The center side solves the deployment decision and the movement decision of the UAVs based on the absolute position information of the ground nodes; Step 3: The information transceiver module in the center side returns the UAV deployment and movement decisions to the UAV communication module in the edge side, Step 4: The UAV flight decision-making module uses a reinforcement learning method to make the flight decision of the UAVs in combination with the adaptive extended Kalman filter distance measurement accuracy improvement strategy. 4.The method of claim 3, wherein, Step 2 is as follows: Step 2.1: The information transceiver module in the center side acquires the absolute position information of the ground nodes, Step 2.2: The UAV deployment decision-making module uses a genetic algorithm to solve the deployment decision of the UAVs, and the specific steps are as follows, Step 2.2.1: Determine the coding rule, adapt the UAV network coverage problem to the genetic algorithm terminology: an individual / chromosome represents the deployment decision of a UAV, a population represents the set of deployment decisions of UAVs, an individual I is represented as: I = { <x1, y1>, <x2, y2>,..., <x k ,y k >}, where x k represents the absolute horizontal coordinate of UAV k, y k represents the absolute vertical coordinate of UAV k, k is the preset number of UAVs, Step 2.2.2: Determine the fitness of each individual, as shown in formula (1), where N covered represents the number of ground nodes that the drone deployment scheme represented by the individual is able to cover, N represents the total number of ground nodes, Step 2.2.3: Generate an initial population using a random strategy, Step 2.2.4: Calculate the fitness of each individual in the population, and perform crossover, mutation, and selection operations on the individuals. Select two individuals from the original population, exchange some dimensions, complete the crossover operation, then randomly generate new values for some dimensions in the individual, complete the mutation operation, and finally, select some individuals with better fitness from the evolved population to replace the individuals with poor fitness in the original population, complete the selection operation, Step 2.2.5: Determine whether the algorithm termination condition is met, if yes, output the optimal individual, go to step 2.3, otherwise go to step 2.2.4, Step 2.3: The UAV movement decision module uses a greedy algorithm to solve the movement decision of the UAV, the specific steps are as follows, Step 2.3.1: Calculate the distance between each UAV in the last UAV deployment decision and each UAV in the current UAV deployment decision, and generate distance pairs and sort them from low to high according to the distance, where t represents the time, represents the deployment position of UAV a at time t-1, represents the deployment position of UAV b at time t, d a,b represents the distance between UAV a at time t-1 and UAV b at time t, Step 2.3.2: Select the element with the smallest distance in the distance pair set, which represents the movement of the last deployed UAV to the current deployment location (i.e., the movement decision of the UAV), then update the distance pair set and delete elements with the same first item as the selected distance pair, Step 2.3.3: Determine whether the distance pair set is empty, if yes, output the movement decision of the UAV, go to step 3, otherwise go to step 2.3.
2. 5.The method of claim 4, wherein, The step 4: as follows: Step 4.1: Combine the adaptive extended Kalman filter ranging accuracy improvement strategy to train the DQN network locally on the UAV; Step 4.2: The UAV cluster determines the movement priority according to the movement distance, moves in order according to the movement priority, selects the action with the maximum Q value on the trained DQN network to complete the movement, and finally reaches the specified location to complete the deployment.
Citation Information
Patent Citations
Maximum coverage deployment method and system for capsule airport and unmanned aerial vehicle application
CN115208454A
Unmanned aerial vehicle data acquisition scheduling system and method based on mobile edge computing
CN116069463A