Optimized control method, device and medium for drone information collection

By building a robust reinforcement learning model, the information collection of drones in a masked environment is optimized, and the information collection optimization problem in drone position/trajectory control is solved, and efficient information collection is achieved in the case of position error, which is suitable for a variety of application scenarios.

CN116859718BActive Publication Date: 2025-09-02NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310653250.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-05
Publication Date
2025-09-02
Estimated Expiration
2043-06-05

AI Technical Summary

Technical Problem

In the collection of drone information, the prior art cannot effectively solve the problem of drone position/trajectory control, especially when the drone cannot accurately obtain its own position information under environmental shading, information collection optimization cannot be achieved.

Method used

By obtaining the dynamic measurement information amount of the drone in the presence of position estimation error, a robust reinforcement learning model is built. The optimization goal is to maximize the collection of information amounts. The optimal control strategy in robust reinforcement learning is used to solve the problem, and the expected cumulative information amount model of the drone in the worst case is generated, and Q-value solutions are performed to obtain the optimal control strategy.

Benefits of technology

When there is error in drone information in a masked environment, it can quickly and effectively collect information, adapt to a variety of scenarios and environments, and has strong expansion and expansion. It is suitable for application scenarios such as drone coverage communication and relay communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116859718B_ABST
    Figure CN116859718B_ABST
Patent Text Reader

Abstract

The present application discloses an optimization control method, device and medium for unmanned aerial vehicle (UAV) information collection, which includes: defining the information collection method of the UAV, establishing a set of UAV flight actions, and constructing an optimization control problem for UAV information collection; obtaining the action information and location information taken by the UAV at a certain moment, and constructing a model for the total amount of information collected by the UAV; constructing a sequential optimization decision model based on the model for the total amount of information collected by the UAV, and modeling the UAV information collection optimization control problem as a problem of finding the optimal strategy for the sequential optimization decision problem; and reconstructing the sequential decision problem as a steady-state reinforcement learning problem and establishing a robust reinforcement learning model; based on the robust reinforcement learning model, using a robust Q-value learning algorithm to solve the problem, and obtaining the optimal control strategy for information collection, which can collect information quickly and effectively, adapt to various scenarios and environments, and have good scalability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of drone communication technology, and more specifically, to an optimized control method, device, and medium for drone information collection. Background Art

[0002] Drones, due to their flexible deployment and low cost, are widely used in military and civilian communications, and hold broad application prospects. Drone communications primarily take the form of coverage communications, relay communications, and information collection. Using drones for information collection offers new technical means and implementation options, with applications in mobile communications, the Internet of Things, and other fields. Furthermore, drones' controllable location and trajectory greatly facilitate information collection.

[0003] Numerous research results have been achieved in leveraging the controllable position and trajectory of drones for information collection, primarily in scenarios where information is known and unknown. In the known case, the drone information collection problem is primarily modeled as an optimization problem (primarily non-convex). Given information such as the drone's position and trajectory and the locations of the nodes being collected, the amount of information collected and energy consumption are used as optimization objectives, and non-convex optimization methods are employed to control the drone's position and trajectory for optimal performance. In the unknown case, drone information collection primarily employs methods such as reinforcement learning, which uses heuristic learning to control the drone's position and trajectory based on limited environmental and reward information to achieve optimal performance.

[0004] When using drones for information collection, the location of the nodes being collected is often unknown, necessitating the use of methods such as reinforcement learning to adjust their position and trajectory using limited environmental and reward information. Furthermore, drones often face environmental obstructions (such as tall buildings in urban environments, trees in jungle environments, and structures indoors or underground) that prevent the drone from accurately acquiring its own location information.

[0005] Existing learning-based methods for controlling the position and trajectory of drones require accurate knowledge of the drone's state information, including its position. When the acquired drone position information contains errors, if the position information of the nodes from which the information is collected is available, the drone's position error can be modeled and the drone's position and trajectory optimized based on this error model. However, when the position information of the collected nodes is unavailable and the drone's position is inaccurate, controlling the drone's position and trajectory through learning to optimize information collection presents a significant challenge, and no publicly available methods exist. Summary of the Invention

[0006] In response to at least one defect or improvement need in the prior art, the present invention provides an optimization control method for drone information collection, characterized by comprising:

[0007] Acquire dynamic measurement information of the UAV in the presence of position estimation error, the dynamic measurement information including the measurement spatial position of the UAV and the state of the transition within the spatial position;

[0008] Maximizing the amount of information collected by the drone about the measurement space position within a given time is taken as the optimization goal, and converting the solution of the optimization goal into the solution of the optimal control strategy in robust reinforcement learning;

[0009] Constructing a model of the expected cumulative amount of information collected by the drone within a given time based on the amount of dynamic measurement information;

[0010] An optimal control strategy solution model is constructed based on the expected cumulative information model, and the Q value is solved to obtain the control strategy corresponding to the maximum total amount of information that the UAV can collect when there is a position estimation error.

[0011] Furthermore, the step of constructing a model of the expected cumulative amount of information collected by the drone during the information collection time based on the dynamic measurement information includes the following steps:

[0012] Obtaining a flight state space set of the UAV and establishing a set of flight actions that can be taken in the flight state space;

[0013] Construct multiple state transition sets of the drone;

[0014] Obtain the instant reward collected when the drone takes a certain action in the current state to transfer to the next state. The instant reward is the normalized sum of the accumulated node information collected by the drone in the current state, generating an instant reward set;

[0015] An expected cumulative information volume model is generated based on the flight state space set, the flight action set, the multiple state transition sets, and the immediate reward set.

[0016] Furthermore, the step of obtaining a plurality of state transition sets of the drone includes the following steps:

[0017] Construct the state transition probability variable of the UAV and generate the randomness measurement factor;

[0018] Based on the state transition probability variable and the randomness measurement factor, a state transition probability matrix is ​​generated to obtain multiple state transition sets.

[0019] Furthermore, the expected cumulative information volume is specifically: V π (s) = minP E P {M};

[0020] Wherein, the E P {M} is the state of the drone in multiple transitions. In the case of P E P {M} is the expected cumulative amount of information collected from nodes in the worst case corresponding to multiple different state transitions P; M is the cumulative amount of information;

[0021] π is the number of times the UAV is in the tth (t=0,1,…,T max -1) is in area S at the moment t Take action A t is defined as the control strategy of the UAV, where T max The total number of moments of information collection for the drone;

[0022] S is the area that the drone can reach. This area can be any area in three-dimensional space or two-dimensional space. The area S is divided into I smaller areas s that do not overlap with each other. i (i=1,2,…,I), and S={s1,s2,…,s I}. S t is the state S at time t t , S t ∈S={s1,s2,…,s I}. The drone is in area s i When you are in the area, you can take different actions to move to other areas;

[0023] A indicates that the drone is in area S t The set of actions that can be taken; A t is the action taken by the drone at time t, A t ∈A={a1,a2,…,a M}; P is the set of multiple state transitions performed by the drone; In robust reinforcement learning, when L state transitions are executed, the state transition is recorded as the realization of the L state transition probability matrix, that is, (P1, P2, ..., P L ),in

[0024] Furthermore, the state transition probability matrix is ​​generated based on the state transition probability variable combined with the randomness measurement factor as follows:

[0025]

[0026] Where R is the randomness measurement factor, and X is the uncertainty vector of the state transition vector.

[0027] Furthermore, the multiple state transition sets are specifically:

[0028]

[0029] in, In robust reinforcement learning, when performing L state transitions, the state transition probability matrix is ​​realized once for each state transition.

[0030] Furthermore, the cumulative amount of information collected by the UAV under a given control strategy π and initial state S0 = s is specifically:

[0031]

[0032] Where π is the number of times the UAV is in area S at time t (t=1,2,…) t Take action A t , defined as the control strategy of the UAV.

[0033] Furthermore, the control strategy corresponding to the maximum value of the total amount of information that can be collected by the UAV in the presence of position estimation error by solving the Q value includes:

[0034] Step 1: Initialize the decision iteration time step T max ,Q0(s,a), a∈A, behavior strategy π b , initial state S0, step size α t and R;

[0035] Step 2: Starting from time t=0, according to the behavior strategy π b Choose the first action A t And execute, enter the second state S t When obtaining the signal-to-noise ratio information of different nodes ρ(S t ,Node k ) k = 1, 2, ..., K and store;

[0036] Step 3: Get the third state S t+1 , and calculate the immediate reward c(S t ,A t );

[0037] Step 4: For any s∈S, take V t (s)=max a∈A Q t (s,a);

[0038] Step 5:

[0039] Step 6: For all (s,a)≠(S t ,A t ), Q t+1 (s,a)←Q t (s,a);

[0040] Step 7: Execute t=t+1. If t≤T max , then go to step 1; otherwise, go to step 8.

[0041] Step 8: Output

[0042] Step 9: Select the corresponding state s The action a corresponding to the maximum value is the optimal strategy π * .

[0043] According to a second aspect of the present invention, an electronic device is also provided, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the processing unit is enabled to perform the steps of any one of the above methods.

[0044] According to a third aspect of the present invention, a storage medium is provided, which stores a computer program executable by an access authentication device. When the computer program runs on the access authentication device, the access authentication device can perform the steps of any of the above methods.

[0045] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0046] The present invention provides an optimization control method, device and medium for drone information collection. By acquiring the dynamic information amount of the drone under a given control strategy and state and establishing a robust state value function, a model of the expected cumulative information amount of nodes that the drone may collect under the worst case scenario is constructed; a corresponding robust action value function is generated through the expected cumulative information amount model, and the solution is performed to obtain the control strategy corresponding to the maximum value of the total amount of information that can be collected under the worst case scenario, so that the drone can quickly and effectively collect information even when there are errors in the drone information in a shielded environment, and can adapt to a variety of scenarios and environments, and at the same time has strong scalability. The present invention proposes a universal modeling method and solution algorithm, which can be applied to application scenarios such as drone coverage communication and relay communication. At the same time, its optimization target can also be changed to indicators including factors such as communication quality and energy consumption, so it has strong scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0048] Figure 1 A flowchart of an optimization control method for collecting drone information provided in an embodiment of the present application;

[0049] Figure 2 A block diagram of an electronic device suitable for implementing the above-described method provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0050] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0051] The terms "first," "second," "third," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0052] Existing learning-based methods for controlling drone position and trajectory require accurate knowledge of the drone's state, including its position. When the acquired drone position information contains errors, if the position information of the nodes being collected is available, the drone's position error can be modeled and optimized based on this error model. However, when the collected node position information is unavailable and the drone's position contains errors, controlling the drone's position and trajectory through learning to optimize information collection presents a significant challenge, and currently no publicly available methods exist.

[0053] This invention addresses the problem of optimizing the control strategy for drone information collection when a drone is unable to obtain its own accurate position information in obscured environmental conditions. This problem is addressed by using heuristic learning (reinforcement learning) methods to transform the original problem from another perspective, obtaining an optimized control strategy for the drone when estimation errors exist. Specifically, rather than attempting to solve a control strategy that maximizes the amount of information that can be collected in the presence of position estimation errors, the invention considers the minimum information that can be obtained each time in the presence of position estimation errors, attempting to find the optimal strategy for the entire information collection process. This technology then uses heuristic learning to obtain environmental and reward information to achieve drone position / trajectory control, thus enabling efficient information collection when estimation errors exist in the drone's position information.

[0054] Specifically, Figure 1 A flowchart of an optimization control method for collecting drone information provided by an embodiment of the present invention is shown as follows: Figure 1 As shown, step S110: obtaining dynamic measurement information of the UAV in the presence of position estimation error; the dynamic measurement information includes the measurement space position (or state) of the UAV and the state of transfer in the space position;

[0055] Step S120: maximizing the amount of information collected by the drone about the measurement space position within a given time is used as an optimization goal, and converting the solution to the optimization goal into the solution to the optimal control strategy in robust reinforcement learning;

[0056] In one embodiment, when a drone collects information in a shielded environment, it obtains its own position information by estimation, and the position information has estimation errors. The total time of information collection is represented as T, and there are K nodes whose information is collected, distributed on the ground, represented as Node k (k=1,2,…,K). The area that the drone can reach is represented as the state space S, which can be any area in three-dimensional space or two-dimensional space. The area S is divided into I smaller areas s that do not overlap with each other. i (i=1,2,…,I), and S={s1,s2,…,s I}.

[0057] At the same time, the drone is in area s i When A={a1,a2,…,a M} indicates that the drone is in area s i The set of actions that can be taken, the action taken by the drone at time t is represented by A t (A t ∈A={a1,a2,…,a M}).

[0058] When collecting information, place the drone in area s i Time and ground node Node k The received signal-to-noise ratio is expressed as ρ(s i ,Node k ), the signal-to-noise ratio includes the influence of factors such as the relative position between the UAV and the ground node, channel fading, etc., and the divided area is small enough so that ρ(s i ,Node k ) can be ignored.

[0059] Therefore, the drone is in area s i Node k The amount of information collected can be expressed as Δt*log2(1+ρ(s i ,Node k )), where Δt is the UAV’s position in area s i The time the drone stays or flies in any area is Δt (without loss of generality, assume T / Δt is an integer). Therefore, the drone can stay or fly for a total of T times in the entire information collection time T. max =T / Δt areas, that is, the drone can fly T max A moment. t (t=0,1,…,T max -1) represents the area where the drone is located at the tth moment, then S t ∈S={s1,s2,…,s I}.

[0060] Furthermore, the basic process of drone information collection is as follows: the drone is in a certain area s i When , you can take action a i (a i ∈A={a1,a2,…,a M}), move to area s j (i=j or i≠j), receive signals from each node, and then in area s j Stay or fly for Δt time and calculate the amount of information received from each node. j Node k The amount of information collected is Δt*log2(1+ρ(s j ,Node k )); Then, the drone is in area s j When the action set A={a1,a2,…,a M} take an action a j, move to the next area and stay or fly for Δt time, and collect information; the above process is repeated until the drone reaches the upper limit of information collection time T.

[0061] A t (A t ∈A={a1,a2,…,a M}) and S t (S t ∈S={s1,s2,…,s I}) represent the action taken by the drone at the tth moment and the area where the drone is located. Therefore, the total amount of information collected by the drone during the entire time T is:

[0062]

[0063] Among them, the inner summation term Indicates that the drone is in area S at time t t The sum of the information collected from all K nodes in time T; the external summation term is the sum of the total number of T nodes that the drone has experienced in time T. max It is worth noting that the areas that the drone has visited within time T may be repeated.

[0064] From the above description, we can see that the area that the drone passes through in time T is related to the initial area that the drone is in, and is related to the actions taken in each area (i.e., control strategy). t Take action A t Defined as the control strategy of the UAV, denoted as π(A t |S t )(abbreviated as π), and the control strategy is random, that is, ∑ At∈A π(A t |S t )=1.

[0065] Therefore, the optimization control problem of UAV information collection can be modeled as follows: Obtain T max The control strategy problem of maximizing the expected value of the sum of the information collected at each moment is to find the optimal strategy for the following sequential optimization decision problem:

[0066]

[0067] Among them, E π {·} is the expected value of the policy π.

[0068] However, the present invention considers that the UAV does not know the location information of the node, that is, it does not know the signal-to-noise ratio (i.e., ρ(S)) between each node and the UAV at all times. t ,Node k)) information, the signal-to-noise ratio (i.e. ρ(S t ,Node k )) information. Therefore, it is impossible to optimize formula (1) to obtain the optimal control strategy before the drone flies. At the same time, the present invention considers that in practice, due to the existence of estimation errors, that is, S t The true value of cannot be obtained, so the above formula cannot be solved in practice. To solve these problems, the present invention provides an optimization control method for solving the above-mentioned problem of optimizing control of drone information collection.

[0069] It is worth noting that the present invention considers that in the absence of external positioning system support, the UAV obtains its own position information (i.e., the corresponding area s) through the position estimation module. i (i=1,2,…,I)), but the estimated position information has errors. In other words, when there is a position estimation error, the drone does not know the actual area it is currently in.

[0070] Therefore, in the embodiment of the present invention, the drone is estimated to obtain the area s i As its state variable (the corresponding actual real area is s p ), when taking action a i After moving, the actual area it is in is transferred to s q (The corresponding estimated area is s j ).

[0071] Generally, the area estimated by the drone is a random variable around its actual area. p Take action a m Transfer to the actual area q When the UAV is estimated (or perceived) in the area is s i Transfer to s j Considering that there is an estimation error, it is impossible for the drone to know its real location, and s j It is around s q A random variable, then s i Transfer to position s j is random (i.e. the estimated area s j is random).

[0072] That is to say, in the embodiment of the present invention, the actual area where the drone is located is s p is an estimated value, which has a certain error compared to the actual area. iWhen moved, its actual location is shifted to area s p At the same time, the estimated area is also determined by s i Transferred to a random variable s j Because of the estimation error, the drone cannot accurately know its true position, and s j is around area s p random variable, so in the region s p Transfer to s j When , the estimated area of ​​the drone will also be a random variable.

[0073] Furthermore, the present invention adds a randomness measurement factor to the state transition probability variable of the drone under a given control strategy, and constructs a state transition probability matrix based on the state transition probability variable combined with the randomness measurement factor, and defines multiple state transition sets. A robust state value function is established through multiple state transition sets, and then a model of the expected cumulative information amount of nodes that the drone may collect in the worst case of state transition is constructed.

[0074] Step S130: constructing a model of the expected cumulative amount of information collected by the drone during the information collection time based on the amount of dynamic measurement information;

[0075] Furthermore, constructing a model of the expected cumulative amount of information collected by the drone during the information collection time based on the dynamic measurement information amount includes the following steps:

[0076] S131: Acquire a flight state space set of the UAV and establish a set of flight actions that can be taken in the flight state space;

[0077] Preferably, the state space is defined as a region S, and S={s1,s2,…,s I}, where s i (i∈{1,2,…,I}) is the state of the drone through position estimation (that is, the state perceived by the drone, which is not necessarily the actual state). The state of the drone at time t is represented by S t , and there is S t ∈S={s1,s2,…,s I}, t∈{0,1,…,T max -1}. Define action space A={a1,a2,…,a M}, the action taken by the drone at time t is represented by A t (A t ∈A={a1,a2,…,a M}).

[0078] S132: Construct multiple state transition sets of the drone;

[0079] Preferably, in robust reinforcement learning, when executing e state transitions, the state transitions can be summarized as the e state transition probability matrix The realization of (P1, P2, ..., P e ), Generate state transition probability variables at the same time in, is a set of probability vectors; the state transition probability matrix after adding the randomness measurement factor R is expressed as: Where R is the randomness measurement factor, and X is the uncertainty vector of the state transition vector. For the convenience of expression, multiple state transitions in robust reinforcement learning are expressed as That is, the multiple transition probability matrix The set of realizations (P1, P2, ...).

[0080] S133: Obtaining the instant reward collected when the drone takes a certain action in the current state to transfer to the next state. The instant reward is the normalized sum of the accumulated node information collected by the drone in the current state, and generating an instant reward set;

[0081] Preferably, the UAV is at the tth (t=0,1,…,T max -1) Take action A t Transfer to state S t+1 , then the instant reward is the drone in state S t The normalization of the sum of all node information collected at the time is:

[0082]

[0083] in, The drone is in state S t The maximum value of the signal-to-noise ratio of the K nodes received at the time, then 0≤c(S t ,A t )≤1, and the immediate reward set is represented by c={c(S t ,A t )|S t ∈S,A t ∈A,t∈{0,1,…,T max -1}}. At the same time, the discount factor γ satisfies γ∈[0,1).

[0084] S134: Generate an expected cumulative information model based on the flight state space set, the flight action set, the multiple state transition set, and the instant reward set.

[0085] Preferably, the robust state value function under a given strategy π and state s (s∈S) is the formula for the expected cumulative amount of information that can be collected in the worst case:

[0086] V π (s) = min P E P {M}

[0087] wherein, wherein, said E P {M} is the state of the drone in multiple transitions. In the case of P E P {M} is the expected cumulative amount of information collected from nodes in the worst case corresponding to multiple different state transitions P; M is the cumulative amount of information;

[0088] P is a set of mathematical expression models for the UAV to implement multiple state transitions;

[0089] Furthermore, in one embodiment, the formula for the cumulative amount of information collected by the drone under a given control strategy and state is:

[0090]

[0091] Among them, π is the UAV in area S at the tth moment t Take action A t Defined as the control strategy for the UAV.

[0092] Therefore, the formula for the expected cumulative information that can be collected in the worst case under a given policy π and state s (s∈S) is:

[0093]

[0094] Furthermore, the corresponding robust action value function formula generated by the expected cumulative information model is:

[0095]

[0096] Furthermore, the sequential optimization decision problem of formula (1) can be reconstructed as the problem of finding the optimal control strategy in the robust reinforcement learning problem (S, A, P, c, γ), that is, finding the following optimal strategy:

[0097]

[0098] It should be noted that the optimization in formula (5) is to find the control strategy corresponding to the maximum amount of information that can be collected in the worst case of state transition when the drone has position estimation error, while formula (1) is to find the control strategy corresponding to the maximum amount of information when the drone has position estimation error. The optimal control strategies corresponding to formula (5) and formula (1) may be different. When the state in formula (1) is understood as a state with estimation error, the optimal control strategies corresponding to formula (1) and formula (5) are equivalent.

[0099] Step S140: constructing an optimal control strategy solution model based on the expected cumulative information volume model, performing Q value solution to obtain a control strategy corresponding to the maximum total amount of information that can be collected by the UAV in the presence of position estimation error;

[0100] In one embodiment, for the robust reinforcement learning problem in formula (5), π * represents the optimal control strategy for this problem, then we have This embodiment proposes to use a robust Q-value learning algorithm to find the optimal control strategy for the robust reinforcement learning problem (i.e., formula (5)), that is, to obtain the optimal control strategy by solving the Q value. The proposed robust Q-value learning algorithm is as follows:

[0101] Initialization: Decision iteration time step T max ,Q0(s,a), a∈A, behavior strategy π b , initial state S0, step size α t and R

[0102] 1: for t=0,1,…,T max -1 do

[0103] 2: Based on behavioral strategy π b (·|S t ) Select action A t And execute, in state S t When the UAV obtains the signal-to-noise ratio information ρ(S t ,Node k )k=1,2,…,K and store

[0104] 3: Get the next state S t+1 , and calculate the instant reward c(S t ,A t )

[0105] 4: For any s∈S, take V t (s)=max a∈A Q t (s,a)

[0106] 5:

[0107] 6: For all (s,a)≠(S t ,A t ), Q t+1 (s,a)←Q t (s,a)

[0108] 7: end for

[0109] Output:

[0110] It can be seen that when the step size α t satisfy and When T max →∞, Converges to the Q value corresponding to the optimal control strategy with probability 1. At the same time, according to the convergence Determine the optimal control strategy π * (i.e. based on Choose the action a that is largest for any state s The action a when the value is the control strategy), then the optimal control strategy is the optimal control strategy found by formula (5).

[0111] This invention provides an optimization control method for drone information collection, which accounts for position estimation errors in obscured environments and is adaptable to a variety of scenarios and environments. This method establishes a state-value function and a robust action-value function, and utilizes an expected cumulative information model to calculate the total amount of information that a drone can collect under worst-case scenarios. Ultimately, using a universal modeling approach and solution algorithm, a control strategy that maximizes the amount of information is found. Furthermore, this method can be applied to other scenarios and can consider different optimization objectives, including factors such as communication quality and energy consumption. Therefore, this method is highly practical and scalable.

[0112] Figure 2 The block diagram schematically shows an electronic device suitable for implementing the method described above according to an embodiment of the present invention. Figure 2 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0113] like Figure 2As shown, the electronic device 1000 described in this embodiment includes: a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage part 1008 into a random access memory (RAM) 1003. The processor 1001 may, for example, include a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 1001 may also include an onboard memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0114] Various programs and data required for the operation of the system 1000 are stored in the RAM 1003. The processor 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. The processor 1001 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 1002 and / or RAM 1003. It should be noted that the programs may also be stored in one or more memories other than the ROM 1002 and RAM 1003. The processor 1001 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.

[0115] According to an embodiment of the present disclosure, electronic device 1000 may further include an input / output (I / O) interface 1005, which is also connected to bus 1004. System 1000 may also include one or more of the following components connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or a modem. Communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in drive 1010 as needed, so that computer programs read therefrom can be installed into storage section 1008 as needed.

[0116] The method flow according to the embodiment of the present disclosure can be implemented as a computer software program. For example, the embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to the embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.

[0117] Embodiments of the present invention further provide a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the methods according to the embodiments of the present disclosure.

[0118] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In an embodiment of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the ROM 1002 and / or RAM 1003 described above.

[0119] It should be noted that the functional modules in the various embodiments of the present invention can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product.

[0120] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0121] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of the present disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in the present disclosure. In particular, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of the present disclosure, and all such combinations and / or couplings fall within the scope of the present disclosure.

[0122] Although the present disclosure has been shown and described with reference to specific exemplary embodiments of the present disclosure, it should be understood by those skilled in the art that various changes in form and detail may be made to the present disclosure without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents. Therefore, the scope of the present disclosure should not be limited to the above-mentioned embodiments, but should be determined not only by the appended claims, but also by the equivalents of the appended claims. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0123] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An optimization control method for drone information collection, characterized in that: include: Acquire dynamic measurement information of the UAV in the presence of position estimation error, the dynamic measurement information including the measurement spatial position of the UAV and the state of the transition within the spatial position; Maximizing the amount of information collected by the drone about the measurement space position within a given time is taken as the optimization goal, and converting the solution of the optimization goal into the solution of the optimal control strategy in robust reinforcement learning; Constructing a model of the expected cumulative amount of information collected by the drone within a given time based on the amount of dynamic measurement information; An optimal control strategy solution model is constructed based on the expected cumulative information model, and the Q value is solved to obtain the control strategy corresponding to the maximum amount of information that can be collected by the UAV in the presence of position estimation error; The control strategy corresponding to the maximum value of the total amount of information that can be collected by the UAV in the presence of position estimation error by solving the Q value includes: Step 1: Initialize the decision iteration time step , , behavioral strategy , initial state , step length and ; Step 2: From At the beginning of the moment, according to the behavioral strategy Select the first action And execute, enter the second state Get the signal-to-noise ratio information of different nodes and store; Step 3: Get the third state , and calculate the instant reward obtained ; Step 4: For any ,Pick ; Step 5: ; Step 6: For all , ; Step 7: Execution ;like , then go to step 1; otherwise, go to step 8; Step 8: Output ; Step 9: Select any state hour The action corresponding to the maximum value , which is the optimal strategy ;in, The area that the drone can reach can be any area in three-dimensional space or two-dimensional space. Divided into smaller, non-overlapping regions ; and there is ; For the moment Status , ; The drone is in the area When you are in the area, you can take different actions to move to other areas; Indicates that the drone is in the area The set of actions that can be taken; For drones at all times The actions taken, .

2. The optimization control method for drone information collection according to claim 1, characterized in that: The method of constructing a model of the expected cumulative information amount collected by the drone during the information collection time based on the dynamic measurement information amount includes the following steps: Obtaining a flight state space set of the UAV and establishing a set of flight actions that can be taken in the flight state space; Construct multiple state transition sets of the drone; Obtain the instant reward collected when the drone takes a certain action in the current state to transfer to the next state. The instant reward is the normalized sum of the accumulated node information collected by the drone in the current state, generating an instant reward set; An expected cumulative information volume model is generated based on the flight state space set, the flight action set, the multiple state transition sets, and the immediate reward set.

3. The optimization control method for drone information collection according to claim 2, characterized in that: The method of obtaining multiple state transition sets of the drone includes the following steps: Construct the state transition probability variable of the UAV and generate the randomness measurement factor; Based on the state transition probability variable and the randomness measurement factor, a state transition probability matrix is ​​generated to obtain multiple state transition sets.

4. The optimization control method for drone information collection according to claim 3, characterized in that: The expected cumulative information amount is specifically: ; Among them, the The drone transitions to multiple states In the case of , the expected cumulative information amount is For different multiple state transitions The expected cumulative amount of information collected from the nodes in the worst case; is the accumulated information amount; For drones In the area at any moment Take action is defined as the control strategy of the UAV, where, The total number of moments of collecting information for the drone, ; The area that the drone can reach can be any area in three-dimensional space or two-dimensional space. Divided into smaller, non-overlapping regions , and there are ; For the moment Status , ; The drone is in the area When you are in the area, you can take different actions to move to other areas; Indicates that the drone is in the area The set of actions that can be taken; For drones at all times The actions taken, ; P is the set of multiple state transitions executed by the drone; For robust reinforcement learning, execution When the state transition occurs, the state transition is recorded as The realization of the sub-state transition probability matrix, that is, ,in .

5. The optimization control method for drone information collection according to claim 4, characterized in that: The state transition probability matrix is ​​generated based on the state transition probability variable combined with the randomness measurement factor as follows: ; Among them, R is the randomness measurement factor, is the state transition vector uncertainty vector.

6. The optimization control method for drone information collection according to claim 5, characterized in that: The multiple state transition sets are specifically: ; in, For robust reinforcement learning, execution The state transition probability matrix is ​​realized once for each state transition.

7. The optimization control method for collecting drone information according to claim 6, characterized in that: The UAV has a given control strategy and the initial state The cumulative amount of information collected is as follows: ; in, For drones In the area at any moment Take action , defined as the control strategy of the UAV, .

8. An electronic device, characterized in that: The method comprises at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the processing unit is enabled to perform the steps of the method according to any one of claims 1 to 7.

9. A storage medium, characterized in that: It stores a computer program executable by an access authentication device. When the computer program runs on the access authentication device, the access authentication device is enabled to execute the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for deploying unmanned aerial vehicle emergency communication system in post-disaster area

    CN112333767A

  • Unmanned aerial vehicle information collection control method and system

    CN114117633A