A radiation source positioning method based on online UAV trajectory planning
Through online drone track planning, five omnidirectional antennas and optimal strategy learning algorithms are used to optimize the drone flight path, solving the problem of the drone platform quickly and accurately locate illegal radiation sources on a large scale, and achieving fast and accurate radiation source positioning.
Patent Information
- Application Number
- CN202310468340.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-26
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-04-26
AI Technical Summary
Existing radiation source positioning technology cannot quickly and effectively locate illegal radiation sources in a short period of time. Especially the drone platform is limited by ground conditions, and cannot quickly plan paths and the operation selection effectiveness is low, making it difficult to ensure the convergence speed of optimized path planning.
Using an online drone track planning method, five omnidirectional antennas are used to receive signals, combined with a random walk model that is minimized for search area and an optimal strategy learning algorithm, the flight path of the drone is optimized and the radiation source is quickly located through the Markov decision-making process and the ε-greedy strategy.
It realizes that the drone quickly plans shorter paths in a short time, effectively avoids the problem of non-convergence of search results, improves the accuracy and speed of radiation source positioning, and can quickly find illegal radiation sources on a large scale.
Smart Images

Figure CN116593962B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of radio monitoring technology, and in particular to a radiation source positioning method based on online unmanned aerial vehicle (UAV) trajectory planning. Background Art
[0002] With the continuous development of radio technology and applications, the number of radio frequency devices is growing exponentially. Criminals are using radio technology and devices to commit crimes in increasingly diverse ways, and the electromagnetic spectrum is becoming increasingly complex. Therefore, quickly and effectively locating illegal signal sources is crucial for timely control of illegal frequency devices and ensuring radio security. Existing radiation source location technologies are limited by ground conditions, making it difficult to quickly plan routes and accurately locate illegal radiation sources in a short period of time.
[0003] In recent years, passive positioning technology based on drone platforms has gradually developed. Drones, due to their high maneuverability and flexible autonomous trajectory planning, have become an excellent choice for searching for radiation sources from ground users. While using drones as carriers and relevant algorithms for trajectory planning is simple, unstable signal strength can cause drones to deviate from the optimal path. Therefore, further algorithm improvements are necessary to maximize resource utilization.
[0004] Patent number CN115686065A discloses a method for controlling dynamic target tracking for unmanned aerial vehicles (UAVs) based on deep reinforcement learning. This method includes the design of a Markov decision process for UAV target tracking, a reward function for UAV target tracking, a targeted deep neural network architecture, training of a velocity command perception controller based on the SAC algorithm, and the use of a UAV dynamic target controller. This end-to-end integrated controller simplifies the UAV dynamic target tracking process, demonstrating robustness, fast real-time response, and adaptability to varying target motion patterns. However, as signal strength fluctuates, the UAV often deviates from its optimal path.
[0005] Patent number CN114337875A discloses a method for optimizing the flight trajectory of a swarm of drones for tracking multiple emitters. The method comprises a building module, an estimation module, a matching module, a positioning module, and a tracking module. The building module is used to solve the trajectory optimization problem for a swarm of drones under multiple constraints. The estimation module uses a deep neural network to map received signal strength and distance. The matching module uses an interactive matrix generation method to generate matching solutions between drones and emitters. The positioning module uses a multi-sphere intersection positioning method to determine the reference positions of the emitters. The tracking module employs deep reinforcement learning to design a flight trajectory optimization algorithm for the swarm of drones. Compared to traditional methods, the proposed method offers significant advantages in terms of average tracking time, task completion rate, and convergence speed. However, the drones cannot quickly traverse the entire system, their action selection is ineffective, and the speed of finding emitters is slow, making it difficult to ensure convergence of the optimized path planning. Summary of the Invention
[0006] The present invention discloses a radiation source positioning method based on online UAV trajectory planning, which can enable a UAV that autonomously plans its trajectory online to quickly plan a shorter path, accurately locate the position of the radiation source in a short time, and promptly find illegal radiation sources.
[0007] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:
[0008] A radiation source positioning method based on online UAV trajectory planning, the radiation source positioning method comprising the following steps:
[0009] S1: Control a drone equipped with five omnidirectional antennas at a set altitude, and the drone receives surrounding radiation signals through the five omnidirectional antennas. The five omnidirectional antennas are placed on the drone in a clockwise direction, and the distance from each omnidirectional antenna to the center of mass of the drone is the same;
[0010] S2, using a random walk model for minimizing the search area to make the drone search the entire world. The step size of the random walk model is specified as the ratio of the normal random parameter u to the square of the standard normal random parameter v, so that the variance of the random walk model shows an exponential relationship with time, and the flight path shows the characteristics of alternating long and short.
[0011] S3, for the process of the drone searching for the radiation source, a Markov decision process model is performed, and the maximum cumulative reward return value of the drone during the search period is used as the objective function. At the same position, different reward values will be received in five directions, and the reward value is related to the received signal strength value of each omnidirectional antenna; the received signal strength value in each direction obtained by the five omnidirectional antennas is processed using the optimal strategy learning algorithm, and the state of the drone and the flight action are determined by the processing result. The current position of the drone is judged according to the distance difference between the drone and the radiation source to determine whether the algorithm termination condition is met; if the condition is met, the location information of the illegal radiation source and the search trajectory information of the drone are returned; if the condition is not met, return to step 2, continue searching and use the five-element omnidirectional antenna to receive signals; wherein, the state of the drone includes the position of the drone, the speed of the drone, and the received signal strength in the five directions of the drone.
[0012] Furthermore, in step S1, the process of controlling the drone equipped with five omnidirectional antennas to be at a set height and receiving surrounding radiation signals through the five omnidirectional antennas includes the following steps:
[0013] S11, controlling a drone equipped with five omnidirectional antennas to be at a set height, and the drone receives surrounding radiation signals through the five omnidirectional antennas;
[0014] S12, the spatial coordinates of the radiation source are marked as (x0, y0, z0), the center of mass of the drone is used as the geometric center, and the spatial coordinates of the drone are marked as (x t ,y t , z t ), and the coordinates of the drone change as the drone moves; the distance between the radiation source target and the drone is recorded as d1:
[0015]
[0016] S13, connect the coordinate point corresponding to the i-th omnidirectional antenna with the coordinate point where the center of mass of the drone is located. The angle between this line and the x-axis in the two-dimensional coordinate system is recorded as i=1,...,5,the coordinates of the i-th antenna are The distance d2 between the i-th omnidirectional antenna and the radiation source target is:
[0017]
[0018] S14, construct the logarithmic path loss model as:
[0019] Pi[dB]=P0[dB]-20log10(d2)-ni+GR;
[0020] Where Pi[dB] is the received signal strength in dB; P0[dB] is the received signal strength at the near-ground reference distance in dB; 20log10(d2) is the spatial path loss in dB; ni is the effect of noise on the received signal strength, which has a mean of 0 and a variance of σ. 2 GR is the gain of the receiving antenna to the signal, in dB;
[0021] S15, taking the inverse function of Pi [dB] according to a logarithmic function to obtain the power of the received signal strength.
[0022] Furthermore, in step S2, the step length of the random walk model is recorded as s, the step length s is a random variable, and the random parameters u and v that obey the normal distribution are set, u~N(0,σ 2 ), v~N(0,1), the calculation formula of the step length s of each iteration satisfies the following formula:
[0023]
[0024] Furthermore, in step S3, the optimal strategy learning algorithm is used to process the received signal strength values in each direction obtained by the five omnidirectional antennas, and the process of determining the state of the drone and the flight action based on the processing results includes the following steps:
[0025] S31, according to the action and state of the drone, the state value function describing the state and action of the drone is recorded as V(s t , a t ), initialize the state value function V(s t , a t ), where s t represents the state of the drone at time t, a t Represents the drone at time t at s t The selected action in the state;
[0026] Set the discount coefficient γ∈[0,1], the learning coefficient α∈[0,1], and the probability threshold ε∈[0,1]. The discount coefficient γ is used to determine the proportion of future reward, the learning coefficient α is used to control the learning rate, and the probability threshold ε is used to determine the probability of selecting a random action. The reward value accumulated by the drone during the search period is used as the objective function. According to the cumulative value of the objective function corresponding to the state at time t, the optimal equation for dynamic programming is obtained:
[0027]
[0028] Where V(s j ) is state s j The state function value obtained under , p is the transition probability in the Markov decision process;
[0029] S32, initialize the drone search time t=0, initialize the drone starting state s0, initialize the state value table, and record the state value table as The state value table is used to temporarily store the state values generated by selecting different actions under possible states at a certain moment. The size of the table depends on the size of the state space and the size of the action space. The size of the state value table is recorded as n S ×n A , where n S represents the number of possible states in the state space, n A Represents the number of possible actions in the action space;
[0030] S33, in the current state of the drone, use the five omnidirectional antennas to obtain the received signal strength values in five directions, and based on the received signal strength values of each omnidirectional antenna, obtain the feedback in the corresponding direction and its state, and update the drone state value table T V ;
[0031] S34, using the ε-greedy strategy, makes a decision on the drone's action in the current state, updates the function and the state according to the dynamic programming optimal equation in step S31, and saves the selected action, the value function V after the decision is executed, and the state after the decision is made;
[0032] S35, updating the drone search time t, calculating and saving the drone search trajectory according to the drone flight speed and time interval;
[0033] S36, judging whether the current position of the drone meets the algorithm termination condition based on the distance difference between the drone and the radiation source. If the condition is met, the location information of the illegal radiation source and the search trajectory information of the drone are returned; if the condition is not met, returning to step S33.
[0034] Furthermore, in step S33, under the current state of the drone, the five omnidirectional antennas are used to obtain the received signal strength values in five directions, and based on the received signal strength values of each omnidirectional antenna, the feedback in the corresponding direction and its state are obtained, and the drone state value table T is updated. V The process includes the following steps:
[0035] S331, normalize the received signal strength values in five directions obtained by the five omnidirectional antennas and integrate them into a five-dimensional vector RSS, where RSS = {RSS1, RSS2, RSS3, RSS4, RSS5}; RSS1, RSS2, RSS3, RSS4, and RSS5 are the normalized values of the received signal strength values in the five directions, respectively;
[0036] S332: Obtain the current state value in the corresponding direction according to each received signal strength value, and update the state space s at time t. t ={s1, s2, s3, s4, s5}, s1, s2, s3, s4, s5 are the current state values in the five directions respectively;
[0037] S333, according to the state value of the current state space, update the t Execute action a in state t The reward return value r(s t , a t ), update the state value table T according to the reward return value V ;r(s t , a t ) is the reward value converted from the received signal strength values in the five directions after processing.
[0038] Furthermore, in step S34, the ε-greedy strategy is adopted to make a UAV action decision in the current state, and the function and state are updated according to the dynamic programming optimal equation in step S31. The process of saving the selected action, the value function V after the decision is executed, and the state after the decision is made includes the following steps:
[0039] S341, observe the state value function V(s t , a t ), use the ε-greedy strategy to select a drone action a t , for ε∈(0,1), the action with the highest reward value is selected, and its probability is 1-ε. The larger the ε value, the higher the probability that the drone can randomly generate an action. According to the maximized state value function, the action to be executed is selected, and the action a that maximizes the state value is found. t :
[0040]
[0041] Among them, the set of 5 actions evenly distributed on the corresponding angular domain is denoted as a t ∈{a1,a2,a3,a4,a5},the current state value s t The corresponding state value is recorded as V(s t ,:), and fill in the status value table;
[0042] S342, the drone performs the selected action a t , and simultaneously process the received signal strength value on the antenna, predict the possible state of the next state space according to the five-dimensional vector RSS = {RSS1, RSS2, RSS3, RSS4, RSS5}, and obtain the possible state value V(s) of the next state t+1,:);
[0043] S343, assuming that in the current state s i Next, take action a i , the reward value of executing this action in this state is r(s i , a i ), the steering angle of the drone is recorded as θ i , execute the corresponding flight step, and record the next state of the drone as si +1 ; Further modify the optimal equation of function dynamic programming and record the update rule of the state value function as:
[0044] V(s i , a i )←V(s i , a i )+α(r(s i , a i )+γmax(V(s i+1 , a i+1 ))-V(s i , a i ))
[0045] where a i+1 For the next state s i+1 Actions taken in.
[0046] Furthermore, in step S35, the process of updating the drone search time t and calculating and saving the drone search trajectory according to the drone flight speed and time interval includes the following steps:
[0047] Update the learning time t = t + 1; update the next state of the drone based on the current drone search state and the execution action generated by the optimal strategy.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] First, the radiation source positioning method based on online UAV trajectory planning of the present invention, the UAV is combined with an efficient flight algorithm for minimizing the search area, and utilizes the characteristics of the efficient flight algorithm for minimizing the search area with alternating step sizes and strong ergodicity, which effectively avoids the problem of non-convergence of search results, and solves the technical problem that when searching for mobile radiation sources over a large area, the UAV platform is increasingly affected by the external environment and the signal-to-noise ratio in the environment is too low.
[0050] Second, the radiation source positioning method based on online UAV trajectory planning of the present invention combines the optimal strategy learning algorithm with the UAV's five-element omnidirectional antenna, making the UAV's action selection for searching for radiation sources more effective and rapid; the signal strength values received in the five directions of the five-element omnidirectional antenna are different, and the direction of the maximum received signal strength value is closer to the target radiation source.
[0051] Third, the radiation source positioning method based on online UAV trajectory planning of the present invention can quickly plan a shorter path in a shorter time to achieve rapid positioning of illegal radiation sources on the ground, thereby quickly finding illegal radiation sources. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 This is a flow chart of the radiation source positioning method based on online UAV trajectory planning of the present invention;
[0053] Figure 2 It is a three-dimensional geometric position relationship diagram between the radiation source target and the UAV;
[0054] Figure 3 It is a simplified two-dimensional geometric position relationship diagram between the radiation source target and the UAV;
[0055] Figure 4 A detailed physical scene diagram for drones to search for radiation sources;
[0056] Figure 5 Schematic diagram of equipping a drone with five omnidirectional antennas;
[0057] Figure 6 Schematic diagram of the five actions of the drone, which correspond to five directions evenly distributed in the angular domain;
[0058] Figure 7 The figure is a flow chart of the radiation source positioning method. DETAILED DESCRIPTION
[0059] The embodiments of the present invention are described in further detail below with reference to the accompanying drawings.
[0060] This embodiment discloses a radiation source positioning method based on online UAV trajectory planning, which includes the following steps:
[0061] S1: Control a drone equipped with five omnidirectional antennas at a set altitude, and the drone receives surrounding radiation signals through the five omnidirectional antennas. The five omnidirectional antennas are placed on the drone in a clockwise direction, and the distance from each omnidirectional antenna to the center of mass of the drone is the same;
[0062] S2, using a random walk model for minimizing the search area to make the drone search the entire world. The step size of the random walk model is specified as the ratio of the normal random parameter u to the square of the standard normal random parameter v, so that the variance of the random walk model shows an exponential relationship with time, and the flight path shows the characteristics of alternating long and short.
[0063] S3, for the process of the drone searching for the radiation source, a Markov decision process model is performed, and the maximum cumulative reward return value of the drone during the search period is used as the objective function. At the same position, different reward values will be received in five directions, and the reward value is related to the received signal strength value of each omnidirectional antenna; the received signal strength value in each direction obtained by the five omnidirectional antennas is processed using the optimal strategy learning algorithm, and the state of the drone and the flight action are determined by the processing result. The current position of the drone is judged according to the distance difference between the drone and the radiation source to determine whether the algorithm termination condition is met; if the condition is met, the location information of the illegal radiation source and the search trajectory information of the drone are returned; if the condition is not met, return to step 2, continue searching and use the five-element omnidirectional antenna to receive signals; wherein, the state of the drone includes the position of the drone, the speed of the drone, and the received signal strength in the five directions of the drone.
[0064] like Figure 1 As shown, this embodiment proposes a radiation source positioning method based on online UAV trajectory planning, and the specific implementation steps are as follows:
[0065] Step S1: receiving a signal emitted by a radiation source;
[0066] The drone equipped with five omnidirectional antennas is controlled to be at a set height, and the drone receives the surrounding radiation signals through the five omnidirectional antennas. Figure 2 As shown in the three-dimensional geometric position relationship diagram of the radiation source target and the UAV, the spatial coordinates of the radiation source are marked as (x0, y0, z0), and the spatial coordinates of the UAV are marked as (x t ,y t , z t ), and the coordinates of the drone change as the drone moves. The distance between the radiation source target and the drone is recorded as d1, and d1 is required to satisfy the following formula:
[0067]
[0068] Considering the volume of the drone, the center of the drone is the geometric center, and the spatial coordinates of the drone are marked as (x t ,y t , z t), among the five omnidirectional antennas, connect the coordinate point corresponding to the i-th omnidirectional antenna with the coordinate point where the center of mass of the drone is located. The angle between this line and the x-axis in the two-dimensional coordinate system is recorded as The coordinates of the i-th antenna are (i=1, 2, 3, 4, 5), the distance d2 between the radiation source target and the i-th omnidirectional antenna satisfies the following formula:
[0069]
[0070] Considering the path loss of radio signals during channel transmission, the received signal strength is equal to the received signal strength at the near-ground reference distance minus the spatial path loss, minus the impact of noise, and finally adding the gain of the receiving signal antenna. The logarithmic path loss model is:
[0071] Pi[dB]=P0[dB]-20log10(d2)-ni+GR;
[0072] Where Pi[dB] is the received signal strength (RSS), in dB; P0[dB] is the received signal strength at the near-ground reference distance, in dB; 20log10(d2) is the spatial path loss, that is, the signal attenuation at a distance d2 from the radiation source to the drone, reflecting the impact of the propagation channel on signal attenuation, also in dB; ni is the impact of noise on the received signal strength, which has a mean of 0 and a variance of σ 2 The received signal strength is calculated by subtracting the spatial path loss from the received signal strength at the near-ground reference distance, minus the noise contribution, and adding the gain of the receiving antenna. The inverse of the logarithmic function Pi [dB] yields the received signal strength power.
[0073] Step S2: Using an efficient flight algorithm for minimizing the search area, the UAV can effectively search the entire world.
[0074] Since the location of the radiation source is unknown a priori, the drone needs to efficiently traverse the entire world to find the approximate location of the radiation source. A random walk model is used to enable the drone to efficiently search the entire world. The variance of this efficient flight algorithm for minimizing the search area is exponential over time, and the flight path is characterized by alternating long and short paths, so it can quickly traverse the entire world. The step length of the flight algorithm is denoted as s. The step length s of the flight algorithm is a random variable, and the random parameters u and v are set to obey the normal distribution, u~N(0,σ 2), v~N(0,1), the step size of the efficient flight algorithm for minimizing the search area is defined as the ratio of the normal random parameter u to the standard normal random parameter v to the power of two-thirds, where two-thirds is a preset constant. That is, the calculation formula for the step size s of each iteration satisfies the following formula:
[0075]
[0076] After introducing an efficient flight algorithm for minimizing the search area, the drone's search range is wider and can effectively avoid falling into local optimality.
[0077] Step S3: Use the optimal strategy learning algorithm to process the received signal strength values in each direction obtained by the five omnidirectional antennas. Considering the size of the drone, the moving direction of the drone is determined based on the normalized values of the received signal strength in the five directions, thereby realizing the positioning operation of the illegal radiation source.
[0078] The optimal strategy learning algorithm does not require modeling the environment; since the transmission power of the radiation source is unknown, the optimal strategy learning algorithm can operate without a received signal strength observation model or prior information about the environment. The optimal strategy learning algorithm is suitable for solving the problem of locating illegal radiation sources when the prior information of the drone is unknown and the radiation source beam scanning pattern is unknown.
[0079] For the process of the drone searching for radiation sources, a Markov decision process model is performed, that is, the next state of the drone depends on the current state and current behavior. The state information of the drone at time t is defined as s t The drone state includes the drone’s location, drone speed, and the drone’s received signal strength in five directions. The flight direction selected by the drone at each moment is recorded as a t , drones in s t Execute a in the state t The action will generate a reward value, recorded as r(s t , a t ), the reward return value is a piecewise function; choosing different flight directions will result in different reward values. The Markov decision process strategy for the drone's search for radiation sources is denoted by π. To find the optimal flight strategy, the reward return value accumulated by the drone during the search period is used as the objective function. Based on the cumulative value of the objective function corresponding to the state at time t, a dynamic programming equation is derived, denoted as the dynamic programming optimal equation:
[0080]
[0081] Where V(s j ) is state s jThe state function value obtained under , γ is the discount coefficient, and p is the transition probability in the Markov decision process.
[0082] In order to solve the UAV path planning problem, the optimal strategy learning algorithm is combined with a five-element omnidirectional antenna to receive signal strength values (RSS) from five directions, denoted as RSS, which is a five-dimensional vector, denoted as RSS = {RSS1, RSS2, RSS3, RSS4, RSS5}. At the same time, a state value table is set to record the UAV status information, and the state value table is denoted as The state value table is used to temporarily store the state values generated by selecting different actions under possible states at a certain moment. The size of the table depends on the size of the state space and the size of the action space. The size of the state value table is recorded as n s ×n A , where n s represents the number of possible states in the state space, n A Represents the number of possible actions in the action space. Use the RSS value as the reward signal to update the state value table Each state-action in the state-value table corresponds to a value function V(s t , a t ) is used to determine the long-term discounted reward of taking a certain action in the current state. When the optimal policy learning algorithm is combined with the five-element omnidirectional antenna, the same location can receive reward values of varying magnitudes in five directions. The reward value is related to the RSS value. A larger reward value indicates that this direction is a good choice to a certain extent, and it can help find the target faster.
[0083] In the UAV path planning problem, since the UAV has been fixed at the controlled height, it can be assumed that the UAV moves on a two-dimensional plane and its position can be expressed by the coordinates in the plane rectangular coordinate system. At this time, the relative position of the UAV and the radiation source can be seen as Figure 3 The position of the drone is recorded as s, and its direction can be expressed by the polar angle θ, the value range of θ is [0, 2π), that is, the angular domain, and a certain direction of the drone is recorded as a i Assume that the angular domain is divided into five directions, corresponding to five action spaces, such as Figure 6 As shown in the figure, the five directions evenly distributed in the angular domain correspond to the five actions of the drone, and its action space is recorded as A = {a1, a2, a3, a4, a5}. i Indicates that the drone's heading angle is θ iAt the same time, the drone determines five states based on the measured received signal strength value. The state value of each drone contains the received signal strength values in five directions. Among the five received signal strength values, there is always a maximum value that determines the current state of the drone. The five received signal strength values correspond to five states. The state space of the drone is recorded as S = {s1, s2, s3, s4, s5}. In the corresponding state, the corresponding action in the action space has a corresponding received signal strength value RSS. The reward return value is recorded as: r(s i , a i )=RSS, where max(RSS) represents the maximum received signal strength in the five directions.
[0084] Therefore, the state value table T V Each state-action in corresponds to a value function V(s i , a i ), used to determine the long-term discounted return of taking an action in the current state. These value functions are stored in an n S ×n A In the matrix, since both the state value vector and the action value vector are five-dimensional, then n s =5,n A =5, that is, T V =0 5×5 At the beginning of the algorithm, this matrix will be initialized to a zero matrix, and then as the actions are executed, it will be continuously filled with updated values.
[0085] Figure 5 This is the radiation pattern of the five-element antenna. Due to the consideration of the size of the UAV, the RSS values measured in different directions are not equal; furthermore, compared with other directions, the direction with the maximum RSS value is closer to the radiation source target to a certain extent.
[0086] Considering the importance of maintaining a balance between exploration and exploitation in learning UAV trajectory planning, on the one hand, excessive exploration will reduce the performance of the optimal strategy learning algorithm; on the other hand, pure exploitation can easily make the system quickly reach the optimal strategy. Therefore, the ε-greedy strategy is used to select an action a for the UAV. t , for ε∈(0,1), the action with the highest action value is selected with a probability of 1-ε. This strategy usually selects the action to be executed based on the maximum value of the state value function, that is, to find the action a that maximizes the Q value. t .
[0087] Step S3 specifically includes the following steps:
[0088] S31: According to the action and state of the drone, the state value function describing the state and action of the drone is recorded as V(st , a t ), initialize the state value function V(s) of the optimal strategy learning algorithm t , a t ), where s t Represents the state of the drone at time t, a t Represents that the UAV is located at s at time t t The selected action in the state; set the discount coefficient γ∈[0,1], the learning coefficient α∈[0,1], and the probability threshold ε∈[0,1]. The discount coefficient γ is used to determine the proportion of future return rewards, the learning coefficient α is used to control the learning rate, and the probability threshold ε is used to determine the probability of selecting a random action.
[0089] S32: At time t=0, initialize the drone to the initial state s o , initialize the drone search time t = 0, initialize the state value table, and record the state value table as The state value table is used to temporarily store the state values generated by selecting different actions under possible states at a certain moment. The size of the table depends on the size of the state space and the size of the action space. The size of the state value table is recorded as n s ×n A , where n s represents the number of possible states in the state space, n A Represents the number of possible actions in the action space.
[0090] S33: Under the current state of the drone, use the five omnidirectional antennas to obtain the received signal strength values in five directions, and based on the received signal strength values of each omnidirectional antenna, obtain the feedback in the corresponding direction and its state, and update the drone state value table T V .
[0091] S34: In order to balance the relationship between the drone's search for radiation sources and the search for the environment around the radiation source, the ε-greedy strategy is adopted. Under the current state, the drone action decision is made, and the function V and the state are updated according to the dynamic programming optimal equation of the optimal strategy learning algorithm. The selected action, the value function V after the decision is executed, and the state after the decision are saved.
[0092] S35: Update the drone search time t, and calculate and save the drone search trajectory according to the drone flight speed and time interval.
[0093] S36: Determine whether the drone's current location meets the algorithm termination criteria based on the distance difference between the drone and the radiation source. If so, the algorithm returns the location of the illegal radiation source and the drone's search trajectory. If not, the algorithm continues with step S33.
[0094] Specifically, step S33 includes the following steps:
[0095] S331: Use five omnidirectional antennas to obtain received signal strength values in the five directions, which are recorded as RSS, which is a five-dimensional vector, recorded as RSS={RSS1, RSS2, RSS3, RSS4, RSS5}.
[0096] S332: After normalizing the received signal strength values in the five directions, obtain five normalized values as the five elements in the five-dimensional state value vector. According to each received signal strength value, obtain the current state value in the corresponding direction, and record the five-dimensional state value as a five-dimensional vector s t ={s1, s2, s3, s4, s5}, and then update the state space s at time t t ={s1, s2, s3, s4, s5}.
[0097] S333: Update the state value in s according to the current state space t Execute action a in state t The reward return value r(s t , a t ), and update the state value table T according to the reward return value V Among them, the drone determines five states according to the measured received signal strength value. The state value of each drone contains the received signal strength values in five directions. The five received signal strength values correspond to five states s t ∈{s1, s2, s3, s4, s5}, r(s t , a t ) is the reward value converted from the received signal strength values in the five directions after processing.
[0098] Specifically, step S34 includes the following steps:
[0099] S341: Observe the state value function V(s t , a t ), use the ε-greedy strategy to select a drone action a t , for ε∈(0,1), the action with the highest reward value is selected with a probability of 1-ε. The larger the ε value, the higher the probability that the drone can randomly generate an action, and the more effectively it can explore the entire space. The action to be executed is selected based on the maximized state value function, that is, the action a that maximizes the state value is found. t :
[0100]
[0101] Among them, the set of 5 actions evenly distributed on the corresponding angular domain is denoted as at ∈{a1,a2,a3,a4,a5},the current state value s t The corresponding state value is recorded as V(s t ,:), and fill in the status value table.
[0102] S342: The drone performs the selected action a t , and simultaneously process the received signal strength value on the antenna, predict the possible state of the next state space according to the five-dimensional vector RSS = {RSS1, RSS2, RSS3, RSS4, RSS5}, and obtain the possible state value of the next state: V(s t+1 ,:).
[0103] S343: Assume that in the current state s i Next, take action a i , the reward value of executing this action in this state is r(s i , a i ), the steering angle of the drone is recorded as θ i , execute the corresponding flight step, and record the next state of the drone as s i+1 Considering the cumulative rewards of executing this strategy in the long term, it is necessary to calculate the maximum possible reward value of the next state after executing the action, which will also affect the update of the value function V. For the optimal equation of the dynamic programming of the objective function, we make further corrections to consider the long-term impact of the action execution. We write the update rule of the state value function as follows:
[0104] V(s i , a i )←V(s i , a i )+α(r(s i , a i )+γmax(V(s i+1 , a i+1 ))-V(s i , a i ))
[0105] Where α is the learning coefficient, γ is the discount coefficient, and a i+1 For the next state s i+1 The actions that can be taken in.
[0106] Specifically, step S35 includes the following steps:
[0107] S351: Update the learning time t=t+1.
[0108] S352: Based on the current UAV search state, including the UAV's location, flight speed, and state value in the current state space, and the execution action generated according to the optimal strategy, update the UAV's next state, including the location, flight speed, and predicted state value in the next state space.
[0109] The overall algorithm flow chart of step S3 is shown in Figure 7 .
[0110] The radiation source positioning method based on online UAV trajectory planning provided in the present invention is verified by combining specific examples.
[0111] The drone's initial position is (0, 0, 200), meaning it's flying at an altitude of 200 meters, with a flight step size of 1. The ground radiation source's position is (550, 900, 0). The discount factor γ is 0.6, the learning coefficient α is 0.5, and the probability threshold ε is 0.2. The optimal policy learning algorithm is repeated 5000 times. When the noise is Gaussian white noise, with a standard deviation σ of 1, the drone plans a smooth and short path, resulting in satisfactory results.
[0112] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiment of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal translation scripting language JavaScript, etc.
[0113] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0114] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions for executing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0116] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0117] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A radiation source positioning method based on online UAV trajectory planning, characterized in that: The radiation source positioning method comprises the following steps: S1: Control a drone equipped with five omnidirectional antennas at a set altitude, and the drone receives surrounding radiation signals through the five omnidirectional antennas. The five omnidirectional antennas are placed on the drone in a clockwise direction, and the distance from each omnidirectional antenna to the center of mass of the drone is the same; S2, using a random walk model for minimizing the search area to make the drone search the entire world. The step size of the random walk model is specified as the ratio of the normal random parameter u to the square of the standard normal random parameter v, so that the variance of the random walk model shows an exponential relationship with time, and the flight path shows the characteristics of alternating long and short. S3, for the process of the drone searching for the radiation source, a Markov decision process model is performed, and the maximum cumulative reward return value of the drone during the search period is used as the objective function. At the same position, different reward values will be received in five directions, and the reward value is related to the received signal strength value of each omnidirectional antenna; the received signal strength value in each direction obtained by the five omnidirectional antennas is processed using the optimal strategy learning algorithm, and the state of the drone and the flight action are determined by the processing result. The current position of the drone is judged based on the distance difference between the drone and the radiation source to determine whether the algorithm termination condition is met; if the condition is met, the location information of the illegal radiation source and the search trajectory information of the drone are returned; if the condition is not met, the process returns to step 2, continues the search and uses the five-element omnidirectional antenna to receive the signal; wherein, the state of the drone includes the position of the drone, the speed of the drone, and the received signal strength in the five directions of the drone.
2. The radiation source positioning method based on online UAV trajectory planning according to claim 1 is characterized in that: In step S1, the process of controlling the drone equipped with five omnidirectional antennas to be at a set height and receiving the surrounding radiation signals through the five omnidirectional antennas includes the following steps: S11, controlling a drone equipped with five omnidirectional antennas to be at a set height, and the drone receives surrounding radiation signals through the five omnidirectional antennas; S12, the spatial coordinates of the radiation source are marked as (x0, y0, z0), the center of mass of the drone is used as the geometric center, and the spatial coordinates of the drone are marked as (x t ,y t ,z t ), and the coordinates of the drone change as the drone moves; the distance between the radiation source target and the drone is recorded as d1: S13, connect the coordinate point corresponding to the i-th omnidirectional antenna with the coordinate point where the center of mass of the drone is located. The angle between this line and the x-axis in the two-dimensional coordinate system is recorded as i=1,…,5, the coordinates of the i-th antenna are The distance d2 between the i-th omnidirectional antenna and the radiation source target is: S14, construct the logarithmic path loss model as: Pi[dB]=P0[dB]-20log10(d2)-ni+GR; Where Pi[dB] is the received signal strength in dB; P0[dB] is the received signal strength at the near-ground reference distance in dB; 20log10(d2) is the spatial path loss in dB; ni is the effect of noise on the received signal strength, which has a mean of 0 and a variance of σ. 2 GR is the gain of the receiving antenna to the signal, in dB; S15, taking the inverse function of Pi [dB] according to a logarithmic function to obtain the power of the received signal strength.
3. The radiation source positioning method based on online UAV trajectory planning according to claim 1 is characterized in that: In step S2, the step length of the random walk model is recorded as s, the step length s is a random variable, and the random parameters u and v are set to obey the normal distribution, u~N(0,σ 2 ),v~N(0,1), the calculation formula of the step size s of each iteration satisfies the following formula:
4. The radiation source positioning method based on online UAV trajectory planning according to claim 1 is characterized in that: In step S3, the optimal strategy learning algorithm is used to process the received signal strength values in each direction obtained by the five omnidirectional antennas. The process of determining the state of the drone and the flight action based on the processing results includes the following steps: S31, according to the action and state of the drone, the state value function describing the state and action of the drone is recorded as V(s t ,a t ), initialize the state value function V(s t ,a t ), where s t represents the state of the drone at time t, a t Represents the drone at time t at s t The selected action in the state; Set the discount coefficient γ∈[0,1], the learning coefficient α∈[0,1], and the probability threshold ε∈[0,1]. The discount coefficient γ is used to determine the proportion of future reward, the learning coefficient α is used to control the learning rate, and the probability threshold ε is used to determine the probability of selecting a random action. The reward value accumulated by the drone during the search period is used as the objective function. According to the cumulative value of the objective function corresponding to the state at time t, the optimal equation for dynamic programming is obtained: Where V(s j ) is state s j The state function value obtained under the condition of Markov decision process is p, and r(s t ,a t ) is the UAV in s t Execute a in the state t The reward value generated by the action; S32, initialize the drone search time t=0, initialize the drone starting state s0, initialize the state value table, and record the state value table as The state value table is used to temporarily store the state values generated by selecting different actions under possible states at a certain moment. The size of the table depends on the size of the state space and the size of the action space. The size of the state value table is recorded as n S ×n A , where n S represents the number of possible states in the state space, n A Represents the number of possible actions in the action space; S33, in the current state of the drone, use the five omnidirectional antennas to obtain the received signal strength values in five directions, and based on the received signal strength values of each omnidirectional antenna, obtain the feedback in the corresponding direction and its state, and update the drone state value table T V ; S34, using the ε-greedy strategy, makes a decision on the drone's action in the current state, updates the function and the state according to the dynamic programming optimal equation in step S31, and saves the selected action, the value function V after the decision is executed, and the state after the decision is made; S35, updating the drone search time t, calculating and saving the drone search trajectory according to the drone flight speed and time interval; S36, judging whether the current position of the drone meets the algorithm termination condition based on the distance difference between the drone and the radiation source. If the condition is met, the location information of the illegal radiation source and the search trajectory information of the drone are returned; if the condition is not met, returning to step S33.
5. The radiation source positioning method based on online UAV trajectory planning according to claim 4 is characterized in that: In step S33, under the current state of the drone, the five omnidirectional antennas are used to obtain the received signal strength values in five directions, and based on the received signal strength values of each omnidirectional antenna, the corresponding direction of the report and its state are obtained, and the drone state value table T is updated. V The process includes the following steps: S331, normalize the received signal strength values in five directions obtained by the five omnidirectional antennas and integrate them into a five-dimensional vector RSS, where RSS = {RSS1, RSS2, RSS3, RSS4, RSS5}; RSS1, RSS2, RSS3, RSS4, RSS5 are the normalized values of the received signal strength values in the five directions respectively; S332: Obtain the current state value in the corresponding direction according to each received signal strength value, and update the state space s at time t. t ={s1,s2,s3,s4,s5}, s1,s2,s3,s4,s5 are the current state values in the five directions respectively; S333, according to the state value of the current state space, update the t Execute action a in state t The reward return value r(s t ,a t ), update the state value table T according to the reward return value V ;r(s t ,a t ) is the reward value converted from the received signal strength values in the five directions after processing.
6. The radiation source positioning method based on online UAV trajectory planning according to claim 4 is characterized in that: In step S34, the ε-greedy strategy is adopted to make a UAV action decision in the current state. According to the dynamic programming optimal equation in step S31, the function and state are updated, and the selected action, the value function V after the decision is executed, and the state after the decision are saved. The process includes the following steps: S341, observe the state value function V(s t ,a t ), use the ε-greedy strategy to select a drone action a t , for ε∈(0,1), the action with the highest reward value is selected, and its probability is 1-ε. The larger the ε value, the higher the probability that the drone can randomly generate an action. According to the maximized state value function, the action to be executed is selected, and the action a that maximizes the state value is found. t : Among them, the set of 5 actions evenly distributed on the corresponding angular domain is denoted as a t ∈{a1,a2,a3,a4,a5}, the current state value s t The corresponding state value is recorded as V(s t ,:), and fill in the status value table; S342, the drone performs the selected action a t , and simultaneously process the received signal strength value on the antenna, predict the possible state of the next state space according to the five-dimensional vector RSS = {RSS1, RSS2, RSS3, RSS4, RSS5}, and obtain the possible state value V(s) of the next state t+1 ,:); S343, assuming that in the current state s i Next, take action a i , the reward value of executing this action in this state is r(s i ,a i ), the steering angle of the drone is recorded as θ i , execute the corresponding flight step, and record the next state of the drone as s i+1 ; Further modify the optimal equation of function dynamic programming and record the update rule of the state value function as: V(s i ,a i )←V(s i ,a i )+α(r(s i ,a i )+γmax(V(s i+1 ,a i+1 ))-V(s i ,a i )) where a i+1 For the next state s i+1 Actions taken in.
7. The radiation source positioning method based on online UAV trajectory planning according to claim 4 is characterized in that: In step S35, the drone search time t is updated. The process of calculating and saving the drone search trajectory according to the drone flight speed and time interval includes the following steps: Update the learning time t = t + 1; update the next state of the drone based on the current drone search state and the execution action generated by the optimal strategy.
Citation Information
Patent Citations
Unmanned aerial vehicle group flight path optimization method for multi-radiation source tracking
CN114337875A
Unmanned aerial vehicle dynamic target tracking control method based on deep reinforcement learning
CN115686065A
Method of unmanned aerial vehicle for searching illegal broadcasting station based on reinforcement learning
CN108387866A
Radiation source positioning device and method based on UAV (Unmanned Aerial Vehicle) platform loaded with co-prime linear array
CN110297213A