Short-wave autonomous direction-finding positioning method and system based on depth deterministic strategy gradient algorithm
By applying the reinforcement learning method of the deep deterministic strategy gradient algorithm in short-wave direction finding positioning, the problem of time-consuming, manual and high computational cost in traditional methods is solved, and real-time and automatic short-wave signal direction finding positioning is achieved.
Patent Information
- Application Number
- CN202510196371.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional short-wave direction finding and positioning methods are time-consuming, labor-intensive, rely on manual experience, have a large amount of calculations, and cannot meet real-time requirements.
The reinforcement learning method based on the deep deterministic strategy gradient algorithm is adopted, and the reinforcement learning direction finding environment of hybrid action space is set through space-time high-resolution time-frequency diagrams, reward functions are designed to improve positioning accuracy, and intelligent short-wave direction finding positioning model is trained to achieve real-time positioning.
Automatic direction finding and positioning is realized, manual annotation is reduced, and the intelligent direction finding and positioning ability of short-wave signals is improved, which can meet the real-time positioning needs.
Smart Images

Figure CN120044472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of short-wave signal processing, and particularly to a short-wave autonomous direction finding and positioning method and system based on a deep deterministic policy gradient algorithm. Background Art
[0002] Short-wave communication can achieve long-distance over-the-horizon transmission through the ionosphere. By performing spatio-temporal high-resolution processing on the received short-wave signals, the signals can be direction-found and positioned, playing a significant role in analyzing the situation. Traditional short-wave direction finding and positioning work is usually carried out by staff who listen to and distinguish signals based on their own experience and operate the direction finding equipment to repeatedly compare the multi-direction clustering results for accurate direction finding. Traditional direction finding and positioning work often consists of two steps, that is, first direction finding the short-wave signal and then positioning. The most commonly used short-wave positioning method based on parameter estimation is multi-station angle measurement intersection positioning. This method requires each observation station to install an array antenna and uses the direction of arrival (DOA) estimation method to calculate the azimuth angle and elevation angle of the incident signal, and then uses the DOA positioning parameter estimation method to determine the target position. The positioning method based on parameter estimation has a large amount of calculation, so it cannot meet the real-time requirements in short-wave direction finding and positioning. Summary of the Invention
[0003] Therefore, the present invention provides a short-wave autonomous direction finding and positioning method and system based on a deep deterministic policy gradient algorithm to solve the problems of time-consuming, laborious, relying on manual experience, large amount of calculation, and inability to meet the real-time requirements of traditional direction finding and positioning. A reinforcement learning direction finding environment with a mixed action space is set using a spatio-temporal high-resolution time-frequency map, a reward function is designed according to the positioning accuracy, and a deep reinforcement learning algorithm is used to train the intelligent short-wave direction finding and positioning environment to improve the real-time performance of short-wave direction finding and positioning work in actual deployment applications.
[0004] According to the design scheme provided by the present invention, on the one hand, a short-wave autonomous direction finding and positioning method based on a deep deterministic policy gradient algorithm is provided, including:
[0005] Receiving short-wave signal array data by using a plurality of observation stations, and performing spatio-temporal high-resolution processing on the short-wave signal array data to obtain a spatio-temporal high-resolution time-frequency map;
[0006] Input the spatio-temporal high-resolution time-frequency map into the trained short-wave direction-finding and positioning model, and use the short-wave direction-finding and positioning model to obtain the direction-finding and positioning results of the target signal source. The short-wave direction-finding and positioning model is trained based on the deep deterministic policy gradient algorithm, and the deep deterministic policy gradient algorithm is used to learn actions in different environmental state spaces based on the state space, action space, and reward function. Among them, the state space of the deep deterministic policy gradient algorithm is set based on the spatio-temporal high-resolution time-frequency map of the array signal, the action space is constructed based on the executable actions of the rectangular box used to select the direction-finding and positioning area of the signal source, and the reward function is set according to the signal source positioning accuracy.
[0007] As the short-wave autonomous direction-finding and positioning method based on the deep deterministic policy gradient algorithm of the present invention, further, use several observation stations to receive short-wave signal array data, including:
[0008] Set the frequency of the observation station receiver according to the direction-finding parameters, and obtain the short-wave signal array data based on the signal model. The signal model is constructed based on the signal source, the array manifold matrix, and the array additive noise.
[0009] As the short-wave autonomous direction-finding and positioning method based on the deep deterministic policy gradient algorithm of the present invention, further, the action space includes: a discrete action space and a continuous action space. The discrete action space consists of the moving actions in the up, down, left, and right directions of the rectangular box, and the continuous action space consists of the adjustment actions of the length and / or width of the rectangular box.
[0010] As the short-wave autonomous direction-finding and positioning method based on the deep deterministic policy gradient algorithm of the present invention, further, set the reward function according to the signal source positioning accuracy, including:
[0011] Use the direction-finding data of the selected spatio-temporal high-resolution time-frequency map to represent the direction-finding error. Based on the direction-finding error and use the synchronous direction-finding intersection of each observation station to estimate the position of the short-wave signal source on the earth's surface, and set the reward function according to the area of the convex polygon composed of the synchronous direction-finding intersection points of the observation stations.
[0012] As the short-wave autonomous direction-finding and positioning method based on the deep deterministic policy gradient algorithm of the present invention, further, the direction-finding error includes the azimuth observation error and the elevation observation error, and both the azimuth observation error and the elevation observation error are represented as Gaussian distribution vectors with a mean of 0 and a standard deviation of σ. Among them, Y over is the overlapping area between the rectangular box and the signal source area, Y agent is the area of the rectangular box.
[0013] As the short-wave autonomous direction-finding and positioning method based on the deep deterministic policy gradient algorithm of the present invention, further, the reward function is expressed as: Among them, V is the area of the convex polygon, Vmax is the maximum allowable error area of the convex polygon.
[0014] As the short-wave autonomous direction-finding and positioning method based on the deep deterministic policy gradient algorithm of the present invention, further, the short-wave direction-finding and positioning model is trained according to the following steps:
[0015] Initialize the experience replay pool, value network parameters, and policy network parameters;
[0016] Input the space-time high-resolution time-frequency map as the current moment state into the policy network to obtain the corresponding rectangular box execution action for the selected signal source direction-finding and positioning area; calculate the reward value of the deep deterministic policy gradient algorithm based on the signal source direction-finding and positioning area selected by the rectangular box execution action and use the reward value to update the next moment state. Store the current moment state, corresponding action, reward value, and next moment state as historical experience in the experience replay pool to collect historical experience data and store it in the experience replay pool by interacting with different environmental states through iterative loops until the maximum iterative loop condition is met. And in each iterative loop, based on the next moment state, obtain the corresponding action at the next moment using the policy network, and calculate the value corresponding to the next moment state and the corresponding action at the next moment using the value network. Update according to the value acquisition network update gradient;
[0017] Randomly sample historical experience from the experience replay pool as training samples. Calculate the value of the policy network predicting the current action based on the training samples. Update the policy network and value network based on the value acquisition network loss gradient and according to the network loss gradient. Return to resample historical experience until the training convergence condition is met. Use the final policy network as the short-wave direction-finding and positioning model.
[0018] On the other hand, the present invention also provides a short-wave autonomous direction-finding and positioning system based on the deep deterministic policy gradient algorithm, including: a signal receiving module and a direction-finding and positioning module, where,
[0019] The signal receiving module is used to receive short-wave signal array data using a number of observation stations and perform space-time high-resolution processing on the short-wave signal array data to obtain a space-time high-resolution time-frequency map;
[0020] The direction finding and positioning module is used to input the spatio-temporal high-resolution time-frequency map into the trained short-wave direction finding and positioning model, and obtain the direction finding and positioning result of the target signal source by using the short-wave direction finding and positioning model. The short-wave direction finding and positioning model is trained based on the deep deterministic policy gradient algorithm, and the deep deterministic policy gradient algorithm is used to learn actions in different environmental state spaces based on the state space, action space and reward function. Among them, the state space of the deep deterministic policy gradient algorithm is set based on the spatio-temporal high-resolution time-frequency map of the array signal, the action space is constructed based on the executable actions of the rectangular frame used to select the direction finding and positioning area of the signal source, and the reward function is set according to the signal source positioning accuracy.
[0021] Advantages of the present invention:
[0022] Based on the reinforcement learning environment of the spatio-temporal high-resolution time-frequency map and the agent with a hybrid action space, the present invention designs a reward function according to the accuracy of the positioning result, establishes an observable Markov decision process, and uses the multi-level non-refined evaluation of deep reinforcement learning to achieve automatic direction finding and positioning, reduce manual annotation, realize autonomous evolution while working and learning, and gradually improve the intelligent direction finding and positioning ability of short-wave signals. The simulation experiment results show that the deep deterministic policy gradient algorithm of the solution in this case has better performance than other deep reinforcement learning algorithms. Finally, the trained model is saved and tested, and the test results show that the solution in this case can realize the real-time direction finding and positioning work of short-wave signals. Description of the drawings
[0023] Figure 1 Schematic diagram of the short-wave autonomous direction finding and positioning process based on the deep deterministic policy gradient algorithm in the embodiment;
[0024] Figure 2 Schematic diagram of the principle of the short-wave autonomous direction finding and positioning algorithm in the embodiment;
[0025] Figure 3 Schematic diagram of the signal and the agent overlap in the embodiment;
[0026] Figure 4 Schematic diagram of the spectrogram environment before vector quantization processing in the embodiment;
[0027] Figure 5 Schematic diagram of the spectrogram environment after vector compression in the embodiment;
[0028] Figure 6 Schematic diagram of the network architecture of the DDPG algorithm and the A2C algorithm structure in the embodiment;
[0029] Figure 7 Schematic diagram of the average return comparison curve of different algorithms in the embodiment;
[0030] Figure 8 Schematic diagram of the average step length comparison curve of different algorithms in the embodiment;
[0031] Figure 9 Schematic diagram of the comparison of the average running time of different algorithms in the embodiments;
[0032] Figure 10 Schematic diagram of the comparison of the positioning error of different algorithms with the change of the direction finding error in the embodiments. Detailed implementation manners
[0033] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and technical solutions.
[0034] As another research hotspot in the field of artificial intelligence, reinforcement learning has been applied to aspects such as unmanned aerial vehicles, communication networks, and intelligent manufacturing. The basic idea of reinforcement learning is to continuously learn the optimal strategy to complete the goal by maximizing the long-term reward obtained by the agent from the environment, and it has advantages such as interacting with the environment, continuous state-action, and online learning. Therefore, the reinforcement learning method focuses more on learning the strategy to solve problems. With the rapid development of human society, in complex real-world tasks, deep reinforcement learning can be used to autonomously learn the abstract features of sample data and optimize the strategy to solve problems. In the embodiments of the present invention, refer to Figure 1 as shown, a short-wave autonomous direction finding and positioning method based on a deep deterministic policy gradient algorithm is provided, including:
[0035] S101. Use a number of observation stations to receive short-wave signal array data, and perform space-time high-resolution processing on the short-wave signal array data to obtain a space-time high-resolution time-frequency diagram.
[0036] Assume that N observation stations are used to receive D short-wave signals, and the receiver is frequency-set according to the direction finding parameters to obtain array signal data. The signal model is:
[0037] X(t) = A(Θ)S(t) + N(t)
[0038] In the formula, X(t) = [x 1 (t), x 2 (t),... x M (t)] is the array output vector, N(t) = [n 1 (t), n 2 (t),..., n M (t)] is the array additive noise vector, S(t) = [s 1 (t), s 2 (t),... s D (t)] is the received signal source vector, A(Θ) = [a(Θ 1 ) a(Θ 2 )... a(Θ D )] is the array manifold matrix, is the array direction vector.
[0039] Perform space-time high-resolution processing on the received signal, remove interference and noise, and obtain corresponding multiple space-time high-resolution spectrograms to utilize joint signal direction finding and positioning using multiple time-frequency diagrams.
[0040] S102. Input the space-time high-resolution time-frequency diagram into the trained short-wave direction finding and positioning model, and use the short-wave direction finding and positioning model to obtain the direction finding and positioning result of the target signal source. The short-wave direction finding and positioning model is trained based on the deep deterministic policy gradient algorithm, and the deep deterministic policy gradient algorithm is used to learn actions in different environmental state spaces based on the state space, action space, and reward function. Among them, the state space of the deep deterministic policy gradient algorithm is set based on the space-time high-resolution time-frequency diagram of the array signal, the action space is constructed based on the executable actions of the rectangular frame used to select the direction finding and positioning area of the signal source, and the reward function is set according to the signal source positioning accuracy.
[0041] Specifically, the action space can be designed to include: a discrete action space and a continuous action space. The discrete action space consists of the moving actions in the up, down, left, and right directions of the rectangular frame, and the continuous action space consists of the adjustment actions of the length and / or width of the rectangular frame.
[0042] Among them, setting the reward function according to the signal source positioning accuracy can be designed to include:
[0043] Use the direction finding data of the selected space-time high-resolution time-frequency diagram to represent the direction finding error. Based on the direction finding error and using the synchronous direction finding intersection of each observation station to estimate the position of the short-wave signal source on the earth's surface, set the reward function according to the area of the convex polygon formed by the synchronous direction finding intersection points of the observation stations.
[0044] In the direction finding and positioning environment based on the space-time high-resolution time-frequency diagram, multiple direction finding stations receive signals synchronously. As Figure 2 shown in the reinforcement learning environment, first initialize a (400, 300, 3) window, which represents the single-channel time-frequency diagram of the signal received by the direction finding station. The positions of the short-wave signal and the rectangular frame agent are both randomly initialized. In the simulated space-time high-resolution time-frequency diagram, as Figure 4 shown, the yellow area represents the target signal, the red area represents the interference signal, and the rectangular frame represents the agent. Reinforcement learning is a Markov decision process composed of a state space, an action space, a state transition matrix, a reward function, and a discount factor, etc.
[0045] For the rectangular frame agent, the discrete action can be designed as {up, down, left, right}, and the continuous action space is the length w and width h of the rectangular frame, and the value ranges are (0, w max and (0, h max .
[0046] When the rectangular frame and the signal are at random initial positions, the observation state of the single-channel time-frequency RGB image of the received signal (400, 300, 3) is grayscaled, as Figure 5 shown; then the image shape is adjusted to (10, 10); then the two-dimensional vector is flattened into a one-dimensional vector, and is concatenated with the coordinates (x, y), length w, and width h of the agent to form a one-dimensional observation vector.
[0047] When obtaining the direction-finding error based on the selected time-frequency diagram for direction finding, during the process of simulating spatial spectrum direction finding, the accuracy of direction finding is inversely proportional to the signal-to-noise ratio of the selected area. That is, the smaller the signal-to-noise ratio, the higher the accuracy of direction finding; the larger the signal-to-noise ratio, the lower the accuracy of direction finding. When the agent selects a signal, as Figure 3 shown, the overlapping area Y over between the agent and the signal and the area Y agent of the agent's rectangular frame, the ratio is denoted as the standard deviation σ of the direction-finding error, and can be expressed as:
[0048]
[0049] Therefore, given the longitude and latitude coordinates of the direction-finding station, the true longitude and latitude of a single or multiple targets are required to calculate the true azimuth angle α r and true elevation angle β r , to obtain the true angle matrix θ r =[α r ,β r , which can simulate the actual direction-finding work. The direction-finding angle θ = [α, β] of each direction-finding station relative to the target is the true angle of each direction-finding station relative to the target plus the direction-finding error. The azimuth angle observation error σ θ and the elevation angle observation error σ β are both N×1 column vectors and follow a Gaussian distribution with a mean of 0 and a standard deviation of σ.
[0050] θ = θ r +σ
[0051] where σ represents the observation error, which is a 2N×1 column vector, and the covariance matrix is Ω = E[σσ T .
[0052] Considering that the earth is a spherical surface, under the condition of no prior information about the target, the size of the volume of the convex polyhedron formed by the intersection points of direction finding by multiple direction-finding stations is used as an evaluation index to design a task-based reward function. The value range of the reward function is set to 0 - 100.
[0053]
[0054] where V is the area of the convex polygon, V maxis the maximum allowable error area of the convex polygon. According to the positioning accuracy requirements, the positioning fails when the positioning error is greater than 100 km. Therefore, set V max = 10 4 .
[0055] Estimate the position of the signal source by synchronous direction-finding intersection positioning under any number of observation stations When, first convert the longitude and latitude of the direction-finding station to the Earth-centered Earth-fixed coordinates. Consider using N direction-finding stations to locate the short-wave radiation source on the Earth's surface. The longitude and latitude of the nth direction-finding station are (ω n , ρ n ), so the Earth-centered Earth-fixed coordinates of this direction-finding station are
[0056]
[0057] In the formula R e = 6378.160 km is the equivalent radius of the Earth; e = 0.081819643716348 is the eccentricity. The coordinate transformation matrix corresponding to this direction-finding station is
[0058]
[0059] Assume that the longitude and latitude of the short-wave radiation source are (ω, ρ), so the Earth-centered Earth-fixed coordinates of this short-wave radiation source are
[0060]
[0061] In the formula The vector u satisfies the equality constraint
[0062]
[0063] In the formula
[0064]
[0065] Specifically, the short-wave direction-finding positioning model is trained according to the following steps:
[0066] Initialize the experience replay pool, value network parameters, and policy network parameters;
[0067] The spatio-temporal high-resolution time-frequency map is used as the current moment state and input into the policy network to obtain the corresponding rectangular box for the signal source direction-finding and positioning area, and perform actions; based on the signal source direction-finding and positioning area selected by the actions of the rectangular box and using the reward function to calculate the reward value of the deep deterministic policy gradient algorithm, and using the reward value to update the next moment state. The current moment state, corresponding actions, reward value, and next moment state are stored in the experience replay pool as historical experience, so as to collect historical experience data and store them in the experience replay pool by interacting with different environmental states through iterative loops until the maximum iterative loop condition is satisfied. And in each iterative loop, based on the next moment state and using the policy network to obtain the corresponding actions at the next moment, and using the value network to calculate the value corresponding to the next moment state and the corresponding actions at the next moment, and update the gradient according to the value acquisition network.
[0068] Randomly sample historical experience from the experience replay pool as training samples, calculate the value of the current actions predicted by the policy network based on the training samples, calculate the loss gradient of the value acquisition network, and update the policy network and value network according to the network loss gradient. Return to resample historical experience until the training convergence condition is satisfied, and use the final policy network as the short-wave direction-finding and positioning model.
[0069] As Figure 6 shown, the output of the deterministic policy network μ(s; θ) is set as a d-dimensional vector a as the action. The input state s is a matrix or tensor, and μ is composed of several convolutional layers, fully connected layers, etc. The deterministic policy can be regarded as a special case of the stochastic policy. The output of the deterministic policy μ(s; θ) is a d-dimensional vector, and its i-th element is denoted as where the stochastic policy can be expressed as:
[0070]
[0071] This stochastic policy is a multivariate normal distribution with a mean of μ(s; θ) and a covariance matrix of diag(σ 1 ,…σ d ), and the deterministic stochastic policy can be regarded as a special case of the above stochastic policy when σ = [σ 1 ,…σ d is a zero vector.
[0072] In deep reinforcement learning, the agent interacts with the environment, records the observed states, actions, and rewards, and uses these experiences to learn a policy function π(a|s). The action value function Q π (s t ,a t ) depends on s t and a t and does not depend on the states and actions at time t+1 and later.
[0073]
[0074] Where the discounted return U t = R t + γ·R t+1 + γ 2 ·R t+2 + γ 3 ·R t+3 …, γ ∈ [0, 1] is the discount rate. It can be seen from the above formula that Q π (s t , a t ) depends on the policy function π(a|s). To exclude the influence of the policy function π, the solution is the optimal action-value function Q * (s t , a t ).
[0075]
[0076] The state-value function V π (s t ) is a prediction of future rewards, indicating the expected reward obtained by performing action a in state s t .
[0077]
[0078] The state-value function V π (s t ) is also the expectation of the return U t :
[0079]
[0080] The larger the state-value function, the greater the expected return. The state-value function can be used to measure the quality of the policy function π and the state s t . The policy that controls the interaction between the agent and the environment is called the behavior policy. The role of the behavior policy is to collect experience, that is, the observed environment, actions, and rewards. In the embodiments of this case, the ε-greedy behavior policy that balances the importance of "exploration" and "exploitation" is adopted, which is expressed as:
[0081]
[0082] In the experiment, at the beginning, the exploration rate ε is made relatively large, and during the training process, ε is gradually decayed, and after hundreds of thousands of steps, it decays to a smaller value (such as ε = 0.01), and then ε = 0.01 is fixed.
[0083] The action space in the embodiments of this case includes a continuous action space and a discrete action space. Therefore, the Deep Deterministic Policy Gradient (DDPG) and the Advantaged Actor-Critic (A2C) algorithms can be used to handle the discrete action space and the continuous action space. The DDPG algorithm and the A2C algorithm use neural networks to approximate the action-value function Q π , that is, the value network, denoted as q(s,a;ω), where ω represents the trainable parameters of the neural network. The A2C algorithm uses the neural network π(a|s;θ) to approximate the policy function π(a|s), which is called the policy network, and θ represents the parameters of the neural network. The deterministic policy network of the DDPG algorithm in the solution of this case is μ(s;θ), and θ represents the parameters of the neural network.
[0084] First, initialize the experience replay pool T. The value network is denoted as q(s,a;ω), the deterministic policy of the DDPG algorithm is μ(s;θ), and the current parameters of the neural network are denoted as ω now and θ now ; obtain the current state s according to the randomly initial positions of the rectangular box and the signal t ; the rectangular box determines according to the current state s t , and executes the action a using the policy network t ; obtain the direction finding error according to the direction finding of the selected time-frequency diagram, estimate the position of the signal source , calculate the reward r t , and obtain the new state s t+1 ; the policy network makes a prediction and gets the value network makes a prediction and gets and calculate the error according to the prediction results and update the value network and the policy network according to the error. Among them, the update of the value network is expressed as: The update of the policy network is expressed as:
[0085]
[0086] The value network q(s,a;ω) is an approximation of the action-value function Q π (s,a). The inputs of the value network are the state s and the action a, and the output is the value which is a real number and can reflect the quality of the action; the better the action a is, the greater the value is. Therefore, the value network can evaluate the performance of the policy network. During the training process, the value network helps to train the policy network; after the training is completed, the value network is discarded and the policy network controls the agent.
[0087] Furthermore, based on the above method, an embodiment of the present invention further provides a short-wave autonomous direction-finding and positioning system based on the deep deterministic policy gradient algorithm, including: a signal receiving module and a direction-finding and positioning module, where,
[0088] The signal receiving module is used to receive short-wave signal array data by using a number of observation stations, and perform space-time high-resolution processing on the short-wave signal array data to obtain a space-time high-resolution time-frequency diagram;
[0089] The direction-finding and positioning module is used to input the space-time high-resolution time-frequency diagram into a trained short-wave direction-finding and positioning model, and use the short-wave direction-finding and positioning model to obtain the direction-finding and positioning result of the target signal source. The short-wave direction-finding and positioning model is trained based on the deep deterministic policy gradient algorithm. The deep deterministic policy gradient algorithm is used to learn actions in different environmental state spaces based on the state space, action space, and reward function. Among them, the state space of the deep deterministic policy gradient algorithm is set based on the space-time high-resolution time-frequency diagram of the array signal, the action space is constructed based on the executable actions of the rectangular box used to frame the direction-finding and positioning area of the signal source, and the reward function is set according to the signal source positioning accuracy.
[0090] To verify the effectiveness of the solution of this case, the following further explains with experimental data:
[0091] Suppose there are 4 observation stations for positioning a short-wave signal source. The longitude of the first observation station is 116.03° east longitude and the latitude is 34.3° north latitude. The longitude of the second observation station is 104.63° east longitude and the latitude is 35.26° north latitude. The longitude of the third observation station is 122.25° east longitude and the latitude is 33.24° north latitude; the longitude of the fourth observation station is 119.39° east longitude and the latitude is 45.48° north latitude; the longitude of the short-wave target source is 115° east longitude and the latitude is 20° north latitude. Now, uniform circular arrays are installed at all 4 observation stations. The virtual heights of the ionosphere experienced by the short-wave target source signal reaching the above 4 observation stations are 280 km, 300 km, 290 km, and 370 km respectively. It is assumed that there is no ionospheric height error.
[0092] Figure 7 and Figure 8 respectively show the change curves of the average reward and average step length of the A2C and DDPG algorithms for training the intelligent direction-finding and positioning environment. The abscissa represents the number of episodes, and the ordinate represents the average reward and average step length per 20 episodes. The shaded part represents the standard deviation of the average reward and average step length. It can be seen that compared with the A2C algorithm, the DDPG algorithm has a faster convergence speed and a higher average return than the A2C algorithm, but there are fluctuations in the training process. The average step length of the DDPG algorithm using discrete actions and continuous actions can converge within 10 steps, and the convergence speed is faster than that of the A2C algorithm, that is, it reaches a better target with a shorter path.
[0093] Figure 9 It shows the comparison of the average running time between the solution of this case and the traditional two-step direction-finding positioning algorithm under the specified scenario. Figure 10 It shows the comparison between the solution of this case, the traditional two-step direction-finding positioning algorithm and the CRLB under different standard deviations of direction-finding errors. It can be seen from this that the average running time of the traditional two-step direction-finding positioning algorithm is still greater than that of the solution of this case. Under the set positioning scenario, the solution of this case is still close to the CRLB. Therefore, the solution of this case can be positioned under different scenarios and has good generalization ability.
[0094] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0095] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0096] The units and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.
[0097] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software function module. The present invention is not limited to any specific form of combination of hardware and software.
[0098] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the technical field of the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A shortwave autonomous direction finding and positioning method based on a deep deterministic policy gradient algorithm, characterized in that: Include: Using several observation stations to receive shortwave signal array data, and performing space-time high-resolution processing on the shortwave signal array data to obtain space-time high-resolution time-frequency diagrams; The spatial-temporal high-resolution time-frequency graph is input into the trained shortwave direction finding positioning model, and the target signal source direction finding positioning result is obtained by using the shortwave direction finding positioning model. The shortwave direction finding positioning model is trained based on a deep deterministic policy gradient algorithm, and the deep deterministic policy gradient algorithm is used to learn actions in different environmental state spaces based on state space, action space and reward function, wherein the state space of the deep deterministic policy gradient algorithm is set based on the spatial-temporal high-resolution time-frequency graph of the array signal, the action space is constructed based on a rectangular box executable action for selecting a signal source direction finding positioning area, and the reward function is set according to the signal source positioning accuracy.
2. The shortwave autonomous direction finding and positioning method based on the deep deterministic policy gradient algorithm according to claim 1 is characterized in that: Several observatories are used to receive shortwave signal array data, including: The receiver of the observation station is frequency-set according to the direction-finding parameters, and the shortwave signal array data is obtained based on the signal model, wherein the signal model is constructed according to the signal source, the array flow matrix and the array additive noise.
3. The shortwave autonomous direction finding and positioning method based on the deep deterministic policy gradient algorithm according to claim 1 is characterized in that: The action space includes: a discrete action space and a continuous action space. The discrete action space is composed of movement actions in four directions of the rectangular frame: up, down, left, and right. The continuous action space is composed of adjustment actions of the length and / or width of the rectangular frame.
4. The shortwave autonomous direction finding and positioning method based on the deep deterministic policy gradient algorithm according to claim 1 is characterized in that: The reward function is set according to the signal source positioning accuracy, including: The direction finding error is represented by the selected space-time high-resolution time-frequency diagram direction finding data. The position of the shortwave signal source on the earth's surface is estimated based on the direction finding error and the synchronous direction finding intersection of each observation station. The reward function is set according to the area of the convex polygon formed by the intersection points of the synchronous direction finding of the observation stations.
5. The shortwave autonomous direction finding and positioning method based on the deep deterministic policy gradient algorithm according to claim 4 is characterized in that: The direction finding error includes an azimuth observation error and an elevation observation error, and both the azimuth observation error and the elevation observation error are expressed as Gaussian distribution vectors with a mean of 0 and a standard deviation of σ, where: Y over is the overlapping area between the rectangular frame and the signal source area, Y agent is the area of the rectangular frame.
6. The shortwave autonomous direction finding and positioning method based on deep deterministic policy gradient algorithm according to claim 4 is characterized in that: The reward function is expressed as: Where V is the area of the convex polygon, V max is the maximum allowable error area of a convex polygon.
7. The shortwave autonomous direction finding and positioning method based on deep deterministic policy gradient algorithm according to claim 1 is characterized in that: The shortwave direction finding positioning model is trained according to the following steps: Initialize the experience replay pool, value network parameters, and strategy network parameters; The high-resolution spatial-temporal frequency graph is input into the policy network as the current state, and the corresponding rectangular box is obtained when the signal source direction finding and positioning area is selected; the signal source direction finding and positioning area selected by the rectangular box execution action is calculated based on the reward function, and the reward value of the deep deterministic policy gradient algorithm is used to update the next state, and the current state, the corresponding action, the reward value and the next state are stored in the experience replay pool as historical experience, so as to collect historical experience data and store it in the experience replay pool through iterative loops and interact with different environmental states until the maximum iteration loop condition is met, and in each iteration loop, the corresponding action at the next moment is obtained based on the next state and the policy network, and the value network is used to calculate the value corresponding to the next state and the corresponding next action, and the network update gradient update is obtained according to the value; Historical experience is randomly sampled from the experience revisit pool as training samples. The policy network is calculated based on the training samples to predict the value of the current action. The network loss gradient is obtained based on the value and the policy network and value network are updated according to the network loss gradient. The historical experience is returned and resampled until the training convergence conditions are met. The final policy network is used as the shortwave direction finding positioning model.
8. A shortwave autonomous direction finding and positioning system based on a deep deterministic policy gradient algorithm, characterized in that: It includes: a signal receiving module and a direction finding and positioning module, wherein: A signal receiving module is used to receive shortwave signal array data using a number of observation stations, and perform space-time high-resolution processing on the shortwave signal array data to obtain a space-time high-resolution time-frequency diagram; A direction finding and positioning module is used to input the spatial-temporal high-resolution time-frequency diagram into a trained shortwave direction finding and positioning model, and use the shortwave direction finding and positioning model to obtain the direction finding and positioning result of the target signal source. The shortwave direction finding and positioning model is trained based on a deep deterministic policy gradient algorithm. The deep deterministic policy gradient algorithm is used to learn actions in different environmental state spaces based on state space, action space and reward function. The state space of the deep deterministic policy gradient algorithm is set based on the spatial-temporal high-resolution time-frequency diagram of the array signal, the action space is constructed based on a rectangular box executable action for selecting a signal source direction finding and positioning area, and the reward function is set according to the signal source positioning accuracy.
9. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 can be implemented.