Array beam spatial domain synthesis method based on deep reinforcement learning
Through the deep reinforcement learning array beam airspace synthesis method, the optimal amplitude phase distribution of the array beam is directly output, solving the beam integration difficulties during the full airspace scanning of traditional phased array antennas, and achieving high-precision airspace scanning coverage and beam quality improvement.
Patent Information
- Application Number
- CN202510477271.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-01
AI Technical Summary
The beam comprehensive effect of traditional phased array antennas is poor when scanning the entire airspace. The existing methods are time-consuming and labor-intensive and cannot completely eliminate the directional map error caused by the antenna unit due to differences in physical structure, installation position, etc.
Using the array beam airspace synthesis method based on deep reinforcement learning, the array beam comprehensive training platform is built, and the TD3 algorithm and deep reinforcement learning training model are used to directly output the optimal amplitude phase distribution of the array beam to achieve high-precision airspace scanning.
Achieve high-precision beam scanning coverage in the entire airspace, significantly improving the problem of large-angle beam deterioration without wave-by-wave-by-wave-by-wave correction. It is suitable for various array types and improves the beam quality of the airspace coverage.
Smart Images

Figure CN120409216A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of phased array beam synthesis, and in particular to an array beam spatial domain synthesis method based on deep reinforcement learning. Background Art
[0002] Phased array antennas feature flexible and fast beam scanning, possess strong vitality, and have been widely used in fields such as radar and communications. In specific applications, phased array antennas need to be calibrated first to eliminate the amplitude and phase differences between array element channels, reducing the impact of inter-channel amplitude and phase errors on the main lobe gain, main-to-sidelobe ratio, pointing accuracy, and other performance of the array element beam, thereby obtaining a high-precision pointing beam. However, in engineering projects, array calibration is usually only performed in the normal direction, and the normal correction data is used for the rest of the spatial beam synthesis. This can only eliminate the amplitude and phase errors between channels and in the antenna normal direction, but cannot eliminate the amplitude and phase deviations of antenna units in the entire spatial domain, as well as the quantitative deviations of components such as amplitude and phase control. This results in a deterioration in the beam synthesis effect in the rest of the spatial domain, especially in the large-angle spatial domain.
[0003] In order to obtain high-precision beam scanning coverage in the entire airspace, the existing solution is to perform full airspace wave position correction, or to perform wave position correction at fixed intervals within the airspace range and then fit the correction value of all coverage angles. This is not only time-consuming and labor-intensive, but also requires a large amount of storage space for multi-wave position correction data storage and indexing. Therefore, researchers have proposed some beam correction error compensation methods, which perform phase compensation through recursive compensation feed phase, interpolation fitting, comparative rounding, etc. However, these methods are all aimed at the problem of phased array beam deterioration due to the quantization of wave control devices, and cannot completely eliminate the directional pattern error caused by differences in antenna units due to physical structure, installation position, process processing, shell deformation, etc. In addition, relevant researchers have also proposed to directly perform array beam comprehensive optimization through optimization algorithms such as genetic algorithms and particle swarm optimization, and have achieved good results. However, these optimization algorithms have problems such as premature convergence or easy to fall into local optimal solutions. Summary of the Invention
[0004] The present invention aims to provide an array beam spatial integration method based on deep reinforcement learning to solve the problem of large workload and difficulty in beam correction for full spatial scanning of traditional phased array systems, while the conventional normal correction cannot solve the problem of beam deterioration during spatial scanning.
[0005] The present invention provides an array beam spatial domain synthesis method based on deep reinforcement learning, comprising:
[0006] Build an array beam comprehensive training platform;
[0007] Construct mathematical equations for array beam synthesis optimization problems, determine deep reinforcement learning training data and array beam spatial domain synthesis network models;
[0008] Update the training data at fixed angle intervals according to the airspace scanning angle range, perform the training and learning of the network model parameters, and output the trained array beam airspace synthesis network model;
[0009] Call the trained array beam airspace synthesis network model to perform wave position scanning in the array airspace to obtain the airspace scanning coverage beam.
[0010] In some embodiments, the array beam synthesis training platform includes an agent and an environment;
[0011] The agent is a deep reinforcement learning algorithm program responsible for task decision-making;
[0012] The environment includes an array antenna, a directive antenna, and a vector network analyzer. The excitation signal is generated by the vector network analyzer, and the environment completes the task execution;
[0013] When performing array beam synthesis training, set the state s k as the array beam pointing angle and gain, and the action a k as the array amplitude and phase distribution. Execute the action in the environment, that is, control the array to synthesize the corresponding beam according to the array phase distribution. At the same time, the excitation signal is generated by the vector network analyzer and the array beam gain is measured to obtain the next state s k+1 , and calculate the reward r k according to the array beam gain. Maximize the cumulative reward to obtain the array amplitude and phase distribution of the desired beam, thereby realizing the beam synthesis of the array at different airspace angles.
[0014] In some embodiments, the construction of the mathematical solution equation for the array beam synthesis optimization problem includes: using the TD3 algorithm for beam synthesis intelligent optimization, adjusting the array beam by finding the optimal amplitude-phase control strategy to minimize the objective function used to calculate the difference between the actual beam performance and the desired requirements, that is, the TD3 algorithm needs to make an optimal action strategy according to the current desired beam performance, determine the amplitude and phase distribution of each array element, so that the array forms an array beam with optimal performance in the target direction
[0015] In some embodiments, the determination of the deep reinforcement learning training data includes: defining the training data in each iteration of the TD3 algorithm, that is, setting the state space, action space, and immediate reward involved in the deep reinforcement learning. The training data for each iteration consists of the current state, action, next action, and immediate reward.
[0016] In some embodiments, the determination of the deep reinforcement learning training data includes:
[0017] S21. The state space is the set of all states for each training. The real-time gain G at the k-th iteration of each training is k and the beam pointing is defined as the current state s k , where the beam pointing includes the azimuth angle θ s and the elevation angle . The corresponding state space is
[0018] S22. Consider a two-dimensional phased array of M×N. The corresponding action space is expressed as:
[0019]
[0020] where represents the action at the k-th iteration, which is the random amplitude distribution value and phase distribution value of the array antenna. The action dimension is the number of array elements M×N, that is: is the amplitude of the array element antenna at the k-th moment, is the phase of the array element antenna at the k-th moment,
[0021] S23. The immediate reward is obtained by calculating the beam gain. By comparing the real-time beam gain G k at the k-th iteration and the expected gain G of the target direction s , the formula for the immediate reward r k is as follows:
[0022]
[0023] where ε is the gain discount factor. During training, when the beam gain gradually approaches the expected gain, the reward increases. When G k ≥εG s , terminate this training episode, output the array phase distribution value at this time to control the array synthesized beam, and give a positive value of 1 to strengthen the training reward.
[0024] In some embodiments, the array beam spatial domain synthesis network model is a TD3 network model. The TD3 network model includes 1 actor network, 1 actor target network, 2 critic networks and 2 critic target networks. Among them, the actor network is a policy network for policy function estimation, and the critic network is a Q-value network for Q-value estimation.
[0025] In some embodiments, the training data is updated at fixed angles at intervals according to the spatial domain scanning angle range, including: the state in the training data participating in the training of the array beam spatial domain synthesis network model consists of beam gain and beam pointing. In each round of training iteration, spatial domain beam synthesis training is performed by incrementally covering the beam scanning range at a preset fixed angle interval.
[0026] In some embodiments, performing network model parameter training and learning, and outputting the trained array beam spatial domain synthesis network model, including: training the TD3 network model, and the specific training process includes: network parameter initialization, network parameter training and update, and outputting the trained array beam spatial domain synthesis network model.
[0027] In some embodiments, the network parameter initialization includes: initializing 1 actor network parameter and 2 critic network parameters in the TD3 network model with random parameters, and copying the initialized parameters to the corresponding 1 actor target network and 2 critic target networks for parameter initialization, that is, the target network and the corresponding network use the same network parameters, and at the same time initializing the experience buffer D.
[0028] In some embodiments, the network parameter training and update includes: performing episode training on the network parameters, and updating the network parameters by calculating the loss function and the gradient, specifically including:
[0029] S31, setting training parameters, including the number of episodes, the number of training iterations per episode, the learning rate, the discount factor, and Gaussian noise;
[0030] S32, randomly generating beam pointing angles and array amplitude-phase distribution values, and obtaining the beam gain as the initialized state and action;
[0031] S33, adding exploration noise to the action output by the actor network to select an action, observing the reward and the next moment state feedback by the environment, and forming training data to be stored in the experience buffer D;
[0032] S34, repeatedly executing step S33. When the training episode is greater than the preset exploration times, sampling M training data from the experience buffer D for network training. The action is output by the actor target network, and then the smaller Q value is selected from the target Q values output by the two critic target networks to calculate the target value, and the gradient of the loss function is calculated. The two critic network parameters are updated according to the gradient descent to minimize the loss function;
[0033] S35, after the critic network is updated for d steps, updating the actor network parameters by maximizing the critic network Q value output through the gradient ascent algorithm; at the same time, updating the corresponding target network parameters;
[0034] S36. Repeat steps S33 - S35. Each time a training round is executed, the iteration count is incremented by 1 until the iteration count reaches the set number of iteration for round training, or when the beamforming gain approaches the desired beam gain value, the current training round terminates.
[0035] S37. Repeat steps S32 - S36 until the number of round training is reached. The network training ends and outputs the trained model.
[0036] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are as follows:
[0037] 1. Aiming at the problem that traditional normal correction cannot improve the deterioration of the spatial domain beam, the present invention adopts an intelligent optimization method for array beam synthesis based on deep reinforcement learning. Instead of correcting each wave position, it utilizes the powerful decision-making optimization ability and non-linear fitting ability of the deep reinforcement learning algorithm. Through the reinforcement learning working mechanism of perception - action - reward, it directly outputs the amplitude and phase distribution parameters of the array elements of the scanning beam in the current electromagnetic environment, automatically compensates for the amplitude and phase errors between elements without correction, and can obtain high-precision array spatial domain scanning beams at any angle in the entire airspace, and can significantly improve the problem of beam deterioration at large angles in the spatial domain.
[0038] 2. The present invention can be used as a new means for array beam correction or synthesis. Moreover, the method of the present invention is data-driven and independent of the antenna layout configuration, and can be extended and applied to non-regular arrays such as conformal arrays, curved surface arrays, and sparse arrays, with high flexibility and universality, and has strong engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of an array beam spatial domain synthesis method based on deep reinforcement learning provided by an embodiment of the present invention.
[0040] Figure 2 is a schematic diagram of an array beam synthesis training platform built in an embodiment of the present invention.
[0041] Figure 3 is a flowchart of performing network model parameter training and learning in an embodiment of the present invention.
[0042] Figure 4 is a graph of the cumulative reward of the training round of the array beam spatial domain synthesis method based on deep reinforcement learning in an embodiment of the present invention.
[0043] Figure 5 is a graph of the beam switching time of the test round of the array beam spatial domain synthesis method based on deep reinforcement learning in an embodiment of the present invention.
[0044] Figure 6 is a graph of the test spatial domain scanning beam of the array beam spatial domain synthesis method based on deep reinforcement learning in an embodiment of the present invention.
[0045] Figure 7 This is a comparison diagram of the beam and the normal correction beam of the array beam spatial domain synthesis method based on deep reinforcement learning in the embodiments of the present invention. Specific implementation manners
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0047] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] The embodiments of the present invention propose an array beam spatial domain synthesis method based on deep reinforcement learning. Without performing spatial domain wave position correction, by constructing a deep reinforcement learning algorithm training model, using the non-linear fitting and decision optimization capabilities of the deep reinforcement learning algorithm, directly give the optimal array element amplitude-phase distribution value of the expected beam at the current spatial domain scanning angle of the array, perform amplitude-phase regulation on the array to synthesize the expected pointing beam, so as to obtain a high-precision spatial domain scanning coverage beam.
[0049] As Figure 1 shown, the array beam spatial domain synthesis method based on deep reinforcement learning includes the following steps:
[0050] S101, build an array beam synthesis training platform;
[0051] S102, construct a mathematical solution equation for the array beam synthesis optimization problem, and determine the deep reinforcement learning training data and the array beam spatial domain synthesis network model;
[0052] S103, update the training data at fixed angles according to the spatial domain scanning angle range interval, perform network model parameter training and learning, and output the trained array beam spatial domain synthesis network model;
[0053] S104, call the trained array beam spatial domain synthesis network model to perform array spatial domain wave position scanning, and obtain a spatial domain scanning coverage beam.
[0054] In some embodiments, as Figure 2As shown, the array beam synthesis training platform in step S101 includes an agent (decision area) and an environment (radiation area). The agent is a deep reinforcement learning algorithm program responsible for task decision-making; the environment includes an array antenna, a directive antenna, and a vector network analyzer. The excitation signal is generated by the vector network analyzer, and the environment completes task execution. When performing array beam synthesis training, the state s is set k as the array beam pointing angle and gain, and the action a k is the array amplitude and phase distribution. Executing the action in the environment, that is, controlling the array to synthesize the corresponding beam according to the array phase distribution. At the same time, the vector network analyzer generates an excitation signal and measures the array beam gain to obtain the next state s k+1 and calculates the reward r k based on the array beam gain, and maximizes the cumulative reward to obtain the array amplitude and phase distribution of the desired beam, thereby realizing the beam synthesis of the array at different spatial angles.
[0055] In some embodiments, constructing the mathematical solution equation of the array beam synthesis optimization problem in step S102 includes: using the TD3 algorithm for beam synthesis intelligent optimization, adjusting the array beam by finding the optimal amplitude-phase control strategy to minimize the objective function for calculating the difference between the actual beam performance and the desired requirements, that is, the TD3 algorithm needs to make an optimal action strategy according to the current desired beam performance, determine the amplitude and phase distribution of each array element, so that the array forms an array beam with optimal performance in the target direction The mathematical solution equation of this array beam synthesis optimization problem can be expressed as follows:
[0056]
[0057] [[ID=z19]]where K is the number of training iterations, G k is the real-time gain of the array beam pointing to the target direction at the kth iteration, θ s is the azimuth angle, is the elevation angle, and G s is the target gain that is expected to be achieved at this beam pointing.
[0058] In some embodiments, determining the deep reinforcement learning training data in step S102 includes: defining the training data in each iteration of the TD3 algorithm, that is, setting the state space, action space, and immediate reward involved in deep reinforcement learning. The training data for each iteration is composed of the current moment state, action, next moment action, and immediate reward. Specifically as follows:
[0059] S21, the state space is the set of all states in each training. The real-time gain G kThe sum beam pointing is defined as the current state s k , where the beam pointing includes the azimuth angle θ s and the elevation angle The corresponding state space is as follows:
[0060]
[0061] S22. Considering a two-dimensional phased array of M×N, the corresponding action space can be expressed as:
[0062]
[0063] where represents the action at the k-th iteration moment, which is the random amplitude distribution value and phase distribution value of the array antenna. The action dimension is the number of array elements M×N, that is: m∈M, n∈N. is the amplitude of the array element antenna at the k-th moment, is the phase of the array element antenna at the k-th moment,
[0064] S23. The immediate reward is obtained by calculating the beam gain. By comparing the real-time beam gain G k at the k-th iteration moment and the expected gain G of the target direction s , the expected gain G s is obtained through antenna three-dimensional simulation software and on-site calibration. The calculation formula of the immediate reward r k is as follows:
[0065]
[0066] where ε is the gain discount factor. During training, when the beam gain gradually approaches the expected gain, the reward increases. When G k ≥εG s , terminate this training episode, output the array phase distribution value at this time to control the array synthesized beam, and give a positive value of 1 to strengthen the training reward.
[0067] In some embodiments, the array beam spatial domain synthesis network model in step S102 is a TD3 network model. The TD3 network model includes 1 actor network, 1 actor target network, 2 critic networks and 2 critic target networks. Among them, the actor network is a policy network for policy function estimation, and the critic network is a Q-value network for Q-value estimation.
[0068] In some embodiments, in step S103, the training data is updated at fixed - angle intervals according to the spatial - domain scanning angle range, including: The state in the training data participating in the training of the array beam spatial - domain synthesis network model consists of beam gain and beam direction. In each round of training iteration, spatial - domain beam synthesis training is performed by incrementally covering the beam scanning range at a preset fixed - angle interval.
[0069] In some embodiments, as Figure 3 shown, in step S103, network model parameter training and learning are performed, and the trained array beam spatial - domain synthesis network model is output, including: Training the TD3 network model. The specific training process includes: network parameter initialization, network parameter training and update, and outputting the trained array beam spatial - domain synthesis network model.
[0070] The above - mentioned network parameter initialization includes: Initializing 1 actor network π φ parameters φ and 2 critic networks parameters θ1, θ2 with random parameters, and copying the initialized parameters to the corresponding 1 actor target network and 2 critic target networks for parameter initialization: θ1'←θ1, θ2'←θ2, φ'←φ, that is, the target network and the corresponding network use the same network parameters, and at the same time, the experience buffer D is initialized.
[0071] The above - mentioned network parameter training and update include: Performing episode training on the network parameters, and updating the network parameters by calculating the loss function and gradient, specifically as follows:
[0072] S31, Set training parameters, including the number of episodes T, the number of training iterations K per episode, the learning rate τ, the discount factor r, and Gaussian noise n~N(0,σ), where N is a normal distribution with a mean of 0 and a standard deviation of σ;
[0073] S32, Randomly generate beam - pointing angles and array amplitude - phase distribution values, and obtain the beam gain, as the initialized state s0 and action a0;
[0074] S33, Select an action by adding exploration noise to the action output by the actor network, that is, a k =π φ (s)+n, Observe the reward r k fed back by the environment and the next - moment state s k+1 , and form the training data (s k ,a k ,r k ,s k+1 ) and store it in the experience buffer D;
[0075] S34. Repeat step S33. When the number of training rounds is greater than the preset exploration times, sample M training data (s, a, r, s') from the experience buffer D for network training. The actor target network outputs an action. Then, select the smaller Q value from the target Q values output by the two critic target networks to calculate the target value, i.e.: And calculate the gradient of the loss function. Update the parameters of the two critic networks by minimizing the loss function according to gradient descent. The formula is as follows:
[0076]
[0077] S35. After the critic network is updated d steps, update the actor network parameters by maximizing the Q value output of the critic network through the gradient ascent algorithm. The formula is as follows:
[0078]
[0079] Meanwhile, update the corresponding target network parameters, i.e.: θ1'←τθ1+(1 - τ)θ1', θ2'←τθ2+(1 - τ)θ2', φ'←τφ+(1 - τ)φ'.
[0080] S36. Repeat steps S33 - S35. Each time a training round is executed, the iteration count is incremented by 1 until the iteration count reaches the set number of round training iterations, i.e.: k = K, or when the beamforming gain is close to the desired beam gain value, i.e., when the current iteration termination condition G k ≥εG s is satisfied, the current training round terminates;
[0081] S37. Repeat steps S32 - S36 until the number of round training reaches T. The network training ends and outputs the trained model.
[0082] In some embodiments, in step S104, calling the trained array beam spatial domain synthesis network model for array spatial domain wave - by - wave scanning includes: After obtaining the trained array beam spatial domain synthesis network model, any given scanning angle can be input into the trained array beam spatial domain synthesis network model, and the optimal amplitude - phase distribution of the corresponding desired beam can be directly output, thereby realizing phased array beam synthesis and spatial domain scanning. At the same time, since the spatial domain beams are all synthesized according to the optimal amplitude - phase distribution output by the trained array beam spatial domain synthesis network model, the problem of spatial domain scanning beam deterioration in engineering applications can be significantly improved, and the quality of the spatial domain coverage beam can be enhanced.
[0083] A specific example:
[0084] Refer to Figure 1, an array beam synthesis training platform is built. The array is an 8-element linear array, and the test frequency point is 12 GHz. The array synthesized beam is controlled according to the initial scanning angle and phase distribution. An excitation signal is generated by a vector network analyzer and connected to the array for signal transmission. After the indicator antenna receives the signal, it is connected to the vector network for signal acquisition to obtain the beam gain, and the training reward is calculated based on the beam gain and the expected gain. Among them, the expected gain can be obtained through the array normal calibration gain and the antenna three-dimensional software simulation.
[0085] The beam scanning range is set to [-40°, 40°]. During training, the beam switches between -40° and 40°, and the training angle interval has a certain fixed step. The scanning angles participating in the training are -40°, -35°, -25°, -15°, -5°, 0°, 5°, 15°, 25°, 35°, 40° in turn. Each scanning of an angle is a round of training. In each training round, when the beam pointing gain reaches the expected gain, the training round terminates. The number of training rounds is 1000, and the number of iterations per round is 10000. After training is completed, a beam scanning switching test is carried out. The beam scanning angle step is 2.5°, and the number of test rounds is 200.
[0086] Plot the cumulative reward of the training rounds, as Figure 4 shown, and plot the beam scanning switching time of the test rounds, as Figure 5 shown. It can be seen that when the number of training rounds is greater than 200, it basically reaches convergence, and the cumulative reward approaches 1. During testing, the beam is scanned according to a 2.5° scanning step, and beam synthesis is carried out according to the amplitude-phase distribution values given by the trained array beam spatial domain synthesis network model. The beam switching time can reach the millisecond level.
[0087] Plot the scanned beam of the spatial domain synthesis as Figure 6 shown, where the solid line plotted beam is the pointing beam participating in the model training, and the dashed line plotted beam is the pointing beam not participating in the model training. Through the plotted spatial domain scanned beam, the feasibility and correctness of the array beam spatial domain synthesis method based on deep reinforcement learning can be verified. This method can directly output the optimal amplitude-phase distribution values of the spatial domain scanned pointing beam without performing wave position-by-wave correction, and can comprehensively obtain the spatial domain scanned coverage beam.
[0088] Furthermore, to verify the performance of the array beam spatial domain synthesis method based on deep reinforcement learning of the present invention and its improvement of the spatial domain angle beam, the synthesized beams in the directions of 0°, 10°, 20°, 30°, and 40° are respectively selected. Among them, 0° and 40° are the beam angles participating in the training, and 10°, 20°, and 30° are the beam angles not participating in the model training. Compare the beam synthesis results with the beams obtained by conventional normal calibration, as Figure 7As shown, where the solid line is the comprehensive radiation pattern of the method of the present invention, and the dashed line is the normal correction comprehensive radiation pattern.
[0089] In this test, the array is a transmitting array, and the performance of the transmitting beam mainly considers the main lobe gain. It can be seen that the method of the present invention can obtain the desired beam in the specified direction. In the normal direction, the beam synthesis effect is equivalent to that of the normal correction comprehensive beam. However, as the beam scanning angle increases, the synthesis effect of the present invention has a certain performance improvement in the main lobe gain compared with the normal correction comprehensive beam, and the improvement of the beam performance for large angles is more obvious. Table 1 lists the main lobe gain values of the method of the present invention (intelligent optimization) and the normal correction test beam at different scanning angles.
[0090] Table 1, Comparison table of the main lobe gain of the method of the present invention and the normal correction test:
[0091]
[0092] As can be seen from the above table, except for the normal direction, compared with the normal correction, the method of the present invention has a gain improvement of about 0.5 dB in the directions of 10°, 20°, and 30°, and the gain improvement can reach 1.5 dB in the 40° direction, which is very helpful for improving the overall radiation power level of the system airspace scan. Therefore, it can be verified that the array beam airspace synthesis method based on deep reinforcement learning of the present invention can significantly improve the problem of beam deterioration in airspace scanning.
[0093] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An array beam spatial domain synthesis method based on deep reinforcement learning, characterized in that Including: Building an array beam synthesis training platform; Constructing a mathematical solution equation for the array beam synthesis optimization problem to determine the deep reinforcement learning training data and the array beam spatial domain synthesis network model; Updating the training data at fixed angles according to the spatial scanning angle range interval, performing network model parameter training and learning, and outputting the trained array beam spatial domain synthesis network model; Invoking the trained array beam spatial domain synthesis network model to perform wave-by-wave scanning in the array space to obtain the spatial scanning coverage beam.
2. The method for array beam spatial domain synthesis based on deep reinforcement learning according to claim 1, wherein The array beam synthesis training platform includes an agent and an environment; The agent is a deep reinforcement learning algorithm program responsible for task decision-making; The environment includes an array antenna, a directive antenna, and a vector network analyzer. The excitation signal is generated by the vector network analyzer, and the environment completes task execution; When performing array beam synthesis training, set the state s k as the array beam pointing angle and gain, and the action a k as the array amplitude and phase distribution. Execute the action in the environment, that is, control the array to synthesize the corresponding beam according to the array phase distribution. At the same time, generate an excitation signal by the vector network analyzer and measure the array beam gain to obtain the next state s k+1 and calculate the reward r based on the array beam gain k Maximize the cumulative reward to obtain the array amplitude and phase distribution of the desired beam, so as to realize the beam synthesis of the array at different spatial angles.
3. The method for array beam spatial domain synthesis based on deep reinforcement learning according to claim 1, wherein The mathematical solution equation for the constructed array beam synthesis optimization problem includes: using the TD3 algorithm for beam synthesis intelligent optimization, adjusting the array beam by finding the optimal amplitude-phase control strategy to minimize the objective function for calculating the difference between the actual beam performance and the desired requirements, that is, the TD3 algorithm needs to make an optimal action strategy according to the current desired beam performance, determine the amplitude-phase distribution of each array element, so that the array forms an array beam with optimal performance in the target direction form an array beam with optimal performance.
4. The array beam spatial domain synthesis method based on deep reinforcement learning according to claim 1, characterized in that, The determination of the deep reinforcement learning training data includes: for the TD3 algorithm, defining the training data in each round of iteration, that is, setting the state space, action space, and immediate reward involved in deep reinforcement learning. The training data for each iteration at each moment is jointly composed of the current moment state, action, next moment action, and immediate reward.
5. The array beam spatial domain synthesis method based on deep reinforcement learning according to claim 4, characterized in that The determination of the deep reinforcement learning training data includes: S21, the state space is the set of all states for each training. Define the real-time gain G at the k-th iteration of each training k and the beam pointing as the current state s k , where the beam pointing includes the azimuth angle θ s and the elevation angle The corresponding state space is S22, considering a two-dimensional phased array of M×N, the corresponding action space is expressed as: Among them, represents the action at the k-th iteration moment, which is the random amplitude distribution value and phase distribution value of the array antenna. The dimension of the action is the number of array elements M×N, that is: m∈M, n∈N; is the amplitude of the array element antenna at the k-th moment, is the phase of the array element antenna at the k-th moment, S23, the immediate reward is obtained by beam gain calculation. By comparing the real-time beam gain G at the k-th iteration moment k and the target direction of the expected gain G s , the calculation formula of the immediate reward r k is as follows: where ε is the gain discount factor; during training, when the beam gain gradually approaches the desired gain, the reward increases. When G k ≥ εG s the training episode is terminated, the array phase distribution value at this time is output to control the array synthesized beam, and a positive value of 1 is given to strengthen the training reward.
6. The array beam spatial domain synthesis method based on deep reinforcement learning according to claim 1, wherein, The array beam spatial domain synthesis network model is a TD3 network model. The TD3 network model includes 1 actor network, 1 actor target network, 2 critic networks, and 2 critic target networks. Among them, the actor network is a policy network for policy function estimation, and the critic network is a Q-value network for Q-value estimation.
7. The method for array beam spatial domain synthesis based on deep reinforcement learning according to claim 1, wherein The update of the training data at fixed angles according to the spatial scanning angle range interval includes: the state in the training data participating in the training of the array beam spatial domain synthesis network model is composed of beam gain and beam pointing. In each round of training iteration, spatial beam synthesis training is performed by incrementing the coverage beam scanning range at a preset fixed angle interval.
8. The array beam spatial domain synthesis method based on deep reinforcement learning according to claim 6, characterized in that, The execution of network model parameter training and learning, and the output of the trained array beam spatial domain synthesis network model includes: training the TD3 network model. The specific training process includes: network parameter initialization, network parameter training and update, and output of the trained array beam spatial domain synthesis network model.
9. The method for array beam spatial domain synthesis based on deep reinforcement learning according to claim 8, wherein The network parameter initialization includes: initializing the parameters of 1 actor network and 2 critic networks in the TD3 network model with random parameters, and copying the initialized parameters to the corresponding 1 actor target network and 2 critic target networks for parameter initialization, that is, the target network and the corresponding network use the same network parameters, and at the same time initializing the experience buffer D.
10. The method for array beam spatial domain synthesis based on deep reinforcement learning according to claim 9, wherein, The network parameter training and update includes: performing episode training on the network parameters, and updating the network parameters by calculating the loss function and gradient. Specifically, it includes: S31, setting training parameters, including the number of episodes, the number of iteration times in episode training, the learning rate, the discount factor, and Gaussian noise; S32, randomly generating beam pointing angles and array amplitude-phase distribution values, and obtaining the beam gain as the initialized state and action; S33. Add exploration noise to the action output by the actor network to select an action, observe the reward and the state at the next moment feedback by the environment, and store them as training data in the experience buffer D; S34. Repeat step S33. When the number of training rounds is greater than the preset number of explorations, sample M training data from the experience buffer D for network training. Output an action by the actor target network, then select the smaller Q value from the target Q values output by the two critic target networks to calculate the target value, and calculate the gradient of the loss function. Update the parameters of the two critic networks by minimizing the loss function according to gradient descent; S35. After the critic network is updated d steps, update the parameters of the actor network by maximizing the Q value output of the critic network through the gradient ascent algorithm; meanwhile, update the corresponding target network parameters; S36. Repeat steps S33 - S35. Each time a training round is executed, the iteration count is incremented by 1 until the iteration count reaches the set number of training round iterations, or when the beamforming gain approaches the desired beam gain value, the current training round terminates; S37. Repeat steps S32 - S36 until the number of training rounds is reached, the network training ends and the trained model is output.
Citation Information
Cited By
Reinforced learning-based phased array phase control method and device
CN121710974A
Phased array phase control method and device based on reinforcement learning
CN121710974B