Phased array phase control method and device based on reinforcement learning

By using reinforcement learning and an improved particle swarm optimization algorithm, the problems of high computational complexity and low policy network update efficiency in traditional phased array phase control methods are solved, achieving efficient and accurate phased array phase control, adapting to complex communication environments, and improving the performance of communication systems.

CN121710974BActive Publication Date: 2026-05-08SICHUAN BOPU MICROWAVE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN BOPU MICROWAVE TECH CO LTD
Filing Date
2026-02-12
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Traditional phased array phase control methods have high computational complexity, making it difficult to meet the requirements of low-latency communication. Furthermore, existing reinforcement learning methods have slow convergence speeds during policy network updates and are prone to getting trapped in local optima, resulting in insufficient beam control accuracy.

Method used

A phased array phase control method based on reinforcement learning is adopted. The optimal policy is learned through the interaction between the policy network and the environment. The updated samples are updated by combining the experience replay pool and the first-in-first-out policy. An improved particle swarm optimization algorithm is used to improve the policy network update effect, including nonlinear adjustment weight learning, fuzzy information exchange and sinusoidal Lévy flight policy, so as to achieve rapid global exploration and precise control.

Benefits of technology

It improves the adaptability and accuracy of phased array phase control, enabling it to adapt to environmental changes and device aging, significantly enhances communication performance, reduces computational latency, and improves beam pointing control accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121710974B_ABST
    Figure CN121710974B_ABST
Patent Text Reader

Abstract

The application discloses a phased array phase control method based on reinforcement learning, belongs to the technical field of wireless communication, and controls the phased array phase based on reinforcement learning through a deep reinforcement learning algorithm, can effectively improve control adaptability and accuracy, can adapt the control of the phased array to environmental changes and device aging, and further adopts an improved particle swarm algorithm to update a strategy network, improves the updating effect of the strategy network, and further enhances the phase control accuracy of the phased array.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication technology, specifically relating to a phase control method and apparatus for phased arrays based on reinforcement learning. Background Technology

[0002] Traditional phased array phase control methods typically rely on accurate Channel State Information (CSI) and utilize beamforming algorithms (such as codebook search or convex optimization methods) to calculate optimal phase weights. However, in real-world communication environments, channel states are often time-varying and difficult to obtain accurately, especially in high-speed mobile scenarios, where channel estimation feedback introduces significant overhead and latency. Furthermore, traditional optimization algorithms are generally computationally complex and struggle to meet the demands of low-latency communication. Therefore, this invention employs reinforcement learning to improve control accuracy. Reinforcement learning learns the optimal policy through agent-environment interaction and reward feedback, requiring no prior knowledge of the precise channel model and exhibiting strong adaptability. However, existing reinforcement learning methods often suffer from slow convergence and susceptibility to local optima during training, particularly in the parameter update phase of the policy network. Traditional gradient descent methods have limited search capabilities when facing complex non-convex optimization problems, resulting in low beam pointing control accuracy of phased arrays and limited improvements in communication performance. Summary of the Invention

[0003] This invention provides a phased array phase control method and apparatus based on reinforcement learning, which solves the problem that traditional optimization algorithms are usually computationally complex and difficult to meet the requirements of low-latency communication. At the same time, it solves the problems of low policy network update efficiency and easy getting trapped in local optima in existing reinforcement learning technology, resulting in insufficient beam control accuracy, and further improves the control accuracy.

[0004] On one hand, the present invention provides a phase control method for a phased array based on reinforcement learning, comprising:

[0005] The real-time status of the phased array is collected, and based on the real-time status, a strategy network is used to select the phase control actions of each antenna element of the phased array from the action space.

[0006] Based on the phase control actions of each antenna element of the phased array, the beam pointing of the phased array is adjusted, and the adjusted phased array is used to obtain communication performance indicators.

[0007] Rewards are obtained based on the communication performance indicators, and the state at the next moment is obtained. Based on the real-time state, phase control actions, rewards, and the state at the next moment, an experience sample is constructed.

[0008] An experience replay pool is constructed using experience samples, and the experience samples in the experience replay pool are updated using a first-in-first-out strategy.

[0009] Once the number of experience samples in the experience replay pool reaches a preset number, multiple experience samples are randomly sampled from the experience replay pool, and the policy network is updated using the sampled experience samples.

[0010] The phase of the phased array is controlled by the updated policy network, thus completing the phase control of the phased array based on reinforcement learning.

[0011] Furthermore, the real-time status includes channel state information fed back by the receiver, the received signal-to-noise ratio, and / or the phase control action of the previous moment.

[0012] Furthermore, the phase control action includes the phase offset of each antenna element in the phased array.

[0013] Furthermore, obtaining rewards based on the aforementioned communication performance metrics includes:

[0014] ;

[0015] In the formula, The reward at the current time t, Let be the communication performance index at the current time t. Let be the phase vector at the current time t. Let be the phase vector of the previous time t-1. For communication weighting coefficients, These are the smoothing weighting coefficients.

[0016] Furthermore, the policy network is updated using sampled empirical samples, including:

[0017] Based on the empirical samples, the network parameters of the policy network are updated using the gradient descent method.

[0018] Furthermore, the policy network is updated using sampled empirical samples, including:

[0019] Based on the network parameters of the policy network, a particle swarm is generated;

[0020] For any particle in the particle swarm, the loss function value corresponding to the particle is obtained based on the sampled empirical samples, and the particle with the smallest loss function value is determined as the instantaneous optimal particle.

[0021] Based on the instantaneous optimal particle, a dual-optimal learning strategy improved by nonlinearly adjusting the weight learning coefficient is used to quickly explore the solution space of the particles in the particle swarm, and determine the particles after the quick exploration of the solution space.

[0022] A nonlinear information exchange strategy with improved fuzzy information content is used to perform enhanced exploration of the solution space for the particles after rapid exploration of the solution space, and the particles after enhanced exploration of the solution space are determined.

[0023] A time-adaptive, improved sinusoidal Levy flight strategy is used to perform a global greedy exploration of the particles after the enhanced exploration of the solution space, thereby determining the particles after the global greedy exploration.

[0024] Obtain the number of iterations and determine whether the number of iterations is greater than or equal to the preset maximum number of iterations. If so, determine the globally optimal particle based on the particles after the global greedy exploration, and use the network parameters in the globally optimal particle as the network parameters of the policy network to complete the update. Otherwise, return to the step of obtaining the instantaneous optimal particle.

[0025] Furthermore, based on the instantaneously optimal particle, a dual-optimal learning strategy improved by nonlinearly adjusting the weight learning coefficients is used to rapidly explore the solution space of the particles in the particle swarm, determining the particles after the rapid exploration of the solution space, including:

[0026] The nonlinear adjustment function is constructed as follows:

[0027] ;

[0028] In the formula, It is a nonlinear adjustment function. x Let e ​​be a variable, and let e be a natural constant.

[0029] Based on the aforementioned nonlinear adjustment function, the velocity adjustment weight, the first learning factor, and the second learning factor are obtained as follows:

[0030] ;

[0031] ;

[0032] ;

[0033] In the formula, Adjust the weights for speed. As the first learning factor, As the second learning factor, To adjust the parameters, This represents the number of iterations already performed. This represents the maximum number of iterations.

[0034] Based on the velocity adjustment weight, the first learning factor, and the second learning factor, the particles in the particle swarm are subjected to a rapid exploration of the solution space to determine the particles after the rapid exploration of the solution space:

[0035] ;

[0036] ;

[0037] In the formula, For the first iter During the nth iteration i One particle, For the first i Particles after rapid exploration of the solution space. For the first iter During the +1st iteration, the... i The update rate of each particle For the first iter During the nth iteration i The update rate of each particle i =1,2,...,M, where M is the total number of particles. The first random number between (0,1) The second random number between (0,1) for The historical best value, It is the instantaneously optimal particle.

[0038] Furthermore, a nonlinear information exchange strategy with improved fuzzy information content is employed to perform enhanced solution space exploration on the particles after the rapid exploration of the solution space, determining the particles after the enhanced solution space exploration, including:

[0039] Obtain the loss function value of the particles after the fast exploration of the solution space, and obtain the degree of belonging of the particles after the fast exploration of the solution space relative to the instantaneous optimal particle based on the loss function value;

[0040] ;

[0041] In the formula, For the first j The degree of affiliation of a particle after a rapid exploration of the solution space. The loss function value of the instantaneously optimal particle. Let be the maximum loss function value corresponding to the particle after fast exploration of all solution spaces. Let be the loss function value of the particle after fast exploration of the j-th solution space;

[0042] Based on the aforementioned affiliation degree, the degree to which a particle, after a rapid exploration of the solution space, is not located at the current optimal position of the instantaneously optimal particle is determined as follows:

[0043] ;

[0044] In the formula, For the first jThe degree to which a particle, after a rapid exploration of the solution space, is no longer in the current optimal position of the instantaneously optimal particle. It is an adjustable parameter, and its value is between [0.8, 1].

[0045] Based on the degree of affiliation and the degree of affiliation, the amount of fuzzy information obtained is as follows:

[0046] ;

[0047] In the formula, K For fuzzy information content, max is the function to find the maximum value;

[0048] Based on the amount of fuzzy information, the first nonlinear exchange coefficient and the second nonlinear exchange coefficient are obtained as follows:

[0049] ;

[0050] ;

[0051] In the formula, The first nonlinear commutation coefficient, The second nonlinear exchange coefficient, A third random number that is either 1 or -1. It is a fourth random number that is either 1 or -1;

[0052] Based on the first nonlinear exchange coefficient and the second nonlinear exchange coefficient, the particles after the rapid exploration of the solution space are subjected to enhanced exploration of the solution space to determine the particles after the enhanced exploration of the solution space as follows:

[0053] ;

[0054] In the formula, For the first iter During the nth iteration j Particles after rapid exploration of the solution space. For the first j Particles after enhanced exploration of the solution space. Pi To and Different random solution spaces enhance the exploration of particles.

[0055] Furthermore, a time-adaptive adjusted improved sinusoidal Lévy flight strategy is used to perform a global greedy exploration of the particles after the enhanced exploration of the solution space, determining the particles after the global greedy exploration, including:

[0056] To obtain the golden angle:

[0057] ;

[0058] In the formula, For the golden angle, The formula for calculating the golden ratio is as follows: ;

[0059] The Levy flight factor is generated, and a global search is performed on the particles after the enhanced exploration of the solution space based on the golden angle and the Levy flight factor to obtain the global search particles:

[0060] ;

[0061] In the formula, For the first iter During the nth iteration m Particles after enhanced exploration of the solution space. For the first m A global search particle, The fifth random number between (0,1) This is the step size scaling factor. It is a vector with a random step size, and its dimension is the same as that of the vector. same;

[0062] Determine whether the loss function value of the particle after the enhanced exploration of the solution space is less than the loss function value of its corresponding global search particle. If so, the particle after the enhanced exploration of the original solution space is taken as the particle after the global greedy exploration; otherwise, the global search particle is taken as the particle after the global greedy exploration.

[0063] On the other hand, the present invention provides a phase control device for a phased array based on reinforcement learning, comprising:

[0064] The action control module is used to collect the real-time status of the phased array and, based on the real-time status, use a policy network to select the phase control actions of each antenna element of the phased array from the action space.

[0065] The index acquisition module is used to adjust the beam pointing of the phased array based on the phase control actions of each antenna element of the phased array, and to obtain communication performance indicators using the adjusted phased array.

[0066] The sample construction module is used to obtain rewards based on the communication performance indicators, obtain the state at the next moment, and construct experience samples based on the real-time state, phase control actions, rewards, and the state at the next moment.

[0067] The sample update module is used to construct an experience replay pool using experience samples and to update the experience samples in the experience replay pool using a first-in-first-out strategy.

[0068] The network update module is used to randomly sample multiple experience samples from the experience replay pool after detecting that the number of experience samples in the experience replay pool has reached a preset number, and to update the policy network using the sampled experience samples.

[0069] The reinforcement control module is used to control the phase of the phased array using the updated policy network, thereby completing the phase control of the phased array based on reinforcement learning.

[0070] This invention provides a phase control method for phased arrays based on reinforcement learning. By using a deep reinforcement learning algorithm to control the phase of a reinforcement learning-based phased array, the method can effectively improve the control adaptability and accuracy, enabling the phased array control to adapt to environmental changes and device aging. Furthermore, an improved particle swarm optimization algorithm is used to update the policy network, improving the update effect of the policy network and further enhancing the phase control accuracy of the phased array. Attached Figure Description

[0071] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0072] Figure 1 A flowchart of a phase control method for a phased array based on reinforcement learning provided by the present invention;

[0073] Figure 2 A schematic diagram of a phased array phase control device based on reinforcement learning provided by the present invention;

[0074] Among them, 201-Action Control Module, 202-Indicator Acquisition Module, 203-Sample Construction Module, 204-Sample Update Module, 205-Network Update Module, and 206-Enhancement Control Module.

[0075] The accompanying drawings have illustrated specific embodiments of the invention, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the invention in any way, but rather to illustrate the concept of the invention to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0076] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0077] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0078] like Figure 1 As shown in the figure, an embodiment of the present invention provides a phase control method for a phased array based on reinforcement learning, comprising:

[0079] S101. Collect the real-time status of the phased array, and based on the real-time status, use a strategy network to select the phase control actions of each antenna element of the phased array from the action space.

[0080] S102. Based on the phase control actions of each antenna element of the phased array, adjust the beam pointing of the phased array, and use the adjusted phased array to obtain communication performance indicators.

[0081] S103. Obtain the reward based on the communication performance indicators, and at the same time obtain the state at the next moment. Construct an experience sample based on the real-time state, phase control action, reward and state at the next moment.

[0082] S104. An experience replay pool is constructed using experience samples, and the experience samples in the experience replay pool are updated using a first-in-first-out strategy.

[0083] S105. After the number of experience samples in the experience replay pool reaches a preset number, multiple experience samples are randomly sampled from the experience replay pool, and the policy network is updated using the sampled experience samples.

[0084] S106. The phase of the phased array is controlled by the updated policy network to complete the phased array phase control based on reinforcement learning.

[0085] This invention utilizes reinforcement learning, which does not require a precise pre-existing channel mathematical model. It learns the optimal policy through interaction with the environment, making it particularly suitable for scenarios with complex channel environments and significant nonlinear effects. Once the policy network is trained, only one forward propagation is needed to obtain the phase control result in practical applications, with computational latency significantly lower than traditional convex optimization iterative methods. This invention employs an experience replay pool and a first-in-first-out (FIFO) strategy to store and update experience samples. By shuffling data correlations and randomly sampling for training, the correlation between sample sequences is eliminated, reducing variance during training and improving the stability and sample utilization of the policy network updates.

[0086] In this embodiment of the invention, the real-time state includes channel state information (CSI), received signal-to-noise ratio (SNR), and / or the phase control action of the previous moment fed back by the receiver.

[0087] In this embodiment of the invention, the phase control action includes the phase offset of each antenna element in the phased array. In practical applications, a phase offset range can be preset by the staff to determine the action space, and the policy network can select actions within the action space.

[0088] In this embodiment of the invention, obtaining a reward based on the communication performance indicators includes:

[0089] ;

[0090] In the formula, The reward at the current time t, This represents the communication performance index (such as signal-to-interference-plus-noise ratio) at the current time t. Let be the phase vector at the current time t. Let be the phase vector of the previous time t-1. For communication weighting coefficients, These are the smoothing weighting coefficients.

[0091] In designing the reward function, this invention not only considers current communication performance metrics but also introduces a smoothing term for phase vector changes. This design prompts the reinforcement learning agent to not only maximize communication quality when selecting phase control actions but also consider the smoothness of phase changes between adjacent time slots. This effectively avoids drastic jitter or jumps in the phased array beam between adjacent time slots, ensuring the stability of the communication link, reducing the hardware's response stress to large phase changes, and extending equipment lifespan.

[0092] In this embodiment of the invention, updating the policy network using sampled empirical samples includes: updating the network parameters of the policy network using gradient descent based on the empirical samples.

[0093] Existing reinforcement learning methods often suffer from slow convergence and a tendency to get trapped in local optima during training, especially in the parameter update phase of the policy network. Traditional gradient descent methods have limited search capabilities when facing complex non-convex optimization problems, resulting in low beam pointing control accuracy of phased arrays and limited improvement in communication performance. Therefore, this invention also proposes a Synergistic Particle Swarm Optimization (SPSO) algorithm, which improves the performance of the original Particle Swarm Optimization algorithm and enhances the training effect of the policy network.

[0094] In this embodiment of the invention, updating the policy network using sampled empirical samples includes:

[0095] Based on the network parameters of the policy network, a particle swarm is generated;

[0096] For any particle in the particle swarm, the loss function value corresponding to the particle is obtained based on the sampled empirical samples, and the particle with the smallest loss function value is determined as the instantaneous optimal particle.

[0097] To facilitate understanding in this field, the embodiments of the present invention also introduce the application scenarios of policy networks. Policy networks are mainly applied in the Actor-Critic architecture, which combines policy gradient and value function methods in deep reinforcement learning, with the core being division of labor and collaboration. The Actor (policy network) is responsible for decision-making, outputting actions or action probability distributions based on the current state, and maximizing long-term cumulative rewards by adjusting parameters. The Critic (value network) is responsible for evaluation, calculating the value (Q-value or V-value) of the current state or action, evaluating the Actor's performance, and calculating the temporal difference (TD) error as a feedback signal. During training, the Actor optimizes its policy based on the score provided by the Critic (such as the advantage function), increasing the probability of high-quality actions; the Critic improves the accuracy of the evaluation by minimizing the prediction error. This structure effectively reduces the variance of the policy gradient, balancing variance and bias, and performs well in high-dimensional complex tasks such as continuous control. Therefore, the temporal difference error can be used as the loss function value. However, it is worth noting that other existing deep reinforcement learning architectures and corresponding loss functions can also be used to obtain the loss function value, thereby ensuring that the policy network can achieve accurate updates.

[0098] Based on the instantaneous optimal particle, a dual-optimal learning strategy improved by nonlinearly adjusting the weight learning coefficient is used to quickly explore the solution space of the particles in the particle swarm, and determine the particles after the quick exploration of the solution space.

[0099] A nonlinear information exchange strategy with improved fuzzy information content is used to perform enhanced exploration of the solution space for the particles after rapid exploration of the solution space, and the particles after enhanced exploration of the solution space are determined.

[0100] A time-adaptive, improved sinusoidal Levy flight strategy is used to perform a global greedy exploration of the particles after the enhanced exploration of the solution space, thereby determining the particles after the global greedy exploration.

[0101] Obtain the number of iterations and determine whether the number of iterations is greater than or equal to the preset maximum number of iterations. If so, determine the globally optimal particle based on the particles after the global greedy exploration, and use the network parameters in the globally optimal particle as the network parameters of the policy network to complete the update. Otherwise, return to the step of obtaining the instantaneous optimal particle.

[0102] In this embodiment of the invention, based on the instantaneously optimal particle, a dual-optimal learning strategy improved by nonlinearly adjusting the weight learning coefficients is used to rapidly explore the solution space of the particles in the particle swarm, determining the particles after the rapid exploration of the solution space, including:

[0103] The nonlinear adjustment function is constructed as follows:

[0104] ;

[0105] In the formula, It is a nonlinear adjustment function. x Let e ​​be a variable, and let e be a natural constant.

[0106] Based on the aforementioned nonlinear adjustment function, the velocity adjustment weight, the first learning factor, and the second learning factor are obtained as follows:

[0107] ;

[0108] ;

[0109] ;

[0110] In the formula, Adjust the weights for speed. As the first learning factor, As the second learning factor, To adjust the parameter (which can be set to 10), This represents the number of iterations already performed. This represents the maximum number of iterations.

[0111] Based on the velocity adjustment weight, the first learning factor, and the second learning factor, the particles in the particle swarm are subjected to a rapid exploration of the solution space to determine the particles after the rapid exploration of the solution space:

[0112] ;

[0113] ;

[0114] In the formula, For the first iter During the nth iteration i One particle, For the first i Particles after rapid exploration of the solution space. For the first iter During the +1st iteration, the... i The update rate of each particle For the first iter During the nth iteration i The update rate of each particle i=1,2,...,M, where M is the total number of particles. The first random number between (0,1) The second random number between (0,1) for The historical best value, It is the instantaneously optimal particle.

[0115] In this embodiment of the invention, Decrease from 0.35 to 0.05, Decrease from 0.45 to 0.05, Increasing the weight from 0.1 to 0.5, this improvement in non-linear weight setting better balances the algorithm's exploration and development capabilities, thereby enhancing its adaptability at different stages. In the early stages of the algorithm, a larger weight... and smaller Encourage particles to explore themselves and maintain group diversity; while in later stages, smaller particles... and larger The value then prompts the particles to converge rapidly toward the global optimal solution.

[0116] An improved dual-optimization learning strategy is adopted by introducing nonlinearly adjusted weights and learning factors. The velocity weights and learning factors are dynamically adjusted using a nonlinear adjustment function. In the early stages of iteration, larger weights and learning factors help particles to explore the entire solution space quickly, avoiding blind searches. In the later stages of iteration, smaller weights and learning factors help particles to refine their exploration in local regions. This dynamic balancing mechanism significantly accelerates the convergence speed of the policy network parameters while improving the accuracy of parameter optimization.

[0117] In this embodiment of the invention, a nonlinear information exchange strategy with improved fuzzy information content is used to perform enhanced solution space exploration on the particles after the rapid exploration of the solution space, and the particles after the enhanced solution space exploration are determined, including:

[0118] Obtain the loss function value of the particles after the fast exploration of the solution space, and obtain the degree of belonging of the particles after the fast exploration of the solution space relative to the instantaneous optimal particle based on the loss function value;

[0119] ;

[0120] In the formula, For the first j The degree of affiliation of a particle after a rapid exploration of the solution space. The loss function value of the instantaneously optimal particle. Let be the maximum loss function value corresponding to the particle after fast exploration of all solution spaces. Let be the loss function value of the particle after fast exploration of the j-th solution space;

[0121] Based on the aforementioned affiliation degree, the degree to which a particle, after a rapid exploration of the solution space, is not located at the current optimal position of the instantaneously optimal particle is determined as follows:

[0122] ;

[0123] In the formula, For the first j The degree to which a particle, after a rapid exploration of the solution space, is no longer in the current optimal position of the instantaneously optimal particle. It is an adjustable parameter, and its value is between [0.8, 1].

[0124] Based on the degree of affiliation and the degree of affiliation, the amount of fuzzy information obtained is as follows:

[0125] ;

[0126] In the formula, K For fuzzy information content, max is the function to find the maximum value;

[0127] In the early stages of the algorithm's operation, due to the relatively dispersed nature of the particles, the convergence of the particle swarm is low, and the value of fuzzy information is also small. This means that the algorithm assigns a larger step size to the particles at this stage, allowing them to explore the solution space more extensively and find regions that may contain the global optimum. In the later stages of the algorithm's operation, when the behavior of all particles tends to stagnate and the convergence of the particle swarm reaches its maximum, the value of fuzzy information reaches its maximum. At this point, the algorithm performs a fine-grained local search using the smallest particle step size to ensure that the most accurate optimal solution is found.

[0128] Based on the amount of fuzzy information, the first nonlinear exchange coefficient and the second nonlinear exchange coefficient are obtained as follows:

[0129] ;

[0130] ;

[0131] In the formula, The first nonlinear commutation coefficient, The second nonlinear exchange coefficient, A third random number that is either 1 or -1. It is a fourth random number that is either 1 or -1;

[0132] Based on the first nonlinear exchange coefficient and the second nonlinear exchange coefficient, the particles after the rapid exploration of the solution space are subjected to enhanced exploration of the solution space to determine the particles after the enhanced exploration of the solution space as follows:

[0133] ;

[0134] In the formula, For the first iter During the nth iteration j Particles after rapid exploration of the solution space. For the first j Particles after enhanced exploration of the solution space. Pi To and Different random solution spaces enhance the exploration of particles.

[0135] By introducing a nonlinear information exchange strategy improved by fuzzy information, the affiliation degree and deviation degree of particles from the instantaneous optimal particles are calculated, and the nonlinear exchange coefficients are adaptively adjusted, so that particles can flexibly exchange information with other particles according to the current distribution state of the population, thereby enhancing the diversity of the population and preventing premature convergence.

[0136] In this embodiment of the invention, a time-adaptive adjusted improved sinusoidal Lévy flight strategy is used to perform a global greedy exploration of the particles after the solution space enhancement exploration, and the particles after the global greedy exploration are determined, including:

[0137] To obtain the golden angle:

[0138] ;

[0139] In the formula, For the golden angle, The formula for calculating the golden ratio is as follows: ;

[0140] The Levy flight factor is generated, and a global search is performed on the particles after the enhanced exploration of the solution space based on the golden angle and the Levy flight factor to obtain the global search particles:

[0141] ;

[0142] In the formula, For the first iter During the nth iteration m Particles after enhanced exploration of the solution space. For the first m A global search particle, The fifth random number between (0,1) This is the step size scaling factor. It is a vector with a random step size, and its dimension is the same as that of the vector. same;

[0143] Determine whether the loss function value of the particle after the enhanced exploration of the solution space is less than the loss function value of its corresponding global search particle. If so, the particle after the enhanced exploration of the original solution space is taken as the particle after the global greedy exploration; otherwise, the global search particle is taken as the particle after the global greedy exploration.

[0144] The sinusoidal Lévy flight strategy, combined with the golden angle and long-tailed Lévy flight distribution, enables particles to perform large-step jump searches globally, greatly enhancing the algorithm's ability to escape local optima. Simultaneously, a greedy strategy is employed for adaptive selection, ensuring the algorithm's iteration speed.

[0145] By training the policy network using SPSO as described above, it is ensured that the policy network can converge to the global optimum, thereby making the beam pointing control of the phased array more precise and significantly improving the performance indicators of the communication system, such as throughput and signal-to-noise ratio.

[0146] like Figure 2 As shown, this embodiment of the invention also provides a phase control device for a phased array based on reinforcement learning, comprising:

[0147] The action control module 201 is used to collect the real-time status of the phased array and, based on the real-time status, use a strategy network to select the phase control action of each antenna element of the phased array from the action space.

[0148] The indicator acquisition module 202 is used to adjust the beam pointing of the phased array according to the phase control action of each antenna element of the phased array, and to obtain the communication performance indicators using the adjusted phased array.

[0149] The sample construction module 203 is used to obtain a reward based on the communication performance index, obtain the state at the next moment, and construct an experience sample based on the real-time state, phase control action, reward and the state at the next moment.

[0150] The sample update module 204 is used to construct an experience replay pool using experience samples and to update the experience samples in the experience replay pool using a first-in-first-out strategy.

[0151] The network update module 205 is used to randomly sample multiple experience samples from the experience replay pool after detecting that the number of experience samples in the experience replay pool has reached a preset number, and to update the policy network using the sampled experience samples.

[0152] The reinforcement control module 206 is used to control the phase of the phased array using the updated policy network, thereby completing the phase control of the phased array based on reinforcement learning.

[0153] The principle and beneficial effects of this reinforcement learning-based phased array phase control device are similar to those of the reinforcement learning-based phased array phase control method, and will not be elaborated here.

[0154] This invention also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps in the method disclosed in this invention.

[0155] This invention also provides a computer program product that, when run on an electronic device, causes a processor to execute the steps in the method disclosed in this invention.

[0156] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0157] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0158] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0160] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0161] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0162] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A phase control method for a phased array based on reinforcement learning, characterized in that, include: The real-time status of the phased array is collected, and based on the real-time status, a strategy network is used to select the phase control actions of each antenna element of the phased array from the action space. Based on the phase control actions of each antenna element of the phased array, the beam pointing of the phased array is adjusted, and the communication performance index is obtained using the adjusted phased array; the communication performance index is the signal-to-interference-plus-noise ratio (SINR). Rewards are obtained based on the communication performance indicators, and the state at the next moment is obtained. Based on the real-time state, phase control actions, rewards, and the state at the next moment, an experience sample is constructed. Rewards are obtained based on the aforementioned communication performance metrics, including: ; In the formula, The reward at the current time t, Let be the communication performance index at the current time t. Let be the phase vector at the current time t. Let be the phase vector of the previous time t-1. For communication weighting coefficients, For smoothing weighting coefficients; An experience replay pool is constructed using experience samples, and the experience samples in the experience replay pool are updated using a first-in-first-out strategy. Once the number of experience samples in the experience replay pool reaches a preset number, multiple experience samples are randomly sampled from the experience replay pool, and the policy network is updated using the sampled experience samples. The phase of the phased array is controlled by the updated policy network, thus completing the phase control of the phased array based on reinforcement learning.

2. The phase control method for phased arrays based on reinforcement learning according to claim 1, characterized in that, The real-time status includes channel state information fed back by the receiver, received signal-to-noise ratio, and / or phase control action from the previous moment.

3. The phase control method for phased arrays based on reinforcement learning according to claim 1, characterized in that, The phase control action includes the phase offset of each antenna element in the phased array.

4. The phase control method for phased arrays based on reinforcement learning according to claim 1, characterized in that, Updating the policy network using sampled empirical samples includes: Based on the empirical samples, the network parameters of the policy network are updated using the gradient descent method.

5. The phase control method for phased arrays based on reinforcement learning according to claim 1, characterized in that, Updating the policy network using sampled empirical samples includes: Based on the network parameters of the policy network, a particle swarm is generated; For any particle in the particle swarm, the loss function value corresponding to the particle is obtained based on the sampled empirical samples, and the particle with the smallest loss function value is determined as the instantaneous optimal particle. Based on the instantaneous optimal particle, a dual-optimal learning strategy improved by nonlinearly adjusting the weight learning coefficient is used to quickly explore the solution space of the particles in the particle swarm, and determine the particles after the quick exploration of the solution space. A nonlinear information exchange strategy with improved fuzzy information content is used to perform enhanced exploration of the solution space for the particles after rapid exploration of the solution space, and the particles after enhanced exploration of the solution space are determined. A time-adaptive, improved sinusoidal Levy flight strategy is used to perform a global greedy exploration of the particles after the enhanced exploration of the solution space, thereby determining the particles after the global greedy exploration. Obtain the number of iterations and determine whether the number of iterations is greater than or equal to the preset maximum number of iterations. If so, determine the globally optimal particle based on the particles after the global greedy exploration, and use the network parameters in the globally optimal particle as the network parameters of the policy network to complete the update. Otherwise, return to the step of obtaining the instantaneous optimal particle.

6. The phase control method for phased arrays based on reinforcement learning according to claim 5, characterized in that, Based on the instantaneously optimal particle, a dual-optimal learning strategy improved by nonlinearly adjusting the weight learning coefficients is used to rapidly explore the solution space of the particles in the particle swarm, determining the particles after the rapid exploration of the solution space, including: The nonlinear adjustment function is constructed as follows: ; In the formula, It is a nonlinear adjustment function. x Let e ​​be a variable, and let e be a natural constant. Based on the aforementioned nonlinear adjustment function, the velocity adjustment weight, the first learning factor, and the second learning factor are obtained as follows: ; ; ; In the formula, Adjust the weights for speed. As the first learning factor, As the second learning factor, To adjust the parameters, This represents the number of iterations already performed. This represents the maximum number of iterations. Based on the velocity adjustment weight, the first learning factor, and the second learning factor, the particles in the particle swarm are subjected to a rapid exploration of the solution space to determine the particles after the rapid exploration of the solution space: ; ; In the formula, For the first iter During the nth iteration i One particle, For the first i Particles after rapid exploration of the solution space. For the first iter During the +1st iteration, the... i The update rate of each particle For the first iter During the nth iteration i The update rate of each particle i =1,2,...,M, where M is the total number of particles. The first random number between (0,1) The second random number between (0,1) for The historical best value, It is the instantaneously optimal particle.

7. The phase control method for phased arrays based on reinforcement learning according to claim 6, characterized in that, A nonlinear information exchange strategy with improved fuzzy information content is used to perform enhanced solution space exploration on the particles after the rapid exploration of the solution space, and the particles after the enhanced solution space exploration are determined, including: Obtain the loss function value of the particles after the fast exploration of the solution space, and obtain the degree of belonging of the particles after the fast exploration of the solution space relative to the instantaneous optimal particle based on the loss function value; ; In the formula, For the first j The degree of affiliation of a particle after a rapid exploration of the solution space. The loss function value of the instantaneously optimal particle. Let be the maximum loss function value corresponding to the particle after fast exploration of all solution spaces. Let be the loss function value of the particle after fast exploration of the j-th solution space; Based on the aforementioned affiliation degree, the degree to which a particle, after a rapid exploration of the solution space, is not located at the current optimal position of the instantaneously optimal particle is determined as follows: ; In the formula, For the first j The degree to which a particle, after a rapid exploration of the solution space, is no longer in the current optimal position of the instantaneously optimal particle. It is an adjustable parameter, and its value is between [0.8, 1]. Based on the degree of affiliation and the degree of affiliation, the amount of fuzzy information obtained is as follows: ; In the formula, K For fuzzy information content, max is the function to find the maximum value; Based on the amount of fuzzy information, the first nonlinear exchange coefficient and the second nonlinear exchange coefficient are obtained as follows: ; ; In the formula, The first nonlinear commutation coefficient, The second nonlinear exchange coefficient, A third random number that is either 1 or -1. It is a fourth random number that is either 1 or -1; Based on the first nonlinear exchange coefficient and the second nonlinear exchange coefficient, the particles after the rapid exploration of the solution space are subjected to enhanced exploration of the solution space to determine the particles after the enhanced exploration of the solution space as follows: ; In the formula, For the first iter During the nth iteration j Particles after rapid exploration of the solution space. For the first j Particles after enhanced exploration of the solution space. Pi To and Different random solution spaces enhance the exploration of particles.

8. The phase control method for phased arrays based on reinforcement learning according to claim 7, characterized in that, A time-adaptive, improved sinusoidal Lévy flight strategy is used to perform a global greedy exploration of the particles after the enhanced exploration of the solution space, determining the particles after the global greedy exploration, including: The golden angle is obtained as follows: ; In the formula, For the golden angle, The formula for calculating the golden ratio is as follows: ; The Levy flight factor is generated, and a global search is performed on the particles after the enhanced exploration of the solution space based on the golden angle and the Levy flight factor to obtain the global search particles: ; In the formula, For the first iter During the nth iteration m Particles after enhanced exploration of the solution space. For the first m A global search particle, The fifth random number between (0,1) This is the step size scaling factor. It is a vector with a random step size, and its dimension is the same as that of the vector. same; Determine whether the loss function value of the particle after the enhanced exploration of the solution space is less than the loss function value of its corresponding global search particle. If so, the particle after the enhanced exploration of the original solution space is taken as the particle after the global greedy exploration; otherwise, the global search particle is taken as the particle after the global greedy exploration.

9. A reinforcement learning-based phased array phase control device, wherein the reinforcement learning-based phased array phase control device is used to execute the reinforcement learning-based phased array phase control method according to any one of claims 1 to 8, characterized in that, include: The action control module is used to collect the real-time status of the phased array and, based on the real-time status, use a policy network to select the phase control actions of each antenna element of the phased array from the action space. The index acquisition module is used to adjust the beam pointing of the phased array based on the phase control actions of each antenna element of the phased array, and to obtain communication performance indicators using the adjusted phased array. The sample construction module is used to obtain rewards based on the communication performance indicators, obtain the state at the next moment, and construct experience samples based on the real-time state, phase control actions, rewards, and the state at the next moment. The sample update module is used to construct an experience replay pool using experience samples and to update the experience samples in the experience replay pool using a first-in-first-out strategy. The network update module is used to randomly sample multiple experience samples from the experience replay pool after detecting that the number of experience samples in the experience replay pool has reached a preset number, and to update the policy network using the sampled experience samples. The reinforcement control module is used to control the phase of the phased array using the updated policy network, thereby completing the phase control of the phased array based on reinforcement learning.

Citation Information

Patent Citations

  • Array beam spatial domain synthesis method based on deep reinforcement learning

    CN120409216A

  • Intelligent beam forming system and method for millimeter wave communication

    CN120582653A