Intelligent adaptive control method and system for 5G wireless communication antenna
Through multi-agent depth deterministic strategy gradient algorithm and digital twin technology, dynamic collaborative optimization of 5G wireless communication antenna beamforming and RIS phase configuration is achieved, solving the problem of improving system performance in complex scenarios, and significantly improving the real-time and robustness of the system.
Patent Information
- Application Number
- CN202510538000.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In complex 5G wireless communication scenarios, how to achieve dynamic joint adjustment of antenna beamforming and reconstructible intelligent surface (RIS) phase configuration to improve system performance.
Multi-agent depth deterministic strategy gradient (MADDPG) algorithm is used to achieve collaborative optimization of beamforming and RIS phase through technical means such as agent division and task allocation, state space and action space design, reward function design, distributed training and optimization, dynamically encoded electromagnetic fingerprint map, digital twin construction and strategy verification.
Through the perception-decision-verification closed loop, the system's real-time, robustness and energy efficiency are improved, and are suitable for complex 5G communication scenarios, significantly improving system performance.
Smart Images

Figure CN120074591A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mobile communications, and specifically to an intelligent adaptive control method and system for 5G wireless communication antennas. Background Art
[0002] In a 5G wireless communication system, the combination of large-scale MIMO (multiple-input multiple-output) antennas and RIS is considered a key technology to improve spectral efficiency and coverage. Beamforming optimizes the signal directivity by adjusting the transmission weights of the antenna array, while RIS changes the signal propagation path by dynamically adjusting the phase of its reflection units.
[0003] However, in complex scenarios of 5G communication (such as environments with high mobility, dense multi-users, and significant multipath effects), how to achieve dynamic joint adjustment of antenna beamforming and phase configuration of reconfigurable intelligent surfaces (RIS) to improve system performance (such as throughput, coverage, and anti-interference ability) remains an unsolved problem. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides an intelligent adaptive control method and system for 5G wireless communication antennas, which solves the problems mentioned in the above background.
[0005] The present invention provides the following technical solutions: The present invention discloses an intelligent adaptive control method for 5G wireless communication antennas, including the following steps: S1: Agent division and task assignment: Divide the beamforming of the base station and the phase configuration of the reconfigurable intelligent surface in wireless communication into a base station agent and a reconfigurable intelligent surface agent respectively. The base station agent is responsible for optimizing the beamforming weight vector to maximize the signal-to-interference-plus-noise ratio (SINR) of the target user, and the reconfigurable intelligent surface agent is responsible for adjusting the reconfigurable intelligent surface phase matrix to enhance the main path signal and suppress multipath interference; S2: State space and action space design: Construct a state space, including real-time channel state information, electromagnetic characteristics, and user location information; construct an action space, including the base station agent adjusting the amplitude and phase of the beam, and the reconfigurable intelligent surface agent selecting discrete or continuous phase values; S3: Reward function design: Considering SINR, energy consumption, and interference comprehensively, define a reward function to guide the optimization process of the agent. The reward function is as follows: ; where α, β, and γ are weight coefficients, which are used to control the influence of SINR, power, and interference respectively. SINR k is the signal-to-interference-plus-noise ratio of the k-th user, W represents the beamforming weight vector of the base station, and ∣W∣ 2 represents the transmission power, and Interference represents the interference power within the system; S4: Distributed training and optimization: The multi-agent deep deterministic policy gradient algorithm is adopted. During the training process, the policy is updated through local observations, and the global reward signal is shared regularly to ensure coordination. S5: Dynamically encoding the electromagnetic fingerprint spectrum: The electromagnetic environment is dynamically encoded through the multi-sensor data acquisition function and the feature extraction method, and the obtained electromagnetic features are converted into a phase codebook to guide the actions of multi-agent deep reinforcement learning. S6: Digital twin construction and policy verification: Through the high-precision virtual channel model and random perturbation, a digital twin is established, and multiple groups of multi-agent deep reinforcement learning policies are rehearsed and the performance data of the physical system are instantaneously fed back. S7: Real-time feedback and iterative optimization: The real-time performance feedback of the physical system is used to update the simulation model parameters to ensure the effectiveness and stability of the optimization process.
[0006] As a preferred solution, the real-time channel state information in the state space is the base station-reconfigurable intelligent surface channel matrix G and the reconfigurable intelligent surface-user channel matrix H. The state space is represented by combining the real-time channel state information and the electromagnetic fingerprint feature F, and the formula is: ; where Encoder(F) represents the encoding result of the electromagnetic fingerprint.
[0007] As a preferred solution, the generation process of the phase codebook in step S5 is as follows: S51: Generating the electromagnetic feature tensor: ; where K represents the feature dimension and T represents the time window length; S52: Using the variational autoencoder to compress the electromagnetic feature tensor F into a low-dimensional feature vector Z: ; where Z is the encoded low-dimensional feature vector with dimension d; And the loss function of the variational autoencoder is defined as: ; where F′ represents the reconstructed electromagnetic feature generated by the encoder, μ z , σ z respectively represent the mean and standard deviation of the feature vector Z, λ is the regularization weight used to control the influence of the KL divergence, and KL is the Kullback-Leibler divergence; S53: Generating the RIS phase codebook according to the encoded low-dimensional feature vector Z: ; Among them, C represents the phase codebook of the reconfigurable intelligent surface, which contains multiple possible phase configurations, and represents the Mth phase configuration of the reconfigurable intelligent surface.
[0008] As a preferred solution, the phase selection strategy of the reconfigurable intelligent surface agent selects the phase configuration that best matches the current electromagnetic feature codebook based on the cosine similarity metric function, and the cosine similarity metric function is expressed as: .
[0009] As a preferred solution, in the digital twin, multiple multi-agent deep reinforcement learning strategies are simulated in parallel. The evaluation metrics include cumulative reward, throughput, and bit error rate, and the optimal strategy is selected for physical system execution. The optimal strategy is represented by the following formula: ; where, represents the reward generated by the base station and the reconfigurable intelligent surface actions at time t, γ represents the discount factor, which is used to control the influence of future rewards, and E sim represents the expectation calculated in the digital twin simulation environment.
[0010] As a preferred solution, when the action space of the reconfigurable intelligent surface agent is to select discrete phase values, the discrete phase values are: .
[0011] As a preferred solution, the base station agent increases the signal coverage range by adjusting the beam direction.
[0012] As a preferred solution, the digital twin is constructed based on ray tracing technology and neural networks to build a high-precision virtual channel model, including a base station, a reconfigurable intelligent surface, users, and environmental obstacles.
[0013] As a preferred solution, the digital twin is iterated through real-time feedback with the physical system, and the cycle control is within 10 ms.
[0014] The present invention also provides an intelligent adaptive control system for a 5G wireless communication antenna, which is used to implement the intelligent adaptive control method.
[0015] Compared with the prior art, the present invention has the following beneficial effects: Through the innovative combination of a multi-agent DRL framework, dynamic encoding of electromagnetic fingerprint maps, and digital-twin virtual-real linkage, the present invention realizes the explicit collaborative optimization of beamforming and RIS phase in 5G wireless communication antennas. It not only solves the limitations of traditional methods in dynamic environments but also improves the real-time performance, robustness, and energy efficiency of the system through a sensing-decision-validation closed-loop, with significant innovation and practical value, and is applicable to complex 5G communication scenarios (such as urban dense areas, high-speed mobile scenarios, etc.). BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flowchart of agent division in the intelligent adaptive control method of the present invention; Figure 2 It is a flowchart of dynamic encoding and digital twin in the intelligent adaptive control method of the present invention; Figure 3 It is a schematic flowchart of the intelligent adaptive control system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] Please refer to Figure 1-2 The intelligent adaptive control method for the 5G wireless communication antenna in this embodiment includes the following steps: S1: Agent division and task assignment: Divide the beamforming of the base station and the phase configuration of the reconfigurable intelligent surface (RIS) in wireless communication into a base station agent and a reconfigurable intelligent surface agent respectively. The base station agent is responsible for optimizing the beamforming weight vector to maximize the signal-to-interference-plus-noise ratio (SINR) of the target user, and the reconfigurable intelligent surface agent is responsible for adjusting the reconfigurable intelligent surface phase matrix to enhance the main path signal and suppress multipath interference. The goal of optimizing the phase configuration of the RIS is to enhance the signal quality and reduce interference; In a 5G wireless communication system, the base station and the reconfigurable intelligent surface (RIS) are two important components. To improve the performance of the system, tasks and optimization goals need to be reasonably allocated to different agents. In this step, the division and task assignment of agents lay the foundation for subsequent optimization work.
[0019] S2: State Space and Action Space Design: Construct a state space that includes real-time channel state information, electromagnetic characteristics, and user location information; construct an action space that includes the base station agent adjusting the amplitude and phase of the beam, and the reconfigurable intelligent surface agent selecting discrete or continuous phase values; State Space: In wireless communication, the state space contains important information related to the operating conditions of the communication system, such as: Real-time Channel State Information (CSI): Provides signal quality information of the current communication link, including signal strength, noise interference, user location, etc.; Electromagnetic Characteristics: Such as the electromagnetic wave propagation characteristics of the environment, affected by factors such as weather and obstacles; User Location Information: The real-time location of the user equipment, which helps the base station and RIS adjust their configurations.
[0020] Action Space: Provides a decision-making space for the behavior of the agent. The action spaces of the base station agent and the RIS agent are respectively: Base Station Agent: Can adjust the amplitude and phase of the beam to optimize the signal transmission quality; RIS Agent: Can select different phase values, which may be discrete or continuous, to adjust the reflection characteristics of the RIS and optimize the signal path.
[0021] S3: Reward Function Design: Considering signal-to-interference-plus-noise ratio (SINR), energy consumption, and interference comprehensively, define a reward function to guide the agent optimization process. The reward function is as follows: ; where α, β, and γ are weight coefficients, used to control the impacts of SINR, power, and interference respectively. SINR k is the signal-to-interference-plus-noise ratio of the k-th user, W represents the beamforming weight vector of the base station, and ∣W∣ 2 represents the transmit power, and Interference represents the interference power within the system; Signal-to-interference-plus-noise ratio is the core optimization goal in wireless communication. The goal is to maximize SINR, ensure the maximization of signal strength, and minimize interference; consider energy efficiency during the optimization process, reduce unnecessary energy consumption, and meet the requirements of green communication; improve communication quality by reducing multipath interference and other signal interferences.
[0022] S4: Distributed Training and Optimization: Adopt the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm. During the training process, update the policy through local observations and regularly share the global reward signal to ensure coordination. Ensure that each agent independently updates the policy through local observations and regularly shares the global reward signal. The multi-agent deep deterministic policy gradient algorithm is a common multi-agent reinforcement learning method and will not be elaborated here; Each agent can update its policy through local observations. However, to ensure the synergy of the system, the agents will regularly share the global reward signal. This distributed learning method allows the agents to independently optimize without complete information, while promoting cooperation through shared information, thereby improving the performance of the overall system.
[0023] S5: Dynamic Encoding of Electromagnetic Fingerprint Map: Dynamically encode the electromagnetic environment through multi-sensor data acquisition functions and feature extraction methods, and convert the obtained electromagnetic features into a phase codebook to guide the actions of multi-agent deep reinforcement learning (DRL); To further enhance the adaptability and accuracy of the communication system, the electromagnetic environment is dynamically encoded through various sensor data (such as RF sensors, millimeter-wave radars, and cameras). Specifically: By collecting electromagnetic data in the environment and combining feature extraction methods, real-time electromagnetic environment information can be obtained; convert these feature information into a phase codebook to guide the action selection in the process of multi-agent deep reinforcement learning. This dynamic encoding of the electromagnetic fingerprint map helps the agents to respond promptly to the changing electromagnetic environment and optimize the strategy.
[0024] S6: Digital Twin Construction and Policy Verification: Establish a digital twin through a high-precision virtual channel model and random perturbations, and pre-enact multiple groups of multi-agent deep reinforcement learning policies and instantaneously feedback the performance data of the physical system; Digital twin technology can simulate the actual communication environment by constructing a high-precision virtual channel model and introduce random perturbations to verify the system performance, which provides a virtual test platform for the optimization process. Through the digital twin, different multi-agent deep reinforcement learning policies can be pre-enacted in the simulation environment, and the simulation results can be real-time feedback to help adjust the model parameters and policies to ensure the actual effect in the physical system.
[0025] S7: Real-time Feedback and Iterative Optimization: Use the real-time performance feedback of the physical system to update the parameters of the simulation model to ensure the effectiveness and stability of the optimization process; Finally, the real-time performance feedback of the physical system is used for continuous optimization. By real-time monitoring the performance data of the actual communication system, the feedback information is transmitted back to the simulation model: Update the parameters of the simulation model to make it more in line with the actual situation, while ensuring the stability and effectiveness of the optimization process, and achieve adaptive dynamic optimization. In this way, through the iterative process, the performance of the system is continuously improved.
[0026] The base station agent independently optimizes the beamforming weight vector: ; where N BS$N$ is the number of base station antennas, and the goal is to maximize the signal-to-interference-plus-noise ratio (SINR) of the target user. The reconfigurable intelligent surface (RIS) agent is responsible for adjusting the RIS phase matrix: ; This matrix contains the phase settings of each element of the RIS, where $N$ RIS is the number of RIS elements, and $\theta_{n}$ represents the phase value in the form of a complex exponential, which is used to adjust the phase response of the RIS element. The goal is to enhance the main path signal and suppress multipath interference; In summary, through the innovative combination of the multi-agent deep reinforcement learning (DRL) framework, electromagnetic fingerprint dynamic coding, and digital twin virtual-real linkage, this solution realizes the explicit cooperative optimization of beamforming and RIS phase in 5G wireless communication antennas. It not only solves the limitations of traditional methods in dynamic environments but also improves the real-time performance, robustness, and energy efficiency of the system through the perception-decision-validation closed-loop, showing significant innovation and practical value, and is applicable to complex 5G communication scenarios (such as urban dense areas, high-speed mobile scenarios, etc.).
[0027] Furthermore, the real-time channel state information in the state space is the base station-RIS channel matrix $G$ and the RIS-user channel matrix $H$. The state space is represented by combining the real-time channel state information and the electromagnetic fingerprint feature $F$, and the formula is: ; where $Encoder(F)$ represents the encoding result of the electromagnetic fingerprint; In the intelligent adaptive control method of the 5G wireless communication antenna described above, the core idea is to construct a dynamically changing state space model based on the real-time channel state information and the electromagnetic fingerprint feature to optimize the performance of the communication system; Among them, the real-time channel state information (CSI) refers to the channel matrix between the base station and the reconfigurable intelligent surface (RIS), and the channel matrix between the reconfigurable intelligent surface and the user. It is a very important part of the communication system because it determines the quality of signal propagation; Base station-RIS channel matrix $G$: This represents the wireless communication channel state between the base station and the RIS. The RIS is a technology that can optimize signal propagation by adjusting its surface state and can play roles such as enhancing signals and improving coverage in wireless communication; Reconfigurable intelligent surface-user channel matrix, $H$: This represents the wireless communication channel state between the reconfigurable intelligent surface and the user. By adjusting the reflection units on the RIS, the propagation path of the signal can be affected, thereby optimizing the communication quality; These matrices G and H together constitute the real-time channel state information of the communication system, which can accurately describe the wireless channel quality between the base station, RIS and users, facilitating the adaptive optimization of communication strategies.
[0028] Electromagnetic fingerprints are obtained by analyzing the wireless signals emitted by devices to extract their unique electromagnetic characteristics. The electromagnetic fingerprint of each device is unique and can be used to identify devices or optimize communication channels. where the state space s t is a comprehensive data structure that contains all real-time channel state information and features related to electromagnetic fingerprints. It provides comprehensive input information for intelligent adaptive control algorithms, which can help the system dynamically adjust communication strategies according to the current channel state and electromagnetic characteristics to achieve optimized communication effects. Based on the above state space s t , the behavior of the system can be adaptively adjusted through intelligent algorithms. These algorithms can make predictions and decisions based on real-time updated channel information and electromagnetic fingerprint features, optimizing operations such as signal transmission, power allocation, and resource scheduling. By combining traditional channel state information and electromagnetic fingerprint features and comprehensively optimizing through intelligent algorithms, more precise and flexible communication control can be achieved, improving the performance and reliability of 5G wireless communication systems.
[0029] Furthermore, the generation process of the phase codebook in step S5 is as follows: S51: Generate the electromagnetic feature tensor: ; where K represents the feature dimension, that is, the number or complexity of electromagnetic fingerprint features extracted at each moment, and T represents the time window length, that is, the number of frames within the observed time range. Assuming there is a set of electromagnetic features at each moment, then T can be regarded as the set of electromagnetic features at different time points within this time window. The purpose of generating this electromagnetic feature tensor is to capture the electromagnetic features in the environment, which may include reflections of buildings, interference in signal propagation, etc. This tensor is raw high-dimensional data that needs to be further processed to extract useful information.
[0030] S52: Use a variational autoencoder to compress the electromagnetic feature tensor F into a low-dimensional feature vector Z: ; Z is the encoded low-dimensional feature vector with dimension d, and d is much smaller than K. This is because through the encoder part of the VAE, we compress the high-dimensional data into a low-dimensional space. The low-dimensional representation can retain the important information in the input data while reducing redundancy and noise.
[0031] And define the loss function of the variational autoencoder as: ; where F′ represents the reconstructed electromagnetic features generated by the encoder, μ z , σ z represent the mean and standard deviation of the feature vector Z respectively, λ is the regularization weight used to control the influence of the KL divergence, and KL is the Kullback-Leibler divergence used to measure the difference between the distribution of the latent variable and the standard normal distribution; is the mean square error between the input electromagnetic feature tensor F and the output F′ after reconstruction by the autoencoder, aiming to enable the autoencoder to reconstruct the input data as accurately as possible; is the Kullback-Leibler divergence between the distribution of the latent variable Z and the standard normal distribution, used to ensure that the distribution of the latent variable Z learned by the encoder is close to the standard normal distribution, so that the generative model can generate real and diverse samples; Through this process, the variational autoencoder can not only compress high-dimensional electromagnetic feature data, but also map it to a latent space, retaining important electromagnetic information, which is suitable for subsequent RIS phase codebook generation.
[0032] S53: Generate the RIS phase codebook according to the encoded low-dimensional feature vector Z: ; where C represents the reconfigurable intelligent surface phase codebook, which contains multiple possible phase configurations, represents the Mth reconfigurable intelligent surface phase configuration; The phase codebook is used to adjust the reflection phase of the RIS according to different channel environments to optimize signal propagation. By inputting the low-dimensional feature vector Z into a generative network or phase codebook generation model, the system dynamically generates a phase codebook suitable for the current environment according to the current channel state and electromagnetic features. The advantage of doing this is that the generation process of the RIS phase codebook is optimized based on the features of the current environment, thereby improving the quality of signal transmission.
[0033] In this embodiment, the variational autoencoder can automatically extract the most important low-dimensional features from high-dimensional features, and through the normal distribution constraint of the latent space, it can generate more diverse features. This low-dimensional representation is not only convenient for calculation, but also can effectively compress the redundant information in the electromagnetic features while retaining the useful signal features.
[0034] In summary, this method compresses electromagnetic features through a variational autoencoder, maps high-dimensional features to a low-dimensional space, and generates a RIS phase codebook based on this. In this way, the system can adaptively adjust the RIS phase according to changes in the electromagnetic environment, thereby improving the efficiency and performance of the 5G wireless communication system. This method can enhance the adaptability of the communication system in complex and dynamic environments and maximize the utilization efficiency of resources.
[0035] Furthermore, the phase selection strategy of the reconfigurable intelligent surface agent selects the phase configuration that best matches the current electromagnetic feature codebook based on the cosine similarity metric function, and the cosine similarity metric function is expressed as: ; By calculating the similarity between the current electromagnetic feature vector Z and each possible phase configuration vector , the phase configuration that best matches the current environment is selected. In this way, the RIS can automatically select the optimal phase according to the electromagnetic feature information, thereby optimizing the signal propagation effect.
[0036] This metric method does not require accurate modeling of the detailed channel model, but selects the appropriate configuration through similarity measurement, which simplifies the computational complexity.
[0037] Furthermore, in the digital twin, multiple multi-agent deep reinforcement learning strategies are simulated in parallel, and the evaluation metrics include cumulative reward, throughput, and bit error rate, and the optimal strategy is selected for physical system execution. The optimal strategy is represented by the following formula: ; where, represents the reward generated by the actions of the base station and the reconfigurable intelligent surface at time t, γ represents the discount factor (0 ≤ γ ≤ 1), which is used to adjust the influence degree of future rewards and control the influence of future rewards, and E sim represents the expectation calculated in the digital twin simulation environment, that is, in all simulation environments, a strategy π is selected to maximize the cumulative reward; In this method, the cooperation and competition among multiple agents are the core ideas. Each agent learns how to make optimal decisions according to the current environmental state, and can share information and cooperate with other agents through the reward mechanism. Deep reinforcement learning approximates the value function or policy function through a neural network to enable effective learning and decision-making in complex and continuous high-dimensional state spaces.
[0038] In the 5G communication environment, the goal of multiple agents is to optimize the system performance by coordinating their actions (such as adjusting antenna phase, spectrum allocation, etc.). The decisions of these agents not only depend on the current system state but also take into account the decisions and actions of other agents.
[0039] Furthermore, when the action space of the reconfigurable intelligent surface agent is to select discrete phase values, the discrete phase values are as follows: ; The above discrete phase set provides a specific operation space for the phase adjustment of the reconfigurable intelligent surface (RIS). These phase values allow the RIS to adjust the phase of the reflected signal according to different environmental conditions, thereby improving the signal strength and quality, ensuring that the signal enhanced by the intelligent surface is in phase with the original signal sent by the base station, or achieving directional adjustment of the beam through phase adjustment, thereby improving the overall signal quality.
[0040] Furthermore, the base station agent increases the signal coverage range by adjusting the beam direction, which is achieved through beamforming technology. This technology allows the base station to dynamically adjust the beam direction and shape of the antenna according to different changes in the environment to more accurately cover the target users.
[0041] Furthermore, the construction of the digital twin is based on ray tracing technology and neural networks to build a high-precision virtual channel model, including the base station, reconfigurable intelligent surface, users, and environmental obstacles. This high-precision virtual channel model not only includes the base station, reconfigurable intelligent surface, and user equipment, but also includes obstacles in the environment (such as buildings, trees, etc.), which have a significant impact on signal propagation.
[0042] Furthermore, the digital twin iterates through real-time feedback with the physical system, and the cycle control is within 10 ms, ensuring that the digital twin can obtain information in a very short time and make rapid adjustments to the physical system, which is crucial for 5G communication systems due to their high real-time and low-latency requirements.
[0043] Reference Figure 3 , the present invention also provides an intelligent adaptive control system for 5G wireless communication antennas for implementing the intelligent adaptive control method. The intelligent adaptive control system includes a base station agent module, a reconfigurable intelligent surface agent module, a state space and action space design module, a reward function calculation module, a distributed training and optimization module, an electromagnetic fingerprint map encoding and feature extraction module, a digital twin module, a policy evaluation and screening module, and a real-time feedback and iterative optimization module; Base station agent module: This module is responsible for optimizing the weight vector of beamforming with the aim of maximizing the signal-to-interference-plus-noise ratio (SINR) of the target user. The base station agent module increases the signal coverage range by adjusting the beam direction and dynamically adjusts the amplitude and phase of the beam according to the real-time requirements of the system.
[0044] Reconfigurable Intelligent Surface (RIS) Agent Module: This module is responsible for adjusting the phase configuration of the RIS to enhance the main path signal and suppress multipath interference. The RIS agent selects the optimal phase values by optimizing the RIS phase matrix and combining the current electromagnetic characteristics to improve communication performance.
[0045] State Space and Action Space Design Module: This module provides an optimization strategy for the system by designing appropriate state space and action space. The state space includes real-time channel state information, electromagnetic characteristics, and user location information; the action space includes the strategies for the base station agent to adjust the amplitude and phase of the beam and the RIS agent to select phase values.
[0046] Reward Function Calculation Module: This module designs the reward function, comprehensively considering factors such as signal-to-interference-plus-noise ratio (SINR), energy consumption, and system interference, to guide each agent to optimize and improve the overall system performance during the training and optimization process through the reward function.
[0047] Distributed Training and Optimization Module: This module is based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm to ensure that each agent independently updates its policy through local observations and regularly shares the global reward signal to achieve multi-agent collaborative optimization.
[0048] Electromagnetic Fingerprint Map Encoding and Feature Extraction Module: This module dynamically encodes the electromagnetic environment through multi-sensor data acquisition and feature extraction methods, and provides guidance for deep reinforcement learning (DRL) by generating the RIS phase codebook to help the agent select the optimal control strategy.
[0049] Digital Twin Module: This module uses ray tracing technology to construct a virtual channel model including the base station, RIS, and users to simulate and evaluate the effects of different strategies. The digital twin validates the DRL strategy in advance through simulation and feedbacks the real-time performance data of the physical system.
[0050] Policy Evaluation and Screening Module: This module sorts different policies according to the cumulative reward, selects the top 3 optimal policies, and applies them to the physical system for execution to ensure the effectiveness of policy selection.
[0051] Real-Time Feedback and Iterative Optimization Module: This module is used to receive the real-time performance feedback from the physical system and use the feedback to update the simulation model parameters. Through the closed-loop feedback mechanism of real-time data, the system performance is continuously optimized.
[0052] The base station agent module and the reconfigurable intelligent surface agent module cooperate with each other by sharing the data of the state space and action space to optimize beamforming and RIS phase configuration; The reward function calculation module and the distributed training and optimization module ensure that each agent optimizes towards the global goal by calculating the reward signal in real time and updating the policies of each agent; The electromagnetic fingerprint encoding and feature extraction module provides the necessary electromagnetic environment data for the base station agent module and the RIS agent module; The digital twin module and the policy evaluation and screening module evaluate different DRL policies through the virtual channel model, and jointly act with the policy feedback and real-time feedback and iterative optimization module to ensure the effectiveness of the policy selection and optimization process of the physical system.
[0053] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
An intelligent adaptive control method for a 1.5G wireless communication antenna is characterized by: The following steps are involved: S1: Agent division and task allocation: The beamforming of the base station and the phase configuration of the reconfigurable smart surface in wireless communication are divided into the base station agent and the reconfigurable smart surface agent respectively. The base station agent is responsible for optimizing the beamforming weight vector to maximize the signal-to-interference-noise ratio of the target user, and the reconfigurable smart surface agent is responsible for adjusting the reconfigurable smart surface phase matrix to enhance the main path signal and suppress multipath interference. S2: State space and action space design: Construct the state space, which includes real-time channel state information, electromagnetic characteristics and user location information; construct the action space, which includes the base station agent adjusting the amplitude and phase of the beam, and the reconfigurable smart surface agent selecting discrete or continuous phase values; S3: Reward function design: Considering the signal-to-noise ratio, energy consumption and interference, a reward function is defined to guide the agent optimization process. The reward function is as follows: ; Among them, α, β, and γ are weight coefficients, which are used to control the influence of signal to interference noise ratio, power, and interference, respectively. k is the signal-to-interference-noise ratio of the kth user, W is the beamforming weight vector of the base station, |W| 2 Indicates the transmit power, Interference indicates the interference power within the system; S4: Distributed training and optimization: Using a multi-agent deep deterministic policy gradient algorithm, the training process updates the strategy through local observations and regularly shares global reward signals to ensure coordination; S5: Dynamically coded electromagnetic fingerprint: Dynamically encode the electromagnetic environment through multi-sensor data acquisition functions and feature extraction methods, convert the acquired electromagnetic features into a phase codebook to guide the actions of multi-agent deep reinforcement learning; S6: Digital twin construction and strategy verification: Through high-precision virtual channel models and random perturbations, digital twins are established to preview multiple groups of multi-agent deep reinforcement learning strategies and provide real-time feedback on the performance data of the physical system; S7: Real-time feedback and iterative optimization: Use the real-time performance feedback of the physical system to update the simulation model parameters to ensure the effectiveness and stability of the optimization process.
2. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, characterized in that: The real-time channel state information in the state space is the base station-reconfigurable smart surface channel matrix G and the reconfigurable smart surface-user channel matrix H. The state space is represented by combining the real-time channel state information and the electromagnetic fingerprint feature F, and the formula is: ; Wherein, Encoder(F) represents the encoding result of the electromagnetic fingerprint.
3. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 2 is characterized in that: The generation process of the phase codebook in step S5 is as follows: S51: Generate electromagnetic eigentensor: ; Among them, K represents the feature dimension, and T represents the time window length; S52: Use variational autoencoder to compress the electromagnetic feature tensor F into a low-dimensional feature vector Z: ; Among them, Z is the encoded low-dimensional feature vector with dimension d; And define the loss function of the variational autoencoder as: ; where F′ represents the reconstructed electromagnetic feature generated by the encoder, μ z , σ z They are respectively represented as the mean and standard deviation of the eigenvector Z, λ is the regularization weight, which is used to control the influence of KL divergence, KL is the Kullback-Leibler divergence; S53: Generate a RIS phase codebook according to the encoded low-dimensional feature vector Z: ; Where C represents the reconfigurable smart surface phase codebook, which contains multiple possible phase configurations. Represented as the Mth reconfigurable smart surface phase configuration.
4. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 3 is characterized in that: The phase selection strategy of the reconfigurable intelligent surface agent selects the phase configuration that best matches the current electromagnetic signature codebook based on the cosine similarity metric function. The cosine similarity metric function is expressed as: 。 5. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, characterized in that: In the digital twin, multiple multi-agent deep reinforcement learning strategies are simulated in parallel. The evaluation indicators include cumulative reward, throughput, and bit error rate, and the optimal strategy is selected for physical system execution. The optimal strategy is expressed by the following formula: ; in, is represented by the reward generated by the base station and the reconfigurable smart surface action at time t, γ is represented by the discount factor used to control the impact of future rewards, E sim Expressed as expectations computed in the digital twin simulation environment.
6. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, characterized in that: When the action space of the reconfigurable intelligent surface agent is to select a discrete phase value, the discrete phase value is: 。 7. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, characterized in that: The base station agent increases the coverage of the signal by adjusting the beam direction.
8. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, characterized in that: The digital twin is constructed based on ray tracing technology and neural networks to build a high-precision virtual channel model, including base stations, reconfigurable smart surfaces, users and environmental obstacles.
9. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 8, characterized in that: The digital twin is iterated through real-time feedback with the physical system, with the cycle controlled within 10ms.
10. An intelligent adaptive control system for a 5G wireless communication antenna, characterized in that: Used to implement the intelligent adaptive control method as described in any one of claims 1-9.
Citation Information
Patent Citations
EH-RIS assisted physical layer secure transmission method for space-ground integrated network
CN117354787A
Intelligent reflection surface communication system deduction optimization method and system based on digital twinning
CN117793754A
RIS-assisted MU-MISO communication system intelligent beam forming method based on deep reinforcement learning
CN118764055A
Multi-agent system networking and resource optimization method
CN118921712A
Beam management for communication via network controlled repeaters and reconfigurable intelligent surfaces
WO2023160802A1
Cited By
5G base station intelligent beam dynamic optimization system
CN120711412A