Intelligent Adaptive Control Method and System for 5G Wireless Communication Antenna
By dividing agents in the 5G wireless communication system and combining multi-agent deep learning and digital twin technology, the phase configuration of base stations and RIS is optimized, and the system performance improvement problem in complex scenarios is solved, achieving significant real-time performance and energy efficiency improvement.
Patent Information
- Application Number
- CN202510538000.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In 5G wireless communication systems, how to achieve dynamic joint adjustment of antenna beamforming and reconstructible intelligent surface phase configuration in complex scenarios to improve system performance, such as throughput, coverage and anti-interference capability.
The beamforming and phase configuration of the reconstructible intelligent surface of the base station are divided into base station agents and reconstructible intelligent surface agents. Through the multi-agent depth deterministic strategy gradient algorithm and digital twin technology, combined with real-time channel state information and electromagnetic characteristics, the signal-to-interference noise ratio and energy consumption are optimized, and multi-sensor data acquisition and high-precision virtual channel model are used for dynamic encoding and feedback optimization.
It realizes explicit collaborative optimization of beamforming and RIS phase in 5G wireless communication antennas, improving the real-time, robustness and energy efficiency of the system, and is suitable for complex 5G communication scenarios such as urban dense areas and high-speed mobile scenarios.
Smart Images

Figure CN120074591B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mobile communications, and specifically to an intelligent adaptive control method and system for 5G wireless communication antennas. Background Art
[0002] In a 5G wireless communication system, the combination of large-scale MIMO (Multiple-Input Multiple-Output) antennas and RIS is considered a key technology to improve spectral efficiency and coverage ability. Beamforming optimizes the signal directivity by adjusting the transmission weights of the antenna array, while RIS changes the signal propagation path by dynamically adjusting the phase of its reflection units.
[0003] However, in complex scenarios of 5G communication (such as environments with high mobility, multi-user density, and significant multipath effects), how to achieve dynamic joint adjustment of antenna beamforming and the phase configuration of reconfigurable intelligent surfaces (RIS) to improve system performance (such as throughput, coverage, and anti-interference ability) remains an unsolved problem. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides an intelligent adaptive control method and system for 5G wireless communication antennas, which solves the problems mentioned in the above background.
[0005] The present invention provides the following technical solutions: The present invention discloses an intelligent adaptive control method for 5G wireless communication antennas, including the following steps:
[0006] S1: Agent Division and Task Allocation: Divide the beamforming of the base station and the phase configuration of the reconfigurable intelligent surface in wireless communication into a base station agent and a reconfigurable intelligent surface agent respectively. The base station agent is responsible for optimizing the beamforming weight vector to maximize the signal-to-interference-plus-noise ratio (SINR) of the target user, and the reconfigurable intelligent surface agent is responsible for adjusting the reconfigurable intelligent surface phase matrix to enhance the main path signal and suppress multipath interference;
[0007] S2: State Space and Action Space Design: Construct a state space, including real-time channel state information, electromagnetic characteristics, and user location information; construct an action space, including the base station agent adjusting the amplitude and phase of the beam, and the reconfigurable intelligent surface agent selecting discrete or continuous phase values;
[0008] S3: Reward Function Design: Considering SINR, energy consumption, and interference comprehensively, define a reward function to guide the optimization process of the base station agent. The reward function is as follows:
[0009] ;
[0010] where α, β, and γ are weight coefficients, which are used to control the influence of SINR, power, and interference respectively, and SINR kis the signal-to-interference-plus-noise ratio (SINR) of the k-th user, \(W\) represents the beamforming weight vector of the base station, and \(|W|\) 2 represents the transmit power, and Interference represents the interference power within the system;
[0011] S4: Distributed training and optimization: The multi-agent deep deterministic policy gradient algorithm is adopted. During the training process, the policy is updated through local observations, and the global reward signal is shared regularly to ensure coordination;
[0012] S5: Dynamic encoding of electromagnetic fingerprint maps: The electromagnetic environment is dynamically encoded through multi-sensor data acquisition functions and feature extraction methods, and the obtained electromagnetic features are converted into phase codebooks to guide the actions of multi-agent deep reinforcement learning;
[0013] S6: Digital twin construction and policy verification: Through a high-precision virtual channel model and random perturbations, a digital twin is established to preview multiple groups of multi-agent deep reinforcement learning policies and instantaneously feedback the performance data of the physical system;
[0014] S7: Real-time feedback and iterative optimization: The real-time performance feedback of the physical system is used to update the parameters of the simulation model to ensure the effectiveness and stability of the optimization process.
[0015] As a preferred solution, the real-time channel state information in the state space is the base station-reconfigurable intelligent surface channel matrix \(G\) and the reconfigurable intelligent surface-user channel matrix \(H\). The state space is represented by combining the real-time channel state information and the electromagnetic feature \(F\), and the formula is:
[0016] ;
[0017] where \(Encoder(F)\) represents the encoding result of the electromagnetic feature.
[0018] As a preferred solution, the generation process of the phase codebook in step S5 is as follows:
[0019] S51: Generate an electromagnetic feature tensor:
[0020] ;
[0021] where \(K\) represents the feature dimension and \(T\) represents the time window length;
[0022] S52: Use a variational autoencoder to compress the electromagnetic feature tensor \(F\) into a low-dimensional feature vector \(Z\):
[0023] ;
[0024] And the loss function of the variational autoencoder is defined as:
[0025] ;
[0026] Among them, F′ represents the reconstructed electromagnetic feature generated by the encoder, μz and σz respectively represent the mean and standard deviation of the feature vector Z, λ is the regularization weight used to control the influence of the KL divergence, and KL is the Kullback-Leibler divergence;
[0027] S53: Generate the RIS phase codebook according to the encoded low-dimensional feature vector Z:
[0028] ;
[0029] Among them, C represents the reconfigurable intelligent surface phase codebook, which contains multiple possible phase configurations, and φ M represents the Mth reconfigurable intelligent surface phase configuration.
[0030] As a preferred solution, the phase selection strategy of the reconfigurable intelligent surface agent selects the phase configuration that best matches the current phase codebook based on the cosine similarity metric function, and the cosine similarity metric function is expressed as:
[0031] .
[0032] As a preferred solution, in the digital twin, multiple multi-agent deep reinforcement learning strategies are simulated in parallel, and the evaluation metrics include cumulative reward, throughput, and bit error rate, and the optimal strategy is selected for physical system execution. The optimal strategy is represented by the following formula:
[0033] ;
[0034] Among them, r t (W, φ) represents the reward generated by the base station and the reconfigurable intelligent surface actions at time t, γ represents the discount factor used to control the influence of future rewards, and E sim represents the expectation calculated in the digital twin simulation environment.
[0035] As a preferred solution, when the action space of the reconfigurable intelligent surface agent is to select discrete phase values, the discrete phase values are:
[0036] .
[0037] As a preferred solution, the base station agent increases the signal coverage range by adjusting the beam direction.
[0038] As a preferred solution, the digital twin is constructed based on ray tracing technology and neural networks to build a high-precision virtual channel model, including the base station, reconfigurable intelligent surface, user, and environmental obstacles.
[0039] As a preferred solution, the digital twin is iterated through real-time feedback with the physical system, and the cycle control is within 10 ms.
[0040] The present invention also provides an intelligent adaptive control system for a 5G wireless communication antenna, which is used to implement the intelligent adaptive control method.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] Through the innovative combination of the multi-agent DRL framework, electromagnetic fingerprint dynamic coding, and digital twin virtual-real linkage, the present invention realizes the explicit collaborative optimization of beamforming and RIS phase in 5G wireless communication antennas. It not only solves the limitations of traditional methods in dynamic environments but also improves the real-time performance, robustness, and energy efficiency of the system through the perception-decision-validation closed-loop. It has significant innovation and practical value and is applicable to complex 5G communication scenarios (such as urban dense areas, high-speed mobile scenarios, etc.). BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a working flow chart of agent division in the intelligent adaptive control method of the present invention;
[0044] Figure 2 It is a working flow chart of dynamic coding and digital twin in the intelligent adaptive control method of the present invention;
[0045] Figure 3 It is a schematic flow chart of the intelligent adaptive control system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] Please refer to Figure 1-2 The intelligent adaptive control method for a 5G wireless communication antenna in this embodiment includes the following steps:
[0048] S1: Agent Division and Task Allocation: Divide the beamforming of the base station and the phase configuration of the reconfigurable intelligent surface (RIS) in wireless communication into a base station agent and a reconfigurable intelligent surface agent respectively. The base station agent is responsible for optimizing the beamforming weight vector to maximize the signal-to-interference-plus-noise ratio (SINR) of the target user, and the reconfigurable intelligent surface agent is responsible for adjusting the phase matrix of the reconfigurable intelligent surface to enhance the main path signal and suppress multipath interference. The goal of optimizing the phase configuration of the RIS is to enhance the signal quality and reduce interference.
[0049] In a 5G wireless communication system, the base station and the reconfigurable intelligent surface (RIS) are two important components. To improve the system performance, tasks and optimization goals need to be reasonably allocated to different agents. In this step, the division of agents and task allocation lay the foundation for subsequent optimization work.
[0050] S2: State Space and Action Space Design: Construct a state space that includes real-time channel state information, electromagnetic characteristics, and user location information; construct an action space that includes the base station agent adjusting the amplitude and phase of the beam, and the reconfigurable intelligent surface agent selecting discrete or continuous phase values.
[0051] State Space: In wireless communication, the state space contains important information related to the operating conditions of the communication system. For example:
[0052] Real-time Channel State Information (CSI): Provides signal quality information of the current communication link, including signal strength, noise interference, user location, etc.
[0053] Electromagnetic Characteristics: Such as the electromagnetic wave propagation characteristics of the environment, affected by factors such as weather and obstacles.
[0054] User Location Information: The real-time location of the user equipment, which helps the base station and RIS adjust their configurations.
[0055] Action Space: Provides a decision space for the behavior of the agent. The action spaces of the base station agent and the RIS agent are respectively:
[0056] Base Station Agent: Can adjust the amplitude and phase of the beam to optimize the signal transmission quality.
[0057] RIS Agent: Can select different phase values, which may be discrete or continuous, to adjust the reflection characteristics of the RIS and optimize the signal path.
[0058] S3: Reward Function Design: Considering the signal-to-interference-plus-noise ratio, energy consumption, and interference comprehensively, define a reward function to guide the optimization process of the base station agent. The reward function is as follows:
[0059] ;
[0060] where α, β, and γ are weight coefficients, used to control the impacts of signal-to-interference-plus-noise ratio, power, and interference respectively, and SINR k is the signal-to-interference-plus-noise ratio of the k-th user, W represents the beamforming weight vector of the base station, and ∣W∣ 2 represents the transmission power, and Interference represents the interference power within the system;
[0061] The signal-to-interference-plus-noise ratio is the core optimization objective in wireless communication. The goal is to maximize SINR, ensure the maximization of signal strength, and minimize interference; energy efficiency is considered during the optimization process to reduce unnecessary energy consumption, meeting the requirements of green communication; by reducing multipath interference and other signal interferences, the communication quality is improved.
[0062] S4: Distributed training and optimization: The multi-agent deep deterministic policy gradient (MADDPG) algorithm is adopted. During the training process, the policy is updated through local observations, and the global reward signal is shared regularly to ensure coordination. It is ensured that each agent independently updates the policy through local observations and regularly shares the global reward signal. The multi-agent deep deterministic policy gradient algorithm is a common multi-agent reinforcement learning method and will not be elaborated here;
[0063] Each agent can update the policy through local observations. However, to ensure the synergy effect of the system, the agents will regularly share the global reward signal. This distributed learning method allows the agents to independently optimize without complete information, while promoting cooperation through sharing information, thereby improving the performance of the overall system.
[0064] S5: Dynamic encoding of electromagnetic fingerprint maps: The electromagnetic environment is dynamically encoded through multi-sensor data acquisition functions and feature extraction methods, and the obtained electromagnetic features are converted into a phase codebook to guide the actions of multi-agent deep reinforcement learning (DRL);
[0065] To further improve the adaptability and accuracy of the communication system, the electromagnetic environment is dynamically encoded through various sensor data (such as radio frequency sensors, millimeter-wave radars, and cameras). Specifically:
[0066] By collecting electromagnetic data in the environment and combining feature extraction methods, real-time electromagnetic environment information can be obtained; these feature information are converted into a phase codebook to guide the action selection during the multi-agent deep reinforcement learning process. This dynamic encoding of electromagnetic fingerprint maps helps the agents to respond promptly to the changing electromagnetic environment and optimize the policy.
[0067] S6: Digital twin construction and policy verification: Through a high-precision virtual channel model and random perturbations, a digital twin is established to pre-enact multiple groups of multi-agent deep reinforcement learning policies and immediately feedback the performance data of the physical system;
[0068] Digital twin technology can simulate the actual communication environment by constructing a high-precision virtual channel model and introduce random perturbations to verify the system performance, which provides a virtual test platform for the optimization process. Through the digital twin, different multi-agent deep reinforcement learning strategies can be rehearsed in the simulation environment, and the simulation results can be fed back in real time to help adjust the model parameters and strategies to ensure the actual effect in the physical system.
[0069] S7: Real-time feedback and iterative optimization: Use the real-time performance feedback of the physical system to update the simulation model parameters to ensure the effectiveness and stability of the optimization process;
[0070] Finally, the real-time performance feedback of the physical system is used for continuous optimization. By real-time monitoring the performance data of the actual communication system, the feedback information is sent back to the simulation model:
[0071] Update the parameters of the simulation model to make it more in line with the actual situation, while ensuring the stability and effectiveness of the optimization process, and achieve adaptive dynamic optimization. In this way, through the iterative process, the performance of the system is continuously improved.
[0072] The base station agent independently optimizes the beamforming weight vector:
[0073] ;
[0074] where N BS is the number of base station antennas, and the goal is to maximize the signal-to-interference-plus-noise ratio of the target user. The reconfigurable intelligent surface agent is responsible for adjusting the RIS phase matrix:
[0075] ;
[0076] This matrix contains the phase settings of each unit of the RIS, where N RIS is the number of RIS units, is the phase value represented in complex exponential form, which is used to adjust the phase response of the RIS unit. The goal is to enhance the main path signal and suppress multipath interference;
[0077] In summary, this solution realizes the explicit collaborative optimization of beamforming and RIS phase in 5G wireless communication antennas through the innovative combination of multi-agent DRL framework, electromagnetic fingerprint dynamic coding, and digital twin virtual-real interaction. It not only solves the limitations of traditional methods in dynamic environments but also improves the real-time performance, robustness, and energy efficiency of the system through the perception-decision-validation closed-loop, with significant innovation and practical value, and is applicable to complex 5G communication scenarios (such as urban dense areas, high-speed mobile scenarios, etc.).
[0078] Furthermore, the real-time channel state information in the state space is the base station-reconfigurable intelligent surface channel matrix G and the reconfigurable intelligent surface-user channel matrix H. The state space is represented by combining the real-time channel state information and the electromagnetic feature F, and the formula is:
[0079] ;
[0080] where Encoder(F) represents the encoding result of the electromagnetic feature;
[0081] In the intelligent adaptive control method of the 5G wireless communication antenna described above, the core idea is to construct a dynamically changing state space model based on the real-time channel state information and the electromagnetic feature to optimize the performance of the communication system;
[0082] Among them, the real-time channel state information (CSI) refers to the channel matrix between the base station and the reconfigurable intelligent surface (RIS), and the channel matrix between the reconfigurable intelligent surface and the user. It is a very important part of the communication system because it determines the quality of signal propagation;
[0083] Base station-reconfigurable intelligent surface channel matrix G: This represents the wireless communication channel state between the base station and the RIS. The RIS is a technology that can optimize signal propagation by adjusting its surface state and can play roles such as enhancing signals and improving coverage in wireless communication;
[0084] Reconfigurable intelligent surface-user channel matrix, H: This represents the wireless communication channel state between the reconfigurable intelligent surface and the user. By adjusting the reflection units on the RIS, the propagation path of the signal can be affected, thereby optimizing the communication quality;
[0085] These matrices G and H together constitute the real-time channel state information of the communication system. They can accurately describe the wireless channel quality between the base station, the RIS, and the user, facilitating the adaptive optimization of communication strategies.
[0086] The electromagnetic fingerprint is obtained by analyzing the wireless signals emitted by the device to extract its unique electromagnetic feature. The electromagnetic feature of each device is unique and can be used to identify the device or optimize the communication channel;
[0087] where the state space s t is a comprehensive data structure that contains all the real-time channel state information and the features related to the electromagnetic feature. It provides comprehensive input information for the intelligent adaptive control algorithm and can help the system dynamically adjust the communication strategy according to the current channel state and electromagnetic feature to achieve an optimized communication effect;
[0088] Based on the above state space s t, the behavior of the system can be adaptively adjusted through intelligent algorithms. These algorithms can make predictions and decisions based on real-time updated channel information and electromagnetic characteristics, optimize operations such as signal transmission, power allocation, and resource scheduling, combine traditional channel state information and electromagnetic characteristics, and conduct comprehensive optimization through intelligent algorithms. In this way, more accurate and flexible communication control can be achieved, improving the performance and reliability of the 5G wireless communication system.
[0089] Further, the generation process of the phase codebook in step S5 is as follows:
[0090] S51: Generate an electromagnetic feature tensor:
[0091] ;
[0092] Among them, K represents the feature dimension, that is, the number or complexity of the electromagnetic features extracted at each moment, T represents the time window length, that is, the number of frames within the observed time range. Assuming there is a set of electromagnetic features at each moment, then T can be regarded as the set of electromagnetic features at different time points within this time window;
[0093] The purpose of generating this electromagnetic feature tensor is to capture the electromagnetic features in the environment, which may include reflections of buildings, interference in signal propagation, etc. This tensor is the original high-dimensional data and needs to be further processed to extract useful information.
[0094] S52: Use a variational autoencoder to compress the electromagnetic feature tensor F into a low-dimensional feature vector Z:
[0095] ;
[0096] Z is the encoded low-dimensional feature vector with dimension d, and d is much smaller than K. This is because through the encoder part of the VAE, we compress the high-dimensional data into a low-dimensional space. The low-dimensional representation can retain the important information in the input data but reduce redundancy and noise.
[0097] And define the loss function of the variational autoencoder as:
[0098] ;
[0099] Among them, F′ represents the reconstructed electromagnetic feature generated by the encoder, μz and σz respectively represent the mean and standard deviation of the feature vector Z, λ is the regularization weight used to control the influence of the KL divergence, and KL is the Kullback-Leibler divergence used to measure the difference between the distribution of the latent variable and the standard normal distribution;
[0100] is the mean square error between the input electromagnetic feature tensor F and the output F′ after reconstruction by the autoencoder, aiming to enable the autoencoder to reconstruct the input data as accurately as possible;
[0101] is the Kullback-Leibler divergence between the distribution of the latent variable Z and the standard normal distribution, which is used to ensure that the distribution of the latent variable Z learned by the encoder is close to the standard normal distribution, so that the generative model can generate real and diverse samples;
[0102] Through this process, the variational autoencoder can not only compress high-dimensional electromagnetic feature data, but also map it to a latent space, retaining important electromagnetic information, which is suitable for subsequent RIS phase codebook generation.
[0103] S53: Generate the RIS phase codebook according to the encoded low-dimensional feature vector Z:
[0104] ;
[0105] where C represents the reconfigurable intelligent surface phase codebook, which contains multiple possible phase configurations, and φ M represents the Mth reconfigurable intelligent surface phase configuration;
[0106] The phase codebook is used to adjust the reflection phase of the RIS according to different channel environments to optimize signal propagation. By inputting the low-dimensional feature vector Z into a generative network or a phase codebook generation model, the system dynamically generates a phase codebook adapted to the current environment according to the current channel state and electromagnetic features. The advantage of this is that the generation process of the RIS phase codebook is optimized based on the features of the current environment, thereby improving the quality of signal transmission.
[0107] In this embodiment, the variational autoencoder can automatically extract the most important low-dimensional features from high-dimensional features, and through the normal distribution constraint of the latent space, it can generate more diverse features. This low-dimensional representation is not only convenient for calculation, but also can effectively compress the redundant information in electromagnetic features while retaining useful signal features.
[0108] In summary, this method compresses electromagnetic features through a variational autoencoder, maps high-dimensional features to a low-dimensional space, and generates an RIS phase codebook based on this. In this way, the system can adaptively adjust the RIS phase according to changes in the electromagnetic environment, thereby improving the efficiency and performance of the 5G wireless communication system. This method can improve the adaptability of the communication system in complex and dynamic environments and maximize the utilization efficiency of resources.
[0109] Further, the phase selection strategy of the reconfigurable intelligent surface agent selects the phase configuration that best matches the current phase codebook based on the cosine similarity metric function, and the cosine similarity metric function is expressed as:
[0110] ;
[0111] By calculating the similarity between the current electromagnetic feature vector Z and each possible phase configuration vector φ M a phase configuration that best matches the current environment is selected. In this way, the RIS can automatically select the optimal phase according to the electromagnetic feature information, thereby optimizing the signal propagation effect.
[0112] This metric method does not require accurate modeling of the detailed channel model, but selects the appropriate configuration through similarity measurement, which simplifies the computational complexity.
[0113] Further, in the digital twin, multiple multi-agent deep reinforcement learning strategies are simulated in parallel. The evaluation metrics include cumulative reward, throughput, and bit error rate, and the optimal strategy is selected for physical system execution. The optimal strategy is represented by the following formula:
[0114] ;
[0115] where r t (W, φ) represents the reward generated by the base station and the reconfigurable intelligent surface actions at time t, γ represents the discount factor (0 ≤ γ ≤ 1), which is used to adjust the influence degree of future rewards and control the influence of future rewards, and E sim represents the expectation calculated in the digital twin simulation environment, that is, in all simulation environments, a strategy π is selected to maximize the cumulative reward;
[0116] In this method, the cooperation and competition among multiple agents are the core ideas. Each agent learns how to make optimal decisions according to the current environmental state, and can share information and cooperate with other agents through the reward mechanism. Deep reinforcement learning approximates the value function or policy function through a neural network to perform effective learning and decision-making in a complex, continuous high-dimensional state space.
[0117] In the 5G communication environment, the goal of multiple agents is to optimize the system performance by coordinating their actions (such as adjusting antenna phases, spectrum allocation, etc.). The decisions of these agents not only depend on the current system state, but also consider the decisions and actions of other agents.
[0118] Further, when the action space of the reconfigurable intelligent surface agent is to select discrete phase values, the discrete phase values are:
[0119] ;
[0120] The above discrete phase set provides a specific operation space for controlling the phase adjustment of the reconfigurable intelligent surface (RIS). These phase values allow the RIS to adjust the phase of the reflected signal according to different environmental conditions, thereby improving the signal strength and quality, ensuring that the signal enhanced by the intelligent surface is in phase with the original signal sent by the base station, or achieving directional adjustment of the beam through phase adjustment, thereby improving the overall signal quality.
[0121] Furthermore, the base station agent increases the signal coverage range by adjusting the beam direction, which is achieved through beamforming technology. This technology allows the base station to dynamically adjust the beam direction and shape of the antenna according to different changes in the environment to more accurately cover the target users.
[0122] Furthermore, the construction of the digital twin is based on ray tracing technology and neural networks to build a high-precision virtual channel model, including the base station, reconfigurable intelligent surface, user, and environmental obstacles. This high-precision virtual channel model not only includes the base station, reconfigurable intelligent surface, and user equipment, but also includes obstacles in the environment (such as buildings, trees, etc.), which have a significant impact on signal propagation.
[0123] Furthermore, the digital twin performs real-time feedback iteration with the physical system, and the cycle control is within 10 ms, ensuring that the digital twin can obtain information in a very short time and make rapid adjustments to the physical system, which is crucial for the 5G communication system due to its high real-time and low-latency requirements.
[0124] Reference Figure 3 , the present invention also provides an intelligent adaptive control system for a 5G wireless communication antenna for implementing the intelligent adaptive control method. The intelligent adaptive control system includes a base station agent module, a reconfigurable intelligent surface agent module, a state space and action space design module, a reward function calculation module, a distributed training and optimization module, an electromagnetic fingerprint map encoding and feature extraction module, a digital twin module, a policy evaluation and screening module, and a real-time feedback and iterative optimization module;
[0125] Base station agent module: This module is responsible for optimizing the weight vector of beamforming with the aim of maximizing the signal-to-interference-plus-noise ratio (SINR) of the target user. The base station agent module increases the signal coverage range by adjusting the beam direction and dynamically adjusts the amplitude and phase of the beam according to the real-time requirements of the system.
[0126] Reconfigurable Intelligent Surface (RIS) Agent Module: This module is responsible for adjusting the phase configuration of the RIS to enhance the main path signal and suppress multipath interference. The RIS agent selects the optimal phase values by optimizing the RIS phase matrix and combining the current electromagnetic characteristics to improve communication performance.
[0127] State Space and Action Space Design Module: This module provides an optimization strategy for the system by designing appropriate state space and action space. The state space includes real-time channel state information, electromagnetic characteristics, and user location information; the action space includes the strategies for the base station agent to adjust the amplitude and phase of the beam and the RIS agent to select phase values.
[0128] Reward Function Calculation Module: This module designs the reward function, comprehensively considering factors such as Signal-to-Interference-plus-Noise Ratio (SINR), energy consumption, and system interference, to guide each agent to optimize and improve the overall system performance during the training and optimization process through the reward function.
[0129] Distributed Training and Optimization Module: This module is based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, ensuring that each agent independently updates its policy through local observations and regularly shares the global reward signal to achieve multi-agent collaborative optimization.
[0130] Electromagnetic Fingerprint Encoding and Feature Extraction Module: This module dynamically encodes the electromagnetic environment through multi-sensor data acquisition and feature extraction methods, and provides guidance for Deep Reinforcement Learning (DRL) by generating the RIS phase codebook to help the agent select the optimal control strategy.
[0131] Digital Twin Module: This module uses ray tracing technology to build a virtual channel model containing the base station, RIS, and users, which is used to simulate and evaluate the effects of different strategies. The digital twin validates the DRL strategy in advance through simulation and feedbacks the real-time performance data of the physical system.
[0132] Policy Evaluation and Screening Module: This module ranks different policies according to the cumulative reward, selects the top 3 optimal policies, and applies them to the physical system for execution to ensure the effectiveness of policy selection.
[0133] Real-Time Feedback and Iterative Optimization Module: This module is used to receive the real-time performance feedback from the physical system and use the feedback to update the simulation model parameters, continuously optimizing the system performance through the closed-loop feedback mechanism of real-time data.
[0134] The base station agent module and the reconfigurable intelligent surface agent module cooperate with each other by sharing the data of the state space and action space to optimize beamforming and RIS phase configuration;
[0135] The reward function calculation module and the distributed training and optimization module ensure that each agent optimizes towards the global goal by calculating the reward signal in real time and updating the policies of each agent;
[0136] The electromagnetic fingerprint spectrum encoding and feature extraction module provides the necessary electromagnetic environment data for the base station agent module and the RIS agent module;
[0137] The digital twin module and the policy evaluation and screening module evaluate different DRL policies through the virtual channel model, and through policy feedback and real-time feedback, act together with the iterative optimization module to ensure the effectiveness of the policy selection and optimization process of the physical system.
[0138] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. 5G wireless communication antenna intelligent adaptive control method, characterized in that: It includes the following steps: S1: Agent division and task assignment: Divide the beamforming of the base station and the phase configuration of the reconfigurable intelligent surface in wireless communication into a base station agent and a reconfigurable intelligent surface agent respectively. The base station agent is responsible for optimizing the beamforming weight vector to maximize the signal-to-interference-plus-noise ratio (SINR) of the target user, and the reconfigurable intelligent surface agent is responsible for adjusting the reconfigurable intelligent surface phase matrix to enhance the main path signal and suppress multipath interference; S2: State space and action space design: Construct a state space, including real-time channel state information, electromagnetic characteristics, and user location information; construct an action space, including the base station agent adjusting the amplitude and phase of the beam, and the reconfigurable intelligent surface agent selecting discrete or continuous phase values; S3: Reward function design: Considering SINR, energy consumption, and interference comprehensively, define a reward function to guide the optimization process of the base station agent. The reward function is as follows: ; Among them, α, β, and γ are weight coefficients, which are used to control the effects of signal-to-interference-plus-noise ratio, power, and interference, respectively. SINR k is the signal-to-interference-plus-noise ratio of the k-th user, W represents the beamforming weight vector of the base station, and ∣W∣ 2 represents the transmit power, and Interference represents the interference power within the system; S4: Distributed training and optimization: Adopt the multi-agent deep deterministic policy gradient algorithm. In the training process, update the policy through local observations and regularly share the global reward signal to ensure coordination; S5: Dynamically encode the electromagnetic fingerprint map: Dynamically encode the electromagnetic environment through the multi-sensor data acquisition function and feature extraction method, and convert the obtained electromagnetic characteristics into a phase codebook to guide the actions of multi-agent deep reinforcement learning; The generation process of the phase codebook in step S5 is as follows: S51: Generate an electromagnetic feature tensor: ; where K represents the feature dimension and T represents the time window length; S52: Use a variational autoencoder to compress the electromagnetic feature tensor F into a low-dimensional feature vector Z: ; And define the loss function of the variational autoencoder as: ; where F′ represents the reconstructed electromagnetic feature generated by the encoder, μz and σz represent the mean and standard deviation of the feature vector Z respectively, λ is the regularization weight used to control the influence of the KL divergence, and KL is the Kullback-Leibler divergence; S53: Generate the RIS phase codebook according to the encoded low-dimensional feature vector Z: ; Among them, C represents the phase codebook of the reconfigurable intelligent surface, and φ M represents the M-th phase configuration of the reconfigurable intelligent surface; S6: Digital twin construction and policy verification: Establish a digital twin through a high-precision virtual channel model and random perturbations, pre-act multiple sets of multi-agent deep reinforcement learning policies, and real-time feedback the performance data of the physical system; S7: Real-time feedback and iterative optimization: Use the real-time performance feedback of the physical system to update the simulation model parameters to ensure the effectiveness and stability of the optimization process.
2. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, wherein: The real-time channel state information in the state space is the base station - reconfigurable intelligent surface channel matrix G and the reconfigurable intelligent surface - user channel matrix H. The state space is represented by combining the real-time channel state information and the electromagnetic feature F, and the formula is: ; where Encoder(F) represents the encoding result of the electromagnetic feature.
3. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, characterized in that: The phase selection strategy of the reconfigurable intelligent surface agent selects the most matching phase configuration in the current phase codebook based on the cosine similarity metric function. The cosine similarity metric function is expressed as: 。 4. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, characterized in that: In the digital twin, multiple multi-agent deep reinforcement learning policies are simulated in parallel. The evaluation metrics include cumulative reward, throughput, and bit error rate, and the optimal policy is selected for physical system execution. The optimal policy is represented by the following formula: ; where r t (W, φ) represents the reward generated by the actions of the base station and the reconfigurable intelligent surface at time t, and γ represents the discount factor used to control the impact of future rewards, E sim represents the expectation calculated in the digital twin simulation environment.
5. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, characterized in that: When the action space of the reconfigurable intelligent surface agent is to select discrete phase values, the discrete phase values are: 。 6. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, wherein: The base station agent increases the coverage range of the signal by adjusting the beam direction.
7. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 1, wherein: The digital twin is constructed based on ray tracing technology and neural networks to build a high-precision virtual channel model, including base stations, reconfigurable intelligent surfaces, users, and environmental obstacles.
8. The intelligent adaptive control method of the 5G wireless communication antenna according to claim 7, characterized in that: The digital twin is iterated through real-time feedback with the physical system, and the cycle control is within 10 ms. An intelligent adaptive control system for 9.5G wireless communication antennas, which is used to implement the intelligent adaptive control method according to any one of claims 1-8.
Citation Information
Patent Citations
EH-RIS assisted physical layer secure transmission method for space-ground integrated network
CN117354787A
Intelligent reflection surface communication system deduction optimization method and system based on digital twinning
CN117793754A