Channel optimization method and system applied to multi-band antenna array

By jointly training generative adversarial networks and deep reinforcement learning networks, combined with multi-band antenna arrays and physical constraints, channel optimization and resource allocation coordination of wireless communication systems in complex dynamic environments are achieved, solving the problems of insufficient adaptability and real-time feedback in existing systems and improving the overall performance of the communication system.

CN120357931BActive Publication Date: 2025-09-12SICHUAN JIUZHOU SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510851738.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-12
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Existing wireless communication systems have difficulty achieving coordinated optimization of channel processing and resource allocation in complex and changeable wireless channel environments. They have poor adaptability, delayed resource allocation, and non-real-time feedback mechanisms, and are unable to meet the high reliability requirements in high-speed mobile and dynamic scenarios.

Method used

A joint training method of generative adversarial networks and deep reinforcement learning networks is adopted to obtain channel state information through a multi-band antenna array, combine physical constraints to perform channel optimization, and adjust resource configuration parameters in real time, including transmit power, operating frequency band, and modulation method, to achieve end-to-end system performance optimization.

Benefits of technology

It improves the accuracy and authenticity of channel optimization, realizes dynamic intelligent resource allocation, enhances the robustness and adaptability of the system, improves spectrum efficiency and energy efficiency, and ensures communication quality in complex dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357931B_ABST
    Figure CN120357931B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of communication technology and discloses a channel optimization method and system for multi-band antenna arrays. The method includes: obtaining channel state information; optimizing the channel state information by using a generative adversarial network in combination with physical constraints; inputting the optimized channel state information into a deep reinforcement learning network to dynamically determine the resource allocation strategy and output resource configuration parameters; feeding back the resource configuration parameters in a closed loop to the front-end signal acquisition and noise reduction module in real time; wherein, the generative adversarial network and the deep reinforcement learning network are jointly trained, and the intermediate layer feature information of the generative adversarial network discriminator is used as additional state information and input into the policy network of the deep reinforcement learning network. The present invention improves the reliability and spectrum efficiency of wireless communication systems, reduces the risk of transmission interruption in high-mobility scenarios, and provides more stable and efficient communication guarantees for millimeter wave and Sub-6GHz hybrid networking scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a channel optimization method and system applied to a multi-band antenna array. Background Art

[0002] With the commercial deployment of 5G technology and the evolution of 6G, wireless communication systems are facing unprecedented challenges. On the one hand, the expansion of spectrum resources to higher frequency bands (such as millimeter waves) provides greater bandwidth, but also comes with increased path loss, signal blocking, and susceptibility to hardware nonlinear distortion. On the other hand, the booming development of applications such as the Internet of Things and the Internet of Vehicles requires communication systems to support massive connections, ultra-low latency, and high reliability, especially in dynamic scenarios with high-speed mobility and rapidly changing channel environments.

[0003] Traditional wireless communication system optimization methods typically employ a separate design: channel estimation and signal processing (such as filtering and noise reduction) and resource allocation (such as power control, frequency band selection, and modulation and coding) are optimized as independent modules. For example, traditional methods such as Kalman filtering are used for channel tracking and noise reduction, while algorithms based on fixed rules or offline optimization are used for resource allocation. These methods have the following limitations:

[0004] 1. Poor adaptability: It is difficult to effectively cope with complex and changing wireless channel environments, especially under mixed frequency bands, high mobility and non-Gaussian noise interference, where performance degrades sharply.

[0005] 2. Insufficient physical constraints: Although data-driven deep learning noise reduction methods (such as basic GAN) are powerful, they may generate signals that do not conform to the actual physical channel characteristics (such as delay spread and Doppler shift), affecting the accuracy of subsequent processing.

[0006] 3. Resource allocation lag: Static or slowly adaptive resource allocation strategies cannot match rapidly changing channel conditions in real time, resulting in low spectrum efficiency, high bit error rates, or connection interruptions.

[0007] 4. Feedback mechanism limitations: Traditional closed-loop real-time feedback mechanisms may have problems such as extended time, high transmission error rate, or fixed feedback cycle, and cannot meet the needs of fast and reliable policy updates in dynamic environments.

[0008] 5. Lack of collaborative optimization: The separate optimization of channel processing and resource allocation ignores the intrinsic connection between the two, making it difficult to achieve system-level end-to-end performance optimization.

[0009] While previous studies have attempted to apply deep learning to wireless communication optimization, such as using GANs alone for channel generation or noise reduction, or reinforcement learning for resource allocation, these approaches often fail to fully address all of the aforementioned issues. In particular, there remains room for improvement in areas such as physical realism constraints, multi-module collaborative optimization, and real-time robust feedback. Therefore, there is an urgent need for an intelligent communication approach that can jointly optimize channel processing and resource allocation and adapt to complex dynamic environments in real time. Summary of the Invention

[0010] In order to solve the problems existing in the above-mentioned prior art, the technical solutions provided by the present invention include:

[0011] A channel optimization method applied to a multi-band antenna array includes the following steps:

[0012] S1. Acquire channel state information through the multi-band antenna array in the signal acquisition and noise reduction module;

[0013] S2. Using a generative adversarial network, combined with physical constraints, to optimize the channel state information;

[0014] S3. Input the optimized channel state information into the deep reinforcement learning network, dynamically determine the resource allocation strategy, and output resource configuration parameters; the resource configuration parameters include transmit power, operating frequency band, and modulation mode;

[0015] S4. The resource configuration parameters are fed back to the front-end signal acquisition and noise reduction module in a closed loop in real time;

[0016] The generative adversarial network and the deep reinforcement learning network are jointly trained, and the intermediate layer feature information of the generative adversarial network discriminator is input into the policy network of the deep reinforcement learning network as additional state information.

[0017] Preferably, the multi-band antenna array includes:

[0018] Multi-channel signal fusion module, used to output a multi-dimensional channel state information matrix containing delay spread, Doppler frequency shift and angle domain characteristics;

[0019] Noise floor calibration circuit to eliminate hardware nonlinear distortion using a reference signal.

[0020] Preferably, the training method of the generative adversarial network includes:

[0021] Inputting the channel state information into the generative adversarial network, extracting noise features through a dynamic attention mechanism, and generating a preliminary noise reduction signal;

[0022] Inputting the preliminary noise reduction signal into the channel impulse response model to verify whether the difference in the delay spread distribution between the delay spread and the Doppler frequency shift exceeds a preset tolerance. If so, triggering retraining of the generative adversarial network and increasing the physical constraint loss weight;

[0023] Based on the real-time signal quality, the weight coefficients of adversarial loss, perceptual loss and physical constraint loss are dynamically adjusted through a hybrid loss function.

[0024] Preferably, the resource allocation strategy includes:

[0025] Divide the transmit power into discrete action sets ranging from 10dBm to 30dBm, dynamically adjust the power allocation in steps of 1dBm, and keep the interval between adjacent power switching at no less than 50ms;

[0026] Seamless switching between mmWave and Sub-6GHz dual-bands with a switching latency of no more than 150μs;

[0027] Dynamically select one of the QPSK, 16QAM, or 64QAM modulation modes based on the target channel signal-to-noise ratio to keep the bit error rate below 1e-4.

[0028] Preferably, the joint training method includes:

[0029] Calculating loss terms and assigning weights to each loss term, the loss terms including at least: a denoising loss associated with the optimized channel state information output by the generative adversarial network, and a physical constraint loss associated with satisfying physical constraints; and a resource allocation loss associated with resource configuration parameters output by the deep reinforcement learning network;

[0030] Updating parameters of the generative adversarial network based on the denoising loss and the physical constraint loss;

[0031] Based on the resource allocation loss, parameters of the policy network of the deep reinforcement learning network are updated, wherein the update of the policy network also utilizes intermediate layer feature information coupled from the discriminator of the generative adversarial network as additional state information.

[0032] Preferably, the joint training method also includes: adopting an alternating training method to set up M times of deep reinforcement learning network training after completing N times of generative adversarial network training, where N>M and the learning rate of the generative adversarial network is greater than the learning rate of the deep reinforcement learning network.

[0033] Preferably, the closed-loop real-time feedback method includes:

[0034] Using a priority experience replay buffer to screen update samples for updating the policy network of the deep reinforcement learning network, the update samples representing channel mutation events and having high weights;

[0035] Polar codes are used to encode policy parameters related to the policy network, and the policy parameters are transmitted through a dedicated control channel with a bit error rate lower than 1e-6.

[0036] According to the Doppler frequency shift estimation value, the feedback period of the closed-loop real-time feedback is dynamically adjusted to match the time-varying characteristics of the channel.

[0037] Preferably, the method for calculating the feedback cycle includes:

[0038] The channel coherence time and Doppler frequency shift are calculated in real time. When the Doppler frequency shift exceeds a first preset threshold, the feedback cycle is shortened to quickly respond to channel changes. Otherwise, the feedback cycle is extended to reduce system overhead.

[0039] The present invention also provides a channel optimization system applied to a multi-band antenna array, and the channel optimization system applied to a multi-band antenna array is used to implement the above method.

[0040] Beneficial effects

[0041] 1. Improving channel optimization accuracy and authenticity: By introducing a cascaded physical constraint module into the generative adversarial network, physical channel characteristics such as delay spread and Doppler shift are verified during the signal optimization phase. This ensures that the optimized channel state information not only suppresses noise but also better conforms to the laws of the real physical channel, significantly improving signal authenticity and the accuracy of subsequent resource allocation decisions. Combined with a dynamic attention mechanism, it can more accurately extract and process noise characteristics.

[0042] 2. Dynamic and intelligent resource allocation: Leveraging the adaptive decision-making capabilities of deep reinforcement learning, the system dynamically and jointly determines the optimal combination of transmit power, operating frequency band, and modulation scheme based on real-time optimized channel state information, effectively improving spectral and energy efficiency while ensuring service quality under varying channel conditions. This system also enables seamless and rapid switching between millimeter wave and sub-6 GHz frequency bands.

[0043] 3. Enhanced overall system performance and robustness: A cross-network joint training framework couples the intermediate-layer feature information of the GAN discriminator to the DRL policy network, achieving deep integration and coordinated optimization of the channel optimization and resource allocation modules. This improves end-to-end system performance, outperforming separate optimization methods. The system is more robust and adaptable to complex dynamic environments.

[0044] 4. Improving the real-time performance and reliability of closed-loop real-time feedback: An efficient closed-loop real-time feedback mechanism has been designed, using priority experience replay to accelerate learning responses to channel mutation events. Highly robust polar code encoding ensures reliable transmission of policy parameters with low bit error rates, and dynamically adjusts the feedback period based on real-time Doppler shift estimation. This ensures rapid response to channel changes while minimizing system overhead, making it particularly suitable for high-speed mobile scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 The figure is a flow chart of a channel optimization method applied to a multi-band antenna array provided in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described below with reference to the accompanying drawings. In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inner", "outer", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present invention and simplifying the description. They do not indicate or imply that the devices or components referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, they should not be understood as limiting the present invention.

[0047] This invention provides a channel optimization method and system for multi-band antenna arrays, aiming to address the problems of poor signal quality, inefficient resource allocation, and insufficient system adaptability faced by existing wireless communication systems in complex dynamic environments, particularly in high-band transmission, high-speed mobility, and hardware nonlinear interference scenarios. This method significantly improves the overall performance of communication systems by deeply integrating the powerful signal processing capabilities of generative adversarial networks (GANs) with the intelligent decision-making advantages of deep reinforcement learning (DRL), specifically Deep Q-learning (DQN). This method is combined with innovative physical constraint mechanisms, efficient closed-loop real-time feedback, and an end-to-end joint training framework.

[0048] Example 1

[0049] like Figure 1 As shown, the present invention provides a channel optimization method applied to a multi-band antenna array, comprising the following steps:

[0050] S1. Acquire channel state information through the multi-band antenna array in the signal acquisition and noise reduction module.

[0051] This step uses the multi-band antenna array in the signal acquisition and noise reduction module to obtain raw, potentially noisy and distorted channel state information. To adapt to complex wireless environments and obtain rich channel characteristics, the multi-band antenna array preferably employs an advanced design. In a specific embodiment, the multi-band antenna array includes reconfigurable intelligent surface (RIS) antenna units. These RIS units can support broadband operation, such as simultaneous coverage of signal reception in millimeter wave bands such as 28 GHz and sub-6 GHz bands such as 3.5 GHz. Furthermore, the RIS units can support dynamic adjustment of the beamforming angle, for example, within a range of 1° to 15°, thereby flexibly focusing signal energy and improving received signal quality.

[0052] To extract comprehensive channel characteristics from received signals, the antenna array can also integrate a multi-channel signal fusion module. This module processes signals from different antenna elements or different frequency bands, ultimately outputting a multidimensional channel state information (CSI) matrix. This CSI matrix can contain a rich set of channel characteristics, such as delay spread, Doppler shift, and angular domain features (e.g., angle of arrival, angle of departure). For example, the CSI matrix can be organized into [time × frequency × space] dimensions, where the spatial dimension can further include information on horizontal and vertical polarization components, providing a detailed basis for subsequent channel optimization and resource allocation.

[0053] Furthermore, to account for the potential for nonlinear distortion introduced by actual hardware, which can affect CSI accuracy, multi-band antenna arrays can also include a noise floor calibration circuit. This circuit injects a reference signal of known power (e.g., in the range of -40dBm to -20dBm) to evaluate and calibrate the hardware's inherent nonlinear characteristics, effectively suppressing nonlinear distortion to a low level (e.g., below -30dBc) and ensuring the purity of the acquired CSI.

[0054] S2. Using a generative adversarial network and combining it with physical constraints, the channel state information is optimized.

[0055] The purpose of this step is to eliminate noise, correct distortion, and improve the consistency of CSI with the actual physical channel characteristics. In a preferred embodiment, this process is implemented by a cascaded generative adversarial network and a physical constraint module. The generative adversarial network consists of a generator and a discriminator. The generator's task is to learn the mapping from noisy / distorted CSI to pure / ideal CSI, while the discriminator's task is to distinguish between the real pure CSI and the CSI generated by the generator. Specifically, the training method of the generative adversarial network may include the following aspects:

[0056] First, the (possibly noisy) channel state information obtained in step S1 is input into the generator of the generative adversarial network. To more effectively extract and remove noise features, a dynamic attention mechanism is integrated into the generator. This mechanism adaptively focuses on the parts of the CSI data that are more strongly correlated with noise, thereby generating a preliminary denoised CSI signal. Preferably, a possible generator network structure includes multiple convolutional modules in the encoder stage, a dynamic attention module (including spatial attention and channel attention), and multiple deconvolution modules in the decoder stage. The input is a noisy CSI feature matrix of dimension [1024 × 256 × 64], for example, and the output is a preliminary denoised CSI matrix of the same dimension. The discriminator can be a convolutional classification network used to determine the authenticity of the input CSI.

[0057] Secondly, physical constraints are incorporated into the GAN optimization process. Specifically, the preliminary noise reduction signal output by the generator is input into the channel impulse response model. Through this model, the physical characteristics of the channel represented by the generated signal, such as delay spread and Doppler shift, can be analyzed. These characteristics are then compared with the expected physical characteristics obtained from real measurements or theoretical models. For example, it is verified whether the difference between the delay spread distribution of the generated signal and the delay spread distribution of the measured signal exceeds a preset tolerance. If the difference is too large, it indicates that the generated CSI has a large deviation from the physical reality. At this time, the system will trigger the retraining of the generator of the generative adversarial network and, crucially, increase the weight of the physical constraint loss in the overall loss function. This mechanism forces the generator to pay more attention to the physical rationality of the generated signal while learning to reduce noise, avoiding the generation of CSI that appears clean but does not conform to the actual channel propagation laws.

[0058] Physical Constraint Loss Function It can be defined as the mean square error between the generated signal and the real signal in key physical parameters such as delay spread τ and Doppler shift ν:

[0059] ;

[0060] in, Indicates the total number of samples used for verification, that is, the number of channel samples participating in the physical constraint calculation each time; represents the delay spread characteristic corresponding to the 𝑖th generated signal sample, that is, the delay dispersion degree of the received signal caused by the multipath signal propagation path in the channel; represents the actual delay spread characteristic corresponding to the 𝑖th actual measurement signal, which is calculated by collecting data from the actual wireless environment and used as a comparison benchmark for the generated signal; represents the Doppler frequency shift characteristics of the 𝑖th generated signal, that is, the frequency offset of the received signal caused by the dynamic changes of the mobile terminal or the environment, which characterizes the dynamic change characteristics of the channel; It represents the real Doppler frequency shift characteristic corresponding to the 𝑖th actual measurement signal, which is the value measured by the channel data acquisition device in the actual scenario.

[0061] Finally, the entire optimization process is guided by a hybrid loss function. This hybrid loss function includes at least the GAN's own adversarial loss (which encourages the generator to deceive the discriminator and encourages the discriminator to accurately discriminate) and the aforementioned physical constraint loss. In some embodiments, a perceptual loss may also be included to further improve the perceived quality of the generated signal. Importantly, the weighting coefficients of these different loss terms can be dynamically adjusted based on the real-time signal quality. For example, when the signal is noisy, the weights of the adversarial and perceptual losses can be appropriately increased to enhance the noise reduction effect; when physical consistency is poor, the weight of the physical constraint loss can be increased. In some preferred embodiments, the weighting ratio between the adversarial and physical constraint losses can be adjusted between 1:1 and 3:1 based on the real-time signal quality. This dynamically weighted hybrid loss function achieves a better balance between different optimization objectives.

[0062] S3. Input the optimized channel state information into the deep reinforcement learning network, dynamically determine the resource allocation strategy, and output resource configuration parameters; the resource configuration parameters include transmission power, operating frequency band, and modulation mode.

[0063] After being optimized by the GAN and physical constraint modules, the high-quality CSI is fed into a deep reinforcement learning (DRL) network to dynamically determine the optimal communication resource allocation strategy. In some preferred embodiments, the DRL network utilizes the Deep Q-Learning (DQN) algorithm. Its core components include a state space, an action space, and a reward function.

[0064] The state space is primarily composed of the optimized CSI feature matrix. For example, the [1024×256×64]-dimensional CSI matrix optimized by the GAN can be reduced to [128×32×16] through the feature compression module before being input into the DQN network.

[0065] The action space contains various resource configuration options that can be adjusted by the communication system. It can be broken down into:

[0066] Transmit power: Transmit power can be divided into discrete sets ranging from 10dBm to 30dBm, dynamically adjusted in 1dBm steps, with a total of 21 selectable actions. For system stability, the interval between two consecutive power switches can be set to no less than 50ms.

[0067] Operating frequency band: Seamless switching between mmWave (e.g., 28 GHz) and Sub-6 GHz (e.g., 3.5 GHz) bands, with two selectable actions. Switching latency is strictly controlled, for example, to no more than 150 μs, to ensure communication continuity.

[0068] Modulation mode: Dynamically selects one of three modulation modes: QPSK, 16QAM, or 64QAM, based on the target channel signal-to-noise ratio. The goal is to ensure that the bit error rate is below a preset threshold, such as 1e-4.

[0069] Reward function: This function is used to evaluate the quality of the actions selected by the DQN and guide the learning process. The design goal of the reward function is to comprehensively consider multiple aspects of communication performance. In some preferred embodiments, the reward function is designed as follows:

[0070] ;

[0071] in, 、 、 They are the weight coefficients corresponding to throughput, bit error rate and energy consumption, and are initially set to .

[0072] The DQN network architecture can be designed to include several convolutional layers (for extracting features from the CSI matrix) and fully connected layers (for approximating the Q-value function). Specifically, its output layer can be designed as multiple parallel strategy heads, each outputting a Q-value corresponding to the transmit power, operating frequency band, and modulation scheme. When making decisions, the action combination that maximizes the overall Q-value is selected. During training, the classic Bellman equation is used to update Q-values, combined with an ε-greedy strategy to balance exploration and exploitation, and an experience replay mechanism to improve sample utilization efficiency and training stability.

[0073] S4. Feedback the resource configuration parameters in a closed loop to the signal acquisition and noise reduction module at the front end in real time.

[0074] After the DRL network determines the optimal resource configuration parameters, it needs to use a closed-loop real-time feedback mechanism to feed these parameters back to the front-end signal acquisition and noise reduction modules in real time to adjust the system's operating status. The efficiency and reliability of the feedback loop are crucial to the dynamic responsiveness of the entire system. Specifically, closed-loop real-time feedback methods can include the following aspects:

[0075] Prioritized experience replay for policy network updates: During DQN training, a prioritized experience replay buffer can be utilized. This buffer selects samples that indicate sudden channel changes (i.e., samples that result in large time-delay errors) and assigns higher weights to these samples, prioritizing them for updating the deep reinforcement learning network's policy network. Preferably, the update frequency of high-weighted samples can be increased to three times that of normal samples, and the buffer capacity can be set to 1,000-5,000 sets of samples. This enables the policy network to learn and adapt more quickly to dramatic channel changes.

[0076] Robust parameter encoding and transmission: To ensure accurate and reliable transmission of feedback information, DRL-generated policy parameters (such as power, frequency band, and modulation scheme combinations) can be encoded. Polar codes are preferably used for encoding due to their excellent error correction performance. For example, the parameter combination can be encoded into a 256-bit data frame, with CRC check bits accounting for 12.5% ​​to achieve a coding gain of at least 4dB. The encoded parameters are transmitted via a dedicated control channel, aiming for an extremely low bit error rate (BER), for example, below 1e-6.

[0077] Dynamic feedback cycle adjustment based on Doppler shift estimation: To adapt to the time-varying characteristics of the channel, the feedback cycle is not fixed but is dynamically adjusted based on the real-time Doppler shift estimate. The magnitude of the Doppler shift directly reflects the speed of channel variation. Algorithms such as Kalman filtering can be used to estimate the Doppler shift in real time, keeping the estimation error within a small range (e.g., ±2 Hz). The specific adjustment logic may include real-time calculation of the channel coherence time and Doppler shift. When the monitored Doppler shift exceeds a preset first threshold (e.g., 100 Hz, indicating a significant channel variation), the feedback cycle is shortened (e.g., to 10 ms) to achieve rapid response to channel changes. In this case, the parameter update delay can be compressed to, for example, 10-200 ms. Conversely, when the channel stabilizes (e.g., when the Doppler shift fluctuates less than 10 Hz over five consecutive feedback cycles), the feedback cycle can be appropriately extended (e.g., to 200 ms) to reduce unnecessary system overhead and computational load. To avoid frequent oscillations in the feedback cycle, the switching process can be smoothed using a sliding window algorithm. For example, if the sliding window length is 50ms, the jitter amplitude of the switching process is limited to ±5ms.

[0078] It should be noted that a core technical idea of ​​the present invention is that the generative adversarial network and the deep reinforcement learning network do not work independently and then simply cascade, but adopt a joint training method to perform deep coupling and collaborative optimization. This joint training aims to achieve end-to-end overall performance improvement. The key feature of joint training is that the intermediate layer feature information of the generative adversarial network (specifically its discriminator) is input into the policy network of the deep reinforcement learning network as an additional state information. This means that when the DRL network makes a decision, it not only considers the optimized CSI itself, but also refers to the deeper features extracted by the GAN discriminator when evaluating the quality and authenticity of the CSI. These features from the discriminator can reveal some potential information that is not easy to see directly from the CSI, such as subtle differences in the statistical characteristics of the generated signal and the real signal, thereby providing DRL with richer decision-making basis. In some preferred embodiments, the joint training method can further include:

[0079] Defining a collaborative objective function and loss allocation: To collaboratively optimize GAN and DRL, a unified or related objective is required. A joint loss function can be designed, or separate loss functions can be defined but linked via weights. The loss terms include at least: a denoising loss associated with the optimized CSI output by the GAN (e.g., the sum of the generation loss and adversarial loss of the GAN), and a physical constraint loss associated with satisfying physical constraints; and a resource allocation loss associated with the resource configuration parameters output by the DRL network (e.g., the Q-value prediction error loss of the DQN). An example of a collaborative objective function is given in some preferred embodiments: ;

[0080] in, is the overall denoising loss function for the generative adversarial network; is the Q-value prediction error loss of the reinforcement learning policy network; is a physical constraint loss function used to constrain the generated signal to meet physical consistency requirements. The weight ratio of the three loss terms is fixed at 6:3:1 to ensure the effectiveness and stability of cross-network collaborative optimization.

[0081] Parameter Update Strategy: Based on the aforementioned loss terms, the parameters of the GAN and DRL networks are updated separately. Specifically, the parameters of the generative adversarial network are updated based on the denoising loss and the physical constraint loss; while the parameters of the policy network of the deep reinforcement learning network are updated based on the resource allocation loss and incorporating the intermediate layer feature information coupled from the GAN discriminator as additional state.

[0082] Alternating Training Mechanism: To ensure the stability and efficiency of joint training, an alternating training approach can be employed. Specifically, after every N iterations of GAN training (e.g., generator parameter updates), the DRL network is trained M times (policy network parameter updates). Typically, N > M. For example, after every three GAN generator network parameter updates, the DRL policy network parameter update is performed once. Furthermore, to balance the learning rates of the different networks, the GAN learning rate can be set higher than that of the DRL network (e.g., the DRL policy network learning rate can be set to 1 / 10 of the GAN generator network learning rate). This alternating and differentiated learning rate configuration helps the two networks adapt to each other and evolve together.

[0083] Example 2

[0084] The present invention also provides a channel optimization system for a multi-band antenna array, which is configured to implement the method described in the first embodiment. The system may specifically include:

[0085] A signal acquisition and noise reduction module, including a multi-band antenna array (eg, including a multi-channel signal fusion module and a noise floor calibration circuit), is used to perform the function of step S1, namely, to obtain the original channel state information.

[0086] A channel optimization module is configured with a generative adversarial network and a physical constraint processing unit, and is used to perform the function of step S2, that is, to optimize the channel state information.

[0087] A resource allocation module, whose core is a deep reinforcement learning network, is used to perform the function of step S3, that is, dynamically determine resource configuration parameters based on the optimized channel state information.

[0088] A closed-loop real-time feedback module is responsible for executing the functions of step S4, including encoding the policy parameters and transmitting them through the control channel, and dynamically adjusting the feedback cycle.

[0089] A joint training control module is used to coordinate and manage the joint training process of the generative adversarial network and the deep reinforcement learning network, including feature sharing, loss calculation and distribution, and the execution of alternating training strategies.

[0090] These modules can be a combination of hardware, software or firmware, and interact with data and control signaling through internal interfaces, together forming an efficient and intelligent joint channel optimization communication system.

[0091] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A channel optimization method applied to a multi-band antenna array, characterized in that: The following steps are involved: S1. Acquire channel state information through the multi-band antenna array in the signal acquisition and noise reduction module; S2. Using a generative adversarial network, combined with physical constraints, to optimize the channel state information; S3. Input the optimized channel state information into the deep reinforcement learning network, dynamically determine the resource allocation strategy, and output resource configuration parameters; the resource configuration parameters include transmit power, operating frequency band, and modulation mode; S4. The resource configuration parameters are fed back to the front-end signal acquisition and noise reduction module in a closed loop in real time; The generative adversarial network and the deep reinforcement learning network are jointly trained, and the intermediate layer feature information of the generative adversarial network discriminator is used as additional state information and input into the policy network of the deep reinforcement learning network; The training method of the generative adversarial network includes: Inputting the channel state information into the generative adversarial network, extracting noise features through a dynamic attention mechanism, and generating a preliminary noise reduction signal; Inputting the preliminary noise reduction signal into the channel impulse response model to verify whether the difference in the delay spread distribution between the delay spread and the Doppler frequency shift exceeds a preset tolerance. If so, triggering retraining of the generative adversarial network and increasing the physical constraint loss weight; Dynamically adjust the weight coefficients of adversarial loss, perception loss, and physical constraint loss through a hybrid loss function based on real-time signal quality; The joint training method includes: Calculating loss terms and assigning weights to each loss term, the loss terms including at least: a denoising loss associated with the optimized channel state information output by the generative adversarial network, and a physical constraint loss associated with satisfying physical constraints; and a resource allocation loss associated with resource configuration parameters output by the deep reinforcement learning network; Updating parameters of the generative adversarial network based on the denoising loss and the physical constraint loss; Based on the resource allocation loss, parameters of the policy network of the deep reinforcement learning network are updated, wherein the update of the policy network also utilizes intermediate layer feature information coupled from the discriminator of the generative adversarial network as additional state information.

2. The channel optimization method for a multi-band antenna array according to claim 1, wherein: The multi-band antenna array comprises: Multi-channel signal fusion module, used to output a multi-dimensional channel state information matrix containing delay spread, Doppler frequency shift and angle domain characteristics; Noise floor calibration circuit to eliminate hardware nonlinear distortion using a reference signal.

3. The channel optimization method for a multi-band antenna array according to claim 1, wherein: The resource allocation strategy includes: Divide the transmit power into discrete action sets ranging from 10dBm to 30dBm, dynamically adjust the power allocation in steps of 1dBm, and keep the interval between adjacent power switching at no less than 50ms; Seamless switching between mmWave and Sub-6GHz dual-bands with a switching latency of no more than 150μs; Dynamically select one of the QPSK, 16QAM, or 64QAM modulation modes based on the target channel signal-to-noise ratio to keep the bit error rate below 1e-4.

4. The channel optimization method for a multi-band antenna array according to claim 1, wherein: The joint training method also includes: using an alternating training method to set up M times of deep reinforcement learning network training after completing N times of generative adversarial network training, where N>M and the learning rate of the generative adversarial network is greater than the learning rate of the deep reinforcement learning network.

5. The channel optimization method for a multi-band antenna array according to claim 1, wherein: The closed-loop real-time feedback method includes: Using a priority experience replay buffer to screen update samples for updating the policy network of the deep reinforcement learning network, the update samples representing channel mutation events and having high weights; Polar codes are used to encode policy parameters related to the policy network, and the policy parameters are transmitted through a dedicated control channel with a bit error rate lower than 1e-6. According to the Doppler frequency shift estimation value, the feedback period of the closed-loop real-time feedback is dynamically adjusted to match the time-varying characteristics of the channel.

6. The channel optimization method for a multi-band antenna array according to claim 5, wherein: The calculation method of the feedback cycle includes: The channel coherence time and Doppler frequency shift are calculated in real time. When the Doppler frequency shift exceeds a first preset threshold, the feedback cycle is shortened to quickly respond to channel changes. Otherwise, the feedback cycle is extended to reduce system overhead.

7. A channel optimization system for a multi-band antenna array, characterized in that: The channel optimization system applied to a multi-band antenna array is used to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • A noise robust face recognition method based on a cascade deep convolutional neural network

    CN109948573A

  • NOMA multi-beam satellite communication method based on deep reinforcement learning

    CN118075877A