Channel optimization method and system applied to multi-band antenna array

By combining the joint training of the generative adversarial network and deep reinforcement learning network, combined with physical constraints, optimize channel state information and dynamically determine resource configuration parameters, the problem of channel processing and resource allocation collaborative optimization of wireless communication systems in complex dynamic environments is solved, signal quality and spectrum efficiency are improved, and spectrum efficiency is adapted to high-speed mobile scenarios.

CN120357931AActive Publication Date: 2025-07-22SICHUAN JIUZHOU SOFTWARE CO LTD

Patent Information

Application Number
CN202510851738.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the complex and changeable dynamic environment, especially under high frequency band, high mobility and non-Gaussian noise interference, existing wireless communication systems are difficult to achieve coordinated optimization of channel processing and resource allocation, resulting in poor signal quality, low resource allocation efficiency, insufficient system adaptability, and the feedback mechanism has problems such as time delay and high bit error rate.

Method used

The joint training method of generative adversarial network and deep reinforcement learning network is adopted, combined with physical constraints, optimize channel state information, dynamically determine resource configuration parameters, and end-to-end collaborative optimization of channel optimization and resource allocation through an efficient closed-loop feedback mechanism.

Benefits of technology

It significantly improves channel optimization accuracy and resource allocation accuracy, improves spectrum efficiency and system robustness, adapts to complex dynamic environments, reduces the risk of transmission interruption, and is especially suitable for high-speed mobile scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357931A_ABST
    Figure CN120357931A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of communication, and discloses a channel optimization method and system applied to a multi-band antenna array, and the method comprises the steps: obtaining channel state information; a generative adversarial network is adopted, physical constraints are combined, and channel state information is optimized; inputting the optimized channel state information into a deep reinforcement learning network, dynamically determining a resource allocation strategy, and outputting resource configuration parameters; the resource configuration parameters are fed back to the front-end signal acquisition and noise reduction module in a closed-loop and real-time manner; wherein the generative adversarial network and the deep reinforcement learning network adopt a joint training mode, and intermediate layer feature information of a discriminator of the generative adversarial network serves as additional state information and is input into a strategy network of the deep reinforcement learning network. According to the method, the reliability and the spectrum efficiency of a wireless communication system are improved, the transmission interruption risk in a high-mobility scene is reduced, and a more stable and more efficient communication guarantee is provided for a millimeter wave and Sub-6GHz hybrid networking scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and particularly to a channel optimization method and system applied to a multi-band antenna array. Background Art

[0002] With the commercial deployment of 5G technology and the evolution of 6G technology, wireless communication systems are facing unprecedented challenges. On the one hand, the spectral resources are extended to higher frequency bands (such as millimeter waves), which brings larger bandwidth but also is accompanied by more serious path loss, signal blocking, and sensitivity to hardware non-linear distortion. On the other hand, the booming development of applications such as the Internet of Things and the Internet of Vehicles requires that the communication system can support massive connections, ultra-low latency, and high reliability, especially in dynamic scenarios with high-speed movement and rapidly changing channel environments.

[0003] Traditional optimization methods for wireless communication systems usually adopt a separated design: channel estimation and signal processing (such as filtering and noise reduction) and resource allocation (such as power control, frequency band selection, modulation and coding) are optimized as independent modules. For example, traditional methods such as Kalman filtering are used for channel tracking and noise reduction, and algorithms based on fixed rules or offline optimization are used for resource allocation. These methods have the following limitations: 1. Poor adaptability: It is difficult to effectively cope with complex and changing wireless channel environments, especially in mixed frequency bands, high mobility, and non-Gaussian noise interference, where the performance drops sharply.

[0004] 2. Insufficient physical constraints: Although deep learning noise reduction methods based on data driving (such as basic GAN) are powerful, they may generate signals that do not conform to the actual physical channel characteristics (such as delay spread, Doppler shift), affecting the accuracy of subsequent processing.

[0005] 3. Lagging resource allocation: Static or slow-adaptive resource allocation strategies cannot match the rapidly changing channel conditions in real time, resulting in low spectral efficiency, high bit error rate, or connection interruption.

[0006] 4. Limitations of the feedback mechanism: Traditional closed-loop real-time feedback mechanisms may have problems such as long time delay, high transmission error rate, or fixed feedback period, and cannot meet the requirements of fast and reliable strategy updates in dynamic environments.

[0007] 5. Lack of collaborative optimization: The separated optimization of channel processing and resource allocation ignores the internal connection between the two, and it is difficult to achieve system-level end-to-end performance optimization.

[0008] Although existing research has attempted to apply deep learning to wireless communication optimization, such as using GAN alone for channel generation or noise reduction, or using reinforcement learning for resource allocation, it often fails to fully address all of the above problems, especially in terms of physical authenticity constraints, multi-module collaborative optimization, and real-time robust feedback. Therefore, there is an urgent need for an intelligent communication method that can jointly optimize channel processing and resource allocation and can adapt to complex dynamic environments in real time. Summary of the Invention

[0009] To solve the problems existing in the above-mentioned prior art, the technical solutions provided by the present invention include: A channel optimization method applied to a multi-band antenna array, comprising the following steps: S1. Obtain channel state information through a multi-band antenna array in a signal acquisition and noise reduction module; S2. Use a generative adversarial network, combined with physical constraints, to optimize the channel state information; S3. Input the optimized channel state information into a deep reinforcement learning network to dynamically determine a resource allocation strategy and output resource configuration parameters; the resource configuration parameters include transmit power, operating frequency band, and modulation method; S4. Closed-loop and real-time feedback the resource configuration parameters to the front-end signal acquisition and noise reduction module; Among them, the generative adversarial network and the deep reinforcement learning network are jointly trained, and the intermediate layer feature information of the generative adversarial network discriminator is used as additional state information and input into the policy network of the deep reinforcement learning network.

[0010] Preferably, the multi-band antenna array includes: A multi-channel signal fusion module for outputting a multi-dimensional channel state information matrix including delay spread, Doppler frequency shift, and angle domain characteristics; A noise floor calibration circuit for eliminating hardware nonlinear distortion through a reference signal.

[0011] Preferably, the training method of the generative adversarial network includes: Input the channel state information into the generative adversarial network, extract noise features through a dynamic attention mechanism, and generate a preliminary noise reduction signal; Input the preliminary noise reduction signal into a channel impulse response model to verify whether the difference in the delay spread distribution of the delay spread and the Doppler frequency shift exceeds a preset tolerance. If so, trigger retraining of the generative adversarial network and increase the weight of the physical constraint loss; Dynamically adjust the weight coefficients of the adversarial loss, perceptual loss, and physical constraint loss based on the real-time signal quality through a mixed loss function.

[0012] Preferably, the resource allocation strategy includes: Dividing the transmit power into a discrete action set from 10 dBm to 30 dBm, dynamically adjusting the power allocation with a step size of 1 dBm, and the adjacent power switching interval is not less than 50 ms; Performing seamless handover between the millimeter wave and Sub-6 GHz dual bands, and the handover delay is not higher than 150 μs; Dynamically selecting one of QPSK, 16QAM or 64QAM modulation modes according to the target channel signal-to-noise ratio to make the bit error rate lower than 1e-4.

[0013] Preferably, the method of joint training includes: Calculating the loss terms and assigning weights to each loss term. The loss terms at least include: the denoising loss related to the optimized channel state information output by the generative adversarial network, and the physical constraint loss related to satisfying the physical constraints; the resource allocation loss related to the resource configuration parameters output by the deep reinforcement learning network; Updating the parameters of the generative adversarial network based on the denoising loss and the physical constraint loss; Updating the parameters of the policy network of the deep reinforcement learning network based on the resource allocation loss, where the update of the policy network also uses the intermediate layer feature information coupled from the discriminator of the generative adversarial network as additional state information.

[0014] Preferably, the method of joint training further includes: setting to perform M times of training of the deep reinforcement learning network after every N times of training of the generative adversarial network in an alternating training manner, where N > M and the learning rate of the generative adversarial network is greater than the learning rate of the deep reinforcement learning network.

[0015] Preferably, the method of closed-loop real-time feedback includes: Using a priority experience replay buffer to screen the update samples for updating the policy network of the deep reinforcement learning network. The update samples characterize channel mutation events and have high weights; Encoding the policy parameters related to the policy network using polar codes and transmitting them through a dedicated control channel with a transmission bit error rate lower than 1e-6; Dynamically adjusting the feedback period of the closed-loop real-time feedback according to the Doppler frequency shift estimate to match the time-varying characteristics of the channel.

[0016] Preferably, the calculation method of the feedback period includes: Calculating the channel coherence time and the Doppler frequency shift amount in real time. When the Doppler frequency shift amount exceeds the first preset threshold, shortening the feedback period to quickly respond to channel changes, otherwise lengthening the feedback period to reduce system overhead.

[0017] The present invention also provides a channel optimization system applied to a multi-band antenna array, and the channel optimization system applied to the multi-band antenna array is used to implement the above method.

[0018] Beneficial effects 1. Improve the accuracy and authenticity of channel optimization: By introducing a cascaded physical constraint module into the generative adversarial network, physical channel characteristics such as delay spread and Doppler frequency shift are verified during the signal optimization stage, ensuring that the optimized channel state information not only suppresses noise but also conforms more to the real physical channel law, significantly improving the authenticity of the signal and the accuracy of subsequent resource allocation decisions. Combined with the dynamic attention mechanism, it can more accurately extract and process noise features.

[0019] 2. Achieve dynamic and intelligent resource allocation: Utilize the adaptive decision-making ability of deep reinforcement learning to dynamically and jointly determine the optimal combination of transmit power, operating frequency band, and modulation method according to the real-time optimized channel state information, effectively improving the spectral efficiency and energy efficiency, and ensuring the quality of service under different channel conditions. Achieved seamless and fast switching between the millimeter-wave and Sub-6GHz frequency bands.

[0020] 3. Enhance the overall system performance and robustness: Adopt a cross-network joint training framework to couple the intermediate layer feature information of the GAN discriminator to the DRL policy network, realizing the deep integration and collaborative optimization of the two modules of channel optimization and resource allocation, achieving end-to-end system performance improvement, which is better than the method of separate optimization. The system has stronger robustness and adaptability to complex dynamic environments.

[0021] 4. Improve the real-time performance and reliability of closed-loop real-time feedback: Design an efficient closed-loop real-time feedback mechanism, adopt priority experience replay to accelerate the learning response to channel mutation events, use highly robust polar code encoding to ensure reliable transmission of policy parameters at low bit error rates, and dynamically adjust the feedback period according to real-time Doppler frequency shift estimation, taking into account both the fast response to channel changes and the system overhead, which is especially suitable for high-speed mobile scenarios. Description of the drawings

[0022] Figure 1 It is a schematic flow chart of a channel optimization method applied to a multi-band antenna array provided in a preferred embodiment of the present invention. Specific implementation manners

[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings. In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.

[0024] The present invention provides a channel optimization method and system for a multi-band antenna array, aiming to solve the problems of poor signal quality, low resource allocation efficiency, and insufficient system adaptability faced by existing wireless communication systems in complex dynamic environments, especially in high-frequency band transmission, high-speed mobility, and hardware non-linear interference scenarios. The present invention deeply integrates the powerful signal processing capabilities of the Generative Adversarial Network (GAN) and the intelligent decision-making advantages of Deep Reinforcement Learning (DRL, specifically Deep Q-Network, DQN can be adopted), and combines an innovative physical constraint mechanism, efficient closed-loop real-time feedback, and an end-to-end joint training framework, thereby significantly improving the overall performance of the communication system.

[0025] Embodiment 1 As Figure 1 shown, the present invention provides a channel optimization method for a multi-band antenna array, including the following steps: S1. Obtain channel state information through the multi-band antenna array in the signal acquisition and noise reduction module.

[0026] This step obtains the original channel state information that may contain noise and distortion through the multi-band antenna array in the signal acquisition and noise reduction module. To adapt to the complex wireless environment and obtain rich channel characteristics, the multi-band antenna array preferably adopts an advanced design. In a specific embodiment, the multi-band antenna array includes Reconfigurable Intelligent Surface (RIS) antenna units. These RIS units can support broadband operation, such as receiving signals in millimeter wave bands such as 28 GHz and Sub-6 GHz bands such as 3.5 GHz simultaneously. Further, the RIS units can support dynamic adjustment of the beamforming angle, for example, adjusting within the range of 1° to 15°, so as to be able to flexibly focus the signal energy and improve the received signal quality.

[0027] To extract comprehensive channel characteristics from the received signals, the antenna array can also integrate a multi-channel signal fusion module. This module is responsible for processing signals from different antenna elements or different frequency bands and finally outputs a multi-dimensional channel state information (CSI) matrix. This CSI matrix can contain rich channel characteristics such as delay spread, Doppler shift, and angular domain characteristics (such as angle of arrival, angle of departure). For example, this CSI matrix can be organized in the dimensions of [time × frequency × space], where the space dimension can further include horizontal and vertical polarization component information, providing a detailed basis for subsequent channel optimization and resource allocation.

[0028] In addition, considering that the actual hardware may introduce nonlinear distortion, affecting the accuracy of CSI, the multi-band antenna array can also include a noise floor calibration circuit. This circuit evaluates and calibrates the nonlinear characteristics of the hardware itself by injecting a reference signal with a known power (such as a reference signal in the range of -40 dBm to -20 dBm), thereby effectively suppressing the nonlinear distortion to a lower level (such as below -30 dBc) to ensure the purity of the acquired CSI.

[0029] S2. Use a generative adversarial network, combined with physical constraints, to optimize the channel state information.

[0030] The purpose of this step is to eliminate noise, correct distortion, and improve the consistency between the CSI and the true physical channel characteristics. In a preferred implementation, this process is achieved through a cascaded generative adversarial network and physical constraint module. The generative adversarial network consists of a generator and a discriminator. The task of the generator is to learn the mapping from the noisy / distorted CSI to the clean / ideal CSI, while the task of the discriminator is to distinguish the true clean CSI from the CSI generated by the generator. Specifically, the training method of the generative adversarial network can include the following aspects: First, the (possibly noisy) channel state information obtained in step S1 is input into the generator of the generative adversarial network. To more effectively extract and remove noise features, a dynamic attention mechanism is integrated inside the generator. This mechanism can adaptively focus on the parts of the CSI data that are more strongly correlated with the noise, thereby generating a preliminarily denoised CSI signal. Preferably, a possible generator network structure includes multiple convolutional modules in the encoder stage, a dynamic attention module (including spatial attention and channel attention), and multiple deconvolutional modules in the decoder stage. The input is a noisy CSI feature matrix with dimensions such as [1024 × 256 × 64], and the output is a preliminarily denoised CSI matrix with the same dimensions. The discriminator can be a convolutional classification network for judging the authenticity of the input CSI.

[0031] Secondly, physical constraints are incorporated into the optimization process of the GAN. Specifically, the preliminary noise-reduced signal output by the generator is input into the channel impulse response model. Through this model, the physical characteristics of the channel represented by the generated signal can be analyzed, such as delay spread and Doppler shift. Then, these characteristics are compared with the expected physical characteristics obtained from real measurements or theoretical models. For example, verify whether the difference between the delay spread distribution of the generated signal and that of the measured signal exceeds a preset tolerance. If the difference is too large, it indicates that there is a large deviation in the physical authenticity of the generated CSI. At this time, the system will trigger the retraining of the generator of the generative adversarial network, and crucially, increase the weight of the physical constraint loss in the overall loss function. This mechanism forces the generator to pay more attention to the physical rationality of the generated signal while learning to reduce noise, avoiding the generation of CSI that seems clean but does not conform to the actual channel propagation law.

[0032] Physical constraint loss function can be defined as the mean square error of the generated signal and the real signal in key physical parameters such as delay spread τ and Doppler shift ν: ; where represents the total number of samples for verification, that is, the number of channel samples participating in the physical constraint calculation each time; represents the delay spread feature corresponding to the 𝑖-th generated signal sample, that is, the degree of delay dispersion of the received signal caused by the propagation paths of multipath signals in the channel; represents the real delay spread feature corresponding to the 𝑖-th actual measurement signal, which is calculated by collecting data in the actual wireless environment and is used as a comparison benchmark for the generated signal; represents the Doppler shift feature of the 𝑖-th generated signal, that is, the frequency offset of the received signal due to the dynamic changes of the mobile terminal or the environment, representing the dynamic change characteristics of the channel; represents the real Doppler shift feature corresponding to the 𝑖-th actual measurement signal, which is the value measured by the channel data acquisition device in the actual scenario.

[0033] Finally, the entire optimization process is guided by a hybrid loss function. The hybrid loss function at least includes the adversarial loss of the GAN itself (encouraging the generator to deceive the discriminator and the discriminator to accurately discriminate), and the aforementioned physical constraint loss. In some embodiments, a perceptual loss may also be included to further improve the perceptual quality of the generated signal. Importantly, the weight coefficients of these different loss terms can be dynamically adjusted according to the real-time signal quality. For example, when the signal noise is large, the weights of the adversarial loss and the perceptual loss can be appropriately increased to strengthen the noise reduction effect; when the physical consistency is poor, the weight of the physical constraint loss is increased. In some preferred embodiments, the weight adjustment range of the adversarial loss and the physical constraint loss can be set between 1:1 and 3:1 according to the real-time signal quality. Through this dynamically weighted hybrid loss function, a better balance can be achieved among different optimization objectives.

[0034] S3. Input the optimized channel state information into a deep reinforcement learning network to dynamically determine the resource allocation strategy and output resource configuration parameters; the resource configuration parameters include transmit power, operating frequency band, and modulation method.

[0035] The high-quality CSI optimized by the GAN and the physical constraint module will be input into a deep reinforcement learning (DRL) network to dynamically determine the optimal communication resource allocation strategy. In some preferred embodiments, the DRL network adopts the deep Q-learning (DQN) algorithm. Its core components include the state space, action space, and reward function.

[0036] Among them, the state space is mainly composed of the optimized CSI feature matrix. For example, the [1024×256×64] - dimensional CSI matrix optimized by the GAN can be first reduced in dimension to [128×32×16] through a feature compression module and then input into the DQN network.

[0037] The action space contains various resource configuration options that can be adjusted by the communication system. It can be specifically refined as follows: Transmit power: The transmit power can be divided into a discrete set from 10 dBm to 30 dBm and dynamically adjusted in steps of 1 dBm, with a total of 21 optional actions. For system stability, the interval between two adjacent power switches can be set to not less than 50 ms.

[0038] Operating frequency band: Perform seamless switching between the millimeter wave (such as 28 GHz) and Sub - 6 GHz (such as 3.5 GHz) frequency bands, with a total of 2 optional actions. The switching delay is strictly controlled, for example, not higher than 150 μs, to ensure communication continuity.

[0039] Modulation method: Dynamically select one of QPSK, 16QAM, or 64QAM according to the target channel signal-to-noise ratio. There are a total of 3 optional actions. The goal of the selection is to ensure that the bit error rate is lower than a preset threshold, such as 1e-4.

[0040] Reward function: Used to evaluate the quality of the actions selected by the DQN and guide the learning process. The design goal of the reward function is to comprehensively consider multiple aspects of communication performance. In some preferred embodiments, the reward function is designed as follows: ; Among them, , , are the weight coefficients corresponding to throughput, bit error rate, and energy consumption respectively, and are initially set to .

[0041] The structure of the DQN network can be designed to include several convolutional layers (for extracting features from the CSI matrix) and fully connected layers (for approximating the Q-value function). In particular, its output layer can be designed as multiple parallel policy heads, corresponding to the Q-value outputs of transmit power, operating frequency band, and modulation method respectively. During decision-making, select the action combination that maximizes the total Q value. During the training process, the classic Bellman equation is used to update the Q value, and the ε-greedy strategy is combined to balance exploration and exploitation, as well as the experience replay mechanism to improve sample utilization efficiency and training stability.

[0042] S4. Closed-loop and real-time feedback the resource configuration parameters to the front-end signal acquisition and noise reduction module.

[0043] After the DRL network determines the optimal resource configuration parameters, it is necessary to feedback these parameters back to the front-end signal acquisition and noise reduction module in real time through a closed-loop real-time feedback mechanism to timely adjust the working state of the system. The efficiency and reliability of the feedback loop are crucial for the dynamic response ability of the entire system. Specifically, the methods of closed-loop real-time feedback can include the following aspects: Priority experience replay for policy network update: During the training process of DQN, a priority experience replay buffer can be used. This buffer can screen out those samples that characterize channel mutation events (i.e., samples that cause large TD-errors), and assign higher weights to these samples, making them preferentially used to update the policy network of the deep reinforcement learning network. Preferably, the update frequency of high-weight samples can be increased to 3 times that of ordinary samples, and the buffer capacity can be set to 1000 - 5000 groups of samples. This enables the policy network to learn and adapt to drastic changes in the channel faster.

[0044] Robust Parameter Encoding and Transmission: To ensure the accurate and reliable transmission of feedback information, the policy parameters generated by DRL (such as the combination of power, frequency band, and modulation method) can be encoded. Preferably, polar codes are used for encoding because of their excellent error correction performance. For example, the parameter combination can be encoded into a data frame with a length of 256 bits, where the CRC check bits account for 12.5%, to achieve an encoding gain of no less than 4 dB. The encoded parameters are transmitted through a dedicated control channel, and an extremely low transmission error rate is sought, such as less than 1e-6.

[0045] Dynamic Feedback Period Adjustment Based on Doppler Shift Estimation: To match the time-varying characteristics of the channel, the feedback period is not fixed but dynamically adjusted according to the real-time Doppler shift estimation value. The magnitude of the Doppler shift directly reflects the speed of channel change. Algorithms such as Kalman filtering can be used to estimate the Doppler shift in real time, and the estimation error is controlled within a small range (such as ±2 Hz). The specific adjustment logic can be: calculate the channel coherence time and the Doppler shift amount in real time. When the monitored Doppler shift amount exceeds a preset first threshold (such as 100 Hz, indicating that the channel changes violently), the feedback period is shortened (such as to 10 ms) to achieve a fast response to channel changes, and the parameter update delay can be compressed to, for example, 10 - 200 ms. On the contrary, when the channel tends to be stable (such as the Doppler shift amount fluctuates less than 10 Hz within 5 consecutive feedback periods), the feedback period can be appropriately extended (such as to 200 ms) to reduce unnecessary system overhead and computational load. To avoid frequent oscillations of the feedback period, the switching process can be smoothed through a sliding window algorithm, such as a sliding window length of 50 ms, and the jitter amplitude of the switching process is limited within the range of ±5 ms.

[0046] It should be noted that a core technical idea of the present invention is that the generative adversarial network and the deep reinforcement learning network do not work independently and then are simply cascaded, but are deeply coupled and co-optimized in a joint training manner. This joint training aims to achieve an overall end-to-end performance improvement. The key feature of the joint training is that the intermediate layer feature information of the generative adversarial network (specifically its discriminator) is used as an additional state information and input into the policy network of the deep reinforcement learning network. This means that when making decisions, the DRL network not only considers the optimized CSI itself but also refers to the deeper features extracted by the GAN discriminator when evaluating the quality and authenticity of the CSI. These features from the discriminator can reveal some potential information that is not easily seen directly from the CSI, such as the subtle differences in statistical characteristics between the generated signal and the real signal, thus providing richer decision-making basis for DRL. In some preferred embodiments, the method of joint training can further include: Define the collaborative objective function and loss allocation: To collaboratively optimize GAN and DRL, a unified or related objective is needed. A joint loss function can be designed, or their respective loss functions can be defined separately but correlated through weights. The loss terms should at least include: the denoising loss related to the optimized CSI output by GAN (such as the sum of the generation loss and adversarial loss of GAN), and the physical constraint loss related to satisfying physical constraints; it also includes the resource allocation loss related to the resource configuration parameters output by the DRL network (such as the Q-value prediction error loss of DQN). An example of a collaborative objective function is given in some preferred embodiments: ; where is the overall denoising loss function of the generative adversarial network; is the Q-value prediction error loss of the reinforcement learning policy network; is the physical constraint loss function used to constrain the generated signal to meet the physical consistency requirements. The weight ratio of the three loss terms is fixed at 6:3:1 to ensure the effectiveness and stability of cross-network collaborative optimization.

[0047] Parameter update strategy: Based on the above loss terms, the parameters of the GAN and DRL networks are updated separately. Specifically, the parameters of the generative adversarial network are updated based on the denoising loss and physical constraint loss; while the parameters of the policy network of the deep reinforcement learning network are updated based on the resource allocation loss and combined with the intermediate layer feature information coupled from the GAN discriminator as an additional state.

[0048] Alternating training mechanism: To ensure the stability and efficiency of joint training, an alternating training method can be adopted. That is, after every N times of training of the generative adversarial network (for example, updating the generator parameters), M times of training of the deep reinforcement learning network (updating the policy network parameters) are executed. Usually, N > M. For example, after every 3 times of updating the parameters of the GAN generator network, 1 time of updating the parameters of the reinforcement learning policy network is executed. In addition, to balance the learning rates of different networks, the learning rate of the generative adversarial network can be set to be greater than that of the deep reinforcement learning network (for example, the learning rate of the DRL policy network is set to 1 / 10 of the learning rate of the GAN generator network). This alternating and differential learning rate setting helps the two networks adapt to each other and co-evolve.

[0049] Embodiment 2 The present invention also provides a channel optimization system applied to a multi-band antenna array, which is configured to implement the method described in the foregoing Embodiment 1. Specifically, the system may include: A signal acquisition and denoising module, which contains a multi-band antenna array (such as including a multi-channel signal fusion module and a noise floor calibration circuit), and is used to perform the function of step S1, that is, to obtain the original channel state information.

[0050] A channel optimization module, configured with a generative adversarial network and a physical constraint processing unit, for performing the functions of step S2, that is, optimizing the channel state information.

[0051] A resource allocation module, with a deep reinforcement learning network at its core, for performing the functions of step S3, that is, dynamically determining resource configuration parameters according to the optimized channel state information.

[0052] A closed-loop real-time feedback module, responsible for performing the functions of step S4, including encoding policy parameters and transmitting them through the control channel, and dynamically adjusting the feedback period.

[0053] A joint training control module, for coordinating and managing the joint training process of the generative adversarial network and the deep reinforcement learning network, including feature sharing, loss calculation and allocation, and execution of the alternating training strategy.

[0054] These modules can be a combination of hardware, software or firmware, and interact with each other through internal interfaces for data and control signaling, jointly constituting an efficient and intelligent joint channel optimization communication system.

[0055] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A channel optimization method applied to a multi-band antenna array, characterized in that, It includes the following steps: S1. Obtain channel state information through the multi-band antenna array in the signal acquisition and noise reduction module; S2. Use a generative adversarial network, combined with physical constraints, to optimize the channel state information; S3. Input the optimized channel state information into a deep reinforcement learning network to dynamically determine a resource allocation strategy and output resource configuration parameters; the resource configuration parameters include transmit power, operating frequency band, and modulation method; S4. Closed-loop and real-time feedback the resource configuration parameters to the front-end signal acquisition and noise reduction module; Among them, the generative adversarial network and the deep reinforcement learning network are jointly trained, and the intermediate layer feature information of the discriminator of the generative adversarial network is used as additional state information and input into the policy network of the deep reinforcement learning network.

2. The channel optimization method applied to a multi-band antenna array according to claim 1, wherein The multi-band antenna array includes: A multi-channel signal fusion module for outputting a multi-dimensional channel state information matrix containing delay spread, Doppler frequency shift, and angle domain characteristics; A noise floor calibration circuit for eliminating hardware nonlinear distortion through a reference signal.

3. The channel optimization method applied to a multi-band antenna array according to claim 1, characterized in that The training method of the generative adversarial network includes: Input the channel state information into the generative adversarial network, extract noise features through a dynamic attention mechanism, and generate a preliminary denoised signal; Input the preliminary denoised signal into a channel impulse response model to verify whether the difference in the delay spread distribution of the delay spread and Doppler frequency shift exceeds a preset tolerance. If so, trigger retraining of the generative adversarial network and increase the weight of the physical constraint loss; Dynamically adjust the weight coefficients of the adversarial loss, perceptual loss, and physical constraint loss based on the real-time signal quality through a hybrid loss function.

4. The channel optimization method applied to a multi-band antenna array according to claim 1, characterized in that The resource allocation strategy includes: Divide the transmit power into a discrete action set of 10 dBm to 30 dBm, dynamically adjust the power allocation in steps of 1 dBm, and the adjacent power switching interval is not less than 50 ms; Perform seamless switching between the millimeter wave and Sub-6 GHz dual frequency bands, and the switching delay is not higher than 150 μs; Dynamically select one of the QPSK, 16QAM, or 64QAM modulation methods according to the target channel signal-to-noise ratio to make the bit error rate lower than 1e-4.

5. The channel optimization method applied to a multi-band antenna array according to claim 1, characterized in that The method of joint training includes: Calculate loss terms and allocate weights to each loss term. The loss terms at least include: a denoising loss related to the optimized channel state information output by the generative adversarial network, and a physical constraint loss related to meeting physical constraints; a resource allocation loss related to the resource configuration parameters output by the deep reinforcement learning network; Update the parameters of the generative adversarial network based on the denoising loss and the physical constraint loss; Update the parameters of the policy network of the deep reinforcement learning network based on the resource allocation loss. Among them, the update of the policy network also uses the intermediate layer feature information coupled from the discriminator of the generative adversarial network as additional state information.

6. The channel optimization method applied to the multi-band antenna array according to claim 5, wherein The method of joint training also includes: adopting an alternating training method, setting that after every N times of training of the generative adversarial network, M times of training of the deep reinforcement learning network are executed, where N > M and the learning rate of the generative adversarial network is greater than the learning rate of the deep reinforcement learning network.

7. The channel optimization method applied to a multi-band antenna array according to claim 1, characterized in that, The method of closed-loop real-time feedback includes: Using a prioritized experience replay buffer to screen update samples for updating the policy network of the deep reinforcement learning network, where the update samples represent channel mutation events and have high weights; Encoding the policy parameters related to the policy network using polar codes and transmitting them through a dedicated control channel with a transmission error rate lower than 1e-6; Dynamically adjusting the feedback period of the closed-loop real-time feedback according to the Doppler frequency shift estimate to match the time-varying characteristics of the channel.

8. The channel optimization method applied to a multi-band antenna array according to claim 7, characterized in that, The calculation method of the feedback period includes: Calculating the channel coherence time and the Doppler frequency shift amount in real time. When the Doppler frequency shift amount exceeds a first preset threshold, shortening the feedback period to quickly respond to channel changes; otherwise, lengthening the feedback period to reduce system overhead.

9. Channel optimization system applied to a multi-band antenna array, characterized in that, The channel optimization system applied to the multi-band antenna array is used to implement the method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • A noise robust face recognition method based on a cascade deep convolutional neural network

    CN109948573A

  • CondenseNet algorithm fused with attention selection mechanism

    CN111160488A

  • Channel estimation enhancement method, device and system based on generative adversarial network

    CN111355675A

  • Deep learning training method and device for computing equipment

    CN112183718A

  • Method and system for extracting region of interest of video

    CN114782676A

Cited By

  • System and method for cooperative control of dynamic frequency bands of multimode communication terminal

    CN120692616A

  • A system and method for dynamic frequency band coordinated control of multi-mode communication terminals

    CN120692616B