Hybrid beamforming method and system against channel aging for MIMO-OFDM system
By using a deep learning end-to-end network architecture to predict channel states and optimize beamforming in MIMO-OFDM systems, the problem of beam inaccuracy and spectral efficiency degradation caused by channel aging is solved. This achieves efficient channel prediction and feedback, and improves system robustness and spectral efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING JIAOTONG UNIV
- Filing Date
- 2026-04-20
- Publication Date
- 2026-06-02
AI Technical Summary
In MIMO-OFDM systems, channel aging in high-mobility environments leads to beam misalignment and reduced spectral efficiency. Existing linear tracking schemes have large computational delays and nonlinear models are difficult to adapt to complex scattering environments, while static convolutional neural networks discard time-related information, resulting in spectral efficiency loss.
An end-to-end collaborative network architecture based on deep learning is adopted. The channel state is predicted by the learnable pilot sensing and bidirectional long short-term memory network at the user end. Combined with multi-scale convolution and channel attention calibration, a binary bit stream is generated and fed back to the base station. The base station reconstructs the hybrid beamforming matrix to optimize the parameters of the entire network.
It achieves adaptive prediction and efficient feedback of the channel in high mobility scenarios, improves beam alignment accuracy and spectral efficiency, reduces feedback overhead, and adapts to complex non-stationary scattering environments.
Smart Images

Figure CN122137436A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication, specifically to a hybrid beamforming method and system for resisting channel aging in MIMO-OFDM systems. Background Technology
[0002] Frequency division duplex massive multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) technology has become the core transmission solution for high-mobility scenarios such as vehicle-to-everything (V2X) communication in sixth-generation mobile communications due to its extremely high spectrum utilization. In this system, the base station transmits pilot signals, and the user terminal estimates channel state information and feeds it back to the base station to generate a hybrid beamforming matrix. However, the severe Doppler frequency shift in high-mobility environments, coupled with the inherent physical delays of pilot transmission and feedback calculation, causes the channel state information acquired by the base station to be severely aged by the time the actual data is transmitted, resulting in beam misalignment and a sharp decline in spectrum efficiency.
[0003] In existing technologies, linear tracking schemes based on multidimensional matrix beams estimate Doppler phase rotation and extrapolate channel states by solving generalized eigenvalues and singular value decompositions. Their advantage is that they do not require a large amount of training data, but their disadvantage is that each baseband processing unit must manually perform complex matrix inversion and decomposition operations, significantly increasing computational latency in high-mobility scenarios. Furthermore, rigid linear models struggle to adapt to nonlinear evolution under complex scattering environments. Channel state information compression feedback schemes based on static convolutional neural networks use independently trained autoencoders for feature extraction. Their advantage is that they effectively reduce feedback overhead, but their disadvantage is that they discard temporal correlation information of the channel in order to compress dimensionality, causing base station beam alignment to rely on outdated static snapshots, inevitably resulting in a loss of spectral efficiency. Summary of the Invention
[0004] To address the technical problems mentioned in the background section, this invention proposes a hybrid beamforming method and system for resisting channel aging in MIMO-OFDM systems.
[0005] Therefore, the technical solution adopted by the present invention is as follows: A hybrid beamforming method for channel aging resistance in MIMO-OFDM systems, characterized in that the method includes: S1: The user terminal receives the pilot signal sent by the base station, extracts and aggregates the channel semantic features, and further generates fused timing features; the nonlinear mapping value of the fused timing features is amplitude controlled to generate Doppler compensation increments, and the predicted channel state is obtained by superposition. S2: The predicted channel state is converted to the angle and time delay domain, the corresponding feature map is generated, and multi-scale fusion features are extracted. Channel attention calibration is performed on the multi-scale fusion features, and the calibrated features are quantized into binary bit streams and fed back to the base station. S3: Based on the binary bit stream, the hybrid beamforming matrix is reconstructed through a dual-head parallel decoding architecture. The first decoding head generates an analog pre-encoder that satisfies constant mode constraints, and the second decoding head reconstructs a digital pre-encoder. Finally, the end-to-end negative sum rate loss function is calculated, and the network parameters are optimized through the loss function.
[0006] Furthermore, the base station uses a learnable pilot matrix to perform active space detection and transmit pilot signals, which are received by the user terminal to obtain complex observation signals; The user terminal performs feature mapping and aggregation on the complex observation signals using a two-layer linear multilayer perceptron to obtain a historical channel feature matrix. Specifically, 1) Decompose the complex observation signal into real and imaginary parts and concatenate them along the channel dimension to obtain a real tensor; 2) The real tensor is used as the input of a two-layer linear multilayer perceptron. The first fully connected network maps the input to a high-dimensional feature space; the second fully connected network then reduces the high-dimensional features to a low-dimensional semantic space and outputs a channel semantic feature vector. 3) Aggregate the channel semantic feature vectors along the subcarriers to form a historical channel feature matrix, and store it in the historical observation sequence cache pool.
[0007] Furthermore, extract the historical observation sequence cache pool with a length of... Historical observation sequence , is represented as: in, Indicates the first Channel feature matrix of time slot, Indicates the first One user device; The historical observation sequence is input into a bidirectional long short-term memory network to extract the dual hidden states of forward and backward directions, and the fused temporal features are obtained by vector concatenation. .
[0008] Furthermore, the fused temporal features are matched to the same dimension as the channel feature matrix using a multilayer perceptron; Introducing learnable residual threshold parameters and constrained within the interval Within, amplitude control coefficients are generated to control the amplitude of the fused temporal features after dimension matching; The Doppler compensation increment is superimposed on the channel feature matrix to obtain the time slot. Predicted channel state.
[0009] Furthermore, the feature matrix of the predicted channel state is separated into two real matrices, one real and one imaginary, and then a discrete Fourier transform is performed along the subcarrier dimension to obtain the feature map. , is represented as: in, Represents the th element in the characteristic matrix The feature dimension at the th feature dimension Values on each subcarrier; The characteristic dimension of the subcarrier; Indicates the total number of subcarriers; The feature map is fed into two convolutional branches with different receptive fields for multi-scale feature extraction. The outputs of the two convolutional branches are concatenated to generate a concatenated vector, which is then subjected to grouping normalization. Finally, the multi-scale fused features are obtained through the Mish activation function.
[0010] Furthermore, average pooling is performed on the multi-scale fusion features to compress the spatial dimension of each channel into a scalar value. The global average pooling output of each channel is represented as: in, Represents the spatial dimensions of the multi-scale fusion features; concatenates the pooling results of each channel into a vector. Perform a one-dimensional convolution operation on the vector to obtain the dependencies between adjacent channels. ; By using the Sigmoid activation function to Mapped to The interval is used to obtain the attention weight for each channel. Pay attention weights The recalibrated feature map is obtained by multiplying the original multi-scale fusion feature map channel by channel. ; The recalibrated feature map is reduced in dimension and compressed using a fully connected layer to obtain a dimension of [dimensionality]. low-dimensional vectors Using differentiable layers to transform vectors Convert to binary bit stream .
[0011] Furthermore, the binary bitstream is aggregated into a matrix. The base station dequantizes and decodes the matrix to recover the coarse-grained feature tensor. ; The first decoding head uses a fully connected layer to... The mapping is performed as a phase angle matrix, and an analog precoder is constructed based on the phase angle matrix. ; The second decoding head will The image is projected into a high-dimensional space and reshaped into a spatial tensor. This tensor is then processed using a convolutional thinning network containing convolutional and ECA residual blocks, and separated into real and imaginary matrices. This process is then used to reconstruct the digital precoder for all subcarriers. It is then cascaded with an analog pre-encoder to form a complete hybrid beamforming matrix. .
[0012] Furthermore, the negative sum rate loss function is expressed as: in, This represents the training dataset; Indicates the total number of samples; All network parameters; Indicates the first The user equipment in the first The signal-to-interference-plus-noise ratio on each subcarrier; The negative sum rate loss gradient is calculated and backpropagated to complete the joint optimization update of the network parameters in this round. Then, the process returns to the pilot signal acquisition stage and starts the next round of communication scheduling based on the updated network parameters.
[0013] A hybrid beamforming system for MIMO-OFDM systems to combat channel aging, characterized in that the system comprises: At the user end, pilot signals sent by the base station are received, channel semantic features are extracted and aggregated, and further fused timing features are generated. The nonlinear mapping value of the fused timing features is amplitude controlled to generate Doppler compensation increments, and the predicted channel state is obtained by superposition. The predicted channel state is converted to the angle and time delay domains to generate corresponding feature maps, and multi-scale fused features are extracted. Channel attention calibration is performed on the multi-scale fused features, and the calibrated features are quantized into binary bit streams and fed back to the base station. The base station receives the binary bit stream and reconstructs the hybrid beamforming matrix through a dual-head parallel decoding architecture, wherein the first decoding head generates an analog pre-encoder and the second decoding head reconstructs a digital pre-encoder; the analog pre-encoder and the digital pre-encoder are applied to the downlink transmit link to perform beam alignment and data distribution. The joint optimization module aims to maximize the downlink speed of the system. It calculates the end-to-end negative sum rate loss function, performs joint optimization and updates on the network parameters, and distributes the updated parameters to user terminals and base stations.
[0014] Compared with the prior art, the advantages of the present invention are as follows: 1. This invention uses a unified network architecture based on deep learning end-to-end collaboration. By integrating pilot sensing, timing prediction and quantization feedback into a deep learning process, it solves the problems of severe information sharing difficulties among modules in traditional separate communication architectures and excessive redundant overhead due to decoupling design.
[0015] 2. This invention uses a residual compensation mechanism based on time sequence modeling. By dynamically controlling the Doppler frequency shift characteristics during the prediction process, it solves the problems of beam misalignment caused by the lack of capture of time evolution in existing static deep networks, and the precipitous drop in accuracy of traditional linear tracking algorithms in complex non-stationary scattering environments due to rigid model assumptions.
[0016] 3. This invention uses a hybrid beam generation technique that combines multi-scale convolution with dual-head reconstruction networks. By applying constant modulus hardware constraints and interference cancellation strategies through parallel paths, it solves the problems of existing technologies struggling to achieve near-full-digital capacity when faced with stringent analog phase shifter hardware limitations, and lacking an effective mechanism to balance feedback overhead and overall system spectral efficiency performance. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of the hybrid beamforming anti-channel aging method of the present invention; Figure 2 This is a flowchart of the predicted channel state generation process of the present invention; Figure 3 This is a flowchart of the binary bit stream generation process of the present invention. Detailed Implementation
[0019] To achieve the above objectives, this invention provides a hybrid beamforming method for resisting channel aging in MIMO-OFDM systems. Please refer to [link to relevant documentation]. Figures 1-3 The method includes: S1: The user terminal receives the pilot signal transmitted by the base station, extracts and aggregates the channel semantic features, and further generates fused timing features; the nonlinear mapping value of the fused timing features is amplitude controlled to generate Doppler compensation increments, and the predicted channel state is obtained by superposition. This step utilizes learnable pilot sensing (LPS) to achieve end-to-end optimization from active base station detection to nonlinear feature mapping at the user end, providing high-quality channel semantic features for subsequent time series prediction. The base station maintains a parameterized learnable pilot matrix. ,in This represents the total number of base station antennas. The length of the pilot sequence is given. To ensure that the transmit power meets the normalization constraint, the matrix is normalized during forward propagation. The base station transmits the pilot signal using an active space detection strategy. This signal is fading through a broadband non-stationary physical channel and superimposed with environmental noise before being received by the user terminal, resulting in a complex observation signal. , Indicates the first Each user device. Indicates the first Subcarriers, Indicates the current time slot; To prevent severe noise amplification and the curse of dimensionality from traditional linear estimation algorithms, the user terminal, upon receiving a complex observation signal, no longer attempts to directly reconstruct the physical channel. Instead, it performs nonlinear semantic feature extraction. Specifically, the complex observation signal is decomposed into real and imaginary parts and concatenated along the channel dimension to obtain a real tensor, represented as: This tensor serves as the input to a two-layer linear multilayer perceptron (MLP). The first fully connected layer maps the input to a high-dimensional feature space, which then undergoes a nonlinear transformation using the Mish activation function. The second fully connected layer then reduces the high-dimensional features to a low-dimensional semantic space, outputting a channel semantic feature vector. , is represented as: in, Represents the upward projection weight matrix; Represent the downward projection weight matrix; and This represents the bias term; compared to the traditional ReLU function, Mish retains a small amount of gradient information in the negative region, which can improve the richness of feature representation and alleviate the problem of neuron "death". Subsequently, the user terminal aggregates the semantic feature vectors on all subcarriers along the subcarrier dimension to form a complete channel feature matrix. This feature matrix is stored in the local historical observation sequence cache pool. The aforementioned nonlinear sensing mechanism not only effectively filters out environmental noise, but also extracts low-dimensional robust semantic features that are highly correlated with the physical channel state, providing high-quality initial input for time series prediction.
[0020] This embodiment predicts the channel state at future times based on historical observation sequences, thereby providing forward-looking compensation for channel aging. Unlike traditional schemes that only feed back the current estimated channel, this embodiment utilizes the historical observation sequences stored at the user terminal to predict the channel state at future times. Extract the length of the local cache. The historical observation sequence is represented as follows: in, Indicates the first The channel feature matrix of the time slot; the sequence is input into a bidirectional long short-term memory network (Bi-LSTM); Bi-LSTM contains two branches, forward LSTM and backward LSTM, which process the sequence in forward and reverse order of time, respectively, thereby capturing the temporal dependencies of the past and the future at the same time; Configure the forward LSTM in the time slot The hidden state of the output is The hidden state output by the LSTM is The fused temporal feature vector is obtained by concatenating the vectors of the forward and backward hidden states, and is represented as: at this time, It only contains timing information and has not yet established any connection with the current channel state. In order to achieve stable and adaptive prediction in a drastically dynamic environment, a learnable residual threshold parameter is introduced. This parameter is used in the end-to-end training along with the Bi-LSTM network.
[0021] Through a multilayer perceptron Mapping to the same dimensional space as the channel feature matrix yields nonlinear mapping values. , Then, the learnable residual threshold is set using the Sigmoid function. Constraints in interval Within, generate amplitude control coefficients. This coefficient determines the intensity of the incremental compensation for the predicted branch output; the final Doppler compensation increment is expressed as: This increment and Superimposing these values yields the future time (i.e., the time slot). The predicted channel state is represented as: The aforementioned residual threshold mechanism functions as both a stabilizer and a compensator. In static or low-speed moving environments, due to the small Doppler frequency shift, the network, through training, enables... It tends to 0, and thus Approaching 0, the prediction results mainly depend on This avoids unnecessary fluctuations; however, in high-mobility scenarios, the Doppler shift is drastic, and the network... Tend to 1, As the value approaches 1, the output of the prediction branch is amplified sufficiently, thereby achieving accurate incremental compensation for channel phase transitions. This design enables the network to adaptively adjust the prediction strength according to the current dynamic level of the channel without the need to manually set a switching threshold. After completing the prediction, the user equipment will predict the channel state and the threshold parameters for participating in this round of compensation. Transmitted together, along with threshold parameters It is also sent to the optimizer, so that in subsequent backpropagation, the optimizer can accurately calculate the compensation weight gradient containing only the high mobility increment, thereby collaboratively eliminating beam misalignment caused by channel aging in end-to-end training.
[0022] S2: Transform the predicted channel state into the angle and time delay domains, generate the corresponding feature map, extract multi-scale fusion features, perform channel attention calibration on the multi-scale fusion features, and quantize the calibrated features into a binary bit stream and feed it back to the base station. This step efficiently compresses the predicted future channel state to accommodate the limited bandwidth constraints of the uplink feedback link, while ensuring that the base station can accurately recover the channel backbone characteristics from the feedback information with extremely low bit overhead.
[0023] The user terminal receives the feature matrix (prediction feature matrix) for predicting the channel state. ,in For each subcarrier's feature dimension, The feature matrix represents the total number of subcarriers and indicates future time slots. The channel state information is obtained by considering that the original frequency domain features have different sparsity levels in the angle and time delay domains. That is, the channel energy is usually concentrated in a few angle and time delay taps, while it is sparse in other regions. Directly inputting the frequency domain features into the convolutional network will make it difficult for the convolutional kernel to take into account both local details and global structure. Therefore, the predicted features are subjected to domain transformation processing, specifically, The predicted feature matrix is separated into two real matrices, one real and one imaginary. Then, a Discrete Fourier Transform (DFT) is performed along the subcarrier dimension (i.e., the frequency domain dimension) to transform the signal from the frequency domain to the angle-delay domain. The transformed angle-delay domain feature map is then obtained. Represented as: in, Represents the th element in the characteristic matrix The feature dimension at the th feature dimension The values on each subcarrier; to facilitate processing by convolutional networks, the transformed amplitude is usually taken or it is treated as the real and imaginary parts of a complex tensor concatenated. Since the angle-delay domain has natural sparsity, that is, the multipath components of the channel are only distributed on a few angle and delay taps, the values of most regions in the transformed feature map are close to zero, while the dominant multipath components are concentrated in several obvious peaks. This sparsity characteristic enables the subsequent convolutional network to capture the physical backbone information of the channel with extremely high efficiency. After the domain transformation is completed, the feature maps are fed in parallel into two convolutional branches with different receptive fields to achieve multi-scale feature extraction. The first branch uses... The convolution kernel, with its relatively small receptive field, is primarily used to capture fine local details and rapid phase changes between adjacent subcarriers. Let the kernel weight of this branch be... ,in, and Let the number of input and output channels be respectively, then the convolution output is represented as: in, Represents a two-dimensional convolution operation; The second branch adopts The convolution kernel has a large receptive field and is used to capture long-range information dependent on the global channel structure. The output is represented as: in, The weights are represented by two branches that are computed in parallel without interfering with each other, thus simultaneously extracting local fine features and global structural features.
[0024] To fuse the heterogeneous features of the two branches, the two outputs are concatenated along the channel dimension to obtain the concatenated feature tensor. Subsequently, Group Normalization (GN) is performed on the spliced tensor. Unlike batch normalization, group normalization divides the channels into several groups and calculates the mean and variance independently within each group, thereby eliminating the dependence on batch size, which is particularly suitable for small-batch training scenarios. Set the output of group normalization to The calculation process is as follows: After normalization, the feature distribution is stabilized to near zero mean and unit variance, which is beneficial for the nonlinear mapping of the subsequent activation function. Finally, the multi-scale fused features are obtained through the Mish activation function, expressed as: in As a non-linear activation function, it can retain a small amount of gradient in the negative region, thus enhancing the richness of feature representation.
[0025] To further enhance the representation of key channel components and suppress the interference of quantization noise on weak signals, an efficient channel attention (ECA) mechanism is embedded within the encoder. Unlike traditional fully connected layer attention, this mechanism uses one-dimensional convolution to capture cross-channel interaction information, achieving dynamic calibration of channel weights with extremely low computational overhead. Specifically, right Performing Global Average Pooling (GAP) compresses the spatial dimension of each channel into a scalar value, setting... The space dimensions are (corresponding angle and time delay dimension), then the first The global average pooling output of each channel is represented as: The pooling results of each channel are concatenated into a vector. Then, perform a one-dimensional convolution operation on the vector, with a kernel size of . (Usually a value of 3 or 5), thereby capturing the dependencies between adjacent channels. Next, the Sigmoid activation function is used to activate... Mapped to The interval is used to obtain the attention weight for each channel. Pay attention weights The recalibrated feature map is obtained by multiplying the original fused feature map channel by channel. ; After completing attention calibration, recalibrate the feature maps. The data is fed into a fully connected layer for dimensionality reduction and compression, which maps the features of the high-dimensional space to a dimension of 2. low-dimensional vectors ,in The preset number of feedback bits is then used to convert the vector using a differentiable layer. Converted to a binary bitstream, the quantization operation performs a deterministic rounding function during forward propagation, represented as: in, Round each element to either 0 or 1; Will Compress to The interval is used to ensure quantization interval matching, because The function is not differentiable at zero, making direct backpropagation gradient calculation impossible. This embodiment introduces a pass-through estimator to approximate the gradient. Specifically, during forward propagation, a pass-through estimator is used. It generates discrete outputs; during backpropagation, the gradient of the quantization operation is approximated as an identity mapping, so that the gradient can flow unimpeded through the quantization layer and update the network parameters at the front end.
[0026] Finally, the user terminal transmits the generated binary bit stream back to the base station via the uplink air interface link. Since the bit length is much smaller than the dimension of the original channel feature matrix (i.e., ...), ... This stage achieves an extremely high compression ratio, significantly reducing uplink feedback overhead, while ensuring the fidelity of the channel backbone information through multi-scale feature extraction and attention calibration.
[0027] S3: Based on the binary bitstream, a hybrid beamforming matrix is reconstructed using a dual-head parallel decoding architecture. The first decoding head generates an analog pre-encoder that satisfies constant mode constraints, and the second decoding head reconstructs the digital pre-encoder. Finally, an end-to-end negative sum rate loss function is calculated, and the network parameters are optimized using this loss function. This step decodes and reconstructs the binary bitstream uploaded by the user, generates a hybrid precoding matrix that meets hardware constraints, and achieves collaborative updates of network parameters through end-to-end joint optimization.
[0028] After receiving the binary bit stream, the base station first performs aggregation processing and sets the... The bitstream fed back by each user equipment is The base station then aggregates the bitstreams of all user equipment into a matrix. By performing dequantization, discrete binary values are restored to dense, continuous spatial features. The dequantization operation employs a mapping function symmetric to the quantization function at the user end, expressed as: in, Indicates the scaling factor, used to... The quantized values of the interval are mapped back to the original continuous range; after dequantization, the continuous feature vector of the user equipment is obtained. After concatenating the feature vectors of all user devices, the concatenation is fed into a deep fully connected network to recover the high-dimensional continuous spatial feature tensor. The deep fully connected network consists of two cascaded fully connected layers. The first layer contains 4096 neurons, and the second layer contains 2048 neurons. Each layer is followed by a Mish activation function; the dequantized output vector is... (The result of splicing all user devices) is then the coarse-grained feature tensor after passing through the fully connected network. Represented as: in, and This represents the weights and biases of the first fully connected layer; and This represents the weights and biases of the second fully connected layer.
[0029] The dual-head parallel decoding architecture generates analog and digital precoders separately, and both decoding heads will... Each input has its own independent network structure and parameters, enabling targeted optimization for different physical constraints (constant modulus constraint in the analog domain, and elimination of multi-user interference in the digital domain). The first path is the Analog Head, used to generate the analog pre-encoder. This branch uses a fully connected layer to map coarse-grained features into a phase angle matrix, specifically, Input features First, it goes through a containment A fully connected layer with output neurons yields an unconstrained phase angle vector. Then it is compressed to its final size using the element-wise sigmoid function. The interval is linearly scaled to... The range is represented as: in, and This represents the weights and biases of the simulated fully connected layer; Indicates the first Line number The phase angles of the columns; thus, the complete phase angle matrix is obtained. ; The analog precoder is constructed based on the phase angle matrix and is expressed as: Each element in this construction method The modulus is constant. Therefore, it naturally satisfies the constant modulus constraint of hardware phase shifters, requiring no additional projection or regularization operations; at the same time, all phase shifters share a unified amplitude factor. This ensures the normalization of transmission power.
[0030] The second path is the digital head, used to generate the digital precoder. ,in Representing the subcarrier index, unlike analog precoders, digital precoders need to independently process frequency-selective fading on each subcarrier; this branch first projects coarse-grained features to a high-dimensional space and reshapes them into a spatial tensor, specifically, First, use a fully connected layer to... Mapped to a space tensor, then utilizing the inclusion... The convolutional and ECA residual blocks are processed by a convolutional thinning network, with the input of the convolutional thinning network set as... The output after passing through several residual blocks is: in, Indicates the first The convolution, normalization, and attention operations are performed in each residual block; the refined feature map is separated into real and imaginary matrices, and then the baseband digital precoder for all subcarriers is reconstructed. It is then cascaded with an analog pre-encoder to form a complete hybrid beamforming matrix. .
[0031] After the reconstruction of the analog and digital precoders is completed, the joint optimization phase begins. This embodiment abandons the decoupling approach of designing each module independently in traditional communication systems, and unifies all trainable parameters (including the learnable pilot matrix on the base station side, the semantic extraction network on the user side, the Bi-LSTM prediction network, the multi-scale compression coding network, and the decoding and reconstruction network on the base station side) into a parameter set. The goal of optimization is to optimize the data transmission process at the actual data transmission time. Maximize the downlink sum rate. No. The user equipment in the first The signal-to-interference-plus-noise ratio on each subcarrier is denoted as . , Training employs an unsupervised learning strategy, defining the average negative sum rate as the loss function, expressed as: in, This represents the training dataset; This represents the total number of samples; the loss function directly uses the core performance indicators (and rate) of the communication system as the optimization objective, avoiding the drawback of fitting intermediate labels module by module in traditional methods; Using stochastic gradient descent (SGD) or its adaptive variants (such as the Adam optimizer) to adjust the loss function By minimizing the process, deep collaboration can be achieved in steps such as pilot sensing, timing prediction, feature compression, and beam reconstruction. After each round of parameter updates, the updated global network parameters are distributed to the base station and each user equipment, and the process returns to the pilot acquisition stage to start the next round of communication scheduling. The above process will continue to iterate until the loss function converges or the preset number of training rounds is reached. In online deployment mode (i.e., actual communication application after training), the base station no longer needs to perform backpropagation. It only needs to perform forward propagation to quickly reconstruct the hybrid precoding matrix from the binary bit stream fed back by the user and apply it to the downlink transmission link to achieve accurate beam alignment and data distribution. Since the computational complexity of neural network forward propagation is much lower than that of traditional matrix inversion or singular value decomposition methods, this embodiment can meet the stringent requirements for low latency in high mobility scenarios such as vehicle networking.
[0032] A hybrid beamforming system for MIMO-OFDM systems to combat channel aging, characterized in that the system comprises: At the user end, pilot signals sent by the base station are received, channel semantic features are extracted and aggregated, and further fused timing features are generated. The nonlinear mapping value of the fused timing features is amplitude controlled to generate Doppler compensation increments, and the predicted channel state is obtained by superposition. The predicted channel state is converted to the angle and time delay domains to generate corresponding feature maps, and multi-scale fused features are extracted. Channel attention calibration is performed on the multi-scale fused features, and the calibrated features are quantized into binary bit streams and fed back to the base station. The base station receives the binary bit stream and reconstructs the hybrid beamforming matrix through a dual-head parallel decoding architecture, wherein the first decoding head generates an analog pre-encoder and the second decoding head reconstructs a digital pre-encoder; the analog pre-encoder and the digital pre-encoder are applied to the downlink transmit link to perform beam alignment and data distribution. The joint optimization module aims to maximize the downlink speed of the system. It calculates the end-to-end negative sum rate loss function, performs joint optimization and updates on the network parameters, and distributes the updated parameters to user terminals and base stations.
[0033] This invention proposes a hybrid beamforming method and system for MIMO-OFDM systems to combat channel aging. By constructing an end-to-end deep learning framework, it integrates pilot sensing, timing prediction, feature compression, and hybrid beam reconstruction. Specifically, the user end first utilizes learnable pilot sensing and a bidirectional long short-term memory network to extract historical channel semantic features, and introduces a learnable residual threshold mechanism to generate Doppler compensation increments, achieving adaptive prediction of future channel states. Subsequently, the predicted channel is converted to the angle-delay domain, quantized into a binary bitstream after multi-scale convolution and efficient channel attention calibration, and fed back to the base station. The base station employs a dual-head parallel decoding architecture to reconstruct an analog precoder that satisfies constant modulus constraints and a digital precoder for frequency-selective fading, respectively. With the goal of maximizing system and rate, it achieves joint optimization of all network parameters through a negative sum rate loss function.
[0034] In summary, this invention effectively solves the problems of beam misalignment and spectral efficiency degradation caused by channel aging in high mobility scenarios. Compared with traditional linear tracking and static compressed feedback schemes, it has stronger nonlinear modeling capabilities, lower feedback overhead, and higher beam alignment accuracy, significantly improving the robustness and spectral efficiency of MIMO-OFDM systems in complex dynamic environments such as vehicle-to-everything (V2X) networks.
[0035] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A hybrid beamforming method for resisting channel aging in MIMO-OFDM systems, characterized in that, The method includes: S1: The user terminal receives the pilot signal sent by the base station, extracts and aggregates the channel semantic features, and further generates fused timing features; the nonlinear mapping value of the fused timing features is amplitude controlled to generate Doppler compensation increments, and the predicted channel state is obtained by superposition. S2: The predicted channel state is converted to the angle and time delay domain, the corresponding feature map is generated, and multi-scale fusion features are extracted. Channel attention calibration is performed on the multi-scale fusion features, and the calibrated features are quantized into binary bit streams and fed back to the base station. S3: Based on the binary bit stream, the hybrid beamforming matrix is reconstructed through a dual-head parallel decoding architecture. The first decoding head generates an analog pre-encoder that satisfies constant mode constraints, and the second decoding head reconstructs a digital pre-encoder. Finally, the end-to-end negative sum rate loss function is calculated, and the network parameters are optimized through the loss function.
2. The hybrid beamforming method for resisting channel aging in MIMO-OFDM systems according to claim 1, characterized in that, The base station transmits pilot signals for active space detection using a learnable pilot matrix, which are received by the user terminal to obtain complex observation signals. The user terminal performs feature mapping and aggregation on the complex observation signals using a two-layer linear multilayer perceptron to obtain a historical channel feature matrix. Specifically, 1) Decompose the complex observation signal into real and imaginary parts and concatenate them along the channel dimension to obtain a real tensor; 2) The real tensor is used as the input of a two-layer linear multilayer perceptron. The first fully connected network maps the input to a high-dimensional feature space; the second fully connected network then reduces the high-dimensional features to a low-dimensional semantic space and outputs a channel semantic feature vector. 3) Aggregate the channel semantic feature vectors along the subcarriers to form a historical channel feature matrix, and store it in the historical observation sequence cache pool.
3. The hybrid beamforming method for resisting channel aging in MIMO-OFDM systems according to claim 2, characterized in that, Extract the historical observation sequence from the cache pool with a length of Historical observation sequence , is represented as: in, Indicates the first Channel feature matrix of time slot, Indicates the first One user device; The historical observation sequence is input into a bidirectional long short-term memory network to extract the dual hidden states of forward and backward directions, and the fused temporal features are obtained by vector concatenation. .
4. The hybrid beamforming method for resisting channel aging in MIMO-OFDM systems according to claim 3, characterized in that, The fused temporal features are matched to the same dimension as the channel feature matrix using a multilayer perceptron. Introducing learnable residual threshold parameters and constrained within the interval Within, amplitude control coefficients are generated to control the amplitude of the fused temporal features after dimension matching; The Doppler compensation increment is superimposed on the channel feature matrix to obtain the time slot. Predicted channel state.
5. The hybrid beamforming method for resisting channel aging in MIMO-OFDM systems according to claim 4, characterized in that, The feature matrix of the predicted channel state is separated into two real matrices, one real and one imaginary. Then, a discrete Fourier transform is performed along the subcarrier dimension to obtain the feature map. , represented as: in, Represents the th element in the characteristic matrix The feature dimension at the th feature dimension Values on each subcarrier; The characteristic dimension of the subcarrier; Indicates the total number of subcarriers; The feature map is fed into two convolutional branches with different receptive fields for multi-scale feature extraction. The outputs of the two convolutional branches are concatenated to generate a concatenated vector, which is then subjected to grouping normalization. Finally, the multi-scale fused features are obtained through the Mish activation function.
6. The hybrid beamforming method for resisting channel aging in MIMO-OFDM systems according to claim 5, characterized in that, Average pooling is performed on the multi-scale fusion features to compress the spatial dimension of each channel into a scalar value. The global average pooling output of each channel is represented as: in, Represents the spatial dimensions of the multi-scale fusion features; concatenates the pooling results of each channel into a vector. Perform a one-dimensional convolution operation on the vector to obtain the dependencies between adjacent channels. ; By using the Sigmoid activation function to Mapped to The interval is used to obtain the attention weight for each channel. Pay attention weights The recalibrated feature map is obtained by multiplying the original multi-scale fusion feature map channel by channel. ; The recalibrated feature map is reduced in dimension and compressed using a fully connected layer to obtain a dimension of [dimensionality]. low-dimensional vectors Using differentiable layers to transform vectors Convert to binary bit stream .
7. The hybrid beamforming method for resisting channel aging in MIMO-OFDM systems according to claim 6, characterized in that, The binary bit stream is aggregated into a matrix. The base station dequantizes and decodes the matrix to recover the coarse-grained feature tensor. ; The first decoding head uses a fully connected layer to... The mapping is performed as a phase angle matrix, and an analog precoder is constructed based on the phase angle matrix. ; The second decoding head will The image is projected into a high-dimensional space and reshaped into a spatial tensor. This tensor is then processed using a convolutional thinning network containing convolutional and ECA residual blocks, and separated into real and imaginary matrices. This process is then used to reconstruct the digital precoder for all subcarriers. It is then cascaded with an analog pre-encoder to form a complete hybrid beamforming matrix. .
8. The hybrid beamforming method for resisting channel aging in MIMO-OFDM systems according to claim 7, characterized in that, The negative sum rate loss function is expressed as: in, This represents the training dataset; Indicates the total number of samples; All network parameters; Indicates the first The user equipment in the first The signal-to-interference-plus-noise ratio on each subcarrier; The negative sum rate loss gradient is calculated and backpropagated to complete the joint optimization update of the network parameters in this round. Then, the process returns to the pilot signal acquisition stage and starts the next round of communication scheduling based on the updated network parameters.
9. A hybrid beamforming system for MIMO-OFDM systems to resist channel aging, characterized in that, The system includes: At the user end, pilot signals sent by the base station are received, channel semantic features are extracted and aggregated, and further fused timing features are generated. The nonlinear mapping value of the fused timing features is amplitude controlled to generate Doppler compensation increments, and the predicted channel state is obtained by superposition. The predicted channel state is converted to the angle and time delay domains to generate corresponding feature maps, and multi-scale fused features are extracted. Channel attention calibration is performed on the multi-scale fused features, and the calibrated features are quantized into binary bit streams and fed back to the base station. The base station receives the binary bit stream and reconstructs the hybrid beamforming matrix through a dual-head parallel decoding architecture, wherein the first decoding head generates an analog pre-encoder and the second decoding head reconstructs a digital pre-encoder; the analog pre-encoder and the digital pre-encoder are applied to the downlink transmit link to perform beam alignment and data distribution. The joint optimization module aims to maximize the downlink speed of the system. It calculates the end-to-end negative sum rate loss function, performs joint optimization and updates on the network parameters, and distributes the updated parameters to user terminals and base stations.