A robust beam tracking method for U6G ultra-massive MIMO hybrid field
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]发明目的:为了克服现有技术在混合场环境下场域切换适应性差、低信噪比鲁棒性不足、多径合并增益损失等关键问题,提供一种面向U6G超大规模MIMO混合场的鲁棒波束追踪方法,通过CNN-Transformer架构实现从含噪角度谱序列到最优复合波束矢量的映射,直接输出综合所有多径分量的最优相干合并波束,有效利用LoS径和NLoS径的信息,在低信噪比条件下仍能实现稳定的波束追踪,且每个追踪周期的导频开销降至常数级别
[0042]1、本发明方法通过CNN特征提取模块的层次化特征抽象和残差连接,有效抑制低信噪比条件下角度谱中的噪声干扰,在宽泛的信噪比范围内实现稳定的波束追踪性能。
Smart Images

Figure CN122553955A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication and relates to robust beam tracking technology, specifically a robust beam tracking method for U6G ultra-large-scale MIMO mixed fields. Background Technology
[0002] Ultra-large-scale MIMO (U6G) technology is one of the key enabling technologies for fully unleashing the potential of the U6G band. By densely deploying large-scale antenna arrays at the base station, the short-wavelength advantage of the U6G band can be transformed into higher spectral efficiency and stronger spatial resolution. However, with the rapid increase in antenna array aperture, the Rayleigh distance increases proportionally to the square of the array aperture, making it more likely that users within the actual communication distance will fall into the near-field region. The channel exhibits a mixed field characteristic with both far-field and near-field path components. Beam training requires a large number of pilots, resulting in high training overhead. Rapidly tracking position changes based on the historical position information of users or scatterers can reduce the use of pilots and thus reduce overhead. In mixed-field environments, beam tracking must simultaneously process the far-field path determined only by angle parameters and the near-field path determined jointly by angle and range parameters, significantly increasing complexity. Moreover, under low signal-to-noise ratio conditions (0dB and below), the performance of traditional channel estimation methods degrades severely. Taking least squares (LS) channel estimation as an example, its estimation error power is proportional to the noise power. When SNR=0dB, the estimation error is on the same order of magnitude as the channel's energy, resulting in severe distortion of the reconstructed channel vector. Moreover, when the user or scatterer is in motion, the channel parameters continuously change over time. If a complete beam training is re-executed in each coherence time interval, the accumulated pilot overhead will severely erode the effective data transmission time.
[0003] Preliminary research has been conducted on near-field beam management methods for XL-MIMO systems. However, existing beam tracking schemes suffer from the following shortcomings: First, traditional channel estimation methods exhibit drastic performance degradation under low signal-to-noise ratio (SNR) conditions. In practical deployments, users are often located at the coverage edge, and relying on channel estimation results to drive beam updates will lead to severe beam inaccuracies. Second, existing schemes are primarily designed for single far-field or near-field scenarios, lacking adaptability to field switching in mixed-field environments. Kalman filter-based schemes assume path parameters follow a linear dynamic model, while neighbor-based beam search schemes rely on the topological continuity of beam indices in the codebook. Both of these assumptions fail when field switching occurs. Third, most existing beam tracking schemes neglect multipath effects, outputting only the optimal beam codeword for a single path. They cannot coherently combine the energy of all path components in a mixed-field channel, resulting in multipath combining gain loss.
[0004] In summary, how to achieve robust tracking of dynamic mixed-field multipath channels in U6G ultra-large-scale MIMO systems, while also having the ability to handle field switching and low signal-to-noise ratio conditions, has become an urgent technical problem to be solved. Summary of the Invention
[0005] Purpose of the invention: To overcome the key problems of existing technologies, such as poor adaptability to field switching in mixed field environments, insufficient robustness at low signal-to-noise ratios, and multipath merging gain loss, this invention provides a robust beam tracking method for U6G ultra-large-scale MIMO mixed fields. It realizes the mapping from noisy angle spectrum sequence to the optimal composite beam vector through a CNN-Transformer architecture, directly outputs the optimal coherently combined beam that integrates all multipath components, effectively utilizes the information of LosS and NLoS paths, and can still achieve stable beam tracking under low signal-to-noise ratio conditions, while reducing the pilot overhead of each tracking cycle to the constant level.
[0006] Technical Solution: To achieve the above objectives, this invention provides a robust beam tracking method for U6G ultra-large-scale MIMO mixed fields, comprising the following steps:
[0007] S1: Based on the random motion model of the user and the scatterer in two-dimensional space, a hybrid field time-varying channel containing both far-field and near-field path components is generated, and a dynamic hybrid field channel model is established.
[0008] S2: Based on the dynamic hybrid field channel model, training data is constructed and preprocessed using a data generation strategy based on the geometric physical model;
[0009] S3: Extract angular spectral space features using a well-designed CNN feature extraction module;
[0010] S4: Based on the angular spectrum space characteristics, the temporal dependencies are captured by the well-designed Transformer temporal coding module, and the coding features are output;
[0011] S5: Based on the coding characteristics, the predictor outputs the optimal composite beam vector that integrates all multipath components.
[0012] Furthermore, the dynamic hybrid field channel model in step S1 is expressed as follows:
[0013] ;
[0014] Channel vector modeling is a superposition of far-field path components and near-field path components: the far-field path adopts a planar wavefront steering vector. The description states that the near-field path uses a spherical wavefront steering vector. Description: The user and the scatterer move with random velocity and direction in a two-dimensional plane, and their positions are updated as follows. ,in For speed, It is a unit vector in direction. The sampling interval is given by the scatterer or user at time [time]. Location and the position of the base station reference antenna Positional changes are directly mapped to temporal variations in channel angle and distance parameters through geometric relationships. ; at every moment The user sends a pilot symbol, and the base station receives the received signal vector. Perform a DFT on the received signal to obtain the time. Angular spectrum.
[0015] Furthermore, the method for constructing training data based on the data generation strategy of the geometric physical model in step S2 includes: simulating the random motion of the user and the scatterer in two-dimensional space, calculating the channel angle, distance, and gain parameters according to the real-time position, and then generating the channel vector by the steering vector model; randomly and uniformly sampling the signal-to-noise ratio over a wide range to enhance generalization ability; the motion parameters of each path are independent of each other, and the training data naturally includes field switching events; no minimum spacing constraint is imposed on the angle parameters of the four paths, and the training set covers angle spectrum overlapping scenarios.
[0016] Furthermore, the generation of training data in step S2 includes: generating a mixed-field time-varying channel trajectory through a physical model, dividing the training samples using a sliding window mechanism, with each sample containing the angle spectrum of the previous time and the current time and the optimal beam vector of the previous time as input, and using the optimal beam vector of the current time as a label.
[0017] Furthermore, the data preprocessing in step S2 includes: normalizing the angle spectrum to eliminate absolute amplitude differences; representing complex data with real and imaginary parts to avoid discontinuities in phase transitions; and dividing the normalized angle spectrum into 8 feature channels, namely the normalized amplitude, real part, and imaginary part of the angle spectrum at the previous moment, the normalized amplitude, real part, and imaginary part of the angle spectrum at the current moment, and the real and imaginary parts of the beam vector at the previous moment.
[0018] Furthermore, the CNN feature extraction module in step S3 includes an Initial Conv1D layer and multiple cascaded ResBlocks. Each ResBlock contains two layers of one-dimensional convolution, batch normalization, GELU activation, and residual connections. The sequence length is halved step by step through convolution with a downsampling stride of 2. The residual connections ensure that global spectral information can be obtained under low signal-to-noise ratio conditions.
[0019] Furthermore, the operation of the CNN feature extraction module in step S3 includes:
[0020] The CNN feature extraction module first uses the Initial Conv1D layer to extract the feature. The number of input channels has been expanded to Each channel, kernel size The sequence length is maintained. Unchanged; then connected A cascaded ResBlock, via After one ResBlock, the output feature map dimension is ,in This represents the final number of channels. The length of the compressed sequence; the feature map output by the CNN is flattened and then passed through a linear projection layer. and bias vector Mapped to The embedding vector is obtained by applying layer normalization to the 3D space. .
[0021] Furthermore, in step S4, the Transformer temporal coding module is composed of multiple identical coding layers stacked together, and each coding layer includes a multi-head self-attention sub-layer and a feedforward network sub-layer;
[0022] The multi-head self-attention mechanism maps the input to multiple subspaces through multiple sets of different projection matrices, calculates attention separately, and then concatenates and fuses the results. Different attention heads can capture motion features of different path components.
[0023]
[0024]
[0025] in For the number of attention heads, , , Here are the projection matrices for each head. To output the projection matrix;
[0026] The feedforward network in each coding layer consists of two linear transformations and a GELU activation function.
[0027]
[0028] in, , This is the weight matrix. The hidden layer dimension of the feedforward network, and The bias vector is used; the role of the feedforward sub-layer is to perform nonlinear transformation and feature purification on the features captured by the multi-head attention sub-layer, amplify the feature signals that are useful for prediction, and suppress useless noise features.
[0029] The output of each sublayer is processed by residual concatenation and LayerNorm.
[0030]
[0031] The Sublayer is either an MHSA or FFN sublayer; LayerNorm normalizes the feature dimensions of each sample.
[0032]
[0033] in and These are the mean and variance along the feature dimension, respectively; and For learnable scaling and offset parameters, It is the numerical stability constant;
[0034] After processing by the Transformer encoder, the output of the last time slot in the sequence (i.e., the position corresponding to the current time t) is taken as the final encoded feature representation, with a dimension of . .
[0035] Furthermore, the process of the prediction head outputting the optimal composite beam vector that integrates all multipath components in step S5 includes:
[0036] The prediction head will output the Transformer encoder's output. The dimensional encoded features are mapped to the beam vector space and consist of two fully connected layers: the first layer will... Dimensional encoded feature mapping to The second layer will use GELU activation. Dimension mapping to Dimension; Output 3D real vector The former Each component and the last Each component is used as the real and imaginary parts of the complex beam vector, respectively, and reconstructed as... Complex beam vector;
[0037]
[0038] in and These represent the first and second halves of the prediction head output, respectively; ultimately, for Normalization is performed to meet power constraints.
[0039] .
[0040] Furthermore, for the training of the network model consisting of the CNN feature extraction module, the Transformer temporal coding module, and the prediction head, the loss function adopts a cosine similarity-based form, which is invariant to the common phase factor between complex vectors; the training uses the Adam optimizer in conjunction with a cosine annealing restart scheduler to dynamically adjust the learning rate.
[0041] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0042] 1. The method of the present invention effectively suppresses noise interference in the angle spectrum under low signal-to-noise ratio conditions through hierarchical feature abstraction and residual connection of the CNN feature extraction module, and achieves stable beam tracking performance over a wide signal-to-noise ratio range.
[0043] 2. The method of the present invention, through the self-attention mechanism of the Transformer temporal coding module, does not make any prior assumptions about the functional form of channel evolution, and can adaptively handle various nonlinear dynamic changes, including field switching, thus overcoming the fundamental defects of traditional Kalman filtering schemes and neighboring beam search schemes in field switching scenarios.
[0044] 3. The method of this invention breaks the traditional paradigm of path-by-path parameter estimation and beam synthesis. The network directly outputs the optimal composite beam vector that integrates all multipath components, making full use of the multipath merging gain in the mixed field channel and avoiding error accumulation in multi-stage parameter estimation.
[0045] 4. The method of the present invention requires only one pilot to acquire the angle spectrum in each tracking cycle in the uplink system, reducing the pilot overhead to the constant level, which is independent of the number of antennas and the codebook size, and has good scalability. Attached Figure Description
[0046] Figure 1 This is a flowchart of the method of the present invention;
[0047] Figure 2 This is a schematic diagram of a large-scale MIMO dynamic mixing field scenario;
[0048] Figure 3 A schematic diagram illustrating the construction of a sliding window and the generation of training samples;
[0049] Figure 4 This is a schematic diagram of the CNN-Transformer network architecture;
[0050] Figure 5 The graph shows how the achievable rate changes with SNR.
[0051] Figure 6 This is a diagram showing how the achievable rate varies with the number of antennas. Detailed Implementation
[0052] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading this invention, any modifications of the invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0053] Example 1:
[0054] This embodiment provides a robust beam tracking method for U6G ultra-large-scale MIMO mixed fields, such as... Figure 2 As shown, the robust beam tracking method in this embodiment considers the U6G band ultra-large-scale MIMO uplink system, with carrier frequency... , wavelength is Assume the BS is equipped with one Uniform linear array (ULA) of antenna elements, with an antenna spacing of half a wavelength. The user terminal is equipped with a single antenna. The channel includes... There are multiple path components, one of which is a Loss-of-Support (LoS) path. NLoS path.
[0055] like Figure 1 As shown, the robust beam tracking method includes the following steps:
[0056] S1: Based on the random motion model of the user and the scatterer in two-dimensional space, a hybrid field time-varying channel containing both far-field and near-field path components is generated, and a dynamic hybrid field channel model is established.
[0057] The dynamic hybrid field channel model is expressed as follows:
[0058]
[0059] Channel vector modeling is a superposition of far-field path components and near-field path components: the far-field path adopts a planar wavefront steering vector. The description states that the near-field path uses a spherical wavefront steering vector. Description: The user and the scatterer move with random velocity and direction in a two-dimensional plane, and their positions are updated as follows. ,in For speed, It is a unit vector in direction. The sampling interval is given by the scatterer or user at time [time]. Location and the position of the base station reference antenna Positional changes are directly mapped to temporal variations in channel angle and distance parameters through geometric relationships. ; at every moment The user sends a pilot symbol, and the base station receives the received signal vector. Perform a DFT on the received signal to obtain the time. Angular spectrum.
[0060] S2: Based on the dynamic hybrid field channel model, training data is constructed and preprocessed using a data generation strategy based on the geometric physical model;
[0061] The method for constructing training data based on the data generation strategy of geometric physics model includes: simulating the random motion of users and scatterers in two-dimensional space, calculating the channel angle, distance and gain parameters according to the real-time position, and then generating the channel vector by the steering vector model; randomly and uniformly sampling the signal-to-noise ratio over a wide range to enhance generalization ability; the motion parameters of each path are independent of each other, and the training data naturally includes field switching events; no minimum spacing constraint is imposed on the angle parameters of the four paths, and the training set covers the scene of overlapping angle spectra.
[0062] The generation of training data includes: generating time-varying channel trajectories of mixed fields through a physical model, dividing training samples using a sliding window mechanism, with each sample containing the angle spectrum of the previous and current time moments and the optimal beam vector of the previous time moment as input, and using the optimal beam vector of the current time moment as the label.
[0063] like Figure 3 As shown, multiple complete user motion trajectories are generated through a physical model, and each trajectory contains... The continuous channel evolution process at each time step is described. A sliding window mechanism is used to divide the continuous trajectory into training samples, with a window step size of 1. For time step... ,enter The output label is composed of the angle spectrum at the previous time step, the angle spectrum at the current time step, and the optimal beam vector at the previous time step. MRC optimal beam vector calculated based on perfect channel state information (CSI) A length of The trajectory can be generated Training samples.
[0064] Data preprocessing includes: normalizing the angle spectrum to accelerate network training convergence and improve numerical stability. This eliminates the absolute amplitude difference under different signal-to-noise ratios and channel gain conditions; complex data is represented by real and imaginary parts to avoid discontinuities in phase transitions; the normalized angle spectrum is divided into 8 characteristic channels, namely the normalized amplitude, real part, and imaginary part of the angle spectrum at the previous moment, the normalized amplitude, real part, and imaginary part of the angle spectrum at the current moment, and the real and imaginary parts of the beam vector at the previous moment.
[0065] S3: Extract angular spectral space features using a well-designed CNN feature extraction module;
[0066] like Figure 4 As shown, the CNN feature extraction module includes an Initial Conv1D layer and multiple cascaded ResBlocks. Each ResBlock contains two layers of one-dimensional convolution, batch normalization, GELU activation, and residual connections. The sequence length is halved step by step through convolutions with a downsampling stride of 2. The residual connections ensure that global spectral information can be obtained under low signal-to-noise ratio conditions.
[0067] The operation of the CNN feature extraction module includes:
[0068] The CNN feature extraction module first uses the Initial Conv1D layer to extract the feature. The number of input channels has been expanded to Each channel, kernel size The sequence length is maintained. Unchanged; then connected A cascaded ResBlock, via After one ResBlock, the output feature map dimension is ,in This represents the final number of channels. The length of the compressed sequence; the feature map output by the CNN is flattened and then passed through a linear projection layer. and bias vector Mapped to The embedding vector is obtained by applying layer normalization to the 3D space. :
[0069]
[0070] in, Let be the projection weight matrix. This is the bias vector.
[0071] S4: Based on the angular spectrum space characteristics, the temporal dependencies are captured by the well-designed Transformer temporal coding module, and the coding features are output;
[0072] like Figure 4 As shown, the Transformer temporal coding module is composed of multiple identical coding layers stacked together. Each coding layer contains a multi-head self-attention sub-layer and a feedforward network sub-layer.
[0073] In this embodiment, the embedding vectors of each time slot in the sequence are... Stacked as a sequence embedding matrix Add learnable positional coding Later obtained Positional encoding enables the model to distinguish and The temporal sequence relationship. The embedding matrix after adding positional encoding. Sent to The Transformer encoder consists of stacked identical coding layers, each containing two core sub-layers: a Multi-Head Self-Attention (MHSA) sub-layer and a Feed-Forward Network (FFN) sub-layer. Each sub-layer is followed by a residual connection and a LayerNorm. In the self-attention computation, the query, key, and value matrices are all... We obtain this through linear projection: , , Self-attention mechanisms utilize computation. and The similarity between them determines how much effort should be invested in specific historical states for current beam prediction, and utilizes the corresponding Weighted fusion is performed. A single attention function may not be able to capture information from different representation subspaces.
[0074] The multi-head self-attention mechanism maps the input to multiple subspaces through multiple sets of different projection matrices, calculates attention separately, and then concatenates and fuses the results. Different attention heads can capture motion features of different path components.
[0075]
[0076]
[0077] in For the number of attention heads, , , Here are the projection matrices for each head. The output projection matrix is used; multi-head attention allows the model to simultaneously focus on information from different subspaces. For example, in beam tracking tasks, different attention heads can capture the motion features of different path components. The self-attention mechanism adaptively determines which information to extract and fuse from historical and current time slots through data-driven attention weights, without making any prior assumptions about the functional form of channel evolution, and is naturally adapted to nonlinear dynamic changes, including field switching.
[0078] The feedforward network in each coding layer consists of two linear transformations and a GELU activation function.
[0079]
[0080] in, , This is the weight matrix. The hidden layer dimension of the feedforward network, and The bias vector is used; the role of the feedforward sub-layer is to perform nonlinear transformation and feature purification on the features captured by the multi-head attention sub-layer, amplify the feature signals that are useful for prediction, and suppress useless noise features.
[0081] The output of each sublayer is processed by residual concatenation and LayerNorm.
[0082]
[0083] The Sublayer is either an MHSA or FFN sublayer; LayerNorm normalizes the feature dimensions of each sample.
[0084]
[0085] in and These are the mean and variance along the feature dimension, respectively; and For learnable scaling and offset parameters, It is the numerical stability constant;
[0086] After processing by the Transformer encoder, the output of the last time slot in the sequence (i.e., the position corresponding to the current time t) is taken as the final encoded feature representation, with a dimension of . .
[0087] S5: Based on the coding characteristics, the predictor outputs the optimal composite beam vector that integrates all multipath components.
[0088] The prediction head will output the Transformer encoder's output. The dimensional encoded features are mapped to the beam vector space and consist of two fully connected layers: the first layer will... Dimensional encoded feature mapping to The second layer will use GELU activation. Dimension mapping to Dimension; Output 3D real vector The former Each component and the last Each component is used as the real and imaginary parts of the complex beam vector, respectively, and reconstructed as... Complex beam vector;
[0089]
[0090] in and These represent the first and second halves of the prediction head output, respectively; ultimately, for Normalization is performed to meet power constraints.
[0091]
[0092] The optimal beam vector is physically equivalent to the normalization of the channel vector (MRC solution), which comprehensively considers the optimal coherent combining after weighting the complex gain of all paths, without the need to explicitly estimate the physical parameters of each path.
[0093] S6: For training the network model consisting of the CNN feature extraction module, the Transformer temporal coding module, and the prediction head, the loss function adopts a cosine similarity-based form, which is invariant to the common phase factor between complex vectors; the training uses the Adam optimizer in conjunction with a cosine annealing restart scheduler to dynamically adjust the learning rate, specifically including:
[0094] All generated mixed-field channel trajectory data were randomly divided into training, validation, and test sets in a 6:2:2 ratio, ensuring no trajectory overlap between the three subsets and guaranteeing the objectivity of the evaluation results. The training set was used for optimizing and updating network parameters, the validation set for early stopping detection and hyperparameter selection during training, and the test set for final performance evaluation. The loss function adopted was based on cosine similarity.
[0095]
[0096] in To take the modulus of a complex number. When and When fully aligned, When the two are completely orthogonal, This loss function is invariant to the common phase factor between complex vectors, making it suitable for beam tracking tasks that only focus on orientation alignment rather than absolute phase.
[0097] To accelerate network training convergence and avoid getting trapped in local optima, cosine annealing is used to restart the scheduler and dynamically adjust the learning rate. The core formula is:
[0098]
[0099] in The current learning rate, This is the current round number. This represents the total number of rounds in the current cycle. and These represent the upper and lower bounds of the learning rate, respectively. Training employs the Adam optimizer, coupled with a cosine annealing restart scheduler to dynamically adjust the learning rate. At the end of each epoch, the learning rate is approached to the minimum for refined searching, and the scheduler restarts at the beginning of the next epoch to escape local optima. The validation set loss is monitored during training, and training terminates when the validation loss no longer decreases over several consecutive epochs.
[0100] Example 2:
[0101] To verify the effectiveness of the method of the present invention, the following simulation comparison experiments and data analysis were conducted in this embodiment:
[0102] In this embodiment, achievable rate is used as a performance indicator. The formula for calculating achievable rate is as follows:
[0103]
[0104] in It is the beam vector predicted by the model. It is a mixed field channel vector. For transmission power, This represents noise power.
[0105] The simulation parameters are set as follows: Number of base station antennas carrier frequency Antenna spacing half wavelength, number of paths (1 Loss path, 3 NLoS paths), user movement speed range 1 to 6 m / s, training data sampling interval The movement distance ranges from 50 to 300 meters. The training dataset consists of... The trajectory consists of several user movement tracks, each track being [length missing]. The dataset is divided into training, validation, and test sets in a 6:2:2 ratio, with each set being independent at the trajectory level to ensure the effectiveness of the test set in generalizing to the training set.
[0106] The method of this invention is compared with the perfect CSI beamforming method, the MRC beamforming method based on least squares channel estimation, the beam tracking method based on deep neural network (DNN), and the near-field beam tracking (LNBT) method based on long short-term memory network.
[0107] Figure 5The achievable rate of each method is shown as a function of SNR. In the -5dB to 0dB range, the achievable rate of the method described in this invention reaches approximately 90% of that based on perfect CSI beamforming, verifying its beam alignment accuracy in low SNR environments. Within this range, the DNN method gradually catches up with the method described in this invention because the single-slot angle spectrum is already relatively clear under high SNR, and preceding time-series information introduces interference. In the -15dB to -5dB range, the method described in this invention significantly outperforms both DNN and LNBT. This advantage is mainly due to the robust feature extraction of the noisy angle spectrum by the CNN module, which effectively suppresses noise perturbations, and the Transformer's use of the self-attention mechanism of historical sequences to compensate for the decrease in single-time-series observation quality. In the -20dB to -15dB range, due to severe noise overwhelming the signal, all deep learning methods converge towards 1 bps / Hz. The performance of the method described in this invention and the LNBT method is similar in this range, but still significantly better than the DNN method. This indicates that at extremely low SNR, single-time spatial information is severely degraded, but methods with temporal modeling capabilities significantly outperform DNNs without temporal modeling. The MRC beamforming method based on least-squares channel estimation achieves extremely low rates across the entire SNR range, experimentally confirming the failure of traditional channel estimation methods at low SNR.
[0108] Figure 6 The achievable rate of different methods varies with the number of antennas, with the SNR fixed at 0 dB. As the number of antennas increases from 128 to 512, the achievable rate of the method in this invention approaches 90% of that based on the perfect CSI beamforming method, indicating that the method in this invention has good scalability for very large-scale arrays. The growth trend tends to level off as the number of antennas gradually increases; the performance loss mainly stems from the increased dimension of the network output beam vector as the number of antennas increases, within the model capacity... While remaining constant, the information compression ratio of the output layer increases accordingly. The performance of the DNN method is lower than that of the method in this invention, mainly because it lacks temporal modeling capabilities and can only rely on single-time observations for each tracking cycle. As the number of antennas increases, the achievable rate difference between the method in this invention and the DNN method gradually narrows. This is because as the array size increases, the equivalent signal-to-noise ratio improvement brought about by the array gain makes the signal components of the angular spectrum clearer, and at this time, a relatively accurate beam can be obtained by relying only on single-time-slot spatial features. The achievable rate growth of the LNBT method is significantly lower than that of other deep learning schemes because the codewords it uses are independent of the number of antennas. The achievable rate of the MRC beamforming method based on least squares channel estimation remains at an extremely low level under all antenna numbers, further illustrating that traditional channel estimation is completely ineffective under low signal-to-noise ratios, and the spatial resolution improvement brought by the array cannot be effectively utilized.
Claims
1. A robust beam tracking method for U6G ultra-massive MIMO hybrid field, characterized in that, Includes the following steps: S1: Based on the random motion model of the user and the scatterer in two-dimensional space, a hybrid field time-varying channel containing both far-field and near-field path components is generated, and a dynamic hybrid field channel model is established. S2: Based on the dynamic hybrid field channel model, training data is constructed and preprocessed using a data generation strategy based on the geometric physical model; S3: Extract angular spectral space features using a well-designed CNN feature extraction module; S4: Based on the angular spectrum space characteristics, the temporal dependencies are captured by the well-designed Transformer temporal coding module, and the coding features are output; S5: Based on the coding characteristics, the predictor outputs the optimal composite beam vector that integrates all multipath components.
2. The robust beam tracking method for U6G ultra-massive MIMO hybrid field according to claim 1, characterized in that, The dynamic hybrid field channel model in step S1 is expressed as follows: ; Channel vector modeling is a superposition of far-field path components and near-field path components: the far-field path adopts a planar wavefront steering vector. The description states that the near-field path uses a spherical wavefront steering vector. Description: The user and the scatterer move with random velocity and direction in a two-dimensional plane, and their positions are updated as follows. ,in For speed, It is a unit vector in direction. The sampling interval is given by the scatterer or user at time [time]. Location and the position of the base station reference antenna Positional changes are directly mapped to temporal variations in channel angle and distance parameters through geometric relationships. ; at every moment The user sends a pilot symbol, and the base station receives the received signal vector. Perform a DFT on the received signal to obtain the time. Angular spectrum.
3. A robust beam tracking method for U6G ultra-large-scale MIMO mixed fields according to claim 2, characterized in that, The method for constructing training data based on the data generation strategy of the geometric physical model in step S2 includes: simulating the random motion of the user and the scatterer in two-dimensional space, calculating the channel angle, distance and gain parameters according to the real-time position, and then generating the channel vector by the steering vector model; randomly and uniformly sampling the signal-to-noise ratio over a wide range to enhance generalization ability; the motion parameters of each path are independent of each other, and the training data naturally includes field switching events; no minimum spacing constraint is imposed on the angle parameters of the four paths, and the training set covers the scene of overlapping angle spectra.
4. The robust beam tracking method for U6G ultra-massive MIMO hybrid field according to claim 3, characterized in that, The generation of training data in step S2 includes: generating a mixed-field time-varying channel trajectory through a physical model, dividing the training samples using a sliding window mechanism, with each sample containing the angle spectrum of the previous time and the current time and the optimal beam vector of the previous time as input, and using the optimal beam vector of the current time as the label.
5. The robust beam tracking method for U6G ultra-massive MIMO hybrid field according to claim 4, characterized in that, The data preprocessing in step S2 includes: normalizing the angle spectrum to eliminate absolute amplitude differences; representing complex data with real and imaginary parts to avoid discontinuities in phase transitions; and dividing the normalized angle spectrum into 8 feature channels, namely the normalized amplitude, real part, and imaginary part of the angle spectrum at the previous moment, the normalized amplitude, real part, and imaginary part of the angle spectrum at the current moment, and the real and imaginary parts of the beam vector at the previous moment.
6. The robust beam tracking method for U6G ultra-massive MIMO hybrid field according to claim 5, characterized in that, The CNN feature extraction module in step S3 includes an Initial Conv1D layer and multiple cascaded ResBlocks. Each ResBlock contains two layers of one-dimensional convolution, batch normalization and GELU activation, as well as residual connections. The sequence length is halved step by step through convolution with a downsampling stride of 2.
7. The robust beam tracking method for U6G ultra-massive MIMO hybrid field according to claim 6, characterized in that, The operation of the CNN feature extraction module in step S3 includes: The CNN feature extraction module first uses the Initial Conv1D layer to extract features. The number of input channels has been expanded to Each channel, kernel size The sequence length is maintained. Unchanged; then connected A cascaded ResBlock, via After one ResBlock, the output feature map dimension is ,in This represents the final number of channels. The length of the compressed sequence; the feature map output by the CNN is flattened and then passed through a linear projection layer. and bias vector Mapped to The embedding vector is obtained by applying layer normalization to the dimensional space. .
8. The robust beam tracking method for U6G ultra-massive MIMO hybrid field according to claim 7, characterized in that, In step S4, the Transformer temporal coding module is composed of multiple identical coding layers stacked together. Each coding layer contains a multi-head self-attention sub-layer and a feedforward network sub-layer. The multi-head self-attention mechanism maps the input to multiple subspaces through multiple different projection matrices, calculates attention separately for each subspace, and then concatenates and fuses the results. Different attention heads can capture motion features of different path components. ; ; wherein is the number of attention heads, , , is the projection matrix for each head, is the output projection matrix; The feedforward network in each coding layer consists of two linear transformations and a GELU activation function. ; wherein, , is a weight matrix, is a hidden layer dimension of the feedforward network, and is a bias vector; The output of each sub-layer is processed by residual connection and LayerNorm: ; The Sublayer is either an MHSA or FFN sublayer; LayerNorm normalizes the feature dimensions of each sample: ; in and These are the mean and variance along the feature dimension, respectively; and For learnable scaling and offset parameters, It is the numerical stability constant; After processing by the Transformer encoder, the output of the last time slot in the sequence is taken as the final encoded feature representation, with a dimension of . .
9. The robust beam tracking method for U6G ultra-massive MIMO hybrid field according to claim 8, characterized in that, The process of the predictor outputting the optimal composite beam vector that integrates all multipath components in step S5 includes: The prediction head will output the Transformer encoder's output. The dimensional encoded features are mapped to the beam vector space, consisting of two fully connected layers: the first layer will... Dimensional encoded feature mapping to The second layer will use GELU activation. Dimension mapping to Dimension; Output 3D real vector The former Each component and the last Each component is used as the real and imaginary parts of the complex beam vector, respectively, and reconstructed as... Complex beam vector; ; wherein and are the first and second halves of the prediction head output, respectively; and is normalized to satisfy the power constraint 。