Wireless channel prediction method based on Mama architecture
Through the channel prediction model based on the Mamba architecture, the accuracy and computing complexity problems of wireless channel prediction in high-speed mobile and high-frequency band communication scenarios are solved, and efficient and real-time channel prediction is achieved, which is suitable for computing-constrained terminal devices and complex communication scenarios.
Patent Information
- Application Number
- CN202510564337.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-05
AI Technical Summary
The existing wireless channel prediction methods have problems such as low accuracy, high computational complexity, difficulty in dealing with long sequences and high deployment costs in high-speed mobile or high frequency band communication scenarios, which limits their application in future wireless communication systems.
The channel prediction model based on the Mamba architecture is adopted to model the wireless communication system through a uniform planar array, combining dual-path processing of frequency domain and time domain information, parameter sharing within the Mamba block, and bidirectional Mamba structure, a channel prediction model is built, which significantly reduces the computing volume and memory requirements, and is suitable for terminal devices with limited computing power and power consumption.
It improves the accuracy and real-time performance of channel prediction, can handle inputs from long historical channel sequences and large-scale systems, reduces inference time, enhances the scalability and deployment convenience of the model, and adapts to the needs of complex communication scenarios.
Smart Images

Figure CN120433872A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless communications, and in particular to a wireless channel prediction method. Background Art
[0002] With the development of next-generation wireless communication systems (such as 6G), the demand for key performance indicators such as ultra-high speed and ultra-low latency is increasing, making it crucial to model and control wireless channels with higher precision. As a core means to improve system capacity and reliability, the performance of multiple-input multiple-output technology is highly dependent on accurate and timely channel state information (CSI). However, in high-speed mobile or high-frequency communication scenarios, the channel exhibits rapid time-varying characteristics, resulting in high overhead and significant feedback delay problems for traditional pilot-based CSI estimation methods, which seriously restricts the performance of key technologies such as beamforming and resource allocation. Therefore, predicting future channel states based on historical channel information has become an important research direction to overcome channel aging and improve system performance.
[0003] Deep learning methods have shown great potential in solving wireless channel prediction problems. Compared to traditional statistical model-based approaches, deep learning models can automatically learn the complex nonlinear characteristics of channel data and achieve significant improvements in prediction accuracy and robustness. For example, models based on recurrent neural networks can effectively capture the time series dependencies of channels. However, traditional deep learning models struggle to fully exploit the deep patterns inherent in complex wireless channel data, and their sequential processing nature limits the parallelism of the training process.
[0004] In recent years, Transformer has been widely used in channel prediction tasks due to its excellent performance in capturing long-range dependencies of sequences through its self-attention mechanism. However, the self-attention mechanism at the core of the Transformer architecture has an O(N 2 ) severely restricts its ability to process long sequence data and limits its feasibility in high-dimensional application scenarios such as large-scale multi-antenna systems. At the same time, the resulting high inference latency and training cost problems further hinder its deployment and application in actual production environments. Faced with the inherent limitations of the Transformer architecture, the research community has been actively exploring more efficient sequence modeling paradigms. Although a variety of Transformer variants have emerged to reduce complexity (such as models based on sparse or linear attention), these methods often come at the expense of some modeling performance.
[0005] The invention patent application number 202110788681.2 provides a channel prediction method and system for large-scale MIMO systems based on joint time-frequency correlation. Based on the weak time-domain correlation and strong frequency-domain correlation characteristics of measured data, a time-frequency combined channel prediction method based on a convolutional long short-term memory network is proposed. Convolutional LSTM is a deep learning model that can simultaneously extract time-domain and frequency-domain features. It extracts the frequency-domain characteristics of the channel through the input convolution structure and uses the internal LSTM structure to extract the time-domain characteristics of the channel. The channel characteristics in the frequency domain are applied to the channel prediction in the time domain, thereby achieving the effect of joint time-frequency channel prediction. This method combines the characteristics of the time and frequency domains to predict the channel, and uses the strong autocorrelation in the frequency domain to improve the accuracy of time-domain prediction. Compared with existing methods that only use time-domain correlation for channel prediction, it has higher channel prediction accuracy. However, it has the problem of difficulty in processing long sequences.
[0006] Therefore, to overcome the problems of low channel prediction accuracy, high computational complexity, difficulty in processing long sequences, and high deployment costs, a new technical solution is urgently needed that can significantly improve computational efficiency and model scalability while ensuring prediction accuracy, so as to better meet the real-time and resource efficiency requirements of future wireless communication systems. Summary of the Invention
[0007] In response to the technical problems of existing channel prediction methods such as low precision, high computational complexity, difficulty in processing long sequences and high deployment cost, the present invention proposes a wireless channel prediction method based on the Mamba architecture, applies the Mamba architecture to the field of channel prediction, and proposes a channel prediction model based on the Mamba architecture. The method can not only efficiently process inputs containing longer historical information or from systems with a large number of antennas and subcarriers, but also significantly reduce the amount of computation and memory requirements, making it easier to deploy on terminal devices or edge nodes with limited computing power and power consumption, and has high prediction accuracy.
[0008] In order to achieve the above object, the technical solution of the present invention is achieved as follows:
[0009] A wireless channel prediction method based on the Mamba architecture includes the following steps:
[0010] S1. A uniform planar array is provided at a base station end of a wireless communication system, and a channel model of the wireless communication system is modeled by the uniform planar array; based on the channel model, a data representation of the wireless communication system is obtained;
[0011] S2. Based on the channel model and data representation of the wireless communication system, a historical CSI sequence is obtained; a channel prediction model is constructed and trained, and the historical CSI sequence is input into the trained channel prediction model as input data to predict the wireless channel state to be detected, thereby obtaining a channel prediction result;
[0012] The channel prediction model includes an embedding module, a Mamba-based channel prediction core module, and an output layer, which are connected in sequence.
[0013] Furthermore, the uniform planar array is regularly arranged on a two-dimensional plane, and for a frequency f c , there are K orthogonal subcarriers; when the uniform planar array UPA is located in the xz plane, including N t antenna elements, with N rows x , the number of columns is N z , the row spacing is d x , the column spacing is d z ;
[0014] The uniform planar array describes the response of a signal from a specific direction (φ, θ) through the array steering vector a(φ, θ), which is expressed as:
[0015]
[0016] Where φ is the azimuth angle, θ is the elevation angle, and m is the two-dimensional index of the antenna element (m x ,m z ), where λ is the carrier wavelength.
[0017] Furthermore, the process of modeling the channel model of the wireless communication system by using a uniform planar array includes: for time t, assuming that the channel is a superposition of Q discrete propagation paths, and the qth path has a time-varying complex gain α q [t], Doppler frequency shift f q , delay τ q and departure angle (φ q ,θ q ), then the frequency domain CSI matrix observed at time t is expressed as:
[0018]
[0019] Among them, T step is the time step, f(τ q ) is the frequency response vector, a H (φ q ,θ q ) is the steering vector of the spatial response.
[0020] Furthermore, the acquisition of data representation of the wireless communication system includes: the frequency domain CSI is generated by the time domain impulse response corresponding to the frequency domain CSI, and the time domain CIR vector Among them, L CIR is the maximum number of delay taps;
[0021] The frequency domain CSI is obtained by performing a K-point discrete Fourier transform on the time domain CIR of each antenna, which is expressed as: in, is a partial DFT matrix;
[0022] The estimated values of CSI are:
[0023]
[0024] in, is the estimated value of CSI, W[t] is the additive white Gaussian noise vector;
[0025] Construct a three-dimensional input tensor Where P represents the length of the historical observation window, and the input tensor X captures the temporal evolution, frequency distribution, and spatial structure of the CSI;
[0026] For a given channel state sequence of P moments in a historical observation window, the input tensor is represented as:
[0027] X={H[t-P+1],H[t-P+2],...,H[t]};
[0028] Predict the channel state Y = {H[t+1], H[t+2], ..., H[t+L]} at the next L time moments.
[0029] Furthermore, the step of inputting the historical CSI sequence as input data into the trained channel prediction model to predict the wireless channel state to be detected to obtain the channel prediction result includes:
[0030] S2.1. Convert the complex historical CSI sequence into real representations in the frequency and time domains through the embedding module and perform independent embedding to obtain frequency-domain embedding vectors and time-domain embedding vectors.
[0031] S2.2. Input the frequency domain embedding vector and the time domain embedding vector into the Mamba-based channel prediction core module to obtain enhanced features;
[0032] S2.3. Map the enhanced features into predictions of channel state information for one or more future time steps through the output layer.
[0033] Furthermore, the step of obtaining the frequency domain embedding vector and the time domain embedding vector includes:
[0034] The frequency domain CSI in the complex frequency domain CSI sequence is converted to a real number representation to obtain a real number frequency domain CSI sequence. At the same time, a K-point inverse discrete Fourier transform is applied along the subcarrier dimension to the frequency domain CSI to generate the time domain impulse response corresponding to the frequency domain CSI. This is converted to a real number representation to obtain a real number time domain impulse response sequence.
[0035] The real-number frequency-domain CSI sequence and the time-domain impulse response sequence form a multidimensional CSI tensor. The multidimensional CSI tensor at each time t is flattened along the subcarrier and antenna dimensions and combined across all time steps to generate a frequency-domain vector and a time-domain vector.
[0036] Then we enter the independent embedding part, using two independent linear embedding layers to process the frequency domain and time domain vectors respectively, and apply Dropout, and finally convert the frequency domain vector and time domain vector into frequency domain embedding vector and time domain embedding vector respectively.
[0037] Furthermore, the Mamba-based channel prediction core module includes: a shared time dependency Mamba layer, a concat layer, a first linear projection layer, a bidirectional Mamba block, a second linear projection layer and a lightweight attention mechanism module connected in sequence.
[0038] Furthermore, the method of inputting the frequency domain embedding vector and the time domain embedding vector into the Mamba-based channel prediction core module to obtain enhanced features includes: the frequency domain embedding vector and the time domain embedding vector enter a shared time-dependent Mamba layer, and the shared time-dependent Mamba layer includes: a first normalization layer, a first Mamba block, and a first Dropout layer connected in sequence;
[0039] In the shared time-dependent Mamba layer, the frequency domain embedding vector and the time domain embedding vector are normalized by the shared first normalization layer, and then the frequency domain original input and the time domain original input are obtained and passed through the first Mamba block together;
[0040] After being processed by the first Mamba block, the frequency domain output features and the time domain output features are obtained. After being processed by the first Dropout layer, they are added to the original frequency domain input and the original time domain input respectively to obtain the frequency domain features and time domain features that have encoded time dependencies.
[0041] The frequency domain features and time domain features are connected through the concat layer, and the projected frequency domain features and projected time domain features are obtained through the first linear projection layer and sent to the bidirectional Mamba block;
[0042] In the bidirectional Mamba block, layer normalization is first performed in the second normalization layer to obtain fused features. After the fused features are transposed, they are processed using a bidirectional Mamba structure consisting of an independent forward and a reverse Mamba block to capture global feature dimension dependencies. The output of the bidirectional Mamba structure is further transposed and sent to the second Dropout layer, where it is combined with the fused features through a residual connection to obtain residual connection features. The residual connection features are then processed in sequence through the third normalization layer and the feedforward neural network layer to obtain nonlinear features.
[0043] The nonlinear features are projected through the second linear projection layer to obtain the nonlinear features, which are then sent to the lightweight attention mechanism module for processing. The lightweight attention mechanism module outputs enhanced features.
[0044] Furthermore, the step of mapping the enhanced features into a prediction of channel state information for one or more future time steps through the output layer includes:
[0045] In the output layer, the enhanced features are first normalized by the fourth normalization layer, and the feature dimensions are projected back to the original flattened dimensions by the FFN layer, and the output feature matrix O = FFN (LayerNorm (H final ));
[0046] Among them, LayerNorm represents layer normalization processing, FFN(·) represents the feedforward neural network layer;
[0047] To get an L-step forecast, we map the time dimension from P to L and get the forecast sequence:
[0048]
[0049] in, is the channel prediction sequence of the channel prediction model for the next L moments, Transpose(·) is the transposition operation, and Linear(·) is the linear layer operation;
[0050] Then the prediction sequence Perform the reverse dimension reshaping operation to transform the predicted sequence into Convert to complex form and finally get the predicted output sequence is the predicted frequency domain CSI.
[0051] Furthermore, when training the channel prediction model, the loss function between the real channel and the predicted channel is minimized, and the normalized mean square error is used as the objective function, which is expressed as conditional probability modeling:
[0052]
[0053] Among them, NMSE is the normalized mean square error, It is the expectation of the joint distribution of historical observations and future true channels, is the Frobenius norm squared of the matrix, i is the summation index, and f(X) is the mapping function of the input tensor X.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] 1. The present invention applies the Mamba architecture to the field of channel prediction. Based on channel modeling and data-driven, a channel prediction model based on the Mamba architecture is proposed. When processing long historical channel sequences or a large number of antennas and subcarriers, the computational complexity and memory requirements are significantly reduced, making the channel prediction model easier to deploy on terminal devices or edge nodes with limited computing power and power consumption, thereby enhancing the practicality and economy of the technology.
[0056] 2. This invention significantly reduces the inference time required for channel prediction, improving the real-time nature of prediction and better meeting the demand for immediate channel information in dynamic scenarios such as high-speed mobility. This enables the channel prediction model to efficiently process inputs containing longer histories or from systems with a large number of antennas and subcarriers, providing excellent scalability and adaptability to the demands of more complex communication scenarios in the future.
[0057] 3. Through a specially designed network structure, including dual-path processing of frequency and time domain information, parameter sharing within Mamba blocks, and modeling of correlations between feature dimensions using a bidirectional Mamba structure, this invention can more comprehensively and deeply capture the complex joint dynamic characteristics of the channel in the time, frequency, and space dimensions. This invention achieves high prediction accuracy, providing a more accurate and reliable channel information foundation for subsequent key wireless technologies such as beamforming and resource allocation. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0059] Figure 1 Flow chart of the method of the present invention.
[0060] Figure 2 FIG. 4 is a diagram of a channel model architecture of a wireless communication system in an embodiment of the present invention.
[0061] Figure 3This is a diagram of the channel prediction model architecture of the present invention. DETAILED DESCRIPTION
[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.
[0063] A wireless channel prediction method based on Mamba architecture, such as Figure 1 As shown, the following steps are included:
[0064] S1. Equip the base station of the wireless communication system with a uniform planar array and use it to model the channel model of the wireless communication system. Based on the channel model, obtain the data representation of the wireless communication system and provide structured input data containing time-space-frequency (time, space, frequency) information for the subsequent channel prediction model.
[0065] The uniform planar array (UPA) equipped at the base station is regularly arranged on a two-dimensional plane. In this embodiment, for frequency f c , there are K orthogonal subcarriers. The uniform planar array UPA is located in the xz plane, including N t antenna elements, with N rows x , the number of columns is N z , the row spacing is d x , the column spacing is d z .
[0066] The uniform planar array (UPA) enables the base station to distinguish the departure angle of the signal, which includes the azimuth angle φ∈[-π,π) and the elevation angle The response of a uniform planar array (UPA) to a signal from a specific direction (φ, θ) is described by the array steering vector a(φ, θ), which is expressed as:
[0067]
[0068] Where m is the two-dimensional index of the antenna element (m x ,m z ), where λ is the carrier wavelength.
[0069] like Figure 2 As shown in Figure 1, the process of defining the channel statistics, propagation path, antenna array response, and noise impact and modeling the channel model of a wireless communication system using a uniform planar array includes:
[0070] For time t, based on the widely used geometric stochastic channel model (GSCM), the channel is assumed to be a superposition of Q discrete propagation paths, with the qth path having a time-varying complex gain α q [t], Doppler frequency shift f q , delay τ q and departure angle (φ q ,θ q ), therefore, the frequency domain CSI matrix observed at time t Expressed as:
[0071]
[0072] Among them, through The term reflects the Doppler effect, T step is the time step, f(τ q ) is the frequency response vector, a H (φ q ,θ q ) is the steering vector of the spatial response.
[0073] Obtaining data representation of a wireless communication system includes:
[0074] The user terminal (UE) moves within the time τ, which causes the propagation path between the user terminal and the base station to change. In actual broadband communication systems, due to multipath propagation, the channel has a sparse impulse response structure. The frequency domain CSI is generated by the time domain impulse response (CIR) corresponding to the frequency domain CSI. The time domain CIR vector Among them, L CIR is the maximum number of time-delay taps, satisfying the wide-sense stationary uncorrelated scattering (WSSUS) assumption.
[0075] Therefore, the frequency domain CSI is obtained by performing a K-point discrete Fourier transform (DFT) on the time domain CIR of each antenna, which is expressed as: in, is a partial DFT matrix.
[0076] In a real system, CSI is obtained by pilot estimation, and there is noise interference. Therefore, the estimated value of CSI for:
[0077]
[0078] in, is an additive white Gaussian noise vector, each element of which independently obeys a complex Gaussian distribution. Represents the noise power, which is determined by the receiver thermal noise and quantization error.
[0079] To facilitate deep learning modeling, construct a three-dimensional input tensor Here, P represents the length of the historical observation window, and the input tensor X captures the temporal evolution, frequency distribution, and spatial structure of the CSI.
[0080] Then for a given channel state sequence of P moments in the historical observation window, the input tensor is expressed as:
[0081] X={H[t-P+1],H[t-P+2],...,H[t]};
[0082] The channel prediction task aims to establish the mapping function f: Predict the channel state Y = {H[t+1], H[t+2], ..., H[t+L]} at the next L time moments.
[0083] S2. Based on the channel model and data representation of the wireless communication system, obtain the historical CSI sequence; construct Figure 3 The channel prediction model shown is trained, and the historical CSI sequence is input as input data into the trained channel prediction model to predict the wireless channel state to be detected to obtain a channel prediction result.
[0084] The channel prediction model includes an embedding module, a Mamba-based channel prediction core module, and an output layer connected in sequence. The historical CSI sequence is in plural form.
[0085] S2.1. To adapt the original channel data to the Mamba model and fully utilize the characteristics of the Mamba model in different domains, the complex historical CSI sequence is converted into real number representations in the frequency and time domains through an embedding module. Independent embedding is performed to obtain frequency domain embedding vectors and time domain embedding vectors. This enables the channel prediction model to learn and distinguish feature representations specific to the frequency and time domains at an early stage.
[0086] According to the data representation of the wireless communication system in step S1, the frequency domain CSI sequence X of the historical observation window P moments is taken as input. In the embedding module, the frequency domain CSIH[t′] in the complex form of the frequency domain CSI sequence X is converted into a real number representation to obtain the real form of the frequency domain CSI sequence Where t′ is any specific discrete time step in the historical observation window P.
[0087] At the same time, in order to utilize the time domain (delay domain) information, the K-point inverse discrete Fourier transform (IDFT) is applied along the subcarrier dimension through the frequency domain CSIH[t′] to generate the corresponding time domain impulse response (CIR) h[t′] (assuming that the CIR length is equal to the number of orthogonal subcarriers K, or zero padding / truncation), and converted to real number representation to obtain the real number form of the time domain impulse response sequence
[0088] The frequency domain CSI sequence X in real form f [t′] and the time domain impulse response sequence X d [t′] forms a multi-dimensional CSI tensor. The multi-dimensional CSI tensor of each time t is flattened along the subcarrier and antenna dimensions and all time steps are combined to generate a frequency domain vector suitable for sequence processing. and time domain vector Among them, B represents the batch size.
[0089] Next, we enter the independent embedding part. The independent embedding layer allows the channel prediction model to learn and distinguish domain-specific feature representations at an early stage. The present invention uses two independent linear embedding layers to process frequency domain and time domain vectors respectively, and transforms K×N t ×2-dimensional feature maps to the core hidden dimension d of the channel prediction model model , and applying Dropout will eventually convert the frequency domain vector x f [t′] and the time domain vector x d [t′] is converted into a frequency domain embedding vector and the time domain embedding vector
[0090] S2.2. Input the frequency domain embedding vector and the time domain embedding vector into the Mamba-based channel prediction core module to obtain the enhanced feature H final .
[0091] The Mamba-based channel prediction core module consists of a sequentially connected shared time dependency Mamba layer, a concat layer, a first linear projection layer, a bidirectional Mamba block, a second linear projection layer, and a lightweight attention mechanism module. The shared time dependency Mamba layer extracts temporal dependencies, fuses dual-path information, and processes the fused features using the bidirectional Mamba block to model inter-dimensional correlations. The lightweight attention mechanism module globally refines features in a low-dimensional space.
[0092] The frequency domain embedding vector and the time domain embedding vector enter a shared time-dependent Mamba layer, which includes: a first normalization layer, a first Mamba block, and a first Dropout layer connected in sequence.
[0093] In the shared time dependency Mamba layer, the frequency domain embedding vector E f and the time domain embedding vector E d After normalization through the shared first normalization layer, the original input H in the frequency domain is obtained inf And the original time domain input H ind, to stabilize the input of the subsequent Mamba block, then the feature sequence after normalization in both frequency domain and time domain is the original input in frequency domain H inf And the original time domain input H ind Together through the first Mamba block.
[0094] The feature fusion within the first Mamba block is that the network uses the same Mamba block at each layer to update the representations of the frequency domain and time domain paths respectively. This parameter sharing strategy significantly reduces the number of parameters in the time-dependent modeling part of the channel prediction model. Based on the assumption that the temporal evolution of the channel in different domains shares the underlying dynamics, it encourages the model to learn more general and robust temporal features.
[0095] After processing by the first Mamba block, the frequency domain output feature H is obtained f_TD And the time domain output feature H d_TD , after being processed by the first Dropout layer, they are respectively combined with the original input H in the frequency domain inf And the original time domain input H ind Add together to get the frequency domain feature TD of the encoded time dependency f and time domain features TD d .
[0096] In order to integrate the information extracted by the two paths, the frequency domain feature TD is converted into f and time domain features TD d Connect and obtain the projected frequency domain feature TD′ through the first linear projection layer f And the projected time domain feature TD′ d And sent to the bidirectional Mamba block, in the bidirectional Mamba block, first perform layer normalization in the second normalization layer to obtain the fusion feature Expressed as:
[0097] X fused =LayerNorm(TD′ f +TD′ d ).
[0098] Among them, LayerNorm represents layer normalization processing.
[0099] To model the fusion feature X fused Hidden dimension d at the core model The complex internal dependencies (the core hidden dimension is a mixture of original space, frequency, and delay information) are processed using the idea of Mamba to deal with the correlation between variables. First, the fusion feature X fused Transpose (B,d model,P), then a bidirectional Mamba structure (Bidirectional) is applied for processing, which consists of an independent forward and a reverse Mamba block to capture the global feature dimension dependency, and then the output result is transposed back to (B,d model ,P), then sent to the second Dropout layer and fused with the feature X fused The residual connection feature X is obtained by combining the residual connections r .
[0100] Connect the residual to the feature X r It is processed sequentially through the third normalization layer and the feedforward neural network layer. Each block performs a standard nonlinear transformation to further enhance the representation ability of the channel prediction model. The output is a nonlinear feature.
[0101] Nonlinear characteristic H FFN After passing the second linear projection layer, the projected nonlinear feature H′ is obtained FFN , then the nonlinear feature H′ after projection FFN Send it to the lightweight attention mechanism module.
[0102] In order to further integrate global context information based on the efficiency of Mamba, a lightweight attention module is used as an enhancement. In the lightweight attention mechanism module, the nonlinear feature H′ after projection is FFN First, the projected nonlinear feature H′ is transformed into FFN The dimension of d model Greatly reduced to dimension d tf (d tf <<d model );
[0103] Then, in this low-dimensional space, a standard Transformer Encoder layer is applied. Each Transformer Encoder layer mainly consists of two sub-layers (multi-head self-attention and FFN), and the built-in multi-head self-attention mechanism of the Transformer Encoder layer is used for global feature interaction and refinement.
[0104] Since all attention computations and internal FFN operations are performed in low-dimensional d tf The above is done, and the number of TransformerEncoder layers is L TF Small, lightweight attention mechanism module has a computational complexity of O(P 2 d tf )) is much lower than the original d modelThe O(P 2 d model )) complexity; after L TF After processing by the Transformer Encoder layer, the final enhanced features are obtained
[0105] S2.3, through the output layer to enhance the feature H final The mapping is a prediction of the channel state information for one or more time steps in the future.
[0106] The output layer will enhance the feature H final Converted to channel prediction for the next L moments. The output layer first normalizes the enhanced feature H through the fourth layer final Normalize and then pass the FFN layer to reduce the feature dimension from d tf Project back to the original flattened dimension K×N t ×2, the output feature matrix O is expressed as:
[0107] O=FFN(LayerNorm(H final ));
[0108] Here, FFN(·) represents a feed-forward neural network layer.
[0109] To obtain an L-step prediction, the time dimension needs to be mapped from P to L. By transposing the feature matrix O into (B, K×N t ×2,P), and then apply another linear layer and transpose it back, which is represented as:
[0110]
[0111] in, is the channel prediction model's prediction sequence (flattened form) for the channel at the next L moments, Transpose(·) is the transposition operation, and Linear(·) is the linear layer operation.
[0112] Then the prediction sequence Perform the inverse reshape operation and then transform the predicted sequence into Convert to complex form and finally get the predicted output sequence is the predicted frequency domain CSI.
[0113] The channel prediction model is trained by minimizing the loss function between the true channel and the predicted channel, using the normalized mean square error (NMSE) as the objective function, expressed as conditional probability modeling:
[0114]
[0115] in, It is the expectation of the joint distribution of historical observations and future true channels, is the Frobenius norm squared of the matrix, i is the summation index, and f(X) is the mapping function of the input tensor X.
[0116] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A wireless channel prediction method based on Mamba architecture, characterized in that: The following steps are involved: S1. A uniform planar array is provided at a base station end of a wireless communication system, and a channel model of the wireless communication system is modeled by the uniform planar array; based on the channel model, a data representation of the wireless communication system is obtained; S2. Based on the channel model and data representation of the wireless communication system, a historical CSI sequence is obtained; a channel prediction model is constructed and trained, and the historical CSI sequence is input into the trained channel prediction model as input data to predict the wireless channel state to be detected, thereby obtaining a channel prediction result; The channel prediction model includes an embedding module, a Mamba-based channel prediction core module, and an output layer, which are connected in sequence.
2. The wireless channel prediction method based on Mamba architecture according to claim 1, characterized in that The uniform planar array is regularly arranged on a two-dimensional plane. For the frequency f c , there are K orthogonal subcarriers; when the uniform planar array UPA is located in the xz plane, including N t antenna elements, with N rows x , the number of columns is N z , the row spacing is d x , the column spacing is d z ; The uniform planar array describes the response of a signal from a specific direction (φ, θ) through the array steering vector a(φ, θ), which is expressed as: Where φ is the azimuth angle, θ is the elevation angle, and m is the two-dimensional index of the antenna element (m x ,m z ), where λ is the carrier wavelength.
3. The wireless channel prediction method based on Mamba architecture according to claim 2, characterized in that: The process of modeling the channel model of the wireless communication system by using a uniform planar array includes: for time t, assuming that the channel is a superposition of Q discrete propagation paths, and the qth path has a time-varying complex gain α q [t], Doppler frequency shift f q , delay τ q and departure angle (φ q ,θ q ), then the frequency domain CSI matrix observed at time t is expressed as: Among them, T step is the time step, f(τ q ) is the frequency response vector, a H (φ q ,θ q ) is the steering vector of the spatial response.
4. The wireless channel prediction method based on Mamba architecture according to claim 3, characterized in that: The method of obtaining the data representation of the wireless communication system includes: generating the frequency domain CSI by the time domain impulse response corresponding to the frequency domain CSI, and the time domain CIR vector Among them, L CIR is the maximum number of delay taps; The frequency domain CSI is obtained by performing a K-point discrete Fourier transform on the time domain CIR of each antenna, which is expressed as: in, is a partial DFT matrix; The estimated values of CSI are: in, is the estimated value of CSI, W[t] is the additive white Gaussian noise vector; Construct a three-dimensional input tensor Where P represents the length of the historical observation window, and the input tensor X captures the temporal evolution, frequency distribution, and spatial structure of the CSI; For a given channel state sequence of P moments in a historical observation window, the input tensor is represented as: X={H[t-P+1],H[t-P+2],...,H[t]}; Predict the channel state Y = {H[t+1], H[t+2], ..., H[t+L]} at the next L time moments.
5. The wireless channel prediction method based on Mamba architecture according to claim 4, characterized in that: The step of inputting the historical CSI sequence as input data into the trained channel prediction model to predict the wireless channel state to be detected to obtain the channel prediction result includes: S2.
1. Convert the complex historical CSI sequence into real representations in the frequency and time domains through the embedding module and perform independent embedding to obtain frequency-domain embedding vectors and time-domain embedding vectors. S2.
2. Input the frequency domain embedding vector and the time domain embedding vector into the Mamba-based channel prediction core module to obtain enhanced features; S2.
3. Map the enhanced features into predictions of channel state information for one or more future time steps through the output layer.
6. The wireless channel prediction method based on Mamba architecture according to claim 5, characterized in that: The steps of obtaining the frequency domain embedding vector and the time domain embedding vector include: The frequency domain CSI in the complex frequency domain CSI sequence is converted to a real number representation to obtain a real number frequency domain CSI sequence. At the same time, a K-point inverse discrete Fourier transform is applied along the subcarrier dimension to the frequency domain CSI to generate the time domain impulse response corresponding to the frequency domain CSI. This is converted to a real number representation to obtain a real number time domain impulse response sequence. The real-number frequency-domain CSI sequence and the time-domain impulse response sequence form a multidimensional CSI tensor. The multidimensional CSI tensor at each time t is flattened along the subcarrier and antenna dimensions and combined across all time steps to generate a frequency-domain vector and a time-domain vector. Then we enter the independent embedding part, using two independent linear embedding layers to process the frequency domain and time domain vectors respectively, and apply Dropout, and finally convert the frequency domain vector and time domain vector into frequency domain embedding vector and time domain embedding vector respectively.
7. The wireless channel prediction method based on Mamba architecture according to claim 5, characterized in that: The Mamba-based channel prediction core module includes: a shared time dependency Mamba layer, a concat layer, a first linear projection layer, a bidirectional Mamba block, a second linear projection layer and a lightweight attention mechanism module connected in sequence.
8. The wireless channel prediction method based on Mamba architecture according to claim 7, characterized in that: The method of inputting the frequency domain embedding vector and the time domain embedding vector into a Mamba-based channel prediction core module to obtain enhanced features includes: the frequency domain embedding vector and the time domain embedding vector enter a shared time-dependent Mamba layer, and the shared time-dependent Mamba layer includes: a first normalization layer, a first Mamba block, and a first Dropout layer connected in sequence; In the shared time-dependent Mamba layer, the frequency domain embedding vector and the time domain embedding vector are normalized by the shared first normalization layer, and then the frequency domain original input and the time domain original input are obtained and passed through the first Mamba block together; After being processed by the first Mamba block, the frequency domain output features and the time domain output features are obtained. After being processed by the first Dropout layer, they are added to the original frequency domain input and the original time domain input respectively to obtain the frequency domain features and time domain features that have encoded time dependencies. The frequency domain features and time domain features are connected through the concat layer, and the projected frequency domain features and projected time domain features are obtained through the first linear projection layer and sent to the bidirectional Mamba block; In the bidirectional Mamba block, layer normalization is first performed in the second normalization layer to obtain fused features. After the fused features are transposed, they are processed using a bidirectional Mamba structure consisting of an independent forward and a reverse Mamba block to capture global feature dimension dependencies. The output of the bidirectional Mamba structure is further transposed and sent to the second Dropout layer, where it is combined with the fused features through a residual connection to obtain residual connection features. The residual connection features are then processed in sequence through the third normalization layer and the feedforward neural network layer to obtain nonlinear features. The nonlinear features are projected through the second linear projection layer to obtain the nonlinear features, which are then sent to the lightweight attention mechanism module for processing. The lightweight attention mechanism module outputs enhanced features.
9. The wireless channel prediction method based on Mamba architecture according to claim 5, characterized in that: The step of mapping the enhanced features into a prediction of channel state information for one or more future time steps through the output layer includes: In the output layer, the enhanced features are first normalized by the fourth normalization layer, and the feature dimensions are projected back to the original flattened dimensions by the FFN layer, and the output feature matrix O = FFN (LayerNorm (H final )); Among them, LayerNorm represents layer normalization processing, FFN(·) represents the feedforward neural network layer; To get an L-step forecast, we map the time dimension from P to L and get the forecast sequence: in, is the channel prediction sequence of the channel prediction model for the next L moments, Transpose(·) is the transposition operation, and Linear(·) is the linear layer operation; Then the prediction sequence Perform the reverse dimension reshaping operation to transform the predicted sequence into Convert to complex form and finally get the predicted output sequence is the predicted frequency domain CSI.
10. The wireless channel prediction method based on Mamba architecture according to claim 1 or 5, characterized in that: When training the channel prediction model, the loss function between the real channel and the predicted channel is minimized. The normalized mean square error is used as the objective function, which is expressed as conditional probability modeling: Among them, NMSE is the normalized mean square error, It is the expectation of the joint distribution of historical observations and future true channels, is the Frobenius norm squared of the matrix, i is the summation index, and f(X) is the mapping function of the input tensor X.
Citation Information
Patent Citations
Channel prediction method and system based on time-frequency joint correlation for massive MIMO systems
CN113595666B
Cited By
Dynamic channel prediction method based on feature extraction and multi-scale attention
CN122027060A