A method for designing and training a wireless base large model
By designing a large-scale wireless base model based on a mask autoencoder, the accuracy and adaptability issues of channel estimation and prediction tasks in MIMO-OFDM systems were solved, achieving efficient channel estimation and prediction and reducing model deployment costs.
Patent Information
- Application Number
- CN202411805649.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-12-10
AI Technical Summary
Existing deep learning techniques suffer from limitations in channel estimation and prediction tasks in MIMO-OFDM systems, including limited model size, insufficient accuracy, poor adaptability due to a single training dataset, and high model deployment overhead.
We designed a large-scale wireless base model based on a masked autoencoder, pre-trained it on a large-scale dataset using a self-supervised training method, constructed a network architecture that adapts to CSI data of different sizes, and applied it to channel estimation and channel prediction tasks.
It improves the accuracy of channel estimation and prediction, enhances the generalization ability of the model, and reduces the model deployment overhead on the base station side.
Smart Images

Figure CN119583263B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication technology, specifically relating to a design and training method for a large wireless base model. It is a base model design scheme based on a mask autoencoder applied to MIMO-OFDM wireless communication systems. The constructed base model is pre-trained in a self-supervised manner on a large-scale CSI dataset, thereby achieving high-precision and strong generalization of channel estimation and channel prediction. Background Technology
[0002] Massive Multiple-Input Multiple-Output (mMIMO) and Orthogonal Frequency-Division Multiplexing (OFDM) technologies are two fundamental technologies widely used in 5G mobile communication systems. In MIMO-OFDM systems, accurate Channel State Information (CSI) is crucial for tasks including precoding, beamforming, power allocation, modulation selection, and transmit antenna selection. Channel estimation and channel prediction are two basic methods for obtaining CSI. The former typically estimates CSI using received pilot signals, while the latter predicts unknown CSI in both the time and frequency domains based on known partial CSI.
[0003] Deep learning technology, due to its strong nonlinear fitting capabilities, has been widely applied in channel estimation and prediction, such as convolutional neural networks (CNNs) and Transformers. However, existing deep learning technologies have some drawbacks in channel estimation and prediction tasks: 1) Limited by model size, estimation and prediction accuracy are insufficient; 2) Training is performed on a single dataset with fixed configuration, requiring retraining when data distribution or system parameters change; 3) The separate design of each task necessitates the deployment of multiple models simultaneously at the base station, significantly increasing storage and computational overhead.
[0004] Foundation models have achieved great success in fields such as natural language processing. These models demonstrate superior few-shot learning capabilities on downstream tasks, surpassing even specially designed small models, through self-supervised pre-training on large datasets. Foundation models offer a novel solution to the challenges of existing deep learning techniques in channel estimation and prediction tasks. Summary of the Invention
[0005] This invention proposes a design and training method for a large wireless base station model. It is a design scheme for a large wireless base station model based on masked autoencoders (MAE). The network architecture of the large base station model and a self-supervised training method based on MAE are designed. The pre-trained large wireless base station model can be directly used for inference in channel estimation tasks, time-domain and frequency-domain channel prediction tasks of wireless communication systems, which can improve the accuracy of channel estimation and prediction and reduce the operating cost.
[0006] The network architecture of the base-based large model proposed in this invention can adapt to CSI data of different sizes. The self-supervised training method proposed in this invention can effectively capture the three-dimensional spatial-temporal-frequency relationship of channel state information of wireless communication systems. The pre-trained base-based large model can be directly applied to channel estimation and channel prediction tasks.
[0007] To achieve the above objectives, this invention designs a network architecture based on a masked autoencoder (MAE), including encoder and decoder modules. Simultaneously, various pre-training tasks, such as mask reconstruction, are designed to enable the network to undergo effective self-supervised learning (pre-training), resulting in a pre-trained base model. Finally, the pre-trained base model is directly applied to channel estimation and channel prediction tasks, capable of handling CSI data of varying sizes.
[0008] The base large model design scheme based on the mask autoencoder of the present invention includes the following steps:
[0009] 1) Construct a large-scale wireless base station network architecture based on MAE
[0010] The network architecture comprises an embedding module, a masking module, an encoder module, a decoder module, and an output module. CSI data is first converted into a series of tokens by the embedding module. Then, the masking module selectively retains a portion of the tokens and inputs them into the encoder module. The encoder module's output is sequentially concatenated with some learnable masking features and input into the decoder. The decoder's output is then processed by the output module to reconstruct the complete CSI data.
[0011] 2) Several pre-training tasks, such as mask reconstruction and interpolation denoising, were designed to pre-train the large model of the wireless base station;
[0012] Two types of pre-training tasks were designed. The first type is mask reconstruction, which involves partially masking tokens in a mask module according to a certain method and proportion, and then training by reconstructing the complete CSI. The second type is interpolation denoising, which involves downsampling the original complete CSI and then interpolating to obtain a coarse estimate, which is then used to estimate the original accurate CSI using a base large model.
[0013] 3) Apply the pre-trained base model to channel estimation and channel prediction tasks.
[0014] For channel estimation tasks, they can be transformed into interpolation denoising tasks. For time-domain and frequency-domain channel prediction tasks, they can be transformed into time-domain and frequency-domain mask recovery tasks, respectively.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0016] The proposed large-scale wireless base station model design scheme based on masked autoencoders enables efficient self-supervised pre-training on large-scale datasets with inconsistent CSI sizes. After training, it can be directly applied to channel estimation and prediction tasks. This significantly improves the model's accuracy in channel estimation and prediction tasks, while also exhibiting strong generalization ability, capable of handling CSI data of different sizes and distributions.
[0017] The large-scale base model design scheme based on mask autoencoder proposed in this invention has the following technical advantages:
[0018] (i) Improved estimation and prediction accuracy: This base model has obtained strong CSI representation capabilities through pre-training on massive datasets, which can significantly improve the accuracy of channel estimation and prediction.
[0019] (ii) Enhanced generalization ability: This base model can achieve strong performance on CSI distributions not seen in the training set without retraining.
[0020] (iii) Reduced model deployment overhead: The large base model can simultaneously process CSI data with different configurations and different types of channel estimation and prediction tasks, which greatly reduces the number of models required on the base station side, thereby reducing model deployment overhead. Attached Figure Description
[0021] Figure 1 This is a structural block diagram of the large-scale wireless base model based on a mask autoencoder constructed in this invention. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0023] This invention provides a design and training method for a large wireless base station model. It designs the network architecture of the large base station model and a self-supervised training method based on MAE. The self-supervised training method can effectively capture the three-dimensional relationship of space-time-frequency. The pre-trained large wireless base station model can be directly used for inference in channel estimation tasks, time domain and frequency domain channel prediction tasks of wireless communication systems, which can improve the accuracy of channel estimation and prediction and reduce the operating cost.
[0024] In practical implementation, this invention is applied to a MIMO-OFDM system, where the base station is equipped with a planar array of multiple antennas, while the user side is equipped with a single antenna. The proposed large-scale base model is deployed on the base station side and can simultaneously handle channel estimation and channel prediction tasks for the three-dimensional CSI (Time-Frequency-Spatial). Channel estimation refers to recovering the complete CSI from known partial location CSI estimates. The channel prediction task includes time-domain channel prediction and frequency-domain channel prediction, which respectively refer to predicting unknown CSI in the time and frequency directions using partial CSI.
[0025] This invention comprises three steps: network construction, pre-training task design, and deployment. The specific steps are as follows:
[0026] S10: Construct a large-scale wireless base network architecture based on MAE, including an embedding module, a masking module, an encoder module, a decoder module, and an output module.
[0027] S20: Design various self-supervised pre-training tasks such as mask reconstruction and interpolation denoising. The proposed base model can obtain strong CSI representation capabilities after pre-training.
[0028] S30: The pre-trained wireless base station large model is applied to channel estimation and channel prediction tasks, achieving good estimation and prediction accuracy, while also having strong generalization ability.
[0029] In step S10: Figure 1 The diagram shows the network structure of the constructed MAE-based large-scale wireless base model. In specific implementation, this invention obtains a 3D CSI dataset of the wireless communication system through simulation or actual measurement. The 3D CSI samples of the dataset are sequentially processed through an embedding module, a masking module, an encoder module, a decoder module, and an output module, ultimately outputting the reconstructed CSI. This step includes the following processes S11–S15:
[0030] S11: The embedding module's function is to process the input "time-space-frequency" 3D CSI samples into blocks, converting them into a series of tokens. Assume the 3D CSI samples are H∈C. T×N×KWhere C represents the set of complex numbers, and T, N, and K represent the number of time sampling points, the number of antennas, and the number of subcarriers of the CSI sample, respectively. First, the real and imaginary parts of the three-dimensional CSI sample H are separated and converted into a real tensor H. in ∈C 2×T×N×K ,Right now
[0031] H in [1,:,:,:]=Re{H}
[0032] H in [2,:,:,:]=Im{H}
[0033] Among them, H in [1,:,:,:] and H in [2,:,:,:] represent H respectively. in The subtensors corresponding to the first-dimensional indices 1 and 2 are Re and Im, which represent taking the real and imaginary parts of the tensor, respectively.
[0034] We employ 3D convolution operations to implement block segmentation and embedding operations. Specifically, the 3D samples are divided into blocks of size (t, n, k) and embedded into the feature space to obtain a series of one-dimensional tokens. Here, t, n, and k represent the temporal, spatial, and frequency sampling lengths of each cube, respectively. Let H... in The resulting 3D token tensor after 3D convolution is H. conv Its expression can be written as
[0035]
[0036] Here, Conv3d represents a 3D convolution operation with 2 input channels and D output channels, and both the kernel size and stride are (t, n, k). H... conv The last three dimensions are merged to obtain the CSI token tensor H. emb ∈R D×L ,in That is, the total number of tokens. The resulting CSI token tensor H emb This is the output of the embedded module.
[0037] S12: Transfer the CSI token tensor H obtained from the embedded module emb Perform a masking operation on the input H. emb Perform masking according to the masking rules, that is, preserve the original H according to the masking rules. emb The L blocks of tokens contain W tokens, where the specific masking rules depend on the pre-training task performed (see steps S21-S24). Let H be the tensor formed by the tokens retained after the masking operation. mask And H mask ∈R D×W ,Right now
[0038] H mask =MASK{H emb}
[0039] MASK stands for masking operation. The masking operation is different for each pre-training task, as detailed in step S20.
[0040] S13: H mask Input encoder module. First, add encoder position encoding to the masked W token. Assume the encoder position encoding is P. e ∈R D×W Then the token with added location encoding is
[0041] H pos =H mask +P e ∈R D×W
[0042] The positional encoding of each token depends on the token dimension before merging in H. conv The three-dimensional position within H. mask The i-th token in the array, assuming its three-dimensional position is (Pt) i ,Pn i ,Pk i ), whose corresponding position code is P ei =P e [:,i]∈R D×1 P ei Encoded by time location Spatial location coding and frequency position coding It is pieced together. Among them, This indicates a floor operation. All three methods use SinCos absolute position encoding, commonly used in transformers.
[0043] like Figure 1 As shown, the main body of the token input encoder after adding position encoding consists of M transformer blocks. Each transformer block adopts the same block structure as the classic vision transformer (ViT) architecture. The output token tensor of the overall encoder module after passing through these transformer blocks is H. enc And H enc ∈R D×W .
[0044] S14: Convert the encoder result H enc The input decoder module processes the data. First, it uses LW mask tokens M∈R. D×(L-W)The output of the encoder is concatenated with the output of the concatenated tensor H to obtain the concatenated tensor H. con =[H enc ,M]∈R D×L Each mask token is a learnable variable. Then, for H... con Adding decoder position encoding yields a token with added position encoding.
[0045] H dec_pos =H con +P d ∈R D×L
[0046] The decoder position encoding and the encoder position encoding in step S13 are set in the same way, both of which are SinCos absolute position encoding based on the three-dimensional position of each token.
[0047] The position-encoded token is input into the main body of the decoder, which consists of N ViT transformer blocks. This yields the output of the decoder module, i.e., the output token is H. dec And H dec ∈R D×L .
[0048] S15: Transfer the decoder's output token H dec The final reconstructed CSI is obtained through the output module. First, the H... dec The fully connected layer of the input / output module obtains the final output H. pred ∈R (2tnk)×L H pred After dimensional rearrangement, it becomes H. pred1 ∈ Then, the dimensions are further rearranged into a 7-dimensional tensor. Then, the 7-dimensional tensor H... pred2 The 2nd and 3rd dimensions, the 4th and 5th dimensions, and the 6th and 7th dimensions are merged to obtain the dimension-merged tensor H. pred3 And H pred3 ∈R 2×T×N×K H pred3 The first dimension (of size 2) represents the real and imaginary parts of the final reconstructed CSI. Let H be the CSI obtained from the final reconstruction. final And H final ∈C T×N×K Then the following relationship is satisfied:
[0049] H final =H pred3 [1,:,:,:]+1j×H pred3 [2,:,:,:]
[0050] Here, 1h represents the imaginary unit. H pred3[1,:,:,:] and H pred3 [2,:,:,:] represent tensors H respectively. pred3 The subtensors corresponding to the first-dimensional indices 1 and 2.
[0051] Step S20: Several self-supervised pre-training tasks, such as mask reconstruction and interpolation denoising, are designed, i.e., pre-training is performed based on the large base model constructed in S10. S21 to S25 respectively introduce five pre-training tasks: random mask reconstruction, temporal mask reconstruction, spatial mask reconstruction, frequency domain mask reconstruction, and interpolation denoising.
[0052] S21: Pre-training task based on random mask reconstruction. Based on the large wireless base model constructed in S10, input 3D CSI samples H∈C. T×N×K The final output of the wireless base station large model network is H final The loss function for training the network model is the output H. final The mean square error (MSE) between the input H and the input H is...
[0053]
[0054] The masking module employs a random masking method. Specifically, in step S12, H is randomly selected. emb ∈R D×L W tokens from the L block of tokens are reserved as token H. mask ∈R D×W .
[0055] S22: Pre-training task based on temporal mask reconstruction. Its training process is the same as S21, the difference being that the mask module uses a temporal masking method. For ease of operation, firstly, H... emb Dimensional rearrangement in and These represent the number of blocks in the three dimensions of time, space, and frequency, respectively. Then, all time-dimension coordinates less than [a certain value] are retained. The token, assuming the result after retention is H maskres ,but
[0056]
[0057] Then for H maskres Perform dimensional rearrangement to obtain the temporal mask result H. mask ∈R D×W ,in
[0058] S23: Pre-training task based on frequency-domain mask reconstruction. Its training process is the same as that of S21, except that the mask module uses the frequency-domain mask method. That is, all tokens with frequency dimension coordinates less than are retained. Assuming the result after retention is H maskres , then
[0059]
[0060] Then, perform dimension rearrangement on H maskres to obtain the result of the frequency-domain mask H mask ∈R D×W , where
[0061] S24: Pre-training task based on spatial-domain mask reconstruction. Its training process is the same as that of S21, except that the mask module uses the spatial-domain mask method. Different from the frequency-domain mask and the time-domain mask, the spatial-domain mask randomly retains some tokens according to a certain ratio Q (0 < Q < 1) according to the spatial coordinates, and the result after retention is , where
[0062] Then, perform dimension rearrangement on H maskres to obtain the result of the spatial-domain mask H mask ∈R D×W , where
[0063] S25: Pre-training task based on interpolation denoising. Based on the base large model constructed in S10, input the three-dimensional CSI sample H ∈ C T×N×K , first downsample at intervals of it, in, and ik in the time domain, spatial domain, and frequency domain. Assuming the result after downsampling is H ds , then
[0064] H ds = H[1:it:T, 1:in:N, 1:ik:K]
[0065] Perform linear interpolation on H ds to obtain H us ∈C T×N×K , that is, the same size as H. Input H us into the network, where the mask module in the network does not perform masking, that is
[0066] H mask = H emb
[0067] The final output of the network is H final , and the loss function for network training is the MSE between the output H final and the input H, that is
[0068]
[0069] Step S30: The pre-trained base model is applied to the channel estimation and channel prediction tasks. Steps S31 to S33 describe the methods for applying the pre-trained network to channel estimation, time-domain channel prediction, and frequency-domain channel prediction, respectively.
[0070] S31: The pre-trained wireless base station large model is used for the channel estimation task. Channel estimation refers to recovering the complete CSI from known partial location CSI estimates. First, linear interpolation is performed on the known partial location CSI estimates to obtain the interpolated CSI, i.e., H∈C. T×N×K Then input the data into the network. The final output is the CSI estimation result.
[0071] S32: Apply the pre-trained wireless base station large model to the time-domain channel prediction task. Time-domain prediction refers to predicting unknown CSIs using partial CSIs over time. Assume the known CSI is H. t0 ∈C T0×N×K The CSI that needs to be predicted is H. t1 ∈C (T -T0)×N×K First, H t0 By padding with zeros, the overall CSI is obtained as H∈C. T×N×K Then input it into the network. The masking module uses a time-domain masking method, that is, it retains the tokens with time-domain coordinates 1:T0. Finally, the output result H is taken. final The time-domain prediction portion is used as the prediction result, i.e.
[0072] H t1 =H final [T0+1:T,:,:]
[0073] S33: Apply the pre-trained large-scale wireless base station model to the frequency domain channel prediction task. Frequency domain prediction refers to predicting unknown CSIs using partial CSIs at the frequency level. Assume the known CSI is H. k0 ∈C T×N×K0 The CSI that needs to be predicted is H. k1 ∈C T ×N×(K-K0) First, H k0 By padding with zeros, the overall CSI is obtained as H∈C. T×N×K Then input it into the network. The masking module uses frequency domain masking, that is, it retains the tokens with frequency domain coordinates 1:K0. Finally, the output result H is taken. final The frequency domain prediction portion is used as the prediction result, i.e., H. f1 =H final [:,:,K0+1:K].
[0074] Through the above steps, this invention designs a network architecture for a large-scale wireless base model and a self-supervised training method based on MAE. The self-supervised training method can effectively capture the three-dimensional relationship between space, time, and frequency. The pre-trained large-scale wireless base model can be directly used for inference in channel estimation, time-domain and frequency-domain channel prediction tasks of wireless communication systems, which can improve the accuracy of channel estimation and prediction and reduce the operating cost.
[0075] It should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art will understand that various substitutions and modifications are possible without departing from the scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection of the present invention is defined by the scope of the claims.
Claims
1. A design and training method for a large-scale wireless base station model, characterized in that, Design a masked autoencoder (MAE) based network architecture for channel estimation and channel prediction tasks; design a masked reconstruction pre-training task to enable the network to perform effective self-supervised learning and obtain a pre-trained base model; The pre-trained base model is used to process Channel State Information (CSI) data of different sizes to achieve channel estimation and channel prediction; including: 1) Construct a large-scale wireless base network architecture based on MAE, including an embedding module, a masking module, an encoder module, a decoder module, and an output module; Acquire a time-space-frequency three-dimensional CSI dataset of a wireless communication system; convert the CSI data into a series of tokens through an embedding module; selectively retain a portion of the tokens through a masking module and input them into an encoder module; concatenate the learnable mask features in sequence using the output of the encoder module and input them into a decoder module; the output of the decoder module is then processed by an output module to recover the complete CSI data. 2) Design various self-supervised pre-training tasks to pre-train the large wireless base model; the self-supervised pre-training tasks include: mask reconstruction and interpolation denoising; mask reconstruction includes random mask reconstruction, temporal mask reconstruction, spatial mask reconstruction and frequency domain mask reconstruction; For the pre-training task of mask reconstruction, the input 3D CSI samples are fed into the constructed wireless base large model, and the loss function for training the network model is the mean square error between the output and the input. The mask module adopts random masking, time-domain masking, frequency-domain masking, and spatial masking, respectively, corresponding to the pre-training tasks based on random mask reconstruction, time-domain mask reconstruction, frequency-domain mask reconstruction, and spatial mask reconstruction. For the pre-training task of interpolation denoising, 3D CSI samples are input into the constructed wireless base station large model, and downsampling is performed at intervals in the time, spatial and frequency domains to obtain the downsampled results; the downsampled results are then obtained by linear interpolation; the results obtained by linear interpolation are then input into the wireless base station large model network, where the masking module in the network does not perform masking; the loss function for network training is the mean square error between the final output and the input; 3) Apply the pre-trained base model to channel estimation and channel prediction tasks; The channel estimation task is transformed into an interpolation denoising task for processing; the time-domain and frequency-domain channel prediction tasks are transformed into time-domain and frequency-domain mask recovery tasks for processing, respectively. By following the steps described above, channel estimation and channel prediction can be performed on the wireless communication system.
2. The design and training method for the large-scale wireless base model as described in claim 1, characterized in that, Specifically, the time-space-frequency three-dimensional CSI dataset is obtained through simulation or actual measurement.
3. The design and training method for the large-scale wireless base model as described in claim 1, characterized in that, Let the samples of the time-space-frequency three-dimensional CSI dataset be denoted as H∈C T×N×K , where C represents the set of complex numbers, and T, N and K represent the number of time sampling points, the number of antennas and the number of subcarriers of the CSI sample, respectively.
4. The design and training method for the large-scale wireless base model as described in claim 3, characterized in that, The large-scale wireless base station network architecture constructed in step 1) specifically includes: S11. The embedding module separates the real and imaginary parts of the CSI data samples and converts them into real tensors; it uses three-dimensional convolution to implement block and embedding operations and embeds them into the feature space to obtain a three-dimensional token tensor; it merges the last three dimensions of the obtained three-dimensional token tensor, and the resulting CSI token tensor is the output of the embedding module. S12: Mask the CSI token tensor obtained through the embedding module, retaining a portion of the tokens to obtain the masked tensor; each pre-training task corresponds to a different masking operation. S13: Input the masked tokens from the masking module into the encoder module; add encoder position codes to the masked tokens; the position code of each token depends on the three-dimensional position of the token in the three-dimensional token tensor before the token dimensions are merged; input the tokens with added position codes into the encoder module to obtain the output of the encoder module; S14: Input the output of the encoder module into the decoder module for processing, including: concatenating a portion of the mask tokens with the encoder output to obtain a concatenated tensor; where each mask token is a learnable variable; and then adding decoder position encoding to the concatenated tensor to obtain a token with added position encoding; where the decoder position encoding and encoder position encoding are set to be the same. The position-encoded token is input into the decoder module to obtain the output of the decoder module, i.e., the output token; S15: Obtain the final reconstructed CSI by passing the output of the decoder module through the output module; Input the output token of the decoder module into the fully connected layer to obtain the output token H. pred ; First, H pred After dimensional rearrangement, it becomes H. pred1 Then rearrange the dimensions to H pred2 ; Then for H pred2 The second and third, fourth and fifth, and sixth and seventh dimensions are merged respectively to obtain H. pred3 ;where H pred3 The first dimension represents the real part of the reconstructed CSI, and the second dimension represents the imaginary part of the reconstructed CSI; converting the real and imaginary parts into their corresponding complex numbers yields the final reconstructed CSI.
5. The design and training method for the large-scale wireless base model as described in claim 4, characterized in that, In S11, the embedded module specifically includes: First, the real and imaginary parts of the 3D CSI sample are separated and converted into a real tensor H. in ∈C 2×T×N×K , is represented as: A in [1,:,:,:]=Re{H} H in [2,:,:,:]=Im{H} Among them, H in [1,:,:,:] and H in [2,:,:,:] represent H respectively. in The subtensors corresponding to the first-dimensional indices 1 and 2, Re and Im represent taking the real and imaginary parts of the tensor, respectively; The block segmentation and embedding operations are implemented using three-dimensional convolution operations. Specifically, the three-dimensional samples are divided into blocks of size (t,n,k) and embedded into the feature space to obtain a series of one-dimensional tokens. Here, t, n, and k represent the temporal, spatial, and frequency sampling lengths of each block, respectively. The result after 3D convolution is H. conv , is represented as: Where Conv3d represents a 3D convolution operation, with 2 input channels and D output channels, and the kernel size and stride are both (t,n,k); The total number of tokens; H conv The last three dimensions are merged to obtain H. emb ∈R D×L ,in That is, the total number of blocks; the resulting H emb This is the output of the embedded module; In S12, the embedded token is masked, specifically by masking the input H. emb Perform masking, that is, preserve the original H according to the masking rules. emb The L blocks of tokens contain W tokens; let the masked tensor be H. mask ∈R D×W ,Right now H mask =MASK{H emb } MASK stands for masking operation.
6. The design and training method for a large-scale wireless base model as described in claim 5, characterized in that, In S13, encoder position encoding is added to the W tokens after the mask, specifically as follows: Let the encoder position code be P e ∈R D×W Then the token with added position encoding is represented as: H pos =H mask +P e ∈R D×W For H mask The i-th token in the array has a 3D position (Pt). i ,Pn j ,Pk i ), whose corresponding position code is P ei =P e [:,i]∈R D×1 ; P ei Encoded by time location Spatial location coding and frequency position coding It is pieced together; among them, This indicates a floor operation; all positional encodings use the SinCos absolute positional encoding commonly used in transformers.
7. The design and training method for a large-scale wireless base model as described in claim 6, characterized in that, The encoder's main body consists of M transformer blocks; each transformer block adopts a vision transformer (ViT) block structure; after passing through the transformer blocks, the output token tensor of the encoder module is H. enc And H enc ∈R D×W .
8. The design and training method for a large-scale wireless base model as described in claim 7, characterized in that, In S14, the decoder module first uses LW mask tokens M∈R D×(L-W) The output of the encoder is concatenated with the output of the concatenated tensor H to obtain the concatenated tensor H. con =[H enc ,M]∈R D×L Each mask token is a learnable variable; then H con Adding positional encoding to the decoder yields a position-encoded token, represented as follows: H dec_pos =H con +P d ∈R D×L The main body of the decoder consists of N ViT transformer blocks; In S15, specifically, the decoder's output token H is... dec The input is a fully connected layer to obtain the final output H. pred ∈R (2tnk)×L ;Including: First, after dimensional rearrangement, it is transformed into Then the dimensions are rearranged as follows: Then, the 1st and 3rd dimensions, the 3rd and 4th dimensions, and the 5th and 6th dimensions are merged to obtain H. pred1 ∈R 2×T×N×K ;where H pred1 The first dimension represents the real part of the reconstructed CSI, and the second dimension represents the imaginary part of the reconstructed CSI; the final reconstructed CSI is H. final ∈C T×N×K , represented as: H final =H pred1 [1,:,:,:]+1j×H pred1 [2,:,:,:], where 1j represents the imaginary unit, H pred3 [1,:,:,:] and H pred3 [2,:,:,:] represent tensors H respectively. pred3 The subtensors corresponding to the first-dimensional indices 1 and 2.
9. The design and training method for a large-scale wireless base model as described in claim 8, characterized in that, In step 2), for the pre-training task of random mask reconstruction, the loss function for network training is expressed as: For the pre-training task of temporal mask reconstruction, firstly, H... emb Dimensional rearrangement in and These represent the number of blocks in the three dimensions of time, space, and frequency, respectively; then, all time dimension coordinates less than [a certain value] are retained. The token; the result H after retention maskres Represented as: Then for H maskres Perform dimensional rearrangement to obtain the temporal mask result H. mask ∈R D×W ,in For the pre-training task of frequency domain mask reconstruction, the preserved result H maskres Represented as: Then for H maskres Perform dimensional rearrangement to obtain the frequency domain mask result H. mask ∈R D×W ,in For the pre-training task of spatial mask reconstruction, the spatial mask randomly retains some tokens according to the spatial coordinates at a ratio of Q, resulting in the following retained result: in, Then for H maskres Perform dimensional rearrangement to obtain the spatial mask result H. mask ∈R D×W ,in For the pre-training task of interpolation denoising, the downsampled result H ds Represented as: H ds =H[1:it:T,1:in:N,1:ik:K] For H ds H is obtained through linear interpolation. us ∈C T×N×K That is, the same size as H; H us Input network, where the masking module in the network does not perform masking, i.e., H mask =H emb .
10. The design and training method for a large-scale wireless base model as described in claim 9, characterized in that, In step 3), the pre-trained network is used for channel estimation, time-domain channel prediction, and frequency-domain channel prediction, specifically as follows: S31: Use the pre-trained large wireless base station model for the channel estimation task: First, linear interpolation is performed on the CSI estimates using known partial locations to obtain the interpolated CSI, i.e., H∈C. T ×N×K Then input the pre-trained wireless base large model, and the final output result is the CSI estimation result; S32: Applying the pre-trained large wireless base model to the time-domain channel prediction task: First, the known CSI is padded with zeros to obtain the overall CSI, denoted as H∈C. T×N×K Then, the pre-trained wireless base model is input; the masking module uses a temporal masking method, that is, it retains the token with temporal coordinates of 1:T0; finally, the temporal prediction part of the output result is taken as the predicted CSI result, that is, H. t1 =H final [T0+1:T,:,:]; S33: The pre-trained wireless base station large model is used for the frequency domain channel prediction task. The process is as follows: First, the known CSI is padded with zeros to obtain the overall CSI, denoted as H∈C. T×N×K Then, the pre-trained wireless base model is input; the masking module uses frequency domain masking, that is, it retains the token with frequency domain coordinates 1:K0; finally, the frequency domain prediction part of the output result is taken as the prediction result, that is, H. f1 =H final [:,:,K0+1:K].
Citation Information
Patent Citations
Electromyographic signal self-supervision pre-training method based on mask auto-encoder
CN116049635A
Time series data self-supervision pre-training model, construction method, equipment and storage medium
CN116522099A