Near-field channel estimation method based on perceptual auxiliary mutual attention in super-large scale MIMO (Multiple Input Multiple Output) sensing integration

Through the CAT-CENet neural network driven by the cross-modal attention mechanism, the radar-perceived scatterer position information and channel matrix data are integrated to solve the complexity and noise interference problems of high-dimensional channel estimation in ultra-large-scale MIMO ISAC systems, and achieve efficient and accurate channel estimation.

CN120710831APending Publication Date: 2025-09-26NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510844818.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In ultra-large-scale MIMO ISAC systems, the high dimensionality of the channel matrix leads to increased computational complexity, and the channel estimation accuracy is affected by noise and interference. Existing methods are inefficient in high-dimensional channel estimation.

Method used

The CAT-CENet neural network driven by the cross-modal attention mechanism is combined with a convolutional neural network and a Transformer encoder. The radar-perceived scatterer position information and channel matrix data are fused through the cross-modal attention mechanism. The feature channel and spatial channel attention mechanism are used to extract channel features, reduce noise interference, and improve estimation accuracy.

Benefits of technology

It significantly improves the accuracy and stability of channel estimation, reduces computational complexity, adapts to different channel environments, improves the versatility and flexibility of the system, and is suitable for various practical communication scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120710831A_ABST
    Figure CN120710831A_ABST
Patent Text Reader

Abstract

The invention discloses a near-field channel estimation method based on perceptual auxiliary mutual attention in super-large scale MIMO (Multiple Input Multiple Output) sensing integration. Comprising the steps of obtaining a training data set and a test data set through channel characteristics in a super-large-scale MIMO ISAC system, position information characteristics of a scatterer in a radar sensing communication process, model construction of a radar sensing target and preprocessing of signals received by a base station; a deep convolutional neural network model is constructed based on a convolutional layer and a cross-modal attention encoder, the convolutional layer is used for extracting local features of a matrix and remodeling data types of the local features and the cross-modal attention encoder, and the encoder fuses matrix data through cross-modal attention and enhances extraction and utilization of important features by using double attention; training and testing the deep convolutional neural network model; and obtaining an estimated channel matrix by adopting a deep convolutional neural network model. According to the method, the accuracy and the stability of channel estimation are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communications, and in particular relates to a near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration. Background Art

[0002] ISAC (Integrated Sensing and Communication) and XL-MIMO (Extra-Large Multiple Input Multiple Output) technologies have become innovative in sixth-generation wireless communications. Frequency band sharing between radar and communication systems effectively avoids the inefficient use of otherwise fixed spectrum resources, thereby improving system efficiency. XL-MIMO utilizes a large number of antennas at the base station, significantly increasing communication capacity and reliability.

[0003] XL-MIMO significantly improves communication capacity and reliability by deploying a large number of antennas at base stations. Benefiting from the widespread use of antennas, it excels in spectral efficiency, energy efficiency, and interference mitigation. ISAC systems employ MIMO technology to improve spatial multiplexing and interference mitigation capabilities. However, in MIMO systems, the increased dimensionality of the channel matrix significantly increases the computational effort and complexity of channel estimation. Furthermore, in XL-MIMO, the base station needs to obtain channel state information from a large number of antennas, resulting in a high-dimensional channel matrix. Compared to traditional massive MIMO systems, the channel matrix in XL-MIMO not only has a higher dimensionality but also exhibits more challenging characteristics. When dealing with large-scale antenna systems, the computational load increases dramatically, requiring efficient algorithms and powerful computing resources for channel estimation. In XL-MIMO ISAC systems, due to the large antenna array, near-field effects are more pronounced, impacting estimation accuracy.

[0004] Traditional channel estimation methods typically rely on linear models and statistical assumptions. Least square (LS)-based approaches can only achieve low received signal-to-noise ratios and incur considerable pilot overhead. To reduce the high pilot overhead in channel estimation, in current massive MIMO systems, some compressed sensing-based methods, such as orthogonal matching pursuit and sparse Bayesian learning, leverage channel sparsity in the angular domain to estimate high-dimensional channels with low pilot overhead. In the study of ISAC near-field channel estimation, radar sensing and channel estimation were jointly analyzed. However, in XL-MIMO ISAC systems, these methods typically have higher complexity when processing high-dimensional channel matrices. Summary of the Invention

[0005] In response to the deficiencies of the above-mentioned prior art, the present invention provides a near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO interawareness integration, and proposes a neural network (CAT-CENet) for perception-assisted ultra-large-scale MIMO ISAC channel estimation driven by a cross-modal attention mechanism. The method is suitable for ISAC near-field users with different channels; it combines convolutional neural networks and Transformer encoder structures to enhance feature extraction and representation capabilities, significantly improve the accuracy and stability of channel estimation, and can solve the problems of complexity and noise interference in high-dimensional channel estimation in existing XL-MIMO ISAC systems.

[0006] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:

[0007] The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO interawareness integration includes:

[0008] Step 1: Obtain training and test datasets based on the channel characteristics of the ultra-large-scale MIMO ISAC system, the construction of a radar target perception model, and the preprocessing of the signals received by the base station;

[0009] Step 2: Build a deep convolutional neural network model based on convolutional layers and cross-modal attention encoders. The convolutional layers are used to extract local features of the noisy channel vector and radar array response vector, and reshape their data types. The encoder is used to fuse the radar array response vector and the noisy channel vector data through cross-modal attention and use dual attention to fuse and extract features from the vector data in different channel and spatial dimensions, thereby enhancing the extraction and utilization of important features.

[0010] Step 3: Use the training dataset and test dataset to train and test the deep convolutional neural network model;

[0011] Step 4: Input the noisy channel vector and radar array response vector into the deep convolutional neural network model obtained in step 3, and output the estimated channel matrix.

[0012] To optimize the above technical solutions, specific measures taken also include:

[0013] The above step 1 includes the following sub-steps:

[0014] Step 1.1: Model the channel characteristics in the ultra-large-scale MIMO ISAC system:

[0015] Assume that the base station of an XL-MIMO ISAC system with M antennas serves a single-antenna user. There are L paths between the base station and the user. The radar senses K targets. There is partial overlap between the K radar-sensed targets and the L communication scatterers. Let the overlap number be X, 0≤X≤min(K,L), the sensed communication scatterer position information be R, and the arrival angle information be AoA. Thus, the radar array response vector containing the X overlapping scatterer position information can be calculated. The overlapping scatterer information is used as auxiliary information for the communication channel estimation task to perform ISAC joint channel estimation. It is assumed that the position information of the radar-perceived target can be estimated using radar array processing technology.

[0016] The channel characteristic model in the ultra-large-scale MIMO ISAC system is:

[0017] y=hx+η

[0018] Where y is the uplink signal received by the base station, h is the channel vector between the base station and the user, x represents the pilot signal transmitted by the user in a time slot, and η represents additive Gaussian noise;

[0019] When the scattering distance is less than the Rayleigh distance D Ray When , the near-field channel model is:

[0020]

[0021] in, λ is the carrier wavelength, represents the array response vector of the base station antenna array to the lth path, φ l represents the azimuth of the lth path, g l is the path gain, r l represents the distance from the lth scatterer to the center of the base station antenna array, represents the distance between the lth scatterer and the mth antenna of the base station, where d is the distance between adjacent antennas, and the parameter

[0022] Step 1.2: Build radar perception target model:

[0023] Radar sensing is used to detect the presence of a target and estimate the target's AoA and distance r k , radar cross section RCS parameters, so in the pth DP symbol duration of radar target detection, the sensing radar BS transmits a DP The corresponding received signal is expressed as:

[0024]

[0025] in, is the radar receiving signal, is the radar channel matrix, Dedicated pilot symbol With variance Additive Gaussian white noise;

[0026] The radar channel matrix depends on the target's AoA, distance r k and radar cross section RCS, modeled as:

[0027]

[0028] in r k ,and are the AoA, distance and RCS of the k-th target, is the radar array response vector containing the position information of X overlapping scatterers;

[0029] For a half-wavelength spatially uniform linear array, the array response vector for:

[0030]

[0031] Step 1.3: Preprocessing of the signal received by the base station:

[0032] Assuming that the power of the pilot signal x is known to be P, the uplink received signal y received by the base station is simplified to:

[0033]

[0034] Based on the above formula, the least squares estimate complex vector of the noise-free true channel vector h between the base station and the user is:

[0035]

[0036] in, is the noisy channel vector;

[0037] Step 1.4: Based on steps 1.1-1.3, the channel vector with noise Radar array response vector containing X overlapping scatterer position information The training dataset and test dataset are generated by taking the noise-free real channel vector between the base station and the user as the input data and the output label.

[0038] The input of the deep convolutional neural network model in step 2 is the tensor H c-input 、H s-input :

[0039] The noisy channel vector Expand to: (Z,M,1), then split it and connect it, reshape it into a The tensor H c-input ; Expand the radar array response vector containing X overlapping scatterer position information to: (Z, M, 1 × X), then split and connect it and reshape it into a The tensor H s-input ; Where Z is the number of samples in the training dataset, and M is the number of antennas configured in the base station of the XL-MIMO ISAC system.

[0040] The above-mentioned deep convolutional neural network model includes a two-dimensional convolutional layer 1, a two-dimensional convolutional layer 2, and an encoder 1, an encoder 2, two-dimensional convolutional layers 3 to 5 and a layer normalization and residual connection module connected in sequence;

[0041] The two-dimensional convolution layer 1 and the two-dimensional convolution layer 2 respectively process the input tensor H of the deep convolutional neural network model. c-input 、H s-input After extracting the features, the tensors H1, H2 are reshaped to obtain sizes of M×64, where M is the number of antennas of the base station in the XL-MIMO system;

[0042] H1 and H2 are input into encoder 1 to obtain the output H4. H4 and H2 are then input into encoder 2 to obtain the output tensor H5. After reshaping, H5 is input into two-dimensional convolution layers 3 to 5 to obtain H6.

[0043] H6 is input to the layer normalization and residual connection module to transform the input H c-input Subtracting H6 gives output H output .

[0044] A batch normalization layer is added after each of the above two-dimensional convolutional layers to accelerate the network training process; a rectified linear unit activation function is used on the output of each two-dimensional convolutional layer to introduce nonlinear features to make the network more expressive; a dropout layer is introduced between each convolutional layer to randomly inactivate a certain proportion of neurons to prevent network overfitting.

[0045] The structures of the above encoders 1 and 2 are the same, and both include a cross-modal attention mechanism module, a feature channel attention mechanism module, a spatial channel attention mechanism module, two fully connected layers, a dropout layer, and a normalization layer.

[0046] In the encoder 1, the cross-modal attention mechanism module generates a reconstruction matrix with the same dimension as the tensors H1 and H2 through the multi-head mutual attention mechanism, and sums the reconstruction matrix generated by the multi-head mutual attention mechanism with H1 and H2 element by element to obtain the final output tensor H m , size is M×64, H mrepresents the weighted sum of all high-dimensional feature maps;

[0047] The feature channel attention mechanism module applies a multi-head self-attention mechanism to generate and tensor H m The reconstruction matrix with the same dimension is finally combined with the reconstruction matrix generated by the multi-head self-attention mechanism m Perform element-by-element summation to obtain the final output tensor H c , size is M×64, H c represents the weighted sum of all high-dimensional feature maps;

[0048] The spatial channel attention mechanism module uses a multi-head self-attention mechanism to c Transpose to exchange dimensions and get the attention output matrix. Finally, transpose the attention output matrix again and compare it with H one by one. c Sum and get the output tensor H3, which is the weighted sum of all position features and the original features of each position, with a size of M×64;

[0049] The tensor H3 passes through two fully connected layers, a Dropout layer, and a normalization layer in sequence to obtain H4 and output it.

[0050] The formula of the above multi-head mutual attention mechanism is:

[0051] head i =Attention(Q i ,K i ,V i ),i=1,...,h,

[0052]

[0053] Among them, h represents the number of "heads", head i Represents the attention matrix obtained by the i-th “head”, Attention(Q i ,K i ,V i ) is the attention mechanism formula corresponding to the i-th “head”; Q i , K i and V i For Q, K, V through Represents the query, key and value, Q, K, V are the matrices of query, key and value respectively, Represents multiple weight matrix groups corresponding to different heads in the multi-head mutual attention mechanism.

[0054] The loss function used in step 3 above is:

[0055]

[0056] in, is the estimated channel matrix and the real channel matrix, θ is the network model parameter, N tr Indicates the number of training samples in the training data set, and the superscript (n) indicates the nth sample. represents the square of the Frobenius norm of a matrix.

[0057] In step 3 above, the network model is optimized using the Adam optimizer. The Adam optimizer combines momentum and adaptive learning rate techniques to accelerate convergence. Its update rules are as follows:

[0058]

[0059]

[0060] Among them, m t is the momentum vector at the tth iteration; v t is the squared gradient momentum at the tth iteration; since in the initial stage of training, m t and v t The initial value is 0. and is the momentum vector and squared gradient momentum at the tth iteration after correction; β1 is the attenuation coefficient of the momentum term; β2 is the attenuation coefficient of the squared gradient momentum; is the gradient of the loss function L(θ) with respect to the model parameters θ; α is the learning rate; θ t is the parameter vector of the model at the tth iteration; ∈ is a constant.

[0061] In the test phase of step 3 above, the estimated channel matrix output by the calculation model is With the real channel matrix The normalized mean square error NMSE between is used to evaluate the channel estimation result of the model. The smaller the normalized mean square error, the more accurate the channel estimation result. The calculation formula of the normalized mean square error is:

[0062]

[0063] in For expectations, represents the square of the Frobenius norm of a matrix.

[0064] The present invention has the following beneficial effects:

[0065] The present invention uses a cross-modal attention mechanism to effectively fuse the position information of radar-sensed scatterers with the pilot channel matrix data, reducing the impact of noise and multipath effects on the channel matrix, effectively extracting data features in the channel matrix, and using the feature channel attention mechanism and the spatial channel attention mechanism to extract the channel and spatial features of the channel matrix. The convolutional layer and residual connection optimization model are used to more accurately restore the channel matrix, maintain a low estimation error even under complex channel conditions, and significantly improve the reliability of channel estimation.

[0066] The modeling method of radar target perception and communication joint modeling in the present invention can adapt to different channel environments and can effectively cope with the multipath effect of the complex near-field of ISAC. As the number of near-field paths continues to increase, CAT-CENet can still maintain a high estimation performance. This adaptability enables the present invention to be widely used in various actual communication scenarios, improving the versatility and flexibility of the system.

[0067] This invention effectively reduces computational complexity. The deep convolutional network employed significantly improves computational efficiency through its parallel computing capabilities and efficient feature extraction mechanism. Compared with traditional methods, this invention utilizes batch normalization, residual connections, and dropout techniques during training, which not only increases training speed but also optimizes the model's computational overhead. This enables fast and efficient computation while maintaining high accuracy, making it suitable for efficient deployment in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 This is a neural network model diagram of the near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration of the present invention;

[0069] Figure 2 is a curve showing the change of the normalized mean square error value performance with the signal-to-noise ratio under different near-field conditions in the present invention;

[0070] Figure 3 This is a curve showing the change of normalized mean square error performance with signal-to-noise ratio under different scatterer overlap numbers in the present invention;

[0071] Figure 4 This is a curve showing how the normalized mean square error performance of various schemes changes with the number of near-field paths L when the signal-to-noise ratio is 15 dB in the present invention. DETAILED DESCRIPTION

[0072] The embodiments of the present invention are described in further detail below with reference to the accompanying drawings.

[0073] The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration of the present invention specifically includes the following steps:

[0074] Step 1: Obtain training and test datasets based on the channel characteristics of the ultra-large-scale MIMO ISAC system, the construction of a radar target perception model, and the preprocessing of the signals received by the base station;

[0075] In this embodiment, the channel characteristics in the ultra-large-scale MIMO ISAC system are modeled, a channel matrix is ​​constructed, a radar perception target is modeled, and the signals received by the base station are preprocessed to obtain a training data set and a test data set, which specifically includes the following sub-steps:

[0076] Step 1.1: Model the channel characteristics in the ultra-large-scale MIMO ISAC system:

[0077] Assume that the base station of the XL-MIMO ISAC system with M antennas serves a single-antenna user. There are L paths between the base station and the user. The radar senses K targets. There is partial overlap between the K radar-sensed targets and the L communication scatterers. Let the overlap number be X, 0≤X≤min(K,L), the sensed communication scatterer position information is R, and the arrival angle information is AoA. Thus, the position information of X overlapping scatterers can be calculated. The radar array response vector uses this overlapping scatterer information as auxiliary information for the communication channel estimation task to perform ISAC joint channel estimation. Because this paper mainly studies the perception-assisted near-field channel estimation task, it is assumed that the position information of the radar perception target can be estimated using radar array processing technology.

[0078] The channel characteristic model in the ultra-large-scale MIMO system is:

[0079] y=hx+η

[0080] Where y is the uplink signal received by the base station, h is the channel vector between the base station and the user, x represents the pilot signal transmitted by the user in a time slot, and η represents additive Gaussian noise.

[0081] Very large-scale antenna arrays can operate in both far-field and near-field conditions, which are determined by the Rayleigh distance D Ray Sure:

[0082]

[0083] Among them, D a is the array aperture, λ is the carrier wavelength;

[0084] For a uniform linear array, if the spacing between adjacent antennas is Rayleigh distance D Ray Expressed as:

[0085]

[0086] When the scattering distance is less than D Ray When , the near-field channel model is:

[0087]

[0088] in represents the array response vector of the base station antenna array to the lth path, φ l represents the azimuth of the lth path, g l is the path gain, r l represents the distance from the lth scatterer to the center of the base station antenna array, represents the distance between the lth scatterer and the mth antenna of the base station, where

[0089] Step 1.2: Build radar perception target model:

[0090] The core function of radar sensing is to detect the presence of a target and estimate the target's AoA and distance r k , radar cross section (RCS) and other related parameters. To achieve this, during the pth DP symbol duration of radar target detection, the BS transmits a DP The corresponding received signal can be expressed as:

[0091]

[0092] in, is the radar receiving signal, is the radar channel matrix, Dedicated pilot symbol With variance Additive Gaussian white noise;

[0093] The radar channel matrix depends on the target's AoA, distance r k and radar cross section (RCS), which can be modeled as:

[0094]

[0095] in r k ,and is the AoA and distance r of the kth target k and RCS, is the radar array response vector containing the position information of X overlapping scatterers;

[0096] For a half-wavelength spatially uniform linear array (ULA), the array response vector is:

[0097]

[0098] The present invention considers a monostatic radar architecture for sensing the common location of the radar (i.e., BS) transmit and receive arrays, so that the departure angles of the DPs and the arrival angles of the echo signals are the same.

[0099] Step 1.3: Pre-process the signal received by the base station:

[0100] Preprocess the data. x is a pilot signal with known power. Assuming its power is P, simplify y to:

[0101]

[0102] Based on the above formula, the least squares estimate of h is expressed as:

[0103]

[0104] Step 1.4: Based on steps 1.1-1.3, the channel vector with noise Radar array response vector containing X overlapping scatterer position information As input data, the noise-free real channel vector between the base station and the user is used as the output label to generate training data sets and test data sets. The noisy channel vector is estimated by the proposed network model to obtain the estimated channel vector and the noise-free true channel vector Comparisons were made to verify the estimated effects.

[0105] Step 2: Construct a deep convolutional neural network model based on convolutional layers and cross-modal attention encoders. The convolutional layers are used to extract local features of the noisy channel matrix and the radar-sensing scatterer position information matrix, and reshape their data types. The encoder is used to fuse the radar-sensing scatterer position information matrix with the noisy channel matrix data through cross-modal attention and use dual attention to fuse and extract features from the matrix data in different channel and spatial dimensions, thereby enhancing the extraction and utilization of important features.

[0106] In the embodiment, based on the convolutional network and the encoder, Figure 1The deep convolutional neural network model shown in the figure extracts local features of the channel matrix and the radar-sensing scatterer position information matrix through the convolution layer, and reshapes the data types of both for subsequent input to the encoder. The batch normalization layer is used to accelerate the convergence of the network, and the residual connection and layer normalization are used to improve the stability of the network. The encoder uses cross-modal attention to fuse the radar-sensing scatterer position information with the channel matrix data, and then uses the mixed domain attention of feature channels and spatial channels to fuse and extract features of the channel matrix data in three dimensions: data, channel, and space, thereby enhancing the extraction and utilization of important features.

[0107] (1) Construct a convolutional network structure, extract the local features of the channel matrix through the convolution layer, use the batch normalization layer to accelerate the convergence of the network, and improve the stability of the network through residual connection and layer normalization; the details are as follows:

[0108] The deep convolutional neural network model is CAT-CENet, and the input is the tensor H c-input 、H s-input :

[0109] The noisy channel vector Expand to: (Z,M,1), then split it and connect it, reshape it into a The tensor H c-input ;

[0110] The radar array response vector containing X overlapping scatterer position information is expanded to: (Z, M, 1 × X), and then split and connected to reshape it into a The tensor H s-input ;

[0111] Where Z is the number of samples in the training dataset, and M is the number of antennas configured in the base station of the XL-MIMO ISAC system.

[0112] CAT-CENet uses two two-dimensional convolution layers with a convolution kernel size of 3×3 and 64 convolution kernels to perform a convolution on the tensor H. c-input 、H s-input After extracting the features, they are reshaped into tensors H1 and H2 of size M×64 suitable for encoder input.

[0113] A batch normalization layer is added after each convolutional layer to accelerate the network training process. By normalizing small batches of data, the input of each layer remains stable and prevents gradient vanishing and gradient exploding problems.

[0114] The rectified linear unit (Relu) activation function is used on the output of the convolutional layer to introduce nonlinear features and make the model more expressive.

[0115] A Dropout layer is introduced between convolutional layers to randomly inactivate a certain proportion of neurons to prevent the model from overfitting.

[0116] After H1 and H2 are input into the first layer encoder, the output H4 and H2 are input into the second layer encoder to obtain the output tensor H5. After reshaping, they are input into three two-dimensional convolution layers with a convolution kernel size of 3×3 and a convolution kernel number of 2 to obtain H6. Then, layer normalization and residual connection are used to improve the stability of the network. In order to avoid the gradient disappearance problem, residual connection is used to input H input Same as H6. At the same time, layer normalization is used to normalize the output of each layer to improve training stability.

[0117] (2) Design the encoder structure and integrate the cross-modal attention mechanism module, feature channel attention mechanism module, and spatial channel attention mechanism module to obtain the importance score matrix of different information and enhance the extraction and utilization of important features. The details are as follows:

[0118] The Transformer encoder structure in the attention mechanism is applied to the system model network. The multi-head mutual attention mechanism is used in the cross-modal attention mechanism of the encoder structure. The mutual attention mechanism is used to calculate the dependency between different data for radar perception data and communication channel matrix data. The two inputs are recorded as: Q, K, V are the matrices of query, key, and value respectively, and the attention weight formula is:

[0119] Q=W Q X1, K=W k X2, V = W v X2,

[0120]

[0121] Among them, W q 、W k 、W v Denotes the corresponding weight matrix of Q, K, V, d k is the dimension of vector V;

[0122] The subspace in the multi-head mutual attention mechanism is generated as follows:

[0123]

[0124] The above formulas Q, K, and V are transformed by their respective linear transformation matrices Represents query Q i , key K i Sum V i ,in, Represents multiple weight matrix groups corresponding to different heads in the multi-head mutual attention mechanism;

[0125] Then the formula of the multi-head mutual attention mechanism is:

[0126] head i =Attention(Q i ,K i ,V i ),i=1,...,h,

[0127] Among them, h represents the number of "heads", head i Represents the attention matrix obtained by the i-th “head”, Attention(Q i ,K i ,V i ) is the attention mechanism formula corresponding to the i-th “head”.

[0128] The cross-modal attention mechanism utilizes the multi-head mutual attention mechanism and applies the multi-head self-attention mechanism in the feature channel and spatial domain.

[0129] In the cross-modal attention mechanism, a reconstruction matrix with the same dimension as the input tensors H1 and H2 is generated. Finally, the reconstruction matrix obtained by the mutual attention mechanism is summed element by element with H1 and H2 to obtain the final output tensor H m , size is M×64.

[0130] Then H m Input the feature channel attention module, process the feature map in the channel domain, capture the channel characteristics of the channel matrix, and output H c ;

[0131] Finally, H c After transposing to exchange dimensions, the spatial domain attention module is input to extract the spatial feature weight matrix. Finally, the attention output matrix is ​​transposed again to obtain the output H3, which is M × 64. H3 is the weighted sum of all position features and the original features of each position.

[0132] Step 3: Use the training dataset and test dataset to train and test the deep convolutional neural network model;

[0133] In the embodiment, a training data set and a test data set are used to train and test the deep convolutional neural network model;

[0134] (1) Train the entire network including the convolutional layer and the cross-modal attention encoder until the model converges, which includes the following sub-steps:

[0135] Step 3.1, use Mean Squared Error (MSE) as the loss function. The estimated channel matrix is The real channel matrix is The mean square error loss function is defined as:

[0136]

[0137] in, is the estimated channel matrix and the real channel matrix, θ is the network model parameter, N tr Indicates the number of training samples in the training data set, and the superscript (n) indicates the nth sample. represents the square of the Frobenius norm of a matrix.

[0138] Step 3.2: Use the Adam optimizer to optimize the network. The Adam optimizer combines momentum and adaptive learning rate techniques to converge quickly within a small number of iterations. The update rules are as follows:

[0139]

[0140] Among them, m t is the momentum vector at the tth iteration, representing the weighted average of the gradient, v t is the squared gradient momentum at the tth iteration, which represents the weighted average of the squared gradient and is the attenuation coefficient of the momentum term. Since in the initial stage of training, m t and v t The initial values ​​are all 0, and deviation correction is required. and is the corrected m t and v t ; β1 is the attenuation coefficient of the momentum term, β2 is the attenuation coefficient of the squared gradient momentum, is the gradient of the loss function L(θ) with respect to the model parameters θ, α is the learning rate, θ is the parameter vector of the model, and ∈ is a small constant used to avoid division by zero.

[0141] Step 3.3: Update the network weights and biases through continuous iteration until the loss function converges and the training process terminates.

[0142] (2) In the testing phase, the noisy channel matrix is ​​input into the trained deep neural network model, and the estimated channel matrix is ​​output to evaluate the channel recovery effect of the model, as follows:

[0143] During the testing phase, the noisy channel matrix is ​​input into the trained deep neural network. The model estimates the channel matrix through forward propagation. Calculate the estimated channel matrix With the real channel matrix Normalized Mean Squared Error (NMSE) between:

[0144]

[0145] This indicator is used to measure the accuracy of the model in channel recovery. The smaller the normalized mean square error, the more accurate the channel estimation result.

[0146] Step 4: Input the noisy channel matrix and perception matrix into the deep convolutional neural network model obtained in step 3, and output the estimated channel matrix.

[0147] This paper uses a radar-aware, cross-modal attention-driven, ultra-large-scale MIMO (ISAC) system for channel estimation. It employs a neural network consisting of convolutional layers and a cross-modal attention encoder to achieve high-precision channel estimation under a variety of channel conditions. By extracting local features from the channel matrix using a deep convolutional network and combining it with a cross-modal attention mechanism, the accuracy and robustness of channel estimation are effectively improved, overcoming the limitations of traditional methods in complex ISAC channel conditions. This paper not only excels in accuracy, flexibility, and efficiency, but is also applicable to dynamically changing channel environments, demonstrating broad application prospects and practical value.

[0148] Example 1

[0149] Set the simulation scenario parameters: number of base station antennas M = 256, wavelength λ = 0.01 meters, average path gain σ 2 =1, r l ~U(10,80) meters.

[0150] In the simulation, the training set has 45,000 samples and the validation set has 5,000 samples, all of which are data generated by the near-field signal ISAC channel model, containing 6 near-field paths, and the test set has 2,000 samples.

[0151] Apply Adam as the optimizer.

[0152] Specifically, the proposed method is compared with the traditional LS estimation method, LMMSE estimation method, HY-OMP and MAT-CNet.

[0153] Figure 2The normalized mean square error performance of near-field users versus signal-to-noise ratio is depicted, with two subgraphs corresponding to 3 and 6 near-field paths, respectively. Compared to other methods, CAT-CENet demonstrates superior performance in near-field conditions, significantly reducing the estimation error, especially in low signal-to-noise ratio conditions, showing higher accuracy. Figure 2 The number of paths in (b) increases, but the estimation accuracy remains high.

[0154] Figure 3 The figure shows the relationship between the normalized mean square error performance and the signal-to-noise ratio under different scatterer overlap numbers, including 6 near-field paths. Figure 3 It shows that the channel estimation error decreases significantly with the increase in the number of scatterers perceived by the radar. This is a significant advantage of CAT-CENet in making full use of the radar-perceived scatterer information.

[0155] Example 2

[0156] Set the simulation scenario parameters: number of base station antennas M = 256, wavelength λ = 0.01 meters, average path gain σ 2 =1, r l ~U(10,80) meters.

[0157] In the simulation, the training set has 45,000 samples and the validation set has 5,000 samples, both of which are data generated by the ISAC near-field channel model, which includes 6 near-field paths, and the test set has 2,000 samples.

[0158] Apply Adam as the optimizer.

[0159] Specifically, the proposed method is compared with the traditional LS estimation method, LMMSE estimation method, HY-OMP and MAT-CENet.

[0160] In order to better demonstrate the performance of the method proposed in this invention, Figure 4 The normalized mean square error performance of various schemes is shown as the number of near-field paths increases, where the signal-to-noise ratio is 15dB and there are 6 paths in the near field. As shown in the figure, as the number of paths increases, the complexity and difficulty of the channel estimation task increase, but the performance of CAT-CENet is far lower than other schemes in the scenario when there is only one scatterer overlapping. In addition, when all L scatterers overlap under the corresponding L near-field paths, on the one hand, the model complexity of CAT-CENet hardly increases, and on the other hand, its performance is significantly improved. Moreover, CAT-CENet also shows stable performance in other ISAC near-field channel estimation, highlighting the excellent performance and strong generalization ability of this invention.

[0161] The radar-aware assisted ultra-large-scale MIMO ISAC channel estimation method driven by a cross-modal attention mechanism in this paper applies neural networks to the near-field channel estimation of ultra-large-scale MIMO ISAC. It not only has excellent performance, showing high sensitivity and accuracy, but also has strong generalization ability, which is specifically reflected in its robustness to ISAC near-field scenarios. Regardless of the number of near-field paths, regardless of the influence of multipath effects and interference noise, and regardless of high or low signal-to-noise ratios, CAT-CENet always significantly outperforms other solutions.

[0162] In summary, the present invention models the channel characteristics in the ultra-large-scale MIMO ISAC system, constructs a channel matrix, models the position information of scatterers in the radar perception communication process, and jointly analyzes the perception-assisted communication channel estimation, and then preprocesses the signal received by the base station; a deep convolutional neural network model is constructed based on a convolutional network and an encoder, wherein the convolutional network extracts local features of the channel matrix and the radar perception scatterer position information matrix through a convolutional layer, and reshapes the data types of the two for subsequent input into the encoder, uses a batch normalization layer to accelerate the convergence of the network, and improves the stability of the network through residual connections and layer normalization; the encoder uses cross-modal attention to fuse the radar perception modal information with the communication pilot modal information, and then uses dual attention: feature channel attention and spatial channel attention to fuse and extract features of the channel matrix data through different channel and spatial dimensions, thereby enhancing the extraction and utilization of important features therein; the deep convolutional neural network model is trained and tested; the noisy channel matrix and perception matrix are input into the deep convolutional neural network model, and an estimated channel matrix is ​​output. The present invention utilizes a cross-modal attention mechanism to fully utilize the position information of perceived scatterers, enhances feature extraction and representation capabilities, and significantly improves the accuracy and stability of channel estimation.

[0163] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration, characterized by: include: Step 1: Obtain training and test datasets based on the channel characteristics of the ultra-large-scale MIMO ISAC system, the construction of a radar target perception model, and the preprocessing of the signals received by the base station; Step 2: Build a deep convolutional neural network model based on convolutional layers and cross-modal attention encoders. The convolutional layers are used to extract local features of the noisy channel vector and radar array response vector, and reshape their data types. The encoder is used to fuse the radar array response vector and the noisy channel vector data through cross-modal attention and use dual attention to fuse and extract features from the vector data in different channel and spatial dimensions, thereby enhancing the extraction and utilization of important features. Step 3: Use the training dataset and test dataset to train and test the deep convolutional neural network model; Step 4: Input the noisy channel vector and radar array response vector into the deep convolutional neural network model obtained in step 3, and output the estimated channel matrix.

2. The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration according to claim 1 is characterized in that The step 1 includes the following sub-steps: Step 1.1: Model the channel characteristics in the ultra-large-scale MIMO ISAC system: Assume that the base station of an XL-MIMO ISAC system with M antennas serves a single-antenna user. There are L paths between the base station and the user. The radar senses K targets. There is partial overlap between the K radar-sensed targets and the L communication scatterers. Let the overlap number be X, 0≤X≤min(K,L), the sensed communication scatterer position information be R, and the arrival angle information be AoA. Thus, the radar array response vector containing the X overlapping scatterer position information can be calculated. The overlapping scatterer information is used as auxiliary information for the communication channel estimation task to perform ISAC joint channel estimation. It is assumed that the position information of the radar-perceived target can be estimated using radar array processing technology. The channel characteristic model in the ultra-large-scale MIMO ISAC system is: y=hx+η Where y is the uplink received signal received by the base station, h is the noise-free real channel vector between the base station and the user, x represents the pilot signal transmitted by the user in a time slot, and η represents additive Gaussian noise; When the scattering distance is less than the Rayleigh distance D Ray When , the near-field channel model is: in, λ is the carrier wavelength, represents the array response vector of the base station antenna array to the lth path, φ l represents the azimuth of the lth path, g l is the path gain, r l represents the distance from the lth scatterer to the center of the base station antenna array, represents the distance between the lth scatterer and the mth antenna of the base station, where d is the distance between adjacent antennas, and the parameter Step 1.2: Build radar perception target model: Radar sensing is used to detect the presence of a target and estimate the target's AoA and distance r k , radar cross section RCS parameters, so in the pth DP symbol duration of radar target detection, the sensing radar BS transmits a The corresponding received signal is expressed as: in, is the radar receiving signal, is the radar channel matrix, Dedicated pilot symbol With variance Additive Gaussian white noise; The radar channel matrix depends on the target's AoA, distance r k and radar cross section RCS, modeled as: in r k ,and are the AoA, distance and RCS of the k-th target, is the radar array response vector containing the position information of X overlapping scatterers; For a half-wavelength spatially uniform linear array, for: Step 1.3: Preprocessing of the signal received by the base station: Assuming that the power of the pilot signal x is known to be P, the uplink received signal y received by the base station is simplified to: Based on the above formula, the least squares estimate complex vector of the noise-free true channel vector h between the base station and the user is: in, is the noisy channel vector; Step 1.4: Based on steps 1.1-1.3, the channel vector with noise Radar array response vector containing X overlapping scatterer position information is the input data, the noise-free true channel vector h is the output label, and the training data set and test data set are generated.

3. The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration according to claim 1 is characterized in that The input of the deep convolutional neural network model in step 2 is the tensor H c-input 、H s-input : The noisy channel vector Expand to: (Z,M,1), then split it and connect it, reshape it into a The tensor H c-input ; Expand the radar array response vector containing X overlapping scatterer position information to: (Z, M, 1 × X), then split and connect it and reshape it into a The tensor H s-input ; Where Z is the number of samples in the training dataset, and M is the number of antennas configured in the base station of the XL-MIMO ISAC system.

4. The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration according to claim 1 is characterized in that The deep convolutional neural network model includes a two-dimensional convolutional layer 1, a two-dimensional convolutional layer 2, and an encoder 1, an encoder 2, two-dimensional convolutional layers 3 to 5 and a layer normalization and residual connection module connected in sequence; The two-dimensional convolution layer 1 and the two-dimensional convolution layer 2 respectively process the input tensor H of the deep convolutional neural network model. c-input 、H s-input After extracting the features, the tensors H1, H2 are reshaped to obtain sizes of M×64, where M is the number of antennas of the base station in the XL-MIMO system; H1 and H2 are input into encoder 1 to obtain the output H4. H4 and H2 are then input into encoder 2 to obtain the output tensor H5. After reshaping, H5 is input into two-dimensional convolution layers 3 to 5 to obtain H6. H6 is input to the layer normalization and residual connection module to transform the input H c-input Subtracting H6 gives output H output .

5. The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration according to claim 4 is characterized in that: A batch normalization layer is added after each two-dimensional convolutional layer to accelerate the network training process; a rectified linear unit activation function is used on the output of each two-dimensional convolutional layer to introduce nonlinear features to make the network more expressive; a dropout layer is introduced between each convolutional layer to randomly inactivate a certain proportion of neurons to prevent network overfitting.

6. The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration according to claim 1 is characterized in that The structures of the encoder 1 and the encoder 2 are the same, and both include a cross-modal attention mechanism module, a feature channel attention mechanism module, a spatial channel attention mechanism module, two fully connected layers, a dropout layer and a normalization layer. In the encoder 1, the cross-modal attention mechanism module generates a reconstruction matrix with the same dimension as the tensors H1 and H2 through the multi-head mutual attention mechanism, and sums the reconstruction matrix generated by the multi-head mutual attention mechanism with H1 and H2 element by element to obtain the final output tensor H m , size is M×64, H m represents the weighted sum of all high-dimensional feature maps; The feature channel attention mechanism module applies a multi-head self-attention mechanism to generate and tensor H m The reconstruction matrix with the same dimension is finally combined with the reconstruction matrix generated by the multi-head self-attention mechanism m Perform element-by-element summation to obtain the final output tensor H c , size is M×64, H c represents the weighted sum of all high-dimensional feature maps; The spatial channel attention mechanism module uses a multi-head self-attention mechanism to c Transpose to exchange dimensions and get the attention output matrix. Finally, transpose the attention output matrix again and compare it with H one by one. c Sum and get the output tensor H3, which is the weighted sum of all position features and the original features of each position, with a size of M×64; The tensor H3 passes through two fully connected layers, a Dropout layer, and a normalization layer in sequence to obtain H4 and output it.

7. The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration according to claim 6 is characterized in that: The formula of the multi-head mutual attention mechanism is: head i =Attention(Q i ,K i ,V i ),i=1,...,h, Q i =QW i Q ,K i =KW i K ,V i =VW i V Among them, h represents the number of "heads", head i Represents the attention matrix obtained by the i-th "head", Attention(Q i ,K i ,V i ) is the attention mechanism formula corresponding to the i-th "head"; Q i , K i and V i For Q, K, V through W iQ 、W iK 、W iV Represents the query, key and value, Q, K, V are the matrices of query, key and value respectively, Represents multiple weight matrix groups corresponding to different heads in the multi-head mutual attention mechanism.

8. The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration according to claim 1 is characterized in that The loss function used in step 3 is: in, is the estimated channel matrix and the real channel matrix, θ is the network model parameter, N tr Indicates the number of training samples in the training data set, and the superscript (n) indicates the nth sample. represents the square of the Frobenius norm of the matrix.

9. The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration according to claim 1, characterized in that In step 3, the Adam optimizer is used to optimize the network model. The Adam optimizer combines momentum and adaptive learning rate technology to accelerate convergence. The update rule is as follows: Among them, m t is the momentum vector at the tth iteration; v t is the squared gradient momentum at the tth iteration; since in the initial stage of training, m t and v t The initial value is 0. and is the momentum vector and squared gradient momentum at the tth iteration after correction; β1 is the attenuation coefficient of the momentum term; β2 is the attenuation coefficient of the squared gradient momentum; is the gradient of the loss function L(θ) with respect to the model parameters θ; α is the learning rate; θ t is the parameter vector of the model at the tth iteration; ∈ is a constant.

10. The near-field channel estimation method based on perception-assisted mutual attention in ultra-large-scale MIMO synaesthesia integration according to claim 1, characterized in that: In the test phase, step 3 calculates the estimated channel matrix output by the model With the real channel matrix The normalized mean square error NMSE between is used to evaluate the channel estimation result of the model. The smaller the normalized mean square error, the more accurate the channel estimation result. The calculation formula of the normalized mean square error is: in For expectations, represents the square of the Frobenius norm of the matrix.

Citation Information

Cited By

  • Data classification method based on self-attention mechanism in ISCC system

    CN121434913A