Orthogonal time-frequency-space double-fractional channel estimation method and system based on deep neural network
Through the encoder-decoder convolutional neural network of deep neural network, the channel estimation problem of orthogonal time-frequency air conditioning system in high mobility scenarios is solved, and high-precision and low-complexity channel estimation is achieved, which is suitable for vehicle networking and satellite communications.
Patent Information
- Application Number
- CN202510494084.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-19
- Publication Date
- 2025-07-11
AI Technical Summary
In the orthogonal time-frequency air conditioning system in high mobility scenarios, fractional delay and Doppler shift lead to channel response sparseness failure. Traditional methods cannot effectively estimate phase continuity, and the calculation complexity is high, so they cannot meet the real-time communication needs.
The encoder-decoder convolutional neural network is adopted based on deep neural networks, and the multi-resolution interactive module and phase-constrained mixed loss function is used to realize end-to-end dual-fraction channel response estimation, reducing the computational complexity and improving the estimation accuracy.
It realizes high-precision and low-complexity channel estimation, supports real-time communication requirements, and is suitable for high-mobility scenarios such as Internet of Vehicles and satellite communications, reducing mean square error in the signal-to-noise ratio range.
Smart Images

Figure CN120301748A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a channel estimation method and system for an orthogonal time-frequency space modulation system in a high-mobility scenario. Background Art
[0002] Orthogonal Time Frequency Space (OTFS) modulation maps signals to the delay-Doppler domain to utilize the quasi-static characteristics and sparse characteristics represented in this domain, etc., effectively combating Doppler spread and multipath interference in a high-speed mobile environment, and becoming a key technology for scenarios such as 6G vehicle-to-everything (V2X) and satellite communication. Its core system architecture and channel estimation problem are modeled as follows:
[0003] (1) First, complete the transformation from the Delay-Doppler (DD) domain symbol to the Time-Frequency (TF) domain, that is, the Inverse Symplectic Finite Fourier Transform (ISFFT):
[0004]
[0005] where N is the number of time slots, M is the number of subcarriers, and x[k, l] ∈ C N×M is the DD domain symbol matrix; X[n, m] ∈ C N×M is the TF domain symbol matrix.
[0006] (2) Convert the obtained TF domain symbol into a continuous time-domain waveform available for transmission by the radio frequency module. This process is the Heisenberg Transform (HT):
[0007]
[0008] where T is the symbol period, and Δ f is the subcarrier spacing.
[0009] (3) The process of the signal passing through a channel with double-dispersion characteristics. The DD domain impulse response of this process is expressed as: Also known as the original channel response, where is the path gain, τ i ∈(0, τ max ) is the time delay, and ν i ∈(0, ν max ) is the Doppler frequency shift.
[0010] (4) Through this channel, the received signal is converted into a continuous TF-domain signal through the Wigner Transform (WT) and discretized. Finally, the transceiver relationship in the DD domain satisfies:
[0011]
[0012] where h w [·] is also called the effective channel response, P is the number of scattering paths, including w ν (·) and w τ (·) are discrete window functions for digital reception, used to achieve the matching of the continuous DD domain to the discrete DD domain. Based on different model assumptions, w v (·) and w τ (·) have different mathematical expressions. For the actual signal transmission process, that is, the sampling points are generally not on the integer multiple grid points of delay and Doppler, which is the delay-Doppler double fractional case that this technical solution focuses on solving:
[0013]
[0014] It should be noted that the double integer or the delay-Doppler single fractional model are both degenerations of the double fractional case.
[0015] (5) Finally, discuss the channel estimation task of the double fractional model. First, transform the transceiver relationship in the DD domain into a matrix form, and adopt an embedded pilot scheme with a single pilot plus a guard interval:
[0016]
[0017] x[k p ,l p is the pilot signal, H DD =diag(h i )∈C P×P , N is the DD domain noise term that combines data interference and channel background noise. Therefore, the matrix form of the equivalent channel response can be expressed as:
[0018]
[0019] where H∈C N×M , the core of the channel estimation task is to obtain the estimation of the channel matrix according to Y DD The existing technical solutions face the following challenges:
[0020] 1. Fractional Delay Doppler Shift: In the actual channel, the delay and Doppler parameters usually deviate from the discrete grid, resulting in energy leakage to adjacent grid points and destroying the sparsity of the channel response.
[0021] 2. Phase Rotation Mismatch: Fractional offset causes linear phase changes. Traditional gridding methods (such as Orthogonal Matching Pursuit (OMP), Sparse Bayesian Learning (SBL)) and deep learning methods based on the Mean Squared Error (MSE) loss function have insufficient optimization effects due to the continuous phase characteristics, resulting in discontinuous estimated phases.
[0022] 3. Computational Complexity and Real - time Performance: Traditional iterative algorithms (such as SBL based on oversampling, OMP - like algorithms, etc.) require multiple matrix inversions, with high computational complexity, resulting in too long inference time and unable to meet the immediate communication requirements of high - moving - speed double - diffusion channels. Existing deep - learning solutions generally introduce residual networks on the basis of traditional iterative estimation schemes to optimize the estimation accuracy, such as the OMP + ResNet cascade framework. These methods further increase the system complexity and separate the sparse reconstruction and denoising tasks, and cannot achieve end - to - end estimation. Summary of the Invention
[0023] The object of the present invention is to propose an orthogonal time - frequency - space double - fractional channel estimation method and system based on a deep neural network (denoted as DFCE - Net) to achieve high - precision and low - complexity estimation of the fractional delay - Doppler channel in the OTFS system.
[0024] The orthogonal time - frequency - space double - fractional channel estimation method based on a deep neural network provided by the present invention includes obtaining a delay - Doppler (DD) domain received signal from the receiving end of an orthogonal time - frequency - space (OTFS) transmission system; using an encoder - decoder convolutional neural network (CNN) to estimate the double - fractional channel response from the received signal; the neural network includes a symmetric encoder module and a decoder module, the encoder module extracts multi - scale features through convolution and downsampling, and the decoder module restores the channel response through transposed convolution and skip connections; outputting the estimated double - fractional channel response to a signal equalization module for subsequent signal detection.
[0025] Wherein:
[0026] The received signal is input in complex form, and its real and imaginary parts are separated into two-channel data, and the feature map size consistency is maintained by symmetric padding;
[0027] The double fractional channel response is a two-dimensional complex matrix, and its dimension matches the delay-Doppler grid of the OTFS system;
[0028] The encoder module includes at least two levels of downsampling convolutional blocks, and each level includes: two 3×3 convolutional layers; a batch normalization (BatchNorm) layer; a ReLU activation function; a 2×2 max pooling layer;
[0029] The decoder module includes at least two levels of upsampling convolutional blocks, and each level includes: a 2×2 transposed convolutional layer; skip connection feature splicing corresponding to the encoder level; two 3×3 convolutional layers; a batch normalization (BatchNorm) layer; a ReLU activation function;
[0030] The encoder-decoder neural network further includes a multi-resolution interaction module for dynamically fusing local details and global context features to solve the energy leakage problem caused by fractional time-delay Doppler shift;
[0031] The training of the neural network uses a phase-constrained hybrid loss function, including: mean square error loss; sparsity constraint (L1 regularization); phase continuity loss based on cosine similarity;
[0032] The neural network realizes real-time inference through GPU acceleration, and the single estimation delay is less than 2 milliseconds; this neural network supports model compression and quantization, adapts to FPGA or embedded device deployment, and the model size is compressed to less than 500KB.
[0033] Based on the above double fractional channel estimation method, the present invention also provides an orthogonal time-frequency-space double fractional channel estimation system based on a deep neural network, including: a signal receiving module for obtaining the delay-Doppler domain received signal of the OTFS system; an encoder-decoder convolutional neural network module configured to estimate the double fractional channel response from the signal; an output module for transmitting the estimation result to the signal equalization module; wherein the neural network module includes a multi-resolution interaction module, skip connections and a phase-constrained hybrid loss function; this system is deployed on a vehicle-mounted communication terminal or a satellite communication device and supports real-time channel estimation in 5G / 6G high-mobility scenarios.
[0034] Wherein:
[0035] The encoder-decoder neural network is adapted to the following channel types: time-delay Doppler double-integer channel: both the delay and the Doppler shift are integer multiples of the grid interval; single-fraction channel: either the delay or the Doppler is a fractional multiple of the grid interval; double-fraction channel: both the delay and the Doppler are fractional multiples of the grid interval;
[0036] The encoder-decoder neural network supports simplified layer structures, including the following cases: reducing the number of layers in the encoder or decoder; removing some convolutional layers or pooling layers; provided that the skip connection mechanism is retained to achieve diffusion texture feature learning;
[0037] The phase-constrained hybrid loss function can independently adjust its constituent terms, including the following cases: only using the mean square error loss and the phase continuity loss; only using the mean square error loss and the sparsity constraint; the neural network structure remains unchanged.
[0038] The present invention realizes end-to-end learning of the channel response in the delay-Doppler domain through an encoder-decoder convolutional neural network; the network uses a multi-resolution interaction module to dynamically fuse local details and global context features, solves the energy leakage problem caused by fractional time delay and Doppler frequency shift, and combines the mean square error loss, sparsity constraint and phase continuity loss to jointly optimize the loss function, significantly reducing the phase rotation mismatch error; the present invention supports lightweight design, training and deployment with low latency, is applicable to high-mobility scenarios such as vehicle-to-everything (V2X), satellite communication, and high-speed rail communication, and the normalized mean square error is reduced by 3-8 dB compared with traditional methods (OMP / SBL) in the 0-30 dB SNR range, and has robustness in multi-path dense scenarios. Description of the Drawings
[0039] Figure 1 It is a structural diagram of an orthogonal time-frequency-space OTFS transceiver system processed by the present invention.
[0040] Figure 2 It is a schematic diagram of the neural network structure designed by the present invention.
[0041] Figure 3 It is a topological diagram of the neural network layer structure of the present invention.
[0042] Figure 4 It is a normalized mean square error NMSE performance diagram of the present invention under the channel condition of 0-30 dB signal-to-noise ratio.
[0043] Figure 5 It is a NMSE performance diagram of the present invention under the channel condition of 2-10 scattering paths. Detailed Implementation Manner
[0044] The present invention will be further described below with reference to the accompanying drawings.
[0045] The orthogonal time-frequency-space double-fractional channel estimation method based on a deep neural network provided by the present invention has the following specific operation steps:
[0046] Step (1), randomly generate channel parameters P, h i , τ i , v i , which respectively represent the number of channel scattering paths, the single-path equivalent complex gain value, the single-path delay parameter, and the single-path Doppler parameter, and obtain the complex matrix H according to the channel model; then pass the randomly generated pilot complex signal through the double-fractional OTFS system model, and superimpose additive white Gaussian noise to obtain the delayed Doppler domain grid receiving complex signal matrix Y DD , and thus obtain a data set for neural network training, so as to realize that by inputting Y DD the estimation result of the channel matrix can be obtained Divide the data set into a training set, a validation set, and a test set according to a ratio; the expressions of H and Y DD are shown in the background section.
[0047] Step (2), neural network structure design. First is the multi-scale encoder-decoder architecture. The design of the l-th level feature map of the encoder is as follows:
[0048]
[0049] where * represents the convolution operation, the kernel size is 3×3, is 2×2 max pooling, and the number of channels increases from 64 to 128; BN(·) represents batch normalization, which is used to accelerate convergence; ReLU is the rectified linear unit activation function; is the learnable convolution kernel parameter of the encoder, l is the encoder block number, and can take values of 1 or 2; C in is the number of input channels, which is determined by the output of the previous layer; C out is the number of output channels, which determines the depth of the feature map of this layer;
[0050] The corresponding l-th level feature generation module of the decoder is designed as follows:
[0051]
[0052] where, is the 2×2 transposed convolution with a stride of 2, Crop(·) represents feature cropping; Concat[·,·] represents channel dimension concatenation; is the decoder convolution kernel parameter matrix, and the matrix structure is the same as that of the encoder convolution kernel;
[0053] The connection layer between the encoding and decoding is designed as a cycle of two convolutional layers, BN, and ReLU to connect the output of the second encoding module and the input of the first decoding module; Skip connection 1 performs a Concat operation on the ReLU output of the encoding block 1 and the transposed convolution output of the decoding block 2 to establish a direct connection channel; Skip connection 2 performs a Concat operation on the ReLU output of the encoding block 2 and the transposed convolution output of the decoding block 1 to establish a direct connection channel; The skip connection realizes multi-scale feature fusion, that is, it retains high-dimensional feature information such as diffusion texture.
[0054] Step (3), loss function design; Based on the physical characteristics of the time-delay Doppler domain channel response, a composite loss function with phase preservation and sparse constraints is constructed; This design breaks through the amplitude-phase coupling optimization limit of the traditional mean square error (MSE), and realizes high-precision channel estimation through a decoupled learning mechanism; The composite loss function is defined as:
[0055] L Total =L MSE +λ1L Sparse +λ2L Phase , (3)
[0056] where λ1 and λ2 are the sparse term coefficient and the phase loss term coefficient respectively; The basic mean square error term L MSE maintains the global convergence of the amplitude response estimation:
[0057]
[0058] where and H[k, l] are the elements of the channel estimation matrix and the corresponding true value elements respectively, and |·| represents the calculation of the complex amplitude; The L1 regularization term is introduced to encourage the sparse output of the model:
[0059]
[0060] To eliminate the phase continuity breakdown phenomenon in the low signal-to-noise ratio region, a phase similarity loss term is introduced:
[0061]
[0062] where N is the number of time slots, M is the number of subcarriers, and [k, l] is the receiving grid index in the delay Doppler domain; This loss term extracts the real part of the complex signal through the Re{·} operator, and H * is the matrix conjugate; The phase loss term establishes a pure phase gradient update channel decoupled from the amplitude, and its gradient form effectively avoids the problem of phase estimation failure of weak path components; Finally, it effectively guarantees the continuity and accuracy of the model-estimated phase response.
[0063] The actual application process of the DFCE-Net described in the present invention in the OTFS system refersFigure 1 As shown in Figure 1 , the receiving end first performs time-frequency domain transformation and discretization processing on the radio frequency signal to generate a delay-Doppler domain received signal matrix. This signal matrix is converted into a two-channel real tensor through real and imaginary part separation and input into the DFCE-Net network. The encoder module extracts the spatial correlation features of the signal through multi-level convolution and downsampling operations, capturing the path energy distribution characteristics at different resolutions. The decoder module then restores the channel response details through transposed convolution and skip connections, solving the energy diffusion problem caused by fractional delay Doppler. The finally output complex channel matrix is directly input into the subsequent equalization module to complete signal detection. The entire process adopts a full data-driven method, avoiding the computational bottlenecks of grid parameter matching and iterative optimization in traditional methods, and is especially suitable for real-time tracking of time-varying channels.
[0064] Figure 2 The mathematical operation processes of the encoder and decoder modules are described in detail. The generation of the l-th level features in the encoder follows the formula:
[0065]
[0066] where and are 3×3 convolution kernels, and represents the max pooling operation. The generation of the l-th level features in the decoder follows the formula:
[0067]
[0068] where is a 2×2 transposed convolution, and Crop(·) realizes the spatial alignment of the feature maps of the encoder and decoder.
[0069] The skip connection injects the local texture information of the encoder into the global context of the decoder through channel concatenation, effectively learning the diffusion features caused by energy leakage. The output layer uses a 1×1 convolution to map the number of feature channels to the real and imaginary components, and reconstructs the complex channel matrix through regression.
[0070] Figure 3The network architecture topology of DFCE-Net is shown; the input layer receives the real and imaginary part separated signal with a dimension of N×M×2; the encoder module contains two levels of downsampling units: the first level extracts local detail features through two 3×3 convolutions (64 filters), and compresses the spatial dimension to N / 2×M / 2×64 through 2×2 max pooling; the second level captures the global energy distribution through the convolution operation with 128 filters, and outputs the feature map of N / 4×M / 4×128; the connection layer uses 256 filters to enhance the feature expression ability; the decoder module restores the resolution through two levels of upsampling units: the first level transposed convolution expands the feature map to N / 2×M / 2×128, and after splicing with the features of the second level of the encoder, the texture is refined through two convolutions; the second level transposed convolution restores to the input size of N×M×64, and after fusing with the features of the first level of the encoder, the final output is generated; the output layer realizes the channel dimension compression through 1×1 convolution, and directly regresses the real and imaginary part values; the skip connection structure refers to Figure 2 。
[0071] Figure 4 The normalized mean square error performance of DFCE-Net under different signal-to-noise ratios is shown; the simulation conditions are: under the simulation conditions of a subcarrier spacing of 15 kHz, a carrier frequency of 5 GHz, N = 20, and M = 32, the Adam optimizer is used, with an initial learning rate of 1×10 -4 The initial learning rate, a batch size of 32, a validation round of 200, 25,200 simulation samples for training data, and the training set, validation set, and test set are set according to 8:1:1; as shown in the figure, the method of the present invention achieves a NMSE gain of 3-8 dB compared with the traditional OMP / SBL algorithm in the 0-30 dB SNR range.
[0072] Figure 5 The channel estimation results in a multi-path dense scenario are compared, and the simulation conditions are as Figure 4 ; for various scattering environments with the number of scattering paths ranging from 2 to 10, DFCE-Net has a NMSE performance gain superior to that of the traditional OMP / SBL algorithm.
[0073] The channel estimation network DFCE-Net for the dual fractional OTFS system proposed by the present invention significantly improves the channel estimation accuracy while meeting the requirements of the actual system complexity and real-time performance, and has good robustness.
[0074] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the premise of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An orthogonal time-frequency-space double fractional channel estimation method based on a deep neural network, characterized in that, Including obtaining a delay-Doppler (DD) domain received signal from the receiving end of an orthogonal time-frequency-space (OTFS) transmission system; using an encoder-decoder convolutional neural network (CNN) to estimate a bi-fractional channel response from the received signal; the neural network includes a symmetric encoder module and a decoder module, the encoder module extracts multi-scale features through convolution and downsampling, and the decoder module restores the channel response through transposed convolution and skip connections; outputting the estimated bi-fractional channel response to a signal equalization module for subsequent signal detection; wherein: The received signal is input in complex form, its real and imaginary parts are separated into dual-channel data, and the feature map size consistency is maintained through symmetric padding. The bi-fractional channel response is a two-dimensional complex matrix, and its dimensions match the delay-Doppler grid of the OTFS system.
2. The orthogonal time-frequency-space bi-fractional channel estimation method according to claim 1, characterized in that: The encoder module includes at least two levels of downsampling convolutional blocks, and each level includes: two 3×3 convolutional layers; a batch normalization layer; a ReLU activation function; a 2×2 max pooling layer. The decoder module includes at least two levels of upsampling convolutional blocks, and each level includes: a 2×2 transposed convolutional layer; splicing of skip connection features corresponding to the encoder at the same level. Two 3×3 convolutional layers; a batch normalization layer; a ReLU activation function.
3. The orthogonal time-frequency-space double fractional channel estimation method according to claim 1, characterized in that The encoder-decoder neural network further includes a multi-resolution interaction module for dynamically fusing local details and global context features to solve the energy leakage problem caused by fractional time delay-Doppler shift.
4. The orthogonal time-frequency-space double fractional channel estimation method according to claim 1, characterized in that The training of the neural network uses a phase-constrained hybrid loss function, including: mean squared error loss, sparsity constraint, and phase continuity loss based on cosine similarity.
5. The orthogonal time-frequency-space double fractional channel estimation method according to claim 1, characterized in that The neural network realizes real-time inference through GPU acceleration, and the single estimation delay is less than 2 milliseconds; this neural network supports model compression and quantization, and is adapted to be deployed on FPGA or embedded devices, and the model size is compressed to less than 500KB.
6. An orthogonal time-frequency-space bi-fractional channel estimation system based on a deep neural network for the method according to any one of claims 1-5, characterized in that, Including: A signal receiving module for obtaining the delay-Doppler domain received signal of the OTFS system; an encoder-decoder convolutional neural network module configured to estimate the bi-fractional channel response from the signal; an output module for transmitting the estimation result to the signal equalization module; wherein the neural network module includes a multi-resolution interaction module, skip connections, and a phase-constrained hybrid loss function; the system is deployed on a vehicle-mounted communication terminal or a satellite communication device and supports real-time channel estimation in 5G / 6G high-mobility scenarios.
7. The orthogonal time-frequency-space double fractional channel estimation system according to claim 6, wherein The encoder-decoder neural network is adapted to the following channel types: time-delay-Doppler double-integer channel: both the delay and the Doppler shift are integer multiples of the grid interval; single-fractional channel: one of the delay or the Doppler is a fractional multiple of the grid interval; bi-fractional channel: both the delay and the Doppler are fractional multiples of the grid interval.
8. The orthogonal time-frequency-space double fractional channel estimation system according to claim 6, characterized in that, The encoder-decoder neural network supports simplification of the layer structure, including the following situations: reducing the number of levels of the encoder or decoder; removing some convolutional layers or pooling layers. Provided that the skip connection mechanism is retained to achieve diffusion texture feature learning.
9. The orthogonal time-frequency-space double fractional channel estimation system according to claim 6, characterized in that, The phase-constrained hybrid loss function can independently adjust its constituent terms, including the following cases: only using the mean squared error loss and the phase continuity loss; only using the mean squared error loss and the sparsity constraint; the neural network structure remains unchanged.