A channel prediction method based on multi-scale complex time-frequency feature fusion
Patent Information
- Application Number
- CN202610759998.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-01
AI Technical Summary
[0005]因此,本发明提供了一种基于多尺度复数时频特征融合的信道预测方法解决固定尺度的特征提取机制难以兼顾时间分辨率与频率分辨率的矛盾需求——快时变信道需窄时间窗,强频选信道需高频率分辨率,而现有方法往往顾此失彼,即便部分工作尝试引入复值神经网络,但在跨尺度或跨特征融合阶段仍常退化为实数运算,未能有效利用复共轭内积等操作捕捉不同尺度间的相位一致性与能量协同关系问题
[0016] The beneficial effects of this invention are as follows: Based on the channel coherence time and coherence bandwidth, a learnable multi-scale complex time-frequency basis function is constructed, and a complex convolution transformation is performed on the historical complex value channel state information sequence. This achieves adaptive alignment between the time-frequency analysis window and the channel's physical dynamic characteristics, overcoming the mismatch problem of traditional fixed basis functions in high-speed or broadband scenarios. This results in obtaining high-fidelity, multi-resolution complex time-frequency feature maps. The energy entropy of each scale feature map is calculated, and activation confidence is generated accordingly. Scales with high information concentration are dynamically selected to form a sparse multi-scale feature set. This not only reduces subsequent computational complexity and memory overhead but also enables the system to adaptively focus on key time-frequency structures according to the channel environment, improving robustness and energy efficiency.
Smart Images

Figure CN122678818A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, and in particular to a channel prediction method based on multi-scale complex time-frequency feature fusion. Background Technology
[0002] With the evolution of 5G / 6G and the increasing demands for low latency and high reliability from services such as the Industrial Internet and the Internet of Vehicles, wireless links require more refined adaptive scheduling and beamforming. This relies on the accurate acquisition and timely maintenance of channel state information. 3GPP, in TR 38.901, provides a standardized channel modeling and evaluation framework covering 0.5–100 GHz, clearly characterizing the impact of factors such as multipath, shadowing fading, and Doppler on the time-varying characteristics of the channel, providing a unified benchmark for system design. Classical wireless communication theory and textbooks generally point out that mobility and scattering environments cause the channel to change at multiple time scales. Once CSI lags, it will weaken the effectiveness of adaptive modulation and coding, power control, and precoding. Therefore, "predicting the channel in several future time slots" is one of the important means to improve link efficiency.
[0003] Fixed-scale feature extraction mechanisms struggle to balance the conflicting demands of time resolution and frequency resolution—fast time-varying channels require narrow time windows, while strong frequency-selective channels require high frequency resolution. Existing methods often compromise on one aspect while neglecting the other. Even when some work attempts to introduce complex-valued neural networks, they often degenerate into real-number operations during cross-scale or cross-feature fusion stages, failing to effectively utilize operations such as complex conjugate inner products to capture phase consistency and energy synergy between different scales. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a channel prediction method based on multi-scale complex time-frequency feature fusion to solve the contradiction between time resolution and frequency resolution in fixed-scale feature extraction mechanisms. Fast time-varying channels require narrow time windows, while strong frequency-selective channels require high frequency resolution. Existing methods often compromise on one aspect while addressing the other. Even though some works attempt to introduce complex-valued neural networks, they often degenerate into real number operations in the cross-scale or cross-feature fusion stage, failing to effectively utilize operations such as complex conjugate inner products to capture the phase consistency and energy coordination relationship between different scales.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a channel prediction method based on multi-scale complex time-frequency feature fusion, which includes: acquiring a historical complex-valued channel state information sequence of a target communication link, estimating the channel coherence time and coherence bandwidth from the historical complex-valued channel state information sequence, and obtaining physical parameters; Based on physical parameters, a learnable multi-scale complex time-frequency basis function is constructed, and the historical complex value channel state information sequence is transformed to obtain a complex time-frequency feature map. The activation confidence of each scale is calculated based on the complex time-frequency feature maps at multiple scales, and the complex time-frequency feature maps at some scales are selected based on the activation confidence to form a sparse multi-scale feature set. For each scale of complex time-frequency features in the sparse multi-scale feature set, query vector, key vector and value vector are extracted respectively. Cross-scale attention weights are calculated in the complex domain by the conjugate inner product of query vector and key vector to characterize the amplitude-phase coupling relationship between different scales. By using cross-scale attention weights to perform weighted summation on the value vectors in the sparse multi-scale feature set, complex context features are obtained, and based on the complex context features, prediction results of complex numerical channel state information for future time slots are generated. The training objective is constructed based on the predicted results of the complex-valued channel state information and the actual feedback channel state information. The multi-scale complex time-frequency basis function, the activation confidence calculation logic, and the complex mapping parameters are jointly updated.
[0007] As a preferred embodiment of the channel prediction method based on multi-scale complex time-frequency feature fusion described in this invention, the specific steps for obtaining the historical complex channel state information sequence of the target communication link are as follows: At the receiving end, using the pilot signals periodically transmitted by the transmitting end, the complex-valued channel response on each orthogonal frequency division multiplexing subcarrier within multiple consecutive historical time units is calculated by the least squares estimation method. The calculated complex-valued channel responses are arranged in chronological order and subcarrier index order to form a two-dimensional data structure; Each row of the two-dimensional data structure corresponds to a historical time unit, each column corresponds to a subcarrier, and each element consists of a real part and an imaginary part, reflecting the channel amplitude attenuation and phase rotation under the combination of time unit and subcarrier, respectively.
[0008] As a preferred embodiment of the channel prediction method based on multi-scale complex time-frequency feature fusion described in this invention, the specific steps for estimating the channel coherence time and coherence bandwidth from historical complex channel state information sequences to obtain physical parameters are as follows: The normalized autocorrelation function is calculated by sliding along the time dimension for the historical complex-valued channel state information sequence, and the autocorrelation value is first decayed to 0. The corresponding time interval is determined as the channel coherence time. ; The frequency domain cross-correlation coefficient is calculated for the complex channel responses of adjacent subcarriers, and the reciprocal of the subcarrier spacing corresponding to the first drop of the cross-correlation coefficient to 0.5 is determined as the channel coherence bandwidth. ; Therefore, from and The physical parameters that constitute it.
[0009] As a preferred embodiment of the channel prediction method based on multi-scale complex time-frequency feature fusion described in this invention, the steps of constructing a learnable multi-scale complex time-frequency basis function based on physical parameters and transforming the historical complex value channel state information sequence to obtain a complex time-frequency feature map are as follows: Two learnable adjustment factors are introduced for each scale, one to control the widening of the scale time window and the other to control the modulation position of the scale in the frequency dimension. The degree of time window widening is inversely proportional to the channel coherence bandwidth, meaning that the more drastic the channel frequency changes, the narrower the corresponding time window. The frequency modulation position is inversely proportional to the channel coherence time, so the faster the channel time changes, the higher the corresponding modulation frequency. Based on the adjusted time-frequency parameters, a complex exponential basis function for Gaussian window modulation is constructed. The complex time-frequency feature map is generated by performing a complex convolution transformation along the time dimension on the historical complex numerical channel state information sequence using complex exponential basis functions. Repeat the process for all preset scales to obtain a set of complex time-frequency feature maps covering different time-frequency resolutions, which are then used as multi-scale representation inputs to subsequent processing stages.
[0010] As a preferred embodiment of the channel prediction method based on multi-scale complex time-frequency feature fusion described in this invention, the specific steps are as follows: Calculating the activation confidence of each scale based on complex time-frequency feature maps of multiple scales, and selecting complex time-frequency feature maps of some scales according to the activation confidence to form a sparse multi-scale feature set. Calculate the energy distribution at all time-frequency locations for the complex time-frequency feature map at the k-th scale, and normalize the energy distribution to a probability mass function; The energy concentration scale is calculated based on the probability mass function, and Shannon entropy is used as a measure. The lower the entropy value, the more concentrated the energy. By introducing a learnable positive real parameter, the Shannon entropy is mapped to the activation confidence of the scale, so that the scale with more concentrated energy obtains a higher confidence. Calculate the average activation confidence across all scales, and retain the complex time-frequency feature maps corresponding to the scales with activation confidence not lower than the average value, while discarding the feature maps of the remaining scales. The retained complex time-frequency feature maps form a sparse multi-scale feature set, which is used for subsequent cross-scale fusion processing.
[0011] As a preferred embodiment of the channel prediction method based on multi-scale complex time-frequency feature fusion described in this invention, the following steps are taken: query vector, key vector, and value vector are extracted from the complex time-frequency features of each scale in the sparse multi-scale feature set, and cross-scale attention weights are calculated in the complex domain using the conjugate inner product of the query vector and the key vector to characterize the amplitude-phase coupling relationship between different scales. Feature maps of each activation scale Flattened into a vector ; Query vectors are generated using three independent complex-valued linear mappings. Key vector AND value vector ; For any two activation scales and The cross-scale attention weights are calculated using the following expression: ; in, To flatten the feature vectors, , , For complex-valued learnable mapping matrices, For query vector, For key vectors, This indicates the conjugate transpose. This indicates taking the real part of a complex number. For feature dimension, To sum dummy variables, This represents the attention weight.
[0012] As a preferred embodiment of the channel prediction method based on multi-scale complex time-frequency feature fusion described in this invention, the steps of using cross-scale attention weights to perform weighted summation on the value vectors in the sparse multi-scale feature set to obtain complex context features, and generating complex numerical channel state information prediction results for future time slots based on the complex context features, are as follows: For each query scale With its corresponding attention weight For all value vectors The expression is: ; in, For the first The context-enhanced complex numerical feature vector obtained after cross-scale attention fusion at each scale Represents a set All activation scale indexes Perform a summation operation. scale scale Cross-scale attention weights For scale The corresponding value vector; All Concatenate into a joint context vector ; Will Input complex-valued fully connected networks, mapped to the future Each time slot Complex-valued channel state information prediction results for each subcarrier .
[0013] As a preferred embodiment of the channel prediction method based on multi-scale complex time-frequency feature fusion described in this invention, the specific steps of constructing a training target based on the prediction result of the complex numerical channel state information and the actual feedback channel state information, and jointly updating the multi-scale complex time-frequency basis function, the activation confidence calculation logic, and the complex mapping parameters are as follows: Obtaining real channel state information through pilot signals ; based on and The complex mean square error loss is calculated using the following expression: ; in, The loss function for complex-valued channel prediction. The number of future time slots to be predicted. The total number of subcarriers in the system. Indicates the range from 1 to 1 Summing is performed on each future time slot. Indicates the range from 1 to 1 Summing the subcarriers The model predicts the first The first future time slot, the first Complex-valued channel state information on each subcarrier This refers to the actual complex-valued channel state information at the corresponding location obtained through actual measurement of pilot signals.
[0014] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the channel prediction method based on multi-scale complex time-frequency feature fusion as described in the first aspect of the present invention.
[0015] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the channel prediction method based on multi-scale complex time-frequency feature fusion as described in the first aspect of the present invention.
[0016] The beneficial effects of this invention are as follows: Based on the channel coherence time and coherence bandwidth, a learnable multi-scale complex time-frequency basis function is constructed, and a complex convolution transformation is performed on the historical complex value channel state information sequence. This achieves adaptive alignment between the time-frequency analysis window and the channel's physical dynamic characteristics, overcoming the mismatch problem of traditional fixed basis functions in high-speed or broadband scenarios. This results in obtaining high-fidelity, multi-resolution complex time-frequency feature maps. The energy entropy of each scale feature map is calculated, and activation confidence is generated accordingly. Scales with high information concentration are dynamically selected to form a sparse multi-scale feature set. This not only reduces subsequent computational complexity and memory overhead but also enables the system to adaptively focus on key time-frequency structures according to the channel environment, improving robustness and energy efficiency. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of the channel prediction method based on multi-scale complex time-frequency feature fusion in Example 1; Figure 2 This provides an overview of the complex-valued time-frequency fusion prediction framework in Example 2; Figure 3 This is a schematic diagram of learnable multi-scale STFT feature fusion in Example 2; Figure 4 This is a schematic diagram of the complex-valued multi-head self-attention structure in Example 2; Figure 5 This is a schematic diagram of the bidirectional time-frequency cross-domain complex-valued attention structure in Example 2. Detailed Implementation
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0022] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides a channel prediction method based on multi-scale complex time-frequency feature fusion, comprising the following steps: S1. Obtain the historical complex-valued channel state information sequence of the target communication link.
[0023] Furthermore, at the receiving end, using the pilot signals periodically transmitted by the transmitting end, the complex-valued channel response on each orthogonal frequency division multiplexing subcarrier within multiple consecutive historical time units is calculated using the least squares estimation method. The calculated complex-valued channel responses are arranged in chronological order and subcarrier index order to form a two-dimensional data structure; Each row of the two-dimensional data structure corresponds to a historical time unit, each column corresponds to a subcarrier, and each element consists of a real part and an imaginary part, reflecting the channel amplitude attenuation and phase rotation under the combination of time unit and subcarrier, respectively.
[0024] It should be noted that by using pilot signals for least-squares channel estimation and organizing the complex numerical response into a two-dimensional time-frequency structure, the joint selective fading characteristics of the channel in the time and frequency dimensions can be accurately captured. This provides high-fidelity, structured input data for subsequent multi-scale modeling based on physical priors, avoiding prediction bias caused by data distortion or missing dimensions.
[0025] S2. Estimate the channel coherence time and coherence bandwidth from the historical complex value channel state information sequence to obtain the physical parameters.
[0026] Furthermore, a normalized autocorrelation function is calculated by sliding along the time dimension on the historical complex-valued channel state information sequence, and the autocorrelation value is initially decayed to... The corresponding time interval is determined as the channel coherence time. ; The frequency domain cross-correlation coefficient is calculated for the complex channel responses of adjacent subcarriers, and the reciprocal of the subcarrier spacing corresponding to the first drop of the cross-correlation coefficient to 0.5 is determined as the channel coherence bandwidth. ; Therefore, from and The physical parameters that constitute it.
[0027] It should be noted that the method of determining the coherence bandwidth by using the autocorrelation function to decay to the coherence time and the cross-correlation coefficient to drop to 0.5 can adaptively extract physical parameters that match the current communication environment from actual channel data, avoiding reliance on fixed empirical values or idealized assumptions, thereby improving the pertinence and effectiveness of subsequent time-frequency basis function design.
[0028] S3. Construct learnable multi-scale complex time-frequency basis functions based on physical parameters, and transform the historical complex value channel state information sequence to obtain complex time-frequency feature maps.
[0029] Furthermore, two learnable adjustment factors are introduced for each scale, one to control the widening of the scale time window and the other to control the modulation position of the scale in the frequency dimension. The degree of time window widening is inversely proportional to the channel coherence bandwidth, meaning that the more drastic the channel frequency changes, the narrower the corresponding time window. The frequency modulation position is inversely proportional to the channel coherence time, so the faster the channel time changes, the higher the corresponding modulation frequency. Based on the adjusted time-frequency parameters, a complex exponential basis function for Gaussian window modulation is constructed. The complex time-frequency feature map is generated by performing a complex convolution transformation along the time dimension on the historical complex numerical channel state information sequence using complex exponential basis functions. Repeat the process for all preset scales to obtain a set of complex time-frequency feature maps covering different time-frequency resolutions, which are then used as multi-scale representation inputs to subsequent processing stages.
[0030] It should be noted that by dynamically adjusting the multi-scale basis functions in a way that the time window width is inversely proportional to the coherent bandwidth and the modulation frequency is inversely proportional to the coherent time, the time-frequency resolution of each scale can be automatically adapted to the actual dynamic characteristics of the channel. In high-speed mobile scenarios, a narrow time window is enabled to capture rapid changes, and in broadband scenarios, a high frequency resolution is enabled to resolve fine frequency selection, thereby achieving physically guided adaptive feature extraction and overcoming the mismatch problem of traditional fixed basis functions.
[0031] S4. Calculate the activation confidence of each scale based on the complex time-frequency feature maps of multiple scales, and select some complex time-frequency feature maps of some scales according to the activation confidence to form a sparse multi-scale feature set.
[0032] Furthermore, the energy distribution at all time-frequency locations is calculated for the complex time-frequency feature map at the k-th scale, and the energy distribution is normalized to a probability mass function; The energy concentration scale is calculated based on the probability mass function, and Shannon entropy is used as a measure. The lower the entropy value, the more concentrated the energy. By introducing a learnable positive real parameter, the Shannon entropy is mapped to the activation confidence of the scale, so that the scale with more concentrated energy obtains a higher confidence. Calculate the average activation confidence across all scales, and retain the complex time-frequency feature maps corresponding to the scales with activation confidence not lower than the average value, while discarding the feature maps of the remaining scales. The retained complex time-frequency feature maps form a sparse multi-scale feature set, which is used for subsequent cross-scale fusion processing.
[0033] It should be noted that calculating activation confidence based on energy distribution Shannon entropy and selecting high-information scales accordingly can remove redundant or noise-dominated scales while retaining key time-frequency structures, effectively reducing the complexity of subsequent attention calculations and the risk of overfitting, and improving model inference efficiency and generalization ability. This is especially suitable for deployment on resource-constrained terminal devices.
[0034] S5. Extract query vector, key vector and value vector from the complex time-frequency features of each scale in the sparse multi-scale feature set, and calculate cross-scale attention weight in the complex domain by the conjugate inner product of query vector and key vector to characterize the amplitude-phase coupling relationship between different scales.
[0035] Furthermore, the feature maps of each activation scale... Flattened into a vector ; Query vectors are generated using three independent complex-valued linear mappings. Key vector AND value vector ; For any two activation scales and The cross-scale attention weights are calculated using the following expression: ; in, To flatten the feature vectors, , , For complex-valued learnable mapping matrices, For query vector, For key vectors, This indicates the conjugate transpose. This indicates taking the real part of a complex number. For feature dimension, To sum dummy variables, This represents the attention weight.
[0036] It should be noted that calculating cross-scale attention weights by using the conjugate inner product of the query vector and the key vector in the complex domain can simultaneously measure the degree of alignment in amplitude and phase between different scales, accurately capture the amplitude-phase coupling relationship between multi-scale features, avoid the loss of phase information caused by real domain fusion, and thus generate a more physically consistent contextual representation.
[0037] S6. Use cross-scale attention weights to perform weighted summation on the value vectors in the sparse multi-scale feature set to obtain complex context features, and generate complex numerical channel state information prediction results for future time slots based on the complex context features.
[0038] Furthermore, for each query scale With its corresponding attention weight For all value vectors The expression is: ; in, For the first The context-enhanced complex numerical feature vector obtained after cross-scale attention fusion at each scale Represents a set All activation scale indexes Perform a summation operation. scale scale Cross-scale attention weights For scale The corresponding value vector; All Concatenate into a joint context vector ; Will Input complex-valued fully connected networks, mapped to the future Each time slot Complex-valued channel state information prediction results for each subcarrier .
[0039] It should be noted that by using cross-scale attention weights to weight and fuse value vectors, it is possible to dynamically emphasize other scale features that are most relevant to the semantics of the current query scale, thereby achieving intelligent aggregation of multi-resolution information. Finally, the future channel prediction results are generated through a complex-valued fully connected network, which fully preserves the amplitude and phase information of the complex domain and improves the prediction accuracy of channel evolution trends in high-frequency and high-speed scenarios.
[0040] S7. Construct training objectives based on the prediction results of complex-valued channel state information and the actual feedback channel state information, and jointly update the multi-scale complex time-frequency basis functions, activation confidence calculation logic, and complex mapping parameters.
[0041] Furthermore, real channel state information can be obtained through pilot signals. ; based on and The complex mean square error loss is calculated using the following expression: ; in, The loss function for complex-valued channel prediction. The number of future time slots to be predicted. The total number of subcarriers in the system. Indicates the range from 1 to 1 Summing is performed on each future time slot. Indicates the range from 1 to 1 Summing the subcarriers The model predicts the first The first future time slot, the first Complex-valued channel state information on each subcarrier This refers to the actual complex-valued channel state information at the corresponding location obtained through actual measurement of pilot signals.
[0042] It should be noted that by using the complex mean square error as the loss function, and jointly optimizing the learnable time-frequency basis parameters, confidence scaling factor, and complex-valued mapping matrix, an end-to-end closed-loop training mechanism is formed. This enables the entire system to continuously and adaptively adjust according to the actual channel feedback, which not only improves prediction accuracy but also enhances the model's robustness and transferability to different communication scenarios.
[0043] This embodiment also provides a computer device applicable to the channel prediction method based on multi-scale complex time-frequency feature fusion, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the channel prediction method based on multi-scale complex time-frequency feature fusion as proposed in the above embodiment.
[0044] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0045] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the channel prediction method based on multi-scale complex time-frequency feature fusion as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0046] In summary, this invention achieves adaptive alignment between the time-frequency analysis window and the channel's physical dynamic characteristics by constructing learnable multi-scale complex time-frequency basis functions based on channel coherence time and coherence bandwidth, and performing complex convolution transformation on historical complex channel state information sequences. This overcomes the mismatch problem of traditional fixed basis functions in high-speed or broadband scenarios, thereby obtaining high-fidelity, multi-resolution complex time-frequency feature maps. The energy entropy of each scale feature map is calculated and activation confidence is generated accordingly. Scales with high information concentration are dynamically selected to form a sparse multi-scale feature set. This not only reduces subsequent computational complexity and memory overhead but also enables the system to adaptively focus on key time-frequency structures according to the channel environment, improving robustness and energy efficiency.
[0047] Example 2, refer to Figure 2 — Figure 5 This is a second embodiment of the present invention, which provides a method for predicting array channel measurement and direction-finding correlation sequences based on learnable multi-scale time-frequency fusion completion, comprising: Step 1: Input sequence construction and complex value mapping: Align the input CSI according to its real / imaginary parts to form a complex tensor, and flatten the frequency domain dimension as needed for the task to obtain a sequence format suitable for the network input; for example, organize the input and target separately as follows: , The first 5 TTIs are used as input, and the last 15 TTIs are used as the prediction target.
[0048] To reduce computational cost and highlight relevant information, the real and imaginary parts can be linearly reduced and then fused. The dimensionality reduction can be achieved using the following two equations: ; ; in , It is a learnable projection matrix.
[0049] Step 2: Learnable multi-scale STFT time-frequency feature extraction and cross-scale fusion completion: To adapt to the different spectral characteristics of CSI at different time resolutions, multi-scale STFT is performed on the input complex value sequence, and adaptive fusion is performed after alignment in the frequency-time two-dimensional plane to obtain a unified time-frequency complex value representation.
[0050] Specifically, for the first STFT mapping is performed at each scale, and the expression is: ; in For window length, For FFT points, and Related to the position / step of the sliding window.
[0051] Time-frequency features generated at different scales Alignment and completion are performed on the dimension, and the data is fused using learnable normalized weights to obtain unified time-frequency complex-valued features, expressed as: ; ; ; in For element-wise multiplication, For the first Learnable fusion weight parameters at each scale This indicates that features at various scales are completed to a unified standard. .
[0052] Step 3, Complex Value Transformer Encoding: The fused time-frequency complex-valued features obtained in step two are input into the complex-valued Transformer encoder, and long-term dependencies and non-stationarity are modeled through complex-valued multi-head self-attention and complex-valued feedforward networks. Let the input sequence be Then the Query / Key / Value is obtained by linear mapping of three sets of complex values: ; in It is a learnable complex-valued matrix.
[0053] Will Divided into One point of attention: ; ; ; Each subspace dimension ,and .
[0054] In the Within each attention head, the correlation matrix is calculated using the complex-valued Hermitian inner product:
[0055] in This is the conjugate transpose.
[0056] To correlate attention allocation with the complex magnitude, a normalization method called "modulo softmax" is introduced: ; The output is thus obtained, expressed as: ; ; ; in This is for outputting the projection matrix.
[0057] To enhance stability and avoid gradient degradation, complex-valued residuals and complex-valued normalization are introduced: ; in This represents a complex-valued normalization operator based on amplitude and phase decoupling.
[0058] Step 4: Bidirectional Time-Frequency Cross-Domain Complex Value Attention (BCDA-CVS): To enhance the information flow and coupling modeling between the time domain and the frequency domain, a bidirectional cross-domain attention module is introduced, including two interaction paths: "Time Domain to Frequency Domain (T2F)" and "Frequency Domain to Time Domain (F2T)".
[0059] Its calculation can be performed in the following form: ; ; ; Subscript These represent the time-domain and frequency-domain branches, respectively. , Indicates feature splicing, This is a learnable mapping.
[0060] Step 5: Complex Value Prediction Projection and Output Organization: The encoded features are mapped to the prediction window dimension, and the frequency domain CSI prediction results for multiple future TTIs are generated through complex-valued linear projection. The real and imaginary parts are output respectively, thus preserving amplitude-phase consistency.
[0061] The above steps constitute a complete CSI sequence prediction method of "multi-scale time-frequency fusion completion + complex-valued attention + time-frequency cross-domain interaction", which can be directly used for channel prediction tasks in highly maneuverable scenarios and maintain explicit modeling of the physical structure of complex-valued signals.
[0062] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A channel prediction method based on multi-scale complex time-frequency feature fusion, characterized in that: include: Obtain the historical complex-valued channel state information sequence of the target communication link, estimate the channel coherence time and coherence bandwidth from the historical complex-valued channel state information sequence, and obtain the physical parameters; Based on physical parameters, a learnable multi-scale complex time-frequency basis function is constructed, and the historical complex value channel state information sequence is transformed to obtain a complex time-frequency feature map. The activation confidence of each scale is calculated based on the complex time-frequency feature maps of multiple scales, and the complex time-frequency feature maps of some scales are selected according to the activation confidence to form a sparse multi-scale feature set. For each scale of complex time-frequency features in the sparse multi-scale feature set, query vector, key vector and value vector are extracted respectively. Cross-scale attention weights are calculated in the complex domain by the conjugate inner product of query vector and key vector to characterize the amplitude-phase coupling relationship between different scales. By using cross-scale attention weights to perform weighted summation on the value vectors in the sparse multi-scale feature set, complex context features are obtained, and based on the complex context features, prediction results of complex numerical channel state information for future time slots are generated. The training objective is constructed based on the predicted results of the complex-valued channel state information and the actual feedback channel state information. The multi-scale complex time-frequency basis function, the activation confidence calculation logic, and the complex mapping parameters are jointly updated.
2. The channel prediction method based on multi-scale complex time-frequency feature fusion as described in claim 1, characterized in that: The specific steps for obtaining the historical complex-valued channel state information sequence of the target communication link are as follows: At the receiving end, using the pilot signals periodically transmitted by the transmitting end, the complex-valued channel response on each orthogonal frequency division multiplexing subcarrier within multiple consecutive historical time units is calculated by the least squares estimation method. The calculated complex-valued channel responses are arranged in chronological order and subcarrier index order to form a two-dimensional data structure; Each row of the two-dimensional data structure corresponds to a historical time unit, each column corresponds to a subcarrier, and each element consists of a real part and an imaginary part, reflecting the channel amplitude attenuation and phase rotation under the combination of time units and subcarriers, respectively.
3. The channel prediction method based on multi-scale complex time-frequency feature fusion as described in claim 1, characterized in that: The steps for estimating the channel coherence time and coherence bandwidth from the historical complex-valued channel state information sequence to obtain physical parameters are as follows: The normalized autocorrelation function is calculated by sliding along the time dimension for the historical complex-valued channel state information sequence, and the autocorrelation value is first decayed to 0. The corresponding time interval is determined as the channel coherence time. ; The frequency domain cross-correlation coefficient is calculated for the complex channel responses of adjacent subcarriers, and the reciprocal of the subcarrier spacing corresponding to the first drop of the cross-correlation coefficient to 0.5 is determined as the channel coherence bandwidth. ; Therefore, from and The physical parameters that constitute it.
4. The channel prediction method based on multi-scale complex time-frequency feature fusion as described in claim 2, characterized in that: The steps for constructing a learnable multi-scale complex time-frequency basis function based on physical parameters and transforming the historical complex numerical channel state information sequence to obtain a complex time-frequency feature map are as follows: Two learnable adjustment factors are introduced for each scale, one to control the widening of the scale time window and the other to control the modulation position of the scale in the frequency dimension. The degree of time window widening is inversely proportional to the channel coherence bandwidth, meaning that the more drastic the channel frequency changes, the narrower the corresponding time window. The frequency modulation position is inversely proportional to the channel coherence time, so the faster the channel time changes, the higher the corresponding modulation frequency. Based on the adjusted time-frequency parameters, a complex exponential basis function for Gaussian window modulation is constructed. The complex time-frequency feature map is generated by performing a complex convolution transformation along the time dimension on the historical complex numerical channel state information sequence using complex exponential basis functions. Repeat the process for all preset scales to obtain a set of complex time-frequency feature maps covering different time-frequency resolutions, which are then used as multi-scale representation inputs to subsequent processing stages.
5. The channel prediction method based on multi-scale complex time-frequency feature fusion as described in claim 3, characterized in that: The specific steps are as follows: Calculating the activation confidence of each scale based on complex time-frequency feature maps at multiple scales, and selecting complex time-frequency feature maps at a subset of scales based on the activation confidence to form a sparse multi-scale feature set. Calculate the energy distribution of the complex time-frequency feature map at all time-frequency locations for the k-th scale, and normalize the energy distribution to a probability mass function; The energy concentration scale is calculated based on the probability mass function, with Shannon entropy as the measure. The lower the entropy value, the more concentrated the energy. By introducing a learnable positive real parameter, Shannon entropy is mapped to the activation confidence of the scale, so that the scale with more concentrated energy obtains a higher confidence. Calculate the average activation confidence across all scales, and retain the complex time-frequency feature maps corresponding to the scales with activation confidence not lower than the average value, while discarding the feature maps of the remaining scales. The retained complex time-frequency feature maps form a sparse multi-scale feature set, which is used for subsequent cross-scale fusion processing.
6. The channel prediction method based on multi-scale complex time-frequency feature fusion as described in claim 4, characterized in that: The steps involve extracting query vectors, key vectors, and value vectors from the complex time-frequency features at each scale of the sparse multi-scale feature set, and calculating cross-scale attention weights in the complex domain using the conjugate inner product of the query vectors and key vectors to characterize the amplitude-phase coupling relationship between different scales. Feature maps of each activation scale Flattened into a vector ; Query vectors are generated using three independent complex-valued linear mappings. Key vector AND value vector ; For any two activation scales and The cross-scale attention weights are calculated using the following expression: ; in, To flatten the feature vectors, , , For complex-valued learnable mapping matrices, For query vector, For key vectors, This indicates the conjugate transpose. To take the real part of a complex number, For feature dimension, To sum dummy variables, For attention weights.
7. The channel prediction method based on multi-scale complex time-frequency feature fusion as described in claim 5, characterized in that: The method involves using cross-scale attention weights to perform a weighted summation of the value vectors in the sparse multi-scale feature set to obtain complex context features, and then generating complex-valued channel state information prediction results for future time slots based on these complex context features. The specific steps are as follows: For each query scale With its corresponding attention weight For all value vectors The expression is: ; in, For the first The context-enhanced complex numerical feature vector obtained after cross-scale attention fusion at each scale Represents a set All activation scale indexes Perform a summation operation. scale scale Cross-scale attention weights For scale The corresponding value vector; All Concatenate into a joint context vector ; Will Input complex-valued fully connected networks, mapped to the future Each time slot Complex-valued channel state information prediction results for each subcarrier .
8. The channel prediction method based on multi-scale complex time-frequency feature fusion as described in claim 6, characterized in that: The specific steps for constructing a training objective based on the predicted results of the complex-valued channel state information and the actual feedback channel state information, and jointly updating the multi-scale complex time-frequency basis function, activation confidence calculation logic, and complex mapping parameters are as follows: Obtaining real channel state information through pilot signals ; based on and The complex mean square error loss is calculated using the following expression: ; in, The loss function for complex-valued channel prediction. The number of future time slots to be predicted. The total number of subcarriers in the system. Indicates the range from 1 to 1 Summing is performed on each future time slot. Indicates the range from 1 to 1 Summing the subcarriers The model predicts the first The first future time slot, the first Complex-valued channel state information on each subcarrier This refers to the actual complex-valued channel state information at the corresponding location obtained through actual measurement of pilot signals.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the channel prediction method based on multi-scale complex time-frequency feature fusion as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the channel prediction method based on multi-scale complex time-frequency feature fusion as described in any one of claims 1 to 8.