TCN-BigBird-based channel equalization method, apparatus and device, and medium
By using the TCN-BigBird cascade structure and combining the TCN and BigBird models, the problem of degraded channel equalization performance in complex multipath environments is solved, achieving efficient channel equalization and improving channel equalization performance and robustness.
Patent Information
- Application Number
- CN202511004354.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-31
AI Technical Summary
In complex multipath environments and high-speed scenarios, existing channel equalization methods struggle to effectively handle multipath interference, especially short-range and long-range inter-symbol interference, leading to a decline in channel equalization performance.
A TCN-BigBird cascaded structure is adopted. The TCN module extracts multi-scale short-range interference features and combines them with the BigBird model to capture long-range interference features. The feature fusion is performed using gated residual connections to construct a composite channel equalization network, which performs signal preprocessing and nonlinear transformation, and finally performs channel equalization.
It significantly improves channel equalization performance, increasing the gain by about 6dB at the same bit error rate, which is 2-4dB higher than traditional methods. It has stronger robustness and computational efficiency, and is suitable for large-scale channel modeling tasks.
Smart Images

Figure CN120880846A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication, and specifically relates to a channel equalization method, apparatus, device and medium based on TCN-BigBird. Background Technology
[0002] Multipath propagation is a common phenomenon in wireless channel environments. When a signal encounters an obstacle during propagation, multiple paths are created. Each path may experience phase shifts, attenuation, and delays, leading to inter-symbol interference (ISI). Especially in high-speed mobile scenarios, the time-varying characteristics of the signal become more pronounced due to the Doppler effect, making signal equalization at the receiver more difficult. Therefore, channel equalization is a key technology for improving signal quality and has significant research value.
[0003] The main purpose of channel equalization is to compensate for channel distortion, thereby reducing the impact of inter-symbol interference (ISI) on the received signal. One method utilizes zero-forcing (ZF) equalization, which completely eliminates ISI by reversing the channel frequency response. Its main advantage is the complete elimination of ISI, theoretically enabling distortion-free transmission. However, zero-forcing equalization is highly sensitive to noise, especially when there are zeros in the channel frequency response, leading to noise amplification and degraded system performance. Another method utilizes minimum mean square error (MMSE) equalization, which considers both noise and signal distortion. Its advantage is achieving a balance between noise and distortion, improving the system's signal-to-noise ratio (SNR). However, MMSE equalization requires precise noise statistics, resulting in high computational complexity. Decision feedback equalization (DFE) is a nonlinear equalization method that uses feedback from decided symbols to eliminate ISI caused by previous symbols. Its main advantage is the effective elimination of interference caused by previous symbols, outperforming linear equalizers. However, DFE is highly sensitive to decision errors, and error propagation can severely degrade performance. Adaptive equalizers can automatically adjust their parameters based on the characteristics of the received signal to adapt to channel variations. Common adaptive algorithms include the Least Mean Square (LMS) algorithm and the Recursive Least Squares (RLS) algorithm. Their advantage is that they can track channel changes in real time and are suitable for time-varying channels. However, the convergence speed and steady-state error of adaptive equalizers are key design issues. Blind equalization does not require training sequences; it directly extracts statistical characteristics from the received signal for equalization. One approach is the Modified Constant Modulus (MCMA) algorithm based on a fractional-interval decision feedback structure. This method uses MCMA to correct carrier frequency and phase offsets while recovering the constellation diagram, achieving fast convergence and a small steady-state error. Its advantage is that it does not require additional training sequences, saving bandwidth resources. However, blind equalization algorithms typically have slow convergence speeds, and initial conditions have a significant impact on performance.
[0004] With the development of artificial intelligence, deep learning has shown great potential in the field of channel equalization. This paper presents a channel equalization algorithm based on Convolutional Neural Networks (CNNs). This method automatically extracts local features in the channel using CNNs and uses a multi-layer convolutional structure for signal recovery. Experiments show that CNNs can effectively eliminate multipath interference in the frequency domain, significantly improving equalization accuracy. Compared with traditional methods, this algorithm has a faster convergence speed and maintains high robustness in various channel environments. However, while CNNs have strong local perception capabilities, they may perform poorly when processing data with global dependencies. Another paper presents an equalizer framework based on Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This method uses CNNs to learn channel characteristics, then inputs them into an RNN for time-domain modeling, and finally classifies the signal. Experimental results show that this method improves performance by 2–4 dB compared to the traditional recursive least squares method while achieving the same bit error rate.
[0005] Channel equalization is achieved through deep learning, with CNN, RNN, and LSTM being commonly used models. CNN can effectively extract local features of time-frequency signals and is suitable for handling spatially correlated channel characteristics. RNN is suitable for handling short-term dependent channel state information and can be used to predict channel states that change over time. However, CNN mainly processes local features and may struggle to capture long-range correlated channel changes. RNN suffers from the vanishing gradient problem and is poorly adapted to channel state changes over long time spans. Although LSTM overcomes the vanishing gradient problem of traditional RNN to some extent, it may still face information loss when dealing with very long temporal dependencies. Furthermore, RNN and LSTM require sequential processing of sequences, making parallel computation difficult and resulting in slow inference speeds. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a channel equalization method, apparatus, device, and medium based on TCN-BigBird, which significantly improves channel equalization performance in complex multipath environments and high-speed scenarios.
[0007] In a first aspect, embodiments of this application provide a channel equalization method based on TCN-BigBird, comprising the following steps:
[0008] Step 1: Obtain signal sample data in the simulated multipath channel and construct a dataset.
[0009] Step 2: Perform signal preprocessing on the input signal sample data to construct a two-dimensional real number input tensor.
[0010] Step 3: Input the constructed two-dimensional real tensor into the TCN module of the composite channel equalization network to extract multi-scale short-range interference features.
[0011] Step 4: Input the extracted multi-scale short-range interference features into the BigBird model to extract long-range interference features.
[0012] Step 5: The acquired multi-scale short-range and long-range interference features are fused using gated residual connections to obtain fused features.
[0013] Step 6: Perform a nonlinear transformation on the obtained fused features to obtain the fused multi-scale temporal features.
[0014] Step 7: Map the multi-scale temporal features to the target output space through a fully connected layer to obtain the channel-equalized signal output by the composite channel equalization network.
[0015] Step 8: Optimize and train the composite channel equalization network based on the constructed dataset.
[0016] In one possible implementation, the signal sample data is signal sample data at different signal-to-noise ratios obtained through simulation; the signal sample data includes baseband signals, short-range inter-symbol interference signals in multipath channels, long-range inter-symbol interference signals, and time delay and frequency fading signals.
[0017] In one possible implementation, the signal preprocessing method specifically includes the following steps:
[0018] First, the received signal sample data is decomposed into real part (I) and imaginary part (Q); second, the real part (I) data is used as one dimension and the imaginary part (Q) data is used as another dimension, together forming a two-dimensional real input tensor, which is used as the input data for subsequent processing.
[0019] In one possible implementation, the TCN module is formed by several TCN networks, each TCN network being formed by stacking several residual blocks, each residual block containing dilated causal convolution, normalization, activation function, and regularization. A two-dimensional real-number input tensor is input in parallel to several TCN networks, and then the outputs of the several TCN networks are concatenated dimensionally to extract multi-scale short-range interference features.
[0020] In one possible implementation, the BigBird model operates as follows:
[0021] First, the input multi-scale short-range interference features are projected and positionally encoded using an encoding layer. Then, sparse attention interactions are performed to obtain long-range interference features. Local attention focuses on the interference patterns of adjacent symbols, global attention preserves the global connections at key positions (the beginning and end of the signal sequence), and random attention randomly selects 10% of the symbol positions to establish global associations, thus modeling long-range temporal dependencies.
[0022] In one possible implementation, the gated residual connection is specifically implemented as follows:
[0023] First, feature dimension alignment is performed to ensure that the output of the TCN module is consistent with the output dimension of the BigBird model. Then, gate weight calculation is introduced to dynamically adjust the contribution ratio of multi-scale short-range and long-range interference features.
[0024] Secondly, after adaptively fusing multi-scale short-range and long-range interference features based on the adjusted contribution ratio through gated weighting, residual connections are made with the original TCN features, and layer normalization is applied to achieve multi-source feature fusion and obtain the final fused features.
[0025] In one possible implementation, the nonlinear transformation process is specifically implemented as follows:
[0026] First, the FFN module, consisting of two linear transformation layers and a nonlinear activation function, is used to perform nonlinear mapping on the fused features to capture higher-order interference patterns in the signal sequence.
[0027] Secondly, the output of the FFN module is filtered using a gating mechanism to obtain the fused multi-scale temporal features.
[0028] In one possible implementation, the optimization training uses a signal without channel interference as the supervision target, employs the mean squared error loss function, and improves training stability and convergence through the AdamW optimizer, learning rate scheduling, and gradient pruning.
[0029] Secondly, embodiments of this application provide a channel equalization device based on TCN-BigBird, comprising the following modules:
[0030] Dataset building module: Used to acquire signal sample data in simulated multipath channels and build datasets.
[0031] Preprocessing module: Used to preprocess the input signal sample data and construct a two-dimensional real number input tensor.
[0032] Short-range interference feature extraction module: The TCN module is used to extract multi-scale short-range interference features based on the constructed two-dimensional real tensor.
[0033] Long-range interference feature extraction module: The BigBird model is used to extract long-range interference features based on the extracted multi-scale short-range interference features.
[0034] Feature fusion module: The acquired multi-scale short-range and long-range interference features are fused through gated residual connections to obtain fused features.
[0035] Nonlinear transformation module: Performs nonlinear transformation on the obtained fused features to obtain the fused multi-scale temporal features.
[0036] Output module: The multi-scale temporal features are mapped to the target output space through a fully connected layer to obtain the channel-equalized signal output by the composite channel equalization network.
[0037] Model optimization module: Based on the dataset constructed by the dataset construction module, optimize and train a composite channel equalization network including the TCN module, BigBird model, gated residual connection, nonlinear transformation and fully connected mapping.
[0038] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory;
[0039] The memory is used to store computer programs;
[0040] When the processor executes the program stored in the memory, it implements any of the channel equalization methods described in this application.
[0041] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the channel equalization methods described in this application.
[0042] Fifthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the channel equalization methods described in this application.
[0043] The beneficial effects of this invention are as follows:
[0044] 1. It solves the problem of significant degradation in time-varying channel tracking capability under complex multipath environments.
[0045] 2. A composite channel equalization network with local sensing and global modeling capabilities was constructed. Through the designed TCN-BigBird cascade structure, the local features of short-range inter-symbol interference were efficiently extracted, the rapid changes in short-time channel dynamics were captured, and the information loss problem caused by long-range inter-symbol interference was effectively addressed.
[0046] 3. Under the same bit error rate, it improves the gain by about 6dB compared with the traditional equalization method and by about 2~4dB compared with the single deep learning model equalization method, which verifies its robustness and effectiveness in complex interference environment.
[0047] In summary, this invention is characterized by high effectiveness and robustness, and it also maintains high channel equalization performance even at low signal-to-noise ratios.
[0048] Compared to existing deep learning models, the TCN-BigBird cascaded structure proposed in this invention has stronger long-term dependency modeling capabilities. This is because TCN employs extended causal convolution, enabling simultaneous modeling of both short-term and long-term dependencies, making it more stable than RNN structures. BigBird utilizes a sparse attention mechanism, effectively capturing global channel changes over long distances, making it suitable for time-varying channels such as frequency-hopping and fast-fading channels. Furthermore, it has lower computational complexity and supports parallel computing. In addition, the proposed model is more suitable for large-scale channel modeling tasks because BigBird can handle large-scale time-frequency data, making it suitable for high-bandwidth, multi-antenna channel equalization, which CNN, RNN, and LSTM struggle to scale to. It not only accurately solves short-range interference problems but also further optimizes channel equalization performance through long-range dependency modeling. Attached Figure Description
[0049] Figure 1 This is a schematic diagram of a composite channel equalization network structure according to an embodiment of the present invention.
[0050] Figure 2 This is a schematic diagram of the TCN network structure according to an embodiment of the present invention.
[0051] Figures 3(a), 3(b), and 3(c) are schematic diagrams of the analog signal generation process in an embodiment of the present invention.
[0052] Figure 4 This is a schematic diagram showing the time-domain waveform comparison and multipath effect of an embodiment of the present invention.
[0053] Figures 5(a), 5(b), 5(c), 5(d), 5(e), 5(f), 5(g), and 5(h) are constellation diagrams of various channel equalization methods in embodiments of the present invention (SNR=12dB).
[0054] Figures 6(a) and 6(b) show the experimental results of channel equalization in an embodiment of the present invention. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0056] This application proposes a TCN-BigBird-based channel equalization method to improve channel equalization performance. The TCN-BigBird model fully utilizes the structural advantages of causal dilated convolution in TCN to efficiently extract local features of short-range inter-symbol interference and capture rapidly changing short-term channel dynamics. Simultaneously, the introduction of the BigBird model can handle long-term temporal dependencies and reduce computational complexity, making it suitable for channel modeling of long-time-series data and effectively addressing the information loss problem caused by long-range inter-symbol interference. The two are cascaded and fused to construct a composite channel equalization network with both local sensing and global modeling capabilities. Under the same bit error rate, the proposed TCN-BigBird-based equalization method achieves a gain of approximately 6 dB compared to traditional equalization methods and an improvement of approximately 2–4 dB compared to single deep learning model equalization methods, verifying its robustness and effectiveness in complex interference environments.
[0057] Step 1: Obtain signal sample data in the simulated multipath channel and construct a dataset.
[0058] Taking QPSK signals as an example, the baseband signal of the QPSK signal is first modulated and generated in MATLAB. Then, the baseband signal is passed through a Rayleigh channel and Ricean fly-through noise and Gaussian white noise are introduced to obtain short-range inter-symbol interference (ISI) signals and long-range ISI signals. Finally, time delays from the multipath channel are added, with delays of 0 microseconds, 1 microsecond, 2.3 microseconds, and 12 microseconds, which can be regarded as reflection or scattering paths; the path with a delay of 0 is the direct path, which can represent the main path in line-of-sight propagation.
[0059] The signal sample data contains six signal-to-noise ratios (SNRs): 0dB, 4dB, 8dB, 12dB, 16dB, and 20dB. Each SNR has 100 signal groups, each with a length of 16384 kHz. The sampling rate is 1MHz, the code rate is 250kHz, the roll-off factor is 0.35, the Rice K-factor is 4, the Doppler shift is 50Hz, and the multipath delays are 0µs, 1µs, 2.3µs, and 12µs, respectively. The multipath gains are 0dB, -3dB, -6dB, and -9dB, respectively.
[0060] Step 2: Perform signal preprocessing on the input signal sample data to construct a two-dimensional real number input tensor.
[0061] The input data is signal sample data. Signal sequence length The input tensor is decomposed into a real part (I) and an imaginary part (Q). Then, the real part (I) is used as one dimension and the imaginary part (Q) is used as another dimension, together forming a two-dimensional real input tensor. , as input data for subsequent processing;
[0062] (1)
[0063] In the formula, B is the batch size, which is 256 in this paper, and 2 represents the I and Q dual channels.
[0064] Step 3: Input the constructed two-dimensional real tensor into the TCN module of the composite channel equalization network to extract multi-scale short-range interference features.
[0065] In this embodiment, the TCN module is composed of four identical TCN networks in parallel, and each TCN network is composed of four layers of residual blocks stacked together.
[0066] In a TCN network, the input sequence is represented by the mathematical formula for residual blocks as follows:
[0067] (2)
[0068] In the formula, For convolution kernel weights, Given the input sequence, Ensure temporal causality. This is the output of the residual block.
[0069] By stacking residual blocks, each layer expands the temporal receptive field with an exponentially increasing inflation rate. The expansion rate of the layer is It captures multi-scale short-range interference features in the input sequence. For the first TCN network output
[0070] Finally, the outputs of the four TCN networks are spliced together to cover interference modes from the microscopic symbol level to the medium time scale, thus obtaining the output of the TCN module.
[0071] (3)
[0072] In the formula Output for the TCN module.
[0073] Step 4: Input the extracted multi-scale short-range interference features into the BigBird model to extract long-range interference features.
[0074] First, feature projection and position encoding are performed at the encoding layer. Linear projection can be represented as:
[0075] (4)
[0076] In the formula, the parameters For feature projection parameters, for peacekeeping Column vectors, For projection output.
[0077] Positional encoding is embedded in the sine wave position. The output is:
[0078] (5)
[0079] In the formula, To learn the position vector, For L-dimensional peace A vector of columns.
[0080] Then, sparse attention interaction is performed to obtain long-range interference features. Local attention focuses on the interference patterns of adjacent symbols, global attention preserves the global connections of key positions (the beginning and end of the signal sequence), and random attention randomly selects 10% of the symbol positions to establish global associations, thus modeling long-range temporal dependencies.
[0081] Step 5: The acquired multi-scale short-range and long-range interference features are fused using gated residual connections to obtain fused features.
[0082] Introducing gated residual connections into the cascaded structure of TCN and BigBird aims to address the gradient vanishing problem caused by the difference in feature distribution between the two modules, while also enhancing the complementarity of cross-scale features.
[0083] First, feature dimension alignment is performed to ensure the output of the TCN module. Output of the BigBird model Dimensions must be consistent. If dimensions do not match (e.g....), Projection is performed using a 1×1 convolution:
[0084] (6)
[0085] In the formula, for Weihe A vector of columns.
[0086] Then, gating weight calculation is introduced, using learnable gating vectors. Dynamically adjust the contribution ratio of multi-scale short-range and long-range disturbance features:
[0087] (7)
[0088] In the formula, For the Sigmod function, For parameter matrices, This represents a learnable bias vector.
[0089] Then, the two features are adaptively fused using gated weighting to obtain preliminary fused features:
[0090] (8)
[0091] In the formula This indicates element-wise multiplication, and the gating mechanism allows the model to autonomously choose to depend on local details or the global context.
[0092] Finally, the preliminary fused features are residually connected to the original TCN features, and layer normalization is applied to obtain the final fused features:
[0093] (9)
[0094] The complete gated residual connection formula can be expressed as:
[0095] (10)
[0096] TCN excels at capturing local temporal patterns, such as short-term dependencies and causality, while BigBird extracts long-term global dependencies through a sparse attention mechanism. This gating module dynamically adjusts the contribution ratio of the two types of features using learnable weights, achieving complementarity between local and global information.
[0097] Step 6: Perform a nonlinear transformation on the obtained fused features to obtain the fused multi-scale temporal features.
[0098] The FFN module consists of two linear transform layers and a nonlinear activation function. It performs nonlinear mapping on the gated fused features, capturing higher-order interference patterns in the signal sequence, such as nonlinear distortion caused by multipath propagation. It also suppresses negative noise and preserves positive key features through the ReLU activation function. The mathematical formula can be expressed as:
[0099] (11)
[0100] in, , , Usually The FFN module will take the input feature dimension as an example. Expand to This increases model capacity to encode complex channel responses. A second linear layer is then used to reduce the dimensionality to [the required level]. Redundant information is removed, and features that are strongly correlated with signal recovery are retained.
[0101] By using a gating mechanism to filter the output of the FFN module, important features are retained and irrelevant signals are suppressed, resulting in fused multi-scale temporal features.
[0102] Step 7: Map the multi-scale temporal features to the target output space through a fully connected layer to obtain the channel-equalized signal output by the composite channel equalization network.
[0103] TCN captures local temporal patterns through dilated convolutions, while BigBird captures global dependencies through sparse attention; however, the output features of both may still be at a low to mid-order representation. Directly using these features for classification or regression may not fully uncover complex temporal relationships. By employing a multi-layer fully connected network to perform high-order nonlinear transformations on the fused features, the model's ability to fit complex patterns is improved, and high-dimensional features are compressed or expanded to the target dimension through weight matrices.
[0104] In this embodiment, the selected signal is QPSK, with a signal range of [-1, 1]. The Tanh activation function is used to constrain the model output to this range.
[0105] Step 8: Optimize and train the composite channel equalization network based on the constructed dataset.
[0106] Signals that have not undergone multiple path channels are used as the supervision sequence. Mean squared error (MSE) is used as the loss function to calculate the difference between the model output and the supervision sequence. AdamW is selected as the optimizer, and combined with learning rate scheduling and gradient pruning, the model is optimized to minimize the MSE. AdamW can adjust the parameter update magnitude through adaptive learning rate, the gradient of MSE tends to flatten as it approaches the optimal solution, learning rate scheduling can prevent skipping the optimal solution due to an excessively large learning rate in the later stages, and gradient pruning can ensure the stability of long sequence training and prevent gradient explosion.
[0107] TCN-BigBird model, for example Figure 1 As shown, the TCN model excels at capturing local temporal dependencies through its convolutional structure, making it particularly suitable for handling short-range interference. The BigBird model, on the other hand, effectively models long-range dependencies through its sparse self-attention mechanism, suppressing long-range interference caused by long delays and multipath propagation. Therefore, by introducing a gated residual structure and dynamically adjusting the contribution ratio of the two modules' features using learnable sigmoid gate weights: enhancing the local feature weights of the TCN in regions with rapid interference transients (such as burst noise), and relying on the global context information of the BigBird model in regions of stable fading, these two models are combined to form a TCN-BigBird cascaded structure. This not only accurately solves the short-range interference problem but also further optimizes the channel equalization effect through long-range dependency modeling.
[0108] Temporal Convolutional Networks (TCNs) are deep learning models used for sequence modeling. This invention employs a TCN module formed by multiple TCN networks. The specific structure of the TCN network is as follows: Figure 2 As shown, in channel equalization, the TCN network ensures that only current and historical information is used through causal convolution, avoiding future information leakage and improving real-time performance. Simultaneously, it expands the receptive field using dilated convolution, reducing network depth while capturing short-term dependencies, thus improving computational efficiency and generalization ability, making it suitable for handling short-range inter-symbol interference.
[0109] Example:
[0110] 1. Dataset
[0111] The example uses a QPSK signal. First, the QPSK signal is modulated and generated in MATLAB. To avoid excessively large signal bandwidth, pulse shaping is performed on the QPSK symbols. A 4x raised cosine filter is used to interpolate and filter the symbols, generating a smooth baseband signal with a length of 16384. Then, the signal is passed through a Rayleigh channel, and Ricean fly-through noise and Gaussian white noise are introduced to simulate short-range inter-symbol interference, long-range inter-symbol interference, time delay, and frequency fading in multipath channels. The signal generation process is shown in Figures 3(a), 3(b), and 3(c). Figure 3(a) shows the original QPSK signal constellation, Figure 3(b) shows the 4x sampled QPSK constellation, and Figure 3(c) shows the constellation after multipath processing and noise addition.
[0112] Time-domain waveform comparison and multipath effect diagram as shown in the figure Figure 4 As shown in the figure, from top to bottom, the first two figures visually illustrate the interference of multipath on the signal. The signal after Gaussian white noise is superimposed on the signal exhibits random fluctuations in amplitude, which can simulate short-range inter-symbol interference. The signal after passing through the Rice-Rayleigh channel is significantly different from the time-domain waveform of the original signal, reflecting that the signal will undergo significant distortion after passing through the multipath channel, which can simulate long-range interference. Figure 4 The third figure shows the time delay in the multipath channel, with delays of 1 microsecond, 2.3 microseconds, and 12 microseconds, which can be regarded as reflection or scattering paths; the path with a delay of 0 is the direct path, which can represent the main path in line-of-sight propagation. Figure 4 The fourth figure illustrates the time-varying nature of the channel. The gain of the direct path changes dynamically over time, and the Doppler effect causes the frequency of the received signal to shift (which is shown as periodic fluctuations in gain in the figure).
[0113] The dataset in this embodiment contains six signal-to-noise ratios (SNRs): 0dB, 4dB, 8dB, 12dB, 16dB, and 20dB. Each SNR has 100 signal groups, each with a length of 16384, a sampling rate of 1MHz, a code rate of 250kHz, a roll-off factor of 0.35, a Rice K-factor of 4, a Doppler shift of 50Hz, multipath delays of 0µs, 1µs, 2.3µs, and 12µs, and multipath gains of 0dB, -3dB, -6dB, and -9dB, respectively.
[0114] 2. Experimental Environment and Network Configuration
[0115] In this example, the server GPU is an NVIDIA TITAN RTX 3090, the deep learning framework used is PyTorch, and the model is implemented in Python using Sublime Text 3. Since CPU-based model training is relatively slow, we trained the network framework using a GPU. The GPU used is an NVIDIA TITAN RTX 3090, the model training epochs are 100, and the loss function is the cross-entropy loss function.
[0116] 3. Performance Simulation
[0117] Mean Squared Error (MSE) and Bit Error Rate (BER) are selected as evaluation metrics. The formula for calculating MSE can be expressed as:
[0118] (12)
[0119] In the formula, The length of the signal sequence. This represents the true value of the signal after it has not passed through a multipath channel. This is the output value of the model.
[0120] The equalized sequence is then demodulated. The demodulation steps mainly include symbol synchronization, carrier synchronization, and symbol decision. Finally, the bit error rate is calculated.
[0121] (13)
[0122] In the formula, It is the number of error symbols. It represents the total number of code elements.
[0123] This embodiment compares the proposed equalization method based on TCN-BigBird with traditional equalization methods: Zero Forcing Equalization (ZF), Minimum Mean Square Error Equalization (MMSE), Decision Feedback Equalization (DFE), as well as equalization methods based on RNN, TCN, BigBird, and Transformer. The constellation diagrams and experimental results after signal equalization are shown in Figures 5(a), 5(b), 5(c), 5(d), 5(e), 5(f), 5(g), and 5(h), and Figures 6(a) and 6(b).
[0124] Figure 5(a) shows ZF channel equalization, Figure 5(b) shows MMSE channel equalization, Figure 5(c) shows DFE channel equalization, Figure 5(d) shows RNN channel equalization, Figure 5(e) shows TCN channel equalization, Figure 5(f) shows Transformer channel equalization, Figure 5(g) shows BigBird channel equalization, and Figure 5(h) shows TCN-BigBird channel equalization. Figure 6(a) shows the mean square error, and Figure 6(b) shows the bit error rate. It can be seen that the TCN-BigBird equalization method exhibits better constellation convergence compared to other methods and maintains superior performance under different signal-to-noise ratio conditions. Regarding the mean square error, at 4dB, the TCN-BigBird mean square error is approximately an order of magnitude lower than that of traditional equalization methods; at 12dB, the TCN-BigBird mean square error also shows a significant reduction compared to Transformer and BigBird, indicating that the TCN structure effectively complements BigBird, resulting in a substantial improvement in equalization accuracy. In terms of bit error rate, under the same bit error rate, the equalization method based on TCN-BigBird improves the gain by about 6dB compared with the traditional method, improves the gain by about 4dB compared with the equalization method based on RNN, and improves the gain by about 2dB compared with the single TCN, Transformer, BigBird model method.
[0125] This application also provides a channel equalization device based on TCN-BigBird, including the following modules:
[0126] Dataset building module: Used to acquire signal sample data in simulated multipath channels and build datasets.
[0127] Preprocessing module: Used to preprocess the input signal sample data and construct a two-dimensional real number input tensor.
[0128] Short-range interference feature extraction module: The TCN module is used to extract multi-scale short-range interference features based on the constructed two-dimensional real tensor.
[0129] Long-range interference feature extraction module: The BigBird model is used to extract long-range interference features based on the extracted multi-scale short-range interference features.
[0130] Feature fusion module: The acquired multi-scale short-range and long-range interference features are fused through gated residual connections to obtain fused features.
[0131] Nonlinear transformation module: Performs nonlinear transformation on the obtained fused features to obtain the fused multi-scale temporal features.
[0132] Output module: The multi-scale temporal features are mapped to the target output space through a fully connected layer to obtain the channel-equalized signal output by the composite channel equalization network.
[0133] Model optimization module: Based on the dataset constructed by the dataset construction module, optimize and train a composite channel equalization network including the TCN module, BigBird model, gated residual connection, nonlinear transformation and fully connected mapping.
[0134] In one possible implementation, within the dataset building module:
[0135] The signal sample data is obtained through simulation under different signal-to-noise ratios; the signal sample data includes baseband signals, short-range inter-symbol interference signals in multipath channels, long-range inter-symbol interference signals, and time delay and frequency fading signals.
[0136] In one possible implementation, the signal preprocessing method of the preprocessing module specifically includes the following steps:
[0137] First, the received signal sample data is decomposed into real part (I) and imaginary part (Q); second, the real part (I) data is used as one dimension and the imaginary part (Q) data is used as another dimension, together forming a two-dimensional real input tensor, which is used as the input data for subsequent processing.
[0138] In one possible implementation, the TCN module of the short-range interference feature extraction module is formed by several TCN networks, each of which is formed by stacking several residual blocks. Each residual block contains dilated causal convolution, normalization, activation function, and regularization. A two-dimensional real-number input tensor is input in parallel to several TCN networks, and then the outputs of several TCN networks are concatenated dimensionally to extract multi-scale short-range interference features.
[0139] In one possible implementation, the BigBird model of the long-range interference feature extraction module is as follows:
[0140] First, the input multi-scale short-range interference features are projected and positionally encoded using an encoding layer. Then, sparse attention interactions are performed to obtain long-range interference features. Local attention focuses on the interference patterns of adjacent symbols, global attention preserves the global connections at key positions (the beginning and end of the signal sequence), and random attention randomly selects 10% of the symbol positions to establish global associations, thus modeling long-range temporal dependencies.
[0141] In one possible implementation, the gated residual connection of the feature fusion module is specifically implemented as follows:
[0142] First, feature dimension alignment is performed to ensure that the output of the TCN module is consistent with the output dimension of the BigBird model. Then, gate weight calculation is introduced to dynamically adjust the contribution ratio of multi-scale short-range and long-range interference features.
[0143] Secondly, after adaptively fusing multi-scale short-range and long-range interference features based on the adjusted contribution ratio through gated weighting, residual connections are made with the original TCN features, and layer normalization is applied to achieve multi-source feature fusion and obtain the final fused features.
[0144] In one possible implementation, the nonlinear transformation processing of the nonlinear transformation module is specifically implemented as follows:
[0145] First, the FFN module, consisting of two linear transformation layers and a nonlinear activation function, is used to perform nonlinear mapping on the fused features to capture higher-order interference patterns in the signal sequence.
[0146] Secondly, the output of the FFN module is filtered using a gating mechanism to obtain the fused multi-scale temporal features.
[0147] In one possible implementation, the model optimization module uses a signal without channel interference as the supervision target, employs the mean squared error loss function, and improves training stability and convergence through the AdamW optimizer, learning rate scheduling, and gradient pruning.
[0148] This application also provides an electronic device, which includes a processor and a memory.
[0149] The memory is used to store computer programs.
[0150] When the processor executes a program stored in the memory, it implements any of the methods described in this application.
[0151] In one possible implementation, the electronic device of this application embodiment further includes a communication interface and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.
[0152] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.
[0153] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0154] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0155] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0156] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements any of the methods described in this application.
[0157] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the methods described in this application.
[0158] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0159] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0160] The various embodiments in this specification are described in a related manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other.
[0161] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. A channel equalization method based on TCN-BigBird, characterized in that, Includes the following steps: Step 1: Obtain signal sample data from the simulated multipath channel and construct a dataset; Step 2: Perform signal preprocessing on the input signal sample data to construct a two-dimensional real number input tensor; Step 3: Input the constructed two-dimensional real tensor into the TCN module of the composite channel equalization network to extract multi-scale short-range interference features; Step 4: Input the extracted multi-scale short-range interference features into the BigBird model to extract long-range interference features; Step 5: The acquired multi-scale short-range and long-range interference features are fused using gated residual connections to obtain fused features; Step 6: Perform a nonlinear transformation on the obtained fused features to obtain the fused multi-scale temporal features; Step 7: Map the multi-scale temporal features to the target output space through a fully connected layer to obtain the channel-equalized signal output by the composite channel equalization network. Step 8: Optimize and train the composite channel equalization network based on the constructed dataset.
2. The channel equalization method based on TCN-BigBird according to claim 1, characterized in that, The signal sample data is obtained through simulation under different signal-to-noise ratios; the signal sample data includes baseband signals, short-range inter-symbol interference signals in multipath channels, long-range inter-symbol interference signals, and time delay and frequency fading signals.
3. The channel equalization method based on TCN-BigBird according to claim 1, characterized in that, The signal preprocessing method specifically includes the following steps: First, the received signal sample data is decomposed into a real part I and an imaginary part Q. Second, the real part I data is used as one dimension and the imaginary part Q data is used as another dimension to form a two-dimensional real input tensor, which is used as the input data for subsequent processing.
4. The channel equalization method based on TCN-BigBird according to claim 1, characterized in that, The TCN module consists of several TCN networks, each of which is formed by stacking several residual blocks. Each residual block includes dilated causal convolution, normalization, activation function, and regularization. A two-dimensional real input tensor is input to several TCN networks in parallel, and then the outputs of several TCN networks are concatenated to extract multi-scale short-range interference features.
5. The channel equalization method based on TCN-BigBird according to claim 1, characterized in that, The BigBird model is operated as follows: First, the input multi-scale short-range interference features are projected and positionally encoded through the Encoding layer. Then, sparse attention interaction is performed to obtain long-range interference features. Local attention focuses on the interference patterns of adjacent symbols, global attention preserves the global connections of key positions, and random attention randomly selects 10% of the symbol positions to establish global associations, thus modeling long-range temporal dependencies.
6. The channel equalization method based on TCN-BigBird according to claim 1, characterized in that, The gated residual connection is specifically implemented as follows: First, feature dimension alignment is performed to ensure that the output dimension of the TCN module is consistent with that of the BigBird model; then, gate weight calculation is introduced to dynamically adjust the contribution ratio of multi-scale short-range interference features and long-range interference features. Secondly, after adaptively fusing multi-scale short-range and long-range interference features based on the adjusted contribution ratio through gated weighting, residual connections are made with the original TCN features, and layer normalization is applied to achieve multi-source feature fusion and obtain the final fused features.
7. The channel equalization method based on TCN-BigBird according to claim 1, characterized in that, The aforementioned nonlinear transformation processing is specifically implemented as follows: First, the FFN module, consisting of two linear transformation layers and a nonlinear activation function, is used to perform nonlinear mapping on the fused features to capture higher-order interference patterns in the signal sequence. Secondly, the output of the FFN module is filtered using a gating mechanism to obtain the fused multi-scale temporal features.
8. A channel equalization device based on TCN-BigBird, characterized in that, Includes the following modules: Dataset building module: Used to acquire signal sample data in simulated multipath channels and build datasets; Preprocessing module: Used to preprocess the input signal sample data and construct a two-dimensional real number input tensor; Short-range interference feature extraction module: The TCN module is used to extract multi-scale short-range interference features based on the constructed two-dimensional real tensor; Long-range interference feature extraction module: The BigBird model is used to extract long-range interference features based on the extracted multi-scale short-range interference features; Feature fusion module: The acquired multi-scale short-range and long-range interference features are fused through gated residual connections to obtain fused features; Nonlinear transformation module: Performs nonlinear transformation on the obtained fused features to obtain the fused multi-scale temporal features; Output module: The multi-scale temporal features are mapped to the target output space through a fully connected layer to obtain the channel-equalized signal output by the composite channel equalization network; Model optimization module: Based on the dataset constructed by the dataset construction module, optimize and train a composite channel equalization network including the TCN module, BigBird model, gated residual connection, nonlinear transformation and fully connected mapping.
9. An electronic device, characterized in that, Including processor and memory; The memory is used to store computer programs; When the processor executes the program stored in the memory, it implements the channel equalization method according to any one of claims 1-7.
10. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the channel equalization method according to any one of claims 1-7.