A Channel Estimation Method, Device, and Equipment Based on a Deep Stacked Network Structure

Through the channel estimation method of deep stacked network structure, combined with 1D CNN, LSTM, self-attention module and cross-attention module, the channel estimation results are optimized, and the problems of low channel estimation accuracy and high complexity in the non-continuous spectrum OFDM communication system are solved, achieving more efficient channel estimation.

CN120090903BActive Publication Date: 2025-07-01SICHUAN JIUZHOU ELECTRIC GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510558746.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-01
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

When facing the discontinuous spectrum OFDM communication system, the existing channel estimation method has high noise sensitivity and great influence on multipath effect, resulting in low channel estimation accuracy. The pre-data processing of deep learning methods is complex and has average performance.

Method used

The channel estimation method based on the deep stacked network structure is adopted, including 1D CNN, LSTM, self-attention module, cross-attention module and DNN. The channel estimation results are optimized through local feature extraction, timing modeling and global feature fusion.

Benefits of technology

Improve the robustness and accuracy of channel estimation, reduce the complexity of pre-data processing, and enhance the efficiency and performance of channel estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090903B_ABST
    Figure CN120090903B_ABST
Patent Text Reader

Abstract

The present application discloses a channel estimation method, device and equipment based on a deep stacking network structure, which relates to the field of communication technologies. The method first extracts local features from the input data through 1D CNN, and then uses LSTM to perform temporal modeling on the features output by the convolution. At the same time, in order to further improve the feature representation ability, a self-attention mechanism is introduced to capture the dependencies between different positions in the convolution output features; then, through a cross-attention mechanism, the LSTM output and the self-attention features are fused to form a richer multi-dimensional feature representation; finally, the fused multi-dimensional feature representation, the LSTM output and the self-attention features are concatenated to generate a high-dimensional fusion vector, which is input into the DNN for deep non-linear mapping, so as to obtain the final channel estimation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technologies, and particularly relates to a channel estimation method, apparatus, and device based on a deep stacked network structure. Background Art

[0002] The current wireless spectrum allocation mechanism mainly adopts a static authorization mode, that is, authorized users within a specific geographical area have exclusive use rights to exclusive frequency band resources. This rigid licensing mechanism directly leads to low efficiency in spectrum resource allocation. Through empirical analysis, it is found that a considerable proportion of authorized frequency bands exhibit characteristics of persistent low utilization in actual applications, resulting in a dual waste of spectrum resources. Therefore, it is very necessary to study the related technologies of non-continuous spectrum OFDM (Orthogonal Frequency Division Multiplexing) communication systems. In a non-continuous spectrum OFDM communication system under communication-sensing integration, not only the spectrum and environment need to be sensed, but also an accurate channel estimation algorithm is required to ensure the normal operation of the communication system.

[0003] A reliable estimation of the channel response is crucial for subsequent equalization, demodulation, and decoding operations of the receiver, which greatly affects the system performance. The actual communication system may encounter various non-ideal noises and unknown effects, which cannot be well captured by classical estimators, such as the pilot-based Least Squares (LS) estimation. The traditional LS channel estimation method is sensitive to noise and cannot eliminate the multipath effect, resulting in relatively low channel estimation accuracy. The Minimum Mean Squared Error (MMSE) channel estimation method relies on prior information, and it is relatively difficult to obtain the statistical characteristics of the channel in actual applications.

[0004] Deep Learning (DL) has received attention due to its success in computer vision, automatic speech recognition, and natural language processing. In addition, they can achieve low computational complexity, making them very suitable for physical layer applications of communication, especially channel estimation. However, the pre-data processing in the channel estimation method based on deep learning is relatively complex, and the performance is average. Summary of the Invention

[0005] This application proposes a channel estimation method, apparatus, and device based on a deep stacked network structure, which adopts a network for multi-level feature extraction and fusion, takes into account the deep mining of temporal and global features, and finally realizes the correction and optimization of the LS channel estimation error through layer-by-layer optimization, providing higher robustness and better channel estimation performance for the actual communication system.

[0006] This application is implemented through the following technical solutions:

[0007] A channel estimation method based on a deep stacked network structure, comprising:

[0008] Performing preliminary channel estimation on a discontinuous spectrum OFDM communication system by using a least squares channel estimation method, and preprocessing the preliminary channel estimation result; the preprocessing includes: separating the complex-form estimation result into a real part and an imaginary part, and splicing and constructing input data;

[0009] Inputting the preprocessed data into a pre-trained deep stacked network structure for optimization and correction to obtain corrected data; the deep stacked network structure includes 1D CNN, LSTM, self-attention module, cross-attention module, splicing module and DNN; performing local feature extraction on the input data through 1D CNN and respectively inputting the extracted features into LSTM and the self-attention module, performing temporal modeling on the input features through LSTM, and setting a return sequence parameter to retain sequence information, and at the same time capturing the dependencies between different positions in the input features through the self-attention module; fusing the output of LSTM and the output of the self-attention module through the cross-attention module to form a multi-dimensional feature representation; generating a high-dimensional fusion vector by the multi-dimensional feature representation, the output of LSTM and the output of the self-attention module through the splicing module; inputting the high-dimensional fusion vector into DNN for deep non-linear mapping to obtain corrected data;

[0010] Restoring the corrected data to a complex form to obtain the corrected channel estimation result.

[0011] In some embodiments, the training process of the deep stacked network structure includes:

[0012] Establishing a discontinuous spectrum OFDM communication system model, performing preliminary channel estimation by using a least squares channel estimation method, and preprocessing the preliminary channel estimation result, separating the complex-form estimation result into a real part and an imaginary part, and splicing and constructing a training data set;

[0013] Constructing a stacked network, and training the constructed stacked network by using the training data set to obtain a deep stacked network structure;

[0014] The constructed stacked network is composed of 1D CNN, LSTM, DNN, self-attention module, cross-attention module and splicing module.

[0015] In some embodiments, the 1D CNN employs multiple filters with a convolution kernel size of 1 and uses the ReLU activation function. Among them, the convolution kernel with a scale of 1 is used for point-to-point non-linear mapping of each channel, thereby realizing the reconstruction of the feature dimension and the preliminary extraction of local features. The multiple filters are connected to the channel feature map through a set of weights, spanning the feature map along the time axis and calculating the convolution result. Each filter processes the data on different channels, and the data is convolved and summed through a sliding window. The output obtained by convolution is used as the input of the LSTM and the self-attention module. Among them, the number of filters is the same as the dimension of the input data.

[0016] In some embodiments, the LSTM sets multiple hidden units and uses the ReLU activation function, and at the same time sets to return the entire sequence. Among them, the number of hidden units is the same as the dimension of the output data of the 1D CNN.

[0017] The working process of the LSTM includes:

[0018] Combining the information of the current moment and the previous moment to obtain the value of the forget gate, multiplying the obtained value of the forget gate by the cell state of the previous moment, discarding some past information, and retaining important information.

[0019] The information of the previous moment and the current moment passes through the input gate to obtain a gated value. At the same time, the information of the previous moment and the current moment passes through tanh to obtain an information state of the current moment. Multiplying the information state by the gated value determines how much of the input at the current moment is retained and finally flows to the cell state, integrating the information of the current moment into the cell storage state to obtain the cell state of the current moment.

[0020] Combining the information of the current moment and the previous moment to obtain the value of the output gate, multiplying the cell state of the current moment after being activated by tanh by the value of the output gate to determine the final output value.

[0021] The weight parameters and bias parameters in the forget gate, input gate, cell state, and output gate in the LSTM are learned and updated through training.

[0022] In some embodiments, the working process of the self-attention module includes:

[0023] Based on the output of the 1D CNN, combined with the corresponding linear weight matrix, the query vector, key vector, and value vector are respectively calculated.

[0024] Calculating the dot product between the query vector and the key vector and scaling it by a scaling factor to obtain the attention score.

[0025] The attention scores are input into the Softmax function for normalization into a probability distribution, and the obtained attention weights are used to perform weighted summation on the value vectors to obtain the final output;

[0026] The difference between the working process of the cross-attention module and the self-attention module is that in the cross-attention module, the input of the query vector is the output of the LSTM, while the inputs of the key vector and the value vector are the outputs of the self-attention module, and the other processes are the same.

[0027] In some embodiments, the DNN sequentially transmits the input features from the neurons of one layer to the neurons of the next layer. This transmission process is repeated multiple times. In each step, information is extracted and transmitted to the next layer. Each neuron receives weighted inputs from multiple other neurons. These inputs are summed and transmitted to the internal activation function. A hyperparameter is selected to optimize the performance of the model; after the information is transmitted through all layers, an output is generated; the loss function is used to calculate the error between the output value and the true value, and an optimization method is used to minimize this error.

[0028] In a second aspect, the present application proposes a channel estimation device based on a deep stacked network structure. The channel estimation device includes:

[0029] A preprocessing unit that performs preliminary channel estimation on a non-continuous spectrum OFDM communication system using the least squares channel estimation method and preprocesses the preliminary channel estimation result; the preprocessing includes: separating the complex-form estimation result into a real part and an imaginary part, and splicing and constructing input data;

[0030] A correction unit for inputting the preprocessed data into a pre-trained deep stacked network structure for optimization and correction to obtain corrected data; the deep stacked network structure includes 1D CNN, LSTM, self-attention module, cross-attention module, splicing module, and DNN; 1D CNN is used to extract local features from the input data and input the extracted features into the LSTM and self-attention module respectively. The LSTM is used to perform temporal modeling on the input features, and the return sequence parameter is set to retain sequence information. At the same time, the self-attention module captures the dependencies between different positions in the input features; the cross-attention module is used to fuse the LSTM output and the self-attention module output to form a multi-dimensional feature representation; the multi-dimensional feature representation, the LSTM output, and the self-attention module output generate a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into the DNN for deep non-linear mapping to obtain corrected data;

[0031] And an output unit for restoring the corrected data into a complex form to obtain the corrected channel estimation result.

[0032] In some embodiments, the channel estimation device further includes:

[0033] A model construction unit for constructing a stacked network and training to obtain the deep stacked network structure.

[0034] In some embodiments, the channel estimation device is disposed at the receiving end of a non - linear OFDM communication system.

[0035] In a third aspect, the present application proposes an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above - mentioned methods are implemented.

[0036] A channel estimation method based on a deep stacked network structure proposed in the present application first performs local feature extraction on input data through 1D CNN, then uses LSTM to perform temporal modeling on the features output by the convolution. At the same time, in order to further improve the feature representation ability, a self - attention mechanism is introduced to capture the dependencies between different positions in the convolution output features; then, through a cross - attention mechanism, the LSTM output and the self - attention features are fused to form a richer multi - dimensional feature representation; finally, the fused multi - dimensional feature representation, the LSTM output, and the self - attention features are concatenated to generate a high - dimensional fusion vector, and it is input into DNN for deep non - linear mapping, so as to obtain the final channel estimation result. The present application takes into account the in - depth mining of both temporal and global features, and through layer - by - layer optimization, finally realizes the correction and optimization of the LS channel estimation error, improves the channel estimation efficiency at the same time, and provides higher robustness and better channel estimation performance for the actual communication system.

[0037] Correspondingly, a channel estimation device and an electronic device based on a deep stacked network structure proposed in the present application also have the same above - mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings described herein are used to provide a further understanding of the embodiments of the present application, form a part of the present application, and do not limit the embodiments of the present application. In the drawings:

[0039] Figure 1 is a schematic flowchart of the channel estimation method proposed in the embodiments of the present application;

[0040] Figure 2 is a schematic diagram of the stacked network architecture constructed in the embodiments of the present application;

[0041] Figure 3 is a schematic diagram of the structure of the channel estimation device proposed in the embodiments of the present application;

[0042] Figure 4 is a schematic diagram of the principle of the channel estimation system proposed in the embodiments of the present application;

[0043] Figure 5 Schematic diagram of the principle of the electronic device proposed in the embodiment of the present application;

[0044] Figure 6 Schematic diagram of the computer-readable storage medium proposed in the embodiment of the present application;

[0045] Figure 7 Training results of the channel estimation method proposed in the embodiment of the present application in different signal-to-noise ratio datasets;

[0046] Figure 8 Comparison test results between the channel estimation method proposed in the embodiment of the present application and the existing channel estimation methods;

[0047] Figure 9 Comparison results of the running time between the channel estimation method proposed in the embodiment of the present application and the MMSE channel estimation method;

[0048] Reference numerals and corresponding component names:

[0049] 200 - Evaluation device, 201 - Acquisition unit, 202 - First evaluation unit, 203 - Second evaluation unit, 204 - Fusion unit, 300 - Prediction system, 301 - Input device, 302 - Output device, 303 - Processor A, 304 - Memory A, 400 - Electronic device, 410 - Memory B, 420 - Processor B, 411 - Computer program A, 500 - Computer-readable storage medium, 511 - Computer program B. Detailed implementation manners

[0050] In the following, the term "comprise" or "may comprise" that may be used in various embodiments of the present application indicates the presence of the invented functions, operations or elements, and does not limit the addition of one or more functions, operations or elements. Further, as used in various embodiments of the present application, the terms "comprise", "have" and their cognates are only intended to indicate specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be construed as precluding the existence or addition of the possibility of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items.

[0051] In various embodiments of the present application, the expression "or" or "at least one of A or / and B" includes any combination or all combinations of the recited words. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B.

[0052] Expressions (such as "first", "second", etc.) used in various embodiments of the present application may modify various constituent elements in the various embodiments, but do not limit the corresponding constituent elements. For example, the above expressions do not limit the order and / or importance of the elements. The above expressions are only for the purpose of distinguishing one element from other elements. For example, the first user device and the second user device indicate different user devices, although both are user devices. For example, without departing from the scope of the various embodiments of the present application, the first element may be referred to as the second element, and similarly, the second element may also be referred to as the first element.

[0053] It should be noted that: if it is described that one constituent element is "connected" to another constituent element, the first constituent element may be directly connected to the second constituent element, and a third constituent element may be "connected" between the first constituent element and the second constituent element. Conversely, when one constituent element is "directly connected" to another constituent element, it can be understood that there is no third constituent element between the first constituent element and the second constituent element.

[0054] The terms used in the various embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the various embodiments of the present application. As used herein, the singular form is intended to also include the plural form, unless the context clearly indicates otherwise. Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the various embodiments of the present application belong. The terms (such as those defined in a commonly used dictionary) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning, unless clearly defined in the various embodiments of the present application.

[0055] To make the purpose, technical solutions, and advantages of the present application clearer and more understandable, the following further details the present application in conjunction with embodiments and drawings. The illustrative embodiments and descriptions of the present application are only for explaining the present application and do not serve as a limitation to the present application.

[0056] Embodiment 1: To solve the technical problems existing in the existing channel estimation methods, this embodiment proposes a channel estimation method based on a deep stacked network structure. It uses the deep stacked network structure to enhance and correct the traditional channel estimation results, combines a variety of attention mechanisms at the same time, reduces the complexity of the pre-data processing of the existing model, and solves the inaccuracy problem of the least squares channel estimation under low signal-to-noise ratio to a certain extent.

[0057] As Figure 1 shown, the channel estimation method proposed in this embodiment includes the following steps:

[0058] Step 110, perform preliminary channel estimation on the discontinuous spectrum OFDM communication system using the least squares channel estimation method, and preprocess the estimation result. The preprocessing mainly includes: separating the estimation result in complex form into real and imaginary parts, and splicing and constructing the input data.

[0059] Step 120, input the preprocessed data into the pre-trained deep stacked network structure for optimization and correction to obtain the corrected data.

[0060] Among them, the deep stacked network structure includes 1D CNN (one-dimensional convolutional neural network), LSTM (long short-term memory network), self-attention module, cross-attention module, splicing module, and DNN (deep neural network); local features of the input data are extracted by 1D CNN, and then the extracted features are modeled in time series by LSTM, and the return sequence parameter is set to retain sequence information; at the same time, the self-attention module captures the dependencies between different positions in the local features; then, the output of LSTM and the output of the self-attention module are fused through the cross-attention module to form a multi-dimensional feature representation; the fused multi-dimensional feature representation, the output of LSTM, and the output of the self-attention module generate a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into DNN for deep non-linear mapping to obtain the corrected data.

[0061] Step 130, restore the corrected data to complex form to obtain the corrected channel estimation result.

[0062] Optionally, the training process of the above deep stacked network structure includes:

[0063] Establish a discontinuous spectrum OFDM communication system model, perform preliminary channel estimation using the least squares channel estimation method, and preprocess the estimation result. Separate the estimation result in complex form into real and imaginary parts, and splice and construct the training dataset. For the OFDM communication system under discontinuous spectrum, assuming that each OFDM symbol contains N subcarrier input data, the frequency-domain signal can be obtained through modulation first , at this time, the transmitted is serial high-speed data, and the signal is easily affected by inter-symbol interference, thus reducing the communication quality. Therefore, it is necessary to perform serial-to-parallel conversion and insert pilot information first, and then send the signal after IFFT transformation. The time-domain signal is:

[0064]

[0065] Among them, N is the length of IFFT. After IFFT transformation, a cyclic prefix (CP) needs to be inserted into the signal to prevent inter-symbol interference (ISI), and after serial-to-parallel conversion, the newly generated signal It is transmitted through a multipath channel with noise. Thus, the received signal can be obtained :

[0066]

[0067] Where: is the Gaussian white noise in the channel, represents the channel impulse response. The receiver removes the CP in the received signal to obtain , and the frequency-domain signal after FFT transformation is as follows:

[0068]

[0069] Where n = 0, 1, 2, …, N -1. The pilot signal is extracted from the signal after FFT transformation, and then the channel state information at the pilot is obtained. The interpolation algorithm can make it obtain the complete channel state information , and the original data information is accurately restored through signal correction.

[0070] The LS algorithm is used for preliminary channel estimation. The LS algorithm mainly applies it to channel estimation on the premise of ignoring the influence of Gaussian noise on the channel, so as to minimize the following function and obtain the channel state information at the pilot.

[0071]

[0072] Where: represents the channel frequency response at the pilot, represents the received pilot signal, represents the transmitted pilot signal. Let , and thus the channel estimation obtained by the LS algorithm can be obtained:

[0073]

[0074] It can be found from this that the LS algorithm does not need to know the channel statistical prior knowledge during the application of channel estimation. It only needs to use the pilot signal known to both the transmitter and the receiver to perform channel estimation. This algorithm does not need to consider the influence of noise and is relatively simple to implement. However, it is difficult to ignore noise in the actual application process. In the same signal-to-noise ratio environment, the estimation accuracy of this algorithm is not ideal. In order to further improve the performance of LS channel estimation, this embodiment uses the deep stacked network algorithm to accurately estimate the estimation result of the LS algorithm.

[0075] Preprocess the LS channel estimation results. Specifically, separate the complex-valued estimation results into real and imaginary parts, and construct a vector of shape 1×128 as the input data. Since the scale differences of the input variables may increase the training difficulty of the modeling problem, the input data is normalized to zero mean and unit variance.

[0076] Construct a stacked network and use the training dataset to train the constructed network to obtain a deep stacked network structure.

[0077] The constructed stacked network is mainly composed of 1D CNN, LSTM, DNN, self-attention module, cross-attention module, and splicing module, as Figure 2 shown. Among them, 1D CNN uses 128 filters (i.e., the same dimension as the input data), the convolutional kernel size is 1, and the ReLU activation function is used. The convolutional kernel of scale 1 acts like a point-to-point non-linear mapping for each channel, thus realizing the reconstruction of the feature dimension and the preliminary extraction of local features; the 128 filters are connected to the channel feature map through a set of weights, spanning the image along the horizontal direction (time axis) and calculating the convolution result. Each filter processes the data on different channels, and performs convolution summation on the data through a sliding window. Let W be the convolutional filter, then the transformation formula of CNN is:

[0078]

[0079] Among them, b is the bias, f is the activation function, * represents the convolution operation, x is the input data, is the convolutional output. The output obtained by 1DCNN serves as the common feature basis for the subsequent two branches (LSTM and self-attention module).

[0080] LSTM is set with 128 hidden units (i.e., the same dimension as the output data of 1D CNN), and the ReLU activation function is used, and at the same time, return all sequences is set. LSTM includes a forget gate, an input gate, a cell state, an output gate, and a hidden unit vector, which are mainly used to capture the temporal dependencies and dynamic features in the input sequence. The role of the forget gate is to determine how much of the output information from the previous moment needs to be discarded, the role of the input gate is to judge how much of the useful part of the input information at the current moment needs to be retained, and the output gate decides which information to output after integrating the current moment information and the past moment information. The cell state can be regarded as a repository that can retain the previously extracted information for a long time for prediction at the next moment, and the hidden unit vector is the information input to the next moment. And the parameters of the neural network are the same at each time step.

[0081]

[0082] Among them, among them, 、 、 、 are all weight parameters of the LSTM network, 、 、 、 are all bias parameters of the LSTM network, and these parameters are all learned and updated through training. is element-wise multiplication; is the sigmoid function. Based on the above formula, the working principle of LSTM can be divided into three steps:

[0083] In the first step, the value of the forget gate is obtained by combining the information of the current moment and the previous moment , and this value is multiplied by the cell state of the previous moment . A sigmoid function is used to determine how much past information to discard and retain important information.

[0084] In the second step, the gated value is obtained after the information of the previous moment and the current moment pass through the input gate. At the same time, the information of the previous moment and the current moment passes through tanh to obtain an information state of the current moment, and then it is multiplied by the gated value of the input gate to determine how much of the input at the current moment to retain. Finally, it flows to the cell state, integrating the information of the current moment into the cell storage state to obtain the cell state of the current moment.

[0085] In the third step, the value of the output gate is obtained by combining the information of the current moment and the previous moment . The cell state of the current moment after passing through the tanh activation is multiplied by the value of the output gate to determine the final output value . The output of LSTM retains the information change of the input in the time dimension, reflecting the possible temporal correlation in channel estimation.

[0086] At the same time, the output of the CNN is also input into the self-attention module. By calculating the similarity weights between different positions, the important features in the input data are dynamically weighted, strengthening the expression of key local information. The feature output obtained after passing through this module can adaptively capture the dependencies between global features, making up for the deficiency of pure convolution in capturing long-range dependencies. Assuming that the length of the input sequence is L and the dimension of each vector is D, for the input sequence (i.e., the output of the CNN), through three linear transformation matrices of the science department 、 and , respectively calculate the query vector Q , key vector K Sum value vector V :

[0087]

[0088] in, X is the embedding matrix of the input sequence, with dimension L × D ,and X The multiplication is the weight matrix, the dimension is D × D . Then calculate the query vector Q and key vector K The dot product between them and the scaling factor Scaling to get the attention scores:

[0089]

[0090] The attention score is input into the Softmax function to normalize it into a probability distribution, and the obtained attention weight pair value vector is used V Perform weighted summation to get the final output.

[0091] In order to fully integrate the temporal dynamics captured by LSTM and the global features extracted by the self-attention module, the cross-attention module is used to deeply cross-align the LSTM output and the self-attention module output. By calculating the attention weights between each other, the two-way fusion of information is achieved to generate new cross-feature outputs. In this way, the temporal features can be utilized, and the local and global information can be supplemented with self-attention to form a richer feature expression. In cross-attention, the difference from self-attention is that the input element of the query vector Q is the LSTM output, and the input of the key vector K and the value vector V is the self-attention module output. That is, for the LSTM output sequence and the self-attention module output sequence, there are:

[0092]

[0093] in, X is the embedding matrix of the LSTM output sequence, Y is the embedding matrix of the output sequence of the self-attention module.

[0094] Then calculate the attention score and Softmax normalization and output the sum of the weighted value vector of each element in the self-attention module output. The above cross-attention mechanism captures the dependency between the two sequences, enhances the fusion of context information, and improves the accuracy and generalization ability of the task.

[0095] Next, the features from three different sources, namely the LSTM output, the self-attention module output, and the cross-attention module, are concatenated to form a high-dimensional fusion vector. The fused features are then subjected to feature learning through a DNN. First, they undergo a non-linear mapping through a 128-unit layer, and then, through the output layer, the features are mapped to the final 128-dimensional result, which serves as the new channel estimation output. This mapping process occurs in multiple connected layers, each layer containing multiple neurons. Each neuron is a mathematical processing unit that combines with other neurons to learn the relationship between the input features and the output. The DNN sequentially passes the input feature data from the neurons in one layer to the neurons in the next layer. This process is repeated multiple times. At each step, information is extracted and passed to the next layer, and each neuron receives weighted inputs from multiple other neurons. These inputs are summed and passed to an internal activation function, and a hyperparameter is selected to optimize the performance of the model. Once the information has passed through all the layers, as described above, the model generates an output, which is compared with the actual label values in the training dataset.

[0096] Let L be the number of hidden layers and the output layer of the DNN, with nodes in each layer, where 1 ≤ l ≤ L. Consider layer 0 as the input layer, which has nodes. Each artificial neuron or node located at the n th position, such that 1 ≤ n ≤ in the l th layer, receives an input weighted by the vector , plus a bias , and applies an activation function , producing an output:

[0097]

[0098] The output of all the nodes in the l th layer can be expressed as:

[0099]

[0100] where is the weight matrix between layer l −1 and layer l , , is the bias vector, and represents a stack of activation functions. The training process of the network aims to find the optimal weights and biases to approximate a non-linear function for classification or regression tasks. After selecting the network architecture and initializing the weights, we first apply forward propagation to obtain the output values Then, according to an appropriate loss function calculate the error, which represents the difference between the output value and the actual value in the label. Then minimize this error through an optimization method, namely gradient descent with backpropagation.

[0101] Optionally, the channel estimation method proposed in this embodiment can be applied to the receiving end of an OFDM communication system.

[0102] This embodiment also proposes an implementation manner of a channel estimation device 200 based on a deep stacked network structure, as Figure 3 shown. The channel estimation device 200 includes:

[0103] A preprocessing unit 201 that performs preliminary channel estimation on a discontinuous spectrum OFDM communication system using the least squares channel estimation method and preprocesses the estimation result. The preprocessing mainly includes: separating the complex-form estimation result into a real part and an imaginary part, and splicing and constructing the input data.

[0104] A correction unit 202 for inputting the preprocessed data into a pre-trained deep stacked network structure for optimization and correction to obtain corrected data. Among them, the deep stacked network structure includes 1D CNN, LSTM, a self-attention module, a cross-attention module, a splicing module, and DNN; local features of the input data are extracted through 1D CNN, and then the extracted features are modeled in time series using LSTM, and the sequence information is retained by setting the return sequence parameter; at the same time, the self-attention module captures the dependencies between different positions in the local features; then, the output of LSTM and the output of the self-attention module are fused through the cross-attention module to form a multi-dimensional feature representation; the fused multi-dimensional feature representation, the output of LSTM, and the output of the self-attention module generate a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into DNN for deep non-linear mapping to obtain corrected data.

[0105] And an output unit 203 for restoring the corrected data to a complex form to obtain the corrected channel estimation result.

[0106] Optionally, the channel estimation device 200 further includes:

[0107] A model construction unit 204 for constructing a stacked network and training to obtain a deep stacked network structure, and the specific process is as described in the above method and will not be elaborated here.

[0108] Optionally, the channel estimation device 200 can be set at the receiving end of a discontinuous spectrum OFDM communication system or can also be set in other devices of an independent discontinuous spectrum OFDM communication system.

[0109] This embodiment also proposes a channel estimation system 300 based on a deep stacked network structure. As Figure 4 shown, the channel estimation system 300 proposed in this embodiment includes:

[0110] an input device 301, an output device 302, a processor A 303, and a memory A 304; among them, the number of the processor A 303 and the memory A 304 can be one or more, Figure 4 and one processor A 303 and one memory A 304 are taken as examples for illustration. The input device 301, the output device 302, the processor A 303, and the memory A 304 can be connected through a bus or other means, Figure 4 and taking connection through a bus as an example.

[0111] Among them, by invoking the operation instructions stored in the memory A 304, the processor A 303 is used to perform the following steps:

[0112] Perform preliminary channel estimation on the discontinuous spectrum OFDM communication system using the least squares channel estimation method, and preprocess the estimation result; the preprocessing mainly includes: separating the complex-form estimation result into real and imaginary parts, and splicing and constructing the input data;

[0113] Input the preprocessed data into the pre-trained deep stacked network structure for optimization and correction to obtain corrected data;

[0114] Among them, the deep stacked network structure includes 1D CNN, LSTM, self-attention module, cross-attention module, splicing module, and DNN; perform local feature extraction on the input data through 1D CNN, then use LSTM to perform temporal modeling on the extracted features, and set the return sequence parameter to retain the sequence information; at the same time, capture the dependency relationships between different positions in the local features through the self-attention module; then fuse the output of LSTM and the output of the self-attention module through the cross-attention module to form a multi-dimensional feature representation; the fused multi-dimensional feature representation, the output of LSTM, and the output of the self-attention module generate a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into DNN for deep non-linear mapping to obtain corrected data;

[0115] Restore the corrected data to the complex form to obtain the corrected channel estimation result.

[0116] Optionally, by invoking the operation instructions stored in the memory A 304, the processor A 303 is further used to execute any one of the corresponding embodiments in the above channel estimation method.

[0117] This embodiment also proposes an electronic device 400. As Figure 5As shown, the electronic device 400 includes: a memory B410, a processor B420, and a computer program A411 stored on the memory B410 and executable on the processor B420. When the processor B420 executes the computer program A411, the following steps are implemented:

[0118] Adopt the least squares channel estimation method to perform preliminary channel estimation on the discontinuous spectrum OFDM communication system, and preprocess the estimation result; the preprocessing mainly includes: separating the complex-form estimation result into real and imaginary parts, and splicing and constructing the input data;

[0119] Input the preprocessed data into a pre-trained deep stacked network structure for optimization and correction to obtain corrected data;

[0120] Among them, the deep stacked network structure includes 1D CNN, LSTM, self-attention module, cross-attention module, splicing module, and DNN; use 1D CNN to extract local features from the input data, then use LSTM to perform temporal modeling on the extracted features, and set the return sequence parameter to retain sequence information; at the same time, capture the dependencies between different positions in the local features through the self-attention module; then fuse the output of LSTM and the output of the self-attention module through the cross-attention module to form a multi-dimensional feature representation; the fused multi-dimensional feature representation, the output of LSTM, and the output of the self-attention module generate a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into DNN for deep non-linear mapping to obtain corrected data;

[0121] Restore the corrected data to complex form to obtain the corrected channel estimation result.

[0122] Optionally, when the processor B420 executes the computer program A411, any implementation manner in the corresponding embodiment of the above channel estimation method can be implemented.

[0123] It should be noted that the electronic device proposed in this embodiment is the device used to implement the above channel estimation method. Therefore, based on the above channel estimation method proposed in this embodiment, those skilled in the art can understand the specific implementation manner and various variation forms of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic system implements the above channel estimation method will not be described in detail here. As long as the electronic device used by those skilled in the art to implement the above channel estimation method belongs to the scope protected by this application.

[0124] In another embodiment, this embodiment also proposes a computer-readable storage medium 500, as Figure 6 shown, a computer program B511 is stored on the computer-readable storage medium 500, and when the computer program B511 is executed by a processor, the following steps are implemented:

[0125] The least - squares channel estimation method is used to perform preliminary channel estimation on the discontinuous - spectrum OFDM communication system, and the estimation result is pre - processed; the pre - processing mainly includes: separating the complex - form estimation result into a real part and an imaginary part, and splicing and constructing the input data;

[0126] The pre - processed data is input into a pre - trained deep stacked network structure for optimization and correction to obtain corrected data;

[0127] Among them, the deep stacked network structure includes 1D CNN, LSTM, self - attention module, cross - attention module, splicing module and DNN; 1D CNN is used to extract local features of the input data, then LSTM is used to perform temporal modeling on the extracted features, and the return sequence parameter is set to retain sequence information; at the same time, the self - attention module captures the dependencies between different positions in the local features; then, the cross - attention module is used to fuse the output of LSTM and the output of the self - attention module to form a multi - dimensional feature representation; the fused multi - dimensional feature representation, the output of LSTM and the output of the self - attention module generate a high - dimensional fusion vector through the splicing module; the high - dimensional fusion vector is input into DNN for deep non - linear mapping to obtain corrected data;

[0128] The corrected data is restored to the complex form to obtain the corrected channel estimation result.

[0129] Optionally, when the computer program B511 is executed by a processor, it can implement any one of the implementation manners in the corresponding embodiments of the above - mentioned channel estimation method.

[0130] It should be noted that in the above - mentioned embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0131] Embodiment 2:

[0132] In this embodiment, a simulation test is performed on the channel estimation method proposed in the above - mentioned embodiment.

[0133] In this embodiment, NMSE (Normalized Mean Square Error) is used to evaluate the effects of different channel estimation methods. Among them, NMSE is obtained by dividing the original root - mean - square error by the maximum value (or data range) of the signal, and its calculation formula is as follows:

[0134]

[0135] Among them, is the true value, is the predicted value, Nis the number of samples, and max( signal ) represents the maximum value of the signal.

[0136] Since the performance of the network highly depends on the SNR (Signal-to-Noise Ratio) considered during training. Comparatively, the model can provide the best performance when training with data of a certain SNR value. If the same network is trained at different SNR values, the model will not be able to learn meaningful behaviors. Therefore, it is necessary to fix the SNR value for each trained network model. To explore the SNR for training the best model, in this embodiment, datasets with 0 - 40 dB are respectively input into the network for training, and the NMSE results of each model are as Figure 7 shown. As Figure 7 can be seen, when training with a high SNR value, the network can better learn the channel because within this SNR range, the influence of the channel is greater than that of the noise. Thanks to the good generalization property of the network, even when the noise increases, that is, at a low SNR value, it can still estimate the channel. Therefore, even if the expected maximum SNR value is unknown, a relatively high value can be considered as a safety margin. As Figure 7 can be known, the robustness is the best when this value is set to 30 dB, and good estimation performance can be obtained accordingly.

[0137] In this embodiment, the LS channel estimation method, the MMSE channel estimation method, the CRNN channel estimation method, the CLSTM channel estimation method, and the DNN channel estimation method are used as comparative examples, and their NMSE values are compared with those obtained by the channel estimation method (ChanAttNet) proposed in the embodiment of the present application, as Figure 8 shown. As Figure 8 can be seen, the channel estimation method proposed in this embodiment is far superior to the traditional LS channel estimation method and MMSE channel estimation method, and is also superior to other channel estimation methods based on learning networks. The comparison networks constructed are as shown in Table 1 below.

[0138] Table 1 Structures of the comparison models

[0139]

[0140] Finally, the running times of the MMSE channel estimation method and the channel estimation method proposed in the embodiment of the present application are compared, as Figure 9As shown. Since the channel estimation method proposed in the embodiments of this application requires a large amount of initialization work when running for the first time, it will additionally increase the time. Through testing, it can be obtained that the average running time of the MMSE channel estimation method on different signal-to-noise ratio data sets is 2.6322 s, and the average running time of the channel estimation method proposed in this embodiment is 1.053 s. From this, it can be seen that the channel estimation method proposed in the embodiments of this application improves the channel estimation speed to a certain extent.

[0141] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0142] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 a process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.

[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements in the process Figure 1 a process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.

[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 a process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.

[0145] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above description is only the specific embodiments of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A channel estimation method based on a deep stacked network structure, characterized in that: include: The least squares channel estimation method is used to perform preliminary channel estimation on the non-continuous spectrum OFDM communication system, and the preliminary channel estimation results are preprocessed; The preprocessing includes: separating the complex number estimation result into real part and imaginary part, and concatenating them to construct input data; The preprocessed data is input into a pre-trained deep stacked network structure for optimization and correction to obtain corrected data; the deep stacked network structure includes 1D CNN, LSTM, self-attention module, cross-attention module, splicing module and DNN; local features of the input data are extracted by 1D CNN and the extracted features are respectively input into LSTM and self-attention modules, time series modeling of the input features is performed by LSTM, and sequence information is retained by setting return sequence parameters, and the dependency between different positions in the input features is captured by the self-attention module; the LSTM output and the self-attention module output are fused through the cross-attention module to form a multi-dimensional feature expression; the multi-dimensional feature expression and the LSTM output and the self-attention module output are generated into a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into the DNN for deep nonlinear mapping to obtain corrected data; The corrected data is restored to a complex form to obtain a corrected channel estimation result.

2. A channel estimation method based on a deep stacked network structure according to claim 1, characterized in that: The training process of the deep stacked network structure includes: A discontinuous spectrum OFDM communication system model is established, and the least squares channel estimation method is used for preliminary channel estimation. The preliminary channel estimation results are preprocessed, and the complex estimation results are separated into real and imaginary parts, which are then spliced ​​to construct a training data set. Constructing a stacked network, and using the training data set to train the constructed stacked network to obtain a deep stacked network structure; The constructed stacked network consists of 1D CNN, LSTM, DNN, self-attention module, cross-attention module and splicing module.

3. A channel estimation method based on a deep stacked network structure according to claim 1 or 2, characterized in that: The 1D CNN uses multiple filters, a convolution kernel size of 1, and a ReLU activation function; The convolution kernel with a scale of 1 is used to perform point-to-point nonlinear mapping on each channel, thereby realizing the reconstruction of feature dimensions and the preliminary extraction of local features; multiple filters are connected to the channel feature map through a set of weights, spanning the feature map along the time axis and calculating the convolution results, each filter processes data on different channels, and convolves and sums the data through a sliding window; the output obtained by convolution is used as the input of the LSTM and self-attention modules; wherein the number of filters is the same as the dimension of the input data.

4. A channel estimation method based on a deep stacked network structure according to claim 1 or 2, characterized in that: The LSTM is configured with multiple hidden units, and uses a ReLU activation function, and is configured to return all sequences; wherein the number of hidden units is the same as the dimension of the 1D CNN output data; The working process of the LSTM includes: Combine the information of the current moment and the previous moment to get the value of the forget gate, multiply the value of the forget gate by the cell state of the previous moment, discard some of the past information, and retain the important information; The information of the previous moment and the current moment passes through the input gate to obtain a gated value. At the same time, the information of the previous moment and the current moment passes through tanh to obtain an information state of the current moment. The information state and the gated value are multiplied to determine how much input is retained at the current moment, and finally flow to the cell state, integrating the information of the current moment into the cell storage state to obtain the cell state of the current moment; Combine the information of the current moment and the previous moment to get the value of the output gate, multiply the cell state at the current moment by the value of the output gate after tanh activation to determine the final output value; The weight parameters and bias parameters in the forget gate, input gate, cell state, and output gate in LSTM are learned and updated through training.

5. A channel estimation method based on a deep stacked network structure according to claim 1 or 2, characterized in that: The working process of the self-attention module includes: Based on the 1D CNN output and the corresponding linear weight matrix, the query vector, key vector and value vector are calculated respectively; Calculate the dot product between the query vector and the key vector and scale it by a scaling factor to obtain an attention score; The attention score is input into the Softmax function to be normalized into a probability distribution, and the obtained attention weight is used to perform weighted summation on the value vector to obtain the final output; The difference between the working process of the cross-attention module and the self-attention module is that in the cross-attention module, the input of the query vector is the LSTM output, and the input of the key vector and the value vector is the output of the self-attention module, and the other processes are the same.

6. A channel estimation method based on a deep stacked network structure according to claim 1 or 2, characterized in that: The DNN sequentially transmits input features from neurons in one layer to neurons in the next layer. This transmission process is repeated many times. In each step, information is extracted and passed to the next layer. Each neuron receives weighted inputs from multiple other neurons. These inputs are summed and passed to the internal activation function. A hyperparameter is selected to optimize the performance of the model. After the information is passed through all layers, an output is generated. The loss function is used to calculate the error between the output value and the true value, and the error is minimized through optimization methods.

7. A channel estimation device based on a deep stacked network structure, characterized in that: The channel estimation device comprises: A preprocessing unit, which uses a least squares channel estimation method to perform preliminary channel estimation on a non-continuous spectrum OFDM communication system, and preprocesses the preliminary channel estimation result; the preprocessing includes: separating the complex number estimation result into a real part and an imaginary part, and splicing and constructing input data; A correction unit is used to input the preprocessed data into a pre-trained deep stacked network structure for optimization and correction to obtain corrected data; the deep stacked network structure includes 1D CNN, LSTM, self-attention module, cross-attention module, splicing module and DNN; local feature extraction is performed on the input data through 1D CNN and the extracted features are respectively input into LSTM and self-attention modules, time series modeling is performed on the input features through LSTM, and sequence information is retained by setting return sequence parameters, and the dependency relationship between different positions in the input features is captured through the self-attention module; the LSTM output and the self-attention module output are fused through the cross-attention module to form a multi-dimensional feature expression; the multi-dimensional feature expression and the LSTM output and the self-attention module output are generated into a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into the DNN for deep nonlinear mapping to obtain corrected data; And, an output unit is used to restore the corrected data into a complex form to obtain a corrected channel estimation result.

8. A channel estimation device based on a deep stacked network structure according to claim 7, characterized in that: The channel estimation device also includes: The model building unit is used to build a stacked network and train it to obtain the deep stacked network structure.

9. A channel estimation device based on a deep stacked network structure according to claim 7, characterized in that: The channel estimation device is arranged at the receiving end of the nonlinear OFDM communication system.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • OFDM underwater acoustic communication receiving method based on deep learning

    CN119254593A

  • Deep fusion network production line fault prediction method based on deep learning

    CN119357769A