Channel estimation method, device and equipment based on deep cascading network structure

Through a channel estimation method based on a deep stacked network structure, combined with a multi-level feature extraction and fusion network, the problem of low accuracy of existing channel estimation methods in a noisy environment is solved, and higher channel estimation accuracy and robustness are achieved.

CN120090903AActive Publication Date: 2025-06-03SICHUAN JIUZHOU ELECTRIC GROUP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510558746.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing channel estimation method has low accuracy in a noisy environment and cannot effectively eliminate the multipath effect, resulting in large channel estimation errors and affecting the performance of the communication system.

Method used

The channel estimation method based on the deep stacked network structure is adopted, and a multi-level feature extraction and fused network is optimized layer by layer by layer by layer by layer, combining 1D CNN, LSTM, self-attention module, cross-attention module and DNN.

Benefits of technology

Improves the robustness and accuracy of channel estimation and enhances the performance of communication systems, especially in low signal-to-noise ratio environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090903A_ABST
    Figure CN120090903A_ABST
Patent Text Reader

Abstract

The invention discloses a channel estimation method, device and equipment based on a deep cascading network structure, and relates to the technical field of communication, the method comprises the following steps: firstly, carrying out local feature extraction on input data through a 1D CNN, then carrying out time sequence modeling on features output by convolution by using LSTM, and meanwhile, in order to further improve the feature representation capability, carrying out time sequence modeling on the features output by convolution; a self-attention mechanism is introduced to capture a dependency relationship among different positions in convolution output features; lSTM output and self-attention features are fused through a cross attention mechanism, and richer multi-dimensional feature expression is formed; and finally, splicing the fused multi-dimensional feature expression, the LSTM output and the self-attention features to generate a high-dimensional fusion vector, and inputting the high-dimensional fusion vector to a DNN for deep nonlinear mapping, thereby obtaining a final channel estimation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technologies, and particularly relates to a channel estimation method, apparatus, and device based on a deep stacked network structure. Background Art

[0002] The current wireless spectrum allocation mechanism mainly adopts a static authorization mode, that is, authorized users within a specific geographical area have exclusive use rights to exclusive frequency band resources. This rigid licensing mechanism directly leads to low efficiency in spectrum resource allocation. Through empirical analysis, it is found that a considerable proportion of authorized frequency bands exhibit characteristics of persistent low utilization in actual applications, resulting in a double waste of spectrum resources. Therefore, it is very necessary to study the related technologies of non-continuous spectrum OFDM (Orthogonal Frequency Division Multiplexing) communication systems. In a non-continuous spectrum OFDM communication system under communication-perception integration, not only is it necessary to sense the spectrum and environment, but also an accurate channel estimation algorithm is required to ensure the normal operation of the communication system.

[0003] A reliable estimation of the channel response is crucial for subsequent equalization, demodulation, and decoding operations of the receiver, which greatly affects the system performance. Practical communication systems may encounter various non-ideal noises and unknown influences, which cannot be well captured by classical estimators, such as pilot-based Least Squares (LS) estimation. Traditional LS channel estimation methods are sensitive to noise and cannot eliminate the multipath effect, resulting in relatively low channel estimation accuracy. And the Minimum Mean Squared Error (MMSE) channel estimation method relies on prior information, and it is relatively difficult to obtain the statistical characteristics of the channel in practical applications.

[0004] Deep Learning (DL) has received attention due to its success in computer vision, automatic speech recognition, and natural language processing. In addition, they can achieve low computational complexity, making them very suitable for physical layer applications of communication, especially channel estimation. However, the pre-data processing in deep learning-based channel estimation methods is relatively complex, and the performance is generally average. Summary of the Invention

[0005] This application proposes a channel estimation method, apparatus, and device based on a deep stacked network structure, which adopts a network for multi-level feature extraction and fusion, takes into account the deep mining of temporal and global features, and finally realizes the correction and optimization of LS channel estimation errors through layer-by-layer optimization, providing higher robustness and better channel estimation performance for practical communication systems.

[0006] This application is realized through the following technical solutions: A channel estimation method based on a deep stacked network structure, comprising: Performing preliminary channel estimation on a discontinuous spectrum OFDM communication system by using a least squares channel estimation method, and preprocessing the preliminary channel estimation result; the preprocessing includes: separating the estimation result in complex form into a real part and an imaginary part, and splicing and constructing input data; Inputting the preprocessed data into a pre-trained deep stacked network structure for optimization and correction to obtain corrected data; the deep stacked network structure includes 1D CNN, LSTM, a self-attention module, a cross-attention module, a splicing module, and DNN; locally extracting features of the input data through 1D CNN and respectively inputting the extracted features into LSTM and the self-attention module, performing temporal modeling on the input features through LSTM, and setting a return sequence parameter to retain sequence information, and simultaneously capturing the dependencies between different positions in the input features through the self-attention module; fusing the output of LSTM and the output of the self-attention module through the cross-attention module to form a multi-dimensional feature representation; generating a high-dimensional fusion vector from the multi-dimensional feature representation, the output of LSTM, and the output of the self-attention module through the splicing module; inputting the high-dimensional fusion vector into DNN for deep non-linear mapping to obtain corrected data; Restoring the corrected data to complex form to obtain the corrected channel estimation result.

[0007] In some embodiments, the training process of the deep stacked network structure includes: Establishing a discontinuous spectrum OFDM communication system model, performing preliminary channel estimation by using a least squares channel estimation method, and preprocessing the preliminary channel estimation result, separating the estimation result in complex form into a real part and an imaginary part, and splicing and constructing a training dataset; Constructing a stacked network, and training the constructed stacked network by using the training dataset to obtain a deep stacked network structure; The constructed stacked network is composed of 1D CNN, LSTM, DNN, a self-attention module, a cross-attention module, and a splicing module.

[0008] In some embodiments, the 1D CNN uses multiple filters, the convolution kernel size is 1, and the ReLU activation function is used; the convolution kernel with a scale of 1 is used to perform point-to-point non-linear mapping on each channel, thereby realizing the reconstruction of the feature dimension and the preliminary extraction of local features; multiple filters are connected to the channel feature map through a set of weights, spanning the feature map along the time axis and calculating the convolution result, each filter processes the data on different channels, and performs convolution summation on the data through a sliding window; the output obtained by convolution is used as the input of LSTM and the self-attention module; wherein, the number of filters is the same as the dimension of the input data.

[0009] In some embodiments, the LSTM is provided with a plurality of hidden units, and the ReLU activation function is adopted, and at the same time, all sequences are set to be returned; wherein, the number of the hidden units is the same as the dimension of the output data of the 1D CNN; The working process of the LSTM includes: Combining the information of the current moment and the previous moment to obtain the value of the forget gate, multiplying the obtained value of the forget gate by the cell state of the previous moment, discarding some past information, and retaining important information; The information of the previous moment and the current moment passes through the input gate to obtain a gated value. At the same time, the information of the previous moment and the current moment passes through the tanh to obtain an information state of the current moment. Multiplying the information state by the gated value determines how much of the input at the current moment is retained and finally flows to the cell state, integrating the information of the current moment into the cell storage state to obtain the cell state of the current moment; Combining the information of the current moment and the previous moment to obtain the value of the output gate, multiplying the cell state of the current moment after passing through the tanh activation by the value of the output gate to determine the final output value; The weight parameters and bias parameters in the forget gate, input gate, cell state, and output gate in the LSTM are learned and updated through training.

[0010] In some embodiments, the working process of the self-attention module includes: Based on the output of the 1D CNN, combining with the corresponding linear weight matrix, calculating and obtaining the query vector, key vector, and value vector respectively; Calculating the dot product between the query vector and the key vector, and scaling it by a scaling factor to obtain the attention score; Inputting the attention score into the Softmax function to be normalized into a probability distribution, and using the obtained attention weights to perform weighted summation on the value vectors to obtain the final output; The difference between the working process of the cross-attention module and the self-attention module is that in the cross-attention module, the input of the query vector is the output of the LSTM, and the inputs of the key vector and the value vector are the outputs of the self-attention module, and the other processes are the same.

[0011] In some embodiments, the DNN sequentially transmits the input features from the neurons of one layer to the neurons of the next layer. This transmission process is repeated multiple times. In each step, information is extracted and transmitted to the next layer. Each neuron receives weighted inputs from multiple other neurons. These inputs are summed and transmitted to the internal activation function. A hyperparameter is selected to optimize the performance of the model; after the information is transmitted through all layers, an output is generated; the loss function is used to calculate the error between the output value and the true value, and the optimization method is used to minimize this error.

[0012] In a second aspect, the present application proposes a channel estimation device based on a deep stacked network structure. The channel estimation device includes: A preprocessing unit that performs preliminary channel estimation on a discontinuous spectrum OFDM communication system using the least squares channel estimation method and preprocesses the preliminary channel estimation result. The preprocessing includes: separating the complex-form estimation result into a real part and an imaginary part, and splicing and constructing input data; A correction unit that inputs the preprocessed data into a pre-trained deep stacked network structure for optimization and correction to obtain corrected data. The deep stacked network structure includes 1D CNN, LSTM, a self-attention module, a cross-attention module, a splicing module, and DNN. The 1D CNN extracts local features from the input data and inputs the extracted features into the LSTM and the self-attention module respectively. The LSTM performs temporal modeling on the input features and retains sequence information by setting the return sequence parameter. At the same time, the self-attention module captures the dependencies between different positions in the input features. The cross-attention module fuses the output of the LSTM and the output of the self-attention module to form a multi-dimensional feature representation. The multi-dimensional feature representation, the output of the LSTM, and the output of the self-attention module generate a high-dimensional fusion vector through the splicing module. The high-dimensional fusion vector is input into the DNN for deep non-linear mapping to obtain corrected data; And an output unit that restores the corrected data to a complex form to obtain the corrected channel estimation result.

[0013] In some embodiments, the channel estimation device further includes: A model construction unit that constructs a stacked network and trains to obtain the deep stacked network structure.

[0014] In some embodiments, the channel estimation device is disposed at the receiving end of a non-linear OFDM communication system.

[0015] In a third aspect, the present application proposes an electronic device including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above methods are implemented.

[0016] A channel estimation method based on a deep stacked network structure proposed in this application first extracts local features from the input data through 1D CNN, then uses LSTM to perform temporal modeling on the features output by the convolution. At the same time, in order to further improve the feature representation ability, a self-attention mechanism is introduced to capture the dependencies between different positions in the convolution output features; then through a cross-attention mechanism, the LSTM output and the self-attention features are fused to form a richer multi-dimensional feature representation; finally, the fused multi-dimensional feature representation, the LSTM output and the self-attention features are concatenated to generate a high-dimensional fusion vector, and it is input into the DNN for deep non-linear mapping, so as to obtain the final channel estimation result. This application takes into account the in-depth mining of temporal and global features, and through layer-by-layer optimization, finally realizes the correction and optimization of the LS channel estimation error, and at the same time improves the channel estimation efficiency, providing higher robustness and better channel estimation performance for the actual communication system.

[0017] Correspondingly, a channel estimation device and an electronic device based on a deep stacked network structure proposed in this application also have the same above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the embodiments of the present application, form a part of the present application, and do not constitute a limitation on the embodiments of the present application. In the drawings: Figure 1 It is a schematic flow chart of the channel estimation method proposed in the embodiment of the present application; Figure 2 It is a schematic diagram of the stacked network architecture constructed in the embodiment of the present application; Figure 3 It is a schematic diagram of the structure of the channel estimation device proposed in the embodiment of the present application; Figure 4 It is a schematic diagram of the principle of the channel estimation system proposed in the embodiment of the present application; Figure 5 It is a schematic diagram of the principle of the electronic device proposed in the embodiment of the present application; Figure 6 It is a schematic diagram of the computer-readable storage medium proposed in the embodiment of the present application; Figure 7 It is the training result of the channel estimation method proposed in the embodiment of the present application on different signal-to-noise ratio data sets; Figure 8 It is the comparison test result between the channel estimation method proposed in the embodiment of the present application and the existing channel estimation methods; Figure 9 It is the comparison result of the running time between the channel estimation method proposed in the embodiment of the present application and the MMSE channel estimation method; Reference numerals and corresponding component names: 200 - Evaluation device, 201 - Acquisition unit, 202 - First evaluation unit, 203 - Second evaluation unit, 204 - Fusion unit, 300 - Prediction system, 301 - Input device, 302 - Output device, 303 - Processor A, 304 - Memory A, 400 - Electronic device, 410 - Memory B, 420 - Processor B, 411 - Computer program A, 500 - Computer - readable storage medium, 511 - Computer program B. Detailed implementation manners

[0019] In the following, the term "comprising" or "may comprise" that may be used in various embodiments of the present application indicates the presence of an invented function, operation, or element, and does not limit the addition of one or more functions, operations, or elements. Further, as used in various embodiments of the present application, the terms "comprising", "having", and their cognates are only intended to indicate a specific feature, number, step, operation, element, component, or a combination of the foregoing items, and should not be construed as precluding the existence or addition of one or more other features, numbers, steps, operations, elements, components, or a combination of the foregoing items.

[0020] In various embodiments of the present application, the expression "or" or "at least one of A or / and B" includes any combination or all combinations of the recited words. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B.

[0021] Expressions (such as "first", "second", etc.) used in various embodiments of the present application may modify various constituent elements in the various embodiments, but do not limit the corresponding constituent elements. For example, the above - mentioned expressions do not limit the order and / or importance of the elements. The above - mentioned expressions are only for the purpose of distinguishing one element from other elements. For example, the first user device and the second user device indicate different user devices, although both are user devices. For example, without departing from the scope of various embodiments of the present application, the first element may be referred to as the second element, and similarly, the second element may also be referred to as the first element.

[0022] It should be noted that: If it is described that one constituent element is "connected" to another constituent element, the first constituent element may be directly connected to the second constituent element, and a third constituent element may be "connected" between the first constituent element and the second constituent element. Conversely, when one constituent element is "directly connected" to another constituent element, it can be understood that there is no third constituent element between the first constituent element and the second constituent element.

[0023] The terms used in various embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the various embodiments of the present application. As used herein, the singular forms are also intended to include the plural forms unless the context clearly indicates otherwise. Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of the present application belong. The terms (such as those defined in a general-use dictionary) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present application.

[0024] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the following further details the present application in conjunction with embodiments and drawings. The illustrative embodiments and descriptions of the present application are only for explaining the present application and do not serve as a limitation to the present application.

[0025] Embodiment 1: To solve the technical problems existing in the existing channel estimation methods, this embodiment proposes a channel estimation method based on a deep stacked network structure. It uses the deep stacked network structure to enhance and correct the traditional channel estimation results, combines various attention mechanisms at the same time, reduces the complexity of the pre-data processing of the existing model, and solves the inaccuracy problem of the least squares channel estimation under low signal-to-noise ratio to a certain extent.

[0026] As Figure 1 shown, the channel estimation method proposed in this embodiment includes the following steps: Step 110, perform preliminary channel estimation on the non-continuous spectrum OFDM communication system using the least squares channel estimation method, and preprocess the estimation results. The preprocessing mainly includes: separating the complex-form estimation results into real and imaginary parts, and splicing and constructing the input data.

[0027] Step 120, input the preprocessed data into the pre-trained deep stacked network structure for optimization and correction to obtain corrected data.

[0028] Among them, the deep stacked network structure includes 1D CNN (one-dimensional convolutional neural network), LSTM (long short-term memory network), self-attention module, cross-attention module, splicing module, and DNN (deep neural network); the 1D CNN is used to extract local features from the input data, and then the LSTM is used to perform temporal modeling on the extracted features, and the sequence information is retained by setting the return sequence parameter; at the same time, the self-attention module is used to capture the dependencies between different positions in the local features; then the cross-attention module is used to fuse the output of the LSTM and the output of the self-attention module to form a multi-dimensional feature representation; the fused multi-dimensional feature representation, the output of the LSTM, and the output of the self-attention module generate a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into the DNN for deep non-linear mapping to obtain the corrected data.

[0029] Step 130, restore the corrected data to the complex form to obtain the corrected channel estimation result.

[0030] Optionally, the training process of the above deep stacked network structure includes: Build a non-continuous spectrum OFDM communication system model, use the least squares channel estimation method for preliminary channel estimation, and preprocess the estimation result, separate the complex form of the estimation result into real and imaginary parts, and splice and construct a training data set. For the OFDM communication system under non-continuous spectrum, assuming that each OFDM symbol contains N subcarrier input data, first, the frequency-domain signal can be obtained through modulation , at this time, the transmitted data is serial high-speed data, and the signal is easily affected by inter-symbol interference, thus reducing the communication quality. Therefore, it is necessary to perform serial-to-parallel conversion and insert pilot information first, and then send the signal after IFFT transformation. The time-domain signal is:

[0031] where N is the length of the IFFT. After the IFFT transformation, a cyclic prefix (CP) needs to be inserted into the signal to prevent inter-symbol interference (ISI), and after serial-to-parallel conversion, the newly generated signal is sent through a multi-path channel with noise. Thus, the received signal can be obtained:

[0032] where: is the Gaussian white noise in the channel, represents the channel impulse response. The receiving end removes the CP in the received signal to obtain , and the frequency-domain signal after FFT transformation is as follows:

[0033] where n = 0, 1, 2, …, N −1. The pilot signal is extracted from the signal after FFT transformation, and then the channel state information at the pilot is obtained. The interpolation algorithm can be used to obtain the complete channel state information and the original data information is accurately restored through signal correction.

[0034] The LS algorithm is used for preliminary channel estimation. The LS algorithm mainly ignores the influence of Gaussian noise on the channel and applies it to channel estimation, so as to minimize the following function and obtain the channel state information at the pilot.

[0035]

[0036] where: represents the channel frequency response at the pilot, represents the received pilot signal, represents the transmitted pilot signal. Let , and the channel estimation obtained by the LS algorithm can be obtained:

[0037] It can be found from this that the LS algorithm does not need to know the prior knowledge of channel statistics during the application of channel estimation. It only needs to use the pilot signals known at both the transmitter and the receiver to perform channel estimation. This algorithm does not need to consider the influence of noise and is relatively simple to implement. However, it is difficult to ignore noise in actual applications, and the estimation accuracy of this algorithm is not ideal in the same signal-to-noise ratio environment. In order to further improve the performance of LS channel estimation, this embodiment uses the deep stacked network algorithm to accurately estimate the estimation result of the LS algorithm.

[0038] The LS channel estimation result is preprocessed. Specifically, the complex-form estimation result is separated into real and imaginary parts, and a vector with a shape of 1×128 is constructed as the input data. Since the scale difference of the input variables may increase the training difficulty of the modeling problem, the input data is normalized to zero mean and unit variance.

[0039] A stacked network is constructed, and the constructed network is trained using the training data set to obtain the deep stacked network structure.

[0040] The constructed stacked network is mainly composed of 1D CNN, LSTM, DNN, self-attention module, cross-attention module and splicing module, etc., as Figure 2As shown. Among them, the 1D CNN uses 128 filters (i.e., the same as the dimension of the input data), the convolution kernel size is 1, and the ReLU activation function is used. The convolution kernel with a scale of 1 acts like a point-to-point non-linear mapping for each channel, thus realizing the reconstruction of the feature dimension and the preliminary extraction of local features; the 128 filters are connected to the channel feature map through a set of weights, spanning the image along the horizontal direction (time axis) and calculating the convolution result. Each filter processes the data on different channels, and performs convolution summation on the data through a sliding window. Let W be the convolution filter, then the transformation formula of the CNN is:

[0041] Among them, b is the bias, f is the activation function, * represents the convolution operation, x is the input data, is the convolution output. The output obtained by the 1D CNN serves as the common feature basis for the subsequent two branches (LSTM and self-attention module).

[0042] The LSTM is set with 128 hidden units (i.e., the same as the dimension of the 1D CNN output data), and the ReLU activation function is adopted, and at the same time, the return of all sequences is set. The LSTM includes a forget gate, an input gate, a cell state, an output gate, and a hidden unit vector, mainly used to capture the temporal dependencies and dynamic features in the input sequence. The role of the forget gate is to determine how much of the output information from the previous moment needs to be discarded, the role of the input gate is to judge how much of the useful part in the input information at the current moment needs to be retained, and the output gate decides which information to output after integrating the current moment information and the past moment information. The cell state can be regarded as a repository that can retain the previously extracted information for a long time for prediction at the next moment, and the hidden unit vector is the information input to the next moment. And the parameters of the neural network are the same at each time step.

[0043]

[0044] Among them, among them, , , , are all weight parameters of the LSTM network, , , , are all bias parameters of the LSTM network, and these parameters are all learned and updated through training. is the element-wise multiplication; is the sigmoid function. Based on the above formula, the working principle of the LSTM can be divided into three steps: In the first step, the value of the forget gate is obtained by combining the information at the current moment and the previous moment , and this value is multiplied by the cell state at the previous moment . A sigmoid function is used to determine how much past information to discard and retain important information.

[0045] In the second step, the information at the previous moment and the current moment pass through the input gate to obtain a gated value . At the same time, the information at the previous moment and the current moment pass through tanh to obtain a kind of information state at the current moment, and then multiply it by the gated value of the input gate to determine how much of the input at the current moment to retain. Finally, it flows into the cell state, integrating the information at the current moment into the cell storage state to obtain the cell state at the current moment .

[0046] In the third step, the value of the output gate is obtained by combining the information at the current moment and the previous moment , and the cell state at the current moment after being activated by tanh is multiplied by the value of the output gate to determine the final output value . The output of the LSTM retains the information changes of the input in the time dimension, reflecting the possible temporal correlation in channel estimation.

[0047] At the same time, the output of the CNN is also input into the self-attention module. By calculating the similarity weights between different positions, the important features in the input data are dynamically weighted, strengthening the expression of key local information. The feature output obtained after passing through this module can adaptively capture the dependencies between global features, making up for the deficiency of pure convolution in capturing long-range dependencies. Assuming the input sequence length is L and the dimension of each vector is D, for the input sequence (i.e., the output of the CNN), through three linear transformation matrices in the scientific department , and , the query vector Q , the key vector K and the value vector V are calculated respectively:

[0048] Among them, X is the embedding matrix of the input sequence, with a dimension of L × D , and the matrix multiplied by X is the weight matrix, with a dimension of D × D . Then calculate the dot product between the query vector Q and the key vector K , and pass it through the scaling factor Scaling to get the attention scores:

[0049] The attention score is input into the Softmax function to normalize it into a probability distribution, and the obtained attention weight pair value vector is used V Perform weighted summation to get the final output.

[0050] In order to fully integrate the temporal dynamics captured by LSTM and the global features extracted by the self-attention module, the cross-attention module is used to deeply cross-align the LSTM output and the self-attention module output. By calculating the attention weights between each other, the two-way fusion of information is achieved to generate new cross-feature outputs. In this way, the temporal features can be utilized, and the local and global information can be supplemented with self-attention to form a richer feature expression. In cross-attention, the difference from self-attention is that the input element of the query vector Q is the LSTM output, and the input of the key vector K and the value vector V is the self-attention module output. That is, for the LSTM output sequence and the self-attention module output sequence, there are:

[0051] in, X is the embedding matrix of the LSTM output sequence, Y is the embedding matrix of the output sequence of the self-attention module.

[0052] Then calculate the attention score and Softmax normalization and output the sum of the weighted value vector of each element in the self-attention module output. The above cross-attention mechanism captures the dependency between the two sequences, enhances the fusion of context information, and improves the accuracy and generalization ability of the task.

[0053] Next, the features from three different sources, namely the LSTM output, the self-attention module output, and the cross-attention module, are concatenated to form a high-dimensional fusion vector. The fused features are then subjected to feature learning through a DNN. First, they undergo a non-linear mapping through a 128-unit layer, and then, through the output layer, the features are mapped to the final 128-dimensional result, which serves as the new channel estimation output. This mapping process occurs in multiple connected layers, each layer containing multiple neurons. Each neuron is a mathematical processing unit that combines with other neurons to learn the relationship between the input features and the output. The DNN sequentially passes the input feature data from the neurons in one layer to the neurons in the next layer. This process is repeated multiple times. In each step, information is extracted and passed to the next layer, and each neuron receives weighted inputs from multiple other neurons. These inputs are summed and passed to an internal activation function, and a hyperparameter is selected to optimize the performance of the model. Once the information has passed through all the layers, as described above, the model generates an output, which is compared with the actual label values in the training dataset.

[0054] Let L be the number of hidden layers and the output layer of the DNN, and each layer has nodes, where 1 ≤ l ≤ L. Consider layer 0 as the input layer, which has nodes. Each artificial neuron or node located at the n th position, such that 1 ≤ n ≤ of the l th layer, receives inputs weighted by the vector , plus a bias , and applies an activation function , producing an output:

[0055] The outputs of all the nodes in the l th layer can be represented as:

[0056] where is the weight matrix between layer l −1 and layer l , , is the bias vector, and represents a stack of activation functions. The training process of the network aims to find the optimal weights and biases to approximate a non-linear function for classification or regression tasks. After selecting the network architecture and initializing the weights, we first apply forward propagation to obtain the output value . Then, based on an appropriate loss function Calculation error, which represents the difference between the output value and the actual value in the label. Then, this error is minimized through an optimization method, namely gradient descent of backpropagation.

[0057] Optionally, the channel estimation method proposed in this embodiment can be applied to the receiving end of an OFDM communication system.

[0058] This embodiment also proposes an implementation manner of a channel estimation device 200 based on a deep stacked network structure, as Figure 3 shown. The channel estimation device 200 includes: A preprocessing unit 201, which performs preliminary channel estimation on a non - continuous spectrum OFDM communication system using the least - squares channel estimation method and preprocesses the estimation result. The preprocessing mainly includes: separating the complex - form estimation result into real and imaginary parts, and splicing and constructing the input data.

[0059] A correction unit 202, which is used to input the preprocessed data into a pre - trained deep stacked network structure for optimization and correction to obtain corrected data. Among them, the deep stacked network structure includes 1D CNN, LSTM, self - attention module, cross - attention module, splicing module, and DNN; 1D CNN is used to extract local features from the input data, then LSTM is used to perform temporal modeling on the extracted features, and the return sequence parameter is set to retain sequence information; at the same time, the self - attention module captures the dependencies between different positions in the local features; then, the cross - attention module fuses the output of LSTM and the output of the self - attention module to form a multi - dimensional feature representation; the fused multi - dimensional feature representation, the output of LSTM, and the output of the self - attention module generate a high - dimensional fusion vector through the splicing module; the high - dimensional fusion vector is input into DNN for deep non - linear mapping to obtain corrected data.

[0060] And an output unit 203, which is used to restore the corrected data to a complex form to obtain the corrected channel estimation result.

[0061] Optionally, the channel estimation device 200 further includes: A model construction unit 204, which is used to construct a stacked network and train to obtain a deep stacked network structure. The specific process is as described in the above method and will not be elaborated here.

[0062] Optionally, the channel estimation device 200 can be set at the receiving end of a non - continuous spectrum OFDM communication system, or can be set in other devices of an independent non - continuous spectrum OFDM communication system.

[0063] This embodiment also proposes a channel estimation system 300 based on a deep stacked network structure, as Figure 4As shown in the figure, the channel estimation system 300 proposed in this embodiment includes: An input device 301, an output device 302, a processor A 303, and a memory A 304; among them, the number of the processor A 303 and the memory A 304 can be one or more. Figure 4 Here, one processor A 303 and one memory A 304 are taken as examples for illustration. The input device 301, the output device 302, the processor A 303, and the memory A 304 can be connected through a bus or other means. Figure 4 Here, the connection through the bus is taken as an example.

[0064] Among them, by invoking the operation instructions stored in the memory A 304, the processor A 303 is used to perform the following steps: Adopt the least squares channel estimation method to perform preliminary channel estimation on the non - continuous spectrum OFDM communication system, and pre - process the estimation result; the pre - processing mainly includes: separating the complex - form estimation result into real and imaginary parts, and splicing and constructing the input data. Input the pre - processed data into the pre - trained deep stacked network structure for optimization and correction to obtain the corrected data. Among them, the deep stacked network structure includes 1D CNN, LSTM, self - attention module, cross - attention module, splicing module, and DNN; the 1D CNN is used to extract local features of the input data, then the LSTM is used to perform temporal modeling on the extracted features, and the return sequence parameter is set to retain the sequence information; at the same time, the self - attention module captures the dependencies between different positions in the local features; then, the cross - attention module is used to fuse the output of the LSTM and the output of the self - attention module to form a multi - dimensional feature representation; the fused multi - dimensional feature representation, the output of the LSTM, and the output of the self - attention module generate a high - dimensional fusion vector through the splicing module; the high - dimensional fusion vector is input into the DNN for deep non - linear mapping to obtain the corrected data. Restore the corrected data to the complex form to obtain the corrected channel estimation result.

[0065] Optionally, by invoking the operation instructions stored in the memory A 304, the processor A 303 is further used to execute any one of the corresponding embodiments in the above - mentioned channel estimation method.

[0066] This embodiment also proposes an electronic device 400, as Figure 5 shown, the electronic device 400 includes: a memory B 410, a processor B 420, and a computer program A 411 stored on the memory B 410 and operable on the processor B 420. When the processor B 420 executes the computer program A 411, the following steps are implemented: The least - squares channel estimation method is used to perform preliminary channel estimation on the discontinuous - spectrum OFDM communication system, and the estimation result is pre - processed; the pre - processing mainly includes: separating the estimation result in complex form into real and imaginary parts, and splicing and constructing the input data; The pre - processed data is input into a pre - trained deep stacked network structure for optimization and correction to obtain corrected data; Among them, the deep stacked network structure includes 1D CNN, LSTM, self - attention module, cross - attention module, splicing module and DNN; 1D CNN is used to extract local features of the input data, then LSTM is used to perform temporal modeling on the extracted features, and the return sequence parameter is set to retain sequence information; at the same time, the self - attention module captures the dependencies between different positions in the local features; then, the cross - attention module is used to fuse the output of LSTM and the output of the self - attention module to form a multi - dimensional feature representation; the fused multi - dimensional feature representation, the output of LSTM and the output of the self - attention module generate a high - dimensional fusion vector through the splicing module; the high - dimensional fusion vector is input into DNN for deep non - linear mapping to obtain corrected data; The corrected data is restored to complex form to obtain the corrected channel estimation result.

[0067] Optionally, when the processor B420 executes the computer program A411, any implementation manner in the corresponding embodiment of the above - mentioned channel estimation method can be implemented.

[0068] It should be noted that the electronic device proposed in this embodiment is the device used to implement the above - mentioned channel estimation method. Therefore, based on the above - mentioned channel estimation method proposed in this embodiment, those skilled in the art can understand the specific implementation manner and various variation forms of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic system implements the above - mentioned channel estimation method will not be introduced in detail here. As long as the electronic device used by those skilled in the art to implement the above - mentioned channel estimation method belongs to the scope protected by this application.

[0069] In another embodiment, this embodiment also proposes a computer - readable storage medium 500, as Figure 6 shown. A computer program B511 is stored on the computer - readable storage medium 500. When the computer program B511 is executed by a processor, the following steps are implemented: The least - squares channel estimation method is used to perform preliminary channel estimation on the discontinuous - spectrum OFDM communication system, and the estimation result is pre - processed; the pre - processing mainly includes: separating the estimation result in complex form into real and imaginary parts, and splicing and constructing the input data; The pre - processed data is input into a pre - trained deep stacked network structure for optimization and correction to obtain corrected data; Among them, the deep stacking network structure includes 1D CNN, LSTM, self-attention module, cross-attention module, splicing module and DNN; the 1D CNN is used to extract local features of the input data, and then the LSTM is used to perform temporal modeling on the extracted features, and the sequence information is retained by setting the return sequence parameter; at the same time, the self-attention module is used to capture the dependencies between different positions in the local features; then the cross-attention module is used to fuse the output of the LSTM and the output of the self-attention module to form a multi-dimensional feature representation; the fused multi-dimensional feature representation, the output of the LSTM and the output of the self-attention module generate a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into the DNN for deep non-linear mapping to obtain the corrected data; The corrected data is restored to the complex form to obtain the corrected channel estimation result.

[0070] Optionally, when the computer program B511 is executed by a processor, any implementation manner in the corresponding embodiment of the above channel estimation method can be implemented.

[0071] It should be noted that in the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailedly described in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0072] Embodiment 2: In this embodiment, a simulation test is carried out on the channel estimation method proposed in the above embodiment.

[0073] In this embodiment, NMSE (Normalized Mean Square Error) is used to evaluate the effects of different channel estimation methods. Among them, NMSE is obtained by dividing the original root mean square error by the maximum value (or data range) of the signal, and its calculation formula is as follows:

[0074] Among them, is the true value, is the predicted value, N is the number of samples, max( signal ) represents the maximum value of the signal.

[0075] Since the performance of the network highly depends on the SNR (Signal-to-Noise Ratio) considered during training. Comparatively, the model can provide the best performance when training with data at a certain SNR value. If the same network is trained at different SNR values, the model will not be able to learn meaningful behaviors. Therefore, it is necessary to fix the SNR value for each trained network model. To explore the SNR for training the best model, in this embodiment, datasets with SNR values ranging from 0 to 40 dB are respectively input into the network for training, and the NMSE results of each model are as Figure 7 shown. It can be seen from Figure 7 that when training with a high SNR value, the network can better learn the channel because the influence of the channel is greater than that of the noise within this SNR range. Thanks to the good generalization property of the network, even when the noise increases, that is, at a low SNR value, it can still estimate the channel. Therefore, even if the expected maximum SNR value is unknown, a relatively high value can be considered as a safety margin. It can be known from Figure 7 that when this value is set to 30 dB, the robustness is the best, and good estimation performance can be obtained accordingly.

[0076] In this embodiment, the LS channel estimation method, the MMSE channel estimation method, the CRNN channel estimation method, the CLSTM channel estimation method, and the DNN channel estimation method are used as comparative examples, and their NMSE values are compared with those obtained by the channel estimation method (ChanAttNet) proposed in the embodiment of the present application, as Figure 8 shown. It can be seen from Figure 8 that the channel estimation method proposed in this embodiment is far superior to the traditional LS channel estimation method and MMSE channel estimation method, and is also superior to other channel estimation methods based on learning networks. The comparison networks constructed are shown in Table 1 below.

[0077] Table 1 Structures of the comparison models

[0078] Finally, the running times of the MMSE channel estimation method and the channel estimation method proposed in the embodiment of the present application are compared, as Figure 9 shown. Since a large amount of initialization work needs to be carried out when the channel estimation method proposed in the embodiment of the present application runs for the first time, the time will be additionally increased. Through testing, it can be obtained that the average running time of the MMSE channel estimation method on datasets with different signal-to-noise ratios is 2.6322 s, and the average running time of the channel estimation method proposed in this embodiment is 1.053 s. Thus, it can be known that the channel estimation method proposed in the embodiment of the present application improves the speed of channel estimation to a certain extent.

[0079] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0080] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0081] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0083] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above description is only the specific embodiments of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A channel estimation method based on a deep stacked network structure, characterized in that: include: The least squares channel estimation method is used to perform preliminary channel estimation on the non-continuous spectrum OFDM communication system, and the preliminary channel estimation results are preprocessed; The preprocessing includes: separating the complex number estimation result into real part and imaginary part, and concatenating them to construct input data; The preprocessed data is input into a pre-trained deep stacked network structure for optimization and correction to obtain corrected data; the deep stacked network structure includes 1D CNN, LSTM, self-attention module, cross-attention module, splicing module and DNN; local features of the input data are extracted by 1D CNN and the extracted features are respectively input into LSTM and self-attention modules, time series modeling of the input features is performed by LSTM, and sequence information is retained by setting return sequence parameters, and the dependency between different positions in the input features is captured by the self-attention module; the LSTM output and the self-attention module output are fused through the cross-attention module to form a multi-dimensional feature expression; the multi-dimensional feature expression and the LSTM output and the self-attention module output are generated into a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into the DNN for deep nonlinear mapping to obtain corrected data; The corrected data is restored to a complex form to obtain a corrected channel estimation result.

2. A channel estimation method based on a deep stacked network structure according to claim 1, characterized in that: The training process of the deep stacked network structure includes: A discontinuous spectrum OFDM communication system model is established, and the least squares channel estimation method is used for preliminary channel estimation. The preliminary channel estimation results are preprocessed, and the complex estimation results are separated into real and imaginary parts, which are then spliced ​​to construct a training data set. Constructing a stacked network, and using the training data set to train the constructed stacked network to obtain a deep stacked network structure; The constructed stacked network consists of 1D CNN, LSTM, DNN, self-attention module, cross-attention module and splicing module.

3. A channel estimation method based on a deep stacked network structure according to claim 1 or 2, characterized in that: The 1D CNN uses multiple filters, a convolution kernel size of 1, and a ReLU activation function; The convolution kernel with a scale of 1 is used to perform point-to-point nonlinear mapping on each channel, thereby realizing the reconstruction of feature dimensions and the preliminary extraction of local features; multiple filters are connected to the channel feature map through a set of weights, spanning the feature map along the time axis and calculating the convolution results, each filter processes data on different channels, and convolves and sums the data through a sliding window; the output obtained by convolution is used as the input of the LSTM and self-attention modules; wherein the number of filters is the same as the dimension of the input data.

4. A channel estimation method based on a deep stacked network structure according to claim 1 or 2, characterized in that: The LSTM is configured with multiple hidden units, and uses a ReLU activation function, and is configured to return all sequences; wherein the number of hidden units is the same as the dimension of the 1D CNN output data; The working process of the LSTM includes: Combine the information of the current moment and the previous moment to get the value of the forget gate, multiply the value of the forget gate by the cell state of the previous moment, discard some of the past information, and retain the important information; The information of the previous moment and the current moment passes through the input gate to obtain a gated value. At the same time, the information of the previous moment and the current moment passes through tanh to obtain an information state of the current moment. The information state and the gated value are multiplied to determine how much input is retained at the current moment, and finally flow to the cell state, integrating the information of the current moment into the cell storage state to obtain the cell state of the current moment; Combine the information of the current moment and the previous moment to get the value of the output gate, multiply the cell state at the current moment by the value of the output gate after tanh activation to determine the final output value; The weight parameters and bias parameters in the forget gate, input gate, cell state, and output gate in LSTM are learned and updated through training.

5. A channel estimation method based on a deep stacked network structure according to claim 1 or 2, characterized in that: The working process of the self-attention module includes: Based on the 1D CNN output and the corresponding linear weight matrix, the query vector, key vector and value vector are calculated respectively; Calculate the dot product between the query vector and the key vector and scale it by a scaling factor to obtain an attention score; The attention score is input into the Softmax function to be normalized into a probability distribution, and the obtained attention weight is used to perform weighted summation on the value vector to obtain the final output; The difference between the working process of the cross-attention module and the self-attention module is that in the cross-attention module, the input of the query vector is the LSTM output, and the input of the key vector and the value vector is the output of the self-attention module, and the other processes are the same.

6. A channel estimation method based on a deep stacked network structure according to claim 1 or 2, characterized in that: The DNN sequentially transfers input features from neurons in one layer to neurons in the next layer. This transfer process is repeated multiple times. In each step, information is extracted and passed to the next layer. Each neuron receives weighted inputs from multiple other neurons. These inputs are summed and passed to an internal activation function. A hyperparameter is selected to optimize the performance of the model. After the information is passed through all layers, an output is generated. The loss function is used to calculate the error between the output value and the true value, and the error is minimized through an optimization method.

7. A channel estimation device based on a deep stacked network structure, characterized in that: The channel estimation device comprises: A preprocessing unit, which uses a least squares channel estimation method to perform preliminary channel estimation on a non-continuous spectrum OFDM communication system, and preprocesses the preliminary channel estimation result; the preprocessing includes: separating the complex number estimation result into a real part and an imaginary part, and splicing and constructing input data; A correction unit is used to input the preprocessed data into a pre-trained deep stacked network structure for optimization and correction to obtain corrected data; the deep stacked network structure includes 1D CNN, LSTM, self-attention module, cross-attention module, splicing module and DNN; local feature extraction is performed on the input data through 1D CNN and the extracted features are respectively input into LSTM and self-attention modules, time series modeling is performed on the input features through LSTM, and sequence information is retained by setting return sequence parameters, and the dependency relationship between different positions in the input features is captured through the self-attention module; the LSTM output and the self-attention module output are fused through the cross-attention module to form a multi-dimensional feature expression; the multi-dimensional feature expression and the LSTM output and the self-attention module output are generated into a high-dimensional fusion vector through the splicing module; the high-dimensional fusion vector is input into the DNN for deep nonlinear mapping to obtain corrected data; And, an output unit is used to restore the corrected data into a complex form to obtain a corrected channel estimation result.

8. A channel estimation device based on a deep stacked network structure according to claim 7, characterized in that: The channel estimation device also includes: The model building unit is used to build a stacked network and train it to obtain the deep stacked network structure.

9. The channel estimation device based on a deep stacked network structure according to claim 7, characterized in that: The channel estimation device is arranged at the receiving end of the nonlinear OFDM communication system.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • OFDM underwater acoustic communication receiving method based on deep learning

    CN119254593A

  • Deep fusion network production line fault prediction method based on deep learning

    CN119357769A

  • Driver fatigue detection method and system based on combining a pseudo-3d convolutional neural network and an attention mechanism

    US20230154207A1