Channel buffer compression based on deep learning
By using neural networks to compress and decompress channel estimates, the problem of large storage requirements for channel estimates is solved, achieving efficient storage of channel estimates and compression of channel buffers, which is suitable for 5G new radio systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2021-09-03
- Publication Date
- 2026-05-29
AI Technical Summary
In coherent detection, channel estimation requires a large amount of memory, especially in high-bandwidth systems. Existing technologies struggle to effectively compress channel estimation data to save storage space.
A compressed neural network learning method is used to estimate the channel of the reference signal. An autoencoder is used to compress and decompress the channel estimate and then interpolate it to reduce storage requirements.
By compressing channel estimation data, memory requirements are reduced and the efficiency of channel estimation is improved. This method is applicable to various electronic devices, especially channel buffer compression in new 5G radio systems.
Smart Images

Figure CN114257478B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is based on and claims priority to U.S. Provisional Patent Application Serial No. 63 / 083,296, filed September 25, 2020, pursuant to 35 U.S.SC §119(e), the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to deep learning using compressed neural networks. Background Technology
[0004] In coherent detection, the receiver needs to perform a block of channel estimation (CE) (i.e., estimating the channel gain that affects the signal transmitted from the transmitter to the receiver). Furthermore, in systems with multiple receive antennas, CE is necessary for interference cancellation or diversity combining. While incoherent detection methods (e.g., differential phase shift keying) can avoid channel estimation, they result in a 3-4 dB loss in signal-to-noise ratio (SNR).
[0005] CE (Continuous Encoding) is performed with the aid of pilot signals (e.g., demodulation reference signals (DMRS)). In orthogonal frequency division multiplexing (OFDM) systems, the DMRS is inserted into the subcarriers of a symbol at the transmitter. The receiver estimates the channel at the location of the DMRS symbol based on the received signal and stores these estimates in a buffer. To detect the Physical Downlink Shared Channel (PDSCH), data in the CE buffer is extracted and interpolated in the time domain. As the number of subcarriers increases, CE data requires a larger amount of memory. Therefore, when allocating large bandwidths, data compression plays a crucial role in saving memory space. Summary of the Invention
[0006] According to one embodiment, a method for learning using compressed neural networks includes performing channel estimation on a reference signal (RS), compressing the channel estimation of the RS using a neural network, decompressing the compressed channel estimation using a neural network, and interpolating the decompressed channel estimation.
[0007] According to one embodiment, a system for learning using compressed neural networks includes a memory and a processor configured to perform channel estimation on an RS, compress the channel estimation of the RS using a neural network, decompress the compressed channel estimation using a neural network, and interpolate the decompressed channel estimation. Attached Figure Description
[0008] The above and other aspects, features, and advantages of certain embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:
[0009] Figure 1 A schematic diagram of PDSCH-DMRS according to an embodiment is shown;
[0010] Figure 2 A schematic diagram of an automatic encoder according to an embodiment is shown;
[0011] Figure 3 A schematic diagram of an automatic encoder scheme according to an embodiment is shown;
[0012] Figure 4 A schematic diagram of an automatic encoder scheme according to an embodiment is shown;
[0013] Figure 5 A schematic diagram of a network according to an embodiment is shown;
[0014] Figure 6 A schematic diagram of a network according to an embodiment is shown;
[0015] Figure 7 A histogram of the distribution according to an embodiment is shown;
[0016] Figure 8 A histogram of the input features in polar coordinates according to an embodiment is shown;
[0017] Figure 9 A flowchart of a channel estimation method according to an embodiment is shown; and
[0018] Figure 10 A block diagram of an electronic device in a network environment according to an embodiment is shown. Detailed Implementation
[0019] In the following description, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that although the same elements are shown in different drawings, they will be designated by the same reference numerals. In the following description, specific details such as detailed configurations and components are provided merely to aid in a comprehensive understanding of the embodiments of the present disclosure. Therefore, it will be apparent to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Furthermore, for clarity and brevity, descriptions of well-known functions and structures have been omitted. The terminology described below is defined in consideration of the functions in this disclosure and may vary depending on the user, the user's intent, or habits. Therefore, the definitions of the terms should be determined based on the content throughout this specification.
[0020] This disclosure can have various modifications and various embodiments, which will be described in detail below with reference to the accompanying drawings. However, it should be understood that this disclosure is not limited to the embodiments, but includes all modifications, equivalents and substitutions within the scope of this disclosure.
[0021] Although ordinal terms such as first, second, etc., can be used to describe various elements, structural elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, a first structural element may be referred to as a second structural element without departing from the scope of this disclosure. Similarly, a second structural element may also be referred to as a first structural element. As used herein, the term "and / or" includes any and all combinations of one or more related items.
[0022] The terminology used herein is for the purpose of describing various embodiments of this disclosure only and is not intended to limit this disclosure. Singular forms are intended to include plural forms unless the context clearly indicates otherwise. In this disclosure, it should be understood that the terms “comprising” or “having” indicate the presence of features, numbers, steps, operations, structural elements, components, or combinations thereof, and do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, structural elements, components, or combinations thereof.
[0023] Unless otherwise defined, all terms used herein have the same meaning as understood by one of ordinary skill in the art to which this disclosure pertains. Terms such as those defined in commonly used dictionaries shall be interpreted as having the same meaning as in the context of the relevant field, and shall not be construed as having an ideal or overly formal meaning unless expressly defined in this disclosure.
[0024] An electronic device according to one embodiment can be one of various types of electronic devices. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computers, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. According to one embodiment of this disclosure, the electronic device is not limited to those described above.
[0025] The terminology used in this disclosure is not intended to limit the disclosure, but rather to include various variations, equivalents, or alternatives to corresponding embodiments. Regarding the description of the drawings, similar reference numerals may be used to refer to similar or related elements. The singular form of a noun corresponding to an item may include one or more things unless the relevant context clearly indicates otherwise. As used herein, each of the phrases such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B, or C,” “at least one of A, B, and C,” and “at least one of A, B, or C” may include all possible combinations of the items listed together in the corresponding phrase. As used herein, terms such as “first,” “second,” “first,” and “second” may be used to distinguish a corresponding component from another component, but are not intended to limit the components in other respects (e.g., importance or order). The intention is that if an element (e.g., a first element) is referred to as “coupled to another element (e.g., a second element),” “coupled to another element,” “connected to another element,” or “connected to another element” with or without the terms “operably” or “communicably”, this indicates that the element can be coupled to another element directly (e.g., wired), wirelessly, or via a third element.
[0026] As used herein, the term "module" can include units implemented in hardware, software, or firmware, and is used interchangeably with other terms such as "logic," "logic block," "part," and "circuit." A module can be a single integrated component suitable for performing one or more functions, or its smallest unit or part. For example, according to one embodiment, a module can be implemented as an application-specific integrated circuit (ASIC).
[0027] The system and method disclosed in this paper provide an efficient channel buffer compression method for fifth-generation (5G) new radio (NR) baseband modem applications using deep learning techniques. Specifically, the system utilizes an autoencoder to compress / decompress CE data. CE values are stored in a buffer at DMRS symbol locations. To detect PDSCH, CE values are loaded from the buffer and interpolated in the time domain. Since the number of stored channel estimates increases with the number of subcarriers, user equipment (UE) requires a large amount of storage to support high bandwidth. Therefore, CE buffer compression plays a crucial role in saving the space required to store CE data when processing a large number of subcarriers. The system provides training and testing of the autoencoder, which is then split into a compressor and a decompressor.
[0028] Data compression involves representing information with fewer bits, and it falls into categories such as lossless compression and lossy compression. Lossless compression removes redundancy so that no data is lost. Therefore, the original data can be recovered. However, lossy compression achieves better compression by eliminating some less important information. Although the decompressed result of lossy compression is not exactly the same as the original data due to the loss of some information, an acceptable version of the original data can be recovered. In other words, lossy compression achieves better compression at the cost of quality loss. For example, lossy compression results in a significant reduction in the file size of images.
[0029] Autoencoders (i.e., neural networks that copy input to output) are utilized. Autoencoders have been widely used for compression of image, video, time series, and quantum data. For ease of implementation, the autoencoder described in this paper has only one hidden layer. Hyperparameters of the autoencoder, such as the activation function, number of input nodes, batch size, etc., are selected through simulation. A final simulation is performed using the selected hyperparameters to present the results. One challenge is training the autoencoder using quantization values at the hidden layer to mimic the digital storage constraints of the CE (Computational Encoder). In other words, the quantization function lies above the activation function at the hidden layer. However, the gradient of the quantization function is almost zero, except for some points with infinite values. The quantization function can be replaced with a substitution function in backpropagation to obtain acceptable results.
[0030] The following definitions are used throughout the description.
[0031] The two's complement of a bit sequence b is obtained by inverting the bits of b and then adding one.
[0032] The signed fixed-point representation of a real number x (word length W and fractional length F) is a bit sequence b = (b W ,b W-1 ,…,b F ,…,b1), which consists of the symbol bit b W Integer part (b) W-1 ,…,b F+1 ) and the decimal part (b F It consists of (b, ..., b1). It is represented as b′=(b′ W ,b′ W-1 ,…,b′ F ,…,b′1) is the two's complement of b. The decimal value of b is given by equation (1).
[0033]
[0034] The range of b is [-2]. W-F-1 ,2 W-F-1 -2 -FFor example, if b = (1, 0, 1) (decimal length F = 2), then the decimal value of b is 0.25(-4+1) = -0.75. The decimal value of the bit sequence b (W = 3 and F = 2) is in the range [-1, 0.75]. For the real number x, the fixed-point quantization function is shown in equation (2):
[0035]
[0036] And, as shown in equation (3).
[0037]
[0038] In one example, a real number x (word length W and mantissa length is...) The floating-point representation of ) includes the exponent. Sum of last digits bit sequence The two exponent bits are unsigned. Therefore, the decimal value of the exponent is as shown in equation (4).
[0039]
[0040] The mantissa is signed, and if the sign bit is b... l =0, then it is positive. Otherwise, it is negative. When the mantissa is negative, this disclosure uses two's complement to interpret the mantissa. Then, as shown in equation (5).
[0041]
[0042] Therefore, the decimal value of b is given by equation (6).
[0043] 2 E M (6)
[0044] because and The range of b is shown in equation (7).
[0045]
[0046] For example, if the two bits of the exponent wl = 2 and the nine bits of the mantissa (1 = 9) are allocated, then the range of b is [-2048, 2040]. For a real number x, floating-point quantization is as shown in equation (8).
[0047]
[0048] Compression is defined in equation (9).
[0049]
[0050] Vector X = (X1, ..., X2) n ) and its prediction The mean square error (MSE) is defined in equation (10).
[0051]
[0052] Vector X = (X1, ..., X2) n ) and its prediction The normalized mean square error (NMSE) is defined as in equation (11).
[0053]
[0054] Here, the error vector magnitude (EVM) is also used, as defined in equation (12).
[0055]
[0056] EVM is the square root of NMSE. EVM can be expressed as a percentage in equation (13).
[0057]
[0058] According to the 3GPP specification, a Resource Element (RE) is the smallest element in an OFDM system, consisting of a subcarrier in the frequency domain and a symbol in the time domain, while a Resource Block (RB) consists of twelve consecutive subcarriers in the frequency domain within an NR. (Using N...) RB This represents the total number of redundancies (RBs) in the channel bandwidth. Therefore, the total number of subcarriers per symbol is 12 × N. RB First, this disclosure provides details of the system model for NR, followed by details of the system model for LTE. For LTE, Normal Cyclic Prefix (NCP), Extended Cyclic Prefix (ECP), and Multicast Service Single Frequency Network (MBSFN) Reference Signal (MRS) are considered.
[0059] Figure 1 A schematic diagram of PDSCH-DMRS according to an embodiment is shown. For NR, a normal cyclic prefix is considered, therefore each time slot includes 14 symbols. Furthermore, PDSCH-DMRS is used to estimate the channel. Figure 1An example is shown with 6 PDSCH-DMRS subcarriers per RB (in configuration type 1) and 2 PDSCH-DMRS symbols per time slot. For CE estimation, the channel is first estimated for REs with PDSCH-DMRS (for PDSCH-DMRS configuration type 1 in a symbol, this is a total of 2 PDSCH-DMRS symbols x 6 REs per PDSCH-DMRS symbol). These estimates are then interpolated in the frequency domain to estimate the channel of PDSCH in these symbols. Finally, 12 complex channel coefficients are stored per RB. For demodulation, firstly, for each symbol with PDSCH-DMRS, a total of 12 complex channel coefficients are loaded per RB. These coefficients are then interpolated in the time domain to estimate the channel coefficients of PDSCH in the other symbols.
[0060] For each symbol with PDSCH-DMRS, the 12 complex elements per RB are compressed, and the compressed representation of the estimated channel is stored in the CE buffer. For demodulation, the estimated channel values are loaded from the CE buffer and decompressed, and then interpolated in the time domain.
[0061] For a fixed point (FXP) (12s, 12s), each channel coefficient C i =R i +j I i Contains 24 bits, real part R i 12 bits, and the imaginary part I i 12 bits, with a decimal length of zero. Each R i (or I) i The system has 12 bits, ranging from {-2048, ..., -1, 0, 1, ..., 2047}. The system for C... i Polar coordinates are used, which is derived from Cartesian representation.
[0062] For floating-point (FLP) (2u, 9s, 9s), each complex channel coefficient C i It consists of 20 bits. The first two bits are E. i The real and imaginary parts share an exponent, and the next nine bits are M. R,i The mantissa of the real part, and the following nine bits are M. I,i The number of the imaginary part. Each R i (or I) i It has 12 bits and a range of {-2048, ..., -1, 0, 1, ..., 2040}.
[0063] FLP(2u, 9s, 9s) implies that a total of 24 bits are compressed to 20 bits because the real and imaginary parts share a common exponent to suppress the dynamic range of the signal.
[0064] For LTE NCP, the Cell-Specific Reference Signal (CRS) symbols are fixed at the positions of the 0, 4, 7, and 11 OFDM symbols, and are the opposite of the configurable PDSCH-DMRS symbols in NR. In LTE, two CRSs from different antenna ports do not overlap, while in NR, the PDSCH-DMRS use Orthogonal Covering Code (OCC), which allows PDSCH-DMRSs from different antenna ports to overlap.
[0065] Channel coefficients are represented in FXP(11s, 11s) format. Each channel coefficient C i =R i +j I i Contains 22 bits, real part R i 11 bits and imaginary part I i 11 bits. This is represented by FXP(11s, 11s). Each R i (or I) i It has 11 bits and a range of {-1024, ..., -1, 0, 1, ..., 1024}.
[0066] For LTE ECP, the CRS frequency density is the same as LTE NCP, but the position of the CRS symbols differs because there are 12 OFDM symbols per subframe instead of 14 in the case of LTE NCP. In Multimedia Broadcast Multicast Service (MBMS) applications, the MBSFN area always has ECP mode in Multimedia Broadcast Multicast Service Single Frequency Network (MBSFN). The channel coefficients for LTE ECP are represented using the FXP(11s, 11s) format, the same as those for LTE NCP.
[0067] Figure 2 A schematic diagram of an autoencoder 200 according to an embodiment is shown. The autoencoder 200 includes a compressor 202 and a decompressor 204. An autoencoder can be defined as a neural network that attempts to copy its input to its output. Values at the hidden layers can be considered as codes representing the input. For ease of implementation, an autoencoder with one hidden layer is described herein. An important characteristic of this type of neural network is that it only approximates (i.e., only copies inputs that are statistically similar to the training data). A typical use of the autoencoder 200 is to learn low-dimensional representations at the hidden layers from the input, which can be useful for compressor / decompressor pairs.
[0068] The autoencoder 200 can first determine the weights W at the hidden layer. i,j (1) (1≤i≤N and 1≤j≤H), deviation b j (1) The activation function f1(·) will convert the input vector X = (X1, ..., X...) into...N The values mapped to the hidden layer are shown in equation (14).
[0069]
[0070] The autoencoder 200 can then quantize the value at the hidden layer based on fixed-point quantization with word length W and fraction length F: Z j =q FXP (Y j This is uniform scalar quantization. This step is performed because of the digital memory used for the values at the hidden layer. Typical autoencoders do not perform this step.
[0071] The autoencoder 200 can then be based on the weights W at the output layer. i,j (2) Deviation b i (2) The decompressed version of the activation function f2(·) that maps Z to X. As shown in equation (15).
[0072]
[0073] A clipping function (as defined in equation (5)) can be used on top of another function so that f2(·) only produces values within the range of the input. The autoencoder 200 is trained to be configured to result in a small compression loss. Then, it is divided into a compressor 202 that performs the first and second steps described above and a decompressor 204 that performs the third step described above.
[0074] Several autoencoder schemes can be used. For NR and LTE, the same compressor / decompressor is used for all RBs and all symbols. The only difference between NR and LTE compression is the range of channel coefficients. This difference only affects the normalization factor. As described below, for symbols with PDSCH-DMRS, using C1, ..., C 12 Denotes the channel coefficients across one RB, where for 1≤i≤12, C i =R i +jI i Use them separately R represents i I i C i The decompressed version is indicated by R. i I i Normalized before compression and denormalized after decompression. Normalization factors of 2048 and 1024 are applied to LTE and NR respectively, depending on the range of channel coefficients.
[0075] 24 values (12 R values)i and 12I i s)) can be compressed / decompressed using a 24-input autoencoder together, or compressed / decompressed separately using a 12-input autoencoder. Furthermore, C i It can be in FXP or FLP format. Therefore, four options are typically provided. When the input data is in FLP format, the use of a 12-input autoencoder can be ruled out due to G2's preference and the fact that a 24-input autoencoder achieves better performance.
[0076] Figure 3 A schematic diagram of an autoencoder scheme according to an embodiment is shown. The autoencoder scheme 300 includes a compressor 302 and a decompressor 304, and represents 24 inputs C in FXP(12s, 12s) format. i An automatic encoder. A compressor 302 with 24 inputs is used to compress R1, ..., R2 together. 12 and I1,...,I 12 Similarly, a decompressor 304 with 24 outputs is used to decompress 12 real parts and 12 imaginary parts together. This is because 12 C... i Each 24 bits is mapped to H values, each value being W bits, so the compression is as shown in equation (16).
[0077]
[0078] Input C has 12 FXP (12s, 12s) format. i In another autoencoder scheme, a compressor with 12 inputs is used to compress R1, ..., R2 separately. 12 Or I1, ..., I 12 Similarly, a decompressor with 12 outputs is used to decompress either the 12 real parts or the 12 imaginary parts separately. This is because 12 C... i (Each 24 bits) is mapped to 2H values, each value being W bits, so the compression is as shown in equation (17).
[0079]
[0080] Figure 4 A schematic diagram of an autoencoder scheme according to an embodiment is shown. The autoencoder scheme 400 includes a compressor 402 and a decompressor 404. The autoencoder scheme includes 24 inputs C with FLP(2u, 9s, 9s) format. i An automatic encoder. A compressor 402 with 24 inputs is used to compress R1, ..., R2 together. 12 and I1,...,I 12Similarly, a decompressor 404 with 24 outputs is used to decompress 12 real parts and 12 imaginary parts together. The channel coefficients are expanded from 20 bits to 22 bits before compression. This expansion will be compensated for by the overall compression. Due to the 12 C... i (Each 20 bits) bits are mapped to H values, each value being W bits, so the compression is as shown in equation (18).
[0081]
[0082] because, It is FLP(2u, 9s, 9s), from and estimate It's not insignificant. The shared exponent and mantissa portions come from... and calculate.
[0083] When the input is in FLP(2u, 9s, 9s) format, there are six possible compression candidates. Simulation results show that the third scheme achieves the minimum decompression loss. Due to the fact that the LTE channel coefficients are in FXP(11s, 11s) format, the compression ratios of the first and second schemes become as shown in equations (19) and (20), respectively.
[0084]
[0085]
[0086] For quantization, the autoencoder performs quantization at the hidden layer. FXP(12s, 12s) is used in the hidden layer. For quantization, the system can be trained and tested by considering the quantization at the hidden layer only during forward propagation (i.e., ignoring the quantization function in backpropagation). For forward propagation, precise quantization q is used. FXP (x), but for backpropagation, use the following function instead of q. FXP The derivative of (x) is shown in equation (21).
[0087]
[0088] According to equation (3), the quantization function q FXP The range of (x) is [-2]. W-F-1 ,2 W-F-1 -2 -F Therefore, during backpropagation, if the input x of g(x) exceeds the quantizer q... FXP The range of (x), then W i,j (1) It won't change. Otherwise, W i,j (1) It will change and will not be affected by quantification.
[0089] The inputs and outputs of the auto encoder are quantized in the hardware implementation. There are two different implementation formats: 1) FXP (12s, 12s) for inputs and outputs; or, 2) FLP (2u, 9s, 9s) for inputs and outputs.
[0090] Figure 5 A schematic diagram of a network according to an embodiment is shown. Network 500 represents an architecture for FXP(12s, 12s) mode at both input and output. Network 500 includes a channel estimator 502, an autoencoder 504, and an interpolator 506. The autoencoder 504 includes a compressor 510, a channel estimation buffer 512, and a decompressor 514. The output (wideband or narrowband) of the channel estimator 502 is FXP(12s, 12s). The output of the decompressor 514 is also FXP(12s, 12s) before interpolation 506. The number of nodes at the hidden layer is denoted by H, and the word length (W) at each node at the hidden layer is 11. Other values of W are also possible. Interpolation can be performed in time or frequency.
[0091] Figure 6 A schematic diagram of a network according to an embodiment is shown. Network 600 represents an architecture for FLP(2u, 9s, 9s) mode at input and output. Network 600 includes a channel estimator 602, an autoencoder 604, and an interpolator 606. The autoencoder 604 includes a compressor 610, a channel estimation buffer 612, and a decompressor 614. Network 600 also includes a fixed-point to floating-point converter 620 and a floating-point to fixed-point converter 622. Compared to FXP mode, the output format of the channel estimator 602 and the input format of the interpolator 606 are FXP(12s, 12s) format, while the input and output of the autoencoder 604 are FLP(2u, 9s, 9s) format. After decompression, the system converts each pair of real and imaginary parts to FLP(2u, 9s, 9s). This requires calculating a common exponent between the real and imaginary parts. Interpolation can be performed in time or frequency.
[0092] The output of the decompressor is denormalized to obtain the values for 1 ≤ i ≤ 12. and Next, from R i and I i Estimate C i This includes calculating R. i and I i A common index between them. C i The estimated value (or decompressed version) is used In order to express. In order to extract from each pair and Obtain get They are E i M R,i M I,i The estimate. Once the common index is set. There is only one possibility, as shown in equations (22) and (23).
[0093]
[0094]
[0095] The clipping functions in equations (22) and (23) are due to the fact that the mantissa ranges from [-256, 255]. The optimal method yields the minimum error, where the error is shown in equation (24):
[0096]
[0097] in, Get the best One method is brute force, which involves trying... All four possibilities. Therefore, for a single pair and This method may not take a long time. However, since for each C i This must be done, therefore given a large number of samples (10 5 -10 7 In this case, a faster alternative is considered. The activation function of the autoencoder's output is set to produce values within the input range. Therefore,
[0098] Therefore, alternative methods may include estimation. The index and ignore As in equation (25):
[0099]
[0100] As shown in equation (26).
[0101]
[0102] Then, alternative methods can include ignoring estimate The exponent, as in equation (27).
[0103]
[0104] The alternative method can be set as follows, such as equation (28).
[0105]
[0106] Compression can be performed as Cartesian coordinate compression. The input to the neural network used for compression comes from the estimated channel outputs (beam-based channel estimation (BBCE) / non-beam-based channel estimation (NBCE)), which are in I (in-phase) and Q (quadrature) formats. Furthermore, the I and Q signals within a single RB are not expected to change significantly. Therefore, performing compression in Cartesian coordinates is natural. Secondly, in the sense of the signal's dynamic range, the I-axis and Q-axis signals are not completely independent.
[0107] For the loss function in Cartesian coordinates, we use the MSE of equation (10) and the EVM of equation (12). This is because the Euclidean distance is chosen as the "closeness" between two points in Cartesian coordinates.
[0108] Figure 7 Histograms of the distributions according to an embodiment are shown. Histogram 702 is a histogram of the I distribution, and histogram 704 is a histogram of the Q distribution. Figure 7 As shown, the distributions of I and Q are Gaussian. Therefore, there are two methods for Cartesian coordinate compression: 12-node-based compression and 24-node-based compression. For the 12-node-based compression scheme, the neural network is trained by randomly selecting / mixing observations from I and Q, and this network is used to predict / compress both I and Q separately. For the 24-node-based compression, both axes are trained together in the neural network.
[0109] In polar compression, the amplitude and phase of the channel coefficients are compressed, rather than their real and imaginary parts. A single-tap channel with a frequency offset can be fully characterized by three independent variables: slope, the y-intercept of the phase, and amplitude. In other words, for a large number of subcarriers, if the channel model is single-tap, the polar representation (phase and amplitude) of the channel coefficients forms a compressed representation. The same idea can be extended to general channels with multiple taps. The phase and amplitude of the channel coefficients of adjacent subcarriers are correlated.
[0110] For polarization compression, consider two loss functions. The first approach is to simply apply the MSE directly to the amplitude and phase of the polarization representation, as shown in equation (29).
[0111]
[0112] The second approach is to convert the polarization representation to a Cartesian representation and then apply the Euclidean distance to measure the MSE, as shown in equation (30).
[0113]
[0114] Figure 8 Histograms of input features in polar coordinates according to an embodiment are shown. Histogram 802 is a histogram of amplitude, and histogram 804 is a histogram of phase.
[0115] For training, three networks were considered: one for NR-Frequency Range (FR) 1 and LTE-NCP, one for NR-FR2, and one for LTE-ECP. To train these datasets, one of the three datasets described below was selected for its own network, and the samples were shuffled. The selected subset was then split into training and testing sets. The portions allocated to training and testing were 90% and 10%, respectively.
[0116] For LTE and NR coexistence, from a memory perspective, it is beneficial to reuse the weights / biases of a neural network trained on NR-FR1 in LTE Normal Cyclic Prefix (NCP) mode. Simulation results show supporting evidence for the use of the same neural network in both NR-FR1 and LTE-NCP scenarios.
[0117] The NR-FR1 neural network was trained on both BBCE and NBCE with an SNR of {10, 12, ..., 42} dB using the NR-FR1 dataset and random seeds (8 random seeds for subcarrier spacing (SCS) = 15 kHz and 4 seeds 0-3 for CSS = 30 kHz). Table 1 summarizes the training results.
[0118] Table 1
[0119] W F H EVM % Number of samples Compression ratio 10 8 8 0.7290 16698586 3.6 10 8 9 0.5668 16698586 3.2 10 8 10 0.3311 16698586 2.88 11 9 8 0.7088 16698586 3.27 11 9 9 0.5333 16698586 2.91 11 9 10 0.2683 16698586 2.62
[0120] Polar compression and Cartesian compression share the same dataset (NR-FR1 dataset), but the normalization differs. The original samples are in Cartesian format and converted to polar representation using equations (31) and (32):
[0121]
[0122]
[0123] Then it is fed into a 24-node network. In particular, for FXP(12s, 12s), the amplitude is normalized (divided by 2048), as in equation (33):
[0124]
[0125] This makes the amplitude range in [-1, 1], while the phase is normalized as in equation (34).
[0126]
[0127] The phase range is also in [-1, 1].
[0128] As mentioned above, polarization compression employs two types of loss functions. The training results for these two types will be described in the following subsections. Furthermore, due to the different amplitude and phase distributions, this network can only be used with a 24-node network.
[0129] The neural network was trained on both BBCE and NBCE using the NR-FR1 dataset with SNRs of {10, 12, ..., 42} dB on the BBCE / NBCE dataset using random seeds (8 seeds 0-7 for SCS = 15 kHz and 4 seeds 0-3 for SCS = 30 kHz), which is the same as its counterpart in Cartesian coordinates. Table 2 summarizes the training results. The loss function used here is in equation (30).
[0130] Table 2
[0131] H W F EVM % 8 10 8 7.2403 12 10 8 3.8593 16 10 8 0.2433 20 10 8 0.1382 24 10 8 0.1314
[0132] The neural network was trained using the same settings as in the previous section, but with the polarization loss function described above. Table 3 summarizes the training results. The results are similar to those described above regarding the Cartesian loss function.
[0133] Table 3
[0134] H W F EVM % 8 10 8 6.772 12 10 8 3.6256 16 10 8 0.2047 20 10 8 0.1257 24 10 8 0.1166
[0135] The dataset used for the NR-FR2 network was generated from an NR-FR2 setup (SCS = 120 kHz, EPA channel, carrier frequency 28 GHz) with BBCE having SNR ∈ {24, 26, ..., 40}. For NR-FR2, different weights / biases of the neural network need to be trained based on FR2 channel estimation samples, rather than shared with FR1, because compression performance is poor when the NR-FR1 network is directly applied to NR-FR2 simulations.
[0136] Table 4 lists the details of the parameters and training results for the NR-FR2 dataset, including word length, decimal bits, number of hidden layer nodes, number of samples, and compression ratio.
[0137] Table 4
[0138] W F H EVM % Number of samples Compression ratio 10 8 8 0.6916 9785319 3.6 10 8 9 0.6668 9785319 3.2 10 8 10 0.6257 9785319 2.88 11 9 8 0.6557 9785319 3.27 11 9 9 0.6257 9785319 2.91 11 9 10 0.5985 9785319 2.62
[0139] The MRS was configured in LTE ECP mode. Table 5 lists details about the information regarding the training of LTE MRS test cases, including word length, fractional length, number of hidden layer nodes, number of samples, and compression ratio. The EVM values are larger than those trained with PDSCH-DMRS, meaning compression is more difficult for the MRS.
[0140] Table 5
[0141] W F H EVM % Compression ratio 8 6 9 2.65 3.67 8 6 10 1.78 3.3 9 7 9 2.61 3.26 9 7 10 1.72 2.93 10 8 9 2.6 2.93 10 8 10 1.7 2.64 11 9 9 2.59 2.67 11 9 10 1.69 2.4 12 10 9 2.58 2.44 12 10 10 1.69 2.2
[0142] Furthermore, in Table 5, the compression ratio differs from that of the PDSCH-DMRS 24-input node network with FXP mode because the number of bits per input to the network is 11. Therefore, the normalization factor applied to the input during training is 1024 instead of 2048, because the MRS with FXP mode is (11s, 11s) instead of (12s, 12s).
[0143] For modem development, channel estimation is inevitably required for the detection and decoding of transmitted signals. In practice, under pipelined operation, the channel estimation process is quite complex and requires finer timing adjustments to properly coordinate many blocks. To accommodate this need in modem development, channel estimation requires buffers to store intermediate results and control the timing of many client blocks.
[0144] According to this disclosure, the number of elements is reduced via an autoencoder. The neural network identifies how the numbers are related to each other. Thus, six elements can be reduced to, for example, two elements, which can then be represented using an appropriate bit width. As a result, both the number of elements and the bit width can be reduced to improve the compression ratio.
[0145] This disclosure provides per-RB compression. In NR 5G resource allocation, there may be many units for allocating resources, such as Resource Block Groups (RBGs), bundles, REs, RBs, Bandwidth Parts (BWPs), etc. There are trade-offs for each resource allocation unit to be compressed. The system utilizes RB-based compression, which has many reasonable supports. Because the compression unit is a single RB, it can be applied to DMRS, Tracking Reference Signal (TRS), Supplementary Synchronization Signal (SSS), etc.
[0146] This disclosure provides a general application for beam-based and non-beam-based channel estimation. The system is generally applicable to both beam-based and non-beam-based channel estimation algorithms, and is also generally applicable to all RS 202 (DMRS, TRS, SSS, etc.). That is, regardless of the beam configuration, the system is generally applicable to both channel estimation algorithms.
[0147] This disclosure utilizes an autoencoder to provide channel compression, including Cartesian compression and polarization compression. A deep learning architecture, i.e., an autoencoder, is applied per RB to channel buffer compression in 5G NR and is generally applied to both beam-based and non-beam-based channel estimation and to all reference signals. Deep learning algorithms can be applied to compressed channel estimation in 5G modems.
[0148] This disclosure provides neural networks that can be constructed using 24-input or 12-input architectures, influencing network size / complexity, etc. If Cartesian compression is performed, compression per RB (i.e., 12 subcarriers) can be achieved via an autoencoder network with either 24-node inputs (i.e., both real and imaginary values are taken together and the network is used once) or 12-node inputs (i.e., both real and imaginary values are taken sequentially and the network is used twice, once for real values and once for imaginary values). The same applies to polarization compression with amplitude and phase inputs instead of real and imaginary values.
[0149] Figure 9 A flowchart 900 of a channel estimation method according to an embodiment is shown. Any component or combination of components described (i.e., in the device diagram) may be used to perform one or more operations of flowchart 900. The operations described in flowchart 900 are example operations and may involve various additional steps not explicitly provided in flowchart 900. The order of operations described in flowchart 900 is exemplary and not exclusive, as the order may vary depending on the implementation.
[0150] At 902, the system performs channel estimation on the RS. At 904, the system compresses the channel estimation via a neural network. The system can compress the channel estimation at each RB of the RS. At 906, the system decompresses the compressed channel estimation. At 908, the system interpolates the decompressed channel estimation.
[0151] Figure 10 A block diagram of an electronic device 1001 in a network environment 1000 according to one embodiment is shown. (Refer to...) Figure 10In network environment 1000, electronic device 1001 can communicate with electronic device 1002 via a first network 1098 (e.g., a short-range wireless communication network), or with electronic device 1004 or server 1008 via a second network 1099 (e.g., a long-range wireless communication network). Electronic device 1001 can communicate with electronic device 1004 via server 1008. Electronic device 1001 may include processor 1020, memory 1030, input device 1050, sound output device 1055, display device 1060, audio module 1070, sensor module 1076, interface 1077, haptic module 1079, camera module 1080, power management module 1088, battery 1089, communication module 1090, subscriber identification module (SIM) 1096, or antenna module 1097. In one embodiment, at least one component (e.g., display device 1060 or camera module 1080) may be omitted from electronic device 1001, or one or more other components may be added to electronic device 1001. In one embodiment, some of the components may be implemented as a single integrated circuit (IC). For example, sensor module 1076 (e.g., fingerprint sensor, iris sensor, or illuminance sensor) may be embedded in display device 1060 (e.g., display).
[0152] Processor 1020 can execute software (e.g., program 1040) to control at least one other component (e.g., hardware or software component) of electronic device 1001 coupled to processor 1020, and can perform various data processing or calculations. As at least part of data processing or calculation, processor 1020 can load commands or data received from another component (e.g., sensor module 1076 or communication module 1090) into volatile memory 1032, process the commands or data stored in volatile memory 1032, and store the resulting data in non-volatile memory 1034. Processor 1020 may include a main processor 1021 (e.g., central processing unit (CPU) or application processor (AP)) and an auxiliary processor 1023 (e.g., graphics processing unit (GPU), image signal processor (ISP), sensor hub processor, or communication processor (CP)), which may operate independently of main processor 1021 or in conjunction with main processor 1021. Additionally or alternatively, the auxiliary processor 1023 may be adapted to consume less power than the main processor 1021, or to perform specific functions. The auxiliary processor 1023 may be implemented independently of the main processor 1021 or as part of the main processor 1021.
[0153] When the main processor 1021 is inactive (e.g., in sleep) mode, the auxiliary processor 1023, instead of the main processor 1021, can control at least some of the functions or states associated with at least one component of the electronic device 1001 (e.g., display device 1060, sensor module 1076, or communication module 1090). Alternatively, when the main processor 1021 is active (e.g., executing an application), the auxiliary processor 1023 can work with the main processor 1021 to control at least some of the functions or states associated with at least one component of the electronic device 101. According to one embodiment, the auxiliary processor 1023 (e.g., an image signal processor or a communication processor) can be implemented as part of another component (e.g., camera module 1080 or communication module 1090) functionally associated with the auxiliary processor 1023.
[0154] Memory 1030 may store various data used by at least one component of electronic device 1001 (e.g., processor 1020 or sensor module 1076). The various data may include, for example, input or output data of software (e.g., program 1040) and associated commands. Memory 1030 may include volatile memory 1032 or non-volatile memory 1034.
[0155] Program 1040 can be stored as software in memory 1030 and may include, for example, an operating system (OS) 1042, middleware 1044, or application 1046.
[0156] Input device 1050 can receive commands or data from outside electronic device 1001 (e.g., a user) that will be used by other components of electronic device 1001 (e.g., processor 1020). Input device 1050 may include, for example, a microphone, mouse, or keyboard.
[0157] The sound output device 1055 can output sound signals to the outside of the electronic device 1001. The sound output device 1055 may include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as playing multimedia or playing records, while the receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0158] Display device 1060 can visually provide information to the outside of electronic device 1001 (e.g., to a user). Display device 1060 may include, for example, a display, a holographic device, or a projector, and control circuitry that controls a corresponding one of the display, holographic device, and projector. According to one embodiment, display device 1060 may include touch circuitry adapted to detect touch, or sensor circuitry adapted to measure the intensity of the force caused by touch (e.g., a pressure sensor).
[0159] The audio module 1070 can convert sound into electrical signals and vice versa. According to one embodiment, the audio module 1070 can obtain sound via the input device 1050, or output sound via the sound output device 1055 or headphones of the external electronic device 1002 directly (e.g., wired) or wirelessly coupled to the electronic device 1001.
[0160] Sensor module 1076 can detect the operating state of electronic device 1001 (e.g., power or temperature) or the environmental state outside electronic device 1001 (e.g., user state), and then generate an electrical signal or data value corresponding to the detected state. Sensor module 1076 may include, for example, a gesture sensor, gyroscope sensor, atmospheric pressure sensor, magnetic sensor, accelerometer, grip sensor, proximity sensor, color sensor, infrared (IR) sensor, biosensor, temperature sensor, humidity sensor, or illuminance sensor.
[0161] Interface 1077 may support one or more specified protocols used by electronic device 1001 to couple directly (e.g., wired) or wirelessly with external electronic device 1002. According to one embodiment, interface 1077 may include, for example, a High Definition Multimedia Interface (HDMI), a Universal Serial Bus (USB) interface, a Secure Digital Card (SD) interface, or an audio interface.
[0162] Connection terminal 1078 may include a connector via which electronic device 1001 can be physically connected to external electronic device 1002. According to one embodiment, connection terminal 1078 may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0163] The tactile module 1079 can convert electrical signals into mechanical stimuli (e.g., vibration or motion) or electrical stimuli, which can be recognized by a user via touch or kinesthesia. According to one embodiment, the tactile module 1079 may include, for example, a motor, a piezoelectric element, or an electrical stimulator.
[0164] Camera module 1080 can capture still or moving images. According to one embodiment, camera module 1080 may include one or more lenses, an image sensor, an image signal processor, or a flash.
[0165] The power management module 1088 can manage the power supplied to the electronic device 1001. The power management module 1088 can be implemented as at least part of, for example, a power management integrated circuit (PMIC).
[0166] The battery 1089 can supply power to at least one component of the electronic device 1001. According to one embodiment, the battery 1089 may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0167] Communication module 1090 can support the establishment of a direct (e.g., wired) or wireless communication channel between electronic device 1001 and external electronic devices (e.g., electronic device 1002, electronic device 1004, or server 1008), and perform communication via the established communication channel. Communication module 1090 may include one or more communication processors (e.g., APs) that can operate independently of processor 1020, and supports direct (e.g., wired) or wireless communication. According to one embodiment, communication module 1090 may include wireless communication module 1092 (e.g., cellular communication module, short-range wireless communication module, or Global Navigation Satellite System (GNSS) communication module) or wired communication module 1094 (e.g., local area network (LAN) communication module, or power line communication (PLC) module). One of these communication modules can communicate with an external electronic device via a first network 1098 (e.g., a short-range communication network such as Bluetooth™, Wi-Fi Direct, or the Infrared Data Association (IrDA) standard) or a second network 1099 (e.g., a long-range communication network such as a cellular network, the Internet, or a computer network (e.g., a LAN or a wide area network (WAN))). These various types of communication modules can be implemented as a single component (e.g., a single IC) or as multiple components that are separate from each other (e.g., multiple ICs). The wireless communication module 1092 can use subscriber information (e.g., International Mobile Subscriber Identity (IMSI)) stored in the subscriber identification module 1096 to identify and authenticate the electronic device 1001 in the communication network (e.g., the first network 1098 or the second network 1099).
[0168] Antenna module 1097 can transmit or receive signals or power to or from the exterior of electronic device 1001 (e.g., external electronic device). According to one embodiment, antenna module 1097 may include one or more antennas, and thus, for example, communication module 1090 (e.g., wireless communication module 1092) can select at least one antenna suitable for a communication scheme used in a communication network (e.g., a first network 1098 or a second network 1099). Signals or power can then be transmitted or received between communication module 1090 and the external electronic device via the selected at least one antenna.
[0169] At least some of the aforementioned components may be coupled to each other and transmit signals (e.g., commands or data) between them via a peripheral communication scheme (e.g., bus, general purpose input and output (GPIO), serial peripheral interface (SPI), or mobile industrial processor interface (MIPI)).
[0170] According to one embodiment, commands or data can be sent or received between electronic device 1001 and external electronic device 1004 via server 1008 coupled to a second network 1099. Each of electronic devices 1002 and 1004 can be a device of the same or different type as electronic device 1001. All or some operations to be performed on electronic device 1001 can be performed on one or more of external electronic devices 1002, 1004, or 1008. For example, if electronic device 1001 is required to automatically or in response to a request from a user or another device to perform a function or service, electronic device 1001 may request one or more external electronic devices to perform at least a portion of the function or service without performing the function or service, or electronic device 1001 may request one or more external electronic devices to perform at least a portion of the function or service in addition to performing the function or service. The one or more external electronic devices receiving the request may perform at least a portion of the requested function or service, or additional functions or services associated with the request, and transmit the result of the execution to electronic device 1001. Electronic device 1001 may provide the result, with or without further processing, as at least part of a response to the request. For this purpose, cloud computing, distributed computing, or client-server computing technologies can be used, for example.
[0171] One embodiment may be implemented as software (e.g., program 1040) comprising one or more instructions stored in a machine-readable storage medium (e.g., internal memory 1036 or external memory 1038). For example, a processor of electronic device 1001 may invoke at least one of the one or more instructions stored in the storage medium and execute it with or without one or more other components under the control of the processor. Thus, the machine may be operated to perform at least one function according to the invoked at least one instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. The term "non-transitory" indicates that the storage medium is a tangible device and does not include signals (e.g., electromagnetic waves), but the term does not distinguish between cases where data is semi-permanently stored in the storage medium and cases where data is temporarily stored in the storage medium.
[0172] According to one embodiment, the methods disclosed herein can be included and provided in a computer program product. The computer program product can be traded as a product between a seller and a buyer. The computer program product can be distributed in the form of a machine-readable storage medium (e.g., an optical disc read-only memory (CD-ROM)) or via an app store (e.g., the Play Store). TM The computer program product may be distributed online (e.g., downloaded or uploaded) or directly between two user devices (e.g., smartphones). If distributed online, at least a portion of the computer program product may be temporarily generated or at least temporarily stored in a machine-readable storage medium (such as the memory of a manufacturer's server, an app store's server, or a relay server).
[0173] According to one embodiment, each of the above-described components (e.g., a module or program) may include a single entity or multiple entities. One or more of the above-described components may be omitted, or one or more other components may be added. Alternatively or additionally, multiple components (e.g., modules or programs) may be integrated into a single component. In this case, the integrated component may still perform one or more functions of each of the multiple components in the same or similar manner as they were performed by the corresponding one of the multiple components before integration. Operations performed by a module, program, or other component may be performed sequentially, in parallel, repeatedly, or heuristically, or one or more of the operations may be performed in a different order or omitted, or one or more other operations may be added.
[0174] Although certain embodiments of this disclosure have been described in the detailed description thereof, modifications may be made in various forms without departing from the scope of this disclosure. Therefore, the scope of this disclosure should not be determined solely based on the described embodiments, but rather on the appended claims and their equivalents.
Claims
1. A method for deep learning using compressed neural networks, comprising: Channel estimation (CE) is performed on the resource block RB of the reference signal RS to obtain the CE of the RB; The CE of RB is compressed using a neural network to generate RB compressed CE; The RB is compressed and stored in the buffer; A neural network is used to decompress the RB-compressed CE to generate a decompressed CE; and Interpolate the decompressed CE to detect the channel.
2. The method according to claim 1, wherein, The neural network includes an autoencoder.
3. The method according to claim 2, wherein, The autoencoder is a 24-input neural network.
4. The method according to claim 2, wherein, The autoencoder is a 12-input neural network.
5. The method according to claim 1, further comprising: The RB-compressed CE is extracted from the buffer before decompressing the RB-compressed CE.
6. The method according to claim 1, wherein, The CE is compressed using Cartesian compression.
7. The method according to claim 6, wherein, Cartesian compression is performed using mean square error (MSE) and error vector magnitude (EVM).
8. The method according to claim 1, wherein, The CE is compressed via polarization compression.
9. The method according to claim 1, further comprising: When the CE of the RB of RS is represented in floating-point format, estimate the shared exponent.
10. A system for deep learning using compressed neural networks, comprising: buffer, Memory, and The processor is configured as follows: Channel estimation (CE) is performed on the resource block RB of the reference signal RS to obtain the CE of the RB; The CE of RS is compressed using a neural network to generate RB compressed CE; The RB is compressed and stored in the buffer; A neural network is used to decompress the RB-compressed CE to generate a decompressed CE; and Interpolate the decompressed CE to detect the channel.
11. The system according to claim 10, wherein, The neural network includes an autoencoder.
12. The system according to claim 11, wherein, The autoencoder is a 24-input neural network.
13. The system according to claim 11, wherein, The autoencoder is a 12-input neural network.
14. The system according to claim 10, wherein, The processor is also configured to extract the RB-compressed CE from the buffer before decompressing the RB-compressed CE.
15. The system according to claim 10, wherein, The CE is compressed using Cartesian compression.
16. The system according to claim 15, wherein, Cartesian compression is performed using mean square error (MSE) and error vector magnitude (EVM).
17. The system according to claim 10, wherein, The CE is compressed via polarization compression.
18. The system according to claim 10, wherein, The processor is also configured to estimate the shared exponent when the CE of the RB of the RS is represented in floating-point format.