Reinforcement learning assisted 5G NR indoor communication high-order signal demodulation method and system

Through the reinforcement learning-assisted deep Q network demodulation model, the demodulation strategy of 5G NR indoor communication is optimized, which solves the high bit error rate problem caused by multipath loss and noise interference, and achieves higher signal demodulation accuracy and bit error rate reduction.

CN120602292APending Publication Date: 2025-09-05CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510952270.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In 5G NR indoor communications, due to the high demodulation bit error rate caused by multipath loss and environmental random noise, existing methods cannot adapt to the needs of fast interaction in terms of computational complexity.

Method used

A reinforcement learning-assisted method is used to compensate the signal through the least squares algorithm. Then, a high-order modulation and demodulation model based on the deep Q network is designed. The observation state and action space are used in combination with the reward value for training to optimize the demodulation strategy and eliminate signal errors.

Benefits of technology

Significantly reduces the bit error rate under different signal-to-noise ratios, improves signal demodulation accuracy, and adapts to the fast interaction requirements of 5G NR indoor communications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602292A_ABST
    Figure CN120602292A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of 5G NR, particularly discloses a reinforcement learning assisted 5G NR indoor communication high-order signal demodulation method and system, and provides a deep Q network-based 5G NR indoor communication high-order modulation and demodulation model, and the model designs an observation state according to the independent characteristics of signal errors after LS compensation, and provides a deep Q network-based 5G NR indoor communication high-order modulation and demodulation model. And generating an action space according to the modulation information, and exciting the model to adaptively remove adaptive residual noise by taking demodulation accuracy as a reward value to demodulate a symbol of ideal mapping. And taking the initial compensation signal matrix of the system as a feature set needing to process noise errors, and according to the mutual independence of data in the feature set, performing demodulation mapping processing by using the features of a single resource element and further eliminating the errors. Simulation results show that compared with a traditional algorithm, the method is greatly improved, it is guaranteed that the bit error rate under 5G NR indoor communication under different signal-to-noise ratios is greatly reduced, and the accuracy of 5G NR indoor communication signals is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 5G NR technology, and in particular to a reinforcement learning-assisted high-order signal demodulation method and system for 5G NR indoor communications. Background Art

[0002] As the demand for high-speed, high-accuracy information in intelligent communications for 5G NR indoor communications increases, the amount of data for various network services increases, leading to a corresponding increase in the modulation order of 5G signals. This significantly reduces the reliability of data transmission. In real-world indoor communication environments, crowded and densely populated with obstacles, 5G NR signals experience non-stationary signal loss due to reflections from multiple paths and interference from environmental noise. Furthermore, the high frequency band and large bandwidth of 5G NR indoor communication channels, coupled with the large amount of data carried by the subcarriers, exacerbate this dense multipath interference with environmental noise, significantly increasing the demodulation bit error rate (BER) of the received signal.

[0003] To address the low accuracy of demodulated received signals in complex indoor communication environments of 5G NR, practical systems often use the stable and fast least squares (LS) algorithm to quickly estimate interference in the transmission path through two-dimensional interpolation. Artificial intelligence (AI) has flourished in wireless communications in recent years, offering significant advantages in processing large amounts of data and demanding computational tasks. Therefore, existing research focuses on optimizing the LS algorithm to address the fading of multipath reflections and superposition signals. This approach then uses a hard demodulation method based on minimizing the Euclidean distance between each orthogonal frequency division multiplexing (OFDM) symbol in the received frame signal and the actual OFDM symbol. A super-resolution reconstruction channel estimation method is proposed to recover some of the channel information estimated by the LS algorithm and construct an equalizer subnetwork based on this information. Simulation results demonstrate that this method effectively improves LS estimation performance, reduces the impact of Gaussian white noise, and enhances OFDM receiver performance. A SA-CLDNN digital signal demodulator based on single-carrier modulation is proposed, which uses a self-attention mechanism to enhance the neural network's ability to extract amplitude and phase features from low-oversampled data. Simulation experiments show significant improvements in demodulation accuracy for BPSK and M-QAM modulation under Gaussian white noise. Referenced to the 3GPP protocol for OFDM system configuration, a paper proposed a deep convolutional estimation network based on CENet. This network uses the LS algorithm's estimated features as a preliminary feature matrix to learn features and recover LS performance. Simulation results demonstrate that its demodulation constellation mapping is ideal. Another paper proposed an image restoration network based on LS combined with a residual connection for 5G NR OFDM systems. Simulation results show that the mean squared error (MSE) of the signal is significantly reduced under time-varying channels. Another paper proposed a data-driven approach to optimize the MLP, ResNet, and Transformer architectures, respectively, to directly recover the RF input into bit information in an end-to-end manner. Simulations show that this significantly reduces the bit error rate. Another paper proposed a holistic representation demodulation algorithm based on the Transformer architecture combined with a multi-head attention mechanism. This algorithm globally captures the noise interference characteristics caused by Ricean and Rayleigh fading in high-order modulation, improving the demodulation performance of high-order modulation in OFDM systems. Simulation results demonstrate that its ability to compensate for the non-idealities of high-order modulation signals in OFDM systems surpasses that of traditional algorithms. A paper proposes an improved weight optimization algorithm, using an adaptive feedback equalizer to perform demodulation based on the Euclidean distance of the constellation diagram. Simulation results show that under low signal-to-noise ratio conditions, this method significantly improves bit error rate performance compared to coherent demodulation. However, this method is highly dependent on the accuracy of the optimization model, and the computational complexity increases rapidly with network scale, making it unsuitable for the high-speed interactive demands of 5G NR indoor communications. Summary of the Invention

[0004] The present invention provides a reinforcement learning-assisted high-order signal demodulation method and system for 5G NR indoor communications, which solves the technical problem of high demodulation bit error rate caused by multipath loss and environmental random noise under the high-speed interaction requirements of 5G NR indoor communications.

[0005] To solve the above technical problems, the present invention provides a reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communications, including:

[0006] Receive a radio frame signal generated in accordance with the 3GPP protocol and sent by a 5G NR indoor communication base station;

[0007] Compensating the wireless frame signal using a least squares mathematical algorithm to obtain a compensated signal;

[0008] Designing observation states based on the error-independent characteristics of the compensation signal, generating an action space based on the modulation information, and then designing a high-order modulation and demodulation model based on a deep Q network with demodulation accuracy as the reward value;

[0009] Training the designed high-order modulation and demodulation model;

[0010] The compensated signal is demodulated using the trained high-order modulation and demodulation model to obtain a demodulated signal.

[0011] Furthermore, the observed state includes the real and imaginary parts of the frequency domain data S1, the real and imaginary parts of the actual mapped symbol S2, and the real and imaginary parts of the error S3 between the frequency domain data S1 and the actual mapped symbol S2; the action space is a set of actions for selecting the corresponding 5G NR high-order demodulation constellation symbol according to the current state.

[0012] Furthermore, the reward value is designed to be Indicates the correct number of mapping bits corresponding to the randomly selected subcarrier f and the time domain symbol t index, B QAM The number of bits mapped to 64QAM.

[0013] Furthermore, the high-order modulation and demodulation model includes an intelligent agent, a main Q-value network and a predicted Q-value network based on a deep Q network, and an experience recycling pool; the intelligent agent is a high-order signal demodulator after equalization of the communication channel frequency response of a 5G NR receiver;

[0014] During the training process, the agent randomly selects a small batch of samples from the offline deployed experience replay pool to obtain the current state and action s t ,a t Input into the main Q value network to get the current state s t 、Action a t The judgment value Q undert (s t ,a t ; θ), θ represents the latest parameters of the main Q value network; get the next state s t+1 And all its actions are input into the predicted Q value network to generate the next state s t+1 Q-predictions for all actions The network parameter θ of the predicted Q value network clones the main Q value every step length C;

[0015] The agent is also used to t (s t ,a t ;θ), the reward return value r(s) of the current state-action t ,a t ), the maximum Q prediction value of the next state action Learning rate α, decay factor β of future rewards, calculate the current state s t Execute the current action a t Thus entering the new state s t+1 The maximum cumulative reward Q obtained after * (s t ,a t );

[0016] The agent is also used to minimize Q t (s t ,a t ;θ) and Q * (s t ,a t ) is used to update the parameters θ of the main Q-value network.

[0017] Furthermore, the maximum cumulative reward Q * (s t ,a t ) is calculated as:

[0018]

[0019] Furthermore, the parameter θ of the main Q value network is updated by the following formula:

[0020] θ * ←θ-αΔL(θ)

[0021] θ * represents the updated parameter θ, and ΔL(θ) represents the gradient of L(θ).

[0022] Furthermore, the main Q-value network and the predicted Q-value network adopt the same network architecture, including an input layer, a hidden layer and an output layer, the input layer is a state space, and the output layer is an action space; in the inference stage, the observation state corresponding to the compensation signal is input to the main Q-value network, and the main Q-value network outputs the corresponding action to obtain the selected 5G NR high-order demodulation constellation symbol.

[0023] Furthermore, the input layer dimension is 6, the hidden layer dimension is 12, and the output layer dimension is 64.

[0024] Furthermore, each RB of the wireless frame signal generated by the 5G NR indoor communication base station has a total of N scs ×N symbol resource elements, including received symbol information and demodulation reference information, N scs is the frequency domain number, N symbol is the time domain number.

[0025] The present invention also provides a reinforcement learning-assisted high-order signal demodulation system for 5G NR indoor communication, the key of which is: including a signal receiver, a least squares compensator, and a demodulator;

[0026] The signal receiver is used to receive a radio frame signal generated with reference to the 3GPP protocol and sent by a 5G NR indoor communication base station;

[0027] The least square compensator is used to compensate the wireless frame signal using a least square mathematical algorithm to obtain a compensated signal;

[0028] The demodulator is used to design an observation state based on the error-independent characteristics of the compensation signal, generate an action space based on the modulation information, and then use the demodulation accuracy as a reward value to design a high-order modulation and demodulation model based on a deep Q network; and train the designed high-order modulation and demodulation model; and use the trained high-order modulation and demodulation model to demodulate the compensation signal to obtain a demodulated signal.

[0029] The reinforcement learning-assisted high-order signal demodulation method and system for 5G NR indoor communication provided by the present invention refers to the 5G NR 3GPP standard protocol, builds a 5G NR indoor system for data acquisition, and adopts LS two-dimensional interpolation to preliminarily compensate the signal. The preliminary compensation signal matrix of the system is used as the feature set that needs to process the noise error. According to the mutual independence of the data in the feature set, the characteristics of the single resource element are used to perform demodulation mapping processing and further eliminate the error. Therefore, a high-order modulation and demodulation model for 5G NR indoor communication based on deep Q network (Deep Q-Network, DQN) is proposed. The model designs the observation state according to the independent characteristics of the signal error after LS compensation, and generates the action space according to the modulation information. Then, the demodulation accuracy is used as the reward value to motivate the model to adaptively demodulate the ideal mapping symbol by adaptively removing the residual noise. The simulation results show that the present invention has a significant improvement over the traditional algorithm, ensuring that the bit error rate of 5G NR indoor communication under different signal-to-noise ratios is greatly reduced, and the accuracy of 5G NR indoor communication signals is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Schematic diagram of a 5G NR indoor communication environment provided by an embodiment of the present invention;

[0031] Figure 2 is a schematic diagram of a 5G NR resource grid provided by an embodiment of the present invention;

[0032] Figure 3 is a schematic diagram of a high-order modulation and demodulation model provided by an embodiment of the present invention;

[0033] Figure 4 is a structural diagram of a modulation and demodulation system provided by an embodiment of the present invention;

[0034] Figure 5 is a training convergence diagram under different learning rates provided by an embodiment of the present invention;

[0035] Figure 6 This is a comparison chart of BER values ​​of three algorithms provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The following describes the embodiments of the present invention in detail with reference to the accompanying drawings. The embodiments are provided for illustrative purposes only and are not to be construed as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of protection of the present invention. Many changes may be made to the present invention without departing from the spirit and scope of the present invention.

[0037] Figure 1 Figure 1 is a schematic diagram of the 5G NR indoor communication (IC) scenario. Figure 1As shown, the 5G NRIC scenario includes a base station (BS) and a user equipment (UE). The base station sends 5G NR signals and the user equipment receives data. The signal in the IC scenario sends frame signals according to the 3GPP standard protocol, and adopts single-input single-output (SISO) Rayleigh fading combined with Gaussian white noise channel model to transmit signals. Through propagation under multipath reflection paths, the noise interference between frame symbols increases, the amplitude and phase decay significantly, and the LS traditional signal compensation signal effect is poor, resulting in excessive noise error in the received OFDM modulation symbol frequency domain matrix data, and the demodulation bit error rate soars. Therefore, the present invention mainly considers the problem of high demodulation bit error rate caused by poor compensation of multipath and noise interference by the LS algorithm in the IC scenario.

[0038] The radio frame signal sent by the base station is generated according to the 3GPP protocol. Its signal data is carried by resource elements (RE) and is modulated by OFDM to form Figure 2 Multi-carrier resource grid transmission.

[0039] Figure 2 A RB in scs ×N symbol RE, where N scs =12, N symbol =14. The blue RE grid represents the received symbol information, and the green RE represents the demodulation reference signal (DMRS), i.e., the pilot signal. The signal data matrix transmitted in the channel is as follows: Figure 2 Data is loaded in the form of . The X direction is the time domain OFDM symbol index, and the Y direction is the frequency domain subcarrier index. After the UE obtains the resource information, it parses the frame signal, then performs OFDM demodulation to extract the received data in the frequency domain. Based on this data matrix, the known symbol data at the DMRS RE during transmission is extracted and the LS algorithm is used to compensate the signal in two-dimensional interpolation. The compensated signal still has high noise and residual multipath fading errors, and the constellation points under high-order modulation are very dense. The demodulation mapping deviation between symbols is too large, and the output bit error rate BER is still high. Moreover, this part of the error changes with the indoor communication environment, and it is difficult to learn this feature in a supervised manner. Therefore, a reinforcement learning network is used to autonomously learn its optimal demodulation.

[0040] The demodulation target of the present invention is to obtain the number of erroneous bits at all subcarriers and OFDM symbols. The sum of the two bits is divided by the total number of bits sent, B, to obtain a lower bit error rate to evaluate the demodulation performance, as shown in formula (1):

[0041]

[0042] In order to achieve a lower bit error rate, an embodiment of the present invention provides a reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communication, which specifically includes the following steps:

[0043] Receive a radio frame signal generated in accordance with the 3GPP protocol and sent by a 5G NR indoor communication base station;

[0044] The wireless frame signal is compensated using a least squares (LS) mathematical algorithm to obtain a compensated signal;

[0045] The observation state is designed based on the error-independent characteristics of the compensation signal, and the action space is generated based on the modulation information. Then, the demodulation accuracy is used as the reward value to design a high-order modulation and demodulation model based on the deep Q network;

[0046] Train the designed high-order modulation and demodulation model;

[0047] The trained high-order modulation and demodulation model is used to demodulate the compensation signal to obtain a demodulated signal.

[0048] Based on a reinforcement learning network-based demodulation model, the system analyzes the characteristics of equalized symbol data, sets appropriate state and action spaces, and designs a reward method. Through interaction between the environment and the agent, the system gradually acquires the optimal demodulation strategy for removing small noise. Ultimately, the system converges to a minimum bit error rate (BER), achieving excellent 5G NR high-order demodulation performance.

[0049] The LS algorithm can eliminate some noise in NR channel transmission, but there is still a small amount of noise. The 5G NR transmitter uses high-order modulation, and its constellation mapping points are relatively dense and the demodulation threshold is small, which leads to demodulation misjudgment and high BER value. To solve the problem of high demodulation BER due to misjudgment of mapped symbols by small noise, the present invention proposes the following method: Figure 3 The DQN-based demodulation model shown in Figure 1. The DQN algorithm adaptively optimizes the demodulation strategy through interaction with the environment. By using experience replay and the target network, the correlation of data and the non-smooth convergence of demodulation can be weakened, and the optimal demodulation strategy π is finally obtained. * .

[0050] Depend on Figure 3 It can be seen that the present invention first constructs a Markov decision process (MDP) model for 5G NR high-order signal demodulation, and then combines the DQN algorithm to solve the optimal demodulation strategy π *The high-order signal demodulator after equalization of the 5G NR receiver communication channel frequency response is used as the intelligent agent, and the optimal demodulation result is found through continuous training. The specific definitions of state S, action A, and reward R are as follows.

[0051] 1) State S: S = {Re(S1),Im(S1),Re(S2),Im(S2),Re(S3),Im(S3)}, where S1 represents the frequency domain data obtained by equalizing the response guessed from the LS channel, S2 represents the actual mapped symbol, and S3 represents the error S3 between the equalized frequency domain data S1 and the actual mapped symbol S2. This error is the small noise that was not eliminated in the previous step. Re(·) and Im(·) represent the real and imaginary parts of the complex data. A total of six state spaces are set, and a subcarrier in a frame is randomly selected for state selection.

[0052] 2) Action A: The slave agent selects a set of corresponding 5G NR high-order demodulation constellation symbols based on the current state, i.e., the action set.

[0053] 3) Reward R: The agent obtains the reward by performing the corresponding action in the current state. Value is returned as a reward. Indicates the correct number of mapping bits corresponding to the randomly selected subcarrier f and the time domain symbol t index, B QAM =6 is the number of bits mapped by 64QAM.

[0054] The demodulation model consists of a DQN main Q-value network and a DQN predicted Q-value network. The former is used to generate Q estimates for actions and evaluate the current NR downlink channel demodulation error performance, while the latter is used to predict the error performance of the previous main Q-value network under the same parameters. This method creates two such networks to improve stability and convergence, and uses the mean square error loss function to update the target network.

[0055] Experience revisit pool storage process: First restart the environment and randomly obtain the corresponding state s under a subcarrier t , and then select the current state s through the greedy strategy selection action function of the intelligent agent t The next action a t , a t Return to the environment step function to find the reward value r t and the next state S t+1 , and then (s t ,a t ,r t ,S t+1 ) is stored in the replay pool. When the pool is full, the latest experience is saved based on the first-in-first-out method.

[0056] The model learning process is as follows.

[0057] First, the agent randomly selects a small batch of samples from the offline deployed experience replay pool to obtain the current state and action s t ,a t Input into the DQN main Q value network to obtain the judgment value Q under the current state t (s t ,a t ;θ),Q t (s t ,a t ) represents the current evaluation Q value network, and θ represents the latest network parameters. The predicted Q value network does not update the network parameters immediately, but clones the main Q value network weights at a certain step length C to generate the predicted Q value network evaluation value Represents the Q prediction value of all actions in the next state, and the network parameters are the same as the main Q value network. Then, execute strategy a in the environment t , get instant reward r t With the next state S t+1 Finally, (s t ,a t ,r t ,S t+1 ) is stored in the experience replay pool to form new experience for the next training step.

[0058] Evaluate the optimal state s during the learning process t Next take action a t The value of Q * (s t ,a t ), is to maximize the future discounted benefits, indicating the current state s t Execute current strategy a t Thus entering the new state s t+1 The maximum cumulative reward obtained after , its update formula is shown in (2):

[0059]

[0060] Among them, α is the learning rate, β is the decay factor of future rewards, r(s t ,a t ) is the reward return value of the current state-action, Indicates the next state s t+1 The optimal action a * The corresponding maximum Q prediction value.

[0061] The present invention updates the parameters of the main Q value estimation network by minimizing the mean square error MSE, that is, formula (3):

[0062] L(θ)=E[Q * (s t ,a t)-Q(s t ,a t ;θ)] 2 (3)

[0063] The update of the network parameters of the main Q value network through back propagation and stochastic gradient descent can be expressed as formula (4):

[0064] θ * ←θ-αΔL(θ) (4)

[0065] According to the ε-greedy greedy strategy, we randomly select actions in the current state and explore new actions to find possible optimal demodulation symbols. The strategy expression is shown in (5):

[0066]

[0067] Where A is the action space set, Rand(·) represents randomly selecting an action in the action space, ε represents the greedy probability factor, It means finding the action corresponding to the maximum Q value of all possible actions in the current state. Represents the best action to choose.

[0068] The 5G NR transmitter 64QAM modulated signal is transmitted over a Rayleigh channel. After the LS algorithm equalizes the channel response, the data is demodulated using the DQN algorithm. The specific DQN demodulation algorithm process is shown in Table 1 below, which specifically includes the following steps:

[0069] 1. Initialize the 5G NR communication environment and get the current state s t (initial state s0);

[0070] 2. From the current state s t , select an action a according to the ε-greedy greedy strategy t ;

[0071] 3. 5G NR communication environment execution action a t ;

[0072] 4. Calculate execution action a t Instant Rewards t ;

[0073] 5. Restart the 5G NR communication environment and randomly select the state of the frame signal subcarrier as the next state S t+1 ;

[0074] 6. The experience generated (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool;

[0075] 7. Set the next state S t+1 As the current state, repeat steps 2 to 6. When the experience replay pool meets the minimum batch size, proceed to the next step;

[0076] 8. Randomly draw Γ experience samples from the experience replay pool;

[0077] 9. The agent inputs the experience sample into the main Q value network and the predicted Q value network to obtain the Q value Q under the current state-action respectively. t (s t ,a t ;θ), and the Q prediction value of all actions in the next state

[0078] 10. According to Q t (s t ,a t ;θ), the reward return value r(s) of the current state-action t ,a t ), the maximum Q prediction value of the next state action Learning rate α, decay factor β of future rewards, calculate the current state s t Execute the current action a t Thus entering the new state s t+1 The maximum cumulative reward Q obtained after * (s t ,a t ), as shown in formula (2);

[0079] 11. By minimizing Q t (s t ,a t ;θ) and Q * (s t ,a t ) to update the parameters of the main Q value network, as shown in Equations (3) and (4);

[0080] 12. Follow steps 3 to 6.

[0081] 13. When the parameter update number of the main Q value network reaches the step size C, the predicted Q value network copies the weight parameters of the main Q value network;

[0082] 14. After the learning of this batch of samples is completed, return to step 1 and enter the next round of learning until the number of training rounds n is reached episode or time stop flag env done ;

[0083] 15. After training is completed, the main Q value network outputs the current state s t The optimal action.

[0084] Table 1 High-order signal demodulation algorithm based on DQN

[0085]

[0086]

[0087] It should be noted that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This embodiment is not limited here.

[0088] Corresponding to the above-mentioned reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communication, the present invention also provides a reinforcement learning-assisted high-order signal demodulation system for 5G NR indoor communication, which is characterized by comprising a signal receiver, a least squares compensator, and a demodulator.

[0089] Corresponding to the method, the signal receiver is used to receive a radio frame signal generated with reference to the 3GPP protocol and sent by a 5G NR indoor communication base station;

[0090] The least square compensator is used to compensate the wireless frame signal using a least square mathematical algorithm to obtain a compensated signal;

[0091] The demodulator is used to design an observation state based on the error-independent characteristics of the compensation signal, generate an action space based on the modulation information, and then design a high-order modulation and demodulation model based on a deep Q network with demodulation accuracy as a reward value; and train the designed high-order modulation and demodulation model; and use the trained high-order modulation and demodulation model to demodulate the compensation signal to obtain a demodulated signal.

[0092] The embodiments described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components.

[0093] Computer programs for implementing the methods and systems of the present invention can be written in any combination of one or more programming languages ​​and stored in a computer-readable storage medium. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0094] A computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.

[0095] The following section focuses on the design of DQN high-order signal demodulation after compensating for signal errors affected by the channel based on LS estimation of channel state information. Simulation experiments are also conducted. The hardware platform used in the experiment is a PC with an Intel(R) Xeon(R) Gold 6242R CPU @ 3.10GHz, an NVIDIA RTX3080Ti GPU, and 64GB of memory.

[0096] This embodiment takes the 5G NR standard to generate subcarriers as an example. The modulation and demodulation system modeling process is as follows: Figure 4 As shown. The simulation model is developed using the Pytorch_2.3.1 framework combined with Python_3.11.9. A resource grid is generated to transmit data between the user equipment (UE) and the base station (BS). Each resource grid contains 12 subcarriers with a subcarrier spacing of 15kHz, a carrier frequency of 2.4GHz, and a bandwidth of 180kHz. The pilot symbols are based on Figure 2 The position setting is as follows. According to the protocol, each time slot can be configured with a maximum of 106 resource blocks, which is set to 84 resource blocks in this paper. The frequency domain data matrix size for one time slot transmission in this paper is a complex number of 1008×14.

[0097] like Figure 4As shown in the figure, the resource blocks generated by the transmitter undergo 64QAM modulation and resource mapping, and then OFDM modulation is performed using the inverse fast Fourier transform (IFFT). According to the protocol, the number of FFT points is set to 1024 to ensure spectral resolution. During the modulation process, a cyclic prefix (CP) with a length of 144 is added to prevent multipath interference between symbols. In order to simulate the actual frequency fading situation, this paper adopts the Rayleigh fading channel model with a complex Gaussian distribution, and the standard deviation is set to 0.5. During the channel transmission process, the amplitude of the signal will attenuate and will be affected by the Gaussian white noise n k,t The receiver removes the cyclic prefix (CP) and performs an FFT operation to demodulate the signal in the frequency domain, thereby estimating the channel. A least squares algorithm is then used to compensate the signal, and a deep Q-network (DQN) is used to perform deep denoising and demodulation to restore the original data.

[0098] A high-order signal demodulation model based on the least squares method was built using the Pytorch framework. Each round was set to 50 time steps. The state was updated once every time step, and the reward value was accumulated once every round. The hidden layer of the main Q-value network and the predicted Q-value network was set to a dimension of 12. The input layer was the state space (dimension 6), and the output layer was the action space (dimension 64). Other training parameters are shown in Table 2:

[0099] Table 2 DQN high-order signal demodulation model training parameter settings

[0100]

[0101] During the network learning process, frame signals with different signal-to-noise ratios can be sent flexibly. This paper uses a signal-to-noise ratio of 25dB for network learning and randomly sends frame signals for training. By adjusting the learning rate, the training convergence can be obtained. Figure 5 As shown in . Because the number of time steps is set to 50, and the maximum reward value each time is 1, the round reward value limit is 50, Figure 5 It can be seen that when the learning rate is set to 0.0001, the learning convergence value is larger and more stable. Therefore, the model with a learning rate of 0.0001 is saved for algorithm performance analysis.

[0102] The performance of different 5G NR demodulation algorithms is compared based on the preserved DQN high-order signal demodulation model. 100 frames of signal are sent online, and the average bit error rate is calculated every 10 frames. The bit error rate is calculated as shown in Equation (1). Figure 6The comparison of different signal-to-noise ratios (SNRs) for the LS-based hard demodulator LS_HardDecision, the LMMSE-based hard demodulator LMMSE_HardDecision, and the proposed DQN-based high-order signal demodulator is shown. It can be seen that while the BER value of the LS_HardDecision algorithm decreases as the SNR increases, it remains too high. The LMMSE_HardDecision algorithm further removes some noise based on the LS algorithm, resulting in a lower BER compared to the LS-based hard demodulator, reaching a BER of 0.150 at a SNR of 25dB. However, the BER is too high when the noise is high. The proposed DQN_Decision demodulator achieves a significantly lower SNR than the other two algorithms. Comparing LS_HardDecision with LMMSE_HardDecision, the average BER reduction for different SNRs reaches 76.4% and 65.0%, respectively.

[0103] In summary, to address the issues of excessive bit error rates (BERs) and inaccurate received information in high-order signal demodulation for 5G NR indoor communications, embodiments of the present invention provide a reinforcement learning-assisted high-order signal demodulation method and system for 5G NR indoor communications. This method first compensates the transmitted signal based on the traditional LS algorithm, and then eliminates residual errors in the signal at the demodulation end using a DQN demodulation model. The DQN demodulation model analyzes the error characteristics of the LS-compensated signal for indoor communications. Based on the mutual independence of these error characteristics, a single-subcarrier demodulator design is designed, which online collects all subcarriers of the frame signal for experience pool learning. The DQN demodulation model can eliminate errors at the demodulation end in the presence of multipath effects and random noise interference in indoor communications. Simulation experiments demonstrate that the present invention achieves a lower BER (bit error rate), reducing the BER by 76.4% and 65.0% compared to the LS hard demodulation and LMMSE hard demodulation algorithms. This provides more accurate reception performance for 5G NR indoor communications.

[0104] The above embodiments are preferred implementations of the present invention, but the implementations of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communications, characterized in that: include: Receive a radio frame signal generated in accordance with the 3GPP protocol and sent by a 5G NR indoor communication base station; Compensating the wireless frame signal using a least squares mathematical algorithm to obtain a compensated signal; Designing observation states based on the error-independent characteristics of the compensation signal, generating an action space based on the modulation information, and then designing a high-order modulation and demodulation model based on a deep Q network with demodulation accuracy as the reward value; Training the designed high-order modulation and demodulation model; The compensated signal is demodulated using the trained high-order modulation and demodulation model to obtain a demodulated signal.

2. The reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communication according to claim 1, characterized in that: The observed state includes the real and imaginary parts of the frequency domain data S1, the real and imaginary parts of the actual mapped symbol S2, and the real and imaginary parts of the error S3 between the frequency domain data S1 and the actual mapped symbol S2; the action space is a set of actions for selecting the corresponding 5G NR high-order demodulation constellation symbol according to the current state.

3. The reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communication according to claim 2, characterized in that: The reward value is designed to be Indicates the correct number of mapping bits corresponding to the randomly selected subcarrier f and the time domain symbol t index, B QAM The number of bits mapped to 64QAM.

4. The reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communication according to claim 3, characterized in that: The high-order modulation and demodulation model includes an intelligent agent, a main Q-value network and a predicted Q-value network based on a deep Q network, and an experience recycling pool; the intelligent agent is a high-order signal demodulator after equalization of the communication channel frequency response of a 5G NR receiver; During the training process, the agent randomly selects a small batch of samples from the offline deployed experience replay pool to obtain the current state and action s t ,a t Input into the main Q value network to get the current state s t 、Action a t The judgment value Q under t (s t ,a t ; θ), θ represents the latest parameters of the main Q value network; Get the next state s t+1 And all its actions are input into the predicted Q value network to generate the next state s t+1 Q-predictions for all actions The network parameter θ of the predicted Q value network clones the main Q value every step length C; The agent is also used to t (s t ,a t ;θ), the reward return value r(s) of the current state-action t ,a t ), the maximum Q prediction value of the next state action Learning rate α, decay factor β of future rewards, calculate the current state s t Execute the current action a t Thus entering the new state s t+1 The maximum cumulative reward Q obtained after * (s t ,a t ); The agent is also used to minimize Q t (s t ,a t ;θ) and Q * (s t ,a t ) is used to update the parameters θ of the main Q-value network.

5. The reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communication according to claim 4, characterized in that: Maximum cumulative reward Q * (s t ,a t ) is calculated as:

6. The reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communication according to claim 5, characterized in that: The parameters θ of the main Q-value network are updated by the following formula: i * ←θ-αΔL(θ) θ * represents the updated parameter θ, and ΔL(θ) represents the gradient of L(θ).

7. The reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communication according to claim 6, characterized in that: The main Q-value network and the predicted Q-value network adopt the same network architecture, including an input layer, a hidden layer, and an output layer, the input layer is a state space, and the output layer is an action space; in the inference stage, the observation state corresponding to the compensation signal is input to the main Q-value network, and the main Q-value network outputs the corresponding action to obtain the selected 5G NR high-order demodulation constellation symbol.

8. The reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communication according to claim 7, characterized in that: The input layer dimension is 6, the hidden layer dimension is 12, and the output layer dimension is 64.

9. The reinforcement learning-assisted high-order signal demodulation method for 5G NR indoor communication according to any one of claims 1 to 8, characterized in that: Each RB of the radio frame signal generated by the 5G NR indoor communication base station has a total of N scs ×N symbol resource elements, including received symbol information and demodulation reference information, N scs is the frequency domain number, N symbol is the time domain number.

10. A reinforcement learning-assisted high-order signal demodulation system for 5G NR indoor communications, featuring: Including signal receiver, least square compensator, and demodulator; The signal receiver is used to receive a radio frame signal generated with reference to the 3GPP protocol and sent by a 5G NR indoor communication base station; The least square compensator is used to compensate the wireless frame signal using a least square mathematical algorithm to obtain a compensated signal; The demodulator is configured to design an observation state based on the error-independent characteristics of the compensation signal, generate an action space based on the modulation information, and then design a high-order modulation and demodulation model based on a deep Q network using demodulation accuracy as a reward value; and train the designed high-order modulation and demodulation model; Furthermore, the compensation signal is demodulated using the trained high-order modulation and demodulation model to obtain a demodulated signal.