A method, device, electronic device and storage medium for optimizing detection of desired propagation signal
By optimizing the expected propagation algorithm through a reinforcement learning model and combining it with an orthogonal time-frequency-space system, the instability and complexity of signal detection in high-speed mobile scenarios are solved, achieving more efficient signal recovery.
Patent Information
- Application Number
- CN202310535833.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-05-12
AI Technical Summary
The signal detection of the expectation propagation algorithm in high-speed mobile scenarios has problems such as instability, large number of iterations and high complexity. In particular, the Doppler frequency shift problem in high-speed mobile scenarios has not been effectively solved.
The reinforcement learning model, especially the dominant actor-commentator algorithm model, is used to optimize the expectation propagation algorithm. By adjusting the damping factor and the number of iterations, the orthogonal time-frequency-space system model is combined for signal detection, and the Bayesian estimation method and moment matching technology are used to restore the original signal.
The stability and reliability of signal detection are improved, the number of iterations is reduced, the complexity is reduced, and the anti-interference performance is improved.
Smart Images

Figure CN116489051B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, device, electronic equipment and storage medium for optimizing detection of a desired propagation signal, and belongs to the technical field of signal detection. Background Art
[0002] In the context of digital communications, digital receivers are required to detect received signals and recover the original signal. Since signals are subject to various unknown interferences during propagation, such as noise and shadowing, recovering the original signal requires signal detection. Signal detection is a method that uses a specific algorithm to restore the received signal to its original state. Digital receivers must be designed to provide the probability of each possible symbol being transmitted based on the received signal. Currently, there are many classic linear and nonlinear detection algorithms, such as maximum likelihood detection, maximum ratio combining, zero-forcing equalization, and minimum mean square error. Compared to these classic algorithms, the expectation propagation (EP) algorithm seeks a Gaussian posterior probability density function that minimizes the KL divergence with respect to the true posterior probability density function, thereby achieving better signal detection performance. The expectation propagation (EP) algorithm is a method for approximating the marginal distribution of the posterior probability of a random variable. Proposed in 2001, it has been widely used in the fields of artificial intelligence and machine learning and has attracted considerable attention in the field of signal detection. The core concept of the EP algorithm is to iteratively approximate the posterior probability distribution with polynomial complexity. Therefore, this algorithm has great potential in signal detection.
[0003] To address the wireless communication issues of fast time-varying channels in high-speed mobile scenarios, a new waveform called Orthogonal Time-Frequency Space (OTFS) has emerged to solve the Doppler shift problem of traditional waveforms in high-speed mobile scenarios. This system modulates and demodulates signals in the delay-Doppler domain, thereby meeting the signaling requirements of high Doppler shift.
[0004] Since the iterative algorithm of expectation propagation (EP) has uncertain damping factors and iteration times, a fixed value can only be determined based on simulation experience, which leads to problems such as unstable detection, large number of iterations, and high complexity. Summary of the Invention
[0005] The object of the present invention is to provide a method, device, electronic device and storage medium for optimizing detection of a desired propagation signal, which can improve the stability and reliability of signal detection.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for optimizing detection of a desired propagation signal, comprising:
[0008] Using the pre-built reinforcement learning model, the pre-acquired expectation propagation algorithm is optimized to obtain an improved expectation propagation algorithm;
[0009] Applying the improved expected propagation algorithm to a pre-built orthogonal time-frequency-space system model to perform signal detection and achieve expected propagation signal detection optimization;
[0010] The reinforcement learning model is an advantaged actor-commentator algorithm model obtained by improving the actor-commentator algorithm model.
[0011] Combined with the first aspect, the pre-acquired expectation propagation algorithm is further optimized by using the dominant actor-commentator algorithm model. The improved expectation propagation algorithm includes:
[0012] The signal-to-noise ratio, running speed and current iteration number of the expectation propagation algorithm are used as the state information of the dominant actor-critic algorithm model, and the baseline function in the dominant actor-critic algorithm model is set to the signal detection error value when the expectation propagation algorithm uses a fixed damping factor;
[0013] Each iteration of the expectation propagation algorithm is regarded as a layer of network, and each layer of network is connected in series to obtain the series network flow graph;
[0014] Add the training parameter α to the same position in each layer of the series network flow graph l , γ l , β l , and the training parameter α l , γ l , β l Initialized to 0, 1, 1;
[0015] Combine the state information and the baseline function to train the damping factor parameter β l , and use the training parameter α l and γ l The convergence speed of the expectation propagation algorithm is accelerated. With the goal of minimizing the loss function, the batch gradient descent algorithm is used to train the expectation propagation algorithm in different signal transmission processes to obtain an improved expectation propagation algorithm.
[0016] In combination with the first aspect, further, the different signal transmission processes include signal transmission processes under different signal-to-noise ratios, different operating speeds and different numbers of iterations.
[0017] In combination with the first aspect, further, the expression of the orthogonal time-frequency space system model is shown in formula (1):
[0018]
[0019] In formula (1), y is the received signal, y∈RMN×1 , H is the channel matrix, H∈R MN×MN , x is the transmitted signal, x∈R MN ×1 , w is noise, w∈R MN×1 , M is the total number of points in the delay dimension, and N is the total number of points in the Doppler domain;
[0020] Perform ISFFT transformation on the transmitted signal x to obtain the frequency domain expression of the transmitted signal x. The frequency domain expression of the transmitted signal x is shown in formula (2):
[0021]
[0022] In formula (2), X[n, m] is the frequency domain expression of the transmitted signal x at time domain grid point n and frequency domain grid point m, x[k, l] is the frequency domain expression of the transmitted signal x at time domain grid point k and frequency domain grid point l, n is the time domain grid point n, m is the frequency domain grid point m, k is the time domain grid point k, and l is the frequency domain grid point l;
[0023] Perform Wiener transform on the received signal y to transform the received signal y into the time-frequency domain received signal Y(f, t). The frequency domain expression of the received signal y is shown in formula (3):
[0024] Y[m,n]=Y(f,t)| t=nT,f=mΔf (3)
[0025] In formula (3), Y[m,n] is the frequency domain expression of the received signal, T is the sampling event, and Δf is the sampling frequency;
[0026] Perform SFFT transformation on the time-frequency domain received signal Y(f, t) to obtain the delayed Doppler domain received signal. The expression of the delayed Doppler domain received signal is shown in formula (4):
[0027]
[0028] In formula (4), y[k, l] is the expression of the received signal in the delay Doppler domain.
[0029] In combination with the first aspect, further, applying the improved desired propagation algorithm to a pre-built orthogonal time-frequency-space system model to perform signal detection, and realizing the desired propagation signal detection optimization includes:
[0030] The Bayesian estimation method is used to obtain the true prior probability density function of the transmitted signal and the transition probability density function of the channel in the orthogonal time-frequency-space system model;
[0031] Perform moment matching on the outer edge distributions of the true prior value and the approximate value of the transmitted signal according to the true prior probability density function of the transmitted signal and the transition probability density function of the channel, to obtain the mean and variance of the Gaussian approximation of the true prior probability density function of the transmitted signal;
[0032] According to the mean and variance of the Gaussian approximation of the true prior probability density function of the transmitted signal, the improved expected propagation algorithm is used to perform signal detection, restore the transmitted signal, and achieve the expected propagation signal detection optimization.
[0033] In a second aspect, the present invention provides a device for detecting and optimizing a desired propagation signal, comprising:
[0034] Algorithm optimization module: used to optimize the pre-acquired expectation propagation algorithm using the pre-built reinforcement learning model to obtain an improved expectation propagation algorithm;
[0035] Signal detection module: used to apply the improved expected propagation algorithm to the pre-built orthogonal time-frequency-space system model to perform signal detection and realize the optimization of expected propagation signal detection;
[0036] The reinforcement learning model is an advantaged actor-commentator algorithm model obtained by improving the actor-commentator algorithm model.
[0037] In a third aspect, the present invention provides an electronic device including a processor and a storage medium;
[0038] The storage medium is used to store instructions;
[0039] The processor is configured to operate according to the instructions to execute the steps of the method according to any one of the first aspects.
[0040] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods described in the first aspect.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] The expected propagation signal detection optimization method provided by the present invention optimizes the traditional expected propagation algorithm through a reinforcement learning model to obtain an improved expected propagation algorithm, and applies the improved expected propagation algorithm to an orthogonal time-frequency space system model, which can realize the expected propagation signal detection optimization; while iterating, the damping factor is continuously changed each time according to the environment, and a new damping factor is introduced to accelerate the convergence of the original algorithm, thereby improving the stability of detection, reducing the number of iterations, and reducing complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1is a flow chart of a method for detecting and optimizing a desired propagation signal provided by an embodiment of the present invention;
[0044] Figure 2 Schematic diagram showing the comparison results of the A2C-EP algorithm, the EP algorithm, and the MMSE algorithm provided in an embodiment of the present invention under 4QAM modulation and a mobile speed of 200 km / h;
[0045] Figure 3 This is a schematic diagram showing the comparison results of the A2C-EP algorithm provided by an embodiment of the present invention, the EP algorithm, and the MMSE algorithm under 16QAM modulation and a moving speed of 200 km / h. DETAILED DESCRIPTION
[0046] The technical solution of this patent is further described in detail below in conjunction with specific implementation methods.
[0047] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and are not to be construed as limiting the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0048] Example 1:
[0049] Figure 1 This is a flow chart of a method for optimizing detection of a desired propagation signal provided by the first embodiment of the present invention. This flow chart only shows the logical sequence of the method of this embodiment. In other possible embodiments of the present invention, different methods may be used without conflict. Figure 1 The steps shown or described are accomplished in the order shown.
[0050] The desired propagation signal detection optimization method provided in this embodiment can be applied to a terminal and can be executed by a desired propagation signal detection optimization device. The device can be implemented in software and / or hardware and can be integrated into a terminal, such as any tablet computer or computer device with communication functions. Figure 1 The method of this embodiment specifically includes the following steps:
[0051] Step 1: Use the pre-built reinforcement learning model to optimize the pre-acquired expectation propagation algorithm to obtain an improved expectation propagation algorithm;
[0052] The reinforcement learning model is an advantaged actor-critic algorithm model obtained by improving the actor-critic algorithm model.
[0053] In the actor-critic algorithm, the policy network acts as the actor, taking actions based on the state (here, the state refers to the signal-to-noise ratio, movement speed, and the number of current algorithm iterations). The value network, on the other hand, acts as the critic, scoring the actor's performance and quantifying the quality of the action taken given the current state. Unlike the actor-critic algorithm, the advantage actor-critic algorithm introduces advantage to measure the difference between the predicted baseline and the actual baseline.
[0054] The Advantage Actor-Critic algorithm has two neural networks, one for the actor network and the other for the critic network. The actor is directly responsible for outputting the probability of each action, with as many outputs as there are actions. The critic outputs the action value Q, which is used by the critic to learn the reward and punishment mechanisms. After learning, the actor is responsible for modifying parameters, and the critic provides guidance and evaluation based on the actor's modifications. This is because the critic can understand the potential reward of the current state by learning the relationship between the environment and rewards. The Advantage Actor-Critic algorithm also adds a baseline function to the action value Q. The baseline function is characterized by reducing its variance without changing the policy gradient. We use the Q value minus this baseline to judge the quality of the current logic.
[0055] Using the dominant actor-commentator algorithm model, the pre-acquired expectation propagation algorithm is optimized. The improved expectation propagation algorithm includes the following steps:
[0056] Step A: The signal-to-noise ratio, running speed, and current iteration number of the expectation propagation algorithm are used as the state information of the dominant actor-critic algorithm model, and the baseline function in the dominant actor-critic algorithm model is set to the signal detection error value when the expectation propagation algorithm uses a fixed damping factor;
[0057] Step B: Treat each iteration of the expectation propagation algorithm as a layer of network, and connect each layer of network in series to obtain the series network flow graph;
[0058] Step C: Add the training parameter α to the same position in each layer of the series network flow graph l , γ l , β l , and the training parameter α l , γ l , βl Initialized to 0, 1, 1;
[0059] Step D: Combine state information and baseline function to train the damping factor parameter β l , and use the training parameter α l and γ l Accelerate the convergence of the expectation propagation algorithm. With the goal of minimizing the loss function, use the batch gradient descent algorithm to train the expectation propagation algorithm in different signal transmission processes to obtain an improved expectation propagation algorithm.
[0060] Among them, different signal transmission processes include signal transmission processes under different signal-to-noise ratios, different operating speeds and different numbers of iterations.
[0061] Since the expectation propagation algorithm (EP algorithm) is an iterative algorithm, each iteration is regarded as a layer in the network, and each layer is connected in series. In this way, the iterative EP algorithm is expanded into a series of network flow graphs between layers. These network flow graphs can be used to implement the original iterative approximate message passing signal detection algorithm. At the same time, a trainable parameter α is added to the same position in each layer. l , γ l , β l , and initialize it to 0, 1, 1. Among them, β l The effect is the same as the original fixed damping factor, α l , γ l In order to make the original EP algorithm converge faster, this EP algorithm combined with the Advantage Actor-Critic algorithm is called the A2C-EP algorithm, that is, the improved expectation propagation algorithm.
[0062] In order to make the EP detection algorithm have better detection performance, it is necessary to continuously train to adjust the training parameters α of each layer l , γ l , β l This training method uses offline training. While there are many deep learning training methods, this network is trained using batch gradient descent. The advantage of batch gradient descent is that it can achieve a global optimum and is easy to implement, thus determining parameter values that maximize network detection performance. After extensive training, the original EP algorithm can achieve the optimal damping factor under different environments (different signal-to-noise ratios, movement speeds, and iteration times) to reduce bit error rates and improve algorithm detection performance.
[0063] The Critic in the Advantage Actor-Critic algorithm is a value function learned by collecting data through interaction with the environment. This function aims to evaluate the actor's performance and guide its next action. The actor, in turn, can use policy gradient learning to determine the optimal policy steps by interacting with the environment and combining the value function provided by the Critic.
[0064] The Actor and Critic are different neural networks with tanh activation function. The network parameters, i.e., the sets of neuron weights, are θ and w, respectively. The Xavier method in deep learning is used here to obtain the initial values. The advantage of this method is that it can avoid the gradient vanishing and gradient exploding problems by keeping the variance of the input and output consistent (obeying the same distribution), allowing the signal to be transmitted deeper in the neural network and remain within a reasonable range (not too small or too large) after passing through multiple layers of neurons.
[0065] At the same time, the state S is defined as the set of signal-to-noise ratio, motion speed and number of iterations of the current algorithm, and the initial feature vector is defined is [0,200km / h,0], and then Input Actor network, output action Action, get new state S new and And feedback the reward value R.
[0066] Will get and Then input it into the Critic network, and the corresponding output action values Q and Q can be obtained respectively. new In the Advantage Actor-Critic algorithm, Q is decomposed into a state-value function V(S) and an advantage function A(S,Action). The state-value function represents the expected future benefits of the agent when starting from state S. The value function can be used to easily evaluate the pros and cons of different strategies. The advantage function expresses the advantage of an action relative to the average in state S. From a quantitative perspective, it is the deviation of a random variable from the mean.
[0067] The purpose of decomposition is to overcome the problem of large variance in the linear function of the Critic network in the Actor-Critic algorithm, which often leads to convergence instability. Next, we calculate the TD error:
[0068] δ=R+ρV(S new )-V(S)
[0069] Among them, ρ is the attenuation factor and R is the reward value, which is based on the error between the simulated signal and the real signal.
[0070] After obtaining the TD error, we need to use the gradient to update the network parameters of the Actor and Critic. The update method is as follows:
[0071]
[0072] Where ∈ and τ are the step sizes of the two network training respectively, and π θ (S, Action) is the policy function of the Advantage Actor-Critic algorithm, which means the probability of choosing action Action under a given state S.
[0073] Step 2: Apply the improved expected propagation algorithm to the pre-built orthogonal time-frequency-space system model to perform signal detection and optimize the expected propagation signal detection;
[0074] The Orthogonal Time-Frequency-Space (OTFS) system model effectively mitigates signal losses and inter-symbol interference (ISI) caused by multipath propagation. At the transmitter, a two-dimensional delay-Doppler domain network of size M×N is constructed, where M is the total number of points in the delay dimension and N is the total number of points in the Doppler domain. The goal of both the transmitter and receiver is to implement QAM modulation in the delay-Doppler domain. All QAM signals placed in the delay-Doppler domain experience the same two-dimensional convolutional channel, mitigating the effects of time- and frequency-selective fading, and are ultimately received at the receiver.
[0075] The expression of the orthogonal time-frequency space system model is shown in formula (1):
[0076] y=Hx+w (1)
[0077] In formula (1), y is the received signal, y∈R MN×1 , H is the channel matrix, H∈R MN×MN , x is the transmitted signal, x∈R MN ×1 , w is noise, w∈R MN×1 , M is the total number of points in the delay dimension, and N is the total number of points in the Doppler domain.
[0078] Perform ISFFT transformation on the transmitted signal x to obtain the frequency domain expression of the transmitted signal x. The frequency domain expression of the transmitted signal x is shown in formula (2).
[0079]
[0080] In formula (2), X[n, m] is the frequency domain expression of the transmitted signal x at time domain grid point n and frequency domain grid point m, x[k, l] is the frequency domain expression of the transmitted signal x at time domain grid point k and frequency domain grid point l, n is the time domain grid point n, m is the frequency domain grid point m, k is the time domain grid point k, and l is the frequency domain grid point l.
[0081] The transmitted signal x is processed using the Heisenberg transform to obtain s(t). After s(t) passes through the transmit filter, it enters the wireless channel. Considering the high-speed mobile scenario, the signals have their own time delay and Doppler frequency deviation, so the received signal must have errors with the original transmitted signal. In the case of not knowing the transmitted signal, the original signal must be restored by the received signal. This is also the role of the following signal detection. After transmission through the wireless channel, the received signal r(t) is obtained. Among them, there is additive white Gaussian noise w(t) in the channel, whose elements are independent and identically distributed zero-mean cyclic symmetric complex Gaussian random variables with variance mean σ 2 The delayed Doppler signal can be obtained by performing an inverse transform of the response at the receiving end. First, the received signal is transformed into the time-frequency domain through the Wiener transform. The received signal y is transformed into the time-frequency domain received signal Y(f, t) by performing the Wiener transform. The frequency domain expression of the received signal y is shown in formula (3):
[0082] Y[m,n]=Y(f,t)| t=nT,f=mΔf (3)
[0083] In formula (3), Y[m,n] is the frequency domain expression of the received signal, T is the sampling event, and Δf is the sampling frequency.
[0084] Perform SFFT transformation on the time-frequency domain received signal Y(f, t) to obtain the delayed Doppler domain received signal. The expression of the delayed Doppler domain received signal is shown in formula (4):
[0085]
[0086] In formula (4), y[k, l] is the expression of the received signal in the delay Doppler domain.
[0087] After receiving the signal y, the improved expectation propagation algorithm is used to restore the original transmitted signal x.
[0088] Applying the improved expectation propagation algorithm to the pre-built orthogonal time-frequency-space system model for signal detection and optimizing the expected propagation signal detection includes the following steps:
[0089] Step a: Using the Bayesian estimation method, obtain the true prior probability density function of the transmitted signal and the transition probability density function of the channel in the orthogonal time-frequency-space system model;
[0090] Step b: Based on the true prior probability density function of the transmitted signal and the transition probability density function of the channel, moment matching is performed on the outer edge distribution of the true prior value and the approximate value of the transmitted signal to obtain the mean and variance of the Gaussian approximation of the true prior probability density function of the transmitted signal;
[0091] Step c: Based on the mean and variance of the Gaussian approximation of the true prior probability density function of the transmitted signal, an improved expected propagation algorithm is used to perform signal detection, restore the transmitted signal, and achieve the expected propagation signal detection optimization.
[0092] The signal-to-noise ratio, speed, and current iteration number of the EP algorithm are regarded as state information S(SNR, Speed, l). Each iteration calculation is regarded as a process of interaction between the agent and the environment. The agent's action is to select the parameters of the current iteration of the EP algorithm, and the reward is set to Among them, k represents the kth symbol number, represents the bit error rate of the EP algorithm when using a fixed damping factor β and at the lth iteration, Represents the bit error rate using a non-fixed damping factor. The output data after the current iteration will be used as the input for the next state. When the maximum number of iterations is reached, an episode ends. An episode can be understood as a round or cycle of learning in reinforcement learning. In this algorithm, the number of episodes is set to 500.
[0093] Assume that the OTFS transmit signal x is a column vector x=(x1,x2,…,x M×N ) T The value of the element is the data bit stream passing through the constellation diagram The mapping is generated, where Mod is the modulation order. Let is the channel matrix of OTFS.
[0094] The purpose of the A2C-EP algorithm is the same as that of the traditional EP signal detection algorithm. The task of the A2C-EP algorithm is to estimate the original transmitted signal x based on the received signal y. The recovered estimated value is used express, In the algorithm, it can be represented by the mean of the Gaussian approximation of the posterior probability density distribution. Like the traditional EP algorithm, the improved algorithm still uses the Bayesian estimation method, which is a common parameter estimation method in signal detection.
[0095] In signal detection, the posterior mean E[x|y] of the transmitted signal x is used as the estimated value, and the kth element x of the signal k The estimated value of is:
[0096]
[0097] Among them, p(x k |y) is the marginal distribution of the posterior probability density distribution p(x|y), which is given by the Ye-Bas formula:
[0098]
[0099] The transfer probability density function of the channel can be obtained as:
[0100] p(y|x)∝CN(y;Hx,σ 2 I)
[0101] Among them, CN represents the mean Hx and the covariance matrix is σ 2 The probability density function of the cyclically symmetric complex Gaussian vector y of I;
[0102] p(x) is the true prior probability density function of the transmitted signal vector x:
[0103]
[0104] When calculating p(x|y) according to the Ye-Bas formula, the complexity will reach an exponential level. Therefore, the EP algorithm uses the Gaussian approximation q(x) of the posterior probability density distribution p(x|y) to participate in the detection of signal x. The expression of Gaussian approximation q(x) is as follows:
[0105]
[0106] is a true prior p(x k ), assuming and are their mean and covariance respectively, forming vectors and
[0107] to get To determine the mean and variance of a data set, we need to use moment matching. Moment matching is a commonly used data distribution matching method that maps a random variable with a known distribution to a random variable with another distribution. Moment matching can transform a complex distribution into a simple one, such as a Gaussian distribution, thereby simplifying data processing and modeling.
[0108] According to the Gaussian approximation calculation formula of the posterior probability distribution, its value in the l-th iteration process can be obtained:
[0109]
[0110] Assume that the mean and covariance matrix of its Gaussian distribution are μ [l] and S [l], we can get the distribution:
[0111] q [l] (x)∝CN(x;μ [l] ,S [l] )
[0112] The values of the mean and covariance matrix are:
[0113]
[0114] In the algorithm, the mean μ of the Gaussian approximation of the posterior probability density distribution is used [l] Represents the estimated value of the kth element of the signal Right now
[0115] As the number of iterations l = 1, 2, ..., 10 increases, the outer edge distribution can be introduced The method makes and Do moment matching to continuously update The mean at the lth iteration and variance The outer marginal distribution refers to the distribution obtained by marginalizing one or more variables in a joint distribution. Its function is to simplify the problem, transforming a multidimensional problem into a single-dimensional problem, making it easier to analyze and apply. The outer marginal distribution can be used to calculate the probability distribution of a variable or for model selection and comparison. The calculation formula is as follows:
[0116]
[0117] According to the formula, the mean and variance of the outer edge distribution can be obtained as follows:
[0118]
[0119] The A2C-EP algorithm includes the following steps:
[0120] Step 1: Initialize the parameters of the A2C-EP algorithm;
[0121] When the number of iterations l is 0, all means (marginal distribution means) The mean of the Gaussian approximation to the true prior Set to all zero vectors, and all variances (marginal distribution variances Variance of the Gaussian approximation of the true prior ) is set to a vector of all 1s.
[0122] Calculate the mean μ of the Gaussian approximation of the posterior probability distribution when the number of iterations l is 0 [0] and the covariance matrix S [0], get the estimated value of the kth element of signal x
[0123] Given the channel matrix H and the received signal y from the receiving end, the following iterative process begins.
[0124] Step 2: Set the variable H, y, μ [l] and S [l] In the first iteration of the algorithm, the meaning of each variable is shown in Table 1.
[0125] Table 1 Meaning of algorithm input variables
[0126]
[0127] Step 1: Utilization Participation in Moment Matching:
[0128]
[0129] In the above formula is the non-normalized Gaussian distribution corresponding to the posterior probability density distribution p(x|y)), and its mean is calculated and covariance matrix
[0130] Step ii: Set non-normalized Gaussian distribution:
[0131]
[0132] Make the mean equal to Variance is equal to Thus we get:
[0133]
[0134] It can be seen that the above two formulas are for l+1 iterations The mean and variance
[0135] Step iii: Bring it into the current Actor network, output the action Action, and obtain the damping factor β trained during the lth iteration l , damp the result of step ii and update the mean:
[0136]
[0137] Step iv: Update the outer edge distribution at time l+1 The mean and variance And get the training damping factor α from Action l and γ l Perform damped updates to the mean and variance:
[0138]
[0139] Step V: Get the estimated signal value at time l+1 make And get the current reward value:
[0140]
[0141] In the formula Represents the signal estimate obtained using a fixed damping factor.
[0142] Step ⅵ: Use the reward value R obtained in step ⅴ l , calculate the TD error, update the Actor and Critic networks, and continue with step ②, entering the next iteration of the A2C-EP algorithm at time l+1, until the iteration is terminated when l is 10.
[0143] Whether it is data-driven deep learning or reinforcement learning, it is necessary to combine data for a lot of training. In the network training of this algorithm, the back propagation technology of gradient and gradient descent method in deep learning are used to train the damping factor α in the lth iteration. l , γ l , β l , in the initial case [α l , γ l , β l ] is [0, 1, 1].
[0144] When training a neural network, the weights are randomly initialized. Randomization refers to randomizing the signal-to-noise ratio, speed, and number of iterations. Obviously, initializing these parameters will not generally yield good results. However, during training, we hope to achieve a very low loss function at the end of training. This also makes it possible to improve the network because we can modify the function by adjusting the weights.
[0145] Through the above steps, the iterative algorithm of A2C-EP has been expanded into a structure similar to a multi-layer fully connected network structure and connected to the neural network. Next, the neural network needs to be trained to achieve the optimization goal. Specifically, in the first round of training, the loss function is set to To minimize it, the parameter α e , γ e , β lAdjustments are made by using a batch gradient descent algorithm. The input of the loss function is the network prediction value and the true target value, and then a distance value is calculated to measure the effectiveness of the network in this current situation. After completing the first round of training, the training parameters of the first round are used as the initial values of the training parameters of the next round, so the loss function of the next round becomes Similarly, continue to use this training method to make adjustments.
[0146] Through extensive training, all network parameters can be fully trained. The number of training episodes for this network is set to 500. Using batch gradient descent, the problem of vanishing training gradients can be overcome, thus achieving efficient training. This network integrates well with the original EP detection algorithm, inheriting its superior detection performance while also producing an A2C-EP detection algorithm with even better detection performance through training of the damping factor.
[0147] The simulation results are as follows Figure 2 、 Figure 3 The simulation system has a total number of carriers M of 64, a number of multi-carrier symbols N of 16, a sampling frequency of 1×106 Hz, a mobile speed of 200 km / h, a carrier frequency of 3.5×109 Hz, a Rayleigh channel model, Doppler frequency deviations of [0, -5, -7, -9, -11, -14], delays of [1, 3, 6, 8, 9, 11], and modulation orders of 4QAM and 16QAM.
[0148] The expected propagation signal detection optimization method provided in this embodiment optimizes the traditional expected propagation algorithm through a reinforcement learning model to obtain an improved expected propagation algorithm, and applies the improved expected propagation algorithm to the orthogonal time-frequency space system model, which can realize the expected propagation signal detection optimization; while iterating, the damping factor is continuously changed each time according to the environment, and a new damping factor is introduced to accelerate the convergence of the original algorithm, thereby improving the detection stability, reducing the number of iterations, and reducing the complexity. The reinforcement learning used by the algorithm is a network that focuses on both the iterative intermediate layer and the iterative final layer, and by designing intermediate incentives, it avoids sparse rewards while achieving the purpose of improving the convergence performance of the iterative intermediate layer. What is superior to the traditional EP signal detection algorithm is that not only the complexity is not increased, but the anti-interference performance is also optimized to a certain extent.
[0149] Example 2:
[0150] This embodiment provides a desired propagation signal detection and optimization device, including:
[0151] Algorithm optimization module: used to optimize the pre-acquired expectation propagation algorithm using the pre-built reinforcement learning model to obtain an improved expectation propagation algorithm;
[0152] Signal detection module: used to apply the improved expected propagation algorithm to the pre-built orthogonal time-frequency-space system model to perform signal detection and optimize the expected propagation signal detection;
[0153] Among them, the reinforcement learning model is the advantage actor-commentator algorithm model obtained by improving the actor-commentator algorithm model.
[0154] The desired propagation signal detection and optimization device provided in the embodiment of the present invention can execute the desired propagation signal detection and optimization method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0155] Example 3:
[0156] This embodiment provides an electronic device, including a processor and a storage medium;
[0157] The storage medium is used to store instructions;
[0158] The processor is configured to operate according to the instructions to execute the steps of the method in the first embodiment.
[0159] Example 4:
[0160] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the method in the first embodiment are implemented.
[0161] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0162] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0163] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0164] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0165] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for optimizing detection of a desired propagation signal, characterized in that: include: Using the pre-built reinforcement learning model, the pre-acquired expectation propagation algorithm is optimized to obtain an improved expectation propagation algorithm; Applying the improved expected propagation algorithm to a pre-built orthogonal time-frequency-space system model to perform signal detection and achieve expected propagation signal detection optimization; Wherein, the reinforcement learning model is an advantage actor-commentator algorithm model obtained by improving the actor-commentator algorithm model; The pre-acquired expectation propagation algorithm is optimized using the dominant actor-commentator algorithm model. The improved expectation propagation algorithm includes: The signal-to-noise ratio, running speed and current iteration number of the expectation propagation algorithm are used as the state information of the dominant actor-critic algorithm model, and the baseline function in the dominant actor-critic algorithm model is set to the signal detection error value when the expectation propagation algorithm uses a fixed damping factor; Each iteration of the expectation propagation algorithm is regarded as a layer of network, and each layer of network is connected in series to obtain the series network flow graph; Add training parameters to the same position in each layer of the series network flow graph 、 、 , and the training parameters 、 、 Initialized to 0, 1, 1; Combine the state information and the baseline function to train the damping factor parameter , and using the training parameters and The convergence speed of the expectation propagation algorithm is accelerated. With the goal of minimizing the loss function, the batch gradient descent algorithm is used to train the expectation propagation algorithm in different signal transmission processes to obtain an improved expectation propagation algorithm.
2. The method for optimizing detection of desired propagation signals according to claim 1, wherein: The different signal transmission processes include signal transmission processes under different signal-to-noise ratios, different operating speeds and different numbers of iterations.
3. The method for detecting and optimizing the desired propagation signal according to claim 1, wherein: The expression of the orthogonal time-frequency space system model is shown in formula (1): (1); In formula (1), To receive the signal, , is the channel matrix, , To send a signal, , is noise, , is the total number of points in the delay dimension, is the total number of points in the Doppler domain; Sending signal Perform ISFFT transformation to obtain the transmitted signal The frequency domain expression of the transmitted signal The frequency domain expression of is shown in formula (2): (2); In formula (2), To send a signal In the time domain grid and frequency domain grid The frequency domain expression at To send a signal In the time domain grid and frequency domain grid The frequency domain expression at Time domain grid , is the frequency domain grid , Time domain grid , is the frequency domain grid ; To receive the signal Perform Wiener transform to receive the signal Transform the received signal into the time-frequency domain , receiving signal The frequency domain expression of is shown in formula (3): (3); In formula (3), is the frequency domain expression of the received signal, is the sampling event, is the sampling frequency; Received signal in time-frequency domain Perform SFFT transformation to obtain the delayed Doppler domain received signal. The expression of the delayed Doppler domain received signal is shown in formula (4): (4); In formula (4), is the expression of the received signal in the delay-Doppler domain.
4. The method for detecting and optimizing the desired propagation signal according to claim 1, wherein: Applying the improved expected propagation algorithm to a pre-built orthogonal time-frequency-space system model to perform signal detection and achieve expected propagation signal detection optimization includes: The Bayesian estimation method is used to obtain the true prior probability density function of the transmitted signal and the transition probability density function of the channel in the orthogonal time-frequency-space system model; Perform moment matching on the outer edge distributions of the true prior value and the approximate value of the transmitted signal according to the true prior probability density function of the transmitted signal and the transition probability density function of the channel, to obtain the mean and variance of the Gaussian approximation of the true prior probability density function of the transmitted signal; According to the mean and variance of the Gaussian approximation of the true prior probability density function of the transmitted signal, the improved expected propagation algorithm is used to perform signal detection, restore the transmitted signal, and achieve the expected propagation signal detection optimization.
5. A device for detecting and optimizing a desired propagation signal, characterized in that: include: Algorithm optimization module: used to optimize the pre-acquired expectation propagation algorithm using the pre-built reinforcement learning model to obtain an improved expectation propagation algorithm; Signal detection module: used to apply the improved expected propagation algorithm to the pre-built orthogonal time-frequency-space system model to perform signal detection and realize the optimization of expected propagation signal detection; Wherein, the reinforcement learning model is an advantage actor-commentator algorithm model obtained by improving the actor-commentator algorithm model; The pre-acquired expectation propagation algorithm is optimized using the dominant actor-commentator algorithm model. The improved expectation propagation algorithm includes: The signal-to-noise ratio, running speed and current iteration number of the expectation propagation algorithm are used as the state information of the dominant actor-critic algorithm model, and the baseline function in the dominant actor-critic algorithm model is set to the signal detection error value when the expectation propagation algorithm uses a fixed damping factor; Each iteration of the expectation propagation algorithm is regarded as a layer of network, and each layer of network is connected in series to obtain the series network flow graph; Add training parameters to the same position in each layer of the series network flow graph 、 、 , and the training parameters 、 、 Initialized to 0, 1, 1; Combine the state information and the baseline function to train the damping factor parameter , and using the training parameters and The convergence speed of the expectation propagation algorithm is accelerated. With the goal of minimizing the loss function, the batch gradient descent algorithm is used to train the expectation propagation algorithm in different signal transmission processes to obtain an improved expectation propagation algorithm.
6. An electronic device, characterized in that: including processors and storage media; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Dynamic detection method and device for signals in MIMO (Multiple Input Multiple Output) system
CN114915321A
Low-complexity OTFS signal detection method for high-speed moving scene
CN115412416A