Power amplifier linearization thermal compensation method based on reinforcement learning

By applying a neural network model based on reinforcement learning in the amplifier linearization technology, the temperature variable is introduced to adjust the predistorted signal, and the linearization performance degradation caused by environmental temperature changes is solved, and efficient and low-cost temperature adaptive amplifier linearization is achieved.

CN120067610APending Publication Date: 2025-05-30XIDIAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510036366.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the ambient temperature changes, the mismatch between the predistortion model and the amplifier model leads to a degradation of linearization performance, and the existing solutions have problems such as high computational cost, complex structure and difficult hardware deployment.

Method used

Using a neural network model based on reinforcement learning, the state space and action space are constructed, temperature variables are introduced, and the predistorted signals output by the digital predistortion model are adjusted to achieve temperature adaptive amplifier linearization.

Benefits of technology

Under different ambient temperatures, the reinforcement learning network model can effectively adjust the predistorted signal, ensure the stability of the amplifier linearization performance, reduce computing costs, and simplify hardware deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067610A_ABST
    Figure CN120067610A_ABST
Patent Text Reader

Abstract

The invention discloses a power amplifier linearization thermal compensation method based on reinforcement learning, and the method comprises the steps: constructing a digital pre-distortion model, and obtaining a data set with a temperature label; building a neural network model based on reinforcement learning; constructing a state space and an action space of the neural network model, introducing the temperature into the state space and the action space, and adjusting a pre-distortion signal output by the digital pre-distortion model; defining parameters related to training of a neural network model, inputting the data set into the neural network model for training, and obtaining a loss function convergence reinforcement learning network model; and the reinforcement learning network model is cascaded with the digital pre-distortion model and the power amplifier, so that a temperature self-adaptive power amplifier linearization function is realized. By sensing the environment temperature and adjusting the pre-distortion signal output by the digital pre-distortion model, the pre-distortion signal can be matched with the behavior characteristics of the power amplifier more accurately, and the linearization performance of the pre-distorter is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and relates to a thermal compensation method for power amplifier linearization based on reinforcement learning. Background Art

[0002] With the increasing development of mobile communication networks and intelligent terminals, the issues of data throughput and spectral efficiency have imposed higher performance requirements on wireless communication systems. As one of the key components in wireless communication systems, ensuring the linearity of power amplifiers at a certain efficiency level has always been a hot topic in academic research and applications. For example, as the signal bandwidth continues to increase and the modulation method becomes more complex, more serious in-band and out-of-band distortions will occur after the signal is transmitted through the power amplifier. Therefore, effective power amplifier linearization technologies are required. Digital predistortion has become the current mainstream power amplifier linearization technology due to its advantages such as moderate complexity and bandwidth, and strong linearization correction ability. Digital predistortion is mainly divided into lookup table-based, polynomial-based, and neural network-based digital predistortion. They all achieve the modeling of the predistortion model by testing the baseband data corresponding to the input and output signals of the power amplifier at a certain temperature. However, when the environmental temperature changes and causes the characteristics of the power amplifier to change, the mismatch between the predistortion model and the power amplifier model will cause the deterioration of the linearization performance.

[0003] There are mainly two existing solutions: One is at the circuit level. A temperature compensation circuit is designed for the power amplifier to ensure the stability of the power amplifier characteristics, thereby avoiding the mismatch between the predistortion model and the power amplifier model when the temperature changes. For example, the gate voltage of the power amplifier is linearly adjusted by simply connecting a diode in series or an external chip is added. The other is at the algorithm level. An adaptive predistortion architecture can be selected in the architecture. When the environmental temperature changes, the parameters of the predistortion model are updated in a timely manner to adapt to the power amplifier model at different temperatures; or a digital predistortion model based on deep learning is implemented. Temperature tags corresponding to data at different temperatures are added to construct a temperature-dependent training, verification, and test set to expand the predistortion model and make it have thermal compensation characteristics.

[0004] Existing temperature compensation circuits often cannot meet the requirements of high compensation accuracy, easy debugging, and simple structure, so there are certain limitations in applications; the real-time update characteristic of the adaptive predistortion architecture requires that after a certain number of frames of new data pass through the power amplifier, the feature matrix and model coefficients of the predistortion model need to be calculated according to the current data, or a new predistortion network model needs to be retrained, which will greatly increase the computational cost and at the same time impose very high requirements on the feedback correction loop; for the predistortion network model that can sense temperature, first, a large amount of training data at different temperatures is required, and in order to ensure the generalization and accuracy of the model, the number of model parameters is large, which will increase the complexity and response time of the network and is not conducive to the deployment and application of the network model on the hardware platform.

[0005] To ensure the performance of wireless communication systems, linearization techniques for power amplifiers (PAs) are currently important and practical methods. Among them, the mainstream digital predistortion (DPD) linearization technique models the inverse model of the power amplifier using the baseband input signal of the power amplifier and the power amplifier output signal down-converted to the baseband, so as to achieve the effect of predistortion. Therefore, one of the keys to ensuring the linearization performance of digital predistortion lies in the modeling accuracy of its power amplifier behavior model. However, for both silicon-based and GaN-based power amplifiers, changes in the operating temperature will affect the DC parameters of the power amplifier transistors, thereby changing the behavior characteristics of the power amplifier. In practical applications, the data required to model the predistorter is often measured under specific temperature conditions. When the ambient temperature changes and causes changes in the power amplifier characteristics, the predistorter and the power amplifier cannot be accurately matched, resulting in a significant decrease in its linearization performance. Summary of the Invention

[0006] The present invention aims to solve the technical problem of how the predistorter can have better linearization performance at temperatures different from a specific temperature. The present invention provides a power amplifier linearization thermal compensation method based on reinforcement learning, and the technical solution adopted is:

[0007] A power amplifier linearization thermal compensation method based on reinforcement learning, comprising the steps of:

[0008] S1. Construct a digital predistortion model and obtain a data set with temperature tags;

[0009] S2. Build a neural network model based on reinforcement learning;

[0010] S3. Construct the state space and action space of the neural network model, introduce temperature into the state space and the action space, and adjust the predistortion signal output by the digital predistortion model;

[0011] S4. Define the parameters related to the training of the neural network model, input the data set into the neural network model for training, and obtain a reinforcement learning network model with a converged loss function;

[0012] S5. Cascade the reinforcement learning network model with the digital predistortion model and the power amplifier to realize the temperature-adaptive power amplifier linearization function.

[0013] In one embodiment of the present invention, the step S1 includes:

[0014] S11: Test and save the input and output data of the power amplifier at different ambient temperatures, and perform temporal alignment through the cross-correlation method to obtain the aligned data;

[0015] S12: For different types of the digital predistortion models, perform corresponding processing on the aligned data, and calculate the correlation coefficient of the digital predistortion model to obtain digital predistortion models at different temperatures;

[0016] S13: Add the same excitation signal into the digital predistortion models at different temperatures to obtain accurate predistortion signals at different temperatures;

[0017] S14: Package the predistortion signals into a data set with temperature labels.

[0018] In one embodiment of the present invention, the step S12 includes:

[0019] S121: The digital predistortion model is based on a polynomial. According to the power amplifier output data obtained by testing, obtain the corresponding feature matrix, and obtain the final coefficient vector by inverting the feature matrix;

[0020] S122: The digital predistortion model is based on a neural network. Process the test data into a form suitable for network training, and complete the training and verification of the network model.

[0021] In one embodiment of the present invention, the step S14 includes:

[0022] S141: Use the predistortion signal at room temperature as the original data;

[0023] S142: Select the predistortion signals of several groups of typical temperature data as label data within the entire temperature range, and add the corresponding ambient temperature to the label data;

[0024] S143: Merge the label data, and process the temperature corresponding to the label data into a temperature vector;

[0025] S144: Process the formats of the original data and the label data, and package the original data, label data, and temperature vector after format processing into a data set with temperature labels.

[0026] In one embodiment of the present invention, in the step S2, the neural network model based on reinforcement learning includes:

[0027] A feature extraction module, configured to pre-extract features of input data;

[0028] An Actor network module, configured to generate a reinforcement learning policy π;

[0029] A Critic network module, configured to generate a reinforcement learning value function v;

[0030] The feature extraction module is respectively connected to the Actor network module and the Critic network module.

[0031] In an embodiment of the present invention, the feature extraction module includes two convolutional layers. The number of input channels of the first convolutional layer of the feature extraction module is 1, which matches the dimension of the picture type data in the dataset, and the number of output channels is 16. The number of input channels and the number of output channels of the second convolutional layer of the feature extraction module are both 16;

[0032] The Actor network module includes two convolutional layers and a Softmax activation layer. The number of input channels and the number of output channels of the first convolutional layer of the Actor network module are both 16. The number of input channels of the second convolutional layer of the Actor network module is 16, and the number of output channels is 3. The output of the second convolutional layer of the Actor network module is passed through the Softmax activation layer to obtain the probability distribution of 3 discrete actions;

[0033] The Critic network module includes two convolutional layers. The number of input channels and the number of output channels of the first convolutional layer of the Critic network module are both 16. The number of input channels of the second convolutional layer of the Critic network module is 16, and the number of output channels is 1.

[0034] In an embodiment of the present invention, in step S3, the state space for constructing the neural network model includes: using the current input data and the corresponding temperature label as state data. The addition of the temperature label enables the agent in reinforcement learning to perceive the change in temperature, and the temperature is introduced into the policy formulation of the Actor network and the value evaluation of the Critic network.

[0035] In an embodiment of the present invention, in step S3, the action space for constructing the neural network model includes:

[0036] Two parts, discrete actions and continuous actions. Two filtering methods and the action "hold" are selected as the discrete action space of the agent in reinforcement learning. The two filtering methods are bilateral filtering and Gaussian filtering respectively;

[0037] The initial value of the sigmaColor parameter of the bilateral filtering function is set to 0.1, the initial value of sigmaSpace is set to 5, and the window size d is set to 3;

[0038] The initial values of the sigmaX and sigmaY parameters of the Gaussian filtering function are both set to 0.5, and the Gaussian kernel size is [3, 3];

[0039] Set the temperature correlation coefficient function, and the temperature correlation coefficient function is expressed as:

[0040]

[0041] Among them, T represents the temperature label of the current input data, and T 0 represents the reference temperature, that is, the ambient temperature corresponding to the input data of reinforcement learning, and f T represents the temperature correlation coefficient;

[0042] Introduce the temperature correlation coefficient function into the bilateral filtering function, sigmaColor = 0.1×f T , sigmaSpace = 5×f T , and keep the window size unchanged;

[0043] Introduce the temperature correlation coefficient function into the Gaussian filtering function, sigmaX = sigmaY = 0.5×f T , and keep the Gaussian kernel size unchanged;

[0044] The continuous action is set as:

[0045] I adjusted = I filtered ×(1 + a)+b

[0046] Among them, I adjusted represents the image data after adjusting the continuous action, I filtered represents the image data after the filtering operation, a represents the scaling factor, and b represents the bias adjustment value, where a is regarded as a linear adjustment and b is regarded as a non-linear adjustment;

[0047] When the temperature is low, the power gain of the power amplifier is large, and the corresponding pre-distortion signal has a small amplitude at low temperatures. Express a and b as: a = 0.1 + 0.01×(T - T 0 ), b = 0.05×(T - T 0 ).

[0048] In an embodiment of the present invention, in step S4, the parameters defining the neural network model include:

[0049] Adjust the initial pre-distortion signal at the reference temperature to the corresponding pre-distortion signal after the ambient temperature changes, and use the improvement degree of the deviation between the input and output data as the measurement standard of the current state action reward;

[0050] The agent in the neural network model operates on 5×5 pixel elements in the input data, and the reward is expressed as:

[0051] R t =(I target - s t ) 2 -(I target-s t+1 ) 2

[0052] Among them, I target represents the corresponding ideal pre-distortion output after the temperature change, s t represents the pre-distortion output in the current state, s t+1 represents the pre-distortion output in the next state. As the agent continuously explores, when the pre-distortion output in the next state is closer to the ideal value, R t is larger. Therefore, the agent is encouraged to continue exploring in this direction and update the weight parameters of the Actor network and the Critic network accordingly. Maximizing the total reward is equivalent to minimizing the mean squared error between the final network output and the ideal pre-distortion output;

[0053] The update formulas for the policy function of the Actor network and the value function of the Critic network based on the temporal difference method are expressed as:

[0054]

[0055] Among them, δ i = r i+1 + γV(s i+1 ) - V(s i ) represents the temporal difference error of the agent, r i+1 and V(s t ) represent the reward and state value in the corresponding state, γ represents the discount factor, which is set to 0.95. That is, the reinforcement learning network not only considers the current reward but also incorporates the value of future rewards into the decision-making. θ a represents the parameters of the Actor network, represents the probability distribution of the a i action executed in the state s i . β represents the entropy regularization coefficient, which controls the influence of the entropy term on the loss function. A larger value makes the agent more inclined to explore and is set to 0.01. Among them, can be expressed as represents the entropy of the action distribution, which is used to evaluate the exploration ability of the current policy. θ v represents the parameters of the Critic network.

[0056] In an embodiment of the present invention, the policy loss function and the value loss function of the neural network model are expressed as:

[0057]

[0058] Among them, L π represents the policy loss function, and L v represents the value loss function;

[0059] Adjust the contributions of the policy loss and the value loss to the total loss through a weighting coefficient adjustment strategy, L total = ω π ·L π + ω v ·L v , where ω π and ω v respectively represent the weight coefficients of the policy loss and the value loss, which are set to 1.0 and 0.5, and then use L total to participate in backpropagation and gradient update to optimize the parameters of the model.

[0060] Advantages of the present invention:

[0061] The power amplifier linearization thermal compensation method based on reinforcement learning of the present invention utilizes the powerful adjustment ability of reinforcement learning and the characteristics of interacting and learning with the environment. By sensing the environmental temperature and adjusting the predistortion signal output by the digital predistortion model, it can accurately match the behavioral characteristics of the power amplifier, and at the cost of low computational cost, ensure that the predistorter still has good linearization performance at temperatures different from a specific temperature. Description of the drawings

[0062] Figure 1 is a flowchart of the power amplifier linearization thermal compensation method based on reinforcement learning provided by an embodiment of the present invention;

[0063] Figure 2 is a flowchart of obtaining a dataset with temperature labels provided by an embodiment of the present invention;

[0064] Figure 3 is a schematic diagram of the construction of a dataset with temperature labels provided by an embodiment of the present invention;

[0065] Figure 4 is a schematic diagram of converting one-dimensional complex signal data into two-dimensional image type data provided by an embodiment of the present invention;

[0066] Figure 5 is a schematic diagram of the structure of a neural network model based on reinforcement learning provided by an embodiment of the present invention;

[0067] Figure 6 is a schematic diagram of the cascade of a reinforcement learning network model, a digital predistortion model and a power amplifier provided by an embodiment of the present invention. Detailed implementation manners

[0068] The present invention will be described in detail below with reference to the drawings and specific implementation manners.

[0069] The present invention provides a power amplifier linearization thermal compensation method based on reinforcement learning. Referring to the attached Figure 1 , it includes the following steps:

[0070] S1. Build a digital predistortion model and obtain a dataset with temperature labels;

[0071] S2. Build a neural network model based on reinforcement learning;

[0072] S3. Construct the state space and action space of the neural network model, introduce temperature into the state space and action space, and adjust the predistortion signal output by the digital predistortion model;

[0073] S4. Define the parameters related to the training of the neural network model, input the dataset into the neural network model for training, and obtain a reinforcement learning network model with convergent loss function;

[0074] S5. Cascade the reinforcement learning network model with the digital predistortion model and the power amplifier to achieve the temperature-adaptive power amplifier linearization function.

[0075] The power amplifier linearization thermal compensation method based on reinforcement learning of the present invention can adjust the output of the digital predistortion model according to the ambient temperature, so as to ensure the linearization performance of the digital predistortion model under different ambient temperatures.

[0076] The present invention first clarifies how to construct a dataset with temperature labels. This construction method has strong versatility and is applicable not only to reinforcement learning but also to supervised learning in deep learning. Since the reinforcement learning adjustment network actually processes the initial predistortion signal output by the digital predistortion network, it is necessary to first complete the construction of the digital predistortion model before constructing the dataset of the reinforcement learning network. Refer to the appendix Figure 2 , the acquisition process of the dataset with temperature labels of the present invention includes:

[0077] S11: Through the construction of a simulation or actual test platform, test and save the input and output data of the power amplifier at different ambient temperatures. Since the power amplifier itself has a certain delay, the cross-correlation method is used for time series alignment to obtain the aligned data (the input and output data mentioned below in the specification are all aligned data).

[0078] S12: Use the input and output data of the power amplifier tested at different operating temperatures, perform corresponding processing on the aligned data for different types of digital predistortion models, and calculate the correlation coefficient of the digital predistortion model to obtain digital predistortion models at different temperatures.

[0079] For example, if the digital pre-distortion model is based on a polynomial, the corresponding feature matrix is obtained according to the power amplifier output data obtained from the test, and the final coefficient vector is obtained by inverting the feature matrix; if the digital pre-distortion model is based on a neural network, the test data is processed into a form suitable for network training, and the training and verification of the network model are completed.

[0080] S13: Add the same excitation signal to the digital pre-distortion models at different temperatures to obtain accurate pre-distortion signals at different temperatures.

[0081] S14: Package the pre-distortion signals into a data set with temperature tags. Specifically, the pre-distortion signals at room temperature are used as the original data; the pre-distortion signals of several groups of typical temperature data are selected as the tag data within the entire temperature range, and the corresponding ambient temperature is added to the tag data; the tag data is merged, and the temperature corresponding to the tag data is processed into a temperature vector; the formats of the original data and the tag data are processed, and the original data, tag data, and temperature vector after format processing are packaged into a data set with temperature tags.

[0082] Appendix Figure 3 This is a schematic diagram of the construction of the data set with temperature tags according to the present invention. The present invention takes into account the influence of temperature changes on the linearization technology and combines the reinforcement learning algorithm to achieve the effect of intelligently adjusting the output of the digital pre-distortion model according to the ambient temperature, which can effectively ensure the performance of the linearization technology under a large range of ambient temperature changes, and solves the problems of excessive computational cost of the adaptive pre-distortion architecture and excessive demand for training data of the traditional pre-distortion neural network model.

[0083] In an embodiment of the present invention, if the operating temperature range of the power amplifier is -40 to 85 °C, initially, the power amplifier operates at 27 °C (room temperature), corresponding to the digital pre-distortion model at room temperature. Therefore, the pre-distortion signals at room temperature are used as the original data in the data set. Then, several groups of typical temperature data are selected within the entire temperature range. For example, the pre-distortion signals at -40 °C, 0 °C, and 75 °C are used as the tag data in the data set, and the corresponding ambient temperature is added to the tag data. Specifically, in the current example, the original data at 27 °C (room temperature) is expanded three times to align with the length of the tag data. Then, the tag data at -40 °C, 0 °C, and 75 °C are merged, and the temperature corresponding to the tag data is processed into a temperature vector, indicating at what temperature this group of data is measured.

[0084] Next, it is also necessary to process the formats of the original data and the label data to make them conducive to the training of the reinforcement learning network. Since the present invention is inspired by the application of reinforcement learning in the field of image processing, it is based on the pixel point method and uses the A3C (Asynchronous Advantage Actor-Critic) algorithm to adjust the image type data of the input and output. Since the pre-distortion signal obtained by testing is a complex signal or two-channel real number signals of I / Q, it is necessary to adjust it into the data form of the picture type, and convert the one-dimensional complex signal data into the two-dimensional image type data. The image size can be adjusted according to the memory depth M and the non-linear order k, as shown in the appendix Figure 4 shown. Where M represents the memory depth. In each group of original data and label data, there are I / Q components, envelope correlation terms and the time delay terms of both. In addition, different envelope correlation terms can be considered by adjusting the non-linear order k. In the present invention, M is set to 5 and the non-linear order is set to 4. Therefore, the dimension of the processed data is (N, 1, 5, 5), where N represents the number of samples. On the one hand, this maps the pre-distortion signal from one-dimensional data to two-dimensional data, which is convenient for inputting into the reinforcement learning network for training. On the other hand, more relevant features of the current data are added, which is beneficial to improving the accuracy of network adjustment. Finally, the original data, label data and temperature vector with adjusted formats are encapsulated into the final data set.

[0085] The preparation process of the data set reflects the advantages of the present invention. Different from the pre-distortion network model for sensing temperature, the present invention only needs a few groups of data under typical temperatures, and can have good temperature calibration performance by combining the powerful exploration and adjustment capabilities of reinforcement learning itself; although it is necessary to first obtain the pre-distortion models at different temperatures, currently most of the digital pre-distortion models have low complexity and do not need to be updated in real time, so the increased computational cost is acceptable.

[0086] The present invention adopts the A3C reinforcement learning architecture, which uses multi-threading to asynchronously execute multiple agents on the basis of the AC (Actor-Critic) architecture, thereby accelerating the training process. Therefore, it has a similar network structure to the AC architecture. Referring to the appendix Figure 5 , the neural network model based on reinforcement learning includes: a feature extraction module, an Actor network module and a Critic network module. Among them, the feature extraction module is used to pre-extract the features of the input data, the Actor network module is used to generate the reinforcement learning policy π, and the Critic network module is used to generate the value function v of reinforcement learning; the feature extraction module is respectively connected to the Actor network module and the Critic network module.

[0087] Among them, the feature extraction module is at the front end of the Actor network module and the Critic network module, and consists of two convolutional layers. The input channel number of the first convolutional layer is 1, which matches the dimension of the image type data in the dataset. The output channel number is 16, the convolutional kernel size is 3*3, and the padding is 1. The input and output channel numbers of the second convolutional layer are both 16, the convolutional kernel size is 3*3, and the padding is 1. The role of the feature extraction module is to pre-extract the features of the input data, facilitating the subsequent Actor network module and Critic network module to extract more features of the input data.

[0088] The Actor network module includes two convolutional + ReLU layers and a Softmax activation layer, which is used to generate the policy function π of reinforcement learning. The input and output channel numbers of the first convolutional layer are both 16, the convolutional kernel size is 3*3, and the padding is 1. The input channel number of the second convolutional layer is 16, and the output channel number is the number of discrete actions included in the action space. In the present invention, the number of discrete actions is 3, so the output channel number is 3, corresponding to the different tendencies of the agent to adopt 3 discrete actions in the next step under the current state. Finally, the output of the second convolutional layer passes through the Softmax activation layer to obtain the probability distribution of 3 discrete actions, and the range of each value is [0,1].

[0089] The Critic network module contains two convolutional + ReLU layers, which is used to generate the value function v of reinforcement learning. The input and output channel numbers of the first convolutional layer are both 16, the convolutional kernel size is 3*3, and the padding is 1. The input channel number of the second convolutional layer is 16, and the output channel number is 1, and the output corresponds to the long-term reward for taking a certain action in the given state.

[0090] Because the size of the input image type data is 1*5*5, the image size is small, and at the same time, the most critical I / Q and envelope related items at the current moment are in the first column of the two-dimensional image data, which belong to the edge information of the image. In order to avoid losing important information and key edge information during image reduction, the padding of the convolutional layers in the above three modules is set to 1.

[0091] In an embodiment of the present invention, the state space and action space of the neural network based on reinforcement learning are constructed. First, the state space is constructed: the current input data and the corresponding temperature label are used as state data. The addition of the temperature label enables the reinforcement learning agent to perceive the change of temperature and introduces the temperature into the policy formulation of the Actor network and the value evaluation of the Critic network.

[0092] The construction of the action space includes two parts: discrete actions and continuous actions, and the parameter settings are shown in Table 1.

[0093] Table 1: Action Space Construction and Parameter Settings

[0094]

[0095] The present invention is a pixel-based algorithm. Therefore, different image filtering operations are used as optional actions in the discrete action space. Specifically, two filtering methods and the action of "holding" (i.e., not making additional actions) are selected as the discrete action space of the agent in reinforcement learning. The two filtering operations are bilateral filtering and Gaussian filtering respectively.

[0096] It is easy to observe that the pre-distortion signals output by the digital pre-distortion model at different temperatures do not differ much. Therefore, the construction of the action space focuses more on fine-tuning. The discrete action space should not be too large to avoid increasing the difficulty of network training. At the same time, the filtering intensity needs to be limited to a small range. Therefore, the initial value of the sigmaColor parameter of the bilateral filtering function is set to 0.1, the initial value of sigmaSpace is set to 5, and the window size d is set to 3; the initial values of the sigmaX and sigmaY parameters of the Gaussian filtering function are both set to 0.5, and the Gaussian kernel size is [3,3].

[0097] In addition, a temperature correlation coefficient function is set to adjust the intensity of the filtering action according to the temperature, thereby increasing the accuracy of network adjustment. The temperature correlation coefficient function can be expressed as:

[0098]

[0099] In formula (1), T represents the temperature label of the current input data, T 0 represents the reference temperature (i.e., the environmental temperature corresponding to the input data of reinforcement learning), and f T represents the temperature correlation coefficient. The larger the absolute value of the difference between the current temperature and the reference temperature, the greater the temperature change, and the greater the difference between the corresponding pre-distortion signals. Therefore, the temperature correlation coefficient is also greater. The parameter values in the bilateral filtering function are introduced with temperature-related terms, that is, sigmaColor = 0.1×f T , sigmaSpace = 5×f T , and the window size remains unchanged; in the Gaussian filtering function, sigmaX = sigmaY = 0.5×f T , and the Gaussian kernel size remains unchanged.

[0100] By observing the characteristic curves of the power amplifier at different temperatures, it can be seen that the power gain at different temperatures changes. Reflecting on the time-domain signal, it is that the signal amplitude is slightly different. Therefore, the pre-distortion signals at different temperatures are regarded as the results of the pre-distortion signal at the reference temperature scaled by different scaling factors. Therefore, the continuous action is set as:

[0101] I adjusted = I filtered × (1 + a) + b (2)

[0102] In formula (2), I adjusted represents the image data after continuous action adjustment, I filtered represents the image data after filtering operation, a represents the scaling factor, b represents the deviation adjustment value, where a can be regarded as a linear adjustment and b can be regarded as a non - linear adjustment. The adjustment of continuous actions also requires precise control, so the ranges of a and b are controlled within [-0.1, 0.1].

[0103] The influence of environmental temperature also needs to be introduced into the continuous action space. Through actual observation, it can be found that the power gain of the power amplifier is larger at lower temperatures. Therefore, the amplitude of the corresponding predistortion signal is smaller at lower temperatures. So a and b can be expressed as: a = 0.1 + 0.01×(T - T 0 ), b = 0.05×(T - T 0 ). After the state space and action space are defined, the intelligent agent can select different discrete actions according to the current state and further adjust the filtered image data through continuous actions to obtain a predistortion output that precisely matches the corresponding temperature.

[0104] In addition to the construction of the state space and action space having a great influence on the performance of the reinforcement learning model, the setting of the reward function is also crucial. The algorithm implemented in the present invention is a temperature - adaptive adjustment algorithm, which can adjust the initial predistortion signal at the reference temperature to the corresponding predistortion signal after the environmental temperature changes. Therefore, the improvement degree of the deviation between the input and output data is used as the measurement standard for the current state - action reward.

[0105] The intelligent agent in the neural network model based on reinforcement learning operates on 5×5 pixel elements in the input data. Then the reward can be expressed as:[[]]

[0106] R t = (I target - s t ) 2 - (I target - s t+1 ) 2 (3)

[0107] In formula (3), I target represents the corresponding ideal predistortion output after temperature change, s t represents the predistortion output in the current state, s t+1 represents the predistortion output in the next state. As the intelligent agent continuously explores, when the predistortion output in the next state is closer to the ideal value, R tLarger, thus encouraging the agent to continue exploring in this direction and updating the weight parameters of the Actor network and the Critic network accordingly. Maximizing the total reward is equivalent to minimizing the squared error between the final network output and the ideal pre-distortion output.

[0108] Based on the Temporal-Difference (TD) method, the update formula for the policy function of the Actor network and the update formula for the value function of the Critic network can be expressed as:

[0109]

[0110] In formulas (4) and (5), δ i = r i+1 + γV(s i+1 ) - V(s i ) represents the temporal-difference error of the agent, r i+1 and V(s t ) represent the reward and state value in the corresponding state, γ represents the discount factor, which is set to 0.95 here. That is, the reinforcement learning network not only considers the current reward but also incorporates the value of future rewards into the decision-making. θ a represents the parameters of the Actor network, represents the probability distribution of the a i action executed in state s i , β represents the entropy regularization coefficient, which controls the influence of the entropy term on the loss function. A larger value makes the agent more inclined to explore, and it is set to 0.01 here. Among them, can be expressed as represents the entropy of the action distribution, which is used to evaluate the exploration ability of the current policy. θ v represents the parameters of the Critic network.

[0111] The policy loss function and the value loss function of the neural network model based on reinforcement learning are:

[0112]

[0113] In formula (6), L π represents the policy loss function, and L v represents the value loss function;

[0114] By adjusting the contributions of the policy and value losses to the total loss through the weighting coefficient, L total = ω π ·L π + ω v ·L v , where ω π and ω vrespectively represent the weight coefficients of the policy and value loss, set to 1.0 and 0.5, and then use L total to participate in backpropagation and gradient update to optimize the parameters of the model.

[0115] Input the dataset with temperature labels obtained in step S1 into the reinforcement learning network model for training. After multiple rounds of iteration, a reinforcement learning network model with a converged loss function is obtained.

[0116] The reinforcement learning network model with a converged loss function is cascaded with the digital predistortion model and the power amplifier. Refer to the appendix Figure 6 as shown. Appendix Figure 6 In it, the vertical axis represents the gain, and the horizontal axis represents the input power. When the temperature changes from T1 to T2, the gain characteristic curve of the power amplifier changes. The reinforcement learning network model adjusts the output of the DPD model to make it more accurately match the changed power amplifier characteristics.

[0117] The current ambient temperature can be sensed by a temperature sensor. The data to be input into the power amplifier is first input into the digital predistortion model, and then the reinforcement learning network can adjust the predistorted signal according to the sensed ambient temperature. The adjusted predistorted signal is converted into I / Q two-channel signals or complex signals and finally input into the power amplifier to implement the temperature adaptive power amplifier linearization technology and ensure the linearization effect of the digital predistortion technology under temperature changes.

[0118] The present invention takes the ambient temperature as a variable and introduces it into the state space and action space of the reinforcement learning model. Under the idea of image enhancement, the discrete actions are set to two filtering methods, bilateral filtering and Gaussian filtering. At the same time, according to the designed temperature-related coefficient function, the filtering intensity is dynamically adjusted according to the ambient temperature; then the filtered data is processed through continuous actions, and the linear and non-linear adjustment amounts in the continuous actions are also dynamically adjusted according to the ambient temperature.

[0119] When the characteristics of the input signal change or the characteristics of the power amplifier change due to ambient temperature changes, the adaptive predistortion architecture requires continuous updating of the predistortion model parameters to match the current power amplifier model. However, the present invention only needs to train a reinforcement learning network model based on the data at different temperatures, and can dynamically adjust the output of the digital predistortion model within a large temperature change range without frequently updating the digital predistortion model parameters, thus greatly reducing the computational cost. The network model that can perform thermal compensation implemented by traditional deep learning requires a large amount of training data at each temperature. In order to ensure the generalization and adjustment accuracy of the model, the model usually has a large number of parameters. The present invention takes advantage of the characteristics that reinforcement learning can interact with the ambient temperature and continuously explore and learn, and the demand for training data and the number of model parameters are small.

[0120] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be covered within the protection scope of the present invention.

Claims

1. A power amplifier linearization thermal compensation method based on reinforcement learning, characterized in that: Includes steps: S1, building a digital pre-distortion model and obtaining a data set with temperature labels; S2. Build a neural network model based on reinforcement learning; S3, constructing the state space and action space of the neural network model, introducing temperature into the state space and the action space, and adjusting the predistortion signal output by the digital predistortion model; S4, defining parameters related to the training of the neural network model, inputting the data set into the neural network model for training, and obtaining a reinforcement learning network model with converged loss function; S5. The reinforcement learning network model is cascaded with the digital pre-distortion model and the power amplifier to realize a temperature-adaptive power amplifier linearization function.

2. A power amplifier linearization thermal compensation method based on reinforcement learning according to claim 1, characterized in that: The step S1 comprises: S11: testing and saving the input and output data of the power amplifier under different ambient temperatures, and performing timing alignment through a cross-correlation method to obtain aligned data; S12: for different types of digital pre-distortion models, the aligned data are processed accordingly, and correlation coefficients of the digital pre-distortion models are calculated to obtain digital pre-distortion models at different temperatures; S13: adding the same excitation signal to the digital predistortion model at different temperatures to obtain accurate predistortion signals at different temperatures; S14: Encapsulate the predistorted signal into a data set with a temperature label.

3. A power amplifier linearization thermal compensation method based on reinforcement learning according to claim 2, characterized in that: The step S12 comprises: S121: The digital pre-distortion model is based on a polynomial, and a corresponding characteristic matrix is ​​obtained according to the power amplifier output data obtained by the test, and a final coefficient vector is obtained by inverting the characteristic matrix; S122: The digital pre-distortion model is based on a neural network, processes the test data into a form suitable for network training, and completes the training and verification of the network model.

4. A power amplifier linearization thermal compensation method based on reinforcement learning according to claim 2, characterized in that: The step S14 comprises: S141: taking the predistortion signal at room temperature as original data; S142: Selecting several groups of pre-distorted signals of typical temperature data within the entire temperature range as label data, and adding the corresponding ambient temperature to the label data; S143: merging the tag data, and processing the temperature corresponding to the tag data into a temperature vector; S144: Process the formats of the original data and the label data, and encapsulate the format-processed original data, label data, and temperature vector into a data set with a temperature label.

5. The power amplifier linearization thermal compensation method based on reinforcement learning according to claim 1, characterized in that: In step S2, the neural network model based on reinforcement learning includes: A feature extraction module is used to extract features of input data in advance; Actor network module, used to generate reinforcement learning strategy π; Critic network module, used to generate the value function v of reinforcement learning; The feature extraction module is connected to the Actor network module and the Critic network module respectively.

6. A power amplifier linearization thermal compensation method based on reinforcement learning according to claim 5, characterized in that: The feature extraction module includes two convolutional layers, the first convolutional layer of the feature extraction module has an input channel number of 1, matching the dimension of the image type data in the data set, and an output channel number of 16, and the second convolutional layer of the feature extraction module has an input channel number and an output channel number of 16; The Actor network module includes two convolutional layers and a Softmax activation layer. The number of input channels and the number of output channels of the first convolutional layer of the Actor network module are both 16. The number of input channels of the second convolutional layer of the Actor network module is 16, and the number of output channels is 3. The output of the second convolutional layer of the Actor network module is passed through the Softmax activation layer to obtain the probability distribution of three discrete actions. The critic network module includes two convolutional layers, the number of input channels and the number of output channels of the first convolutional layer of the critic network module are both 16, the number of input channels of the second convolutional layer of the critic network module is 16, and the number of output channels is 1.

7. The power amplifier linearization thermal compensation method based on reinforcement learning according to claim 5, characterized in that: In step S3, constructing the state space includes: taking the current input data and the corresponding temperature label as the state data. The addition of the temperature label enables the reinforcement learning agent to perceive the temperature change, and introduces the temperature into the strategy formulation of the Actor network and the value evaluation of the Critic network.

8. The power amplifier linearization thermal compensation method based on reinforcement learning according to claim 5, characterized in that: In step S3, constructing the action space includes: There are two parts: discrete action and continuous action. Two filtering methods and action "maintenance" are selected as the discrete action space of the agent in reinforcement learning. The two filtering methods are bilateral filtering and Gaussian filtering. The initial value of the sigmaColor parameter of the bilateral filter function is set to 0.1, the initial value of sigmaSpace is set to 5, and the window size d is set to 3; The initial values ​​of the sigmaX and sigmaY parameters of the Gaussian filter function are both set to 0.5, and the Gaussian kernel size is [3,3]; A temperature correlation coefficient function is set, and the temperature correlation coefficient function is expressed as: Among them, T represents the temperature label of the current input data, T0 represents the reference temperature, that is, the ambient temperature corresponding to the reinforcement learning input data, and f T represents the temperature correlation coefficient; The temperature correlation coefficient function is introduced into the bilateral filter function, sigmaColor = 0.1 × f T , sigmaSpace=5×f T , the window size remains unchanged; The temperature correlation coefficient function is introduced into the Gaussian filter function, sigmaX = sigmaY = 0.5 × f T , the Gaussian kernel size remains unchanged; The continuous action is set as: I adjusted =I filtered ×(1+a)+b Among them, I adjusted Represents the image data after continuous action adjustment, I filtered represents the image data after the filtering operation, a represents the scaling factor, and b represents the deviation adjustment value, where a is regarded as a linear adjustment and b is regarded as a nonlinear adjustment; When the temperature is low, the power gain of the power amplifier is large, and the corresponding predistortion signal has a smaller amplitude at a low temperature. a and b are expressed as: a=0.1+0.01×(T-T0), b=0.05×(T-T0).

9. The power amplifier linearization thermal compensation method based on reinforcement learning according to claim 5, characterized in that: In step S4, defining the parameters of the neural network model includes: The initial predistortion signal at the reference temperature is adjusted to the corresponding predistortion signal after the ambient temperature changes, and the improvement degree of the deviation between the input and output data is used as the measure of the current state action reward; The agent in the neural network model operates on 5×5 pixel elements in the input data, and the reward is expressed as: R t =(I target -s t ) 2 -(I target -s t+1 ) 2 Among them, I target represents the corresponding ideal predistortion output after temperature change, s t Indicates the predistortion output in the current state, s t+1 represents the predistorted output of the next state. As the agent continues to explore, the predistorted output of the next state is closer to the ideal value, R t is larger, thus encouraging the agent to continue exploring in this direction and update the weight parameters of the Actor network and the Critic network accordingly. Maximizing the total reward is equivalent to minimizing the square of the error between the final network output and the ideal pre-distorted output. The strategy function update formula of the Actor network based on the temporal difference method and the value function update formula of the Critic network are expressed as follows: Among them, δ i =r i+1 +γV(s i+1 )-V(s i ) represents the temporal difference error of the agent, r i+1 and V(s t ) represents the reward and state value under the corresponding state, γ represents the discount factor, which is set to 0.95, that is, the reinforcement learning network not only considers the current reward, but also incorporates the value of future rewards into the decision, θ a Represents the parameters of the Actor network, Indicates that in state s i The following is executed i The probability distribution of the action, β represents the entropy regularization coefficient, which controls the effect of the entropy term on the loss function. A larger value makes the agent more inclined to explore, and is set to 0.

01. It can be expressed as Represents the entropy of the action distribution, which is used to evaluate the exploration ability of the current strategy, θ v Represents the parameters of the Critic network.

10. A power amplifier linearization thermal compensation method based on reinforcement learning according to claim 9, characterized in that: The strategy loss function and value loss function of the neural network model are expressed as: Among them, L π represents the policy loss function, L v represents the value loss function; The contribution of strategy loss and value loss to the total loss is adjusted by the weighted coefficient, L total =ω π ·L π +ω v ·L v , where ω π and ω v Represent the weight coefficients of strategy loss and value loss, set to 1.0 and 0.5, and then use L total Participate in back-propagation and gradient updates to optimize the model's parameters.

Citation Information

Cited By

  • Satellite-ground joint nonlinear and linear distortion correction method and device based on reinforcement learning

    CN120729398A

  • Satellite-ground joint nonlinear linear distortion correction method and device based on reinforcement learning

    CN120729398B