Model architecture search and hardware optimization
Model architecture search techniques optimize DPD configurations in RF transceivers to enhance linearity and efficiency of power amplifiers, addressing nonlinear behavior and distortion challenges in high-power applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ANALOG DEVICES INC
- Filing Date
- 2022-05-09
- Publication Date
- 2026-05-11
AI Technical Summary
Existing RF systems face challenges in achieving linearity and efficiency in power amplifiers due to nonlinear behavior, particularly in high-power applications, which leads to distortion and reduced modulation accuracy, and conventional DPD methods struggle with increasing sampling rates and physical constraints.
Implementing model architecture search techniques, such as neural architecture search (NAS), to optimize DPD configurations in RF transceivers by training parameterized models on target hardware, leveraging differentiable building blocks to improve linearity and efficiency of power amplifiers.
Enhances the linearity and efficiency of power amplifiers by optimizing DPD configurations, reducing distortion and improving modulation accuracy, while addressing the limitations of conventional DPD methods in high-frequency RF systems.
Smart Images

Figure 0007856675000030 
Figure 0007856675000031 
Figure 0007856675000032
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 187,536, filed on 12 May 2021, entitled “DIGITAL PREDISTORTION FOR POWER AMPLIFIER LINEARIZATION USING NEURAL NETWORKS,” and U.S. Non-Provisional Patent Application No. 17 / 732,809, filed on 29 April 2022, entitled “MODEL ARCHITECTURE SEARCH AND OPTIMIZATION FOR HARDWARE,” respectively, which are incorporated herein by reference in their entirety as described below, and for all applicable purposes.
[0002] This disclosure relates, in general, to electronics, and more specifically to configuring hardware blocks (e.g., digital pre-distortion (DPD) hardware for linearizing power amplifiers) using model architecture search techniques (e.g., neural architecture search (NAS)). [Background technology]
[0003] RF systems are systems that transmit and receive signals in the form of electromagnetic waves in the RF range of approximately 3 kilohertz (kHz) to 300 gigahertz (GHz). RF systems are commonly used in wireless communication, with cellular / wireless mobile technology being a prominent example, but they can also be used in cable communications such as cable television. In both of these types of systems, the linearity of the various components within them plays a crucial role.
[0004] The linearity of RF components or systems, such as RF transceivers, is theoretically easy to understand. That is, linearity generally refers to the ability of a component or system to provide an output signal that is directly proportional to the input signal. In other words, if a component or system is perfectly linear, the relationship between the ratio of the output signal to the input signal is a straight line. Achieving this behavior in actual components and systems is far more complex, and often requires overcoming many challenges to linearity, often at the expense of other performance parameters such as efficiency and / or output power.
[0005] Power amplifiers (PAs), fabricated from inherently nonlinear semiconductor materials and required to operate at relatively high power levels, are typically the first components to analyze when considering the design of an RF system from a linearity perspective. PA outputs with nonlinear distortion can result in reduced modulation accuracy (e.g., reduced error vector magnitude (EVM)) and / or out-of-band radiation. Therefore, both wireless RF systems (e.g., Long-Term Evolution (LTE) and millimeter-wave or fifth-generation (5G) systems) and cable RF systems have stringent specifications regarding PA linearity. [Overview of the Initiative] [Means for solving the problem]
[0006] DPD (Digital-to-Phonetic Distortion) can be applied to improve the linearity of a PA system. Typically, DPD involves applying pre-distortion to the signal provided as input to the PA in the digital domain to reduce and / or cancel the distortion expected to be caused by the PA. Pre-distortion can be characterized by a PA model, which can be updated based on feedback from the PA (i.e., based on the PA's output). The more accurate the PA model is in predicting the distortion introduced by the PA, the more effective the pre-distortion of the input to the PA will be in reducing the impact of distortion caused by the amplifier.
[0007] Implementing DPD in RF systems is not an easy task, as various factors can affect its cost, quality, and robustness. Physical constraints, such as space / surface area, and regulations can also impose further limitations on DPD requirements or specifications. As sampling rates used in state-of-the-art RF systems continue to increase, DPD becomes particularly challenging, and therefore, trade-offs and ingenuity are required in DPD design.
[0008] To provide a more complete understanding of this disclosure and its features and advantages, please refer to the following description in conjunction with the attached drawings, where similar reference numerals represent similar components. [Brief explanation of the drawing]
[0009] [Figure 1A] This disclosure provides schematic block diagrams of exemplary radio frequency (RF) transceivers in which parameterized model-based digital pre-distortion (DPD) may be implemented according to several embodiments of this disclosure. [Figure 1B] This disclosure provides schematic block diagrams of exemplary indirect learning architecture-based DPDs in which parameterized model-based configurations can be implemented according to several embodiments of this disclosure. [Figure 1C] This disclosure provides schematic block diagrams of exemplary direct learning architecture-based DPDs in which parameterized model-based configurations can be implemented according to several embodiments of this disclosure. [Figure 2A] This disclosure provides illustrative diagrams of schemes for offline training and online adaptation and operation of indirect learning architecture-based DPDs according to several embodiments of this disclosure. [Figure 2B] This disclosure provides illustrative diagrams of offline training, online adaptation, and operation of direct learning architecture-based DPDs according to several embodiments of this disclosure. [Figure 3] This disclosure provides illustrative diagrams of exemplary implementations of lookup table (LUT)-based DPD actuator circuits according to several embodiments of this disclosure. [Figure 4] Provide an illustrative diagram of an exemplary implementation of a LUT-based DPD actuator circuit according to some embodiments of the present disclosure. [Figure 5] Provide an illustrative diagram of an exemplary implementation of a LUT-based DPD actuator circuit according to some embodiments of the present disclosure. [Figure 6] Provide an illustrative diagram of an exemplary software model derived from a hardware design having a one-to-one function mapping according to some embodiments of the present disclosure. [Figure 7] Provide an illustrative diagram of an exemplary method for training a parameterized model for DPD operation according to some embodiments of the present disclosure. [Figure 8] Provide a schematic illustrative diagram of an exemplary parameterized model that models DPD operation as a sequence of differentiable function blocks according to some embodiments of the present disclosure. [Figure 9] A flowchart illustrating an exemplary method for training a parameterized model for DPD operation according to some embodiments of the present disclosure. [Figure 10] A flowchart illustrating an exemplary method for performing DPD operation for online operation and adaptation according to some embodiments of the present disclosure. [Figure 11] Provide a schematic illustrative diagram of an exemplary mapping of a sequence of hardware blocks to a sequence of differentiable function blocks according to some embodiments of the present disclosure. [Figure 12] Provide a schematic illustrative diagram of an exemplary mapping of a sequence of hardware blocks to a sequence of differentiable function blocks according to some embodiments of the present disclosure. [Figure 13] Provide a flowchart illustrating a method for training a parameterized model mapped to target hardware according to some embodiments of the present disclosure. [Figure 14]This disclosure provides flowcharts illustrating methods for performing operations on target hardware configured based on a parameterized model, according to several embodiments of this disclosure. [Figure 15] This disclosure provides block diagrams illustrating exemplary data processing systems that may be configured to implement or control at least part of a hardware block configuration using a neural network, according to some embodiments of this disclosure. [Modes for carrying out the invention]
[0010] overview Each of the systems, methods, and devices described herein has several innovative embodiments, and no single embodiment alone represents all of the desirable attributes disclosed herein. Details of one or more implementations of the subject matter described herein are given in the following description and accompanying drawings.
[0011] For the purpose of illustrating DPD using the neural networks proposed herein, it may be useful to first understand the phenomena that may affect RF systems. The following basic information can be considered as a basis for properly explaining this disclosure. Such information is provided for illustrative purposes only and should not be construed as limiting the broad scope and potential uses of this disclosure.
[0012] As explained above, the PA (Power Amplifier) is typically the first component to analyze when considering the design of an RF system from the perspective of linearity. A linear and efficient PA is essential for wireless and cable RF systems. While linearity is also important for small-signal amplifiers such as low-noise amplifiers, the challenges of linearity are particularly pronounced for PAs because such amplifiers typically need to generate relatively high levels of output power and are therefore particularly prone to certain operating conditions where nonlinear behavior can no longer be ignored. On the one hand, the nonlinear behavior of the semiconductor materials used to form the amplifier tends to worsen when the amplifier operates with high-power signal levels (operating conditions commonly referred to as "operating at saturation"), increasing the amount of nonlinear distortion in its output signal, which is highly undesirable. On the other hand, an amplifier operating at relatively high power levels (i.e., operating at saturation) also typically operates at its highest efficiency, which is highly desirable. As a result, linearity and efficiency (or power level) are two performance parameters for which an acceptable trade-off must often be found, in that improvement in one of these parameters often comes at the cost of the other parameter being suboptimal. To this end, the term "backoff" is used in this art to describe a measure of how much the input power (i.e., the power of the signal supplied to the amplifier being amplified) should be reduced in order to achieve the desired output linearity (for example, backoff can be measured as the ratio of the input power providing maximum power to the input power providing the desired linearity). Thus, reducing the input power may result in an improvement in terms of linearity, but it will result in a decrease in the efficiency of the amplifier.
[0013] As explained above, DPD can pre-distort the input to the PA to reduce and / or cancel distortion caused by the amplifier. To achieve this function, at a high level, DPD involves forming a model of how the PA may affect the input signal, and this model defines the coefficients of a filter applied to the input signal in an attempt to reduce and / or cancel distortion caused by the amplifier (such coefficients are referred to as "DPD coefficients"). In this way, DPD attempts to compensate an amplifier for applying undesirable nonlinear modifications to the transmitted signal by applying modifications corresponding to the input signal provided to the amplifier.
[0014] The models used in DPD algorithms are typically adaptive models, meaning they are formed through an iterative process by gradually adjusting coefficients based on a comparison between data entering an amplifier's input and data exiting the amplifier's output. Estimating DPD coefficients is based on the acquisition of a finite sequence of input and output data (i.e., inputs to and outputs from the PA), commonly referred to as "capture," and the formation of a feedback loop to which the model is adapted based on the analysis of the capture. More specifically, conventional DPD algorithms are based on a general memory polynomial (GMP) model, which involves forming a set of polynomial equations commonly referred to as "update equations," and updating the model of the PA by searching for suitable solutions to the equations in a broad solution space. To this end, the DPD algorithm solves an inverse problem, which is the process of calculating the random factors that generated a set of observations from a set of observations.
[0015] Solving inverse problems in the presence of nonlinear effects can be difficult and may involve inappropriate settings. In particular, the inventors of this disclosure recognize that GMP-based PA models may have limitations due to the signal dynamics required to store polynomial data and limited memory depth, especially in the context of constantly increasing sampling rates used in state-of-the-art RF systems.
[0016] Solid-state devices usable at high frequencies are extremely important in modern semiconductor technology. Partly due to their large bandgap and high mobility, III-N transistors (i.e., transistors using compound semiconductor materials having a first sublattice of at least one element from Group III of the periodic table (e.g., Al, Ga, In) and a second sublattice of nitrogen (N) as channel material), GaN-based transistors, and so on, can be particularly advantageous for high-frequency applications. In particular, PAs can be constructed using GaN transistors.
[0017] GaN transistors possess desirable characteristics in terms of cutoff frequency and efficiency, but their behavior is complicated by an effect known as charge trapping, where defect sites within the transistor channel trap charge carriers. The density of trapped charge is highly dependent on the gate voltage, which is typically proportional to the signal amplitude. To further complicate matters, the opposite effect can simultaneously compete with the charge trapping effect: when some charge carriers are trapped by a defect site, other charge carriers are released from the trap, for example, due to thermal activation. These two effects have very different time constants; whenever the gate voltage increases, the defect site can be quickly filled with trapped charge, while the release of trapped charge occurs much more slowly. The release time constant can range from 10 microseconds to as many as milliseconds, and its effect is typically very pronounced on the symbol duration timescale of 4G or 5G data, especially for data containing bursts.
[0018] Various embodiments of this disclosure provide systems and methods aimed at improving one or more of the drawbacks described above in providing linear and efficient amplifiers (such as PAs) to RF systems (including, but not limited to, wireless RF systems using millimeter-wave / 5G technologies). In particular, aspects of this disclosure provide techniques for behaviorally modeling hardware operation using differentiable building blocks and performing model architecture searches (e.g., differentiable neural architecture searches (DNAS)) using datasets collected on target hardware. While aspects of this disclosure describe techniques for applying model architecture searches to optimize DPD placement for linearizing power amplifiers in RF transceivers, the techniques disclosed herein are suitable for use in optimizing configurations for any suitable hardware blocks and / or subsystems.
[0019] According to one aspect of the present disclosure, a computer implementation system may implement a method for performing a model architecture search to optimize the configuration of target hardware to perform a specific data transformation. The data transformation may include linear and / or nonlinear operations and may generally include operations that change the representation of a signal from one form to another. The target hardware may include a pool of processing units capable of performing at least arithmetic operations and / or signal selection operations (e.g., multiplexing and / or demultiplexing). The model architecture search may be performed in a search space, including the pool of processing units and associated capabilities, desired hardware resource constraints, and / or hardware operations associated with the data transformation. The model architecture search may also be performed to achieve a specific desired performance metric associated with the data transformation, for example, to minimize an error metric associated with the data transformation.
[0020] As used herein, a pool of processing units may include, but is not limited to, digital hardware blocks (e.g., digital circuits including combinational logic and gates), general processors, digital signal processors, and / or microprocessors that execute instruction code (e.g., software and / or firmware), analog circuits, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and the like. Generally, a processing unit (or simply a hardware block) may be a circuit having defined inputs, outputs, and / or control signals. Furthermore, multiple processing units (e.g., circuit blocks) can be connected in a defined manner to form subsystems for performing data transformations, for example, including sequences of transformations. Hardware configuration optimization can be performed at the functional level (e.g., using input / output correspondences) and / or at the subsystem level (e.g., including sequences of operations).
[0021] To perform a model architecture search, the computer implementation system may receive information associated with a pool of processing units. The received information may include hardware resource constraints, hardware operations, and / or hardware functions associated with the pool of processing units. The computer implementation system may further receive a dataset associated with a data transformation operation. The dataset may be collected on target hardware and may include input data, output data, control data, etc. The computer implementation system may use the received hardware information and the received dataset to train a parameterized model associated with the data transformation. Training may include updating at least one parameter of the parameterized model associated with configuring at least a subset of the processing units in the pool (performing the data transformation). The computer implementation system may output one or more configurations for at least a subset of the processing units in the pool.
[0022] In some embodiments, the computer implementation system may further generate parameterized models by, for example, generating mappings between different differentiable function blocks from each of the processing units in the pool. That is, there is a one-to-one correspondence between each processing unit in the pool and each of the differentiable function blocks.
[0023] In some embodiments, the data transformation operation may include a sequence of at least a first and a second data transformation, and training may include computing a first parameter associated with the first data transformation (e.g., a first learnable parameter) and a second parameter associated with the second data transformation (e.g., a second learnable parameter). In some embodiments, computing the first parameter associated with the first data transformation and the second parameter associated with the second data transformation may further be based on backpropagation and loss functions. In some embodiments, the first or second data transformation in the sequence may be associated with executable instruction code. In other words, the parameterized model can model hardware operation implemented by a processor-executable digital circuit and / or analog circuit and / or instruction code (e.g., firmware).
[0024] In certain embodiments, the data conversion operation may be associated with a DPD for pre-distorting an input signal to a nonlinear electronic component (e.g., a PA). In one example, the data conversion may correspond to a DPD operation. In this regard, a first data conversion in the sequence may include selecting memory terms from the input signal based on a first parameter. A second data conversion in the sequence may include generating feature parameters associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory terms, based on a second parameter. The sequence associated with the data conversion operation may further include a third data conversion, which includes generating a pre-distorted signal based on the feature parameters. In another example, the data conversion may correspond to a DPD adaptation. In this regard, a first data conversion in the sequence may include selecting memory terms from a nonlinear electronic component or a feedback signal representing the output of the input signal, based on a first parameter. A second data conversion in the sequence may include generating features associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory terms, based on a second parameter. The sequence associated with the data transformation operation may further include a third data transformation, which involves updating coefficients based on the features and a second signal (e.g., a pre-distorted signal for indirect learning DPD, or the difference between the input signal and the feedback signal for direct learning DPD).
[0025] In some embodiments, the computer implementation system may include a memory for storing instructions and one or more computer processors, and when executed by one or more computer processors, the instructions cause one or more computer processors to perform a model architecture lookup method for configuring target hardware. In other embodiments, the model architecture lookup method may be in the form of an encoded instruction in a non-temporary computable-readable storage medium, which, when executed by one or more computer processors, causes one or more computer processors to perform the method.
[0026] In a further aspect of this disclosure, the device may include an input node for receiving an input signal and a pool of processing units for performing one or more arithmetic operations and / or one or more signal selection operations (e.g., multiplexing and / or demultiplexing). Each processing unit in the pool may be associated with at least one parameterized model (e.g., a NAS model) corresponding to a data transformation (e.g., including linear and / or nonlinear operations). The device may further include a control block for configuring and / or selecting at least a first subset of processing units for processing an input signal to generate a first signal based on the first parameterized model. In some aspects, the first parameterized model may be trained offline from each of the processing units in the pool based on a mapping between one of several differentiable building blocks and at least one of an input dataset or an output dataset collected on target hardware or hardware constraints. For example, training may be based on a NAS across several differentiable building blocks.
[0027] In some embodiments, data conversion may include a sequence of data conversions. For example, a data conversion may include a first data conversion followed by a second data conversion, where the first data conversion converts an input signal to a first signal, and the second data conversion converts the first signal to a second signal. In some embodiments, the sequence of data conversions may be carried out by a combination of digital hardware blocks (e.g., digital circuits) and a processor that execute instruction codes (e.g., software or firmware). For example, a first subset of processing units may include digital hardware blocks (e.g., digital circuits) for carrying out the first conversion, and a control block may further constitute a second subset of processing units in a pool for executing instruction codes to carry out the second conversion.
[0028] In certain embodiments, the device may be a DPD device for pre-distorting an input signal to a nonlinear electronic component (e.g., a PA). For example, an input signal received at an input node may correspond to an input signal to a nonlinear electronic component, and the first signal may correspond to a pre-distorted signal. The device may further include a memory for storing one or more lookup tables (LUTs) associated with one or more nonlinear characteristics of the nonlinear electronic component, based on a first parameterized model and DPD coefficients. The device may further include a DPD block, which includes a first subset of processing units. In the case of DPD operation, the first subset of processing units may select a first memory term from the input signal based on a first parameterized model. The first subset of processing units may further generate a pre-distorted signal based on one or more LUTs and the selected first memory term. In some embodiments, in the case of DPD adaptation using an indirect learning architecture, the first subset of processing units may further select a second memory term from a feedback signal associated with the output of the nonlinear electronic component, based on a first parameterized model. The control block may further configure a second subset of processing units to execute instruction code for calculating or updating DPD coefficients based on a first parameterized model, a selected second memory term, a set of basis functions, and input signals. The instruction code may also cause the second subset of processing units to update at least one of one or more LUTs based on the calculated coefficients and set of basis functions. In other embodiments, for DPD adaptation using a direct learning architecture, the control block may further configure a second subset of processing units to execute instruction code for calculating or updating DPD coefficients based on a first parameterized model, a selected first memory term, a set of basis functions, and the difference between input and feedback signals, and for updating at least one of one or more LUTs based on the calculated coefficients and set of basis functions.
[0029] The systems, schemes, and mechanisms described herein advantageously leverage NAS technology to search for the optimal configuration for configuring hardware to perform specific data transformations. Because using heuristic search in finding the optimal configuration is complex, time-consuming, and therefore can increase the cost and / or time to market for deploying new hardware, this disclosure may be particularly advantageous for optimizing the configuration of hardware that performs complex data transformations such as DPD. Furthermore, this disclosure may be particularly advantageous for optimizing configurations with additional system constraints and / or multiple optimization targets. Moreover, model-architecture search-based hardware configurations may be particularly advantageous for certain transformations (input-to-output) that cannot be readily expressed by mathematical functions or that otherwise require very high-order polynomials.
[0030] Exemplary RF transceiver with DPD configuration Figure 1A provides a schematic block diagram of an exemplary RF transceiver 100 in which a parameterized model-based DPD may be implemented according to several embodiments of the present disclosure. As shown in Figure 1A, the RF transceiver 100 may include a DPD circuit 110, a transmitter circuit 120, a PA 130, an antenna 140, and a receiver circuit 150.
[0031] The DPD circuit 110 is configured to receive an input signal 102, represented by x, which may be a sequence of digital samples and may be a vector. Generally, as used herein, each of the lowercase, bold, italicized single-letter labels used in the figures (e.g., labels x, z, y, and y' shown in Figure 1A) refers to a vector. In some embodiments, the input signal 102x may include one or more active channels in the frequency domain, but for brevity, an input signal having only one channel (i.e., a single frequency range of in-band frequencies) is described. In some embodiments, the input signal x may be a baseband digital signal. The DPD circuit 110 is configured to generate an output signal 104, which may be represented by z, based on the input signal 102x. The DPD output signal 104z may be further provided to the transmitter circuit 120. For this purpose, the DPD circuit 110 may include a DPD actuator 112 and a DPD adaptive circuit 114. In some embodiments, the actuator 112 may be configured to generate an output signal 104z based on an input signal 102x and a DPD coefficient c calculated by a DPD adaptive circuit 114, as will be described in more detail below.
[0032] The transmitter circuit 120 may be configured to upconvert the signal 104z from a baseband signal to a higher frequency signal such as an RF signal. The RF signal generated by the transmitter 120 may be supplied to the PA 130, which may be implemented as a PA array containing N individual PAs. The PA 130 may be configured to amplify the RF signal generated by the transmitter 120 (thus the PA 130 may be driven by a drive signal based on the output of the DPD circuit 110) and output an amplified RF signal 131 which can be represented by y (e.g., a vector).
[0033] In some embodiments, the RF transceiver 100 may be a wireless RF transceiver, in which case it would also include an antenna 140. In the context of wireless RF systems, an antenna is a device that acts as an interface between radio waves propagating wirelessly through space and electric currents moving within a metal conductor used in a transmitter, receiver, or transceiver. During transmission, the transmitter circuit of the RF transceiver may supply an electrical signal, which is amplified by a PA, and the amplified version of the signal is provided to the antenna's terminals. The antenna can then radiate energy from the signal output by the PA as radio waves. Antennas are essential components of all wireless equipment and are used in radio broadcasting, broadcast television, two-way radio, communication receivers, radar, mobile phones, satellite communications, and other devices.
[0034] An antenna with a single antenna element typically broadcasts a radiation pattern that radiates equally in all directions within the spherical wavefront. A phased antenna array generally refers to a collection of antenna elements used to concentrate electromagnetic energy in a specific direction, thereby creating a primary beam; this is a process commonly referred to as "beamforming." Phased antenna arrays offer many advantages over single antenna systems, including high gain, the ability to perform directional steering, and simultaneous communication. Therefore, phased antenna arrays are more frequently used in a wide variety of applications, including mobile / cellular radio technology, military applications, aircraft radar, automotive radar, industrial radar, and Wi-Fi technology.
[0035] In an embodiment where the RF transceiver 100 is a wireless RF transceiver, an amplified RF signal 131y can be provided to an antenna 140, which may be implemented as an antenna array containing multiple antenna elements, for example, N antenna elements. The antenna 140 is configured to wirelessly transmit the amplified RF signal 131y.
[0036] In embodiments where the RF transceiver 100 is a wireless RF transceiver of a phase antenna array system, the RF transceiver 100 may further include a beamformer configuration configured to change the input signals provided to individual PAs of the PA array 130 in order to manipulate the beam generated by the antenna array 140. Such beamformer configurations are not specifically shown in Figure 1, for this reason, as they can be implemented in different ways, for example, as an analog beamformer (i.e., the input signals amplified by the PA array 130 are modified in the analog domain, i.e., after these signals have been converted from the digital domain to the analog domain), as a digital beamformer (i.e., the input signals amplified by the PA array 130 are modified in the digital domain, i.e., before these signals have been converted from the digital domain to the analog domain), or as a hybrid beamformer (i.e., the input signals amplified by the PA array 130 are modified partially in the digital domain and partially in the analog domain).
[0037] Ideally, the amplified RF signal 131y from PA130 should be an upconverted and amplified version of the output of transmitter circuit 120, e.g., an upconverted, amplified, and beamformed version of input signal 102x. However, as discussed above, the amplified RF signal 131y may have distortion outside the main signal component. Such distortion may arise from nonlinearity in the response of PA130. As discussed above, it may be desirable to reduce such nonlinearity. Therefore, RF transceiver 100 may further include a feedback path (or observation path) that enables the RF transceiver to analyze the amplified RF signal 131y from PA130 (in the transmission path). In some embodiments, the feedback path may be implemented as shown in Figure 1A, and the feedback signal 151y' may be provided from PA130 to receiver circuit 150. However, in other embodiments, the feedback signal may be a signal from a probe antenna element configured to sense a radio RF signal transmitted by antenna 140 (not specifically shown in Figure 1A).
[0038] Therefore, in various embodiments, at least a portion of the output of PA130 or the output of antenna 140 may be provided to receiver circuit 150 as a feedback signal 151. The output of receiver circuit 150 is coupled to DPD circuit 110, in particular DPD adaptive circuit 114. In this way, the output signal 151(y') of receiver circuit 150, which is based on the feedback signal 151 and indicates the output signal 131(y) from PA130, may be provided to DPD adaptive circuit 114 via receiver circuit 150. DPD adaptive circuit 114 may process the received signal and update the DPD coefficient c applied by DPD actuator circuit 112 to the input signal 102x to generate actuator output 104z. The signal based on actuator output z is provided as input to PA130, meaning that the operation of PA130 can be controlled using the DPD actuator output z.
[0039] According to aspects of this disclosure, the DPD circuit 110, including the DPD actuator circuit 112 and / or the DPD adaptive circuit 114, may be configured based on a parameterized model 170. The parameterized model 170 may be generated and trained offline by a parameterized model training system 172 (e.g., a computer implementation system such as the data processing system 2300 shown in Figure 15) using a model architecture retrieval technique (e.g., DNAS), as will be considered more fully below with reference to Figures 2A-2B and 3-14. Furthermore, the DPD actuator circuit 112 and / or the DPD adaptive circuit 114 may be configured to implement DPD using an indirect learning architecture as shown in Figure 1B or a direct learning architecture as shown in Figure 1C.
[0040] As further shown in Figure 1A, in some embodiments, the transmitter circuit 120 may include a digital filter 122, a digital-to-analog converter (DAC) 124, an analog filter 126, and a mixer 128. In such a transmitter, the pre-distorted signal 104z may be filtered in the digital domain by the digital filter 122 to produce a filtered pre-distorted input, i.e., a digital signal. The output of the digital filter 122 may then be converted to an analog signal by the DAC 124. The analog signal produced by the DAC 124 may then be filtered by the analog filter 126. The output of the analog filter 126 may then be upconverted to RF by the mixer 128, which receives a signal from a local oscillator (LO) 162 and converts the filtered analog signal from the analog filter 126 from baseband to RF. Other methods of implementing the transmitter circuit 120 are also possible and within the scope of this disclosure. For example, in another implementation (not illustrated in these drawings), the output of the digital filter 122 may be directly converted into an RF signal by the DAC 124 (for example, in a direct RF architecture). In such an implementation, the RF signal provided by the DAC 124 may then be filtered by the analog filter 126. Since the DAC 124 directly combines the RF signals in this implementation, the mixer 128 and local oscillator 162 illustrated in Figure 1A can be omitted from the transmitter circuit 120 in such an embodiment.
[0041] As further shown in Figure 1A, in some embodiments, the receiver circuit 150 may include a digital filter 152, an analog-to-digital converter (ADC) 154, an analog filter 156, and a mixer 158. In such a receiver, the feedback signal 151 may be down-converted to baseband by the mixer 158, which may receive a signal from a local oscillator (LO) 160 (which may be the same as or different from the local oscillator 160) to convert the feedback signal 151 from RF to baseband. The output of the mixer 158 may then be filtered by the analog filter 156. Next, the output of the analog filter 156 may be converted to a digital signal by the ADC 154. The digital signal generated by the ADC 154 may then be filtered in the digital domain by the digital filter 152 to produce a filtered and down-converted feedback signal 151y', which may be a sequence of digital values representing the output y of PA 130 and may be modeled as a vector. The feedback signal 151y' can be provided to the DPD circuit 110. Other methods of implementing the receiver circuit 150 are also possible and within the scope of this disclosure. For example, in another implementation (not illustrated in these drawings), the RF feedback signal 151y' can be directly converted to a baseband signal by the ADC 154 (e.g., in a direct RF architecture). In such an implementation, the downconverted signal provided by the ADC 154 can then be filtered by the digital filter 152. Since the ADC 154 directly synthesizes the baseband signal in this implementation, the mixer 158 and local oscillator 160 illustrated in Figure 1A can be omitted from the receiver circuit 150 in such an embodiment.
[0042] Further modifications are possible to the RF transceiver 100 described above. For example, while up-conversion and down-conversion have been described with respect to the baseband frequency, in other embodiments of the RF transceiver 100, an intermediate frequency (IF) may be used instead. The IF may be used in a superheterodyne radio receiver in which the received RF signal is shifted to the IF before final detection of the information of the received RF signal takes place. Conversion to IF may be useful for several reasons. For example, if several stages of filtering are used, they can all be set to a fixed frequency, which makes it easier to construct and tune the filters. In some embodiments, the mixer of the RF transmitter 120 or receiver 150 may include several such stages of IF conversion. In another example, a single-path mixer is shown for each of the transmit (TX) path (i.e., the signal path of the signal processed by the transmitter 120) and receive (RX) path (i.e., the signal path of the signal processed by the receiver 150) of the RF transceiver 100. However, in some embodiments, the TX path mixer 128 and the RX path mixer 158 may be implemented as quadrature upconverters and downconverters, respectively, in which case each of them would include a first mixer and a second mixer. For example, in the case of the RX path mixer 158, the first RX path mixer may be configured to perform downconversion to generate a common-mode (I) downconverted RX signal by mixing the feedback signal 151 with the common-mode component of the local oscillator signal provided by the local oscillator 160. A second RX path mixer may be configured to perform down-conversion and generate a quadrature (Q) down-converted RX signal by mixing the feedback signal 151 with the quadrature component of the local oscillator signal provided by the local oscillator 160 (the quadrature component is a component offset by 90 degrees in phase from the common-mode component of the local oscillator signal). The output of the first RX path mixer may be supplied to the I signal path, and the output of the second RX path mixer may be supplied to the Q signal path, which may be substantially out of phase with the I signal path by 90 degrees.Generally, the transmitter circuit 120 and the receiver circuit 150 may utilize a zero-IF architecture, a direct-conversion RF architecture, a complex IF architecture, a high (real) IF architecture, or any suitable RF transmitter and / or receiver architecture.
[0043] Generally, the RF transceiver 100 can be any device / apparatus or system configured to support the transmission and reception of signals in the form of electromagnetic waves within the RF range of approximately 3 kHz to 300 GHz. In some embodiments, the RF transceiver 100 may be used for wireless communication in base station (BS) or user equipment (UE) devices of any suitable cellular wireless communication technology, such as Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), or LTE. In further examples, the RF transceiver 100 may be used as, or within, a BS or UE device of millimeter-wave wireless technology such as 5G wireless (i.e., having frequencies in the range of approximately 20 to 60 GHz corresponding to wavelengths in the high-frequency / short-wavelength spectrum, e.g., in the range of approximately 5 to 15 millimeters). In yet another example, the RF transceiver 100 may be used for wireless communication using Wi-Fi technology (e.g., a 2.4 GHz frequency band corresponding to a wavelength of about 12 cm, or a 5.8 GHz frequency band corresponding to a wavelength of about 5 cm, spectrum) in Wi-Fi-enabled devices such as desktops, laptops, video game consoles, smartphones, tablets, smart TVs, digital audio players, cars, and printers. In some implementations, the Wi-Fi-enabled device may be, for example, a node in a smart system configured to communicate data with other nodes, such as smart sensors. In yet another example, the RF transceiver 100 may be used for wireless communication using Bluetooth technology (e.g., a frequency band of about 2.4 to about 2.485 GHz corresponding to a wavelength of about 12 cm). In another embodiment, the RF transceiver 100 may be used to transmit and / or receive wireless RF signals for purposes other than communication, for example, in automotive radar systems or in medical applications such as magnetic resonance imaging (MRI). In yet another embodiment, the RF transceiver 100 may be used for cable communication, for example, in a cable television network.
[0044] Figure 1B provides a schematic block diagram of an exemplary indirect learning architecture-based DPD 180 in which a parameterized model-based configuration can be implemented according to several embodiments of the present disclosure. In some embodiments, the DPD circuit 110 in Figure 1A may be implemented as shown in Figure 1B, and a parameterized model training system 172 may train the parameterized model 170 to configure the DPD circuit 110 for indirect learning-based adaptation. For simplification, the transmitter circuit 120 and receiver circuit 150 are not shown in Figure 1B, and only the elements relevant to implementing the DPD are shown.
[0045] In the case of indirect learning, the DPD adaptive circuit 114 can use the observed received signal (e.g., feedback signal 151y') as a reference to predict the PA input sample corresponding to the reference. The function used to predict the input sample is known as the inverse PA model (for linearizing PA 130). When the prediction of the input sample corresponding to the observed data is good (e.g., when the error between the predicted input sample and the pre-distorted signal 104z satisfies a certain criterion), the estimated inverse PA model is used to pre-distort the data to be transmitted to PA 130 (e.g., input signal 102x). That is, the DPD adaptive circuit 114 can pre-distort the input signal 102x by computing the inverse PA model used by the DPD actuator circuit 112. To this end, the DPD adaptive circuit 114 may observe or capture N samples of the PA input (from the pre-distorted signal 104z) and N samples of the PA output (from the feedback signal 151y'), calculate a set of M coefficients that can be represented by c corresponding to the inverse PA model, and update the DPD actuator circuit 112 with coefficient c as indicated by the dotted arrow. In some examples, the DPD adaptive circuit 114 may use least-squares approximation to find the set of coefficients c.
[0046] Figure 1C provides a schematic block diagram of an exemplary direct learning architecture-based DPD 190 in which a parameterized model-based configuration may be implemented according to several embodiments of the present disclosure. In some embodiments, the DPD circuit 110 in Figure 1A may be implemented as shown in Figure 1C, and a parameterized model training system 172 may train the parameterized model 170 to configure the DPD circuit 110 for direct learning. For simplification, the transmitter circuit 120 and receiver circuit 150 are not shown in Figure 1B, and only elements relevant to the implementation of the DPD are shown.
[0047] In the case of direct learning, the DPD adaptive circuit 114 can use the input signal 102x as a reference to minimize the error between the observed received data (e.g., the feedback signal 151y') and the transmitted data (e.g., the input signal 102x). In some examples, the DPD adaptive circuit 114 can use iterative techniques to compute a set of M coefficients, which can be represented by c, used by the DPD actuator circuit 112 to pre-distort the input signal 102x. For example, the DPD adaptive circuit 114 can compute the current coefficient based on previously computed coefficients (in previous iterations) and the currently estimated coefficient. The DPD adaptive circuit 114 can compute the coefficient to minimize the error that represents the difference between the input signal 102x and the feedback signal 151y'. The DPD adaptive circuit 114 can update the DPD actuator circuit 112 with the coefficient c, as indicated by the dotted arrow.
[0048] In some embodiments, the DPD actuator circuit 112 of the indirect learning-based DPD 180 in Figure 1B or the direct learning-based DPD 190 in Figure 1C may implement DPD operation using the Volterra series or the GMP model (a subset of the Volterra series), as shown below.
[0049]
number
[0050] In the formula, z[n] represents the nth sample of the pre-distorted signal 104z, and f k (.) represents the kth function of the DPD model (e.g., a set of M basis functions), and c ijk ||x[ni]|| represents a set of DPD coefficients (for example, to combine a set of M basis functions), where x[ni] and x[nj] represent samples of the input signal 102 delayed by i and j samples, respectively, and ||x[ni]|| represents the envelope or amplitude of sample x[ni]. In some cases, the values of the sample delays i and j may depend on the nonlinear characteristics of PA130 of interest regarding pre-distortion, and x[ni] and x[nj] may be referred to as the i,j cross-memory terms. Equation (1) illustrates, but is not limited to, the application of the GMP model to the envelope or amplitude of the input signal 102x. In general, the DPD actuator circuit 112 may apply DPD operation to the input signal 102x directly or after pre-processing the input signal 102x according to a pre-processing function represented by P(), which may be an amplitude function, amplitude square, or any preferred function.
[0051] In some embodiments, the DPD operating circuit 112 may implement equation (1) using one or more lookup tables (LUTs). For example, term Σ k c ijk f k (||x[ni]||) can be stored in the LUT, and the LUT for the i,j cross-memory term is represented as follows:
[0052]
number
[0053] Thus, as will be more fully considered below with reference to FIGS. 3-5, the operation of the DPD operation circuit 112 can include selecting a first memory term (e.g., x[n-i] and x[n-j]) from the input signal 102x and generating a pre-distorted signal 104z based on the LUT and the selected first memory term. In the case of DPD adaptation using the direct learning architecture shown in FIG. 1C, the operation of the DPD adaptation circuit 114 can include calculating a DPD coefficient (e.g., a set of coefficients c k based on the selected first memory term and the set of basis functions f k ) and updating one or more LUTs based on the calculated coefficients. On the other hand, in the case of DPD adaptation using the indirect learning architecture shown in FIG. 1B, the operation of the DPD adaptation circuit 114 can include selecting a second memory term (e.g., y’[n-i] and y’[n-j]) from the feedback signal 151y’ and calculating a DPD coefficient (e.g., a set of coefficients c k based on the selected second memory term and the set of basis functions f k ) and updating one or more LUTs based on the calculated coefficients. Thus, the DPD circuit 110 can include various circuits such as a memory for storing LUTs of various cross-memory terms, a multiplexer for memory term selection, multipliers, adders, and various other digital circuits and / or processors for executing instructions for performing DPD operations (e.g., operation and adaptation).
[0054] According to aspects of this disclosure, the parameterized model training system 172 may train the parameterized model 170 to configure the DPD actuation circuit 112 and / or DPD adaptation circuit 114 to perform these DPD actuation and adaptation (indirect and / or direct learning) operations. Mechanisms for training the parameterized model 170 (e.g., while offline) and configuring the DPD hardware for actuation and adaptation according to the trained parameterized model 170 (e.g., while online) are discussed in more detail below with reference to Figures 2A-2B and 3-14. For brevity, Figures 2A-2B and 3-14 are discussed using the same signal representation as in Figures 1A-1C. For example, the symbol x may refer to the input signal to the DPD actuator circuit that linearizes the PA, the symbol z may refer to the output signal (pre-distorted signal) provided by the DPD, the symbol y may refer to the output of the PA, the symbol y' may refer to the observed received signal indicating the output of the PA, and the symbol c may refer to the DPD coefficients for combining basis functions associated with the characteristics or nonlinearity of the PA. Furthermore, the input signal 102x and the pre-distorted signal 104z may be referred to as the transmission data (TX), and the feedback signal 151y' may be referred to as the observation data (ORx).
[0055] Search for exemplary hardware model architectures Aspects of the present disclosure provide DPD arrangements configured to discover optimal kernels that can be mapped to DPD hardware blocks using NNs and deep learning algorithms (e.g., DNAS). Such DPD arrangements may be particularly suitable for LUT-based DPD actuators designed for GMP models (e.g., as shown in equation (1)). In some embodiments, the LUT-based DPD actuator may include a multiplexer that selects one signal from among several input signals (e.g., for memory selection). In some embodiments, the LUT-based DPD actuator may include a LUT (e.g., as shown in equation (2)) configured to take one signal as input and produce an output according to the input, as will be considered more fully below with reference to Figures 3-6.
[0056] In contrast to conventional NAS, where optimization is performed on a differentiable superset of candidate neural network architectures and backpropagation and gradient descent search are used to select one neural network architecture for the given problem, aspects of the present disclosure model hardware behaviorally using differentiable building blocks and perform DNAS using a dataset collected on target hardware. The hardware may be a circuit with defined inputs / outputs and control signals, or a subsystem having multiple circuit blocks connected in a defined manner. Optimization may be performed not only at the functional level (i.e., input / output correspondence) but also on the sequence of operations performed on the subsystem. In one embodiment, an implementation for using DNAS may include a DPD actuator, a transmit signal path, and a DPD adaptive engine operating on a microprocessor.
[0057] Figures 2A and 2B are discussed in relation to Figures 1A-1C to illustrate a model architecture retrieval mechanism applied to DPD hardware. Figure 2A provides an illustrative diagram of scheme 200 for offline training and online adaptation and operation of an indirect learning architecture-based DPD (e.g., DPD180) according to several embodiments of the present disclosure. Scheme 200 includes offline training shown on the left side of Figure 2A and online adaptation and operation of the DPD shown on the right side of Figure 2A.
[0058] In some embodiments, the offline training system (e.g., parameterized model training system 172) may include a transceiver system, a processor, and a memory system. The transceiver system may be substantially analogous to the target system on which DPD operation and adaptation are implemented. For example, the transceiver system may include a PA (e.g., PA 130), a transmission path (input signal 102x may be pre-distorted by DPD actuator circuit 112 and transmitted via PA 130), and an observation path (a feedback signal 151y' indicating the output of PA 130 may be received) substantially similar to the RF transceiver 100 in Figure 1A.
[0059] A processor and memory system (e.g., a computer implementation system such as the data processing system 2300 shown in Figure 15) may be configured to perform multiple captures of the transceiver system's transmission and observation data, indicated as capture 202, which includes the measured and / or signals. In particular, in the case of an indirect learning DPD, capture 202 may include a pre-distorted signal 104z and a feedback signal 151y' captured from the target hardware, and / or a desired pre-distorted signal and / or feedback signal for the corresponding input signal. More specifically, captures may be performed at specific intervals (e.g., every 0.5 seconds, 1 second, 2 seconds or more), and each capture may include L samples of the input signal 102x, M consecutive samples of the pre-distorted signal 104z, and / or N samples of the feedback signal 151y', where L, M, and N may be the same or different.
[0060] The processor and memory system may generate a parameterized model 170 that maps one-to-one with hardware blocks or circuits in the actuator circuit 112 and the DPD adaptive circuit 114. The processor and memory system may further generate the parameterized model 170 based on hardware constraints 204 (e.g., target resource utilization or power consumption) associated with the actuator circuit 112 and the DPD adaptive circuit 114. The processor and memory system may further implement an optimization algorithm that takes in the transmitted and observed captures 202 and optimizes the actuator model parameters and adaptive model parameters for the parameterized model 170.
[0061] After optimization is complete, the processor and memory system can convert the optimized parameterized model 170 (with optimized parameters) into configurations such as actuator configuration 212 and adaptive engine configuration 214, which can be loaded onto firmware for configuring the corresponding hardware for online operation. In some examples, the parameterized model 170 may be trained on a specific type of PA 130 having specific nonlinear characteristics, and therefore the actuator configuration 212 and adaptive engine configuration 214 may include parameters for configuring the DPD actuator and DPD adaptive engine, respectively, to pre-compensate for those nonlinear characteristics. In some examples, the actuator configuration 212 may show information for configuring a LUT for DPD operation, and the adaptive engine configuration 214 may show information associated with a basis function used to adapt the coefficients used by the DPD actuator.
[0062] In some embodiments, the on-chip DPD subsystem for operation and adaptation may include, as shown on the right side of Figure 2A, a DPD actuator circuit 112, a PA 130, a transmission path (an input signal 102x may be pre-distorted by the DPD actuator circuit 112 and transmitted via the PA 130), an observation path (a feedback signal 151y' indicating the output of the PA 130 may be received), a capture buffer 220, and a processor and memory system (e.g., including a processor core 240). The DPD actuator 112 may include a LUT (e.g., equation (2)) and a memory term programmable delay and multiplexer. The processor and memory system may be configured to configure the memory term programmable delay and multiplexer of the DPD actuator 112 according to offline trained parameters (e.g., indicated by actuator configuration 212). The processor and memory system may perform memory term selection and basis function generation (as shown by feature generation 232 in the DPD adaptive circuit 114) according to offline-trained parameters (e.g., as shown by the adaptive engine configuration 214) and data in the capture buffer 220. In particular, in the case of indirect learning DPD, the processor and memory system can capture the pre-distorted signal 104z output by the DPD actuator circuit 112 and the feedback signal 151y' in the capture buffer 220. The processor and memory system may further use the selected memory terms and generated basis functions to obtain a set of linear combination coefficients (as shown by the solver and actuator mapping 230 in the DPD adaptive circuit 114). In some examples, the solver and actuator mapping 230 may use least-squares approximation techniques to obtain the set of linear combination coefficients. The processor and memory system may further generate LUT entries from the obtained coefficients and basis functions according to the offline-trained parameters (e.g., as shown by the adaptive engine configuration 214) and map them to the corresponding memory term LUTs.Furthermore, in some embodiments, the DPD actuator circuit 112 may be implemented by a digital hardware block or circuit, and the DPD adaptive circuit 114 may be implemented by a processor core 240 that executes instruction code (e.g., firmware) that performs feature generation 232, as well as solver and actuator mapping 230.
[0063] Figure 2B provides an illustrative diagram of Scheme 250 for offline training and online DPD adaptation and operation of a direct learning architecture-based DPD (e.g., DPD 190) according to several embodiments of the present disclosure. Scheme 250 in Figure 2B is similar in many respects to Scheme 200 in Figure 2A, and for brevity, the consideration of these elements will not be repeated, as these elements may take any form of the embodiments disclosed herein.
[0064] As mentioned above with reference to Figure 1C, in the case of direct learning DPD, the DPD adaptive circuit 114 can calculate coefficients to minimize the error representing the difference between the input signal 102x and the feedback signal 151y'. Thus, in scheme 250, for offline training on the left side of Figure 2B, an offline processor and memory system (e.g., a computer implementation system such as the data processing system 2300 shown in Figure 15) can perform multiple captures of the input signal 102x and the feedback signal 151y' from the target hardware. That is, capture 202 may include the input signal 102x and the feedback signal 151y' collected from the target hardware, as well as / or a desired feedback signal for the corresponding input signal. Furthermore, as shown on the right side of Figure 2B, an on-chip DPD subsystem for operation and adaptation can capture the input signal 102x and the feedback signal 151y' in the capture buffer 220. Feature generation 232 may be based on the input signal 102x and the feedback signal 151y'. Furthermore, in some examples, the solver and actuator mapping 230 may use an iterative solution approach to determine the set of linear combination coefficients used by the DPD actuator circuit 112.
[0065] Exemplary Offline Parameterized Model Training for Target Hardware Using Model Architecture Search Techniques Therefore, in certain embodiments, as shown in the offline training of Figures 2A and 2B, a computer implementation system (e.g., the parameterized model training system 172 in Figure 1A and / or the data processing system 2300 in Figure 15) may implement an offline training method for performing specific data transformations (e.g., DPD operation and / or adaptation) by performing a model architecture search to optimize the hardware configuration of the target hardware. The data transformations may include linear and / or nonlinear operations. The target hardware (e.g., the on-chip DPD subsystem shown on the right side of Figure 2A) may include a pool of processing units capable of performing at least arithmetic operations and / or signal selection operations (e.g., multiplexing and / or demultiplexing). The model architecture search may be performed on a search space that includes the pool of processing units and their associated capabilities, desired hardware resource constraints (e.g., HW constraints 204), and / or hardware operations associated with the data transformations. The model architecture search can also be performed to achieve specific desired performance metrics associated with the data transformations, for example, minimizing error metrics associated with the data transformations.
[0066] To perform a model architecture search, the computer implementation system may receive information associated with a pool of processing units. The received information may include hardware resource constraints (e.g., HW constraints 204), hardware operations (e.g., signal selection, multiplication, addition, address generation for table lookups, etc.), and / or hardware capabilities associated with the pool of processing units (e.g., speed, delay, etc.). The computer implementation system may further receive a dataset (e.g., capture 202) associated with data transformation operations. The dataset may be collected on target hardware and may include input data, output data, control data, etc. That is, the dataset may include signals measured from the target hardware and / or desired signals. In some examples, the dataset may include, for example, captures of input signal 102x, pre-distorted signal 104z, and / or feedback signal 151y', depending on whether a direct learning DPD architecture or an indirect learning DPD architecture is used. The computer implementation system can use the received hardware information and the received dataset to train a parameterized model associated with the data transformation (e.g., parameterized model 170). Training may involve updating (performing data transformations) at least one parameter of a parameterized model associated with configuring at least a subset of processing units in the pool. The computer implementation system may output one or more configurations (e.g., actuator configuration 212 and adaptive engine configuration 214) for at least a subset of processing units in the pool.
[0067] In some embodiments, the computer implementation system may further generate parameterized models by, for example, generating mappings between different differentiable function blocks from each processing unit in the pool. That is, there is a one-to-one correspondence between each processing unit in the pool and each differentiable function block.
[0068] In some embodiments, the data transformation operation may include a sequence of at least a first and a second data transformation, and training may include computing a first parameter associated with the first data transformation (e.g., a first learnable parameter) and a second parameter associated with the second data transformation (e.g., a second learnable parameter). In some embodiments, computing the first parameter associated with the first data transformation and the second parameter associated with the second data transformation may be further based on backpropagation and loss functions (e.g., by using gradient descent search). In some embodiments, the first or second data transformation in the sequence may be associated with executable instruction code. In other words, the parameterized model can model hardware operation implemented by a processor-executable digital circuit and / or analog circuit and / or instruction code (e.g., firmware).
[0069] In an example modeling DPD operation for offline training, the first data transformation in the sequence may include selecting a memory term from an input signal (e.g., input signal 102x) based on a first parameter. The second data transformation in the sequence may include selecting a set of basis functions (e.g., f) based on a second parameter. k (.)) and selected memory terms may be used to generate feature parameters associated with the nonlinear characteristics of PA130. The sequence associated with the data transformation operation may further include a third data transformation, which involves generating a pre-distorted signal (e.g., a pre-distorted signal 104z) based on the feature parameters.
[0070] In an example modeling DPD adaptation for offline training, the first data transformation in the sequence may include selecting a memory term from a feedback signal (e.g., feedback signal 151y') that represents the output of a nonlinear electronic component or input signal, based on a first parameter. The second data transformation in the sequence may include selecting a set of basis functions (e.g., f) based on a second parameter.k ), DPD coefficient (for example, c k ), and using selected memory terms, may include generating features associated with the nonlinear properties of PA130. The sequence associated with the data transformation operation may further include a third data transformation, which includes updating coefficients based on the features and the second signal. The second signal may correspond to the pre-distorted signal 104z when using an indirect learning DPD (for example, as shown in Figure 1B). Alternatively, the second signal may correspond to the difference between the input signal 102x and the feedback signal 151y' when using a direct learning DPD (for example, as shown in Figure 1C).
[0071] Exemplary online hardware behavior based on parameterized models trained using model architecture search techniques. In certain embodiments, the device may be configured based on a parameterized model (e.g., parameterized model 170) trained for online operation as described herein. For example, the device may include an input node for receiving input signals and a pool of processing units for performing one or more arithmetic operations (e.g., multiplication, addition, etc.) and / or one or more signal selection operations (e.g., multiplexing and / or demultiplexing, address generation, etc.). Each processing unit in the pool may be associated with at least one parameterized model (e.g., NAS model) corresponding to a data transformation (e.g., linear operation, nonlinear operation, DPD operation, etc.). The device may further include control blocks (e.g., control registers) for processing input signals and generating first signals, by configuring and / or selecting at least a first subset of processing units based on a first parameterized model (e.g., parameterized model 170). In some embodiments, a first parameterized model may be trained offline based on a mapping between different sets of differentiable building blocks from each of the processing units in the pool, and at least one of an input dataset or output dataset collected on target hardware or hardware constraints (e.g., target resource utilization / or power consumption). For example, training may be based on a NAS across multiple differentiable building blocks, as considered herein.
[0072] In some embodiments, data conversion may include a sequence of data conversions. For example, a data conversion may include a first data conversion followed by a second data conversion, where the first data conversion converts an input signal to a first signal, and the second data conversion converts the first signal to a second signal. In some embodiments, the sequence of data conversions may be carried out by a combination of a processor and digital hardware blocks (e.g., digital circuits) that execute instruction codes (e.g., software or firmware). For example, a first subset of processing units may include digital hardware blocks (e.g., digital circuits) for carrying out the first conversion, and a control block may be further configured to execute instruction codes for carrying out a second conversion on a second subset of processing units in the pool.
[0073] In a particular embodiment, the device may be a DPD device (e.g., DPD circuit 110) for pre-distorting an input signal to a nonlinear electronic component. For example, the received input signal may correspond to input signal 102x, the nonlinear electronic component may correspond to PA 130, and the first signal may correspond to the pre-distorted signal 104z. The device may further include a memory for storing one or more lookup tables (LUTs) associated with one or more nonlinear characteristics of the nonlinear electronic component based on a first parameterized model. The device may further include a DPD block including a first subset of processing units for selecting a first memory term (e.g., the [ni] and x[nj] terms shown in equation (1)) from the input signal based on a first parameterized model. The first subset of processing units includes one or more LUTs (e.g., the L shown in equation (2)). i、jBased on the selected first memory term (e.g., for DPD operation), a pre-distorted signal may be further generated. In some embodiments, a first subset of processing units may further select a second memory term (e.g., y'[ni] and y'[nj]) from a feedback signal (e.g., feedback signal 151y') associated with the output of a nonlinear electronic component, based on the first parameterized model. The control block executes instruction codes based on the first parameterized model to generate DPD coefficients (e.g., a set of coefficients c) based on the selected second memory term and a set of basis functions. k A second subset of the processing units may be further configured to calculate the coefficients and update at least one of one or more LUTs (for example, for DPD adaptation) based on the calculated set of coefficients and basis functions.
[0074] Exemplary LUT-based DPD actuator implementation As discussed above, in some embodiments, a LUT-based DPD actuator may include a multiplexer that selects one of several input signals. In some embodiments, a LUT-based DPD actuator includes a LUT configured to take one signal as input and produce an output according to the input. Figures 3-5 illustrate various implementations for a LUT-based DPD actuator.
[0075] Figure 3 provides an illustrative diagram of exemplary implementations of a LUT-based DPD actuator circuit 300 according to several embodiments of the present disclosure. For example, the DPD actuator circuits 112 in Figures 1A-1C and 2A-2B may be implemented as shown in Figure 3. As shown in Figure 3, the LUT-based DPD actuator circuit 300 may include a complex / size conversion circuit 310, tap delay lines 312, a plurality of LUTs 320, 322, 324, 326, a complex multiplier 330, and an adder 340. For simplification, Figure 3 illustrates three delay taps 312. However, the LUT-based DPD actuator circuit 300 may be scaled to include any preferred number of delay taps 312 (e.g., 1, 2, 3, 4, 5, 10, 100, 200, 500, 1000 or more).
[0076] The LUT-based DPD actuator circuit 300 may receive an input signal 102x containing, for example, a block of samples x[n], where N can vary from 0 to (N-1), x0, x1, ..., x N-1 It can be expressed as follows. In some cases, the input signal 102x may be a digital baseband complex common-mode orthogonal-phase (IQ) signal. The complex / magnitude conversion circuit 310 can calculate the absolute value or magnitude for each complex sample x[n]. The tap delay line 312 is a delayed version of the magnitude of the input signal 102x, e.g., |x0|, |x1|, ..., |x N-1 | can be generated. LUT320 (for example,
number
number
number
number
number
number
number
number
[0077] Figure 3 illustrates LUTs 320, 322, 324, and 326 as separate LUTs, each corresponding to a specific i,j cross-memory term (for example, modeling a specific nonlinear characteristic of PA130). However, generally, the LUT-based DPD actuator circuit 400 can store LUTs 320, 322, 324, and 326 in any preferred form.
[0078] Figure 4 provides an illustrative diagram of exemplary implementations of the LUT-based DPD actuator circuit 400 according to several embodiments of the present disclosure. For example, the DPD actuator circuits 112 in Figures 1A-1C and 2A-2B can be implemented as shown in Figure 4. The LUT-based DPD actuator circuit 400 in Figure 4 is similar in many respects to the LUT-based DPD actuator circuit 300 in Figure 3, and for brevity, the discussion of these elements will not be repeated, as these elements may take any form of the embodiments disclosed herein.
[0079] In Figure 4, as with the LUT-based DPD actuator circuit 300 in Figure 3, the LUT-based DPD actuator circuit 400 can utilize multiple LUTs to generate a pre-distorted signal sample z[n] for each pre-distorted sample z[n] instead of a single LUT. For simplicity, Figure 4 illustrates a LUT-based DPD actuator circuit 400 that uses two LUTs, LUT A420 and LUT B422, to generate each pre-distorted sample z[n]. However, the LUT-based DPD actuator circuit 400 can be scaled using any suitable number of LUTs (e.g., about 3, 4 or more) to generate each pre-distorted sample z[n]. Furthermore, to avoid disrupting the diagram in Figure 4, Figure 4 illustrates only LUT A420 and LUT B422 for the first two samples, x0 and x1, but LUT A420 and LUT B422 can also be used for delayed samples x2, x3, ..., x N-1 Each of these may be included.
[0080] As shown in Figure 4, for a sample x[n], LUT A420 can take the magnitude of the signal |x[n]| as input and generate an output.
number
number
number
number
[0081] Figure 4 illustrates LUTs 420 and 422 as separate LUTs, each corresponding to a specific i,j cross-memory term. However, generally, the LUT-based DPD actuator circuit 400 can store LUTs 420 and 422 in any preferred form.
[0082] Figure 5 provides an illustrative diagram of exemplary implementations of a LUT-based DPD actuator circuit 500 according to several embodiments of the present disclosure. For example, the DPD actuator circuits 112 in Figures 1A-1C and 2A-2B may be implemented as shown in Figure 5. The LUT-based DPD actuator circuit 500 in Figure 5 is similar in many respects to the LUT-based DPD actuator circuit 300 in Figure 3, and for brevity, the discussion of these elements will not be repeated, as these elements may take any form of the embodiments disclosed herein. As shown in Figure 5, the LUT-based DPD actuator circuit 500 may include a tap delay line 312, a plurality of signal multiplexers 510, a plurality of preprocessing circuits 514 (represented, for example, by a preprocessing function P(.)), a plurality of LUTs 520, a plurality of signal multiplexers 512, a plurality of multipliers 330, and an adder 340. To avoid confusion with the diagram in Figure 5, Figure 5 illustrates only the signal multiplexer 510, preprocessing circuit 514, LUT 520, and signal multiplexer 512 for the first sample x0. However, the signal multiplexer 510, preprocessing circuit 514, LUT 520, and signal multiplexer 512 are used for delayed samples x1, x2, ..., x N-1 Each of these can be arranged in the same way as sample x0.
[0083] As shown in Figure 5, the tap delay line 312 is, for example, x0, x1, ..., x N-1 It generates a delayed version of the input signal 102x, which is represented as . Each multiplexer 510 selects one signal x from all possible inputs based on the selection signal 511. i Each signal multiplexer 512 selects one signal x from all possible inputs based on the selected signal 513. j Select the selected signal x. Each preprocessing circuit 514 processes the selected signal x. i The signal is preprocessed. Preprocessing may involve complex envelope or amplitude calculations, magnitude squaring, scaling functions, or any suitable preprocessing function. Each LUT520 processes the processed signal P(x i ) takes as input and outputs L i、j (P(x iThe output of LUT520 is then multiplied by the respective signals selected by the signal multiplexer 512 in the complex multiplier 330. The products from the outputs of the complex multiplier 330 are summed in the adder 340 to provide the output z[n] to the actuator 300, the output which may correspond to the pre-distorted signal 104z.
[0084] The hardware implementations for the LUT-based DPD actuator shown in Figures 3-5 can be used to drive the model architecture search of the DPD block. In some embodiments, the selection signal 511 for multiplexer 510, the selection signal 513 for multiplexer 512, and LUT 520 can be mapped to learnable parameters trained as part of the model architecture search, as will be discussed in more detail below.
[0085] Mapping hardware blocks to parameterized model elements. According to aspects of this disclosure, a computer implementation system can create a software model of DPD actuator hardware (e.g., a LUT-based DPD actuator shown in Figures 3-5) that captures relevant hardware constraints (e.g., allowed memory terms, LUTs, model size, etc.). The software model may include an adaptive step (e.g., a linear least-squares adaptation for indirectly learned DPD, or an iterative solution for directly learned DPD) within the model that determines a set of DPD coefficients c (e.g., as shown in equation (1) above). In some embodiments, a nonlinear LUT basis function (e.g., f k (.)) can be arbitrary (GMP restricts them to polynomials). For example, a sequence of NN layers may be used. In some embodiments, memory term multiplexing can be modeled using a vector dot product parameterized by weight w.
[0086] The non-linear function can be optimized along with the selection of memory terms in an offline pre-training stage and can be used in the post-deployment operation without adaptation (i.e., any means of changing the pre-trained parameters).
[0087] In some embodiments, a "learnable" multiplexing layer can be used to enable optimization of the selection of memory terms. By doing so, an "N choose M" (M < N) operation using learnable parameters can be implemented.
[0088] In some embodiments, the parameters of the LUT basis function and the "learnable" multiplexing layer can be trained to minimize the final least-squares error. For example, in some embodiments, this can be done using gradient descent along with backpropagation.
[0089] In some examples, generating a software model can include replicating hardware operations with specially designed differentiable building blocks, reproducing a series of hardware events as a differentiable computational graph, and optimizing the hardware configuration offline using hardware functions and constraints.
[0090] FIG. 6 is discussed in relation to FIG. 5 and, for the sake of simplicity, the same reference numbers may be used to refer to the same elements as in FIG. 5. FIG. 6 provides an illustrative diagram of a software model 600 derived from a hardware design having a one-to-one functional mapping, according to some embodiments of the present disclosure. As shown in FIG. 6, the software model 600 may include a first portion 602 that models the LUT operation and a second portion 604 that models the memory term selection operation.
[0091] In the first portion 602, the LUT 520 operation of the DPD hardware takes the magnitude of the input signal as the input |x[i]| and the output L i(|x[i]|) can be generated. The software model 600 can represent LUT520 as an arbitrary function 620 with respect to its input |x[i]|. The operation of LUT520 can be represented as an NN layer in the parameterized model 170, and in some examples it can be trained with another NN, as will be discussed more fully below with reference to Figures 7-9.
[0092] In the second part 604, the operation of the multiplexer 510 on the DPD hardware can select one signal x[ni] from among several input signals x[n], x[n-1], x[n-2], ..., x[nM] based on the selection signal 511. The software model 600 can represent the operation of the multiplexer 510 as a set of weights 624 obtained by multiplying the input signals [n], x[n-1], x[n-2], ..., x[nM] by w[0], w[1], w[2], ..., w[M] in the multiplier 622, and the products are summed in the adder 626 as shown by 610. As further shown, weight w[i]=1 and all other weights w[k≠i]=0, and therefore the output of the adder 626 corresponds to the selected signal x[ni]. The software model 600 can also model the signal selection operation of the multiplexer 512 in Figure 5 using similar operation as shown by 610. In general, the software model 600 can model the operation of the multiplexers 510 and / or 512 in a wide variety of ways to provide the same signal selection function. The operation of the multiplexer 510 (for memory term selection) can be represented as an NN layer in the parameterized model 170, where the weights w[0], w[1], w[M], ..., w[M] can be trained as part of the model architecture search, which is discussed in more detail below with reference to Figures 7-9.
[0093] Automated model detection of exemplary DPD placements Figure 7 provides an illustrative diagram of exemplary method 700 for training a parameterized model for DPD operation, according to several embodiments of the present disclosure. Method 700 may be implemented by a computer implementation system (e.g., parameterized model training system 172 in Figure 1A and / or data processing system 2300 shown in Figure 15). In some embodiments, Method 700 may be implemented as part of offline training shown in Figures 2A and / or 2B. At a high level, Method 700 performs model architecture lookup to optimize the hardware configuration of the DPD hardware (e.g., DPD actuator circuit 112 and DPD adaptive circuit 114 in Figures 1A-1C and 2A-2B) and pre-compensate for nonlinearities of the PA (e.g., PA 130 in Figures 1A-1C and 2A-2B). Method 700 may replicate the hardware operation of the DPD actuator circuit 112 and / or DPD adaptive circuit 114.
[0094] In some embodiments, the computer implementation system may include a memory for storing instructions and one or more computer processors, and when the instructions are executed by one or more computer processors, they cause one or more computer processors to perform the operation of method 700. In other embodiments, the operation of method 700 may be in the form of an encoded instruction in a non-temporary computable-readable storage medium, which, when executed by one or more computer processors of the computer implementation system, causes one or more computer processors to perform method 700.
[0095] In 710, the computer implementation system may receive a capture of the measurement signal and / or the desired signal collected on the target hardware. The target hardware may be analogous to the RF transceiver 100 in Figure 1, the indirect learning DPD 180 in Figure 1B, and / or the direct learning DPD 190 in Figure 1C. The capture may be analogous to capture 202. In particular, the capture may include an input signal 102x, a pre-distorted signal 104z, and / or a feedback signal 151y' captured from the target hardware and / or the corresponding desired signal.
[0096] In 712, the computer implementation system may generate a delayed version of the captured signal (to replicate the tap delay line 312), select memory terms from the captured signal (to replicate the operation of multiplexers 510 and 512), and / or align the captured signal (according to a specific reference sample time).
[0097] In 714, the computer implementation system can perform DPD feature generation. DPD feature generation can perform various basis functions 716 (e.g., f i、j (P(x i This may include applying ))) to generate features (or nonlinear characteristics) associated with PA130. DPD function generation can output all possible features.
[0098] In 718, the computer implementation system may perform feature selection. For example, feature selection may select one or more features from the possible features output by DPD feature generation. Selection may be based on a specific criterion or threshold, for example, if a particular order or combination of nonlinearities exceeds a threshold. In some examples, if the features output by feature generation in 714 indicate the presence of cubic nonlinearity but not quintic nonlinearity, feature selection may select features associated with cubic nonlinearity. Furthermore, a set of basis functions may be generated for cubic nonlinearity. In another example, if the features output by feature generation indicate the presence of both cubic and quintic nonlinearity, feature selection may select features associated with cubic nonlinearity and features associated with quintic nonlinearity. Furthermore, a set of basis functions may be generated for cubic nonlinearity, and another set of basis functions may be generated for quintic nonlinearity. As a further example, if the features output by feature generation show a correlation between cubic and quintic nonlinearity, feature selection may select features associated with cubic and quintic nonlinearity. Furthermore, the set of basis functions can be generated for both cubic and quintic nonlinearity.
[0099] In 720, the computer implementation system can calculate the sum of products and generate a pre-distorted signal 104, as shown in Figure 5, based on the selected features (e.g., memory terms and basis functions) output by feature selection and DPD coefficients.
[0100] In 722, the computer implementation system can determine the mean squared error (MSE) loss (e.g., the difference between the target or desired transmit signal and the pre-distorted signal).
[0101] The computer implementation system may perform backpropagation to adjust feature selection in 718, feature generation in 714, and / or delay / memory term selection in 712, and method 700 may be repeated until the MSE loss meets a specific criterion (e.g., a threshold).
[0102] In some embodiments, a computer implementation system can train the NN732 to generate basis functions such as those shown by 730. To this end, the NN732 takes various memory terms of the input signal 102x and / or the feedback signal 151y' as inputs to the NN732, and the NN732 generates basis functions f i、j (|x i |) can be generated. In this case, the basis function can be any arbitrary function and is not necessarily mathematically expressible as a polynomial.
[0103] Figure 8 provides a schematic illustrative diagram of an exemplary parameterized model 800 that models DPD operation as a sequence of differentiable function blocks, according to several embodiments of the present disclosure. Model 800 may be generated by a computer implementation system (e.g., the parameterized model training system 172 in Figure 1A and / or the data processing system 2300 shown in Figure 15). Model 800 may be similar to the parameterized model 170. In some embodiments, the parameterized model 800 may be generated as part of offline training shown in Figure 2A / or 2B. At a high level, the DPD hardware (e.g., the DPD circuit 110) may include a pool of processing units (e.g., including digital hardware blocks or circuits, analog circuits, ASICs, FPGAs, and / or processors running firmware), and Model 800 may map each of the processing units to a different of several differentiable function blocks. In some embodiments, the computer implementation system may generate the parameterized model 800 using a mechanism substantially similar to the offline training in Figures 2A-2B and / or the method 700 in Figure 7.
[0104] As shown in Figure 8, Model 800 models DPD operation as a feature generator 830 parameterized by learnable parameters or weights θ, and a matrix multiplier 850 running on a digital hardware block 804. For example, digital hardware block 804 may correspond to the digital circuit in the DPD actuator circuit 112. Model 800 further models DPD adaptation as a differentiable function block that replicates the DPD adaptation procedure running on a digital hardware block 806 and firmware 808 operating on a processor. For example, the processor running digital hardware block 806 and firmware 808 may correspond to the digital circuit and processor in the DPD adaptation circuit 114. In some embodiments, digital hardware block 804 and digital hardware block 806 may correspond to the same digital hardware block. In other embodiments, at least one of the digital hardware blocks in digital hardware block 804 is not part of digital hardware block 806.
[0105] As further shown in Figure 8, model 800 may receive a dataset 802. Dataset 802 may be substantially similar to the captures in captures 202 and / or 710 in Figures 2A-2B. For example, dataset 802 may include captures of input signal 102x and / or feedback signal 151y' measured from target hardware (e.g., RF transceiver 100). In DPD operation, model 800 may perform feature generation 830 (parameterized by learnable parameters or weights θ) based on the input signal 102x to output a feature matrix A. Model 800 may perform matrix multiplication 850 between the feature matrix A and a set of coefficients.
number
[0106] Model 800 can model capture operations 810 and preprocessing operations 820 performed on a digital hardware block 806. Preprocessing operation 820 may preprocess the input signal 102x and the feedback signal 151y', respectively, and output preprocessed signals x' and y''. In some embodiments, preprocessing operation 820 may include time-aligning the feedback signal 151y' with the input signal 102x. Preprocessing operation 820 may depend on whether a direct learning DPD or an indirect learning DPD is used. In the case of a direct learning DPD, the output preprocessed signal x' may correspond to the input signal 102x, and the output preprocessed signal y'' may correspond to the difference between the aligned input signal 102x and the feedback signal 151y'. In the case of an indirect learning DPD, the output preprocessed signal x' may correspond to the feedback signal 151y', and the output preprocessed signal y'' may correspond to the input signal 102x. In the case of DPD adaptation, Model 800 may include performing feature generation 840 and solver 860, which are part of the DPD adaptation firmware (e.g., instruction code to be executed on a processor for online operation). As shown in the figure, Model 800 performs feature generation 840 (parameterized by the same learnable parameters or weights θ as feature generation 830 for DPD operation) based on the output and preprocessed signal x' to generate a feature matrix
number
number
number
number
number
number
[0107] Figures 9 and 10 are considered in relation to each other to illustrate a model architecture retrieval procedure performed offline for target DPD hardware and the corresponding online DPD operation on the target DPD hardware. The target hardware may include a pool of processing units capable of performing arithmetic operations and / or signal selection operations (e.g., multiplexing and / or demultiplexing) to perform DPD operation and DPD adaptation. The pool of processing units may include digital circuits, analog circuits, processors, ASICs, FPGAs, etc. In certain embodiments, the target DPD hardware may include digital circuits and at least a processor capable of executing instruction codes.
[0108] Figure 9 is a flowchart illustrating exemplary methods 900 for training a parameterized model of a DPD according to several embodiments of the present disclosure. Method 900 may be implemented by a computer implementation system (e.g., the parameterized model training system 172 in Figure 1A and / or the data processing system 2300 shown in Figure 15). In some embodiments, Method 900 may be implemented as part of offline training shown in Figures 2A and / or 2B. At a high level, Method 900 performs model architecture lookup to optimize the hardware configuration of the DPD hardware (e.g., the DPD actuator circuit 112 and DPD adaptive circuit 114 in Figures 1A-1C and 2A-2B) to pre-compensate for nonlinearities of the PA (e.g., the PA 130 in Figures 1A-1C and 2A-2B). Method 900 can replicate the hardware operation of the DPD actuator circuit 112 and / or the DPD adaptive circuit 114 and implement a model architecture for configuring the actual DPD actuator circuit 112 and / or the DPD adaptive circuit 114 for online operation. Method 900 can utilize mechanisms similar to Method 700 in Figure 7 and Model 800 in Figure 8.
[0109] In some embodiments, the computer implementation system may include a memory for storing instructions and one or more computer processors, and when the instructions are executed by one or more computer processors, they cause one or more computer processors to perform the operation of method 900. In other embodiments, the operation of method 900 may be in the form of an encoded instruction in a non-temporary computable-readable storage medium, which, when executed by one or more computer processors of the computer implementation system, causes one or more computer processors to perform method 900.
[0110] In 910, the computer implementation system may receive an input containing a measurement signal and / or a desired signal collected from the target hardware. The target hardware may be analogous to the RF transceiver 100 in Figure 1, the indirect learning DPD 180 in Figure 1B, and / or the direct learning DPD 190 in Figure 1C. The input may be analogous to the capture 202 and / or data 802. In one example, the input may include an input signal 102x (for input to PA 130) and a capture of an observed received signal or feedback signal 151y' indicating the output of PA 130 and / or a desired signal (e.g., a desired PA input and / or output signal).
[0111] In 912, the computer implementation system can select a programmable delay based on learnable weights w. For example, the programmable delay may correspond to the tap delay line 312 in Figures 3-5, and the learnable weights w may correspond to the weights w used to model signal selection in multiplexers 510 and 512, as shown in Figure 6.
[0112] In 914, the computer implementation system can select memory terms based on learnable weights w. The memory terms can correspond to combinations of i,j cross-memory terms discussed above, with reference to Figures 2A-2B and 3-5.
[0113] In 916, the computer implementation system may generate features A using a basis function with learnable parameters θ, for example, similar to the feature generation 840 in Figure 8. As described above, the feature generation 840 may be implemented by executing firmware or instruction code on the processor during online operation.
[0114] In 918, the computer implementation system is, for example, similar to solver 860 in Figure 8,
number
[0115] In 920, the computer implementation system trains the learnable parameters w and θ,
number
[0116] In one embodiment, the operations in 910, 912, 914, 916, 918, and 920 of method 900 can be viewed as a sequence of data transformations 902, 903, 904, 905, and 906 that can be mapped to a sequence of NN layers. Thus, the computer implementation system may further perform a backpropagation 922 from 920 back to 912 to adjust or update the learnable parameters w and θ. As part of the backpropagation 922, the computer implementation system may update the learnable parameters w for memory term selection. For example, if the gradient with respect to parameter w is oriented in a particular direction, the backpropagation 922 may optimize parameter w in that direction. Similarly, if the gradient with respect to parameter θ is oriented in a particular direction, the backpropagation 922 may optimize parameter θ in that direction. After this backpropagation, method 900 may be repeated as needed, followed by another backpropagation 922. In general, this process may continue until the error in 920 meets a certain criterion that the learnable parameters w and θ are considered trained. The trained parameterized model is then ready to be used for inference (e.g., to construct DPD hardware). In other words, Method 900 can train a parameterized model (e.g., parameterized model 170) represented by parameters w and θ by replicating the operation of DPD hardware, and use the trained parameterized model (e.g., trained parameters w and θ) to construct actual DPD hardware circuitry for online adaptation and operation.
[0117] Figure 10 provides a flowchart illustrating exemplary method 1000 for performing DPD operation for online operation and adaptation according to several embodiments of the present disclosure. Method 1000 may be implemented by DPD devices (e.g., DPD circuit 110, indirect learning DPD 180, and / or direct learning DPD 190). The DPD devices may be LUT-based, including, for example, a LUT-based DPD actuator similar to the LUT-based DPD actuator 500 in Figure 5. In some embodiments, Method 1000 may be implemented as part of online adaptation and operation as shown in Figures 2A and / or 2B. As seen below, the operation of Method 1000 corresponds to the operation of Method 900, which is performed on the DPD device and used to train a parameterized model (e.g., learnable parameters w and θ), and the trained parameters w and θ are used directly in the online operation.
[0118] In 1002, the DPD device receives input of the measured and / or desired signal. In some examples, the input may include input signal 102x received from the input node of the DPD device. In some examples, the input may be obtained from the capture buffer of the DPD device (e.g., capture buffer 220). The input may include input signal 102x (for input to PA130) and a capture of the observed received signal or feedback signal 151y' indicating the output and / or desired signal of PA130 (e.g., the desired PA input and / or output signal).
[0119] In 1004, the DPD device may generate memory terms and delayed samples based on trained weights w corresponding to the operation in 914 of method 900, for example. The memory terms may correspond to combinations of i,j cross-memory terms discussed above with reference to Figures 2A-2B and 3-5.
[0120] In the case of DPD operation, in 1012, the DPD device may configure the actuators of the DPD device (e.g., DPD actuator circuit 112) based on selected memory terms. For example, the actuators may be implemented as shown in Figure 5, and the DPD device may configure programmable delays (e.g., delay 312) and multiplexers (e.g., multiplexers 510 and 512) based on selected memory terms. In some cases, the DPD device may also configure the LUT to store combinations of basis functions and coefficients (e.g., as shown in equation (2)).
[0121] In the case of DPD adaptation, in 1006, the DPD device may generate feature A using a basis function having trained parameters θ corresponding to the operation in 916 of method 900, for example. In 1008, the DPD device may, for example, the operation corresponding to 918 of method 900,
number
number
[0122] Exemplary differentiable sequential operation Sequential operation can be modeled as a differentiable computation graph. The same sequence of operation can be reproduced in differentiable function blocks. The mapping of differentiable function blocks of sequential hardware operation discussed above with reference to Figures 6-10 is considered in the context of DPD operation and DPD adaptation, but similar techniques may be applicable to any suitable hardware operation. Several examples of differentiable sequential operation are shown in Figures 11 and 12.
[0123] Figure 11 provides a schematic illustrative diagram of exemplary mapping 1100 of a sequence of hardware blocks to a sequence of differential function blocks in some embodiments of the present disclosure. As shown in Figure 11, target hardware 1102 may perform a quadratic function 1110 on an input signal x, and subsequently perform a finite impulse response (FIR) filter 1112 (indicated as Conv 1D), which is a one-dimensional (1D) convolution. The FIR filter 1112 may include filter coefficients that can be represented by h. These hardware operations can be mapped to a parameterized model 1104 having a one-to-one correspondence between hardware operations and differentiable function blocks. As shown, the sequences of quadratic functions 1110 and FIR 1112 on the hardware are mapped to sequences of differentiable function blocks 1120 and 1122 in the parameterized model 1104, respectively. The gradient descent algorithm is represented by the dotted arrows, and the loss function
number
number
number
number
[0124] Figure 12 provides a schematic illustrative diagram of exemplary mapping 1200 of a sequence of hardware blocks to a sequence of differentiable function blocks according to some embodiments of the present disclosure. As shown in Figure 12, target hardware 1202 can perform pre-compensation 1210 (e.g., DPD) to linearize certain nonlinearities of downstream nonlinear components 1220 (e.g., PA 130). Pre-compensation 1210 may include a sequence of data transformations B_1 1212, B_2 1214, B3_1216, ..., B_N1218 applied to an input signal x. At least some of these data transformations may be configured based on learnable parameters. These hardware operations may be mapped to a parameterized model 1204 having a one-to-one correspondence between hardware operations and differentiable function blocks. As illustrated, the sequence of data transformations B_1 1212, B_2 1214, B3_1216, ..., B_N1218 on the hardware maps to the sequence of differentiable function blocks 1222, 1224, 1226, ..., 1228 in the parameterized model 1204, respectively. As indicated by the dotted arrows, the gradient descent algorithm can be backpropagated to optimize the learnable parameters in the parameterized model 1204. After the parameterized model 1104 is trained, the target hardware 1102 can be configured according to the parameterized model 1104 (e.g., the trained parameters).
[0125] Exemplary methods for training a parameterized model mapped to target hardware and applying the trained parameterized model to the target hardware. Figure 13 provides a flowchart illustrating a method 1300 for training a parameterized model mapped to target hardware, according to several embodiments of the present disclosure. Method 1300 may be implemented by a computer implementation system (e.g., the parameterized model training system 172 in Figure 1A and / or the data processing system 2300 shown in Figure 15). In some embodiments, Method 1300 may be implemented as part of offline training shown in Figures 2A and / or 2B. At a high level, Method 1300 performs a model architecture lookup to optimize the hardware configuration of the DPD hardware (e.g., the DPD actuator circuits 112 and DPD adaptive circuits 114 in Figures 1A-1C and 2A-2B) to pre-compensate for nonlinearities of the PA (e.g., the PA 130 in Figures 1A-1C and 2A-2B). Although the operations are illustrated in Figure 13, each once and in a specific order, the operations may be performed in parallel, rearranged, and / or repeated as desired.
[0126] In some embodiments, the computer implementation system may include a memory for storing instructions and one or more computer processors, and when the instructions are executed by one or more computer processors, they cause one or more computer processors to perform the operation of method 1300. In other embodiments, the operation of method 1300 may be in the form of instructions encoded in a non-temporary computable-readable storage medium, which, when executed by one or more computer processors of the computer implementation system, causes one or more computer processors to perform method 1300.
[0127] In 1302, the computer implementation system may receive information associated with a pool of processing units. The pool of processing units may be target hardware on which a parameterized model (e.g., parameterized model 170) is trained using model architecture search (e.g., DNAS) techniques. The pool of processing units may include digital circuits, analog circuits, processors for executing instruction code (e.g., firmware), ASICs, FPGAs, etc. The pool of processing units may perform one or more arithmetic operations and one or more signal selection operations. This information may include hardware constraints, hardware operations, and hardware capabilities.
[0128] In 1304, the computer implementation system may receive a dataset associated with a data conversion operation (e.g., nonlinear operation, linear operation). The dataset may include an input signal, an output signal corresponding to the input signal measured from the target hardware, and / or a desired signal. In one example, the data conversion operation may be a DPD operation, and the dataset may include a capture of the observed received signal or feedback signal 151y' indicating the input signal 102x (for input to PA130) and the output of PA130 and / or the desired signal (e.g., the desired PA input and / or output signal).
[0129] In 1306, a computer implementation system may train a parameterized model associated with data transformation operations based on information associated with a dataset and a pool of processing units. Training may include updating at least one parameter (e.g., a learnable parameter) of the parameterized model associated with constituting at least a subset of processing units in the pool.
[0130] In 1308, the computer implementation system may, based on training, output one or more configurations for at least a subset of processing units in the pool. For example, one or more configurations may represent information associated with at least several learnable parameters updated from training.
[0131] In some embodiments, method 1300 may further generate parameterized models. Generating these models may include, for example, generating mappings between one of several differentiable function blocks from each of the processing units in the pool, as considered above with reference to Figures 8-12.
[0132] In some embodiments, the data conversion operation may include a sequence of at least a first and a second data conversion, and training in 1306 may include computing a first parameter (e.g., a learnable parameter) associated with the first data conversion and a second parameter (e.g., a learnable parameter) associated with the second data conversion. In some embodiments, computing the first parameter associated with the first data conversion and the second parameter associated with the second data conversion is further based on backpropagation and a loss function. In some embodiments, the first or second data conversion in the sequence is associated with an executable instruction code. In some examples, the first data conversion may be performed by the digital circuitry of the target hardware, and the second data conversion may be implemented in firmware executed by the processor of the target hardware.
[0133] In certain embodiments, the data transformation operation is associated with a DPD (e.g., DPD circuit 110, direct learning DPD 180, and / or DPD 190) for pre-distorting the input signal to a nonlinear electronic component. For example, the input signal may correspond to input signal 102x, and the nonlinear electronic component may correspond to PA 130, as discussed herein. In the first example, the data transformation may correspond to a DPD operation. Thus, the first data transformation in the sequence may include selecting memory terms (e.g., i,j cross-memory terms discussed above) from the input signal based on a first parameter (e.g., a learnable weight w, discussed above with reference to Figure 9). The second data transformation in the sequence may include generating features (e.g., feature matrix A, discussed above with reference to Figures 8-9) associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory terms, and the generation may be based on a second parameter (e.g., a learnable parameter θ, discussed above with reference to Figures 8-9). The sequence associated with the data conversion operation may further include a third data conversion, which involves generating a pre-distorted signal based on the features.
[0134] In the second example, the data transformation may correspond to DPD adaptation. Thus, the first data transformation in the sequence may include selecting a memory term from a feedback signal (e.g., feedback signal 151y') that represents the output of a nonlinear electronic component or input signal, the selection being based on a first parameter (e.g., a learnable weight w, as considered above with reference to Figure 9). The second data transformation in the sequence may include generating a feature (e.g., a feature matrix A, as considered above with reference to Figures 8-9) associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory term, the generation being based on a second parameter (e.g., a learnable parameter θ, as considered above with reference to Figures 8-9). The sequence associated with the data transformation operation may further include a third data transformation, which includes updating a coefficient (e.g., a coefficient c, as considered above with reference to Figures 1A-1C, 2A-2B, 8-9) based on the feature and the second signal. In an example of an indirectly learned DPD, the first data transformation may include selecting memory terms from a feedback signal based on a first parameter, and the third data transformation may include updating coefficients based on a pre-distorted signal (e.g., signal 104z) generated from the DPD, as discussed above with reference to Figure 1B. In an example of a directly learned DPD, the first data transformation may include selecting memory terms from an input signal based on a first parameter, and the third data transformation may include updating coefficients based on the difference between the feedback signal and the input signal, as discussed above with reference to Figure 1C. In some embodiments, as part of training a model parameterized in 1306, the computer implementation system may perform backpropagation to update a second parameter and generate a set of basis functions, as discussed above with reference to, for example, Figures 7 and 9. In some embodiments, as part of outputting one or more configurations in 1308, the computer implementation system may output one or more configurations, further showing at least one of the lookup table (LUT) configurations associated with the selection of memory terms or the set of basis functions.For example, one or more configurations may be similar to the actuator configuration 212 and / or adaptive engine configuration 214, as discussed above with reference to Figures 2A-2B.
[0135] Figure 14 provides a flowchart illustrating a method 1400 for performing an operation on target hardware (e.g., a device) configured based on a parameterized model, according to several embodiments of the present disclosure. In some embodiments, method 1400 may be implemented by a DPD device (e.g., a DPD circuit 110, an indirectly learned DPD 180, and / or a directly learned DPD 190) on which the parameterized model is trained. In some embodiments, method 1400 may be implemented as part of an online adaptation and operation as shown in Figures 2A and / or 2B. Although the operations are illustrated in Figure 14, each once and in a specific order, the operations may be performed in parallel, rearranged, and / or repeated as desired.
[0136] In 1402, the device can receive an input signal.
[0137] In 1404, the device may constitute at least a first subset of processing units in a pool of processing units based on a parameterized model (e.g., parameterized model 170) associated with a data transformation (e.g., nonlinear operation, linear operation, DPD operation, etc.). The pool of processing units may include digital hardware blocks or digital circuits, analog hardware blocks or analog circuits, processors, ASICs, FPGAs, etc. The first subset of processing units may perform one or more signal selections and one or more arithmetic operations.
[0138] In 1406, the apparatus may perform data conversion on an input signal, and performing data conversion may include processing the input signal using a first subset of processing units to generate a first signal.
[0139] In some embodiments, method 1400 may include devices that constitute a second subset of processing units in a pool of processing units based on a parameterized model, and performing data conversion may further include processing a first signal using the second subset of processing units to generate a second signal. That is, data conversion may include a sequence of data conversions. In some embodiments, the first subset of processing units may include digital hardware blocks (e.g., digital circuits), and the second subset of processing units may include one or more processors, and performing data conversion may include processing an input signal using the digital hardware blocks to generate a first signal, and executing instruction codes on one or more processors to process the first signal to generate a second signal. In some embodiments, method 1400 may include devices that constitute a third subset of processing units in a pool of processing units based on a parameterized model, and performing data conversion may further include processing a second signal using the third subset of processing units to generate a third signal. In some embodiments, the first subset of processing units is the same as the third subset of processing units. In other embodiments, at least one processing unit in the first subset of processing units is not in the third subset of processing units.
[0140] In some embodiments, a parameterized model for constructing a first subset of processing units may be trained from each of the processing units in the pool based on mappings between different of multiple differentiable building blocks (as considered above with reference to, for example, Figures 2A-2B, 6, 7-12) and at least one of input datasets collected on target hardware, output datasets collected on target hardware, or hardware constraints. In some embodiments, the parameterized model for constructing the first subset of processing units is further trained based on NAS across multiple differentiable building blocks.
[0141] In certain embodiments, the apparatus may be a DPD apparatus (e.g., DPD circuit 110, indirect learning DPD 180, and / or direct learning DPD 190) for performing DPD operation and DPD adaptation. In the example of DPD operation, the input signal may be associated with the input (e.g., input signal 102x) of a nonlinear electronic component (e.g., PA 130). Processing the input signal to generate a first signal in 1406 may include selecting a first memory term (e.g., i,j cross-memory terms) from the input signal based on a parameterized model (e.g., based on trained weights w of the parameterized model), and generating a pre-distorted signal (e.g., output signal 104z) based on one or more LUTs (e.g., LUTs 320, 322, 324, 326, 420, 422, and / or 520) associated with one or more nonlinear characteristics of the nonlinear electronic component, and the first selected memory term, wherein the first signal may correspond to the pre-distorted signal. In some examples, one or more LUTs may be configured based on a parameterized model (e.g., based on trained parameters θ in the parameterized model). In an example of DPD adaptation using indirect learning (e.g., as shown in Figure 1B), the device may further select a second memory term from feedback signals associated with a nonlinear electronic component based on a parameterized model (e.g., trained weights w of the parameterized model). The device may further configure a second subset of processing units in a pool of processing units to execute instruction code to compute DPD coefficients (e.g., coefficients c) based on the selected second memory term and set of basis functions, based on the parameterized model. The instruction code may also cause the second subset of processing units to update at least one of one or more LUTs based on the computed coefficients and set of basis functions.In an example of DPD adaptation using direct learning (as shown, for example, in Figure 1C), the device may further configure a second subset of processing units in a pool of processing units to execute instruction code to compute DPD coefficients (e.g., coefficient c) based on a selected first set of memory terms and basis functions, based on a parameterized model. The instruction code may also cause the second subset of processing units to update at least one of one or more LUTs based on the computed coefficients and basis functions.
[0142] Figure 15 provides a block diagram illustrating an exemplary data processing system 2300, which may be configured to implement or control at least a portion of a hardware block configuration using a neural network, according to several embodiments of the present disclosure. In one example, the data processing system 2300 may be configured to train a parameterized model (e.g., parameterized model 170) for configuring target hardware using a model architecture retrieval technique (e.g., DNAS), as considered herein. In another example, the data processing system 2300 may be configured to configure DPD hardware based on the configuration provided by the trained parameterized model, as considered herein.
[0143] As shown in Figure 15, the data processing system 2300 may include at least one processor 2302, for example, a hardware processor 2302, coupled to the memory element 2304 via a system bus 2306. Thus, the data processing system can store program code in the memory element 2304. Furthermore, the processor 2302 can execute program code accessed from the memory element 2304 via the system bus 2306. In one embodiment, the data processing system may be implemented as a computer suitable for storing and / or executing program code. However, it should be understood that the data processing system 2300 may be implemented in the form of any system including a processor and memory capable of performing the functions described in this disclosure.
[0144] In some embodiments, the processor 2302 can execute software or algorithms to perform activities such as those considered in this disclosure, particularly activities related to performing DPD using a neural network as described herein. The processor 2302 may include, in non-limiting examples, any combination of hardware, software, or firmware that provides programmable logic, including a microprocessor, a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA), an application-specific integrated circuit (IC) (ASIC), or a virtual machine processor. The processor 2302 may be communicatively coupled to a memory element 2304 in a direct memory access (DMA) configuration, for example, so that the processor 2302 can read from or write to the memory element 2304.
[0145] In general, the memory element 2304 may include any suitable volatile or non-volatile memory technology, including double data-rate (DDR) random access memory (RAM), synchronous RAM (SRAM), dynamic RAM (DRAM), flash, read-only memory (ROM), optical media, virtual memory area, magnetic or tape memory, or any other suitable technology. Unless otherwise specified, any memory element considered herein should be interpreted as being encompassed within the broad term "memory." Information measured, processed, tracked, or transmitted to any component of the data processing system 2300 may be provided to any database, register, control list, cache, or storage structure, all of which may be referenced within any suitable time frame. Any such storage option, as used herein, may be encompassed within the broad term "memory." Similarly, any potential processing elements, modules, and machines described herein should be interpreted as being encompassed within the broad term "processor." Each of the elements shown in this figure, for example, any element illustrating a DPD configuration for performing DPD using a neural network as shown in Figures 1-13, may also include a suitable interface for receiving, transmitting, and / or otherwise communicating data or information within a network environment, and as a result, they can communicate with, for example, a data processing system 2300.
[0146] In certain exemplary implementations, a mechanism for performing DPD using a neural network as outlined herein may be implemented by logic encoded on one or more tangible media, which may include non-temporary media such as embedded logic provided to an ASIC, DSP instructions, software executed by a processor (potentially including object code and source code), or other similar machines. In some of these examples, a memory element, such as memory element 2304 shown in Figure 15, can store data or information used for the operations described herein. This includes a memory element that can store software, logic, code, or processor instructions executed to perform the activities described herein. The processor can execute any type of instruction associated with the data or information to realize the operations detailed herein. In one example, a processor, such as processor 2302 shown in Figure 15, can transform an element or article (e.g., data) from one state or thing to another. In another example, the activities outlined herein may be implemented in fixed logic or programmable logic (e.g., software / computer instructions executed by a processor), and the elements identified herein may be several types of programmable processors, programmable digital logic (e.g., FPGAs, DSPs, erasable programmable read-only memory (EPROMs), electrically erasable programmable read-only memory (EEPROMs)), or ASICs including digital logic, software, code, electronic instructions, or any preferred combination thereof.
[0147] The memory element 2304 may include, for example, one or more physical memory devices such as local memory 2308 and one or more bulk storage devices 2310. Local memory may refer to RAM or other non-persistent memory devices commonly used during the actual execution of program code. Bulk storage devices may be implemented as hard drives or other persistent data storage devices. The processing system 2300 may also include one or more cache memories (not shown) that provide temporary storage for at least some program code to reduce the number of times program code needs to be retrieved from the bulk storage device 2310 during execution.
[0148] As shown in Figure 15, the memory element 2304 can store the application 2318. In various embodiments, the application 2318 may be stored in local memory 2308, one or more bulk storage devices 2310, or separately from the local memory bulk storage devices. It should be understood that the data processing system 2300 may further run an operating system (not shown in Figure 15) that can facilitate the execution of the application 2318. The application 2318, implemented in the form of executable program code, may be executed by the data processing system 2300, for example, by the processor 2302. In response to the execution of the application, the data processing system 2300 may be configured to perform one or more operational or method steps described herein.
[0149] The input / output (I / O) devices, designated as input device 2312 and output device 2314, may optionally be coupled to a data processing system. Examples of input devices include, but are not limited to, keyboards, pointing devices such as mice, etc. Examples of output devices include, but are not limited to, monitors or displays, speakers, etc. In some embodiments, output device 2314 may be any type of screen display, such as a plasma display, liquid crystal display (LCD), organic light-emitting diode (OLED) display, electroluminescent (EL) display, or any other indicator such as a dial, barometer, or LED. In some implementations, the system may include a driver (not shown) for output device 2314. The input and / or output devices 2312, 2314 may be coupled to a data processing system directly or via an intermediary I / O controller.
[0150] In one embodiment, the input device and the output device may be implemented as a combined input / output device (illustrated in Figure 15 by dashed lines surrounding input device 2312 and output device 2314). An example of such a combined device is a touch-sensitive display, sometimes referred to as a “touchscreen display” or simply a “touchscreen.” In such an embodiment, input to the device may be provided by the movement of a physical object, such as a stylus or a user’s finger, on or near the touchscreen display.
[0151] The network adapter 2316 may also optionally be coupled to the data processing system, enabling it to be coupled to other systems, computer systems, remote network devices, and / or remote storage devices via an intervening private or public network. The network adapter may comprise a data receiver for receiving data transmitted to the data processing system 2300 by such systems, devices, and / or networks, and a data transmitter for transmitting data from the data processing system 2300 to such systems, devices, and / or networks. Modems, cable modems, and Ethernet cards are examples of different types of network adapters that may be used in the data processing system 2300. [Examples]
[0152] Embodiment 1 includes an input node for receiving an input signal, and a pool of processing units for performing one or more arithmetic operations and one or more signal selection operations, each of which is associated with at least one parameterized model corresponding to a data transformation operation, and a control block for configuring a first subset of the processing units in the pool based on a first parameterized model, the first subset of processing units including a device that processes an input signal to generate a first signal.
[0153] In Example 2, the apparatus of Example 1 may optionally include a case in which a first subset of processing units performs at least one first signal selection operation from among one or more signal selection operations.
[0154] In Example 3, the apparatus described in Example 1 or 2 may optionally include a data conversion operation that includes a linear data conversion operation.
[0155] In Example 4, the apparatus described in any one of Examples 1 to 3 may optionally include a case where the data conversion operation includes a nonlinear data conversion operation.
[0156] In Example 5, the apparatus described in any one of Examples 1 to 4 may optionally include a first subset of processing units comprising digital hardware blocks for processing input signals to generate a first signal, and a control block further configuring a second subset of processing units to execute instruction code for processing the first signal to generate a second signal, wherein processing the first signal is associated with a first parameterized model.
[0157] In Example 6, the apparatus described in any one of Examples 1 to 5 may optionally include a third subset of processing units that processes a second signal to generate a third signal, and the third subset of processing units is configured based on a first parameterized model.
[0158] In Example 7, the apparatus described in any one of Examples 1 to 6 may optionally include a case where the first subset of processing units is the same as the third subset of processing units.
[0159] In Example 8, the apparatus described in any one of Examples 1 to 6 may optionally include a case where at least one processing unit in the first subset of processing units is not in the third subset of processing units.
[0160] In Example 9, the apparatus described in any one of Examples 1 to 8 may optionally include a first subset of the processing unit comprising a set of digital hardware blocks for processing an input signal to generate a first signal, and a control block further comprising a second subset of the processing unit comprising another set of digital hardware blocks for processing the first signal to generate a second signal.
[0161] In Example 10, the apparatus described in any one of Examples 1 to 9 may optionally include training a first parameterized model for configuring a first subset of processing units based on a mapping between different sets of differentiable building blocks and at least one of an input dataset collected on target hardware, an output dataset collected on target hardware, or hardware constraints, from each of the processing units in the pool.
[0162] In Example 11, the apparatus described in any one of Examples 1 to 10 may optionally include a case where a first parameterized model for configuring a first subset of processing units is further trained based on neural architecture search across multiple differentiable building blocks.
[0163] In Example 12, the apparatus described in any one of Examples 1 to 11 optionally includes a memory for storing one or more lookup tables (LUTs) associated with one or more nonlinear characteristics of the nonlinear electronic component based on a first parameterized model, and a digital pre-distortion (DPD) block including a first subset of processing units for selecting a first memory term from the input signal based on the first parameterized model, and generating a pre-distorted signal based on one or more LUTs and the selected first memory term, wherein the first signal may correspond to a pre-distorted signal.
[0164] In Example 13, the apparatus described in any one of Examples 1 to 12 may optionally be further configured to further configure a first subset of processing units to further select a second memory term from feedback signals associated with the output of a nonlinear electronic component based on a first parameterized model, and a control block to execute instruction code to perform the following actions: calculate DPD coefficients based on the selected second memory term and a set of basis functions based on the first parameterized model, and update at least one of one or more LUTs based on the calculated coefficients.
[0165] In Example 14, the apparatus described in any one of Examples 1 to 13 may optionally be further configured to execute instruction code to calculate DPD coefficients based on a selected first set of memory terms and basis functions, and to update at least one of one or more LUTs based on the calculated coefficients.
[0166] Example 15 is an apparatus for applying digital pre-distortion (DPD) to an input signal of a nonlinear electronic component, the apparatus comprising a pool of processing units associated with a parameterized model, and a component for selecting at least a subset of the processing units in the pool, a second subset of processing units, the first subset of processing units converting the input signal into a pre-distorted signal based on the parameterized model and DPD coefficients, and the second subset of processing units updating the DPD coefficients at least in part on a feedback signal representing the output of the nonlinear electronic component.
[0167] In Example 16, the apparatus described in Example 15 may optionally include a case in which a first subset of processing units converts an input signal into a pre-distorted signal by generating a first memory term from an input signal based on a parameterized model, and generating a pre-distorted signal based on the first memory term, a set of basis functions, and DPD coefficients.
[0168] In Example 17, the apparatus described in Example 15 or 16 may optionally include a first subset of processing units that further generates a second memory term from a feedback signal or input signal based on a parameterized model, and a second subset of processing units that further updates a set of coefficients based on the second memory term and a set of basis functions.
[0169] In Example 18, the apparatus described in any one of Examples 15 to 17 may optionally include a case in which a second subset of processing units updates a set of coefficients based on the input signal.
[0170] In Example 19, the apparatus described in any one of Examples 15 to 17 may optionally include a case in which a second subset of the processing units updates the set of coefficients based on an error representing the difference between the feedback signal and the input signal.
[0171] In Example 20, the apparatus described in any one of Examples 15 to 19 may optionally include a memory for capturing a feedback signal and at least one of an input signal or a pre-distorted signal, and a first subset of the processing unit may generate a second memory term based on the alignment between the feedback signal and at least one of the input signal or the pre-distorted signal.
[0172] In Example 21, the apparatus described in any one of Examples 15 to 20 may optionally include a first subset of processing units comprising one or more digital hardware blocks for converting input signals into pre-distorted signals, and a second subset of processing units comprising at least a processor for executing instruction codes for updating coefficients.
[0173] In Example 22, the apparatus described in any one of Examples 15 to 21 may optionally include a case where the parameterized model includes multiple differentiable function blocks having a one-to-one correspondence to processing units in a pool, and the parameterized model is trained using gradient descent search.
[0174] Embodiment 23 includes a method comprising: receiving an input signal and configuring at least a first subset of processing units in a pool of processing units based on a parameterized model associated with data conversion, wherein the first subset of processing units performs one or more signal selections and one or more arithmetic operations; and performing data conversion on the input signal, wherein performing data conversion includes processing the input signal using the first subset of processing units to generate a first signal.
[0175] In Example 24, the method described in Example 23 optionally includes configuring a second subset of processing units in a pool of processing units based on a parameterized model, and performing data transformation may further include processing a first signal using the second subset of processing units to generate a second signal.
[0176] In Example 25, the method described in Example 23 or 24 may optionally include a first subset of the processing unit comprising digital hardware blocks, a second subset of the processing unit comprising one or more processors, and the data conversion being performed may include processing an input signal using the digital hardware blocks to generate a first signal, and executing instruction code on one or more processors to process the first signal and generate a second signal.
[0177] In Example 26, the method according to any one of Examples 23-25 may optionally include training a parameterized model to constitute a first subset of processing units based on a mapping between different multiple differentiable building blocks and at least one of the following: an input dataset collected on target hardware, an output dataset collected on target hardware, or hardware constraints, from each of the processing units in the pool.
[0178] In Example 27, the method according to any one of Examples 23 to 26 may optionally include: an input signal associated with the input of a nonlinear electronic component; processing the input signal to generate a first signal; selecting a first memory term from the input signal based on a parameterized model; and generating a pre-distorted signal based on one or more lookup tables (LUTs) associated with one or more nonlinear characteristics of the nonlinear electronic component and the first selected memory term, wherein the first signal may correspond to the pre-distorted signal.
[0179] Example 28 may optionally include configuring the method according to any one of Examples 23 to 27 to execute instruction code to: select a second memory term from feedback signals associated with a nonlinear electronic component based on a parameterized model; calculate DPD coefficients for a second subset of processing units in a pool of processing units based on the selected second memory term and a set of basis functions based on the parameterized model; and update at least one of one or more LUTs based on the calculated coefficients.
[0180] In Example 29, the method according to any one of Examples 23 to 27 may optionally be configured to execute instruction code to perform the following actions based on a parameterized model: to calculate DPD coefficients for a second subset of processing units in a pool of processing units based on a selected first set of memory terms and basis functions; and to update at least one of one or more LUTs based on the calculated coefficients.
[0181] Example 30 includes a method comprising: receiving information associated with a pool of processing units by a computer implementation system; receiving a dataset associated with a data transformation operation by a computer implementation system; training a parameterized model associated with a data transformation operation based on the information associated with the pool of processing units of the dataset, wherein the training includes updating at least one parameter of the parameterized model associated with constituting at least a subset of processing units in the pool; and outputting one or more configurations for at least a subset of processing units in the pool based on the training.
[0182] In Example 31, the method described in Example 30 may optionally include a case in which a pool of processing units performs one or more arithmetic operations and one or more signal selection operations.
[0183] In Example 32, the method according to Example 30 or 31 may optionally include generating a parameterized model, wherein generating includes generating a mapping between one of a plurality of differentiable function blocks from each of the processing units in the pool.
[0184] In Example 33, the method according to any one of Examples 30-32 may optionally include cases where training the parameterized model is further based on hardware resource constraints indicated by information associated with a pool of processing units.
[0185] In Example 34, the method according to any one of Examples 30 to 33 may optionally include a data transformation operation comprising at least a sequence of first and second data transformations, and training comprising calculating a first parameter associated with the first data transformation and a second parameter associated with the second data transformation.
[0186] In Example 35, the method according to any one of Examples 30-34 may optionally include cases where the calculation of the first parameter associated with the first data transformation and the second parameter associated with the second data transformation is further based on backpropagation and loss functions.
[0187] In Example 36, the method according to any one of Examples 30 to 35 may optionally include a case where the first or second data transformation in the sequence is associated with an executable instruction code.
[0188] In Example 37, the method according to any one of Examples 30 to 36 may optionally include a data conversion operation associated with digital pre-distortion (DPD) for pre-distorting an input signal to a nonlinear electronic component, wherein a first data conversion in the sequence includes selecting memory terms from the input signal based on a first parameter, a second data conversion in the sequence includes generating features associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory terms based on a second parameter, and a third data conversion in the sequence associated with the data conversion operation includes generating a pre-distorted signal based on the features.
[0189] In Example 38, the method according to any one of Examples 30 to 37 may optionally include a data conversion operation associated with a digital pre-distortion (DPD) for pre-distorting an input signal to a nonlinear electronic component, wherein a first data conversion in the sequence includes selecting a memory term from a feedback signal indicating the output of a nonlinear electronic component or input signal based on a first parameter, a second data conversion in the sequence includes generating a feature associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory term based on a second parameter, and a third data conversion in the sequence associated with the data conversion operation includes updating coefficients based on the feature and the second signal.
[0190] In Example 39, the method according to any one of Examples 30 to 38 may optionally include a first data transformation comprising selecting a memory term from a feedback signal based on a first parameter, and a third data transformation comprising updating a coefficient based on a pre-distorted signal generated from the DPD.
[0191] In Example 40, the method according to any one of Examples 30 to 39 may optionally include a first data transformation comprising selecting a memory term from an input signal based on a first parameter, and a third data transformation comprising updating a coefficient based on the difference between a feedback signal and an input signal.
[0192] In Example 41, the method described in any one of Examples 30-40 may optionally include training a parameterized model by performing backpropagation to update the first parameter for the selection of memory terms.
[0193] In Example 42, the method described in any one of Examples 30 to 41 may optionally further include training a parameterized model by performing backpropagation to update a second parameter and generate a set of basis functions.
[0194] In Example 43, the method according to any one of Examples 30 to 42 may optionally include outputting one or more configurations that further indicate at least one of the lookup table (LUT) configurations associated with a selection of memory terms or a set of basis functions.
[0195] Embodiment 44 is a computer implementation system comprising a memory containing instructions and one or more computer processors, wherein when an instruction is executed by one or more computer processors, the computer implementation system causes one or more computer processors to perform an operation comprising: receiving information associated with a pool of processing units, the pool of processing units performing one or more arithmetic operations and one or more signal selections; receiving a dataset associated with a data transformation; and training a parameterized model associated with a data transformation based on the dataset and the information associated with the pool of processing units, the training of the parameterized model comprising updating at least one parameter of the parameterized model associated with constituting at least a subset of processing units in the pool; and outputting one or more configurations for at least a subset of processing units in the pool based on the training.
[0196] In Example 45, the computer implementation described in Example 44 may optionally further include the operation generating a parameterized model by generating a mapping between one of several differentiable function blocks from each of the processing units in the pool.
[0197] In Example 46, the computer implementation described in Example 44 or 45 may optionally include a case where the data transformation includes a sequence of at least a first data transformation and a second data transformation, and training the parameterized model includes calculating a first parameter associated with the first data transformation and a second parameter associated with the second data transformation based on backpropagation and a loss function.
[0198] In Example 47, the computer implementation described in any one of Examples 44 to 46 may optionally include a subset of processing units comprising one or more digital hardware blocks associated with a first data conversion and one or more processors for executing instruction codes associated with a second data conversion.
[0199] In Example 48, the computer implementation described in any one of Examples 44 to 47 may optionally include a data conversion operation associated with at least one of digital pre-distortion (DPD) operation or DPD adaptation for pre-distorting input signals to a nonlinear electronic component, and outputting one or more configurations, which may include outputting at least one of the DPD operation configuration or DPD adaptation configuration.
[0200] Embodiment 49 includes a non-temporary computable-readable storage medium which includes instructions causing one or more computer processors to perform operations including receiving information associated with a pool of processing units, the pool of processing units performing one or more arithmetic operations and one or more signal selections, generating a mapping between one of a plurality of differentiable function blocks from each of the processing units in the pool, receiving a dataset associated with a data transformation, and training a parameterized model to configure at least a subset of the processing units in the pool to perform the data transformation, the training being based on the dataset, the information associated with the pool of processing units, and the mapping, and including updating at least one parameter of the parameterized model associated with configuring at least a subset of the processing units in the pool, and outputting one or more configurations for at least a subset of the processing units in the pool based on the training.
[0201] In Example 50, the non-temporary computable-readable storage medium described in Example 49 may optionally include a data transformation that includes at least a sequence of first and second data transformations, and training that further includes updating a first parameter associated with the first data transformation and a second parameter associated with the second data transformation based on backpropagation and a loss function.
[0202] In Example 51, the non-temporary computable-readable storage medium described in Example 49 or 50 may optionally include a data conversion operation associated with digital pre-distortion (DPD) for pre-distorting an input signal to a nonlinear electronic component, wherein a first data conversion in the sequence includes selecting memory terms from the input signal based on a first parameter, a second data conversion in the sequence includes generating features associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory terms based on a second parameter, and the sequence further includes a third data conversion that generates a pre-distorted signal based on the features.
[0203] In Example 52, the non-temporary computable-readable storage medium described in any one of Examples 49 to 51 optionally includes a data transformation associated with a digital pre-distortion (DPD) for pre-distorting an input signal to a nonlinear electronic component, wherein a first data transformation in the sequence includes selecting a memory term from a feedback signal indicating the output of a nonlinear electronic component or input signal based on a first parameter, a second data transformation in the sequence includes generating a feature associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory term based on a second parameter, and the sequence further includes a third data transformation including updating coefficients based on the feature and the second signal, wherein the first and second data transformations are performed by a subset of processing units, and the third data transformation is performed by executing instruction code on at least one other processing unit in the pool.
[0204] Modified form and implementation method Various embodiments of implementing a DPD based on a model trained using a NAS are described herein with reference to the fact that the “input signal for PA” is the drive signal for PA, i.e., the DPD arrangement is a signal generated based on the input signal x described herein, applying pre-distortion based on the DPD coefficients. However, in other embodiments of a DPD based on a model trained using a NAS, the “input signal for PA” may be a bias signal used to bias the PA. Thus, embodiments of the present disclosure also cover DPD arrangements based on a model trained using a NAS similar to those described herein and illustrated in the drawings, but instead of modifying the drive signal for PA, the DPD arrangement may be configured to modify the bias signal for PA, which may be done based on a control signal generated by a DPD adaptive circuit (e.g., the DPD adaptive circuit described herein), and the output of the PA is based on the bias signal used to bias the PA. In other embodiments of the present disclosure, both the drive signal and the bias signal for PA can be tuned as described herein to implement a DPD using a neural network.
[0205] While some of the description is provided herein with reference to PA, generally, various embodiments of DPDs configured based on models using NAS presented herein are applicable to amplifiers other than PA, such as low-noise amplifiers and variable-gain amplifiers, as well as nonlinear electronic components of RF transceivers other than amplifiers (i.e., components that may exhibit nonlinear behavior). Furthermore, while some of the description is provided herein with reference to millimeter-wave / 5G technology, generally, various embodiments of DPDs using neural networks presented herein are applicable to any technology or standard wireless communication system other than millimeter-wave / 5G, any wireless RF system other than wireless communication systems, and / or RF systems other than wireless RF systems.
[0206] While embodiments of the present disclosure have been described above with reference to exemplary implementations shown in Figures 1A-1C, 2A-2B, and 3-15, those skilled in the art will recognize that the various teachings described above are applicable to a wide variety of other implementations.
[0207] In certain contexts, the features considered herein may be applicable to automotive systems, safety-critical industrial applications, medical systems, scientific instrumentation, wireless and wired communications, radio, radar, industrial process control, audio-video equipment, current sensing, instrumentation (which may be high-precision), and other digital processing-based systems.
[0208] In the consideration of the embodiments described above, system components such as multiplexers, multipliers, adders, delay taps, filters, converters, mixers, and / or other components can be readily replaced, substituted, or otherwise modified to meet specific circuit needs. Furthermore, it should be noted that the use of complementary electronic devices, hardware, software, etc., provides equally viable options for implementing the teachings of this disclosure relating to the application of model architecture lookup to hardware configurations in various communication systems.
[0209] As proposed herein, various parts of a system for using model architecture retrieval techniques for hardware configurations may include electronic circuits for performing the functions described herein. In some examples, one or more parts of the system may be provided by a processor specifically configured to perform the functions described herein. For example, the processor may include one or more application-specific components or programmable logic gates configured to perform the functions described herein. The circuits may operate in the analog domain, the digital domain, or the mixed-signal domain. In some examples, the processor may be configured to perform the functions described herein by executing one or more instructions stored in a non-temporary computer-readable storage medium.
[0210] In one exemplary embodiment, any number of electrical circuits in the figures may be mounted on a substrate of the associated electronic device. The substrate may be a general circuit board capable of holding various components of the internal electronic system of the electronic device and further providing connectors to other peripherals. More specifically, the substrate may provide electrical connections that enable other components of the system to communicate electrically. Any suitable processor (including DSPs, microprocessors, supporting chipsets, etc.), computer-readable non-temporary memory elements, etc., can be suitably coupled to the substrate based on specific configuration needs, processing requirements, computer design, etc. Other components such as external storage, additional sensors, audio / video display controllers, and peripheral devices may be attached as plug-in cards, via cables, or mounted on the substrate itself. In various embodiments, the functions described herein may be implemented in emulation form as software or firmware running within one or more configurable (e.g., programmable) elements arranged in a structure that supports these functions. The software or firmware providing the emulation may be provided on a non-temporary computer-readable storage medium containing instructions that enable the processor to perform those functions.
[0211] In another exemplary embodiment, the electrical circuits in this figure may be implemented as standalone modules (e.g., devices having associated components and circuits configured to perform a particular application or function) or as plug-in modules into application-specific hardware of an electronic device. Note that certain embodiments of this disclosure may readily be contained in a system-on-a-chip (SOC) package, either in part or in whole. An SOC represents an IC that integrates components of a computer or other electronic system onto a single chip. It may include digital, analog, mixed-signal, and often RF functions, all of which may be provided on a single chip substrate. Other embodiments may include a multi-chip module (MCM) having multiple distinct ICs arranged within a single electronic package and configured to interact closely with each other through the electronic package.
[0212] Furthermore, it is essential to note that all specifications, dimensions, and relationships outlined herein (for example, the number of components of the apparatus and / or RF transceiver shown in Figures 1A–1C, 2A–2B, 3–5, and 15) are provided for illustrative and teaching purposes only. Such information may vary considerably without departing from the spirit of this disclosure or the scope of the appended claims. It should be understood that the system can be integrated in any preferred manner. Any of the illustrated circuits, components, modules, and elements may be combined in various possible configurations in accordance with alternatives of similar designs, all of which are clearly within the broad scope of this specification. In the foregoing description, exemplary embodiments have been described with reference to specific processor and / or component arrangements. Various modifications and changes may be made to such embodiments without departing from the scope of the appended claims. Therefore, the description and drawings should be considered illustrative, not restrictive.
[0213] It should be noted that in the numerous examples provided herein, interactions may be described in terms of two, three, four, or more electrical components. However, it should be understood that the systems, which are presented for clarity and illustrative purposes only, can be integrated in any preferred manner. In line with alternatives to similar designs, any of the components, modules, and elements illustrated in the drawings can be combined in a variety of possible configurations, all of which are clearly within the broad scope of this specification. In particular cases, it may be easier to describe one or more functions of a given set of flows by referring to only a limited number of electrical elements. It should be understood that the electrical circuits in the drawings and their teachings are readily scalable and can accommodate a large number of components as well as more complex / sophisticated arrangements and configurations. Therefore, the examples provided should not limit the scope of electrical circuits or hinder the broad teaching of electrical circuits, so that they may be potentially applicable to countless other architectures.
[0214] In this specification, references to various features (e.g., elements, structures, modules, components, steps, operations, characteristics, etc.) included in “one embodiment,” “exemplary embodiment,” “embodiment,” “another embodiment,” “several embodiments,” “various embodiments,” “other embodiments,” and “alternative embodiments” are intended to mean that such features are included in one or more embodiments of this disclosure, but may or may not be combined in the same embodiment. Also, as used herein, including in claims, “or” in lists of items (e.g., lists of items beginning with a phrase such as “at least one of” or “one or more of”) means a comprehensive list, such as the list [at least one of A, B, or C] meaning A or B or C or AB or AC or BC or ABC (i.e., A and B and C).
[0215] Various aspects of the illustrative embodiments are described using terminology commonly used by those skilled in the art to convey the nature of their work to others skilled in the art. For example, the term “connected” means a direct electrical connection between things that are connected without any intermediate devices / components, and the term “coupled” means either a direct electrical connection between things that are connected, or an indirect connection via one or more passive or active intermediate devices / components. In another example, the term “circuit” means one or more passive and / or active components arranged to cooperate with each other to provide a desired function. Also, as used herein, terms such as “substantially,” “approximately,” and “about” may be used generally to mean within + / - 20% of a target value, for example, within + / - 10% of a target value, based on the context of specific values described herein or known in the art.
[0216] Many other changes, substitutions, modifications, alterations, and modifications may be apparent to those skilled in the art, and this disclosure is intended to encompass all such changes, substitutions, modifications, alterations, and modifications so as to fall within the scope of the examples and the appended claims. All optional features of the apparatus described above may also be implemented in relation to the methods or processes described herein, and it should be noted that the details of the examples may be used in any one or more embodiments.
Claims
1. It is a method, A computer implementation system receives information about hardware resource constraints associated with a pool of processing units, wherein multiple processing units in the pool perform one or more arithmetic operations and one or more signal selections. The computer implementation system receives a dataset associated with a data conversion operation, the data conversion operation being associated with digital pre-distortion (DPD) for pre-distorting input signals to a nonlinear electronic component, and receives the dataset. Training a parameterized power amplifier (PA) model associated with the data transformation operation based on the data set and the information of the hardware resource constraints associated with the pool of processing units, wherein the training includes updating at least one parameter of the parameterized PA model associated with constituting at least a subset of the plurality of processing units in the pool. A method comprising outputting one or more DPD operating configurations for at least the subset of the plurality of processing units in the pool based on the training.
2. The method according to claim 1, further comprising generating the parameterized PA model, wherein the generation comprises generating a mapping between each of the plurality of processing units in the pool and one of the plurality of differentiable function blocks.
3. The data conversion operation includes at least a sequence of first and second data conversions, The aforementioned training is The method according to claim 1, comprising calculating a first parameter associated with the first data transformation and a second parameter associated with the second data transformation.
4. The method according to claim 3, wherein the calculation of the first parameter associated with the first data transformation and the second parameter associated with the second data transformation is further based on the pre-strain backpropagation and loss function.
5. The method according to claim 3, wherein the first data conversion or the second data conversion in the sequence is associated with an executable instruction code.
6. The first data conversion in the sequence includes selecting a memory term from the input signal based on the first parameter, The second data transformation in the sequence includes generating features associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory terms based on the second parameter, The method according to claim 3, wherein the sequence associated with the data conversion operation further includes a third data conversion, which includes generating a pre-distorted signal based on the features.
7. The first data conversion in the sequence includes selecting a memory term from a feedback signal indicating the output of the nonlinear electronic component or the input signal, based on the first parameter, The second data transformation in the sequence includes generating features associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory terms based on the second parameter, The method according to claim 3, wherein the sequence associated with the data conversion operation further includes a third data conversion, which includes updating the DPD coefficients based on the features and the second signal.
8. Training the parameterized PA model is The method according to claim 7, further comprising performing backpropagation of pre-strain errors to update the second parameter and generate the set of basis functions.
9. Outputting one or more of the above configurations means The method according to claim 7, further comprising outputting one or more configurations that further indicate at least one of the lookup table (LUT) configurations associated with the selection of the memory term or the set of basis functions.
10. A computer implementation system, Memory containing instructions, A computer processor comprising one or more computer processors, When the instruction is executed by one or more computer processors, the one or more computer processors will: Receiving information on hardware resource constraints associated with a pool of processing units, wherein multiple processing units in the pool perform one or more arithmetic operations and one or more signal selections, and receiving information on hardware resource constraints associated with a pool of processing units. Receiving a dataset associated with a data transformation, wherein the data transformation is associated with at least one of digital pre-distortion (DPD) operation or DPD adaptation for pre-distorting an input signal to a nonlinear electronic component; Training a parameterized power amplifier (PA) model associated with the data transformation based on the information of the hardware resource constraints associated with the dataset and the pool of processing units, wherein the training of the parameterized PA model includes updating at least one parameter of the parameterized PA model associated with constituting at least a subset of the plurality of processing units in the pool. A computer implementation system that performs an operation including outputting at least one of a DPD operating configuration or a DPD adaptive configuration for at least the subset of the plurality of processing units in the pool, based on the training.
11. The aforementioned operation is, The computer implementation system according to claim 10, further comprising generating the parameterized PA model by generating a mapping between each of the plurality of processing units in the pool and one of the plurality of differentiable function blocks.
12. The data conversion includes at least a sequence of a first data conversion and a second data conversion, Training the parameterized PA model is The computer implementation system according to claim 10, comprising calculating a first parameter associated with the first data transformation and a second parameter associated with the second data transformation based on the backpropagation and loss function of the pre-strain.
13. The subset of the processing unit is One or more digital hardware blocks associated with the first data conversion, The computer implementation system according to claim 12, comprising one or more processors for executing instruction codes associated with the second data conversion.
14. A non-temporary computer-readable storage medium storing instructions, wherein when an instruction is executed by one or more computer processors, the one or more computer processors... Receiving information on hardware resource constraints associated with a pool of processing units, wherein multiple processing units in the pool perform one or more arithmetic operations and one or more signal selections, and receiving information on hardware resource constraints associated with a pool of processing units. To generate a mapping between each of the multiple processing units in the pool and one of the multiple differentiable function blocks, Receiving a dataset associated with a data transformation, wherein the data transformation is associated with digital pre-distortion (DPD) for pre-distorting input signals to a nonlinear electronic component, Training a parameterized power amplifier (PA) model to constitute at least a subset of the plurality of processing units in the pool to perform the data transformation, wherein the training is based on the dataset, information on the hardware resource constraints associated with the pool of processing units, and the mapping, and includes updating at least one parameter of the parameterized PA model associated with constituting at least a subset of the plurality of processing units in the pool. A non-temporary computer-readable storage medium that causes an operation to be performed, which includes outputting one or more DPD operating configurations for at least the subset of the plurality of processing units in the pool, based on the training described above.
15. The data conversion includes at least a sequence of a first data conversion and a second data conversion, The aforementioned training is The non-temporary computer-readable storage medium according to claim 14, further comprising updating a first parameter associated with the first data transformation and a second parameter associated with the second data transformation based on the backpropagation and loss function of the pre-strain.
16. The first data conversion in the sequence includes selecting a memory term from the input signal based on the first parameter, The second data transformation in the sequence includes generating features associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory terms based on the second parameter, The non-temporary computer-readable storage medium according to claim 15, wherein the sequence further comprises a third data transformation including generating a pre-distorted signal based on the features.
17. The first data conversion in the sequence includes selecting a memory term from a feedback signal indicating the output of the nonlinear electronic component or the input signal, based on the first parameter. The second data transformation in the sequence includes generating features associated with the nonlinear properties of the nonlinear electronic component using a set of basis functions and the selected memory terms based on the second parameter, The sequence further includes a third data transformation which includes updating the DPD coefficients based on the features and the second signal, The first data conversion and the second data conversion are performed by the subset of the processing unit. The non-temporary computer-readable storage medium according to claim 15, wherein the third data conversion is performed by executing instruction code on at least another processing unit in the pool.