Apparatus and method for providing a plurality of output samples based on a plurality of input samples
By using parallel decimation digital convolutional converters and hierarchical tree-structured sample combiner logic, the problem of parallel processing at high sampling rates in digital signal processing is solved, achieving flexible sampling rate conversion and efficient signal processing, suitable for real-time or near-real-time digital signal processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ADVANTEST CORP
- Filing Date
- 2019-12-23
- Publication Date
- 2026-05-12
AI Technical Summary
In existing digital signal processing technologies, when the sampling rate is higher than the clock rate of the digital signal processor, it is difficult to achieve effective parallel processing, resulting in low sampling efficiency.
By employing a parallel decimation digital convolutional converter and utilizing multiple processing cores and a hierarchical tree-structured sample combiner logic, multiple input samples are processed in parallel to provide multiple output samples, achieving flexible sampling rate conversion and efficient signal processing.
It achieves efficient signal processing under high sampling rate conditions, supports flexible sampling rate conversion and high-quality sampling of digital waveforms, and can process complex signals in real time or near real time, reducing processing latency and improving processing efficiency.
Smart Images

Figure CN114128145B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to digital signal processing.
[0002] A further embodiment of the invention relates to real-time waveform processing on a digital signal processor (DSP). More specifically, it relates to real-time waveform processing on a DSP, wherein the rate of the processed data is higher than the clock speed of the DSP, and therefore a parallel data processing architecture is employed.
[0003] Embodiments of the present invention relate to parallel decimation digital convolutions. Background Technology
[0004] Decimation describes a downsampling process to produce an approximation of the sequence that would otherwise be obtained by sampling the signal at a lower rate. This means that the output sampling rate is generally lower than or equal to the input sampling rate.
[0005] A decimator, or decimation convolution, convolves an input waveform given by isometric sampling with a continuous-time impulse response, and produces the result of this operation at its output at a sampling rate lower than or equal to the input rate. The continuous-time impulse response is stretched over time proportionally to the sampling rate. By using an appropriately selected impulse response, the decimator can be designed to suppress spectral content in the input waveform that would otherwise produce unwanted aliasing at the output sampling rate.
[0006] The decimator exhibits an algorithmic architecture suitable for convenient implementation on application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). A traditional decimator can be implemented as a transposed Farrow structure. The impulse response of the transposed Farrow structure is described in a piecewise polynomial manner.
[0007] The implementation of traditional operations for performing decimation convolution or decimation digital convolution on sequential DSPs was proposed by Babic and Hentschel and is summarized below.
[0008] The time accumulator accumulates fractional samples in the half-open interval [0:1) with an increment of Δ. The sampling ratio is 1 / Δ, where Δ is within the half-open interval [0:1). When the time accumulator overflows, the sampler emits an output sample and shifts the output sample one position in the output accumulator.
[0009] Within the output accumulator, multiple output samples are being prepared. The output accumulator accumulates or integrates the results of multiple so-called dot-cores. Each dot-core computes a dot product or scalar vector product between a vector of coefficients and the corresponding output vector of the polynomial evaluator. The coefficients of the dot-cores determine the continuous-time convolution kernel in a piecewise polynomial manner, and thus also determine the decimator's response.
[0010] The number of output samples or the number of corresponding kernels in a set of multiple output samples. M The support is called the Faro extractor, and the number of coefficients in the coefficient vector is... N It is the degree of the Faro extractor.
[0011] The multinomial estimator multiplies the input samples by successive powers of the accumulated fractional time: 0, 1, ... N .
[0012] As a result of the accumulation process, the amplitude of the output waveform is scaled by 1 / Δ. t To match the output amplitude with the input or input amplitude, each output sample is multiplied by Δ. t .
[0013] The traditional Faro implementation processes one sample at a time, meaning its parallelism is 1.
[0014] Whenever the sampling rate exceeds the clock rate of the digital signal processor, parallel processing operations need to be performed (e.g., on a common set of samples), while keeping the effort of combining samples reasonably small.
[0015] This objective is addressed by the subject matter of the independent claims. Summary of the Invention
[0016] One embodiment of the present invention (see, for example, claim 1) is a digital signal processing apparatus, such as a decimator or decimated convolutional unit, for processing core input values based on a plurality of input samples or sets of input values, for example, and providing a plurality of output samples or output values in parallel, for example, P One output sample.
[0017] Digital signal processing apparatuses include multiple processing cores or modified transposed Faro cores configured to perform processing operations, such as decimation operations or decimated digital convolution operations based on individual input samples and associated processing times, to provide a set of processing core output samples, for example, each processing core... M Each core outputs a sample.
[0018] The digital signal processing apparatus also includes a sample combiner logic or structure configured to provide multiple output samples from multiple sets of output samples from multiple processing cores, such as decimation cores or Faro decimators, which perform processing operations associated with different processing times, such as the time associated with the input samples, or the time relative to a reference time, for example... t , t +Δ t , t +2Δ t, … , .
[0019] The sample combiner logic includes a hierarchical tree structure with combiner nodes at multiple levels.
[0020] The highest-level combiner nodes are configured to provide a set of combined output samples based on two or more sets of processing core output samples.
[0021] Furthermore, each combiner node at a given level lower than the highest level is configured to provide a set of combined output samples based on two or more sets of output samples from associated combiner nodes at higher levels.
[0022] Each combiner node is configured to combine its respective set of input samples, and each set of input samples becomes shifted and / or zero-padded based on the time information associated with these sets of input samples.
[0023] In other words, associated with different processing times, for example P One input sample is provided P Each processing core or a modified transposed Faro core. Each processing core provides, for example, the combiner logic. M Each output sample, the combiner logic includes a hierarchical tree structure, consisting of combiner nodes at multiple levels.
[0024] Each combiner node is configured to combine two or more sets of input samples from a given combiner node. Each combiner node at a given level receives input samples from the combiner node at the next higher level and feeds its set of output samples to the combiner node at the next lower level.
[0025] The output sample of the combiner logic, for example P + M A single sample is the output of the combiner node at the lowest level, while the set of inputs to the combiner logic is, for example... M The set of samples is the input set of the combiner node at the highest level.
[0026] According to an embodiment (see, for example, claim 2), the target output sampling rate of the output samples of the digital signal processing device is lower than or equal to the input sampling rate of the input samples of the digital signal processing device.
[0027] Digital signal processing devices are configured to provide output samples that are generally coarser than the input samples. A digital signal processing device produces the result of its operation at its output at a sampling rate lower than or equal to its input rate.
[0028] The following are some typical, but not limiting, use cases and / or applications of this property of digital signal processing devices:
[0029] - Flexible (or virtually arbitrary) sample rate conversion, where the target sample rate is less than or equal to the source sample rate, and / or
[0030] - Digital delay with subsample resolution, a special case of flexible (or nearly arbitrary) sampling rate conversion when the target rate equals the source rate, and / or
[0031] - Sample digitized digital waveforms with a defined sampler frequency response, and / or
[0032] – Track input waveforms with timing jitter, for example, as part of a clock recovery loop.
[0033] In a preferred embodiment (see, for example, claim 3), the digital signal processing apparatus includes a time accumulator.
[0034] The time accumulator is configured to track the global processing time and, whenever the global processing time overflows by a predetermined multiple of the sampling period of the output sample (e.g., ...), ... P When this occurs, multiple output samples are triggered from the output register and / or output accumulator, for example... P One output sample. The output register and / or output accumulator are coupled to the sample combiner logic, for example, via a shift block or shifter.
[0035] Time accumulator with P ×Δ t The increment is in the half-open interval [0: P The accumulator accumulates a fraction of samples. Whenever the time accumulator overflows, the extractor fires, for example... P Each output sample is shifted in the output register and / or accumulator.
[0036] According to the embodiments (see, for example, claim 4), the number of samples in the input sample sets of multiple combiner nodes at the same level of the combiner logic is the same, and / or the number of samples in the output sample sets of multiple combiner nodes at the same level of the combiner logic is the same.
[0037] For example, the number of samples in the input sample set and the number of samples in the output sample set of the first combiner node are equal to the number of samples in the input sample set and the number of samples in the output sample set of the second combiner node at the same level.
[0038] A combiner logic in which combiner nodes at the same hierarchical level have an equal number of samples in their input sample sets and an equal number of samples in their output sample sets. The combiner logic has a modular structure with hierarchical levels constructed from the same modules, which makes the production and / or planning of the combiner logic simpler, cheaper, and / or faster.
[0039] In a preferred embodiment (see, for example, claim 5), the number of samples in the output sample set of a given combiner node is greater than the number of samples in each input sample set provided to the given combiner node by the next higher-level combiner node or by the processing core as input samples.
[0040] A given combiner node combines two or more input samples with an equal number of other samples to form an output sample set.
[0041] The number of output samples of a given combiner node is greater than the number of samples in any input sample set of that given combiner node. The input sample set of a given combiner node contains an equal number of samples that are provided as output sample sets by the next higher-level combiner node or by the processing core.
[0042] According to an embodiment (see, for example, claim 6), the sample combiner logic is configured such that the number of samples provided to the combiner node as input samples by each combiner node of the next higher level gradually increases as the level decreases.
[0043] Combiner logic is a series of combiner nodes, where each combiner node receives two or more output sets as input sample sets from combiner nodes at higher levels, and provides output sample sets to combiner nodes at lower levels.
[0044] The combiner nodes at the highest level receive two or more sets of input samples from their respective two or more processing cores.
[0045] According to the tree structure of combiner logic, from top to bottom, the number of samples in the output sample set of combiner nodes at different levels increases, and the number of samples in the input sample set of combiner nodes at lower and lower levels also increases.
[0046] According to an embodiment (see, for example, claim 7), the number of input samples for each combiner node and / or the number of output samples provided by each combiner node is based on the number of samples in the output sample set of a single processing core, for example, expressed as... M And / or based on the hierarchical level of each combiner node, for example, represented as h And / or based on the number of processing cores, for example, expressed as P Factoring a factor into integer factors, for example, represented as p k .
[0047] There exists a relationship between the number of input samples and the number of output samples for a given combiner node. This relationship depends on the hierarchy level of the given combiner node, the number of output samples for the processing core, and an integer factor of the number of processing cores. For example, defining this relationship through an equation provides a clear and direct understanding of the combiner node and / or the entire combiner logic.
[0048] In a preferred embodiment (see, for example, claim 8), the number of input sample sets for each combiner node depends on the number of processing cores, for example, denoted as P Factoring a factor into integer factors, for example, represented as p k .
[0049] For example P Integer factors of a number are not necessarily prime factors, therefore P Depend on Description. In this formula, P Indicates the number of processing cores. k Represents 0 to ( H The variables between -1) are running variables, and H This represents the total number of factors in the selected integer factorization.
[0050] Combiner nodes at the same level have the same number of samples in their input sample set and provide the same number of output samples.
[0051] According to an embodiment (see, for example, claim 9), a given hierarchical level h The number of input sample sets for each combiner node is, for example, represented as... p h It is the number of processing cores. P Integer factors one.
[0052] p h It is the number of processing cores. PInteger factors (not necessarily prime factors) An element of the set, thus P Depend on The description is as described above.
[0053] p h In h This indicates the hierarchical level of each combiner node. The highest hierarchical level is determined by... h = 0 description, and h It increases as the level decreases.
[0054] In a preferred embodiment (see, for example, claim 10), the number of samples in each input sample set of each combiner node is based on the following equation:
[0055]
[0056] In this equation, This represents the number of samples in each input sample set.
[0057] This represents the number of samples in each input sample set of each combiner node at a given hierarchical level.
[0058] Indicates the number of processing cores P Integer factors of a number are not necessarily prime factors, therefore As mentioned above,
[0059] h This represents the hierarchical level of each combiner node, where the highest level is represented by... h = 0 description, and h It increases as the level decreases, and
[0060] M This represents the number of samples in the output sample set of a single processing core.
[0061] In a preferred embodiment (see, for example, claim 11), the number of output samples for each combiner node is based on the following equation:
[0062]
[0063] In this equation, This indicates the number of output samples provided by each combiner node.
[0064] Indicates the number of processing cores P Integer factors of a number are not necessarily prime factors, therefore As mentioned above,
[0065] h This represents the hierarchical level of each combiner node, where the highest level is represented by... h = 0 description, and h It increases as the level decreases, and
[0066] M This represents the number of samples in the output sample set provided by a single processing core.
[0067] In a preferred embodiment (see, for example, claim 12), each combiner node at each level of the sample combiner logic is configured to provide a set of combined output samples. The set of combined output samples is a combination of the sets of input samples.
[0068] The signal processing device is configured to, prior to combination, process the signal based on time information associated with the input sample set (e.g., int). i The relationship between the input sample sets, such as differences, determines how many samples are shifted relative to each other.
[0069] A given combiner node provides a set of combinations of two or more sets of input samples provided to that given combiner node. Different sets of input samples are associated with different processing times.
[0070] Different processing times result in different sets of input samples, and one sample may be contained in more than one set of input samples.
[0071] According to an embodiment (see, for example, claim 13), each combiner node in each level of the sample combiner logic is configured to provide a set of combined output samples by summing an appropriate zero-padded version of the input sample set, wherein the padding amount and position of a particular input sample set depends on time information associated with the input sample set.
[0072] Summing the selected, appropriately zero-padded versions of the input sample set allows the input sample set to be combined into a single output sample set. The combined input sample set is a larger set of samples than the output sample set. Before combining into a single output sample set, a given number of samples are selected from the zero-padded sample set, starting from a starting index that depends on the time information associated with the input sample set.
[0073] In a preferred embodiment (see, for example, claim 14), the highest-level combiner node is configured to receive individual time information associated with each corresponding set of input samples, such as int. iEach time value, such as int or floor(t+Δt), corresponds to (i.e., is based on or related to) the processing time associated with each set of input samples, such as t+n. Δt.
[0074] The timing information associated with the input sample set of each combiner node is used to calculate the starting index selected from the zero-padded input set before combining the input sample sets into the output sample set. This timing information depends on the processing time associated with each input sample set.
[0075] According to an embodiment (see, for example, claim 15), the processing cores are configured to use processing times (e.g., t+n) associated with each processing core. The fractional part of Δt, for example denoted as frac, is used to determine the processing function. The signal processing unit is configured to use the integer part of each processing time t associated with each processing core, for example, an int, as time information, such as an int associated with each set of input samples. i This information is provided by each processing core to each combiner node at the highest level.
[0076] Fractional portions of each processing time are provided to the processing core. Integer portions of each processing time are provided to the individual combiner nodes at the highest level of the combiner logic.
[0077] In a preferred embodiment (see, for example, claim 16), each combiner node at each level is configured to assign time information of integer values to the combined output samples based on time information associated with the set of input samples.
[0078] The time information associated with the set of combined output samples is an integer value based on the time information of one or more sets of input samples. For example, the time information associated with the set of combined output samples is equal to the integer value of the time information of one of the sets of input samples.
[0079] In a preferred embodiment (see, for example, claim 17), the timing information assigned to the combined output sample is equal to the timing information associated with one of the input sample sets.
[0080] Assigning time information associated with one of the input sample sets to the output sample set is a simple way to assign time information to the output sample set.
[0081] In a preferred embodiment (see, for example, claim 18), the digital signal processing apparatus includes an output register configured to store a plurality of output samples.
[0082] The advantage of storing samples in the output register is that data is not lost through further data processing and / or reuse is allowed, i.e., the same sample is processed more than once, for example, through the accumulation of output samples.
[0083] In a preferred embodiment (see, for example, claim 19), the output register is configured to accumulate and / or integrate the values of the output samples.
[0084] The result of accumulating and / or integrating the output values is a combination of output samples, while keeping the set of output values of the signal processing device smaller and / or more compact.
[0085] In a preferred embodiment (see, for example, claim 20), the output register or output accumulator includes a shift register.
[0086] Since only a finite number of output samples need to be stored, a shift register is sufficient. Shift registers are a feasible solution for storing a finite number of samples; they are widely used, simple to use, and inexpensive.
[0087] Furthermore, the accumulation in the output accumulator uses shift operations, which can be easily performed by a shift register.
[0088] According to an embodiment (see, for example, claim 21), the digital signal processing apparatus includes shift and / or padding logic configured to operate on the output sample set of the last combiner node of the sample combiner logic.
[0089] The shift and / or padding logic appends and / or prepends an appropriate number of zeros to the sample set provided by the combiner logic. A predetermined number of samples are selected from the appropriately zero-padded output samples, starting from the associated index of the timing information associated with the output samples of the combiner logic.
[0090] In a preferred embodiment (see, for example, claim 22), if timing jitter is applied, the processing time associated with the processing core is either equidistant or unequal.
[0091] Because processing time is associated with processing operations, the variability of processing time, whether equidistant or non-equidistant, may result in the execution of variable processing operations with equidistant or non-equidistant processing times.
[0092] In a preferred embodiment (see, for example, claim 23), the signal processing device performs extraction on the input sample.
[0093] Whenever the time accumulator overflows, the digital signal processing device emits a new set of output samples.
[0094] The cumulative time information scores are associated with each processing core, while the integer values of the cumulative time information are associated with the output sample set. As a result, the output sample set is a extraction of the input sample set.
[0095] According to an embodiment (see, for example, claim 24), the digital signal processing apparatus performs convolution.
[0096] Since a given processing core performs a sample combination operation by obtaining a set of input samples and outputting a single set of output samples, which provides a single output element from multiple input elements, the sample combiner logic performs a weighted average operation or a convolution operation.
[0097] In a preferred embodiment (see, for example, claim 25), multiple processing cores implement a transposed faro structure. The transposed faro structure is a widely used extractor implementation, making it an easy-to-apply, readily available, and cost-effective solution.
[0098] According to an embodiment (see, for example, claim 26), the construction of different subtrees is based on the number of processing cores. P Integer factors p k The results are derived from the same or different choices.
[0099] As an example, when P When the number of cores is 16, the number of cores processed can be factored into 16 = (2 × 2 × 2) × 2 for a portion of the tree and / or into 16 = (4 × 2) × 2 for a different portion of the tree.
[0100] According to an embodiment (see, for example, claim 27), the construction of different subtrees is based on the number of processing cores. P Integer factors p k The results are derived from the same or different sorting.
[0101] As an example, when P When the number of cores is 16, the number of cores processed can be factored into 16 = 2 × 4 × 2 for a portion of the tree and / or into 16 = 4 × 2 × 2 for a different portion of the tree.
[0102] A corresponding method was created according to another embodiment of the present invention.
[0103] However, it should be noted that these methods are based on the same considerations as the corresponding apparatus. Furthermore, these methods may be supplemented by any features and / or functions and / or details described herein with respect to the apparatus, either individually or in combination. Attached Figure Description
[0104] In the following, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings, in which:
[0105] Figure 1 A schematic block diagram of a signal processing device is shown, which includes combiner logic and multiple processing cores;
[0106] Figure 2 A schematic block diagram of a signal processing device is shown, which is expanded to include a time accumulator, a shifter, and an accumulator module;
[0107] Figure 3 A schematic block diagram of a combiner node with combiner logic having two sets of input samples is shown.
[0108] Figure 4 A schematic block diagram of the shifter is shown;
[0109] Figure 5 A schematic diagram of a conventional faro extractor (conventional transposed faro structure) is shown;
[0110] Figure 6 A schematic block diagram of the modified Faro core is shown, in which, for example, the "modified Faro core" includes the "Faro core" plus the calculation of "int (integer part)" and "frac (fractional part)";
[0111] Figure 7 An exemplary block diagram of an extended signal processing device is shown. Detailed Implementation
[0112] Different inventive embodiments and aspects will be described below. Further embodiments will be defined by the appended claims.
[0113] It should be noted that any embodiment defined by the claims may be supplemented by any details, features, and / or functions described herein. Furthermore, the embodiments described herein may be used alone or optionally supplemented by any details and / or features and / or functions included in the claims.
[0114] Furthermore, it should be noted that the individual aspects described herein can be used individually or in combination. Thus, details can be added to each of the individual aspects without adding details to the other aspect.
[0115] It should be noted that this disclosure explicitly or implicitly describes features that can be used in signal processing apparatuses. Therefore, any feature described herein can be used in the context of signal processing apparatuses.
[0116] Furthermore, the features and functions disclosed herein related to the method can also be used in apparatuses configured to perform such functions. Additionally, any features or functions disclosed herein regarding the apparatus can also be used in the corresponding method. In other words, the method disclosed herein can be supplemented by any features and functions described regarding the apparatus.
[0117] The invention will be more fully understood through the following detailed description and the accompanying drawings of embodiments thereof; however, the detailed description and drawings should not be construed as limiting the invention to the specific embodiments described, but are merely for illustrative and understanding purposes.
[0118] according to Figure 1 Implementation examples
[0119] Figure 1 A block diagram of a digital signal processing apparatus 100 is shown, which includes combiner logic 110 and multiple processing cores 120. The combiner logic 110 includes multiple combiner nodes 130a-f organized into a hierarchical tree structure 140 with multiple hierarchical levels 140a-c.
[0120] The input sample 150 of the digital signal processing device is provided to multiple processing cores 120.
[0121] Multiple processing cores 120 include processing cores 120a-f. The inputs of processing cores 120a-f are the inputs of digital signal processing device 100. The outputs 125a-f of processing cores 120a-f are coupled to combiner logic 110.
[0122] Processing cores 120a-f are associated with different processing times and are configured to acquire one input sample from input samples 150 and provide output sample sets 125a-f to combiner logic 110 each time, each set having, for example, M One output sample.
[0123] The output sample set 125a-f of the processing cores 120a-f is provided to the combiner logic 110 as input samples, wherein the sample set 125a-f is provided to the highest level 140a. h Combiner nodes 130a-c (= 0). Combiner nodes 130a-c take the input sample set 125a-f as input and provide the combined set 160a-d to the next lower level combiner nodes 130d-e at level 140b. The number of samples in the output sample set at the same level is the same, for example, the output sample set 160a-d at level 140a, or the output sample set 160e-f at level 140b.
[0124] Any given combiner node 130a-f obtains two or more sets of input samples from the next higher level. For example, combiner node 130d obtains the set of input samples 160a-b from combiner nodes 130a-b at level 140a and provides a combined set, such as 160e, to the next lower level combiner node, such as combiner node 130f at level 140c.
[0125] The combiner logic has a hierarchical tree structure 140 of combiner nodes 130a-f, wherein the highest-level combiner nodes 130a-c obtain input sample sets 125a-f from each processing core 120a-f, and each other combiner node 130d-f obtains input sample sets from the next higher level.
[0126] The combiner node 130f at the lowest level 140c is providing output 180, which is the output of combiner logic 110 and the signal processing device. The output of each of the other combiner nodes 130a-e of combinational logic 110 is coupled to one of the inputs of the next lower level combiner nodes 130d-f.
[0127] In other words, the digital signal processing apparatus 100 includes multiple processing cores 120 and combiner logic 110, and is configured to provide multiple output samples 180 from multiple input samples 150. The multiple processing cores 120 perform processing operations in parallel, wherein processing cores 120a-f are associated with different processing times. The output sample set 125a-f of processing cores 120a-f is provided to the combiner logic 110 as an input sample set.
[0128] Combiner logic 110 provides output sample set 180 from input sample set 125a-f by using a hierarchical tree structure 140 of combiner nodes 130a-f organized into hierarchical levels 140a-c.
[0129] Input sample 150 is fed as input to processing cores 120a-f to provide output sample sets 125a-d to combiner logic 110, wherein the number of samples in sets 125a-f is equal for all sets 125a-f.
[0130] Each level 140a-c of the combiner logic 110 includes combiner nodes 130a-f, wherein the combiner node 130a-f of a given level 140a-c takes two or more sets 125a-f, 160a-f of input samples from the next higher level and provides a set 160a-f for the next lower level 140a-c.
[0131] The digital signal processing device 100 or parallel decimation digital convolution 100 described herein can be used as a key building block of a signal processor application-specific integrated circuit (ASIC) and / or as part of other instruments.
[0132] The digital signal processing apparatus described in this paper can be applied on parallel DSPs to address flexible (or almost arbitrary high) sampling rates with real-time or near-real-time response times; for example, the digital signal processing apparatus can handle sampling rates of 100 GSa / s in near real-time. It is an area-efficient implementation of an architecture with a parallel processing core.
[0133] Furthermore, this signal processing device can be used to provide high-quality, flexible (or nearly arbitrary) sample rate conversion for radio frequency (RF) and analog baseband applications in near real-time. The usable bandwidth can be, for example, 75% of the Nyquist rate, and image suppression, for example, can be achieved at 60 dB. The conversion ratio is not explicitly limited to some simple fractions, but is truly flexible (or nearly arbitrary) because it is programmed as a number between 0 and 1 with 64-bit resolution. Sample rates far exceeding the clock rate of a DSP can be resolved.
[0134] In addition, signal processing devices can be used to sample digitized non-return-to-zero (NRZ) digital waveforms and / or pulse-amplitude modulation (PAM) digital waveforms to obtain flexible (or virtually arbitrary) user bit rates.
[0135] Additionally, a clock recovery loop can be used to track drifting digital waveforms.
[0136] An important use case is to provide subsample resolution latency for time-to-digital (TDC) based synchronization mechanisms.
[0137] according to Figure 2 Implementation examples
[0138] Figure 2 A schematic block diagram or high-level block diagram of a signal processing device 200 is shown. Figure 1 An enhanced or extended version of the digital signal processing device 100. The output of the digital signal processing device 200 is coupled to a shifter 270. The shifter 270 has one input and one output, and the output of the shifter 270 is coupled to an accumulator 290.
[0139] Accumulator 290 has two inputs and one output. The first input of accumulator 290 is coupled to shifter 270, and the second input of accumulator 290 is coupled to time accumulator 295. The output of accumulator 290 is the output of extended digital signal processing device 200. Time accumulator 295 is coupled to accumulator 290 and is configured to trigger the transmission of output samples of digital signal processing device 200, and is configured to provide timing information to processing core and / or combiner logic 210.
[0140] Input samples 250 of the signal processing device 200 are provided to a plurality of processing cores 220, including processing cores 220a-f. Processing cores 220a-f, such as processing core 220b, are coupled to combiner logic 210. Processing cores 220a-f expect to receive the input samples as input and provide an output sample set 225a-f as output. The output sample set 225a-f is the input sample set of the combiner logic 210.
[0141] Any of the processing cores 220a-f, such as processing core 220b, has one input and one output. Processing cores 220a-f expect input samples from input sample 250 and provide output sample sets 225a-f. These output sample sets 225a-f are the input sample sets of combiner logic 210.
[0142] and Figure 1 The combiner logic 110 is similar to the combiner logic 210, which includes a hierarchical tree structure 240 of combiner nodes 230a-f organized into multiple hierarchical levels 240a-c.
[0143] The inputs of combiner nodes 230a-c on the highest level 240a of combiner logic 210 are the inputs of combiner logic 210. Combiner nodes 230a-c have two or more inputs, which are coupled to processing cores 220a-f in multiple processing cores 220, which is related to... Figure 1 The multiple processing cores are similar to 120.
[0144] Any combiner node 230a-f of combiner logic 210 has one output and two or more inputs. The input of a given combiner node 230a-f is coupled to another combiner node 230a-f at the next higher level 240a-c, and the output of combiner node 230a-f is coupled to another combiner node 230a-f at the next lower level 240a-c.
[0145] The output sample of combiner node 230f at the lowest level 240c is the output sample of combiner logic 210. Combiner node 230f at the lowest level 240c of combiner logic 210 is coupled to accumulator 290 via shifter 270.
[0146] In other words, as Figure 1 The extended version of the digital signal processing device 100, the digital signal processing device 200, includes the digital signal processing device 100 and is extended by a shifter 270, an accumulator 290, and a time accumulator 295.
[0147] The time accumulator 295 is configured to track the processing time and, whenever the processing time overflows by a predetermined multiple of the sampling period of the output sample (e.g., ...), ... P When ), it triggers the emission of output sample 280 from accumulator 290, for example. P One sample.
[0148] Accumulator 290 is configured to accumulate and / or integrate the samples provided by shifter 270 to provide output sample 280, for example P One output sample. The output sample 280 of the accumulator 290 is the output sample of the extended signal processing device 200.
[0149] Shifter 270 is configured to prepend and / or append zeros to the output samples of combiner logic 210, and select a predetermined number of samples from the zero-padded sample set, for example, 2. P + M Two samples are used to provide the selected sample set as input to accumulator 290.
[0150] Processing cores 220a-f, such as transposed Faro cores, will process samples from input samples 250 (e.g., M A set of samples (e.g., a set of samples) is provided to provide an area-efficient implementation, for example, distributed logic 210.
[0151] The input samples of the combiner logic 210, provided by multiple processing cores 220, are the input samples of combiner nodes 230a-c in the first-level layer 240a, and time information based on accumulated time 298. Each combiner node 230a-f on each level layer 240a-c is configured to assign time information to each output sample set, wherein the time information is based on the processing time tracked by time accumulator 295.
[0152] Each combiner node 230a-f of combiner logic 210 is configured to combine the set of input samples into a set of output samples as input to the next lower-level combiner node 230a-f.
[0153] In addition, each combiner node 230a-f at each level 240a-c is configured to assign time information (based on 298) to the output sample set based on the time information of the input sample set assigned to each combiner node 230a-f.
[0154] Depending on whether timing jitter is applied, the processing time 298 tracked by the time accumulator 295 can be equidistant or non-equidistant.
[0155] The combiner node 230f of the lowest level 240c provides output samples to the accumulator 290 via the shifter 270 so as to accumulate and / or integrate the zero-filled output samples into the output sample set 280.
[0156] The digital signal processing device 200 performs the same and / or similar mathematical operations as, for example, a classic faro decimator (based on a transposed faro structure), but processes multiple (e.g., ...) operations at a time in each clock cycle. P (Number) samples. It generates one sample per clock cycle. P It has a number of time-continuous output samples, so its parallelism is greater than 1.
[0157] Multiple processing cores include P Each processing core consists of an identical processing core, or a modified Faro core. Each processing core includes a point core and a polynomial evaluator used in the modified Faro core or in the modified Faro implementation.
[0158] Time accumulator 295 P ×Δ t The increment is in the half-open interval [0; P Accumulated score samples. The extractor fires whenever the time accumulator 295 overflows. P One output sample.
[0159] P Each input sample is sent to the corresponding P Each processing core, thus providing each M One output sample. Multiple processing cores 220a-f include P The same processing core or a modified Faro core, associated with different processing times, for example t, t+Δt, t+2Δt The processing core 220a-f can be implemented as a modified Faro core (…). Figure 6 The modified Faro core (600) comprises multiple point cores and a polynomial evaluator. Each of the modified Faro cores provides combiner nodes 230a-c of the highest level 240a of the combiner logic 210. MEach output sample. The area-efficient implementation of combiner logic 210 ensures that each modified Faroe core or processing core 220 corresponds to the output accumulator 290. M Contribute to the correct subset of each sample.
[0160] A given combiner node takes two or more sets of input samples, for example... M The set of input samples is used to combine a set of output samples into a set of combinations. This set of combinations of output samples is then used as the input sample set for the combiner node at the next lower level. For example, the output sample of combiner node 230f at the lowest level 240c... P + M One sample is provided to shifter 270 as an input sample.
[0161] The shifter is configured to append and / or lead zeros to its input samples, e.g., P One zero, and select samples from the zero-padded sample set, for example, 2. P + M Two samples.
[0162] The selected sample, for example, 2 P + M Two samples were fed to accumulator 290.2 P + M Two samples are accumulated in the output accumulator 290, i.e. P current sample and P + M Two future samples are needed to provide output sample 280, for example. P A number of output samples, which are used as output samples of the signal processing device.
[0163] The combination of combiner logic or sample sets is performed in two stages: combination and shifting.
[0164] The combination phase processes the output sample set of core 220a-f or the modified Faroe core 220a-f, for example M A set of samples is provided to combiner nodes 230a-c of the first level 240a of the combiner logic to combine the set of input samples. Assume... P = 2 H The combination process involves a hierarchical structure of 240, which is a height of H A perfect binary tree of type 1. Therefore, this process involves... HThere are several levels, with level h having... P / 2 h+1 There are combiner nodes, among which h = 0 ... H S 1. The final combiner node is generated. P + M One time-series continuous sample. These are shifted to the correct position by the next shift block or shifter 270 for accumulation by the accumulator 290.
[0165] The shift performed by shifter 270 includes shifting the input sample set, such as P + M One sample, appended and / or led with zeros, produces a zero-padded set of samples, for example, 3. P + M 3 samples. Select the output sample set from the set of zero-padded samples, for example, 2. P + M Two samples are used to correct the position of the samples for accumulation by accumulator 290.
[0166] Figure 3 The text describes hierarchical levels. h Operations on the "combiner node" Figure 4 The operation of the shifter is described in the document, and Figure 7 An example of one implementation method is given in the document.
[0167] according to Figure 3 combiner node
[0168] Figure 3 A schematic block diagram of combiner node 300 is shown, which is similar to... Figure 1 The combiner node 130. The input to the combiner node 300 includes two sample sets 310a-b, each with its own time information 320a-b. The combiner node 300 provides an output sample set 360 with associated time information 350 for the input sample 310. Figure 3 A specific example is when the number of processing cores is a power of two (i.e., P = 2 H And this number is based on In all p k A part of the binary tree structure generated when factoring is performed with a factor of 2.
[0169] Given a level hThe combiner node 300 is configured to combine the input sample sets 310a-b into an output sample set 360. The input sample sets 310a-b have an equal number of samples, for example... W + M One sample, of which W Depend on W =2 h Description, in which h Represents the hierarchical level of a given combiner node, where h =0 is the highest level, and h It increases by 1 as the level decreases.
[0170] Combiner node 300 appends and / or leads zeros to the input sample sets 310a-b, for example, appending to the first and second sets of input samples. W 330a-b zeros, and prepend the second set of input samples. W 340 zeros. Select a predetermined number of samples (e.g., 370) from the zero-padded input sample set. W + M One sample. The selected set of zero-padded input samples is combined into the output sample set, for example, by addition, for example, having 2 W + M 1 sample.
[0171] Samples padded with zeros, for example, from 3 W + M 1 sample, select 370 samples, for example 2 W + M One sample is selected by starting with an index 320a-b that depends on the time information associated with the input sample set, for example, 2. W + M It was conducted using only one sample.
[0172] The starting index for 370 is obtained, for example, by taking the difference between the time information associated with the input sample set, such as the difference between the time information associated with the second set of input samples and the time information associated with the first set of input samples, or it can be described by the following equation:
[0173] index=int second -int first Or index=int right -int left .
[0174] Furthermore, combiner node 300 is configured to associate timing information 350 with the output sample set 360 provided by a given combiner node 300. At a given hierarchical level of combiner node 300, the timing information 350 associated with the output sample set 360 depends on the timing information 320a-b associated with the input sample sets provided to combiner node 300. For example, the timing information associated with the output sample 360 is equal to the timing information 320a-b associated with one of the input sample sets 310a-b.
[0175] Figure 3 It shows in Figure 1 A block diagram of the combiner node 300 used in the digital signal processing device 100. The combiner node 300 is in... Figure 1 The combiner logic 110 is organized in a hierarchical tree structure in order to... Figure 1 The results from multiple processing cores 120a-f are combined into a common output sample set, and the time information 350 is associated with the output sample 360 based on the time information 320a-b associated with the input sample sets 310a-b. The output sample 360 is used as a combiner node at the next lower level or Figure 2 Input sample of shifter 270.
[0176] according to Figure 4 shifter
[0177] Figure 4 A diagram of shifter 400 is shown; it is Figure 2 Example of shifter 270. The input sample set 420 with associated time information 410 is... Figure 1 The combiner node at the lowest level of combiner logic 110 is provided to shifter 400. And shifter 400 provides the output sample set 460 to... Figure 2 290 accumulators.
[0178] The input sample set is 420, for example P + M One sample is provided to shifter 400. Zeros are appended to 430 and / or prefixed to 440 in the input sample set 420. For example, P One zero is appended and P One zero is fronted into the set of input samples, resulting in a set of zero-padded input samples, for example, 3. P + M A set of 3 samples. 450 output samples are selected from the set of zero-padded input samples by starting at the starting index associated with time information 410, for example, 2. P + M Two samples, for example, the starting index equals time information 410. The selected samples, for example, 2. P + M Two samples were provided. Figure 2 The output sample of accumulator 290 is 460.
[0179] Figure 4 A shifter 400 is shown, which is similar to Figure 2 Shifter 270. Shifter 400 from Figure 2 The combiner logic 210 receives input samples 420 with associated time information 410, and provides... Figure 2 The accumulator 290 corrects the position of the input sample.
[0180] according to Figure 5 Traditional Faro extractor
[0181] Figure 5 A block diagram of a conventional faro extractor 500, also known as a transposed faro structure, is shown. The faro extractor 500 includes an output accumulator 510, a time accumulator 520, and a faro core 530.
[0182] Time accumulator 520 with Δ t The increment accumulates fractional samples in the half-open interval [0; 1). When the time accumulator overflows, it requests a shift and emits an output sample 550 from the output accumulator 510. Whenever the time accumulator 520 overflows, the Faro decimator 500 generates an output sample 550 in each clock cycle. The accumulated fractional time is also provided to the polynomial evaluator 570 of the Faro core 530.
[0183] The modified Faro core 530 includes multiple point cores 560 and a polynomial evaluator unit 570.
[0184] The Faro decimator 500 accepts one input sample per clock cycle. The input of the Faro decimator 500 is the input of the polynomial estimator 570. The polynomial estimator 570 also has an input coupled to the time accumulator 520 and to each point core 560.
[0185] The multinomial evaluator 570 obtains input samples and fractional time inputs from the time accumulator 520, and multiplies the input samples by successive powers of the accumulated fractional time 0, 1, … N, thereby providing a sample set to the point kernel 560.
[0186] Point kernel 560 is coupled to polynomial evaluator 570 and output accumulator 510. Each point kernel 560 computes the dot product (scalar vector product) between the vector of coefficients and the vector of output values of polynomial evaluator 570. The output of the modified Faro core 530 is a sample of the outputs of multiple point kernels 560. The sampled outputs of multiple point kernels 560 are fed to output accumulator 510.
[0187] Output accumulator 510 takes the output of dot product core 560 as its input and outputs an output sample 550, which is the output sample of Faro decimator 500. The output accumulator accumulates and / or integrates the results of dot product core 560. The output accumulator emits output sample 550 and shifts the accumulated dot product value, for example, in a shift register, when time accumulator 520 overflows.
[0188] The time accumulator accumulates fractional time and provides it to the polynomial evaluator 570 of the Faroe Core 530. When the time accumulator 520 overflows, it requests to fire a new output sample 550 and shifts the value held in the output accumulator 510 (e.g., in the form of a shift register) by one position.
[0189] The dot product is provided to the output accumulator 510 by the dot kernel 560 of the Faro core 530. Each dot kernel 560 computes the dot product or scalar vector product between the vector of coefficients and the corresponding output vector of the polynomial evaluator 570 of the modified Faro core 530.
[0190] The polynomial evaluator 570 takes input samples (which are the input samples of the Faro core 530 and the Faro extractor 500) and fractional-time inputs from the time accumulator 520, and multiplies the input samples by successive powers of the accumulated fractional time 0, 1, … N, thereby providing a set of values for the point kernel 560.
[0191] The Faro Extractor 500 is a traditional extractor that processes one sample at a time, with a parallelism of 1. Figure 1 Digital signal processing device 100 relative to Figure 5 The novelty of the conventional Faro decimator 500 lies in its ability to solve the digital signal processing device 100 in real-time or near real-time on a parallel DSP for high sampling rates: for example, Figure 1 The digital signal processing device 100 can handle sampling rates of 100 gigabits per second in real time or approximately in real time.
[0192] Figure 1 The digital signal processing device 100 includes multiple processing cores 120 for parallel processing, wherein Figure 1 The processing core 120 can implement the modified Faro core ( Figure 6 (of the 600), which includes the Faro core 530. Figure 1 Combiner logic 110 will be used as Figure 1 Multiple processing cores 120 Figure 6 The output values of multiple modified Faro Core 600s are combined.
[0193] Furthermore, this signal processing device uses a single time accumulator, for example... Figure 2 Instead of multiple time accumulators 520 for each processing core or Faroe core 530, the 295 allows for Figure 6 The modified Faro Core 600 performs processing operations in parallel. Figure 1 The digital signal processing device 100 includes Figure 1 The processing core 120, they are Figure 6 The modified Faro Core 600.
[0194] according to Figure 6 The modified Faro core
[0195] Figure 6 A block diagram of the modified Faro Core 600 is shown, which includes... Figure 5 The Faro core 530 is used as the Faro core 630. The modified Faro core takes input samples 640 with associated time information 620 as input and provides multiple samples or sample sets 650 and associated time information 510 as output. Each modified Faro core takes one sample and fractional sample time as input and contributes to, for example... M One output sample.
[0196] The modified Faro core 600 includes multiple point cores 660 and a polynomial evaluator unit 670.
[0197] The multinomial evaluator 670 obtains input samples and fractional time inputs 680 based on time information 620, and multiplies the input samples by consecutive powers of cumulative fractional time 0, 1, … N, thereby providing a sample set to the point kernel 660.
[0198] Point kernel 660 is coupled to polynomial evaluator 670. Each point kernel 660 computes a dot product or scalar vector product between a vector of coefficients and the corresponding output vector of polynomial evaluator 670. The output of the modified Faro core 600 is a set 650 of output samples from multiple point kernels 660.
[0199] Additionally, the modified Faro core provides time information 610 associated with the output sample set 650. The integer value of the cumulative score time is provided as the output time information output associated with the output sample set 650, as output time information value 610. The score time value of the cumulative score time 680 is provided to the multinomial evaluator 670.
[0200] Figure 1 The digital signal processing device 100 includes multiple processing cores 120 for parallel processing, wherein Figure 1 The processing core 120 can be a modified Faro core 600. Figure 1 Combiner logic 110 will be used as Figure 1 The output values of multiple processing cores 120 and multiple modified Faro cores 600 are combined.
[0201] Furthermore, this signal processing device uses a single time accumulator, for example... Figure 2 Instead of multiple time accumulators per processing core or modified Faro Core 600, the modified Faro Core 600 performs processing operations in parallel. Figure 1 The digital signal processing device 100 includes Figure 1 Its processing core is 120, which is a modified Faro core 600.
[0202] There can be various variations in the implementation, including:
[0203] – Processing the core or modifying the Faro core does not require following Figure 5 The original implementation or the implementation provided by Babic or Hentschel. Computational or approximate support. M Any implementation of a continuous-time response to an input sample value given a time value input (e.g., 620 or 680) is eligible as a suitable processing core and can be used in a signal processing apparatus. One example alternative is a polyphase implementation where the coefficients are determined from the fractional timing information 680, for example, through mathematical relations, through a lookup table, or through a combination of both.
[0204] –Δt, the reciprocal of the extraction ratio does not necessarily have to be strictly less than 1; it can be equal to 1.
[0205] –Δt does not necessarily have to be a constant;
[0206] – The parallelism P is not limited to being an integer power of 2. If P = p 0 p 1 … p H If 1 is a factorization of P, then the combiner logic can be implemented at the hierarchical level. h Having p h The height of the combiner node for the input sample set is H A hierarchical tree of type 1;
[0207] – p kIt doesn't have to be a prime number; and
[0208] – Consider using different intervals to represent time accumulation or fractional timing information, for example [ 0.5; P 0.5), [ 0.5; 0.5) or [ 1; 1).
[0209] The following section provides specific examples of digital signal processing devices, wherein the number of processing cores is [number missing]. P =16, and each processing core outputs M =15 output samples.
[0210] according to Figure 7 Implementation examples
[0211] Figure 7 The digital signal processing device 700 is shown. Figure 1 An example of a digital signal processing apparatus 100. The digital signal processing apparatus 700 includes a time accumulator 710 configured to accumulate fractional samples in a half-open interval, for example, [0:16), with increments of 16×Δt, where Δt is within the interval, for example, (0:1].
[0212] Accumulated score time Figure 1 As shown, along with the input samples, for example, a total of 16 input samples, are provided to the processing cores, for example, 16 processing cores. A given processing core 760 provides, for example, 15 output samples and associated timing information from the input samples to the combiner node at the highest level 740a. Each combiner node 730 at the highest level is provided with, for example, two sets of input samples, for example, each set of 15 samples, along with associated timing information, and outputs a set of output samples, for example, 16 output samples, along with associated timing information.
[0213] The combiner node 730 on the second highest level 740b receives, for example, two sets of input samples, such as 16 samples in each set, along with associated time information, and provides a set of output samples, such as a set of 18 output samples, along with associated time information.
[0214] The combiner node 730 at the next lower level 740c receives, for example, two sets of input samples, such as 18 samples in each set, along with associated time information, and provides a set of output samples, such as a set of 22 output samples, along with associated time information.
[0215] The combiner node at the lowest level 740d receives, for example, two sets of input samples, such as 22 samples in each set, along with associated time information, and provides a set of output samples, such as a set of 30 output samples, along with associated time information.
[0216] The output of combiner node 730 at the lowest level 740d, for example, 30 samples, is provided to shifter 780 to correct the position of samples for accumulator 790, for example, 30 samples. Shifter 780 provides samples to accumulator 790, for example, 45 samples.
[0217] Accumulator 790 accumulates and / or integrates the samples provided by shifter 780, such as 45 samples, into a set of output samples, such as a set of 16 output samples.
[0218] All samples from the subset provided by the combiner node are offered as input samples to the combiner node at the next level. Different levels of combiner nodes provide 16, 18, 22, or 30 samples as input to the lower-level combiner node or shifter 780. The modified Faroe Core 760 is similar to... Figure 6 The modified Faro Core 600, in this example, produces 15 output samples based on one input sample and timing information from the time accumulator 710.
[0219] Comparing signal processing devices with parallel interpolation digital convolutions
[0220] The “parallel interpolation digital convolutional unit” (e.g., as described in a parallel international patent application filed on the same day as this application by the same inventor) is similar to the signal processing apparatus or decimation convolutional unit described herein.
[0221] The similarity lies in the fact that both inventions allow
[0222] - Applying continuous-time impulse response to the sampled input waveform; and
[0223] - Select an output sampling rate that is different from the input sampling rate.
[0224] Differences may include:
[0225] - Through the interpolator, or in the case of interpolation, the output rate is generally higher than or equal to the input rate, unlike the decimation case described herein, where the output rate is generally lower than or equal to the input rate.
[0226] - In the case of interpolation, the convolution kernel is applied at the input sampling rate. If the kernel is designed to attenuate the image at the input rate, this allows for flexible (almost arbitrary) sampling rate conversion to higher sampling rates.
[0227] Unlike the decimation case described in this paper, the convolution kernel is scaled to fit the output sampling rate. With a properly designed kernel, aliasing caused by resampling at lower rates will be attenuated. This allows for flexible (almost arbitrary) sampling rate conversion to lower sampling rates, along with anti-aliasing filtering.
[0228] Further potential use cases
[0229] The following are further potential use cases of the invention described above:
[0230] - This invention is beneficial to manufacturers of test equipment, such as workbenches or ATEs, or to communication systems, such as radio frequency (RF), baseband, and digital communication systems, because:
[0231] o can achieve very high-speed, highly flexible data rate processing, and / or
[0232] Significant gains in integration density can be achieved because tunable analog sampling clocks and / or switchable analog filter banks for aliasing suppression can be avoided.
[0233] - This invention is beneficial to manufacturers of general-purpose high-speed ADCs that sell converters with integrated DSP processing because:
[0234] o can achieve more flexibility than existing DSP solutions, which only support a set of discrete sampling rate ratios, or limit continuous tuning to a small range of ratios, and / or
[0235] This can provide customers of these ADCs with additional value in terms of integration density.
[0236] - This invention is beneficial for integrated high data rate modems, similar to [Erup93, Figure 13], in which the frequency and phase of the receiver sampling clock are strongly recommended—in some cases must—to be aligned with the transmitter, and the sampling clock is higher than the system clock of the DSP, thus strongly recommending—in some cases must—a parallel architecture.
[0237] - This invention is beneficial for integrated radio devices that support multiple communication standards, some or all of which have recommended or required sampling rates higher than the DSP clock speed, and which are not simple ratios to each other.
[0238] Implementation method replacement
[0239] Although some aspects have been described in the context of the apparatus, it should be clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of a corresponding block, item, or feature of the corresponding apparatus.
[0240] References:
[0241] [Babic02] D. Babic, J. Vesma, T. Saramäki, M. Renfors, “Implementation of the Transposed Farrow Structure,” Proceedings of the IEEE International Symposium on Circuits and Systems, Scottsdale, Phoenix, USA, May 26-29, 2002, pp. IV 5-8
[0242] [Hentschel01]T. Hentschel, G. Fettweis, “Continuous Time Digital Filters for Sample Rate Conversion in Reconfigurable Radio Terminals,” Frequenz, Vol. 55(56), pp. 185-188, 2001
[0243] [Erup93]L. Erup, FM Gardner, RA Harris, “Interpolation in Digital Modems Part II: Implementation and Performance,” IEEE Transactions on Communications, Vol. 41, pp. 998-1008, June 1993.
Claims
1. A signal processing apparatus for providing multiple output samples based on multiple input samples, comprising: Multiple processing cores are configured to perform processing operations based on each input sample and associated processing time in order to provide a set of processing core output samples; as well as The sample combiner logic is configured to provide multiple output samples from multiple sets of output samples from the processing cores of the multiple processing cores, which perform processing operations associated with different processing times. The sample combiner logic described therein includes a hierarchical tree structure with combiner nodes at multiple levels. The highest-level combiner nodes are configured to provide a set of combined output samples based on two or more sets of core output samples. In this context, each combiner node at a given level lower than the highest level is configured to provide a set of combined output samples based on two or more sets of output samples from associated combiner nodes at higher levels. Each combiner node is configured to combine different sets of input samples. Each set of input samples is shifted and / or zero-padded based on the time information associated with the input sample set.
2. The signal processing apparatus according to claim 1, wherein, The target output sample rate of the output sample is lower than or equal to the input sample rate of the input sample.
3. The signal processing apparatus according to claim 1 or 2, further comprising a time accumulator, the time accumulator being configured to: Track global processing time, and Whenever the global processing time overflows by a predetermined multiple of the sampling period of the output sample, multiple output samples are triggered to be emitted from the output register and / or accumulator logically coupled to the sample combiner.
4. The signal processing apparatus according to claim 1 or 2, in, The number of samples in the input sample set of combiner nodes at the same level is the same, and / or In this case, the number of samples in the output sample set of multiple combiner nodes at the same level is the same.
5. The signal processing apparatus according to claim 1 or 2, wherein, The number of samples in the output sample set of a given combiner node is greater than the number of samples in each input sample set provided to that given combiner node by the next higher-level combiner node or by the processing core.
6. The signal processing apparatus according to claim 1 or 2, wherein, The sample combiner logic is configured such that the number of samples provided to the combiner node as input samples by each combiner node at the next higher level gradually increases as the level decreases.
7. The signal processing apparatus according to claim 1 or 2, wherein, The number of input samples and / or output samples of each combiner node is based on the number of samples in the output sample set of a single processing core, and / or based on the hierarchical level of each combiner node, and / or based on the number of processing cores, decomposed into integer factors.
8. The signal processing apparatus according to claim 1 or 2, wherein, The number of input sample sets for each combiner node depends on the factorization of the number of processing cores into integer factors.
9. The signal processing apparatus according to claim 1 or 2, wherein, The number of input sample sets for each combiner node at a given hierarchical level is equal to... in Describes integer factors of P, according to Where P represents the number of processing cores, H represents the total number of factors in the chosen integer factorization, and h represents the hierarchical level of each combiner node.
10. The signal processing apparatus according to claim 1 or 2, wherein, The number of samples in each input sample set of each combiner node is based on the following equation: in This represents the number of samples in each input sample set. This represents the number of input sample sets for each combiner node at a given hierarchical level. Describes integer factors of P, according to , in P represents the number of processing cores. H represents the total number of factors in the chosen integer factorization. h represents the hierarchical level of each combiner node, and M represents the number of samples in the output sample set of a single processing core.
11. The signal processing apparatus according to claim 1 or 2, wherein, The number of output samples for each combiner node is based on the following equation: in Indicates the number of output samples. Describes integer factors of P, according to , in P represents the number of processing cores. H represents the total number of factors in the chosen integer factorization. h represents the hierarchical level of each combiner node, and M represents the number of samples in the output sample set of a single processing core.
12. The signal processing apparatus according to claim 1 or 2, wherein, Each combiner node in each level of the sample combiner logic is configured to provide a set of combined output samples. Wherein, the set of combined output samples is a combination of the set of input samples. The signal processing device is configured to determine, prior to combination, how many samples the input sample set has shifted relative to each other based on the following relationship: The relationship between the time information associated with the input sample set.
13. The signal processing apparatus according to claim 1 or 2, wherein, Each combiner node in each level of the sample combiner logic is configured to provide the set of combined output samples by summing an appropriately zero-padded version of the input sample set. The amount and position of filling a specific set of input samples depend on the time information associated with the set of input samples.
14. The signal processing apparatus according to claim 1 or 2, wherein, The highest-level combiner node is configured as follows: Receive time information associated with each input sample set, where each time information corresponds to the processing time associated with each input sample set.
15. The signal processing apparatus according to claim 1 or 2, wherein, The processing cores are configured to determine processing functionality using a fraction of the processing time associated with each processing core, and The signal processing apparatus is configured to use the integer portion of the processing time associated with each processing core as time information associated with each set of input samples provided to each combiner node at the highest level.
16. The signal processing apparatus according to claim 1 or 2, wherein, Each combiner node at each level is configured to assign time information to the combined output sample based on time information associated with the input sample set.
17. The signal processing apparatus according to claim 1 or 2, wherein, The time information assigned to the combined output sample is equal to the time information associated with one of the input sample sets.
18. The signal processing apparatus according to claim 1 or 2, comprising: The output register is configured to store multiple output samples.
19. The signal processing apparatus according to claim 18, wherein, The output register is configured to accumulate and / or integrate the values of the output samples.
20. The signal processing apparatus according to claim 3, wherein, The output accumulator includes a shift register.
21. The signal processing apparatus according to claim 1 or 2, comprising: Shift and / or padding logic is configured to operate on the output sample set of the last combiner node of the sample combiner logic.
22. The signal processing apparatus according to claim 1 or 2, wherein, The processing time associated with the processing core is either equidistant or unequal.
23. The signal processing apparatus according to claim 1 or 2, wherein, The signal processing device performs extraction on the input sample.
24. The signal processing apparatus according to claim 1 or 2, wherein, The signal processing device performs convolution.
25. The signal processing apparatus according to claim 1 or 2, wherein, The processing core implements a transposed faro structure.
26. The signal processing apparatus according to claim 1 or 2, wherein, The construction of different subtrees is derived from the same or different choices of integer factors of the number of processing cores.
27. The signal processing apparatus according to claim 1 or 2, wherein, The construction of different subtrees is derived from the same or different sorting of the number of processing cores by integer factors.
28. A method for providing multiple output samples based on multiple input samples, comprising: Multiple processing cores are used to perform processing operations based on each input sample and associated processing time in order to provide a set of output samples; and The plurality of output samples are provided from a plurality of sets of output samples from the plurality of processing cores, the plurality of processing cores performing processing operations associated with different processing times. The multiple output samples are provided using a hierarchical tree structure with multiple levels. Among these, the highest-level combinations provide a set of combined output samples based on two or more sets of core output samples. Wherein, each combination at a given level lower than the highest level provides a set of combined output samples based on two or more sets of output samples from associated combinations at higher levels. Each combination combines different sets of input samples. Each set of input samples is shifted and / or zero-padded based on the time information associated with the input sample set.