Signal processing device and signal processing method

By employing a parallel processing architecture and hierarchical tree-like sample distribution logic on a digital signal processor, the processing efficiency problem of sampling rates higher than the clock rate is solved, enabling flexible sampling rate conversion and high-quality real-time waveform generation.

CN114144974BActive Publication Date: 2026-05-29ADVANTEST CORP

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ADVANTEST CORP
Filing Date
2019-12-23
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve parallel processing at sampling rates higher than the clock rate on digital signal processors, resulting in low efficiency.

Method used

It adopts a parallel processing architecture, which divides the input sample set into multiple subsets through sample distribution logic, provides multiple output samples in parallel, and performs interpolation or interpolation convolution operations using a hierarchical tree structure and multiple processing cores.

Benefits of technology

It achieves flexible sampling rate conversion higher than the clock rate of the digital signal processor, improving processing efficiency and supporting real-time waveform generation and high-quality sampling rate conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114144974B_ABST
    Figure CN114144974B_ABST
Patent Text Reader

Abstract

A signal processing apparatus (200) for providing a plurality of output samples (280) based on an input set of samples (250) comprises sample distribution logic (210) configured to provide a plurality of subsets of the input set of samples to a plurality of processing cores (220) performing processing operations associated with different time offsets (298), wherein the sample distribution logic comprises a hierarchical tree structure (240) having a plurality of hierarchy levels (240a, 240b, 240c), wherein each split node (230d, 230e, 230f) of a lowest hierarchy level (240c) is configured to provide two or more subsets from input samples of each split node of the lowest hierarchy level to the plurality of processing cores coupled to each split node of the lowest hierarchy level, wherein each split node of a given hierarchy level higher than the lowest hierarchy level is configured to provide two or more subsets from input samples of each split node of the given hierarchy level to a plurality of sub-trees coupled to each split node of the given hierarchy level, wherein each split node is configured to select each subset to coincide with a range of time offsets associated with the processing cores coupled to the respective sub-tree, and the plurality of processing cores are configured to perform the processing operations associated with the different time offsets in parallel to obtain the output samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to digital signal processing.

[0002] A further embodiment of the invention relates to real-time waveform generation on a digital signal processor. More specifically, it relates to real-time waveform generation on a digital signal processor, wherein the rate at which the data is processed is higher than the clock speed of the digital signal processor, and therefore a parallel processing architecture is employed.

[0003] Embodiments of the present invention relate to parallel interpolation digital convolutions. Background Technology

[0004] Interpolation describes an upsampling and filtering process that produces an approximation of a sequence, where the output samples are generally denser than the input samples, hence the name "interpolated" convolution.

[0005] An interpolator, or interpolation convolutional interpolator, convolves an input waveform given by equidistant sampling with a continuous-time impulse response, and produces the result of this operation at its output with different samples, which can be equidistant or non-equidistant. Interpolators exhibit an algorithmic architecture suitable for convenient implementation on application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). Traditionally, the most commonly used interpolator is the Faro interpolator. The impulse response of a Faro interpolator is described using a piecewise polynomial approach.

[0006] The traditional implementation of interpolated digital convolution on a sequential digital signal processor (DSP) is based on Faroe's work and is summarized as follows. The time accumulator operates at Δ... t The Faro interpolator accumulates fractional samples incrementally. When the time accumulator overflows, it requests an input sample. The most recent input sample and several previous input samples are stored in the input register. The stored input samples are fed into the finite impulse response (FIR) kernel. The coefficients of the FIR kernel determine the continuous-time convolution kernel in a piecewise polynomial manner, and thus also determine the interpolator's response. The result of the FIR operation is used as the coefficients of the polynomial in the polynomial evaluator. The polynomial is evaluated with the fractional part of the accumulated time as the independent variable. The Faro interpolator processes one sample at a time and produces one output sample per clock cycle, thus the parallelism of the standard Faro implementation is 1. The Faro interpolator only supports sequential digital processing.

[0007] Whenever the sampling rate exceeds the clock rate of the digital signal processor, parallel processing operations need to be performed (e.g., on a common set of samples) while keeping the effort of sample distribution reasonably small.

[0008] This objective is addressed by the subject matter of the independent claims. Summary of the Invention

[0009] One embodiment of the present invention (see, for example, claim 1) is a digital signal processing apparatus, such as an interpolator or interpolation convolution, for processing based on a set of input samples or input values, for example, 2 P + M Two samples, providing multiple output samples, or output values, in parallel, for example, from... P The core provided by the Faro P One output sample.

[0010] Digital signal processing apparatus includes sample distribution logic or structure configured to provide multiple subsets of an input sample set to multiple processing cores, such as interpolation cores or Faroe cores, which perform processing operations associated with, for example, different time offsets relative to a reference time, where the reference time is, for example, the time associated with the input sample.

[0011] The sample distribution logic includes a hierarchical tree structure with multiple levels of splitting nodes.

[0012] Each segmentation node at the lowest level is configured to provide two or more subsets of the input samples from each segmentation node at the lowest level to multiple processing cores coupled to each segmentation node at the lowest level.

[0013] Furthermore, each segment node at a given level higher than the lowest level is configured to provide two or more subsets from the input samples of each segment node at that given level to multiple subtrees coupled to each segment node at that given level.

[0014] Furthermore, each split node is configured to select each subset in a manner consistent with the range of time offsets associated with the processing cores coupled to the corresponding subtree, for example, such that the first subset is offset relative to or the same as the second subset, depending on the range of time offsets associated with the processing cores of the first subtree.

[0015] The digital signal processing apparatus also includes multiple processing cores configured to perform processing operations associated with different time offsets, such as interpolation operations or interpolated digital convolution operations, in parallel to obtain output samples.

[0016] In other words, for example P Each output sample is generated by P Each processing core or FARO core is provided. Each processing core receives, for example, from the sample distribution logic. M The output sample distribution logic consists of a hierarchical tree structure composed of multiple levels of segmentation nodes.

[0017] Each segmentation node is configured to provide two or more subsets of the input samples for a given segmentation node. Each segmentation node at a given level receives input samples from the next higher-level segmentation node and feeds its output subset of input samples to the next lower-level segmentation node.

[0018] The input samples for the sample distribution logic, for example P + M A single sample is the input to the highest-level segmentation node, while the output subset of the sample distribution logic is, for example... M A subset of samples is the output subset of the splitting nodes at the lowest level.

[0019] According to an embodiment (see, for example, claim 2), the input sampling rate of the input samples of the digital signal processing device is lower than or equal to the target output sampling rate of the output samples of the digital signal processing device.

[0020] Digital signal processing devices are configured to provide output sampling that is generally more dense than the input sampling.

[0021] The following are some typical, but not limiting, use cases and / or applications of this property of digital signal processing devices:

[0022] - Flexible (or virtually arbitrary) sampling rate conversion, where the target sampling rate is greater than or equal to the source sampling rate, and / or

[0023] - Digital delay with subsample resolution, a special case of flexible (or nearly arbitrary) sampling rate conversion when the target rate equals the source rate, and / or

[0024] – Pulse shaping for digital pattern generation, and / or

[0025] – Introducing timing jitter, for example, for controlled signal conditioning in measuring instruments, and / or

[0026] Timing error compensation for interleaved digital-to-analogue converters (DACs).

[0027] In a preferred embodiment (see, for example, claim 3), the digital signal processing apparatus includes a time accumulator configured to track a time offset and trigger the acquisition of a new input sample in an input register. The input register is coupled to sample distribution logic, for example, via a selection block. Whenever the time offset overflows a predetermined multiple of the sampling period of the input sample, for example… P When this happens, it will trigger the acquisition of new input samples.

[0028] Time accumulator with P ×Δ t The increment is in the half-open interval [0: P The accumulated score sample in ) . Whenever the accumulator overflows, it requests, for example P One input sample.

[0029] According to the embodiments (see, for example, claim 4), the number of samples in the input sample sets of multiple segmentation nodes at the same level of the sample distribution logic is the same, and / or the number of samples in each subset of the input samples provided as output samples by multiple segmentation nodes at the same level of the sample distribution logic is the same.

[0030] For example, the number of samples in the input sample set and the number of samples in the output sample set of the first segmentation node are equal to the number of samples in the input sample set and the number of samples in the output sample set of the second segmentation node at the same level.

[0031] A sample distribution logic in which splitting nodes at the same hierarchical level have an equal number of input samples and an equal number of output subsets of the input samples, the subsets containing an equal number of samples, the sample distribution logic having a modular structure with hierarchical levels constructed from the same modules, which makes the production and / or planning of the sample distribution logic simpler, cheaper and / or faster.

[0032] In a preferred embodiment (see, for example, claim 5), the number of samples in the input sample set of a given segmentation node is greater than the number of samples provided as input samples to the next lower-level segmentation node or to each sample subset provided to the processing core.

[0033] A given split node divides the input sample into two or more sets or subsets of input samples with equal sample sizes and provides them as output samples. These two or more subsets of input samples can overlap.

[0034] The number of input samples for a given segmentation node is greater than the number of samples in any subset of the output of that given segmentation node. The output subset of a given segmentation node contains an equal number of samples, which are provided as the input sample set for the next lower-level segmentation node or as the input sample set for the processing core.

[0035] According to an embodiment (see, for example, claim 6), the sample distribution logic is configured such that the number of samples provided to each subset of the segmentation nodes as input samples by each segmentation node of the next level gradually decreases as the level decreases.

[0036] The sample distribution logic is a series of segmentation nodes, where each segmentation node receives an output subset as input samples from a higher-level segmentation node and feeds the output subset to two or more lower-level segmentation nodes.

[0037] The lowest-level split nodes provide two or more subsets of output to their respective two or more processing cores.

[0038] According to the tree structure of the sample distribution logic, from top to bottom, the number of input samples of the segmentation nodes at different levels decreases, and the number of samples in the output subset of the segmentation nodes at increasingly lower levels also decreases.

[0039] According to an embodiment (see, for example, claim 7), the number of input samples for each segmentation node and / or the number of samples in each subset of input samples provided as output samples by each segmentation node is based on the number of samples in a subset of the input sample set provided to a single processing core (e.g., denoted as...). M ), and / or based on the hierarchical level of each segmentation node (e.g., represented as h ), and / or based on the number of processing cores (e.g., expressed as P Factorize into integer factors (e.g., represented as ) p k Factorization of ).

[0040] There exists a relationship between the number of input samples and the number of output samples for a given segmentation node. This relationship depends on the hierarchical level of the segmentation node, the number of input samples for each processing core, and an integer factor of the number of processing cores. For example, defining this relationship through an equation provides a clear and direct understanding of the segmentation node and / or the entire sample distribution logic.

[0041] In a preferred embodiment (see, for example, claim 8), the number of subsets of input samples provided by each segmentation node depends on the number of processing cores (e.g., denoted as...). P Factorize into integer factors (e.g., represented as ) p k Factorization of ).

[0042] For example P Integer factors of a number are not necessarily prime factors, thereforeP Depend on Description. In this equation, P Indicates the number of processing cores. k Represents 0 to ( H 1 The runtime variables between ) and H This represents the total number of factors in the selected integer factorization.

[0043] A given splitting node divides the set of input samples into subsets, where subsets may overlap. The number of subsets of the input sample set provided by a given splitting node depends on the number of processing cores. P Integer factors p k .

[0044] Since the number of subsets provided by a given split node depends on an integer factor of the number of processing cores, the number of hierarchical levels is an integer.

[0045] Segmentation nodes at the same level have the same number of samples in the input sample set and provide an equal number of subsets, each containing an equal number of samples.

[0046] According to an embodiment (see, for example, claim 9), the number of subsets of input samples provided by each segmentation node at a given hierarchical level is, for example, expressed as: p h And it represents an integer factor of the number P of processing cores. one.

[0047] p h It is the number of processing cores. P Integer factors (not necessarily prime factors) An element of the set, thus P Depend on The description is as described above.

[0048] p h In h This indicates the hierarchical level of each segmentation node. The lowest hierarchical level is determined by... h = 0 description, and h It increases as the level increases.

[0049] In a preferred embodiment (see, for example, claim 10), the number of input samples for each segmentation node is based on the following equation:

[0050]

[0051] In this equation, This indicates the number of input samples.

[0052] Indicates the number of processing cores P Integer factors of a number are not necessarily prime factors, therefore As mentioned above,

[0053] h This represents the hierarchical level of each splitting node, where the lowest hierarchical level is determined by... h = 0 description, and h It increases in size as the level increases, and

[0054] M This represents the number of samples in a subset of the input sample set provided to a single processing core.

[0055] In a preferred embodiment (see, for example, claim 11), the number of samples in each subset of the input samples provided as output samples by each segmentation node is based on the following equation:

[0056]

[0057] In this equation, This represents the number of samples in each subset of the input samples provided by each segmentation node as the output sample.

[0058] This represents the number of subsets of input samples provided by each segmentation node at a given hierarchical level.

[0059] Indicates the number of processing cores P Integer factors, but not necessarily prime factors, therefore As mentioned above,

[0060] h This represents the hierarchical level of each splitting node, where the lowest hierarchical level is determined by... h = 0 description, and h It increases in size as the level increases, and

[0061] M This represents the number of samples in a subset of the input sample set provided to a single processing core.

[0062] In a preferred embodiment (see, for example, claim 12), each segmentation node is configured to assign samples from the input sample set to multiple subtrees or processing cores, wherein each segmentation node at each level of the sample distribution logic is configured to select samples from the input samples and / or output samples, such that identical or different consecutive subsets of the input samples, starting from the same or different sample indices, are provided to each subtree or processing core. Furthermore, the starting index of the subset of input samples provided to each subtree depends on the level of each segmentation node. h And / or depends on the number of processing cores. P The integer factors chosen for factorization p k And / or depending on the time offset And / or depends on the time information assigned to the input sample set. .

[0063] A given split node provides two or more subsets of the set of input samples provided to that split node. These subsets of input samples can overlap, meaning that the same sample can be included in two or more subsets of the input sample set. The different subsets of input samples begin with different sample indices, which are provided to each subtree or processing core.

[0064] Starting with different sample indices results in unequal subsets of the input samples, where one sample may be contained in more than one subtree of the input sample set. The starting index of a subset of the input sample set is provided for each subtree and / or computed by a given split node. An equation with defined starting indices and / or computed starting indices of subsets of the input sample set will provide a reproducible subset of the input sample set.

[0065] According to an embodiment (see, for example, claim 13), an index is provided having each segmentation node. i The starting index of the input sample subset of the subtree is based on the following equation:

[0066] .

[0067] In this equation, Indicates that it is provided to the index as i The starting index of the input sample subset of the subtree, where i = 0 refers to the first subtree.

[0068] This represents the number of subsets of input samples provided by each segmentation node.

[0069] W by Description, in which expressP Integer factors of a number are not necessarily prime factors, therefore As mentioned above,

[0070] h This represents the hierarchical level of each splitting node, where the lowest hierarchical level is determined by... h = 0 description, and h It increases in size as the level increases.

[0071] Indicates the largest integer less than or equal to the parameter.

[0072] This indicates the time information assigned to the input sample set, and

[0073] This indicates a time offset, such as the time offset between samples provided by adjacent processing cores.

[0074] In a preferred embodiment (see, for example, claim 14), each segmentation node at each level is configured to be based on the temporal information of the input samples assigned to each segmentation node. And / or based on the hierarchical level of each segmentation node. h And / or based on the number of processing cores P Integer factors chosen for factorization p k and / or based on time offset This is used to assign time information to each subtree.

[0075] The temporal information of the input samples assigned to each segmentation node is used to calculate the starting index of a subset of the input sample set. The temporal information depends on the hierarchical level of a given segmentation node, and / or on an integer factor of the number of processing cores and / or on the time offset.

[0076] According to an embodiment (see, for example, claim 15), the index assigned to each segmentation node is: i The time information of the subtree, for example, represented as It is based on the following equation:

[0077]

[0078] In this equation, Indicates that the index is assigned to i The time information of the subtree, where i = 0 refers to the first subtree.

[0079] W is derived from the equation discussed above. describe,

[0080] Indicates the largest integer less than or equal to the parameter.

[0081] This indicates the time information assigned to the input sample set, and

[0082] This indicates a time offset, such as the time offset between samples provided by adjacent processing cores.

[0083] In a preferred embodiment (see, for example, claim 16), the digital signal processing apparatus includes an input register configured to store a plurality of input samples.

[0084] Storing samples in the input register allows for the selection of the set of samples to be distributed to processing cores by the dispatch logic. A sample can be selected and / or distributed to one or more processing cores multiple times.

[0085] In a preferred embodiment (see, for example, claim 17), the input register is a shift register.

[0086] Since only a finite number of input samples need to be stored, a shift register is sufficient. Shift registers are a feasible solution for storing a finite number of samples; they are widely used, simple to use, and inexpensive.

[0087] According to an embodiment (see, for example, claim 18), a digital signal processing apparatus includes a selector configured to select a set of input samples for a sample distribution logic from a plurality of input samples.

[0088] The selector selects a set of samples from a plurality of input samples stored in the input register that will be distributed to the processing core by the sample distribution logic, resulting in pre-selection of the input samples.

[0089] In a preferred embodiment (see, for example, claim 19), if timing jitter is applied, the length of the time offset, for example, the length of the time offset for segment nodes at the same level and / or for segment nodes at different levels, is equidistant or non-equidistant.

[0090] Because time offsets are associated with processing operations, the variability in the length of equidistant or non-equidistant time offsets can lead to variable processing operations performed with equidistant or non-equidistant time offsets. Non-equidistant time offsets can be used to compensate for timing errors present in interleaved high-speed DAC implementations.

[0091] In a preferred embodiment (see, for example, claim 20), the signal processing device performs interpolation between input samples.

[0092] Whenever the time offset overflows a predetermined multiple of the sampling period of the input sample in the time accumulator, the digital signal processing device acquires a new input sample and performs processing operations associated with different time offsets via multiple processing cores, outputting an output sample. The time offset associated with the processing operation is a fraction of the sampling period of the input sample, resulting in the output sample being an interpolated sample located between the input samples.

[0093] According to an embodiment (see, for example, claim 21), the digital signal processing apparatus performs convolution.

[0094] Since a given processing core performs a processing operation, it obtains multiple input samples and outputs a single output sample. The processing core performs a weighted average operation or a convolution operation, which provides a single output element from multiple input elements.

[0095] In a preferred embodiment (see, for example, claim 22), multiple processing cores implement a Faro structure. The Faro structure is a widely used interpolator implementation, making it an easy-to-apply, readily available, and cost-effective solution.

[0096] According to an embodiment (see, for example, claim 23), the construction of different subtrees is based on the number of processing cores. P Integer factors p k The results are derived from the same or different choices.

[0097] As an example, when P When = 16, the number of cores processed for a portion of the tree can be factored into 16 = (2). The number of cores processed for the tree and / or for a different portion of the tree can be factored into 16 = (4 × 2) × 2.

[0098] According to an embodiment (see, for example, claim 24), the construction of different subtrees is based on the number of processing cores. P Integer factors p k The results are derived from the same or different sorting.

[0099] As an example, when P When the number of cores is 16, the number of cores processed can be factored into 16 = 2 for a portion of the tree and / or factored into 16 = 4 for a different portion of the tree.

[0100] A corresponding method was created according to a further embodiment of the present invention.

[0101] However, it should be noted that these methods are based on the same considerations as the corresponding apparatus. Furthermore, these methods may be supplemented by any features and / or functions and / or details described herein with respect to the apparatus, either individually or in combination. Attached Figure Description

[0102] In the following, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings, in which:

[0103] Figure 1 A block diagram of a signal processing apparatus is shown, which includes sample distribution logic and multiple processing cores;

[0104] Figure 2 A block diagram of a signal processing device is shown, which is expanded to include a time accumulator, an input register, and a selector.

[0105] Figure 3 A block diagram of the segmentation nodes of the sample distribution logic is shown;

[0106] Figure 4 An exemplary block diagram of a segmentation node is shown, which provides two output subsets of input samples with their own temporal information;

[0107] Figure 5 A block diagram of the Faro interpolator is shown;

[0108] Figure 6 An exemplary block diagram of an extended signal processing device is shown;

[0109] Figure 7 Another exemplary block diagram of the extended signal processing device is shown;

[0110] Figure 8 Another exemplary block diagram of the extended signal processing device is shown. Detailed Implementation

[0111] Different inventive embodiments and aspects will be described below. Further embodiments will be defined by the appended claims.

[0112] It should be noted that any embodiment defined by the claims may be supplemented by any details, features, and / or functions described herein. Furthermore, the embodiments described herein may be used alone or optionally supplemented by any details and / or features and / or functions included in the claims.

[0113] Furthermore, it should be noted that the individual aspects described herein can be used individually or in combination. Thus, details can be added to each of the individual aspects without adding details to the other aspect.

[0114] It should be noted that this disclosure explicitly or implicitly describes features that can be used in signal processing apparatuses. Therefore, any feature described herein can be used in the context of signal processing apparatuses.

[0115] Furthermore, the features and functions disclosed herein related to the method can also be used in apparatuses configured to perform such functions. Additionally, any features or functions disclosed herein regarding the apparatus can also be used in the corresponding method. In other words, the method disclosed herein can be supplemented by any features and functions described regarding the apparatus.

[0116] The invention will be more fully understood through the following detailed description and the accompanying drawings of embodiments thereof; however, the detailed description and drawings should not be construed as limiting the invention to the specific embodiments described, but are merely for illustrative and understanding purposes.

[0117] according to Figure 1 Implementation examples

[0118] Figure 1 A block diagram of a digital signal processing apparatus 100 is shown, which includes sample distribution logic 110 and multiple processing cores 120. The sample distribution logic 110 includes multiple segmented nodes 130a-f organized into a hierarchical tree structure 140 with multiple hierarchical levels 140a-c.

[0119] Input samples 150 from the digital signal processing apparatus 100 are provided as input samples to sample distribution logic 110, wherein input samples 150 are provided to segmentation nodes 130a at the highest level 140a. Segmentation node 130a takes input samples 150 as input and provides two or more subsets 160a, 160b from the input samples 150. The number of samples in subsets at the same level, such as subsets 160a-b at level 140a or subsets 160c-f at level 140b, is the same. These subsets, such as subsets 160a, 160b, are then assigned to different segmentation nodes at the next lower level 140b, such as 130b, 130c.

[0120] Any given segment node 130a-f obtains a set of input samples from the next higher level, for example, segment node 130c obtains input sample 160b from segment node 130a at level 140a, and provides two or more subsets, such as 160c, 160d, to two or more segment nodes at the next lower level, such as 140c (segment node 130f is shown).

[0121] The sample distribution logic has a hierarchical tree structure 140 of segmentation nodes 130a-f, where the highest-level segmentation node 130a receives the input sample 150, and each other segmentation node 130b-f receives the input sample set from the next higher-level node. Segmentation nodes 130d-f at the lowest level 140c are coupled to two or more processing cores, and each other segmentation node 130a-c of the sample distribution logic 110 is coupled to two or more segmentation nodes 130b-f at the next lower-level node.

[0122] Multiple processing cores 120 include processing cores 120a-f, whose inputs are coupled to segmentation nodes 130d-f of the lowest level 140c of the sample distribution logic 110. A processing core 120b is coupled to a single segmentation node 130d of the lowest level 140c of the sample distribution logic 110, wherein the segmentation node 130d of the lowest level 140c of the sample distribution logic 110 is coupled to two or more processing cores 120a, 120b of the digital signal processing apparatus 100. A set of input samples 125b for a given processing core 120b is provided by segmentation nodes 130d of the lowest level 140c of the sample distribution logic 110 coupled to that given processing core 120b. Any given processing core 120a-f is configured to provide a single output sample 180a-f from each set of input samples 125a-f. The multiple processing cores 120 perform processing operations in parallel to provide multiple output samples 180, wherein the processing operations are associated with different time offsets.

[0123] In other words, the digital signal processing apparatus 100 includes multiple processing cores 120 and sample distribution logic 110, and is configured to provide multiple output samples 180 from an input sample set 150. The multiple processing cores 120 perform processing operations in parallel, wherein processing cores 120a-f are associated with different time offsets. The input sample sets 125a-f of the processing cores 120a-f are provided by the sample distribution logic 110.

[0124] The sample distribution logic 110 provides a subset 125a-f of the input sample set 150 by using a hierarchical tree structure 140 of segmentation nodes 130a-f organized into hierarchical levels 140a-c.

[0125] Input sample 150 is distributed into subsets 125a-f, which are fed as input into processing cores 120a-f, wherein the number of samples in subsets 125a-f is equal for all subsets 125a-f.

[0126] Each level 140a-c of the sample distribution logic 110 includes segmentation nodes 130a-f, wherein the segmentation node 130a-f of a given level 140a-c obtains a set of input samples from the next higher level and provides two or more subsets 160a-d, 125a-f for the next lower level 140a-c.

[0127] The digital signal processing device 100 or parallel interpolation digital convolution 100 described herein can be used as a key building block of a signal processor application-specific integrated circuit (ASIC) and / or as part of other instruments.

[0128] The digital signal processing apparatus described in this paper can be applied on a parallel DSP to address flexible (or almost arbitrary high) sampling rates with real-time or near-real-time response times; for example, the digital signal processing apparatus can handle sampling rates of 100 GSa / s in real time. It is an area-efficient implementation of an architecture with a parallel processing core.

[0129] Furthermore, this signal processing device can be used to provide high-quality, flexible (or nearly arbitrary) sample rate conversion in real time for radio frequency (RF) and analog baseband applications. The usable bandwidth can be, for example, 75% of the Nyquist rate, and image suppression, for example, can be achieved at 60 dB. The conversion ratio is not explicitly limited to some simple fraction, but is truly flexible (or nearly arbitrary) because it is programmed as a number, for example, between 0 and 1, with 64-bit resolution. It can handle sample rates far exceeding the clock rates of DSPs.

[0130] In addition, signal processing devices can be used to provide pulse shaping for the generation of non-return-to-zero (NRZ) digital waveforms and / or pulse-amplitude modulation (PAM) digital waveforms to obtain flexible (or virtually arbitrary) user bit rates.

[0131] In cases of non-equidistant sampling, signal processing devices can also be used to provide memory-based timing jitter injection.

[0132] An important use case is to provide fractional subsample latency for time-to-digital (TDC) based synchronization mechanisms.

[0133] according to Figure 2 Implementation examples

[0134] Figure 2 A schematic block diagram or high-level block diagram of a signal processing device 200 is shown. Figure 1An enhanced or extended version of the digital signal processing device 100. The input of the digital signal processing device 200 is coupled to an input register 270, which is a shift register. The input register 270 has one input and one output, where the input is also an input to the digital signal processing device 200, and the output of the input register 270 is coupled to a selector 290.

[0135] Selector 290 has two inputs and one output. The first input of the selector is coupled to input register 270, and the second input of selector 290 is coupled to time accumulator 295. The output of selector 290 is coupled to sample distribution logic 210, which is related to… Figure 1 The sample distribution logic 110 is similar. The time accumulator 295 is configured to trigger new samples for the digital signal processing device 200 and is coupled to the selector 290 and the sample distribution logic 210.

[0136] and Figure 1 The sample distribution logic 110 is similar to the sample distribution logic 210, which includes a hierarchical tree structure of segmentation nodes 230a-f organized into multiple hierarchical levels 240.

[0137] The input of the segment node 230a on the highest level 240a of the sample distribution logic 210 is the input of the sample distribution logic 210 and is coupled to the selector 290. The segment node 230a has two or more outputs, which are coupled to different segment nodes 230b-c on the next lower level (e.g., level 240b).

[0138] Any segmentation node 230a-f of the sample distribution logic 210 has one input and two or more outputs. The input of a given segmentation node 230a-f is coupled to another segmentation node 230a-f at the next higher level 240a-c, and the output of a segmentation node 230a-f is coupled to a different segmentation node 230a-f at the next lower level 240a-c.

[0139] The set of output samples from the lowest-level segmentation node 230d-f of the sample distribution logic 240c is the set of output samples from the sample distribution logic 210. The lowest-level segmentation node 230d-f of the sample distribution logic 210 is coupled to two or more processing cores 220a-f among the multiple processing cores 220, which is related to... Figure 1 The multiple processing cores are similar to 120.

[0140] Any of the processing cores 220a-f, such as processing core 220b, has one input and one output. Processing cores 220a-f expect a set of input samples from the coupled segmentation nodes 230a-f as input and provide a single output sample 280a-f. The single output sample 280a-f is the output sample 280 of the signal processing device 200.

[0141] In other words, as Figure 1 The extended version of the digital signal processing device 100, the digital signal processing device 200, includes the digital signal processing device 100 and is extended by an input register 270, a selector 290, and a time accumulator 295.

[0142] The time accumulator 295 is configured to track the time offset, and whenever the time offset overflows by a predetermined multiple of the sampling period of the input sample (e.g., ...), P When ), it triggers the acquisition of a new input sample 250 in the input register 270.

[0143] Input register 270 is a shift register configured to store multiple input samples 250, for example, 2 P + M Two samples are coupled to sample distribution logic 210 via selector block 290.

[0144] Selector 290 is coupled to both input register 270 and sample distribution logic 210, and is configured to select a set of input samples for sample distribution logic 210 from the input samples stored in input register 270.

[0145] The input sample of the sample distribution logic 210 selected by selector 290 is the input sample of the first segmentation node 230a in the first level 240a, accompanied by time information. Each segmentation node 230a-f on each level 240a-c is configured to assign time information to each subtree or subset of the input sample, wherein the time information is based on the time offset tracked by time accumulator 295.

[0146] Each segmentation node 230a-f of the sample distribution logic 210 is configured to divide the input sample set into subsets and provide these subsets as outputs to the next lower-level segmentation node 230a-f.

[0147] Furthermore, each segmentation node 230a-f on each level 240a-c is configured to assign time information 298 to each subtree based on the time information of the input samples assigned to each segmentation node 230a-f, and / or based on the level 240a-c of each segmentation node 230a-f, and / or based on an integer factor of the number of processing cores 220a-f, and / or based on a time offset 298.

[0148] If timing jitter is applied, the length of the time offset 298 tracked by the time accumulator 295 can be equidistant or non-equidistant.

[0149] The lowest level 240c segmentation node 230d-f supplies to the processing core 220a-f coupled to the given segmentation node 230d-f so that the processing core 220a-f provides output samples 280a-f.

[0150] Each processing core 220a-f, such as the Faroe core, receives input samples stored in the input register 270. M A subset of samples, which is pre-selected by selector 290 and distributed by, for example, an area-efficient implementation of distribution logic 210.

[0151] The digital signal processing device 200 performs the same and / or similar mathematical operations as, for example, a Faro interpolator, but processes multiple operations at a time in each clock cycle (e.g., P (Number) samples. It generates one sample per clock cycle. P It has a time-continuous output sample size, therefore its parallelism is greater than 1. In this particular embodiment where each segmentation node has two outputs, the digital... P It is a power of 2.

[0152] Multiple processing cores include P Each core is an identical processing core, or Faro core. Each core includes an FIR filter core and a polynomial evaluator used in the Faro core or in the Faro implementation.

[0153] Time accumulator 295 P ×Δ t The increment accumulates fractional samples in the half-open interval [0; P). Whenever the time offset overflows by a predetermined multiple, for example... P At that time, the time accumulator requests or triggers the acquisition. P There are 250 input samples. The input samples are stored in input register 270, which can store 2... P + M Two samples, including P current sample and P + M Two past samples. From these two... P + M In the two samples, selector 290 selects... P + M One sample serves as the set of input samples for sample distribution logic 210. Sample distribution logic 210... P + M One input sample in P The processing cores 220a-f are distributed among themselves, with each processing core 220a-f being fed... M One sample. Multiple processing cores 220a-f include P It is the same processing core or Faro core.

[0154] Each processing core or Faro core includes an FIR filter core and a multinomial evaluator for the Faro implementation. Each such core obtains M Input samples and calculate P One of the output samples 280a-f.

[0155] Sample distribution is performed in two stages: selection or pre-selection and segmentation. The selection process, performed by selector 290, involves picking samples from input register 270 that meet the criteria for further processing. P + M A continuous subrange of one sample. This selection is based on the time accumulator in the closed interval [0; P The integer part of the accumulated time offset in 1].

[0156] The segmentation phase divides the selected sub-ranges in such a way that each processing core 220a-f or Faroe core 220a-f receives M The correct consecutive runs for each input sample. P = 2 H In this case, the segmentation process involves a hierarchical structure 240, which is of height 240. H 1 A perfect binary tree. Therefore, this process involves... H Each level, at each level h Above P / 2 h+1 There are 1 "segmentation nodes", among which h = 0 ... H 1 Lowest level h= 0 of 2 H 1 Each split node generates P There are sets, each set has M These are samples. P The correct or required number of samples for each processing core.

[0157] Figure 3 The text describes hierarchical levels. h General operations on "split nodes". Figure 4 The example implementation of a "split node" is given, which is part of a perfect binary tree described in the previous paragraphs (i.e., where...). P = 2 H And for all k = 0 ... H 1 , p k = 2).

[0158] according to Figure 3 Segmentation nodes

[0159] Figure 3 A schematic block diagram of the segmentation node 300 is shown, which is similar to... Figure 1 The segmentation node 130. The input of the segmentation node 300 includes input sample 310 and time information 320. The segmentation node 300 provides two or more subsets 360a-c of the input sample 310 with their respective associated time information 350a-c.

[0160] Given level h The segmentation node 300 is configured to divide the input sample set 310 into multiple subsets 360a-c of the input samples 310. Subsets 360a-c have the same number of samples, for example... W + M One sample, of which W Depend on Description, in which p k Integer factors representing the number of processing cores.

[0161] By starting from the starting index, which depends on the time information 320 provided to the segment node 300, a selection is made that includes... W + M A subset of input sample 310, from 1 sample p h W + M Selecting a subset from 310 input samples W + M One sample. The index provided to each segmentation node is... i The starting index of the subset of input samples of the subtree is based on the following equation:

[0162]

[0163] in 320 indicates the time information associated with the input sample.

[0164] Furthermore, segment node 300 is configured to associate time information 350a-c with subsets 360a-c provided by the given segment node 300. The time information 350a-c associated with subsets 360a-c depends on the time information 320 provided to segment node 300, on the given hierarchical level of segment node 300, and on... Figure 1 The number of processing cores 120 is an integer factor.

[0165] Time information 350a-c is based on the following equation:

[0166]

[0167] Figure 3 It shows in Figure 1 A general block diagram of the segmentation node 300 used in the digital signal processing device 100. The segmentation node 300 is in... Figure 1 The sample distribution logic 110 is organized into a hierarchical tree structure in order to... Figure 1 The input sample 150 is divided into subsets of input samples with equal sample sizes; these subsets are used as... Figure 1 The input samples or sets of input samples of multiple processing cores 120.

[0168] according to Figure 4 Segmentation nodes

[0169] Figure 4 The illustration shows the segment node 400, which is Figure 3 A more general example of a segmentation node 300, wherein the segmentation node 400 takes an input sample set 410 and time information 420 as input, and provides two sets of output samples 430a, 430b with their respective time information 440a, 440b. Figure 4 A specific example is when the number of processing cores is a power of 2 (i.e., P = 2 H And this number is based on In all p kA part of the binary tree structure generated when factoring is performed with a factor of 2.

[0170] Each subset 430a, 430b is configured to contain selections from input sample 410, starting from different indices. W + M -1 samples, where the starting index depends on the time information 420.

[0171] Figure 4 The logic for sample distribution (e.g.) is shown. Figure 1 110 or Figure 6 The segmentation node 400 in (660) is used to segment the input sample 410 (which is the next higher level). Figure 1 The output subset of the segmentation node 130 is divided into two output subsets 430a and 430b. These two output subsets have an equal number of samples and have associated time information 440a and 440b, respectively. The time information 440a and 440b are based on the input time information 420.

[0172] according to Figure 5 Traditional Faro interpolation device

[0173] Figure 5 A block diagram of a conventional Faro interpolator 500 is shown. The Faro interpolator 500 includes an input register 510, a time accumulator 520, and a Faro core 530.

[0174] Time accumulator 520 with Δ t The increment accumulates fractional samples in the half-open interval [0; 1). When the accumulator overflows, it requests an input sample of 540.

[0175] The most recent input sample 540 and previous input samples, for example, M One sample is stored in input register 510. The total number of input samples stored in input register 510 and used for interpolation calculations can be referred to as the support of the Faroe interpolator 500. M .

[0176] Input register 510 and time accumulator 520 are coupled to Faro core 530. Faro core 530 of Faro interpolator 500 produces one output sample 550 per clock cycle, while providing and / or requesting input sample 540 when time accumulator 520 overflows.

[0177] The Faroe core 530 includes multiple Finite Impulse Response (FIR) cores 560 and a polynomial estimator unit 570. An input register 510 is coupled to each FIR core 560 of the Faroe core 530. Each FIR core 560 is coupled to the polynomial estimator 570. The polynomial estimator 570 takes input from the FIR cores 560 and a fractional-time input 580 from a time accumulator 520, and provides an output sample 550 each clock cycle, which is the output of the Faroe interpolator 500.

[0178] The time accumulator 520 accumulates fractional time 580 and provides it to the polynomial evaluator 570 of the Faroe core 530. When the time accumulator 520 overflows, it requests a new input sample 540. The new input sample 540 is stored in the input register 510, which is a shift register. The input register 510 stores the new input sample 540 and the previous input sample, for example... M -1 input sample. The set of input samples, for example, stored in input register 510. M Each input sample is fed to the Faro core 530, and specifically to the FIR core 560 of the Faro core 530. Each FIR core 560 is calculating a weighted average of the input samples stored in the input register 510, wherein the FIR cores may have different weights and / or different coefficients for the weighted average calculation. The weighted average provided by the FIR core 560 is provided to the polynomial evaluator 570. Using the weighted average calculated by the FIR core 560 as the coefficient values ​​of the polynomial and using the fractional time value 580 provided by the time accumulator 520 as the independent variable of the polynomial, the polynomial evaluator 570 calculates the value of the polynomial and outputs this value as an output sample 550, which is the output of the Faro core 530 and / or the output of the Faro interpolator 500.

[0179] The Faro interpolator 500 is a traditional interpolator that processes one sample at a time, with a parallelism of 1. Figure 1 Digital signal processing device 100 relative to Figure 5 The novelty of the conventional Faro interpolator 500 lies in its ability to solve the high sampling rate of the digital signal processing device 100 on a parallel DSP in a real-time or near real-time manner: for example, Figure 1 The digital signal processing device 100 can handle a sampling rate of 100 gigabits per second in real time on a DSP with a clock speed of less than 1 gigahertz.

[0180] Figure 1 The digital signal processing device 100 includes multiple processing cores 120 for parallel processing, wherein Figure 1 The processing core 120 can be Figure 5The Faro core 530. Figure 1 The sample distribution logic 110 will Figure 1 The input value 150 is distributed to the values ​​used as input values. Figure 1 Multiple processing cores 120 and multiple Faro cores 530.

[0181] Furthermore, this signal processing device uses a single time accumulator, for example... Figure 2 The 295, instead of multiple 530 per processing core or Faroe core. Figure 5 The time accumulator 520 allows the Faro core 530 to perform processing operations in parallel. Figure 1 The digital signal processing device 100 includes Figure 1 The processing cores are 120, and they are Faro cores 530.

[0182] There can be various variations in the implementation, including:

[0183] – The processing core or Faro core does not have to follow the original Faro implementation. Any implementation that computes the output samples from zero or more input samples and fractional timing information is acceptable and can be used in signal processing devices; an example alternative is a polyphase FIR filter, where the coefficients are determined from the fractional timing information 580, for example, by mathematical relations, by lookup tables, or by a combination of both;

[0184] – The interpolation ratio does not necessarily have to be strictly greater than 1; it can be equal to 1.

[0185] – The interpolation ratio does not have to be a constant;

[0186] - Output samples do not have to be equidistant. Time accumulators or timing accumulators and segmentation logic or sample distribution logic are allowed to generate non-equidistant time points;

[0187] - Parallelism or the number of processing cores P It is not limited to integer powers of 2, although the latter may produce the most efficient implementation.

[0188] – Individual switches in the “segmentation” or sample distribution phase can be combined (see [link]). Figure 7 ).

[0189] – Consider using different intervals to represent time accumulation or fractional timing information, for example [ 0.5; P 0.5), [ 0.5; 0.5) or [ 1; 1).

[0190] The following provides a specific example of a digital signal processing apparatus in which the number of segmentation nodes in the sample distribution logic and / or the number of input samples, and / or the number of processing cores, and / or the number of faro cores, may be different.

[0191] according to Figure 6 Implementation examples

[0192] Figure 6 A specific digital signal processing device 600 is shown, which is Figure 1 An example of a digital signal processing apparatus 100. Digital signal processing apparatus 600 includes a time accumulator 610 configured to trigger the acquisition of new input samples 620, which are stored in an input register 630. The input register 630 is coupled to a selector unit 640, which provides input samples to a first segmentation node 650. The segmentation node 650 is the first segmentation node of a hierarchical tree structure 660 of segmentation nodes, which in this example is a binary tree. Each segmentation node in the binary tree structure 660 has one input and two outputs, where the input samples 670 of a given segmentation node are partitioned into subsets 680a, 680b of the input samples 670. The hierarchical tree structure—in this case, a binary tree structure 660—provides an equal number of input samples to a processing core 690 or a Faro core 690. Each Faro core 690 provides a single output sample from a given set of input samples provided by the segmentation nodes at the lowest level of the binary tree structure 660.

[0193] In other words, when the time fraction Δ of the increment accumulates t or a multiple thereof 16×Δ t When the time accumulator 610 overflows, 16 new input samples are requested. These 16 new input samples, along with the previous input samples, are stored in the input register 630, which stores a total of 45 samples. The selector unit 640 selects 30 samples from the 45 samples stored in the input register and provides them as the input sample set to the first segmentation node 650. The first segmentation node 650 provides two subsets of 22 samples each from the 30 samples in the input sample set. The segmentation nodes in the next lower level receive 22 input samples, and each of them provides two subsets of 18 samples each as output samples. The segmentation nodes in increasingly lower levels receive fewer and fewer samples as input samples, with the highest level receiving 30 samples as input samples, and subsequent segmentation nodes receiving 22, 18, and 16 samples as input samples for lower levels.

[0194] All samples in the subset provided by the segmentation node are provided as input samples for the segmentation node in the next hierarchical level. The first segmentation node 650 provides two subsets of 22 samples each from the 30 samples in the input sample set. Segmentation nodes in different hierarchical levels provide 22, 18, 16, and 15 samples respectively from their input sample sets. The segmentation node in the lowest level of the sample distribution logic or hierarchical tree structure 660 provides two subsets, each with 15 samples, as input samples for the processor core or Faro core 690. The Faro core 690 is similar to... Figure 5 The Faro Core 530 produces an output sample from a set of input samples (in this example, from 15 input samples).

[0195] according to Figure 7 Implementation examples

[0196] Figure 7 A digital signal processing device 700 is shown, as... Figure 1 A specific example of a digital signal processing apparatus 100. The signal processing apparatus 700 has a time accumulator 710 that triggers the acquisition of an input sample set 720, specifically 16 input samples. The new input samples, together with the previous input samples, for a total of 45 input samples, are stored in an input register 730. A selector unit 740 selects 30 from the 45 input samples and provides them as input samples to a segmentation node 750 or a first segmentation node. The first segmentation node is a segmentation node at the highest level of the hierarchical tree structure 760 of the segmentation node 750.

[0197] From the outside, Figure 6 Digital signal processing device 600 and Figure 7 The digital signal processing unit 700 in the middle can perform the same calculation. The main difference lies in the factorization of the number of processing cores (2×2×2×2 vs. 4×2×2) (in Figure 6 and Figure 7 The values ​​are all 16, and the resulting different tree structures and different numbers of hierarchical levels of partitioning nodes, among which Figure 6 The hierarchical tree structure 660 is a binary tree, while the hierarchical tree structure 760 has only three levels, where the splitting node at the lowest level provides four subsets of the input sample set.

[0198] The lowest-level segmentation nodes obtain the input sample set, with each set containing 18 input samples, and provide four subsets of the input samples, each containing 15 samples, to the four processing cores. Processing core 790 is a Faro core, and these Faro cores are... Figure 5 The Faro Core 530 is similar or identical to it, providing one output sample from each of the 15 input sample sets.

[0199] according to Figure 8 Implementation examples

[0200] Figure 8 An exemplary digital signal processing device 800 is shown, with Figure 1 Similar to the digital signal processing device 100. Due to the overflow of the time accumulator 810, 15 input samples are triggered. The 15 input samples 820, together with the previous input samples—a total of 43 samples—are stored in the input register 830.

[0201] Selector unit 840 selects 29 samples from these 43 samples as input samples to the first segmentation node. The segmentation nodes 850 of the digital signal processing device 800 are organized into a hierarchical tree structure 860. In this particular example, the number of processing cores... P It is not a power of 2, and the hierarchical tree structure of the segmentation nodes includes two levels, where the segmentation node 850 at the highest level provides five subsets of the input samples, each with 17 samples, while the segmentation node 850 at the second highest level—which is also the lowest level—provides three subsets of the input samples, each with 15 samples.

[0202] Fifteen samples were provided to multiple processing cores 890, or Faroe cores, similar to... Figure 5 The Faro core 530. Each Faro core 890 provides a single output sample from 15 input samples, so multiple Faro cores 890 provide 15 output samples 895.

[0203] References:

[0204] [Farrow88]CW Farrow, “A Continuously Variable Digital Delay Element,” Proceedings of the IEEE International Symposium on Circuits and Systems, Espoo, Finland, June 6-9, 1988, pp. 2641-2645

[0205] [Erup93] L. Erup, FM Gardner, RA Harris, “Interpolation in Digital Modems—Part II: Implementation and Performance,” IEEE Transactions on Communications, Vol. 41, pp. 998-1008, June 1993

Claims

1. A signal processing apparatus (100, 200, 600, 700, 800) for providing multiple output samples (180, 280, 550, 695, 895) based on an input sample set (150, 250, 310, 410, 540, 620, 720, 820), comprising: The sample distribution logic (110, 210, 660) is configured to provide multiple subsets (160a-f, 125a-f, 360a-c, 430a-b, 680a-b) of the input sample set to multiple processing cores (120, 120a-f, 220, 220a-f, 530, 690, 790, 890) that perform processing operations associated with different time offsets (298, 580). The sample distribution logic includes a hierarchical tree structure (140, 240, 660, 760, 860) with multiple levels (140a-c, 240a-c). Specifically, each segmentation node (130a-f, 230a-f, 300, 400, 650, 750, 850) at the lowest level (140c, 240c) is configured to provide two or more subsets from the input samples (150, 160a-d, 310, 410, 670, 680a-b) of each segmentation node at the lowest level to multiple processing cores coupled to each segmentation node at the lowest level. In this configuration, each segmentation node at a given level higher than the lowest level is configured to provide two or more subsets from the input samples of each segmentation node at the given level to multiple subtrees coupled to each segmentation node at the given level. Each segmentation node is configured to select each subset in a manner consistent with the range of time offsets associated with the processing core coupled to the corresponding subtree, and The multiple processing cores are configured to perform processing operations associated with different time offsets in parallel to obtain the output samples.

2. The signal processing apparatus according to claim 1, wherein, The input sample rate of the input sample is lower than or equal to the target output sample rate of the output sample.

3. The signal processing apparatus according to claim 1 or 2, comprising a time accumulator (295, 520, 610, 710, 810), the time accumulator being configured to: Track the time offset, and Whenever the time offset overflows by a predetermined multiple of the sampling period of the input sample, a new input sample is triggered to be obtained in the input registers (270, 510, 630, 730, 830) coupled with the sample distribution logic.

4. The signal processing apparatus according to claim 1 or 2, in, The number of samples in the input sample sets of multiple segmentation nodes at the same level is the same, and / or the number of samples in each subset of the input samples provided by multiple segmentation nodes at the same level is the same.

5. The signal processing apparatus according to claim 1 or 2, wherein, The number of samples in the input sample set of a given segmentation node is greater than the number of samples provided as input samples to the next lower-level segmentation node or to each sample subset provided to the processing core.

6. The signal processing apparatus according to claim 1 or 2, wherein, The sample distribution logic is configured such that the number of samples provided to each subset of the segmentation node as input samples from each segmentation node at the next higher level gradually decreases as the level decreases.

7. The signal processing apparatus according to claim 1 or 2, wherein, The number of input samples for each segmentation node and / or the number of samples in each subset of the input samples provided by each segmentation node are based on the number of samples in a subset of the set of input samples provided to a single processing core, and / or based on the hierarchical level of each segmentation node, and / or based on factorization that decomposes the number of processing cores into integer factors.

8. The signal processing apparatus according to claim 1 or 2, wherein, The number of subsets of input samples provided by each splitting node depends on the factorization of the number of processing cores into integer factors.

9. The signal processing apparatus according to claim 1 or 2, wherein, The number of subsets of input samples provided by each splitting node at a given hierarchical level is equal to in express P Integer factors, according to in P Indicates the number of processing cores. H This represents the total number of factors in the chosen integer factorization, and h This indicates the hierarchical level of each segmentation node.

10. The signal processing apparatus according to claim 1 or 2, wherein, The number of input samples for each segmentation node is based on the following equation: in This indicates the number of input samples. express P Integer factors, according to in P Indicates the number of processing cores. H This represents the total number of factors in the selected integer factorization. h This indicates the hierarchical level of each splitting node, and M This represents the number of samples in a subset of the input sample set provided to a single processing core.

11. The signal processing apparatus according to claim 1 or 2, wherein, The number of samples in each subset of the input samples provided by each segmentation node is based on the following equation: in This represents the number of samples in each subset of the input samples provided by each segmentation node. This represents the number of subsets of input samples provided by each segmentation node. express P Integer factors, according to , in P Indicates the number of processing cores. H This represents the total number of factors in the selected integer factorization. h This indicates the hierarchical level of each splitting node, and M This represents the number of samples in a subset of the input sample set provided to a single processing core.

12. The signal processing apparatus according to claim 1 or 2, wherein, Each splitting node is configured to assign samples from the input sample set to multiple subtrees or processing cores. In this context, each segmentation node at each level of the sample distribution logic is configured to select samples from the input samples, such that identical or different subsets of the input samples, starting from the same or different sample indices, are provided to each subtree or processing core. The starting index of the subset of input samples provided to each subtree depends on the hierarchical level of each splitting node, and / or on the integer factor chosen for the factorization of the number of processing cores, and / or on the time offset and / or time information assigned to the set of input samples (298, 320, 350a-c, 420, 440a-b, 580).

13. The signal processing apparatus according to claim 1 or 2, wherein, The index provided to each split node is i The starting index of the subset of input samples of the subtree is based on the following equation: in This indicates that the index is provided as i The starting index of a subset of input samples for a subtree, where the first subtree is... i =0 index, This represents the number of subsets of input samples provided by each segmentation node. W by describe, in express P Integer factors, according to , in P Indicates the number of processing cores. H This represents the total number of factors in the selected integer factorization. h This indicates the hierarchical level of each segmentation node. Indicates the largest integer less than or equal to the parameter. This indicates the time information assigned to the input sample set, and Indicates time offset.

14. The signal processing apparatus according to claim 1 or 2, wherein, Each split node at each level is configured to assign time information to each subtree based on the time information of the input samples assigned to each split node, and / or based on the level of each split node, and / or based on the integer factor chosen for the factorization of the number of processing cores, and / or based on the time offset.

15. The signal processing apparatus according to claim 1 or 2, wherein, The index assigned to each split node is i The time information of the subtree is based on the following equation: in Indicates that the index is assigned to i The time information of the subtree, where the first subtree is composed of i = 0 index, W represents the number of subsets of input samples provided by each segmentation node. describe, in express P Integer factors, according to , in P Indicates the number of processing cores. H This represents the total number of factors in the selected integer factorization. h This indicates the hierarchical level of each segmentation node. Indicates the largest integer less than or equal to the parameter. This indicates the time information assigned to the input sample set, and Indicates time offset.

16. The signal processing apparatus of claim 1 or 2, further comprising an input register configured to store a plurality of input samples.

17. The signal processing apparatus according to claim 16, wherein, The input register is a shift register.

18. The signal processing apparatus of claim 16, further comprising a selector (290, 640, 740, 840) configured to select a set of input samples for the sample distribution logic from the plurality of input samples.

19. The signal processing apparatus according to claim 1 or 2, wherein, The length of the time offset can be equidistant or unequal.

20. The signal processing apparatus according to claim 1 or 2, wherein, The signal processing device performs interpolation between the input samples.

21. The signal processing apparatus according to claim 1 or 2, wherein, The signal processing device performs convolution.

22. The signal processing apparatus according to claim 1 or 2, wherein, The multiple processing cores implement the Faro structure (120, 120a-f, 220, 220a-f, 530, 690, 790, 890).

23. The signal processing apparatus according to claim 1 or 2, wherein, The construction of different subtrees is derived from the same or different choices of integer factors of the number of processing cores.

24. The signal processing apparatus according to claim 1 or 2, wherein, The construction of different subtrees is derived from the same or different sorting of the number of processing cores by integer factors.

25. A signal processing method for providing multiple output samples based on an input sample set, the signal processing method comprising: By utilizing a hierarchical tree structure with multiple levels, multiple subsets of the input sample set are provided to multiple processing operations that perform processing operations associated with different time offsets. In this context, each segmentation operation at the lowest level provides two or more subsets from its input samples to multiple processing cores coupled to that segmentation operation. In this context, each segmentation operation at a given level higher than the lowest level provides two or more subsets from the input samples of each segmentation operation at the given level to multiple subtrees coupled to each segmentation operation at the given level. Each segmentation operation selects a subset whose range is consistent with the time offset associated with the processing operation to the corresponding subtree, and The output samples are obtained by performing processing operations associated with different time offsets in parallel.