Processing node for signal processing in wireless transceivers

JP2026529892APending Publication Date: 2026-09-03VIASAT INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026504803
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-07-28
Filing Date
2024-07-26
Publication Date
2026-09-03

Smart Images

  • Figure 2026529892000001_ABST
    Figure 2026529892000001_ABST
Patent Text Reader

Abstract

The wireless system and transceivers are described in a software-defined physical layer (SDP) architecture that incorporates processing nodes into existing hardware blocks. The processing nodes provide a soft processing design that supports existing functionality and allows for added and / or improved flexibility through future updates. By adding processing nodes to the existing hardware design to support existing functionality, additional processing power can be provided to existing hardware blocks when requested or needed, and extensions can be provided by taking over specific functions performed by existing hardware blocks or by replacing existing hardware blocks in the processing chain. Processing nodes allow the wireless system to add a soft processing design that helps perform core tasks of hardware blocks while maintaining their core design, and ultimately takes over and replaces the functionality of the hardware blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-Reference to Related Applications This application claims the benefit of priority to U.S. Provisional Application No. 63 / 516,434, filed on July 28, 2023, entitled "PROCESSING NODES FOR SIGNAL PROCESSING IN RADIO TRANSCEIVERS". The entire content of this document is incorporated by reference in its entirety for all purposes.

[0002] The present disclosure generally relates to processing nodes for signal processing in radio transceivers.

Background Art

[0003] Description of Related Art Signal processing in a radio transceiver includes various operations for manipulating and enhancing transmitted and received signals. A radio transceiver can receive an analog radio frequency (RF) signal and convert it into a digital signal via analog-to-digital conversion (ADC), and correspondingly, the radio transceiver can convert transmittable digital signals via digital-to-analog conversion (DAC). Digital signal processing includes processing steps performed on digital signals (e.g., after processing by an ADC and / or before processing by a DAC). Digital signal processing can be applied to extract, filter, and enhance useful information within a signal. This processing may include processes or operations such as filtering, equalization, modulation, demodulation, channel coding or decoding, error correction, and noise reduction. Other processes that may be performed include frequency conversion (e.g., conversion to baseband frequency or carrier frequency), gain control, amplification, and the like.

Summary of the Invention

[0004] In some embodiments, the disclosure relates to a processing node (PN) comprising a first digital signal processor (DSP) core, a second DSP core, a plurality of extended direct memory access controllers (DCX), each DCX having a shared memory space, an input packet interface, and an output packet interface, the input packet interface being configured to receive samples from a hardware block separate from the processing node, the shared memory space being configured to store received samples, and the output packet interface being configured to transmit samples processed by the first or second DSP core to a hardware block, and a PN network interconnect configured to communicatively connect the first DSP core, the second DSP core, and the plurality of DCX, each DSP core and DCX being connected to the PN network interconnect via its respective master interface and its respective slave interface, the PN network interconnect further comprising an SDP master interface and an SDP slave interface, each configured to communicate with an SDP network interconnect. The processing node is integrated into the wireless transceiver, which includes hardware blocks, and is configured to operate in conjunction with the hardware blocks to provide configurable processing functions to the wireless transceiver.

[0005] In some embodiments, the PN network interconnect further includes a configuration interface configured to allow processing nodes to configure hardware blocks. In some embodiments, the processing node further includes a queue interface configured to transfer commands or data from a first DSP core to a second DSP core and from the second DSP core to the first DSP core. In some embodiments, the processing node further includes a first queue interface and a second queue interface, the first queue interface configured to transfer commands or data from a first DSP core to a second DSP core, and the second queue interface configured to transfer commands or data from the second DSP core to the first DSP core.

[0006] In some embodiments, each DSP core includes a general-purpose input / output (GPIO) port connected to a configuration register, which is configured to receive inputs for placement in the configuration register and to transmit data stored in the configuration register. In some embodiments, each DSP core is configured to receive interrupt requests from a hardware block separate from the processing node via a PN network interface.

[0007] In some embodiments, a first DCX among a plurality of DCXs is configured to receive a plurality of samples to be processed from a first hardware block via an input packet interface, to temporarily store the plurality of samples in a shared memory space, and to transmit the plurality of samples to a first DSP core via a PN network interconnect. In some embodiments, a first or second DSP core is configured to program the first DCX to deliver the plurality of samples to the first DSP core, to place the plurality of samples in the internal memory space of the first DSP core, to process the plurality of samples, and to place the processed samples in the internal memory space of the first DSP core. In some embodiments, the first DCX is further configured to reformat the received plurality of samples. In some embodiments, the first DCX is configured to reformat the received plurality of samples by sign-extending the samples of the received plurality of samples to increase the number of bits for each sample. In some embodiments, the first DCX is configured to reformat the received plurality of samples by bit-clipping the samples of the received plurality of samples to reduce the resolution of each sample.

[0008] In some embodiments, a first DCX among a plurality of DCXs is configured to receive a plurality of samples to be processed via an input packet interface, to temporarily store the plurality of samples in a shared memory space, and to send the plurality of samples to a hardware block separate from the processing node for processing using an output packet interface. In some embodiments, the first DCX is configured to receive a plurality of processed samples from the hardware block and to temporarily store the processed plurality of samples in a shared memory space.

[0009] In some embodiments, the first DSP core and the second DSP core are configured to be used as separate entities or as a shared dual-core configuration. In some embodiments, the first DSP core and the second DSP core each include two processors, and the multiple DCXs include DCXs for each processor of the first DSP core and the second DSP core. In some embodiments, the SDP master interface and SDP slave interface of the PN network interconnect are configured to communicate with the PN network interconnect of different processing nodes within the wireless transceiver via the SDP network interconnect.

[0010] In some embodiments, a first DSP core, a second DSP core, and each of a plurality of DCXs include a configuration register configured to store data for configuring the associated DSP core or DCX. In some embodiments, the processing node is configured to be implemented within the demodulator of the radio transceiver. In some embodiments, the processing node is configured to be implemented within the decoder of the radio transceiver. In some embodiments, the processing node is configured to be implemented within the modulator or encoder of the transmitter of the radio transceiver.

[0011] In some embodiments, the Disclosure relates to a signal processing architecture comprising a software-defined physical layer (SDP) network interconnection, a plurality of processing nodes connected to the SDP network interconnection and configured to provide processing capabilities configurable for processing receiver and transmitter waveforms in a radio transceiver, each processing node comprising a plurality of digital signal processing (DSP) cores, a plurality of enhanced direct memory access controllers (DCXs), a plurality of DSP cores, a plurality of DCXs, and a PN network interconnection connected to the SDP network interconnection, a capture memory array (CMA) comprising a plurality of memory banks, the plurality of memory banks connected to the SDP network interconnection and providing the plurality of processing nodes with access to the plurality of memory banks, and a CPU subsystem connected to the SDP network interconnection, wherein the SDP network interconnection enables communication between the plurality of processing nodes, the CMA, and the CPU subsystem, thereby enhancing the processing capabilities and functionality in the radio transceiver.

[0012] In some embodiments, one or more of the processing nodes can be dynamically assigned to provide signal processing capabilities to one or more hardware blocks of the wireless transceiver. In some embodiments, each of the processing nodes is configured to operate in connection with one or more separate hardware blocks in both the receiver and transmitter signal processing data paths within the wireless transceiver.

[0013] In some embodiments, a first processing node among a plurality of processing nodes is implemented in the encoder or modulator of the wireless transceiver. In some embodiments, a second processing node among a plurality of processing nodes is implemented in the demodulator of the wireless transceiver. In some embodiments, a third processing node among a plurality of processing nodes is implemented in the decoder of the wireless transceiver.

[0014] In some embodiments, the signal processing architecture further includes external memory connected to a CPU subsystem, and the multiple processing nodes are configured to pass data from individual DSP cores to the external memory via an SDP network interconnect. In some embodiments, the individual processing nodes of the multiple processing nodes are integrated within different parts of a demodulator. In some embodiments, each processing node includes an SDP master interface and an SDP slave interface to the SDP network interconnect, the CMA includes multiple SDP master interfaces to the SDP network interconnect, and the CPU subsystem includes an SDP master interface and an SDP slave interface to the SDP network interconnect. In some embodiments, the signal processing architecture further includes a second SDP network interconnect connected to the SDP network interconnect, and a second set of multiple processing nodes connected to the second SDP network interconnect, each of which processing nodes includes one or more digital signal processing (DSP) cores, one or more enhanced direct memory access controllers (DCXs), and a PN network interconnect connected to the second SDP network interconnect.

[0015] According to several embodiments, the Disclosure provides a method for passing data to a processing node in a signal processing architecture, the signal processing architecture including a software-defined physical layer (SDP) network interconnect connected to the processing node, the processing node including a digital signal processing (DSP) core, an enhanced direct memory access controller (DCX), a packet receiver, a packet transmitter, and a PN network interconnect connected to the SDP network interconnect, the method comprising: managing available memory using a buffer pointer queue, the buffer pointer queue comprising a plurality of buffer pointers, each buffer pointer identifying and managing a buffer address in memory available for storing data; and managing burst data to be processed using a work descriptor queue, the work descriptor queue comprising a plurality of work descriptors, each work descriptor identifying and managing a buffer address in memory containing the burst data to be processed. The present invention relates to a method comprising: a packet receiver, in response to receiving burst data to be processed, obtaining a first buffer pointer from a buffer pointer queue; processing the received burst data; storing the processed burst data in memory at the buffer address specified by the first buffer pointer; and outputting a new work descriptor, the new work descriptor including the buffer address specified by the first buffer pointer; and a work descriptor queue, in response to processing data to be transmitted, obtaining a first work descriptor from the work descriptor queue; obtaining processed data from memory at the buffer address specified by the first work descriptor; and outputting a new buffer pointer, the new buffer pointer corresponding to the buffer address specified by the first work descriptor; and transmitting the processed data by a packet transmitter.

[0016] In some embodiments, the work descriptor further includes a data header length indicating the storage capacity occupied by the data header associated with the burst data. In some embodiments, the work descriptor further includes a burst start flag indicating that the burst data belongs to the first packet of the burst, and a burst end flag indicating that the burst data belongs to the last packet of the burst. In some embodiments, the work indicator indicates that the burst data is a fully contained burst by setting the burst start flag and the burst end flag to true.

[0017] In some embodiments, the method further includes adding a new work descriptor to the work descriptor queue. In some embodiments, the method further includes adding a new buffer pointer to the buffer pointer queue. In some embodiments, in response to receiving the buffer pointer queue and the work descriptor queue, the method further includes: obtaining a second work descriptor from the work descriptor queue; obtaining data from memory at the buffer address indicated by the second work descriptor; obtaining a second buffer pointer from the buffer pointer queue; processing the obtained data to generate output processed data; storing the output processed data in memory at the buffer address indicated by the second buffer pointer; outputting a new work descriptor, the new work descriptor including the buffer address indicated by the second buffer pointer; and outputting a new buffer pointer, the new buffer pointer indicating the buffer address indicated by the second work descriptor. In some embodiments, each work descriptor further includes a burst identifier for the burst data to be processed and a burst length indicating the storage capacity occupied by the burst data to be processed.

[0018] In some embodiments, the method further includes monitoring the fill level of the work descriptor queue by the DCX and, in response to determining that the fill level is below a threshold fill level, adding one or more work descriptors from the work descriptor list to the work descriptor queue.

[0019] In some embodiments, the method further includes monitoring the fill level of the buffer pointer queue by the DCX and, in response to determining that the fill level is below a threshold fill level, adding one or more buffer pointers from the buffer pointer list to the buffer pointer queue.

[0020] According to some embodiments, the disclosure provides a method for converting between a sample stream or symbol stream and a message in a signal processing architecture for storage and processing by a processing node including a digital signal processor (DSP) core and an extended direct memory access controller (DCX), the method comprising: receiving a sample stream containing bursts to be processed by the signal processing architecture; generating a header message containing information related to the bursts; and dividing the sample stream into a plurality of burst messages, the size of each burst message corresponding to the buffer size in the DCX, except for the final burst message corresponding to the end of the burst, the size of the final burst message being less than or equal to the buffer size in the DCX, The method comprises: generating a footer message containing information relating to the sizes of multiple burst messages; forwarding a burst interface packet to a DCX, wherein the burst interface packet comprises a header message, multiple burst messages, and a footer message; and reformatting the burst interface packet into a burst memory packet and storing it in the DCX, wherein the burst memory packet comprises a header message and a footer message in the first part of the burst memory packet, followed by multiple burst messages in the first part of the burst memory packet, the first part of the burst memory packet indicating the number of burst messages in the burst memory packet.

[0021] In some embodiments, the sample stream is received by a component of the processing node. In some embodiments, the method further includes identifying start and end flags in the sample stream to determine the endpoint of a burst in the sample stream. In some embodiments, the method further includes identifying a first start flag and a second start flag in the sample stream to determine the endpoint of a burst in the sample stream, where the endpoint is the first start flag and data preceding the second start flag but not containing the second start flag. In some embodiments, the bit size of each packet in the sample stream is equal to the size of a word in the DCX. In some embodiments, dividing the sample stream into multiple burst messages is in response to identifying a frame start indicator in the sample stream. In some embodiments, dividing the sample stream into multiple burst messages is terminated in response to identifying a frame end indicator in the sample stream. In some embodiments, the header message includes a burst counter that increments in response to identifying the boundary of a burst. In some embodiments, the size of the first part of a burst memory packet is set to be no more than or equal to the size of two words in the memory in the DCX. In some embodiments, reformatting a burst interface packet further includes converting a first word size of data within the burst interface packet to a second word size compatible with DCX, where the second word size is larger than the first word size.

[0022] In some embodiments, the Disclosure relates to a method for accessing memory in a signal processing architecture including a Capture Memory Array (CMA), wherein the CMA includes a plurality of random access memory (RAM) modules, each RAM module being logically divided into a plurality of sequentially arranged memory banks, and the method includes receiving a plurality of requests for access to a RAM module in the CMA, each of which includes a memory address in the CMA corresponding to memory in a particular RAM module; for each request, deriving a particular bank among the plurality of banks in the RAM module from the memory address in the request, the particular bank including the memory address in the request; determining a priority between two requests in response to determining that two of the plurality of requests are requesting access to the same bank in the same RAM module; granting access to the requested bank to the requester of the higher-priority requester of the two; delaying the requester of the lower-priority requester by a clock cycle; and for each request requesting sequential access to the CMA, granting access to a bank among the plurality of memory banks that is sequentially following the bank in the request.

[0023] In some embodiments, multiple banks are interleaved at a low level, and the number of banks is a power of 2. In some embodiments, multiple RAM modules are further divided into multiple channels, and the number of channels is a power of 2. In some embodiments, multiple channels are interleaved at a high level.

[0024] In some embodiments, the method further comprises, in response to determining that two of the plurality of requests result in access requests to different banks within the same RAM module, allowing simultaneous access to the two requests to the respective requested banks. In some embodiments, determining the priority comprises assigning a lower priority to the request that last accessed the requested RAM module.

[0025] For the purpose of summarizing the present disclosure, certain aspects, advantages, and novel features are described herein. It should be understood that not all such advantages may necessarily be achieved in accordance with any particular embodiment. Accordingly, the disclosed embodiments may be carried out in a manner that achieves or optimizes one advantage or group of advantages taught herein, without necessarily achieving other advantages that may be taught or suggested herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] [Figure 1] FIG. 1 is a diagram illustrating an example wireless system comprising a modem incorporating a plurality of processing nodes comprising a plurality of hardware blocks. [Figure 2] FIG. 2 is a diagram illustrating an example processing node comprising a plurality of digital signal processing (DSP) cores and a plurality of extended direct memory access (DMA) controllers (DCX). [Figure 3A] FIG. 3A is a diagram of an example DCX of a processing node (e.g., the processing node of FIG. 2). [Figure 3B] FIG. 3B is a diagram of an example DSP core of a processing node (e.g., the processing node of FIG. 2). [Figure 3C] FIG. 3C is a detailed diagram of the example DCX of FIG. 3A. [Figure 3D] FIG. 3D is a detailed diagram of the example DCX of FIG. 3A. [Figure 4] FIG. 4 is a diagram of data flow through a processing node (e.g., the DCX and DSP cores of FIGS. 3A and 3B). [Figure 5A] Figure 5A shows an example of a shared memory model for DSP cores within a processing node. [Figure 5B] Figure 5B shows an example of a producer-consumer model for a DSP core within a processing node. [Figure 6A] Figure 6A shows an exemplary software-defined physical layer (SDP) architecture. [Figure 6B] Figure 6B shows an exemplary software-defined physical layer (SDP) architecture. [Figure 7A] Figure 7A is a diagram showing an exemplary wireless transceiver including a CPU subsystem, demodulator, high-speed serial module, decoder module, and transmitter module, the exemplary wireless transceiver including an SDP architecture similar to the SDP architecture in Figure 6B. [Figure 7B] Figure 7B shows the demodulator buck module of Figure 7A in more detail, illustrating the connections between the hardware block and the processing node. [Figure 8A] Figure 8A shows an example of a buffer pointer. [Figure 8B] Figure 8B shows an example of a work descriptor. [Figure 9] Figure 9 shows an example of a component that facilitates data processing in an SDP architecture by utilizing buffer pointers and work descriptors. [Figure 10A] Figure 10A shows how data is received at the packet receiver via a portion of the data path in the SDP architecture, and the components of the SDP architecture operate similarly to the components in Figure 9. [Figure 10B] Figure 10B illustrates how data is passed to the packet transmitter and transmitted through a portion of the data path in the SDP architecture, showing that the components of the SDP architecture operate similarly to the components in Figure 9. [Figure 10C]Figure 10C is a diagram illustrating an example of data flow for software-based processing using buffer pointers and work descriptors in the SDP architecture, where the components of the SDP architecture operate similarly to the components in Figure 9. [Figure 10D] Figure 10D is a diagram illustrating an example of data flow stored in on-chip memory using buffer pointers and work descriptors in the SDP architecture, where the components of the SDP architecture operate similarly to the components in Figure 9. [Figure 11A] Figure 11A shows an example of a data formatting module configured to receive digitized data from a high-speed serial receiver module, similar to the HSS / ADC RX or HSS / SDP RX modules described herein with reference to Figure 7A. [Figure 11B] Figure 11B shows an example of a data formatting module configured to receive processed data from a DCX packet transmission interface and prepare the processed data for a high-speed serial transmission module, similar to the HSS / DAC TX or HSS / SDP TX modules described herein with reference to Figure 7A. [Figure 12A] Figure 12A shows the packet format for data in the SDP architecture. [Figure 12B] Figure 12B shows the packet format for data in the SDP architecture. [Figure 13] Figure 13 is a diagram showing the memory module within the DCX, which includes read ports, write ports, a memory arbiter, and memory banks. [Figure 14A]Figure 14A is a diagram showing an exemplary memory module of a Capture Memory Array (CMA) that includes memory banks, which are divided into multiple channels and banks and provide simultaneous multiple access functionality to the memory, and the CMA is similar to the CMA described herein with reference to Figures 6A, 6B, and 7A. [Figure 14B] Figure 14B is a diagram showing a CMA with multiple memory modules, each of which can be selected based on the bank and channel derived from the requested address. [Modes for carrying out the invention]

[0027] The headings used herein, if any, are for convenience only and do not necessarily affect the scope or meaning of the claimed invention.

[0028] overview Wireless systems can be used to send and receive signals in various architectures (e.g., terrestrial wireless systems, satellite systems, hybrid systems of terrestrial and satellite wireless systems). Signal processing in wireless systems involves various operations to manipulate and enhance transmitted and received signals. Signal processing can be performed using specific or specially designed hardware to perform a particular task (e.g., modulation, demodulation, encoding or decoding, error correction, noise reduction, etc.). The physical layer (PHY) of a wireless communication system includes components and processes involved in the transmission and reception of physical signals. This deals with the physical characteristics of transmitted and received signals (e.g., their modulation, coding, and transmission over the physical medium). The physical layer of a wireless system includes hardware components, modulation techniques, coding schemes, transmission media, antennas, and signal conditioning processes involved in the transmission and reception of physical signals.

[0029] In certain wireless systems, hardware can be designed specifically to perform particular signal processing functions, which are part of the wireless system's physical layer. These may include, for example, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and digital signal processors (DSPs). In wireless systems handling different waveforms (e.g., satellite systems communicating with different satellites), each waveform may have its own specific hardware. This means that a wireless system communicating with multiple satellites, each with different waveforms, may have its own ASIC and / or other hardware components. Furthermore, wireless systems containing hardware specifically designed to perform signal processing tasks (e.g., terminals in a satellite system) lack the flexibility to add future functionality or improvements.

[0030] Accordingly, this specification describes wireless systems and transceivers with a software-defined physical layer (SDP) architecture that incorporates processing nodes into existing hardware blocks. The processing nodes provide a software processing design that supports existing functionality and allows for additional and / or improved flexibility through future updates. The processing nodes can also be adapted or reconfigured to process new waveforms, even if existing signal processing hardware (e.g., hardware blocks, processing nodes, etc.) is not specifically configured to operate with the new waveforms, thereby enabling communication with new or different satellites. As used herein, the term hardware block can be used to refer to any signal processing component and / or module that may be used in modem designs for wireless systems and / or transceivers.

[0031] The disclosed systems, devices, architectures, and methods incorporating the disclosed processing nodes support legacy functionality in that the processing nodes are added to and configured to operate as an addition to an existing hardware design. This allows the processing nodes to support existing functionality, provide additional processing power to existing hardware blocks when requested or needed, and provide extensions by taking over specific functions performed by existing hardware blocks or by replacing existing hardware blocks in the processing chain. The disclosed processing nodes can support signal processing during the initial deployment of a terminal and can be fine-tuned over time to provide extended and / or improved functionality. Thus, the disclosed processing nodes allow wireless systems (e.g., terminals) to add, as a soft processing design, that can assist in performing core tasks of hardware blocks and ultimately take over and replace the functionality of the hardware blocks, while maintaining their core design (e.g., the same hardware blocks perform the same functions).

[0032] In some embodiments, the disclosed processing node includes multiple digital signal processor (DSP) cores. The processing node may be configured to operate in conjunction with existing hardware blocks in a wireless system, enabling communication between the processing node and the hardware blocks. In a wireless system, multiple processing nodes can be added to augment or assist certain hardware blocks and support various signal processing tasks. In certain embodiments, each processing node may be structurally and / or architecturally identical, and / or can be programmed to perform different signal processing tasks depending on the hardware block to which the processing node is associated, or depending on the hardware block to which the processing node replaces. Processing nodes may be configured to operate independently and can be added anywhere in a wireless system where the corresponding processing is desired. Processing nodes can be added to provide software-defined functionality, which is advantageous in extending the functionality of existing hardware blocks.

[0033] Each processing node can be configured to provide a control and configuration interface with an existing hardware block, allowing the hardware block to access the processing node's processing and memory. A processing node can provide a bypass processing route (e.g., bypassing a hardware block) or an enhanced processing route (e.g., providing additional functionality to a hardware block). As used herein, a hardware block typically refers to a component that provides the functionality of the digital receiver / transmitter physical layer within a radio transceiver. In some embodiments, a hardware block is a signal processing component in the physical layer of a radio transceiver, separate from the disclosed processing nodes.

[0034] Some wireless systems employ software-defined radio (SDR). SDR refers to a wireless system in which many traditional hardware components of the radio transceiver are replaced or extended by software processing. In SDR, the majority of signal processing functions are implemented in software, providing flexibility, reconfigurability, and the ability to adapt to various communication standards and protocols. A defining characteristic of SDR is its ability to perform RF signal processing using software algorithms rather than relying on fixed-function hardware. The disclosed processing node provides similar functionality to software-defined radio in that it provides a software-defined physical layer (SDP).

[0035] Therefore, as disclosed herein, the SDP architecture is a cluster of DSP processors that can be tightly integrated with existing hardware blocks (e.g., modem codec blocks) and become part of a wireless terminal ASIC. The disclosed SDP architecture provides flexibility to implement some or all of the digital receiver / transmitter physical layer functions of the ASIC in the processing nodes. The disclosed SDP architecture uses a multiprocessor signal processing architecture combined with existing modem designs, providing flexibility to the receiver or transmitter waveform processing algorithms while still leaving ample room to adapt to future system-level design changes and updates. The disclosed SDP architecture further enables existing terminals to be compatible with updated or new wireless systems (e.g., next-generation satellite systems). As a result, wireless systems or terminals incorporating the disclosed SDP architecture can continue to function with existing systems while being ready to communicate with systems with different characteristics in the future.

[0036] The disclosed processing nodes are configured to provide support circuitry around the DSP core to facilitate signal processing. Each processing node includes support circuitry in addition to the DSP core, providing flexibility in strategic locations within the wireless system such as encoders, modulators, demodulators, and decoders. This allows the disclosed processing nodes to be implemented in multiple locations within the wireless system, rather than requiring custom module design at each location. The disclosed processing nodes are easy to design and implement because they can be synthesized once and then stamped in different locations within the wireless system to provide flexible functionality.

[0037] Figure 1 shows an exemplary wireless system comprising a modem 100 incorporating multiple processing nodes 110 and multiple hardware blocks 104. The modem 100 is configured to receive a signal (Rx input), process the received signal using the processing nodes 110 and hardware blocks 104, and output the processed received signal (Rx output). Furthermore, the modem 100 is configured to receive a transmission signal (Tx input), process the transmission signal using the nodes 110 and hardware blocks 104, and output the processed transmission signal (Tx output).

[0038] The processing node 110 is configured to be implemented in the SDP architecture to operate as part of the software-defined physical layer, as described herein. The hardware block 104 includes modules and blocks that provide signal processing functions and can be incorporated into various parts of the signal processing chain. The processing node 110 is implemented as part of the signal processing chain and can provide flexible processing functions to one or more of the hardware blocks 104 and / or bypass processing by one or more of the hardware blocks 104. The hardware block 104 can be implemented as part of the receive signal processing chain and / or the transmit signal processing chain. The hardware block 104 can be implemented in the demodulator block of the modem 100, the transmit block of the modem 100, the decoder block of the modem 100, the CPU subsystem of the modem 100, etc.

[0039] The SDP architecture of modem 100 is configured to flexibly utilize processing nodes 110 in the signal processing chain. Furthermore, the SDP architecture of modem 100 may include shared memory (e.g., a capture memory array (CMA)) and one or more network interconnections linking the hardware blocks 104, processing nodes 110, and shared memory together. In addition to being a centrally connected multicore system, each processing node 110 can be configured to operate in connection with individual hardware blocks 104 of the receiver and / or transmitter signal processing data path. Processing nodes 110 can be deployed to meet the demands and / or requirements of encoders, decoders, modulators, demodulators, etc. Multiple processing nodes 110 may enable the SDP architecture to reallocate processing resources on demand. Furthermore, multiple processing nodes 110 may enable data to be passed between external resources and the SDP architecture.

[0040] Multiple processing nodes 110 can be placed at various locations in the signal processing chain. This allows for the instantiation of specific processing nodes from among the multiple processing nodes 110, or their inclusion as local processing functions. For example, one instantiation of a processing node might be included for each of the following: a transmitter, a demodulation module, a decoder, etc. In this case, each processing node instantiation is programmed to provide a specific processing function based on its use case.

[0041] Processing node Figure 2 shows an exemplary processing node 210 including multiple DSP cores 214a, 214b and multiple extended direct memory access controllers (DCXs) 212a, 212b, 212c, 212d. Since each of the multiple DSP cores 214a, 214b may be compatible or provide compatible functionality, when referring to a DSP core 214, one should assume that one is referring to an individual DSP core of the multiple DSP cores 214a, 214b. Similarly, since each of the multiple DCXs 212a-212d may be compatible or provide compatible functionality, when referring to a DCX 212, one should assume that one is referring to an individual DCX of the multiple DCXs 212a-212d. Each DCX 212 provides direct memory access (DMA) with extended functionality for moving data to and from memory. An example of this is described herein. Data can be moved in and out of the memory of multiple DCX212a~212d units in combination with processing by multiple DSP cores 214a, 214b, and / or by processing by external hardware blocks. Communication between multiple DSP cores 214a, 214b and multiple DCX212a~212d units, and / or with external hardware blocks, can be performed via the processing node (PN) network interconnect 216.

[0042] Each DCX212 and DSP core 214 includes one or more interfaces for communicating with the PN network interconnect 216. In some embodiments, each DCX212 and DSP core 214 includes a master interface and a slave interface coupled to the PN network interconnect 216. This allows various components to utilize individual DCX212 and / or DSP core 214 as slave components. It also allows individual DCX212 and / or DSP core 214 to function as masters to various components coupled to the PN network interconnect 216.

[0043] The PN network interconnect 216 can communicate with other processing nodes by communicating with a higher-level network interconnect (for example, the SDP network interconnect described herein). To operate in conjunction with the SDP interconnect, the PN network interconnect 216 includes an SDP master interface and an SDP slave interface. The SDP master interface is configured to enable communication with other processing nodes, and processing node 210 acts as a master to other processing nodes. In other words, processing node 210 can send data and / or commands to other processing nodes in the SDP architecture via the SDP master interface and utilize the processing and / or memory of other processing nodes. The master component includes any component that can perform read or write operations on another component (slave component). Similarly, the SDP slave interface is configured to provide communication with other processing nodes, and processing node 210 acts as a slave to other processing nodes. In other words, processing node 210 receives data and / or commands from other processing nodes via the SDP slave interface and makes the processing and / or memory of processing node 210 available to other processing nodes. For example, input data from the master processing node on the SDP slave interface allows the processing node 210 to execute code stored in the memory of the DSP core 214.

[0044] The PN network interconnect 216 also includes one or more configuration ports (or configuration interfaces). The configuration ports are connected to hardware blocks. The configuration ports can operate in conjunction with hardware blocks, allowing the processing node 210 to configure the hardware block's configuration to suit its operation with the processing node 210. For example, the configuration ports allow the processing node 210 to write configuration data to the hardware block, which may include parameters that enable data transfer to and from the hardware block. The configuration data may include any configuration supported by the hardware block.

[0045] A hardware block can be connected to one or more of several DCX212a-212d. For example, a hardware block can send data for processing to a processing node 210 via the DCX212's packet Rx interface. Alternatively, the processing node 210 can send processed data to the hardware block via the DCX212's packet Tx interface. In some embodiments, a configuration port is used to configure the hardware block to send data to and from the DCX212 via the DCX212's packet Rx interface. The sample interface (or packet Rx / Tx interface) may be a FIFO interface that allows for flexible options in connecting the hardware block (e.g., a communication component) to input and output data from the data path. As another example, if there are multiple interfaces within the hardware block and it is desirable to access data from the hardware block in parallel, it may be useful to use multiple DCXs. Similarly, if it is desirable to input data into the hardware block in parallel, or to mix transmitted and received data within the hardware block, it may be useful to use multiple DCXs to connect to and operate with the hardware block. This can also be useful in cases where a hardware block is subdivided into smaller functional blocks, and it may be useful to connect and operate with these smaller functional blocks using a dedicated DCX for each individual functional block.

[0046] The configuration port allows for configuring hardware blocks to function correctly. For example, the configuration port can be used to read the status and configuration of hardware blocks and to monitor the hardware block configuration. In addition, if it is desirable to pass data from the hardware block to the processing node 210, the configuration port can be used to configure the hardware block to route data to the processing node 210 and to configure the interaction so that the processing node 210 can receive the data (for example, via the packet interface in the DCX212). Hardware blocks can be decoder blocks, demodulator blocks, etc. Configuration data transmitted via the configuration interface can be configured to instruct the hardware block which data path to follow (for example, to send data to a specific packet interface). Thus, the configuration port functions as a control plane interface with the hardware block, and the hardware block can connect to the respective packet interfaces of multiple DCX212a-212d to transfer data to the processing node 210.

[0047] Therefore, the processing node 210 can be configured as a dual-core node that can be connected to different hardware modules. The processing node 210 provides configuration buses and interfaces to send and receive samples to and from hardware blocks, become part of the data path, and / or perform data capture. Individual DCX212 can be associated with a specific DSP core 214. The PN network interconnect 216 enables high-bandwidth data transfer between multiple DSP cores 214a, 214b and multiple DCX212a-212d. Individual DCX212 can also constitute part of a working descriptor chain, which will be described in more detail herein with respect to Figures 9, 10A, 10B, 10C, and 10D.

[0048] Multiple DCX212a-212d units provide direct memory access (DMA), enabling data transfer to and from memory. This can be achieved using the DSP core 214 and, advantageously, without intervention from an external processor, thus potentially improving the overall speed of the computer. The DSP core 214 can be configured to control and / or configure individual DMAs and set up data paths to be used for desired functions. DMA transfers blocks of data between the DMA memory and another location (e.g., external memory, the DSP core's internal memory, or other DCX memory). Each DCX212 includes a DMA controller that controls the operation of direct access to memory.

[0049] In some embodiments, the processing node 210 includes a first DSP core 214a and a second DSP core 214b. The processing node 210 also includes a plurality of DCXs 212a-212d, each DCX 212 having a shared memory space (not shown in this figure, but examples of which are described herein with reference to Figures 3A and 4), an input packet interface, and an output packet interface. The input packet interface is configured to receive samples from hardware blocks separate from but coupled to the processing node 210. The shared memory space is configured to store the received samples. The output packet interface is configured to send samples processed by the first DSP core 214a or the second DSP core 214b to the hardware blocks. The input and output packet interfaces of the DCXs are configured to provide flexible plug-in interfaces that allow connection to any general-purpose streaming interface in the data path of the hardware block. These streaming interfaces can be simple (e.g., a valid signal and / or associated data bus). These streaming interfaces can also be made more complex by adding start and end signals for framing. Therefore, since these packet interfaces are designed to be general-purpose, module adaptations are minimal or unnecessary so that such modules can connect to and operate with the DCX212.

[0050] The processing node 210 also includes a PN network interconnect 216 configured to communicatively couple a first DSP core 214a, a second DSP core 214b, and a plurality of DCXs 212a-212d. Each DSP core 214 and each DCX 212 are coupled to the PN network interconnect 216 via their respective master and slave interfaces. The PN network interconnect 216 further includes an SDP master interface and an SDP slave interface, each configured to communicate with the SDP network interconnect (examples of which are not shown in this figure but are described herein with reference to Figures 6A, 6B, 7A, and 7B). The processing node 210 is configured to be integrated into and to operate in conjunction with a wireless transceiver including hardware blocks, to provide configurable processing functions to the wireless transceiver.

[0051] The processing node 210 may include a queue interface configured to transfer commands or data from the first DSP core 214a to the second DSP core 214b and from the second DSP core 214b to the first DSP core 214a. The inter-processor queue 218 can be unidirectional (e.g., a first inter-processor queue from the first DSP core 214a to the second DSP core 214b and a second inter-processor queue from the second DSP core 214b to the first DSP core 214a), bidirectional (e.g., a single inter-processor queue that can transfer messages between the first and second DSP cores 214a, 214b), or the inter-processor queue 218 can provide bidirectional functionality using a combination of unidirectional queues. The inter-processor queue 218 is configured to provide inter-processor communication or descriptors (e.g., commands and messages). This can be done to synchronize the operation between multiple DSP cores 214a, 214b and / or to notify multiple DSP cores 214a, 214b of the location of data for processing. The inter-processor queue 218 includes a FIFO queue that travels directly between multiple DSP cores 214a, 214b.

[0052] In some embodiments, each DSP core 214 includes a general-purpose input / output (GPIO) port connected to a configuration register. This port is configured to receive inputs for placement in the configuration register and to transmit data stored in the configuration register. For example, a GPIO port may include 32-bit inputs to each DSP core 214 and / or 32-bit outputs from each DSP core 214. The GPIO port provides flexibility to the processing node 210. In some embodiments, the GPIO port is coupled to one or more configuration registers. The configuration register allows any master component to write to the 32-bit register of the associated DSP core 214 and any slave component to read from the 32-bit register of the associated DSP core 214. This can be advantageous for firmware debugging and / or troubleshooting because the GPIO can output standard values, allowing the operator to pinpoint where errors are occurring. The GPIO port may also be configured to receive hardware events, receive external events from hardware, generate triggers from hardware events, etc. In some embodiments, each of the multiple DSP cores 214a, 214b includes debug and trace ports connected to a CPU subsystem or operating system interface. This allows the performance of the multiple DSP cores 214a, 214b to be debugged using external data.

[0053] In some embodiments, interrupt requests (IRQs) can provide functionality that would otherwise be provided by GPIO ports. For example, each DSP core 214 can be configured to receive interrupt requests. IRQs can be received via dedicated ports or pins (not shown) within the DSP cores 214a, 214b. In some cases, IRQs can also be received via dedicated ports within the CPU subsystem, as described herein. Interrupts can be configured via interconnect ports and / or configuration ports within the DSP cores 214a, 214b and / or the CPU subsystem. In some cases, in addition to receiving IRQs via dedicated ports or pins, polling routines can also be implemented on the DSP processor and / or within the CPU subsystem to read interrupt registers via interconnect ports or configuration ports and detect pending IRQs. Multiple DSP cores 214a, 214b can be configured to receive IRQs by implementing a multiplexing scheme, allowing multiple DSP cores 214a, 214b to filter interrupts. This allows multiple DSP cores 214a and 214b to respond to specific interrupts and ignore others. Configuration to enable or disable interrupts can be performed via communication with hardware blocks (e.g., via configuration ports). This can be used to configure hardware blocks to generate interrupts for certain configured events. Furthermore, configurations can be implemented in multiple DSP cores 214a and 214b to instruct which interrupts each DSP core 214a and 214b responds to. For example, multiple DSP cores 214a and 214b may have access to many interrupts, but the configuration determines which interrupts each DSP core 214a and 214b responds to. When a DSP core 214 receives an interrupt it is configured to respond to, each DSP core 214 uses its configuration port to read registers in the hardware block and determine the nature of the interrupt.Interrupts can be used for a variety of purposes, including, but not limited to, errors, timers, DMA interrupts, forwarding notifications, packet queue / interface full, and packet queue / interface empty.

[0054] In some embodiments, the first DCX212a is configured to receive a plurality of samples to be processed from a hardware block via an input packet interface. The first DCX212a is then configured to temporarily store the plurality of samples in the first DCX212a's shared memory space. The first DCX212a is also configured to transmit the plurality of samples to the first DSP core 214a via a PN network interconnect 216. The first DSP core 214a is programmed to transmit the plurality of samples to the first DSP core 214a, to place the plurality of samples in the first DSP core 214a's internal memory space, and the first DSP core 214a is configured to process the plurality of samples and to place the processed samples in the first DSP core 214a's internal memory space. In some embodiments, the first DCX212a is further configured to reformat the received plurality of samples. In some embodiments, reformatting a group of received samples includes sign-extending the samples of the received group of samples to increase the number of bits for each sample. In some embodiments, reformatting a group of received samples includes bit-clipping the samples of the received group of samples to reduce the resolution of each sample.

[0055] In some embodiments, the first DCX212a is configured to receive multiple samples via the input packet interface, temporarily store the multiple samples in a shared memory space, and use the output packet interface to send the multiple samples to a hardware block separate from the processing node 210 for processing. The first DCX212a can further be configured to receive multiple processed samples from the hardware block and temporarily store the processed multiple samples in a shared memory space.

[0056] The processing node 210 can operate in different modes. For example, the processing node 210 can be configured to receive data streams from hardware blocks. As another example, the processing node 210 can be configured to respond to interrupts that can be used to identify specific data to capture (for example, to capture data for a specific event). As yet another example, the processing node 210 can operate using circular buffers. The processing node 210 may contain, for example, a list of 100 buffers. The processing node 210 can return the data at the top of the list of buffers to its own queue and continue overwriting the data so that the buffer contains the data of the last 100 buffers. At any point, the processing node 210 can move the data in the buffers into memory for processing and analysis. That is, instead of holding the buffer pointer in the buffer pointer queue, it can send the buffer pointer corresponding to the data of the last 100 buffers to the work descriptor queue. An example of this is described in more detail herein with reference to Figures 8A-810D. As yet another example, the processing node 210 can operate in buffer ID mode (using buffer IDs to assign data to specific DSP cores). In this mode, for example, buffer IDs ending in 0 are sent to core 0 of the first DSP core 214a, and buffer IDs ending in 1 are sent to core 1 of the first DSP core 214a. In this mode, it is also possible to configure certain buffer IDs to be ignored (for example, sending buffer pointers back to the buffer pointer queue instead of the work descriptor queue; an example of this is described in more detail herein with reference to Figures 8A-810D).

[0057] Processing node 210 is an example of a processing node that forms a basic component of the disclosed SDP architecture. Each processing node in the SDP architecture can be configured to include multiple dual-core processors (e.g., two DSP processors) and may include multiple DCXs that enable a wide range of data movement options between memory and data paths. Processing node 210 may also include a shared memory space (e.g., within each DCX 212). Multiple DCXs 212a-212d may also have dedicated packet interfaces for transferring samples to and from hardware blocks.

[0058] In some embodiments, the multiple DSP cores 214a, 214b include a VLIW (Vast-Limited Instruction) SIMD (Single Instruction Multiple Data) processor adapted for complex number processing. In some embodiments, the multiple DSP cores 214a, 214b may include a 32-way multiplier-accumulator (MAC), thereby enabling the multiple DSP cores 214a, 214b to perform up to 32 parallel MAC operations (or 8 complex multiplication and accumulation operations) per clock cycle.

[0059] The processing node 210 is configured to be sufficiently flexible to be placed in various locations within the wireless transceiver as needed (e.g., encoder, modulator, demodulator, decoder, etc.). Having a versatile design, the processing node 210 can be implemented in a specific location and programmed to provide functionality based on where it is implemented, rather than requiring a specific or custom design in different locations. This simplifies the design phase.

[0060] Figure 3A shows a diagram of an exemplary extended DMA controller or DCX312 for a processing node, such as the DCX212 of processing node 210 in Figure 2. The DCX312 may include a packet transmission interface 301 for outputting samples from the processing node. The DCX312 may also include a packet reception interface 302 for receiving samples from hardware blocks and storing and / or processing them by the processing node. The DCX312 may also include one or more configuration registers 303 configured to pass data, messages, and / or commands to and from the DCX312. In some embodiments, the DCX312 may include 100 or more configuration registers 303 (for example, each DMA may have 20 or more associated configuration registers). The DCX312 may also include a memory arbiter 304 that controls access to DCX memory 308, which includes RAM banks 1 and 2. The DCX312 also includes two sub-DMA memory modules 306, 307. Advantageously, the DCX312 can be repeated within a processing node, and sets of configurable DMA channels can be linked together (e.g., DMA channels of other processing nodes, and other shared memory (e.g., CMA)).

[0061] The DCX312 includes a memory arbiter 304 that controls read and write operations to the DCX memory 308 via a slave interface. The slave interface means that the memory arbiter 304 receives read and / or write requests to the DCX memory 308 from other processing nodes via a network interconnect. Therefore, the DCX memory 308 can be configured based on the application. For example, the DCX memory 308 can be used to buffer samples received from a hardware block (e.g., a hardware data path). As another example, the DCX memory 308 can be used as private scratch memory for one or more of several DSP cores. The memory arbiter 304 enables a dual read / write interface to the DCX memory 308, which is described in more detail herein with respect to Figure 13.

[0062] The DCX312 may include a DCX Single Atomic Write Control (SAWC) channel (or SAWC305) with a dedicated access port on the network interconnect, which allows the DCX312 to control the state machine, perform single-atom writes to programmable register offsets, return buffer pointers, and send working descriptors to the DSP core. Figures 3C and 3D show further details regarding the DCX312.

[0063] As described, the DCX312 includes the SAWC305 and sub-DMA modules 306 and 307. Each sub-DMA module 306, 307 includes a DMA read component 331 and a DMA write component 332, a working descriptor DMA channel 334, a reformatting FIFO 333, and a working descriptor controller 339. The SAWC305 includes a SAWC write component 341 and a configuration register 343 that includes working descriptors and buffer pointer queues. The SAWC write component includes individual write-only DMA channels arbitrated in a round-robin manner. Buffer pointers or working descriptors are provided to the write-only DMA channels from various blocks within the DCX312. The SAWC305 enables higher quality service for a single write DMA by allowing the network interconnect to handle arbitration between the two sub-DMA modules 306 and 307. Arbitration at the NIC level occurs when a transaction from any of the AXI masters within a processing node (e.g., individual DMAs, SAWCs, DSP cores, etc.) is about to be sent to a common slave memory located inside or outside the processing node. The DCX312 facilitates data transfer to other processing nodes and / or hardware blocks using multiple instances of a DMA read and write engine, including two configurable sub-DMA modules 306, 307 that operate in conjunction with network interconnects and other interfaces (e.g., packet Rx and Tx interfaces). Each sub-DMA module 306, 307 controls a single write channel and a single read channel. Each sub-DMA module 306, 307 includes four read channels in a DMA read component 331 and four write channels in a DMA write component 332, with each read and write channel operating and configured as a single entity.For example, the read channel and the write channel act as a single entity for performing a transfer from address A to address B. The read channel performs a read transaction to address A over the interconnect and temporarily buffers the data in a reformatted FIFO. The write channel then uses this data to perform a write transaction to address B over the network interconnect. Since the read channel can arbitrate with itself, and the write channel can arbitrate with itself, a single read and single write interface between the DMA and the network interconnect can be used efficiently. This allows read and write operations to be queued sequentially, improving bandwidth utilization efficiency from the interface. Each sub-DMA module 306, 307 manages data transfer using work descriptors from the work descriptor queue and buffer pointers from the buffer pointer queue (examples of which are described herein with reference to Figures 8A-10D).

[0064] In some cases, DMA channels can be configured to operate in normal DMA operation, and each channel can be programmed to perform a transfer from address A to address B. Configuration related to data transfer can be configured using configuration registers available for that DMA channel. Configuration data may include, for example, destination address, source address, memory attributes, interconnect transaction attributes, transfer length, and operating mode, without limitation. In some cases, DMA channels can be configured to operate in WQ operating mode. In this mode, three of the four channels are provided with work descriptors and buffer pointers from the WQ controller 339 in a round-robin manner, and the three DMA channels operate sequentially to improve the efficiency of transactions over the network interconnect 316. Typically, DMA transfers are divided into smaller configurable transaction sizes, so that a single DMA channel can maintain access to burst read / write control for a single transaction. The next DMA waiting in the queue then accesses the interface, and this continues until each DMA channel completes its transfer.

[0065] The DCX312 also includes a packet receiving interface 302 and a packet transmitting interface 301 that communicate with the memory arbiter 304, as described herein. The DCX312 also includes a Tx configuration component 355 and an Rx analysis component 354. This example is described herein with reference to the TX configuration module 1125 and the RX analysis module 1175, respectively.

[0066] In some embodiments, the configuration register 303 can pass a work descriptor and a buffer pointer to two sub-DMA modules 306, 307, and can receive a work descriptor queue and a buffer pointer queue from the two sub-DMA modules 306, 307. For example, the sub-DMA modules 306, 307 can fetch a buffer pointer from the buffer pointer queue and, upon receiving a work descriptor, can begin transferring data. Once the transfer is complete, the sub-DMA modules 306, 307 return the buffer pointer from the received work descriptor (indicating that the buffer pointer is available) and provide the SAWC 305 with a new work descriptor containing the fetched buffer pointer (indicating that the buffer pointer is now pointing to data).

[0067] The first sub-DMA module 306 is connected to the network interconnect via a master interface and transfers data to and from memories mapped as slaves on the network interconnect (including memories in other processing nodes and / or other shared memories, such as CMA). The second sub-DMA module 307 connects read and write channels to a hardware block, effectively enabling direct reading and writing to and from the DCX memory 308 by acquiring samples to and from the hardware block. Direct access to the DCX memory 308 allows the DCX memory to function as a capture buffer for packets or samples, and other DMA and DSP cores can access the DCX memory via the network interconnect.

[0068] Figure 3B shows a diagram of an exemplary DSP core 314 in a processing node, such as the DSP core 214 in processing node 210 in Figure 2. DSP core 314 represents one of several DSP cores that may exist within a processing node. For example, DSP core 314 could be one of two DSP cores within a processing node (e.g., processing node 210 in Figure 2). DSP core 314 includes one or more configuration registers 321 that allow control data or command data to be input from an external processing node, for example, via a network interconnect. Control data or command data includes buffer pointers, work descriptors, and messages. Any external master (or DSP core 314 itself) can perform write transactions via the interconnect and write control messages or command messages via the configuration registers 321. In some embodiments, each of the queues shown in the figure has its own dedicated configuration register. In such embodiments, writes to these particular configuration registers push work descriptors or messages into a FIFO, which can then be read by DSP core 314. The configuration register 321 can pass data in the form of a FIFO queue 322. The FIFO queue 322 may contain data passed from different processing nodes to the DSP core 314 (for example, a buffer pointer queue, a message queue, and a work descriptor queue). The queue 322 is passed to the DSP processor 324. The DSP processor 324 includes internal RAM and cache, and processes the data in the queue 322. In some embodiments, the buffer pointer queue contains 32-bit words, while the message queue and work descriptor queue contain 64-bit words.

[0069] In addition, the DSP core 314 can receive messages from other DSP cores in the processing node via the FIFO interface 323. The DSP core 314 can also receive interrupts (IRQs) and send FIFO interface messages via queue 325. In some embodiments, the processing node includes more than two DSP cores, and messages can be passed directly between DSP cores using the message queue (FIFO interface 323). In some embodiments, a configuration register 321 is used to pass messages directly between DSP cores in a particular processing node, as well as between an external processing node and a hardware block. In some embodiments, the DCX returns a buffer pointer to the DSP core 314 via the configuration register 321.

[0070] The configuration register 321 can be configured to store data passed between components of the SDP architecture, as described herein. The configuration register 321 can be configured to allow the DSP core 314 to communicate with different components (e.g., different DSP cores in the processing node, the DCX of the processing node, external processing nodes (e.g., processing nodes different from the processing node containing the DSP core 314), and / or external hardware blocks).

[0071] Exemplary data flow in a processing node Figure 4 shows a diagram of the data flow through an exemplary processing node 410. The processing node includes a DSP core 414 similar to the DSP cores 214 and 314, and a DCX412 similar to the DCX212 and 312. The DSP core 414 and DCX412 are each connected to a PN network interconnect 416 similar to the PN network interconnect 216. The PN network interconnect 416 enables communication between the DCX412, the DSP core 414, hardware blocks (e.g., via interrupts), and external processing nodes. The PN network interconnect 416 includes a configuration interface for configuring external hardware blocks as described herein, an SDP master interface for sending data to external slave processing nodes, and an SDP slave interface for receiving data from external master processing nodes. The DCX412 includes one or more configuration registers that communicate with the PN network interconnect 416 to send and receive data between components of the processing node 410 and between external processing nodes. Similarly, the DSP core 414 includes one or more configuration registers that communicate with the PN network interconnect 416 to send and receive data between components of the processing node 410 and between external processing nodes. In addition, one or more configuration registers of the DSP core 414 can be used to pass buffer pointers, work descriptors, and messages to the DSP core 414 from external processing nodes or from other components of the processing node 410. In some embodiments, the DSP core 414 includes other FIFO queues that can be used to transfer messages between cores of the processing node 410.

[0072] Data can be received from an external hardware block (not shown) at the packet receiver of the processing node 410. The data may be samples captured from a hardware block (e.g., a modem or communication hardware). The hardware block can be configured for a suitable or preferred capture interface using the configuration ports of the PN network interconnect 416, as described herein. The packet receiver is coupled to the DCX412. Data flows from the packet receiver to the write port of the DCX412's memory arbiter. The write port writes the received data to the DCX412's shared memory (DMA), where the received data is temporarily stored before processing. When ready, the DSP core 414 programs the DCX412 to transfer the dataset for processing to the DSP core 414's internal memory (e.g., the DSP core 414's internal data RAM). Typically, the DSP core 414 configures the data path used to send and receive samples. This is achieved by programming the packet receiver to create packets of a configured size and store the packets in shared memory. Next, a corresponding work descriptor is created using the size and buffer pointer information and transferred to the assigned work descriptor queue DMA (or WQ-DMA). Upon receiving the work descriptor, the WQ-DMA is configured to move the data into DSP core memory and create a new work descriptor. This new work descriptor is sent to the DSP core work queue FIFO interface via the DSP configuration interface. The WQ-DMA is also configured to release the buffer pointer for the work descriptor received from the packet receiver, freeing the buffer pointer back into the packet receiver's buffer pointer FIFO so that it can be reused once the data has been moved into DSP memory. Similarly, the DSP core 414 programs the transmit data path to set the work descriptor and buffer pointer flow in reverse, enabling data transmission to the packet transmitter.With this configuration, the DSP core 414 can receive samples from hardware blocks connected via a configuration interface to the hardware blocks. This allows data from the connected hardware blocks to the packet receiver to be placed in the DSP memory. This means that working descriptors are sent to the DSP core 414 without the DSP core 414 having to monitor and manage the data transfer. Processed data can be temporarily stored in the internal memory of the DSP core 414. The processed data is then returned to the shared memory (DMA) of the DCX 412 and can be temporarily stored there. After processing by the processing node 410, when the data is ready to be transferred to the external hardware block, the data is transferred back to the read port of the memory arbiter of the DCX 412, where it is read by the packet transmitter and sent to the external hardware block. Thus, the shared memory of the DCX 412 can be used to buffer samples for processing. In some embodiments, the shared memory of the DCX 412 can be reused as additional memory for the DSP core 414.

[0073] In addition to loading and unloading samples from memory, the DCX412 can be configured to reformat data when moving data through it. For example, each of the DCX412's internal DMA engines can include data reformatting logic around a FIFO connecting the read and write channels. The reformatting logic can be configured to assist with certain operations, such as sign-extending a sample from 8 bits to 16 bits, or performing bit clipping to reduce the sample resolution from 16 bits to 8 bits.

[0074] Exemplary SDP architecture The disclosed SDP architecture allows for flexibility in using DSP cores and supports various configurations and methods for providing multicore functionality to the SDP architecture. Each processing node can be configured to be used as a dual-core configuration, or multiple nodes can be configured to form a larger multicore functional body. By using network interconnects, each processing node can access the DCX and / or DSP core of another processing node. In dual-core configurations (e.g., a single processing node) and multicore configurations (e.g., multiple processing nodes), the DSP cores can operate in a shared memory model (cores share access to shared memory while accessing private memory) and a producer-consumer model (DSP cores can message each other).

[0075] Figure 5A shows an example of a shared memory model for DSP cores 514a and 514b within a processing node. In the shared memory model, DSP cores 514a and 514b access local shared memory 515 (which may reside in the DCX 512, for example, as described herein with reference to Figure 4, or in any memory accessible to DSP cores 514a and 514b via a network interconnect) to enable operation in this mode. Local shared memory 515 can be the shared memory of the DCX, a section of the CMA marked as shared among DSP cores, or any allocated memory in external DRAM / DDR. Private SRAM 513a and 513b can be used for scratch data during object processing, or they can be used as additional memory shared among cores 514a and 514b. This can be extended to multiple DSP cores in different processing nodes. By using network interconnects, each DSP core can access a shared memory space residing within a specific DCX (which may reside in another processing node), while simultaneously accessing private SRAM residing within the processing node in which the DSP core resides.

[0076] Figure 5B shows an example of a producer-consumer model for DSP cores 514a and 514b. DSP cores 514a and 514b have a queue interface 518 (or the queue can reside in shared memory) and dedicated GPIOs that can be used as message queues within a single processing node. This enables messaging between DSP cores 514a and 514b, which can be simple commands or object data. If the queue interface 518 is unidirectional, the first DSP core 514a can function as a producer, and the second DSP core 514b can function as a consumer, with the queue interface 518 acting as a queue for tasks filled by the first DSP core 514a (producer) and popped by the second DSP core 514b (consumer). This can be extended to multiple DSP cores in different processing nodes. For communication and synchronization between multiple processing nodes, additional memory-mapped message queues can be used to connect DSP cores in different processing nodes; memory-mapped message queues are similar to message queues between DSP cores in the same processing node. In some embodiments, the additional memory-mapped message queue uses an SDP network interconnect to transfer data (for example, using the configuration registers described herein).

[0077] Figures 6A and 6B show exemplary SDP architectures 600a and 600b. As shown in Figure 6A, SDP architecture 600a includes processing nodes 610a–610f. Processing nodes 610a–610f are coupled to the SDP network interconnect 640 via their respective SDP master and SDP slave interfaces. SDP architecture 600a also includes shared memory in the form of a capture memory array 630 (or CMA). This shared memory is coupled to the SDP network interconnect 640 via its SDP master and SDP slave interfaces. SDP architecture 600a also includes a CPU subsystem 650 (e.g., a Linux application layer). The CPU subsystem 650 is coupled to the SDP network interconnect 640 via its SDP master and SDP slave interfaces. In some embodiments, each bidirectional arrow connecting corresponding components of the SDP architecture 600 to the SDP network interconnect 640 may represent one of the SDP master-slave interfaces (for example, one SDP master and one SDP slave interface for each component). The processing nodes 610 are configured to communicate with each other via the SDP network interconnect 640, as will be described in more detail herein. This allows individual processing nodes 610 of the SDP architecture 600a to use different processing nodes 610 of the SDP architecture 600a to obtain memory and / or processing. As will be described herein, the processing nodes 610a to 610f may be implemented in different parts of a wireless transceiver or system. In such cases, the processing nodes 610a to 610f may be configured to communicate with each other via the SDP network interconnect 640 and utilize the memory and processing capabilities of other processing nodes.

[0078] The capture memory array 630 is a memory array that enables on-chip storage and can be used for capturing samples for processing, providing scratch areas, and storing lookups for processing. The capture memory array 630 includes a contiguous address space composed of multiple SRAM banks. In some embodiments, the capture memory array 630 can be designed using interleaved single-port RAM to function as read / write memory with a handshake-based interface for efficient area / complexity design. The capture memory array 630 may include a set of large-capacity memory banks connected to a data path, enabling low-latency, high-bandwidth parallel access to the processor without the need to go out to off-chip memory.

[0079] As described herein, the SDP architecture may further include a capture memory array (CMA) and a switch. The CMA can serve as a common location within the modem 100 where data can be stored, such as data samples, intermediate processing results, etc. In some embodiments, the modem 100 can be implemented as a single chip or multiple chips.

[0080] Referring to Figure 6B, the SDP architecture 600b shows an architecture in which the network interconnect is divided into multiple segments (a first SDP network interconnect 640a and a second SDP network interconnect 640b), to which individual processing nodes 610a-610f, a capture memory array 630, and a CPU subsystem 650 are connected. In some embodiments, one processing node (e.g., a first processing node 610a) can be directly connected to the CPU subsystem 650, and the remaining processing nodes 610b-610f can be connected via the first SDP network interconnect 640a or the second SDP network interconnect 640b (or a set of network interconnect segments). Different arrangements of the processing nodes 610, CPU subsystem 650, and capture memory array 630 and SDP network interconnect 640 can also exist.

[0081] The SDP architecture 600 provides flexibility to the waveform processing algorithms of receivers or transmitters while leaving room for future system-level design changes and updates. The SDP architecture 600 also enables software-based signal processing on the chip. In the SDP architecture 600, each processing node 610a-610f can operate in connection with individual hardware modules of both the receiver and transmitter signal processing data paths. Therefore, the SDP architecture 600 allows for the instantiation of DSP processing capabilities at specific critical or desired locations within the radio system (e.g., encoders, modulators, demodulators, and decoders) while also providing sufficient connectivity for reallocating processing resources to different locations. The SDP architecture 600 also advantageously provides sufficient connectivity to pass data between processing nodes 610a-610f, as well as to external resources (e.g., external memory). Furthermore, since processing nodes 610a to 610f are programmable, the SDP architecture 600 enables customizable signal processing, bringing flexibility to modem design and facilitating different types of signal processing for different applications.

[0082] Accordingly, the SDP architecture 600 includes an SDP network interconnect 640 and a plurality of processing nodes 610a to 610f connected to the SDP network interconnect 640, the plurality of processing nodes 610a to 610f configured to provide processing capabilities configurable to process receiver and transmitter waveforms in a wireless transceiver. As described herein, each processing node 610 may include a plurality of digital signal processing (DSP) cores, a plurality of enhanced direct memory access controllers (DCXs), and a plurality of DSP cores, a plurality of DCXs, and a PN network interconnect connected to the SDP network interconnect 640. The SDP architecture 600 may also include a capture memory array 630 containing a plurality of memory banks connected to the SDP network interconnect 640 to provide the plurality of processing nodes 610 with access to a plurality of memory banks. The SDP architecture 600 may also include a CPU subsystem 650 connected to the SDP network interconnect 640. The SDP network interconnect 640 enables communication between multiple processing nodes 610, capture memory arrays 630, and CPU subsystems 650, thereby enhancing the processing power and functionality of the wireless transceiver.

[0083] In some embodiments, one or more of the processing nodes 610 can be dynamically assigned to provide signal processing capabilities to one or more hardware blocks of the wireless transceiver. This enables dynamic allocation of processing capabilities to hardware blocks. In some embodiments, each of the processing nodes 610 can be configured to operate in connection with one or more separate hardware blocks in both the receiver and transmitter signal processing data paths within the wireless transceiver. This allows multiple hardware blocks to access a single processing node and provides flexible storage and memory capabilities.

[0084] In some embodiments, external off-chip memory can be further connected to the CPU subsystem 650. In such embodiments, multiple processing nodes 610 can be configured to pass data from individual DSP cores to the external memory via an SDP network interconnect 640.

[0085] In some embodiments, the capture memory array 630 includes multiple SDP master interfaces to the SDP network interconnect 640. In some embodiments, the CPU subsystem 650 includes SDP master interfaces and SDP slave interfaces to the SDP network interconnect 640.

[0086] Figure 7A shows an exemplary wireless transceiver 700. The wireless transceiver 700 includes a CPU subsystem 750 (similar to the CPU subsystem 650), demodulators 760a-760d divided into multiple demodulator blocks 760, a high-speed serial or HSS module 770, a decoder module 780, and a transmit module 790. The wireless transceiver 700 includes an SDP architecture similar to the SDP architecture 600b in Figure 6B. The CPU subsystem 750 includes a network interconnect switch (NIC switch) and a cache-coherent network (CCN or bus interconnect) and is coupled to external memory 755 (e.g., DDR memory). The network interconnect switch (NIC switch) is coupled to SDP network interconnect 2 (SDP NIC2). The SDP architecture (e.g., SDP network interconnect, processing nodes, and CMA, as described herein with reference to Figure 6B) may be represented and accessible via SDP NIC2, as illustrated.

[0087] The demodulator 760 can be divided into several modules 760a to 760d. Each module has one or more components, such as an SDP network interconnect, a CMA, and / or a hardware block (or HWB). Other components of the radio transceiver may also include hardware blocks. The hardware blocks can be configured to provide multiple signal processing functions, such as signal conditioning, channelization, downsampling, filtering, equalization, despreading, descrambling, etc. The hardware blocks can send and receive data to processing nodes as described herein.

[0088] The HSS module 770 is configured to convert the received analog signal into a digital signal for processing using the HSS / ADC block. The HSS / ADC block includes two read channels (read channel 0 and read channel 1) that are processed in parallel. The HSS module 770 is further configured to pass signals directly between an external chip or system that is part of the software-defined physical layer (referred to as SDP in the diagram, and referring to a dedicated SDP interface connected to the ASIC's High Speed ​​Serial (HSS) module) and the first processing node of the demodulation module 760a via the HSS / SDP RX1 block and the HSS / SDP TX1 block, and also between the SDP and the sixth processing node of the transmit module 790 via the HSS / SDP RX2 block and the HSS / SDP TX2 block. In some embodiments, the SDP interface is directly connected to a packet receiver and a packet transmitter within the processing node. This makes it possible to receive ADC samples directly at the processing node and transmit DAC samples directly from the processing node to the HSS module 770. The HSS module 770 is further configured to use an HSS / DAC TX block coupled to the transmit module 790 (transmit channel 0) to convert the digital signal to an analog signal for transmission.

[0089] The signal digitized by the HSS / ADC is passed to the demodulation modules 760a-760c, and then to the decoder module 780. The digital signal for transmission is passed to the transmit module 790, and then to the HSS / DAC TX block of the HSS module 770.

[0090] The processing nodes communicate with each other via SDP network interconnects 740a and 740b (SDP NIC1 740a and SDP NIC2 740b). Furthermore, the processing nodes can access the CMA via SDP network interconnects 740a and 740b.

[0091] Each processing node can be dynamically included in the signal processing data flow, as described herein. This allows modules 750, 760, 770, and 780 and their hardware blocks to utilize the flexible memory and processing provided by the processing nodes. In addition, the processing nodes are configured to access data processed by the hardware blocks via the DCX interface, as described herein. The processing nodes can be configured to pass data to the processing nodes using configuration ports, as described herein.

[0092] For example, the transmitting module 790 includes a processing node (PN6) coupled to the SDP NIC1 740a via an SDP master interface and an SDP slave interface. In addition, although not explicitly shown in the figures, the processing node PN6 is coupled to a MOD block and an ENC block via a packet interface and a configuration port, as described herein with reference to Figures 2, 3A, and 4. The processing node PN6 can configure the MOD block and / or ENC block to send data to and receive processed data from the processing node PN6. In some cases, this can be done to enhance the signal processing provided by the MOD block and / or ENC block. In some cases, this can be done to bypass the MOD block and / or ENC block.

[0093] Advantageously, the wireless transceiver 700 can dynamically assign the first DCX of the first processing node to be the master for writing data to shared on-chip memory (e.g., CMA), external memory, and / or the memory of another DCX, and assign the second DCX of the second processing node to be the master for reading the results of processing the data written by the first DCX. The placement of the processing nodes can be configured to enhance specific signal processing elements within the wireless transceiver 700. For example, a processing node can be placed in the front module of a demodulator to assist in signal demodulation, and a processing node can be used to enhance or replace (e.g., bypass) the DSC and / or CDM hardware blocks. The placement of the CMA can be changed from those shown herein, as can the placement and configuration of the SDP network interconnect and processing nodes. The placement of these components can be based on several factors, such as signal processing performance, chip layout, connectivity between blocks and other designs, latency, and manufacturing complexity and cost. The wireless transceiver 700, as shown in the diagram, provides a parallel data path for processing two waveforms in parallel, but it should be understood that a single data path can be used, or more than two data paths can be implemented in parallel. Advantageously, the NIC switch and CCN of the CPU subsystem 750 allow the processing node to access external memory. However, since this can be more costly in terms of speed, the wireless transceiver 700 also advantageously provides on-chip memory in the form of CMA. This can also advantageously help in connecting and operating with software running on an external processor.

[0094] Figure 7B shows the demodulation module 760c of Figure 7A in more detail, illustrating the connections between the hardware block and the processing node 710 (e.g., PN4 in Figure 7A), as well as the connections between the processing node 710 and the SDP NIC1 740a implemented within the demodulation module 760b. The processing node 710 is similar to the processing node 210 described herein with reference to Figure 2. The PN network interconnect 716 of the processing node 710 is coupled to the SDP interconnect 740 of the radio transceiver 700 (e.g., SDP NIC1 in Figure 7A), integrating the processing node 710 into the radio transceiver 700. The hardware block of the demodulation module 760c includes a first hardware block 1 761a for receiving channel 0, a first hardware block 2 762a for receiving channel 0, a second hardware block 1 761b for receiving channel 1, and a second hardware block 2 762b for receiving channel 1. Each hardware block of the demodulation module 760c can be configured to use its respective configuration port on the processing node 710. Each hardware block 761a, 761b, 762a, and 762b is coupled to the respective DCX 712a-712d of the processing node 710 via a packet transmission interface and a packet reception interface. In addition, each hardware block 761a, 761b, 762a, and 762b is coupled to the PN network interconnect 716 via its respective configuration port.

[0095] As described above, data can be passed between the processing node 710 and the hardware blocks (blocks 761a, 761b and blocks 762a, 762b) using the packet transmission / reception interfaces of the corresponding DCX712 710 of the processing node. For example, the first hardware block 1 761a can send data to the processing node 710 via the packet transmission interface of the first DCX712a, the first hardware block 2 762a can send data to the processing node 710 via the packet transmission interface of the second DCX712b, the second hardware block 2 762b can send data to the processing node 710 via the packet transmission interface of the third DCX712c, and the second hardware block 1 761b can send data to the processing node 710 via the packet transmission interface of the fourth DCX712d. Similarly, after processing by the first DSP core 714a and / or the second DSP core 714b, the processed data can be returned from the processing node 710 to a specific hardware block using the corresponding DCX 712a~712d packet receiving interface. For example, the first hardware block 1 761a can receive processed data from the processing node 710 via the first DCX 712a packet receiving interface, the first hardware block 2 762a can receive processed data from the processing node 710 via the second DCX 712b packet receiving interface, the second hardware block 2 762b can receive processed data from the processing node 710 via the third DCX 712c packet receiving interface, and the second hardware block 1 761b can receive processed data from the processing node 710 via the fourth DCX 712d packet receiving interface. In this way, the processing node 710 can provide additional and / or alternative processing to the demodulation module 760c.The remaining processing nodes of the wireless transceiver 700 have similar connections and can provide similar functionality to the other modules and hardware blocks of the wireless transceiver 700 in Figure 7A.

[0096] Work descriptors and buffer pointers As described herein, in order for the disclosed SDP architecture to operate as desired, different processing nodes are configured to exchange data seamlessly (for example, with each other, with memory, with hardware blocks, etc.). To achieve this, the disclosed SDP architecture can utilize work descriptors and buffer pointers as part of the data flow. Using the work descriptors and buffer pointers described herein, the SDP architecture can connect different data transfer portions across the architecture to form a highly configurable data path. The disclosed technology allows each hardware interface and DCX to be used as an interchangeable component that can be connected to any other hardware interface, DCX, or DSP core in the SDP architecture using a simple approach. Here, each component receives a work descriptor pointing to the location in memory of the data to be processed and has some pre-allocated memory indicated by a buffer pointer for the processed data the component uses. When a component has finished processing the data, it transfers a work descriptor containing a pointer to the location in memory where the newly processed data resides to the next component in the processing chain. When memory is freed, the location of the newly freed memory is returned to the buffer pointer queue, making it available to other components in the SDP architecture. For example, memory can be freed when data is sent to a hardware block, when it is moved to a new location in internal or external memory, and / or when data is processed and the result is stored in a new location in memory. Using this approach, data can be passed from one component to the next until it reaches its final destination. By interconnecting data across the SDP architecture in this way, virtually any two hardware interfaces, DCXs, and / or DSP cores can be connected to form a desired data path.

[0097] Figure 8A shows an example of a buffer pointer 811. The buffer pointer 811 may include, for example, a buffer address, a buffer pointer identifier, and reserved bits. As illustrated, the buffer pointer 811 may have a length of 32 bits, but other lengths (larger or smaller) may be used as needed. The buffer pointer 811 identifies the starting address of a pre-allocated location in memory that a processing node or component of a processing node can use to store data. This location in memory corresponds to the buffer address of the buffer pointer 811.

[0098] Figure 8B shows an example of a work descriptor 821. A work descriptor 821 may include, for example, a buffer address, a buffer pointer identifier, reserved bits, a burst identifier, a length, a burst start flag, and a burst end flag. The work descriptor 821 is a pointer to a dataset for a processing node or a component within a processing node. In some embodiments, the work descriptor 821 can have a length of 64 bits. The work descriptor 821 is used to point to a location in memory containing data (e.g., a burst) that is ready to be analyzed or processed. The location in memory corresponds to the buffer address in the work descriptor 821. In addition, the length indicated by the work descriptor 821 can be used to determine the amount of memory used to store the data at the buffer address. Additional identifiers and flags in the work descriptor 821 can be used to facilitate the analysis and processing of the data. Examples of identifiers or flags include the message type, DSP ID, and message count. In some embodiments, buffer pointer descriptors and work descriptors can be distributed in a round-robin manner using the BPID (buffer pool ID)[2:0] and burst ID[1:0] in the work descriptor 821. In some embodiments, the burst start flag and burst end flag in the work descriptor 821 are flags tagged by the packet receiver in the DCX when receiving data from a hardware block that uses or relies on some kind of framing. These flags allow software running on the DSP core to identify the set of descriptors that make up its frame without having to parse the packet stored in memory.

[0099] Figure 9 shows an example of component 900 that facilitates data processing in an SDP architecture using buffer pointers and work descriptors. Component 900 can act as a consumer and / or producer in a consumer / producer model. Component 900 can receive a work descriptor queue 920 (or a list of work descriptors) and a buffer pointer queue 910 (or a list of buffer pointers), and output a work descriptor 921 and a buffer pointer 911. The buffer pointer queue 910 and the work descriptor queue 920 are managed by separate queue controllers or other components, respectively, from component 900. Therefore, the buffer pointer queue 910 and the work descriptor queue 920 are available to component 900, but are not managed by component 900. This allows component 900 to pop an element from a particular queue and remove that element from its respective queue so that it is not visible as available to another consumer / producer in the signal processing chain.

[0100] Component 900 can retrieve the next work descriptor from the work descriptor queue 920 to identify the data to be processed (the work descriptor is removed from the work descriptor queue 920), and can retrieve the next buffer pointer from the buffer pointer queue 910 to identify the memory location available for storing the data (the buffer pointer is removed from the buffer pointer queue 910). Component 900 processes the data at the buffer address indicated in the work descriptor to produce processed data and stores the processed data at the buffer address indicated in the buffer pointer. Component 900 then creates a new work descriptor containing the buffer address where the newly processed data is stored, as well as the length of the data stored at the buffer address, and outputs the work descriptor 921 to the queue controller managing the work descriptor queue 920 and / or to the next consumer / producer in the chain. Component 900 can also create a new buffer pointer containing the buffer address where the just processed data is stored, and outputs the buffer pointer 911 to the queue controller managing the buffer pointer queue and / or to the next consumer / producer in the chain. In other words, the memory location of the stored data corresponding to the buffer address in the pre-processing work descriptor is shown here as available memory in the returned buffer pointer 911, while the pre-processing available memory is shown here as holding the newly processed data in the returned work descriptor 921. In this way, data can be passed between components of processing nodes and / or between processing nodes in the SDP architecture.

[0101] Component 900 is configured to retrieve data from the work descriptor queue 920 at the location pointed to by a work descriptor. Component 900 then processes the retrieved data. Component 900 stores the processed data using a buffer pointer from the buffer pointer queue 910. Component 900 then forms a new work descriptor 921 based on the location of the stored processed data. Component 900 then outputs a buffer pointer containing the buffer address from the work descriptor to a pre-configured location for reuse. Component 900 then sends the work descriptor 921 to the next consumer in the chain. Further examples of using buffer pointers and work descriptors are described in more detail herein with reference to Figures 10A–10D.

[0102] In some embodiments, component 900 does not receive a work descriptor queue 920, but instead receives data for storage (for example, a packet receiver on a processing node receives packets from a hardware block). In such embodiments, the work descriptor queue 920 is not provided to component 900, and component 900 operates only as a producer. Component 900 stores the received data in a buffer address of the buffer pointer queue 910 and outputs a work descriptor 921 indicating the location where the received data is stored. In such embodiments, component 900 does not output a buffer pointer 911 because memory was not available to the process.

[0103] Similarly, in some embodiments, component 900 does not receive a buffer pointer queue 910, but instead receives a work descriptor queue 920 containing data to be sent from a processing node (for example, a packet transmitter on the processing node sends a packet to a hardware block). In such embodiments, the buffer pointer queue 910 is not provided to component 900, and component 900 acts only as a consumer. Component 900 retrieves data from the buffer address indicated in the work descriptor in the work descriptor queue 920, sends the data, and outputs a buffer pointer to the buffer pointer queue manager. The buffer pointer corresponds to the buffer address where the sent data was stored, indicating that these locations in memory are now available for storage. In such embodiments, component 900 does not output a work descriptor 921 because the data was not processed and was not stored again in memory.

[0104] For example, a component 900 operating with a buffer pointer queue 910 and a work descriptor queue 920 can retrieve data from a location pointed to by a work descriptor in the work descriptor queue 920 and use the buffer address identified in the buffer pointer of the buffer pointer queue 910 for the result. Component 900 may then form a new work descriptor 921 based on the output generated and stored in the buffer pointer queue 910. The new work descriptor 921 includes a buffer pointer identified in the buffer pointer queue 910, which is used for the generated result. Component 900 may then output the buffer pointer 911 to a pre-configured location for reuse by a different component operating on data at the location pointed to by a subsequent work descriptor. Component 900 sends the new work descriptor 921 to the next consumer in the processing chain. The disclosed buffer pointers and work descriptors can be used in various control plane interfaces. For example, buffer pointers and work descriptors can be used in an interface that streams data samples from memory to a hardware component. A work descriptor can refer to a data packet in memory containing samples that need to be streamed to a hardware component, and a list of work descriptors corresponds to a playlist of data packets that should be streamed to the hardware component.

[0105] Figures 10A, 10B, 10C, and 10D illustrate an example of passing data through a portion of the data path within the SDP architecture 1000. The components of the SDP architecture 1000 operate similarly to the components in Figure 9. Figure 10A shows how data is received by the packet receiver 1030 through a portion of the data path within the SDP architecture 1000. The packet receiver 1030 receives samples, for example, from a hardware block. The packet receiver 1030 acts as a producer in the consumer / producer model, as described herein with reference to Figure 9, and acquires a buffer pointer queue (BP) along with the received data samples. The packet receiver 1030 acquires the received samples and places them into buffer addresses identified as available for data storage in the buffer pointer queue (BP). The buffer pointer queue is stored in DCX memory accessible to the packet receiver 1030. Once in memory, the packet receiver 1030 generates each work descriptor for inclusion in the work descriptor queue (WQ). This is passed to the Work Descriptor Queue DMA (WQ DMA1031), which manages the Work Descriptor Queue (WQ). In some cases, work descriptors located in the DCX's shared memory are accessible, and other DCXs and / or DSP cores can access these work descriptors.

[0106] In addition, a buffer pointer can be connected to a DMA1020 that monitors the fill level of a buffer pointer queue (BP). In response to the DMA1020 determining that the fill level of a buffer pointer queue (BP) (for example, a BP of WQ DMA1031 in a FIFO) has fallen below a threshold (for example, a FIFO threshold), the DMA1020 may refill the buffer pointer queue (BP) by inputting one or more subsequent buffer pointers from a pre-allocated buffer pointer list 1011, thereby keeping the buffer pointer queue at a sufficient level. If the WQ DMA1031 determines that there is space available in the work descriptor list 1012 (FIFO queue), the WQ DMA1031 sends work descriptors to be added to the work descriptor list 1012. If memory is freed in this operation, the WQ DMA1031 sends a buffer pointer corresponding to the freed memory to the buffer pointer queue associated with the packet receiver 1030. If the buffer pointer queue falls below the FIFO threshold, DMA1020 sends one or more buffer pointers from the pre-assigned buffer pointer list 1011 to the buffer pointer queue associated with WQ DMA1031. In some embodiments, a spare DMA channel in the DCX (see, for example, Figure 3C) can be used to perform flow control operations and monitor the filling of the buffer pointer queue and / or work descriptor queue based on a configurable threshold. This spare DMA channel in the DCX can be used as DMA1020.

[0107] Figure 10B shows how data is passed to the packet transmitter 1040 via a portion of the data path within the SDP architecture 1000, preparing it for transmission. The WQ DMA 1041 has an associated work descriptor queue (WQ) and uses the WQ to send work descriptors to the packet transmitter 1040. The packet transmitter 1040 then sends data to a hardware block, for example, based on the data identified in the work descriptor in the work descriptor queue (WQ) associated with the packet transmitter 1040. Once the data is sent, the memory location that contained the data is released, and the packet transmitter 1040 generates a buffer pointer, which is sent to the buffer pointer queue (BP) associated with the WQ DMA 1041.

[0108] In addition, the list of work descriptors may be associated with a DMA1050 that monitors the fill level of the work descriptor queue (WQ) associated with the WQ DMA1041. Whenever the work descriptor queue (e.g., the WQ of the WQ DMA1041 in a FIFO) falls below a threshold (e.g., a FIFO threshold), the DMA1050 may refill the work queue WQ by inputting one or more subsequent work descriptors from the work descriptor list 1021 to keep the work queue sufficiently filled.

[0109] If there is an available buffer pointer in the BP FIFO, the WQ DMA 1041 retrieves the work descriptor in the WQ, reads it from the FIFO memory, moves it to the allocated buffer (indicated by the buffer pointer identified by the work descriptor), and creates a new work descriptor for the next component (for example, the packet transmitter 1040). The new work descriptor is stored in the work queue (WQ) connected to the packet transmitter 1040. The packet transmitter 1040 may then operate similarly. If there is an entry in its work queue, the packet transmitter 1040 uses the work descriptor in it to identify and read the corresponding packet. The packet transmitter 1040 may have a state machine that reads packets according to the work descriptor and sends them to the corresponding component before the memory in the WQ is freed and provided to the WQ DMA 1041 for reuse as a buffer location. Therefore, buffer pointers and work descriptors are used to exchange information about when buffer space will be available for the next packet, and the work descriptor list 1021 is a list of items that should be streamed out by the packet transmitter 1040. Thus, the buffer pointer FIFO operates effectively to control the data flow. Once all buffer pointers have been used and sent to the packet transmitter 1040, the WQ DMA 1041 is stopped and configured to wait for the packet transmitter 1040 to return one of the released buffer pointers. At this point, the WQ DMA 1041 resumes forwarding work descriptors.

[0110] Figure 10C shows an example of data flow for software-based processing using buffer pointers and work descriptors in the SDP architecture 1000. In this example, a sample is received by the packet receiver 1030 of the processing node and forwarded to the DSP core 1032 for processing using the work descriptor queue as described herein. The processed sample is then returned to the packet transmitter 1040 for transmission. This process is essentially a combination of the data paths described herein with reference to Figures 10A and 10B, excluding the buffer pointer list 1011, the work descriptor list 1021, and the associated DMAs 1020 and 1050. Thus, when the DSP core 1032 processes the data, it returns the buffer pointer to the buffer pointer queue associated with WQ DMA 1031 and sends the work descriptor to the work descriptor queue associated with WQ DMA 1041 so that the data can be queued for transmission by the packet transmitter 1040.

[0111] Figure 10D shows an example of data flow stored in on-chip memory using buffer pointers and work descriptors in the SDP architecture 1000. In this example, a sample is received by packet receiver 1030 and placed in a work descriptor queue associated with WQ DMA 1031, which then forwards the work descriptor to a work descriptor queue associated with DSP core 1032. Once processed by DSP core 1032, the processed data is sent to a work descriptor queue associated with WQ DMA 1033, where it can be stored directly in on-chip storage (e.g., CMA). In addition, the processed data can be sent to a work descriptor queue associated with the next consumer 1034 in the signal processing chain, where it can be processed and moved along the signal processing data path. In some embodiments, the work descriptor queue can be stored in the CMA until needed by the next consumer 1034.

[0112] In each of these examples, each WQ DMA can be configured to move data from memory A to memory B. As a result, the WQ DMA creates a working descriptor that has a pointer to memory B. In some embodiments, the WQ DMA also creates a buffer pointer to memory A.

[0113] In some embodiments, the SDP architecture 1000 performs a method for passing data between components (e.g., processing nodes or hardware blocks). This method involves using a buffer pointer queue to manage available memory, the buffer pointer queue comprising a plurality of buffer pointers, each buffer pointer identifying and managing a buffer address in memory available for storing data. The method also involves using a work descriptor queue to manage packets or samples to be processed, the work descriptor queue comprising a plurality of work descriptors, each work descriptor identifying and managing a buffer address in memory containing burst data to be processed. In response that a packet receiver has received burst data to be processed, the method includes obtaining a first buffer pointer from the buffer pointer queue, processing the received burst data, storing the processed burst data in memory at the buffer address identified by the first buffer pointer, and outputting a new work descriptor, the new work descriptor containing the buffer address identified by the first buffer pointer. The method also includes: obtaining a first work descriptor from the work descriptor queue in response to the work descriptor queue having processed data to be transmitted; obtaining the processed data from memory at the buffer address specified by the first work descriptor; freeing the buffer pointer associated with the work descriptor, the freed buffer pointer corresponding to the buffer address specified by the first work descriptor; and transmitting the processed data by a packet transmitter.

[0114] In some embodiments, the work descriptor further includes a length indicating the total length of the packet. The work descriptor may also further include a burst start flag indicating that the burst data belongs to the first packet of the burst, and a burst end flag indicating that the burst data belongs to the last packet of the burst. The work indicator may also indicate that the burst data is a fully contained burst by setting the burst start flag and the burst end flag to true.

[0115] The SDP architecture can also be configured to add new work descriptors to the work descriptor queue and / or add new buffer pointers to the buffer pointer queue.

[0116] In some embodiments, the SDP architecture is further configured to, in response to receiving a buffer pointer queue and a work descriptor queue, retrieve a second work descriptor from the work descriptor queue, retrieve data from memory at the buffer address indicated by the second work descriptor, retrieve a second buffer pointer from the buffer pointer queue, process the retrieved data to generate output processed data, store the output processed data in memory at the buffer address indicated by the second buffer pointer, output a new work descriptor, the new work descriptor including the buffer address indicated by the second buffer pointer, and output a new buffer pointer, the new buffer pointer indicating the buffer address indicated by the second work descriptor. Each work descriptor may further include a burst identifier for the burst data to be processed and a burst length indicating the memory capacity occupied by the burst data to be processed.

[0117] The SDP architecture can also be configured to monitor the fill level of the work descriptor queue by DCX and, in response to determining that the fill level is below a threshold fill level, add one or more work descriptors from the work descriptor list to the work descriptor queue. The SDP architecture can also be configured to monitor the fill level of the buffer pointer queue by DCX and, in response to determining that the fill level is below a threshold fill level, add one or more buffer pointers from the buffer pointer list to the buffer pointer queue.

[0118] Glue from SDP streaming to packetized interface To enable seamless data transfer between components in the disclosed SDP architecture, data can be formatted in an efficient and consistent manner across the architecture. To this end, input RF signals can be digitized and then formatted according to the disclosed data formatting. Once formatted, the data can be transferred through the disclosed SDP architecture using the disclosed working descriptors and buffer pointers. The disclosed data formatting module is configured to create an adaptive layer between streaming and non-streaming, or between packetized data interfaces. The adaptive layer enables connectivity to the disclosed SDP architecture and processing nodes and is computationally efficient.

[0119] In the SDP architecture, data can be input as a burst containing a stream of samples or symbols. Given a burst of sample streams, the burst can be split into multiple messages, which can then be transferred to the DCX for storage and / or processing by the DSP core. In some embodiments, the maximum message size can be configured by the user, and the burst always generates messages of the maximum message size except for the last message. There are two considerations for determining the maximum message size: the convenience of the DMA transfer size and the frequency of message header transfers. The frequency of message header transfers may require frequent updates (e.g., "current frequency offset" or "current symbol timing").

[0120] In the SDP architecture, sample streams received from other locations in the architecture (e.g., on the chip) can be formatted to conform to the target data format, which is referred to herein as streaming mode. Streaming mode is used when the sample stream is provided on the chip (e.g., from a high-speed serial interface as described herein with reference to Figure 7A). With a high-speed serial interface, sample streams can communicate using streaming or non-streaming mode. This enables connection to another FPGA or ASIC with the DCX-based design described herein, where the DCX can directly transmit packets of the DCX configuration over the high-speed serial interface along with the associated packet header metadata. Sample streams or data received via different interfaces or sources can be assumed to have already been formatted using target data formatting, and that data can be received and processed in non-streaming mode.

[0121] The software can configure components of the SDP architecture to receive sample streams (e.g., samples or symbols) from an interface and send sample streams to the interface when operating in streaming mode. The software can configure these components to divide the received sample stream into segments of a size corresponding to a buffer or memory size (e.g., a size specified in the buffer pointer and working descriptor). In some embodiments, when a component of the SDP architecture receives a sample stream, it creates or determines start and end markers based on a known buffer or memory size. As a result, the component divides the input sample stream into smaller segments, numbers the segments, and allows other components of the SDP architecture to manage and process the data using working descriptors and buffer pointers, as described herein.

[0122] In some embodiments, when a component of the SDP architecture receives a packet with determined start and end points, the component can be configured to identify these points and control, for example, when to capture data using a predefined boundary. During transmission, the component can be configured to either pass start and end markers or transmit a valid data signal, at least in part depending on downstream use.

[0123] Therefore, to facilitate flexibility in the disclosed SDP architecture, consistent data formatting is implemented, for example, to implement DSP processing assistance at various points in the signal processing data path, and interfaces between components (e.g., processing nodes and hardware blocks) allow data to be exchanged in a known, favorable format. Accordingly, the disclosed data formatting module provides an adaptive layer configured to process continuous or streaming data, as well as fixed-size or packetized data. The adaptive layer enables the SDP architecture to process both streaming and packetized data by packaging both types of data streams into a common data structure. Advantageously, the disclosed adaptive layer can support user-defined burst data (which may be frames or time slots of any length) because burst boundaries can be preserved during data formatting processing.

[0124] Figure 11A shows an example of a data formatting module 1100 configured to receive digitized data from a high-speed serial receiver module 1110, similar to the HSS / ADC RX or HSS / SDP RX modules described herein with reference to Figure 7A. In some embodiments, streaming data is received from the HSS / ADC RX module, and non-streaming data is received from the HSS / SDP RX module. The data formatting module 1100 is part of a processing node (for example, processing node 210 described herein with reference to Figure 2). The data formatting module 1100 provides a disclosed adaptive layer that enables the storage and processing of streaming and non-streaming data in the SDP architecture. In particular, the data formatting module 1100 formats the data in a way that allows the processing node and hardware blocks to communicate with each other.

[0125] In streaming mode, the high-speed serial receive module 1110 transmits digitized data to the streaming mode component 1120, which then formats the data using the TX configuration module 1125. The TX configuration module 1125 is configured to configure the data to be suitable for transmission to the DCX packet receive interface 1140. The TX configuration module 1125 receives Rx data (e.g., Rx data 1, Rx data 2) and a valid flag indicating that the data is part of a valid burst. In addition, the TX configuration module 1125 receives ready signals or data from the DCX packet receive interface 1140, as well as configuration data from the DCX packet receive interface 1140, as well as data transmitted via the configuration ports described herein. In some embodiments, the data received from the high-speed serial receive module 1110 includes I samples and Q samples.

[0126] The TX configuration module 1125 is configured to package the received data into the format shown in Figure 12A. When received by the DCX packet receiving interface 1140, the data is stored in memory in the format shown in Figure 12B. Along with this data, the TX configuration module 1125 passes the identified frame start (sof) identifier, frame end (eof) identifier, the valid flag received from the high-speed serial receiving module 1110, and the status flag to the DCX packet receiving interface 1140.

[0127] In non-streaming mode, RX data from the high-speed serial receive module 1110 (e.g., HSS / SDP RX module) is assumed to be already formatted in the target data format (shown in Figure 12B). Therefore, the non-streaming mode component 1130 is configured to provide a data reformatting / FIFO module 1135 that converts the word size of the data to the word size expected by the DCX packet receive interface 1140. In some embodiments, the expected word size is 128 bits. The non-streaming mode component 1130 can also be configured to receive a ready flag from the DCX packet receive interface 1140. In addition, the non-streaming mode component 1130 can be configured to receive and pass the identified frame start (sof) identifier, frame end (eof) identifier, and valid flag received from the high-speed serial receive module 1110. The operating mode can be switched between streaming mode and non-streaming mode using a mode flag.

[0128] When received by the DCX packet receiving interface 1140, the sample / symbol data can be passed to other parts of the SDP architecture or to other parts of the processing node, including the DCX packet receiving interface 1140. The TX configuration module 1125 is configured to format the sample / symbol stream into messages to be passed to the DCX, and then pass the data to the SDP architecture. The TX configuration module 1125 is configured to split the burst into multiple messages and forward the multiple messages to the DCX via the DCX packet receiving interface 1140. The TX configuration module 1125 is configured to generate messages of the configured maximum message size, except for the last message.

[0129] Figure 11B shows an example of a data formatting module 1150 configured to receive processed data from the DCX packet transmission interface 1160 and prepare the processed data for the high-speed serial transmission module 1190, similar to the HSS / DAC TX or HSS / SDP TX modules described herein with reference to Figure 7A. The streaming mode component 1170 includes an RX parsing module 1175. The RX parsing module 1175 is configured to process messages received from the DCX packet transmission interface 1160 and reconstruct a sample / symbol stream therefrom. The RX parsing module 1175 is configured to receive data in the format shown in Figure 12B and convert it to the format shown in Figure 12A. The reconstructed sample / symbol stream is then passed to the high-speed serial transmission module 1190 (for example, the HSS / DAC TX module described herein with reference to Figure 7A).

[0130] The non-streaming mode component 1180 includes a data reformatting / FIFO module 1185. The data reformatting / FIFO module 1185 is configured to receive messages from the DCX packet transmission interface 1160 and reformatte the data according to the expected message size. In some embodiments, the FIFO module 1185 is configured to take data coming from the DCX packet transmission interface 1160 and convert that data to the bit width used by the high-speed serial interface. Otherwise, the non-streaming mode component 1180 does not modify the data because the high-speed serial transmission module 1190 is configured to receive data formatted according to the format in Figure 12B. The non-streaming mode data can then be forwarded to the high-speed serial transmission module 1190 (for example, the HSS / SDP TX module described herein with reference to Figure 7A).

[0131] Figures 12A and 12B show the packet format for data in the SDP architecture. Figure 12A shows the packet format for the DCX packet receiving interface 1140 and the DCX packet transmitting interface 1160, as described herein. The packet format shown in Figure 12B is provided by the DCX packet receiving interface 1140 and stored in the DCX of the processing node, or received by the DCX packet transmitting interface 1160 configured with the packet format of Figure 12A.

[0132] Referring to Figure 12A, a packet is received in burst form. The burst has a header, payload data, and a footer. Data reformatting is performed by determining the number of payload messages in the burst, including that information in the footer, and then moving the header and footer together to form the first two words of the burst data (shown in Figure 12B). Advantageously, the component reading the burst data knows that it needs to read the first two words of the data, and from the first two words, the component knows how many payload messages it needs to read to read the entire burst. This provides a more efficient way to read the burst data compared to when the burst data is packaged with an unknown number of payload messages. In some embodiments, the word size is 128 bits. In such embodiments, the combined header and footer are stored in the first 256 bits of the burst data.

[0133] The adaptive layer can create a buffer pointer capable of storing the entire burst because, after reading the header / footer combination, it can determine the exact size in memory for the burst data. For example, when a component receives a footer, it can generate a working descriptor that includes a buffer pointer to the burst data, which includes the header and footer combination and the length of the payload data. In some embodiments, the formatted data includes a burst ID, which includes the lower two bits of the burst counter. In some embodiments, the formatted data includes a burst start flag indicating that the packet belongs to the beginning of a burst. In some embodiments, the formatted data includes a burst end flag indicating that the packet belongs to the end of a burst. In some embodiments, if a packet includes both the burst start flag and the burst end flag, the packet can be considered to contain the entire burst data.

[0134] In some embodiments, the high-speed serial receive module 1110 is configured to decompose incoming data, and the streaming mode component 1120 (via the TX configuration module 1125) is configured to create packets based on the data structure shown in Figure 12A. In some embodiments, the data is divided into messages (e.g., 128-bit words). The high-speed serial receive module 1110 generates a header (e.g., 192 bits), followed by one or more payload messages (e.g., each 128 bits). After the payload data, the high-speed serial receive module 1110 can generate a footer (e.g., 64 bits) to be included after one or more payload messages. A burst counter can be used and can be incremented each time a new burst is received (e.g., each time a new start flag is encountered). A segment counter can be used and can be incremented for each payload message included in the burst data structure. Each burst can be set to any size (e.g., 4 packets, 2 packets, 10 packets, and 1 packet, etc.). The segment counter can be reset each time a new burst is observed.

[0135] Next, the adaptive layer (indicated by the data formatting module 1100) is configured to convert the data into the data structure shown in Figure 12A. Specifically, the RX analysis module 1175 is configured to receive packets in the format shown in Figure 12A and to split them into valid signals. This is the data structure used to store the data within the DCX and / or elsewhere in the SDP architecture. The data structure includes a combined header and footer in the first two words (e.g., 256 bits). This size can be set based on the memory characteristics of the DCX. This allows a component to read all metadata associated with a burst by reading the two words. As a result, the component knows exactly how much data to read from memory to read all the data for a burst. In addition, the RX analysis module 1175 provides a pull interface, and the high-speed serial transmit module 1190 can be configured to control the rate at which data is pulled from this interface.

[0136] In the SDP architecture, messages can be transmitted one word at a time (for example, 128 bits at a time). The data structure can be any value, for example, with a width of 1 word and a depth up to a maximum of 2048. The maximum message depth can be configured in the SDP architecture. For example, the depth can be 16, 32, 64, 128, ..., 2048. The maximum depth includes two header lines, a payload line, and one dead cycle after the end of the burst data.

[0137] In some embodiments, the adaptive layer may perform a method for converting between a sample stream or symbol stream and a message in a signal processing architecture for storage and processing by a processing node including a digital signal processor (DSP) core and an enhanced direct memory access controller (DCX). The method includes receiving a sample stream containing bursts to be processed by the signal processing architecture. The method also includes generating a header message containing information related to the bursts. The method also includes dividing the sample stream into a plurality of burst messages, the size of which each burst message corresponds to the buffer size in the DCX, except for the final burst message corresponding to the end of the burst, the size of which the final burst message is less than or equal to the buffer size in the DCX. The method also includes generating a footer message containing information related to the sizes of the plurality of burst messages. The method also includes forwarding a burst interface packet to the DCX, the burst interface packet comprising the header message, the plurality of burst messages, and the footer message. This method also includes reformatting burst interface packets into burst memory packets and storing them in the DCX, where the burst memory packet includes a header message and a footer message in the first part of the burst memory packet, followed by multiple burst messages in the first part of the burst memory packet. The first part of the burst memory packet indicates the number of burst messages in the burst memory packet.

[0138] This adaptive layer method can also determine the endpoints of bursts in a sample stream by identifying start and end flags in the sample stream. This adaptive layer method can also determine the endpoints of bursts in a sample stream by identifying a first start flag and a second start flag in the sample stream, where the endpoints are the first start flag and data that precedes the second start flag but does not include the second start flag.

[0139] In some embodiments, splitting a sample stream into multiple burst messages is in response to identifying a frame start indicator within the sample stream. In some embodiments, splitting a sample stream into multiple burst messages is terminated in response to identifying a frame end indicator within the sample stream.

[0140] In some embodiments, in non-streaming mode, the adaptive layer is configured to convert a first word size of data in a burst interface packet to a second word size compatible with DCX, where the second word size is larger than the first word size.

[0141] Management and Use of Data Storage Devices To enhance the flexibility offered by the SDP architecture, memory access within the SDP architecture can be strengthened. Strengthening memory access reduces the likelihood of memory access becoming a bottleneck in terms of actual storage capacity and the access speed of the storage device. The memory architecture within the SDP can be divided into banks and / or channels to simultaneously provide multiple memory access capabilities. The memory read / write logic that may be included in SDP memory (e.g., DCX, CMA, or shared memory in other on-chip memory) not only provides very high read / write throughput by performing read / write interleaving across multiple memory banks (e.g., RAM), but also provides flexibility regarding changes in the data format or resolution of data written to and read from memory. For SDP memory to effectively provide the flexibility to support communication functions (e.g., modem functions), it is important that the memory access process does not become a bottleneck in terms of actual storage capacity (supported by data reformatting) and access speed (interleaving across multiple banks).

[0142] Figure 13 shows a memory module 1300 within the DCX. The memory module 1300 includes a read port 1301, a write port 1303, a memory arbiter 1305, and a memory bank 1307. The memory module 1300 is divided into multiple channels and banks, providing simultaneous multiple access to memory. Channels can be interleaved at a higher level, and banks at a lower level. Each memory module can be selected based on the bank and channel derived from the requested address.

[0143] The memory module 1300 can be a single-port memory configured to allow either reading or writing in a single clock cycle. Interleaving multiple of these memory modules can achieve higher read and write bandwidth because each of these memory modules can be read from or written to simultaneously. The memory module 1300 can be configured and operated in a way that effectively creates a multi-port memory (for example, a quad-port memory allowing four simultaneous reads or writes per clock cycle). For large data transfers, the read port 1301 and / or write port 1303 can be configured to access the memory sequentially. The memory module 1300 can perform single random reads and writes. However, for large data transfers, it can be configured to sequentially increment the requested address per clock cycle to achieve higher data transfer rates.

[0144] The memory arbiter 1305 is configured to manage access to the memory bank 1307 when simultaneous access is requested. In some embodiments, the memory arbiter 1305 can determine access priority by assigning higher priority to requests received from hardware blocks of the signal processing architecture. In some embodiments, the memory arbiter 1305 can determine access priority by assigning lower priority to the component that last accessed the memory module 1300. Other requests are delayed by a clock cycle before access is re-determined.

[0145] Figure 14A shows an exemplary memory module 1400 of a capture memory array 1410 (CMA) which includes a memory bank 1407 divided into multiple channels and banks to simultaneously provide multiple access functionality for memory. The CMA 1410 is similar to the CMA described herein with reference to Figures 6A, 6B, and 7A. The channels may be interleaved at a higher level, and the banks may be interleaved at a lower level. The memory module 1400 also includes a read port 1401, a write port 1403, and a memory arbiter 1405. The memory arbiter 1405 operates similarly to the memory arbiter 1305 described herein with reference to Figure 13. Figure 14B shows the CMA 1410 with multiple memory modules 1400. Each memory module 1400 can be selected based on the bank and channel derived from the requested address.

[0146] Each memory bank 1407 may consist of eight smaller single-port RAMs. These RAMs may be arranged in four groups interleaved by the lower address bits, with each consecutive access sent to a separate RAM. Furthermore, these two groups of four RAMs may constitute the upper and lower halves of the memory area. Each of these memories may use a handshake signal to indicate the process (read / write) requesting access. If access requests are sent to the same RAM, the memory arbiter 1405 can be used to determine which request to delay. If access requests are sent to separate RAMs, the read and write processes can be executed in parallel. If read and write requests are received in the same RAM, arbitration logic can be used to determine when the last access to the RAM occurred and which (e.g., read or write) should access it. This ensures that both the read and write processes have fair access to the RAM, even if one of the processes attempts to access the same RAM consecutively, thereby avoiding a lockup state.

[0147] A memory module can be divided into multiple channels and banks, providing simultaneous multiple access to the memory module. Channels can be interleaved at a higher level, while banks can be interleaved at a lower level. Each memory module can be selected based on the bank and channel derived from the requested address.

[0148] Each memory bank 1407 can be a single-port memory, allowing either a read or write process in a single clock cycle. However, interleaving multiple memory modules can achieve higher read / write bandwidth because it allows simultaneous reading or writing from different memory modules.

[0149] In some embodiments, the memory module 1400 may create a quad-port memory that allows four simultaneous reads or writes per clock cycle. The ports may be used for large data transfers and thus, for example, sequential access to the memory. Single random reads and writes may be permitted for large data transfers, but higher data transfer rates can be achieved by sequentially increasing the requested address per clock cycle. The memory module 1400 may be configured to have an interface configured to allow requests from a network interconnect to read and write to the CMA 1410. The memory module 1400 may also have a specific handshake-based interface that allows the DMA to directly access the memory bank 1407.

[0150] In some embodiments, the CMA1410 may include memory banks capable of storing a large number of data samples. These memories may be contiguous within a memory map. Each memory bank may be a dual-port memory and may be designed using multiple interleaved single-port memories and arbiters, as described herein.

[0151] In some embodiments, a memory arbiter (e.g., memory arbiter 1305 or memory arbiter 1405) performs a method for controlling access to memory in a signal processing architecture. The memory may be part of on-chip memory (e.g., a capture memory array (CMA)) or part of a DCX. The memory includes a plurality of random access memory (RAM) modules, each RAM module being logically divided into a plurality of sequentially arranged memory banks. The method includes receiving a plurality of requests for access to a RAM module in memory, each of which includes a memory address in memory corresponding to a memory bank in a particular RAM module. The method also includes, for each request, deriving a specific bank among the plurality of banks in the RAM module from the memory address in the request, where the specific bank includes the memory address in the request. The method also includes, in response to determining that two of the plurality of requests are requesting access to the same bank in the same RAM module, determining a priority between the two requests, granting access to the requested bank to the request with higher priority, and delaying the request with lower priority by a clock cycle. This method also includes, for each request requesting contiguous access to memory, granting access to a memory bank that is contiguous with the bank in the request, among a plurality of memory banks.

[0152] In some embodiments, multiple banks are interleaved at a low level, and the number of banks is a power of 2. In some embodiments, multiple RAM modules are further divided into multiple channels, and the number of channels is a power of 2. In such embodiments, multiple channels can be interleaved at a higher level.

[0153] In some embodiments, the memory arbiter is further configured to allow simultaneous access to the two requested banks in response to determining that two of several requests result in requests to access different banks within the same RAM module. In some embodiments, determining priority includes assigning a lower priority to the request that last accessed the requested RAM module.

[0154] Additional Embodiments and Terminology This disclosure describes a variety of features, but no single one of them alone accounts for the advantages described herein. As will be apparent to those skilled in the art, the various features described herein can be combined, modified, or omitted. Other combinations and partial combinations not specifically described herein will be apparent to those skilled in the art and are intended to form part of this disclosure. Various methods are described herein in relation to various flowchart steps and / or stages. In many cases, it should be understood that certain steps and / or stages can be combined so that multiple steps and / or stages shown in a flowchart can be performed as a single step and / or stage. Also, certain steps and / or stages can be divided into additional subcomponents and performed separately. In some cases, the order of steps and / or stages can be rearranged, and certain steps and / or stages can be omitted entirely. It should also be understood that the methods described herein are open-ended so that additional steps and / or stages can be performed beyond those illustrated and described herein.

[0155] Some aspects of the systems and methods described herein can, advantageously, be implemented using, for example, computer software, hardware, firmware, or any combination of computer software, hardware, and firmware. The computer software may include computer executable code stored in a computer-readable medium (e.g., non-temporary computer-readable medium) which, when executed, performs the functions described herein. In some embodiments, the computer executable code is executed by one or more general-purpose computer processors. As those skilled in the art will see, any feature or function that can be implemented using software to run on a general-purpose computer in consideration of this disclosure can also be implemented using different combinations of hardware, software, or firmware. For example, such a module can be fully implemented in hardware using a combination of integrated circuits. Alternatively or additionally, such features or functions can be fully or partially implemented using a dedicated computer designed to perform the specific functions described herein, rather than a general-purpose computer.

[0156] Multiple distributed computing devices can be used instead of any one of the computing devices described herein. In such a distributed embodiment, the functions of one computing device are distributed (for example, over a network), and some functions are performed on each of the distributed computing devices.

[0157] Some embodiments may be described with reference to formulas, algorithms, and / or flowcharts. These methods may be implemented using computer program instructions executable on one or more computers. These methods may also be implemented individually as computer program products or as components of a device or system. In this regard, each formula, algorithm, block, or step in a flowchart, and combinations thereof, may be implemented by hardware, firmware, and / or software, including one or more computer program instructions embodied in computer-readable program code logic. It will be understood that any such computer program instruction may be loaded onto one or more computers (e.g., general-purpose computers or dedicated computers, without limitation), or other programmable processing devices that generate a machine, and as a result, computer program instructions executed on a computer or other programmable processing device(s) implement the function specified in the formula, algorithm, and / or flowchart. It will also be understood that each formula, algorithm, and / or block in a flowchart, and combinations thereof, may be implemented by a dedicated hardware-based computer system that performs the specified function or step, or by a combination of dedicated hardware and computer-readable program code logic means.

[0158] Furthermore, computer program instructions (e.g., those embodied in computer-readable program code logic) can also be stored in computer-readable memory (e.g., non-temporary computer-readable media) that can instruct one or more computers or other programmable processing devices to function in a particular way, and as a result, instructions stored in computer-readable memory implement functions specified in blocks of flowcharts. Computer program instructions can also be loaded onto one or more computers or other programmable computing devices and perform a series of operational steps on one or more computers or other programmable computing devices to generate a computer implementation process, and as a result, instructions executed on computers or other programmable processing devices provide steps for performing functions specified in expressions, algorithms, and / or blocks of flowcharts.

[0159] Some or all of the methods and tasks described herein may be performed by a computer system and may be fully automated. The computer system may, in some cases, include several separate computers or computing devices (e.g., physical servers, workstations, storage arrays, etc.) that communicate and interoperate over a network to perform the functions described. Each such computing device typically includes a processor (or more processors) that executes program instructions or modules stored in memory or other non-temporary computer-readable storage media or devices. While the various functions disclosed herein may be embodied in such program instructions, some or all of the disclosed functions may, alternatively, be implemented in application-specific circuits (e.g., ASICs or FPGAs) of the computer system. If the computer system includes multiple computing devices, these devices may, though not required, be located in the same place. The results of the disclosed methods and tasks may be permanently stored by converting physical storage devices (e.g., solid-state memory chips and / or magnetic disks) into different states.

[0160] Unless the context clearly indicates otherwise, throughout the specification and claims, words such as “comprise” and “comprising” should be interpreted in a comprehensive sense, i.e., “without limitation,” and not in an exclusive or exhaustive sense. The word “combined,” as used herein, means that two or more elements may be directly connected or connected through one or more intermediate elements. Furthermore, the words “as herein,” “above,” “below,” and similar words, as used in this application, refer to the entire application and not to any particular part thereof. Where contextually permissible, words used singular or plural in the above detailed descriptions may also include plural or singular, respectively. With respect to a list of two or more items, the word “or” means that any item in the list encompasses all items in the list, and all any combination of items in the list. The word “typical” is used herein only to mean “serving as an example, case, or illustration.” Any embodiment described as “typical” in this specification should not be construed as necessarily preferable or advantageous to other embodiments.

[0161] This disclosure is not intended to be limited to the embodiments shown herein. Various modifications to the embodiments described herein may be readily apparent to those skilled in the art, and the general principles set forth herein may be applied to other embodiments without departing from the spirit or scope of this disclosure. The teachings of the present invention provided herein may be applied to other methods and systems, and are not limited to the methods and systems described herein, and further embodiments may be provided by combining elements and operations of the various embodiments described herein. Thus, the new methods and systems described herein may be embodied in various other forms, and furthermore, various omissions, substitutions, and modifications may be made in the forms of the methods and systems described herein without departing from the spirit of this disclosure. The appended claims and their equivalents are intended to cover forms or modifications contained within the scope and spirit of this disclosure.

Claims

1. A processing node (PN), The first digital signal processor (DSP) core, The second DSP core, A plurality of extended direct memory access controllers (DCXs), each having a shared memory space, an input packet interface, and an output packet interface, wherein the input packet interface is configured to receive samples from a hardware block separate from the processing node, the shared memory space is configured to store the received samples, and the output packet interface is configured to transmit samples processed by the first DSP core or the second DSP core to the hardware block, A PN network interconnect configured to communicately connect the first DSP core, the second DSP core, and the plurality of DCXs, wherein each DSP core and DCX is connected to the PN network interconnect via its respective master interface and its respective slave interface, and the PN network interconnect further includes an SDP master interface and an SDP slave interface, each of which is configured to communicate with the SDP network interconnect. Includes, The processing node is integrated into the wireless transceiver including the hardware block, and is configured to operate in connection with the hardware block to provide configurable processing functions to the wireless transceiver. Processing node.

2. The processing node according to claim 1, wherein the PN network interconnect further includes a configuration interface configured to enable the processing node to constitute the hardware block.

3. The processing node according to claim 1, further comprising a queue interface configured to transfer commands or data from the first DSP core to the second DSP core and from the second DSP core to the first DSP core.

4. The processing node according to claim 1, further comprising a first queue interface and a second queue interface, wherein the first queue interface is configured to transfer commands or data from the first DSP core to the second DSP core, and the second queue interface is configured to transfer commands or data from the second DSP core to the first DSP core.

5. The processing node according to claim 1, wherein each DSP core includes a general-purpose input / output (GPIO) port connected to a configuration register, the GPIO port being configured to receive input for placement in the configuration register and to transmit data stored in the configuration register.

6. The processing node according to claim 1, wherein each DSP core is configured to receive interrupt requests from a hardware block separate from the processing node via the PN network interface.

7. The first DCX among the aforementioned multiple DCXs is The process involves receiving multiple samples to be processed from a first hardware block via the input packet interface, The plurality of samples are temporarily stored in the shared memory space, The plurality of samples are transmitted to the first DSP core via the PN network interconnection, A processing node according to claim 1, configured to perform the following:

8. The first DSP core or the second DSP core is Programming the first DCX to transmit the plurality of samples to the first DSP core, The plurality of samples are arranged in the internal memory space of the first DSP core, Processing the aforementioned multiple samples, The processed sample is placed in the internal memory space of the first DSP core, The processing node according to claim 7, configured to perform the following:

9. The processing node according to claim 7, wherein the first DCX is further configured to reformat the received plurality of samples.

10. The processing node according to claim 9, wherein the first DCX is configured to reformat the received plurality of samples by sign-extending the samples of the received plurality of samples to increase the number of bits for each sample.

11. The processing node according to claim 9, wherein the first DCX is configured to reformat the received plurality of samples by bit-clipping the samples of the plurality of received samples to reduce the resolution of each sample.

12. The first DCX among the aforementioned multiple DCXs is Receiving multiple samples to be processed via the aforementioned input packet interface, The plurality of samples are temporarily stored in the shared memory space, The plurality of samples are sent to the hardware block separate from the processing node and subjected to processing using the output packet interface. A processing node according to claim 1, configured to perform the following:

13. The first DCX described above is Receiving the processed samples from the hardware block, The process of temporarily storing the processed samples in the shared memory space, The processing node according to claim 12, configured to perform the following:

14. The processing node according to claim 1, wherein the first DSP core and the second DSP core are configured to be used as separate entities or as a shared dual-core configuration.

15. The processing node according to claim 1, wherein the first DSP core and the second DSP core each include two processors, and the plurality of DCXs include DCXs for each processor of the first DSP core and the second DSP core.

16. The processing node according to claim 1, wherein the SDP master interface and the SDP slave interface of the PN network interconnect are configured to communicate with the PN network interconnects of different processing nodes within the wireless transceiver via the SDP network interconnect.

17. The processing node according to claim 1, wherein each of the first DSP core, the second DSP core, and the plurality of DCXs includes a configuration register configured to store data for configuring the associated DSP core or DCX.

18. The processing node according to claim 1, wherein the processing node is configured to be implemented within the demodulator of the wireless transceiver.

19. The processing node according to claim 1, wherein the processing node is configured to be implemented within the decoder of the wireless transceiver.

20. The processing node according to claim 1, wherein the processing node is configured to be implemented within the modulator or encoder of the transmitter of the wireless transceiver.

21. It is a signal processing architecture, Software-defined physical layer (SDP) network interconnection, A plurality of processing nodes connected to the SDP network interconnect and configured to provide configurable processing capabilities for processing the waveforms of the receiver and transmitter in a wireless transceiver, each processing node is: Multiple digital signal processing (DSP) cores, Multiple Extended Direct Memory Access Controllers (DCX), The plurality of DSP cores, the plurality of DCXs, and the PN network interconnect connected to the SDP network interconnect, including, Multiple processing nodes, A capture memory array (CMA) comprising multiple memory banks, wherein the multiple memory banks are connected to the SDP network interconnection and provide access to the multiple processing nodes, A CPU subsystem connected to the aforementioned SDP network interconnection, Includes, The SDP network interconnection enables communication between the multiple processing nodes, the CMA, and the CPU subsystem, thereby enhancing the processing power and functionality of the wireless transceiver. Signal processing architecture.

22. The signal processing architecture according to claim 21, wherein one or more of the plurality of processing nodes can be dynamically assigned to one or more hardware blocks of the wireless transceiver to provide signal processing capability.

23. The signal processing architecture according to claim 21, wherein each of the plurality of processing nodes is configured to operate in connection with one or more separate hardware blocks in both the receiver and transmitter signal processing data paths within the wireless transceiver.

24. The signal processing architecture according to claim 21, wherein the first processing node among the plurality of processing nodes is implemented in the encoder or modulator of the wireless transceiver.

25. The signal processing architecture according to claim 24, wherein the second processing node among the plurality of processing nodes is implemented in the demodulator of the wireless transceiver.

26. The signal processing architecture according to claim 25, wherein the third processing node among the plurality of processing nodes is implemented in the decoder of the wireless transceiver.

27. The signal processing architecture according to claim 21, further comprising an external memory connected to the CPU subsystem, wherein the plurality of processing nodes are configured to pass data from individual DSP cores to the external memory via the SDP network interconnect.

28. The signal processing architecture according to claim 21, wherein each of the plurality of processing nodes is integrated within a different part of a demodulator.

29. The signal processing architecture according to claim 21, wherein each processing node includes an SDP master interface and an SDP slave interface to the SDP network interconnect, the CMA includes a plurality of SDP master interfaces to the SDP network interconnect, and the CPU subsystem includes an SDP master interface and an SDP slave interface to the SDP network interconnect.

30. A second SDP network interconnection connected to the aforementioned SDP network interconnection, A second plurality of processing nodes connected to the second SDP network interconnect, each of the second plurality of processing nodes comprising: one or more digital signal processing (DSP) cores; one or more extended direct memory access controllers (DCX); and a PN network interconnect connected to the plurality of DSP cores, the plurality of DCX, and the second SDP network interconnect, The signal processing architecture according to claim 21, further comprising:

31. A method for passing data to a processing node in a signal processing architecture, wherein the signal processing architecture includes a software-defined physical layer (SDP) network interconnect connected to the processing node, the processing node includes a digital signal processing (DSP) core, an enhanced direct memory access controller (DCX), a packet receiver, a packet transmitter, and a PN network interconnect connected to the SDP network interconnect, and the method is The method involves using a buffer pointer queue to manage available memory, wherein the buffer pointer queue includes multiple buffer pointers, and each buffer pointer identifies and manages a buffer address in memory available for storing data. The process involves using a work descriptor queue to manage burst data to be processed, wherein the work descriptor queue includes multiple work descriptors, and each work descriptor identifies and manages the memory buffer address containing the burst data to be processed. In response to the packet receiver receiving burst data to be processed, Obtaining a first buffer pointer from the buffer pointer queue, Processing the received burst data, The processed burst data is stored in memory at the buffer address specified by the first buffer pointer, Outputting a new work descriptor, wherein the new work descriptor includes the buffer address identified by the first buffer pointer. In response to the work descriptor queue having processed the data to be sent, Obtaining a first work descriptor from the aforementioned work descriptor queue, Obtaining the processed data from the memory located at the buffer address specified by the first work descriptor, Outputting a new buffer pointer, wherein the new buffer pointer corresponds to the buffer address identified by the first work descriptor. The processed data is transmitted by the packet transmitter, A method that includes this.

32. The method according to claim 31, wherein the work descriptor further includes a data header length indicating the storage capacity occupied by the data header associated with the burst data.

33. The aforementioned work descriptor further, A burst start flag indicating that the burst data belongs to the first packet of the burst, A burst end flag indicating that the burst data belongs to the last packet of the burst, The method according to claim 32, including the method described in claim 32.

34. The method according to claim 33, wherein the work indicator indicates that the burst data is fully contained in a burst by setting the burst start flag and the burst end flag to true.

35. The method according to claim 31, further comprising adding the new work descriptor to the work descriptor queue.

36. The method according to claim 31, further comprising adding the new buffer pointer to the buffer pointer queue.

37. In response to receiving the buffer pointer queue and the work descriptor queue, Obtaining a second work descriptor from the aforementioned work descriptor queue, Retrieving data from the memory located at the buffer address indicated by the second work descriptor, Obtaining a second buffer pointer from the aforementioned buffer pointer queue, The acquired data is processed to generate output processed data, The processed output data is stored in memory at the buffer address indicated by the second buffer pointer, Outputting a new work descriptor, wherein the new work descriptor includes the buffer address indicated by the second buffer pointer. Outputting a new buffer pointer, wherein the new buffer pointer indicates the buffer address indicated by the second work descriptor. The method according to claim 31, further comprising:

38. The method according to claim 37, wherein each work descriptor further includes a burst identifier for the burst data to be processed and a burst length indicating the storage capacity occupied by the burst data to be processed.

39. Monitoring the fill level of the work descriptor queue by the DCX, In response to determining that the filling level is below the threshold filling level, one or more work descriptors are added to the work descriptor queue from the work descriptor list, The method according to claim 31, further comprising:

40. The DCX monitors the fill level of the buffer pointer queue, In response to determining that the fill level is below the threshold fill level, one or more buffer pointers are added to the buffer pointer queue from the buffer pointer list. The method according to claim 31, further comprising:

41. A signal processing architecture for storage and processing by a processing node including a digital signal processor (DSP) core and an extended direct memory access controller (DCX), wherein a method for converting between a sample stream or symbol stream and a message, the method is: The signal processing architecture receives a sample stream containing bursts to be processed, To generate a header message containing information related to the aforementioned burst, The method of dividing the sample stream into multiple burst messages, wherein the size of each burst message corresponds to the buffer size in the DCX, except for the final burst message corresponding to the end of the burst, and the size of the final burst message is less than or equal to the buffer size in the DCX. To generate a footer message containing information related to the size of the aforementioned multiple burst messages, The transfer of a burst interface packet to the DCX, wherein the burst interface packet includes the header message, the plurality of burst messages, and the footer message. The burst interface packet is reformatted into a burst memory packet and stored in the DCX, wherein the burst memory packet includes the header message and the footer message in the first part of the burst memory packet, and the plurality of burst messages following the first part of the burst memory packet. Includes, The first portion of the burst memory packet indicates the number of burst messages in the burst memory packet. method.

42. The method according to claim 41, wherein the sample stream is received by a component of the processing node.

43. The method according to claim 41, further comprising identifying a start flag and an end flag in the sample stream to determine the endpoint of the burst in the sample stream.

44. The method according to claim 41, further comprising identifying a first start flag and a second start flag in the sample stream to determine the endpoint of the burst in the sample stream, wherein the endpoint is the first start flag and data preceding the second start flag but not including the second start flag.

45. The method according to claim 41, wherein the bit size of each packet in the sample stream is equal to the size of the word in the DCX.

46. The method according to claim 41, wherein dividing the sample stream into the plurality of burst messages is in response to identifying a frame start indicator in the sample stream.

47. The method according to claim 41, wherein dividing the sample stream into the plurality of burst messages is terminated in response to identifying a frame end indicator in the sample stream.

48. The method according to claim 41, wherein the header message includes a burst counter that increments in response to identifying the boundary of the burst.

49. The method according to claim 41, wherein the size of the first portion of the burst memory packet is set to be less than or equal to the size of two words in the memory within the DCX.

50. The method according to claim 41, wherein reformatting the burst interface packet further comprises converting a first word size of data within the burst interface packet to a second word size compatible with the DCX, wherein the second word size is larger than the first word size.

51. A method for accessing memory in a signal processing architecture including a capture memory array (CMA), wherein the CMA includes a plurality of random access memory (RAM) modules, each RAM module being logically divided into a plurality of sequentially arranged memory banks, and the method is: Receiving multiple requests for accessing RAM modules within the CMA, wherein each of the multiple requests includes a memory address within the CMA corresponding to a memory in a specific RAM module. For each request, the derivation involves deriving a specific bank among the plurality of banks in the RAM module from the memory address in the request, wherein the specific bank includes the memory address in the request. In response to determining that two of the aforementioned requests are requesting access to the same bank within the same RAM module, To determine the priority between the two aforementioned requests, The request with the higher priority of the two requests above will be granted access to the requested bank, The request with the lower priority of the two requests mentioned above will be delayed by the clock cycle, For each request that requests continuous access to the CMA, access to a bank following the bank in the request is permitted from among the multiple memory banks. A method that includes this.

52. The method according to claim 51, wherein the plurality of banks are interleaved at a low level, and the number of banks in the plurality of banks is a power of 2.

53. The method according to claim 51, wherein the plurality of RAM modules are further divided into a plurality of channels, and the number of channels in the plurality of channels is a power of 2.

54. The method according to claim 53, wherein the plurality of channels are interleaved in a higher order.

55. The method according to claim 51, further comprising allowing simultaneous access to the two requests for the respective requested banks in response to determining that two of the requests result in requests to access different banks within the same RAM module.

56. The method according to claim 51, wherein determining the priority includes assigning a lower priority to the request that last accessed the requested RAM module.