Fully Parallel Turbo Decoding Using Trellis-Stage Processing Elements
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional turbo decoders, such as those using the Log-BCJR algorithm, are limited by their serial nature and data dependencies, resulting in low processing throughputs and high latencies, which are inadequate for next-generation wireless communication standards requiring multi-gigabit throughputs and ultra-low end-to-end latencies.
Innovation Solution
A fully parallel turbo detection circuit is introduced, where processing elements associated with trellis stages operate autonomously and in parallel, combining a priori forward and backward state metrics with soft decision values to generate extrinsic metrics, allowing simultaneous processing of all trellis stages in each clock cycle, thereby reducing data dependencies and increasing processing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the Log-BCJR algorithm is used for iterative decoding, then decoding performance is improved, but processing throughput is limited and latency increases due to serial processing
Solution Approach 1:
The decoding circuit is divided into multiple processing elements, each handling a specific trellis stage independently. Each processing element contains dedicated functional units for forward recursion, backward recursion, and extrinsic information calculation, enabling parallel processing of different stages while maintaining the iterative decoding performance of the Log-BCJR algorithm
Solution Approach 2:
The patent transforms the serial time-domain processing of the Log-BCJR algorithm into a parallel space-domain architecture by mapping trellis stages to spatial processing elements. This dimensional transformation allows all trellis stages to be processed simultaneously in each clock cycle, achieving multi-gigabit throughput while preserving decoding performance
2Reliability
If the Log-BCJR algorithm is used for iterative decoding, then decoding performance is improved, but processing latency increases due to sequential operations
Solution Approach 1:
The decoding circuit is divided into multiple processing elements, each handling a specific trellis stage independently. Each processing element contains dedicated functional units for forward recursion, backward recursion, and extrinsic information calculation, enabling parallel processing of different stages while maintaining the iterative decoding performance of the Log-BCJR algorithm
Solution Approach 2:
The patent pre-computes and stores transition metrics for all trellis stages before decoding begins. During iterative decoding, processing elements directly use these pre-computed metrics without recalculation, reducing computation time and latency while maintaining decoding accuracy
3Productivity
If fully parallel processing is implemented, then processing throughput is improved, but data dependencies must be resolved to enable parallel operation
Solution Approach 1:
The decoding circuit is divided into multiple processing elements, each handling a specific trellis stage independently. Each processing element contains dedicated functional units for forward recursion, backward recursion, and extrinsic information calculation, enabling parallel processing of different stages while maintaining the iterative decoding performance of the Log-BCJR algorithm
Solution Approach 2:
The patent introduces buffer memory as an intermediary between processing elements to store and exchange forward state metrics, backward state metrics, and extrinsic information. This intermediary mechanism resolves data dependencies by providing temporary storage, enabling parallel processing while managing circuit complexity through structured data flow control
Data Source
AI summary
A circuit performs a turbo detection process recovering data symbols from a received signal effected, during transmission, by a Markov process with effect that the data symbols are dependent on preceding data symbols represented as a trellis having a plurality of trellis stages. The circuit comprises processing elements, associated with trellis stages representing these dependencies and each configured to receive soft decision values corresponding to associated data symbols Each processing element configured, in one clock cycle to receive data representing a priori forward and backward state metrics, and a priori soft decision values for data symbols detected for the trellis stage. For each clock cycle of the turbo detection process, the circuit processes, for processing elements representing the trellis stages, the a priori information for associated data symbols detected for the trellis stage, and to provide extrinsic soft decision values corresponding to data symbols for a next clock cycle.


