Parallel Turbo Decoder Scheduling for Non-Uniform Trellis Windows
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current turbo decoders, particularly those using the Log-BCJR algorithm, have limited processing throughput and high processing latency due to their serial nature, which restricts the achievable transmission throughput and end-to-end latency in wireless communication systems, especially for next-generation standards requiring multi-gigabit transmission and ultra-low latency.
Innovation Solution
A turbo decoder circuit with a configurable network for interleaving soft decision values, allowing parallel processing of trellis stages across multiple processing elements, which enables dynamic window sizing and reduces data dependencies, thereby improving decoding rate and throughput.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the Log-BCJR algorithm is used for iterative decoding of turbo codes, then reliable communication with strong error correction capability is achieved, but processing throughput is limited and processing latency is high due to the inherently serial nature of the algorithm
Solution Approach 1:
The trellis structure is divided into multiple independent segments or windows, each processed by separate processing elements. This segmentation allows parallel execution of decoding operations across different trellis segments while maintaining the error correction capabilities of the original Log-BCJR algorithm through proper boundary handling and message passing between segments.
Solution Approach 2:
The patent introduces a spatial dimension for parallel processing by distributing processing elements across multiple processors or cores. The serial time-based processing of the Log-BCJR algorithm is transformed into a parallel space-based processing architecture, where multiple processing elements operate simultaneously on different portions of the decoding task, thereby increasing throughput without compromising reliability.
2Reliability
If the Log-BCJR algorithm is applied alternately to two convolutional codes with numerous consecutive time periods, then sufficient decoding iterations are performed for reliable communication, but end-to-end latency is increased
Solution Approach 1:
The iterative decoding process is segmented into parallel operations where multiple processing elements work simultaneously on different aspects of the decoding task. This reduces the total time required to complete the same number of decoding iterations, thereby lowering latency while maintaining decoding accuracy through coordinated information exchange between parallel elements.
Solution Approach 2:
The patent implements continuous parallel processing where multiple processing elements operate continuously without idle time, performing useful decoding operations throughout the entire processing period. This eliminates the sequential waiting periods inherent in traditional alternating application of the Log-BCJR algorithm, reducing latency while maintaining reliable communication through sustained decoding activity.
3Device complexity
If a fixed number of processing elements is used for parallel turbo decoding, then hardware complexity is controlled, but adaptability to different frame sizes and processing requirements is limited
Solution Approach 1:
The patent implements dynamic configuration of processing elements where the number and arrangement of active processing elements can be adjusted based on the incoming frame size and decoding requirements. This dynamic adaptability allows the same hardware architecture to efficiently handle varying frame sizes without requiring a fixed large number of processing elements, thereby controlling hardware complexity while enhancing versatility.
Solution Approach 2:
The processing elements are designed with universal functionality to handle different frame sizes, code rates, and decoding scenarios. Each processing element can be dynamically allocated to different tasks based on requirements, making the hardware architecture versatile and adaptable to various communication standards and frame configurations without increasing the base number of processing elements.
Data Source
AI summary
A turbo decoder circuit performs a turbo decoding process to recover a frame of data symbols from a received signal comprising soft decision values for each data symbol of the frame. The data symbols of the frame have been encoded with a turbo encoder comprising upper and lower convolutional encoders which can each be represented by a trellis, and an interleaver which interleaves the encoded data between the upper and lower convolutional encoders. The turbo decoder circuit comprises a clock, a configurable network circuitry for interleaving soft decision values, an upper decoder and a lower decoder. Each of the upper and lower decoders include processing elements, which are configured, during a series of consecutive clock cycles, iteratively to receive, from the configurable network circuitry, a priori soft decision values pertaining to data symbols associated with a window of an integer number of consecutive trellis stages representing possible paths between states of the upper or lower convolutional encoder. The processing elements perform parallel calculations associated with the window using the a priori soft decision values in order to generate corresponding extrinsic soft decision values pertaining to the data symbols. The configurable network circuitry includes network controller circuitry which controls a configuration of the configurable network circuitry iteratively, during the consecutive clock cycles, to provide the a priori soft decision values for the upper decoder by interleaving the extrinsic soft decision values provided by the lower decoder, and to provide the a priori soft decision values for the lower decoder by interleaving the extrinsic soft decision values provided by the upper decoder. The interleaving performed by the configurable network circuitry controlled by the network controller is in accordance with a predetermined schedule, which provides the a priori soft decision values at different cycles of the one or more consecutive clock cycles to avoid contention between different a priori soft decision values being provided to the same processing element of the upper or the lower decoder during the same clock cycle. Accordingly the processing elements can have a window size which includes a number of stages of the trellis so that the decoder can be configured with an arbitrary number of processing elements, making the decoder circuit an arbitrarily parallel turbo decoder.


