Low-complexity parallel chien search architecture based on matrix decomposition
By decomposing the Vandermonde matrix into matrices with fewer non-zero entries and reformulating the Chien search architecture, the complexity and power consumption of parallel Chien search are significantly reduced, addressing the inefficiencies of conventional architectures.
Patent Information
- Application Number
- PCT/US2024/059507
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-12-11
- Publication Date
- 2025-06-19
AI Technical Summary
Conventional parallel Chien search architectures for Reed-Solomon (RS) and Bose-Chaudhuri-Hocquenghem (BCH) decoding are complex and inefficient, leading to high gate counts and power consumption, despite efforts to reduce complexity through substructure sharing and modular reduction.
The proposed solution involves decomposing the Vandermonde matrix used in the Chien search into two or more matrices with fewer non-zero entries, allowing for reduced multiplications and complexity in the parallel Chien search architecture. This decomposition is achieved using standard basis representation of finite field elements and reformulating segments of the Vandermonde matrix to share constant multipliers.
The proposed low-complexity parallel Chien search architecture achieves a significant reduction in hardware complexity, with area reductions of 26% to 31% compared to conventional designs for 9-error-correcting RS/BCH codes, while also reducing power consumption.
Smart Images

Figure US2024059507_19062025_PF_FP_ABST
Abstract
Description
Atty. Dkt. No.103361-613WO1 LOW-COMPLEXITY PARALLEL CHIEN SEARCH ARCHITECTURE BASED ON MATRIX DECOMPOSITION CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 608,880, filed December 12, 2023, entitled “LOW-COMPLEXITY PARALLEL CHIEN SEARCH ARCHITECTURE BASED ON MATRIX DECOMPOSITION,” the disclosure of which is expressly incorporated herein by reference in its entirety. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under grant no.2011785 awarded by the National Science Foundation. The government has certain rights in the invention. BACKGROUND
[0003] Reed-Solomon (RS) and BCH codes are widely used in digital communication and storage systems, such as optical transport networks, digital video broadcasting, and flash memory. The Chien search that finds the roots of the error-locator polynomial by exhaustively searching for all finite field elements is one of the crucial steps in RS / BCH decoding. A highly parallel Chien search is needed to achieve high decoding throughput required by many applications, and typically accounts for a significant part of the overall decoder complexity.
[0004] In at least one conventional parallel Chien search architecture, each polynomial coefficient is associated with a column of constant multipliers that share a common input. Substructure sharing can be applied to the multipliers in the same column to reduce the gate count. Certain other Chien search architectures put off the modular reduction for finite field multiplication to simplify the multiplication; however, in these implementations, gate count is larger than that of applying substructure sharing to the multipliers across different columns and rows. Besides, the power consumption of the Chien search can be reduced through factorizing out the roots of the polynomials. Yet other Chien search architectures employ a two-stepAtty. Dkt. No.103361-613WO1 process, where a second step is activated only when, for evaluating a finite field element, the element seems to be a possible root from the first step. However, these designs still do not reduce the gate count of the Chien search architecture.
[0005] The Chien search can be expressed as a Vandermonde matrix multiplication. A Reed- Muller transformation can be used to decompose a Vandermonde matrix to two matrices that have fewer non-zero entries. However, this decomposition requires the entries in the Vandermonde matrix to be in binary order. As a result, it cannot be used as the Chien search architecture for RS / BCH decoding, which requires the finite field elements to be searched in the order of consecutive power of the primitive finite field element. Additionally, this leads to irregular non-zero entries in the decomposed matrices and hence does not reduce the complexity of parallel design. SUMMARY
[0006] One implementation of the present disclosure is a method of ^^^^-error-correcting Reed- Solomon (RS) and Bose–Chaudhuri–Hocquenghem (BCH) decoding. The method includes receiving an RS or BCH code, wherein the RS or BCH code is ^^^^-error-correcting of length nconstructed over finite field ^^^^^^^^(2^^^^)(^^^^ ∈ ^^^^+); determining a row vector ^^^^ representing an error-locator polynomial ^^^^(^^^^) for the RS or BCH code; calculating roots of the row vector ^^^^ using a parallel Chien search, wherein the parallel Chien search comprises multiplying the row vector ^^^^ by a Vandermonde matrix ^^^^, wherein the Vandermonde matrix ^^^^ is decomposed into two or more matrices; and decoding the RS or BCH code based on the calculated roots.
[0007] Another implementation of the present disclosure is a system for Reed-Solomon (RS) and Bose–Chaudhuri–Hocquenghem (BCH) decoding. The system includes at least one processor; and memory having instructions stored thereon that, when executed by the at least one processor, cause the system to receive an RS or BCH code, wherein the RS or BCH code is ^^^^- error-correcting and has a length n; determine a row vector ^^^^ representing an error-locator polynomial ^^^^(^^^^) for the RS or BCH code; calculate roots of the row vector ^^^^ using a parallel Chien search, wherein the parallel Chien search comprises multiplying the row vector ^^^^ by a Vandermonde matrix ^^^^, wherein the Vandermonde matrix ^^^^ is decomposed into two or more matrices; and decode the RS or BCH code based on the calculated roots.Atty. Dkt. No.103361-613WO1
[0008] Additional advantages will be set forth in part in the description which follows or may be learned by practice. The advantages will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive, as claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] FIG.1 is a diagram of a conventional P-parallel Chien search architecture for t-error correcting RS / BCH decoding, according to some implementations.
[0010] FIG.2 is a diagram of a proposed P-parallel Chien search architecture based on Vandermonde matrix decomposition, according to some implementations.
[0011] FIG.3 is a diagram illustrating the decomposition of a Vandermonde matrix, according to some implementations.
[0012] FIG.4 is a chart that compares the complexity of the proposed P-parallel Chien search architecture of FIG.2 with the conventional P-parallel Chien search architecture of FIG.1, according to some implementations.
[0013] FIG.5 is a block diagram of a system for RS / BCH decoding using the proposed P- parallel Chien search architecture, according to some implementations.
[0014] FIG.6 is a flow chart of a process for RS / BCH decoding using the proposed P-parallel Chien search architecture, according to some implementations.
[0015] Various objects, aspects, features, and advantages of the disclosure will become more apparent and better understood by referring to the detailed description taken in conjunction with the accompanying drawings, in which like reference characters identify corresponding elements throughout. In the drawings, like reference numbers generally indicate identical, functionally similar, and / or structurally similar elements.Atty. Dkt. No.103361-613WO1 DETAILED DESCRIPTION
[0016] Referring generally to the figures, a system and methods for RS / BCH decoding using a low-complexity Chien search architecture are shown, according to various implementations. In at least one aspect, a method to decompose a Vandermonde matrix for parallel Chien search in RS / BCH decoding is described herein. The Vandermonde matrix whose elements are in the order of consecutive power of the primitive element is decomposed into matrices with substantially smaller number of non-zero entries by utilizing the standard basis representation of finite field elements. In order to achieve efficient parallel design, additional reformulation is proposed for segments of the Vandermonde matrix so that the same constant multipliers can be utilized to implement the decomposed multiplications of each segment. Accordingly, a low- complexity parallel Chien search architecture is developed. For a 9-error-correcting RS or BCH code over ^^^^^^^^(210), the proposed design with 30, 40 and 50 parallel processing can achieve a 26%, 28%, and 31%, respectively, area reduction from architecture level analysis compared to conventional designs. Chien Search in RS / BCH Decoding
[0017] Consider a t-error-correcting RS or BCH code of length n constructed over finite field^^^^^^^^(2^^^^)(^^^^ ∈ ^^^^+). The decoder first computes 2^^^^ syndromes. Using these syndromes, the keyequation solver (KES) computes the error-locator polynomial, ^^^^(^^^^). If the errors are correctable, ^^^^(^^^^) is in the format of:where 0 ≤ ^^^^1, ^^^^2, … , ^^^^^^^^ are erroneous of ^^^^(^^^^) arecomputed by the Chien search, which is to evaluate ^^^^(^^^^) over n finite field elements,1,^^^^, … ,^^^^^^^^−1, where ^^^^ is a primitive element of ^^^^^^^^(2^^^^). The evaluation has to be carried out onfinite field elements in the order of increasing power of ^^^^ (e.g., 1,^^^^,^^^^2,^^^^3, …).
[0018] To achieve high throughput, a highly parallel Chien search architecture is desirable.For t-error-correcting decoding, ^^^^^^^^^^^^(^^^^(^^^^)) ≤ ^^^^ and ^^^^(^^^^) can be rewritten as Λ ^^^^^^^^^^^^ +Λ^^^^−1^^^^^^^^−1 + ⋯+ Λ0. FIG. 1 shows a conventional P-parallel Chien search architecture. In clockAtty. Dkt. No.103361-613WO1 cycle 0, the coefficients of ^^^^(^^^^) are sent through the multiplexers. Each row in the dotted box consists of t multipliers and computes:where 0 ≤ ^^^^ < ^^^^.in clock cycle 0.
[0019] In later clock cycles, the multiplexers on theof FIG.1 pass through the outputs of the multipliers in the feedback loops. The l-th feedback loop from the right side iterativelycomputes ^^^^^^^^^^^^^^^^^^^^^^^^ in clock cycle ^^^^(1 ≤ ^^^^ < ⌈^^^^ / ^^^^⌉). As a result, the j-th output of the architecturein FIG.1 in clock cycle i is: As aItconsists of ^^^^^^^^ constant multipliers, ^^^^^^^^ adders, ^^^^ + 1 multiplexers, and ^^^^ + 1 registers.
[0020] Each constant finite field multiplier can be described as a ^^^^ × ^^^^ constant binary matrixmultiplication. Substructure sharing can be applied among all the constant multipliers in FIG.1 to reduce the gate count. Vandermonde Matrix Decomposition for Low-Complexity Chien Search
[0021] In RS / BCH decoders, the highly parallel Chien search accounts for a significant portion of the overall decoder complexity. The Chien search can be described as a Vandermonde matrix multiplication. Disclosed herein is a technique to decompose the Vandermonde matrix into two matrices with substantially smaller number of non-zero entries. Accordingly, the number of multiplications in the Chien search can be substantially reduced, thereby reducing the complexity of the Chein architecture as discussed in greater detail below.
[0022] Generally, the Vandermonde matrix ^^^^ describing the Chien search is constructed using finite field elements in the order of consecutive powers of the primitive element. The proposed scheme decomposes such a ^^^^ matrix into two or more matrices so that the number of multiplications needed for the matrix multiplications is reduced. In an example case where ^^^^ isdecomposed into ^^^^ · ^^^^ – which is demonstrated in FIG. 3 – the entries of ^^^^ in row ^^^^ = 2^^^^ andAtty. Dkt. No.103361-613WO1column ^^^^ ≤ ^^^^ < ^^^^ are zeros and ^^^^ is a binary matrix; however, it will be appreciated that otherdecompositions are also possible.
[0023] Let ^^^^ = [^^^^0,^^^^1, … ,^^^^^^^^] be the row vector representing the error-locator polynomialand ^^^^ = [^^^^(1),^^^^(^^^^), … ,^^^^(^^^^^^^^−1]. The Chien search can be described as:where ^^^^ is a Vandermonde matrix in the
[0024] The Vandermondewhere the dimension of ^^^^ is (^^^^ + 1) × ^^^^ and that of ^^^^ is ^^^^ × ^^^^. This technique, in particular,decomposes Vandermonde matrices whose entries are in consecutive power of ^^^^ for the Chien search in RS / BCH decoding.
[0025] A standard basis of ^^^^^^^^(2^^^^) is {1,^^^^, … ,^^^^^^^^−1}. Each ^^^^^^^^ in the second row of ^^^^ can berepresented in standard basis as ^^^^^^^^ = ^^^^(^^^^) (^^^^)(^^^^) ^^^^−1 (^^^^) (^^^^) (^^^^)0+ ^^^^1 ^^^^ + ⋯+ ^^^^^^^^−1^^^^ , where ^^^^0, ^^^^1,…, ^^^^^^^^−1are binary bits. For each row with indices ^^^^ = 2^^^^(^^^^ ≥bewritten as:
[0026] ,…, are rowthe first ^^^^ entries of ^^^^ in row ^^^^ =are set to those of ^^^^. The coefficients in (5) are the same forevery row ^^^^ in the format of 2^^^^. Hence the first ^^^^ entries in j-th column of ^^^^ are set to [^^^^(^^^^),^^^^(^^^^) (^^^^) ^^^^ ^^^^1 , … ,^^^^ ] , where T denotes transpose. Then the entries in row ^^^^ = 2 andAtty. Dkt. No.103361-613WO1^^^^ ≤ ^^^^ < ^^^^ of ^^^^ are set to zero. Accordingly, the product of row ^^^^ = 2^^^^ of ^^^^ and column ^^^^ of ^^^^equals ^^^^^^^^^^^^in ^^^^.
[0027] To derive ^^^^^^^^′^^^^ in row ^^^^′ ≠ 2^^^^of ^^^^, the first q entries in row ^^^^′ of E are set to 1, ^^^^^^^^′,…,^^^^^^^^′(^^^^−1) since the ^^^^ × ^^^^ entries in the upper-left corner of B is an identity matrix. In this case,the entry in row ^^^^′ and column ^^^^ ≤ ^^^^ < ^^^^ of E can be calculated as:to make thethe entries inthe diagonal of B are set to ‘1’ and those in rows ^^^^ ≤ ^^^^ < ^^^^ are set to ‘0’.
[0028] Let ^^^^^^^^,^^^^ and ^^^^^^^^,^^^^ denote the entries of E and B, respectively, in row 0 ≤ ^^^^ ≤ ^^^^ andcolumn 0 ≤ ^^^^ < ^^^^. In summary, the entries of E are as follows:
[0029] B is an
[0030] Take ^^^^ = 7 and ^^^^ = 3 as an example. Consider finite field ^^^^^^^^(23) constructed usingprimitive polynomial ^^^^(^^^^) = ^^^^3 + ^^^^ + 1. Use the standard basis {1,^^^^,^^^^2}, where ^^^^ is the rootof ^^^^(^^^^). According to (6) and (7), E and B can be formed as follows:Atty. Dkt. No.103361-613WO1
[0031] Since B is afield multiplications needed in our proposed design depends on the number of entries in E that are not ‘0’ or ‘1’. For the example above, E has nine entries that are not ‘0’ or ‘1’. Accordingly, the proposed designreduces the number of multiplications in the Chien search from ^^^^(^^^^ − 1) = 18 as needed in thedirect Vandermonde matrix multiplication to nine.
[0032] From (6), the number of ‘1’ in the first q columns of E is ^^^^ + ^^^^. In columns ^^^^ ≤ ^^^^ <^^^^, the entries in the first row are either ‘0’ or ‘1’. Besides, there are ⌈^^^^^^^^^^^^2(^^^^)⌉ rows in E whose indices are powers of two. For these rows, the entries in columns q≤j<n are ‘0’. Therefore, the total number of multiplications needed in our design is at most: Vandermonde
[0033] For longer RS / BCH codes, multiplying the entire E matrix all at once to achieve fully parallel Chien search is impractical due to very high hardware complexity and routing congestion. On the other hand, dividing the E matrix into blocks of columns and multiplying one block of columns in each clock cycle would require general finite field multipliers since the ^^^^^^^^,^^^^entries are different. A general multiplier over ^^^^^^^^(2^^^^) has much higher complexity than a constant multiplier. Below, the Vandermonde matrix in (3) is reformulated such that theAtty. Dkt. No.103361-613WO1 products for each block of columns can be derived by multiplying the same matrix, which only needs constant multipliers, with an iterative update.
[0034] Assuming n is divisible by P, the Vandermonde matrix in (3) can be divided into ^^^^ / ^^^^ sub-matrices as follows: ^^^^0is a Vandermonde matrix inHence theproposed decomposition can also be applied. From (9), ^^^^^^^^ for ^^^^ > 0 can be re-written as:
[0035] Divide the^^^^(^^^^) =[Λ�^^^^^^^^^^^^�,Λ�^^^^^^^^^^^^+1�, … ,Λ�^^^^(^^^^+1)^^^^−1�]. Then ^^^^(^^^^) = Λ ∙ ^^^^^^^^
[0036] ^^^^0 can beproposedmethod. Then (11) becomes: As a result, each ^^^^(^^^^)If n is notdivisible by P, the extra ^^^^ − ^^^^%^^^^ outputs from the last group are ignored.
[0037] Consider the same example used in the previous section with ^^^^ = 7 and ^^^^ = 3. For^^^^ = 4, ^^^^ is divided into two groups as ^^^^ = [^^^^(0), ^^^^(1)] . Using the proposed decomposition, ^^^^0 =^^^^0 ∙ ^^^^0, where:Atty. Dkt. No.103361-613WO1 Thenand ^^^^(1) is derived by [Λ0,Λ^^^^ − ^^^^%^^^^ = 1 output ignored.
[0038] FIG.2 shows the proposed P-parallel Chien search architecture based onVandermonde matrix decomposition. It calculates ^^^^(^^^^) for ^^^^ = 0, 1, … ,^^^^ / ^^^^ − 1 in clock cycles0, 1, … ,^^^^ / ^^^^ − 1, respectively. The entries in thevector, [Λ ^^^^^^^^ ^^^^^^^^^^^^0,Λ1^^^^ , … ,Λ^^^^^^^^ ], aregenerated by the feedback loops located at the bottom in FIG.2 in clock cycle i. This vector is multiplied by the ^^^^0and then ^^^^0matrices. Since the entries of ^^^^0are constant, its multiplication can be implemented by constant finite field multipliers. The multiplication by ^^^^0is implemented by finite field adders since ^^^^0is a binary matrix. Complexity Analysis
[0039] In the proposed design shown in FIG.2, the maximum number of constant multipliersrequired for multiplying ^^^^0 is ^^^^^^^^ − (^^^^ + ⌈^^^^^^^^^^^^2(^^^^)⌉(^^^^ − ^^^^)) from (8). Considering the constantmultipliers in the bottom feedback loops, the proposed P-parallel Chien search architecture has t^^^^ + ^^^^^^^^ −�^^^^ + ⌈^^^^^^^^^^^^2(^^^^)⌉(^^^^ − ^^^^)� = ^^^^^^^^ − ⌈^^^^^^^^^^^^2(^^^^)⌉(^^^^ − ^^^^) constant multipliers. A constant^^^^^^^^(2^^^^) multiplier can be implementedmatrix multiplication. Substructuresharing can be applied to the multipliers in the proposed Chien search architecture to further reduce the gate count. Besides, the number of required finite field adders is decided by thenumber of non-zero entries in ^^^^0 and ^^^^0. As shown in FIG. 2, the proposed design also has ^^^^ +1 multiplexers and ^^^^ + 1 registers.Atty. Dkt. No.103361-613WO1
[0040] The graph shown in FIG.4 lists the hardware complexity of the proposed design for a9-error-correcting RS / BCH code with ^^^^ = 1023 over ^^^^^^^^(210). For ^^^^^^^^(210) constructed usingthe irreducible polynomial ^^^^(^^^^) = ^^^^10 + ^^^^3 + 1, a constant multiplier requires around 40 XORgates to implement on average based on simulations. Each adder over ^^^^^^^^(210) is implemented using q XOR gates and has similar complexity as a q-bit 2-to-1 multiplexer. A q-bit register hasaround the same area as 3^^^^ = 30 XOR gates. Based on these assumptions, the complexity of theoverall Chien search is estimated in terms of the number of XOR gates needed as listed in FIG. 4. The proposed design can finish the Chien search in ⌈^^^^ / ^^^^⌉ clock cycles. The architecture thatimplements the ^^^^0 multiplication has one multiplier and ⌈^^^^^^^^^^^^2(^^^^ + 1)⌉ adders in the data path. A^^^^^^^^(2^^^^) constant multiplier has at most ⌈^^^^^^^^^^^^2^^^^⌉ gates in the data path. In addition, only the first q rows of ^^^^0may have multiple non-zero entries and each of the other rows has at most a singlenon-zero entry. Hence, the data path of its multiplication has at most ⌈^^^^^^^^^^^^2(^^^^ + 1)⌉ adders.Including the multiplexers located at the bottom in FIG.2, the critical path of the proposed P- parallel Chien search architecture for a t-error-correcting RS / BCH code over ^^^^^^^^(2^^^^) has at most⌈^^^^^^^^^^^^2(^^^^ + 1)⌉ + ⌈^^^^^^^^^^^^2^^^^⌉ + ⌈^^^^^^^^^^^^2(^^^^ + 1)⌉ + 1 gates. The actual critical paths for the designswith different P shown in FIG.4 are derived from the actual matrices involved.
[0041] For comparison, the parallel Chien search architecture in FIG.1 is considered. Itconsists of ^^^^^^^^ constant multipliers, ^^^^^^^^ adders, ^^^^ + 1 multiplexers, and ^^^^ + 1 registers. In theproposed design, the entries of ^^^^ ^^^^0in row 2 (^^^^ ≥ 0) and columns ^^^^ ≤ ^^^^ < ^^^^ are zero. Hence thenumber of multipliers required by our parallel Chien search architecture is significantly reduced.Also, the number of ‘1’s in column ^^^^ ≤ ^^^^ < ^^^^ of ^^^^0is dependent on the standard basisrepresentation of ^^^^^^^^. For small j, it typically has a small number of non-zero bits. As a result, the number of adders in our design is also reduced as shown in FIG.4. Overall, this design can reduce the complexity by 26% to 31% compared to the conventional parallel Chien search architecture for example code with parallelism ranging between 30 and 50. The critical path ofthe conventional Chien search architecture is 1 + ⌈^^^^^^^^^^^^2^^^^⌉ + ⌈^^^^^^^^^^^^2(^^^^ + 1)⌉. To achieve the samecritical path, pipelining can be applied between the ^^^^0and ^^^^0multiplication blocks. In this case, the design has one more clock cycle in the latency.
[0042] For larger P, the area reduction achieved by the proposed design is more significant asshown in FIG. 4, since all entries in column ^^^^ ≥ ^^^^ of the rows with indices 2^^^^ in ^^^^0 are zero.Atty. Dkt. No.103361-613WO1 There are a bigger percentage of rows with indices 2^^^^when t is smaller. Hence, the area saving achieved by the proposed design is also more significant for codes with smaller t. System and Methods
[0043] Referring now to FIG.5, a block diagram of a system 500 for RS / BCH decoding using the proposed P-parallel Chien search architecture (e.g., the Chien architecture in FIG.2) is shown, according to some implementations. System 500 is shown to include a processing circuit 502 that includes a processor 504 and a memory 506. Processor 504 can be a general-purpose processor, an application-specific integrated circuit (ASIC), one or more field programmable gate arrays (FPGAs), a group of processing components (e.g., a central processing unit (CPU)), or other suitable electronic processing structures. In some implementations, processor 504 is configured to execute program code stored on memory 506 to cause system 500 to perform one or more operations, as described below in greater detail. It will be appreciated that, in implementations where system 500 is part of another computing device, the components of system 500 may be shared with, or the same as, the host device.
[0044] Memory 506 can include one or more devices (e.g., memory units, memory devices, storage devices, etc.) for storing data and / or computer code for completing and / or facilitating the various processes described in the present disclosure. In some implementations, memory 506 includes tangible (e.g., non-transitory), computer-readable media that stores code or instructions executable by processor 504. Tangible, computer-readable media refers to any physical media that is capable of providing data that causes system 500 to operate in a particular fashion. Example tangible, computer-readable media may include, but is not limited to, volatile media, non-volatile media, removable media and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Accordingly, memory 506 can include random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), hard drive storage, temporary storage, non-volatile memory, flash memory, optical memory, or any other suitable memory for storing software objects and / or computer instructions. Memory 506 can include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described inAtty. Dkt. No.103361-613WO1 the present disclosure. Memory 506 can be communicably connected to processor 504, such as via processing circuit 502, and can include computer code for executing (e.g., by processor 504) one or more processes described herein.
[0045] While shown as individual components, it will be appreciated that processor 504 and / or memory 506 can be implemented using a variety of different types and quantities of processors and memory. For example, processor 504 may represent a single processing device or multiple processing devices. Similarly, memory 506 may represent a single memory device or multiple memory devices. Additionally, in some implementations, system 500 may be implemented within a single computing device (e.g., one server, one housing, etc.). In other implementations, system 500 may be distributed across multiple servers or computers (e.g., that can exist in distributed locations). For example, system 500 may include multiple distributed computing devices (e.g., multiple processors and / or memory devices) in communication with each other that collaborate to perform operations. For example, but not by way of limitation, an application may be partitioned in such a way as to permit concurrent and / or parallel processing of the instructions of the application. Alternatively, the data processed by the application may be partitioned in such a way as to permit concurrent and / or parallel processing of different portions of a data set by the two or more computers.
[0046] Memory 506 is shown to include an RS / BCH decoder 508 configured to decode received RS / BCH codes. In particular, RS / BCH decoder 508 can be configured to decode RS / BCH codes received via a communications channel, from a storage device, etc., where the received code was previously encoded by an RS / BCH encoder. RS / BCH decoder 508 is shown to include a Chien search component 510 which, as discussed above, determines roots of an error-locator polynomial, ^^^^(^^^^) that is computed (e.g., by a KES of RS / BCH decoder 508) for the RS / BCH code. In this regard, Chien search component 510 generally includes or is configured to implement the Chien search architecture described above with respect to FIG.2.
[0047] System 500 is also shown to include a communications interface 512 that facilitates communications between system 500 and any external components or devices. In other words, communications interface 512 can provide means for transmitting data to, or receiving data from, any remote computing device. Accordingly, communications interface 512 can be or can include wired and / or wireless communications interfaces (e.g., jacks, antennas, transmitters, receivers,Atty. Dkt. No.103361-613WO1 transceivers, wire terminals, etc.) for conducting data communications. In some implementations, communications via communications interface 512 are direct (e.g., local wired or wireless communications) or via a network (e.g., a WAN, the Internet, a cellular network, etc.). For example, communications interface 512 may include one or more Ethernet ports for communicably coupling system 500 to a network (e.g., the Internet). In another example, communications interface 512 can include a Wi-Fi transceiver for communicating via a wireless communications network. In yet another example, communications interface 512 may include cellular or mobile phone communications transceivers.
[0048] Referring now to FIG.6, a flow chart of a process 600 for RS / BCH decoding using the proposed P-parallel Chien search architecture is shown, according to some implementations. Process 600 may be implemented by system 500, as described above; however, the present disclosure is not limiting in this regard. It will be appreciated that certain steps of process 600 may be optional and, in some implementations, process 600 may be implemented using less than all of the steps. It will also be appreciated that the order of steps shown in FIG.6 is not intended to be limiting.
[0049] At step 602, an RS or BCH code is received, e.g., from a remote computing device. As discussed above, the RS or BCH code may be an ^^^^-error-correcting of length n constructedover finite field ^^^^^^^^(2^^^^)(^^^^ ∈ ^^^^+). At step 604, a row vector ^^^^ representing an error-locatorpolynomial ^^^^(^^^^) for the RS or BCH code is determined. In some implementations, this determinization can include calculating 2^^^^ syndromes and then calculating the error-locator polynomial ^^^^(^^^^) using a key equation solver (KES). At step 606, roots of the row vector ^^^^ are calculated using a parallel Chien search, e.g., as discussed above with respect to FIG.2. Accordingly, the parallel Chien search generally includes multiplying the row vector ^^^^ by a Vandermonde matrix ^^^^.
[0050] As mentioned above, the Vandermonde matrix ^^^^ describing the Chien search is constructed using finite field elements in the order of consecutive powers of the primitive element. The proposed scheme decomposes such a ^^^^ matrix into two or more matrices so that the number of multiplications needed for the matrix multiplications is reduced. In an examplecase where ^^^^ is decomposed into ^^^^ · ^^^^, the entries of ^^^^ in row ^^^^ = 2^^^^ and column ^^^^ ≤ ^^^^ < ^^^^ arezeros and ^^^^ is a binary matrix; however, it will be appreciated that other decompositions are alsoAtty. Dkt. No.103361-613WO1 possible. At step 608, the RS or BCH code is decoded based on the calculated roots (e.g., using the results of the parallel Chien search). Discussion
[0051] Described herein are a system and technique scheme that decompose a Vandermonde matrix into two matrices with smaller number of non-zero entries such that the number of multiplications in the Chien search can be substantially reduced. Also, a further reformulation on the matrix decomposition is proposed to enable efficient parallel Chien search. Then a low- complexity parallel Chien search architecture is designed. Compared to the conventional parallel Chien search architecture, the proposed design achieves a substantial reduction in hardware complexity, especially for higher parallelisms. Configuration of Certain Implementations
[0052] The construction and arrangement of the systems and methods as shown in the various implementations are illustrative only. Although only a few implementations have been described in detail in this disclosure, many modifications are possible (e.g., variations in sizes, dimensions, structures, shapes, and proportions of the various elements, values of parameters, mounting arrangements, use of materials, colors, orientations, etc.). For example, the position of elements may be reversed or otherwise varied, and the nature or number of discrete elements or positions may be altered or varied. Accordingly, all such modifications are intended to be included within the scope of the present disclosure. The order or sequence of any process or method steps may be varied or re-sequenced according to alternative implementations. Other substitutions, modifications, changes, and omissions may be made in the design, operating conditions, and arrangement of the implementations without departing from the scope of the present disclosure.
[0053] The present disclosure contemplates methods, systems, and program products on any machine-readable media for accomplishing various operations. The implementations of the present disclosure may be implemented using existing computer processors, or by a special purpose computer processor for an appropriate system, incorporated for this or another purpose, or by a hardwired system. Implementations within the scope of the present disclosure include program products including machine-readable media for carrying or having machine-executable instructions or data structures stored thereon. Such machine-readable media can be any available media that can be accessed by a general purpose or special purpose computer or other machineAtty. Dkt. No.103361-613WO1 with a processor. By way of example, such machine-readable media can comprise RAM, ROM, EPROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to carry or store desired program code in the form of machine-executable instructions or data structures, and which can be accessed by a general purpose or special purpose computer or other machine with a processor.
[0054] When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a machine, the machine properly views the connection as a machine-readable medium. Thus, any such connection is properly termed a machine-readable medium. Combinations of the above are also included within the scope of machine-readable media. Machine-executable instructions include, for example, instructions and data which cause a general-purpose computer, special purpose computer, or special purpose processing machines to perform a certain function or group of functions.
[0055] Although the figures show a specific order of method steps, the order of the steps may differ from what is depicted. Also, two or more steps may be performed concurrently or with partial concurrence. Such variation will depend on the software and hardware systems chosen and on designer choice. All such variations are within the scope of the disclosure. Likewise, software implementations could be accomplished with standard programming techniques with rule-based logic and other logic to accomplish the various connection steps, processing steps, comparison steps and decision steps.
[0056] It is to be understood that the methods and systems are not limited to specific synthetic methods, specific components, or to particular compositions. It is also to be understood that the terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting.
[0057] As used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another implementation includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms anotherAtty. Dkt. No.103361-613WO1 implementation. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint.
[0058] “Optional” or “optionally” means that the subsequently described event or circumstance may or may not occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.
[0059] Throughout the description and claims of this specification, the word “comprise” and variations of the word, such as “comprising” and “comprises,” means “including but not limited to,” and is not intended to exclude, for example, other additives, components, integers or steps. “Exemplary” means “an example of” and is not intended to convey an indication of a preferred or ideal implementation. “Such as” is not used in a restrictive sense, but for explanatory purposes.
[0060] Disclosed are components that can be used to perform the disclosed methods and systems. These and other components are disclosed herein, and it is understood that when combinations, subsets, interactions, groups, etc. of these components are disclosed that while specific reference of each various individual and collective combinations and permutation of these may not be explicitly disclosed, each is specifically contemplated and described herein, for all methods and systems. This applies to all aspects of this application including, but not limited to, steps in disclosed methods. Thus, if there are a variety of additional steps that can be performed it is understood that each of these additional steps can be performed with any specific implementation or combination of implementations of the disclosed methods.
Claims
Atty. Dkt. No.103361-613WO1 WHAT IS CLAIMED IS:
1. A method of ^^^^-error-correcting Reed-Solomon (RS) and Bose–Chaudhuri– Hocquenghem (BCH) decoding, the method comprising: receiving an RS or BCH code, wherein the RS or BCH code is ^^^^-error-correcting of length nconstructed over finite field ^^^^^^^^(2^^^^)(^^^^ ∈ ^^^^+);determining a row vector ^^^^ representing an error-locator polynomial ^^^^(^^^^) for the RS or BCH code; calculating roots of the row vector ^^^^ using a parallel Chien search, wherein the parallel Chien search comprises multiplying the row vector ^^^^ by a Vandermonde matrix ^^^^, wherein the Vandermonde matrix ^^^^ is decomposed into two or more matrices; and decoding the RS or BCH code based on the calculated roots.
2. The method of claim 1, wherein the parallel Chien search is defined as: ^^^^(^^^^)=[^^^^0,^^^^1, … ,^^^^^^^^] · ^^^^.
3. The method of claim 2, wherein ^^^^(^^^^) is calculated for ^^^^ = 0, 1, … ,^^^^ / ^^^^ − 1 in clockcycles 0, 1, … ,^^^^ / ^^^^ − 1, respectively.
4. The method of claim 1, wherein a number of constant multipliers of the parallel Chien search is less than ^^^^^^^^.
5. The method of claim 4, wherein the constant multipliers are implemented as a ^^^^ × ^^^^binary matrix multiplication.
6. The method of claim 1, wherein the parallel Chien search is performed in ⌈^^^^ / ^^^^⌉ clock cycles.
7. The method of claim 1, wherein the determining includes calculating 2^^^^ syndromes and then calculating the error-locator polynomial ^^^^(^^^^) using a key equation solver (KES).
8. The method of claim 1, wherein the Vandermonde matrix ^^^^ is decomposed into two ormore matrices ^^^^ · ^^^^, wherein the entries of ^^^^ in row ^^^^ = 2^^^^ and column ^^^^ ≤ ^^^^ < ^^^^ are zeros and^^^^ is a binary matrix.Atty. Dkt. No.103361-613WO1 9. The method of claim 8, wherein a total number of constant finite field multiplications depends on a number of entries in E that are not ‘0’ or ‘1’.
10. The method of claim 8, wherein entries of ^^^^ ^^^^0in row 2 (^^^^ ≥ 0) and columns ^^^^ ≤ ^^^^ < ^^^^are zero to reduce a number of multipliers.
11. A system for Reed-Solomon (RS) and Bose–Chaudhuri–Hocquenghem (BCH) decoding, the system comprising: at least one processor; and memory having instructions stored thereon that, when executed by the at least one processor, cause the system to: receive an RS or BCH code, wherein the RS or BCH code is ^^^^-error-correcting and has a length n; determine a row vector ^^^^ representing an error-locator polynomial ^^^^(^^^^) for the RS or BCH code; calculate roots of the row vector ^^^^ using a parallel Chien search, wherein the parallel Chien search comprises multiplying the row vector ^^^^ by a Vandermonde matrix ^^^^, wherein the Vandermonde matrix ^^^^ is decomposed into two or more matrices; and decode the RS or BCH code based on the calculated roots.
12. The system of claim 11, wherein the parallel Chien search is defined as: ^^^^(^^^^)=[^^^^0,^^^^1, … ,^^^^^^^^] · ^^^^.
13. The system of claim 12, wherein ^^^^(^^^^) is calculated for ^^^^ = 0, 1, … ,^^^^ / ^^^^ − 1 in clockcycles 0, 1, … ,^^^^ / ^^^^ − 1, respectively.
14. The system of claim 11, wherein a number of constant multipliers of the parallel Chien search is less than ^^^^^^^^.
15. The system of claim 14, wherein the constant multipliers are implemented as a ^^^^ × ^^^^binary matrix multiplication.Atty. Dkt. No.103361-613WO1 16. The system of claim 11, wherein the parallel Chien search is performed in ⌈^^^^ / ^^^^⌉ clock cycles.
17. The system of claim 11, wherein the determining includes calculating 2^^^^ syndromes and then calculating the error-locator polynomial ^^^^(^^^^) using a key equation solver (KES).
18. The system of claim 11, wherein the Vandermonde matrix ^^^^ is decomposed into twoor more matrices ^^^^ · ^^^^, wherein the entries of ^^^^ in row ^^^^ = 2^^^^ and column ^^^^ ≤ ^^^^ < ^^^^ are zerosand ^^^^ is a binary matrix.
19. The system of claim 18, wherein a total number of constant finite field multiplications depends on a number of entries in E that are not ‘0’ or ‘1’.
20. The system of claim 18, wherein entries of ^^^^ ^^^^0in row 2 (^^^^ ≥ 0) and columns ^^^^ ≤ ^^^^ <^^^^ are zero to reduce a number of multipliers.
Citation Information
Patent Citations
Apparatus and method for decoding digital data
EP1102406A2
Decomposable forward error correction
US20190190653A1
BCH fast soft decoding beyond the (d-1) / 2 bound
US20230223958A1
Accelerated polynomial coding system and method
US20230378979A1
Real-time BCH error correction code decoding mechanism
US4866716A