A matrix and vector multiplication unit
By simplifying the LDPC encoder structure and optimizing the matrix and vector multiplication calculation process, the problems of complex structure and insufficient flexibility in the prior art are solved, and efficient LDPC encoder operation and fast matrix and vector multiplication operations are realized.
Patent Information
- Application Number
- CN202111023530.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2015-12-28
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2035-12-28
AI Technical Summary
The existing LDPC encoding technology has complex structure, low utilization rate of functional modules, low encoding throughput, and is not flexible enough to be suitable for quasi-cyclic verification matrix with different structures.
Simplify the structure of the LDPC encoder, change complex operations such as matrix inversion into offline software work, and optimize the matrix and vector multiplication calculation process by reasonably allocating hardware and programmable microcode instructions, reuse intermediate results, and reduce the number of executed instructions.
It improves the operating efficiency and throughput of the LDPC encoder, can be applied to quasi-cyclic check matrices with different structures and code rates, and improves the execution speed of matrix and vector multiplication.
Smart Images

Figure CN113708779B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a low density parity check code (LDPC) encoder for matrix and vector multiplication operations and a method thereof, and in particular to a matrix and vector multiplication unit for optimizing the LDPC encoding process. Background Art
[0002] LDPC codes are widely used for channel coding in high-speed wireless communication systems and are expected to be used in high-performance solid-state storage systems. In "Efficient encoding of low-density parity-check codes" ([J]. IEEE TransInformation Theory, 2001, 47(2):638–656), an RU encoding algorithm is proposed. This algorithm transforms the check matrix into a quasi-lower triangular matrix and uses the Gauss method to solve the equations for encoding. Furthermore, in "Online Programmable High-Speed Encoder Structure for QC-LDPC Codes" (Journal of Tsinghua University (Science and Technology), Vol. 49, No. 7, pp. 1025-1018, 2009), a quasi-cyclic low-density parity-check code encoder structure that supports variable parameters is proposed. Summary of the Invention
[0003] Existing LDPC coding technology is structurally complex, with numerous functional modules and control units. The software and hardware functional modules are not clearly defined, resulting in low utilization of each module and low coding throughput. Furthermore, it lacks flexibility and cannot be applied to quasi-cyclic parity check matrices with different structures.
[0004] The present invention simplifies the structure of the LDPC encoder, reduces the cumbersome control flow to instruction control, and replaces complex operations such as matrix inversion with offline software operations, thereby improving the operating efficiency of each functional component and the throughput. The invention can be applied to the encoding of quasi-cyclic check matrices with different structures and code rates.
[0005] One object of the present invention is to efficiently implement a matrix and vector multiplication circuit for LDPC coding and to reasonably distribute the computational process between hardware and programmable microcode instructions.
[0006] Another object of the present invention is to optimize the calculation process of matrix and vector multiplication by reusing the intermediate results in the matrix and vector multiplication process, reducing the number of instructions executed in the process of matrix and vector multiplication, thereby speeding up the execution speed of vector and matrix multiplication.
[0007] According to a first aspect of the present invention, a matrix and vector multiplication method of a first embodiment of the first aspect of the present invention is provided, wherein the matrix M is a sum matrix of cyclically shifted unit matrices, comprising: a first step: initializing a global register; a second step: shifting a vector S by a specified number of bits and performing an XOR operation with the content of the global register, and storing the XOR result in the global register; and a third step: storing the value in the global register.
[0008] According to the first implementation of the first aspect of the present invention, a second implementation of the first aspect of the present invention is provided, wherein when the matrix M is a sum matrix of multiple cyclic shift unit matrices, the second step is repeatedly performed.
[0009] According to the first or second embodiment of the first aspect of the present invention, a third embodiment of the first aspect of the present invention is provided, further comprising: a fourth step: obtaining a vector S from a data memory.
[0010] According to the third implementation of the first aspect of the present invention, a fourth implementation of the first aspect of the present invention is provided, further comprising: a fifth step: loading the vector S in the fourth step into a vector register.
[0011] According to the first, third or fourth implementation of the first aspect of the present invention, a fifth implementation of the first aspect of the present invention is provided, wherein the matrix M=I1+I2+...I m +…+I n , where I m is the cyclic shift identity matrix, and is cyclically shifted from the identity matrix I by d m The cyclic shift identity matrix I is obtained m , where 1≤m≤n; in the second step, for the n cyclic shift unit matrices I1, I2, ... I that constitute the matrix M m ,…I n Each cyclic shift of the identity matrix I m , perform the following operation: shift vector S by d m bits, XOR the shift result with the value of the global register, and store the XOR result in the global register.
[0012] According to the third embodiment of the first aspect of the present invention, there is provided a sixth embodiment of the first aspect of the present invention, wherein the matrix M=I1+I2+...I m +…+I n , where I1, I m is the cyclic shift identity matrix, and I1 is obtained by shifting d1 positions from the identity matrix I, and I1*S is cyclically shifted by d m 'Get the cyclic shift identity matrix I m, where 2≤m≤n; the second step includes: shifting the vector S by d1 bits, XORing the shift result with the value of the global register to obtain a specific value, and storing it in the global register; storing the specific value of the global register in the data memory; for the n-1 cyclic shift unit matrices I2, I3, ... I constituting the matrix M m ,…,I n Each cyclic shift of the identity matrix I m , where 2≤m≤n, perform the following operations: get a specific value from the data memory, shift the specific value by d m ' bit, XOR the shift result with the value of the global register, and store the XOR result in the global register.
[0013] According to the fourth implementation of the first aspect of the present invention, there is provided a seventh implementation of the first aspect of the present invention, wherein the matrix M=I1+I2+...I m +…+I n , where I1, I m is the cyclic shift identity matrix, and I1 is obtained by shifting d1 positions from the identity matrix I, and I1*S is cyclically shifted by d m 'Get the cyclic shift identity matrix I m , wherein 2≤m≤n; the second step (S20) comprises: shifting the vector S by d1 bits, performing an XOR operation on the shift result and the value of the global register and storing the result in the global register; storing the value of the global register in the vector register; for the n-1 cyclic shift unit matrices I2, I3, ... I constituting the matrix M m ,…I n Each cyclic shift of the identity matrix I m , where 2≤m≤n, perform the following operation: shift the value of the vector register by d m ' bit, XOR the shift result with the value of the global register, and store the XOR result in the global register.
[0014] According to the fourth embodiment of the first aspect of the present invention, there is provided an eighth embodiment of the first aspect of the present invention, wherein the n cyclic shift identity matrices {I1, I2, ... I n} is sorted so that Minimum.
[0015] According to the fourth implementation of the first aspect of the present invention, there is provided a ninth implementation of the first aspect of the present invention, wherein the matrix M=K1+K2+...K m +…+K n , where K1, K m is the sum of two cyclically shifted identity matrices, cyclically shifted from K1 by d m Get the matrix K m, circularly shifting the identity matrix I by dI1 bits to obtain a circularly shifted identity matrix I1, and circularly shifting the identity matrix I by dI2 to obtain a circularly shifted identity matrix I2, dI1 and dI2 are consecutive natural numbers, and K1=I1+I2, where 2≤m≤n; the second step includes: shifting the vector S by dI1 bits, XORing the shift result with the value of the global register and storing them in the global register; shifting the vector S by dI2 bits, XORing the shift result with the value of the global register and storing them in the global register; storing the value of the global register in the vector register; for the n-1 matrices K2 constituting the matrix M, ...K m ,…K n Each matrix K in m , where 2≤m≤n, perform the following operation: shift the value of the vector register by d m bits, XORs the shift result with the value of the global register, and stores the XOR result in the global register.
[0016] According to the third implementation of the first aspect of the present invention, there is provided a tenth implementation of the first aspect of the present invention, wherein the matrix M=K1+K2+...K m +…+K n , where K1, K m is the sum of two cyclically shifted identity matrices, cyclically shifted from K1 by d m Get the matrix K m , circularly shifting the identity matrix I by dI1 bits to obtain a circularly shifted identity matrix I1, and circularly shifting the identity matrix I by dI2 to obtain a circularly shifted identity matrix I2, dI1 and dI2 are consecutive natural numbers, and K1=I1+I2, where 2≤m≤n; the second step includes: shifting the vector S by dI1 bits, XORing the shift result with the value of the global register and storing the result in the global register; shifting the vector S by dI2 bits, XORing the shift result with the value of the global register to obtain a specific value, and storing the result in the global register; storing the specific value of the global register in the data memory; for the n-1 matrices K2 constituting the matrix M, ....K m ,…K n Each matrix K in m , where 2≤m≤n, perform the following operations: get a specific value from the data memory, shift the specific value by d m bits, XORs the shift result with the value of the global register, and stores the XOR result in the global register.
[0017] According to the third implementation of the first aspect of the present invention, there is provided an eleventh implementation of the first aspect of the present invention, wherein the matrix M=K1+K2+...K m +…+K n , where K mIt is the sum matrix of p cyclic shift identity matrices, where P is a positive integer, cyclically shifted from K1 by d m Get the matrix K m , where 2≤m≤n, the identity matrix I is cyclically shifted by dj positions to obtain the cyclically shifted identity matrix Ij, where 1≤j≤P, and K1=I1+I2+...Ij+…+I P , dI1, dI2…dIj,…dIp are continuous natural numbers; the second step includes: for each of the p cyclic shift unit matrices constituting the matrix K1, performing the following operations: shifting the value of the vector S by d1j bits, performing an XOR operation on the shift result and the value of the global register to obtain a specific value, and storing it in the global register; storing the specific value of the global register in the data memory; for the n-1 matrices K2,…K constituting the matrix M, m ,...K n Each matrix K in m , where 2≤m≤n, perform the following operations: get a specific value from the data memory, shift the specific value by d m bits, XORs the shift result with the value of the global register, and stores the XOR result in the global register.
[0018] According to the third embodiment of the first aspect of the present invention, there is provided an eleventh embodiment of the first aspect of the present invention, wherein the matrix M=K1+K2+..K m +....+K n , where K m It is the sum matrix of p cyclic shift identity matrices, where P is a positive integer, cyclically shifted from K1 by d m Get the matrix K m , where 2≤m≤n, the identity matrix I is cyclically shifted by dIj positions to obtain the cyclically shifted identity matrix Ij, where 1≤j≤P, and K1=I1+I2+...Ij+...+I P , dI1, dI2...dIj,...dIp are continuous natural numbers; the second step includes: for each of the p cyclic shift unit matrices constituting the matrix K1, performing the following operations: shifting the value of the vector S by d1j bits, performing an XOR operation on the shift result and the value of the global register and storing the result in the global register; storing the value of the global register in the vector memory; for the n-1 matrices K2,...K constituting the matrix M, m ,...K n Each matrix K in m , perform the following operation: shift the value of the vector register by d m bits, XORs the shift result with the value of the global register, and stores the XOR result in the global register.
[0019] According to the fourth implementation of the first aspect of the present invention, there is provided a twelfth implementation of the first aspect of the present invention, wherein the matrix M=K1+K2+..K m +....+K n , where K m It is the sum matrix of p cyclic shift identity matrices, where P is a positive integer, cyclically shifted from K1 by d m Get the matrix K m , where 2≤m≤n, cyclic shift d from the identity matrix I Ij The cyclic shift identity matrix Ij is obtained, where 1≤j≤P, and K1=I1+I2+...Ij+...+I P , d I1 , d I2 ...d Ij ,...d Ip Is a continuous natural number; the second step (S20) comprises: for each of the p cyclic shift unit matrices constituting the matrix K1, performing the following operation: shifting the value of the vector S by d 1j bits, XOR the shift result with the value of the global register and store it in the global register; store the value of the global register in the vector memory; for the n-1 matrices K2,...K that constitute the matrix M m ,...K n Each matrix K in m , perform the following operation: shift the value of the vector register by d m bits, XORs the shift result with the value of the global register, and stores the XOR result in the global register.
[0020] According to the ninth to twelfth embodiments of the first aspect of the present invention, there is provided a thirteenth embodiment of the first aspect of the present invention, wherein the sum matrix {K1, K2, ..., K n} is sorted so that Minimum.
[0021] According to a second aspect of the present invention, a method for calculating the multiplication of a matrix M and a vector S according to a first embodiment of the second aspect of the present invention is provided, wherein the matrix Wherein K(i) is the sum matrix of i cyclic shift identity matrices, and K(i) has i consecutive non-zero rows; and there are f(i) matrices K(i) with i consecutive non-zero rows, and K(i, j(i)) is the j(i)th matrix K(i) in the f(i) matrices; the method comprises: for each value of i, according to the method for calculating matrix and vector multiplication according to the eleventh or twelfth embodiment of the first aspect of the present invention, calculating The result is stored in the data storage;
[0022] Multiple data stored in the data memory The results of XOR are used to obtain the calculation result of M*S.
[0023]
[0024] According to a third aspect of the present invention, a matrix and vector multiplication operation unit is provided, comprising: a shift unit, an XOR unit and a global register, wherein the shift unit is used to shift a vector by a specified number of bits to obtain a shift result; the XOR unit is connected to the shift unit and the global register, and is used to receive the shift result from the shift unit and XOR the shift result with a stored value in the global register to obtain an XOR result; the global register is used to store the XOR result from the XOR unit.
[0025] According to an embodiment of the third aspect of the present invention, it further includes: an instruction memory for storing instructions, the instructions including a first instruction, wherein the first instruction instructs the shift unit to shift the vector by a specified number of bits to obtain a shift result and instructs the XOR unit to XOR the shift result with the stored value.
[0026] According to an embodiment of the third aspect of the present invention, the method further comprises: a data memory connected to the shift unit and the global register, and configured to store vectors.
[0027] According to an embodiment of the third aspect of the present invention, the device further comprises a vector register connected to the data memory and the shift unit, and configured to receive the vector from the data memory and provide the vector to the shift unit.
[0028] According to a fourth aspect of the present invention, a matrix and vector multiplication device of the fourth aspect of the present invention is provided, wherein the matrix M is a sum matrix of cyclically shifted unit matrices, and the device comprises: a module for initializing a global register; a module for shifting a vector S by a specified number of bits and performing an XOR operation with the contents of the global register, and storing the XOR result in the global register; and a module for storing the value in the global register.
[0029] According to a fifth aspect of the present invention, there is provided a computer program comprising computer program code which, when loaded into and executed on a computer system, causes the computer system to perform a method according to an embodiment of the first or second aspect of the present invention.
[0030] According to a sixth aspect of the present invention, there is provided a program comprising program code which, when loaded into and executed on a storage device, causes the storage device to perform a method according to an embodiment of the first or second aspect of the present invention.
[0031] The present invention optimizes the calculation process of matrix and vector multiplication, reduces the number of instructions executed in the process of matrix and vector multiplication by reusing intermediate results in the process of matrix and vector multiplication, and thus speeds up the execution speed of vector and matrix multiplication. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiments below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. Throughout the accompanying drawings, identical reference symbols are used to denote identical components. In the accompanying drawings, letter designations following reference numerals indicate multiple identical components. When referring to these components in general, the last letter designation will be omitted. In the accompanying drawings:
[0033] Figure 1A A schematic structural diagram of a matrix and vector multiplication unit for LDPC coding according to one embodiment of the present invention is shown;
[0034] Figure 1B A flowchart illustrating a matrix and vector multiplication method in an LDPC encoder according to an embodiment of the present invention is provided;
[0035] Figure 2 A schematic structural diagram of a matrix and vector multiplication unit for LDPC coding according to one embodiment of the present invention is shown;
[0036] Figure 3A FIG2 shows a schematic structural diagram of a matrix and vector multiplication unit for LDPC coding according to another embodiment of the present invention;
[0037] Figure 3B A flowchart of a matrix and vector multiplication method in an LDPC encoder according to one embodiment of the present invention is shown;
[0038] Figure 4A FIG2 shows a schematic structural diagram of a matrix and vector multiplication unit for LDPC coding according to another embodiment of the present invention;
[0039] Figure 4B A flowchart of a matrix and vector multiplication method in an LDPC encoder according to another embodiment of the present invention is shown;
[0040] Figure 5A flowchart of a matrix and vector multiplication method in an LDPC encoder according to another embodiment of the present invention is shown;
[0041] Figure 6 A flowchart of a matrix and vector multiplication method in an LDPC encoder according to another embodiment of the present invention is shown.
[0042] In the drawings, the same or similar reference numbers are used to refer to the same or similar elements. DETAILED DESCRIPTION
[0043] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0044] During the LDPC encoding process, the multiplication of matrix M and vector S is a key operation. Matrix M is the sum of n cyclically shifted identity matrices. The multiplication of matrix M and vector S can be decomposed into a shift operation on vector S and a modulo-2 sum of the shifted results. In the LDPC encoder according to the present invention, matrix multiplication by vector S is achieved by executing an instruction sequence. Furthermore, multiple matrix and vector multiplication operations are involved during the LDPC encoding process. Multiple instruction sequences corresponding to multiple matrix and vector multiplication operations are provided, as well as matrix M and vector S corresponding to multiple matrix and vector multiplications. Intermediate and / or final LDPC encoding calculation results are obtained by executing the multiple instruction sequences. Under the control of the multiple instruction sequences, the matrix and vector multiplication operation unit according to an embodiment of the present invention performs multiple matrix and vector multiplication operations and implements LDPC encoding. Therefore, the matrix and vector multiplication operation unit according to an embodiment of the present invention is also an LDPC encoder.
[0045] Figure 1A A schematic diagram of the structure of a matrix-vector multiplication unit for LDPC encoding according to one aspect of the present invention is shown. As shown in Figure 1, the matrix-vector multiplication unit includes a shift unit 140, an XOR unit 160, and a global register 150. The shift unit 140 is used to shift a vector by a specified number of bits to obtain a shift result. It should be noted that the target object that the shift unit 140 can shift can be any type of data, such as binary data, scalars, vectors, etc. In an embodiment of the present invention, the target object shifted by the shift unit 140 is a vector, and the shift result is obtained by shifting the vector.
[0046] The XOR unit 160 is connected to the shift unit 140 and the global register 150 respectively, and is used to receive the shift result from the shift unit 140 and XOR the shift result with the storage value in the global register 150 to obtain an XOR result. The XOR unit 160 receives the shift result from the shift unit 140. It should be noted here that the shift result is the shift result obtained after the shift unit 140 shifts any data. The arbitrary data can be, for example, binary data, a vector, a scalar, etc. The shift result in the present invention is the shift result obtained after the shift unit 140 shifts the vector. The XOR unit 160 XORs the shift result with the storage value in the global register 150 to obtain an XOR result. The storage value in the global register 150 in the present invention is an XOR value obtained based on the principle of matrix and vector multiplication operation, which will be described in detail below.
[0047] Global register 150 is used to store and transmit the XOR result from XOR unit 160. The XOR operation is equivalent to a modulo-2 sum of bits, and the circular shift operation of a vector is equivalent to the multiplication of the circularly shifted identity matrix and the vector. Therefore, by controlling the operands and operation process of the shift and XOR operations, the final result of the matrix-vector multiplication operation is obtained in global register 150. The calculation results in global register 150 can be stored in memory and used for further calculations.
[0048] Figure 1B FIG1 shows a flow chart of a matrix and vector multiplication method in an LDPC encoder according to an embodiment of the present invention. It can be understood that Figure 1B The flowchart shown is merely illustrative, and the steps described therein may be performed in a different order, performed in parallel, omitted, and / or with additional steps. Figure 1B As shown, the method for multiplying a matrix M by a vector S in an LDPC encoder includes step S10: initializing a global register; step S20: shifting vector S by a specified number of bits and performing an exclusive OR operation with the contents of the global register, storing the exclusive OR result in the global register; and step S30: storing the value in the global register. When matrix M is the sum of multiple cyclically shifted identity matrices, step S20 is repeated, where the number of repetitions is related to the number of cyclically shifted identity matrices. In one example, matrix M is composed of n cyclically shifted identity matrices, and step S20 is repeated n times.
[0049] Figure 2 A schematic structural diagram of a matrix and vector multiplication unit for LDPC coding according to an embodiment of the present invention is shown.
[0050] According to one embodiment of the present invention, Figure 2As shown, the matrix and vector multiplication unit for LDPC coding further includes an instruction memory 120, which is used to store instructions. The number of instructions can be multiple, and the types of instructions can be multiple.
[0051] The matrix and vector multiplication operation unit according to the present invention performs LDPC encoding or matrix and vector multiplication operations in LDPC encoding by executing an instruction sequence. When executing an instruction, the shift unit 140 can perform a shift operation of a specified number of bits on the vector, and the shift result is sent to the XOR unit 160, wherein the number of bits of the shift is specified by the instruction. When executing an instruction, data can be loaded into the global register 150, or data of the global register 150 can be stored. When executing an instruction, the XOR unit 160 can perform an XOR operation on the data of the global register 150 and the output data of the shift unit 140, and store the result in the global register 150. By executing multiple instructions in the instruction memory 120, the matrix and vector multiplication operation in the LDPC encoding process is completed.
[0052] Figure 3A A schematic structural diagram of a matrix and vector multiplication unit for LDPC coding according to another embodiment of the present invention is shown.
[0053] According to one embodiment of the present invention, Figure 3A As shown, the matrix and vector multiplication unit for LDPC coding further includes a data memory 110, which is connected to a shift unit 140 and a global register 150 (also called an Rd register). In response to executing an instruction, data at a specified location in the data memory 110 can be loaded into the global register 150, or the contents of the global register 150 can be stored in a specified location in the data memory 110. In response to executing an instruction, the shift unit 140 shifts the data at the specified location in the data memory 110, and calculates an exclusive OR between the shift result of the shift unit 140 and the contents of the global register 150, and stores the exclusive OR result in the global register 150.
[0054] According to one embodiment of the present invention, in response to executing an instruction in the instruction memory 120, a vector in the data memory 110 specified by the instruction is loaded into the global register 150. In one example, to set the initial state of the global register 150 to 0, the number "zero" or the contents of the storage space storing the number "zero" in the data memory 110 are loaded into the global register 150 by executing the instruction.
[0055] According to one embodiment of the present invention, in response to executing an instruction, the content of the global register 150 is stored in the data memory 110 .
[0056] Table 1 shows an instruction list according to an embodiment of the present invention. By combining these instructions, matrix-vector multiplication or LDPC encoding can be implemented by executing the instruction sequence. Therefore, the matrix-vector multiplication unit according to an embodiment of the present invention is also an LDPC encoder. The instruction sequence is stored in instruction memory 120.
[0057] When the LDPC encoder executes a LOAD instruction, data is loaded from the data memory 110 into the global register 150 according to the parameters described in the LOAD instruction. The data loaded by the LOAD instruction may be a vector S. The LOAD instruction may use a variety of addressing modes. In one example, the parameters described in the LOAD instruction indicate the location of the data to be loaded in the data memory 110. In another example, the parameters described in the LOAD instruction indicate the address of the data to be loaded in the data memory 110 obtained from a register. The parameters described in the LOAD instruction may also indicate an offset value relative to a base address.
[0058] When the LDPC encoder executes a STORE instruction, the data in the global register 150 is stored in the data memory 110 according to the parameters described in the STORE instruction. The data stored by the STORE instruction may be the result of shifting and / or performing an XOR on the vector S. The STORE instruction can use multiple addressing modes.
[0059] When the LDPC encoder executes the SHIFT_XOR instruction, the shift unit 140 shifts the specified data in the data memory by a specified number of bits according to the parameters described in the SHIFT_XOR instruction, and sends the shift result to the XOR unit 160. The XOR unit 160 performs XOR operation on the shift result and the value of the global register 150, and stores the result in the global register 150.
[0060] Table 1 Instruction list
[0061]
[0062] In accordance with Figure 3A In another embodiment of the present invention, the LDPC encoder can also execute NOP instructions. The NOP instruction stands for no operation and is used to avoid resource access conflicts during the execution of LDPC encoder instructions.
[0063] Figure 3B A flowchart of a matrix and vector multiplication method in an LDPC encoder according to one embodiment of the present invention is shown.
[0064] like Figure 3BAs shown, the matrix-vector multiplication method in the LDPC encoder includes step S10: initializing a global register. Step S12: retrieving vector S from a data memory. Step S20: shifting vector S by a specified number of bits and performing an exclusive OR operation with the contents of the global register, storing the exclusive OR result in the global register. Step S30: storing the value in the global register. When the matrix M is the sum matrix of multiple cyclically shifted identity matrices, step S20 is repeated, where the number of repetitions is related to the number of cyclically shifted identity matrices.
[0065] Figure 4A A schematic structural diagram of a matrix and vector multiplication unit for LDPC coding according to another embodiment of the present invention is shown.
[0066] According to one embodiment of the present invention, Figure 4A As shown, the matrix and vector multiplication unit for LDPC coding includes a data memory 110, an instruction memory 120, a shift unit 140, a global register (also known as Rd (destination) register) 150, an XOR unit 160, and a vector register (also known as Rs (source) register) 130. The vector register 130 is connected to the data memory 110 and the shift unit 140, and is used to receive the vector from the data memory 110 and provide the vector to the shift unit 140.
[0067] According to one embodiment of the present invention, in response to executing an instruction in the instruction memory 120, the data specified by the instruction is loaded into the vector register 130. In this example, the data loaded into the vector register 130 is a vector used as a multiplier in a matrix-vector multiplication operation.
[0068] Table 2 shows an instruction list according to an embodiment of the present invention. By combining these instructions, matrix and vector multiplication or LDPC coding can be implemented by executing the instruction sequence. Figure 4A The matrix and vector multiplication unit of the embodiment shown is also a LDPC encoder. Instruction sequences are stored in the instruction memory 120.
[0069] When the LDPC encoder executes a LOAD instruction, data is loaded from the data memory 110 into the vector register 130 or the destination register 150 according to the parameters described in the LOAD instruction. The data loaded by the LOAD instruction can be a vector used as a multiplier in a matrix-vector multiplication operation. The LOAD instruction can use a variety of addressing modes. In one example, the parameters described in the LOAD instruction indicate the location of the data to be loaded in the data memory 110. In another example, the parameters described in the LOAD instruction indicate the address of the data to be loaded in the data memory 110 obtained from a register. The parameters described in the LOAD instruction can also indicate an offset value relative to a base address.
[0070] When the LDPC encoder executes a STORE instruction, the data in the destination register 150 is stored in the data memory 110 according to the parameters described in the STORE instruction. The data stored by the STORE instruction can be the result of performing a shift and / or XOR on the vector. The STORE instruction can use multiple addressing modes.
[0071] When the LDPC encoder executes the SHIFT_XOR instruction, the shift unit 140 shifts the contents of the vector register 130 by a specified number of bits according to the parameters described in the SHIFT_XOR instruction, and sends the shift result to the XOR unit 160. The XOR unit 160 performs an XOR operation on the shift result and the value of the destination register 150, and stores the result in the destination register 150.
[0072] Table 2 Instruction list
[0073]
[0074] In accordance with Figure 4A In another embodiment of the present invention, the LDPC encoder can also execute NOP instructions. The NOP instruction stands for no operation and is used to avoid resource access conflicts during the execution of LDPC encoder instructions.
[0075] Figure 4B A flowchart of a matrix and vector multiplication method in an LDPC encoder according to another embodiment of the present invention is shown.
[0076] like Figure 4BAs shown, the matrix-vector multiplication method in the LDPC encoder includes step S10: initializing the global register. Step S12: obtaining vector S from the data memory. Step S14: loading vector S from step S12 into the vector register. Step S20: shifting vector S by a specified number of bits and performing exclusive OR operation with the contents of the global register, and storing the exclusive OR result in the global register. Step S30: storing the value in the global register. When the matrix M is a sum matrix of multiple cyclic shifted identity matrices, step S20 is repeatedly executed, and the number of repetitions is related to the number of cyclic shifted identity matrices.
[0077] The matrix and vector multiplication operation implemented by the LDPC encoder according to the instruction sequence is described in detail below through specific embodiments.
[0078] In the LDPC encoding process, the multiplication of matrix M and vector S is a key operation. Matrix M is the sum of n cyclically shifted identity matrices. The cyclically shifted identity matrix is the matrix obtained by cyclically shifting the identity matrix. For example, Equation (1) is an example of a cyclically shifted identity matrix. The cyclically shifted identity matrix in Equation (1) is the matrix obtained by cyclically shifting the 8*8 identity matrix right by one position.
[0079]
[0080] Since M is the sum matrix of n cyclic shift identity matrices, let M = I1 + I2 + ... + I n , where Im (1≤m≤n) is the cyclic shift identity matrix, and m and n are both positive integers.
[0081] The multiplication operation of matrix M and vector S is M*S = (I1+I2+…+In)*S. The multiplication operation of the unit circular shift identity matrix Im (1≤m≤n) and vector S can be converted into a shift operation on vector S. Furthermore, (I1+I2+…+In)*S can be converted into a shift of vector S and a modulo-2 sum of the shifted results. That is, (I1+I2+…+In)*S can be decomposed into shift(S,d1)xorshift(S,d2)xorshift(S,d3)…xorshift(S,dn), where shift(S,dm) represents a shift of vector S by dm bits; dm represents the circular shift of the identity matrix Im obtained by a circular right shift of dm bits from the identity matrix I; and XOR represents an exclusive-OR operation. Thus, the matrix-vector multiplication (M*S) can be converted into a series of shifts and exclusive-OR operations.
[0082] The process of calculating the multiplication operation of the matrix M and the vector S according to the embodiment of the present invention is described below through a specific example.
[0083] Example 1
[0084] M is the sum matrix of the 8*8 cyclic shift identity matrix, S is the 8*1 vector, the matrix M is as shown in formula (2), and the vector S is as shown in formula (3):
[0085]
[0086] S=(1 0 0 1 0 0 1 0)' (3)
[0087] In formula (2), M = I1 + I2 + I3, where I1 is the identity matrix I obtained by cyclically shifting it right by 1 bit, I2 is the identity matrix I obtained by cyclically shifting it right by 3 bits, and I3 is the identity matrix I obtained by cyclically shifting it right by 7 bits. The calculation process of M*S = I1*S XOR I2*S XOR I3*S can be decomposed into the following operations: Shift(S,1) xorShift(S,3) xorShift(S,7). These operations can be performed by Figure 3A The LDPC encoder is implemented by executing the following instruction sequence stored in the instruction memory 120:
[0088] ①LOAD Rd,0;
[0089] ②Shift_XOR[ADDR1],1;
[0090] ③Shift_XOR[ADDR1],3;
[0091] ④Shift_XOR[ADDR1],7;
[0092] ⑤STORE ADDR2,Rd.
[0093] like Figure 3A As shown, the matrix and vector multiplication unit for the LDPC encoder includes: a data memory 110, an instruction memory 120, a shift unit 140, a global register (Rd) 150 and an XOR unit 160. The data memory 110 stores the matrix M and the vector S. The data memory 110 is connected to the shift unit 140 and the global register 150 respectively. The shift unit 140 is connected to the XOR unit 160, and the XOR unit 160 is connected to the global register 150.
[0094] When instruction ① is executed, the vector value loaded into the global register 150 is 0, that is, the global register 150 is initialized to 0.
[0095] When executing instruction ②, the shift unit 140 shifts the vector S stored at address ADDR1 in the data memory 110 by one bit. The XOR unit 160 performs an XOR operation on the vector S shifted by one bit and the value of the global register 150 (initial value 0). The XOR result is stored in the global register 150. At this time, the value in the global register 150 is I1*S.
[0096] When executing instruction ③, shift unit 140 shifts vector S stored at address ADDR1 in data memory 110 by 3 bits. XOR unit 160 performs an XOR operation on the 3-bit-shifted vector S with the value stored in global register 150 after executing instruction ②, and stores the XOR result in global register 150. At this point, the value in global register 150 is I1*S XOR I2*S.
[0097] When executing instruction 4, shift unit 140 shifts vector S stored at address ADDR1 in data memory 110 by 7 bits. XOR unit 160 performs an XOR operation on vector S shifted by 7 bits with the value stored in global register 150 after executing instruction 3, and stores the XOR result in global register 150. At this point, the value in global register 250 is I1*S XOR I2*S XOR I3*S, which is the result of the calculation of M*S.
[0098] When instruction ⑤ is executed, the value in the global register 150 after executing instruction ④ is stored in the storage location with address ADDR2.
[0099] The multiplication operation method of the matrix M of formula (2) and the vector S of formula (3) corresponding to the execution of the above instructions ①-⑤ is as follows: Step S510: Initialize the global register (Rd) 150 so that the value in the global register (Rd) 150 is zero; Step S512: Get the vector S from the address ADDR1 of the data memory 110; Step S520: Shift the vector S by 1 bit and perform XOR operation with the value of the global register (Rd) 150 (initial value is 0), and the XOR result is stored in the global register (Rd) 150; Shift the vector S by 3 bits and perform XOR operation with the value of the global register (Rd) 150, and the XOR result is stored in the global register (Rd) 150; Shift the vector S by 7 bits and perform XOR operation with the value of the global register (Rd) 150, and the XOR result is stored in the global register (Rd) 150. Step S530: Store the value in the global register (Rd) 150 at the storage location of the data memory 110 at the address ADDR2.
[0100] In embodiment 1 of the present invention, in order to calculate the matrix and vector multiplication M*S, the composition of the matrix M is analyzed, and the matrix M is respectively the sum of several cyclic shift unit matrices. For each multiplication operation of the cyclic shift unit matrix and the vector S, an instruction Shift_XOR[ADDR],offset is generated, where the offset value represents the cyclic shift unit matrix obtained by cyclically shifting the unit matrix I right by the offset bit, and [ADDR] indicates that the operation object of the instruction is the data stored at the ADDR in the data memory. In addition, an instruction for initializing the global register (Rd) and an instruction for saving the calculation result are generated. In such a case, Figure 3A The generated instruction sequence (for example, the instruction sequence ①-⑤ above) is executed in the LDPC encoder shown in FIG. to obtain the calculation result of matrix and vector multiplication M*S. In addition to being applied to the LDPC encoder, as shown in FIG. Figure 3A The illustrated embodiment of the present invention may also be used in other application scenarios where matrix and vector multiplication needs to be calculated.
[0101] The LDPC encoder in this embodiment 1 involves multiple matrix and vector multiplication operations. Instruction memory 120 can store multiple instructions corresponding to the multiple matrix and vector multiplication operations, simplifying the structure of the LDPC encoder, reducing the complex control flow to instruction control, improving the operating efficiency of each functional component, and increasing throughput.
[0102] Example 2
[0103] M is the sum matrix of 8*8 cyclic shift identity matrices, S is an 8*1 vector, the matrix M is as shown in formula (4), and the vector S is as shown in formula (5):
[0104]
[0105] S=(1 0 0 1 0 0 1 0)' (5)
[0106] In formula (4), M = I1' + I2' + I3', where I1' is the cyclic shift identity matrix obtained by cyclically shifting the identity matrix I by 3 positions, I2' is the cyclic shift identity matrix obtained by cyclically shifting the identity matrix I1' by 2 positions, and I3' is the cyclic shift identity matrix obtained by cyclically shifting the identity matrix I1' by 4 positions. In this process, a total of 9 vector shift operations are performed. It can be seen that in Example 2, the matrix and vector multiplication M*S calculation is also completed, which reduces the number of vector shift operations by 2 compared to the calculation process of Example 1.
[0107] In Example 2, these operations can be performed by Figure 3AThe following instruction sequence is executed in the LDPC encoder of the embodiment. The initial state of the global register (Rd) 150 is 0. The vector S is stored in the storage space at the address ADDR1 of the data register 110.
[0108] 1)Shift_XOR[ADDR1],3;
[0109] 2)STORE ADDR1,Rd;
[0110] 3)Shift_XOR[ADDR1],-2;
[0111] 4)Shift_XOR[ADDR1]4;
[0112] 5)STORE ADDR2,Rd.
[0113] The initial state of global register (Rd) 150 is 0. Global register (Rd) 150 can be initialized by executing the instruction LOAD Rd,0. When executing instruction 1), vector S is retrieved from address ADDR1 of data memory 110, shifted by 3 bits (right shift), and XORed with the value of global register (Rd) 150 (initial value 0). The XOR result is stored in global register (Rd) 150. When executing instruction 2), the value in global register (Rd) 150 is stored in the storage space at address ADDR1 of data memory 110. When executing instruction 3), the data at address ADDR1 of data memory 110 is shifted by -2 bits (left shift), and XORed with the value of global register (Rd) 150. The XOR result is stored in global register (Rd) 150. When instruction 4 is executed, the data at address ADDR1 in data memory 110 is shifted 4 bits (right shifted) and XORed with the value in global register (Rd) 150. The XOR result is stored in global register (Rd) 150. At this point, the value in global register (Rd) 150 is the result of the calculation M*S. When instruction 5 is executed, the value in global register (Rd) 150 is stored in the storage location at address ADDR2 in data memory 110.
[0114] The multiplication method of the matrix M of formula (3) and the vector S of formula (4) corresponding to executing the above instructions 1)-5) is:
[0115] Step S610: Initialize the global register (Rd) 150 so that the value in the global register (Rd) 150 is zero;
[0116] Step S612: Obtain vector S from the address ADDR1 of the data memory 110;
[0117] Step S620: Shift the vector S by 3 bits (right shift), and perform XOR operation on it with the value of the global register (Rd) 150 (initial value is 0), and the XOR result is stored in the global register (Rd) 150; store the value in the global register (Rd) 150 in the storage space at the address ADDR1 of the data memory 110; shift the data at the address ADDR1 of the data memory 110 by -2 bits (left shift), and perform XOR operation on it with the value of the global register (Rd) 150, and the XOR result is stored in the global register (Rd) 150; shift the data at the address ADDR1 of the data memory 110 by 4 bits (right shift), and perform XOR operation on it with the value of the global register (Rd) 150, and the XOR result is stored in the global register (Rd) 150. At this time, the value in the global register (Rd) 150 is the calculation result of M*S.
[0118] Step S630 : Store the value of the global register (Rd) 150 in the storage space at the address ADDR2 of the data memory 110 .
[0119] This embodiment reuses the result of multiplying the cyclically shifted identity matrix and vector S to reduce shift operations. Consequently, compared to Example 1, fewer instructions are used during matrix-vector multiplication, resulting in faster computation. Furthermore, reducing shift operations reduces state inversions in storage cells, thereby saving energy during the computation process.
[0120] Example 3
[0121] M is the sum matrix of the 8*8 cyclic shift unit matrix, S is the 8*1 vector, the matrix M is as shown in formula (6), and the vector S is as shown in formula (7):
[0122]
[0123] S=(1 0 0 1 0 0 1 0)' (7)
[0124] In formula (6), the matrix composed of "1" in the brackets (i.e., "(1)") is the sum matrix K1 of matrices obtained by cyclically shifting the unit matrix by 3 and 4 bits respectively. And the matrix composed of "1" without brackets is the sum matrix K2 of matrices obtained by cyclically shifting the unit matrix by 6 and 7 bits respectively. Both matrices K1 and K2 are sum matrices of two cyclically shifted unit matrices, and the two cyclically shifted unit matrices constituting matrix K1 are adjacent in terms of the number of shifts relative to the unit matrix, and the two cyclically shifted unit matrices constituting matrix K2 are adjacent in terms of the number of shifts relative to the unit matrix. Therefore, it is considered that matrix K1 and matrix K2 are matrices with the same structure, or it is said that both matrix K1 and matrix K2 have two consecutive non-zero rows. Similarly, if matrix K m is the sum matrix of m cyclic shift unit matrices, and constitutes the matrix Km The number of shifts of the m cyclic shift identity matrices relative to the identity matrix is adjacent or continuous, then it is called the matrix K m With m consecutive non-zero rows. For matrices K1 and K2 with the same structure, a circular shift of a predetermined number of bits from matrix K1 (in formula (6), a right shift of 3 bits) will yield matrix K2.
[0125] In Example 3, these operations can be performed by Figure 3A The following instruction sequence is executed in the LDPC encoder of the embodiment. The initial state of the global register (Rd) 150 is 0. The vector S is stored in the storage space at the address ADDR1 of the data register 110.
[0126] 6)Shift_XOR[ADDR1],3;
[0127] 7)Shift_XOR[ADDR1],4;
[0128] 8)STORE ADDR1,Rd;
[0129] 9)Shift_XOR[ADDR1],3;
[0130] 10)STORE ADDR2,Rd.
[0131] The initial state of global register (Rd) 150 is 0. When instruction 6 is executed, vector S is retrieved from address ADDR1 of data memory 110, shifted by 3 bits (right shift), and XORed with the value of global register (Rd) 150 (initial value 0). The XOR result is stored in global register (Rd) 150. When instruction 7 is executed, vector S is retrieved from address ADDR1 of data memory 110, shifted by 4 bits (right shift), and XORed with the value of global register (Rd) 150. The XOR result is stored in global register (Rd) 150 (the result is K1*S). When instruction 8 is executed, the value of global register (Rd) 150 is stored in the storage space at address ADDR1 of data memory 110. When executing instruction 9), data (the result of K1*S) is retrieved from address ADDR1 of data memory 110. This data is shifted 4 bits (right-shifted) (to K2*S) and XORed with the value of global register (Rd) 150. The XOR result is stored in global register (Rd) 150 (K1*S XOR K2*S). When executing instruction 10), the value of global register (Rd) 150 is stored in the storage space at address ADDR2 of data memory 110.
[0132] The multiplication method of the matrix M of formula (6) and the vector S of formula (7) corresponding to executing the above instructions 6)-10) is:
[0133] Step S710: Initialize the global register (Rd) 150 so that the value in the global register (Rd) 150 is zero;
[0134] Step S712: Obtain vector S from the address ADDR1 of the data memory 110;
[0135] Step S720: Shift vector S by 3 bits (right shift), and perform XOR operation on the vector S with the value of global register (Rd) 150 (initial value is 0), and store the XOR result in global register (Rd) 150; obtain vector S from address ADDR1 of data memory 110, shift vector S by 4 bits (right shift), and perform XOR operation on the vector S with the value of global register (Rd) 150, and store the XOR result in global register (Rd) 150 (the result is K1*S); store the value of global register (Rd) 150 in the storage space of address ADDR1 of data memory 110; obtain data (the result of K1*S) from address ADDR1 of data memory 110, shift the obtained data by 4 bits (right shift) (to K2*S), and perform XOR operation on the vector S with the value of global register (Rd) 150, and store the XOR result in global register (Rd) 150 (K1*S XOR K2*S).
[0136] Step S730 : Store the value of the global register (Rd) 150 in the storage space at the address ADDR2 of the data memory 110 .
[0137] In embodiment 3, the calculation result of K1*S is reused by performing a shift operation on the data in the storage space at the address ADDR1 of the data memory 110, thereby reducing the instructions required for calculating K*S.
[0138] In Example 3, the matrix M=K1+K2, and both K1 and K2 are sum matrices of two cyclic shifted unit matrices. In another embodiment according to the present invention, both matrices K1 and K2 are sum matrices of n cyclic shifted unit matrices, and the n cyclic shifted unit matrices constituting the matrix K are adjacent or continuous in terms of the number of shifts relative to the unit matrix. Thus, the matrices K1 and K2 have the same structure, and K2*S can be obtained by shifting the calculation result of K1*S. Those skilled in the art will recognize that the matrix M can be decomposed into M=K1+K2+…+K j , where K1, K2, …, K j All have the same structure (for example, the matrix K iThe shift times of multiple cyclic shift unit matrices relative to the unit matrix are adjacent or continuous), so K can be obtained by shifting the calculation result of K1*S. i *S(2≤i≤j).
[0139] In a further embodiment according to the present invention, the composition of the matrix M is analyzed and the matrix M is decomposed into {K1, K2, ..., K j}, where K1, K2, …, K j They all have the same structure, which is the sum matrix of p cyclic shift unit matrices (p is a positive integer), and they form the matrix K i The p cyclic shift unit matrices are adjacent or continuous in number relative to the unit matrix. Thus, K can be obtained by shifting the calculation result of K1*S. i *S(2≤i≤j). {K1,K2,…,K j} is sorted so that the sum of d2,…dj is minimized, where dm represents the matrix Km (2≤m≤j) obtained by shifting dm bits from K1.
[0140] Example 4
[0141] M is the sum matrix of 8*8 cyclic shift identity matrices, S is an 8*1 vector, the matrix M is as shown in formula (8), and the vector S is as shown in formula (9):
[0142]
[0143] S=(1 0 0 1 0 0 1 0)' (9)
[0144] In equation (8), M = I1 + I2 + I3, where I1 is the identity matrix I obtained by cyclically shifting it right by 1, I2 is the identity matrix I obtained by cyclically shifting it right by 3, and I3 is the identity matrix I obtained by cyclically shifting it right by 7. The calculation process of M*S = I1*S XOR I2*S XOR I3*S can be decomposed into the following operations: Shift(S,1) xorShift(S,3) xorShift(S,7). These operations can be implemented by executing the following instruction sequence in the LDPC encoder shown in Figure 4, where the initial states of vector register (Rs) 130 and destination register (Rd) 150 are 0.
[0145] ⑩LOAD Rs,ADDR1;
[0146] Shift_XOR Rs,1;
[0147] Shift_XOR Rs,3;
[0148] Shift_XOR Rs,7;
[0149] STORE ADDR2,Rd.
[0150] like Figure 4A As shown, the matrix and vector multiplication operation unit for LDPC encoding includes: a data memory 110, an instruction memory 120, a vector register (Rs) 130, a shift unit 140, a destination register (Rd) 150 and an XOR unit 160. The data memory 110 is connected to the vector register 130 and the destination register 150 respectively, the vector register 130 is connected to the shift unit 140, the shift unit 140 is connected to the XOR unit 160, and the XOR unit 160 is connected to the destination register 150.
[0151] When instruction ⑩ is executed, vector S is obtained from the address ADDR1 of the data memory 110 and loaded into the vector register 130.
[0152] Execute instructions , the vector S in the vector register 130 is shifted by 1 bit via the shift unit 140, and the XOR unit 160 performs an XOR operation on the vector S shifted by 1 bit and the value of the destination register 150 (initial value is 0), and the XOR result is stored in the destination register 150. At this time, the value in the destination register 150 is I1*S.
[0153] Execute instructions When the shift unit 140 shifts the vector S in the vector register 130 by 3 bits, the XOR unit 160 combines the vector S shifted by 3 bits with the executed instruction stored in the destination register 150. The XOR result is stored in the destination register 150. At this time, the value in the destination register 150 is I1*S XOR I2*S.
[0154] Execute instructions When the shift unit 140 shifts the vector S in the vector register 130 by 7 bits, the XOR unit 160 combines the vector S shifted by 7 bits with the executed instruction stored in the destination register 150. The XOR result is stored in the destination register 150. At this time, the value in the destination register 150 is I1*S XOR I2*S XOR I3*S, that is, the calculation result of M*S.
[0155] Execute instructions When the command is executed The value in the destination register 150 is then stored at the memory location with address ADDR2.
[0156] Execute the above instructions ⑩- The corresponding calculation method for multiplying the matrix M of formula (8) and the vector S of formula (9) is:
[0157] Step S810: Initialize the global register (Rd) 150 so that the value in the global register (Rd) 150 is zero;
[0158] Step S812: Obtain vector S from the address ADDR1 of the data memory 110;
[0159] Step S814: Load the vector S in step S812 into the vector register (Rs) 130;
[0160] Step S820: Shift the vector S in the vector register (Rs) 130 by 1 bit, and perform XOR operation on it with the value of the global register (Rd) 150 (initial value is 0), and store the XOR result in the global register (Rd) 150; shift the vector S in the vector register (Rs) 130 by 3 bits, and perform XOR operation on it with the value of the global register (Rd) 150, and store the XOR result in the global register (Rd) 150; shift the vector S in the vector register (Rs) 130 by 7 bits, and perform XOR operation on it with the value of the global register (Rd) 150, and store the XOR result in the global register (Rd) 150.
[0161] Step S830 : Store the value in the global register (Rd) 150 in the storage location at the address ADDR2 of the data memory 110 .
[0162] In embodiment 2 according to the present invention, in order to calculate the matrix multiplication M*S by the vector, the composition of the matrix M is analyzed, and the matrix M is divided into the sum of several cyclic shift unit matrices. For each multiplication operation of the cyclic shift unit matrix and the vector S, an instruction is generated: shift Rs, offset, where the offset value represents the cyclic shift unit matrix obtained by cyclically shifting the unit matrix I right by offset bits, and Rs indicates that the operation object of the instruction is the data from the vector register 130. Instructions for initializing the vector register (Rs) and the destination register (Rd) and instructions for saving the calculation results are also generated. In the example Figure 4A The generated instruction sequence (for example, the instruction sequence ⑩- ) to obtain the result of matrix and vector multiplication M*S. In addition to being applied to LDPC encoders, such as Figure 4A The illustrated embodiment of the present invention may also be used in other application scenarios where matrix and vector multiplication needs to be calculated.
[0163] Example 5
[0164] M is the sum matrix of 8*8 cyclic shift identity matrices, S is an 8*1 vector, the matrix M is as shown in formula (10), and the vector S is as shown in formula (11):
[0165]
[0166] S=(1 0 0 1 0 0 1 0)' (11)
[0167] In formula (10), M = I1' + I2' + I3', where I1' is the cyclic shift identity matrix obtained by cyclically shifting the identity matrix I by 3 bits to the right, I2' is the cyclic shift identity matrix obtained by cyclically shifting the cyclic shift identity matrix I1' by 2 bits to the left, and I3' is the cyclic shift identity matrix obtained by cyclically shifting the cyclic shift identity matrix I1' by 4 bits to the right. In this process, a total of 9 shift operations on vector S are performed. It can be seen that in Example 4, the matrix and vector multiplication M*S is also completed, which reduces the shift operations on vector S by 2 compared to the calculation process of Example 4.
[0168] In Example 5, these operations can be performed by Figure 4A The following instruction sequence is executed in the LDPC encoder of : The initial states of the global register (Rd) 150 and the vector register (Rs) 130 are 0.
[0169] (100)LOAD Rs,ADDR1;
[0170] (200)Shift_Xor Rs,3;
[0171] (300)Store Rs,Rd;
[0172] (400)Shift_Xor Rs,-2;
[0173] (500)Shift_Xor Rs,4;
[0174] (600)STORE ADDR2,Rd.
[0175] The initial states of vector register (Rs) 130 and global register (Rd) 150 are 0. When instruction (100) is executed, vector S is obtained from address ADDR1 of data memory 110 and loaded into vector register (Rs) 130. When instruction (200) is executed, vector S in vector register (Rs) 130 is shifted by 3 bits (right shift) and XORed with the value of global register (Rd) 150 (initial value is 0), and the XOR result is stored in global register (Rd) 150. When instruction (300) is executed, the value of global register (Rd) 150 is stored in vector register (Rs) 130. When instruction (400) is executed, the data in vector register (Rs) 130 is shifted by -2 bits (left shift) and XORed with the value of global register (Rd) 150, and the XOR result is stored in global register (Rd) 150. When instruction (500) is executed, the data in vector register (Rs) 130 is shifted 4 bits (right shift) and XORed with the value in global register (Rd) 150. The XOR result is stored in global register (Rd) 150. At this time, the value in global register (Rd) 150 is the result of the calculation of M*S. When instruction (600) is executed, the value in global register (Rd) 150 is stored in the storage location at address ADDR2 of data memory 110.
[0176] The multiplication method of the matrix M of formula (10) and the vector S of formula (11) corresponding to executing the above instructions (100)-(600) is:
[0177] Step S910: Initialize the global register (Rd) 150 so that the value in the global register (Rd) 150 is zero;
[0178] Step S912: Obtain vector S from the address ADDR1 of the data memory 110;
[0179] Step S914: Load the vector S in step S912 into the vector register (Rs) 130;
[0180] Step S920: Shift the vector S in the vector register (Rs) 130 by 3 bits (right shift), and perform XOR operation on the vector S with the value of the global register (Rd) 150 (initial value is 0), and store the XOR result in the global register (Rd) 150; store the value of the global register (Rd) 150 in the vector register (Rs) 130; shift the data in the vector register (Rs) 130 by -2 bits (left shift), and perform XOR operation on the vector S with the value of the global register (Rd) 150, and store the XOR result in the global register (Rd) 150; shift the data in the vector register (Rs) 130 by 4 bits (right shift), and perform XOR operation on the vector S with the value of the global register (Rd) 150, and store the XOR result in the global register (Rd) 150;
[0181] Step S930 : Store the value in the global register (Rd) 150 in the storage location with the address ADDR2 of the data memory 110 .
[0182] In Example 5, the result of multiplying the cyclic shifted identity matrix and the vector S is reused to reduce shift operations. Reducing shift operations will reduce the state inversion of storage units, thereby saving energy consumption in the calculation process.
[0183] Example 6
[0184] M is the sum matrix of the 8*8 cyclic shift identity matrix, S is the 8*1 vector, the matrix M is as shown in formula (12), and the vector S is as shown in formula (13):
[0185]
[0186] S=(1 0 0 1 0 0 1 0)' (13)
[0187] In formula (12), the matrix composed of "1" in the brackets (i.e., "(1)") is the sum matrix K1 of matrices obtained by cyclically shifting the unit matrix by 3 and 4 bits respectively. The matrix composed of "1" without brackets is the sum matrix K2 of matrices obtained by cyclically shifting the unit matrix by 6 and 7 bits respectively. Both matrices K1 and K2 are sum matrices of two cyclically shifted unit matrices, and the two cyclically shifted unit matrices constituting matrix K1 are adjacent in terms of the number of shifts relative to the unit matrix, and the two cyclically shifted unit matrices constituting matrix K2 are adjacent in terms of the number of shifts relative to the unit matrix. Therefore, it is considered that matrix K1 and matrix K2 are matrices with the same structure, or it is said that both matrix K1 and matrix K2 have two consecutive non-zero rows. Similarly, if matrix K m is the sum matrix of m cyclic shift unit matrices, and constitutes the matrix K m The number of shifts of the m cyclic shift identity matrices relative to the identity matrix is adjacent or continuous, then it is called the matrix K m With m consecutive non-zero rows. For matrices K1 and K2 with the same structure, a circular shift of a predetermined number of bits from matrix K1 (in formula (12), a right shift of 3 bits) will yield matrix K2.
[0188] In Example 6, these operations can be performed by Figure 4A The following instruction sequence is executed in the LDPC encoder of the embodiment. The initial states of the global register (Rd) 150 and the vector register (Rs) are 0. The vector S is stored in the storage space at the address ADDR1 of the data register 110.
[0189] (110)LOAD Rs,ADDR1;
[0190] (210)Shift_Xor Rs,3;
[0191] (310)Shift_Xor Rs,4;
[0192] (410)STORE Rs,Rd;
[0193] (510)Shift_Xor Rs,3;
[0194] (610)STORE ADDR2,Rd.
[0195] The initial states of vector register (Rs) 130 and global register (Rd) 150 are 0. When instruction (110) is executed, vector S is obtained from address ADDR1 of data memory 110 and loaded into vector register (Rs) 130. When instruction (210) is executed, vector S in vector register (Rs) 130 is shifted 3 bits (right shift) and XORed with the value of global register (Rd) 150 (initial value is 0). The XOR result is stored in global register (Rd) 150. When instruction (310) is executed, vector S in vector register (Rs) 130 is shifted 4 bits (right shift) and XORed with the value of global register (Rd) 150. The XOR result is stored in global register (Rd) 150. At this time, the calculation result of K1*S is stored in global register (Rd) 150. Reusing the calculation result of K1*S and shifting it right by 3 bits will obtain the calculation result of K2*S. When instruction (410) is executed, the value of global register (Rd) 150 (i.e., the calculation result of K1*S) is stored in vector register (Rs) 130. When instruction (510) is executed, vector S in vector register (Rs) 130 is shifted by 3 bits (right shift) and XORed with the value of global register (Rd) 150. The XOR result is stored in global register (Rd) 150. At this time, the value in global register (Rd) 150 is the calculation result of M*S. When instruction (610) is executed, the value in global register (Rd) 150 is stored in the storage location at address ADDR2 of data memory 110.
[0196] The multiplication method of the matrix M of formula (12) and the vector S of formula (13) corresponding to executing the above instructions (110)-(610) is:
[0197] Step S1010: Initialize the global register (Rd) 150 so that the value in the global register (Rd) 150 is zero;
[0198] Step S1012: Obtain vector S from the address ADDR1 of the data memory 110;
[0199] Step S1014: Load the vector S in step S1012 into the vector register (Rs) 130;
[0200] Step S1020: Shift the vector S in the vector register (Rs) 130 by 3 bits (right shift), and perform XOR operation on the vector S with the value of the global register (Rd) 150 (initial value is 0), and store the XOR result in the global register (Rd) 150; shift the vector S in the vector register (Rs) 130 by 4 bits (right shift), and perform XOR operation on the vector S with the value of the global register (Rd) 150, and store the XOR result in the global register (Rd) 150; store the value of the global register (Rd) 150 (i.e., the calculation result of K1*S) in the vector register (Rs) 130; shift the vector S in the vector register (Rs) 130 by 3 bits (right shift), and perform XOR operation on the vector S with the value of the global register (Rd) 150, and store the XOR result in the global register (Rd) 150;
[0201] The third step S30: storing the value in the global register (Rd) 150 in the storage location with the address ADDR2 of the data memory 110.
[0202] In the sixth embodiment, the calculation result of K1*S is reused by performing a shift operation on the data in the storage space at the address ADDR1 of the data memory 110, thereby reducing the instructions required for calculating K*S.
[0203] In Example 6, the matrix M = K1 + K2, and both K1 and K2 are sum matrices of two cyclic shifted unit matrices. In another embodiment according to the present invention, both matrices K1 and K2 are sum matrices of n cyclic shifted unit matrices, and the n cyclic shifted unit matrices constituting the matrix K are adjacent or continuous in terms of the number of shifts relative to the unit matrix. Thus, the matrices K1 and K2 have the same structure, and K2 * S can be obtained by shifting the calculation result of K1 * S. Those skilled in the art will recognize that the matrix M can be decomposed into M = K1 + K2 + ... + K j , where K1, K2, …, K j All have the same structure (for example, the matrix K i The shift times of multiple cyclic shift unit matrices relative to the unit matrix are adjacent or continuous), so K can be obtained by shifting the calculation result of K1*S. i *S(2≤i≤j).
[0204] In a further embodiment according to the present invention, the composition of the matrix M is analyzed and the matrix M is decomposed into {K1, K2, ..., K j}, where K1, K2, …, K jThey all have the same structure, which is the sum matrix of p cyclic shift unit matrices (p is a positive integer), and they form the matrix K i The p cyclic shift unit matrices are adjacent or continuous in number relative to the unit matrix. Thus, K can be obtained by shifting the calculation result of K1*S. i *S(2≤i≤j). {K1,K2,…,K j} is sorted so that the sum of d2,…dj is minimized, where dm represents the matrix Km (2≤m≤j) obtained by shifting dm bits from K1.
[0205] Figure 5 A flowchart of a matrix and vector multiplication method in an LDPC encoder according to one embodiment of the present invention is shown.
[0206] According to the present invention Figure 5 In the embodiment of Figure 3A The LDPC encoder shown in the figure executes a sequence of instructions to calculate the matrix-vector multiplication (M*S). Where the matrix M = K1+K2+…+K m +…+K n (1≤m≤n)(m,n are both positive integers), M is the sum matrix of multiple cyclic shift unit matrices, K m is the sum matrix of p cyclic shift identity matrices (K m =I j1 +I j2 +…+I jp , where I j is the cyclic shift unit matrix), and the matrix K m The p cyclic shift identity matrices are adjacent or continuous in number relative to the identity matrix, then the matrix is called K m has p consecutive non-zero rows. And the matrix K is obtained by cyclic shifting dm bits from K1 m (2≤m≤n). The instruction sequence can be generated offline by processing the matrix M and stored in the instruction memory 120. The instruction sequence is implemented by executing the instruction sequence. Figure 5 Flowchart of the method for matrix and vector multiplication shown in FIG.
[0207] In step S1110, the global register (Rd) 150 (see Figure 3A As an example, the global register (Rd) 150 is initialized to 0.
[0208] In step S1120, for each of the p cyclic shift unit matrices constituting matrix K1, vector S is obtained from the data memory (for example, the address is ADDR1), vector S is shifted by di bits, and the shift result is XORed with the value of global register (Rd) 150 and stored in global register (Rd) 150 (execution instruction SHIFT_Xor[ADDR1],di). Through step S1120, the calculation result (S1) of K1*S is obtained. By shifting S1 by a predetermined number of bits, K can be obtained. m *S. Where K1=I 11 +I 12 +…I 1i +…+I 1p , and the cyclic shift identity matrix I is obtained by cyclic shifting di positions from the identity matrix 1i .
[0209] In step S1130 , the value ( S1 ) of the global register (Rd) 150 is stored in the data memory 110 (eg, the address is ADDR1 ) (by executing the instruction STORE ADDR1 , Rd).
[0210] In step S1140, for the n-1 matrices K constituting the matrix M, m (For example, K2, ..., K m ,…,K n ), obtain S1 from data memory address ADDR1, shift S1 by dm positions, XOR the shifted result with the value of global register (Rd) 150, and store the result in global register (Rd) 150 (execute instruction SHIFT_Xor[ADDR1],dm). The calculation result M*S is obtained in step S1140. Matrix Km is obtained by shifting K1 by dm positions.
[0211] In step S1150 , the value of the global register (Rd) 150 is stored in the data memory 110 (for example, the address is ADDR2 ) (the instruction STORE ADDR2, Rd is executed).
[0212] Figure 6 A flowchart of a matrix and vector multiplication method in an LDPC encoder according to another embodiment of the present invention is shown.
[0213] Combine Figure 4A , Figure 6 The following shows a matrix and vector multiplication method in an LDPC encoder according to another embodiment of the present invention. Wherein the matrix M = K1 + K2 + ... + K m +…+K n (1≤m≤n)(m,n are both positive integers), M is the sum matrix of multiple cyclic shift unit matrices, K mis the sum matrix of p cyclic shift identity matrices (K m =I j1 +I j2 +…+I jp , where I j is the cyclic shift unit matrix), and the matrix K m The p cyclic shift identity matrices are adjacent or continuous in number relative to the identity matrix, then the matrix is called K m has p consecutive non-zero rows. And the matrix K is obtained by cyclic shifting dm bits from K1 m (2≤m≤n). The instruction sequence can be generated offline by processing the matrix M and stored in the instruction memory 120. The instruction sequence is implemented by executing the instruction sequence. Figure 6 Flowchart of the method for matrix and vector multiplication shown in FIG.
[0214] In step S1210, the global register (Rd) 150 (see Figure 4A As an example, the global register (Rd) 150 is initialized to 0.
[0215] In step S1220, for each of the p cyclic shift unit matrices constituting the matrix K1, a vector S is obtained from the data memory (for example, the address is ADDR1), the vector S is loaded into the vector register (Rs) 130, the vector S in the vector register (Rs) 130 is shifted by di bits, the shift result is XORed with the value of the global register (Rd) 150 and stored in the global register (Rd) 150 (execution instruction SHIFT_Xor Rs,di). Through step S1220, the settlement result (S1) of K1*S is obtained. By shifting S1 by a predetermined number of bits, K can be obtained. m *S. Where K1=I 11 +I 12 +…I 1i +…+I 1p , and the cyclic shift identity matrix I is obtained by cyclic shifting di positions from the identity matrix 1i . .
[0216] In step S1230 , the value ( S1 ) of the global register (Rd) 150 is stored in the vector register (Rs) 130 (instruction STORE Rs, Rd is executed).
[0217] In step S1240, for the n-1 matrices K constituting the matrix M, m (For example, K2, ..., K m , ..., K n), obtain S1 from vector register (Rs) 130, shift S1 by dm positions, XOR the shifted result with the value of global register (Rd) 150, and store the result in global register (Rd) 150 (execute instruction SHIFT_XorRs,dm). The calculation result of M*S is obtained in step S1240. Matrix Km is obtained by shifting K1 by dm positions.
[0218] In step S1250, the value of the global register (Rd) 150 is stored in the data memory 110 (for example, the address is ADDR2) (the instruction STOREADDR2, Rd is executed). According to one embodiment of the present invention, in order to calculate the matrix-vector multiplication M*S, the matrix M may have a complex structure. For example, the matrix M is a sum matrix of multiple cyclic shifted identity matrices, Where K(i) is the sum matrix of i cyclic shifted identity matrices, and K(i) has i consecutive non-zero rows; and there are f(i) matrices K(i) with i consecutive non-zero rows, and K(i, j(i)) is the j(i)th matrix K(i) among the f(i) matrices K(i). For example, M=K(1,1)+K(1,2)+K(2,1)+K(2,2)+K(2,3)+K(3,1), where K(1,1) and K(1,2) are cyclic shift identity matrices, K(2,1), K(2,2) and K(2,3) are sum matrices of two cyclic shift identity matrices, and K(2,1), K(2,2) and K(2,3) each have two adjacent non-zero rows, that is, the two cyclic shift identity matrices constituting each of K(2,1), K(2,2) and K(2,3) are adjacent or continuous to each other with respect to the number of shifts of the identity matrix; K(3,1) is the sum matrix of three cyclic shift identity matrices, and K(3,1) each has three adjacent non-zero rows, that is, the three cyclic shift identity matrices constituting K(3,1) are adjacent or continuous with respect to the number of shifts of the identity matrix.
[0219] For multiple K(i, j(i)) with the same number of consecutive non-zero rows, compute according to Figure 5 or Figure 6 In the illustrated embodiment, calculation When i can take q different values, the calculation results corresponding to the q values of i are stored in the data memory 110 (see Figure 3A or 4A). Then, the q data stored in the data memory are Sum the results (for example, execute the instruction SHIFT_Xor ADDRi, 0), calculate
[0220] According to another aspect of the present invention, the present invention also provides a computer program comprising computer program code, which, when loaded into a computer system and executed on the computer system, causes the computer system to perform the method described above.
[0221] According to another aspect of the present invention, there is also provided a program comprising program codes, which, when loaded into a storage device and executed on the storage device, causes the storage device to execute the method described above.
[0222] The present invention optimizes the calculation process of matrix and vector multiplication, reduces the number of instructions executed in the process of matrix and vector multiplication by reusing intermediate results in the process of matrix and vector multiplication, and thus speeds up the execution speed of vector and matrix multiplication.
[0223] It should be understood that each block of the block diagrams and flowcharts, and combinations of blocks in the block diagrams and flowcharts, can be implemented by various means including computer program instructions. These computer program instructions can be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data control device to produce a machine, so that the instructions executed on the computer or other programmable data control device create means for implementing the functions specified in one or more flowchart blocks.
[0224] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data-controlled device to function in a specific manner, so that the instructions stored in the computer-readable memory can be used to manufacture an article of manufacture comprising computer-readable instructions for implementing the functions specified in one or more flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable data-controlled device to cause a series of operational steps to be performed on the computer or other programmable data-controlled device, thereby producing a computer-implemented process, whereby the instructions executed on the computer or other programmable data-controlled device provide steps for implementing the functions specified in one or more flowchart blocks.
[0225] Thus, the blocks of the block diagrams and flow charts support combinations of means for performing the specified functions, combinations of steps for performing the specified functions, and combinations of program instruction means for performing the specified functions. It should also be understood that each block of the block diagrams and flow charts, and combinations of blocks of the block diagrams and flow charts, can be implemented by a hardware-based special-purpose computer system that performs the specified functions or steps, or by a combination of special-purpose hardware and computer instructions.
[0226] At least a portion of the various blocks, operations, and techniques described above may be implemented using hardware, a control device executing firmware instructions, a control device executing software instructions, or any combination thereof. When implemented using a control device executing firmware and software instructions, the software or firmware instructions may be stored on any computer-readable storage medium, such as a disk, optical disk, or other storage medium, in RAM, ROM, or flash memory, on a control device, on a hard disk, optical disk, or on a magnetic disk, etc. Similarly, the software and firmware instructions may be transmitted to a user or system via any known or desired transmission method, including, for example, on a computer-readable disk or other portable computer storage mechanism or via a communication medium. Communication media typically embodies computer-readable instructions, data structures, sequence modules, or other data in a modulated data signal, such as a carrier wave or other transmission mechanism. By way of example, and not limitation, communication media include wired media, such as a wired network or a single-wire connection, and wireless media, such as acoustic, radio frequency, infrared, and other wireless media. Thus, software and firmware instructions may be transmitted to a user or system via a communication channel, such as a telephone line, a DSL line, a cable television line, a fiber optic cable, a wireless channel, the Internet, or the like (providing such software via a portable storage medium is considered equivalent or interchangeable). The software or firmware instructions may include machine-readable instructions that, when executed by a control device, cause the control device to perform various actions.
[0227] When implemented in hardware, the hardware may include one or more discrete components, an integrated circuit, an application specific integrated circuit (ASIC), and the like.
[0228] It should be understood that the present invention can be implemented in pure software, pure hardware, firmware, or various combinations thereof. Hardware can be, for example, a control device, a dedicated integrated circuit, a large-scale integrated circuit, or the like.
[0229] Although the present invention has been described with reference to examples, this is for purposes of illustration only and not limitation, and changes, additions and / or deletions to the embodiments may be made without departing from the scope of the present invention.
[0230] Those skilled in the art to which these embodiments relate and who have benefited from the teachings presented in the above description and the associated drawings will recognize many modifications and other embodiments of the inventions described herein. Therefore, it should be understood that the invention is not limited to the specific embodiments disclosed, and modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. A matrix and vector multiplication unit, characterized in that: include: A data memory, a shift unit, an XOR unit and a global register; the matrix and vector multiplication operation unit performs a multiplication operation of a matrix M and a vector S by executing an instruction sequence, wherein the matrix M is a sum matrix of a cyclic shift unit matrix; Wherein, the matrix M=K1+K2+...K m +…+K n , K m It is the sum matrix of p cyclic shift identity matrices, p is a positive integer, cyclic shift d from K1 m Get the matrix K m , cyclic shift dI from the identity matrix I j The cyclic shift identity matrix I is obtained j , where 1≤j≤p; and K1=I1+I2+...Ij+…+I P , I1, dI2…dIj,…dIp are continuous natural numbers; for each of the p cyclic shift unit matrices constituting the matrix K1, the shift unit shifts the vector S in the data memory by d by executing the Shift_XOR instruction 1j The shift result is sent to the XOR unit, which performs an XOR operation on the data of the global register and the output data of the shift unit, and stores the result in the global register; for the n-1 matrices K2, ...K constituting the matrix M m ,…K n Each matrix K in m , where 2≤m≤n, by executing the Shift_XOR instruction, the shift unit shifts the vector S in the data memory by the specified number of bits d m The shift result is sent to the XOR unit, and the XOR unit performs an XOR operation on the data of the global register and the output data of the shift unit, and stores the result in the global register.
2. The matrix and vector multiplication unit according to claim 1, wherein: Also includes: The vector register is connected to the data memory and the shift unit, and is used for receiving the vector from the data memory and providing the vector to the shift unit.
3. The matrix and vector multiplication unit according to claim 2, wherein: Wherein, the matrix M=K1+K2+...K m +…+K n , where K1, K m is the sum of two cyclically shifted identity matrices, cyclically shifted from K1 by d m Get the matrix K m , the cyclic shift identity matrix I1 is obtained by cyclic shifting dI1 from the identity matrix I, and the cyclic shift identity matrix I2 is obtained by cyclic shifting dI2 from the identity matrix I, dI1 and dI2 are consecutive natural numbers, and K1=I1+I2, where, ; For the n-1 matrices K2 that make up the matrix M, ... . K m ,…K n Each matrix K in m ,in , execute the Shift_XOR instruction, so that the shift unit shifts the content in the data memory by d m Bit, get K m *S, K m *S is XORed with the value of the global register and the XOR result is stored in the global register.
4. The matrix and vector multiplication unit according to any one of claims 1 to 3, characterized in that: in, Vector S is a vector stored in the data memory at address ADDR1; and wherein, By executing the Shift_XOR instruction, the vector S at the address ADDR1 in the data memory is shifted by dI1 bits, the shift result is XORed with the value of the global register and stored in the global register, and the Shift_XOR instruction is executed to shift the vector S at the address ADDR1 in the data memory by dI2 bits, the shift result is XORed with the value of the global register to obtain a specific value, which is K1*S, and stored in the global register.
5. The unit according to any one of claims 1 to 3, characterized in that in, The matrix , where K(i) is the sum matrix of i cyclic shifted identity matrices, and K(i) has i consecutive non-zero rows; and there are f(i) matrices K(i) with i consecutive non-zero rows, and K(i,j(i)) is the j(i)th matrix in the f(i) matrices K(i); For each value of i, calculate The result is stored in the data storage; The calculation results for each value of i in the data memory are summed to obtain the calculation result of M*S.
6. The unit according to any one of claims 1 to 3, characterized in that Also includes: An instruction memory is used to store the instruction sequence; the instructions of the instruction sequence include a LOAD instruction, a Shift_XOR instruction, and a STORE instruction.
7. The unit according to claim 6, characterized in that in Execute a LOAD instruction to load the value of the data memory into the global register; Execute the STORE instruction to store the value of the global register into the data memory.
8. The unit according to claim 6 or 7, characterized in that include: When a LOAD instruction is executed, data is loaded from the data memory into the global register according to the parameters described in the LOAD instruction; wherein the parameters described in the LOAD instruction indicate the address of the data to be loaded in the data memory obtained from the register, or the parameters described in the LOAD instruction indicate the offset value of the address of the data to be loaded in the data memory obtained from the register relative to the base address.
9. The unit according to claim 6 or 7, characterized in that include: When a STORE instruction is executed, data in a global register is stored in a data memory according to parameters described in the STORE instruction; wherein the data stored by the STORE instruction is a result of shifting and / or performing an XOR on a vector S; and wherein the STORE instruction uses multiple addressing modes.
10. The unit according to claim 6 or 7, characterized in that The instructions of the instruction sequence further include: NOP instruction, which stands for no operation.
Citation Information
Patent Citations
Quasi-loop LDPC code encoding device capable of on-line programming
CN101399553A
QC-LDPC code coding method and device
CN105099467A