COMPUTER SYSTEM AND METHOD FOR ARITHMETIC DECODING
The computer system with a dedicated arithmetic decoding circuit addresses the inefficiencies of existing methods by reducing processor cycles and energy consumption, thereby enhancing decoding efficiency.
Patent Information
- Application Number
- FR2023014483
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
Existing arithmetic decoding methods require numerous branches at each iteration, leading to inefficiencies and increased time and energy consumption.
A computer system with a dedicated circuit for arithmetic decoding, which performs comparisons based on coefficients and interval values, and counts identical successive results to determine symbol values, thereby reducing the number of processor cycles and energy consumption.
The proposed solution accelerates arithmetic decoding by reducing the number of processor cycles and energy consumption, while also avoiding numerous branches, thus improving decoding efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: COMPUTER SYSTEM AND ARITHMETIC DECODING METHOD
[0001] Embodiments and implementations relate to arithmetic decoding.
[0002] Arithmetic coding is a data compression technique commonly used in computer science and signal processing. Unlike traditional binary coding, where each symbol is represented by a fixed number of bits, arithmetic coding allows symbols to be represented using fractions of real number intervals.
[0003] Arithmetic coding has the advantage of being able to efficiently compress data based on the probability of the symbols. More frequent symbols will occupy smaller fractions of the interval, while less frequent symbols will occupy larger fractions, thus allowing more efficient compression of the data.
[0004] Arithmetic coding is notably used by the “LC3” (Low Complexity Communication Coded) audio coder-decoder. Arithmetic coding is then used to compress audio data and then arithmetic decoding to restore the original audio data from the compressed representation.
[0005] In particular, arithmetic coding implements different steps. First, probabilities are assigned to the symbols. In particular, before starting the coding, each symbol in the message is associated with a probability. The probabilities can be based on statistics of the frequency of occurrence of the symbols in the message.
[0006] Next, an initial interval is defined at the beginning of the coding process. Then the message is traversed symbol by symbol. At each step, the current interval is divided into sub-intervals, the size of which is proportional to the probabilities of the symbols. The sub-interval corresponding to the symbol currently being processed is selected to represent this symbol. The current interval then becomes the selected sub-interval.
[0007] Once all symbols have been encoded, the binary representation of the real number in the final interval is extracted. This binary representation is the compressed message.
[0008] Once the message has been compressed by arithmetic coding, it is possible to recover the initial message by carrying out arithmetic decoding reversing the steps implemented during arithmetic coding.
[0009] In particular, the arithmetic decoding comprises an initialization of the interval initial used when coding from the probabilities associated with the symbols.
[0010] Once the initial interval is reconstructed, iterative decoding of the symbols can begin. The decoder begins reading the bits of the compressed message's binary representation one by one, and at each step, it readjusts the interval based on the bit sequence read. It must take into account the probabilities associated with the symbols to determine which subinterval corresponds to the symbol being decoded.
[0011] At each iteration, the decoder compares the current interval with the subintervals corresponding to each possible symbol. When a subinterval matches the current interval, the decoder identifies the symbol associated with that subinterval as the decoded symbol.
[0012] More particularly, at each iteration, a comparison is made between the interval start value and a value calculated by a multiplication between the length of the interval and a coefficient. Each coefficient corresponds to a value associated with a symbol. The sequence of tests is carried out until a comparison result corresponding to a test validity condition is obtained. To decode a symbol, the number of tests carried out is a value between 1 and the number of possible symbols.
[0013] Known methods for performing arithmetic decoding require branching at each iteration. These numerous branches increase the time required to perform the arithmetic decoding. Thus, such arithmetic decoding is inefficient.
[0014] There is therefore a need to propose a solution allowing arithmetic decoding to be carried out more quickly.
[0015] According to one aspect, there is provided a computer system comprising: - a data memory configured to store a table of coefficients, - a digital signal processing processor configured to execute a computer program comprising instructions for performing arithmetic decoding from said table of coefficients and an interval of values, - a circuit dedicated to said arithmetic decoding, this circuit being configured to: • make comparisons based on the coefficients of said coefficient table and said range of values, then • counting the results of said comparisons which are identical and successive, the digital signal processing processor being configured to determine a value of a symbol associated with said table of coefficients from the number of said identical and successive results counted by said dedicated circuit.
[0016] The use of a circuit dedicated to arithmetic decoding integrated in said digital signal processing processor makes it possible to accelerate arithmetic decoding, by reducing the number of cycles of said signal processing processor. digital to perform arithmetic decoding. In this way, it also helps to reduce the energy consumption for performing arithmetic decoding. In addition, by first performing all the comparisons associated with each of the coefficients in the table before calculating the value of the symbol, it is possible to avoid making numerous connections with respect to known solutions.
[0017] In an advantageous embodiment, the dedicated circuit comprises a first block configured to calculate a comparison variable making it possible to store the results of said comparisons, the result of each comparison corresponding to a bit of the comparison variable.
[0018] Preferably, the first block is configured to carry out in parallel at least two comparisons from at least two successive coefficients of the coefficient table with the interval of values.
[0019] Thus, in such a computer system, the fact of first carrying out all of the comparisons associated with each of the coefficients of the table before calculating the value of the symbol allows for parallelization of the comparisons. This therefore makes it possible to accelerate the arithmetic decoding.
[0020] Advantageously, the first block comprises at least two parallel branches allowing said two comparisons to be carried out simultaneously, each branch comprising: • a multiplier circuit configured to perform a multiplication between a coefficient among said two successive coefficients and a length of said interval of values shifted by a given number of bits, this given number of bits being in particular between 0 and a number of bits used to define the length of said interval (for example said given number of bits is equal to 10), • a comparison circuit configured to compare a result of the multiplication to a starting value of the range of values.
[0021] In an advantageous embodiment, the first block further comprises a circuit for updating the comparison variable configured to integrate into said comparison variable the results of the comparison carried out by each comparison circuit.
[0022] Preferably, the dedicated circuit also comprises a second block configured to implement a state machine configured to count the successive identical results stored in said comparison variable by analyzing the bits of said comparison variable.
[0023] Advantageously, the results of each comparison carried out by the first block are stored progressively to the right in the comparison variable. The state machine implemented by the second block is then configured to count the number of identical comparison results from the bits located furthest to the right in the comparison variable.
[0024] Alternatively, the results of each comparison performed by the first block are stored progressively to the left in the comparison variable. The state machine implemented by the second block is then configured to count the number of identical comparison results from the leftmost bits in the comparison variable.
[0025] In an advantageous embodiment, said computer program comprises instructions which, when executed by the digital signal processing processor, lead it to: - implement the first block in order to carry out in parallel a comparison from each coefficient of said coefficient table with said interval, the result of each comparison being stored in the comparison variable, then - implement the second block, once all the comparisons have been carried out, to count the successive identical results stored in said comparison variable by analyzing the bits of said comparison variable, then - calculate a value of the symbol associated with said coefficient table from the number of successive identical results counted.
[0026] According to another aspect, there is provided a method implemented by a computer system, the method comprising an execution, by a digital signal processing processor of the computer system, of a computer program comprising instructions for performing an arithmetic decoding from a table of coefficients stored in a data memory of the computer system and an interval of values, the arithmetic decoding comprising: - an implementation of a circuit of the computer system dedicated to said arithmetic decoding for: • make comparisons based on the coefficients of said coefficient table and said range of values, then • counting the results of said comparisons which are identical and successive, then - a determination by the digital signal processing processor of a value of a symbol associated with said table of coefficients from the number of said identical and successive results counted by said dedicated circuit.
[0027] In an advantageous embodiment, the implementation of said dedicated circuit comprises an implementation of a first block of the dedicated circuit for calculating a comparison variable making it possible to store the results of said comparisons, the result of each comparison corresponding to a bit of the comparison variable.
[0028] Preferably, the implementation of the first block is adapted to carry out in parallel at least two comparisons from at least two successive coefficients of the coefficient table with the interval of values.
[0029] Advantageously, the implementation of the first block comprises an implementation of at least two parallel branches adapted to simultaneously carry out said two comparisons, the implementation of each branch comprising: • the implementation of a multiplier circuit to perform a multiplication between a coefficient among said two successive coefficients and a length of said interval of values shifted by a given number of bits, this given number of bits being in particular between 0 and a number of bits used to define the length of said interval, • the implementation of a comparison circuit to compare a result of the multiplication with a starting value of the range of values.
[0030] In an advantageous embodiment, the implementation of the first block further comprises an implementation of a circuit for updating the comparison variable to integrate into said comparison variable the results of the comparison carried out by each comparison circuit.
[0031] Preferably, the implementation of the dedicated circuit also comprises an implementation of a second block for executing a state machine adapted to count the successive identical results stored in said comparison variable by analyzing the bits of said comparison variable.
[0032] Advantageously, the results of each comparison carried out by the first block are stored progressively to the right in the comparison variable. The state machine implemented by the second block is then adapted to count the number of identical comparison results from the bits located furthest to the right in the comparison variable.
[0033] Alternatively, the results of each comparison performed by the first block are stored progressively to the left in the comparison variable. The state machine implemented by the second block is then adapted to count the number of identical comparison results from the bits located furthest to the left in the comparison variable.
[0034] In an advantageous embodiment, the execution of said computer program results in: - an implementation of the first block in order to carry out in parallel a comparison from each coefficient of said coefficient table with said interval, the result of each comparison being stored in the comparison variable, then - an implementation of the second block, once all the comparisons have been carried out, to count the successive identical results stored in said comparison variable by analyzing the bits of said comparison variable, then - a calculation by the digital signal processing processor of a value of the symbol associated with said coefficient table from the number of identical results successive counted.
[0035] Other advantages and characteristics of the invention will appear on examining the detailed description of embodiments, which are in no way limiting, and the appended drawings in which:
[0036] [Fig.l]
[0037] [Fig.2]
[0038] [Fig.3]
[0039] [Fig.4]
[0040] [Fig.5]
[0041] [Fig.6]
[0042] [Fig.7]
[0043] [Fig. 8]
[0044] [Fig.9] illustrate embodiments and implementations of the invention.
[0045] [Fig.l] illustrates a first embodiment of a computer system SYS1. The computer system SYS1 comprises a central processing unit CPU1, a main memory MMEM1 and a digital signal processor DSP1 (also referred to as "Digital signal processor") with a data memory MEM1 and a program memory MEMP1. The DSP1 comprises a control unit CU1 with a bank of data registers RFI (also referred to as "Control Unit and Register File") and an arithmetic and logic unit ALU1 (also referred to as "Arithmetic and Logic Unit"). The arithmetic and logic unit ALU1 comprises a circuit HWC1 dedicated to accelerating arithmetic decoding. The computer system SYS1 may be a system on chip.
[0046] The digital signal processing processor DSP1 is configured to execute a computer program PRG1 comprising instructions for performing arithmetic decoding.
[0047] This computer program PRG1 can be stored in the program memory MEMP1 of the computer system SYS1.
[0048] The digital signal processing processor DSP1 comprises a first register RI 1 configured to store a variable CBITS containing the comparison results.
[0049] The digital signal processing processor DSP1 also comprises a second register R21 configured to store a length RGE of an interval of values and a value LW of the start of this interval of values. The length RGE and the value LW can be concatenated in a same binary word RGELW in the second register R21.
[0050] The digital signal processing processor DSP1 also includes a third register R31 configured to store two concatenated coefficients.
[0051] The digital signal processing processor DSP1 also includes a fourth register R41 configured to store a counter TZC_C of bits with zero on the right.
[0052] The HWC1 circuit can in particular be obtained from an “RTL” (Register Transfer Level) code.
[0053] The dedicated HWC1 circuit comprises a first VMULT1 block configured to perform calculations and tests in a vectorial manner (i.e. in parallel) during arithmetic decoding. This first VMULT1 block is illustrated in [Fig.2].
[0054] In particular, the first block VMULT1 is configured to receive as input the variable CBITS stored in the first register RI 1. The first block VMULT1 is also configured to receive as input the word RGELW stored in the second register R21 comprising the interval length RGE and the interval start value LW. The interval length RGE and the value LW can be concatenated in the same word received as input. The first block VMULT1 is also configured to receive as input the two concatenated coefficients CUM_FREQ2.
[0055] The first block VMULT1 comprises a first shift circuit SFT11. This first shift circuit SFT11 is configured to receive the word RGELW comprising the interval RGE length and the concatenated LW value. This first shift circuit is configured to shift this word by 42 bits to the right. This makes it possible to output the 32 bits associated with the interval start LW value and then to shift the value of the interval RGE length by 10 bits to the right in order to obtain a temporary value TMP.
[0056] The first block VMULT1 also comprises a first AND11 logic gate of the “AND” type. This first AND11 logic gate of the “AND” type is configured to receive the word RGELW comprising the concatenated interval length RGE and the LW value as well as a first MSK11 mask of hexadecimal value 'OxFFFFFFFF'. This first AND11 logic gate of the “AND” type makes it possible to apply the first MSK11 mask to said RGELW word received as input so as to recover the LW value at the start of the interval.
[0057] The first block VMULT1 also comprises a second shift circuit SFT21. This second shift circuit SFT21 is configured to receive the two concatenated coefficients CUM_FREQ2 and to shift these two concatenated coefficients CUM_FREQ2 by 16 bits to the right so as to retain only the odd coefficient C_O1.
[0058] The first block VMULT1 further comprises a second logic gate AND21 of type “AND”. This second logic gate AND21 of type “AND” is configured to receive the two concatenated coefficients CUM_FREQ2 as well as a second MSK21 mask with hexadecimal value 'OxFFFF'. This second AND21 logic gate of type "AND" makes it possible to apply the second MSK21 mask to said value CIM_FREQ2 in order to recover the even coefficient C_E1.
[0059] The first block VMULT1 also comprises two parallel branches BRCH11, BRCH21. A first branch BRCH11 comprises a first multiplier circuit MLT11 and a first comparison circuit CMP11. The second branch BRCH21 comprises a second multiplier circuit MLT21 and a second comparison circuit CMP21. The two branches BRCH11, BRCH21 make it possible to carry out in parallel two multiplications then two comparison tests from the two different coefficients C_O1 and C_E1 resulting from said two concatenated coefficients CUM_FREQ2.
[0060] In particular, the first multiplier circuit MLT11 is configured to receive as input the temporary value TMP generated at the output of the first shift circuit SFT11 and the odd coefficient C_O1. The first multiplier circuit MLT11 is then configured to multiply said temporary value TMP with the odd coefficient C_O1.
[0061] The second multiplier circuit MLT21 is configured to receive as input the temporary value TMP generated at the output of the first shift circuit SFT11 and the even coefficient C_E1. The second multiplier circuit MLT21 is then configured to multiply said temporary value TMP with the even coefficient C_E1.
[0062] The first comparison circuit CMP11 is configured to receive the result of the multiplication carried out by the first multiplication circuit MLT11 as well as the interval start value LW. The first comparison circuit CMP11 is then configured to compare the result of said multiplication with the interval start value LW in order to know whether said interval start value LW is greater than or equal to the result of said multiplication. If said LW value is greater than or equal to the result of said multiplication, then the first comparison circuit CMP11 generates a comparison bit bl equal to 1. If said LW value is less than the result of said multiplication, then the first comparison circuit CMP11 generates a comparison bit bl equal to 0.
[0063] The second comparison circuit CMP21 is configured to receive the result of the multiplication carried out by the second multiplication circuit MLT21 as well as the interval start value LW. The second comparison circuit CMP21 is then configured to compare the result of said multiplication with the interval start value LW in order to know whether said interval start value LW is greater than or equal to the result of said multiplication. If said LW value is greater than or equal to the result of said multiplication, then the first comparison circuit CMP21 generates a comparison bit b0 equal to 1. If said LW value is less than to the result of said multiplication, then the first comparison circuit CMP21 generates a comparison bit bO equal to 0.
[0064] The first block VMULT1 also includes a circuit CUPDT1 for updating the CBITS variable. The circuit CUPDT1 for updating the CBITS variable makes it possible to integrate the comparison bits bl and bO to the right of the old CBITS variable received at the input of the first block VMULTL
[0065] In particular, the update circuit CUPDT1 comprises a third shift circuit SFT31 configured to shift by one bit to the left the variable CBITS received at the input of the first block VMULTL
[0066] The update circuit CUPDT1 also comprises a first logic gate OR11 of the “OR” type configured to perform an “OR” type operation between the variable CBITS shifted by one bit and the comparison bit b0. The first logic gate OR11 of the “OR” type thus makes it possible to integrate the comparison bit b0 into the variable CBITS. This makes it possible to obtain a temporary CIBTS variable CBITS_T
[0067] The update circuit CUPDT1 also comprises a fourth shift circuit SFT41 configured to shift the temporary variable CBITS_T integrating the comparison bit bO by one bit to the left.
[0068] The update circuit CUPDT1 also comprises a second logic gate OR21 of the “OR” type configured to perform an “OR” type operation between the comparison bit bl and the temporary variable CBITS_T integrating the comparison bit bO. The second logic gate OR21 of the “OR” type thus makes it possible to integrate the comparison bit bl into the temporary variable CBITS_T. This makes it possible to obtain an updated CBITS variable integrating the comparison bits bO and bl to the right of the CBITS variable. The new CBITS variable can then be stored in the register RI 1 of the digital signal processing processor DSP1.
[0069] The dedicated circuit HWC1 comprises a second block TZC configured to receive the variable CBITS stored in the first register RI 1. The second block TZC is configured to implement a state machine for executing a method of counting a number of bits to zero on the right in the variable CBITS. Such a method is illustrated in [Fig.3].
[0070] The counting method comprises an initialization step 30. The initialization step 30 makes it possible to initialize an index i to 0, a stop variable STP to 0, a zero bit counter TZC_C to 0, and a variable CB to the value of the variable CBITS.
[0071] The counting method then comprises a first comparison step 31. This first comparison step 31 is adapted to compare whether the value of the index i is less than 32.
[0072] If the value of the index i is less than 32, then the method then comprises a second comparison step 32. This second comparison step 32 is adapted to compare whether the value of the least significant bit CB[0] of the variable CB (i.e. the rightmost bit in the variable CB) is equal to 1.
[0073] If the value of the least significant bit CB[0] of the variable CB is different from 1, in particular equal to 0, then the method then comprises a third comparison step 33. This third comparison step 33 is adapted to compare whether the stop variable STP is equal to 0.
[0074] If the stop variable STP is equal to 0, then the method comprises a step 34 of incrementing the bit counter to zero TZC_C.
[0075] Then, the method comprises a step 36 of incrementing the index i and shifting the variable CB. In this step 36, the value of the index i is incremented by 1 and the variable CB is shifted to the right by one bit.
[0076] Then, the method is repeated from the first comparison step 31.
[0077] If the value of the least significant bit CB[0] of the variable CB is equal to 1 in step 32, then the method comprises a step 35 of updating the stop variable STP. In this step 35, the stop value STP is set to 1. Then, the method resumes at said step 36 of incrementing the index i and shifting the variable CB.
[0078] If the value of the stop variable STP is different from 0, in particular equal to 1, in step 33, then the method resumes at step 36 of incrementing the index i and shifting the variable CB.
[0079] If the value of index i is equal to 32 in step 31 then the value of the bit counter at zero is written into register R41.
[0080] As seen previously, the digital signal processing processor DSP1 is configured to execute a computer program PRG1 comprising instructions which make it possible to carry out arithmetic decoding. In particular, the execution of said instructions leads the digital signal processing processor DSP1 to execute a function DEC_SYMB1. This function DEC_SYMB1 is illustrated in [Fig.4].
[0081] This function DEC_SYMB1 is configured to determine the value of a symbol from a coefficient array CUM_FREQ2_TAB[] associated with this symbol and stored in the data memory MEM1 of the digital signal processing processor DSP1. In order to align each vector of two 16-bit elements of this array to 32-bits, this array is organized according to the parity of the number of symbols NUMSYM as shown in [Fig.9]. A zero element is inserted at the beginning of each array. When NUMSYM is even, an additional element of a hexadecimal value OxFFFF is added to the end of the array.
[0082] More particularly, the decoding method comprises an initialization step 40 before determining the symbol value from the coefficient table CUM_FREQ2_TAB[].
[0083] In this step 40, the pointer CUM_FREQ2_PTR is initialized to point to the address of the coefficient table CUM_FREQ2_TAB[2]. In addition, the variable CBITS is initialized to 1. The length RGE of the value interval and the value LW of the start of the interval are concatenated in a single word RGELW stored in the register RI 1. An index j is defined at 0.
[0084] Once the initialization is complete, the symbol value can be decoded by the following method.
[0085] Said method comprises a test step in which the index j is compared to the number of iterations n_itera to be performed. In particular, the number of iterations n_itera to be performed depends on the number of possible symbols.
[0086] If the number of possible symbols is odd, then the number of iterations to be performed is calculated by the following formula:
[0087] n_itera = (NUMSYM-1) / 2, where n_itera is the number of iterations to perform and numsym is the number of possible symbols.
[0088] If the number of possible symbols is even, then the number of iterations to be performed is calculated by the following formula:
[0089] n_üera = N UMSYM / 2, where n_itera is the number of iterations to perform and numsym is the number of possible symbols.
[0090] If the index j is less than the number of iterations n_itera to be performed, then the method comprises a step 42 making it possible to perform multiplications and comparisons in parallel from two coefficients of the coefficient table.
[0091] Step 42 includes reading a pair of coefficients and then updating the CUM_FREQ2_PTR pointer to point to the coefficient that follows the two coefficients read.
[0092] Step 42 then comprises a calculation of a new CBITS value by implementing the first block VMULT1 of the circuit HWC1. In particular, the first block VMULT1 is implemented by taking as input the current CBITS value, the word RGELW and the two coefficients CUM_FREQ2 previously read.
[0093] The implementation of the first block VMULT1 makes it possible to perform a multiplication between the CUM_FREQ2 coefficients read and the RGE value of the length of the interval shifted 10 bits to the right, then the comparisons between the results of these multiplications and the LW value of the start of the interval. The results of these comparisons correspond to the comparison bits bO and bl which are integrated into the new variable CBITS.
[0094] Step 42 subsequently comprises an incrementation of the index j by 1.
[0095] Then, the method resumes at step 41 so as to repeat the operations implemented implemented by the first VMULT1 block for each coefficient of the CUM_FREQ[] table until the index j reaches the number of increments to be performed.
[0096] When the index j reaches the number of increments to be performed in step 41, then the method comprises a step 43 of calculating the value VAL. This value VAL is notably calculated by implementing the second block TZC in order to determine the number of zeros on the right in the last calculated CBITS variable.
[0097] In particular, if the number of possible symbols is odd, then the value of the symbol is calculated by the formula:
[0098] VAL = ( NUMSYM - 1 ) - TZC{CBITS^ where NUMSYM is an even number of possible symbols, TZC(CBITS) is the function implemented by the second TZC block from the CBITS variable.
[0099] If the number of possible symbols is even, then the value of the symbol is calculated by the formula:
[0100] VAL = NUMSYM - TZC(CBITS) where NUMSYM is an odd number of possible symbols, TZC(CBITS) is the function implemented by the second TZC block from the variable CBITS.
[0101] A symbol can then be decoded from the value VAL calculated from the interval defined by the interval length RGE and the interval start value.
[0102] In such an arithmetic decoding method, the fact of first carrying out all of the comparisons associated with each of the coefficients of the table before calculating the value of the symbol allows parallelization of the comparisons. This therefore makes it possible to accelerate the arithmetic decoding.
[0103] [Fig.5] illustrates a second embodiment of a computer system SYS2. The SYS2 computer system is similar to the SYS1 computer system of [Fig.l]. In particular, the SYS2 computer system comprises a central processing unit CPU2, a main memory MMEM2 and a digital signal processor DSP2 (also referred to as "Digital signal processor") with a data memory MEM2, a program memory MEMP2. The digital signal processor DSP2 comprises a control unit CU2 with a bank of data registers RF2 (also referred to as "Control Unit and Register File") and an arithmetic and logic unit ALU2 (also referred to as "Arithmetic and Logic Unit"). The arithmetic and logic unit ALU2 includes a circuit HWC2 dedicated to accelerating arithmetic decoding is integrated. The SYS2 computer system can be a system on a chip.
[0104] The digital signal processor DSP2 is configured to execute a program PRG2 computer program comprising instructions for performing arithmetic decoding This PRG2 computer program can be stored in the program memory MEMP2 of the SYS2 computer system.
[0105] The digital signal processing processor DSP2 comprises a first register R12 configured to store a comparison variable CBITS.
[0106] The digital signal processing processor DSP2 also comprises a second register R22 configured to store a length RGE of an interval of values and a value LW of the start of this interval of values. The length RGE and the value LW can be concatenated in a same binary word RGELW in the second register R22.
[0107] The digital signal processing processor DSP2 also includes a third register R32 configured to store two concatenated coefficients CUM_FREQ2.
[0108] The digital signal processing processor DSP2 also includes a fourth register R42 configured to store a counter LZC_C of left zero bits in the variable CBITS.
[0109] The HWC2 circuit can in particular be integrated into the DSP2 digital signal processing processor. Such an HWC2 circuit can in particular be obtained from an “RTL” (Register Transfer Level) code.
[0110] The dedicated HWC2 circuit includes a first VMULT2 block configured to perform calculations and tests in a vectorial manner (i.e., in parallel) during arithmetic decoding. This first VMULT2 block is illustrated in [Fig.6].
[0111] In particular, the first block VMULT2 is configured to receive as input the variable CBITS stored in the first register R12. The first block VMULT2 is also configured to receive as input the RGE interval length of values and the LW interval start value. The RGE interval length and the LW value can be concatenated in the same word RGELW stored in the register R22. The first block VMULT2 is also configured to receive as input the two concatenated coefficients CUM_FREQ2.
[0112] The first block VMULT2 comprises a first shift circuit SFT12. This first shift circuit SFT12 is configured to receive the word RGELW comprising the interval RGE length and the concatenated LW value. This first shift circuit SFT12 is configured to shift this word by 42 bits to the right. This makes it possible to output the 32 bits associated with the interval start LW value and then to shift the value of the interval RGE length by 10 bits to the right in order to obtain a temporary value TMP.
[0113] The first VMULT2 block also includes a first AND12 logic gate of type “AND”. This first AND12 logic gate of type “AND” is configured to receive the word RGELW comprising the RGE interval length and the concatenated LW value as well as a first MSK12 mask of hexadecimal value 'OxFFFFFFFF'. This first AND 12 logic gate of type "AND" makes it possible to apply the first MSK12 mask to said word RGELW received at input so as to recover the LW value at the start of the interval.
[0114] The first block VMULT2 also comprises a second shift circuit SFT22. This second shift circuit SFT22 is configured to receive the two concatenated coefficients CUM_FREQ2 and to shift these two concatenated coefficients CUM_FREQ2 by 16 bits to the right so as to retain only the odd coefficient C_O2 of said two concatenated coefficients CUM_FREQ2.
[0115] The first VMULT2 block further comprises a second AND22 logic gate of the “AND” type. This second AND22 logic gate of the “AND” type is configured to receive the two concatenated coefficients CUM_FREQ2 as well as a second mask of hexadecimal value 'OxFFFF'. This second AND22 logic gate of the “AND” type makes it possible to apply the second mask to the two concatenated coefficients CUM_FREQ2 so as to recover the even coefficient C_E2 of said two concatenated coefficients CUM_FREQ2.
[0116] The first block VMULT2 also comprises two parallel branches BRCH12 and BRCH22. A first branch BRCH12 comprises a first multiplier circuit MLT12 and a first comparison circuit CMP 12. The second branch BRCH22 comprises a second multiplier circuit MLT22 and a second comparison circuit CMP22. The two branches BRCH12 and BRCH22 make it possible to carry out in parallel two multiplications then two comparison tests from the two different coefficients C_O2, C_E2 resulting from the two concatenated coefficients CUM_FREQ2.
[0117] In particular, the first multiplier circuit MLT12 is configured to receive as input the temporary value TMP generated at the output of the first shift circuit SFT12 and the odd coefficient C_O2. The first multiplier circuit MLT12 is then configured to multiply said temporary value TMP with the odd coefficient C_O2.
[0118] The second multiplier circuit MLT22 is configured to receive as input the temporary value TMP generated at the output of the first shift circuit SFT12 and the even coefficient C_E2. The second multiplier circuit MLT22 is then configured to multiply said temporary value TMP with the even coefficient C_E2.
[0119] The first comparison circuit CMP 12 is configured to receive the result of the multiplication carried out by the first multiplication circuit MLT12 as well as the interval start LW value. The first comparison circuit CMP12 is then configured to compare the result of said multiplication with the interval start LW value in order to know whether said interval start LW value is greater than or equal to the result of said multiplication. If said value LW is greater than or equal to the result of said multiplication, then the first comparison circuit CMP12 generates a comparison bit bl equal to 1. If said value LW is less than the result of said multiplication, then the first comparison circuit CMP 12 generates a comparison bit bl equal to 0.
[0120] The second comparison circuit CMP22 is configured to receive the result of the multiplication carried out by the second multiplication circuit MLT22 as well as the interval start LW value. The second comparison circuit CMP22 is then configured to compare the result of said multiplication with the interval start LW value in order to know whether said interval start LW value is greater than or equal to the result of said multiplication. If said LW value is greater than or equal to the result of said multiplication, then the first comparison circuit CMP22 generates a comparison bit bl equal to 1. If said LW value is less than the result of said multiplication, then the first comparison circuit CMP22 generates a comparison bit bl equal to 0.
[0121] The first VMULT2 block also includes a CUPDT2 circuit for updating the CBITS variable. The CUPDT2 circuit for updating the CBITS variable makes it possible to integrate the comparison bits bl and bO to the right of the old CBITS variable received at the input of the first VMULT2 block.
[0122] In particular, the update circuit CUPDT2 comprises a third shift circuit SFT32 configured to shift the variable CBITS received as input by one bit to the right.
[0123] The update circuit CUPDT2 includes a fourth shift circuit SFT42 configured to shift the comparison bit b0 by 31 bits to the left.
[0124] The update circuit CUPDT2 also comprises a first OR logic gate 12 of the “OR” type configured to perform an “OR” type operation between the CBITS variable shifted by one bit and the shifted comparison bit bO. The first OR logic gate 12 of the “OR” type thus makes it possible to integrate the comparison bit bO to the left in the CBITS variable. This makes it possible to obtain a temporary CBITS_T variable.
[0125] The update circuit CUPDT2 also comprises a fifth shift circuit SFT52 configured to shift the temporary variable CBITS_T integrating the comparison bit bO by one bit to the right.
[0126] The update circuit CUPDT2 includes a sixth shift circuit SFT62 configured to shift the comparison bit bl by 31 bits to the left.
[0127] The update circuit CUPDT2 also includes a second logic gate OR22 of the “OR” type configured to perform an “OR” type operation between the shifted comparison bit bl and the temporary variable CBITS_T integrating the comparison bit comparison bO shifted one bit to the right. The second OR22 logic gate of the "OR" type thus allows the comparison bit bl to be integrated on the left into the temporary CBITS_T variable. This makes it possible to obtain an updated CBITS variable integrating the comparison bits bO and bl to the left of the CBITS variable. The new CBITS variable can then be stored in the R12 register of the DSP2 digital signal processing processor.
[0128] The dedicated HWC2 circuit comprises a second LZC block configured to receive the CBITS variable stored in the first register R12. The second LZC block is configured to implement a state machine for executing a method of counting a number of leading zeros in the CBITS variable. Such a counting method is illustrated in [Fig.7].
[0129] The counting method comprises an initialization step 70. The initialization step 70 makes it possible to initialize an index i to 0, a stop variable STP to 0, a counter LZC_C of bits with zero on the left to 0, and a variable CB to the value of the variable CBITS.
[0130] The counting method then comprises a first comparison step 71. This first comparison step 71 is adapted to compare whether the value of the index i is less than 32.
[0131] If the value of the index is less than 32, then the method then comprises a second comparison step 72. This second comparison step 72 is adapted to shift the variable CB by a number of bits equal to 31-i and then apply a mask of hexadecimal value 0x1 to said shifted variable CB. The second comparison step 72 is also adapted to then compare whether the value obtained after applying the mask to the shifted variable CB is equal to 1.
[0132] If the value obtained after applying the mask to the shifted variable CB is different from 1, in particular equal to 0, then the method then comprises a third comparison step 73. This third comparison step 73 is adapted to compare whether the stop variable STP is equal to 0.
[0133] If the stop variable STP is equal to 0, then the method comprises a step 74 of incrementing the counter LZC_C of bits to zero on the left.
[0134] Then, the method comprises a step 76 of incrementing the index i. In this step, the value of the index i is incremented by 1.
[0135] Then, the method is repeated from the first comparison step 71.
[0136] If the value obtained after applying the mask to the shifted CB variable is equal to 1 in step 72, then the method comprises a step 75 of updating the stop variable STP. In this step 75, the stop value STP is set to 1. Then, the method resumes at said step 76 of incrementing the index i.
[0137] If the value of the stop variable STP is different from 0, in particular equal to 1, at step 73, then the method resumes at step 76 of incrementing the index i.
[0138] If the value of index i is equal to 32 in step 71 then the value of the left zero bit counter LZC_C is stored in register R42.
[0139] As seen previously, the digital signal processing processor DSP2 is configured to execute a computer program PRG2 comprising instructions which make it possible to carry out arithmetic decoding. In particular, the execution of said instructions leads the digital signal processing processor DSP2 to execute a function DEC_SYMB2. This function DEC_SYMB2 is illustrated in [Fig.8].
[0140] This DEC_SYMB2 function is configured to determine the value of a symbol from a coefficient array CUM_FREQ2_TAB[] associated with this symbol and stored in the data memory MEM2 of the digital signal processing processor DSP2. In order to align each vector of 2 16-bit elements of this array to 32-bit, this array is organized according to the parity of the number of symbols NUMSYM as shown in [Fig.9]. A zero element is inserted at the beginning of each array. When NUMSYM is even, an additional element of a hexadecimal value OxFFFF is added to the end of the array.
[0141] More particularly, the decoding method comprises an initialization step 80 before determining the value of the symbol from the coefficient table CUM_FREQ2_TAB[].
[0142] In this step 80, the pointer CUM_FREQ2_PTR is initialized to point to the address of the coefficient table CUM_FREQ2_TAB[2]. In addition, the variable CBITS is initialized to 1. The length RGE of the value interval and the starting value LW of the interval are concatenated in the same word RGELW. An index j is defined at 0.
[0143] Once the initialization is complete, the symbol can be decoded by the following method.
[0144] The method comprises a test step 81 in which the index j is compared to the number of iterations n_itera to be performed. In particular, the number of iterations n_itera to be performed depends on the number of possible symbols.
[0145] If the number of possible symbols is odd, then the number of iterations to be performed is calculated by the following formula:
[0146] n_itera = (numsym- 1) / 2, where n_itera is the number of iterations to perform and numsym is the number of possible symbols.
[0147] If the number of possible symbols is even, then the number of iterations to be performed is calculated by the following formula:
[0148] n_itera — numsym12, where n_itera is the number of iterations to perform and numsym is the number of possible symbols.
[0149] If the index j is less than the number of iterations to be performed, then the method includes a step 82 allowing multiplications and comparisons to be carried out in parallel from two coefficients in the coefficient table.
[0150] Step 82 includes an update of the CUM_FREQ2_PTR pointer in order to point to the coefficient which follows the two coefficients read.
[0151] Step 82 then comprises a calculation of a new CBITS value by implementing the first block VMULT2 of the circuit HWC2. In particular, the first block is implemented by taking as input the current CBITS value, the word RGELW and the two previously read CUM_FREQ2 coefficients.
[0152] The implementation of the first VMULT2 block makes it possible to perform a multiplication between the coefficients read and the value of the length of the interval shifted 10 bits to the right, then the comparisons between the results of these multiplications and the start value of the interval. The results of these comparisons correspond to the comparison bits bO and bl which are integrated into the new variable CBITS.
[0153] Step 82 subsequently comprises a step of incrementing the index j. In this step, the index j is incremented by 1.
[0154] Then, the method resumes at step 81 so as to reiterate the operations implemented by the first block VMULT2 for each coefficient of the table CUM_FREQ[] until the index j reaches the number of increments to be performed.
[0155] When the index j reaches the number of increments to be performed in step 81, then the method comprises a step 83 of calculating the value VAL. This value VAL is notably calculated by implementing the second LZC block in order to determine the number of zeros on the right in the last calculated CBITS variable.
[0156] In particular, if the number of possible symbols is odd, then the value of the symbol is calculated by the formula:
[0157] VAL = ( NUMSYM - 1 ) - LZC(CBITS), where NUMSYM is an even number of possible symbols, LZC(CBITS) is the function implemented by the second LZC block from the variable CBITS.
[0158] If the number of possible symbols is even, then the value of the symbol is calculated by the formula:
[0159] VAL = NUMSYM - LZCiCBITS), where NUMSYM is an odd number of possible symbols, LZC(CBITS) is the function implemented by the second LZC block from the variable CBITS.
[0160] A symbol can then be decoded from the value VAL calculated from the interval defined by the interval length RGE and the interval start value.
[0161] In such an arithmetic decoding method, the fact of first carrying out all of the comparisons associated with each of the coefficients of the table before calculating the symbol value allows for parallel comparisons. This therefore speeds up arithmetic decoding.
[0162] Of course, the embodiments are susceptible to various variants and modifications which will appear to those skilled in the art. In particular, it is possible for the digital signal processing processor DSP2 to be configured to implement itself the function performed by the second LZC block. Thus, in this case it is not necessary to provide such a second block in the dedicated circuit. This then allows a saving of space in the computer system SYS2.
[0163] In addition, the first blocks VMULT1 and VMULT2 may have more branches than the two branches BRCH11, BRCH21 and BRCH12, BRCH22 respectively in order to perform a greater number of multiplications and comparisons at a time in order to process more quickly all the coefficients of the coefficient table CUM_FREQ2_TAB[].
[0164] The arithmetic decoding methods previously described can be implemented in the context of decoding an audio signal, normally for an LC3 decoder. The symbols to be decoded then correspond to audio data.
Claims
Claims
1. Computer system comprising: - a data memory (MEM) configured to store a table of coefficients (CUM_FREQ2_TAB[]), - a digital signal processing processor (DSP1, DSP2) configured to execute a computer program (PRG1, PRG2) comprising instructions for performing arithmetic decoding from said table of coefficients and an interval of values, - a circuit (HWC1, HWC2) dedicated to said arithmetic decoding, this circuit being configured to: • perform comparisons from the coefficients of said table of coefficients and said interval of values, then • count the results of said comparisons which are identical and successive, the digital signal processing processor (DSP1, DSP2) being configured to determine a value of a symbol associated with said table of coefficients from the number of said identical and successive results counted by said dedicated circuit.
2. System according to claim 1, in which the dedicated circuit (HCW1, HCW2) comprises a first block (VMULT1, VMULT2) configured to calculate a comparison variable (CBITS) making it possible to store the results of said comparisons, the result of each comparison corresponding to a bit of the comparison variable (CBITS).
3. System according to claim 2, in which the first block (VMULT1, VMULT2) is configured to carry out in parallel at least two comparisons from at least two successive coefficients (C_O1, C_E1, C_O2, C_E2) of the coefficient table with the interval of values.
4. System according to claim 3, in which the first block (VMULT1, VMULT2) comprises at least two parallel branches (BRCH11, BRCH21, BRCH12, BRCH22) allowing said two comparisons to be carried out simultaneously, each branch comprising: • a multiplier circuit (MLT11, MLT21, MLT12, MLT22) configured to carry out a multiplication between a coefficient among said two successive coefficients (C_O1, C_E1, C_O2, C_E2) and a length (RGE) of said interval of values shifted by a given number of bits, this given number of bits being in particular between 0 and a number of bits used to define the length of said interval, • a comparison circuit (CMP11, CMP21, CMP12, CMP22) configured to compare a result of the multiplication with a start value (LW) of the interval of values.
5. System according to claim 4, in which the first block further comprises a circuit (CUPDT1, CUPDT2) for updating the comparison variable configured to integrate into said comparison variable (CBITS) the results of the comparison carried out by each comparison circuit.
6. System according to one of claims 2 to 5, in which the dedicated circuit (HCW1, HCW2) also comprises a second block (TZC, LZC) configured to implement a state machine configured to count the successive identical results stored in said comparison variable (CBITS) by analyzing the bits of said comparison variable (CBITS).
7. The system of claim 6, wherein the results of each comparison performed by the first block (VMULT1) are stored progressively to the right in the comparison variable (CBITS), and wherein the state machine implemented by the second block (TZC) is configured to count the number of identical comparison results from the rightmost bits in the comparison variable (CBITS).
8. The system of claim 6, wherein the results of each comparison performed by the first block (VMULT2) are stored progressively to the left in the comparison variable (CBITS), and wherein the state machine implemented by the second block (LZC) is configured to count the number of identical comparison results from the leftmost bits in the comparison variable (CBITS).
9. System according to one of claims 6 to 8, in which said computer program (PRG1, PRG2) comprises instructions which when executed by the digital signal processing processor (DSP1, DSP2) cause the latter to: - implement the first block (VMULT1, VMULT2) in order to carry out in parallel a comparison from each coefficient of said coefficient table with said interval, the result of each comparison being stored in the comparison variable (CBITS), then - implement the second block (TZC, LZC), once all the comparisons have been carried out, to count the successive identical results stored in said comparison variable (CBITS) by analyzing the bits of said comparison variable (CBITS), then - calculate a value (VAL) of the symbol associated with said coefficient table from the number of successive identical results counted.
10. A method implemented by a computer system (SYS1, SYS2), the method comprising an execution, by a digital signal processing processor (DSP1, DSP2) of the computer system (SYS1, SYS2), of a computer program (PRG1, PRG2) comprising instructions for performing arithmetic decoding from a table of coefficients stored in a data memory of the computer system (SYS1, SYS2) and from an interval of values, the arithmetic decoding comprising: - an implementation of a circuit (HWC1, HWC2) of the computer system dedicated to said arithmetic decoding for: • performing comparisons from the coefficients of said table of coefficients and said interval of values, then • counting the results of said comparisons which are identical and successive, then - a determination by the digital signal processing processor (DSP1,DSP2) of a value of a symbol associated with said coefficient table from the number of said identical and successive results counted by said dedicated circuit.,
11. Method according to claim 10, wherein the implementation of said dedicated circuit (HCW1, HCW2) comprises an implementation of a first block (VMULT1, VMULT2) of the dedicated circuit for calculating a comparison variable (CBITS) making it possible to store the results of said comparisons, the result of each comparison corresponding to a bit of the comparison variable (CBITS).
12. Method according to claim 11, in which the implementation of the first block (VMULT1, VMULT2) is adapted to carry out in parallel at least two comparisons from at least two successive coefficients (C_O1, C_E1, C_O2, C_E2) of the coefficient table with the interval of values.
13. The method of claim 12, wherein the implementation of the first block (VMULT1, VMULT2) comprises an implementation of at least two parallel branches (BRCH11, BRCH21, BRCH12, BRCH22) adapted to simultaneously carry out said two comparisons, the implementation of each branch comprising: • the implementation of a multiplier circuit (MLT11, MLT21, MLT12, MLT22) to carry out a multiplication between a coefficient among said two successive coefficients (C_O1, C_E1, C_O2, C_E2) and a length (RGE) of said interval of values shifted by a given number of bits, this given number of bits being in particular between 0 and a number of bits used to define the length of said interval, • the implementation of a comparison circuit (CMP11, CMP21, CMP12, CMP22) to compare a result of the multiplication with a start value (LW) of the interval of values.
14. Method according to claim 13, in which the implementation of the first block (VMULT1, VMULT2) further comprises an implementation of a circuit (CUPDT1, CUPDT2) for updating the comparison variable to integrate into said comparison variable (CBITS) the results of the comparison carried out by each comparison circuit.
15. Method according to one of claims 11 to 14, wherein the implementation of the dedicated circuit (HCW1, HCW2) also comprises an implementation of a second block (TZC, LZC) for executing a state machine adapted to count the successive identical results stored in said comparison variable (CBITS) by analyzing the bits of said comparison variable (CBITS).
16. Method according to claim 15, in which the results of each comparison carried out by the first block (VMULT1) are stored progressively to the right in the comparison variable (CBITS), and in which the state machine implemented by the second block (TZC) is adapted to count the number of identical comparison results from the bits located furthest to the right in the comparison variable (CBITS).
17. Method according to claim 15, in which the results of each comparison carried out by the first block (VMULT2) are stored progressively to the left in the comparison variable (CBITS), and in which the state machine implemented by the second block (LZC) is adapted to count the number of identical comparison results from the bits located furthest to the left in the comparison variable (CBITS).
18. A method according to one of claims 15 to 17, wherein the execution said computer program causes: - an implementation of the first block (VMULT1, VMULT2) in order to carry out in parallel a comparison from each coefficient of said coefficient table with said interval, the result of each comparison being stored in the comparison variable (CBITS), then - an implementation of the second block (TZC, LZC), once all the comparisons have been carried out, to count the successive identical results stored in said comparison variable (CBITS) by analyzing the bits of said comparison variable (CBITS), then - a calculation by the digital signal processing processor (DSP1, DSP2) of a value (VAL) of the symbol associated with said coefficient table from the number of successive identical results counted.
Citation Information
Patent Citations
Parallel CABAC decoding for video decompression
US7932843B2
Hardware-based cabac decoder with parallel binary arithmetic decoding
WO2008002804A1
Parallel arithmetic coding techniques
WO2017074539A1