VVC dependent quantization hardware pipeline architecture design method

By designing the VVC-dependent quantization hardware pipeline architecture, the parallel quantization problem of relying on quantization in the H.266/VVC standard is solved, and the real-time encoding of 4K videos is achieved and the encoding speed is significantly improved.

CN120111222APending Publication Date: 2025-06-06CENT SOUTH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510343119.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

There are two main challenges in the hardwareization that relies on quantization in the H.266/VVC standard: the strong dependence between complex rate distortion cost calculations and transform coefficients, which makes it difficult to quantize in parallel.

Method used

A VVC-dependent quantization hardware pipeline architecture is designed, which is divided into initialization stage, pipeline quantization stage and optimal quantization path output stage. Pre-quantization, distortion calculation, quantization candidate value calculation, context index calculation, code rate estimation and Viterbi calculation are performed through 8-level pipeline units, and ping-pong storage is used for maximum quantization throughput.

Benefits of technology

Real-time encoding of 4K video is realized, which significantly improves the encoding speed, and is greatly improved compared with software encoding. At the same time, the performance of the hardware pipeline is improved when the encoding efficiency loss is acceptable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111222A_ABST
    Figure CN120111222A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video compression coding and decoding, and discloses a VVC dependent quantization hardware pipeline architecture design method, which comprises the following steps of: introducing a predefined code table in an initialization stage to initialize the predefined code table; an eight-level assembly line unit is adopted in the assembly line quantization stage, and four pieces of quantization state information are finally obtained by performing pre-quantization, distortion calculation, quantization candidate value calculation, context index calculation, code rate estimation and viterbi calculation on a transformation coefficient through the eight-level assembly line unit; a first grid memory and a second grid memory are introduced in an optimal quantization path output stage, ping-pong storage operation is carried out through the two grid memories, and quantization throughput can be maximized; the next TB block can be quantized without waiting for the completion of the output of the quantization result; according to the method, the problem that parallel quantization is difficult to realize between the conversion coefficients in the existing dependent quantization is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video compression coding and decoding, and in particular to a VVC dependent quantization hardware pipeline architecture design method. Background Art

[0002] H.266 / VVC is the latest generation of video codec standards jointly developed by the International Telecommunication Union (ITU-T) and the International Organization for Standardization (ISO / IEC) and officially released in July 2020. H.266 / VVC continues the block-based hybrid coding architecture of the H.26x series of standards, but optimizes multiple coding tools on this basis. The optimization of H.266 / VVC on coding tools mainly includes: larger block size and more flexible block division method, multiple intra-frame prediction modes, affine mode of inter-frame prediction, more transform size types, improved quantization and entropy coding, and loop filters with better filtering effects. After optimizing the above coding tools, H.266 / VVC has a 50% improvement in coding efficiency compared to H.265 / HEVC, but at the cost of a 5-fold increase in coding complexity compared to H.265 / HEVC. Among them, the introduction of quantization-dependent coding tools in quantization is one of the important reasons for increasing coding complexity.

[0003] Before the H.266 / VVC standard is widely used, one of the problems that needs to be solved is to hardwareize the algorithm proposed by the standard. For dependent quantization, there are two main challenges in its hardwareization: 1) Complex rate-distortion cost calculation. The transform coefficients need to perform a large amount of distortion calculation and bit rate estimation in the process of selecting quantization candidate values. The huge amount of calculation will result in low hardware throughput and main frequency. 2) Strong dependencies between transform coefficients. In the DQ quantization process, there are dependencies between quantization states and context dependencies during bit rate estimation. These two dependencies make it difficult to quantize transform coefficients in parallel. Summary of the invention

[0004] The present invention provides a VVC dependent quantization hardware pipeline architecture design method to solve the problem that it is difficult to parallel quantize coefficients in the existing dependent quantization.

[0005] In order to achieve the above object, the present invention is implemented by the following technical solutions: The present invention provides a VVC dependent quantization hardware pipeline architecture design method, comprising: dividing the dependent quantization hardware pipeline architecture into three stages, namely: an initialization stage, a pipeline quantization stage, and an optimal quantization path output stage; Introducing a predefined code table in an initialization phase, initializing the predefined code table in the initialization phase to obtain an initialized predefined code table; In the pipeline quantization stage, an 8-stage pipeline unit is used to input the transform coefficients into the 8-stage pipeline unit in the order of anti-diagonal scanning. The transform coefficients are pre-quantized, distortion calculated, quantization candidate value calculated, context index calculated, bit rate estimated, and Viterbi calculated by the 8-stage pipeline unit to finally obtain 4 quantization state information; The first grid memory and the second grid memory are introduced in the output stage of the optimal quantization path. The quantization state information calculated by different TB blocks is saved by the first grid memory and the second grid memory. The quantization throughput can be maximized by performing ping-pong storage operations through the two grid memories. After all the coefficients of a TB block are quantized through the pipeline quantization stage, the next TB block can be quantized without waiting for the quantization result output to be completed.

[0006] Optionally, the initializing the predefined code table to obtain an initialized predefined code table includes: Acquire code table data and syntax elements in a predefined code table, determine a maximum size according to the code table data, and determine a clock cycle required for initializing the code table data according to the maximum size; After the code table data is initialized based on the clock cycle, the initialized code table data is matched with the syntax element according to a matching rule of the predefined code table to obtain an initialized predefined code table.

[0007] Optionally, the pre-quantization includes: The transform coefficients are pre-quantized by a quantization algorithm model, wherein the quantization algorithm model satisfies the following relationship:

[0008] In the formula, is the pre-quantization result, is the absolute value of the transformation coefficient, is the quantized scaling factor, is the quantization offset factor, is the quantization shift factor.

[0009] The prequantization result required for calculating the candidate quantization value is obtained by prequantizing the transform coefficients.

[0010] Optionally, the distortion calculation includes: The distortion calculation model is used to calculate the distortion of the quantization candidate values ​​obtained through pre-quantization, wherein the distortion calculation model satisfies the following relationship: ; In the formula, is the original transform coefficient, is the quantized index value, is the quantization state of the dependent quantization, represents the quantization step size, sgn(x) is the sign function, Indicates distortion.

[0011] Optionally, the quantization candidate value calculation includes: ; In the formula, represents the quantized candidate value, Represents the pre-quantization result.

[0012] Optionally, the context index calculation includes: The average value of the candidate quantization values ​​obtained by pre-quantization of the transform coefficients is calculated, and the quantization result of the transform coefficients is predicted. The calculation method satisfies the following relationship: ; In the formula, represents the average value of the quantized candidate values, Indicates the quantization candidate value 0, represents the quantization candidate value 1, represents the quantization candidate value 2, represents the quantization candidate value 3; The sum of the context templates is calculated, and the variable f(T) related to the context template in the context index is calculated. The calculation method satisfies the following relationship: ; ; In the formula, , Represents an intermediate result that depends on the context template, Represents a context template, represents the partial reconstruction quantization index; The variable ctxOffset related to the diagonal position of the transform coefficient is calculated by the comparator, and the context index of each syntax element is obtained by calculating the sum of the variable f(T) and the variable ctxOffset.

[0013] Optionally, the bit rate estimation includes: Decomposing the quantization candidate values ​​obtained by pre-quantization to obtain corresponding syntax elements, and querying the estimated bit rate corresponding to the syntax elements through the initialized predefined code table; ; ; In the formula, is the quantized index value, represents the partially reconstructed quantization index value, , , , , rem represents the syntax element after the transformation coefficient is decomposed; The estimated code rate is added to the distortion of the corresponding quantization candidate value to obtain the rate-distortion cost of the quantization candidate value.

[0014] Optionally, the Viterbi calculation includes: The bit rate and distortion obtained after bit rate estimation are accumulated to obtain the cumulative rate-distortion cost, and the cumulative rate-distortion costs of multiple quantization candidate values ​​of each quantization coefficient are compared. Only the four groups of quantization candidate values ​​with the smallest rate-distortion cost are retained through the Viterbi algorithm. After X comparisons, four quantization paths are obtained for each TB block, where X is the number of quantization coefficients.

[0015] Beneficial effects: The VVC dependent quantization hardware pipeline architecture design method provided by the present invention, in terms of algorithm, this paper proposes a hardware-friendly bit rate estimation scheme, including an improved context index calculation method, fixed Rice parameters and removal of syntax elements decAbslevel, which facilitates the implementation of the hardware pipeline. In terms of implementation, this paper implements the dependent quantization hardware module on the FPGA through the Verilog HDL hardware language, and achieves the superior performance of real-time encoding of 4K video. Overall, the scheme has a significant improvement in encoding speed compared to software encoding when the loss of encoding efficiency is acceptable. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A scheme diagram of a VVC dependent quantization hardware pipeline architecture design method according to a preferred embodiment of the present invention; Figure 2 It is a top-level architecture diagram of the dependency quantization hardware pipeline of a preferred embodiment of the present invention; Figure 3 This is a diagram of an 8-stage pipeline module architecture of a preferred embodiment of the present invention; Figure 4 A diagram of the hardware architecture of context index calculation according to a preferred embodiment of the present invention; Figure 5 The timing diagram of the design pipeline hardware architecture of the preferred embodiment of the present invention; Figure 6 A hardware architecture diagram of a bit rate estimation module designed for a preferred embodiment of the present invention; Figure 7 A hardware architecture diagram of a Viterbi algorithm module designed for a preferred embodiment of the present invention; Figure 8 A quantified illustration of a single grid memory according to a preferred embodiment of the present invention; Fig. 9 A quantitative illustration of a dual-grid memory according to a preferred embodiment of the present invention; Fig.10 FIG. 4 is a hardware architecture diagram of a dual-grid memory module according to a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0017] The technical solution of the present invention is described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0018] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the usual meanings understood by persons with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one" or "one" do not indicate quantity restrictions, but indicate the existence of at least one. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship also changes accordingly.

[0019] See also Figure 1 , an embodiment of the present application provides a VVC dependent quantization hardware pipeline architecture design method, comprising: dividing the dependent quantization hardware pipeline architecture into three stages, namely: an initialization stage, a pipeline quantization stage, and an optimal quantization path output stage; Introducing a predefined code table in an initialization phase, initializing the predefined code table in the initialization phase to obtain an initialized predefined code table; In the pipeline quantization stage, an 8-stage pipeline unit is used to input the transform coefficients into the 8-stage pipeline unit in the order of anti-diagonal scanning. The transform coefficients are pre-quantized, distortion calculated, quantization candidate value calculated, context index calculated, bit rate estimated, and Viterbi calculated by the 8-stage pipeline unit to finally obtain 4 state information; Among them, when estimating the bit rate, the predefined code table initialized in the initialization phase is called to perform bit rate estimation; The first grid memory and the second grid memory are introduced in the output stage of the optimal quantization path. The quantization state information calculated by different TB blocks is saved by the first grid memory and the second grid memory. The quantization throughput can be maximized by performing ping-pong storage operations through the two grid memories. After all the coefficients of a TB block are quantized through the pipeline quantization stage, the next TB block can be quantized without waiting for the quantization result output to be completed.

[0020] In the above embodiment, in view of the problem that the transform coefficients are dependent on the quantization results of adjacent coefficients in the dependent quantization rate estimation process, a hardware-friendly rate estimation scheme is proposed, which is conducive to the implementation of the pipeline parallel quantization hardware architecture; a multi-state parallel quantization hardware structure and a dual-grid ping-pong storage structure are proposed, both of which can greatly improve the computational throughput of the quantization process; the dependent quantization hardware architecture is divided into three stages for implementation, namely: an initialization stage, a pipeline quantization stage, and an optimal quantization path output stage, wherein the pipeline stage is implemented using an eight-stage pipeline, Figure 2 The figure shows the top-level architecture of the quantization-dependent hardware pipeline designed in this paper, which includes three processes: initialization, pipeline quantization, and optimal quantization path output.

[0021] After initialization, the transform coefficients enter the 8-stage pipeline unit in the order of anti-diagonal scanning. Figure 3 As shown. The first stage pre-quantizes the transform coefficients ( ); The 2nd to 5th stages perform distortion calculations ( - ), quantization candidate value calculation ( - ) and context index calculations ( - ); stage 6-7 is to estimate the bit rate ( - ); Stage 8 implements the Viterbi algorithm ( ).

[0022] In the embodiment, the optimal quantization path output stage mainly selects the quantization paths selected by the four Viterbi algorithms, and finally the quantization path with the lowest rate distortion cost is output as the optimal path. Two grid memories are introduced in the optimal quantization path output stage. In the TB block quantization process, after the current TB block is initialized, there is no need to wait for all coefficients of the previous TB block to be output from the grid memory. The quantization state information after passing through the 8-stage pipeline can be directly stored in another grid memory, and the quantization time interval between the two TBs can be reduced as much as possible through the ping-pong operation.

[0023] Optionally, the initializing the predefined code table to obtain an initialized predefined code table includes: Acquire code table data and syntax elements in a predefined code table, determine a maximum size according to the code table data, and determine a clock cycle required for initializing the code table data according to the maximum size; After the code table data is initialized based on the clock cycle, the initialized code table data is matched with the syntax element according to a matching rule of the predefined code table to obtain an initialized predefined code table.

[0024] In the above embodiment, in the dependent quantization module, the code rate estimation is performed through a predefined code table when calculating the rate distortion cost, and the context update only updates the index of the code table. Due to the simple and easy implementation of the predefined code table, this design continues to use this method, so before officially entering the pipeline quantization, the dependent quantization needs to perform operations such as code table initialization.

[0025] The syntax elements of the quantization index are shown in Table 1. The code tables that need to be initialized are shown in Table 2. The maximum size indicates the maximum capacity required by the code table. Since the context model of the syntax elements of the luminance block is generally more than the context model of the chrominance TB block, the maximum size refers to the number of context models of the luminance TB block. The initialization process of the code table is parallel. goRiceTable is a fixed constant code table and does not need to be initialized; sigRateTable can be evenly split into 3 code tables for initialization according to the dependent quantization state, and the maximum size of each code table is 12. The cycle required for initialization depends on the maximum size of each code table in Table II. It can be seen from Table II that after splitting sigRateTable, it takes 32 clock cycles to initialize lastXRateTable and lastYRateTable, which is the maximum value among the code tables. Therefore, the cycle required to initialize all code tables is 32 cycles.

[0026] Table 1

[0027] Table 2

[0028] Optionally, the pre-quantization includes: The transform coefficients are pre-quantized by a quantization algorithm model, wherein the quantization algorithm model satisfies the following relationship:

[0029] In the formula, is the pre-quantization result, is the absolute value of the transformation coefficient, is the quantized scaling factor, is the quantization offset factor, is the quantization shift factor.

[0030] The quantization candidate values ​​corresponding to the transform coefficients are obtained by pre-quantizing the transform coefficients.

[0031] Optionally, the distortion calculation includes: The distortion calculation model is used to calculate the distortion of the quantization candidate values ​​obtained through pre-quantization, wherein the distortion calculation model satisfies the following relationship: ; In the formula, is the original transform coefficient, is the quantized index value, is the quantization state of the dependent quantization, Indicates the quantization step size.

[0032] In the above embodiment, the present invention continues to use the distortion calculation scheme of VTM18.0, and directly calculates the distortion of each quantization candidate value in the frequency domain. Since the distortion involves a large number of multiplication, addition and shift operations, in order to improve the hardware main frequency, this paper gradually decomposes these operations with large calculation amount, and calculates the distortion of each candidate value in the frequency domain through 4 clock cycles ( - ) to complete the distortion calculation.

[0033] Optionally, the quantization candidate value calculation includes: ; In the above embodiment, for the calculation of the quantization candidate value, although the calculation is relatively simple, one cycle ( ), but in order to synchronize it with the calculation of distortion and context index, registers are used to save it to 3 cycles ( - ) before synchronous output.

[0034] Optionally, the context index calculation includes: The average value of the candidate quantization values ​​obtained by pre-quantization of the transform coefficients is calculated, and the quantization result of the transform coefficients is predicted. The calculation method satisfies the following relationship:

[0035] The sum of the context templates is calculated, and the variable f(T) related to the context template in the context index is calculated. The calculation method satisfies the following relationship:

[0036]

[0037] The variable ctxOffset related to the diagonal position of the transform coefficient is calculated by the comparator, and the context index of each syntax element is obtained by calculating the sum of the variable f(T) and the variable ctxOffset.

[0038] In the above embodiment, the context index calculation adopts the algorithm proposed in this paper, so that it can be independent of the quantization results of adjacent transform coefficients in the context model, so that the pipeline module can proceed smoothly without conflict. Figure 4 shown. In the first stage, the average value of the four candidate quantization values ​​of the transform coefficient is calculated to predict the quantization result of the transform coefficient, and the result is saved in the shift register. The dark part in the figure is the register. The stage calculates the sum of context templates; The variable f(T) related to the context template in the context index is calculated in the stage. At the same time, the variable ctxOffset related to the diagonal position d where the coefficient is located is also calculated through the comparator in this stage; The context index of each syntax element can be obtained by calculating the sum of f(T) and ctxOffset in this stage. For the syntax element sig, the context index calculated in this stage does not include the part related to the quantization state Qstate, because the quantization state corresponding to each quantization candidate value is unknown at this stage, and Qstate will be considered in the bit rate estimation stage.

[0039] Figure 5 is the pipeline timing diagram, - is the transform coefficient in the quantization process, - It is the process of distortion calculation, quantization candidate value calculation and context index calculation. The pre-quantization is done in the stage. After obtaining the pre-quantization result, - The distortion and context index of each syntax element are calculated in the stage. Because this paper adopts an improved context index algorithm, when the quantized transform coefficients When The transformation coefficients can be obtained at this stage - The quantization result prediction value of , so the context index of each syntax element can be directly calculated without waiting for the previous transform coefficients to pass through their respective When quantizing the first 5 transform coefficients of a TB, since the preceding coefficients of each transform coefficient are less than 5 transform coefficients, the insufficient quantization result prediction value is filled with 0. After the stage is finished, the context index of each grammatical element can be obtained.

[0040] Optionally, the bit rate estimation includes: Decomposing the quantization candidate values ​​obtained by pre-quantization to obtain corresponding syntax elements, and querying the estimated bit rate corresponding to the syntax elements through the initialized predefined code table; The estimated code rate is added to the distortion of the corresponding quantization candidate value to obtain the rate-distortion cost of the quantization candidate value.

[0041] In the above embodiment, After the stage is completed, Figure 6 As shown, in The syntax elements obtained by decomposing the four quantization candidate values ​​are estimated by looking up the table in the stage to estimate the bit rate. The result of the rate estimation and the distortion of the corresponding quantization candidate value are added to obtain the rate distortion cost of each quantization candidate value. The state that each quantization candidate value finally corresponds to is unknown, and the corresponding relationship between the quantization candidate value and the quantization state is a 2-to-2 relationship. For example, The value of will be used as the state and In order to avoid being affected by the context index's dependence on Qstate, The stage will calculate the code rate of the two corresponding quantization states for each quantization candidate value, corresponding to Figure 6 , sigRateTable will output the rate lookup results for the two quantization states. In the same stage, each quantization candidate value will get 2 rate-distortion costs corresponding to the quantization state.

[0042] Optionally, the Viterbi calculation includes: The bit rate and distortion obtained after bit rate estimation are accumulated to obtain the cumulative rate-distortion cost, and the cumulative rate-distortion costs of multiple quantization candidate values ​​of each quantization coefficient are compared. Only the four groups of quantization candidate values ​​with the smallest rate-distortion cost are retained through the Viterbi algorithm. After X comparisons, four quantization paths are obtained for each TB block, where X is the number of quantization coefficients.

[0043] In the above embodiment, The pruning function of the Viterbi algorithm is realized in the stage, that is, the grid paths with large rate-distortion cost are removed by pruning, so as to find a path with the minimum rate-distortion cost in the grid graph. Figure 7 As shown, in After the calculation of the rate-distortion cost of each quantization candidate value is completed in the stage, a total of 8 rate-distortion costs will be generated for the 4 quantization candidate values, and then enter After the stage, the branches with larger cumulative rate distortion are pruned and the cumulative rate distortion is updated, and only four branches, i.e., four grid paths, are retained each time. These four grid paths are represented by four cumulative rate distortions and four state indices.

[0044] Pipeline module After the stage is completed, the information of the four states corresponding to each coefficient will be saved in the grid memory, and then a set of quantization indexes with the lowest rate distortion cost will be selected for output. When there is only one grid memory, the quantization of the next TB cannot be performed during the output of the optimal path, because the quantization result of the next TB entering the grid memory will affect the normal output of the current TB. Figure 8 As shown, even if the current pipeline module is idle, After initialization, the first transform coefficient cannot enter the pipeline module immediately because Occupies the only grid storage structure, If the quantized data is stored in the grid memory, it will interfere with the quantized data. Output. Therefore, Can only wait until The output can enter the pipeline module only after it is completed, which will reduce the utilization rate of the pipeline module.

[0045] In order to solve the above problems, this paper introduces a grid storage. After the introduction, there is no need to wait Output completed, To start quantification, just The quantized grid data is stored in another grid memory, and the two grid memories are repeatedly read through ping-pong operation to improve the utilization rate of the hardware pipeline module. Fig. 9 As shown, after introducing another grid storage structure, the reduction in pipeline utilization caused by the gap is eliminated compared to the single grid storage structure.

[0046] The hardware structure diagram of introducing two grid memories is as follows Fig.10As shown. There are two Trellis units in the figure, namely Trellis0 and Trellis1. The Trellis unit consists of a register Rdcost_reg for storing the cumulative rate distortion cost and a ram for storing the quantization state and quantization index. The ping-pong operation is controlled by the Controller module. In the stage of inputting the Trellis unit, the Controller module will input the address information to the Trellis_ram module in the quantization order. Every time the 4 state information of a transform coefficient is input, the cumulative rate distortion cost will be updated once. The cumulative rate distortion cost corresponding to the 4 quantization states of the last transform coefficient is the total rate distortion cost of the 4 grid paths. In the output stage, the comparator Cmp will compare the total rate distortion cost of the 4 grid paths. The comparator will output the state index of the grid path with the smallest rate distortion cost. The Dec module can reversely push out an entire grid path through the state index, and finally control the Sel module through the Controller to output the optimal quantization result.

[0047] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A VVC dependent quantization hardware pipeline architecture design method, characterized in that: include: The hardware pipeline architecture that relies on quantization is divided into three stages: initialization stage, pipeline quantization stage, and optimal quantization path output stage. Introducing a predefined code table in an initialization phase, initializing the predefined code table in the initialization phase to obtain an initialized predefined code table; In the pipeline quantization stage, an 8-stage pipeline unit is used to input the transform coefficients into the 8-stage pipeline unit in the order of anti-diagonal scanning. The transform coefficients are pre-quantized, distortion calculated, quantization candidate value calculated, context index calculated, bit rate estimated, and Viterbi calculated by the 8-stage pipeline unit to finally obtain 4 quantization state information; The first grid memory and the second grid memory are introduced in the output stage of the optimal quantization path. The quantization state information calculated by different TB blocks is saved by the first grid memory and the second grid memory. The quantization throughput can be maximized by performing ping-pong storage operations through the two grid memories. After all the coefficients of a TB block are quantized through the pipeline quantization stage, the next TB block can be quantized without waiting for the quantization result output to be completed.

2. The VVC dependent quantization hardware pipeline architecture design method according to claim 1, characterized in that: The initializing the predefined code table to obtain an initialized predefined code table includes: Acquire code table data and syntax elements in a predefined code table, determine a maximum size according to the code table data, and determine a clock cycle required for initializing the code table data according to the maximum size; After the code table data is initialized based on the clock cycle, the initialized code table data is matched with the syntax element according to a matching rule of the predefined code table to obtain an initialized predefined code table.

3. The VVC dependent quantization hardware pipeline architecture design method according to claim 1, characterized in that: The pre-quantization comprises: The transform coefficients are pre-quantized by a quantization algorithm model, wherein the quantization algorithm model satisfies the following relationship: In the formula, is the pre-quantization result, is the absolute value of the transformation coefficient, is the quantized scaling factor, is the quantization offset factor, is the quantization shift factor The quantization candidate values ​​corresponding to the transform coefficients are obtained by pre-quantizing the transform coefficients.

4. The VVC dependent quantization hardware pipeline architecture design method according to claim 1, characterized in that: The distortion calculation includes: The distortion calculation model is used to calculate the distortion of the quantization candidate values ​​obtained through pre-quantization, wherein the distortion calculation model satisfies the following relationship: ; In the formula, is the original transform coefficient, is the quantized index value, is the quantization state of the dependent quantization, represents the quantization step size, sgn(x) is the sign function, Indicates distortion.

5. The VVC dependent quantization hardware pipeline architecture design method according to claim 1, characterized in that: The quantization candidate value calculation includes: ; In the formula, represents the quantized candidate value, Represents the pre-quantized value.

6. The VVC dependent quantization hardware pipeline architecture design method according to claim 1, characterized in that: The context index calculation includes: The average value of the candidate quantization values ​​obtained by pre-quantization of the transform coefficients is calculated, and the quantization result of the transform coefficients is predicted. The calculation method satisfies the following relationship: ; In the formula, represents the average value of the quantized candidate values, Indicates the quantization candidate value 0, represents the quantization candidate value 1, represents the quantization candidate value 2, represents the quantization candidate value 3; The sum of the context templates is calculated, and the variable f(T) related to the context template in the context index is calculated. The calculation method satisfies the following relationship: ; ; In the formula, , Represents an intermediate result that depends on the context template, Represents a context template, represents the partial reconstruction quantization index; The variable ctxOffset related to the diagonal position of the transform coefficient is calculated by the comparator, and the context index of each syntax element is obtained by calculating the sum of the variable f(T) and the variable ctxOffset.

7. The VVC dependent quantization hardware pipeline architecture design method according to claim 1, characterized in that: The bit rate estimation includes: Decomposing the quantization candidate values ​​obtained by pre-quantization to obtain corresponding syntax elements, and querying the estimated bit rate corresponding to the syntax elements through the initialized predefined code table; ; ; In the formula, is the quantized index value, represents the partially reconstructed quantization index value, , , , , rem represents the syntax element after the transformation coefficient is decomposed; The estimated code rate is added to the distortion of the corresponding quantization candidate value to obtain the rate-distortion cost of the quantization candidate value.

8. The VVC dependent quantization hardware pipeline architecture design method according to claim 1, characterized in that: The Viterbi calculation includes: The bit rate and distortion obtained after bit rate estimation are accumulated to obtain the cumulative rate-distortion cost, and the cumulative rate-distortion costs of multiple quantization candidate values ​​of each quantization coefficient are compared. Only the four groups of quantization candidate values ​​with the smallest rate-distortion cost are retained through the Viterbi algorithm. After X comparisons, four quantization paths are obtained for each TB block, where X is the number of quantization coefficients.

Citation Information

Cited By

  • Code rate estimation device and method, video encoder, electronic equipment, storage medium and computer program product

    CN121309820A