Efficient hardware architecture system of low-frequency inseparable quadratic transformation multiplier

By optimizing the combined design of the control module, selection module, core module, accumulation module and rounding module of the hardware architecture, the problems of low utilization rate of multiplier and waste of hardware resources in low-frequency inseparable quadratic transformation are solved, and efficient matrix multiplication operation and real-time encoding capabilities are achieved.

CN120371258APending Publication Date: 2025-07-25上海芯开技术有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510516834.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing hardware implements the multiplier of medium and low frequency inseparable quadratic transformation with low frequency, has low utilization, has invalid multiplication operation, high hardware area and power consumption, and the processing time is inconsistent when switching different transformation block sizes, which affects the real-time processing capability of the encoding system.

Method used

The combination design of control module, selection module, core module, accumulation module and rounding module is adopted. By processing matrix multiplication operations of 8×16 sub-regions every cycle, a coefficient approximation scheme is introduced, the hardware architecture is optimized, and a fusion calculation unit is used to replace part of the multiplier, sharing data input and constant input, and optimizing the multiplier utilization and calculation efficiency.

Benefits of technology

It improves the utilization rate of multiplier, reduces hardware resource waste and power consumption, improves computing efficiency and throughput, meets the needs of high-resolution real-time coding, and reduces hardware area and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371258A_ABST
    Figure CN120371258A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video coding, and discloses a low-frequency inseparable quadratic transformation multiplier efficient hardware architecture system, which comprises a control module, a selection module, a core module, an accumulation module and a rounding module, and is characterized in that the control module generates a sub-region index and a constant selection signal; the selection module extracts corresponding transformation constant input; the core module executes matrix multiplication operation on one 8 * 16 sub-region in each processing period, and each column of multipliers share the same group of input data; the accumulation module accumulates the multiplication result in the row direction; and the rounding module carries out rounding processing on the accumulation result and outputs a final conversion result. Through the measures of balancing the utilization rate of the multiplier under different transformation unit sizes, reducing invalid multiplication operation, introducing a coefficient approximation scheme to optimize a hardware circuit structure and the like, the calculation efficiency is improved, the hardware area and power consumption are reduced, and the high-resolution real-time coding requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video coding, and specifically to a multiplier-efficient hardware architecture system for low-frequency non-separable second-order transformation. Background Art

[0002] With the continuous increase in video resolution and frame rate, the amount of original video data has expanded rapidly, posing more stringent requirements on the performance of the coding system. As a new generation of international video compression standard, multi-functional video coding achieves significant bitrate compression while maintaining image quality, becoming an important foundation for ultra-high-definition video transmission. In this standard, to further improve the compression efficiency, a variety of enhanced tools are introduced, among which the low-frequency non-separable second-order transformation is widely used to optimize the residual signal structure after the main transformation, effectively improving the energy concentration.

[0003] This transformation process further removes spatial redundancy by performing additional transformation steps on the low-frequency coefficient region after the main transformation. Since the non-separable structure involves large-scale matrix multiplication, its hardware implementation poses new challenges to resource scheduling and computational efficiency. Especially when performing transformation operations on small-sized blocks, the system is prone to uneven resource allocation, affecting the overall efficiency.

[0004] Specifically, most of the current mainstream hardware implementation architectures continue the traditional row-based multiplication processing unit in terms of structure design. When processing small blocks such as 4×4 or 8×8, some hardware units are idle, resulting in resource waste. Further, to reduce the hardware complexity, a strategy of reducing the second-order transformation is proposed in the coding standard to reduce the required transformation scale. However, the existing circuit designs do not effectively adapt to this strategy and still perform fixed-scale calculations according to a unified structure, resulting in redundancy and inefficiency in some transformation processes, and thus the utilization rate of multipliers in the existing designs is relatively low. In terms of hardware resources, multiple current implementations still rely on a large number of general-purpose multipliers for matrix operations. Such multipliers have a high cost in terms of area and power consumption. Especially under the goal of high throughput rate, a large number of resources need to be stacked, further pushing up the implementation threshold. Although some solutions attempt to reduce the no-load power consumption through gating technology, the problems of complex overall structure and low reuse rate are still prominent.

[0005] The existing hardware implementations also face the bottleneck problem of system throughput. In video processing applications with multiple resolutions and multiple modes, the transformation module needs to flexibly adapt to different input sizes. However, the existing solutions fail to ensure the consistency of processing time when switching between different transformation block sizes. Especially in the case of smaller transformation sizes, the accumulated delay affects the overall pipeline efficiency, thereby limiting the real-time processing ability of the coding system. Summary of the Invention

[0006] In view of the deficiencies of the prior art, the present invention provides an efficient hardware architecture system for a multiplier of low-frequency inseparable quadratic transformation, which solves the problems of low utilization rate of multipliers, existence of invalid multiplication operations, high hardware area and high power consumption in the hardware architecture of low-frequency inseparable quadratic transformation.

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: An efficient hardware architecture system for a multiplier of low-frequency inseparable quadratic transformation, including: A control module, configured to output a constant selection signal according to the sub-region index generated by the size of the transformation kernel, the transformation subset index, the transformation kernel index, and the current cycle count; A selection module, connected to the control module, for extracting the constant input corresponding to the transformation kernel sub-region from a predefined constant set based on the constant selection signal; A core module, connected to the selection module, for performing matrix multiplication operations on an 8×16 sub-region in each processing cycle. Each multiplier has a data input terminal and a constant input terminal, and multiple multipliers in each column share the same set of input data; An accumulation module, for summing the multiplication results output by the core module in the row direction to generate an intermediate transformation coefficient value; A rounding module, for performing rounding processing on the intermediate transformation coefficient value to output the final transformation result.

[0008] Preferably, the control module includes: An index unit, configured to control the selection module to select the transformation constants of the corresponding sub-region, and control the core module to perform matrix multiplication operations on the corresponding sub-region; Preferably, the selection module includes: A constant selection unit, for selecting the transformation constants required for the current processing sub-region from the constant set according to the constant selection signal from the control module and sending them to the core module.

[0009] Preferably, the core module includes a plurality of multipliers arranged in an 8-row and 16-column structure. Each column of multipliers shares a set of input data, and each multiplier is connected to a multiplexer, which is used to select the corresponding transformation constant as the input from multiple constant channels. The number of constant channels is 64, and the selection control signal is generated by the selection module according to the size of the transformation kernel, the transformation subset index, the transformation kernel type, and the sub-region index.

[0010] Preferably, the accumulation module includes an adder tree structure distributed in each row. The adder tree is composed of multiple two-input adders cascaded, and is used to compress 16 multiplication results in each row into a single output value.

[0011] Preferably, the rounding module includes: An offset adder for adding an intermediate transform coefficient to a preset rounding offset; A shift unit for shifting the output of the adder to the right by a certain number of bits to achieve fixed-point precision conversion.

[0012] Preferably, the multiplier in the core module is a configurable structure. In a scenario with limited area, the multiplier can be implemented by a fused computing unit as an alternative. The fused computing unit includes shift logic, addition logic, and selection logic, and is used to convert the multiplication operation between a fixed constant and input data into a combined operation of multiple shift operations and at most one addition operation, so as to realize the equivalent approximate calculation of the time-multiplexed multi-constant multiplication operation.

[0013] Preferably, the constants used for multiplication approximation in the fused computing unit come from a set of constants after approximate reconstruction, including: The first set R0, including all constants that can be generated by shifting the value 1 to the left by any number of bits; The second set R1, including constants generated by performing one addition or subtraction on two constants in R0; The third set R2, including constants generated by performing one addition or subtraction on two constants in R0 or R1; The fourth set R3, including constants generated by performing one addition or subtraction on two constants in R0, R1, or R2; Among them, if the original constant does not belong to R0 or R1, it is replaced by the constant in R0 or R1 with the smallest difference from it, and the number of adders required for the operation after replacement does not exceed one. If there are multiple candidate replacement constants, the one with a smaller value is preferably selected.

[0014] Preferably, the control signal output by the control module is determined according to the size of the transform kernel, the transform subset index, the transform kernel index, and the sub-region index generated by the current cycle count. The control signal is used for the selection module to complete constant loading and input the corresponding constant into the multiplier group of the corresponding sub-region of the core module.

[0015] The present invention provides a high-efficiency hardware architecture system for multipliers of low-frequency non-separable quadratic transforms. It has the following beneficial effects: 1. The present invention adopts a matrix multiplication operation method of processing 8×16 sub-regions per cycle, which balances the multiplier utilization rate under different transform unit sizes and avoids the idle state of the multiplier. Compared with the problem of low multiplier resource utilization rate caused by the row-based processing unit in the prior art, it solves the deficiency of resource waste.

[0016] 2. By analyzing the reduction factor in the secondary transform, the present invention optimizes matrix multiplication operations, avoids invalid calculations when processing 4×4 and 8×8 transform blocks, and improves the computational efficiency and reduces the waste of hardware resources compared with the existing design that does not fully utilize the reduction secondary transform strategy, resulting in unnecessary computational overhead.

[0017] 3. The present invention introduces a coefficient approximation scheme, replaces the general multiplier with a fused circuit composed of an adder, a shifter, and a multiplexer, and limits the number of adders to one, achieving a significant reduction in hardware area and power consumption compared with the prior art that uses a general multiplier, resulting in a higher hardware area and power consumption.

[0018] 4. By optimizing the design of the core unit, the present invention reduces the cycle for processing 4×4 transform blocks to one clock cycle while maintaining the same total number of multipliers as the prior art. Compared with the problem of longer processing cycles due to uneven utilization of multipliers in the prior art, the throughput is significantly improved, meeting the requirements of high-resolution real-time coding.

[0019] 5. By introducing a coefficient approximation scheme, the present invention optimizes the computational complexity of the transform coefficients. Different from the traditional method that directly uses exact coefficients for calculation, this scheme reduces the computational complexity while ensuring that the coding efficiency is hardly affected. Therefore, the goal of reducing the computational overhead while maintaining high coding quality is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a framework diagram of the system of the present invention; Figure 2 is the hardware architecture of the low-frequency non-separable secondary transform proposed by the present invention; Figure 3 is a timing diagram of the reference literature and the design proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] Please refer to the appended Figure 1 - appended Figure 3 , the embodiments of the present invention provide a multiplier-efficient hardware architecture system for low-frequency non-separable secondary transform, including: A control module for outputting a constant selection signal according to the sub-region index generated by the size of the transform kernel, the transform subset index, the transform kernel index, and the current cycle count; The present invention relates to a control module, whose main function is to output a constant selection signal according to the sub-region index generated by the size of the transform kernel, the transform subset index, the transform kernel index, and the current cycle count.

[0023] In this embodiment, the main task of the index unit is to calculate the index of the sub-region based on the size of the transform kernel, the transform subset index, the transform kernel index, and the cycle count. Specifically, the sub-region index is determined by the following factors: Size of the transform kernel: The size of the matrix is usually , corresponding to width and height respectively.

[0024] Transform subset index: The transform subset index implicitly inferred according to the prediction mode.

[0025] Transform kernel index: The index indicating a specific part of the transform kernel.

[0026] Current cycle count: A certain transform core within the transform subset.

[0027] Based on the above parameters, the index unit calculates the selection signal value, and its formula is: sel = (lfnst_8x8_flg == 0)? (4 * lfnstSetIdx + 2 * lfnstIdx + subRegionIdx) : (16 + 12 * lfnstSetIdx + 6 * lfnstIdx + subRegionIdx); Among them, lfnst_8x8_flg is implicitly inferred from the size of the transform kernel, lfnstSetIdx represents the transform subset index, lfnstIdx represents the transform kernel index, and subRegionIdx represents the sub-region index generated by the current cycle count.

[0028] The output of these control signals ensures that the selection module and the core module can work precisely and coordinately. The selection module loads the correct constants according to the control signals, while the core module activates the corresponding multiplier group according to the control signals.

[0029] Through the design of the control module in this embodiment, the calculation efficiency of matrix multiplication can be effectively improved, unnecessary calculations can be avoided, and power consumption can be reduced.

[0030] A selection module, connected to the control module, is used to extract the constant input corresponding to the transform kernel sub-region from a predefined constant set based on the constant selection signal; In this embodiment, the selection module is connected to the control module and is responsible for extracting the constant input corresponding to the transformed kernel sub-region from the predefined constant set based on the constant selection signal from the control module and transmitting it to the core module. The selection module plays a crucial role in the matrix transformation process, ensuring the provision of accurate constant input in each cycle, thereby guaranteeing the accuracy and efficiency of matrix multiplication. Through this design, the entire transformation process can maintain high efficiency while avoiding unnecessary calculations and resource waste.

[0031] The constant selection unit is another important part of the selection module and is responsible for selecting the transformation constants required for the current sub-region from the constant set according to the sub-region index signal transmitted by the control module and transmitting it to the core module.

[0032] The working principle of the control module is based on signals such as cycle count, transformation kernel index, and sub-region index to determine the transformed kernel sub-region being processed currently. Assume that the currently processed transformation kernel is , and its size is , and the transformation subset index lfnstSetIdx of the sub-region being processed. The control module generates a constant selection signal based on the cycle count and sub-region index. The constant selection unit extracts constants from the constant memory according to these signals. The constant selection signal calculates the offset of the current sub-region according to the following formula: sel = (lfnst_8x8_flg == 0)? (4 * lfnstSetIdx + 2 * lfnstIdx + subRegionIdx) : (16 + 12 * lfnstSetIdx + 6 * lfnstIdx + subRegionIdx); where, lfnst_8x8_flg is implicitly inferred from the size of the transformation kernel, lfnstSetIdx represents the transformation subset index, lfnstIdx represents the transformation kernel index, and subRegionIdx represents the sub-region index generated by the current cycle count.

[0033] The constant selection unit uses an efficient multiplexer (MUX) structure to select the correct constant from the constant set according to the constant selection signal output by the control module.

[0034] In this embodiment, through the constant selection mechanism, the selection module can ensure the transmission of accurate constant input in each cycle and coordinate the work of the control module and the core module. Through the optimized constant selection unit, the present invention can significantly improve the extraction speed of constant data, ensure the correct execution of the calculation for each sub-region in the transformation kernel, and thus improve the accuracy and performance of matrix operations.

[0035] Generally speaking, the selection module plays an important role in the matrix transformation process. By optimizing the constant selection and storage mechanism, the present invention can provide efficient and accurate matrix multiplication operations, while reducing the consumption of computing resources and improving the overall system performance.

[0036] The core module, connected to the selection module, is used to perform matrix multiplication operations on an 8×16 sub-region in each processing cycle. Each multiplier has a data input terminal and a constant input terminal, and multiple multipliers in each column share the same set of input data; In this embodiment, the core module is connected to the selection module, and its main function is to perform matrix multiplication operations on the 8×16 sub-region in each processing cycle. The core module includes multiple multipliers, and improves the computing efficiency by sharing data input and constant input. Each multiplier selects the corresponding transformation constant from multiple constant channels through a multiplexer (MUX), and the selection of these constants is dynamically generated by the selection module through control signals. In addition, in order to meet the application requirements with limited area, this embodiment also proposes a scheme of using fused computing units to replace some multipliers. The fused computing unit approximates the multiplication calculation through the combination of shift operations and addition operation multiplexers, thereby effectively reducing the occupation of hardware resources.

[0037] In this embodiment, the core module includes multiple multiplier arrays arranged in an 8-row and 16-column structure. The multipliers in each column share the same set of input data, while the constant input of each multiplier is dynamically selected by a multiplexer (MUX). This design can efficiently process the data in matrix multiplication and accelerate the calculation by parallel processing multiple data channels.

[0038] Specifically, the input data of each multiplier comes from the input terminal of the core module, and the input data can be expressed as , ,…, (each data input terminal processes 16 data points). These data are processed in parallel by different multipliers, and the constant input comes from the constant set of the selection module, expressed as , ,…, (a total of 128 constants). Through the multiplexer, the multiplier can select the appropriate constant for operation, ensuring that each multiplier can dynamically select the required constant according to the output of the control module.

[0039] In the specific implementation, the constant selection signal is transmitted to the selection module, and the corresponding constant input is selected through the multiplexer.

[0040] The multiplier in the core module is connected to the constant selection unit through a multiplexer, and each multiplier selects the required constant input. The selection module generates a constant selection signal according to the control signal, and the control signal is based on information such as the current transform unit size, the current transform kernel type, the cycle count, and the sub-region index. This signal ultimately determines which constant the multiplexer selects from the constant channels.

[0041] In a possible implementation, the selection signal generation process is as follows: sel = (lfnst_8x8_flg == 0)? (4 * lfnstSetIdx + 2 * lfnstIdx + subRegionIdx) : (16 + 12 * lfnstSetIdx + 6 * lfnstIdx + subRegionIdx); Among them, lfnst_8x8_flg is implicitly inferred from the size of the transform kernel, lfnstSetIdx represents the transform subset index, lfnstIdx represents the transform kernel index, and subRegionIdx represents the sub-region index generated by the current cycle count.

[0042] To adapt to the hardware environment with limited area, this embodiment proposes to replace the general multiplier with a fused computing unit. The fused computing unit combines a set of equivalent operations through shift operations, addition operations, and constant selection logic to approximate the multiplication operation.

[0043] Specifically, in the low-frequency non-separable quadratic transform, each multiplier is used to perform the multiplication operation of the input signal x and the constant c, where c is a constant selected from a predefined set. In different clock cycles, c may be different.

[0044] To reduce resource consumption, we time-multiplex the multi-constant multiplication algorithm and convert the multiplication operation into a combinational circuit composed of adders, shifters, and multiplexers. In addition, to further reduce the number of adders in the fused circuit, we propose a coefficient approximation scheme: replacing the constant that requires more adders with a neighboring constant that requires fewer adders.

[0045] First, the relevant definitions are given: The first set R0 includes all constants that can be generated by shifting the value 1 by any number of bits to the left; The second set R1 includes the constants generated by performing an addition or subtraction operation once on two constants in R0; The third set R2 includes the constants generated by performing an addition or subtraction operation once on two constants in R0 or R1; The fourth set R3 includes the constants generated by performing an addition or subtraction operation once on two constants in R0, R1, or R2; We analyzed the coverage rates of all constants in LFNST in the sets of R0, R1, R2, and R3. The results show that: The coverage rate of R0 is 49.38% The coverage rate of R1 is 32.87% The coverage rate of R2 is 16.75% The coverage rate of R3 is only 1% Although using only the constants in the R0 set can completely eliminate the adder, its limited coverage rate will significantly reduce the performance. While using the constants in the R0, R1, and R2 sets simultaneously can achieve a high coverage rate, two adders are still required, thus limiting the area optimization space. Therefore, this design proposes to approximately replace the constants in the R2 and R3 sets with the constants in the R0 and R1 sets, thereby limiting the number of adders in the fusion circuit to only one.

[0046] Specifically, for a constant c belonging to the set C (i.e., C = {c|c ∈ R3 ∪ R2}), we replace it with the number n in R0 and R1 that is closest to c and minimize |c - n|. If there are multiple n that satisfy the condition, the one with the smallest numerical value is selected.

[0047] The fusion computing unit is particularly suitable for scenarios with limited hardware resources and can effectively replace the multiplier for approximate calculations. For example, in some low-power or area-limited applications, using the fusion computing unit can significantly reduce the hardware overhead. Specifically, the fusion computing unit can perform an equivalent approximation of the multiplication operation between the constant and the input data that originally requires a multiplier through shift and addition operations, thereby reducing the demand for traditional multiplication hardware.

[0048] Multiplier sharing of data input: In this embodiment, multiple multipliers are arranged in columns so that the multipliers in each column share the same set of input data. This structure can effectively reduce the redundancy of input data and improve the computing efficiency of the system.

[0049] Constant selection module and multiplexer: The close cooperation between the selection module and the multiplexer ensures that in each cycle, the multiplier can accurately select the appropriate constant from the constant channel for calculation. The selection module generates a constant selection signal according to the control signal, thereby guiding the multiplexer to select the correct constant.

[0050] Implementation of the fusion computing unit: For the area-limited scenario, some multipliers in the core module are replaced by the fusion computing unit, and these computing units convert the constant multiplication into an equivalent calculation through shift and addition operations. The use of specific adders is controlled within one to ensure the effective utilization of hardware resources.

[0051] Through the above optimization design, the core module can execute matrix multiplication efficiently and flexibly adjust hardware resources according to actual requirements. By sharing input data and selecting constant inputs, the core module improves data reuse during the calculation process and reduces unnecessary computations. The introduction of the fused computing unit enables efficient completion of multiplication operations even in scenarios with limited hardware resources, significantly reducing area and power consumption. At the same time, constant reconstruction and approximation techniques ensure calculation accuracy and performance, reduce the use of adders, and further optimize the system's hardware resources.

[0052] An accumulation module, configured to sum the multiplication results output by the core module in the row direction to generate intermediate transformation coefficient values; In this embodiment, the accumulation module is used to sum the multiplication results output by the core module in the row direction, thereby generating intermediate transformation coefficient values. The main function of the accumulation module is to efficiently process multiple multiplication results and compress the multiplication results of each row into a single output value through an adder tree structure for subsequent processing.

[0053] Generally, the accumulation module consists of an adder tree structure distributed in each row. The adder tree in each row is formed by cascading multiple two-input adders, which are used to gradually accumulate the 16 multiplication results (i.e., input data) in each row and finally output a single result value.

[0054] Specifically, the hierarchical structure of the adder tree is composed of multiple two-input adders. The input end of each two-input adder is connected to the output of the previous-level adder or the original multiplication result, and the output of each level is passed to the next layer step by step until the accumulation result of the row is finally formed. The design of the adder tree adopts a hierarchical method, which can effectively reduce the demand for computing resources and avoid the computing bottleneck that may be caused when a single adder processes too many inputs.

[0055] For example, assume there are 16 multiplication results in a row . The first layer of the adder tree divides the 16 results into two groups, with 8 in each group, and then the second layer divides the 8 results into two groups, with 4 in each group. Through such a structure, the accumulation task can be completed only after 4 to 5 layers of adder calculations. The output of each level of the adder will finally be combined into the final accumulation result of the row.

[0056] By summing the 16 multiplication results in each row, the accumulation module finally outputs a value as the accumulation result of the row. These results will then be used as the input of the subsequent processing module to ensure the continuity and consistency of the entire system during the data processing process.

[0057] A rounding module, configured to perform rounding processing on the intermediate transformation coefficient values to output the final transformation results.

[0058] In this embodiment, the rounding module is used to perform rounding processing on the intermediate transformation coefficient values generated in the foregoing steps to output the final transformation result. The rounding module consists of an offset adder and a shift unit.

[0059] Experiment 1: Purpose of the experiment: To evaluate the impact of the proposed coefficient approximation scheme on video coding performance. By comparing the changes in coding efficiency with and without enabling the low-frequency non-separable quadratic transformation, the effectiveness of the coefficient approximation scheme is analyzed.

[0060] Experimental steps: Preparation of the experimental environment: Use the multi-functional video coding (VVC) reference software VTM-19.2 for coding tests.

[0061] Select multiple test video sequences to ensure coverage of different scenarios and complexities.

[0062] Configuration of coding parameters: Set the quantization parameter (QP) values to 22, 27, 32, and 37, corresponding to different coding qualities respectively.

[0063] During the coding process, disable the low-frequency non-separable quadratic transformation and introduce coefficient approximation in the low-frequency non-separable quadratic transformation respectively, and record the coding performance in both cases.

[0064] Collection of performance metrics: Calculate and record the bitrate and peak signal-to-noise ratio (PSNR) for each configuration.

[0065] Calculate coding efficiency metrics such as BD-rate loss to evaluate the performance changes after enabling the coefficient approximation scheme.

[0066] Data analysis: Compare the differences in coding efficiency and video quality between disabling the low-frequency non-separable quadratic transformation and introducing coefficient approximation in the low-frequency non-separable quadratic transformation.

[0067] Analyze the impact of different QP values on coding performance and evaluate the effectiveness of the coefficient approximation scheme under different quality settings.

[0068] Experimental tabular data: Table 1 Performance of the proposed coefficient approximation scheme Summary: In Experiment 1, we conducted coding tests on multiple video sequences to evaluate the impact of the proposed coefficient approximation scheme on coding efficiency.

[0069] The results show that the BD-rate loss introduced by directly disabling the low-frequency non-separable quadratic transform is 1.38% under the all-intra configuration and 0.96% under the random access configuration. In contrast, the average BD-rate loss of the coefficient approximation scheme is 0.25% under the all-intra configuration and only 0.16% under the random access configuration. This loss is almost negligible, verifying the effectiveness of the coefficient approximation scheme in significantly optimizing hardware resources while maintaining coding efficiency.

[0070] Experiment 2: Comparing the performance of different hardware architectures for low-frequency non-separable quadratic transforms Purpose: To evaluate the performance of the proposed hardware architecture in terms of throughput, power consumption, area, etc., and compare it with the existing technologies.

[0071] Method: Experimental setup: The proposed design was implemented in RTL and synthesized using DesignCompiler based on different process libraries.

[0072] The following hardware architectures were selected for comparison: The design in [1]: I. Farhat, W. Hamidouche, A. Grill, D. M´enard, and O. D´eforges, “Lightweight hardware transform design for the versatile video coding 4k asic decoders,” IEEE Transactions on Consumer Electronics, vol. 67, no. 4, pp. 329–340, 2021. The design in [2]: J. Goebel, V. Costa, L. Agostini, B. Zatt, and M. Porto, “A high throughput design for the h.266 / vvc low-frequency non-separable transform,” in 2022 IEEE International Symposium on Circuits and Systems (ISCAS), 2022, pp. 1798–1802. Design in [3]: J. Goebel, L. Agostini, B. Zatt, and M. Porto, “Low-frequency non-separable transform hardware system design for the vvc encoder,” in 2022 35th SBC / SB Micro / IEEE / ACM Symposium on Integrated Circuits and Systems Design (SBCCI), 2022, pp. 1–6. This design (including the coefficient approximation scheme and the original design).

[0073] Performance evaluation: Compare the following metrics: Process node (nm), operating frequency (MHz), number of multipliers, pixel parallelism, throughput rate, storage requirement (bit), number of equivalent two-input NAND gates (K), number of logic gates in the core unit (K), power consumption (mW), normalized throughput rate, normalized area.

[0074] Experimental process: Implement the proposed architecture and compare it with the existing design under the same process node.

[0075] Conduct functional verification to ensure that each architecture correctly implements the low-frequency non-separable quadratic transform function.

[0076] Measure and record the above performance metrics.

[0077] Experimental tabular data: Table 2 Comparison of hardware performance of different low-frequency non-separable quadratic transforms Summary: In Experiment 2, we compared the performance of different hardware architectures for low-frequency non-separable quadratic transforms, focusing on metrics such as process node, operating frequency, number of multipliers, pixel parallelism, throughput rate, storage requirement, number of equivalent two-input NAND gates, number of logic gates in the core unit, power consumption, normalized throughput rate, and normalized area.

[0078] As shown in the table results: The number of multipliers required for this design is 128, and the storage requirement is 0 bit. At the 28nm process node, at a frequency of 200MHz, the area of the overall architecture is 92.4K gates, and the power consumption is 13.55mW. It is worth noting that the maximum operating frequency of this design can reach 901MHz. At this time, it can support ultra-high frame rate encoding of 7680×4320@289fps, with a corresponding area of 184.9K gates and a power consumption of 160.65mW. For fair comparison, normalized area (area / throughput) and normalized throughput (throughput / area) are introduced as evaluation criteria.

[0079] Compared with the designs in references [1], [2], and [3], this design shows advantages in many aspects. For example, the design in reference [1] achieved a throughput of 4K@193fps at the 28nm process, but did not provide power consumption data; the design in reference [2] achieved a throughput of 4K@60fps at the 40nm process, with a power consumption of 38.5mW; the design in reference [3] also achieved a throughput of 4K@60fps at the 40nm process, with a power consumption of 32.22mW.

[0080] In contrast, this design achieved a higher throughput at the same or lower power consumption, and had a smaller normalized area, reflecting the efficient use of hardware resources. Compared with the state-of-the-art solution in reference [2], this design achieved a 72% reduction in normalized area and a 3.0-fold increase in normalized throughput. This design (with a coefficient approximation scheme) achieved a throughput of 8K@64fps at the 28nm process node with a lower power consumption (12.45mW), and the normalized area was 7.3, showing excellent performance and area efficiency. These data indicate that the proposed scheme reduces the complexity and cost of hardware implementation while maintaining high performance, and has high practical application value.

[0081] Working principle: First, the control module outputs a constant selection signal according to the sub-region index generated by the size of the transform kernel, the transform subset index, the transform kernel index, and the current cycle count. These signals help the selection module extract the required constants from the preset constant set.

[0082] Next, the selection module passes these constants to the core module. There are multiple multipliers inside the core module, which process data in parallel. The multiple multipliers in each column share the same input, optimizing the use of computing resources. During the calculation process, the multipliers obtain constants from the constant channel through a multiplexer to ensure accurate and efficient calculation for each sub-region.

[0083] The multiplication results are processed by the accumulation module. The adder tree structure compresses the results of each row into one output.

[0084] The rounding module rounds the accumulated result and outputs the final transformation result.

[0085] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An efficient hardware architecture system for a multiplier of low-frequency inseparable quadratic transformation, characterized in that, Comprising: A control module for outputting a constant selection signal according to a sub-region index generated based on the size of a transform kernel, a transform subset index, a transform kernel index, and a current cycle count; A selection module connected to the control module for extracting a constant input corresponding to a transform kernel sub-region from a predefined constant set based on the constant selection signal; A core module connected to the selection module for performing a matrix multiplication operation on an 8×16 sub-region in each processing cycle. Each multiplier has a data input terminal and a constant input terminal, and multiple multipliers in each column share the same set of input data; An accumulation module for summing the multiplication results output by the core module in the row direction to generate an intermediate transform coefficient value; A rounding module for rounding the intermediate transform coefficient value to output a final transform result.

2. The high-efficiency hardware architecture system of the multiplier with low-frequency inseparable secondary transformation according to claim 1, characterized in that, The control module includes: An index unit for controlling the selection module to select transform constants for a corresponding sub-region and controlling the core module to perform a matrix multiplication operation on the corresponding sub-region.

3. The high-efficiency hardware architecture system of the multiplier with low-frequency inseparable secondary transformation according to claim 1, characterized in that The selection module includes: A constant selection unit for selecting transform constants required for the current processing sub-region from the constant set according to a sub-region index signal from the control module and sending them to the core module.

4. The multiplier high-efficiency hardware architecture system with non-separable second transformation at low frequency according to claim 1, characterized in that The core module includes a plurality of multipliers arranged in an 8-row and 16-column structure. Each column of multipliers shares a set of input data, and each multiplier is connected to a multiplexer for selecting a corresponding transform constant as input from multiple constant channels. The number of constant channels is 64, and the selection control signal is generated by the selection module according to the size of the transform kernel, the transform subset index, the transform kernel type, and the sub-region index.

5. The high-efficiency hardware architecture system of a multiplier with non-separable second transformation at low frequency according to claim 1, characterized in that, The accumulation module includes an adder tree structure distributed in each row. The adder tree is formed by cascading a plurality of two-input adders for compressing 16 multiplication results in each row into a single output value.

6. The high-efficiency hardware architecture system of a multiplier with non-separable second transformation at low frequency according to claim 1, characterized in that, The rounding module includes: An offset adder for adding an intermediate transform coefficient to a preset rounding offset; A shift unit for shifting the output of the adder to the right by a certain number of bits to achieve fixed-point precision conversion.

7. The multiplier high-efficiency hardware architecture system with non-separable second transformation at low frequency according to claim 1, characterized in that, The multiplier in the core module has a configurable structure. In a scenario with limited area, the multiplier can be replaced by a fused computing unit. The fused computing unit includes shift logic, addition logic, and selection logic for converting the multiplication operation between a fixed constant and input data into a combined operation of multiple shift operations and at most one addition operation, thereby realizing an equivalent approximate calculation of a time-multiplexed multi-constant multiplication operation.

8. The multiplier high-efficiency hardware architecture system for low-frequency inseparable secondary transformation according to claim 7, characterized in that The constants used for multiplication approximation in the fused computing unit come from a constant set after approximate reconstruction, which includes: The first set R0, including all constants that can be generated by shifting the value 1 to the left by any number of bits; The second set R1, including constants generated by performing an addition or subtraction on two constants in R0; The third set R2, including constants generated by performing an addition or subtraction on two constants in R0 or R1; The fourth set R3, including constants generated by performing an addition or subtraction on two constants in R0, R1, or R2; Among them, if the original constant does not belong to R0 or R1, it is replaced by the constant in R0 or R1 with the smallest difference from it, and the number of adders required for the operation after replacement does not exceed one. If there are multiple candidate replacement constants, the one with the smaller value is preferentially selected.

9. The high-efficiency hardware architecture system of a multiplier with non-separable low-frequency secondary transformation according to claim 1, characterized in that, The control signal output by the control module is determined according to the size of the transform kernel, the transform subset index, the transform kernel index, and the sub-region index generated by the current cycle count. The control signal is used for the selection module to complete constant loading and input the corresponding constant into the multiplier group of the corresponding sub-region of the core module.

Citation Information

Cited By

  • Multi-operator-oriented configurable accelerator and configuration method thereof

    CN121579410A