Quantization method and device in video coding, electronic equipment and storage medium
By performing a low-frequency non-separable transform on the initial transform unit in video coding, frequency domain information redundancy is reduced, and RDOQ processing is skipped when the number of non-zero coefficients is less than a threshold. This solves the problem of high computational cost of RDOQ and improves coding efficiency and resource utilization.
Patent Information
- Application Number
- CN202310126881.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-01
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2043-02-01
AI Technical Summary
Existing rate-distortion optimized quantization (RDOQ) techniques involve large computational demands in video coding, resulting in low efficiency and significant waste of computational resources.
By performing a low-frequency non-separable transform (LFNST) on the initial transform unit in video coding, frequency domain information redundancy is reduced, and the number of non-zero coefficients is determined. If it is less than a preset threshold, RDOQ processing is skipped; otherwise, rate-distortion optimized quantization is performed.
It reduces the amount of RDOQ computation in the video encoding process, improves encoding efficiency, and avoids wasting computing resources.
Smart Images

Figure CN116156193B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of video coding, and particularly relates to a quantization method in video coding, a device, an electronic device and a storage medium. BACKGROUND
[0002] At present, mainstream video coding standards are all based on a hybrid video coding framework, and quantization plays an important role and is also a source of video loss. Taking rate-distortion optimization quantization (RDOQ) technology as an example, it has been applied to many video coding standards such as high efficiency video coding (HEVC), versatile video coding (VVC) and the like, for improving video coding performance and saving video code rate.
[0003] However, the existing RDOQ technology has a large amount of calculation, resulting in problems of low efficiency and serious waste of computing resources. SUMMARY
[0004] The present disclosure provides a quantization method in video coding, a device, an electronic device and a storage medium, to reduce the amount of calculation of RDOQ, improve coding efficiency and avoid waste of computing resources. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, a quantization method in video coding is provided, and the method comprises: obtaining an initial transform unit of a video to be coded, performing transform processing on the initial transform unit to obtain a target transform unit; the initial transform unit comprises a plurality of initial transform coefficients; the target transform unit comprises a plurality of target transform coefficients; the initial transform coefficients and the target transform coefficients are both used to reflect frequency domain information of the video to be coded; the redundancy of the frequency domain information reflected by the plurality of target transform coefficients is less than the redundancy of the frequency domain information reflected by the plurality of initial transform coefficients; determining the number of non-zero coefficients in the plurality of target transform coefficients, and in the case that the number of non-zero coefficients is less than a preset threshold, determining the quantization value of each target transform coefficient in the target transform unit as a preset quantization value, and skipping the rate-distortion optimization quantization processing on each target transform coefficient in the target transform unit.
[0006] Optionally, the method further comprises: in the case that the number of non-zero coefficients is greater than or equal to the preset threshold, performing rate-distortion optimization quantization processing on each target transform coefficient in the target transform unit.
[0007] Optionally, the transform processing on the initial transform unit to obtain the target transform unit comprises: performing low frequency non-separable transform (LFNST) on the initial transform unit to obtain the target transform unit.
[0008] Optionally, the initial transform unit is subjected to a low-frequency non-separable transform (LFNST) to obtain a target transform unit, including: obtaining an intra prediction mode of the initial transform unit, and determining a target transform set corresponding to the initial transform unit from a mapping relationship including a plurality of intra prediction modes and a plurality of transform sets, the target transform set including a first LFNST transform kernel and a second LFNST transform kernel; in a case where a width of the initial transform unit is less than a preset length or a length of the initial transform unit is less than the preset length, performing LFNST on the initial transform unit according to the first LFNST transform kernel to obtain the target transform unit; in a case where the width of the initial transform unit is greater than or equal to the preset length or the length of the initial transform unit is greater than or equal to the preset length, performing LFNST on the initial transform unit according to the second LFNST transform kernel to obtain the target transform unit.
[0009] Optionally, in a case where the number of non-zero coefficients is greater than or equal to a preset threshold, performing rate-distortion optimization quantization processing on each target transform coefficient in the target transform unit, including: in a case where the number of non-zero coefficients is greater than or equal to the preset threshold, determining a candidate quantization set of each target transform coefficient; the candidate quantization set includes a plurality of optional quantization values; determining an optimal quantization value of each target transform coefficient from the candidate quantization set to complete the rate-distortion optimization quantization processing.
[0010] Optionally, the method further includes: obtaining size information of the target transform unit; the size information is used to reflect the size of the target transform unit; determining the preset threshold according to the size information; the preset threshold is positively correlated with the size of the target transform unit.
[0011] Optionally, determining the preset threshold according to the size information includes: determining the preset threshold from a mapping relationship including a plurality of transform unit size information and a plurality of thresholds according to the size information.
[0012] Optionally, the video to be encoded includes pixel information; the pixel information includes luminance information and chrominance information; the method further includes: performing a prediction transform on the pixel information to obtain the initial transform unit; the prediction transform includes any one of a luminance intra prediction transform, a luminance inter prediction transform, a chrominance intra prediction transform, or a chrominance inter prediction transform.
[0013] According to a second aspect of the embodiments of the present disclosure, a quantization device in video coding is provided. The device comprises an obtaining unit and a processing unit. The obtaining unit is configured to obtain an initial transform unit of a video to be coded. The processing unit is configured to perform transform processing on the initial transform unit to obtain a target transform unit. The initial transform unit comprises a plurality of initial transform coefficients. The target transform unit comprises a plurality of target transform coefficients. The initial transform coefficients and the target transform coefficients are used to reflect frequency domain information of the video to be coded. The frequency domain information reflected by the plurality of target transform coefficients has less redundancy than the frequency domain information reflected by the plurality of initial transform coefficients. The processing unit is further configured to determine a number of non-zero coefficients in the plurality of target transform coefficients, and in a case where the number of non-zero coefficients is less than a preset threshold, determine a quantization value of each target transform coefficient in the target transform unit as a preset quantization value, and skip performing rate-distortion optimization quantization processing on each target transform coefficient in the target transform unit.
[0014] Optionally, the processing unit is further configured to, in a case where the number of non-zero coefficients is greater than or equal to the preset threshold, perform rate-distortion optimization quantization processing on each target transform coefficient in the target transform unit.
[0015] Optionally, the processing unit is specifically configured to perform low-frequency non-separable transform (LFNST) on the initial transform unit to obtain the target transform unit.
[0016] Optionally, the processing unit is specifically configured to obtain an intra prediction mode of the initial transform unit, and determine a target transform set corresponding to the initial transform unit from a mapping relationship comprising a plurality of intra prediction modes and a plurality of transform sets. The target transform set comprises a first LFNST transform kernel and a second LFNST transform kernel. In a case where a width of the initial transform unit is less than a preset length or a length of the initial transform unit is less than the preset length, perform LFNST on the initial transform unit according to the first LFNST transform kernel to obtain the target transform unit. In a case where the width of the initial transform unit is greater than or equal to the preset length or the length of the initial transform unit is greater than or equal to the preset length, perform LFNST on the initial transform unit according to the second LFNST transform kernel to obtain the target transform unit.
[0017] Optionally, the processing unit is specifically configured to, in a case where the number of non-zero coefficients is greater than or equal to the preset threshold, determine a candidate quantization set of each target transform coefficient. The candidate quantization set comprises a plurality of optional quantization values. The processing unit is further configured to determine an optimal quantization value of each target transform coefficient from the candidate quantization set to complete the rate-distortion optimization quantization processing.
[0018] Optionally, the obtaining unit is further configured to obtain size information of the target transform unit. The size information is used to reflect a size of the target transform unit. The obtaining unit is further configured to determine the preset threshold according to the size information. The preset threshold is positively correlated with the size of the target transform unit.
[0019] Optionally, the obtaining unit is specifically configured to: determine the preset threshold according to the size information from a mapping relationship comprising a plurality of transform unit size information and a plurality of thresholds.
[0020] Optionally, the video to be encoded comprises pixel information; the pixel information comprises luminance information and chrominance information; and the processing unit is further configured to: perform a prediction transform on the pixel information to obtain an initial transform unit; the prediction transform comprises any one of a luminance intra prediction transform, a luminance inter prediction transform, a chrominance intra prediction transform, or a chrominance inter prediction transform.
[0021] According to a third aspect of the embodiments of the present disclosure, a computer device is provided, comprising: a processor, and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the quantization method in the video encoding of the first aspect.
[0022] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, and the computer readable storage medium stores instructions, when the instructions in the computer readable storage medium are executed by a processor of a computer device, the computer device is enabled to execute the quantization method in the video encoding of the first aspect.
[0023] The technical solutions provided by the present disclosure at least bring the following beneficial effects: the computer device acquires an initial transform unit comprising a plurality of initial transform coefficients, and performs transform processing on the initial transform unit to obtain a target transform unit comprising a plurality of target transform coefficients, so as to reduce the complexity of the initial transform unit. Further, the computer device determines the number of non-zero coefficients in the plurality of target transform coefficients, and in the case that the number of non-zero coefficients is less than a preset threshold, determines the quantization value of each target transform coefficient in the target transform unit as a preset quantization value. In this way, the present disclosure realizes the skipping of the RDOQ process of part of the transform units for the target transform unit whose number of non-zero coefficients is less than the preset threshold, so as to reduce the calculation amount of RDOQ in the video encoding process, improve the encoding efficiency, and avoid the waste of computing resources.
[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0025] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure, and do not constitute an undue limitation on the present disclosure.
[0026] Figure 1 is an application environment diagram of the quantization method provided by the present disclosure according to an exemplary illustration;
[0027] Figure 2 is a flowchart of a quantization method according to an example embodiment;
[0028] Figure 3 is a schematic diagram of intra angular prediction of HEVC according to an example embodiment;
[0029] Figure 4 is a schematic diagram of intra angular prediction of VVC according to an example embodiment;
[0030] Figure 5 is a schematic diagram of the principle of a quantization method according to an example embodiment;
[0031] Figure 6 is a flowchart of a quantization method according to an example embodiment;
[0032] Figure 7 is a schematic diagram of the scope of LFNST according to an example embodiment;
[0033] Figure 8 is a schematic diagram of forward transform of LFNST according to an example embodiment;
[0034] Figure 9 is a schematic diagram of inverse transform of LFNST according to an example embodiment;
[0035] Figure 10 is a flowchart of a quantization method according to an example embodiment;
[0036] Figure 11 is a flowchart of a quantization method according to an example embodiment;
[0037] Figure 12 is a schematic diagram of TU region setting according to an example embodiment;
[0038] Figure 13 is a schematic diagram of a quantization device according to an example embodiment;
[0039] Figure 14 is a schematic diagram of a computer device according to an example embodiment. DETAILED DESCRIPTION
[0040] In order to make the ordinary people in the art better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings.
[0041] It should be noted that the terms "first", "second", etc. in the description of the present disclosure and claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The implementation described in the following exemplary embodiments does not represent all implementations consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0042] In addition, in the description of the embodiments of the present disclosure, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B. "And / or" herein is only a description of the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present disclosure, "multiple" means two or more than two.
[0043] It should be noted that the user information (including but not limited to user equipment information, user personal information, user behavior information, etc.) and data (including but not limited to program code, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties.
[0044] Before explaining the embodiments of the present disclosure in detail, some related technical terms and related technologies involved in the embodiments of the present disclosure are introduced.
[0045] Video encoding refers to a way of converting a file in an original video format into another video format file through compression technology. Video is a sequence of continuous images, composed of continuous frames, and a frame is an image. Due to the persistence of vision of the human eye, when the frame sequence is played at a certain rate, the viewer sees a video with continuous action. Since the similarity between consecutive frames is very high, in order to facilitate storage and transmission, the original video needs to be encoded and compressed to remove spatial and temporal redundancy.
[0046] Quantization refers to the process of mapping the continuous values (or a large number of possible discrete values) of a signal to a finite number of discrete amplitude values, which is a many-to-one mapping. In the video encoding process, after the residual signal is subjected to discrete cosine transform (DCT), the transform coefficients usually have a large range, so quantizing the transform coefficients can effectively reduce the signal value space and thus obtain a better bit rate. However, due to the many-to-one mapping property, the quantization process inevitably introduces data loss. Quantization is an important source of video distortion in video encoding.
[0047] High Efficiency Video Coding (HEVC) is the next generation video coding standard after H.264, and its core goal is to improve the compression efficiency by 1 times on the basis of H.264 / AVC High Profile, i.e., to reduce the bit rate of video stream by 50% under the premise of ensuring the same video image quality. HEVC adopts a hybrid coding framework based on blocks.
[0048] Versatile Video Coding (VVC) is the next generation video coding standard formulated by the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T), and compared with the previous generation HEVC video coding standard, in order to improve the compression performance, more than 30 new coding tools are added in the VVC standard, covering every module in the hybrid video coding system framework, and improvements are made to a certain extent in block partitioning, intra and inter prediction, residual coding, transform and quantization, entropy coding, loop filtering, etc.
[0049] Low-frequency non-separable transform (LFNST) is a new coding tool adopted in the VVC standard, which performs a secondary transform on the low-frequency components by using the transformed coefficients after the angle mode of the intra prediction, so as to improve the efficiency of video coding.
[0050] Rate-distortion optimization (RDO) is a method to improve the quality of video compression. The name refers to optimizing the distortion amount (loss of video quality) with respect to the amount of data (bit rate) required for video coding.
[0051] Rate distortion optimized quantization (RDOQ) is a coefficient optimization algorithm. Specifically, in video coding, distortion and bit rate are both factors affecting coding performance. Among them, distortion reflects video quality (quantization is an important source of distortion), and bit rate reflects compression rate. Reducing distortion generally increases bit rate; reducing bit rate generally increases distortion. Therefore, to balance distortion and bit rate, video coding needs to balance distortion and bit rate, thereby introducing the rate distortion optimized quantization (RDOQ) technology, which combines the quantization process with the RDO (rate distortion optimization) criterion. For a transform coefficient, multiple selectable quantization values are given, and the optimal quantization value (quantization level) is selected by the RDO criterion.
[0052] A new block concept is introduced for HEVC. For example, a coding unit (CU) can refer to a sub-partitioning of a video frame into rectangular blocks of the same or variable size. In HEVC, a CU can replace the macroblock structure of previous standards. Depending on the mode of inter- or intra-prediction, a CU can include one or more prediction units (PUs), each of which can act as a basic prediction unit. For example, for intra-prediction, an 8x8 CU can be symmetrically divided into four 4x4 PUs. For another example, for inter-prediction, a 64x64 CU can be asymmetrically divided into a 16x64 PU and a 48x64 PU. Similarly, a PU can include one or more transform units (TUs), each of which can act as a basic unit for transformation and / or quantization. For example, a 32x32 PU can be symmetrically divided into four 16x16 TUs. Multiple TUs in one PU can share the same prediction mode, but can be individually transformed. In this disclosure, the term "block" can generally refer to any of a macroblock, a CU, a PU, or a TU. RDO is the basis for HEVC to decide the optimal prediction mode.
[0053] In a video coding algorithm, quantization is a core module to achieve good compression effect. In the field of video coding, quantization is a process of approximating continuous values of video signals to a finite number of discrete values. The rate-distortion optimization quantization algorithm (i.e., RDOQ) is a quantization algorithm adopted by many coding standards such as HEVC, VVC, and ECM (ECM is another advanced standard after VVC), which combines the quantization process with RDO. Through RDO, the optimal quantization value is selected for each coefficient, so the coding performance can be significantly improved.
[0054] However, due to the addition of RDO calculation in the process of quantizing the coefficients, the algorithm has high computational complexity. For example, in the fast mode of the x265 encoder of the HEVC coding standard, the computational complexity of RDOQ accounts for 20% to 30% of the total encoder complexity of the encoder. Such high complexity also leads to the fact that RDOQ cannot be effectively applied in some use scenarios with high requirements for computational complexity, and therefore there is an urgent need to design a quantization method in video coding to reduce the computational complexity of RDOQ.
[0055] To solve the above problems, the present disclosure provides a quantization method in video coding, which analyzes the frequency domain information of the coefficients in the TU, and decides whether to skip RDOQ according to the analysis result, thereby reducing the overall computational complexity of RDOQ in the encoder.
[0056] An exemplary application environment of the present disclosure is provided below. For example, the quantization method provided by the embodiments of the present disclosure can be applied to a computer device as shown in Figure 1 .
[0057] The computer device 10 can be configured to access the content (e.g., video) and services of a server.
[0058] The computer device 10 can include an electronic device carrying or externally connected to a display panel, such as a mobile device, a tablet device, a laptop computer, a workstation, a virtual reality device, a game device, a digital streaming media device, a vehicle user terminal, a smart television, a set-top box, etc., and can also include a virtualized computing instance. The virtualized computing instance can include a virtual machine, such as a computer system, an operating system, a server, etc.
[0059] The computer device 10 can be associated with one or more users. A single user can also use one or more of the computer devices 10 to access a server. The computer device 10 can travel to various locations and use different networks to access the server. The computer device 10 can include a plurality of client programs, such as a video codec for providing encoding and decoding services. The video codec can encode and compress a video or image for easy transmission or storage.
[0060] Figure 2 is a flowchart of a quantization method according to some exemplary embodiments. In some embodiments, the quantization method described above can be applied to a computer device as shown in Figure 1 , and can also be applied to other similar devices.
[0061] As shown in Figure 2 , the quantization method provided by the embodiments of the present disclosure includes the following S201-S203.
[0062] S201, the computer device obtains an initial transform unit.
[0063] The transform unit includes a plurality of initial transform coefficients.
[0064] As a possible implementation, the computer device obtains the initial transform unit of the video to be encoded from the server.
[0065] It should be noted that in the video encoding process, the transform unit is the basic unit for independent transformation and quantization, and its size is also flexible. The data in the transform unit is the transform coefficient. These transform coefficients can be used to represent the luma component or the Cb component and the Cr component of the video frame to be encoded.
[0066] VVC is the next generation coding standard of HEVC, and there are some differences in the implementation of the two. In order to facilitate understanding, the differences related to the implementation of the present disclosure are described here.
[0067] Figure 3 An intra-frame angle prediction schematic diagram of HEVC is shown. Figure 4 An intra-frame angle prediction schematic diagram of VVC is shown. Reference Figure 3 And Figure 4 There are 35 angle prediction modes in HEVC, and the angle prediction mode of VVC is further expanded, with a number of 95.
[0068] In HEVC, the size of the TU block has four sizes of 4x4, 8x8, 16x16, and 32x32, and all are square blocks with equal width and height. The size of the TU block in VVC is more diverse, denoted as NxM TU block, (N, M can be equal to 2, 4, 8, 16, 32, 64, 128).
[0069] In HEVC, the concepts of CU and TU are different. In the division, intra / inter prediction module, it is called CU, that is, the CU is predicted. Then in the transform module, the CU is processed by transformation, which is called TU. However, it should be noted that in the transform module, since the CU can be recursively divided by quad-tree, a CU can be divided into multiple TUs. Therefore, it is possible that a CU contains multiple TUs.
[0070] In VVC, the concepts of CU and TU are no longer strictly emphasized, because VVC does not allow CU to be divided into TU in the transform module, so the size of a CU is equal to that of its corresponding TU.
[0071] S202, the computer device performs transform processing on the initial transform unit to obtain a target transform unit.
[0072] Among them, the target transform unit includes a plurality of target transform coefficients; the complexity of the target transform unit is less than the complexity of the initial transform unit. The initial transform coefficients and the target transform coefficients are both used to reflect the frequency domain information of the video to be encoded, but the redundancy of the frequency domain information reflected by the plurality of target transform coefficients is less than the redundancy of the frequency domain information reflected by the plurality of initial transform coefficients.
[0073] As a possible implementation manner, the computer device performs transform processing on the initial transform unit according to a preset change algorithm to obtain the target transform unit.
[0074] It should be noted that the change algorithm is set in the computer device by the operation and maintenance personnel in advance, for example, the change algorithm can be LFNST algorithm.
[0075] The LFNST algorithm is a tool in the VVC transform, which is used to perform a secondary transform on a TU block, convert the spatial domain information of the TU to the frequency domain, and further remove the frequency domain redundancy between the TU block transform coefficients, so as to improve the coding performance. Therefore, the target transform unit obtained through the transform processing has a complexity less than that of the initial transform unit.
[0076] In S203, the computer device determines the number of non-zero coefficients in the plurality of target transform coefficients, and determines the quantization values of the target transform coefficients in the target transform unit as preset quantization values in a case where the number of non-zero coefficients is less than a preset threshold.
[0077] The non-zero coefficients in the target transform coefficients represent the pixels of the video to be encoded, and the zero coefficients represent the pixels of zero. The quantization value of the target transform coefficient is the specific value of the quantization parameter (QP) corresponding to the quantization step (Qstep) adopted in the quantization process.
[0078] As a possible implementation manner, the computer device counts the number of non-zero coefficients in the target transform unit. Further, the computer device compares the counted number of non-zero coefficients with the preset threshold. In a case where the number of non-zero coefficients is less than the preset threshold, the computer device determines the quantization values of the target transform coefficients in the target transform unit as preset quantization values.
[0079] It should be noted that the preset threshold and the preset quantization value are both set in the computer device by the operation and maintenance personnel in advance. The preset threshold is a non-negative integer. For example, the preset threshold is 3, and the preset quantization value is 0.
[0080] Since the initial data that is the encoding target of a picture is a pixel value in the spatial domain, each pixel group having a predetermined size can be used as a data unit that is the encoding target. In addition, a transform on the pixel values of the pixel group in the spatial domain is performed for video encoding, thereby generating transform coefficients in the transform domain, and in this regard, the transform coefficients maintain a coefficient group having the same size as the pixel group in the spatial domain. Therefore, the coefficients of the transform coefficients in the transform domain can also be used as data units for encoding of a picture.
[0081] Therefore, in the entire spatial domain and the transform domain, a data group having a predetermined size can be used as a data unit for encoding. Here, the size of the data unit can be defined as the total number of data included in the data unit. For example, the total number of pixels in the spatial domain or the total number of transform coefficients in the transform domain can indicate the size of the data unit.
[0082] It should be noted that the initial transform coefficients and the target transform coefficients in the present disclosure are both used to reflect the image frame information of the video to be encoded in the frequency domain; the number of target transform coefficients is different from the number of initial transform coefficients.
[0083] The technical solution provided in this disclosure offers at least the following advantages: A computer device acquires an initial transform unit comprising multiple initial transform coefficients and performs transform processing on the initial transform unit to obtain a target transform unit comprising multiple target transform coefficients, thereby reducing the complexity of the initial transform unit. Furthermore, the computer device determines the number of non-zero coefficients among the multiple target transform coefficients. If the number of non-zero coefficients is less than a preset threshold, the quantization value of each target transform coefficient in the target transform unit is determined as a preset quantization value. In this way, for target transform units where the number of non-zero coefficients is less than a preset threshold, this disclosure achieves the skipping of the RDOQ process for some transform units, thereby reducing the computational load of RDOQ during video encoding, improving encoding efficiency, and avoiding waste of computational resources.
[0084] In one design, in order to quantize each target transformation coefficient in the target transformation unit, embodiments of this disclosure provide a quantization method, which further includes the following S204:
[0085] S204. When the number of non-zero coefficients is greater than or equal to a preset threshold, the computer equipment performs rate distortion optimization quantization processing on each target transformation coefficient in the target transformation unit.
[0086] As one possible implementation, the computer device counts the number of non-zero coefficients in the target transformation unit. Further, the computer device compares the counted number of non-zero coefficients with a preset threshold. If the number of non-zero coefficients is greater than or equal to the preset threshold, the computer device performs rate-distortion optimization quantization processing on each target transformation coefficient in the target transformation unit.
[0087] For example, such as Figure 5 As shown, before quantization of the TU block, a second transformation is first performed on the TU using LFNST, the implementation of which is the same as the current LFNST in VVC. Next, the number of non-zero coefficients after the second transformation is counted, denoted as numSig. Finally, the relationship between numSig and a preset threshold Threshold is determined, and the decision on whether the TU needs to be quantized using RDOQ is based on this relationship. If numSig ≥ Threshold, the TU coefficients are quantized using RDOQ; if numSig < Threshold, RDOQ quantization is skipped, and all TU coefficients are directly set to 0.
[0088] In one design, in order to obtain the target transformation unit, such as Figure 6 As shown, the above-described S202 provided in this embodiment of the present disclosure specifically includes the following S2021-S2022:
[0089] S2021, the computer device performs primary transformation on the initial transform unit to obtain an intermediate transform unit.
[0090] The complexity of the intermediate transform unit is less than or equal to the complexity of the initial transform unit and greater than the complexity of the target transform unit.
[0091] As a possible implementation, the computer device performs primary transformation on the initial transform unit, and trims the initial transform unit to obtain the intermediate transform unit.
[0092] It can be understood that the complexity of the trimmed initial transform unit (i.e., the intermediate transform unit) is less than or equal to the complexity of the initial transform unit.
[0093] S2022, the computer device performs low-frequency non-separable secondary transformation on the intermediate transform unit to obtain a target transform unit.
[0094] As a possible implementation, the computer device determines the size (i.e., length* width) of the intermediate transform unit. In the case that the width of the intermediate transform unit is less than a preset length, or the length of the intermediate transform unit is less than a preset length, the computer device performs LFNST on the intermediate transform unit according to a first LFNST transformation kernel to obtain the target transform unit; in the case that the width of the intermediate transform unit is greater than or equal to the preset length, or the length of the intermediate transform unit is greater than or equal to the preset length, the computer device performs LFNST on the intermediate transform unit according to a second LFNST transformation kernel to obtain the target transform unit.
[0095] As shown in Figure 7 , LFNST is used between the initial transformation and quantization of the encoding process, and between the inverse quantization and inverse initial transformation of the decoding process.
[0096] For the convenience of understanding, the principle of secondary transformation of LFNST on the TU block is specifically introduced.
[0097] When the computer device performs LFNST calculation, for a TU block of NxM size, according to the size of N and M, there are two processing conditions: if min(N, M) = 4, at this time, the TU block needs to be first unfolded into a 16x1 vector, then multiplied by a 16x16 LFNST transformation kernel (LFNST kernel) to obtain 16 LFNST coefficients (LFNST coeff), and finally the 16 coefficients are backfilled into the TU block to obtain the target transform unit. For a TU block of min(N, M) >= 8, as shown in Figure 8As shown, first, the top-left 3x3 block is selected; second, it is unfolded into a 9x1 vector, then multiplied by a 16x9 LFNST kernel to get 16 LFNST coefficients, and finally, the 16 coefficients are backfilled into the top-left 3x3 block. It is noted that for a TU block with min(N,M)>=8, after the LFNST secondary transform, the remaining TU block will be set to 0 except the top-left 3x3 block.
[0098] Figure 9 The inverse transform of LFNST is shown, which is exactly the opposite of the forward transform of LFNST, and thus is not described herein. Figure 8 The inverse transform of LFNST is shown, which is exactly the opposite of the forward transform of LFNST, and thus is not described herein.
[0099] In addition, for the selection of the LFNST kernel, the computer device can determine from a preset transform set.
[0100] For example, LFNST has 4 transform sets, each of which has 2 transform kernels, and each transform set corresponds to an intra prediction mode. Referring to Table 1 (VVC intra prediction angle and LFNST transform set correspondence table), if the intra prediction angle mode of the CU is intraPredMode<0 or 56≤intraPredMode≤80, the first LFNST transform set is selected; and so on. Figure 3 The corresponding transform set of each CU can be determined.
[0101] Table 1
[0102] intraPredMode transform set intraPredMode<0 1 0 <= intraPredMode <= 1 0 2 <= intraPredMode <= 12 1 13 <= intraPredMode <= 23 2 24 <= intraPredMode <= 44 3 45 <= intraPredMode <= 55 2 56 <= intraPredMode <= 80 1
[0103] In VVC, each transform set has two transform kernels, so if LFNST is used for secondary transform, each CU block needs to select an optimal transform kernel in the encoder through rate-distortion optimization. For example, for an intra prediction CU block with IntraPredMode = 60, the block needs to calculate three rate-distortion costs, the first time is the rate-distortion cost cost0 without using LFNST transform kernel for secondary transform, and the second and third times are the rate-distortion costs cost1 and cost2 derived by using the first and second transform kernels of the first LFNST transform set for secondary transform. Then, according to the size relationship of the three rate-distortion costs, the selection is made: if cost0 = min(cost0, cost1, cost2), the current CU block does not use LFNST for secondary transform; if cost1 = min(cost0, cost1, cost2), the current CU block selects the first LFNST transform kernel for transform, and the rest is similar.
[0104] Since there are only 35 angular prediction modes in HEVC, the correspondence between the intra prediction angle and the transform set of LFNST needs to refer to Table 2.
[0105] Table 2
[0106]
[0107] In one design, to perform rate-distortion optimization quantization processing on each target transform coefficient in a target transform unit, as shown in FIG. 2, the disclosure provides the S204, which specifically includes the following S2041-S2042. Figure 10
[0108] S2041, in a case where the number of non-zero coefficients is greater than or equal to a preset threshold, the computer device determines a candidate quantization set of each target transform coefficient.
[0109] In the candidate quantization set, multiple selectable quantization values are included.
[0110] As a possible implementation, in a case where the number of non-zero coefficients is greater than or equal to a preset threshold, the computer device calculates a pre-quantization value of each target transform coefficient according to a pre-quantization formula. Further, the computer device determines the selectable quantization value according to the pre-quantization value.
[0111] For example, the computer device calculates the pre-quantization value of each target transform coefficient using the following formula.
[0112]
[0113] Wherein, Ci is the i-th transform coefficient of the TU, li is the pre-quantization value obtained by pre-quantization of Ci, Q represents the quantization step, which can be calculated from the quantization parameter (QP), here 0.5 is used as the quantization offset for quantization, round(*) represents rounding, and |*| represents taking the absolute value.
[0114] Further, the computer device determines the selectable quantization value according to the size of |li|, which is shown in Table 3 (the selectable quantization values corresponding to different |li|):
[0115] Table 3
[0116] S2042, the computer device determines the optimal quantization value of each target transform coefficient from the candidate quantization set.
[0117] As a possible implementation, the computer device traverses all coefficients of the current TU, and for each non-zero coefficient, traverses its selectable quantization values, and determines the optimal quantization value of each coefficient using the RDO criterion.
[0118] For example, for transform coefficients ci where the prequantization coefficient li is not 0, for optional quantization values l i,k Its rate-distortion cost is:
[0119] J(l i,k )=D(ci,l i,k )+λ*R(l i,k )
[0120] Where D(ci,li,k) and R(l) i,k ) indicates that ci is quantized as l i,k At that time, the total distortion and total number of encoded bits of the current TU are given; the computer equipment calculates the rate-distortion cost of all optional quantization values according to the formula, and finally selects the optional quantization value with the minimum rate-distortion cost as the optimal quantization value.
[0121] In the HEVC standard, when a computer device performs entropy encoding on the current TU, it divides it into several 4x4 coefficient groups (CGs). If the current CG is all zero, it is only necessary to mark the CG as all zero; otherwise, it is necessary to encode the identifier and all coefficient magnitude identifiers.
[0122] The computer device uses the RDO criterion to determine the position of the last non-zero coefficient. It iterates through the non-zero coefficients in the TU after steps S2041-S2042, calculates the total rate distortion cost of the current TU when each coefficient is the "last non-zero coefficient", and selects the non-zero coefficient with the smaller total rate distortion cost as the "last non-zero coefficient".
[0123] In one design, to determine a preset threshold, such as Figure 11 As shown, the present disclosure provides a quantization method, which further includes the following steps S301-S302:
[0124] S301, The computer device acquires the size information of the target transformation unit.
[0125] The size information is used to reflect the size of the target transformation unit.
[0126] S302. The computer equipment determines the preset threshold based on the size information.
[0127] As one possible implementation, the computer device determines a preset threshold based on the size information from a mapping relationship between multiple variable unit size information and multiple thresholds.
[0128] For example, such as Figure 12As shown, the embodiment of the present disclosure proposes to set different threshold values for different regions by the TU, specifically, the image is divided into two regions A and B, the region A is the upper left quarter region of the TU, and the region B is the remaining three quarter region. The two regions correspond to different threshold values: the region A corresponds to the threshold value Th1, and the region B corresponds to the threshold value Th2. Assuming that the number of non-zero coefficients in the region A and the region B is numSig1 and numSig2 respectively, only when numSig1 <= Th1 and numSig2 <= Th2, the current TU skips the RDOQ.
[0129] For the values of Th1 and Th2, the embodiment of the present disclosure adaptively adjusts according to the size information of the TU (in HEVC, the TU size has four sizes: 4x4, 8x8, 16x16 and 32x32), and the specific size and the corresponding values of Th1 and Th2 are shown in Table 4 (size and Th1 and Th2 size correspondence).
[0130] Table 4
[0131] 4x4 8x8 16x16 32x32 Th1 0 1 3 3 Th2 1 2 3 5
[0132] The above embodiments mainly introduce the schemes provided by the embodiments of the present disclosure from the perspective of the device (apparatus). It can be understood that, in order to implement the above method, the device or apparatus contains the corresponding hardware structure and / or software module for executing each method process, and these corresponding hardware structure and / or software module for executing each method process can constitute an electronic device. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.
[0133] The embodiments of the present disclosure can divide the device or apparatus into functional modules according to the above method examples, for example, the device or apparatus can correspond to each functional module, or two or more functions can be integrated into one processing module. The above integrated module can be realized in the form of hardware or software functional module. It should be noted that the division of modules in the embodiments of the present disclosure is illustrative, and is only a logical functional division. There can be another division manner when actually implemented.
[0134] Figure 13 is a structural schematic diagram of a quantization device according to an example embodiment. Referring to Figure 13As shown, the quantization device 40 provided by the embodiment of the present disclosure comprises an acquisition unit 401 and a processing unit 402.
[0135] The acquisition unit 401 is configured to acquire an initial transform unit of a video to be encoded; the processing unit 402 is configured to perform transform processing on the initial transform unit to obtain a target transform unit; the initial transform unit comprises a plurality of initial transform coefficients; the target transform unit comprises a plurality of target transform coefficients; the initial transform coefficients and the target transform coefficients are both used to reflect frequency domain information of the video to be encoded; the redundancy of the frequency domain information reflected by the plurality of target transform coefficients is smaller than the redundancy of the frequency domain information reflected by the plurality of initial transform coefficients; the processing unit 402 is further configured to determine the number of non-zero coefficients in the plurality of target transform coefficients, and in the case that the number of non-zero coefficients is less than a preset threshold, determine the quantization value of each target transform coefficient in the target transform unit as a preset quantization value, and skip performing rate-distortion optimization quantization processing on each target transform coefficient in the target transform unit.
[0136] Optionally, the processing unit 402 is further configured to, in the case that the number of non-zero coefficients is greater than or equal to the preset threshold, perform rate-distortion optimization quantization processing on each target transform coefficient in the target transform unit.
[0137] Optionally, the processing unit 402 is specifically configured to perform low-frequency non-separable transform (LFNST) on the initial transform unit to obtain the target transform unit.
[0138] Optionally, the processing unit 402 is specifically configured to acquire an intra prediction mode of the initial transform unit, and determine a target transform set corresponding to the initial transform unit from a mapping relationship comprising a plurality of intra prediction modes and a plurality of transform sets, the target transform set comprising a first LFNST transform kernel and a second LFNST transform kernel; in the case that the width of the initial transform unit is less than a preset length or the length of the initial transform unit is less than the preset length, perform LFNST on the initial transform unit according to the first LFNST transform kernel to obtain the target transform unit; in the case that the width of the initial transform unit is greater than or equal to the preset length or the length of the initial transform unit is greater than or equal to the preset length, perform LFNST on the initial transform unit according to the second LFNST transform kernel to obtain the target transform unit.
[0139] Optionally, the processing unit 402 is specifically configured to, in the case that the number of non-zero coefficients is greater than or equal to the preset threshold, determine a candidate quantization set of each target transform coefficient; the candidate quantization set comprises a plurality of optional quantization values; and determine an optimal quantization value of each target transform coefficient from the candidate quantization set to complete the rate-distortion optimization quantization processing.
[0140] Optionally, the obtaining unit 401 is further configured to: obtain size information of the target transform unit; the size information is used to reflect a size of the target transform unit; determine the preset threshold according to the size information; the preset threshold is positively correlated with the size of the target transform unit.
[0141] Optionally, the obtaining unit 401 is specifically configured to: determine the preset threshold from a mapping relationship including a plurality of transform unit size information and a plurality of thresholds according to the size information.
[0142] Optionally, the video to be encoded includes pixel information; the pixel information includes luminance information and chrominance information; and the processing unit is further configured to: perform a prediction transform on the pixel information to obtain an initial transform unit; the prediction transform includes any one of a luminance intra prediction transform, a luminance inter prediction transform, a chrominance intra prediction transform, or a chrominance inter prediction transform.
[0143] Figure 14 is a structural schematic diagram of a computer device provided by the present disclosure. As shown in Figure 14 The computer device 50 can include at least one processor 501 and a memory 502 for storing processor-executable instructions, wherein the processor 501 is configured to execute the instructions in the memory 502 to implement the quantization method in the above embodiments.
[0144] In addition, the computer device 50 can further include a communication bus 503 and at least one communication interface 504.
[0145] The processor 501 can be a central processing unit (CPU), a micro processing unit, an ASIC, or one or more integrated circuits for controlling program execution of the present disclosure.
[0146] The communication bus 503 can include a path for transmitting information between the above components.
[0147] The communication interface 504 uses any transceiver-like device for communicating with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0148] Memory 502 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory may exist independently and be connected to the processor via a bus. Memory may also be integrated with the processor.
[0149] The memory 502 stores instructions for executing the scheme of this disclosure, and the processor 501 controls the execution of these instructions. The processor 501 executes the instructions stored in the memory 502 to implement the functions of the quantization method of this disclosure.
[0150] As an example, combined Figure 13 The functions implemented by the acquisition unit 401 and the processing unit 402 in the quantization device 50 are the same as those of the acquisition unit 401 and the processing unit 402. Figure 14 The processor 501 in it has the same function.
[0151] In a specific implementation, as one example, the processor 501 may include one or more CPUs, for example... Figure 14 CPU0 and CPU1 in the CPU.
[0152] In a specific implementation, as one example, the computer device 50 may include multiple processors, such as... Figure 14 Processors 501 and 507 are shown in the diagram. Each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor here can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0153] In a specific implementation, as an embodiment, the computer device 50 can further include an output device 505 and an input device 506. The output device 505 is in communication with the processor 501 and can display information in a variety of manners. For example, the output device 505 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, a projector, or the like. The input device 506 is in communication with the processor 501 and can accept input from a user in a variety of manners. For example, the input device 506 can be a mouse, a keyboard, a touch screen device, a sensor device, or the like.
[0154] Those skilled in the art can understand that the structure shown in the above embodiments does not constitute a limitation on the computer device 50, and the computer device 50 can include more or fewer components than shown, or combine certain components, or adopt a different arrangement of components. Figure 14
[0155] In addition, the present disclosure also provides a computer readable storage medium, when the instructions in the computer readable storage medium are executed by the processor of the computer device, the computer device can execute the quantification method provided by the above embodiments.
[0156] In addition, the present disclosure also provides a computer program product, including computer instructions, when the computer instructions are run on the computer device, the computer device executes the quantification method provided by the above embodiments.
[0157] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure following the general principles thereof and including such departures from the present disclosure that come within known, accepted, or customary practice in the art to which the present disclosure pertains. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the present disclosure are indicated by the claims.
Claims
1. A quantization method in video coding, characterized in that, The method includes: An initial transform unit of the video to be encoded is obtained, and the initial transform unit is transformed to obtain a target transform unit. The initial transform unit includes multiple initial transform coefficients, and the target transform unit includes multiple target transform coefficients. Both the initial transform coefficients and the target transform coefficients are used to reflect the frequency domain information of the video to be encoded. The redundancy of the frequency domain information reflected by the multiple target transform coefficients is less than the redundancy of the frequency domain information reflected by the multiple initial transform coefficients. The number of non-zero coefficients among the plurality of target transformation coefficients is determined. If the number of non-zero coefficients is less than a preset threshold, the quantization value of each target transformation coefficient in the target transformation unit is determined as a preset quantization value, and the rate distortion optimization quantization processing of each target transformation coefficient in the target transformation unit is skipped. The method further includes: Obtain the size information of the target transformation unit; the size information is used to reflect the size of the target transformation unit; The preset threshold is determined based on the size information; the preset threshold is positively correlated with the size of the target transformation unit. The target transformation unit includes multiple regions, and different regions in the target transformation unit correspond to different preset thresholds; if the number of non-zero coefficients in the multiple regions is less than the corresponding preset threshold, then the rate distortion optimization quantization process for each target transformation coefficient in the target transformation unit is skipped.
2. The quantization method in video encoding according to claim 1, characterized in that, The method further includes: When the number of non-zero coefficients is greater than or equal to the preset threshold, rate distortion optimization quantization processing is performed on each of the target transformation coefficients in the target transformation unit.
3. The quantization method in video encoding according to claim 1, characterized in that, The transformation process of the initial transformation unit to obtain the target transformation unit includes: The initial transform unit is subjected to a low-frequency non-separable transform (LFNST) to obtain the target transform unit.
4. The quantization method in video encoding according to claim 3, characterized in that, The step of performing a low-frequency non-separable transform (LFNST) on the initial transform unit to obtain the target transform unit includes: The intra-frame prediction mode of the initial transform unit is obtained, and the target transform set corresponding to the initial transform unit is determined from the mapping relationship including multiple intra-frame prediction modes and multiple transform sets. The target transform set includes a first LFNST transform kernel and a second LFNST transform kernel. If the width of the initial transformation unit is less than the preset length, or if the length of the initial transformation unit is less than the preset length, the initial transformation unit is subjected to LFNST according to the first LFNST transformation kernel to obtain the target transformation unit; If the width of the initial transformation unit is greater than or equal to the preset length, or if the length of the initial transformation unit is greater than or equal to the preset length, the initial transformation unit is subjected to LFNST according to the second LFNST transformation kernel to obtain the target transformation unit.
5. The quantization method in video encoding according to claim 2, characterized in that, When the number of non-zero coefficients is greater than or equal to the preset threshold, rate-distortion optimization quantization processing is performed on each of the target transformation coefficients in the target transformation unit, including: If the number of non-zero coefficients is greater than or equal to the preset threshold, a candidate quantization set for each target transform coefficient is determined; the candidate quantization set includes multiple selectable quantization values. From the candidate quantization set, the optimal quantization value of each target transform coefficient is determined to complete the rate-distortion optimized quantization process.
6. The quantization method in video coding according to claim 1, characterized in that, Determining the preset threshold based on the size information includes: Based on the size information, the preset threshold is determined from the mapping relationship between multiple transformation unit size information and multiple thresholds.
7. The quantization method in video coding according to any one of claims 1-6, characterized in that, The video to be encoded includes pixel information; The pixel information includes luminance information and chrominance information; the method further includes: The pixel information is subjected to a predictive transformation to obtain the initial transformation unit; the predictive transformation includes any one of the following: intra-frame predictive transformation of luminance, inter-frame predictive transformation of luminance, intra-frame predictive transformation of chrominance, or inter-frame predictive transformation of chrominance.
8. A quantization device for video encoding, characterized in that, The device includes an acquisition unit and a processing unit; The acquisition unit is configured to acquire the initial transformation unit of the video to be encoded; The processing unit is configured to perform transformation processing on the initial transformation unit to obtain a target transformation unit; the initial transformation unit includes multiple initial transformation coefficients; the target transformation unit includes multiple target transformation coefficients; both the initial transformation coefficients and the target transformation coefficients are used to reflect the frequency domain information of the video to be encoded; the redundancy of the frequency domain information reflected by the multiple target transformation coefficients is less than the redundancy of the frequency domain information reflected by the multiple initial transformation coefficients. The processing unit is further configured to determine the number of non-zero coefficients among the plurality of target transform coefficients, and if the number of non-zero coefficients is less than a preset threshold, to determine the quantization value of each target transform coefficient in the target transform unit as a preset quantization value, and to skip the rate distortion optimization quantization processing performed on each target transform coefficient in the target transform unit. The acquisition unit is further configured to acquire the size information of the target transformation unit; the size information is used to reflect the size of the target transformation unit; The processing unit is further configured to determine the preset threshold based on the size information; the preset threshold is positively correlated with the size of the target transformation unit. The target transformation unit includes multiple regions, and different regions in the target transformation unit correspond to different preset thresholds; if the number of non-zero coefficients in the multiple regions is less than the corresponding preset threshold, then the rate distortion optimization quantization process for each target transformation coefficient in the target transformation unit is skipped.
9. A computer device, characterized in that, include: A processor and a memory for storing instructions executable by the processor; wherein the processor is configured to execute instructions to implement the quantization method in video encoding as described in any one of claims 1-7.
10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of a computer device, the computer device is enabled to perform the quantization method in video coding as described in any one of claims 1-7.