Coding management method based on high-efficiency video coding
By obtaining the correlation calculation results of the basic unit in HEVC video encoding to determine whether to divide it, the problem of high encoder complexity is solved, and the effect of reducing calculation costs and improving encoding efficiency is achieved.
Patent Information
- Application Number
- CN201910512079.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-06-13
AI Technical Summary
In the existing high-efficiency video encoding technology, the encoder has a high complexity, resulting in an increase in computing costs.
By obtaining the correlation calculation results of the HEVC basic unit before and after division, it is determined whether the basic unit is divided to reduce the complexity of the encoder.
By using the correlation results of the basic unit as a division decision condition, the complexity and calculation cost of the encoder are significantly reduced, while improving the encoding efficiency.
Smart Images

Figure CN112087624B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing, and in particular to a coding management method based on high-efficiency video coding. Background Art
[0002] With the rapid development of the video industry, video resolution has increased from standard definition, high definition, ultra-high definition to 4K / 8K, and frame rates have increased from 30 frames, 60 frames, 90 frames to 120 frames. The amount of information contained is constantly expanding, which is bound to bring great pressure to network bandwidth. How to improve the encoding quality of video streams becomes very important.
[0003] In order to improve the coding quality, the International Video Coding Standards Organization proposed the High Efficiency Video Coding (HEVC) standard, also known as H.265, which introduced the Coding Tree Unit (CTU) and adopted a quadtree structure for image block division. This block division method can achieve better coding efficiency than H.264 / AVC (Advanced Video Coding), because it is necessary to calculate the cost of rate-distortion optimization (RDO) for each size of coding unit (CU), prediction unit (PU) and transform unit (TU) to obtain the optimal division, so the complexity of the encoder is very large. Summary of the invention
[0004] In order to solve the above technical problems, the present application provides a coding management method based on high-efficiency video coding, which can reduce the complexity of the encoder.
[0005] In order to achieve the purpose of this application, this application provides a coding management method based on high-efficiency video coding HEVC, including:
[0006] Obtaining correlation calculation results of HEVC basic units before and after division, wherein the correlation results include spatial domain correlation results of one basic unit before division and N basic units generated after division, wherein N is an integer greater than 1;
[0007] According to the correlation calculation result, it is determined whether to perform a division operation on the basic unit.
[0008] In an exemplary embodiment, the spatial domain correlation result includes:
[0009] The spatial correlation α between one basic unit before division and the N basic units generated after divisions ; and / or, the spatial correlation β between the N basic units generated after the division s .
[0010] 3. The method according to claim 2, characterized in that:
[0011] The spatial correlation α between the 1 basic unit before the division and the N basic units generated after the division s It is obtained through the following methods, including:
[0012]
[0013] Where N is the total number of basic units generated after division; d is the depth of the basic unit before division; d+1 represents the depth of the basic unit after division; D(X) d Indicates the size of spatial correlation before basic unit division; Represents the spatial correlation size of the i-th basic unit generated after the basic unit division, where i = 1, 2, 3, ..., N;
[0014] The spatial correlation β between the N basic units generated after the division s It is obtained through the following methods, including:
[0015]
[0016] Where N is the total number of basic units generated after division; d is the depth of the basic unit before division; d+1 represents the depth of the basic unit after division; D(X) d Indicates the size of spatial correlation before basic unit division; It represents the spatial correlation size of the i-th basic unit generated after the basic unit division, where i = 1, 2, 3, ..., N.
[0017] In an exemplary embodiment, the correlation calculation result further includes: a time domain correlation between one basic unit before division and N basic units generated after division.
[0018] In an exemplary embodiment, the time domain correlation result includes:
[0019] The temporal correlation α between one basic unit before division and the N basic units generated after division t ; and / or, the temporal correlation β between the N basic units generated after the division t .
[0020] In an exemplary embodiment, the time domain correlation α between the one basic unit before the division and the N basic units generated after the division ist It is obtained through the following methods, including:
[0021]
[0022] Where N is the total number of basic units generated after division; d is the depth of the basic unit before division; d+1 represents the depth of the basic unit after division; D(Y) d Indicates the magnitude of time domain correlation before basic unit division; Represents the time domain correlation size of the i-th basic unit generated after the basic unit is divided, where i = 1, 2, 3, ..., N;
[0023] The temporal correlation β between the N basic units generated after the division t It is obtained through the following methods, including:
[0024]
[0025] Where N is the total number of basic units generated after division; d is the depth of the basic unit before division; d+1 represents the depth of the basic unit after division; D(Y) d Indicates the magnitude of time domain correlation before basic unit division; It represents the time domain correlation size of the i-th basic unit generated after the basic unit is divided, where i = 1, 2, 3, ..., N.
[0026] In an exemplary embodiment, judging whether to perform a division operation on the basic unit according to the correlation calculation result includes:
[0027] Obtaining the video frame type corresponding to the basic unit;
[0028] If the basic unit is an intra-frame video frame, judging whether to perform a division operation on the basic unit according to the spatial domain correlation result;
[0029] If the basic unit is a unidirectional prediction coding frame or a bidirectional prediction coding frame, it is determined whether to perform a division operation on the basic unit according to the spatial domain correlation result and the temporal domain correlation result.
[0030] In an exemplary embodiment, judging whether to perform a division operation on the basic unit according to the correlation calculation result includes:
[0031] If the basic unit is an intra-frame video frame, further judging whether to perform a division operation on the basic unit according to the spatial domain correlation result before the division of the basic unit;
[0032] If the basic unit is a unidirectional prediction coding frame or a bidirectional prediction coding frame, it is determined whether to perform a division operation on the basic unit according to the spatial domain correlation result and the temporal domain correlation result before the division of the basic unit.
[0033] In an exemplary embodiment, after determining whether to perform a division operation on the basic unit according to the correlation calculation result, the method further includes:
[0034] After obtaining the result of whether the basic unit is divided, counting the correlation results of the basic units that perform the division operation and the correlation results of the basic units that do not perform the division operation;
[0035] Based on the correlation results of the basic units that perform the split operation and the correlation results of the basic units that do not perform the split operation, a threshold used for the next execution to determine whether to perform the split operation on the basic unit is determined, wherein the threshold includes a threshold for performing the split operation and / or a threshold for not performing the split operation.
[0036] In an exemplary embodiment, after determining whether to perform a division operation on the basic unit according to the correlation calculation result, the method further includes:
[0037] After determining that the basic unit is not over-divided, calculating the residual information of the basic unit;
[0038] If the obtained residual data meets the preset residual judgment condition, the basic unit is divided.
[0039] The embodiment provided in the embodiment of the present application obtains the correlation calculation results of the HEVC basic unit before and after the division, and judges whether to perform the division operation on the basic unit based on the correlation calculation results, so as to use the correlation results of the basic unit before and after the division as the division decision condition, thereby reducing the complexity of the judgment operation.
[0040] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings are used to provide further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0042] Figure 1A flowchart of a coding management method based on high-efficiency video coding provided in an embodiment of the present application;
[0043] Figure 2 A flowchart of the coding management method for CU depth division provided in Embodiment 1 of the present application;
[0044] Figure 3 A flowchart of a frame encoding method for CU depth division provided in Embodiment 2 of the present application;
[0045] Figure 4 A flowchart of a method for managing CU depth partitioning based on threshold training provided in Embodiment 3 of the present application. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solution and advantages of the present application more clear, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other arbitrarily without conflict.
[0047] The steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions. Also, although a logical sequence is shown in the flowchart, in some cases, the steps shown or described can be performed in a sequence different from that shown here.
[0048] Taking CU as an example, the technical solution of this application is analyzed and explained:
[0049] In HEVC, the size of the coding block (CB) ranges from the smallest 8x8 to the largest 64x64. On the one hand, a large CB can greatly improve the coding efficiency of flat areas, and on the other hand, a small CB can handle local details of the image well, thus making the prediction of complex images more accurate. A luminance component CB and the corresponding chrominance component CB and related syntax elements together constitute a coding unit CU.
[0050] An image can be divided into several non-overlapping CTUs. Within the CTU, a quadtree-based cyclic hierarchical structure is used. Coding units on the same level have the same segmentation depth. A CTU may contain only one CU, that is, no division is performed, or it may be divided into multiple CUs.
[0051] Whether the coding unit is further divided depends on the split flag Split flag. For the coding unit CU d, assuming the size is 2Nx2N, the depth is d, and the corresponding Split flag value is 0, CU d will no longer be quadtree divided. On the contrary, the corresponding Split flag value is 1, and the coding unit CU d will be divided into 4 independent coding units CU d+1.
[0052] The value of the split flag Split flag is determined by calculating the rate-distortion cost of the current CU d and the 4 sub-CU d+1 after division. If the best mode cost of the current CU d is Best Cost d, the sum of the best mode costs of the 4 sub-CU d+1 is Best Cost d+1. If Best Cost d<=Best Cost d+1, the current CU d is not divided, and the corresponding Split flag=0; otherwise, Best Cost d>Best Cost d+1, the current CU d is divided, and the corresponding Splitflag=1.
[0053] In HEVC, intra prediction supports 4 CU sizes: 8x8, 16x16, 32x32 and 64x64. Each CU size has 35 prediction modes for the corresponding PU. Inter prediction uses motion search based on block motion compensation. These two types of prediction are the most time-consuming modules on the encoding side and are also necessary modules for calculating the best mode. In the process of CU division, each time the division is determined, 4+1 intra-frame and inter-frame mode searches are required, which has a high computational complexity.
[0054] The purpose of analyzing CU division is to distinguish and use different sizes of CB for encoding according to the texture complexity of different local areas of the image. For intra-frame mode, the more complex the texture of the coding block, the greater the change in pixel value, and the division tends to be smaller CU. Conversely, the smoother the coding block, the smaller the change in pixel value, and the division tends to be larger CU. For inter-frame mode, the smaller the similarity between the current frame area and the reference frame area of the coding block, the greater the difference in pixel value, and the division tends to be smaller CU. Conversely, the greater the correlation between the current frame area and the reference frame area of the coding block, the smaller the difference in pixel value, and the division tends to be larger CU.
[0055] During the encoding process, the HEVC standard test platform includes some CU division early decision algorithms, such as early termination strategy (Early_CU), early jump out strategy (Early_SKIP) and fast CBF strategy (CBF_Fast). These conditions are relatively strict and can reduce complexity to a limited extent. On this basis,
[0056] According to the above analysis, the inventors perform simple preprocessing before CU division to determine the basic situation of the current CU block. The intra-frame mode is the size of the spatial domain correlation, and the inter-frame mode is the size of the temporal domain correlation. Then, based on this information, they decide whether to divide the CU, that is, to add a CU division decision condition. Therefore, if the CU division method can be predicted in advance, some nodes in the quadtree can be effectively skipped directly, which can significantly reduce the complexity of the encoder.
[0057] Figure 1 A flowchart of a coding management method based on high-efficiency video coding provided in an embodiment of the present application. Figure 1 The methods shown include:
[0058] Step 101, obtaining correlation calculation results of HEVC basic units before and after division, wherein the correlation results include spatial domain correlation results of one basic unit before division and N basic units generated after division, wherein N is an integer greater than 1;
[0059] In an exemplary embodiment, the basic unit may be a coding unit (Coding Unit, CU), a prediction unit (Prediction Unit, PU) or a transform unit (Transform Unit, TU).
[0060] In an exemplary embodiment, the spatial correlation result includes: the spatial correlation α between one basic unit before division and the N basic units generated after division. s ; and / or, the spatial correlation β between the N basic units generated after the division s .
[0061] Step 102: Determine whether to perform a division operation on the basic unit according to the correlation calculation result.
[0062] Different from the division decision conditions of the related technology, the decision conditions provided in the embodiment of the present application are to determine the cost relationship of the division operation based on the correlation before and after the division of the basic unit to decide whether to perform the division. The required computing operation amount is the calculation of the correlation result, which reduces the computational complexity compared with the related technology.
[0063] The method provided in the embodiment of the present application obtains the correlation calculation results of the HEVC basic unit before and after the division, and judges whether to perform the division operation on the basic unit based on the correlation calculation results, so as to use the correlation results of the basic unit before and after the division as the division decision condition, thereby reducing the complexity of the judgment operation.
[0064] The method provided in the embodiment of the present application is further described below:
[0065] In an exemplary embodiment, the spatial correlation α between the one basic unit before the division and the N basic units generated after the division is s It is obtained through the following methods, including:
[0066]
[0067] Where N is the total number of basic units generated after division; d is the depth of the basic unit before division; d+1 represents the depth of the basic unit after division; D(X) d Indicates the size of spatial correlation before basic unit division; It represents the spatial correlation size of the i-th basic unit generated after the basic unit division, where i = 1, 2, 3, ..., N.
[0068] In an exemplary embodiment, the spatial correlation β between the N basic units generated after the division is s It is obtained through the following methods, including:
[0069]
[0070] Where N is the total number of basic units generated after division; d is the depth of the basic unit before division; d+1 represents the depth of the basic unit after division; D(X) d Indicates the size of spatial correlation before basic unit division; It represents the spatial correlation size of the i-th basic unit generated after the basic unit division, where i = 1, 2, 3, ..., N.
[0071] In an exemplary embodiment, the correlation calculation result further includes: a time domain correlation between one basic unit before division and N basic units generated after division.
[0072] In this exemplary embodiment, after the intra-frame mode of the basic unit is determined by spatial domain correlation, the inter-frame mode of the basic unit is determined by temporal domain correlation, and the division of the basic unit is judged.
[0073] In an exemplary embodiment, the time domain correlation result is obtained by:
[0074] The temporal correlation α between one basic unit before division and the N basic units generated after division t ; and / or, the temporal correlation β between the N basic units generated after the division t .
[0075] In an exemplary embodiment, the time domain correlation α between the one basic unit before the division and the N basic units generated after the division ist It is obtained through the following methods, including:
[0076]
[0077] Where N is the total number of basic units generated after division; d is the depth of the basic unit before division; d+1 represents the depth of the basic unit after division; D(Y) d Indicates the magnitude of time domain correlation before basic unit division; It represents the time domain correlation size of the i-th basic unit generated after the basic unit is divided, where i = 1, 2, 3, ..., N.
[0078] In an exemplary embodiment, the temporal correlation β between the N basic units generated after the division is t It is obtained through the following methods, including:
[0079]
[0080] Where N is the total number of basic units generated after division; d is the depth of the basic unit before division; d+1 represents the depth of the basic unit after division; D(Y) d Indicates the magnitude of time domain correlation before basic unit division; It represents the time domain correlation size of the i-th basic unit generated after the basic unit is divided, where i = 1, 2, 3, ..., N.
[0081] In an exemplary embodiment, judging whether to perform a division operation on the basic unit according to the correlation calculation result includes:
[0082] Obtaining the video frame type corresponding to the basic unit;
[0083] If the basic unit is an intra-frame video frame, judging whether to perform a division operation on the basic unit according to the spatial domain correlation result;
[0084] If the basic unit is a unidirectional prediction coding frame or a bidirectional prediction coding frame, it is determined whether to perform a division operation on the basic unit according to the spatial domain correlation result and the temporal domain correlation result.
[0085] In this exemplary embodiment, the type of video frame corresponding to the basic unit is identified to determine the correlation result to be used, and the calculation scale of the correlation is effectively controlled while ensuring the implementation of the division operation.
[0086] In an exemplary embodiment, judging whether to perform a division operation on the basic unit according to the correlation calculation result includes:
[0087] If the basic unit is an intra-frame video frame (I frame), whether to perform a division operation on the basic unit is also determined according to the spatial domain correlation result before the basic unit is divided;
[0088] If the basic unit is a unidirectional predictive coding frame or a bidirectional predictive coding frame (P frame or B frame), it is determined whether to perform a division operation on the basic unit according to the spatial domain correlation result and the temporal domain correlation result before the division of the basic unit.
[0089] In this exemplary embodiment, by comparing the current spatial or temporal correlation values of the basic units and combining the obtained correlation results before and after the division as the conditions for the division decision, the accuracy of the judgment can be effectively improved.
[0090] In an exemplary embodiment, after determining whether to perform a division operation on the basic unit according to the correlation calculation result, the method further includes:
[0091] After obtaining the result of whether the basic unit is divided, counting the correlation results of the basic units that perform the division operation and the correlation results of the basic units that do not perform the division operation;
[0092] Based on the correlation results of the basic units that perform the split operation and the correlation results of the basic units that do not perform the split operation, a threshold used for the next execution to determine whether to perform the split operation on the basic unit is determined, wherein the threshold includes a threshold for performing the split operation and / or a threshold for not performing the split operation.
[0093] In this exemplary embodiment, the threshold value may be recalculated at regular intervals or when the application scenario changes, so as to make a more accurate judgment.
[0094] In an exemplary embodiment, after determining whether to perform a division operation on the basic unit according to the correlation calculation result, the method further includes:
[0095] After determining that the basic unit is not over-divided, calculating the residual information of the basic unit;
[0096] If the obtained residual data meets the preset residual judgment condition, the basic unit is divided.
[0097] In this exemplary embodiment, for a basic unit that is determined not to be subjected to a split operation, a residual of the basic unit is calculated, and then it is determined whether to perform a split operation, thereby improving the accuracy of encoding.
[0098] The following is an explanation of the method provided in the embodiment of the present application:
[0099] Example 1
[0100] Take the best mode cost of CU d as Best Cost d, which is divided into 4 sub-CUs d+1 as an example for explanation:
[0101] In the related art, during the encoding process, the way of dividing CU leads to high computational complexity at the encoding end.
[0102] The exemplary embodiment of the present application proposes an intelligent coding method for fast CU depth division based on HEVC correlation information. Before CUd division, the preprocessing first obtains the spatial correlation and temporal correlation information of CUd, and then obtains the spatial correlation and temporal correlation information of CUd+1 after division, establishes the cost relationship between the two, and makes intelligent CU division decisions in advance, so as to achieve the purpose of making full use of the correlation of video content and significantly reduce the coding complexity.
[0103] Figure 2 This is a flowchart of the encoding management method for CU depth partitioning provided in Example 1 of the present application. Figure 2 The method shown comprises the following steps:
[0104] Step 201, first obtain the spatial correlation of the current CU, and calculate the spatial correlation of the 1 coding unit before and the 4 coding units after the CU is divided.
[0105] Step 202, then obtain the time domain correlation of the current CU, and calculate the time domain correlation of the 1 coding unit before and the 4 coding units after the CU is divided.
[0106] Step 203, finally determining whether the current CU is to be divided, includes:
[0107] For the I frame, it can be judged only through the spatial domain correlation. If the correlation before division is greater than the correlation after division, no division is performed. Otherwise, it is determined to perform a division operation on the basic unit.
[0108] For the P / B frame, only if both the spatial domain correlation and the temporal domain correlation satisfy that the correlation before division is greater than the correlation after division, no division is performed; otherwise, it is determined that the division operation is performed on the basic unit.
[0109] The method provided in the first embodiment of the present application can improve the encoding speed of video images.
[0110] Example 2
[0111] The processing of CU d is described by taking the HM encoder and IPPP encoding structure as an application scenario:
[0112] Figure 3This is a flowchart of the frame encoding management method for CU depth division provided in Example 2 of the present application. Figure 3 The method shown comprises the following steps: Step 301, inputting a coded video image;
[0113] The video image is a video image waiting for encoding processing, and may be a video sequence;
[0114] Step 302: Input the coding unit as an object to be processed for division decision;
[0115] Step 303: Calculate CU spatial domain correlation;
[0116] First, the spatial correlation size of the first and fourth coding units before CU division is defined, and the spatial correlation is defined by the variance size within the coding unit.
[0117] Then, the spatial correlation size before the CU partition with a depth of d is calculated, and the average value is calculated, as shown in formula (1).
[0118]
[0119] Where n represents the number of pixels contained in the coding unit, and xi represents the pixel value. The variance is calculated as shown in formula (2).
[0120]
[0121] In this way, the spatial correlation size of the current CU is obtained, which is recorded as D(X) CU=d In the same way, the spatial correlation of the four coding units after division is calculated and recorded as and Define the perception factor of the correlation between the two spatial domains, as shown in equations (3) and (4).
[0122]
[0123]
[0124] in, Indicates the correlation between 1 CUd before the division and 4 CUd+1 after the division. The smaller the value, the greater the correlation within CUd. Indicates the correlation between the 4 CUd+1s after division. The smaller the value, the greater the correlation between the sub-CUd+1s.
[0125] Step 304: Calculate CU time domain correlation;
[0126] First, the temporal correlation of the first and fourth coding units before CU division is defined, and the temporal correlation is defined by the variance of the coding unit and the reference unit.
[0127] Then, the time domain correlation size before the CU partition with a depth of d is calculated. The reference unit obtains the spatial domain candidate list similar to the Merge mode in the MV prediction technology. The MVP with the highest priority is obtained according to the spatial domain information of the current coding unit, and then rounded, and the corresponding area of the reference frame is shifted by integer pixels to obtain the reference unit.
[0128] Calculate the variance of the coding unit and the reference unit, as shown in formula (5).
[0129]
[0130] Where n represents the number of pixels contained in the coding unit, yc represents the pixel value of the coding unit, and yr represents the pixel value of the reference unit. The temporal correlation size of the current CU is obtained, which is recorded as D(Y) CU=d In the same way, the time domain correlation of the four coding units after division is calculated and recorded as and Define two time domain correlation perception factors as shown in equations (6) and (7).
[0131]
[0132]
[0133] in, Indicates the correlation between 1 CUd before the division and 4 CUd+1 after the division. The smaller the value, the greater the correlation within CUd. Indicates the correlation between the 4 CUd+1s after division. The smaller the value, the greater the correlation between the sub-CUd+1s.
[0134] Step 305: perform quadtree partitioning according to the correlation information;
[0135] Based on the spatial domain and temporal domain correlations of the current coding unit obtained in step 303 and step 304, judgment is performed, including:
[0136] For I frames, it can be judged only by spatial correlation;
[0137] For P / B frames, judgment is made based on spatial domain correlation and temporal domain correlation.
[0138]
[0139]
[0140] When equation (8a) is satisfied, the CU is not divided, and when equation (9a) is satisfied, the CU is divided.
[0141] Here They are all thresholds and their values can be different. The first 6 represent no threshold division, recorded as data 1, and the next 6 represent the division threshold, recorded as data 2;
[0142] Step 306: perform encoding operation according to the judgment result;
[0143] Step 307: determine whether the encoding of the current frame is completed;
[0144] If it is finished, execute step 308, otherwise, execute step 309 to obtain the next CU;
[0145] Step 308, determine whether the current sequence has been encoded;
[0146] If it is finished, the process ends, otherwise, step 310 is executed to obtain the next video frame.
[0147] Example 3
[0148] The processing of CU d is described by taking the HM encoder and IPPP encoding structure as an application scenario:
[0149] Figure 4 A flowchart of a method for managing CU depth partitioning based on threshold training provided in Embodiment 3 of the present application. Figure 4 The method shown comprises the following steps:
[0150] Step 401: input a coded video image;
[0151] The video image is a video image waiting for encoding processing, and may be a video sequence;
[0152] Step 402: Input the coding unit as an object to be processed for division decision;
[0153] Step 403: Calculate CU spatial domain correlation;
[0154] First, the spatial correlation size of the first and fourth coding units before CU division is defined, and the spatial correlation is defined by the variance size within the coding unit.
[0155] Then, the spatial correlation size before the CU partition with a depth of d is calculated, and the average value is calculated, as shown in formula (1).
[0156]
[0157] Where n represents the number of pixels contained in the coding unit, and xi represents the pixel value. The variance is calculated as shown in formula (2).
[0158]
[0159] In this way, the spatial correlation size of the current CU is obtained, which is recorded as D(X) CU=d In the same way, the spatial correlation of the four coding units after division is calculated and recorded as and Define the perception factor of the correlation between the two spatial domains, as shown in equations (3) and (4).
[0160]
[0161]
[0162] in, Indicates the correlation between 1 CUd before the division and 4 CUd+1 after the division. The smaller the value, the greater the correlation within CUd. Indicates the correlation between the 4 CUd+1s after division. The smaller the value, the greater the correlation between the sub-CUd+1s.
[0163] Step 404: Calculate CU time domain correlation;
[0164] First, the temporal correlation of the first and fourth coding units before CU division is defined, and the temporal correlation is defined by the variance of the coding unit and the reference unit.
[0165] Then, the time domain correlation size before the CU partition with a depth of d is calculated. The reference unit obtains the spatial domain candidate list similar to the Merge mode in the MV prediction technology. The MVP with the highest priority is obtained according to the spatial domain information of the current coding unit, and then rounded, and the corresponding area of the reference frame is shifted by integer pixels to obtain the reference unit.
[0166] Calculate the variance of the coding unit and the reference unit, as shown in formula (5).
[0167]
[0168] Where n represents the number of pixels contained in the coding unit, yc represents the pixel value of the coding unit, and yr represents the pixel value of the reference unit. The temporal correlation size of the current CU is obtained, which is recorded as D(Y) CU=d In the same way, the time domain correlation of the four coding units after division is calculated and recorded as and Define two time domain correlation perception factors as shown in equations (6) and (7).
[0169]
[0170]
[0171] in, Indicates the correlation between 1 CUd before the division and 4 CUd+1 after the division. The smaller the value, the greater the correlation within CUd. Indicates the correlation between the 4 CUd+1s after division. The smaller the value, the greater the correlation between the sub-CUd+1s.
[0172] Step 405: perform quadtree partitioning according to the correlation information;
[0173] Based on the spatial domain and temporal domain correlations of the current coding unit obtained in step 403 and step 404, judgment is performed, including:
[0174] For I frames, it can be judged through spatial correlation.
[0175] For P / B frames, judgment is made through spatial domain correlation and temporal domain correlation.
[0176] In order to prevent extreme conditions, when making a partition decision for a CU with a depth of d, a CU correlation size limit is added separately, as shown in equations (8) and (9).
[0177]
[0178]
[0179] Among them, when the formula (8b) is satisfied, the CU is not divided, and when the formula (9b) is satisfied, the CU is divided. They are all thresholds and their values can be different. The first 6 represent no threshold division, recorded as data 1, and the last 6 represent the threshold division, recorded as data 2.
[0180] There will be two processes next, a threshold training process and an intelligent encoding process.
[0181] In order to obtain these 12 thresholds, a training process is required. Start with N frames of images, such as N = 100, and follow the conventional CU division process. When the depth is d and the CU is not divided, the statistics of 1 are counted separately. Distribution, as long as it meets most of the conditions, such as 80%, it meets formula (8) and obtains the corresponding threshold. When the depth is d, when CU is divided, the statistical data 2 are respectively The distribution, as long as it meets most of the cases, such as 80%, meets formula (9), the corresponding threshold is obtained.
[0182] In the subsequent encoding process, it can be updated at regular intervals or when the scene is switched. In the remaining time periods, the above method can be used to make CU division decisions of different depths d to save encoding time. In addition, in order to prevent extreme conditions, such as CU not being divided, the current coding unit residual can also be judged. If the codeword is particularly large, division is forced.
[0183] Step 406: perform encoding operation according to the judgment result;
[0184] Step 407: determine whether the encoding of the current frame is completed;
[0185] If it is finished, execute step 408, otherwise, execute step 409 to obtain the next CU;
[0186] Step 408, determine whether the current sequence has been encoded;
[0187] If it is finished, the process ends, otherwise, step 410 is executed to obtain the next video frame.
[0188] In the above-mentioned example embodiments one to three, the method of calculating correlation is not limited to variance, but can also be mean square error, covariance, Hadamard transform coefficient size, sine / cosine transform size of residual coefficient, Sobel gradient size, etc., one or a combination of several of them.
[0189] The acquisition of correlation information is not limited to the first and second steps above, and can be obtained directly from the pixel value samples of the coding unit through deep learning or machine learning.
[0190] CU partitioning decision can also be extended to PU and TU partitioning decision.
[0191] The pixel value can be the Y brightness component, the U / V chrominance component, or a combination thereof.
[0192] The training process of threshold acquisition is not limited to conventional methods, and deep learning or machine learning can be used.
[0193] The method provided in the embodiment of the present application is not limited to H.265, but can also be applied to other video coding standards such as H.264, H.266, AV1, VP8, VP9, AVS2, AVS3, etc.
[0194] An embodiment of the present application provides a HEVC-based encoding management device, including a processor and a memory, wherein the memory stores a computer program, and the processor calls the computer program in the memory to implement any of the methods described above.
[0195] The device provided in the embodiment of the present application obtains the correlation calculation results of the HEVC basic unit before and after the division, and judges whether to perform the division operation on the basic unit based on the correlation calculation results, so as to use the correlation results of the basic unit before and after the division as the division decision condition, thereby reducing the complexity of the judgment operation.
[0196] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
Claims
1. A coding management method based on high efficiency video coding HEVC, include: Obtaining correlation results of HEVC basic units before and after division, wherein the correlation results include spatial domain correlation results of one basic unit before division and N basic units generated after division, wherein N is an integer greater than 1; According to the correlation result, determining whether to perform a division operation on the basic unit; The spatial domain correlation results include: The spatial correlation α between one basic unit before division and the N basic units generated after division s , and the spatial correlation β between the N basic units generated after division s ; Among them, the spatial correlation α between the 1 basic unit before the division and the N basic units generated after the division is s It is obtained through the following methods, including: Among them, N is the total number of basic units generated after division; d is the depth before the division of the basic unit; d + 1 represents the depth after the basic unit is divided; D(X) d represents the size of the spatial domain correlation before the division of the basic unit; represents the size of the spatial domain correlation of the i-th basic unit generated after the division of the basic unit, where i = 1, 2, 3,..., N; The spatial correlation β between the N basic units generated after the division s It is obtained through the following methods, including:
2. The method according to claim 1, It is characterized in that The correlation result also includes: a time domain correlation result between one basic unit before division and N basic units generated after division.
3. The method according to claim 2, It is characterized in that The time domain correlation results include: The temporal correlation α between one basic unit before partitioning and the N basic units generated after partitioning t , and the temporal correlation β between the N basic units generated after partitioning t ; The time domain correlation α between the one basic unit before the division and the N basic units generated after the division is t It is obtained through the following methods, including: Where N is the total number of basic units generated after division; d is the depth of the basic unit before division; d+1 represents the depth of the basic unit after division; D(Y) d Indicates the magnitude of time domain correlation before basic unit division; Represents the time domain correlation size of the i-th basic unit generated after the basic unit is divided, where i = 1, 2, 3, ..., N; The time domain correlation β between the N basic units generated after the division t It is obtained through the following methods, including:
4. The method according to any one of claims 1 to 3, It is characterized in that The step of judging whether to perform a division operation on the basic unit according to the correlation result includes: Obtaining the video frame type corresponding to the basic unit; If the basic unit is an intra-frame video frame, judging whether to perform a division operation on the basic unit according to the spatial domain correlation result; If the basic unit is a unidirectional prediction coding frame or a bidirectional prediction coding frame, it is determined whether to perform a division operation on the basic unit according to the spatial domain correlation result and the temporal domain correlation result.
5. The method according to claim 4, It is characterized in that The step of judging whether to perform a division operation on the basic unit according to the correlation result includes: If the basic unit is an intra-frame video frame, further judging whether to perform a division operation on the basic unit according to the spatial domain correlation result before the division of the basic unit; If the basic unit is a unidirectional prediction coding frame or a bidirectional prediction coding frame, it is determined whether to perform a division operation on the basic unit according to the spatial domain correlation result and the temporal domain correlation result before the division of the basic unit.
6. The method according to claim 5, It is characterized in that After determining whether to perform a division operation on the basic unit according to the correlation result, the method further includes: After obtaining the result of whether the basic unit is divided, counting the correlation results of the basic units that perform the division operation and the correlation results of the basic units that do not perform the division operation; Based on the correlation results of the basic units that perform the split operation and the correlation results of the basic units that do not perform the split operation, a threshold used for the next execution to determine whether to perform the split operation on the basic unit is determined, wherein the threshold includes a threshold for performing the split operation and / or a threshold for not performing the split operation.
7. The method according to claim 1, It is characterized in that After determining whether to perform a division operation on the basic unit according to the correlation result, the method further includes: After determining that the basic unit is not over-divided, calculating the residual information of the basic unit; If the obtained residual data meets the preset residual judgment condition, the basic unit is divided.
Citation Information
Patent Citations
Interframe mode fast selecting method and system of video compressed coding
CN106454342A