Video Coding Block Partitioning Method and Video Coding Block Partitioning Prediction Model Training Method

By using the video coded block division prediction model in video coded block division, the probability of CTU-level sub-block boundary division is predicted and the probability value of the CU-level block division mode is calculated, the problem of video coded block division taking too long in the inter prediction mode is solved, and the effect of shortening time and reducing encoding loss is achieved.

CN114173120BActive Publication Date: 2025-06-13BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111464395.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-06-13
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

Video encoding block division is highly complex in inter-frame prediction mode, resulting in excessive time.

Method used

By obtaining the encoding information of the current CTU and its isometric CTU, the video coded block division prediction model predicts the probability of sub-block boundary division at the CTU level, and calculates the probability value of the block division mode at the CU level, removing the block division mode that does not meet the preset threshold, and only traversing the remaining modes to perform rate distortion optimization decisions.

Benefits of technology

It significantly shortens the time of video encoding block division, reduces encoding loss, and improves encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114173120B_ABST
    Figure CN114173120B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method for video coding block partitioning in an inter prediction mode and a method for training a video coding block partitioning prediction model. The method for video coding block partitioning in the inter prediction mode includes: obtaining a first prediction residual and a second prediction residual; inputting the first prediction residual and the second prediction residual into the video coding block partitioning prediction model to obtain a first probability, where the first probability is the probability that each side of each minimum sub-block of the CTU to be coded serves as a partitioning boundary of all possible block partitioning modes; for each CU of the CTU to be coded, perform the following operations: based on the first probability, determine a second probability corresponding to when the current CU is partitioned according to each possible block partitioning mode, remove the block partitioning modes whose second probabilities do not meet a preset threshold from all possible block partitioning modes, and perform a rate-distortion optimization decision on the remaining block partitioning modes to obtain the block partitioning mode of the current CU.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of audio - video processing, and more particularly, to a method and apparatus for video coding block partitioning in an inter - prediction mode, and a method and apparatus for training a video coding block partitioning prediction model. Background Art

[0002] Currently, most video codings adopt a block - based hybrid coding method. Each frame of a video is first divided into multiple coding units in units of blocks for intra - frame or inter - frame prediction, then the prediction residuals are transformed and quantized, and finally, the mode information, the quantized residuals, etc. are entropy - coded to obtain an encoded bitstream. In order to adapt to a variety of video contents and video characteristics, in the latest video coding standard H.266 / VVC, a combined partitioning method of QT (Quadro Tree) and MTT (Multi - Type Tree) is adopted for block partitioning, which can significantly improve the coding efficiency, but also leads to a significant increase in computational complexity, and further leads to an excessive time consumption for video coding block partitioning. Summary of the Invention

[0003] The present disclosure provides a method and apparatus for video coding block partitioning in an inter - prediction mode, and a method and apparatus for training a video coding block partitioning prediction model, so as to at least solve the problems in the above - related technologies.

[0004] According to a first aspect of an embodiment of the present disclosure, a method for video coding block partitioning in an inter - prediction mode is provided, including: obtaining a first prediction residual and a second prediction residual, where the first prediction residual is the prediction residual of a CTU to be encoded in a current video frame, and the second prediction residual is obtained by performing a preset process on the first prediction residual according to the block partitioning mode of a co - located CTU of the CTU to be encoded, and the co - located CTU is the corresponding part of the reference frame of the current video frame corresponding to the CTU to be encoded; inputting the first prediction residual and the second prediction residual into a video coding block partitioning prediction model and obtaining a first probability, where the first probability is the probability that each side of each minimum sub - block of the CTU to be encoded serves as a partitioning boundary of all possible block partitioning modes; for each CU of the CTU to be encoded, perform the following operations: based on the first probability, determine a second probability corresponding to when the current CU is partitioned according to each possible block partitioning mode, remove the block partitioning modes whose second probabilities do not meet a preset threshold from all possible block partitioning modes, and perform a rate - distortion optimization decision on the remaining block partitioning modes to obtain the block partitioning mode of the current CU.

[0005] Optionally, the preset processing of the first prediction residual according to the block partitioning mode of the co-located CTU of the CTU to be encoded includes: taking an average value of the first prediction residual according to the block partitioning mode of the co-located CTU of the CTU to be encoded; using the first prediction residual after taking the average value as the second prediction residual.

[0006] Optionally, the block partitioning mode includes a quadtree partitioning mode, a ternary tree partitioning mode, and a binary tree partitioning mode, and both the binary tree partitioning mode and the ternary tree partitioning mode each include a horizontal partitioning mode and a vertical partitioning mode; removing the block partitioning modes whose second probability does not meet the preset threshold from all possible block partitioning modes includes: respectively removing the horizontal partitioning mode and / or the vertical partitioning mode whose second probability does not meet the preset threshold from the binary tree partitioning mode and the ternary tree partitioning mode.

[0007] Optionally, the preset threshold includes a first preset threshold and a second preset threshold. Taking the second probability corresponding to the horizontal partitioning mode as the horizontal probability and taking the second probability corresponding to the vertical partitioning mode as the vertical probability; respectively removing the horizontal partitioning mode and / or the vertical partitioning mode whose second probability does not meet the preset threshold from the binary tree partitioning mode and the ternary tree partitioning mode includes: for the binary tree partitioning mode or the ternary tree partitioning mode, performing the following operations: respectively comparing the horizontal probability and the vertical probability with the first preset threshold;

[0008] In the case where both the horizontal probability and the vertical probability are less than the first preset threshold, removing the horizontal partitioning mode and the vertical partitioning mode; in the case where one of the horizontal probability and the vertical probability is less than the first preset threshold and the other is greater than or equal to the first preset threshold, removing the block partitioning mode with a probability value less than the first preset threshold among the horizontal probability and the vertical probability; in the case where both the horizontal probability and the vertical probability are greater than or equal to the second preset threshold, taking the difference between the horizontal probability and the vertical probability, and in the case where the absolute value of the difference is greater than or equal to the second preset threshold, removing the block partitioning mode with a smaller probability value among the horizontal probability and the vertical probability.

[0009] According to a second aspect of an embodiment of the present disclosure, a training method for a video coding block partitioning prediction model is provided, comprising: obtaining a video training sample, wherein the video training sample comprises a first prediction residual, a second prediction residual and a true value vector, the first prediction residual is a prediction residual of a first CTU to be encoded in a current video frame, the second prediction residual is obtained by pre-setting the first prediction residual according to a block partitioning mode of a second CTU in a reference frame of the current video frame, the second CTU is a co-located CTU of the first CTU, and the true value vector represents each edge of each minimum sub-block in the first CTU as a real block The invention relates to a method for preparing a video coding block partition prediction model and a method for preparing a video coding block partition prediction model. The method comprises the steps of: inputting the first prediction residual and the second prediction residual into the video coding block partition prediction model to obtain an estimated vector composed of the probability that each edge of each minimum sub-block in the first CTU is a partition boundary of all possible block partition patterns; superimposing corresponding weights on the probabilities of preset edges of multiple minimum sub-blocks in the first CTU; calculating the value of a loss function based on the estimated vector superimposed with corresponding weights and the true value vector; and training the video coding block partition prediction model by adjusting the model parameters of the video coding block partition prediction model according to the value of the loss function.

[0010] Optionally, the first prediction residual is preset according to the block division mode of the second CTU in the reference frame of the current video frame, including: taking the average of the first prediction residual according to the block division mode of the second CTU in the reference frame of the current video frame; and using the first prediction residual after taking the average as the second prediction residual.

[0011] Optionally, the probability of preset edges of multiple minimum sub-blocks in the first CTU is superimposed with corresponding weights, including: multiplying the probability of edges of multiple minimum sub-blocks as preset partition boundaries by corresponding weights to enhance the sensitivity of the video coding block partition prediction model to video coding loss, wherein the value of the weight is greater than 1, the preset partition boundary is obtained by quadtree partitioning the first CTU, and the preset partition boundary does not include the edge of the first CTU.

[0012] Optionally, the video coding block partition prediction model is obtained by pre-setting the intra-frame prediction mode block partition prediction network structure, and the intra-frame prediction mode block partition prediction network structure is implemented by a ResNet structure; the pre-setting the intra-frame prediction mode block partition prediction network structure includes: reducing the number of convolutional layers of the intra-frame prediction mode block partition prediction network structure to half of the original.

[0013] According to a third aspect of the embodiments of the present disclosure, there is provided a video coding block partitioning device in an inter prediction mode, including: a residual obtaining unit configured to obtain a first prediction residual and a second prediction residual, where the first prediction residual is the prediction residual of a CTU to be encoded in a current video frame, and the second prediction residual is obtained by performing a preset process on the first prediction residual according to the block partitioning mode of a co-located CTU of the CTU to be encoded, and the co-located CTU is a part corresponding to the CTU to be encoded in a reference frame of the current video frame; a probability obtaining unit configured to input the first prediction residual and the second prediction residual into a video coding block partitioning prediction model and obtain a first probability, where the first probability is the probability that each side of each minimum sub-block of the CTU to be encoded serves as a partitioning boundary of all possible block partitioning modes; a processing unit configured to, for each CU of the CTU to be encoded, perform the following operations: based on the first probability, determine a second probability corresponding to when the current CU is partitioned according to each possible block partitioning mode, remove from all possible block partitioning modes the block partitioning modes for which the second probability does not meet a preset threshold, and perform a rate-distortion optimization decision on the remaining block partitioning modes to obtain the block partitioning mode of the current CU.

[0014] Optionally, the residual obtaining unit is configured to: take an average value of the first prediction residual according to the block partitioning mode of the co-located CTU of the CTU to be encoded; and use the first prediction residual after taking the average value as the second prediction residual.

[0015] Optionally, the block partitioning modes include a quadtree partitioning mode, a ternary tree partitioning mode, and a binary tree partitioning mode, and both the binary tree partitioning mode and the ternary tree partitioning mode respectively include a horizontal partitioning mode and a vertical partitioning mode; the processing unit is configured to respectively remove from the binary tree partitioning mode and the ternary tree partitioning mode the horizontal partitioning mode and / or the vertical partitioning mode for which the second probability does not meet a preset threshold.

[0016] Optionally, the preset threshold includes a first preset threshold and a second preset threshold, and the second probability corresponding to the horizontal division mode is used as the horizontal probability, and the second probability corresponding to the vertical division mode is used as the vertical probability; the processing unit is configured to: for the binary tree division mode or the ternary tree division mode, perform the following operations: compare the horizontal probability and the vertical probability with the first preset threshold respectively; when the horizontal probability and the vertical probability are both less than the first preset threshold, remove the horizontal division mode and the vertical division mode; when one of the horizontal probability and the vertical probability is less than the first preset threshold and the other is greater than or equal to the first preset threshold, remove the block division mode whose probability value is less than the first preset threshold in the horizontal probability and the vertical probability; when the horizontal probability and the vertical probability are both greater than or equal to the second preset threshold, subtract the horizontal probability and the vertical probability, and when the absolute value of the difference is greater than or equal to the second preset threshold, remove the block division mode whose probability value is smaller in the horizontal probability and the vertical probability.

[0017] According to a fourth aspect of an embodiment of the present disclosure, a training device for a video coding block partition prediction model is provided, comprising: a sample acquisition unit, configured to: acquire a video training sample, wherein the video training sample comprises a first prediction residual, a second prediction residual and a true value vector, the first prediction residual is a prediction residual of a first CTU to be encoded in a current video frame, the second prediction residual is obtained by pre-setting the first prediction residual according to a block partition mode of a second CTU in a reference frame of the current video frame, the second CTU is a co-located CTU of the first CTU, and the true value vector represents a situation where each edge of each minimum sub-block in the first CTU serves as a partition boundary of a true block partition mode; an estimation vector acquisition unit , configured to: input the first prediction residual and the second prediction residual into the video coding block partition prediction model to obtain an estimated vector composed of the probability of each edge of each minimum sub-block in the first CTU as the partition boundary of all possible block partition modes; a weight superposition unit, configured to: superimpose corresponding weights on the probabilities of preset edges of multiple minimum sub-blocks in the first CTU; a loss function calculation unit, configured to: calculate the value of the loss function based on the estimated vector superimposed with corresponding weights and the true value vector; a model parameter adjustment unit, configured to: train the video coding block partition prediction model by adjusting the model parameters of the video coding block partition prediction model according to the value of the loss function.

[0018] Optionally, the sample acquisition unit is configured to: average the first prediction residual according to the block partitioning mode of the second CTU in the reference frame of the current video frame; and use the first prediction residual after averaging as the second prediction residual.

[0019] Optionally, the weight superposition unit is configured to: multiply the probabilities of the edges of multiple minimum sub-blocks that are preset partitioning boundaries by corresponding weights to enhance the sensitivity of the video coding block partitioning prediction model to video coding loss, where the value of the weight is greater than 1, the preset partitioning boundary is obtained by performing a quadtree partitioning on the first CTU, and the preset partitioning boundary does not include the edges of the first CTU.

[0020] Optionally, the video coding block partitioning prediction model is obtained by performing preset processing on an intra prediction mode block partitioning prediction network structure, and the intra prediction mode block partitioning prediction network structure is implemented by a ResNet structure; the preset processing on the intra prediction mode block partitioning prediction network structure includes: reducing the number of convolutional layers of the intra prediction mode block partitioning prediction network structure to half of the original.

[0021] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: at least one processor; at least one memory storing computer-executable instructions, where when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the video coding block partitioning method or the training method of the video coding block partitioning prediction model according to the present disclosure in an inter prediction mode.

[0022] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing instructions, which when run by at least one processor, cause the at least one processor to execute the video coding block partitioning method or the training method of the video coding block partitioning prediction model according to the present disclosure in an inter prediction mode.

[0023] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer program product, where the instructions in the computer program product can be executed by a processor of a computer device to complete the video coding block partitioning method or the training method of the video coding block partitioning prediction model according to the present disclosure in an inter prediction mode.

[0024] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0025] The video coding block partitioning method and apparatus in the inter-frame prediction mode according to the present disclosure are based on the coding information of the current CTU and the coding information of its co-located CTU, and predict the boundary partitioning probabilities of each sub-block in the current CTU at the CTU level through a video coding block partitioning prediction model, and calculate the corresponding CU partitioning mode probability values at the CU level using these partitioning probabilities. According to the CU partitioning mode probability values, a part of the possible block partitioning modes are removed, and only the remaining possible block partitioning modes need to be traversed to perform the rate-distortion optimization decision, thereby avoiding the huge computational amount brought by traversing all possible block partitioning modes to perform the rate-distortion optimization decision, and thus greatly shortening the time of video coding block partitioning.

[0026] In addition, for the training method and apparatus of the video coding block partitioning prediction model according to the present disclosure, when training the video coding block partitioning prediction model, corresponding weights can be superimposed on the boundary probabilities of the predicted partitioned sub-blocks according to the coding loss sensitivity, so that when the video coding block partitioning prediction model performs video coding block partitioning prediction, the block partitioning modes with higher sensitivity to the coding loss can be retained to a greater extent, thereby reducing the coding loss.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0029] Figure 1 It is a schematic structural diagram showing a block partitioning mode according to an exemplary embodiment of the present disclosure.

[0030] Figure 2 It is a flowchart showing a video coding block partitioning method in the inter-frame prediction mode according to an exemplary embodiment of the present disclosure.

[0031] Figure 3 It is a schematic diagram showing the process of obtaining a second prediction residual according to an exemplary embodiment of the present disclosure.

[0032] Figure 4 (a) It is a schematic diagram showing a probability vector according to an exemplary embodiment of the present disclosure.

[0033] Figure 4 (b) It is a schematic diagram showing the process of obtaining a truth value vector according to an exemplary embodiment of the present disclosure.

[0034] Figure 5It is a schematic diagram showing the determination of a second probability corresponding to a horizontal block partitioning pattern based on a first probability according to an exemplary embodiment of the present disclosure.

[0035] Figure 6 It is a flowchart showing a method for training a video coding block partitioning prediction model according to an exemplary embodiment of the present disclosure.

[0036] Figure 7 It is a schematic diagram showing the structure of a video coding block partitioning prediction model according to an exemplary embodiment of the present disclosure.

[0037] Figure 8 It is a schematic diagram showing a partitioning boundary obtained by performing a quadtree partitioning on a first CTU according to an exemplary embodiment of the present disclosure.

[0038] Figure 9 It is a block diagram showing a video coding block partitioning device in an inter-frame prediction mode according to an exemplary embodiment of the present disclosure.

[0039] Figure 10 It is a block diagram showing a training device for a video coding block partitioning prediction model according to an exemplary embodiment of the present disclosure.

[0040] Figure 11 It is a block diagram showing an electronic device 1100 according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0041] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0043] It should be noted here that "at least one of a number of items" in this disclosure means that it includes three parallel cases: "any one of the number of items", "a combination of any multiple of the number of items", and "all of the number of items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example, "executing at least one of step one and step two" means the following three parallel cases: (1) executing step one; (2) executing step two; (3) executing step one and step two.

[0044] In the video coding standard H.265 / HEVC, a series of CUs (Coding Units) are iteratively partitioned from a CTU (Coding Tree Unit) using a QT structure. To adapt to diverse video content and video characteristics, the latest video coding standard H.266 / VVC currently adopts a combined partitioning method of QT and MTT. MTT includes two partitioning methods: BT (Binary Tree) and TT (Ternary Tree). After the CTU is partitioned according to the QT structure, the leaf nodes are further partitioned by MTT, and the partitioning types include four types: horizontal BT, vertical BT, horizontal TT, and vertical TT. Figure 1 is a structural schematic diagram showing a block partitioning mode according to an exemplary embodiment of the present disclosure, which can be referred to Figure 1 to clarify the specific structure of each partitioning mode.

[0045] Although more complex partitioning methods can significantly improve the coding efficiency, they also cause a significant increase in computational complexity, resulting in a very long time-consuming for the video coding block partitioning process. This is because the increase in partitioning methods leads to a significant increase in the number of times the CTU attempts block partitioning modes, and each attempt of a block partitioning mode requires an execution of an RDO (Rate-Distortion Optimization) decision. Statistics show that during the entire video coding process, the time ratio of CTU recursive block partitioning can exceed 90%.

[0046] In order to solve the problem that the video coding block division process takes a long time, an acceleration method for the video coding block division in the intra prediction mode and the inter prediction mode has been proposed. For example, in a latest acceleration scheme for the video coding block division in the intra prediction mode, a neural network is used to predict the probability of each side of each 4x4 sub-block in the 64x64 CTU being divided according to the original pixel value of the CTU of size 64x64, and the division probability of the current CU in each block division mode is calculated according to the probability value of each side being divided. If a certain division probability is lower than a preset value, the RDO decision process in the division mode corresponding to the division probability is skipped, thereby shortening the time of video coding block division. However, for video sequences with low precision, it is difficult to control the coding loss by using this video coding block division method. At present, the acceleration methods for block division in the inter prediction mode are mainly concentrated on the traditional acceleration scheme.

[0047] In order to shorten the block division time in the video encoding process and reduce the coding loss, the present disclosure proposes a new video coding block division method and device in the inter-frame prediction mode and a training method and device for a video coding block division prediction model. Specifically, based on the coding information of the current CTU and the coding information of its co-located CTU, the video coding block division prediction model is used to predict the sub-block boundary division probability in the current CTU at the CTU level, and the corresponding CU division mode probability value is calculated using these division probabilities at the CU level. According to the CU division mode probability value, a part of the possible block division modes is removed, and only the remaining possible block division modes need to be traversed to perform rate-distortion optimization decisions, thereby avoiding the huge amount of calculation caused by traversing all possible block division modes to perform rate-distortion optimization decisions, thereby greatly shortening the video coding block division time. In addition, according to the training method and device of the video coding block partition prediction model disclosed in the present invention, when training the video coding block partition prediction model, the boundary probability of each predicted partition sub-block can be superimposed with a corresponding weight according to the coding loss sensitivity, so that when the video coding block partition prediction model performs video coding block partition prediction, the block partition mode with higher sensitivity to coding loss can be retained to a greater extent, thereby reducing coding loss. Figures 2 to 11 A method and apparatus for dividing a video coding block in an inter-frame prediction mode and a method and apparatus for training a video coding block division prediction model according to exemplary embodiments of the present disclosure are described in detail.

[0048] Figure 2 is a flowchart illustrating a video encoding block partitioning method in an inter-prediction mode according to an exemplary embodiment of the present disclosure.

[0049] Reference Figure 2, at step 201, a first prediction residual and a second prediction residual can be obtained. Here, the first prediction residual is the prediction residual of the CTU to be encoded in the current video frame, and the second prediction residual is obtained by performing a preset process on the first prediction residual according to the block partition mode of the co-located CTU of the CTU to be encoded. The co-located CTU is the corresponding part of the CTU to be encoded in the reference frame of the current video frame. Here, the size of the CTU to be encoded can be 64×64, 32×32, etc., and there is no limitation on this. In the following description, a CTU to be encoded with a size of 64×64 is used for illustration.

[0050] The first prediction residual is the prediction residual of the CTU to be encoded in the current video frame, and can be obtained through two steps: motion estimation and motion compensation. Specifically, in the motion estimation step, a suitable matching region is found for the current CTU to be encoded in the reference frame of the current video frame (for example, it can be a frame before or after the current video frame), and in the motion compensation step, the difference between the current CTU to be encoded and its matching region is found (that is, the prediction residual. For example, it can be obtained by subtracting the pixel values of the current CTU from the pixel values of its matching region). In order to make full use of the information of the encoded CTU, the second prediction residual can be obtained by performing a preset process on the prediction residual of the CTU to be encoded according to the block partition mode of the co-located CTU of the CTU to be encoded. According to an exemplary embodiment of the present disclosure, the average value of the first prediction residual can be taken according to the block partition mode of the co-located CTU, and the first prediction residual after taking the average value is used as the second prediction residual. Here, the specific process of obtaining the second prediction residual can be understood with reference to Figure 3 Figure Figure 3 is a schematic diagram showing the process of obtaining the second prediction residual according to an exemplary embodiment of the present disclosure. Figure 3 (a) shows the prediction residual of a 64×64 CTU. Here, each small square represents an 8×8 sub-block, and the number in each small square represents the difference between the pixel value of this region and the pixel value of the corresponding region in its co-located CTU (that is, the prediction residual). Figure 3 (b) shows the block partition mode of the co-located CTU of this 64×64 CTU (shown by color blocks of different grayscales). Referring to Figure 3 (b), the average value of the prediction residual in Figure 3 (a) can be taken in each shown block. For example, according to the 16×16 block in the upper left corner of Figure 3 (b), after taking the average value of the prediction residual at the corresponding position in Figure 3 (a), the values change from 3, 4, 3, 3 to 3, 3, 3, 3.

[0051] In step 202, the first prediction residual and the second prediction residual may be input into a video coding block partition prediction model to obtain a first probability, which is the probability that each edge of each minimum sub-block of the CTU to be coded serves as a partition boundary of all possible block partition patterns.

[0052] Here, the video coding block partition prediction model may be implemented by an artificial neural network (e.g., a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), etc.). In a specific implementation process, the video coding block partition prediction model may be pre-trained. The specific structure and training process of the model will be described in detail later with reference to Figures 6 to 8 and will not be elaborated here for the time being. The minimum sub-blocks of the CTU to be coded may be 8×8 sub-blocks, 4×4 sub-blocks, etc., and there is no limitation thereto. In the following description, 8×8 sub-blocks are used as the minimum sub-blocks for illustration. The first probability is a probability vector composed of the probabilities that each edge of each minimum sub-block of the CTU to be coded serves as a partition boundary of all possible block partition patterns. The correspondence between each value in the probability vector and the edge of each minimum sub-block may be referred to Figure 4 (a) for description. Figure 4 (a) is a schematic diagram showing the probability vector according to an exemplary embodiment of the present disclosure. In Figure 4 (a), each edge of each 8×8 minimum sub-block corresponds to a probability value, representing the probability that the edge serves as a partition boundary of video coding block partitioning. For example, in Figure 4 (a), the probability corresponding to the bottom edge of the top-left minimum sub-block is 0.6.

[0053] In step 203, for each CU of the CTU to be coded, the following operations may be performed: Based on the first probability, determine the second probability corresponding to when the current CU is partitioned according to each possible block partition pattern, remove the block partition patterns whose second probabilities do not meet the preset threshold from all possible block partition patterns, and perform rate-distortion optimization decision on the remaining block partition patterns to obtain the block partition pattern of the current CU.

[0054] Here, the block partitioning modes include a quadtree partitioning mode (QT), a ternary tree partitioning mode (TT), and a binary tree partitioning mode (BT). The binary tree partitioning mode (BT) and the ternary tree partitioning mode (TT) each include a horizontal partitioning mode (BTH (horizontal binary tree partitioning mode) and TTH (horizontal ternary tree partitioning mode)) and a vertical partitioning mode (BTV (vertical binary tree partitioning mode) and TTV (vertical ternary tree partitioning mode)). That is, for each CU, all possible block partitioning modes are QT, BTH, BTV, TTH, and TTV. According to an exemplary embodiment of the present disclosure, since QT is more sensitive to coding loss, in order to better control coding loss, the present disclosure is configured to always retain the RDO decision for QT (i.e., not remove the QT partitioning mode) during the iterative process of video coding block partitioning, and remove the block partitioning modes in BTH, BTV, TTH, and TTV whose second probability does not meet a preset threshold, that is, remove the horizontal partitioning mode and / or vertical partitioning mode whose second probability does not meet the preset threshold from the binary tree partitioning mode and the ternary tree partitioning mode respectively. For each block partitioning mode, the method for determining its corresponding second probability is also different, and it can be combined with Figure 5 to describe. Figure 5 is a schematic diagram showing the determination of the second probability corresponding to the horizontal block partitioning mode based on the first probability according to an exemplary embodiment of the present disclosure. Refer to Figure 5 , S1, S2, S3, S4 represent the average probability of the partitioning boundaries when partitioning according to BTH and TTH (which can be obtained by adding the first probabilities of the smallest sub-blocks included in the partitioning boundary and then dividing by the number of the smallest sub-blocks). The second probability corresponding to BTH is the smaller probability of S2 and S3, and the second probability corresponding to TTH is the smaller probability of S1 and S4. The method for determining the second probability corresponding to the vertical block partitioning mode is similar thereto, and the second probability corresponding to QT is the smaller probability of BTH and BTV. In some embodiments, after obtaining the first probability, traditional machine learning methods such as decision trees can also be used to obtain the second probability corresponding to each block partitioning mode. Specifically, the second probability corresponding to each block partitioning mode can be obtained by inputting the first probability into a neural network including a convolutional layer (for example, a ResNet structure, etc.).

[0055] According to an exemplary embodiment of the present disclosure, the preset threshold may include a first preset threshold and a second preset threshold. Here, for the convenience of description, the second probability corresponding to the horizontal division mode may be used as the horizontal probability, and the second probability corresponding to the vertical division mode may be used as the vertical probability. When determining whether to remove a certain block division mode, the following operations may be performed for the binary tree division mode or the ternary tree division mode: compare the horizontal probability and the vertical probability with the first preset threshold respectively; when both the horizontal probability and the vertical probability are less than the first preset threshold, remove the horizontal division mode and the vertical division mode; when one of the horizontal probability and the vertical probability is less than the first preset threshold and the other is greater than or equal to the first preset threshold, remove the block division mode with the probability value less than the first preset threshold among the horizontal probability and the vertical probability; when both the horizontal probability and the vertical probability are greater than or equal to the second preset threshold, subtract the horizontal probability from the vertical probability, and when the absolute value of the difference is greater than or equal to the second preset threshold, remove the block division mode with the smaller probability value among the horizontal probability and the vertical probability. In some embodiments, since the sensitivities of BT and TT to video coding loss are different (BT is more sensitive to video coding loss), therefore, to better control the coding loss of video coding, the first preset threshold may be further divided into a first preset threshold for BT (for example, with a value of 0.1 to 0.3) and a first preset threshold for TT (for example, with a value less than 0.1), and the second preset threshold may be further divided into a second preset threshold for BT (for example, with a value of 0.1 to 0.15) and a second preset threshold for TT (for example, with a value less than 0.1). When performing the judgment operation of whether to remove a certain block division mode, for BT, it is judged by comparing the horizontal probability corresponding to HBT and the vertical probability corresponding to VBT with the first preset threshold for BT and the second preset threshold for BT; while for TT, it is judged by comparing the horizontal probability corresponding to HTT and the vertical probability corresponding to VTT with the first preset threshold for TT and the second preset threshold for TT, thereby avoiding the problem of excessive coding loss that may be caused by using a unified preset threshold.

[0056] According to an exemplary embodiment of the present disclosure, the video coding block partition prediction model provided by the present disclosure can be further improved. The improved model includes multiple parallel outputs, for example, each side of the 16x16 sub-block and the 8x8 sub-block is output as the first probability of the partition boundary of all possible block partition modes. When determining the block partition mode of each CU of the CTU to be encoded, the second probability corresponding to each possible block partition mode can be determined based on the first probability corresponding to each side of the 16x16 sub-block, and some block partition modes can be removed according to the aforementioned method, and the rate-distortion optimization decision is performed on the remaining block partition modes to obtain the block partition mode of the current CU; then, for the CU with a scale of 16x16, the second probability corresponding to each possible block partition mode can be determined based on the first probability corresponding to each side of the 8x8 sub-block, and some block partition modes can be removed according to the aforementioned method. Through this model, the time for block partitioning can be shortened to a certain extent, and the coding loss can be better controlled.

[0057] Figure 6 is a flowchart illustrating a training method of a video coding block partition prediction model according to an exemplary embodiment of the present disclosure.

[0058] Reference Figure 6 In step 601, a video training sample may be obtained, wherein the video training sample includes a first prediction residual, a second prediction residual, and a true value vector. The first prediction residual is a prediction residual of a first CTU to be encoded in a current video frame; the second prediction residual is obtained by presetting the first prediction residual according to a block partitioning mode of a second CTU in a reference frame of the current video frame, and the second CTU is a co-located CTU of the first CTU; the true value vector represents a situation where each edge of each minimum sub-block in the first CTU is a partitioning boundary of a true block partitioning mode. Here, the size of the first CTU and the second CTU may be 64×64, 32×32, etc., without limitation. In the following description, a CTU of size 64×64 is used for explanation.

[0059] According to an exemplary embodiment of the present disclosure, an expected number of video sequences (for example, 100 video sequences) can be provided. For each video sequence, a frame is extracted every 4 to 5 frames to calculate the prediction residual of the CTU to be encoded in the frame (i.e., the first CTU), and the prediction residual is used as the first prediction residual. In order to make full use of the information of the encoded CTU, the second prediction residual can be obtained by pre-processing the first prediction residual according to the block division mode of the second CTU (i.e., the co-located CTU of the first CTU) in the reference frame of the extracted video frame. According to an exemplary embodiment of the present disclosure, the first prediction residual can be averaged according to the block division mode of the second CTU, and the first prediction residual after averaging is used as the second prediction residual. The specific process of obtaining the second prediction residual can refer to the aforementioned description aboutFigure 3 Descriptions of (a) and Figure 3 (b) will not be elaborated here. The truth vector is the training label of the video coding block partition prediction model. The truth vector can be obtained from the true block partition pattern of the first CTU in the extracted video frame. For example, reference can be made to Figure 4 (b) to understand this truth vector. Figure 4 (b) is a schematic diagram showing the process of obtaining the truth vector according to an exemplary embodiment of the present disclosure. Figure 4 (b) shows a 64×64 CTU, the size of its smallest sub-block is 8×8, and the thick black solid line represents the partition boundary after block partitioning of the CTU. For the side of the smallest sub-block that serves as the partition boundary in the smallest sub-blocks, it can be denoted as 1, otherwise as 0. Thus, a truth vector can be obtained that reflects the situation of each side of each smallest sub-block in the CTU as the partition boundary of the true block partition pattern.

[0060] In step 602, the first prediction residual and the second prediction residual can be input into the video coding block partition prediction model to obtain an estimated vector composed of the probabilities that each side of each smallest sub-block in the first CTU serves as the partition boundary of all possible block partition patterns.

[0061] According to an exemplary embodiment of the present disclosure, the video coding block partition prediction model can be obtained by performing a preset process on the intra-prediction mode block partition prediction network structure, and the intra-prediction mode block partition prediction network structure can be implemented by a ResNet structure. Specifically, the video coding block partition prediction model shown in the present disclosure can be obtained by reducing the number of convolutional layers of the intra-prediction mode block partition prediction network structure to half of the original, so as to reduce the computational amount during the operation of the video coding block partition prediction model while taking into account the video coding block partition prediction effect. Figure 7 is a schematic diagram showing the structure of the video coding block partition prediction model according to an exemplary embodiment of the present disclosure. Referring to Figure 7 , the video coding block partition prediction model uses 5 convolutional layers to extract features and make predictions for a 64×64 CTU, and finally outputs an estimated vector composed of 112 probability values, corresponding to the probabilities that each side of the 8×8 smallest sub-blocks inside the 64×64 CTU serves as the partition boundary of all possible block partition patterns.

[0062] In step 603, weights can be superimposed on the probabilities of the preset sides of multiple smallest sub-blocks in the first CTU.

[0063] According to an exemplary embodiment of the present disclosure, probabilities of edges of multiple minimum sub-blocks that are preset division boundaries may be multiplied by corresponding weights to enhance the sensitivity of the video coding block division prediction model to video coding loss, where the value of the weight is greater than 1, and the preset division boundary may be obtained by performing quadtree division on a first CTU, and the division boundary does not include the edges of the first CTU. Here, the division boundary obtained after performing quadtree division may refer to Figure 8 the thick black solid line in Figure 8 which is a schematic diagram showing the division boundary obtained by performing quadtree division on a first CTU according to an exemplary embodiment of the present disclosure. Specifically, during the video coding process, the larger the divided block is, the greater the coding loss that may be caused after its removal. Therefore, to enhance the sensitivity of the video coding block division prediction model to video coding loss and thus better control the coding loss during the video coding process, after obtaining the estimated vector, probabilities of edges of the minimum sub-blocks that are the division boundaries obtained by performing quadtree division may be multiplied by corresponding weights, and the value of the weight may be, for example, but not limited to, 1.05.

[0064] In step 604, the value of the loss function may be calculated based on the estimated vector and the ground-truth vector superimposed with the corresponding weights. Here, the loss function may adopt a BCE (Binary Cross Entropy Error Function, cross-entropy loss function for binary classification) loss function, an MSE (Mean-Square Error Function) loss function, etc., and there is no limitation thereto.

[0065] In step 605, the video coding block division prediction model may be trained by adjusting the model parameters of the video coding block division prediction model according to the value of the loss function. That is, the parameters of the video coding block division prediction model may be adjusted by backpropagation of the loss calculated by the loss function. In addition, during the model training process, batch video training samples (for example, 100 video sequences) may be used to adjust (or update) the parameters of the video coding block division prediction model, and the parameters of the video coding block division prediction model may be iteratively adjusted (or updated) with the goal of minimizing the value of the loss function until the video coding block division prediction model converges.

[0066] According to the video coding block partition prediction model trained by the above training method, the probability that the edge of the smallest sub-block in the CTU to be encoded is used as the partition boundary of all possible block partition patterns can be output in the inter-frame prediction mode. Thus, the probabilities of all possible block partition patterns can be obtained based on the obtained probability, avoiding RDO decisions for some block partition patterns with relatively low possibilities, and thereby greatly shortening the block partition time of video coding. Additionally, since the convolutional layer of the video coding block partition prediction model is pruned, the computational load can be reduced during model operation, further shortening the block partition time of video coding. Moreover, since during the training process of the model, corresponding weights are superimposed on the probabilities of the edges of the smallest sub-blocks that are more sensitive to the coding loss, the output of the model can better reflect the sensitivity of the edges of each smallest sub-block to the coding loss (i.e., the probability of retaining the block coding pattern that is more sensitive to the coding loss is greater), thus better controlling the coding loss during the video coding process.

[0067] Figure 9 FIG. is a block diagram showing a video coding block partition device in an inter-frame prediction mode according to an exemplary embodiment of the present disclosure.

[0068] Referring to Figure 9 , the video coding block partition device 900 in an inter-frame prediction mode according to an exemplary embodiment of the present disclosure may include a residual obtaining unit 901, a probability obtaining unit 902, and a processing unit 903.

[0069] The residual obtaining unit 901 may obtain a first prediction residual and a second prediction residual. Among them, the first prediction residual is the prediction residual of the CTU to be encoded in the current video frame, and the second prediction residual is obtained by performing a preset process on the first prediction residual according to the block partition pattern of the co-located CTU of the CTU to be encoded. The co-located CTU is the corresponding part in the reference frame of the current video frame to the CTU to be encoded; the probability obtaining unit 902 may input the first prediction residual and the second prediction residual into the video coding block partition prediction model and obtain a first probability, which is the probability that each edge of each smallest sub-block of the CTU to be encoded is used as the partition boundary of all possible block partition patterns; the processing unit 903 may perform the following operations for each CU of the CTU to be encoded: based on the first probability, determine the second probability corresponding to when the current CU is partitioned according to each possible block partition pattern, remove the block partition patterns whose second probabilities do not meet the preset threshold from all possible block partition patterns, and perform rate-distortion optimization decision on the remaining block partition patterns to obtain the block partition pattern of the current CU.

[0070] Since Figure 2 The video coding block partition method in the inter-frame prediction mode shown can be performed by Figure 9The video coding block partitioning device 900 in the inter-frame prediction mode as shown is executed, and the residual obtaining unit 901, the probability obtaining unit 902, and the processing unit 903 can respectively execute operations corresponding to Figure 2 Steps 201, 202, and 203 in Figure 9 Any relevant details involved in the operations performed by each unit in Figure 2 can be referred to the corresponding description in

[0071] Figure 10 is a block diagram showing a training device for a video coding block partitioning prediction model according to an exemplary embodiment of the present disclosure.

[0072] Referring to Figure 10 , the training device 1000 for a video coding block partitioning prediction model according to an exemplary embodiment of the present disclosure may include a sample obtaining unit 1001, an estimated vector obtaining unit 1002, a weight superposition unit 1003, a loss function calculating unit 1004, and a model parameter adjusting unit 1005.

[0073] The sample obtaining unit 1001 can obtain video training samples, where the video training samples include a first prediction residual, a second prediction residual, and a ground truth vector. The first prediction residual is the prediction residual of the first CTU to be encoded in the current video frame; the second prediction residual is obtained by performing a preset process on the first prediction residual according to the block partitioning mode of the second CTU in the reference frame of the current video frame, and the second CTU is the co-located CTU of the first CTU; the ground truth vector represents the situation where each edge of each smallest sub-block in the first CTU is a partitioning boundary of the true block partitioning mode; the estimated vector obtaining unit 1002 can input the first prediction residual and the second prediction residual into the video coding block partitioning prediction model to obtain an estimated vector composed of the probabilities that each edge of each smallest sub-block in the first CTU is a partitioning boundary of all possible block partitioning modes; the weight superposition unit 1003 superposes corresponding weights on the probabilities of preset edges of multiple smallest sub-blocks in the first CTU; the loss function calculating unit 1004 can calculate the value of the loss function based on the estimated vector and the ground truth vector on which the corresponding weights are superposed; the model parameter adjusting unit 1005 can train the video coding block partitioning prediction model by adjusting the model parameters of the video coding block partitioning prediction model according to the value of the loss function.

[0074] Since Figure 6 the training method of the video coding block partitioning prediction model as shown can be executed by Figure 10 the training device 1000 for the video coding block partitioning prediction model as shown, and the sample obtaining unit 1001, the estimated vector obtaining unit 1002, the weight superposition unit 1003, the loss function calculating unit 1004, and the model parameter adjusting unit 1005 can respectively execute operations corresponding toFigure 6 The operations corresponding to steps 601, 602, 603, 604, and 605 in Figure 10 For any relevant details involved in the operations performed by each unit in Figure 6 Please refer to the corresponding description in

[0075] Figure 11 FIG. 9 is a block diagram of an electronic device 1100 according to an exemplary embodiment of the present disclosure.

[0076] Referring to Figure 11 , the electronic device 1100 includes at least one memory 1101 and at least one processor 1102. A set of computer-executable instructions is stored in the at least one memory 1101. When the set of computer-executable instructions is executed by the at least one processor 1102, a video coding block partitioning method or a training method for a video coding block partitioning prediction model in an inter-frame prediction mode according to an exemplary embodiment of the present disclosure is executed.

[0077] As an example, the electronic device 1100 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 1100 does not have to be a single electronic device, but may also be any assembly of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device 1100 may also be a part of an integrated control system or a system manager, or may be configured as a portable electronic device that can be interconnected locally or remotely (e.g., via wireless transmission).

[0078] In the electronic device 1100, the processor 1102 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0079] The processor 1102 may run instructions or code stored in the memory 1101. Wherein, the memory 1101 may also store data. The instructions and data may also be sent and received via a network interface device over a network, where the network interface device may employ any known transmission protocol.

[0080] The memory 1101 may be integrated with the processor 1102. For example, RAM or flash memory may be disposed within an integrated circuit microprocessor or the like. In addition, the memory 1101 may include separate devices such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory 1101 and the processor 1102 may be operatively coupled or may communicate with each other, for example, via I / O ports, network connections, etc., such that the processor 1102 can read files stored in the memory.

[0081] In addition, the electronic device 1100 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 1100 may be connected to each other via a bus and / or a network.

[0082] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are run by at least one processor, the at least one processor is caused to execute a method for video coding block partitioning or a method for training a video coding block partitioning prediction model in an inter-frame prediction mode according to the present disclosure. Examples of such computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memories, hard disk drives (HDDs), solid state drives (SSDs), cartridge memories (such as multimedia cards, secure digital (SD) cards, or extreme digital (XD) cards), magnetic tapes, floppy disks, magneto-optical data storage devices, optical data storage devices, hard disks, solid state disks, and any other devices that are configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or a computer such that the processor or the computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0083] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, the instructions in which may be executed by a processor of a computer device to complete a video coding block partitioning method in an inter-frame prediction mode or a training method for a video coding block partitioning prediction model according to an exemplary embodiment of the present disclosure.

[0084] According to the video coding block partitioning method and device in the inter-frame prediction mode disclosed in the present invention, based on the coding information of the current CTU and the coding information of its co-located CTU, the video coding block partitioning prediction model is used to predict the sub-block boundary partitioning probability in the current CTU at the CTU level, and the corresponding CU partitioning mode probability value is calculated using these partitioning probabilities at the CU level, and a part of the possible block partitioning modes is removed according to the CU partitioning mode probability value, and only the remaining possible block partitioning modes need to be traversed to perform rate-distortion optimization decisions, thereby avoiding the huge amount of calculation caused by traversing all possible block partitioning modes to perform rate-distortion optimization decisions, thereby greatly shortening the time for video coding block partitioning.

[0085] In addition, according to the training method and device of the video coding block division prediction model disclosed in the present invention, when training the video coding block division prediction model, the corresponding weights can be superimposed on the boundary probabilities of each predicted division sub-block based on the coding loss sensitivity, so that when the video coding block division prediction model performs video coding block division prediction, the block division patterns that are more sensitive to coding loss can be retained to a greater extent, thereby reducing coding loss.

[0086] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0087] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video coding block division method in inter-frame prediction mode, It is characterized in that include: Obtaining a first prediction residual and a second prediction residual, wherein the first prediction residual is a prediction residual of a CTU to be encoded in a current video frame, and the second prediction residual is obtained by performing a preset process on the first prediction residual according to a block partitioning mode of a co-located CTU of the CTU to be encoded, and the co-located CTU is a portion of a reference frame of the current video frame corresponding to the CTU to be encoded; Inputting the first prediction residual and the second prediction residual into a video coding block partition prediction model to obtain a first probability, where the first probability is a probability that each edge of each minimum sub-block of the CTU to be encoded is a partition boundary of all possible block partition modes; For each CU of the CTU to be encoded, perform the following operations: Based on the first probability, determining a second probability corresponding to when the current CU is divided according to each possible block partitioning mode, The block partition mode whose second probability does not meet the preset threshold is removed from all possible block partition modes, and a rate-distortion optimization decision is performed on the remaining block partition modes to obtain the block partition mode of the current CU.

2. The video coding block division method according to claim 1, It is characterized in that The performing preset processing on the first prediction residual according to the block partitioning mode of the co-located CTU of the CTU to be encoded includes: Taking an average value of the first prediction residual according to a block partitioning mode of a co-located CTU of the CTU to be encoded; The first prediction residual after averaging is used as the second prediction residual.

3. The video coding block division method according to claim 1, It is characterized in that The block partitioning mode includes a quadtree partitioning mode, a ternary tree partitioning mode and a binary tree partitioning mode, and the binary tree partitioning mode and the ternary tree partitioning mode each include a horizontal partitioning mode and a vertical partitioning mode; The removing the block division mode whose second probability does not meet the preset threshold from all possible block division modes comprises: The horizontal division pattern and / or the vertical division pattern whose second probability does not meet a preset threshold are removed from the binary tree division pattern and the ternary tree division pattern respectively.

4. The video coding block division method according to claim 3, It is characterized in that The preset threshold includes a first preset threshold and a second preset threshold, the second probability corresponding to the horizontal division mode is used as the horizontal probability, and the second probability corresponding to the vertical division mode is used as the vertical probability; Removing the horizontal division pattern and / or the vertical division pattern whose second probability does not meet the preset threshold from the binary tree division pattern and the ternary tree division pattern respectively includes: For the binary tree partition mode or the ternary tree partition mode, perform the following operations: Comparing the horizontal probability and the vertical probability with the first preset threshold respectively; When both the horizontal probability and the vertical probability are less than the first preset threshold, removing the horizontal division mode and the vertical division mode; When one of the horizontal probability and the vertical probability is less than the first preset threshold and the other is greater than or equal to the first preset threshold, remove the block division mode whose probability value is less than the first preset threshold among the horizontal probability and the vertical probability; When the horizontal probability and the vertical probability are both greater than or equal to the second preset threshold, the horizontal probability and the vertical probability are subtracted, and when the absolute value of the difference is greater than or equal to the second preset threshold, the block division pattern with the smaller probability value among the horizontal probability and the vertical probability is removed.

5. A training method for a video coding block partition prediction model, It is characterized in that include: Acquire a video training sample, wherein the video training sample includes a first prediction residual, a second prediction residual, and a true value vector, the first prediction residual is a prediction residual of a first CTU to be encoded in a current video frame, the second prediction residual is obtained by performing preset processing on the first prediction residual according to a block partitioning mode of a second CTU in a reference frame of the current video frame, the second CTU is a co-located CTU of the first CTU, and the true value vector represents a situation in which each edge of each minimum sub-block in the first CTU serves as a partitioning boundary of a true block partitioning mode; Inputting the first prediction residual and the second prediction residual into the video coding block partition prediction model to obtain an estimated vector composed of the probability that each edge of each minimum sub-block in the first CTU is used as a partition boundary of all possible block partition modes; superimposing corresponding weights on the probabilities of the preset edges of the plurality of minimum sub-blocks in the first CTU; Calculate the value of the loss function based on the estimated vector and the true value vector superimposed with corresponding weights; The video coding block partition prediction model is trained by adjusting the model parameters of the video coding block partition prediction model according to the value of the loss function.

6. The training method of the video coding block partition prediction model as claimed in claim 5, It is characterized in that The presetting the first prediction residual according to the block partitioning mode of the second CTU in the reference frame of the current video frame includes: averaging the first prediction residual according to a block partitioning mode of a second CTU in a reference frame of the current video frame; The first prediction residual after averaging is used as the second prediction residual.

7. The training method of the video coding block partition prediction model according to claim 5, It is characterized in that The superimposing corresponding weights on the probabilities of the preset edges of the plurality of minimum sub-blocks in the first CTU includes: The probability of the edges of multiple minimum sub-blocks as preset partition boundaries is multiplied by corresponding weights to enhance the sensitivity of the video coding block partition prediction model to video coding loss, wherein the value of the weight is greater than 1, the preset partition boundary is obtained by quadtree partitioning the first CTU, and the preset partition boundary does not include the edge of the first CTU.

8. The training method of the video coding block partition prediction model as claimed in claim 5, It is characterized in that The video coding block partition prediction model is obtained by presetting the intra-frame prediction mode block partition prediction network structure, and the intra-frame prediction mode block partition prediction network structure is implemented by a ResNet structure; The presetting process of dividing the intra-frame prediction mode block into prediction network structures includes: The number of convolutional layers of the intra-frame prediction mode block partition prediction network structure is reduced to half of the original number.

9. A video coding block division device in inter-frame prediction mode, It is characterized in that include: A residual acquisition unit is configured to: acquire a first prediction residual and a second prediction residual, wherein the first prediction residual is a prediction residual of a CTU to be encoded in a current video frame, and the second prediction residual is obtained by performing a preset process on the first prediction residual according to a block partitioning mode of a co-located CTU of the CTU to be encoded, and the co-located CTU is a portion of a reference frame of the current video frame corresponding to the CTU to be encoded; a probability acquisition unit, configured to: input the first prediction residual and the second prediction residual into a video coding block partition prediction model and obtain a first probability, where the first probability is a probability that each edge of each minimum sub-block of the CTU to be encoded is a partition boundary of all possible block partition modes; The processing unit is configured to: for each CU of the CTU to be encoded, perform the following operations: Based on the first probability, determining a second probability corresponding to when the current CU is divided according to each possible block partitioning mode, The block partition mode whose second probability does not meet the preset threshold is removed from all possible block partition modes, and a rate-distortion optimization decision is performed on the remaining block partition modes to obtain the block partition mode of the current CU.

10. The video coding block division device according to claim 9, It is characterized in that The residual acquisition unit is configured as follows: Taking an average value of the first prediction residual according to a block partitioning mode of a co-located CTU of the CTU to be encoded; The first prediction residual after averaging is used as the second prediction residual.

11. The video coding block division device according to claim 9, It is characterized in that The block partitioning mode includes a quadtree partitioning mode, a ternary tree partitioning mode and a binary tree partitioning mode, and the binary tree partitioning mode and the ternary tree partitioning mode each include a horizontal partitioning mode and a vertical partitioning mode; The processing unit is configured to remove the horizontal division pattern and / or the vertical division pattern whose second probability does not meet a preset threshold from the binary tree division pattern and the ternary tree division pattern respectively.

12. The video coding block division device according to claim 11, It is characterized in that The preset threshold includes a first preset threshold and a second preset threshold, the second probability corresponding to the horizontal division mode is used as the horizontal probability, and the second probability corresponding to the vertical division mode is used as the vertical probability; The processing unit is configured to: for the binary tree partition mode or the ternary tree partition mode, perform the following operations: Comparing the horizontal probability and the vertical probability with the first preset threshold respectively; When both the horizontal probability and the vertical probability are less than the first preset threshold, the horizontal partitioning mode and the vertical partitioning mode are removed; When one of the horizontal probability and the vertical probability is less than the first preset threshold and the other is greater than or equal to the first preset threshold, the block partitioning mode with a probability value less than the first preset threshold among the horizontal probability and the vertical probability is removed; When both the horizontal probability and the vertical probability are greater than or equal to the second preset threshold, the difference between the horizontal probability and the vertical probability is calculated, and when the absolute value of the difference is greater than or equal to the second preset threshold, the block partitioning mode with a smaller probability value among the horizontal probability and the vertical probability is removed.

13. A training device for a video coding block partitioning prediction model, Characterized in that, Comprising: A sample acquisition unit configured to: acquire video training samples, wherein the video training samples include a first prediction residual, a second prediction residual, and a true value vector, the first prediction residual is the prediction residual of a first CTU to be encoded in a current video frame, the second prediction residual is obtained by performing a preset process on the first prediction residual according to the block partitioning mode of a second CTU in a reference frame of the current video frame, the second CTU is a co-located CTU of the first CTU, and the true value vector represents the situation where each side of each minimum sub-block in the first CTU is a partitioning boundary of a true block partitioning mode; An estimated vector acquisition unit configured to: input the first prediction residual and the second prediction residual into the video coding block partitioning prediction model to obtain an estimated vector composed of the probabilities that each side of each minimum sub-block in the first CTU is a partitioning boundary of all possible block partitioning modes; A weight superposition unit configured to: superpose corresponding weights on the probabilities of preset sides of multiple minimum sub-blocks in the first CTU; A loss function calculation unit configured to: calculate the value of the loss function based on the estimated vector and the true value vector on which the corresponding weights are superposed; A model parameter adjustment unit configured to: train the video coding block partitioning prediction model by adjusting the model parameters of the video coding block partitioning prediction model according to the value of the loss function.

14. The training device for a video coding block partitioning prediction model according to claim 13, Characterized in that, The sample acquisition unit is configured to: Take the average value of the first prediction residual according to the block partitioning mode of the second CTU in the reference frame of the current video frame; Use the first prediction residual after taking the average value as the second prediction residual.

15. The training device for a video coding block partitioning prediction model according to claim 13, Characterized in that, The weight superposition unit is configured to: Multiply the probabilities of the edges of multiple minimum sub-blocks that are preset partitioning boundaries by corresponding weights to enhance the sensitivity of the video coding block partitioning prediction model to video coding loss, where the value of the weight is greater than 1, the preset partitioning boundary is obtained by performing a quadtree partitioning on the first CTU, and the preset partitioning boundary does not include the edges of the first CTU.

16. The training apparatus for a video coding block partitioning prediction model according to claim 13, wherein, the video coding block partitioning prediction model is obtained by performing preset processing on an intra prediction mode block partitioning prediction network structure, and the intra prediction mode block partitioning prediction network structure is implemented by a ResNet structure; the preset processing on the intra prediction mode block partitioning prediction network structure includes: reducing the number of convolutional layers of the intra prediction mode block partitioning prediction network structure to one-half of the original.

17. An electronic device, wherein, it includes: at least one processor; at least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the video coding block partitioning method in the inter prediction mode according to any one of claims 1 to 4 or the training method of the video coding block partitioning prediction model according to any one of claims 5 to 8.

18. A computer-readable storage medium storing instructions, wherein, when the instructions are run by at least one processor, the at least one processor is caused to execute the video coding block partitioning method in the inter prediction mode according to any one of claims 1 to 4 or the training method of the video coding block partitioning prediction model according to any one of claims 5 to 8.

19. A computer program product including computer instructions, wherein, when the computer instructions are executed by at least one processor, the video coding block partitioning method in the inter prediction mode according to any one of claims 1 to 4 or the training method of the video coding block partitioning prediction model according to any one of claims 5 to 8 is implemented.

Citation Information

Patent Citations

  • An encoder, a decoder and corresponding methods for sub-block partitioning mode

    CA3140818A1

  • 3D-HEVC inter-frame rapid method based on Bayesian decision division of a CU

    CN109756719A