Video coding method based on conditional probability and reconstruction distortion propagation optimization

By building a forward-looking image set and a block-level time domain reference relationship chain, the reference architecture between the coding units is optimized, and the problem that the existing time domain distortion propagation model is difficult to simulate the actual encoding process in the preprocessing stage, achieving more efficient rate distortion performance and space-time dependence information fusion.

CN120111227APending Publication Date: 2025-06-06ZHEJIANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311649122.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing time-domain distortion propagation model is difficult to fully simulate the time-domain distortion propagation process in the actual encoding stage in the preprocessing stage, and in the process of measuring time-dependence, the construction form is single and simple, and it is impossible to effectively mine the space-time-domain information between coding units.

Method used

By constructing a forward-looking image set and a block-level time-domain reference relationship chain, setting a reference frame list consistent with the actual encoding process, calculating cumulative distortion offset and time-domain quantization parameter offset, optimizing the reference architecture between coding units, and reconstructing the expression of distortion dependencies.

Benefits of technology

With the controllable code rate, the rate distortion performance of video encoding is improved, the space-time dependence information fusion capability between encoding units is enhanced, and the overall quality of video encoding is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111227A_ABST
    Figure CN120111227A_ABST
Patent Text Reader

Abstract

The invention discloses a video coding method based on conditional probability and reconstruction distortion propagation optimization, which comprises the following steps of: constructing a prospective image set, deciding a coding structure, constructing a block-level time domain reference relation chain, calculating accumulated distortion offset and determining time domain quantization parameter offset, calculating the prediction cost of each image block corresponding to each frame of image in the look-ahead image set based on the reference frame list, determining a reference block according to the prediction cost so as to construct a block-level time domain reference relation chain, and constructing a distortion dependency item based on the block-level time domain reference relation chain, and calculating an accumulated distortion offset based on the distortion dependency so as to calculate a time domain quantization parameter offset of each image block, and coding the to-be-coded reference frame based on the time domain quantization parameter offset. A reference frame list consistent with an actual coding process is set, so that a distortion propagation path of time domain distortion on a reference relation chain in the actual coding process is simulated, an expression of a distortion dependency item is reconstructed, and the rate distortion performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video encoding, and in particular to a video encoding method based on conditional probability and quantization distortion propagation optimization. Background Art

[0002] With the development of network technology and hardware equipment, people require display devices to better restore the fine granularity and true colors of the real world, and ultra-high definition (UHD), panoramic (Panoramic), virtual reality (VR) and three-dimensional (3D) videos have emerged. However, due to the limitations of channel transmission bandwidth and terminal processing equipment performance, we need to compress the video to a level that meets the transmission requirements to output a smooth bitstream. The video compression operation is completed using video encoding methods, which improve the compression rate of the video by effectively removing the temporal redundancy, spatial redundancy, perceptual redundancy, and knowledge redundancy of the video. How to balance the subjective quality consistency of the video and the smoothness of the output bitstream has always been a focus of academic research.

[0003] The rate control algorithm in the video coding method plays a leading role in solving the above problems. The rate control algorithm is responsible for controlling the bit rate so that the overall bit rate of the video is as close as possible to the target bit rate set by external parameters. The traditional rate control algorithm adopts a bit allocation method without time domain dependence, which suppresses the information interaction ability of each coding unit in the time domain chain. At the same time, it focuses on maintaining the stability of the bit rate. It cannot fully explore the content characteristics of different coding units, which limits the rate distortion performance of the codec. In recent years, the rate control algorithm based on the time domain dependence between inter-frame coding units has received widespread attention. The software encoder x264 of the H.264 / AVC[1] standard designed a macroblock tree (MBTree) technology with time domain dependence, and determined the reference value of the current block to the subsequent blocks through the lookahead module. As a new generation of HEVC standard software encoder, x265 inherits this technology and is named coding unit tree (CUTree). In 2017, some scholars constructed a time domain distortion propagation model based on the propagation of CU-level distortion in the time domain chain as a promotion of MBTree technology, and used a distortion function based on SSIM for distortion accumulation. The average BDBR-SSIM gain of this algorithm was 4.4%. In 2018, some scholars found that the above model did not take into account the skip mode in the propagation process, which caused a large target bitrate deviation (TBD). Therefore, the skip probability of CU was introduced into the time domain distortion propagation model, and a new analytical solution for the local quantizer was obtained. In 2019, based on the Source Distortion Temporal Propagation (SDTP) model, scholars proposed a general solution framework based on the weighted Lagrangian multiplier λ, in which the weighting coefficient is related to the strength of the temporal dependence between coding units, and is iteratively estimated at the CTU level and the frame level respectively. Since there is a stepped quantization step between coding units in the time domain chain, the relative size of the quantization step directly affects the degree of distortion accumulation. Some scholars have considered the impact of different quantization steps on distortion transmission during the modeling process, and given corresponding adaptive quantization formulas from the perspective of global rate-distortion optimization.

[0004] The existing time domain distortion propagation model requires that the decisions in the actual coding stage be fully estimated in the preprocessing stage. Existing technologies often enrich the expression of time dependence by estimating more coding configuration information, such as using the probability of the skip mode to measure the impact of the skip mode on the bit rate in the actual coding process, and using the quantization step size to study the impact of the quantization distortion of the reference block on the coding quality of the coded frame.

[0005] However, existing time domain distortion propagation models often ignore the prior conditions of time domain distortion propagation. For example, time domain distortion propagation only occurs through the motion compensation process of inter-frame prediction, and only by fully simulating the actual distortion propagation process in the preprocessing stage can the time domain distortion propagation process consistent with the actual encoding process be performed. There is a problem that the preprocessing stage cannot fully simulate the time domain distortion propagation process in the actual encoding stage. At the same time, in the process of measuring time dependency, a unified representation of multiple estimated coding configurations has not yet been formed, and there is a problem that the form of constructing time domain dependency relationships is single and simple. Summary of the invention

[0006] In order to overcome the shortcomings of the above technologies, the present invention provides a video coding method for optimizing the propagation of time domain distortion based on conditional probability and quantization distortion. In the pre-analysis process, a reference frame list consistent with the actual coding process is set to simulate the distortion propagation path of the time domain distortion in the reference relationship chain in the actual coding process, and reconstruct the expression of the distortion dependency, optimize the reference architecture between coding units during forward analysis, and further explore and integrate the time and space dependency information of different coding units, so as to further improve the rate-distortion performance under the condition of controllable bit rate.

[0007] The technical solution adopted by the present invention to overcome its technical problems is: a video coding method based on conditional probability and reconstruction distortion propagation optimization proposed by the present invention includes: constructing a forward-looking image set: caching a plurality of frame images of a video segment to be encoded as an original image set, and generating a forward-looking image set based on the original image set; deciding a coding structure: deciding a coding structure for the forward-looking image set, and the coding structure at least includes the setting of a frame type for each frame and the setting of a reference frame list; constructing a block-level temporal reference relationship chain: calculating the prediction cost of each image block corresponding to each frame image in the forward-looking image set based on the reference frame list, and determining the reference block based on the prediction cost to construct a block-level temporal reference relationship chain; calculating the cumulative distortion offset: constructing a distortion dependency based on the block-level temporal reference relationship chain, and calculating the cumulative distortion offset based on the distortion dependency; determining a temporal quantization parameter offset: calculating the temporal quantization parameter offset of each image block of a reference frame to be encoded in the forward-looking image set based on the cumulative distortion offset, and encoding the reference frame to be encoded based on the temporal quantization parameter offset.

[0008] Furthermore, the method caches several frames of the video segment to be encoded as an original image set, and generates a forward-looking image set based on the original image set, specifically including: generating four double down-sampled frames in different difference directions for each frame of the original image set, and performing a bracketing operation on the double down-sampled frame in each difference direction, thereby generating four expanded edge down-sampled frames in different interpolation directions; and generating a forward-looking image set based on the original image set for generating the expanded edge down-sampled frames.

[0009] A forward-looking image set is constructed by expanding and downsampling frames in four different interpolation directions to improve the interpolation accuracy and support the search of boundary blocks in the subsequent motion search process.

[0010] Furthermore, the setting of the reference frame list for the forward-looking image set decision specifically includes: based on the reference frame cache list, presetting at least one reference frame list for each image in the forward-looking image set, wherein the reference frame list is used to store at least the POC of each reference frame, the POC of the current image and the corresponding POC difference; based on the coding order, traversing all images in the forward-looking image set, filling the reference frame list of the current image with the reference frames in the reference frame cache list; if the current image frame is determined to be a reference frame based on the frame type, adding it to the reference frame cache list for setting the reference frame list of subsequent images.

[0011] Furthermore, after traversing all the images in the forward-looking image set, the reference frame cache list is rolled back to the state where the encoding of the first encoding structure in the forward-looking image set is completed.

[0012] Furthermore, the method of calculating the prediction cost of each image block corresponding to each frame image in the forward-looking image set based on the reference frame list specifically includes: dividing each frame image in the forward-looking image set into a number of image blocks of the same size; calculating the best inter-frame prediction cost and the best intra-frame prediction cost for each image block; and obtaining a reference with the minimum prediction cost based on the best intra-frame prediction cost and the best inter-frame prediction cost.

[0013] Furthermore, calculating the best inter-frame prediction cost for each image block specifically includes: performing motion search on each image block based on the reference frames in the reference frame list to obtain the best reference block on each reference frame, obtaining pixel residuals based on the pixel values ​​of the image block and the best reference block of each reference frame and performing Hadamard transform, calculating the sum of the absolute values ​​of all coefficients in the transform coefficient matrix, and using it as the inter-frame prediction cost; comparing the inter-frame prediction costs of the current image block for the best reference blocks on each reference frame to obtain the best inter-frame prediction cost, and saving the inter-frame prediction direction, reference frame number and motion vector.

[0014] Furthermore, calculating the optimal intra-frame prediction cost for each image block specifically includes: respectively calculating the Hadamard transform costs of the DC mode and the Planner mode of each image block in the forward-looking image set as the intra-frame prediction cost, and selecting the smaller value of the intra-frame prediction cost in the two modes as the intra-frame prediction cost of the optimal non-angle mode; searching for the angle mode based on a preset step size to obtain the optimal angle mode, and calculating the intra-frame prediction cost of the optimal angle mode; comparing the intra-frame prediction costs of the optimal non-angle mode and the optimal angle mode, and taking the smaller value of the two as the optimal intra-frame prediction cost of the image block.

[0015] The prediction cost is obtained by calculating the best intra-frame prediction cost and the best inter-frame prediction cost, thereby determining the reference block.

[0016] Furthermore, the construction of the distortion dependency based on the block-level temporal reference relationship chain specifically includes: rearranging the images in the forward-looking image set based on the coding order, and reversely traversing each image block in the rearranged forward-looking image set;

[0017] The cumulative distortion ratio is determined based on the direct distortion dependency through the rate-distortion cost on the block-level time domain propagation chain, wherein the cumulative distortion ratio on the block-level time domain reference relationship chain is β is the direct distortion dependency between two image blocks in the direct reference relationship of the block-level temporal reference relationship chain, and n is the number of image blocks in the block-level temporal reference relationship chain.

[0018] Furthermore, the direct distortion dependency term β=ρ*θ*α, wherein ρ is the inter-frame probability, θ is the distortion transfer ratio, and α is the quantization error ratio.

[0019] Further, the temporal quantization parameter offset of each image block of the reference frame to be encoded in the forward-looking image set is calculated based on the cumulative distortion offset, and the reference frame to be encoded is encoded based on the temporal quantization parameter offset, specifically including: converting the cumulative distortion ratio into a quantization parameter offset ΔQP of each image block in the reference frame to be encoded in the forward-looking image set, wherein the temporal quantization parameter offset ΔQP of each image block in the reference frame is expressed as ΔQP=strength*log 2 (1+σ), 1+σ is the cumulative distortion ratio; outputting the time domain quantization parameter offset ΔQP block by block, and encoding the reference frame to be encoded based on the time domain quantization parameter offset ΔQP.

[0020] The beneficial effects of the present invention are:

[0021] 1. Four edge-divided down-sampling frames of the original image with different interpolation directions are used to construct a forward-looking image set, which is used to improve the interpolation accuracy and support the search of boundary blocks in the subsequent motion search process.

[0022] 2. A reference frame list consistent with the actual encoding process is set to simulate the distortion propagation path of the time domain distortion in the reference relationship chain during the actual encoding process.

[0023] 3. Calculate the cumulative distortion ratio 1+σ on the time domain reference relationship chain to indicate the degree of time domain distortion transmission of the current image block in the entire time domain reference relationship chain. By assigning a smaller time domain adaptive quantization offset to an image block with a larger cumulative distortion transmission ratio, it is used to express its importance in the time domain propagation chain.

[0024] 4. The expression of distortion dependency is reconstructed, and a direct distortion dependency β is proposed to represent the distortion ratio between two image blocks with a reference relationship. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A schematic flow chart of a video coding method based on conditional probability and reconstruction distortion propagation optimization according to an embodiment of the present invention;

[0026] Figure 2 The figure is a schematic diagram of the calculation flow of time domain distortion propagation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] CUTree: Coding tree unit, which determines the temporal adaptive quantization offset of each image block on the reference frame by measuring the cumulative distortion ratio of the reference block on the temporal reference chain. The relevant algorithm is implemented on the open source commercial encoder x265.

[0028] Hadamard transform: Hadamard transform is a fast generalized Fourier transform method;

[0029] X265: an open source video encoder based on the H265 standard;

[0030] POC: The playback order index of the video. For example, the poc of the first frame in the video is 0, the poc of the second frame is 1, and so on.

[0031] In order to facilitate those skilled in the art to better understand the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The following is only exemplary and does not limit the protection scope of the present invention.

[0032] like Figure 1 The figure is a flow chart of a video coding method based on conditional probability and reconstruction distortion propagation optimization described in this embodiment, which is described below according to a specific embodiment.

[0033] S1, constructing a forward-looking image set, caching a number of frames of images of the video segment to be encoded as an original image set, and generating a forward-looking image set based on the original image set.

[0034] Several consecutive frame images cached in the display order in the video to be encoded are taken as the original image set, and the forward-looking image set is generated based on the original image set. In one embodiment of the present invention, N consecutive frame images cached in the display order are obtained, which are called the original image set.

[0035] There are many ways to generate the forward-looking image set. It can be completely consistent with the original image set, or it can be a multiple downsampled frame of the original image set, or it can be a reconstructed form of the original image set after one encoding, depending on the time and space complexity requirements of video encoding. The storage order of images in the forward-looking image set is consistent with the display order of video frames. It should be noted that the number of images N in the forward-looking image set is related to the cache capacity of the device and the pre-analysis parameters of the peripherals.

[0036] The forward-looking image set performs preliminary analysis on the video segment to be encoded, estimates the encoding information of the video segment to be encoded, and uses the encoding information to decide the encoding structure of the input video, the frame type of the input video frame, and the key encoding parameters of the input video frame.

[0037] In one embodiment of the present invention, each image in the original image set is subjected to a two-fold downsampling and edge expansion operation, which generates four two-fold downsampling frames in different interpolation directions, and each downsampling frame in the interpolation direction is edge expanded to obtain four edge expanded downsampling frames in different interpolation directions, which are used to improve the interpolation accuracy and support the search of boundary blocks in the subsequent motion search process. Each image in the forward-looking image set includes four edge expanded downsampling frames of the original image set.

[0038] S2, decision coding structure, is the forward-looking image set decision coding structure, and the coding structure at least includes the setting of each frame type in the forward-looking image set and the setting of the reference frame list.

[0039] The frame type can be set to I frame, P frame, GPB frame and B frame, among which the B frame can add a reference flag to indicate whether it can be used as a reference frame for other frames. The frame type decision method can use the fixed structure frame type setting provided by the x265 encoder, or the frame type decision method based on a one-time traversal, or the frame type decision method based on the Viterbi idea.

[0040] The reference frame list can be set in different ways. Taking the X265 nearest single reference frame setting method as an example, the reference frame of each frame only contains one or two frames that are closest in time distance. Taking P frames and B frames as examples, there are two reference situations for B frames. First, if the B frame is a B frame that can be referenced, then the B frame can refer to the two nearest non-B frames; second, if the B frame is a B frame that cannot be referenced, then the B frame bidirectionally references the nearest non-B frame and the nearest B frame that can be referenced. For P frames, the nearest other P frames are always referenced forward in the display order.

[0041] In one embodiment of the present invention, the setting of the reference frame list is described by taking the reference frame list adopting a multiple reference frame setting method as an example.

[0042] In the pre-analysis stage, a reference frame list consistent with the actual encoding process is set for each image in the forward-looking image set to simulate the distortion propagation path of the time domain distortion in the reference relationship chain during the actual encoding process. The specific process is as follows.

[0043] S21, based on a preset reference frame cache list, maintaining a forward and backward reference frame list for each image in the forward-looking image, and storing the difference between the POC of each image and the POC of the reference image.

[0044] A reference frame cache list is maintained in the pre-analysis stage, and a forward and backward reference frame list is maintained for each image in the forward-looking image set, storing the POC of each reference frame, the POC of each image, and the difference between the POC of the current image and the POC of the reference image.

[0045] It should be noted that for B frames, there are forward and backward reference frame lists, and for P frames, only one forward and backward reference frame list is retained, so they are collectively referred to as reference frame lists in the following description.

[0046] S22, traversing all images in the forward-looking image set based on the coding order, and filling the reference frames in the reference frame cache list into the reference frame list of the current image.

[0047] All images in the forward-looking image set are traversed in the coding order, and the reference frames in the reference frame cache list are filled into the forward and backward reference frame lists of the current image.

[0048] It should be noted that due to the existence of B frames, the display order and the encoding order are inconsistent, so the encoding order is based on the images in the forward-looking image set read sequentially from the mapping table.

[0049] S23: If the type of the current image can be referenced, add it to the reference frame cache table as a reference frame cache list for subsequent other images.

[0050] The reference frame cache list is maintained according to the first-in-first-out rule. If the current image can be referenced, it will be added to the reference frame cache list. When the reference frame list is larger than the set maximum storage capacity, the frame with the farthest time distance from the current image will be removed from the reference frame list.

[0051] S24, after the forward-looking image set is traversed, the reference frame cache list is rolled back to the state where the first coding structure in the forward-looking image set is traversed, and after the end of this pre-analysis phase, the reference frames in the reference frame cache list are not cleared to support access to reference frames outside the forward-looking image set for reference in the next pre-analysis process.

[0052] S3, constructing a block-level temporal reference relationship chain, calculating the prediction cost of each image block corresponding to each frame in the forward-looking image set based on the reference frame list, and determining the reference block based on the prediction cost to construct a block-level temporal reference relationship chain.

[0053] Each frame in the forward-looking image set is divided into several image blocks, and the intra-frame prediction cost is calculated based on the image block as the basic unit, and inter-frame prediction is performed based on the reference frame list to obtain the reference block with the minimum prediction cost, which is used to complete the construction of the block-level temporal reference relationship chain. The specific steps include:

[0054] S31, dividing each frame of the forward-looking image set into a number of image blocks of the same size.

[0055] Each image in the look-ahead image set is divided into a number of image blocks of fixed size.

[0056] S32, calculating the best inter-frame prediction cost and the best intra-frame prediction cost for each image block.

[0057] The image blocks of each image in the forward-looking image set P are traversed to calculate the best intra-frame prediction cost and the best inter-frame prediction cost of each image block. First, the best inter-frame prediction cost of each image block is calculated, which includes the following steps.

[0058] S3211, for each image block, search for the reference frame motion in the reference frame list to obtain the best reference block on each reference frame.

[0059] In one embodiment of the present invention, each image block uses a hexagonal search template to perform an integer pixel search on the downsampled frame in the first interpolation direction of the reference frame, and on this basis, performs a weighted 1 / 2 precision pixel search with the downsampled frames in other interpolation directions, wherein the basic unit of motion search is a pixel block of size 8x8.

[0060] S3212, obtaining pixel residuals based on the pixel values ​​of the image block and the best reference block of each reference frame and performing Hadamard transform, calculating the sum of the absolute values ​​of all coefficients in the transform coefficient matrix, and using it as the inter-frame prediction cost.

[0061] The pixel residual is the pixel difference between the image block and the reference block, which is obtained by image block motion search.

[0062] S3213, compare the inter-frame prediction costs of the current image block on each reference frame to obtain the best inter-frame prediction cost, and save the inter-frame prediction direction, reference frame sequence number and motion vector.

[0063] Compare the inter-frame prediction costs of the current image block on all reference frames, and take the minimum inter-frame prediction cost as the best inter-frame prediction cost. And retain the inter-frame prediction direction, reference frame number and motion vector to index the best reference block of the current image block when performing the temporal distortion propagation process later.

[0064] In one embodiment of the present invention, the calculation of the intra prediction cost is as follows.

[0065] S3221, respectively calculate the Hadamard transform cost of the DC mode and the Planner mode of each image block in the forward-looking image set, the Hadamard transform cost is called the intra-frame prediction cost, and the smaller value of the intra-frame prediction cost in the two modes is used as the intra-frame prediction cost of the best non-angle mode.

[0066] S3222, searching for an angle mode based on a preset step size to obtain an optimal angle mode, and calculating an intra-frame prediction cost of the optimal angle mode;

[0067] A first search for the best angle mode is performed based on a first preset step length, and a second search for the best angle mode is performed based on a second preset step length.

[0068] In one embodiment of the present invention, a rough search of the angle pattern is first performed with a step size of 5, and then a fine search of the angle pattern is performed with a step size of 1.

[0069] S3223, comparing the intra prediction costs of the best non-angle mode and the best angle mode, and taking the smaller value of the intra prediction costs as the best intra prediction cost of the image block

[0070] Based on the best intra-frame prediction cost and the best inter-frame prediction cost, a reference block with the minimum prediction cost is obtained, thereby constructing a block-level temporal domain reference relationship chain between the reference block and the current image block.

[0071] It should be noted that, since the reference block can also be referenced by other image blocks, the current image block can also be used as a reference block for other image blocks, so the block-level temporal reference relationship chain includes several image blocks.

[0072] S4, calculate the cumulative distortion offset: traverse the forward-looking image set in reverse, build distortion dependencies based on the block-level temporal reference relationship chain, and calculate the cumulative distortion offset based on the distortion dependencies.

[0073] The images in the forward-looking image set are rearranged in the coding order, and the tail frame p of the rearranged forward-looking image set is NAs the starting point, reversely traverse each image block of each image in the forward-looking image set P, and calculate the cumulative distortion ratio 1+σ on the time domain reference relationship chain with the current image block as the reference starting point. The cumulative distortion ratio indicates the degree of time domain distortion transfer of the current image block in the entire time domain reference relationship chain. By assigning a smaller time domain adaptive quantization offset to an image block with a larger cumulative distortion transfer ratio, it is used to express its importance in the block-level time domain reference relationship chain.

[0074] Since the coding block will produce distortion during the quantization process, in the block-level propagation chain with a reference relationship, through the inter-frame prediction between blocks, part of the distortion generated on the coding block is propagated to the next block through motion compensation. Therefore, the two image blocks with a reference relationship constitute a temporal dependency relationship of coding distortion. Therefore, a direct distortion dependency term β is proposed to represent the distortion ratio between two image blocks with a reference relationship.

[0075] The following formula shows how to determine the relationship between the direct distortion dependency and the cumulative distortion ratio 1+σ.

[0076] S41, taking the rate distortion minimization on the block-level temporal reference relationship chain as the optimization goal.

[0077] The optimization goal is set as Where D represents image distortion and R represents image block bit rate.

[0078] Assume that the distortion between image blocks i and j, which have a direct or indirect reference relationship, has the following dependency relationship (i <j):

[0079]

[0080] Assuming that the distortion dependency between image blocks in the spatial domain and the code rate dependency in the spatiotemporal domain are ignored, that is, changing the quantization parameter of the current image block has no effect on the overall code rate of other image blocks in the spatial domain and the image blocks in the temporal reference relationship chain, the rate-distortion cost in the temporal reference relationship chain can be rewritten as The cumulative distortion offset σ is the cumulative multiplication of the distortion dependency β on the time domain reference relationship chain, expressed as the cumulative distortion offset of image block 0, where image block 0 is the head reference block on the time domain reference relationship chain.

[0081] S42, modeling the distortion dependency term β as the cumulative product of the inter-frame probability ρ, the distortion transfer ratio θ, and the quantization error ratio α.

[0082] The distortion dependency term β is the ratio of the distortion of the current image block to the distortion of the reference image block under the condition of inter-frame coding. The inter-frame probability ρ is used to indicate the probability of the distortion being blocked in the time domain reference relationship chain. When the inter-frame probability is 0, it means that the image block selects intra-frame coding, and the distortion is not propagated in the time domain propagation chain. The distortion dependency term β is modeled as the cumulative multiplication of the inter-frame probability ρ, the distortion transfer ratio θ, and the quantization error ratio α.

[0083] S421, calculate and obtain the inter-frame probability ρ.

[0084] The inter-frame probability ρ is used to represent the probability of distortion being blocked in the time domain reference relationship chain, which is specifically determined by the relative size of the intra-frame prediction cost and the inter-frame prediction cost. There are many functional expressions, such as using the simplest linear relationship, using the softmax function form, and using the sigmoid function form. The formula is as follows.

[0085]

[0086] Where ratio is the ratio of the best inter-frame prediction cost InterCost to the best intra-frame prediction cost IntraCost, and a and b are empirical constants. The empirical constants can be calculated by statistically analyzing the actual inter-frame probabilities of image blocks in the offline statistical sequence. Figure 2 As shown in FIG. 1 , when the inter-frame probability is less than 0.5, the propagation process of the time domain distortion is rejected, and the distortion dependency term β is directly set to 0. When the inter-frame probability is greater than or equal to 0.5, the inter-frame probability is taken as part of the distortion dependency term β and is cumulatively multiplied with the distortion transfer ratio θ and the quantization error ratio α.

[0087] S422, calculate and obtain the distortion transfer ratio θ.

[0088] The calculation is performed using the variance σ and quantization step size Q of the current image block and the reference block, and the residual variance σ0 of the current image block and the reference block.

[0089] In one embodiment of the present invention, the calculation formula is as follows.

[0090]

[0091] Although the distortion transfer ratio θ uses the quantization step size Q to estimate the actual distortion dependency β, during the encoding process, the quantization distortion of the image block needs to be accurately represented through the quantization process. The same statistics representing the frame content may also show different content characteristics, so the additional quantization error ratio α is used to approach the true distortion dependency β.

[0092] S423, calculating the quantization error ratio α.

[0093] It should be noted that the distortion transfer ratio θ utilizes the quantization step size Q to estimate the actual distortion dependency β.

[0094] However, during the encoding process, the quantization distortion of the image block needs to be accurately represented through the quantization process. The same statistics representing the frame content may also show different content characteristics, so the additional quantization error ratio α is used to approach the true distortion dependency β.

[0095] The quantization error ratio is used to indicate the influence of the reconstruction distortion of the reference block on the distortion dependency between the reference block and the image block. The reconstruction distortion can be represented by the relative size of the reconstruction residual and the prediction residual. The value range of the quantization error ratio is [0, 1].

[0096] In one embodiment of the present invention, when the reconstructed residuals of the image block and the reference block are all quantized to 0, it means that the distortion of the reference block can be completely transferred, and the quantization error ratio α is 1. When the reconstructed residuals of the image block and the reference block are completely consistent with the original residuals, it becomes lossless compression, and there is no distortion transfer of the reference block, and the quantization error ratio α is 0.

[0097] Therefore, the specific calculation method of the quantization error ratio α is as follows:

[0098] S4231, calculate the prediction residual R of the current image block and the reference block k , and then the prediction residual R k Perform quantization, transformation, inverse quantization, and inverse transformation operations to obtain the reconstructed residual

[0099] S4232, calculate the prediction residual R k and the reconstruction residual The central moment of the difference is obtained to obtain the quantization error D k0 , expressed as At the same time, calculate the prediction residual R k The central moment Expressed as

[0100] S4233, the quantization error ratio α is expressed as the quantization error D k0 and the prediction error central moment ratio.

[0101] S5, determining a time domain quantization parameter offset, calculating a time domain quantization parameter offset of each image block of a reference frame to be encoded in the forward-looking image set based on the accumulated distortion offset, and encoding the reference frame to be encoded based on the time domain quantization parameter offset.

[0102] The cumulative distortion ratio is converted into the quantization parameter offset ΔQP of each image block in the reference frame to be encoded in the forward-looking image set. Since the Lagrangian factor λ and the quantization parameter QP have a power series relationship, the temporal quantization parameter offset ΔQP of each image block in the reference frame is expressed as ΔQP = strength*log 2 (1+σ), 1+σ is the cumulative distortion ratio. The time domain quantization parameter offset ΔQP is output block by block, and the reference frame to be encoded is encoded based on the time domain quantization parameter offset ΔQP.

[0103] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.

Claims

1. A video coding method based on conditional probability and reconstruction distortion propagation optimization, It is characterized in that include: Constructing a forward-looking image set: caching a number of frames of the video segment to be encoded as an original image set, and generating a forward-looking image set based on the original image set; Decision coding structure: a decision coding structure for a forward-looking image set, wherein the coding structure at least includes the setting of a frame type for each frame and the setting of a reference frame list; Constructing a block-level temporal reference relationship chain: Calculating the prediction cost of each image block corresponding to each frame in the forward-looking image set based on the reference frame list, and determining the reference block based on the prediction cost to construct a block-level temporal reference relationship chain; Calculating the cumulative distortion offset: constructing distortion dependencies based on the block-level temporal reference relationship chain, and calculating the cumulative distortion offset based on the distortion dependencies; Determine a temporal quantization parameter offset: calculate a temporal quantization parameter offset of each image block of a reference frame to be encoded in the forward-looking image set based on the accumulated distortion offset, and encode the reference frame to be encoded based on the temporal quantization parameter offset.

2. The video coding method based on conditional probability and reconstruction distortion propagation optimization according to claim 1, It is characterized in that The step of caching a plurality of frames of the video to be encoded as an original image set and generating a forward-looking image set based on the original image set specifically includes: Generate four double downsampled frames with different difference directions for each frame in the original image set. And perform a bracketing operation on the double down-sampled frame in each difference direction, so as to generate four extended edge down-sampled frames in different interpolation directions; A set of forward-looking images is generated based on the original set of images from which the expanded and downsampled frames are generated.

3. The video coding method based on conditional probability and reconstruction distortion propagation optimization according to claim 1, It is characterized in that The settings for the reference frame list for the forward-looking image set decision include: Based on the reference frame cache list, at least one reference frame list is preset for each image in the forward-looking image set, wherein the reference frame list is used to store at least the POC of each reference frame, the POC of the current image and the corresponding POC difference; Traverse all images in the forward-looking image set based on the coding order, and fill the reference frames in the reference frame cache list into the reference frame list of the current image; If the current image frame is determined to be a reference frame based on the frame type, it is added to the reference frame cache list for setting of subsequent image reference frame lists.

4. The video coding method based on conditional probability and reconstruction distortion propagation optimization according to claim 3, It is characterized in that Also includes: After all images in the forward-looking image set are traversed, the reference frame cache list is rolled back to the state where the encoding of the first encoding structure in the forward-looking image set is completed.

5. The video coding method based on conditional probability and reconstruction distortion propagation optimization according to claim 1, It is characterized in that The step of calculating the prediction cost of each image block corresponding to each frame of the forward-looking image set based on the reference frame list specifically includes: Divide each frame of the forward-looking image set into a number of image blocks of the same size; Calculate the optimal inter-frame prediction cost and the optimal intra-frame prediction cost for each image block; The reference block with the minimum prediction cost is calculated based on the best intra-frame prediction cost and the best inter-frame prediction cost.

6. The video coding method based on conditional probability and reconstruction distortion propagation optimization according to claim 5, It is characterized in that Calculating the optimal inter-frame prediction cost for each image block specifically includes: Perform motion search on each image block based on the reference frames in the reference frame list to obtain the best reference block on each reference frame. Based on the pixel values ​​of the image block and the best reference block of each reference frame, the pixel residual is obtained and Hadamard transform is performed, and the sum of the absolute values ​​of all coefficients in the transform coefficient matrix is ​​calculated and used as the inter-frame prediction cost; The inter-frame prediction cost of the current image block to the best reference block on each reference frame is compared to obtain the best inter-frame prediction cost, and the inter-frame prediction direction, reference frame number and motion vector are saved.

7. The video coding method based on conditional probability and reconstruction distortion propagation optimization according to claim 5, It is characterized in that Calculating the optimal intra prediction cost for each image block specifically includes: The Hadamard transform costs of the DC mode and the Planner mode of each image block in the forward-looking image set are calculated as the intra-frame prediction cost, and the smaller value of the intra-frame prediction cost of the two modes is selected as the intra-frame prediction cost of the optimal non-angle mode; Searching for an angle mode based on a preset step size to obtain an optimal angle mode, and calculating an intra-frame prediction cost of the optimal angle mode; The intra prediction costs of the best non-angle mode and the best angle mode are compared, and the smaller value of the two is taken as the best intra prediction cost of the image block.

8. The video coding method based on conditional probability and reconstruction distortion propagation optimization according to claim 1, It is characterized in that The constructing of distortion dependencies based on the block-level temporal reference relationship chain specifically includes: Rearrange the images in the forward-looking image set based on the coding order, and traverse each image block in the rearranged forward-looking image set in reverse order; The cumulative distortion ratio is determined based on the direct distortion dependency through the rate-distortion cost on the block-level time domain propagation chain, wherein the cumulative distortion ratio on the block-level time domain reference relationship chain is β is the direct distortion dependency between two image blocks in the direct reference relationship of the block-level temporal reference relationship chain, and n is the number of image blocks in the block-level temporal reference relationship chain.

9. The video coding method based on conditional probability and reconstruction distortion propagation optimization according to claim 8, It is characterized in that The direct distortion dependency term β=ρ*θ*α, wherein ρ is the inter-frame probability, θ is the distortion transfer ratio, and α is the quantization error ratio.

10. The video coding method based on conditional probability and reconstruction distortion propagation optimization according to claim 8, It is characterized in that The step of calculating the time domain quantization parameter offset of each image block of the reference frame to be encoded in the forward-looking image set based on the accumulated distortion offset, and encoding the reference frame to be encoded based on the time domain quantization parameter offset, specifically includes: The cumulative distortion ratio is converted into the quantization parameter offset ΔQP of each image block in the reference frame to be encoded in the forward-looking image set, where the temporal quantization parameter offset ΔQP of each image block in the reference frame is expressed as ΔQP = strength*log 2 (1+σ), 1+σ is the cumulative distortion ratio; A time domain quantization parameter offset ΔQP is output block by block, and a reference frame to be encoded is encoded based on the time domain quantization parameter offset ΔQP.

Citation Information

Cited By

  • Video encoding method, video encoder, system and program product

    CN120614461A