Coding, decoding method, transmission method, coding, decoding device and system
By generating virtual frames on the encoding and decoding ends and setting decoding identifiers, the problem of requiring simultaneously encoding motion vectors and residuals in inter-frame prediction mode is solved, and more efficient coding performance is achieved.
Patent Information
- Application Number
- CN201910626135.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-07-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-07-11
AI Technical Summary
In the prior art, motion vectors and prediction residuals need to be encoded simultaneously in the inter prediction mode, resulting in lower encoding performance.
By generating virtual blocks on the encoding and decoding ends based on the pixel values of the encoding tree unit CTU at the same position in the reference frame and its adjacent position pixels, splicing to form a virtual frame, setting a virtual frame decoding identifier, and inserting it into the code stream to determine the decoding method, and avoiding direct encoding of motion vectors and residuals.
The encoding rate is reduced, the encoding performance is improved, and the encoding and transmission of motion vectors and image residuals in traditional encoding modes is avoided.
Smart Images

Figure CN112218086B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data transmission, and more specifically, to a coding and decoding method, a transmission method, a coding and decoding device and a system. Background Art
[0002] Inter-frame prediction has become increasingly popular and has become a crucial module in current mainstream coding platforms such as H.264 / H.265 and AVS2. Inter-frame prediction refers to the use of the correlation in the video time domain to predict the pixels of the current image using the pixels of the adjacent encoded images to effectively remove the temporal redundancy of the video. Since video sequences usually include strong temporal correlation, the prediction residuals are usually "flat", that is, many residual values are close to "0". The residual signal is used as the input of subsequent modules for transformation, quantization, scanning and entropy coding, which can achieve efficient compression of the video signal.
[0003] In the traditional inter-frame prediction mode, each coding unit finds the corresponding reference block through a corresponding motion vector, and then reconstructs it by adding the prediction residual to the reference block pixel value. Therefore, for each coding unit, the motion vector and prediction residual need to be encoded, and this process consumes a large number of codewords. Among them, the use of Merge mode and AMVP technology in HEVC (High Efficiency Video Coding) saves the number of coding bits of motion information to a certain extent and improves the performance of inter-frame prediction. However, after testing and analysis, for some large motion scene sequences, the proportion of motion information coding bits is still not low. It can be seen that although the current continuously improved inter-frame prediction modes and methods have reduced the number of coding bits for motion information to a certain extent, it is still necessary to encode the motion vector and prediction residual at the same time. Therefore, further reducing or even avoiding the number of coding bits for motion information and residual information has become one of the ways to further improve video compression performance. Summary of the invention
[0004] The encoding and decoding methods, transmission methods, encoding and decoding devices and systems provided by the embodiments of the present invention mainly solve the technical problem that in the prior art, the motion vector and the prediction residual need to be encoded simultaneously in the inter-frame prediction mode, and the encoding performance is low.
[0005] In order to solve the above technical problems, an embodiment of the present invention provides an encoding method, which is applied to an encoding end. The encoding method includes:
[0006] Generate virtual blocks according to pixel values of pixels of coding tree units CTU and adjacent positions at the same position in the reference frame, and splice the virtual blocks to form a virtual frame;
[0007] Setting a virtual frame decoding flag of the current frame;
[0008] Inserting the virtual frame decoding identifier into the code stream of the current frame;
[0009] Send the code stream of the current frame to the decoding end.
[0010] In order to solve the above technical problems, an embodiment of the present invention further provides a decoding method, which is applied to a decoding end. The decoding method includes:
[0011] Generate virtual blocks according to pixel values of pixels of coding tree units CTU and adjacent positions at the same position in the reference frame, and splice the virtual blocks to form a virtual frame;
[0012] Determine the decoding mode of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding end;
[0013] The current frame is decoded and reconstructed according to the determined decoding method.
[0014] The embodiment of the present invention further provides a transmission method, the transmission method comprising:
[0015] The encoder generates virtual blocks according to pixel values of pixels of coding tree units CTU and adjacent positions at the same position in the reference frame, and splices the virtual blocks to form a virtual frame; sets a virtual frame decoding identifier of the current frame; inserts the virtual frame decoding identifier into the bitstream of the current frame; and sends the bitstream of the current frame to the decoder;
[0016] The decoding end generates each virtual block according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, and splices the virtual blocks to form a virtual frame; determines the decoding method of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding end; and decodes and reconstructs the current frame according to the determined decoding method.
[0017] The embodiment of the present invention further provides an encoding device, the encoding device comprising: a first generating module, a setting module, an inserting module and a sending module;
[0018] The first generation module is used to generate each virtual block according to the pixel values of each coding tree unit CTU at the same position in the reference frame and the pixels at the adjacent position, and splice the virtual blocks to form a virtual frame;
[0019] The setting module is used to set the virtual frame decoding identifier of the current frame;
[0020] The inserting module is used to insert the virtual frame decoding identifier into the code stream of the current frame;
[0021] The sending module is used to send the code stream of the current frame to the decoding end.
[0022] The embodiment of the present invention further provides a decoding device, the decoding device comprising: a second generating module, a determining module and a decoding module;
[0023] The second generation module is used to generate each virtual block according to the pixel values of each coding tree unit CTU at the same position in the reference frame and the pixels at the adjacent position, and splice the virtual blocks to form a virtual frame;
[0024] The determination module is used to determine the decoding mode of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding end;
[0025] The decoding module is used to decode and reconstruct the current frame according to the determined decoding method.
[0026] The embodiment of the present invention further provides a system, the system comprising: an encoding device and a decoding device;
[0027] The encoding device is used to generate each virtual block according to the pixel values of each coding tree unit CTU and the pixels at the adjacent position in the reference frame, and splice the virtual blocks to form a virtual frame; set the virtual frame decoding identifier of the current frame; insert the virtual frame decoding identifier into the code stream of the current frame; send the code stream of the current frame to the decoding device;
[0028] The decoding device is used to generate each virtual block according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, and splice the virtual blocks to form a virtual frame; determine the decoding method of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding device; and decode and reconstruct the current frame according to the determined decoding method.
[0029] The beneficial effects of the present invention are:
[0030] The encoding, decoding method, transmission method, encoding, decoding device and system provided by the embodiment of the present invention generate virtual blocks according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, splice each virtual block to form a virtual frame, and then set the virtual frame decoding identifier of the current frame, insert the virtual frame decoding identifier into the code stream of the current frame, and further send the code stream of the current frame to the decoding end; the decoding end generates virtual blocks according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, splice each virtual block to form a virtual frame, and then determine the decoding method of the current frame according to the virtual frame decoding identifier received from the encoding end in the current frame code stream, and further, decode and reconstruct the current frame according to the determined decoding method; solve the problem of the need to simultaneously encode the motion vector and the prediction residual in the inter-frame prediction mode in the prior art, and the low encoding performance. That is, the encoding, decoding method, transmission method, encoding, decoding device and system provided by the embodiment of the present invention avoid the encoding and transmission of the motion vector and the image residual in the traditional encoding mode to a certain extent, reduce the encoding bit rate, and improve the encoding performance.
[0031] Other features and corresponding beneficial effects of the present invention are described in the latter part of the specification, and it should be understood that at least part of the beneficial effects become obvious from the description in the specification of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0033] Figure 1 A schematic diagram of a basic flow of an encoding method provided in Embodiment 1 of the present invention;
[0034] Figure 2 A schematic diagram of a reference block selection method provided in Embodiment 1 of the present invention;
[0035] Figure 3 A schematic diagram of generating a virtual block according to a reference block provided in Embodiment 1 of the present invention;
[0036] Figure 4 A schematic diagram of splicing virtual blocks into a complete virtual frame provided in the first embodiment of the present invention;
[0037] Figure 5 A schematic diagram of a basic flow of setting a virtual frame decoding flag of a current frame provided in Embodiment 1 of the present invention;
[0038] Figure 6 A schematic diagram of a basic flow of inserting a corresponding virtual block decoding identifier into a bit stream of a CTU of a current frame provided in the first embodiment of the present invention;
[0039] Figure 7 A basic flow chart of a decoding method provided in Embodiment 1 of the present invention;
[0040] Figure 8 A basic flow chart of a method for determining a decoding mode of a current frame provided in the first embodiment of the present invention is to decode and reconstruct a bit stream of the current frame;
[0041] Fig. 9 A basic flow chart of a frame-level encoding and decoding method provided in Embodiment 2 of the present invention;
[0042] Fig.10 A basic flow chart of a CTU-level encoding and decoding method provided in Embodiment 2 of the present invention;
[0043] Fig.11 A basic flow chart of a frame-level encoding and decoding method provided in Embodiment 3 of the present invention;
[0044] Fig.12 A basic flow chart of a CTU-level encoding and decoding method provided in Embodiment 3 of the present invention;
[0045] Fig.13 A schematic diagram of the structure of an encoding device provided in Embodiment 5 of the present invention;
[0046] Fig.14 A schematic diagram of the structure of a decoding device provided in Embodiment 5 of the present invention;
[0047] Fig.15 This is a schematic diagram of the structure of the system provided in Example 6 of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the following is a further detailed description of the embodiments of the present invention through specific implementation methods combined with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0049] Embodiment 1:
[0050] In order to solve the problem of low coding performance in the prior art that the motion vector and the prediction residual need to be encoded simultaneously in the inter-frame prediction mode, in an embodiment of the present invention, each virtual block is generated according to the pixel values of each coding tree unit CTU and the pixels at the adjacent position in the reference frame, and each virtual block is spliced to form a virtual frame, and then the virtual frame decoding flag of the current frame is set, the virtual frame decoding flag is inserted into the bitstream of the current frame, and the bitstream of the current frame is sent to the decoding end; please refer to Figure 1 As shown, Figure 1 A schematic diagram of the basic flow of the encoding method provided in this embodiment.
[0051] S101: Generate each virtual block based on the pixel values of each coding tree unit (CTU) at the same position in the reference frame and the pixels at adjacent positions, and splice the virtual blocks to form a virtual frame.
[0052] For better understanding, a specific example is used here for illustration. For example, refer to Figure 2 、 3 As shown, taking the center point of the current CTU as the base point, in the two reference frames, namely the previous frame and the next frame of the current frame, select a reference block with a length and width twice that of the current CTU. Through algorithms such as deep learning, generate virtual blocks, and based on this, generate each virtual block; furthermore, refer to Figure 4 As shown, splice the generated virtual blocks to form a complete virtual frame. It should be noted that in practical applications, the number of reference frames and the selection of the length and width of the reference block can be flexibly adjusted according to the specific application scenario.
[0053] Optionally, in this embodiment, when generating the virtual block of the CTU through the deep learning algorithm, first use the deep learning method to pre-train a large number of sample sequences to generate a deep learning network for generating the virtual block of the CTU at the encoding end; it should be clear that the encoding end will send the generated deep learning network to the decoding end so that the decoding end can use this deep learning network to generate the same virtual block of the CTU as the encoding end. It should be noted that here only the deep learning algorithm for generating the virtual block of the CTU is taken as an example. In practical applications, it depends on the specific algorithm adopted. In fact, as long as it can be ensured that the encoding end and the decoding end can generate the same virtual block of the CTU, it is within the protection scope of the present invention, and the present invention does not make specific limitations in this regard.
[0054] S102: Set the virtual frame decoding flag for the current frame.
[0055] In this embodiment, setting the virtual frame decoding flag for the current frame includes at least the following steps. Specifically, refer to Figure 5 :
[0056] S501: Determine the objective quality of the virtual frame compared with the original image of the current frame.
[0057] It can be understood that after splicing the virtual blocks to form a virtual frame, compare it with the original image of the current frame to determine the objective quality of the virtual frame; optionally, use the peak signal-to-noise ratio (PSNR) of the virtual frame as the objective quality of the virtual frame, where the PSNR of the virtual frame can be calculated through Formulas 1 and 2 shown below. For example, assume that the determined PSNR of the virtual frame is A1;
[0058] Formula 1:
[0059] Among them, MSE is the mean square error between the current virtual frame and the current original image, W is the image width, H is the image height, I(i,j) is the pixel value of the current virtual frame, and K(i,j) is the pixel value of the original image of the current frame.
[0060] Formula 2: Where n is the bit depth of the image pixels.
[0061] Optionally, the structural similarity index (SSIM) of the virtual frame is used as the objective quality of the virtual frame, for example, the SSIM of the determined virtual frame is set to A1. It is worth noting that the two common ways of determining the objective quality of the virtual frame are listed here, and the present invention is not limited to these two ways. In practical applications, they can be flexibly adjusted according to specific application scenarios.
[0062] It can also be understood that in some examples of this embodiment, the virtual blocks can be spliced together to form a virtual frame and then compared with the original image of the current frame to determine the distortion value of the virtual frame, where the distortion value of the virtual frame can be calculated by Formula 1 in the above figure; it should be clear that the distortion value of the virtual frame can also be obtained by calculating MSE, SAD, SSD, SSE, and SATD, and the present invention is not limited to this. In actual applications, it can be flexibly adjusted according to specific application scenarios.
[0063] S502: Setting different values of virtual frame decoding flags according to objective quality.
[0064] In this embodiment, different values of the virtual frame decoding flag are set according to the objective quality, including the following two cases:
[0065] In case 1, when the objective quality of the virtual frame is less than or equal to a preset threshold, the value of the virtual frame decoding flag is set to instruct the decoding end to reconstruct the current frame using the virtual frame of the decoding end.
[0066] It can be understood that when the objective quality A1 of the virtual frame is less than or equal to the preset threshold A, the virtual frame decoding flag is optionally set to 1 to instruct the decoding end to reconstruct the current frame using the virtual frame of the decoding end, that is, skip encoding the current frame and reduce the encoding bit rate.
[0067] In the second case, when the objective quality of the virtual frame is greater than a preset threshold, the value of the virtual frame decoding flag is set to the flag to instruct the decoding end to decode and reconstruct the current frame.
[0068] It is understandable that when the objective quality A1 of the virtual frame is greater than the preset threshold A, the virtual frame decoding flag is optionally set to 0 to instruct the decoding end to decode and reconstruct the current frame, that is, the current frame still needs to be encoded.
[0069] Optionally, the preset threshold in this embodiment is calculated based on the PSNR of the forward N-frame reference frame of the current frame and / or the PSNR of the backward M-frame reference frame of the current frame, where N and M are integers, and N and M are greater than or equal to 1. Optionally, the preset threshold can be calculated based on the PSNR of the forward N-frame reference frame of the current frame only, or the preset threshold can be calculated based on the PSNR of the backward N-frame reference frame of the current frame only, or the preset threshold can be calculated based on the PSNR of the forward N-frame reference frame and the PSNR of the backward M-frame reference frame of the current frame.
[0070] Optionally, the preset threshold in this embodiment is calculated based on the SSIM of the forward N-frame reference frame of the current frame and / or the SSIM of the backward M-frame reference frame of the current frame, where N and M are integers, and N and M are greater than or equal to 1. Optionally, the preset threshold can be calculated based on the SSIM of the forward N-frame reference frame of the current frame only, or the preset threshold can be calculated based on the SSIM of the backward N-frame reference frame of the current frame only, or the preset threshold can be calculated based on the SSIM of the forward N-frame reference frame and the SSIM of the backward M-frame reference frame of the current frame.
[0071] It can be understood that in this embodiment, when the preset threshold is obtained by preset calculation based on PSNR, the preset calculation can be to calculate the average value of the PSNR of the forward N frame reference frame and / or the PSNR of the backward M frame reference frame of the current frame; the preset calculation can also be to calculate the weighted average value of the PSNR of the forward N frame reference frame and / or the PSNR of the backward M frame reference frame of the current frame.
[0072] It can also be understood that, in this embodiment, when the preset threshold is obtained by preset calculation based on SSIM, the preset calculation can be to calculate the average value of the SSIM of the forward N frame reference frame and / or the SSIM of the backward M frame reference frame of the current frame; the preset calculation can also be to calculate the weighted average value of the SSIM of the forward N frame reference frame and / or the SSIM of the backward M frame reference frame of the current frame.
[0073] It is worth noting that what are listed here are only two common preset calculations. The present invention is not limited to these two preset calculations. In practical applications, they can be flexibly adjusted according to specific application scenarios.
[0074] S103: Insert a virtual frame decoding identifier into the code stream of the current frame.
[0075] It should be clear that an image frame is composed of one or more slices, wherein a slice is composed of one or more CTUs.
[0076] In this embodiment, in case 1 (i.e., when the objective quality is less than or equal to the preset threshold), inserting the virtual frame decoding identifier into the bitstream of the current frame includes: inserting the virtual frame decoding identifier into the frame header bitstream of the current frame. In this embodiment, the frame header may be a slice header of an image, that is, inserting the virtual frame decoding identifier into the bitstream of the slice header of the current frame.
[0077] In this embodiment, in case 2 (i.e., when the objective quality is greater than a preset threshold), inserting a virtual frame decoding identifier into the bitstream of the current frame includes: inserting the virtual frame decoding identifier into the frame header bitstream of the current frame, and inserting a corresponding virtual block decoding identifier into the bitstream of the CTU of the current frame.
[0078] It should be noted that, in the present embodiment, when the residual frame obtained by using the virtual frame as the predicted image of the current frame is not zero, it is necessary to perform transformation, quantization, entropy coding and other coding steps on the residual frame to generate a bit stream, and the subsequent frame-level coding is completed; when the residual frame obtained by using the virtual frame as the predicted image of the current frame is zero, there is no need to perform transformation, quantization, entropy coding and other coding steps on the residual frame, and the frame-level coding is completed.
[0079] In this embodiment, inserting the corresponding virtual block decoding identifier into the bit stream of the CTU of the current frame includes at least the following steps. Figure 6 As shown:
[0080] S601: Compare the rate-distortion RD cost of the CTU with the rate-distortion RD cost of the virtual block corresponding to the CTU.
[0081] In this embodiment, comparing the rate-distortion RD cost of the CTU with the rate-distortion RD cost of the virtual block corresponding to the CTU includes: calculating the optimal coding mode of the CTU of the current frame and the rate-distortion RD cost of each sub-block in the CTU of the current frame under the coding mode, adding the RD cost of each sub-block of the CTU to obtain the RD cost of the CTU; calculating the RD cost of the virtual block corresponding to the CTU; and then comparing the RD cost of the CTU with the RD cost of the virtual block corresponding to the CTU.
[0082] S602: According to the comparison result, different values of virtual block decoding identifiers are set.
[0083] In this embodiment, different values of the virtual block decoding flag are set according to the comparison result, including the following two cases:
[0084] In case 1, when the RD cost of the virtual block is less than or equal to the RD cost of the CTU, the value of the virtual block decoding flag is set to instruct the decoding end to use the virtual block for decoding and reconstruction.
[0085] It can be understood that when the RD cost of the virtual block is less than or equal to the RD cost of the CTU, the virtual block decoding flag is optionally set to a value of 1 to instruct the decoding end to reconstruct the CTU using the virtual block of the decoding end, that is, to generate a bitstream using the virtual block as the prediction block of the CTU, that is, there is no need to encode the CTU, so as to reduce the encoding bit rate by avoiding encoding of motion information and residual.
[0086] In case 2, when the RD cost of the virtual block is greater than the RD cost of the CTU, the value of the virtual block decoding flag is set to instruct the decoding end to use the bit stream of the CTU of the current frame for decoding and reconstruction.
[0087] It is understandable that when the RD cost of the virtual block is greater than the RD cost of the CTU, the virtual block decoding flag is optionally set to a value of 0 to instruct the decoder to use the bitstream of the CTU of the current frame for decoding and reconstruction, that is, the CTU still needs to be encoded to generate the bitstream.
[0088] In this embodiment, in case 1 (i.e., when the RD cost of the virtual block is less than or equal to the RD cost of the CTU), after inserting the corresponding virtual block decoding identifier, it includes: encoding the residual block obtained by the CTU and the virtual block to generate a code stream of the CTU of the current frame.
[0089] In this embodiment, in case 2 (ie, when the RD cost of the virtual block is greater than the RD cost of the CTU), after inserting the corresponding virtual block decoding identifier, the following steps include: encoding the CTU to generate a code stream of the CTU of the current frame.
[0090] It should be noted that, in this embodiment, when the residual block obtained by using the virtual block as the CTU prediction block of the current frame is not zero, the residual block needs to be transformed, quantized, entropy encoded and other encoding steps to generate a bit stream, and the subsequent CTU-level encoding is completed; when the residual block obtained by using the virtual block as the CTU prediction block of the current frame is zero, there is no need to transform the residual block, quantize, entropy encode and other encoding steps, and the CTU-level encoding is completed.
[0091] S104: Send the code stream of the current frame to the decoding end.
[0092] It can be understood that when the virtual frame decoding identifier is inserted into the frame header code stream of the current frame, the frame header code stream of the current frame needs to be sent to the decoding end; when the virtual block decoding identifier is inserted into the code stream of the CTU of the current frame, the code stream of the CTU of the current frame needs to be sent to the decoding end.
[0093] In order to solve the problem of low coding performance in the prior art that the motion vector and the prediction residual need to be encoded simultaneously in the inter-frame prediction mode, in an embodiment of the present invention, each virtual block is generated according to the pixel values of each coding tree unit CTU and the pixels at the adjacent position in the reference frame, and each virtual block is spliced to form a virtual frame, and then the decoding mode of the current frame is determined according to the virtual frame decoding identifier received in the current frame code stream from the encoding end, and the current frame is decoded and reconstructed according to the determined decoding mode; please refer to Figure 7 As shown, Figure 7 A schematic diagram of the basic flow of the decoding method provided in this embodiment.
[0094] S701: Generate virtual blocks according to pixel values of pixels of coding tree units CTUs at the same position in a reference frame and pixels at adjacent positions, and splice the virtual blocks to form a virtual frame.
[0095] It should be clear that S701 and S101 are the same, the only difference is that one is executed at the decoding end and the other is executed at the encoding end. In other words, both the decoding end and the encoding end will generate each virtual block according to the pixel values of each coding tree unit CTU at the same position in the reference frame and the pixel values of its adjacent position, and splice the virtual blocks to form a virtual frame, which will not be repeated here.
[0096] S702: Determine a decoding method for the current frame according to a virtual frame decoding identifier received from the encoding end in the current frame code stream.
[0097] In this embodiment, there are two situations in which the decoding mode of the current frame is determined according to the virtual frame decoding identifier received in the current frame code stream from the encoder:
[0098] In case 1, when the value of the virtual frame decoding identifier parsed from the bitstream indicates that the decoder reconstructs the current frame using the virtual frame of the decoder, the decoding method of the current frame is determined to be to reconstruct the current frame using the virtual frame.
[0099] Optionally, when the value of the parsed virtual frame decoding flag is 1, the current frame is directly reconstructed using the virtual frame generated by the decoding end. Compared with the prior art, the decoding of the current frame is skipped, which improves the encoding performance to a certain extent.
[0100] It should be noted that, in this embodiment, there are two situations in which the current frame is reconstructed using a virtual frame:
[0101] First, when the residual frame parsed from the bitstream of the current frame is not zero, it is necessary to parse the residual coefficients and perform inverse transformation, inverse quantization, entropy decoding and other decoding steps, and then combine the residual frame and the virtual frame to obtain a reconstructed frame.
[0102] Second, when the residual frame parsed from the code stream of the current frame is zero, the virtual frame is used as the reconstructed image of the current frame to obtain a reconstructed frame.
[0103] In case 2, when the value of the virtual frame decoding identifier parsed from the bitstream is an identifier indicating that the decoding end decodes and reconstructs the current frame, the decoding method of the current frame is determined to be decoding and reconstructing the bitstream of the current frame.
[0104] Optionally, when the value of the parsed virtual frame decoding flag is 0, the existing method is still used, and the code stream of the current frame needs to be decoded and reconstructed.
[0105] In this embodiment, when determining the decoding mode of the current frame is to decode and reconstruct the code stream of the current frame, at least the following steps are included. Figure 8 As shown:
[0106] S801: Determine a decoding mode of a CTU according to a virtual block decoding identifier received in a CTU code stream from an encoding end.
[0107] In this embodiment, the decoding mode of the CTU is determined according to the virtual block decoding identifier received in the CTU code stream from the encoder, including the following two cases:
[0108] In case 1, when the value of the virtual block decoding identifier parsed from the bitstream indicates that the decoder reconstructs the CTU using the virtual block of the decoder, the decoding method of the CTU is determined to be to reconstruct the CTU using the virtual block.
[0109] Optionally, when the value of the parsed virtual block decoding identifier is 1, the virtual block generated by the decoding end is directly used as the prediction block of the CTU, and is reconstructed in combination with the residual block parsed from the bitstream. Compared with the existing method, the decoding of the CTU is skipped, which improves the encoding performance to a certain extent.
[0110] It should be noted that, in this embodiment, there are two situations in which the decoding method of determining the CTU is to use the virtual block as the prediction block of the CTU and to reconstruct the residual block obtained by parsing the bitstream:
[0111] First, when the residual block parsed from the bitstream of the CTU is not zero, it is necessary to parse the residual coefficients and perform inverse transformation, inverse quantization, entropy decoding and other decoding steps, and then combine the residual block and the virtual block to obtain the reconstructed CTU.
[0112] Second, when the residual block parsed from the bitstream of the CTU is zero, the virtual block is used as the prediction block of the current frame to obtain the reconstructed CTU.
[0113] In case 2, when the value of the virtual block decoding identifier parsed from the bitstream indicates that the decoding end uses the bitstream of the CTU of the current frame for decoding and reconstruction, the decoding mode of the CTU is determined to be decoding and reconstruction of the bitstream of the CTU.
[0114] Optionally, when the value of the parsed virtual block decoding identifier is 0, the existing method is still used, for example, the motion vector and the residual block are parsed from the CTU code stream, and then the motion vector is used for motion compensation to obtain the prediction block, and the residual block is combined for decoding and reconstruction.
[0115] S802: Decode and reconstruct the CTU according to the determined decoding method.
[0116] It can be understood that when it is determined that the decoding method of CTU is to reconstruct CTU using virtual blocks, the virtual blocks generated by the decoding end need to be used to reconstruct CTU; when it is determined that the decoding method of CTU is to decode and reconstruct the bitstream of CTU, the bitstream of CTU needs to be decoded and reconstructed.
[0117] S703: Decode and reconstruct the current frame according to the determined decoding method.
[0118] It can be understood that when it is determined that the decoding method of the current frame is to reconstruct the current frame using a virtual frame, the virtual frame generated by the decoding end is required to reconstruct the current frame; when it is determined that the decoding method of the current frame is to decode and reconstruct the code stream of the current frame, the code stream of the current frame needs to be decoded and reconstructed.
[0119] The encoding and decoding method provided by the embodiment of the present invention generates virtual blocks according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, splices the virtual blocks to form a virtual frame, and then sets the virtual frame decoding identifier of the current frame, inserts the virtual frame decoding identifier into the code stream of the current frame, and further sends the code stream of the current frame to the decoding end; the decoding end generates virtual blocks according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, splices the virtual blocks to form a virtual frame, and then determines the decoding method of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding end, and further, decodes and reconstructs the current frame according to the determined decoding method; solves the problem of the need to simultaneously encode the motion vector and the prediction residual in the inter-frame prediction mode in the prior art, and the low encoding performance. That is, the encoding and decoding method provided by the embodiment of the present invention avoids the encoding and transmission of the motion vector and the image residual in the traditional encoding mode to a certain extent, reduces the encoding bit rate, and improves the encoding performance.
[0120] Embodiment 2:
[0121] Based on the first embodiment, the embodiment of the present invention takes a specific encoding and decoding method as an example to further illustrate the embodiment of the present invention.
[0122] First see Fig. 9 As shown, the improvement is for the frame level:
[0123] S901: Generate virtual blocks at the encoding end and the decoding end simultaneously according to the pixel values of the CTUs at the same position in the reference frame and the pixels at the adjacent positions, and splice the generated virtual blocks into a complete virtual frame.
[0124] S902: Determine the objective quality of the virtual frame compared with the original image of the current frame.
[0125] Optionally, the objective quality of the two is determined by comparing the PSNR of the virtual frame.
[0126] S903: When the determined objective quality is less than or equal to a preset threshold, the virtual frame decoding flag is set to 1, the encoding of the current frame is skipped, and the virtual frame decoding flag is added to the bitstream of the current frame; when the decoding end receives the virtual frame decoding flag value of 1, the virtual frame is used to reconstruct the current frame.
[0127] Optionally, a preset threshold value P=f(PSNR1, PSNR2) is obtained according to the PSNRs of two reference frames.
[0128] S904: When the determined objective quality is greater than a preset threshold, the virtual frame decoding flag is set to 0, the current frame is encoded, and the virtual frame decoding flag is added to the bitstream of the current frame; when the decoding end receives that the virtual frame decoding flag has a value of 0, the bitstream of the current frame is decoded and reconstructed.
[0129] See also Fig.10 As shown, it is an improvement for the CTU level (it should be clear that Fig.10 It is a further step based on S904):
[0130] S9041: Calculate the optimal coding mode of the CTU of the current frame and the rate-distortion RD cost of each sub-block in the CTU of the current frame under the coding mode, add the RD cost of each sub-block of the CTU to obtain the RD cost of the CTU; calculate the RD cost of the virtual block corresponding to the CTU.
[0131] S9042: Compare the RD cost of the CTU with the RD cost of the virtual block corresponding to the CTU.
[0132] S9043: When the RD cost of the virtual block is less than or equal to the RD cost of the CTU, the virtual block decoding flag is set to 1 and written into the bitstream, and the residual block obtained according to the CTU and the virtual block is encoded and added to the bitstream of the CTU; when the value of the virtual block decoding flag received by the decoding end is 1, the virtual block is used to reconstruct the CTU.
[0133] S9044: When the RD cost of the virtual block is greater than the RD cost of the CTU, the virtual block decoding flag is set to 0 and written into the bitstream, and the CTU is encoded to generate the bitstream of the CTU; when the value of the virtual block decoding flag received by the decoding end is 0, the bitstream of the CTU is decoded and reconstructed.
[0134] The encoding and decoding method provided in the embodiment of the present invention is based on the existing video encoding framework, adopts the currently emerging deep learning method to generate virtual blocks of each CTU of the current frame, and splices them into a virtual frame of the current frame, and uses virtual frames and virtual blocks to add new encoding methods at the frame level and CTU level to further improve video compression performance. That is, when encoding the current frame, firstly, the pixel values of the coding tree units CTU and the pixels at the same position in the reference frame and the pixels at the adjacent positions are used to generate each virtual block, and the virtual blocks are spliced to form a virtual frame, and the encoding method is decided according to the distortion size of the virtual frame and the original image of the current frame, that is, if the distortion of the virtual frame and the original image of the current frame is less than a certain threshold, the virtual frame is used as the reconstructed image of the frame, and the encoding of the frame ends; otherwise, the CTU is further encoded, and for each CTU, the RDO selection is first performed according to the traditional method to obtain the best encoding mode and its RD cost. In addition, after calculating the RD cost of the CTU and its corresponding virtual block, the RD cost of the CTU is compared with the RD cost of the virtual block. If the former cost is small, the CTU is encoded according to the traditional method, otherwise, the corresponding virtual block is used as the prediction block of the CTU to predict the CTU to obtain the residual and encode the residual. To a certain extent, it avoids the encoding and transmission of motion vectors and image residuals in the traditional encoding mode, reduces the encoding bit rate, and improves the encoding performance.
[0135] Embodiment three:
[0136] Embodiments of the Present Invention Based on the first and second embodiments, further examples are given to illustrate the embodiments of the present invention.
[0137] First, a specific example of inserting a virtual frame decoding identifier to indicate that the decoding end uses the virtual frame of the decoding end to reconstruct the current frame is described. Fig.11 As shown:
[0138] The encoder uses the virtual frame as the predicted image of the current frame, that is, the generated virtual frame and the current frame are subtracted to obtain a residual frame;
[0139] In one example, there are two situations for the obtained residual frame: one is that the residual frame obtained by the encoding end through subtraction processing on the generated virtual frame and the current frame is not zero, and the other is that the residual frame obtained by the encoding end through subtraction processing on the generated virtual frame and the current frame is zero; when the residual frame is not zero, it is necessary to perform transformation, quantization, entropy coding and other coding steps on the residual frame to generate a bit stream, and when the residual frame is zero, the bit stream is directly generated.
[0140] In another example, there are two situations for the obtained residual frame: one is that when the objective quality of the virtual frame compared with the original image of the current frame is greater than a preset threshold, the obtained residual frame is set to be non-zero. At this time, the residual frame needs to be transformed, quantized, entropy encoded and other encoding steps to generate a bit stream; the second is that when the objective quality of the virtual frame compared with the original image of the current frame is less than or equal to a preset threshold, the obtained residual frame is set to zero and the bit stream is directly generated.
[0141] When the decoding end receives the code stream transmitted by the encoding end, it parses the code stream. When no residual frame exists in the analysis, the virtual frame generated by the decoding end is used as the reconstructed image of the current frame to obtain a reconstructed frame; when a residual frame exists in the analysis, the residual coefficients are parsed and then inversely transformed, inversely quantized, entropy decoded and other decoding steps are performed to obtain a residual frame, and then the residual frame and the virtual frame generated by the decoding end are added to obtain a reconstructed frame.
[0142] Next, a specific example is given in which the value of the virtual block decoding identifier is to instruct the decoding end to use the virtual block of the decoding end for decoding and reconstruction. Fig.12 As shown:
[0143] The encoder uses the corresponding virtual block as the prediction block of the CTU of the current frame, that is, performs subtraction processing on the generated virtual block and the CTU of the current frame to obtain a residual block.
[0144] In an example, there are two situations for the obtained residual block: one is that the residual block obtained by the encoder performing subtraction processing on the generated virtual block and the CTU of the current frame is not zero, and the other is that the residual block obtained by the encoder performing subtraction processing on the generated virtual block and the CTU of the current frame is zero; when the residual block is not zero, the residual block needs to be transformed, quantized, entropy encoded and other encoding steps to generate a bit stream, and when the residual block is zero, the bit stream is directly generated.
[0145] In another example, there are two cases for the obtained residual block: one is that when the RD cost of the virtual block is greater than the RD cost of the CTU of the current frame, the obtained residual block is set to be non-zero. At this time, the residual block needs to be transformed, quantized, entropy encoded and other encoding steps to generate a bitstream; the second is that when the RD cost of the virtual block is less than or equal to the RD cost of the CTU of the current frame, the obtained residual block is set to zero and the bitstream is directly generated.
[0146] When the decoder receives the bitstream transmitted by the encoder, it parses the bitstream. When no residual block exists in the analysis, the virtual block generated by the decoder is used as the prediction block of the current frame to obtain the reconstructed CTU. When a residual block exists in the analysis, the residual coefficient is parsed and then inversely transformed, inversely quantized, entropy decoded and other decoding steps are performed to obtain the residual block, and then the residual block and the virtual block generated by the decoder are added to obtain the reconstructed CTU.
[0147] The encoding and decoding method provided by the embodiment of the present invention avoids the encoding and transmission of motion vectors and image residuals in the traditional encoding mode to a certain extent, reduces the encoding bit rate, and improves the encoding performance.
[0148] Embodiment 4:
[0149] The embodiment of the present invention provides a transmission method based on the first, second and third embodiments:
[0150] The encoder generates virtual blocks according to the pixel values of the coding tree units CTU and the pixels at the same position in the reference frame and the pixels at the adjacent positions, and splices the virtual blocks to form a virtual frame; sets the virtual frame decoding identifier of the current frame; inserts the virtual frame decoding identifier into the bitstream of the current frame; and sends the bitstream of the current frame to the decoder.
[0151] The decoding end generates each virtual block according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, and splices each virtual block to form a virtual frame; determines the decoding method of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding end; and decodes and reconstructs the current frame according to the determined decoding method.
[0152] It is worth noting that in order to avoid redundant description, all examples in Embodiments 1, 2 and 3 are not fully described in this embodiment. It should be clear that all examples in Embodiments 1, 2 and 3 are applicable to this embodiment.
[0153] The transmission method provided by the embodiment of the present invention solves the problem in the prior art that the motion vector and the prediction residual need to be encoded simultaneously in the inter-frame prediction mode, resulting in low encoding performance. To a certain extent, it avoids the encoding and transmission of the motion vector and the image residual in the traditional encoding mode, reduces the encoding bit rate, and improves the encoding performance.
[0154] Embodiment five:
[0155] In order to solve the problem that the motion vector and the prediction residual need to be encoded simultaneously in the inter-frame prediction mode in the prior art, and the encoding performance is low, an encoding device is provided in an embodiment of the present invention. Fig.13 As shown:
[0156] The encoding device 13 includes a first generating module 1301, a setting module 1302, an inserting module 1303 and a sending module 1304, wherein:
[0157] The first generation module 1301 is used to generate each virtual block according to the pixel values of each coding tree unit CTU at the same position in the reference frame and the pixels at the adjacent position, and splice the virtual blocks to form a virtual frame;
[0158] The setting module 1302 is used to set the virtual frame decoding flag of the current frame;
[0159] The inserting module 1303 is used to insert the virtual frame decoding identifier into the code stream of the current frame;
[0160] The sending module 1304 is used to send the code stream of the current frame to the decoding end.
[0161] In order to better understand how to generate each virtual block according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, and splice each virtual block to form a virtual frame, a specific example is used here to illustrate, for example, also refer to Figure 2 , 3 As shown, taking the center point of the current CTU as the base point, in the two reference frames of the previous frame and the next frame of the current frame, a reference block with a length and width twice that of the current CTU is selected, and a virtual block is generated through algorithms such as deep learning, and based on this, each virtual block is generated; and then also refer to Figure 4 As shown, the generated virtual blocks are spliced to form a virtual frame. It is worth noting that in practical applications, the number of reference frames and the length and width of the reference blocks can be flexibly adjusted according to specific application scenarios.
[0162] Optionally, in this embodiment, when a virtual block of a CTU is generated by a deep learning algorithm, a deep learning method is first used to generate a deep learning network through pre-training with a large number of sample sequences, so as to generate a virtual block of the CTU at the encoding end; it should be clear that the encoding end will send the generated deep learning network to the decoding end, so that the decoding end can use the deep learning network to generate the same virtual block of the CTU. It is worth noting that here only the deep learning algorithm is used to generate a virtual block of a CTU as an example. In actual applications, it is necessary to follow the specific algorithm used. In fact, as long as it can be guaranteed that the encoding end and the decoding end can generate the same virtual block of the CTU, it is within the protection scope of the present invention, and the present invention does not make specific limitations on this.
[0163] In this embodiment, the setting module 1302 is used to determine the objective quality of the virtual frame; and to set different values of the virtual frame decoding flag according to the objective quality.
[0164] It can be understood that the virtual frame formed by splicing the virtual blocks is compared with the original image of the current frame to determine the objective quality of the virtual frame; optionally, the Peak Signal to Noise Ratio (PSNR) of the virtual frame is used as the objective quality of the virtual frame, where the PSNR of the virtual frame can be calculated by formulas 1 and 2 shown below.
[0165] Formula 1:
[0166] Among them, MSE is the mean square error between the current virtual frame and the current original image, W is the image width, H is the image height, I(i,j) is the pixel value of the current virtual frame, and K(i,j) is the pixel value of the original image of the current frame.
[0167] Formula 2: Where n is the bit depth of the image pixels.
[0168] Optionally, the objective quality of the virtual frame is determined by the structural similarity index (SSIM) of the virtual frame. It should be noted that the two common methods for determining the objective quality of the virtual frame are listed here, and the present invention is not limited to these two methods. In practical applications, they can be flexibly adjusted according to specific application scenarios.
[0169] It can also be understood that in some examples of this embodiment, the virtual blocks can be spliced together to form a virtual frame and then compared with the original image of the current frame to determine the distortion value of the virtual frame, where the distortion value of the virtual frame can be calculated by Formula 1 in the above figure; it should be clear that the distortion value of the virtual frame can also be obtained by calculating MSE, SAD, SSD, SSE, and SATD, and the present invention is not limited to this. In actual applications, it can be flexibly adjusted according to specific application scenarios.
[0170] In this embodiment, the setting module 1302 sets different values of the virtual frame decoding flag according to the objective quality, including the following two cases:
[0171] In case 1, when the objective quality of the virtual frame is less than or equal to a preset threshold, the value of the virtual frame decoding flag is set to instruct the decoding end to reconstruct the current frame using the virtual frame of the decoding end.
[0172] It is understandable that when the objective quality is less than or equal to the preset threshold, the virtual frame decoding flag is optionally set to 1 to instruct the decoding end to reconstruct the current frame using the virtual frame of the decoding end, that is, skip encoding the current frame and reduce the encoding bit rate.
[0173] In the second case, when the objective quality of the virtual frame is greater than a preset threshold, the value of the virtual frame decoding flag is set to the flag to instruct the decoding end to decode and reconstruct the current frame.
[0174] It is understandable that when the objective quality is greater than a preset threshold, the virtual frame decoding flag is optionally set to a value of 0 to instruct the decoding end to decode and reconstruct the current frame, that is, the code stream of the current frame still needs to be encoded.
[0175] Optionally, the preset threshold in this embodiment is calculated based on the PSNR of the forward N-frame reference frame of the current frame and / or the PSNR of the backward M-frame reference frame of the current frame, where N and M are integers, and N and M are greater than or equal to 1. Optionally, the preset threshold can be calculated based on the PSNR of the forward N-frame reference frame of the current frame only, or the preset threshold can be calculated based on the PSNR of the backward N-frame reference frame of the current frame only, or the preset threshold can be calculated based on the PSNR of the forward N-frame reference frame and the PSNR of the backward M-frame reference frame of the current frame.
[0176] Optionally, the preset threshold in this embodiment is calculated based on the SSIM of the forward N-frame reference frame of the current frame and / or the SSIM of the backward M-frame reference frame of the current frame, where N and M are integers, and N and M are greater than or equal to 1. Optionally, the preset threshold can be calculated based on the SSIM of the forward N-frame reference frame of the current frame only, or the preset threshold can be calculated based on the SSIM of the backward N-frame reference frame of the current frame only, or the preset threshold can be calculated based on the SSIM of the forward N-frame reference frame and the SSIM of the backward M-frame reference frame of the current frame.
[0177] It can be understood that in this embodiment, when the preset threshold is obtained by preset calculation based on PSNR, the preset calculation can be to calculate the average value of the PSNR of the forward N frame reference frame and / or the PSNR of the backward M frame reference frame of the current frame; the preset calculation can also be to calculate the weighted average value of the PSNR of the forward N frame reference frame and / or the PSNR of the backward M frame reference frame of the current frame.
[0178] It can also be understood that, in this embodiment, when the preset threshold is obtained by preset calculation based on SSIM, the preset calculation can be to calculate the average value of the SSIM of the forward N frame reference frame and / or the SSIM of the backward M frame reference frame of the current frame; the preset calculation can also be to calculate the weighted average value of the SSIM of the forward N frame reference frame and / or the SSIM of the backward M frame reference frame of the current frame.
[0179] It is worth noting that what are listed here are only two common preset calculations. The present invention is not limited to these two preset calculations. In practical applications, they can be flexibly adjusted according to specific application scenarios.
[0180] It should be noted that an image frame is composed of one or more slices, wherein a slice is composed of one or more CTUs.
[0181] In this embodiment, the inserting module 1303 is used to insert the virtual frame decoding identifier into the bitstream of the current frame in case 1 (i.e., when the objective quality is less than or equal to the preset threshold), including: inserting the virtual frame decoding identifier into the frame header bitstream of the current frame. In this embodiment, the frame header can be an image slice header, that is, the virtual frame decoding identifier is inserted into the bitstream of the slice header of the current frame.
[0182] In this embodiment, the insertion module 1303 is used to insert a virtual frame decoding identifier into the bitstream of the current frame in situation two (i.e., when the objective quality is greater than a preset threshold), including: inserting the virtual frame decoding identifier into the frame header bitstream of the current frame, and inserting the corresponding virtual block decoding identifier into the bitstream of the CTU of the current frame.
[0183] It should be noted that, in the present embodiment, when the residual frame obtained by using the virtual frame as the predicted image of the current frame is not zero, it is necessary to perform transformation, quantization, entropy coding and other coding steps on the residual frame to generate a bit stream, and the subsequent frame-level coding is completed; when the residual frame obtained by using the virtual frame as the predicted image of the current frame is zero, there is no need to perform transformation, quantization, entropy coding and other coding steps on the residual frame, and the frame-level coding is completed.
[0184] In this embodiment, the inserting module 1303 is used to compare the rate-distortion RD cost of the CTU with the rate-distortion RD cost of the virtual block corresponding to the CTU; and set different values of the virtual block decoding flag according to the comparison result.
[0185] Optionally, the insertion module 1303 calculates the optimal coding mode of the CTU of the current frame and the rate-distortion RD cost of each sub-block in the CTU of the current frame under the coding mode, adds the RD cost of each sub-block of the CTU to obtain the RD cost of the CTU; calculates the RD cost of the virtual block corresponding to the CTU; and then compares the RD cost of the CTU with the RD cost of the virtual block corresponding to the CTU.
[0186] In this embodiment, the insertion module 1303 sets different values of the virtual block decoding flag according to the comparison result, including the following two cases:
[0187] In case 1, when the RD cost of the virtual block is less than or equal to the RD cost of the CTU, the value of the virtual block decoding flag is set to instruct the decoding end to use the virtual block for decoding and reconstruction.
[0188] It can be understood that when the RD cost of the virtual block is less than or equal to the RD cost of the CTU, the virtual block decoding flag is optionally set to a value of 1 to instruct the decoding end to reconstruct the CTU using the virtual block of the decoding end, that is, to generate a bitstream using the virtual block as the prediction block of the CTU, that is, there is no need to encode the CTU, so as to reduce the encoding bit rate by avoiding encoding of motion information and residual.
[0189] In case 2, when the RD cost of the virtual block is greater than the RD cost of the CTU, the value of the virtual block decoding flag is set to instruct the decoding end to use the bit stream of the CTU of the current frame for decoding and reconstruction.
[0190] It is understandable that when the RD cost of the virtual block is greater than the RD cost of the CTU, the virtual block decoding flag is optionally set to a value of 0 to instruct the decoder to use the bitstream of the CTU of the current frame for decoding and reconstruction, that is, the CTU still needs to be encoded to generate the bitstream.
[0191] In this embodiment, in case 1 (i.e., when the RD cost of the virtual block is less than or equal to the RD cost of the CTU), after inserting the corresponding virtual block decoding identifier, it includes: encoding the residual block obtained by the CTU and the virtual block to generate a code stream of the CTU of the current frame.
[0192] In this embodiment, in case 2 (ie, when the RD cost of the virtual block is greater than the RD cost of the CTU), after inserting the corresponding virtual block decoding identifier, the following steps include: encoding the CTU to generate a code stream of the CTU of the current frame.
[0193] It should be noted that, in this embodiment, when the residual block obtained by using the virtual block as the CTU prediction block of the current frame is not zero, the residual block needs to be transformed, quantized, entropy encoded and other encoding steps to generate a bit stream, and the subsequent CTU-level encoding is completed; when the residual block obtained by using the virtual block as the CTU prediction block of the current frame is zero, there is no need to transform the residual block, quantize, entropy encode and other encoding steps, and the CTU-level encoding is completed.
[0194] It can be understood that when the virtual frame decoding identifier is inserted into the frame header code stream of the current frame, the frame header code stream of the current frame needs to be sent to the decoding end; when the virtual block decoding identifier is inserted into the code stream of the CTU of the current frame, the code stream of the CTU of the current frame needs to be sent to the decoding end.
[0195] In order to solve the problem that the motion vector and the prediction residual need to be encoded simultaneously in the inter-frame prediction mode in the prior art, and the encoding performance is low, a decoding device is also provided in an embodiment of the present invention. Fig.14 As shown:
[0196] The decoding device 14 includes a second generating module 1401, a determining module 1402 and a decoding module 1403, wherein:
[0197] The second generation module 1401 is used to generate each virtual block according to the pixel values of each coding tree unit CTU at the same position in the reference frame and the pixels at the adjacent position, and splice the virtual blocks to form a virtual frame;
[0198] The determination module 1402 is used to determine the decoding method of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding end;
[0199] The decoding module 1403 is used to decode and reconstruct the current frame according to the determined decoding method.
[0200] It should be clear that the second generation module 1401 is the same as the first generation module 1301, and the only difference is that one is executed at the decoding end and the other is executed at the encoding end. In other words, both the decoding end and the encoding end will generate each virtual block according to the pixel values of each coding tree unit CTU at the same position in the reference frame and the pixel values of its adjacent position pixels, and splice the virtual blocks to form a virtual frame, which will not be repeated here.
[0201] In this embodiment, when the value of the virtual frame decoding identifier parsed from the code stream by the determination module 1402 indicates that the decoding end uses the virtual frame of the decoding end to reconstruct the current frame, the decoding method of the current frame is determined to be to reconstruct the current frame using the virtual frame. Optionally, when the value of the parsed virtual frame decoding identifier is 1, the decoding method of the current frame is determined to be to reconstruct the current frame using the virtual frame, and compared with the existing method, the decoding of the current frame is skipped, which improves the encoding performance to a certain extent.
[0202] It should be noted that, in this embodiment, there are two situations in which the current frame is reconstructed using the virtual frame:
[0203] First, when the residual frame parsed from the bitstream of the current frame is not zero, it is necessary to parse the residual coefficients and perform inverse transformation, inverse quantization, entropy decoding and other decoding steps, and then combine the residual frame and the virtual frame to obtain a reconstructed frame.
[0204] Second, when the residual frame parsed from the code stream of the current frame is zero, the virtual frame is used as the reconstructed image of the current frame to obtain a reconstructed frame.
[0205] In this embodiment, when the value of the virtual frame decoding identifier parsed from the bitstream by the determination module 1402 is an identifier indicating that the decoding end decodes and reconstructs the current frame, the decoding method of the current frame is determined to be decoding and reconstructing the bitstream of the current frame. Optionally, when the value of the parsed virtual frame decoding identifier is 0, the existing method is still used, and the bitstream of the current frame needs to be decoded and reconstructed.
[0206] In this embodiment, when the determination module 1402 determines that the decoding mode of the current frame is to decode and reconstruct the code stream of the current frame, it determines the decoding mode of the CTU according to the virtual block decoding identifier received in the CTU code stream from the encoding end; and decodes and reconstructs the CTU according to the determined decoding mode.
[0207] In this embodiment, when the value of the virtual block decoding identifier parsed from the bitstream by the determination module 1402 indicates that the decoding end uses the virtual block of the decoding end to reconstruct the CTU, the decoding method of the CTU is determined to be to reconstruct the CTU using the virtual block. Optionally, when the value of the parsed virtual block decoding identifier is 1, the virtual block generated by the decoding end is directly used as the prediction block of the CTU, and reconstructed in combination with the residual block parsed from the bitstream, skipping the decoding of the CTU compared to the prior art, which improves the encoding performance to a certain extent.
[0208] It should be noted that, in this embodiment, there are two situations in which the decoding method of determining the CTU is to use the virtual block as the prediction block of the CTU and to reconstruct the residual block obtained by parsing the bitstream:
[0209] First, when the residual block parsed from the bitstream of the CTU is not zero, it is necessary to parse the residual coefficients and perform inverse transformation, inverse quantization, entropy decoding and other decoding steps, and then combine the residual block and the virtual block to obtain the reconstructed CTU.
[0210] Second, when the residual block parsed from the bitstream of the CTU is zero, the virtual block is used as the prediction block of the current frame to obtain the reconstructed CTU.
[0211] In this embodiment, when the value of the virtual block decoding identifier parsed from the bitstream by the determination module 1402 indicates that the decoding end uses the bitstream of the CTU of the current frame for decoding and reconstruction, the decoding method of the CTU is determined to be decoding and reconstruction of the bitstream of the CTU. Optionally, when the value of the parsed virtual block decoding identifier is 0, the existing method is still used, such as parsing the motion vector and the residual block from the CTU bitstream, and then using the motion vector to perform motion compensation to obtain the prediction block, and decoding and reconstruction are performed in combination with the residual block.
[0212] It can be understood that when it is determined that the decoding method of CTU is to reconstruct CTU using virtual blocks, the virtual blocks generated by the decoding end need to be used to reconstruct CTU; when it is determined that the decoding method of CTU is to decode and reconstruct the bitstream of CTU, the bitstream of CTU needs to be decoded and reconstructed.
[0213] It can also be understood that when it is determined that the decoding method of the current frame is to reconstruct the current frame using a virtual frame, the virtual frame generated by the decoding end is required to reconstruct the current frame; when it is determined that the decoding method of the current frame is to decode and reconstruct the code stream of the current frame, the code stream of the current frame needs to be decoded and reconstructed.
[0214] The encoding and decoding device provided by the embodiment of the present invention generates each virtual block according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame through the first generation module, splices each virtual block to form a virtual frame, sets the virtual frame decoding identifier of the current frame by the setting module, inserts the virtual frame decoding identifier into the code stream of the current frame by the insertion module, and further, the sending module sends the code stream of the current frame to the decoding end; the second generation module generates each virtual block according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, splices each virtual block to form a virtual frame, and then the determination module determines the decoding method of the current frame according to the virtual frame decoding identifier received from the current frame code stream of the encoding end, and further, decodes and reconstructs the current frame according to the determined decoding method; solves the problem of the need to simultaneously encode the motion vector and the prediction residual in the inter-frame prediction mode in the prior art, and the low encoding performance. Therefore, compared with the prior art, the encoding and decoding device provided by the embodiment of the present invention avoids the encoding and transmission of the motion vector and the image residual in the traditional encoding mode to a certain extent, reduces the encoding bit rate, and improves the encoding performance.
[0215] Embodiment six:
[0216] Based on the fifth embodiment, the present invention provides a system. Fig.15 As shown, the system 15 includes an encoding device 1501 and a decoding device 1502, wherein:
[0217] The encoding device 1501 is used to generate each virtual block according to the pixel values of each coding tree unit CTU and the pixels at the same position in the reference frame and the adjacent position pixels, and splice the virtual blocks to form a virtual frame; set the virtual frame decoding identifier of the current frame; insert the virtual frame decoding identifier into the code stream of the current frame; and send the code stream of the current frame to the decoding device.
[0218] The decoding device 1502 is used to generate each virtual block according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, and splice the virtual blocks to form a virtual frame; determine the decoding method of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding device; and decode and reconstruct the current frame according to the determined decoding method.
[0219] It is worth noting that in order to avoid redundant description, all examples in the fifth embodiment are not fully described in this embodiment. It should be clear that all examples in the fifth embodiment are applicable to this embodiment.
[0220] The system provided by the embodiment of the present invention includes an encoding device and a decoding device, which solves the problem of low encoding performance in the prior art that the motion vector and the prediction residual need to be encoded simultaneously in the inter-frame prediction mode. Therefore, compared with the prior art, the system provided by the embodiment of the present invention avoids the encoding and transmission of the motion vector and the image residual in the traditional encoding mode to a certain extent, reduces the encoding bit rate, and improves the encoding performance.
[0221] An embodiment of the present invention also provides a storage medium, which stores one or more first programs, and the one or more first programs can be executed by one or more first processors to implement the steps of the encoding method in the above-mentioned embodiment 1, or the storage medium stores one or more second programs, and the one or more second programs can be executed by one or more second processors to implement the steps of the decoding method in the above-mentioned embodiment 1.
[0222] The storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules or other data). Storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable read only memory), flash memory or other memory technology, CD-ROM (Compact Disc Read-Only Memory), digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0223] Obviously, those skilled in the art should understand that all or some steps, systems, and functional modules / units in the above disclosed methods can be implemented as software (which can be implemented with program code executable by a computing device), firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be performed by several physical components in cooperation. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, executed by a computing device, and in some cases, the steps shown or described can be performed in a different order than herein, and the computer-readable medium can include a computer storage medium (or a non-transitory medium) and a communication medium (or a transient medium). As is well known to those of ordinary skill in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, or other data. In addition, it is well known to those of ordinary skill in the art that communication media typically contains computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media. Therefore, the present invention is not limited to any specific combination of hardware and software.
[0224] The above contents are further detailed descriptions of the embodiments of the present invention in combination with specific implementation methods, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A coding method, It is characterized in that Applied to the encoding end, the encoding method includes: Generate virtual blocks according to pixel values of pixels of coding tree units CTU and adjacent positions at the same position in the reference frame, and splice the virtual blocks to form a virtual frame; Setting a virtual frame decoding flag of the current frame; Inserting the virtual frame decoding identifier into the code stream of the current frame; Sending the code stream of the current frame to a decoding end; The step of setting a virtual frame decoding flag of the current frame includes: Determining an objective quality of the virtual frame compared to an original image of the current frame; When the objective quality is greater than or equal to a preset threshold, setting the value of the virtual frame decoding flag to instruct the decoding end to reconstruct the current frame using the virtual frame of the decoding end; When the objective quality is greater than or equal to a preset threshold, the current frame is no longer encoded.
2. The encoding method according to claim 1, It is characterized in that After determining the objective quality of the virtual frame compared with the original image of the current frame, the method further includes: When the objective quality is less than a preset threshold, the value of the virtual frame decoding flag is set to a flag to instruct the decoding end to decode and reconstruct the current frame.
3. The encoding method according to claim 2, It is characterized in that The preset objective quality threshold is calculated based on a peak signal-to-noise ratio (PSNR) of a forward N-frame reference frame of the current frame and / or a peak signal-to-noise ratio (PSNR) of a backward M-frame reference frame of the current frame, where N and M are integers, and N and M are greater than or equal to 1; or, The preset threshold is calculated based on the structural similarity SSIM of the forward N frame reference frames of the current frame and / or the structural similarity SSIM of the backward M frame reference frames of the current frame, where N and M are integers and are greater than or equal to 1.
4. The encoding method according to claim 1, It is characterized in that When the objective quality is greater than or equal to a preset threshold, inserting the virtual frame decoding identifier into the bitstream of the current frame, and sending the bitstream of the current frame to the decoding end, comprises: Inserting the virtual frame decoding identifier into the frame header code stream of the current frame; Sending the frame header code stream of the current frame to the decoding end.
5. The encoding method according to claim 2, It is characterized in that When the objective quality is less than a preset threshold, inserting the virtual frame decoding identifier into the bitstream of the current frame, and sending the bitstream of the current frame to the decoding end, comprises: Inserting the virtual frame decoding identifier into the frame header code stream of the current frame; Sending a frame header stream of the current frame to the decoding end; Inserting a corresponding virtual block decoding identifier into a bitstream of a CTU of the current frame; Send the CTU code stream of the current frame to the decoding end.
6. The encoding method according to claim 5, It is characterized in that The inserting a corresponding virtual block decoding identifier into the bit stream of the CTU of the current frame includes: Comparing the rate-distortion RD cost of the CTU with the rate-distortion RD cost of a virtual block corresponding to the CTU; According to the comparison result, different values of the virtual block decoding identifier are set.
7. The encoding method according to claim 6, It is characterized in that The step of setting different values of the virtual block decoding identifier according to the comparison result includes: When the RD cost of the virtual block is less than or equal to the RD cost of the CTU, setting the value of the virtual block decoding flag to instruct the decoding end to reconstruct the CTU using the virtual block of the decoding end; When the RD cost of the virtual block is greater than the RD cost of the CTU, the value of the virtual block decoding flag is set to instruct the decoding end to use the code stream of the CTU of the current frame for decoding and reconstruction.
8. The encoding method according to claim 7, It is characterized in that When the RD cost of the virtual block is less than or equal to the RD cost of the CTU, after inserting the corresponding virtual block decoding identifier, sending the CTU code stream of the current frame to the decoding end includes: Encoding a residual block obtained by the CTU and the virtual block to generate a bit stream of the CTU of the current frame; Send the CTU code stream of the current frame to the decoding end.
9. The encoding method according to claim 7, It is characterized in that When the RD cost of the virtual block is greater than the RD cost of the CTU, after inserting the corresponding virtual block decoding identifier, sending the CTU code stream of the current frame to the decoding end includes: Encoding the CTU to generate a code stream of the CTU of the current frame; Send the CTU code stream of the current frame to the decoding end.
10. A decoding method, It is characterized in that Applied to a decoding end, the decoding method includes: Generate virtual blocks according to pixel values of pixels of coding tree units CTU and adjacent positions at the same position in the reference frame, and splice the virtual blocks to form a virtual frame; Determining a decoding method for the current frame according to a virtual frame decoding identifier in a current frame code stream received from an encoding end; wherein, when the value of the virtual frame decoding identifier parsed from the code stream indicates that the decoding end uses a virtual frame of the decoding end to reconstruct the current frame, determining the decoding method for the current frame is to use the virtual frame to reconstruct the current frame; The current frame is decoded and reconstructed according to the determined decoding method; wherein, when the decoding method is to reconstruct the current frame using the virtual frame, the current frame is no longer decoded.
11. The decoding method according to claim 10, It is characterized in that The method of determining the decoding mode of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding end also includes: When the value of the virtual frame decoding identifier parsed from the code stream is an identifier indicating that the decoding end decodes and reconstructs the current frame, the decoding mode of the current frame is determined to be decoding and reconstructing the code stream of the current frame.
12. The decoding method according to claim 11, It is characterized in that When the decoding mode of the current frame is to decode and reconstruct the code stream of the current frame, the method includes: Determine the decoding mode of the CTU according to the virtual block decoding identifier received in the CTU code stream from the encoder; The CTU is decoded and reconstructed according to the determined decoding mode.
13. The decoding method according to claim 12, It is characterized in that The determining the decoding mode of the CTU according to the virtual block decoding identifier received in the CTU code stream from the encoding end includes: When the value of the virtual block decoding identifier parsed from the bitstream indicates that the decoding end uses the virtual block of the decoding end to reconstruct the CTU, determining that the decoding mode of the CTU is to reconstruct the CTU using the virtual block; When the value of the virtual block decoding identifier parsed from the bitstream indicates that the decoding end uses the bitstream of the CTU of the current frame for decoding and reconstruction, the decoding mode of the CTU is determined to be decoding and reconstruction of the bitstream of the CTU.
14. A transmission method, It is characterized in that The transmission method comprises: The encoding end generates each virtual block according to the pixel values of each coding tree unit CTU and the pixels at the adjacent position in the same position in the reference frame, and splices the virtual blocks to form a virtual frame; sets the virtual frame decoding identifier of the current frame, wherein the objective quality of the virtual frame compared with the original image of the current frame is determined, and when the objective quality is greater than or equal to a preset threshold, the value of the virtual frame decoding identifier is set to instruct the decoding end to reconstruct the current frame using the virtual frame of the decoding end; inserts the virtual frame decoding identifier into the code stream of the current frame; sends the code stream of the current frame to the decoding end; wherein, when the objective quality is greater than or equal to the preset threshold, the current frame is no longer encoded; The decoding end generates each virtual block according to the pixel values of each coding tree unit CTU and the pixels at the adjacent position in the same position in the reference frame, and splices the virtual blocks to form a virtual frame; determines the decoding method of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding end, wherein, when the value of the virtual frame decoding identifier parsed from the code stream indicates that the decoding end uses the virtual frame of the decoding end to reconstruct the current frame, determines that the decoding method of the current frame is to reconstruct the current frame using the virtual frame; decodes and reconstructs the current frame according to the determined decoding method, wherein, when the decoding method is to reconstruct the current frame using the virtual frame, the current frame is no longer decoded.
15. An encoding device, It is characterized in that The encoding device comprises: a first generating module, a setting module, an inserting module and a sending module; The first generation module is used to generate each virtual block according to the pixel values of each coding tree unit CTU at the same position in the reference frame and the pixels at the adjacent position, and splice the virtual blocks to form a virtual frame; The setting module is used to set the virtual frame decoding flag of the current frame; wherein, the objective quality of the virtual frame compared with the original image of the current frame is determined, and when the objective quality is greater than or equal to a preset threshold, the value of the virtual frame decoding flag is set to instruct the decoding end to reconstruct the current frame using the virtual frame of the decoding end; wherein, when the objective quality is greater than or equal to the preset threshold, the current frame is no longer encoded; The inserting module is used to insert the virtual frame decoding identifier into the code stream of the current frame; The sending module is used to send the code stream of the current frame to the decoding end.
16. The encoding device according to claim 15, It is characterized in that The setting module is further configured to set the value of the virtual frame decoding flag to be a flag to instruct the decoding end to decode and reconstruct the current frame when the objective quality is less than a preset threshold.
17. The encoding device according to claim 16, It is characterized in that When the objective quality is less than a preset threshold, the inserting module is used to insert the corresponding virtual frame decoding identifier into the bit stream of the CTU of the current frame; The sending module is used to send the CTU code stream of the current frame to the decoding end.
18. The encoding device according to claim 17, It is characterized in that The insertion module is used to compare the rate-distortion RD cost of the CTU with the rate-distortion RD cost of the virtual block corresponding to the CTU; According to the comparison result, different values of the virtual block decoding identifier are set.
19. The encoding device according to claim 18, It is characterized in that The inserting module is used for setting the value of the virtual block decoding flag to instruct the decoding end to reconstruct the CTU using the virtual block of the decoding end when the RD cost of the virtual block is less than or equal to the RD cost of the CTU; When the RD cost of the virtual block is greater than the RD cost of the CTU, the value of the virtual block decoding flag is set to instruct the decoding end to use the code stream of the CTU of the current frame for decoding and reconstruction.
20. A decoding device, It is characterized in that The decoding device comprises: a second generating module, a determining module and a decoding module; The second generation module is used to generate each virtual block according to the pixel values of each coding tree unit CTU at the same position in the reference frame and the pixels at the adjacent position, and splice the virtual blocks to form a virtual frame; The determination module is used to determine the decoding mode of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding end; wherein, when the value of the virtual frame decoding identifier parsed from the code stream indicates that the decoding end uses the virtual frame of the decoding end to reconstruct the current frame, the decoding mode of the current frame is determined to be to reconstruct the current frame using the virtual frame; The decoding module is used to decode and reconstruct the current frame according to the determined decoding method, wherein when the decoding method is to reconstruct the current frame using the virtual frame, the current frame is no longer decoded.
21. The decoding device according to claim 20, It is characterized in that The determination module is further configured to determine that the decoding method of the current frame is to decode and reconstruct the code stream of the current frame when the value of the virtual frame decoding identifier parsed from the code stream is an identifier indicating that the decoding end decodes and reconstructs the current frame.
22. The decoding device according to claim 21, It is characterized in that The determining module is used to determine the decoding mode of the CTU according to the virtual block decoding identifier received in the CTU code stream from the encoding end when determining that the decoding mode of the current frame is to decode and reconstruct the code stream of the current frame; The CTU is decoded and reconstructed according to the determined decoding mode.
23. The decoding device according to claim 22, It is characterized in that The determining module is configured to determine that a decoding method of the CTU is to reconstruct the CTU using the virtual block when a value of the virtual block decoding identifier parsed from the bitstream indicates that the decoding end reconstructs the CTU using the virtual block of the decoding end; When the value of the virtual block decoding identifier parsed from the bitstream indicates that the decoding end uses the bitstream of the CTU of the current frame for decoding and reconstruction, the decoding mode of the CTU is determined to be decoding and reconstruction of the bitstream of the CTU.
24. A system, It is characterized in that The system comprises: an encoding device and a decoding device; The encoding device is used to generate each virtual block according to the pixel values of each coding tree unit CTU and the pixels at the adjacent position in the same position in the reference frame, and splice the virtual blocks to form a virtual frame; set the virtual frame decoding identifier of the current frame, wherein the objective quality of the virtual frame compared with the original image of the current frame is determined, and when the objective quality is greater than or equal to a preset threshold, the value of the virtual frame decoding identifier is set to indicate that the decoding end uses the virtual frame of the decoding end to reconstruct the current frame; insert the virtual frame decoding identifier into the code stream of the current frame; send the code stream of the current frame to the decoding device; wherein, when the objective quality is greater than or equal to the preset threshold, the current frame is no longer encoded; The decoding device is used to generate each virtual block according to the pixel values of each coding tree unit CTU and its adjacent position pixels at the same position in the reference frame, and splice the virtual blocks to form a virtual frame; determine the decoding method of the current frame according to the virtual frame decoding identifier received in the current frame code stream from the encoding device, wherein, when the value of the virtual frame decoding identifier parsed from the code stream indicates that the decoding end uses the virtual frame of the decoding end to reconstruct the current frame, determine the decoding method of the current frame to use the virtual frame to reconstruct the current frame; decode and reconstruct the current frame according to the determined decoding method, wherein, when the decoding method is to use the virtual frame to reconstruct the current frame, the current frame is no longer decoded.
Citation Information
Patent Citations
Method for establishing virtual reference frame and equipment
CN106791829A