Processing method of symbolization block, video encoding / decoding method, apparatus, computer program, and computer device
By splitting encoding blocks with shared motion information in intra-block copy prediction, the method addresses inefficiencies in current video encoding standards, reducing bit overhead and enhancing encoding performance for screen content videos.
Patent Information
- Application Number
- JP2024506589
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-31
- Filing Date
- 2022-12-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-28
AI Technical Summary
Current video encoding standards like AV1, AVS3, and VVC restrict intra-block copy prediction modes, leading to additional bit overhead due to dividing large encoding blocks into small sub-blocks with the same motion information, which is inefficient for screen content videos with strong spatial correlation.
A method and apparatus for processing encoding blocks using intra-block copy prediction, where overlapping blocks are split into target blocks with shared motion information, optimizing the encoding and decoding process by determining positional relationships and image sampling formats to reduce bit overhead.
This approach effectively saves encoding and decoding overhead, improves performance, and enhances the prediction ability of screen content videos by allowing blocks with shared motion information to be encoded together, thus reducing bit overhead and improving encoding efficiency.
Smart Images

Figure 0007706638000001 
Figure 0007706638000002 
Figure 0007706638000003
Abstract
Description
Technical Field
[0001] This application claims priority based on a Chinese patent application filed with the China National Intellectual Property Administration on December 31, 2021, with an application number of 202111671659.6 and an invention title of "Processing Method and Apparatus, Video Decoding Method, Video Encoding Method, and Apparatus".
[0002] This application relates to the technical field of computers. Specifically, it relates to a processing method for encoding blocks, a video encoding and decoding method, an apparatus, Computer program , and computer equipment.
Background Art
[0003] In current video encoding standards such as AV1, AVS3, and VVC, the intra-block copy prediction mode is restricted so that the overlapping part between the reference block and the current block is not allowed. However, since screen content videos often have strong spatial correlation, large encoding blocks are often divided into a series of small sub-blocks to predict the current block using adjacent pixels. Some of this series of small sub-blocks have sub-blocks with the same motion information, and encoding this same motion information results in additional bit overhead.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Embodiments of this application can efficiently save the overhead of encoding and decoding and improve the performance of encoding and decoding, including a processing method for encoding blocks, a video encoding and decoding method, an apparatus, Computer program , and computer equipment.
Means for Solving the Problems
[0005] According to one aspect, in an embodiment of the present application, a method for processing an encoded block based on an intra-block copy prediction mode executed by a computer device is provided. This method includes receiving a bitstream, and based on the bitstream, determining a positional relationship between a current encoded block and a reference encoded block; and when the positional relationship overlaps, determining that the current encoded block is a merge block; Based on a predetermined basic division unit, obtaining a plurality of target encoded blocks by splitting the merge block, where the same motion information is included in the plurality of target encoded blocks; obtaining the image sampling format of the current encoded block; obtaining a plurality of target encoded blocks of different chrominances by dividing the chrominance component of the merge block based on the basic division unit; and determining the image sampling ratio of the current encoded block based on the image sampling format, wherein the step of obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit includes obtaining a plurality of target encoded blocks of different chrominances by dividing the chrominance component of the merge block based on the basic division unit, and obtaining a plurality of target encoded blocks of different luminances by dividing the luminance component of the merge block based on the basic division unit and the image sampling ratio; and the like.
[0006] According to another aspect, in an embodiment of the present application, a video decoding method executed by a computer device is provided. This method includes receiving a bitstream of an encoded video including a current image, and based on the bitstream, determining a positional relationship between a current encoded block and a reference encoded block; when the positional relationship overlaps, determining that the current encoded block is a merge block; obtaining a plurality of target encoded blocks by splitting the merge block, where the same motion information is included in the plurality of target encoded blocks; obtaining a predicted value of the current encoded block from the current image based on the block vector information and the intra-block copy prediction mode of the plurality of target encoded blocks; determining a reconstructed value of the current encoded block based on the predicted value, and obtaining a bitstream of the decoded video by decoding the reconstructed value; determining a sequence header flag based on the bitstream of the encoded video; and the like wherein the step of determining the positional relationship between the current encoded block and the reference encoded block based on the bitstream includes, when the flag value of the sequence header flag is a predetermined value, determining the positional relationship between the current encoded block and the reference encoded block, and the flag value of the sequence header flag is set during encoding on the encoding side; .
[0007] According to another aspect, in an embodiment of the present application, a video encoding method executed by a computer device is provided. This method includes a step of determining a positional relationship between a current encoding block and a reference encoding block, and a step of determining that the current encoding block is a merge block when the positional relationships overlap, Based on a predetermined basic division unit, a step of obtaining a plurality of target encoding blocks by splitting the merge block, wherein the same motion information is included in the plurality of target encoding blocks, obtaining the image sampling format of the current encoded block; obtaining a plurality of chrominance target encoded blocks by dividing the chrominance component of the merge block based on the basic division unit; determining the image sampling ratio of the current encoded block based on the image sampling format, wherein the step of obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit includes: obtaining a plurality of chrominance target encoded blocks by dividing the chrominance component of the merge block based on the basic division unit; and obtaining a plurality of luminance target encoded blocks by dividing the luminance component of the merge block based on the basic division unit and the image sampling ratio. a step of obtaining a predicted value of the current encoding block from a current image based on block vector information and an intra-block copy prediction mode of the plurality of target encoding blocks, and determining a reconstructed value of the current encoding block based on the predicted value.
[0008] According to another aspect, in an embodiment of the present application, an apparatus for processing an encoding block based on an intra-block copy prediction mode is provided. This apparatus includes a receiving unit that receives a bitstream and determines a positional relationship between a current encoding block and a reference encoding block based on the bitstream, and a determining unit that determines that the current encoding block is a merge block when the positional relationships overlap, Based on a predetermined basic division unit a splitting unit that obtains a plurality of target encoding blocks by splitting the merge block, wherein the same motion information is included in the plurality of target encoding blocks te and a unit for obtaining the image sampling format of the current encoded block; a unit for obtaining a plurality of chrominance target encoded blocks by dividing the chrominance component of the merge block based on the basic division unit; a unit for determining the image sampling ratio of the current encoded block based on the image sampling format, wherein the division unit further includes: a unit for obtaining a plurality of chrominance target encoded blocks by dividing the chrominance component of the merge block based on the basic division unit; and obtaining a plurality of luminance target encoded blocks by dividing the luminance component of the merge block based on the basic division unit and the image sampling ratio. a unit.
[0009] According to another aspect, in an embodiment of the present application, a video decoding apparatus is provided. This apparatus includes a receiving unit that receives a bitstream of an encoded video including a current image, determines a positional relationship between a current encoding block and a reference encoding block based on the bitstream, and determines that the current encoding block is a merge block when the positional relationships overlap, Based on a predetermined basic division unitA processing unit that obtains a plurality of target encoded blocks by splitting the merge block, wherein the same motion information is included in the plurality of target encoded blocks, A unit for obtaining the image sampling format of the current encoding block; a unit for obtaining a plurality of chrominance target encoding blocks by dividing the chrominance component of the merge block based on the basic division unit; a unit for determining the image sampling ratio of the current encoding block based on the image sampling format; a unit for obtaining a plurality of chrominance target encoding blocks by dividing the chrominance component of the merge block based on the basic division unit; and a unit for obtaining a plurality of luminance target encoding blocks by dividing the luminance component of the merge block based on the basic division unit and the image sampling ratio, and including a processing unit, a prediction unit that obtains a predicted value of the current encoded block from the current image based on block vector information of the plurality of target encoded blocks and an intra-block copy prediction mode, a reconstruction unit that determines a reconstructed value of the current encoded block based on the predicted value, and a decoding unit that obtains a bitstream of a decoded video by decoding the reconstructed value.
[0010] According to another aspect, in an embodiment of the present application, a video encoding apparatus is provided. The apparatus determines a positional relationship between a current encoded block and a reference encoded block, and determines that the current encoded block is a merge block when the positional relationship overlaps, Based on a predetermined basic division unit a determination unit that obtains a plurality of target encoded blocks by splitting the merge block, wherein the same motion information is included in the plurality of target encoded blocks, A unit for obtaining the image sampling format of the current encoding block; a unit for obtaining a plurality of chrominance target encoding blocks by dividing the chrominance component of the merge block based on the basic division unit; a unit for determining the image sampling ratio of the current encoding block based on the image sampling format; a unit for obtaining a plurality of chrominance target encoding blocks by dividing the chrominance component of the merge block based on the basic division unit; and a unit for obtaining a plurality of luminance target encoding blocks by dividing the luminance component of the merge block based on the basic division unit and the image sampling ratio, and including a determination unit, a prediction unit that obtains a predicted value of the current encoded block from the current image based on block vector information of the plurality of target encoded blocks and an intra-block copy prediction mode, and a reconstruction unit that determines a reconstructed value of the current encoded block based on the predicted value.
[0011] According to another aspect, in an embodiment of the present application, a computer-readable storage medium storing a computer program is provided. The computer program is configured to be loaded by a processor and execute the steps of the method in any of the above embodiments.
[0012] According to another aspect, in an embodiment of the present application, a computer device including a processor and a memory is provided. A computer program is stored in the memory, and the processor executes the steps of the method in any of the above embodiments by calling the computer program stored in the memory.
[0013] According to another aspect, in an embodiment of the present application, a computer program product or a computer program including computer instructions is provided. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device is caused to execute the steps of the method in any of the above embodiments.
Brief Description of the Drawings
[0014] To more clearly explain the technical method of the embodiments of the present application, the drawings necessary for the description of the embodiments are briefly introduced below. Obviously, the drawings in the following description only show some embodiments of the present application, and those skilled in the art can also obtain other drawings from these drawings without creative labor.
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Mode for Carrying Out the Invention
[0016] Hereinafter, with reference to the drawings of the embodiments of the present application, the technical methods of the embodiments of the present application will be clearly and completely described. As is clear, the described embodiments are only a part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by those skilled in the art from the embodiments of the present application without creative labor belong to the protection scope of the present application.
[0017] First, some of the nouns or terms appearing in the description of the embodiments of the present application are interpreted as follows.
[0018] Intra prediction: The prediction signal is from the encoded and reconstructed regions within the same image.
[0019] Intra block copy (abbreviated as IBC) prediction mode: The IBC prediction mode is an intra-coding tool adopted in the screen content coding (SCC) extension of HEVC, which significantly improves the coding efficiency of screen content. The IBC prediction mode is a block-level coding mode. On the coding side, the block matching (BM) technology is used to find the optimal matching block for each coding unit (CU).
[0020] In AVS3, VVC, and AV1, the IBC prediction mode technology is similarly adopted to improve the performance of screen content coding. The IBC prediction mode utilizes the spatial correlation of screen content videos, uses the pixels of the encoded image in the current image to predict the pixels of the current block to be encoded, and can efficiently save the bits required for encoding the pixels.
[0021] Current mainstream video coding standards such as HEVC, VVC, AVS3, AV1, and AV2 all adopt a block-based hybrid coding framework. These split the original video data into a series of coding blocks and combine video coding methods such as prediction, transformation, and entropy coding to achieve the compression of video data. Here, motion compensation is a prediction method commonly used in video coding. Motion compensation derives the predicted value of the current coding block from the encoded region based on the redundancy characteristics of video content in the temporal or spatial domain.
[0022] Here, the following series of operations and processes may be performed on the input original video signal.
[0023] Block partition structure: Based on the size of one processing unit, the input image is divided into several non-overlapping processing units, and similar compression operations are performed on each processing unit. This processing unit is called a CTU or LCU. Further dividing the CTU into finer units, one or more basic coding units called CUs can be obtained. Each CU is the most basic element of the coding process. Various coding methods that can be adopted for each CU are described below.
[0024] Predictive Coding: Includes methods such as intra prediction, inter prediction, and IBC prediction. After performing prediction on the original video signal using the selected reconstructed video signal, a residual video signal is obtained. The encoding side needs to determine the most appropriate one of many possible predictive coding modes for the current CU and notify the decoding side.
[0025] Transform Coding & Quantization: Through transformation operations such as discrete Fourier transform (DFT) and discrete cosine transform (DCT), the residual video signal is transformed into the transform domain, and this transformation operation is called transform coefficients. The signal in the transform domain further undergoes an irreversible quantization operation, losing some information, and the quantized signal becomes advantageous for compressed representation. In some video coding standards, there may be multiple selectable transformation methods. Therefore, the encoding side also needs to select one of them for the currently encoded CU and notify the decoding side. The fineness of quantization is usually determined by the quantization parameter (QP). When the value of QP is large, it means that coefficients with a larger value range are quantized as the same output, usually resulting in a larger distortion and a lower bit rate. Conversely, when the value of QP is small, it means that coefficients with a smaller value range are quantized as the same output, usually resulting in a small distortion and corresponding to a high bit rate.
[0026] Entropy Coding or Statistical Coding: For the quantized transform domain signal, statistical coding is performed according to the occurrence frequency of each value, and finally a binary (0 or 1) compressed bit stream is output. Also, for other information generated by coding, such as the selected mode and block vector, etc., entropy coding is necessary to reduce the bit rate. Statistical coding is one of the reversible coding methods and can effectively reduce the bit rate required for the representation of the same signal. Common statistical coding methods include Variable Length Coding (VLC) and Content Adaptive Binary Arithmetic Coding (CABAC).
[0027] Loop Filtering: By performing inverse quantization, inverse transformation, and prediction compensation operations on the coded image, a reconstructed decoded image can be obtained. Since the reconstructed image is affected by quantization compared to the original image, some information is different from the original image and distortion occurs. By performing filtering operations such as deblocking filtering on the reconstructed image, the degree of distortion caused by quantization can be effectively reduced. Since these filtered reconstructed images are used for the prediction of future signals as references for subsequent coded images, the above-mentioned filtering operations are also called loop filtering, that is, filtering operations within the coding loop.
[0028] In current video coding standards such as AV1, AVS3, and VVC, the IBC prediction mode is restricted so that the reference block is not allowed to have a portion overlapping with the current block. However, since screen content videos often have strong spatial correlations, large coding blocks are often divided into a series of small sub-blocks in order to predict the current block using neighboring pixels. This satisfies the restriction that the reference block does not overlap with the current block. Some of the divided small blocks have the same motion information, and coding this same motion information results in additional bit overhead.
[0029] In the embodiments of the present application, there are provided a method and apparatus for processing a coding block based on the IBC prediction mode, a video decoding method and apparatus, a medium, and a device. Thereby, coding overhead can be effectively saved, coding performance can be improved, and the problem of causing additional bit overhead by predicting the current block using neighboring pixels can be alleviated to a certain extent. The embodiments of the present application are applicable to the coding technology of the IBC prediction mode and products of video codecs or video compression using the IBC prediction mode technology.
[0030] Specifically, the method of the embodiments of the present application may be executed by a computer device. Here, the computer device may be a device such as a terminal or a server.
[0031] To better understand the technical methods provided in the embodiments of the present application, hereinafter, applicable application scenarios of the technical methods provided in the embodiments of the present application will be briefly introduced. It should be noted that the application scenarios introduced below are not limiting and are only used to explain the embodiments of the present application. Taking the case where the method for processing a coding block is executed by a computer device as an example. Here, the computer device may be a device such as a terminal or a server.
[0032] The embodiments of the present application may be implemented by a combination with cloud technology or blockchain network technology. For example, in the method for processing encoded blocks disclosed in the embodiments of the present application, these data may be stored in a blockchain. For example, the bitstream, positional relationship, merge block, and target encoded block may all be stored in the blockchain.
[0033] To facilitate the realization of the storage and retrieval of the bitstream, positional relationship, merge block, and target encoded block, optionally, the method for processing the encoded block further includes the steps of: sending the bitstream, positional relationship, merge block, and target encoded block to a blockchain network, causing the nodes of the blockchain network to fill a new block with the bitstream, positional relationship, merge block, and target encoded block, and when reaching a consensus on the new block, adding the new block to the end of the blockchain. In the embodiments of the present application, by adding and storing the bitstream, positional relationship, merge block, and target encoded block in the blockchain, the backup of records can be realized. When it is necessary to obtain the target encoded block, the corresponding target encoded block can be directly and quickly obtained from the blockchain, improving the processing efficiency of the encoded block.
[0034] The following will be described in detail respectively. It should be noted that the order of description of the following embodiments does not limit the priority of the embodiments.
[0035] In each embodiment of the present application, a method for processing an encoded block based on the IBC prediction mode is provided. In the embodiments of the present application, Processing The case where the method is executed by a computer device will be described as an example.
[0036] It should be noted that the method for processing the encoded block based on the IBC prediction mode of the present application is applicable to both the encoding side and the decoding side. In the following embodiments, the decoding side will be described as an example.
[0037] Refer to FIG. 1. FIG. 1 is a schematic diagram of the flow of the processing method provided in the embodiment of the present application. This method includes the following steps.
[0038] In step 110, a bitstream is received, and based on the bitstream, the positional relationship between the current encoded block and the reference encoded block is determined.
[0039] Specifically, the IBC prediction mode is an encoding tool adopted in the Screen Content Coding (SCC for short) extension, which significantly improves the encoding efficiency of screen content. In AVS3, VVC, and AV1, the IBC technology is similarly adopted to improve the performance of screen content encoding. IBC Prediction mode utilizes the spatial correlation of screen content videos, uses the pixels of the encoded image in the current image to predict the pixels of the current encoding target block, and can efficiently save the bits required for encoding the pixels.
[0040] Refer to FIG. 2. FIG. 2 is a schematic diagram of the intra-block copy prediction mode. The gray area is the encoded area, and the reference block is located in the encoded area. The current encoded block is abbreviated as the current block and is located in the white area, that is, the unencoded area. Here, the displacement between the current encoded block and its reference block is called a Block Vector (BV for short).
[0041] As can be understood, a single video contains multiple pictures. Each picture is divided into a plurality of regions, and encoding or decoding is performed on each region. For example, one picture is divided into one or more tiles and / or slices. Each tile and / or slice is divided into one or more coding tree units (CTUs) of the same size. In other words, one picture can be divided into CTUs of the same size, and one or more CTUs can form one tile or one slice. Each CTU can be divided into one or more coding units (CUs). The information applied to each CU is encoded as the syntax of the CU, and the information commonly applied to the CUs included in one CTU is encoded as the syntax of the CTU.
[0042] The bitstream includes a bitstream of an image or a bitstream of a video. On the decoding side, the decoded bitstream is included in the bitstream, and on the encoding side, the encoded bitstream is included in the bitstream. In the following examples, the decoded bitstream will be used as an example for explanation.
[0043] When receiving the bitstream, the current encoded block can be determined, and based on the motion information included in the bitstream, the position information of the reference encoded block can be determined, and further, the positional relationship between the current encoded block and the reference encoded block can be determined.
[0044] Optionally, the step of receiving the bitstream and determining the positional relationship between the current encoded block and the reference encoded block based on the bitstream includes the steps of determining the block vector of the current encoded block based on the bitstream, obtaining the first position information of the reference encoded block based on the block vector, and determining the positional relationship based on the first position information of the reference encoded block and the current encoded block.
[0045] Specifically, when an image is input into the encoder, it is first divided into a series of CTUs. Next, for each CTU, different division structures may be adopted, and division may be performed, for example, in the manner of a quadtree or a binary tree (that is, equally divided into 4 blocks or divided into 2 blocks). For these encoded blocks, various prediction modes will be attempted. However, in each embodiment of the present application, the prediction mode used for the IBC encoded blocks is considered. For the IBC encoded blocks, the encoder searches the reconstructed area of the current image and obtains the optimal matching block as the reference block. This process is called "motion estimation". Here, the displacement between the reference block and the current block is called the "Block Vector (abbreviated as BV)".
[0046] Based on the bitstream, the block vector of the current encoded block is determined. Here, the block vector represents the relative displacement between the current encoded block and the optimal matching block within its reference area. Each of the divided blocks has corresponding motion information (such as a block vector) that needs to be transmitted to the decoding side. On the decoding side, based on the received bitstream and the current encoded block, the block vector of the current encoded block can be obtained.
[0047] Based on the block vector, the relative displacement between the current encoded block and the reference encoded block is obtained, and further, the first position information of the reference encoded block is determined. Here, the reference encoded block includes the optimal matching block within the image where the current encoded block is located.
[0048] When the position information of the current encoded block and the reference encoded block is determined, the positional relationship between the reference encoded block and the current encoded block can be obtained.
[0049] Optionally, a method for obtaining a positional relationship based on the first position information of a reference coding block and the current coding block includes: obtaining second position information of the current coding block; and determining that the positional relationship overlaps when the horizontal coordinate of the lower right corner in the first position information is greater than or equal to the horizontal coordinate of the upper left corner in the second position information, and the vertical coordinate of the lower right corner in the first position information is greater than or equal to the vertical coordinate of the upper left corner in the second position information.
[0050] For example, the second position information of the current coding block is represented as the coordinates (cur_x, cur_y) of the upper left corner. Here, the horizontal coordinate is cur_x, and the vertical coordinate is cur_y. (b_w, b_h) are the values of the width and height of the current coding block. The block vector of the current IBC coding block is (mv_x, mv_y), where mv_x is the x coordinate value of the block vector and mv_y is the y coordinate value of the block vector.
[0051] The calculation of the coordinates (ref_br_x, ref_br_y) of the lower right corner of the first position information of the reference coding block can be
[0052] [Equation 1] ref_br_x = cur_x + b_w - 1 + mv_x
[0053] [Equation 2] ref_br_y = cur_y + b_h - 1 + mv_y represented as.
[0054] The first position information of the reference coding block can be calculated according to Equation 1 and Equation 2.
[0055] When the horizontal coordinate of the lower right corner in the first position information is greater than or equal to the horizontal coordinate of the upper left corner in the second position information, and the vertical coordinate of the lower right corner in the first position information is greater than or equal to the vertical coordinate of the upper left corner in the second position information, that is, when ref_br_x ≥ cur_x and ref_br_y ≥ cur_y, it is determined that the positional relationship overlaps.
[0056] In step 120, if there is an overlap in the positional relationship, it is determined that the current encoded block is a merge block.
[0057] Here, the merge block represents an encoded block having the same block vector.
[0058] Note that when performing motion estimation on the current encoded block, it is permitted that the reference encoded block and the current encoded block overlap.
[0059] When the encoder has completed motion estimation during encoding, several sub-blocks having the same motion information may be searched for. For example, taking a region of size 32×32 as a unit, the motion information of each sub-block is examined to determine whether these sub-blocks have the same block vector and whether the reference encoded block of the sub-block at the lower right corner of the current 32×32 region is within the current 32×32 region. Further, these sub-blocks may be merged. For example, these sub-blocks are merged into a 32×32 merge block. The block vector of the merge block is equal to the block vector of the sub-block.
[0060] In step 130, by splitting the merge block, a plurality of target encoded blocks are obtained, and the same motion information is included in the plurality of target encoded blocks.
[0061] Referring to FIG. 3, optionally, the step of obtaining a plurality of target encoded blocks by splitting the merge block includes step 131 of obtaining a plurality of target encoded blocks by splitting the merge block based on a predetermined basic splitting unit.
[0062] Specifically, the basic division unit may be equal to the minimum width and minimum height of the division allowed by the codec, or may be an integer multiple of the minimum width and minimum height of the division allowed by the codec, or may be other predetermined basic widths and heights. Here, the predetermined basic division unit may be the size of the basic memory cell in the hardware design. For example, the predetermined basic division unit (part_w, part_h) may be values such as (4, 1), (8, 1), (1, 4), or (1, 8).
[0063] For example, the basic division unit is (part_w, part_h), where part_w is the width of the basic division unit and part_h is the height of the basic division unit, and the values of part_w and part_h may be N_w*min_w and N_h*min_h respectively, and both N_w and N_h are integers greater than or equal to 1. min_w and min_h are the minimum width and minimum height of the division allowed by the codec respectively.
[0064] Based on the above-mentioned predetermined basic division unit (part_w, part_h), the current block may be decomposed into target encoding blocks with a size of (part_w, part_h), and subsequent reconstruction may be performed for each target encoding block.
[0065] Optionally, before dividing the merge block based on the predetermined basic division unit, Step 130 is the steps of obtaining the image sampling format of the current encoding block and determining the image sampling ratio of the current encoding block based on the image sampling format are further included, and based on the predetermined basic division unit , ma - The step of obtaining a plurality of target encoding blocks by dividing the merge block includes the step of obtaining a plurality of target encoding blocks of different chrominances by dividing the chrominance component of the merge block based on the predetermined basic division unit, and the step of obtaining a plurality of target encoding blocks of different luminances by dividing the luminance component of the merge block based on the predetermined basic division unit and the image sampling ratio.
[0066] Specifically, when splitting the merge block, the splitting may be performed based on the current image sampling format. Obtain the image sampling format of the current coded block. The image sampling format may include, but is not limited to, the YUV420 format, the YUV444 format, etc. Based on the image sampling format, determine the image sampling ratio of the current coded block. For example, let the ratio of the width of the image sampling ratio between the luminance image and the chrominance image be ratio_w, and the ratio of the height be ratio_h. In the YUV420 format, the image sampling ratio is ratio_w = 2, ratio_h = 2, while in the YUV444 format, ratio_w = 1, ratio_h = 1.
[0067] Once the image sampling format and the image sampling ratio are determined, for the chrominance component, by splitting the chrominance component of the merge block based on a predetermined basic splitting unit, a plurality of target coded blocks of chrominance are obtained. However, for the luminance component, by splitting the luminance component of the merge block based on the predetermined basic splitting unit and the image sampling ratio, a plurality of target coded blocks of luminance may be obtained.
[0068] For example, let the ratio of the width of the image sampling ratio between the luminance image and the chrominance image be ratio_w, and the ratio of the height be ratio_h. The chrominance image is split with the basic splitting unit (part_w, part_h) as a parameter, while the luminance image is split with (part_w * ratio_w, part_h * ratio_h) as a parameter.
[0069] In the YUV420 format, ratio_w = 2 and ratio_h = 2. When the current encoding block contains chrominance components and luminance components, the luminance components of the current encoding block are decomposed into sub-blocks, i.e., a plurality of target encoding blocks of luminance, with (part_w * 2, part_y * 2) as the unit. Accordingly, by decomposing the chrominance components of the current encoding block with (part_w, part_y) as the unit, a plurality of target encoding blocks of chrominance are obtained.
[0070] In the YUV420 format, ratio_w = 2 and ratio_h = 2. When the current encoding block contains only luminance components, a plurality of target encoding blocks of luminance are obtained by decomposing the current encoding block with (part_w, part_h) as the unit.
[0071] In the YUV444 format, ratio_w = 1 and ratio_h = 1. The number of chrominance samples is the same as the number of luminance samples, and a normal sampling format is adopted for both chrominance samples and luminance samples. In this case, a plurality of target encoding blocks are obtained by decomposing the current encoding block with (part_w, part_h) as the unit.
[0072] In this way, based on a predetermined basic division unit and an image sampling ratio, pictures of different sampling formats can be divided with a division size corresponding to the sampling ratio, and furthermore, it is not necessary to adjust chrominance components and luminance components in subsequent encoding or decoding, improving the encoding efficiency.
[0073] Referring to FIG. 4, optionally, based on a predetermined basic division unit, the step of obtaining a plurality of target encoded blocks by dividing a merge block further includes: a step 132 of obtaining a predetermined basic division unit; a step 133 of comparing the basic division unit with the block vector of the current encoded block and determining a target division unit based on the comparison result; and a step 134 of obtaining a plurality of target encoded blocks by dividing the merge block based on the target division unit.
[0074] In one embodiment of the present application, the method of comparing the basic division unit with the block vector of the current encoded block and determining the target division unit based on the comparison result includes: when the y coordinate value of the block vector is less than or equal to the height value of the basic division unit, taking the width value of the current encoded block as the target division unit at the x coordinate of the current encoded block and taking the height value of the basic division unit as the target division unit at the y coordinate of the current encoded block; and when the y coordinate value of the block vector is greater than the height value of the basic division unit, taking the width value of the basic division unit as the target division unit at the x coordinate of the current encoded block and taking the height value of the basic division unit as the target division unit at the y coordinate of the current encoded block.
[0075] Specifically, in a plane coordinate system, the block vector may include an x coordinate and a y coordinate.
[0076] For example, the basic division unit is (part_w, part_h), where part_w is the width value of the basic division unit and part_h is the height value of the basic division unit. The block vector of the current encoded block is (mv_x, mv_y), where mv_x is the x coordinate value of the block vector and mv_y is the y coordinate value of the block vector.
[0077] When mv_y ≤ -part_h, at the x - coordinate of the current encoded block, division is performed based on the value of b_w, and at the y - coordinate of the current encoded block, division is performed based on the value of part_h. That is, with (b_w, part_h) as the target division unit, the current encoded block is divided. When mv_y > -part_h, with the basic division unit (part_w, part_h) as the target division unit, the current encoded block is divided.
[0078] Similarly, when the x - coordinate value of the block vector is less than or equal to the value of the width of the basic division unit, at the x - coordinate of the current encoded block, the value of the width of the basic division unit is used as the target division unit, and at the y - coordinate of the current encoded block, the value of the height of the current encoded block is used as the target division unit.
[0079] When the x - coordinate value of the block vector is greater than the value of the width of the basic division unit, at the x - coordinate of the current encoded block, the value of the width of the basic division unit is used as the target division unit, and at the y - coordinate of the current encoded block, the value of the height of the basic division unit is used as the target division unit.
[0080] In the above example, when mv_x ≤ -part_w, at the x - coordinate of the current encoded block, division is performed based on the value of part_w which is the value of the width of the basic division unit, and at the y - coordinate of the current encoded block, division is performed based on the value of b_h which is the value of the height of the current encoded block. That is, with (part_w, b_h) as the target division unit, the current encoded block is divided. When mv_x > -part_w, at the x - coordinate of the current encoded block, division is performed based on the value of part_w which is the width of the basic division unit, and at the y - coordinate of the current encoded block, division is performed based on the value of part_h which is the height of the basic division unit. That is, with the basic division unit (part_w, part_h) as the target division unit, the current encoded block is divided.
[0081] Note that the coordinate system of the position information in this application may include multiple directions, and this embodiment is also applicable to the case of dividing the coordinate system in two or more directions.
[0082] In other embodiments of the present application, the splitting includes at least splitting of the x coordinate and the y coordinate. The method of comparing the basic splitting unit with the block vector of the current encoded block and determining the target splitting unit based on the comparison result is as follows: when the y coordinate value of the block vector is less than or equal to the height value of the basic splitting unit, at the x coordinate of the current encoded block, using the width value of the current encoded block as the target splitting unit, and at the y coordinate of the current encoded block, using the height value of the basic splitting unit as the target splitting unit; when the y coordinate value of the block vector is greater than the height value of the basic splitting unit, at the x coordinate of the current encoded block, using the width value of the basic splitting unit as the target splitting unit, and at the y coordinate of the current encoded block, using the height value of the current encoded block as the target splitting unit.
[0083] Specifically, in a plane coordinate system, the block vector may include an x coordinate and a y coordinate.
[0084] For example, the basic splitting unit is (part_w, part_h), where part_w is the width of the basic splitting unit and part_h is the height of the basic splitting unit. The block vector of the current encoded block is (mv_x, mv_y), where mv_x is the x coordinate value of the block vector and mv_y is the y coordinate value of the block vector.
[0085] When mv_y ≤ -part_h, at the x - coordinate of the current encoded block, perform division based on the value of b_w, and at the y - coordinate of the current encoded block, perform division based on the value of part_h. That is, using (b_w, part_h) as the target division unit, divide the current encoded block. When mv_y > -part_h, at the x - coordinate of the current encoded block, perform division based on the value of part_w which is the width of the basic division unit, and at the y - coordinate of the current encoded block, perform division based on the value of b_h which is the height of the current encoded block. That is, using (part_w, b_h) as the target division unit, divide the current encoded block.
[0086] Similarly, when the x - coordinate value of the block vector is less than or equal to the value of the width of the basic division unit, at the x - coordinate of the current encoded block, use the value of the width of the basic division unit as the target division unit, and at the y - coordinate of the current encoded block, use the value of the height of the current encoded block as the target division unit.
[0087] When the x - coordinate value of the block vector is greater than the value of the width of the basic division unit, at the x - coordinate of the current encoded block, use the value of the width of the current encoded block as the target division unit, and at the y - coordinate of the current encoded block, use the value of the height of the basic division unit as the target division unit.
[0088] In the above example, when mv_x ≤ -part_w, at the x - coordinate of the current encoded block, perform division based on the value of part_w which is the width of the basic division unit, and at the y - coordinate of the current encoded block, perform division based on the value of b_h which is the height of the current encoded block. That is, using (part_w, b_h) as the target division unit, divide the current encoded block. When mv_x > -part_w, at the x - coordinate of the current encoded block, perform division based on the value of b_w which is the width of the current encoded block, and at the y - coordinate of the current encoded block, perform division based on the value of part_h which is the height of the basic division unit. That is, using (b_w, part_h) as the tar Minutes get division unit, divide the current encoded block.
[0089] Note that the coordinate system of the position information in the present application may include a plurality of directions, and this embodiment is also applicable to the case of dividing the coordinate system in two or more directions.
[0090] In this way, in the present application, a bitstream is received, and based on the bitstream, the positional relationship between the current encoded block and the reference encoded block is determined. When the positional relationship overlaps, it is determined that the current encoded block is a merge block, and by dividing the merge block, a plurality of target encoded blocks are obtained. In the present application, based on the positional relationship in which the current encoded block and the reference encoded block overlap, the current encoded block is determined as a merge block, the merge block is divided, and the same motion information exists in the divided target encoded blocks. At the time of encoding on the encoding side, it is not necessary to encode the motion information of each sub-block respectively, and it is only necessary to encode the motion information of the merge block. Thereby, bit overhead is effectively saved, which is advantageous for improving encoding performance. On the decoding side, in the present application, based on the positional relationship in which the current encoded block and the reference encoded block overlap and information such as the position, width, and height of the current encoded block, it is implicitly derived whether the current encoded block is a merge block, and the merge block can be decomposed by a predetermined division method and decomposed into a plurality of target encoded blocks.
[0091] Also, based on the motion information of the current encoded block, the division unit is determined. Since there are a plurality of division methods, on the premise that the division of the encoded block satisfies the minimum division unit, the current block can be predicted using adjacent pixels as much as possible, and for screen content video, which often has strong spatial correlation, the prediction ability of the IBC prediction mode can be improved to a certain extent.
[0092] All the above-described technical methods can form selectable embodiments of the present application in any combination, but the description of each is omitted here.
[0093] To facilitate a better implementation of the method for processing an encoded block according to an embodiment of the present application, the embodiment of the present application provides the processing of an encoded block based on an IBC prediction mode. Device Further provided is a reference. Referring to FIG. 5, FIG. 5 is a schematic diagram of the configuration of a processing apparatus provided in an embodiment of the present application. Here, this processing apparatus 500 includes a receiving unit 510 that receives a bitstream and determines the positional relationship between a current encoded block and a reference encoded block based on the bitstream, and a determining unit 520 that determines that the current encoded block is a merge block when the positional relationships overlap, and a splitting unit 530 that obtains a plurality of target encoded blocks by splitting the merge block, where the plurality of target encoded blocks include the same motion information.
[0094] Optionally, the receiving unit 510 may determine a block vector of the current encoded block based on the bitstream, obtain first position information of the reference encoded block based on the block vector, and determine the positional relationship based on the first position information and the current encoded block.
[0095] Optionally, the receiving unit 510 may obtain second position information of the current encoded block, and determine that the positional relationship overlaps when the horizontal coordinate of the lower right corner in the first position information is greater than or equal to the horizontal coordinate of the upper left corner in the second position information, and the vertical coordinate of the lower right corner in the first position information is greater than or equal to the vertical coordinate of the upper left corner in the second position information.
[0096] Optionally, the splitting unit 530 may obtain a plurality of target encoded blocks by splitting the merge block based on a predetermined basic splitting unit.
[0097] Optionally, the splitting unit 530 may further obtain the image sampling format of the current encoded block, determine the image sampling ratio of the current encoded block based on the image sampling format, obtain a plurality of target encoded blocks of chrominance by splitting the chrominance component of the merge block based on the basic splitting unit, and obtain a plurality of target encoded blocks of luminance by splitting the luminance component of the merge block based on the predetermined basic splitting unit and the image sampling ratio.
[0098] Optionally, the splitting unit 530 may obtain the basic splitting unit, compare the basic splitting unit with the block vector of the current encoded block, determine the target splitting unit based on the comparison result, and obtain a plurality of target encoded blocks by splitting the merge block based on the target splitting unit.
[0099] Optionally, when the y coordinate value of the block vector is less than or equal to the height value of the basic splitting unit, at the x coordinate of the current encoded block, the width value of the current encoded block may be used as the target splitting unit, and at the y coordinate of the current encoded block, the height value of the basic splitting unit may be used as the target splitting unit. When the y coordinate value of the block vector is greater than the height value of the basic splitting unit, at the x coordinate of the current encoded block, the width value of the basic splitting unit may be used as the target splitting unit, and at the y coordinate of the current encoded block, the height value of the basic splitting unit may be used as the target splitting unit.
[0100] Optionally, when the y - coordinate value of the block vector is less than or equal to the height value of the basic division unit, at the x - coordinate of the current encoded block, the width value of the current encoded block is used as the target division unit, and at the y - coordinate of the current encoded block, the height value of the basic division unit is used as the target division unit. When the y - coordinate value of the block vector is greater than the height value of the basic division unit, at the x - coordinate of the current encoded block, the width value of the basic division unit is used as the target division unit, and at the y - coordinate of the current encoded block, the height value of the current encoded block is used as the target division unit.
[0101] Note that the functions of the respective modules of the processing apparatus 500 in the embodiments of the present application may each refer to any specific implementation form of the above - mentioned method embodiments, and the description thereof will not be repeated here.
[0102] Each unit of the above - mentioned processing apparatus 500 may be realized in whole or in part by software, hardware, or a combination thereof. Each of the above - mentioned units may be embedded in a processor of a computer device in the form of hardware or may be independent. Also, each of the above - mentioned units may be stored in the memory of a computer device in the form of software, thereby facilitating the processor to call and execute the operations corresponding to the above - mentioned units.
[0103] The processing device 500 may be incorporated into, for example, a terminal or a server that has a memory and a processor with computing capabilities, or the processing device 500 may be that terminal or server itself. The terminal may be a device such as a smartphone, a tablet computer, a notebook computer, a smart TV, a smart speaker, a wearable smart device, a personal computer (PC), etc. The terminal may include a client. The client may be a video client, a browser client, an instant communication client, or the like. The server may be an independent physical server, or may be a server cluster or a distributed system composed of multiple physical servers, and may also be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and base cloud computing services such as big data and artificial intelligence platforms.
[0104] FIG. 6 is a schematic configuration diagram of a computer device 600 provided in an embodiment of the present application. As shown in FIG. 6, the computer device 600 may include a communication interface 601, a memory 602, a processor 603, and a communication bus 604. The communication interface 601, the memory 602, and the processor 603 realize mutual communication via the communication bus 604. The communication interface 601 is for the computer device 600 to perform data communication with external devices. The memory 602 can be used to store software programs and modules. The processor 603 executes the software programs and modules stored in the memory 602, for example, the corresponding software programs for the operations in the foregoing method embodiments.
[0105] Optionally, the processor 603 may execute steps of receiving a bitstream by calling software programs and modules stored in the memory 602, determining a positional relationship between a current encoded block and a reference encoded block based on the bitstream, determining that the current encoded block is a merge block if the positional relationships overlap, and obtaining a plurality of target encoded blocks by splitting the merge block, where the same motion information is included in the plurality of target encoded blocks.
[0106] Optionally, the computer device 600 may be incorporated into a terminal or server having a memory and a processor with computing capabilities, or the computer device 600 may be the terminal or server itself. The terminal may be a device such as a smartphone, tablet computer, notebook computer, smart TV, smart speaker, wearable smart device, personal computer, etc. The server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and base cloud computing services such as big data and artificial intelligence platforms.
[0107] Refer to FIG. 7. In the present application, a video decoding method is further provided. This method includes step 710 of receiving a bitstream of an encoded video including a current image, step 720 of determining a positional relationship between a current encoded block and a reference encoded block based on the bitstream, step 730 of determining that the current encoded block is a merge block when the positional relationship overlaps, step 740 of obtaining a plurality of target encoded blocks by splitting the merge block, where the same motion information is included in the plurality of target encoded blocks, step 750 of obtaining a predicted value of the current encoded block from the current image based on the block vector information of the plurality of target encoded blocks and the intra-block copy prediction mode, step 760 of determining a reconstructed value of the current encoded block based on the predicted value, and step 770 of obtaining a bitstream of the decoded video by decoding the reconstructed value.
[0108] Specifically, receive a bitstream of an encoded video including a current image, and for each frame image in the video bitstream, determine the positional relationship between the current encoded block and the reference encoded block based on the bitstream. When the positional relationship overlaps, determine that the current encoded block is a merge block, and obtain a plurality of target encoded blocks by splitting the merge block.
[0109] Optionally, the method of determining the positional relationship between the current encoded block and the reference encoded block based on the bitstream includes a step of determining a block vector of the current encoded block based on the bitstream, a step of obtaining first position information of the reference encoded block based on the block vector, and a step of determining the positional relationship based on the first position information of the reference encoded block and the current encoded block.
[0110] Optionally, the method for determining the positional relationship based on the first position information and the current encoded block includes: obtaining second position information of the current encoded block; and determining that the positional relationship overlaps when the horizontal coordinate of the lower right corner in the first position information is greater than or equal to the horizontal coordinate of the upper left corner in the second position information, and the vertical coordinate of the lower right corner in the first position information is greater than or equal to the vertical coordinate of the upper left corner in the second position information.
[0111] Optionally, the method for obtaining a plurality of target encoded blocks by dividing a merge block includes: obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit.
[0112] Optionally, before obtaining a plurality of target encoded blocks by dividing a merge block, the method further includes: obtaining the image sampling format of the current encoded block; and determining the image sampling ratio of the current encoded block based on the image sampling format. The step of obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit includes: obtaining a plurality of target encoded blocks of different chrominance by dividing the chrominance component of the merge block based on the predetermined basic division unit; and obtaining a plurality of target encoded blocks of different luminance by dividing the luminance component of the merge block based on the predetermined basic division unit and the image sampling ratio.
[0113] Optionally, the method for obtaining a plurality of target encoded blocks by dividing a merge block based on a predetermined basic division unit further includes: obtaining the basic division unit; comparing the basic division unit with the block vector of the current encoded block and determining a target division unit based on the comparison result; and obtaining a plurality of target encoded blocks by dividing the merge block based on the target division unit.
[0114] Optionally, compare the basic division unit with the block vector of the current encoded block , ratio Based on the comparison result, the step of determining the target division unit includes: when the y-coordinate value of the block vector is less than or equal to the height value of the basic division unit, at the x-coordinate of the current encoded block, setting the width value of the current encoded block as the target division unit, and at the y-coordinate of the current encoded block, setting the height value of the basic division unit as the target division unit; and when the y-coordinate value of the block vector is greater than the height value of the basic division unit, at the x-coordinate of the current encoded block, setting the width value of the basic division unit as the target division unit, and at the y-coordinate of the current encoded block, setting the height value of the current encoded block as the target division unit.
[0115] Optionally, compare the basic division unit with the block vector of the current encoded block , ratio Based on the comparison result, the step of determining the target division unit includes: when the y-coordinate value of the block vector is less than or equal to the height value of the basic division unit, at the x-coordinate of the current encoded block, setting the width value of the current encoded block as the target division unit, and at the y-coordinate of the current encoded block, setting the height value of the basic division unit as the target division unit; and when the y-coordinate value of the block vector is greater than the height value of the basic division unit, at the x-coordinate of the current encoded block, setting the width value of the basic division unit as the target division unit, and at the y-coordinate of the current encoded block, setting the height value of the current encoded block as the target division unit.
[0116] The specific embodiments are the same as the above-described method for processing encoded blocks, and further detailed descriptions are omitted here.
[0117] When the target encoding block is determined, in step 750, based on the block vectors Information of the plurality of target encoding blocks and the intra-block copy prediction mode, a predicted value of the current encoding block is obtained from the current image. The intra-block copy prediction mode may include intra-block copy prediction, intra-string copy prediction, etc. In a specific encoding or decoding implementation, these prediction methods can be used alone or in combination. For the encoding blocks using these prediction methods, usually, in the bitstream, one or more two-dimensional displacement vectors indicating the displacement of the current block (or the block at the same position as the current block) with respect to one or more reference blocks need to be encoded explicitly or implicitly. Information In steps 760 and 770, based on the predicted value, a reconstructed value of the current encoding block is determined, and by decoding the reconstructed value, a bitstream of the decoded video is obtained.
[0118] Refer to FIG. 8. FIG. 8 shows an exemplary diagram of the basic flow of a video decoder. On the decoding side, after obtaining the bitstream, the decoder determines each CU, and first performs entropy decoding to obtain various mode information and quantized transform coefficients. By performing inverse quantization and inverse transformation on each coefficient, a residual signal is obtained. On the other hand, based on the known encoding mode information, a predicted signal corresponding to the CU is obtained by a predetermined predictive encoding method. By adding the residual signal and the predicted signal, a reconstructed signal can be obtained. Finally, by decoding the reconstructed value of the image and performing a loop filtering operation, a final output signal is generated.
[0119] Refer to FIG. 8. FIG. 8 shows an exemplary diagram of the basic flow of a video decoder. On the decoding side, after obtaining the bitstream, the decoder determines each CU, and first performs entropy decoding to obtain various mode information and quantized transform coefficients. By performing inverse quantization and inverse transformation on each coefficient, a residual signal is obtained. On the other hand, based on the known encoding mode information, a predicted signal corresponding to the CU is obtained by a predetermined predictive encoding method. By adding the residual signal and the predicted signal, a reconstructed signal can be obtained. Finally, by decoding the reconstructed value of the image and performing a loop filtering operation, a final output signal is generated.
[0120] Meanwhile, at the same time, any of the above-described processing methods for the encoding blocks can be applied to the encoding side to perform encoding processing on the video encoding side.
[0121] Refer to FIG. 9. FIG 9is an exemplary diagram of the basic flow of a video encoder. On the encoding side, a target encoded block is obtained by processing an encoded block according to any of the methods for processing an encoded block based on the IBC prediction mode described above. On the other hand, transform encoding and quantization are performed on the target encoded block. On the other hand, for the target encoded block, intra-image prediction is performed by an intra-mode determiner determined by an encoding mode determiner, and motion compensation prediction is realized based on motion estimation, and finally prediction information is obtained. Based on the encoding mode, intra-prediction mode, and motion information, entropy encoding is performed on the data and prediction information after transform encoding and quantization to output an encoded bitstream.
[0122] Referring to FIG. 10, optionally, before the step of determining the positional relationship between the current encoded block and the reference encoded block based on the bitstream, The video decoding method described above is step 780 of determining a sequence header flag based on the bitstream of the encoded video, and step 790 of determining the positional relationship between the current encoded block and the reference encoded block when the flag value of the sequence header flag is a predetermined value, wherein the flag value of the sequence header flag is set during encoding on the encoding side.
[0123] Specifically, in the bitstream structure, the sequence layer includes a sequence header. On the encoding side, a flag value is set for the sequence header flag to determine whether to apply the method for processing the encoded block of the present application. On the decoding side, before processing the current encoded block, the flag value of the sequence header flag is extracted. The flag value of the sequence header flag is If it is a predetermined value, the current encoded block is processed using the method for processing the encoded block of the present application. Conversely, if the flag value of the sequence header flag is not a predetermined value, the method of the present application may not be used.
[0124] For example, when the flag value of the sequence header flag is 1, it indicates that the processing method of the present application is used, and when the value is 0, it indicates that the processing method of the present application is not used. In this way, the sequence header flag can be used as the switch value of the processing method of the present application.
[0125] Optionally, when it is determined that the current encoded block is a merge block, the video decoding method further includes the step of adjusting the residual coefficient corresponding to the current encoded block to 0.
[0126] Specifically, when the current encoded block is a merge block, encoding of the residual coefficient is skipped, the residual coefficient corresponding to the current encoded block is adjusted to 0, and when determining the reconstructed value, the predicted value is determined as the reconstructed value, and reconstruction may be performed for each target encoded block.
[0127] In this way, on the decoding side, there is no need to decode the residual coefficient, and the reconstructed value can be directly determined based on the predicted value, improving the decoding efficiency to a certain extent.
[0128] All of the above technical methods can form optional embodiments of the present application in any combination, but the description of each is omitted here.
[0129] In the present application, a bitstream of an encoded video including a current image is received, based on the bitstream, the positional relationship between a current encoded block and a reference encoded block is determined, and when the positional relationship overlaps, it is determined that the current encoded block is a merge block, and by splitting the merge block, a plurality of target encoded blocks including the same motion information are obtained, and based on the block vector information and the intra-block copy prediction mode of the plurality of target encoded blocks, a predicted value of the current encoded block is obtained from the current image, a reconstructed value of the current encoded block is determined based on the predicted value, and by decoding the reconstructed value, a bitstream of the decoded video is obtained. On the decoding side, in the present application, based on the positional relationship where the current encoded block and the reference encoded block overlap and information such as the position, width, and height of the current encoded block, it is implicitly derived whether the current encoded block is a merge block, and the merge block can be decomposed by a predetermined splitting method and decomposed into a plurality of target encoded blocks.
[0130] To facilitate better implementation of the video decoding method of the embodiments of the present application, the embodiments of the present application further provide a video decoding apparatus. Referring to FIG. 11. FIG. 11 is a schematic diagram of the configuration of the video decoding apparatus provided in the embodiments of the present application. Here, this video decoding apparatus 800 includes a receiving unit 810 that receives a bitstream of an encoded video including a current image, and based on the bitstream, determines the positional relationship between a current encoded block and a reference encoded block, and when the positional relationship overlaps, determines that the current encoded block is a merge block, and a processing unit 820 that obtains a plurality of target encoded blocks by splitting the merge block, where the same motion information is included in the plurality of target encoded blocks, the processing unit 820, and the block vector of the plurality of target encoded blocks InformationA prediction unit 830 that obtains a predicted value of a current encoded block from a current image based on an intra-block copy prediction mode; a reconstruction unit 840 that determines a reconstructed value of the current encoded block based on the predicted value; and a decoding unit 850 that obtains a bitstream of a decoded video by decoding the reconstructed value. It may include.
[0131] Note that the functions of each module of the video decoding apparatus in the embodiments of the present application may refer to any specific implementation form of the above method embodiments, and the description thereof will not be repeated here.
[0132] Each unit of the above video decoding apparatus may be implemented in whole or in part by software, hardware, and combinations thereof. Each of the above units may be embedded in a processor of a computer device in the form of hardware or may be independent. Also, each of the above units may be stored in the memory of a computer device in the form of software. This facilitates the processor to call and execute the operations corresponding to each of the above units.
[0133] The video decoding device may be incorporated into, for example, a terminal or a server that has a memory and a processor with computing capabilities. Alternatively, the video decoding device may be the terminal or the server itself. The terminal may be a device such as a smartphone, a tablet computer, a notebook computer, a smart TV, a smart speaker, a wearable smart device, or a personal computer (PC). The terminal may include a client. The client may be a video client, a browser client, or an instant communication client, etc. The server may be an independent physical server, or may be a server cluster or a distributed system composed of multiple physical servers. It may also be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms.
[0134] Refer to FIG. 12. In the present application, a video encoding method is further provided. This method includes step 1210 of determining the positional relationship between the current encoding block and the reference encoding block, step 1220 of determining that the current encoding block is a merge block when the positional relationship overlaps, step 1230 of obtaining a plurality of target encoding blocks by splitting the merge block, where the same motion information is included in the plurality of target encoding blocks, step 1240 of obtaining a predicted value of the current encoding block from the current image based on the block vector information and the intra-block copy prediction mode of the plurality of target encoding blocks, and step 1250 of determining a reconstructed value of the current encoding block based on the predicted value.
[0135] Specifically, on the encoding side, for the image of each frame, the positional relationship between the current encoding block and the reference encoding block is determined. When the positional relationship overlaps, it is determined that the current encoding block is a merge block, and by dividing the merge block, a plurality of target encoding blocks are obtained.
[0136] Optionally, the method for determining the positional relationship between the current encoding block and the reference encoding block includes: determining the block vector of the current encoding block based on the bitstream; obtaining the first position information of the reference encoding block based on the block vector; and determining the positional relationship based on the first position information of the reference encoding block and the current encoding block.
[0137] Optionally, the method for determining the positional relationship based on the first position information and the current encoding block includes: obtaining the second position information of the current encoding block; and determining that the positional relationship overlaps when the horizontal coordinate of the lower right corner in the first position information is greater than or equal to the horizontal coordinate of the upper left corner in the second position information, and the vertical coordinate of the lower right corner in the first position information is greater than or equal to the vertical coordinate of the upper left corner in the second position information.
[0138] Optionally, the method for obtaining a plurality of target encoding blocks by dividing the merge block includes: obtaining a plurality of target encoding blocks by dividing the merge block based on a predetermined basic division unit. Optionally, before obtaining a plurality of target encoding blocks by dividing the merge block, the method further includes: obtaining the image sampling format of the current encoding block; and determining the image sampling ratio of the current encoding block based on the image sampling format.
[0139] The step of obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit includes: a step of obtaining a plurality of target encoded blocks of chrominance by dividing the chrominance component of the merge block based on a predetermined basic division unit; and a step of obtaining a plurality of target encoded blocks of luminance by dividing the luminance component of the merge block based on the predetermined basic division unit and the image sampling ratio.
[0140] Optionally, the method of obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit further includes: a step of obtaining a basic division unit; a step of comparing the basic division unit with the block vector of the current encoded block and determining a target division unit based on the comparison result; and a step of obtaining a plurality of target encoded blocks by dividing the merge block based on the target division unit.
[0141] Optionally, comparing the basic division unit with the block vector of the current encoded block , ratio The step of determining a target division unit based on the comparison result includes: when the y coordinate value of the block vector is less than or equal to the height value of the basic division unit, setting the width value of the current encoded block at the x coordinate of the current encoded block as the target division unit, and setting the height value of the basic division unit at the y coordinate of the current encoded block as the target division unit; and when the y coordinate value of the block vector is greater than the height value of the basic division unit, setting the width value of the basic division unit at the x coordinate of the current encoded block as the target division unit, and setting the height value of the basic division unit at the y coordinate of the current encoded block as the target division unit.
[0142] Optionally, comparing the basic division unit with the block vector of the current encoded block , ratioBased on the comparison result, the step of determining the target division unit includes: when the y - coordinate value of the block vector is less than or equal to the height value of the basic division unit, taking the width value of the current encoded block at the x - coordinate of the current encoded block as the target division unit, and taking the height value of the basic division unit at the y - coordinate of the current encoded block as the target division unit; when the y - coordinate value of the block vector is greater than the height value of the basic division unit, taking the width value of the basic division unit at the x - coordinate of the current encoded block as the target division unit, and taking the height value of the current encoded block at the y - coordinate of the current encoded block as the target division unit.
[0143] The specific embodiments are the same as the above - described method for processing the encoded block, and further detailed descriptions are omitted here.
[0144] Optionally, on the encoding side, by setting a flag value of a predetermined value for the sequence header flag, on the decoding side, it is determined whether to apply the decoding method of the present application based on the predetermined value. The specific embodiments are the same as the above - described video decoding method, and further detailed descriptions are omitted here.
[0145] All the above - mentioned technical methods can form optional embodiments of the present application in any combination, and further detailed descriptions are omitted here.
[0146] In the embodiments of the present application, the positional relationship between the current encoded block and the reference encoded block is determined. When the positional relationship overlaps, it is determined that the current encoded block is a merge block. By dividing the merge block, a plurality of target encoded blocks including the same motion information are obtained. Based on the block vector information and the intra - block copy prediction mode of the plurality of target encoded blocks , currentObtain a predicted value of the current encoding block from the existing image, and determine a reconstruction value of the current encoding block based on the predicted value. Thereby, at the time of encoding on the encoding side, it is not necessary to encode the motion information of each sub-block respectively, and it is only necessary to encode the motion information of the merge block, effectively saving bit overhead and being advantageous for improving encoding performance. Also, based on the motion information of the current encoding block, a division unit is determined. With multiple division methods, on the premise that the division of the encoding block satisfies the minimum division unit, the current block can be predicted using adjacent pixels as much as possible, and for screen content video that often has strong spatial correlation, the prediction ability of the intra-block copy prediction mode can be improved to a certain extent.
[0147] To facilitate better implementation of the video encoding method of the embodiments of the present application, the embodiments of the present application further provide a video encoding device. Refer to FIG. 13. FIG. 13 is a schematic diagram of the configuration of the video encoding device provided in the embodiments of the present application. Here, this video encoding device 900 is a determination unit 910 that determines the positional relationship between the current encoding block and the reference encoding block, and when the positional relationship overlaps, determines that the current encoding block is a merge block, and obtains a plurality of target encoding blocks by dividing the merge block, where the same motion information is included in the plurality of target encoding blocks, and based on the block vector information and the intra-block copy prediction mode of the plurality of target encoding blocks , current a prediction unit 920 that obtains a predicted value of the current encoding block from the existing image, and a reconstruction unit 930 that determines a reconstruction value of the current encoding block based on the predicted value, may be included.
[0148] It should be noted that the functions of each module of the video encoding device 900 in the embodiments of the present application may refer to any specific implementation form of the above method embodiments respectively, and the description is omitted here.
[0149] Each unit of the video encoding device 900 may be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above units may be embedded in a processor of a computer device in the form of hardware, or may be independent. Further, each of the above units may be stored in the memory of a computer device in the form of software, thereby facilitating the processor to call and execute the operations corresponding to each of the above units.
[0150] The video encoding device 900 may be incorporated, for example, into a terminal or a server having a memory and a processor with computing capabilities, or the video encoding device 900 may be that terminal or server. The terminal may be a device such as a smartphone, a tablet computer, a notebook computer, a smart TV, a smart speaker, a wearable smart device, or a personal computer (PC). The terminal may include a client. The client may be a video client, a browser client, or an instant communication client, etc. The server may be an independent physical server, or may be a server cluster or a distributed system composed of multiple physical servers, or may be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and base cloud computing services such as big data and artificial intelligence platforms.
[0151] FIG. 14 is a schematic configuration diagram of a computer device 1000 provided in an embodiment of the present application. As shown in FIG. 14, the computer device 1000 may include a communication interface 1100, a memory 1200, a processor 1300, and a communication bus 1400. The communication interface 1100, the memory 1200, and the processor 1300 realize mutual communication via the communication bus 1400. The communication interface 1100 is for the computer device 1000 to perform data communication with an external device. The memory 1200 can be used to store software programs and modules. The processor 1300 executes software programs and modules stored in the memory 1200, for example, the software programs corresponding to the respective operations in the foregoing method embodiments.
[0152] Optionally, when the computer device 1000 is an encoding device, the processor 1300 calls software programs and modules stored in the memory 1200 to determine the positional relationship between the current encoding block and the reference encoding block, and when the positional relationship overlaps, determining that the current encoding block is a merge block, and obtaining a plurality of target encoding blocks by splitting the merge block, where the same motion information is included in the plurality of target encoding blocks, and based on the block vector information and the intra-block copy prediction mode of the plurality of target encoding blocks , current obtaining a predicted value of the current encoding block from the existing image, and determining a reconstructed value of the current encoding block based on the predicted value may be executed.
[0153] Optionally, when the computer device 1000 is a decoding device, the processor 1300 calls software programs and modules stored in the memory 1200 to receive a bitstream of an encoded video including a current image; determine a positional relationship between a current encoded block and a reference encoded block based on the bitstream; when the positional relationship overlaps, determine that the current encoded block is a merge block; obtaining a plurality of target encoded blocks by splitting the merge block, wherein the same motion information is included in the plurality of target encoded blocks; obtaining a predicted value of the current encoded block from the current image based on the block vector information and the intra-block copy prediction mode of the plurality of target encoded blocks; determining a reconstructed value of the current encoded block based on the predicted value, and obtaining a bitstream of a decoded video by decoding the reconstructed value. These steps may be executed.
[0154] Optionally, the computer device 1000 may be a terminal or a server. The terminal may be a device such as a smartphone, a tablet computer, a notebook computer, a smart TV, a smart speaker, a wearable smart device, or a personal computer. The server may be an independent physical server, or may be a server cluster or a distributed system composed of a plurality of physical servers, and may also be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and base cloud computing services such as big data and artificial intelligence platforms.
[0155] Optionally, the present application further provides a computer device comprising a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in each of the above method embodiments are realized.
[0156] The present application further provides a computer-readable storage medium for storing a computer program. The computer-readable storage medium is applicable to a computer device, and the computer program causes the computer device to execute the corresponding flow in the encoding block processing method of the embodiments of the present application. For the sake of simplicity, further description is omitted here.
[0157] The present application further provides a computer program product including computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device is caused to execute the corresponding flow in the encoding block processing method of the embodiments of the present application. For the sake of simplicity, further description is omitted here.
[0158] The present application further provides a computer program including computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device is caused to execute the corresponding flow in the encoding block processing method of the embodiments of the present application. For the sake of simplicity, further description is omitted here.
[0159] The above functions may be stored in a computer-readable storage medium when implemented in the form of software function units and sold or used as independent products. Based on such understanding, the technical method of this application, in essence, in other words, the part that contributes to the prior art, or a part of the technical method, may be embodied in the form of a software product. The computer software product is stored in a storage medium and contains some instructions for causing a computer device (which may be a personal computer or a server) to execute all or part of the steps of the method in each embodiment of this application. The above-mentioned storage medium includes various media that can store program codes, such as USB disks, mobile hard disks, ROMs, RAMs, magnetic disks, or optical disks.
[0160] The above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should all be included within the protection scope of this application. Therefore, the protection scope of this application should follow the scope of the claims.
Claims
1. A method for processing an encoded block based on an intra-block copy prediction mode, executed by a computer device, comprising: receiving a bitstream, and determining a positional relationship between a current encoded block and a reference encoded block based on the bitstream; when the positional relationship overlaps, determining that the current encoded block is a merge block; obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit, wherein the same motion information is included in the plurality of target encoded blocks; obtaining an image sampling format of the current encoded block; obtaining a plurality of target encoded blocks of different chrominance by dividing a chrominance component of the merge block based on the basic division unit; determining an image sampling ratio of the current encoded block based on the image sampling format, and including: the step of obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit includes: obtaining a plurality of target encoded blocks of different chrominance by dividing a chrominance component of the merge block based on the basic division unit; obtaining a plurality of target encoded blocks of different luminance by dividing a luminance component of the merge block based on the basic division unit and the image sampling ratio; A method comprising.
2. The step of determining a positional relationship between a current encoded block and a reference encoded block based on the bitstream includes: determining a block vector of the current encoded block based on the bitstream; obtaining first position information of the reference encoded block based on the block vector; determining the positional relationship based on the first position information and the current encoded block, and including: The method according to claim 1.
3. The step of determining the positional relationship based on the first position information and the current encoded block includes: obtaining second position information of the current encoded block; When the horizontal coordinate of the lower right corner in the first position information is greater than or equal to the horizontal coordinate of the upper left corner in the second position information, and the vertical coordinate of the lower right corner in the first position information is greater than or equal to the vertical coordinate of the upper left corner in the second position information, determining that the positional relationship overlaps, The method according to claim 2.
4. The step of obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit includes: The step of obtaining the basic division unit; Comparing the basic division unit with the block vector of the current encoded block, and determining a target division unit based on the comparison result; Obtaining the plurality of target encoded blocks by dividing the merge block based on the target division unit. The method according to any one of claims 1 to 3.
5. The step of comparing the basic division unit with the block vector of the current encoded block and determining a target division unit based on the comparison result includes: When the y coordinate value indicating the height direction component of the block vector is less than or equal to the height value of the basic division unit, using the width value of the current encoded block as the target division unit in the x coordinate indicating the width direction position of the current encoded block, and using the height value of the basic division unit as the target division unit in the y coordinate indicating the height direction position of the current encoded block; When the y coordinate value of the block vector is greater than the height value of the basic division unit, using the width value of the basic division unit as the target division unit in the x coordinate of the current encoded block, and using the height value of the basic division unit as the target division unit in the y coordinate of the current encoded block. The method according to claim 4.
6. The step of comparing the basic division unit with the block vector of the current encoded block and determining a target division unit based on the comparison result includes: When the y - coordinate value indicating the height - direction component of the block vector is less than or equal to the value of the height of the basic division unit, at the x - coordinate indicating the width - direction position of the current encoded block, the value of the width of the current encoded block is used as the target division unit, and at the y - coordinate indicating the height - direction position of the current encoded block, the value of the height of the basic division unit is used as the target division unit; When the y - coordinate value of the block vector is greater than the value of the height of the basic division unit, at the x - coordinate of the current encoded block, the value of the width of the basic division unit is used as the target division unit, and at the y - coordinate of the current encoded block, the value of the height of the current encoded block is used as the target division unit, including the steps of; The method according to claim 4.
7. A video decoding method executed by a computer device, comprising: Receiving a bitstream of an encoded video including a current image; Determining a positional relationship between a current encoded block and a reference encoded block based on the bitstream; When the positional relationship overlaps, determining that the current encoded block is a merge block; Obtaining a plurality of target encoded blocks by dividing the merge block, wherein the same motion information is included in the plurality of target encoded blocks; Obtaining a predicted value of the current encoded block from the current image based on the block vector information and the intra - block copy prediction mode of the plurality of target encoded blocks; Determining a reconstructed value of the current encoded block based on the predicted value, and obtaining a bitstream of the decoded video by decoding the reconstructed value; Determining a sequence header flag based on the bitstream of the encoded video, including the steps of; The step of determining a positional relationship between a current encoded block and a reference encoded block based on the bitstream is as follows: When the flag value of the sequence header flag is a predetermined value, determining a positional relationship between the current encoded block and the reference encoded block, wherein the flag value of the sequence header flag is set during encoding on the encoding side.
8. If it is determined that the current encoded block is a merge block, the method further includes the step of adjusting the residual coefficients corresponding to the current encoded block to zero. The method according to claim 7.
9. A video encoding method executed by a computer device, comprising: determining a positional relationship between a current encoded block and a reference encoded block; when the positional relationship overlaps, determining that the current encoded block is a merge block; obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit, wherein the same motion information is included in the plurality of target encoded blocks; obtaining an image sampling format of the current encoded block; obtaining a plurality of chrominance target encoded blocks by dividing the chrominance component of the merge block based on the basic division unit; determining an image sampling ratio of the current encoded block based on the image sampling format, and the step of obtaining a plurality of target encoded blocks by dividing the merge block based on a predetermined basic division unit includes: obtaining a plurality of chrominance target encoded blocks by dividing the chrominance component of the merge block based on the basic division unit; obtaining a plurality of luminance target encoded blocks by dividing the luminance component of the merge block based on the basic division unit and the image sampling ratio; obtaining a predicted value of the current encoded block from the current image based on the block vector information and the intra-block copy prediction mode of the plurality of target encoded blocks, and determining a reconstructed value of the current encoded block based on the predicted value; A method including the above steps.
10. An apparatus for processing an encoded block based on an intra-block copy prediction mode, comprising: a receiving unit configured to receive a bitstream and determine a positional relationship between a current encoded block and a reference encoded block based on the bitstream; a determining unit configured to determine that the current encoded block is a merge block when the positional relationship overlaps; A splitting unit that obtains a plurality of target encoded blocks by splitting the merge block based on a predetermined basic splitting unit, wherein the same motion information is included in the plurality of target encoded blocks, a unit that obtains the image sampling format of the current encoded block, a unit that obtains a plurality of target encoded blocks of different chrominance by splitting the chrominance component of the merge block based on the basic splitting unit, and a unit that determines the image sampling ratio of the current encoded block based on the image sampling format. The splitting unit further includes a unit that obtains a plurality of target encoded blocks of different chrominance by splitting the chrominance component of the merge block based on the basic splitting unit, and a unit that obtains a plurality of target encoded blocks of different luminance by splitting the luminance component of the merge block based on the basic splitting unit and the image sampling ratio. An apparatus including the above.
11. A video decoding apparatus, a receiving unit that receives a bitstream of an encoded video including a current image, A processing unit that determines the positional relationship between a current encoded block and a reference encoded block based on the bitstream, determines that the current encoded block is a merge block when the positional relationship overlaps, and divides the merge block based on a predetermined basic division unit to obtain a plurality of target encoded blocks, wherein the same motion information is included in the plurality of target encoded blocks, a unit that obtains the image sampling format of the current encoded block, a unit that divides the chrominance component of the merge block based on the basic division unit to obtain a plurality of chrominance target encoded blocks, a unit that determines the image sampling ratio of the current encoded block based on the image sampling format, a unit that divides the chrominance component of the merge block based on the basic division unit to obtain a plurality of chrominance target encoded blocks, and a unit that divides the luminance component of the merge block based on the basic division unit and the image sampling ratio to obtain a plurality of luminance target encoded blocks. A prediction unit that obtains a predicted value of the current encoded block from the current image based on the block vector information and the intra-block copy prediction mode of the plurality of target encoded blocks. A reconstruction unit that determines a reconstructed value of the current encoded block based on the predicted value. A decoding unit that obtains a bitstream of the decoded video by decoding the reconstructed value. An apparatus including the above. [
12. ] A video encoding apparatus, A determination unit that determines the positional relationship between the current encoded block and the reference encoded block, and if the positional relationship overlaps, determines that the current encoded block is a merge block, and based on a predetermined basic division unit, divides the merge block to obtain a plurality of target encoded blocks. The plurality of target encoded blocks include the same motion information, a unit that obtains the image sampling format of the current encoded block, a unit that divides the chrominance component of the merge block based on the basic division unit to obtain a plurality of chrominance target encoded blocks, a unit that determines the image sampling ratio of the current encoded block based on the image sampling format, a unit that divides the chrominance component of the merge block based on the basic division unit to obtain a plurality of chrominance target encoded blocks, and a unit that divides the luminance component of the merge block based on the basic division unit and the image sampling ratio to obtain a plurality of luminance target encoded blocks, and a determination unit including A prediction unit that obtains a predicted value of the current encoded block from the current image based on the block vector information and the intra-block copy prediction mode of the plurality of target encoded blocks; A reconstruction unit that determines a reconstructed value of the current encoded block based on the predicted value; An apparatus including
13. A computer program that causes a computer to execute the method according to any one of Claims 1, 7, or 9.
14. A computer device including a processor and a memory, wherein the memory stores a computer program, and the processor calls the computer program stored in the memory to execute the method according to any one of Claims 1, 7, or 9.
Citation Information
Patent Citations
Features of the Intra Block Copy Prediction Mode for Video and Image Encoding and Decoding
JP2016539542A
Innovation in Block Vector Prediction and Estimation of Reconstructed Sample Values in Overlap Areas
JP2017511620A
Encoder-side options for intra block copy prediction mode for video and image coding
US20160241858A1