Method and apparatus for video encoding and decoding, computer-readable medium and electronic device

By further dividing the video encoding blocks and adapting to the various residual distributions, the problem of insufficient sub-block transformation performance in the prior art is solved and the encoding and decoding efficiency is improved.

WO2024212676A9PCT designated stage expired Publication Date: 2025-06-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/074646
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-13
Filing Date
2024-01-30
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

When existing video encoding and decoding technologies deal with a variety of residual distributions, the performance of sub-block transformation is insufficient, which affects the encoding and decoding efficiency.

Method used

By further dividing the sub-blocks obtained by coding block division, the performance of sub-block transformation is improved. The specific method includes obtaining block division information from the code stream, decoding the target sub-block and further dividing, performing entropy decoding and inverse quantization processing, and finally performing inverse transformation processing to generate reconstruction residuals.

Benefits of technology

It improves the performance of sub-block transformation, enhances the efficiency of encoding and decoding, and can better adapt to multiple residual distributions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024074646_19062025_PF_FP_ABST
    Figure CN2024074646_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in embodiments of the present application are a method and an apparatus for video encoding and decoding, a computer-readable medium and an electronic device. The video decoding method comprises: obtaining block division information corresponding to a to-be-decoded coding block from a code stream, wherein the block division information comprises information of a target sub-block needing entropy decoding, and the target sub-block is obtained by dividing the coding block according to a residual block corresponding to the coding block; decoding in the code stream to obtain information of at least one transformation sub-block obtained by dividing the target sub-block; performing entropy decoding and inverse quantization on the at least one transformation sub-block based on the information of the at least one transformation sub-block to obtain an inverse quantization coefficient sub-block corresponding to the at least one transformation sub-block; and performing inverse transformation on the inverse quantization coefficient sub-block corresponding to the at least one transformation sub-block, and generating a reconstruction residual error corresponding to the coding block according to results of the inverse transformation, wherein a residual error of the area in the coding block excluding the target sub-block is inferred to be zero.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding and decoding method, device, computer-readable medium, and electronic device

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on April 13, 2023, with application number 202310428833.7, and invention name “Video Coding and Decoding Method, Device, Computer-Readable Medium and Electronic Device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computers and communications technology, and more specifically, to a video encoding and decoding method, apparatus, computer-readable medium, and electronic device. Background Art

[0003] Video coding and decoding technology compresses and decompresses video signals. It reduces the bandwidth and storage required for data transmission and storage by reducing redundancy in video signal data, thereby achieving efficient video data transmission and storage. In relevant audio and video standards, sub-block transform (SBT) technology divides a coding block into multiple sub-blocks according to a specific method for encoding and decoding.

[0004] Technical content

[0005] The embodiments of the present application provide a video encoding and decoding method, apparatus, computer-readable medium, and electronic device, which can further divide the sub-blocks obtained by dividing the coding block to adapt to a variety of residual distribution situations, thereby improving the performance of sub-block transformation, which is conducive to improving encoding and decoding performance.

[0006] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.

[0007] An embodiment of the present application provides a video decoding method, comprising: obtaining block partitioning information corresponding to a coding block to be decoded from a bitstream, the block partitioning information including information of a target subblock that needs to be entropy decoded; wherein the target subblock is obtained by partitioning the coding block according to a residual block corresponding to the coding block; decoding from the bitstream to obtain information of at least one transform subblock obtained by partitioning the target subblock; performing entropy decoding and inverse quantization processing on the at least one transform subblock based on the information of the at least one transform subblock to obtain an inverse quantization coefficient subblock corresponding to each of the at least one transform subblocks; performing inverse transform processing on the inverse quantization coefficient subblock corresponding to each of the at least one transform subblocks, and generating a reconstructed residual corresponding to the coding block based on the inverse transform processing result, wherein the residual of an area of ​​the coding block other than the target subblock is inferred to be zero.

[0008] An embodiment of the present application also provides a video encoding method, including: obtaining a residual block corresponding to a current block to be encoded; determining corresponding block partitioning information based on the residual block, wherein the block partitioning information includes information of a target sub-block that needs to be entropy encoded; wherein the target sub-block is obtained by dividing the encoding block according to the residual block corresponding to the encoding block; dividing the target sub-block to obtain at least one transform sub-block; performing transform processing and quantization processing on the at least one transform sub-block to obtain a quantization coefficient block, so as to perform encoding processing based on the quantization coefficient block.

[0009] An embodiment of the present application further provides a video decoding device, comprising: an acquisition unit, configured to acquire block partitioning information corresponding to a coding block to be decoded from a bitstream, the block partitioning information including information of a target sub-block that needs to be entropy decoded; wherein the target sub-block is obtained by partitioning the coding block according to a residual block corresponding to the coding block; a decoding unit, configured to decode from the bitstream to obtain information of at least one transform sub-block obtained by partitioning the target sub-block, and perform entropy decoding and inverse quantization processing on the at least one transform sub-block based on the information of the at least one transform sub-block to obtain an inverse quantization coefficient sub-block corresponding to each of the at least one transform sub-blocks; a processing unit, configured to perform inverse transform processing on the inverse quantization coefficient sub-block corresponding to each of the at least one transform sub-blocks, and generate a reconstructed residual corresponding to the coding block based on the inverse transform processing result, wherein the residual of an area in the coding block other than the target sub-block is inferred to be zero.

[0010] An embodiment of the present application also provides a video encoding device, including: an acquisition unit, configured to acquire a residual block corresponding to a current block to be encoded; a determination unit, configured to determine corresponding block partitioning information based on the residual block, wherein the block partitioning information includes information of a target sub-block that needs to be entropy encoded; wherein the target sub-block is obtained by dividing the encoding block according to the residual block corresponding to the encoding block; a partitioning unit, configured to divide the target sub-block to obtain at least one transform sub-block; and an encoding unit, configured to perform transform processing and quantization processing on the at least one transform sub-block to obtain a quantization coefficient block, so as to perform encoding processing based on the quantization coefficient block.

[0011] An embodiment of the present application further provides a non-volatile computer-readable medium on which a computer program is stored. When the computer program is executed by a processor, the video decoding method or the video encoding method as described in the above embodiments is implemented.

[0012] An embodiment of the present application also provides an electronic device, comprising: one or more processors; a storage device for storing one or more computer programs, wherein when the one or more computer programs are executed by the one or more processors, the electronic device implements the video decoding method or the video encoding method as described in the above embodiments.

[0013] The present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of an electronic device reads and executes the computer program from the computer-readable storage medium, causing the electronic device to perform the video decoding method or video encoding method provided in the various embodiments described above.

[0014] An embodiment of the present application further provides a non-volatile computer-readable storage medium on which a bit stream is stored. The bit stream is generated by the video encoding method provided by each embodiment of the present application.

[0015] In the technical solutions provided in some embodiments of the present application, block partitioning information corresponding to a coding block to be decoded is obtained to determine information of a target sub-block that needs to be entropy decoded, and then information of at least one transformed sub-block obtained by partitioning the target sub-block is obtained by decoding from a bitstream. Entropy decoding and inverse quantization processing are then performed based on the information of the at least one transformed sub-block to obtain an inverse quantized coefficient sub-block corresponding to the at least one transformed sub-block, and an inverse transform processing is performed on the inverse quantized coefficient sub-block corresponding to the at least one transformed sub-block to generate a reconstructed residual corresponding to the coding block based on the inverse transform processing result. In this way, the sub-blocks obtained by partitioning the coding block can be further divided to adapt to a variety of residual distribution situations, thereby improving the performance of sub-block transformation and facilitating improving encoding and decoding performance.

[0016] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application.

[0017] BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG1 is a schematic diagram showing an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied;

[0019] FIG2 is a schematic diagram showing the placement of a video encoding device and a video decoding device in a streaming transmission system according to an embodiment of the present application;

[0020] FIG3 shows a basic flow chart of a video encoder according to an embodiment of the present application;

[0021] FIG4 shows a schematic diagram of the division method of SBT in an embodiment of the present application;

[0022] FIG5 shows a schematic diagram of a combination of transformation modes in an SBT according to an embodiment of the present application;

[0023] FIG6 shows a flowchart of a video decoding method according to some embodiments of the present application;

[0024] FIG7 shows a flowchart of a video encoding method according to some embodiments of the present application;

[0025] 8A to 8J are schematic diagrams showing sub-block division methods according to some embodiments of the present application;

[0026] FIG9A shows a schematic diagram of a combination of transform modes of a transform block according to some embodiments of the present application;

[0027] FIG9B-1 and FIG9B-2 are schematic diagrams showing other combinations of transform modes of transform blocks in some embodiments of the present application;

[0028] FIG9C-1 and FIG9C-2 are schematic diagrams showing other combinations of transform modes of transform blocks in some embodiments of the present application;

[0029] FIG9D-1 and FIG9D-2 are schematic diagrams showing other combinations of transform modes of transform blocks in some embodiments of the present application;

[0030] FIG9E is a schematic diagram showing some other combinations of transform modes of transform blocks in some embodiments of the present application;

[0031] FIG9F-1 and FIG9F-2 are schematic diagrams showing other combinations of transformation modes of transformation blocks in some embodiments of the present application;

[0032] FIG9G is a schematic diagram showing some other combinations of transform modes of transform blocks in some embodiments of the present application;

[0033] 9H-1 and 9H-2 are schematic diagrams showing other combinations of transform modes of transform blocks in some embodiments of the present application;

[0034] FIG9I-1 and FIG9I-2 are schematic diagrams showing other combinations of transform modes of transform blocks in some embodiments of the present application;

[0035] FIG9J-1 and FIG9J-2 are schematic diagrams showing other combinations of transform modes of transform blocks in some embodiments of the present application;

[0036] FIG10 shows a block diagram of a video decoding apparatus according to some embodiments of the present application;

[0037] FIG11 shows a block diagram of a video encoding apparatus according to some embodiments of the present application;

[0038] FIG12 shows a schematic structural diagram of a computer system suitable for implementing an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0039] Example embodiments will now be described in a more complete manner with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to these examples; rather, these embodiments are provided to make this application more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art.

[0040] In addition, the features, structures or characteristics described in the present application may be combined in one or more embodiments in any suitable manner. In the following description, there are many specific details so that the embodiments of the present application can be fully understood. However, it will be appreciated by those skilled in the art that when implementing the technical solution of the present application, it is not necessary to use all the detailed features in the embodiments, one or more specific details may be omitted, or other methods, elements, devices, steps, etc. may be adopted.

[0041] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0042] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0043] It should be noted that the term "plurality" used in this document refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. The character " / " generally indicates an "or" relationship between the associated objects.

[0044] FIG1 shows a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied.

[0045] 1 , a system architecture 100 includes a plurality of terminal devices that can communicate with each other via, for example, a network 150. For example, the system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via the network 150. In the embodiment of FIG1 , the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.

[0046] For example, the first terminal device 110 can encode video data (such as a video picture stream captured by the terminal device 110) for transmission to the second terminal device 120 via the network 150. The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device 120 can receive the encoded video data from the network 150, decode the encoded video data to restore the video data, and display the video picture based on the restored video data.

[0047] In some embodiments of the present application, the system architecture 100 may include a third terminal device 130 and a fourth terminal device 140 for performing bidirectional transmission of encoded video data, such as during a video conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode video data (e.g., a video picture stream captured by the terminal device) for transmission to the other of the third terminal device 130 and the fourth terminal device 140 via a network 150. Each of the third terminal device 130 and the fourth terminal device 140 may also receive the encoded video data transmitted by the other of the third terminal device 130 and the fourth terminal device 140, decode the encoded video data to recover the video data, and display the video pictures on an accessible display device based on the recovered video data.

[0048] In the embodiment shown in FIG. 1 , the first terminal device 110 , the second terminal device 120 , the third terminal device 130 , and the fourth terminal device 140 may be servers or terminals, but the principles disclosed herein are not limited thereto.

[0049] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. A terminal can be, but is not limited to, a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, intelligent voice interaction device, smartwatch, smart home appliance, vehicle-mounted terminal, aircraft, etc.

[0050] The network 150 shown in FIG1 represents any number of networks, including, for example, wired and / or wireless communication networks, for transmitting encoded video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140. Network 150 may exchange data using circuit-switched and / or packet-switched channels. The network may include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this application, unless otherwise explained below, the architecture and topology of network 150 may not be relevant to the operations disclosed herein.

[0051] In some embodiments of the present application, FIG2 illustrates the placement of a video encoding device and a video decoding device in a streaming environment. The subject matter disclosed herein is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television, and storing compressed video on digital media such as CDs, DVDs, and memory sticks.

[0052] The streaming system may include an acquisition subsystem 213, which may include a video source 201, such as a digital camera, that creates an uncompressed video picture stream 202. In one embodiment, the video picture stream 202 includes samples captured by the digital camera. The video picture stream 202 is depicted as a thicker line to emphasize the higher data volume of the video picture stream compared to the encoded video data 204 (or the encoded video stream 204). The video picture stream 202 may be processed by an electronic device 220, which includes a video encoding device 203 coupled to the video source 201. The video encoding device 203 may include hardware, software, or a combination of hardware and software to implement or embody various aspects of the disclosed subject matter, as described in greater detail below. The encoded video data 204 (or the encoded video stream 204) is depicted as a thinner line to emphasize the lower data volume of the encoded video data 204 (or the encoded video stream 204), which may be stored on a streaming server 205 for future use. One or more streaming client subsystems, such as client subsystem 206 and client subsystem 208 in FIG2 , can access streaming server 205 to retrieve copies 207 and 209 of encoded video data 204. Client subsystem 206 can include, for example, a video decoding device 210 in electronic device 230. Video decoding device 210 decodes the incoming copy 207 of the encoded video data and generates an output video picture stream 211 that can be presented on a display 212 (e.g., a screen) or another presentation device. In some streaming systems, the encoded video data 204, video data 207, and video data 209 (e.g., video code streams) can be encoded according to certain video encoding / compression standards.

[0053] It should be noted that the electronic device 220 and the electronic device 230 may include other components not shown in the figure. For example, the electronic device 220 may include a video decoding device, and the electronic device 230 may also include a video encoding device.

[0054] In some embodiments of the present application, taking the international video coding standards HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), and China's national video coding standard AVS as examples, after a video frame image is input, the video frame image will be divided into several non-overlapping processing units according to a block size, and each processing unit will perform similar compression operations. This processing unit is called CTU (Coding Tree Unit), or LCU (Largest Coding Unit). The CTU can continue to be divided more finely to obtain one or more basic coding units CU (Coding Unit). CU is the most basic element in a coding link.

[0055] In other embodiments, this processing unit may also be referred to as a coding tile (i.e., tile), which is a rectangular area of ​​a multimedia data frame that can be independently decoded and encoded. In the AV1 standard, a coding tile can be further divided into one or more superblocks (SBs). An SB is the starting point for block division and can be further divided to obtain one or more blocks (Bs). Each block is the most basic element in the encoding process. In some embodiments, an SB can contain multiple Bs.

[0056] The above division method for video frame images can be called block partition structure. The following introduces some concepts in the encoding process:

[0057] Predictive Coding: Predictive coding includes intra-frame prediction and inter-frame prediction. The original video signal is predicted by a selected reconstructed video signal to produce a residual video signal. The encoder needs to determine which predictive coding mode to use for the current coding unit (or coding block) and inform the decoder. Intra-frame prediction uses the predicted signal from a previously encoded and reconstructed region within the same image; inter-frame prediction uses the predicted signal from a previously encoded image (called a reference image) that is different from the current image.

[0058] Transform & Quantization: After the residual video signal undergoes transformations such as the Discrete Fourier Transform (DFT) and Discrete Cosine Transform (DCT), the signal is converted to the transform domain, where the coefficients are known as transform coefficients. The transform coefficients are then subjected to a lossy quantization operation, which loses some information, making the quantized signal more suitable for compression. In some video coding standards, more than one transform scheme may be available. Therefore, the encoder must select one for the current coding unit (or coding block) and inform the decoder of this selection. The level of quantization is typically determined by the quantization parameter (QP). A larger QP value means that a wider range of coefficients will be quantized to the same output, which typically results in greater distortion and a lower bitrate. Conversely, a smaller QP value means that a smaller range of coefficients will be quantized to the same output, which typically results in less distortion and a higher bitrate.

[0059] Entropy Coding or Statistical Coding: The quantized transform domain signal is statistically compressed and encoded based on the frequency of occurrence of each value, and finally a binary (0 or 1) compressed code stream is output. At the same time, the encoding generates other information, such as the selected coding mode, motion vector data, etc., which also need to be entropy coded to reduce the bit rate. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to express the same signal. Common statistical coding methods include variable length coding (VLC) or context-based binary arithmetic coding (CABAC).

[0060] The context-based binary arithmetic coding (CABAC) process consists of three main steps: binarization, context modeling, and binary arithmetic coding. After binarization of the input syntax elements, the binary data can be encoded using either the normal coding mode or the bypass coding mode. The bypass coding mode eliminates the need to assign a specific probability model to each binary bit. Instead, the input binary bit bin values ​​are directly encoded using a simple bypass encoder, speeding up both encoding and decoding. Generally, different syntax elements are not completely independent, and even the same syntax elements have some memory. Therefore, according to conditional entropy theory, conditional coding using other coded syntax elements can further improve coding performance compared to independent encoding or memoryless coding. This coded symbol information used as a condition is called context. In the normal coding mode, the binary bits of the syntax elements are sequentially fed into the context modeler. The encoder assigns an appropriate probability model to each input binary bit based on the values ​​of previously coded syntax elements or binary bits. This process is known as context modeling. The context model corresponding to the syntax element can be located using ctxIdxInc (context index increment) and ctxIdxStart (context index Start). After the bin value and the assigned probability model are fed into the binary arithmetic encoder for encoding, the context model needs to be updated based on the bin value, which is the adaptive process in encoding.

[0061] Loop Filtering: The changed and quantized signal will be reconstructed through inverse quantization, inverse transformation and prediction compensation operations to obtain a reconstructed image. Compared with the original image, due to the influence of quantization, some information of the reconstructed image is different from the original image, that is, the reconstructed image will produce distortion. Therefore, the reconstructed image can be filtered, such as deblocking filter (DB), SAO (Sample Adaptive Offset) or ALF (Adaptive Loop Filter) and other filters, which can effectively reduce the degree of distortion caused by quantization. Since these filtered reconstructed images will be used as a reference for subsequent encoded images to predict future image signals, the above filtering operation is also called loop filtering, that is, filtering operation within the encoding loop.

[0062] In some embodiments of the present application, FIG3 shows a basic flow chart of a video encoder, in which intra-frame prediction is used as an example for explanation. k[x,y] and predicted image signal Perform difference operation to obtain the residual signal u k [x,y], residual signal u k [x,y] is transformed and quantized to obtain the quantized coefficients. The quantized coefficients are entropy coded to obtain the encoded bit stream, and the reconstructed residual signal u′ is obtained through inverse quantization and inverse transformation. k [x,y], predicted image signal and reconstructed residual signal u′ k [x,y] superposition generates image signal Image signal On the one hand, it is input to the intra-frame mode decision module and the intra-frame prediction module for intra-frame prediction processing, and on the other hand, the reconstructed image signal s′ is output through loop filtering. k [x,y], reconstructed image signal s′ k [x,y] can be used as the reference image for the next frame for motion estimation and motion compensation prediction. Then based on the result of motion compensation prediction s′ r [x+m x ,y+m y ] and intra prediction results Get the predicted image signal of the next frame And continue to repeat the above process until the encoding is completed.

[0063] Based on the above encoding process, at the decoding end, after obtaining the compressed bitstream (i.e., bitstream), entropy decoding is performed on each coding unit (or coding block) to obtain various mode information and quantization coefficients. The quantized coefficients are then dequantized and inversely transformed to produce a residual signal. Furthermore, based on the known coding mode information, a prediction signal corresponding to the coding unit (or coding block) can be obtained. The residual signal is then added to the prediction signal to produce a reconstructed signal. The reconstructed signal then undergoes loop filtering and other operations to produce the final output signal.

[0064] In the aforementioned encoding process, due to the large error in the prediction method, it is necessary to transmit the residual signal to compensate for the predicted image, thereby improving the quality of the reconstructed image. Therefore, residual processing is an important processing process in the hybrid coding framework. In the process shown in Figure 3, the residual signal is the difference between the original image signal and the predicted image signal, that is, In video coding standards such as HEVC, VVC, and AVS3, the residual signal undergoes the aforementioned transformation (or transformation skipping), quantization processing, and the like.

[0065] The transform process primarily exploits the correlation of the residual signal. Through the transform, the residual signal's energy is concentrated in fewer low-frequency coefficients, resulting in smaller coefficient values. After the subsequent quantization module, these smaller coefficients are reduced to zero, significantly reducing the cost of encoding the residual. Due to the diverse distribution of residual signals, a single DCT cannot accommodate all residual characteristics. Therefore, transform kernels such as DST7 and DCT8 are introduced into the transform module, and different kernels can be used for horizontal and vertical transforms. Taking the Adaptive Multiple Core Transform (AMT) technology as an example, the possible transform combinations for a residual block include: (DCT2, DCT2), (DCT8, DCT8), (DCT8, DST7), (DST7, DCT8), and (DST7, DST7). The specific transform combination selected for a residual block requires rate-distortion optimization (RDO) at the encoder to determine the optimal transform combination. There are also some residual signals that have weak correlation, so skipping the transformation can actually lead to higher coding efficiency, that is, skipping the transformation process of the residual and directly quantizing the residual.

[0066] During transform coding, a sub-block transform (SBT) technique is used. Figure 4 shows the 12 SBT modes. The width and height of the coding block are W and H, respectively, and the sub-block size is 1, 1 / 2, or 1 / 4 of the corresponding size of the coding block. SBT only transforms the gray areas shown in Figure 4, forcing the rest of the area to zero. Gray blocks are not further divided and are directly transformed and quantized.

[0067] Figure 5 shows the transform combinations (i.e., combinations of horizontal and vertical transforms) corresponding to the 12 subblock partitioning modes shown in Figure 4. For example, for mode a, the subblock width is equal to the coding block width W, and the subblock height is equal to 1 / 2 of the coding block height, i.e., H / 2. The corresponding transform combination is DST7 for the horizontal transform and DCT8 for the vertical transform. Accordingly, at the decoder, only the coefficients in the gray area need to be decoded, followed by inverse quantization and inverse transform. The residuals in other areas are set to zero by default.

[0068] However, the partitioning mode of the SBT technology mentioned in the above solution is not flexible enough and cannot effectively handle a variety of residual distribution situations, which in turn affects the encoding and decoding efficiency to a certain extent. Based on this, the embodiments of the present application propose a new video encoding and decoding solution that can further divide the sub-blocks obtained by dividing the coding block to adapt to a variety of residual distribution situations, thereby improving the performance of sub-block transformation and facilitating the improvement of encoding and decoding performance.

[0069] The following is a detailed description of the implementation details of the technical solution of the embodiment of the present application:

[0070] FIG6 shows a flowchart of a video decoding method according to some embodiments of the present application. The video decoding method can be performed by a device with computing and processing capabilities, such as a terminal device or a server. Referring to FIG6 , the video decoding method includes at least the following steps S610 to S640, which are described in detail as follows:

[0071] In step S610, block partition information corresponding to the coding block to be decoded is obtained from the bitstream, where the block partition information includes information of a target sub-block that needs to be entropy decoded, wherein the target sub-block is obtained by dividing the coding block according to the residual block corresponding to the coding block.

[0072] In some embodiments, the target sub-block is the region of the coding block whose residuals need to be encoded and transferred, while the residuals of other regions of the coding block other than the target sub-block do not need to be encoded and transferred. In other words, in the coding block, only the residuals of the target sub-block need to be encoded and transferred to the decoding end, and the decoding end defaults to zero residuals of other regions other than the target sub-block.

[0073] In some embodiments of the present application, a video image frame sequence includes a series of images, each of which can be further divided into slices, which can be further divided into a series of LCUs (or CTUs), and the LCU contains several CUs. When encoding, the video image frame is encoded in blocks. In some new video coding standards, such as the H.264 standard, there are macroblocks (MBs), which can be further divided into multiple prediction blocks (prediction) that can be used for predictive coding. In other standards, such as HEVC, basic concepts such as coding units CU, prediction units (PUs) and transform units (TUs) are used to functionally divide various block units and describe them using a new tree-based structure. For example, a CU can be divided into smaller CUs according to a quadtree, and the smaller CUs can be further divided to form a quadtree structure. The coding block in the embodiment of the present application can be a CU, or a block smaller than a CU, such as a smaller block obtained by dividing a CU.

[0074] In some embodiments, the decoding end can obtain block division information from the code stream, and the block division information includes information about the target sub-blocks obtained by dividing the coding block to be decoded, for example, the information about the target sub-block includes width information and height information of the target sub-block. The width information of the target sub-block includes a first ratio of the width of the target sub-block to the width of the coding block; the height information of the target sub-block includes a second ratio of the height of the target sub-block to the height of the coding block. In some embodiments, the numerical value of the first ratio and the numerical value of the second ratio are any of the following: 1, 1 / 4, 1 / 2, 3 / 4, 1 / 8.

[0075] In some embodiments, the width of the target sub-block may be 1 / 4 of the width of the coding block, and the height of the target sub-block may be 1 / 4 of the height of the coding block.

[0076] In some embodiments, the width of the target sub-block may be ¾ of the width of the coding block, and the height of the target sub-block may be ¾ of the height of the coding block.

[0077] In some embodiments, the width of the target sub-block may be equal to the coding block width, and the height of the target sub-block may be ¾ of the coding block height.

[0078] In some embodiments, the width of the target sub-block may be ¾ of the coding block width, and the height of the target sub-block may be equal to the coding block height.

[0079] In some embodiments, the width of the target sub-block may be 1 / 4 of the width of the coding block, and the height of the target sub-block may be 1 / 2 of the height of the coding block.

[0080] In some embodiments, the width of the target sub-block may be 1 / 4 of the width of the coding block, and the height of the target sub-block may be 3 / 4 of the height of the coding block.

[0081] In some embodiments, the width of the target sub-block may be 1 / 2 of the width of the coding block, and the height of the target sub-block may be 1 / 4 of the height of the coding block.

[0082] In some embodiments, the width of the target sub-block may be 1 / 2 of the width of the coding block, and the height of the target sub-block may be 3 / 4 of the height of the coding block.

[0083] In some embodiments, the width of the target sub-block may be ¾ of the width of the coding block, and the height of the target sub-block may be ¼ of the height of the coding block.

[0084] In some embodiments, the width of the target sub-block may be ¾ of the width of the coding block, and the height of the target sub-block may be ½ of the height of the coding block.

[0085] In some embodiments, the target sub-block information may further include position information of the target sub-block, where the position information is used to indicate the position of the target sub-block in the coding block. For example, the position of the target sub-block in the coding block may include any of the following: the upper left corner of the coding block, the upper right corner of the coding block, the lower left corner of the coding block, the lower right corner of the coding block, above the coding block, below the coding block, to the left of the coding block, or to the right of the coding block.

[0086] In step S620, information of at least one transformed sub-block obtained by dividing the target sub-block is obtained by decoding the bitstream.

[0087] In the embodiment of the present application, the target sub-block obtained by dividing the coding block can be further divided, for example, into smaller transform sub-blocks, to adapt to a variety of residual distribution situations. Of course, the target sub-block can also be divided into only one transform sub-block, that is, the target sub-block is not divided.

[0088] In some embodiments, if the target sub-block is further divided to obtain transform sub-blocks, the size of the transform sub-blocks obtained by dividing the target sub-blocks can meet the following condition: at least one of the height and width of the transform sub-block is an integer power of 2, which can reduce the hardware implementation cost.

[0089] Based on the above conditions, if the width of the target sub-block is 3 / 4 of the width of the coding block and the height of the target sub-block is 3 / 4 of the height of the coding block, then the target sub-block is divided into 4 transform sub-blocks. In some embodiments, the width of the first transform sub-block of the 4 transform sub-blocks is 1 / 2 of the width of the coding block and the height is 1 / 2 of the height of the coding block; the width of the second transform sub-block is 1 / 4 of the width of the coding block and the height is 1 / 2 of the height of the coding block; the width of the third transform sub-block is 1 / 2 of the width of the coding block and the height is 1 / 4 of the height of the coding block; and the width of the fourth transform sub-block is 1 / 4 of the width of the coding block and the height is 1 / 4 of the height of the coding block.

[0090] If the width of the target sub-block is equal to the coding block width and the height of the target sub-block is 3 / 4 of the coding block height, then the target sub-block is divided into three transform sub-blocks in the height direction. In some embodiments, the width of these three transform sub-blocks is the same as the coding block width and the height is 1 / 4 of the coding block height.

[0091] If the width of the target sub-block is 3 / 4 of the coding block width and the height of the target sub-block is equal to the coding block height, then the target sub-block is divided into three transform sub-blocks in the width direction. In some embodiments, the height of these three transform sub-blocks is the same as the height of the coding block and the width is 1 / 4 of the coding block width.

[0092] If the width of the target sub-block is 1 / 4 of the coding block width and the height of the target sub-block is 3 / 4 of the coding block height, the target sub-block is divided into two transform sub-blocks in the height direction. In some embodiments, the width of the first transform sub-block is 1 / 4 of the coding block width and the height is 1 / 2 of the coding block height; the width of the second transform sub-block is 1 / 4 of the coding block width and the height is 1 / 4 of the coding block height.

[0093] If the width of the target sub-block is 1 / 2 of the coding block width and the height of the target sub-block is 3 / 4 of the coding block height, the target sub-block is divided into two transform sub-blocks in the height direction. In some embodiments, the width of the first transform sub-block is 1 / 2 of the coding block width and the height is 1 / 2 of the coding block height; the width of the second transform sub-block is 1 / 2 of the coding block width and the height is 1 / 4 of the coding block height.

[0094] If the width of the target sub-block is 3 / 4 of the coding block width and the height of the target sub-block is 1 / 4 of the coding block height, the target sub-block is divided into two transform sub-blocks in the width direction. In some embodiments, the width of the first transform sub-block is 1 / 2 of the coding block width and the height is 1 / 4 of the coding block height; the width of the second transform sub-block is 1 / 4 of the coding block width and the height is 1 / 4 of the coding block height.

[0095] If the width of the target sub-block is 3 / 4 of the coding block width and the height of the target sub-block is 1 / 2 of the coding block height, the target sub-block is divided into two transform sub-blocks in the width direction. In some embodiments, the width of the first transform sub-block is 1 / 2 of the coding block width and the height is 1 / 2 of the coding block height; the width of the second transform sub-block is 1 / 4 of the coding block width and the height is 1 / 2 of the coding block height.

[0096] In step S630, entropy decoding and inverse quantization are performed on the at least one transformed sub-block based on the information of the at least one transformed sub-block to obtain inverse quantization coefficient sub-blocks corresponding to each of the at least one transformed sub-blocks.

[0097] In an embodiment of the present application, after determining the information of the transform sub-block, entropy decoding and inverse quantization processing can be performed based on the information of the transform sub-block to obtain the inverse quantization coefficient sub-block corresponding to each transform sub-block. In some embodiments, the information of the transform sub-block may also include the width of the transform sub-block, the height of the transform sub-block, the position information of the transform sub-block, etc.

[0098] In step S640, an inverse transform is performed on the inverse quantized coefficient sub-block corresponding to each of the at least one transform sub-blocks, and a reconstructed residual corresponding to the coding block is generated based on the inverse transform result, wherein the residual of the area in the coding block except the target sub-block is inferred to be zero.

[0099] In some embodiments, performing inverse transform processing on the inverse quantized coefficient sub-blocks corresponding to at least one transform sub-block includes: selecting a horizontal transform mode and a vertical transform mode corresponding to each inverse quantized coefficient sub-block from set transform modes, where the set transform modes include: DCT2, DCT5, DCT8, DST1, DST7, and transform skip mode; and then performing inverse transform processing on each inverse quantized coefficient sub-block according to the horizontal transform mode and the vertical transform mode corresponding to each inverse quantized coefficient sub-block. That is, in this embodiment, the transform mode of the inverse quantized coefficient sub-block corresponding to the transform sub-block can be flexibly selected from DCT2, DCT5, DCT8, DST1, DST7, and transform skip mode.

[0100] In some embodiments, if a first dequantized coefficient subblock with a width greater than a set threshold value exists in the dequantized coefficient subblock corresponding to at least one transform subblock, the horizontal transform mode of the first dequantized coefficient subblock is replaced with a DCT2 transform mode; if a second dequantized coefficient subblock with a height greater than a set threshold value exists in the dequantized coefficient subblock corresponding to at least one transform subblock, the vertical transform mode of the second dequantized coefficient subblock is replaced with a DCT2 transform mode. The technical solution of this embodiment enables the use of a DCT2 transform mode that is more suitable for large-size residual block transforms when the size (such as height or width) of the dequantized coefficient subblock corresponding to the transform subblock is large, thereby ensuring the effect of coefficient transformation.

[0101] FIG6 illustrates the technical solution of the embodiment of the present application from the perspective of the decoding end. The technical solution of the embodiment of the present application is further explained from the perspective of the encoding end in conjunction with FIG7 .

[0102] FIG7 shows a flowchart of a video encoding method according to some embodiments of the present application. The video encoding method can be performed by a device with computing and processing capabilities, such as a terminal device or a server. Referring to FIG7 , the video encoding method includes at least the following steps S710 to S740, which are described in detail as follows:

[0103] In step S710, a residual block corresponding to the current block to be encoded is obtained.

[0104] In some embodiments, the residual block is the difference between the original image information (original image block) and the predicted image information (prediction block).

[0105] In step S720, corresponding block partitioning information is determined according to the residual block, where the block partitioning information includes information of a target sub-block that needs to be entropy-coded, wherein the target sub-block is obtained by dividing the block to be coded.

[0106] In some embodiments, the information of the target sub-block includes width information, height information, and position information of the target sub-block. For details, reference may be made to the technical solutions of the aforementioned embodiments.

[0107] In step S730, the target sub-block is divided to obtain at least one transformed sub-block.

[0108] In some embodiments, the target sub-block may be divided according to a predetermined division strategy to obtain at least one transform sub-block. The predetermined division strategy includes: at least one of the height and width of the transform sub-block obtained by division is an integer power of 2. The specific division method may refer to the technical solutions of the aforementioned embodiments.

[0109] In step S740, transform processing and quantization processing are performed on the obtained at least one transform sub-block to obtain a quantized coefficient block, so as to perform encoding processing based on the quantized coefficient block.

[0110] In some embodiments, the quantized coefficient block may be entropy coded, and information of at least one transformed sub-block obtained may be coded to obtain a coded bitstream, which is then transmitted to a decoding end.

[0111] It should be noted that the implementation details of the video encoding method shown in FIG. 7 are similar to the implementation details of the video decoding method described in the aforementioned embodiment, and are not described in detail herein.

[0112] In summary, the technical solution of the embodiment of the present application can further divide the sub-blocks obtained by dividing the coding block to adapt to a variety of residual distribution situations. Specifically, the size of the sub-blocks obtained by dividing the coding block can be a certain proportion of the size of the corresponding direction of the coding block, such as 1, 1 / 4, 1 / 2, 3 / 4, 1 / 8, etc. During encoding, only the residuals of the specified sub-block area need to be encoded and transferred, and the residuals of the remaining areas do not need to be encoded and transferred.

[0113] The following are several embodiments of dividing a coding block into sub-blocks in this application:

[0114] As shown in FIG8A , the width and height of the sub-block are both 1 / 4 of the corresponding directional dimensions of the coding block, and the sub-block can be located at any one of the upper left corner, upper right corner, lower left corner, and lower right corner of the coding block.

[0115] As shown in FIG8B , the width and height of the sub-block are both 3 / 4 of the corresponding directional dimensions of the coding block, and the position of the sub-block can be any one of the upper left corner, upper right corner, lower left corner and lower right corner of the coding block.

[0116] As shown in FIG8C , the width of the sub-block is equal to the width of the coding block, the height of the sub-block is ¾ of the height of the coding block, and the position of the sub-block can be above or below the coding block.

[0117] As shown in FIG8D , the height of the sub-block is equal to the height of the coding block, the width of the sub-block is 3 / 4 of the width of the coding block, and the position of the sub-block can be to the left or right of the coding block.

[0118] As shown in Figure 8E, the width of the sub-block is 1 / 4 of the width of the coding block, the height of the sub-block is 1 / 2 of the height of the coding block, and the position of the sub-block can be any one of the upper left corner, upper right corner, lower left corner and lower right corner of the coding block.

[0119] As shown in Figure 8F, the width of the sub-block is 1 / 4 of the width of the coding block, the height of the sub-block is 3 / 4 of the height of the coding block, and the position of the sub-block can be any one of the upper left corner, upper right corner, lower left corner and lower right corner of the coding block.

[0120] As shown in Figure 8G, the width of the sub-block is 1 / 2 of the width of the coding block, the height of the sub-block is 1 / 4 of the height of the coding block, and the position of the sub-block can be any one of the upper left corner, upper right corner, lower left corner and lower right corner of the coding block.

[0121] As shown in Figure 8H, the width of the sub-block is 1 / 2 of the width of the coding block, the height of the sub-block is 3 / 4 of the height of the coding block, and the position of the sub-block can be any one of the upper left corner, upper right corner, lower left corner and lower right corner of the coding block.

[0122] As shown in Figure 8I, the width of the sub-block is 3 / 4 of the width of the coding block, the height of the sub-block is 1 / 4 of the height of the coding block, and the position of the sub-block can be any one of the upper left corner, upper right corner, lower left corner and lower right corner of the coding block.

[0123] As shown in Figure 8J, the width of the sub-block is 3 / 4 of the width of the coding block, the height of the sub-block is 1 / 2 of the height of the coding block, and the position of the sub-block can be any one of the upper left corner, upper right corner, lower left corner and lower right corner of the coding block.

[0124] It should be noted that the various division modes shown in FIG. 8A to FIG. 8J can be used independently or in combination, and can also be used in combination with the division mode shown in FIG. 4 .

[0125] After the sub-blocks are divided, the transform combinations for the sub-blocks can be selected from DCT2, DCT5, DCT8, DST1, DST7, and TS (transform skip). In addition, the sub-block to be encoded can be transformed as a single transform block or divided into multiple transform blocks.

[0126] The following describes how to further divide sub-blocks in an embodiment of the present application by taking the transformation combination selected from DCT8 and DST7 as an example:

[0127] As shown in Figure 9A, the width and height of a sub-block are both 1 / 4 of the corresponding dimensions of the coding block, so the sub-block can be transformed as a whole block. The transform combination can be selected from DCT8 and DST7.

[0128] As shown in Figure 9B-1, the width and height of a sub-block are both 3 / 4 of the corresponding dimensions of the coding block. Therefore, the sub-block can be transformed as a whole transform block. The transform combination can be selected from DCT8 and DST7.

[0129] As shown in Figure 9B-2, the width and height of the sub-block are both 3 / 4 of the corresponding directional dimensions of the coding block. Then the sub-block can be further divided into 4 transform blocks (the dotted lines in the figure indicate the division boundaries), and the transform combinations of these 4 transform blocks can be selected arbitrarily from DCT8 and DST7.

[0130] As shown in Figure 9C-1, the width of the sub-block is equal to the width of the coding block, and the height of the sub-block is 3 / 4 of the height of the coding block. Therefore, the sub-block can be transformed as a whole block. The transform combination can be selected from DCT8 and DST7.

[0131] As shown in Figure 9C-2, the width of the sub-block is equal to the width of the coding block, and the height of the sub-block is 3 / 4 of the height of the coding block. Then the sub-block can be further divided into 3 transform blocks (the dotted lines in the figure indicate the division boundaries), and the transform combinations of these 3 transform blocks can be arbitrarily selected from DCT8 and DST7.

[0132] As shown in Figure 9D-1, the height of the sub-block is equal to the height of the coding block, and the width of the sub-block is 3 / 4 of the width of the coding block. Therefore, the sub-block can be transformed as a whole block. The transform combination can be selected from DCT8 and DST7.

[0133] As shown in Figure 9D-2, the height of the sub-block is equal to the height of the coding block, and the width of the sub-block is 3 / 4 of the width of the coding block. Then the sub-block can be further divided into 3 transform blocks (the dotted lines in the figure indicate the division boundaries), and the transform combinations of these 3 transform blocks can be arbitrarily selected from DCT8 and DST7.

[0134] As shown in Figure 9E, the width of the sub-block is 1 / 4 of the width of the coding block, and the height of the sub-block is 1 / 2 of the height of the coding block. Therefore, the sub-block can be transformed as a whole block. The transform combination can be selected from DCT8 and DST7.

[0135] As shown in Figure 9F-1, the width of the sub-block is 1 / 4 of the width of the coding block, and the height of the sub-block is 3 / 4 of the height of the coding block. Therefore, the sub-block can be transformed as a whole block. The transform combination can be selected from DCT8 and DST7.

[0136] As shown in Figure 9F-2, the width of the sub-block is 1 / 4 of the width of the coding block, and the height of the sub-block is 3 / 4 of the height of the coding block. Then the sub-block can be further divided into two transform blocks (the dotted line in the figure indicates the division boundary), and the transform combination of these two transform blocks can be selected arbitrarily from DCT8 and DST7.

[0137] As shown in Figure 9G, the width of the sub-block is 1 / 2 of the width of the coding block, and the height of the sub-block is 1 / 4 of the height of the coding block. Therefore, the sub-block can be transformed as a whole block. The transform combination can be selected from DCT8 and DST7.

[0138] As shown in Figure 9H-1, the width of the sub-block is 1 / 2 of the width of the coding block, and the height of the sub-block is 3 / 4 of the height of the coding block. Therefore, the sub-block can be transformed as a whole block. The transform combination can be selected from DCT8 and DST7.

[0139] As shown in Figure 9H-2, the width of the sub-block is 1 / 2 of the width of the coding block, and the height of the sub-block is 3 / 4 of the height of the coding block. Then the sub-block can be further divided into 2 transform blocks (the dotted line in the figure indicates the division boundary), and the transform combination of these 2 transform blocks can be selected arbitrarily from DCT8 and DST7.

[0140] As shown in Figure 9I-1, the width of the sub-block is 3 / 4 of the width of the coding block, and the height of the sub-block is 1 / 4 of the height of the coding block. Therefore, the sub-block can be transformed as a whole block. The transform combination can be selected from DCT8 and DST7.

[0141] As shown in Figure 9I-2, the width of the sub-block is 3 / 4 of the width of the coding block, and the height of the sub-block is 1 / 4 of the height of the coding block. Then the sub-block can be further divided into two transform blocks (the dotted line in the figure indicates the division boundary), and the transform combination of these two transform blocks can be selected arbitrarily from DCT8 and DST7.

[0142] As shown in Figure 9J-1, the width of the sub-block is 3 / 4 of the width of the coding block, and the height of the sub-block is 1 / 2 of the height of the coding block. Therefore, the sub-block can be transformed as a whole block. The transform combination can be selected from DCT8 and DST7.

[0143] As shown in Figure 9J-2, the width of the sub-block is 3 / 4 of the width of the coding block, and the height of the sub-block is 1 / 2 of the height of the coding block. Then the sub-block can be further divided into two transform blocks (the dotted line in the figure indicates the division boundary), and the transform combination of these two transform blocks can be selected arbitrarily from DCT8 and DST7.

[0144] In some embodiments of the present application, the horizontal transform mode and the vertical transform mode of the transform block in the sub-block may both be set to the DCT2 transform mode. Alternatively, the horizontal transform mode and the vertical transform mode of the transform block in the sub-block may both be set to the TS mode, i.e., the transform skip mode.

[0145] In some embodiments of the present application, the transform combination of the transform block in a sub-block may be selected from DCT2, DCT5, DCT8, DST1, DST7, and TS. However, if the size of the transform block in the sub-block is greater than a specified threshold, the transform mode in the corresponding direction may be forcibly changed to DCT2. For example, when the width of the transform block in the sub-block is greater than 8, the horizontal transform mode may be forcibly changed to DCT2; when the height of the transform block in the sub-block is greater than 8, the vertical transform mode may be forcibly changed to DCT2.

[0146] The technical solutions of the above embodiments of the present application expand the division mode of the SBT technology, thereby improving the flexibility of the sub-block transformation technology and further improving the encoding and decoding performance.

[0147] The following describes an embodiment of the device of the present application, which can be used to perform the method in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method in the above embodiment of the present application.

[0148] FIG10 shows a block diagram of a video decoding apparatus according to some embodiments of the present application. The video decoding apparatus may be provided in a device having a computing and processing function, such as a terminal device or a server.

[0149] 10 , a video decoding apparatus 1000 according to some embodiments of the present application includes: an acquisition unit 1002 , a decoding unit 1004 , and a processing unit 1006 .

[0150] Among them, the acquisition unit 1002 is configured to obtain block partitioning information corresponding to the coding block to be decoded, and the block partitioning information includes information of a target sub-block that needs to be entropy decoded; wherein the target sub-block is obtained by dividing the coding block according to the residual block corresponding to the coding block; the decoding unit 1004 is configured to decode from the code stream to obtain information of at least one transform sub-block obtained by dividing the target sub-block, and perform entropy decoding and inverse quantization processing on the at least one transform sub-block based on the information of the at least one transform sub-block to obtain the inverse quantization coefficient sub-block corresponding to each of the at least one transform sub-blocks; the processing unit 1006 is configured to perform inverse transform processing on the inverse quantization coefficient sub-block corresponding to each of the at least one transform sub-blocks, and generate a reconstructed residual corresponding to the coding block based on the inverse transform processing result, wherein the residual of the area other than the target sub-block in the coding block is inferred to be zero.

[0151] In some embodiments of the present application, based on the aforementioned scheme, the information of the target sub-block includes width information and height information of the target sub-block; the width information of the target sub-block includes a first ratio of the width of the target sub-block to the width of the coding block; and the height information of the target sub-block includes a second ratio of the height of the target sub-block to the height of the coding block.

[0152] In some embodiments of the present application, based on the aforementioned solution, the numerical value of the first ratio and the numerical value of the second ratio are any one of the following: 1, 1 / 4, 1 / 2, 3 / 4, 1 / 8.

[0153] In some embodiments of the present application, based on the aforementioned solution, the width of the target sub-block is 1 / 4 of the width of the coding block, and the height of the target sub-block is 1 / 4 of the height of the coding block; or

[0154] The width of the target sub-block is ¾ of the width of the coding block, and the height of the target sub-block is ¾ of the height of the coding block; or

[0155] The width of the target sub-block is equal to the width of the coding block, and the height of the target sub-block is ¾ of the height of the coding block; or

[0156] The width of the target sub-block is ¾ of the width of the coding block, and the height of the target sub-block is equal to the height of the coding block; or

[0157] The width of the target sub-block is 1 / 4 of the width of the coding block, and the height of the target sub-block is 1 / 2 of the height of the coding block; or

[0158] The width of the target sub-block is 1 / 4 of the width of the coding block, and the height of the target sub-block is 3 / 4 of the height of the coding block; or

[0159] The width of the target sub-block is 1 / 2 of the width of the coding block, and the height of the target sub-block is 1 / 4 of the height of the coding block; or

[0160] The width of the target sub-block is 1 / 2 of the width of the coding block, and the height of the target sub-block is 3 / 4 of the height of the coding block; or

[0161] The width of the target sub-block is ¾ of the width of the coding block, and the height of the target sub-block is ¼ of the height of the coding block; or

[0162] The width of the target sub-block is ¾ of the width of the coding block, and the height of the target sub-block is ½ of the height of the coding block.

[0163] In some embodiments of the present application, based on the aforementioned solution, the information of the target sub-block further includes position information of the target sub-block, and the position information is used to indicate the position of the target sub-block in the coding block.

[0164] In some embodiments of the present application, based on the aforementioned scheme, the position of the target sub-block in the coding block includes any one of the following: the upper left corner of the coding block, the upper right corner of the coding block, the lower left corner of the coding block, the lower right corner of the coding block, above the coding block, below the coding block, to the left of the coding block, and to the right of the coding block.

[0165] In some embodiments of the present application, based on the aforementioned solution, the size of the transform sub-block obtained by dividing the target sub-block satisfies the following condition: at least one of the height and width of the transform sub-block is an integer power of 2.

[0166] In some embodiments of the present application, based on the aforementioned scheme:

[0167] If the width of the target sub-block is ¾ of the width of the coding block and the height of the target sub-block is ¾ of the height of the coding block, the target sub-block is divided into 4 transform sub-blocks;

[0168] If the width of the target sub-block is equal to the coding block width and the height of the target sub-block is ¾ of the coding block height, the target sub-block is divided into three transform sub-blocks in the height direction;

[0169] If the width of the target sub-block is ¾ of the width of the coding block and the height of the target sub-block is equal to the height of the coding block, the target sub-block is divided into three transform sub-blocks in the width direction;

[0170] If the width of the target sub-block is 1 / 4 of the width of the coding block and the height of the target sub-block is 3 / 4 of the height of the coding block, the target sub-block is divided into two transform sub-blocks in the height direction;

[0171] If the width of the target sub-block is 1 / 2 of the width of the coding block and the height of the target sub-block is 3 / 4 of the height of the coding block, the target sub-block is divided into two transform sub-blocks in the height direction;

[0172] If the width of the target sub-block is ¾ of the width of the coding block and the height of the target sub-block is ¼ of the height of the coding block, the target sub-block is divided into two transform sub-blocks in the width direction;

[0173] If the width of the target sub-block is ¾ of the width of the coding block and the height of the target sub-block is ½ of the height of the coding block, the target sub-block is divided into two transform sub-blocks in the width direction.

[0174] In some embodiments of the present application, based on the aforementioned scheme, the processing unit 1006 is configured to: select the horizontal transform mode and the vertical transform mode corresponding to each of the dequantized coefficient sub-blocks from the set transform mode, the set transform mode including: DCT2, DCT5, DCT8, DST1, DST7 and transform skip mode; and perform inverse transform processing on each of the dequantized coefficient sub-blocks according to the horizontal transform mode and the vertical transform mode corresponding to each of the dequantized coefficient sub-blocks.

[0175] In some embodiments of the present application, based on the aforementioned scheme, the processing unit 1006 is further configured to: if there is a first dequantization coefficient sub-block with a width greater than a set threshold in the dequantization coefficient sub-block corresponding to the at least one transform sub-block, then the horizontal transform mode of the first dequantization coefficient sub-block is replaced with a DCT2 transform mode; if there is a second dequantization coefficient sub-block with a height greater than a set threshold in the dequantization coefficient sub-block corresponding to the at least one transform sub-block, then the vertical transform mode of the second dequantization coefficient sub-block is replaced with a DCT2 transform mode.

[0176] FIG11 shows a block diagram of a video encoding apparatus according to some embodiments of the present application. The video encoding apparatus may be provided in a device having a computing and processing function, such as a terminal device or a server.

[0177] 11 , a video encoding apparatus 1100 according to some embodiments of the present application includes: an acquiring unit 1102 , a determining unit 1104 , a dividing unit 1106 , and an encoding unit 1108 .

[0178] Among them, the acquisition unit 1102 is configured to obtain the residual block corresponding to the current block to be encoded; the determination unit 1104 is configured to determine the corresponding block division information based on the residual block, and the block division information contains information of the target sub-block that needs to be entropy encoded; wherein the target sub-block is obtained by dividing the coding block according to the residual block corresponding to the coding block; the division unit 1106 is configured to divide the target sub-block to obtain at least one transform sub-block; the encoding unit 1108 is configured to perform transform processing and quantization processing on the at least one transform sub-block to obtain a quantization coefficient block, so as to perform encoding processing based on the quantization coefficient block.

[0179] FIG12 shows a schematic structural diagram of a computer system suitable for implementing an electronic device according to an embodiment of the present application.

[0180] It should be noted that the computer system 1200 of the electronic device shown in FIG12 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0181] As shown in Figure 12, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage part 1208 into the random access memory (RAM) 1203, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 1203. The CPU 1201, ROM 1202 and RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0182] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, and the like; an output section 1207 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1208 including a hard disk; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as needed. Removable media 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1210 as needed, so that computer programs read from the removable media can be installed in the storage section 1208 as needed.

[0183] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1209, and / or installed from a removable medium 1211. When the computer program is executed by the central processing unit (CPU) 1201, the various functions defined in the system of the present application are executed.

[0184] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0185] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and a computer program.

[0186] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.

[0187] As another aspect, embodiments of the present application further provide a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device. The computer-readable medium carries one or more computer programs, and when the one or more computer programs are executed by the electronic device, the electronic device implements the method described in the above embodiments.

[0188] An embodiment of the present application further provides a non-volatile computer-readable storage medium, which stores a bit stream generated by the above-mentioned video encoding method of the present application.

[0189] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0190] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0191] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.

[0192] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A video decoding method, executed by a computer device, comprising: Obtaining block partition information corresponding to a coding block to be decoded from a bitstream, wherein the block partition information includes information of a target sub-block to be entropy decoded; wherein the target sub-block is obtained by partitioning the coding block according to a residual block corresponding to the coding block; Decoding from a bitstream to obtain information of at least one transformed sub-block obtained by dividing the target sub-block; Based on the information of the at least one transformed sub-block, entropy decoding and inverse quantization processing are performed on the at least one transformed sub-block to obtain inverse quantization coefficient sub-blocks corresponding to each of the at least one transformed sub-block; Performing inverse transform processing on the inverse quantized coefficient sub-blocks corresponding to each of the at least one transform sub-blocks, and generating a reconstructed residual corresponding to the coding block according to the inverse transform processing result, wherein the residual of the area in the coding block except the target sub-block is inferred to be zero.

2. The video decoding method according to claim 1, wherein: The information of the target sub-block includes width information and height information of the target sub-block; The width information of the target sub-block includes a first ratio of the width of the target sub-block to the width of the coding block; The height information of the target sub-block includes a second ratio of the height of the target sub-block to the height of the coding block.

3. The video decoding method according to claim 2, wherein: The value of the first ratio and the value of the second ratio are any of the following: 1, 1 / 4, 1 / 2, 3 / 4, 1 / 8.

4. The video decoding method according to claim 3, wherein: The width of the target sub-block is 1 / 4 of the width of the coding block, and the height of the target sub-block is 1 / 4 of the height of the coding block; or The width of the target sub-block is 3 / 4 of the width of the coding block, and the height of the target sub-block is 3 / 4 of the height of the coding block; or The width of the target sub-block is equal to the coding block width, and the height of the target sub-block is 3 / 4 of the coding block height; or The width of the target sub-block is 3 / 4 of the width of the coding block, and the height of the target sub-block is equal to the height of the coding block; or The width of the target sub-block is 1 / 4 of the width of the coding block, and the height of the target sub-block is 1 / 2 of the height of the coding block; or The width of the target sub-block is 1 / 4 of the width of the coding block, and the height of the target sub-block is 3 / 4 of the height of the coding block; or The width of the target sub-block is 1 / 2 of the width of the coding block, and the height of the target sub-block is 1 / 4 of the height of the coding block; or The width of the target sub-block is 1 / 2 of the width of the coding block, and the height of the target sub-block is 3 / 4 of the height of the coding block; or The width of the target sub-block is 3 / 4 of the width of the coding block, and the height of the target sub-block is 1 / 4 of the height of the coding block; or The width of the target sub-block is 3 / 4 of the width of the coding block, and the height of the target sub-block is 1 / 2 of the height of the coding block.

5. The video decoding method according to any one of claims 2 to 4, wherein: The information of the target sub-block also includes position information of the target sub-block, where the position information is used to indicate the position of the target sub-block in the coding block.

6. The video decoding method according to claim 5, wherein: The position of the target sub-block in the coding block includes any one of the following: the upper left corner of the coding block, the upper right corner of the coding block, the lower left corner of the coding block, the lower right corner of the coding block, above the coding block, below the coding block, to the left of the coding block, and to the right of the coding block.

7. The video decoding method according to claim 1, wherein: The information of the at least one transform sub-block includes: width information, height information and position information of each of the at least one transform sub-block.

8. The video decoding method according to claim 1, wherein: The size of the transform sub-block obtained by dividing the target sub-block satisfies the following condition: at least one of the height and width of the transform sub-block is an integer power of 2.

9. The video decoding method according to claim 8, wherein: If the width of the target sub-block is 3 / 4 of the width of the coding block, and the height of the target sub-block is 3 / 4 of the height of the coding block, the target sub-block is divided into 4 transform sub-blocks; If the width of the target sub-block is equal to the coding block width, and the height of the target sub-block is 3 / 4 of the coding block height, the target sub-block is divided into 3 transform sub-blocks in the height direction; If the width of the target sub-block is 3 / 4 of the width of the coding block, and the height of the target sub-block is equal to the height of the coding block, the target sub-block is divided into 3 transform sub-blocks in the width direction; If the width of the target sub-block is 1 / 4 of the width of the coding block, and the height of the target sub-block is 3 / 4 of the height of the coding block, the target sub-block is divided into 2 transform sub-blocks in the height direction; If the width of the target sub-block is 1 / 2 of the width of the coding block, and the height of the target sub-block is 3 / 4 of the height of the coding block, the target sub-block is divided into 2 transform sub-blocks in the height direction; If the width of the target sub-block is 3 / 4 of the width of the coding block, and the height of the target sub-block is 1 / 4 of the height of the coding block, the target sub-block is divided into 2 transform sub-blocks in the width direction; If the width of the target sub-block is 3 / 4 of the width of the coding block and the height of the target sub-block is 1 / 2 of the height of the coding block, the target sub-block is divided into 2 transform sub-blocks in the width direction.

10. The video decoding method according to claim 1, wherein: Performing inverse transform processing on the inverse quantized coefficient sub-block corresponding to the at least one transform sub-block, comprising: Selecting a horizontal transform mode and a vertical transform mode corresponding to each of the dequantized coefficient sub-blocks from a set transform mode, wherein the set transform mode includes: DCT2, DCT5, DCT8, DST1, DST7 and a transform skip mode; According to the horizontal transform mode and the vertical transform mode corresponding to each of the dequantized coefficient sub-blocks, inverse transform processing is performed on each of the dequantized coefficient sub-blocks.

11. The video decoding method according to claim 10, further comprising: If there is a first dequantized coefficient sub-block whose width is greater than a set threshold in the dequantized coefficient sub-block corresponding to the at least one transform sub-block, replacing the horizontal transform mode of the first dequantized coefficient sub-block with a DCT2 transform mode; If there is a second dequantized coefficient subblock whose height is greater than a set threshold in the dequantized coefficient subblock corresponding to the at least one transform subblock, the vertical transform mode of the second dequantized coefficient subblock is replaced with the DCT2 transform mode.

12. A video encoding method, performed by a computer device, comprising: Obtain the residual block corresponding to the current block to be encoded; Determining corresponding block partitioning information according to the residual block, wherein the block partitioning information includes information of a target sub-block to be entropy encoded; wherein the target sub-block is obtained by partitioning the coding block according to the residual block corresponding to the coding block; Dividing the target sub-block to obtain at least one transformation sub-block; The at least one transform sub-block is respectively transformed and quantized to obtain a quantized coefficient block, so as to perform encoding based on the quantized coefficient block.

13. A video decoding device, comprising: An acquisition unit is configured to acquire block partition information corresponding to a coding block to be decoded from a bitstream, wherein the block partition information includes information of a target sub-block to be entropy decoded; wherein the target sub-block is obtained by partitioning the coding block according to a residual block corresponding to the coding block; a decoding unit configured to decode from a bitstream to obtain information of at least one transform sub-block obtained by dividing the target sub-block, and perform entropy decoding and inverse quantization processing on the at least one transform sub-block based on the information of the at least one transform sub-block to obtain inverse quantization coefficient sub-blocks corresponding to each of the at least one transform sub-blocks; A processing unit is configured to perform inverse transform processing on the inverse quantization coefficient sub-blocks corresponding to each of the at least one transform sub-blocks, and generate a reconstructed residual corresponding to the coding block based on the inverse transform processing result, wherein the residual of the area in the coding block except the target sub-block is inferred to be zero.

14. A video encoding device, comprising: An acquisition unit, configured to acquire a residual block corresponding to a current block to be encoded; A determining unit is configured to determine corresponding block partitioning information according to the residual block, wherein the block partitioning information includes information of a target sub-block to be entropy encoded; wherein the target sub-block is obtained by partitioning the coding block according to the residual block corresponding to the coding block; a dividing unit, configured to divide the target sub-block to obtain at least one transformed sub-block; The encoding unit is configured to perform transform processing and quantization processing on the at least one transform sub-block respectively to obtain a quantization coefficient block, so as to perform encoding processing based on the quantization coefficient block.

15. A computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program implements the video decoding method according to any one of claims 1 to 11, or implements the video encoding method according to claim 12.

16. An electronic device, comprising: one or more processors; A memory for storing one or more computer programs, which, when executed by the one or more processors, enables the electronic device to implement the video decoding method as described in any one of claims 1 to 11, or to implement the video encoding method as described in claim 12.

17. A computer program product, comprising computer instructions, wherein the computer instructions are stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer instructions from the computer-readable storage medium, so that the electronic device performs the video decoding method as described in any one of claims 1 to 11, or implements the video encoding method as described in claim 12.

18. A non-volatile computer-readable storage medium having a bit stream stored thereon, wherein the bit stream is generated by the video encoding method according to any one of claims 1 to 11.