Video encoding method and apparatus, video decoding method and apparatus, device, and medium
By using mode indexes to indicate the location distribution of non-zero transform coefficients within transform block sub-blocks in video encoding and decoding, the problem of low efficiency in transform block encoding and decoding is solved, achieving a more efficient and reliable encoding and decoding process.
Patent Information
- Application Number
- PCT/CN2025/108368
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-30
- Filing Date
- 2025-07-14
- Publication Date
- 2026-02-05
AI Technical Summary
In the existing technology, the encoding and decoding efficiency of transform blocks is low and the encoding and decoding reliability is low. In particular, when dealing with the positions of non-zero transform coefficients inside the transform block, there is repetitive scanning, which leads to low efficiency.
By obtaining the encoding information of the target transform block of the current video frame, the distribution pattern of each sub-block is determined, and the position distribution of non-zero transform coefficients is indicated by the mode index. Only the mode index is encoded and decoded, reducing the encoding and decoding operations of each non-zero transform coefficient position within the sub-block.
It improves the encoding and decoding efficiency of transform blocks, enhances the reliability of video encoding and decoding, and reduces the cost of encoding and decoding.
Smart Images

Figure CN2025108368_05022026_PF_FP_ABST
Abstract
Description
Video encoding and decoding methods, devices, equipment and media
[0001] This application claims priority to Chinese Patent Application No. 202411038485.3, filed on July 30, 2024, entitled “Video Coding and Decoding Method, Apparatus, Device and Medium”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of video processing technology, specifically to a video encoding / decoding method, a video encoding / decoding device, an electronic device, and a computer-readable medium. Background Technology
[0003] To accommodate large-scale video data transmission, the original video data is typically encoded at the data sending end to form a compressed data stream. After the data stream is transmitted to the data receiving end, it is then decoded and restored to obtain the predicted and reconstructed video data.
[0004] Technical content
[0005] This application provides a video encoding / decoding method, a video encoding / decoding device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the encoding / decoding efficiency of transform blocks and achieve high reliability in video encoding / decoding.
[0006] This application provides a video decoding method executed by at least one processor. The method includes: acquiring encoding information of a target transform block of a current video frame, the encoding information including a mode index corresponding to each sub-block in the target transform block, the mode index corresponding to each sub-block indicating the distribution mode of the sub-block, and the distribution mode of each sub-block indicating the positional distribution of non-zero transform coefficients within the sub-block; decoding the mode index corresponding to each sub-block from the encoding information; determining the distribution mode corresponding to the sub-block based on the mode index corresponding to each sub-block; and obtaining the positions of the non-zero transform coefficients within the sub-block based on the distribution mode corresponding to each sub-block.
[0007] This application also provides a video encoding method executed by at least one processor. The method includes: for a target transform block comprising multiple sub-blocks in a current video frame, determining a distribution pattern for each sub-block, wherein the distribution pattern of each sub-block indicates the positional distribution of non-zero transform coefficients within the sub-block; determining a mode index corresponding to the distribution pattern of each sub-block; encoding the mode index corresponding to each sub-block; and generating encoding information of the target transform block based on the encoding result of the mode index corresponding to each sub-block.
[0008] This application embodiment also provides a video decoding apparatus, which includes: an acquisition module configured to acquire encoding information of a target transform block of a current video frame, the encoding information including a mode index corresponding to each sub-block in the target transform block, the mode index corresponding to each sub-block being used to indicate the distribution mode of the sub-block, and the distribution mode of each sub-block being used to indicate the positional distribution of non-zero transform coefficients within the sub-block; and a decoding module configured to decode the mode index corresponding to each sub-block from the encoding information; determine the distribution mode corresponding to the sub-block based on the mode index corresponding to each sub-block; and obtain the position of the non-zero transform coefficients within the sub-block based on the distribution mode corresponding to each sub-block.
[0009] This application embodiment also provides a video encoding apparatus, which includes: an acquisition module configured to determine the distribution pattern of each sub-block for a target transform block comprising multiple sub-blocks in the current video frame, wherein the distribution pattern of each sub-block indicates the positional distribution of non-zero transform coefficients within the sub-block; a determination module configured to determine a mode index corresponding to the distribution pattern of each sub-block; and an encoding module configured to encode the mode index corresponding to each sub-block; and generate encoding information of the target transform block based on the encoding result of the mode index corresponding to each sub-block.
[0010] This application also provides an electronic device, including: one or more processors; and a memory for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the video encoding and decoding method described above.
[0011] This application also provides a computer-readable storage medium storing computer-readable instructions and a bit stream thereon. When the computer-readable instructions are executed by a computer's processor, the computer performs the video encoding method described above to generate the bit stream.
[0012] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the video encoding and decoding method described above.
[0013] This application also provides a method for storing a bitstream, comprising: performing the video encoding method described above to generate a bitstream; and storing the bitstream.
[0014] This application also provides a method for transmitting a bit stream, including performing the video encoding method described above to generate a bit stream; and transmitting the bit stream.
[0015] Brief description of the attached figures
[0016] Figure 1 shows a schematic diagram of the system architecture to which the technical solutions of the embodiments of this application can be applied;
[0017] Figure 2 illustrates the placement of the video encoding device and the video decoding device in a streaming environment in some embodiments of this application;
[0018] Figure 3 shows a basic flowchart of the encoding process performed by the video encoder in some embodiments of this application;
[0019] Figure 4A shows a schematic diagram of the transformation block in some embodiments of this application;
[0020] Figure 4B shows a schematic diagram of transform block scanning in some embodiments of this application;
[0021] Figure 5 shows a flowchart of a video decoding method in some embodiments of this application;
[0022] Figure 6 shows a schematic diagram of various encoding modes in some embodiments of this application;
[0023] Figure 7 shows a flowchart of a video encoding method in some embodiments of this application;
[0024] Figure 8 shows a block diagram of a video decoding apparatus in some embodiments of this application;
[0025] Figure 9 shows a block diagram of a video encoding apparatus in some embodiments of this application;
[0026] Figure 10 shows a schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application. Detailed Implementation
[0027] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0028] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0029] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0030] In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0031] The terms "first," "second," "third," and "fourth," etc., used in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0032] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0033] To facilitate understanding of the technical solutions proposed in the embodiments of this application, the video encoding and decoding process will be introduced first.
[0034] Video encoding generally refers to the processing of a sequence of images that form a video or video sequence. In the field of video encoding, the terms "picture," "frame," or "image" can be used synonymously. The video encoding used in the embodiments of this application refers to video encoding or video decoding. Video encoding is performed on the source side and typically involves processing (e.g., by compression) the raw video image to reduce the amount of data required to represent the video image, thereby enabling more efficient storage and / or transmission. Video decoding is performed on the destination side and typically involves inverse processing relative to the encoder to reconstruct the video image. The "encoding" of video frames involved in the embodiments should be understood as involving the "encoding" or "decoding" of a sequence of video images. The combination of encoding and decoding portions is also referred to as encoding and decoding (encoding and decoding).
[0035] Each image in a video sequence is typically segmented into a set of non-overlapping blocks, which are usually encoded at the block level. In other words, the encoder typically processes, i.e., encodes the video at the block (also called image block or video block) level, for example, by generating prediction blocks through spatial (intra-image) and temporal (inter-image) predictions, subtracting the prediction blocks from the current block (the block currently being processed or to be processed) to obtain residual blocks, transforming and quantizing the residual blocks in the transform domain to reduce the amount of data to be transmitted (compressed), while the decoder applies the inverse processing relative to the encoder to the encoded or compressed blocks to reconstruct the current block for representation. Additionally, the encoder replicates the decoder processing loop, such that the encoder and decoder generate the same predictions (e.g., intra-frame and inter-frame predictions) and / or reconstructions for processing, i.e., encoding subsequent blocks.
[0036] The term "block" refers to a portion of an image or frame. In embodiments of this application, the current block refers to the block currently being processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.
[0037] Figure 1 illustrates a system architecture diagram of the technical solutions applicable to embodiments of this application. As shown in Figure 1, system architecture 100 includes multiple terminal devices that can communicate with each other via, for example, a network 150. For example, system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via network 150. In the embodiment of Figure 1, the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.
[0038] For example, the first terminal device 110 can encode video data (e.g., a video image stream captured by the terminal device 110) to transmit it to the second terminal device 120 via the network 150. The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device 120 can receive the encoded video data from the network 150, decode the encoded video data to recover the video data, and display video images based on the recovered video data.
[0039] In some embodiments of this application, system architecture 100 may include a third terminal device 130 and a fourth terminal device 140 that perform bidirectional transmission of encoded video data, such as during a video conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode video data (e.g., a video image stream captured by the terminal device) for transmission over network 150 to the other terminal device. Each of the third terminal device 130 and the fourth terminal device 140 may also receive encoded video data transmitted by the other terminal device, decode the encoded video data to recover the video data, and display the video images on an accessible display device based on the recovered video data.
[0040] In the embodiment of FIG1, the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140 may be servers, personal computers, and smartphones, but the principles disclosed herein are not limited thereto. The embodiments disclosed herein are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 150 refers to any number of networks transmitting encoded video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140, including, for example, wired and / or wireless communication networks. Network 150 may exchange data in circuit-switched and / or packet-switched channels. This network may include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of network 150 may be irrelevant to the operation of this application.
[0041] Figure 2 illustrates the placement of the video encoding and decoding devices in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television (TV), and storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0042] The streaming system may include an acquisition subsystem 213, which may include a video source 201 such as a digital camera, which creates an uncompressed video image stream 202. In an embodiment, the video image stream 202 includes samples captured by a digital camera. The video image stream 202 is depicted as a thick line to emphasize the high data volume of the video image stream compared to encoded video data 204 (or encoded video bitstream 204). The video image stream 202 may be processed by an electronic device 220, which includes a video encoding device 203 coupled to the video source 201. The video encoding device 203 may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data 204 (or encoded video bitstream 204) is depicted as a thin line to emphasize the lower data volume of the encoded video data 204 (or encoded video bitstream 204), which may be stored on a streaming server 205 for future use. One or more streaming client subsystems, such as client subsystems 206 and 208 in FIG. 2, can access streaming server 205 to retrieve copies 207 and 209 of encoded video data 204. Client subsystem 206 may include, for example, a video decoding device 210 in electronic device 230. Video decoding device 210 decodes the incoming copy 207 of the encoded video data and produces an output video picture stream 211 that can be displayed on display 212 (e.g., a screen) or another presentation device. In some streaming systems, the encoded video data 204, video data 207, and video data 209 (e.g., video stream) may be encoded according to certain video encoding / compression standards.
[0043] It should be noted that electronic devices 220 and 230 may include other components not shown in the figures. For example, electronic device 220 may include a video decoding device, and electronic device 230 may also include a video encoding device.
[0044] In some embodiments of this application, taking the international video coding standards HEVC (High Efficiency Video Coding, H.265) and VVC (Versatile Video Coding, H.266), and the Chinese national video coding standard AVS (Audio Video Coding Standard) as examples, after an input video frame image is received, the video frame image is divided into several non-overlapping processing units according to a block size. Each processing unit will perform a similar compression operation. This processing unit is called a CTU (Coding Tree Unit) or LCU (Largest Coding Unit). The CTU can be further subdivided into one or more basic coding units (CUs), and the CU is the most basic element in a coding process.
[0045] Figure 3 shows a basic flowchart of the encoding process performed by the video encoder, which is illustrated using intra-frame prediction as an example.
[0046] Wherein, the original image signal s k [x,y] and the predicted image signal Perform the difference operation to obtain the residual signal u. k [x,y], residual signal u k After transformation and quantization, [x,y] is obtained as quantization coefficients. These coefficients are then used to obtain the encoded bitstream through entropy encoding, and to obtain the reconstructed residual signal u′ through inverse quantization and inverse transform. k [x,y], predict image signal With the reconstructed residual signal u′ k [x,y] superimposed to generate image signals Image signal On one hand, the signal is input to the intra-frame mode decision module and the intra-frame prediction module for intra-frame prediction processing; on the other hand, the reconstructed image signal s′ is output through loop filtering. k [x,y], reconstruct the image signal s′ k [x,y] can be used as a reference image for the next frame for motion estimation and motion compensation prediction. Then, based on the motion compensation prediction result s′ r [x+m x ,y+m y ] and intra-frame prediction results Obtain the predicted image signal for the next frame. And continue repeating the above process until the coding is complete.
[0047] The encoding operations for each CU involved in the above video encoding process are detailed below.
[0048] Predictive coding includes intra-frame prediction and inter-frame prediction. The original video signal is predicted from a selected reconstructed video signal to obtain a residual video signal. The encoder needs to determine which predictive coding mode to choose for the current CU and inform the decoder. Intra-frame prediction refers to the predicted signal coming from a region within the same image that has already been encoded and reconstructed; inter-frame prediction refers to the predicted signal coming from another encoded image (called a reference image) that is different from the current image.
[0049] Transform and Quantization: After the residual video signal undergoes transformation operations such as Discrete Fourier Transform (DFT) and Discrete Cosine Transform (DCT), the signal is transformed into the transform domain, and these are called transform coefficients. The transform coefficients are then subjected to lossy quantization, losing some information to make the quantized signal more suitable for compression. In some video coding standards, there may be more than one transform method to choose from; therefore, the encoder needs to select one of the transform methods for the current CU and inform the decoder. The fineness of quantization is usually determined by the quantization parameter (QP). A larger QP value means that coefficients with a wider range of values will be quantized into the same output, which usually results in greater distortion and a lower bit rate. Conversely, a smaller QP value means that coefficients with a smaller range of values will be quantized into the same output, which usually results in less distortion and a higher bit rate.
[0050] Entropy coding, or statistical coding, involves statistically compressing the quantized transform-domain signal based on the frequency of each value, ultimately outputting a binary (0 or 1) compressed bitstream. Simultaneously, other information generated during encoding, such as the selected coding mode and motion vector data, also requires entropy coding to reduce the bit rate. Statistical coding is a lossless coding method that effectively reduces the bit rate required to represent the same signal. Common statistical coding methods include Variable Length Coding (VLC) and Content Adaptive Binary Arithmetic Coding (CABAC).
[0051] Context-based adaptive binary arithmetic coding mainly involves three steps: binarization, context modeling, and binary arithmetic coding. After binarizing the input syntax elements, the binary data can be encoded using either a regular coding mode or a bypass coding mode. The bypass coding mode does not require assigning a specific probability model to each binary bit; the input binary bit bin value is directly encoded using a simple bypass encoder to speed up the entire encoding and decoding process. Generally, different syntax elements are not completely independent, and the same syntax elements themselves also have a certain degree of memory. Therefore, according to conditional entropy theory, using other encoded syntax elements for conditional coding can further improve coding performance compared to independent coding or memoryless coding. This encoded symbol information used as conditions is called the context. In the regular coding mode, the binary bits of the syntax elements sequentially enter the context modeler. The encoder assigns an appropriate probability model to each input binary bit based on the values of previously encoded syntax elements or binary bits; this process is called context modeling. The context model corresponding to a grammatical element can be located using the context index increment (ctxIdxInc) and the context index start (ctxIdxStart). After the bin value and the assigned probability model are fed into the binary arithmetic encoder for encoding, the context model needs to be updated based on the bin value, which is the adaptive process in encoding.
[0052] Loop Filtering: The transformed and quantized signal undergoes inverse quantization, inverse transform, and prediction compensation to obtain a reconstructed image. Due to the effects of quantization, the reconstructed image differs from the original image in some aspects, resulting in distortion. Therefore, filtering operations can be performed on the reconstructed image, such as deblocking filters (DB), sample adaptive offset (SAO), or adaptive loop filters (ALF), to effectively reduce the distortion caused by quantization. Since these filtered reconstructed images will serve as a reference for subsequent coded images to predict future image signals, the aforementioned filtering operations are also called loop filtering, i.e., filtering operations within the coding loop.
[0053] Based on the above encoding process, at the decoding end, for each CU, after acquiring the compressed bitstream (i.e., bitstream), entropy decoding is performed to obtain various mode information and quantization coefficients. Then, the quantization coefficients undergo inverse quantization and inverse transform processing to obtain the residual signal. On the other hand, based on the known encoding mode information, the prediction signal corresponding to that CU can be obtained. Then, the residual signal and the prediction signal are added together to obtain the reconstructed signal. The reconstructed signal then undergoes loop filtering and other operations to generate the final output signal.
[0054] Understandably, in the transformation and quantization process of video encoding and decoding, the encoding of transform coefficients is involved. Related technologies involve scanning / traversing each position within each sub-block of the transform block and encoding the non-zero transform coefficients scanned / traversed. Correspondingly, the non-zero transform coefficients are decoded. However, this does not take advantage of the repetitive distribution of non-zero transform coefficients within the transform block. During scanning, every position must be traversed, which reduces the encoding and decoding efficiency of the transform block and results in low reliability of video encoding and decoding.
[0055] After transformation and quantization, the residual block yields transform domain coefficients. Typically, the transform block contains non-zero coefficients. Each non-zero transform block coefficient needs to have its sign (+ or -). The first step in coefficient encoding is to identify which positions in the transform block are 0 and which are non-zero coefficients. Since the quantized transform coefficients often exhibit a large distribution of 0s, and these 0s are frequently located in the lower right corner of the transform block, a coefficient scanning method can be used to transform the 2D coefficient matrix into a 1D coefficient sequence. Figure 4A (left) shows an example distribution of non-zero coefficients in a quantized transform coefficient block.
[0056] In the left image of Figure 4A, after diagonal scanning (as shown on the left side of Figure 4B), the 1D coefficient sequence is: 672, -36, 33, 0, -24, 60, 14, 0, 0, 0, 18, 34, -16, 19, -24, 0, 22, 22, -24, 26, 0, 0, 0, 0, 0, 0, -51, 0, ..., 0. By recording the position of the last non-zero coefficient in the entire sequence, all the zero coefficients after it can be ignored. For example, in the example above, -51 is the last non-zero coefficient in the scanning order, and there are 28 coefficients before it in the scanning order. Therefore, after recording (and transmitting) the last 28, the 8x8-28=36 zero coefficients after it do not need to be encoded, and the decoder can synchronously obtain consistent information. Next, each non-zero coefficient in the sequence is replaced with a 1, and the above sequence becomes a binary position sequence of 0s and 1s, which is then encoded separately. The final step is to encode the magnitude and sign (+-) of each non-zero coefficient according to their sequential position. As can be seen from the above process, the entire coefficient encoding process consists of several steps, including coefficient scanning, non-zero position encoding, coefficient absolute value encoding, and sign encoding. Coefficient scanning involves determining the position information of all non-zero coefficients within the transform block.
[0057] As shown in the left part of Figure 4A, the transform block includes multiple non-zero transform coefficients and multiple zero transform coefficients. This transform block is then binarized by replacing the non-zero transform coefficients with 1, resulting in the binarized transform block shown in the right part of Figure 4A. This binarized transform block can then be divided into multiple sub-blocks. Specifically, for the 16 sub-blocks shown in the left part of Figure 4B, scanning can be performed according to the position of each sub-block within the transform block (e.g., diagonal scanning), and scanning can also be performed within each sub-block (e.g., diagonal scanning) to obtain a binarized sequence. The right part of Figure 4B shows the scanning process for the internal position of the upper right sub-block within the transform block.
[0058] Therefore, in order to improve the encoding and decoding efficiency of transform blocks and ensure the reliability of video encoding and decoding, this application provides a video encoding and decoding scheme. Specifically, the encoding end obtains the distribution pattern corresponding to each sub-block for a target transform block containing multiple sub-blocks in the current video frame. The distribution pattern of each sub-block is used to characterize the position distribution of non-zero transform coefficients within the sub-block. Then, based on the distribution pattern corresponding to each sub-block, the encoding end determines the mode index corresponding to the distribution pattern of the sub-block and encodes the mode index. Then, based on the encoding result of the mode index corresponding to each sub-block, the encoding information of the target transform block is generated and sent to the decoding end. Correspondingly, the decoding end receives the encoding information of the target transform block, decodes the mode index corresponding to each sub-block from the encoding information, then determines the distribution pattern of non-zero transform coefficients in each sub-block based on the mode index corresponding to each sub-block, and then obtains the position of non-zero transform coefficients in the target transform block based on the distribution pattern corresponding to each sub-block.
[0059] In this way, by characterizing the distribution pattern of the non-zero transform coefficients within a sub-block, only the pattern index corresponding to the distribution pattern needs to be encoded, decoded, and transmitted. Compared with related technologies that encode, decode, and transmit the non-zero transform coefficient positions one by one within a sub-block, this greatly reduces the cost of encoding and decoding the transform coefficient positions, improves the encoding and decoding efficiency of the transform block, and ensures high reliability of video encoding and decoding.
[0060] It should be noted that in the specific implementation of this application, user-related data is involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0061] The following details the various implementation details of the technical solutions in the embodiments of this application:
[0062] Figure 5 shows a flowchart of a video decoding method in some embodiments of this application. This video decoding method can be executed by at least one processor in a terminal device or server that sends or receives video encoded data. This application embodiment uses a method executed by at least one processor in a terminal device as an example for illustration. The terminal device can be, for example, the video decoding device 210 or the video encoding device 203 shown in Figure 2. As shown in Figure 5, the video decoding method includes at least steps S510 to S540, which are described in detail below:
[0063] S510, obtain the encoding information of the target transform block of the current video frame. The encoding information includes the mode index corresponding to each sub-block in the target transform block. The mode index corresponding to each sub-block is used to indicate the distribution mode of the sub-block. The distribution mode of each sub-block is used to indicate the position distribution of the non-zero transform coefficients in the sub-block.
[0064] In some embodiments, the sub-block size can be determined in two ways: by using a preset default size; or by dynamically specifying it based on syntax elements. For the latter, the codec supports a flexible configuration mechanism. For example, it can be selected from a predefined set of sizes (such as 4x4 or 2x2 sub-blocks). The configuration granularity can support multiple levels, such as based on the current transform block, or based on the coding tree unit (CTU), or based on images / strips / fragments, etc. Alternatively, the width and height of the sub-block can be specified separately, where the length of the sub-block is a proper divisor of the transform block length; and the width of the sub-block is a proper divisor of the transform block width.
[0065] In the embodiments of this application, the target transform block refers to the transform block to be decoded in the current video frame.
[0066] In this embodiment of the application, the pattern index of each sub-block is used to indicate the distribution pattern of the non-zero transform coefficients within the sub-block, wherein each pattern index is used to uniquely identify a distribution pattern.
[0067] The distribution patterns are illustrated below with an example, as shown in Figure 6. Taking a 2×2 sub-block as an example, there are 16 different distribution patterns for its transform coefficients, that is, 16 possible combinations of zero transform coefficients / non-zero transform coefficients. It can be understood that ptn = 0 represents a distribution pattern with all zero transform coefficients, meaning all positions within the sub-block have all zero transform coefficients, while ptn = 1-15 represents a distribution pattern with non-zero transform coefficients within the sub-block, meaning there are 15 possible distribution patterns with non-zero transform coefficients within the sub-block. Correspondingly, each distribution pattern has a pattern index.
[0068] Table 1 shows an example of a mapping table between distribution patterns and pattern indices in some embodiments.
[0069] Table 1
[0070] For example, continuing from the previous example, if the distribution pattern corresponding to the sub-block is ptn=2, then as shown in Table 1, the corresponding pattern index is index=2. The encoding information includes the result of encoding index=2.
[0071] It should be noted that in practical applications, the mapping table between the distribution pattern and the pattern index can be flexibly adjusted according to the specific application scenario.
[0072] In this embodiment, the encoding information includes the encoding results of the pattern indices corresponding to all sub-blocks in the target transform block; for example, if the target transform block includes 16 sub-blocks, then the encoding information includes the encoding results of the pattern indices corresponding to these 16 sub-blocks respectively. The process of generating the encoding information is described below.
[0073] S520, decode the encoding information to obtain the mode index corresponding to each sub-block.
[0074] In this embodiment of the application, the encoded information can be decoded to obtain the pattern index corresponding to each sub-block.
[0075] In some embodiments, the decoder further obtains the scan order of the target transform block, wherein the scan order is used to determine the sub-block processing sequence, and the mode indices corresponding to each sub-block form an ordered sequence according to the scan order; the decoder decodes the mode index corresponding to each sub-block sequentially according to the scan order.
[0076] In some embodiments of this application, the scanning order may include: horizontal row-by-row scanning, vertical column-by-column scanning, diagonal scanning, etc. The process of obtaining the scanning order in S510 may include at least two methods:
[0077] Method 1: Derive the scanning order of the target transform block based on the encoding mode of the current block.
[0078] By deriving the scan order of the target transform block based on the encoding mode of the current block, it is possible to avoid sending flag bits to indicate the scan order, thereby saving signaling overhead.
[0079] Method 2: Obtain the identifier indicating the scanning order from the bitstream.
[0080] That is, the decoding end receives the bitstream sent by the encoding end in real time, obtains the identifier indicating the scanning order from the bitstream, and obtains the scanning order in real time based on the identifier.
[0081] This improves the flexibility and accuracy of transform block decoding by receiving the scan sequence in real time.
[0082] In some embodiments of this application, the process of decoding the pattern index corresponding to each sub-block from the encoded information in S520 may include:
[0083] Obtain a mapping table between the pattern index and its encoding result. The mapping table includes multiple pattern indices and the encoding results corresponding to the multiple pattern indices. The encoding length of the encoding results corresponding to different pattern indices is different.
[0084] Based on the mapping table, the encoding result corresponding to each sub-block is converted into the corresponding pattern index.
[0085] Table 2 shows a mapping table between the pattern index and its encoding result in some embodiments of this application.
[0086] Table 2
[0087] It should be noted that, in practical applications, the mapping table can be flexibly adjusted according to specific application scenarios.
[0088] In this way, the pattern index corresponding to each sub-block can be obtained easily and accurately based on the mapping table between the pattern index and its encoding result.
[0089] In some embodiments of this application, the mapping table is implemented in, but is not limited to, the following two ways:
[0090] Method 1: The mapping table is a static mapping table pre-agreed between the encoding and decoding ends.
[0091] For example, the mapping table is pre-stored in the first storage area of the decoding end, and the stored mapping table is agreed upon by the encoding end and the decoding end; therefore, the mapping table can be retrieved from the first storage area of the decoding end when needed.
[0092] By pre-defining the mapping table between the mode index and its encoding result, the real-time computation overhead can be reduced, further improving the decoding efficiency of the transform block.
[0093] Method 2: The mapping table is a dynamic mapping table constructed by the decoding end according to the same strategy as the encoding end.
[0094] For example, the decoding end generates an initial mapping table according to the same strategy as the encoding end, and then updates the mapping table for each symbol decoded.
[0095] By dynamically constructing an adaptive mapping table, the adaptability to dynamic coding is enhanced, thereby improving the encoding and decoding efficiency of transform blocks.
[0096] S530: Determine the distribution pattern corresponding to each sub-block based on the pattern index corresponding to each sub-block.
[0097] In this embodiment, the pattern index corresponding to each sub-block is obtained, and then the distribution pattern corresponding to the non-zero transformation coefficients in each sub-block can be determined according to the pattern index.
[0098] In some embodiments of this application, the process of determining the distribution pattern corresponding to each sub-block based on the pattern index corresponding to each sub-block in S530 may include:
[0099] Obtain the mapping relationship between distribution patterns and pattern indices;
[0100] Based on the pattern index corresponding to each sub-block, the mapping relationship is queried to obtain the distribution pattern corresponding to each sub-block.
[0101] In some embodiments, the mapping relationship can be a historical information list of distribution patterns; the historical information list records multiple distribution patterns that have appeared historically in encoding and decoding order. The value of the pattern index of each sub-block indicates the position of the distribution pattern of that sub-block in the historical information list.
[0102] In some embodiments, for each sub-block, the distribution pattern at that position in the historical information list is determined based on the position indicated by the pattern index of the sub-block; and the distribution pattern is determined as the distribution pattern corresponding to the sub-block.
[0103] In some embodiments, after the target transform block has been decoded, the historical information list can be updated according to the distribution pattern of each sub-block in the target transform block.
[0104] In some embodiments, updating the historical information list includes: reconstructing the historical information list according to the frequency of occurrence of the distribution patterns of each sub-block in the target transformation block, wherein the distribution patterns in the historical information list are arranged in order of frequency of occurrence from high to low or from low to high.
[0105] After the target transform block is decoded, the frequency of each distribution pattern can be counted, and the historical information list can be reconstructed according to the probability. For high-frequency distribution patterns, they can be placed at the beginning of the historical information list to reduce index query latency. For low-frequency distribution patterns, their positions can be moved to the end or eliminated to free up storage resources.
[0106] For example, based on the frequency of each sub-block distribution pattern, each sub-block distribution pattern can be placed in a different position in the list. For sub-block distribution patterns that occur frequently, they can be placed in a position earlier (or later). This helps to reduce the encoding cost when encoding the flag bits of the sub-block positions in the list.
[0107] In some embodiments, updating the historical information list includes: removing the earliest distribution pattern from the historical information list according to the first-in-first-out principle, and adding the most recently appearing distribution pattern to the historical new list.
[0108] For example, based on the first-in-first-out principle, the oldest sub-block distribution pattern can be removed from the list, and the most recently appearing sub-block distribution pattern can be added to the list.
[0109] In some embodiments, the mapping relationship can be a mapping table between distribution patterns and pattern indices. The mapping table includes multiple distribution patterns and pattern indices corresponding to each distribution pattern, as shown in Table 1.
[0110] For example, continuing from the previous example, if the pattern index corresponding to a certain sub-block is index=2, then by querying Table 1, we can obtain the distribution pattern ptn=2 corresponding to that sub-block.
[0111] In this way, the distribution pattern corresponding to each sub-block can be obtained easily and accurately through the mapping table between the distribution pattern and the pattern index.
[0112] In some embodiments of this application, the process of obtaining the mapping table may include at least two methods:
[0113] Method 1: Obtain the mapping table from the second storage area of the decoding end. The stored mapping table is agreed upon by the decoding end and the encoding end.
[0114] That is, the mapping table is pre-stored in the second storage area of the decoding end, and the stored mapping table is agreed upon by the encoding end and the decoding end; therefore, the mapping table can be obtained from the second storage area of the decoding end when needed.
[0115] By pre-defining the mapping table between the distribution pattern and the pattern index, the decoding efficiency of the transform block is further improved.
[0116] Method 2: Receive the bitstream sent by the encoding end and obtain the mapping table between the distribution mode and the mode index from the bitstream.
[0117] That is, the decoding end receives the bitstream sent by the encoding end in real time and obtains the mapping table from the bitstream, thereby obtaining the mapping table between the distribution mode and the mode index in real time.
[0118] This improves the flexibility and accuracy of transform block decoding by receiving the mapping table between the distribution mode and the mode index in real time.
[0119] In some embodiments of this application, after the target transform block has been decoded, a step of updating the mapping table between the distribution pattern and the pattern index can be further performed.
[0120] For example, after the target transform block is decoded, the frequency of each distribution pattern can be counted. The mapping relationship between each distribution pattern and the pattern index can be adjusted according to the frequency of occurrence. For high-frequency distribution patterns, index values closer to the beginning of the mapping table are prioritized to improve query efficiency. The updated mapping table will be used for decoding subsequent transform blocks.
[0121] Accordingly, after the mapping table is updated, for subsequent transformation blocks, the updated mapping table is queried according to the pattern index corresponding to each sub-block to obtain the distribution pattern corresponding to each sub-block.
[0122] By updating the mapping table between the distribution pattern and the pattern index, storage overhead and access performance are effectively balanced.
[0123] S540, determine the position of the non-zero transformation coefficients within each sub-block according to the distribution pattern corresponding to each sub-block.
[0124] In this embodiment, the distribution pattern corresponding to each sub-block is obtained, and then the position of the non-zero transformation coefficient in the sub-block can be obtained according to the distribution pattern corresponding to each sub-block.
[0125] In this embodiment, the decoding end directly reconstructs the distribution pattern of non-zero transform coefficients of each sub-block by parsing the pattern index in the bitstream (such as the position index of the distribution pattern in the historical information list). Compared with the traditional scheme that requires parsing the position information of coefficients bit by bit, this design reduces the cost of transform coefficient position encoding and improves encoding and decoding efficiency through the index mapping mechanism of a finite pattern set.
[0126] Figure 7 shows a flowchart of a video encoding method in some embodiments of this application. This video encoding method can be executed by at least one processor of a terminal device or server that sends video encoded data. This application embodiment uses a method executed by a terminal device as an example for illustration. The terminal device can be, for example, the video encoding apparatus 203 shown in Figure 2. As shown in Figure 7, the video encoding method includes at least steps S710 to S740, which are described in detail below:
[0127] S710, for a target transform block containing multiple sub-blocks in the current video frame, determine the distribution pattern of each sub-block, wherein the distribution pattern of each sub-block is used to indicate the positional distribution of non-zero transform coefficients within the sub-block.
[0128] In some embodiments of this application, the process of determining the distribution pattern corresponding to each sub-block for a target transform block comprising multiple sub-blocks in step S710 may include:
[0129] The target transformation block is divided into multiple sub-blocks;
[0130] For each sub-block, the distribution pattern of the sub-block is determined based on the distribution of zero and non-zero transform coefficients at various locations within the sub-block.
[0131] For example, in some embodiments, the target transform block is first divided into multiple sub-blocks, and then the multiple sub-blocks are scanned according to the target scanning mode. The transform coefficients of the scanned sub-blocks are binarized to obtain the distribution mode corresponding to the scanned sub-blocks. This process continues until all multiple sub-blocks are scanned, thereby obtaining the distribution mode corresponding to each sub-block.
[0132] In this way, the distribution pattern of each sub-block can be obtained easily and accurately, thus providing strong support for the generation of encoded information.
[0133] In some embodiments of this application, the process of obtaining the target scanning pattern may include at least two methods:
[0134] Method 1: Determine the target scanning mode based on the encoding mode.
[0135] For example, if the current block's encoding mode is intra-frame prediction mode, horizontal scanning, vertical scanning, or diagonal scanning can be used; if the current block's encoding mode is inter-frame prediction mode, diagonal scanning can be used. In this way, the decoder can determine the scanning order based on the current block's encoding mode without needing to indicate it in the bitstream.
[0136] Method 2: Indicate the target scanning mode in the bitstream.
[0137] This improves the coding flexibility and accuracy of transform blocks by indicating the target scanning mode in the bitstream.
[0138] S720, determine the pattern index corresponding to the distribution pattern according to the distribution pattern corresponding to each sub-block.
[0139] In this embodiment of the application, the distribution pattern corresponding to each sub-block is obtained, and then the pattern index corresponding to each distribution pattern can be determined according to the distribution pattern corresponding to each sub-block.
[0140] In some embodiments of this application, the process of determining the pattern index corresponding to the distribution pattern according to the distribution pattern corresponding to each sub-block in S720 may include:
[0141] Create a mapping relationship between distribution patterns and pattern indexes;
[0142] Based on the mapping relationship, determine the pattern index corresponding to the distribution model of each sub-block.
[0143] In some embodiments, the mapping relationship can be a historical information list. The historical information list records historically occurring distribution patterns in the order of encoding and decoding.
[0144] The step of determining the pattern index corresponding to the distribution model of each sub-block based on the mapping relationship includes: determining the pattern index corresponding to the distribution model based on the position of the distribution model of the sub-block in the historical information list. For example, the position index is used to determine the pattern index corresponding to the distribution model.
[0145] In some embodiments, after each transform block is decoded, the historical information list is updated according to the distribution pattern of the sub-blocks within the current transform block.
[0146] In some embodiments, the update includes:
[0147] According to the first-in, first-out principle, the earliest distribution pattern is removed from the historical information list, and the most recently appearing distribution pattern is added to the historical newest list; or
[0148] Based on the frequency of occurrence of the distribution patterns of each sub-block in the target transformation block, the historical information list is reconstructed, wherein in the historical information list, each distribution pattern is arranged in order of frequency of occurrence from high to low or from low to high.
[0149] In some embodiments of this application, the mapping relationship may be a mapping table between distribution patterns and pattern indices, wherein the mapping table includes multiple distribution patterns and pattern indices corresponding to each distribution pattern. Accordingly, determining the pattern index corresponding to the distribution model of each sub-block based on the mapping relationship may include:
[0150] Obtain the mapping table between distribution patterns and pattern indices;
[0151] Based on the distribution pattern corresponding to each sub-block, the mapping table is queried to obtain the pattern index corresponding to each distribution pattern.
[0152] The second storage area of the encoding end pre-stores a mapping table between the distribution mode and the mode index. This mapping table is agreed upon by the encoding end and the decoding end. Therefore, the mapping table can be retrieved from the second storage area of the encoding end when needed.
[0153] In some embodiments, the mapping table between distribution patterns and pattern indexes includes multiple distribution patterns and pattern indexes corresponding to each distribution pattern, as shown in Table 1.
[0154] In this way, the mapping table can be used to easily and accurately obtain the pattern index corresponding to each distribution pattern.
[0155] In some embodiments of this application, after encoding the target transform block is completed, the following may also be included:
[0156] Count the frequency of occurrence of each distribution pattern;
[0157] Adjust the mapping relationship between each distribution pattern and the pattern index according to the frequency of each distribution pattern.
[0158] For example, for high-frequency distribution patterns, index values closer to the beginning of the mapping table are prioritized to improve query efficiency. The updated mapping table will then be used for encoding subsequent transform blocks.
[0159] In some embodiments, the correspondence between distribution patterns and their pattern indices is shown in Table 1, and the order of the generated pattern indices is positively correlated with the order of the distribution patterns.
[0160] In some embodiments, the correspondence between distribution patterns and their pattern indices is shown in Table 3. The order of the generated pattern indices is negatively correlated with the order of the distribution patterns.
[0161] Table 3
[0162] This method sorts multiple distribution patterns based on their respective occurrence counts, thereby generating a pattern index for each distribution pattern. This improves the accuracy of pattern index generation, and the order of multiple pattern indexes is positively or negatively correlated with the order of multiple distribution patterns, offering high flexibility.
[0163] S730 encodes the pattern index corresponding to each sub-block.
[0164] In this embodiment, the pattern index corresponding to each sub-block is obtained, and then the pattern index corresponding to each sub-block can be encoded.
[0165] In some embodiments of this application, encoding the pattern index corresponding to each sub-block in S730 may include:
[0166] The pattern index of each sub-block is encoded according to the variable-length encoding strategy.
[0167] In some embodiments, the encoding lengths corresponding to different mode indices are different, for example, please refer to Table 2.
[0168] In this way, using a variable-length encoding strategy to encode the pattern index improves the encoding efficiency and flexibility of the pattern index.
[0169] In some embodiments of this application, the process of encoding multiple pattern indices according to a variable-length encoding strategy to generate a pattern index encoding result corresponding to each pattern index may include:
[0170] Obtain the number of occurrences of each distribution pattern in the historical transformation block;
[0171] The encoding length of each pattern index is determined based on the occurrence count of each distribution pattern, and the occurrence count is negatively correlated with the encoding length.
[0172] Each pattern index is encoded according to its encoding length.
[0173] In some embodiments, the number of occurrences of a coding pattern is negatively correlated with the coding length of the distribution pattern; that is, the more occurrences of a distribution pattern, the longer the coding length, and vice versa. Please refer to Table 2 for further details.
[0174] By controlling the occurrence of distribution patterns to be negatively correlated with the encoding length of the pattern index, the encoding efficiency of the pattern index is further improved.
[0175] S740 generates the encoding information of the target transform block based on the encoding result of the mode index corresponding to each sub-block.
[0176] In this embodiment, the pattern index encoding result corresponding to each sub-block is obtained, and then the encoding information of the target transform block can be generated. For example, if the target transform block includes 16 sub-blocks, the encoding information is generated according to the pattern index encoding results corresponding to these 16 sub-blocks. The process of obtaining the position of the non-zero transform coefficients corresponding to the target transform block according to the encoding information is described above.
[0177] In this embodiment, the encoder analyzes the distribution pattern of non-zero transform coefficients within each sub-block of the target transform block and encodes only the pattern index (e.g., indicating the position of the distribution pattern in the historical information list). Compared to the traditional method of encoding by coefficient position, this design significantly reduces the code rate while maintaining decoding reliability through indexing of a finite set of patterns.
[0178] Figure 8 is a block diagram illustrating a video decoding apparatus according to some embodiments of this application. As shown in Figure 8, the apparatus includes:
[0179] The acquisition module 801 is configured to acquire the encoding information of the target transform block of the current video frame. The encoding information includes the mode index corresponding to each sub-block in the target transform block. The mode index corresponding to each sub-block is used to indicate the distribution mode of the sub-block. The distribution mode of each sub-block is used to indicate the positional distribution of the non-zero transform coefficients within the sub-block.
[0180] The decoding module 802 is configured to decode the encoding information to obtain the mode index corresponding to each sub-block; determine the distribution mode corresponding to the sub-block according to the mode index corresponding to each sub-block; and obtain the position of the non-zero transform coefficients in the sub-block according to the distribution mode corresponding to each sub-block.
[0181] Figure 9 is a block diagram illustrating a video encoding apparatus according to some embodiments of this application. As shown in Figure 9, the apparatus includes:
[0182] The acquisition module 901 is configured to determine the distribution pattern of each sub-block for a target transform block containing multiple sub-blocks in the current video frame, wherein the distribution pattern of each sub-block indicates the positional distribution of non-zero transform coefficients within the sub-block.
[0183] The determination module 902 is configured to determine the pattern index corresponding to the distribution pattern of each sub-block;
[0184] The encoding module 903 is configured to encode the pattern index corresponding to each sub-block; and generate the encoding information of the target transform block based on the encoding result of the pattern index corresponding to each sub-block.
[0185] It should be noted that the video encoding and decoding apparatus provided in the above embodiments and the video encoding and decoding method provided in the above embodiments belong to the same concept. The specific ways in which each module and unit performs operations have been described in detail in the method embodiments, and will not be repeated here. In practical applications, the video encoding and decoding apparatus provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above, and this is not a limitation here.
[0186] Embodiments of this application also provide an electronic device, including: one or more processors; and a memory for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the video encoding / decoding methods provided in the above embodiments.
[0187] Figure 10 is a schematic diagram of the structure of a computer system suitable for implementing the embodiments of the present application. It should be noted that the computer system 1000 of the electronic device shown in Figure 10 is only an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0188] As shown in Figure 10, the computer system 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1002 or programs loaded from storage portion 1008 into Random Access Memory (RAM) 1003, such as performing the methods described in the above embodiments. The RAM 1003 also stores various programs and data required for system operation. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.
[0189] The following components are connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a local area network (LAN) card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 1010 as needed so that computer programs read from it can be installed into storage section 1008 as needed.
[0190] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by central processing unit (CPU) 1001, it performs various functions defined in the system of this application.
[0191] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. Computer programs contained on computer-readable media can be transmitted using any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0192] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0193] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0194] Another aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the video encoding / decoding method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.
[0195] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video encoding / decoding methods provided in the various embodiments described above.
[0196] Some implementations may involve systems, methods, and / or computer-readable media at any possible level of integration technical detail. The computer-readable medium may include a computer-readable non-transitory storage medium (or multiple media) having computer-readable program instructions on it for causing a processor to perform operations, and may also include storage of a bitstream (or video stream) generated according to the above-described encoding method. When executed by a processor, the computer program / instructions may implement the steps of the video encoding method to generate the bitstream (or video stream), or implement the steps of the video decoding method to decode the bitstream (or video stream).
[0197] The above content is merely a preferred exemplary embodiment of this application and is not intended to limit the implementation of this application. Those skilled in the art can easily make corresponding modifications or alterations based on the main concept and spirit of this application. Therefore, the scope of protection of this application should be determined by the scope of protection claimed in the claims.
[0198] It is understood that in the specific implementation of this application, the input data and other related data required to perform the deduction of the local model are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
Claims
1. A method of video decoding, performed by at least one processor and comprising: obtaining coding information of a target transform block of a current video frame, the coding information comprising a mode index corresponding to each sub-block in the target transform block, the mode index corresponding to each sub-block being used to indicate a distribution mode of the sub-block, the distribution mode of each sub-block being used to indicate a position distribution of non-zero transform coefficients within the sub-block; decoding the mode index corresponding to each sub-block from the coding information; determining the distribution mode corresponding to each sub-block according to the mode index corresponding to the sub-block; obtaining the position of non-zero transform coefficients within each sub-block according to the distribution mode corresponding to the sub-block.
2. The method of claim 1, wherein, a value of the mode index of each sub-block indicates a position of the distribution mode of the sub-block in a history information list, and the history information list records a plurality of distribution modes that have appeared in history in a coding order.
3. The method of claim 2, wherein, the determining the distribution mode corresponding to each sub-block according to the mode index corresponding to the sub-block comprises: for each sub-block, determining the distribution mode at the position indicated by the mode index of the sub-block in the history information list, and determining the distribution mode as the distribution mode corresponding to the sub-block.
4. The method of any of claims 2-3, further comprising: after the target transform block is completely decoded, updating the history information list according to the distribution modes of the sub-blocks in the target transform block.
5. The method of claim 4, wherein, the updating the history information list comprises: reconstructing the history information list according to frequencies of appearance of the distribution modes of the sub-blocks in the target transform block, wherein the history information list arranges the distribution modes in a descending or ascending order of the frequencies of appearance.
6. The method of claim 4, wherein, the updating the history information list comprises: removing the earliest distribution mode from the history information list according to a first-in-first-out principle, and adding the latest distribution mode into the history information list.
7. The method according to any one of claims 1 to 6, wherein, the decoding the mode index corresponding to each sub-block from the coding information comprises: obtaining a mapping table between mode indexes and coding results of the mode indexes, the mapping table comprising a plurality of mode indexes and coding results corresponding to the mode indexes respectively, and coding lengths of the coding results corresponding to different mode indexes being different; converting the coding result corresponding to each sub-block into the mode index corresponding to the sub-block according to the mapping table.
8. The method of claim 7, wherein, the mapping table is a static mapping table agreed by an encoding end and a decoding end in advance, or the mapping table is a dynamic mapping table constructed by the decoding end according to the same strategy as the encoding end.
9. The method of any of claims 1-8, further comprising: obtaining a scanning order of the target transform block; the decoding the mode index corresponding to each sub-block from the coding information comprises: decoding the mode index corresponding to each sub-block from the coding information in a sequence according to the scanning order.
10. The method of claim 9, wherein, the obtaining the scanning order of the target transform block comprises: deriving the scanning order of the target transform block according to a coding mode of a current block; or obtaining an identifier indicating the scanning order from a bitstream.
11. A method of video encoding, performed by at least one processor and comprising: For a target transform block of a current video frame, the target transform block including a plurality of sub-blocks, a distribution pattern of each sub-block is determined, the distribution pattern of each sub-block indicating a position distribution of non-zero transform coefficients within the sub-block; According to the distribution pattern of each sub-block, a mode index corresponding to the distribution pattern is determined; The mode index corresponding to each sub-block is encoded; According to the encoding result of the mode index corresponding to each sub-block, the encoding information of the target transform block is generated.
12. The method of claim 11, wherein, The determination of the mode index corresponding to the distribution pattern of each sub-block includes: A history information list is created, the history information list recording the distribution patterns appeared in history in a coding order; According to the position of the distribution pattern of each sub-block in the history information list, the mode index of the distribution pattern is determined, wherein the value of the mode index of each sub-block indicates the position of the distribution pattern of the sub-block in the history information list.
13. The method of claim 12, further comprising: After the encoding of the target transform block, the history information list is updated according to the distribution patterns of the sub-blocks within the target transform block.
14. The method of claim 13, wherein, The updating of the history information list includes: According to the frequency of appearance of the distribution patterns of the sub-blocks in the target transform block, the history information list is reconstructed, wherein in the history information list, each distribution pattern is arranged in order of frequency of appearance from high to low or from low to high.
15. The method of claim 13, wherein, The updating of the history information list includes: According to the first-in-first-out principle, the earliest distribution pattern is removed from the history information list, and the latest appeared distribution pattern is put into the history information list.
16. A video decoding apparatus, comprising: An acquisition module configured to acquire encoding information of a target transform block of a current video frame, the encoding information including a mode index corresponding to each sub-block in the target transform block, the mode index corresponding to each sub-block being used to indicate a distribution pattern of the sub-block, the distribution pattern of each sub-block being used to indicate a position distribution of non-zero transform coefficients within the sub-block; A decoding module configured to decode the mode index corresponding to each sub-block from the encoding information, determine a distribution pattern corresponding to each sub-block according to the mode index corresponding to the sub-block, and obtain the position of the non-zero transform coefficients within the sub-block according to the distribution pattern corresponding to each sub-block.
17. A video encoding apparatus, comprising: An acquisition module configured to determine a distribution pattern of each sub-block for a target transform block of a current video frame, the target transform block including a plurality of sub-blocks, the distribution pattern of each sub-block indicating a position distribution of non-zero transform coefficients within the sub-block; A determination module configured to determine a mode index corresponding to the distribution pattern of each sub-block according to the distribution pattern of each sub-block; An encoding module configured to encode the mode index corresponding to each sub-block; According to the encoding result of the mode index corresponding to each sub-block, the encoding information of the target transform block is generated.
18. An electronic device, comprising: comprising: one or more processors; a memory storing one or more programs, when executed by the electronic device, cause the electronic device to implement the method of any one of claims 1-15.
19. A computer readable medium having stored thereon a computer program and a bitstream, wherein, The computer program, when executed by a processor, implements the video encoding method of any one of claims 11-15 to generate the bitstream.
20. A method of storing a bitstream, comprising: The video encoding method of any one of claims 11-15 is executed to generate the bitstream. and storing the bitstream.
21. A method of transmitting a bitstream, comprising, executing the video encoding method of any one of claims 11-15 to generate the bitstream; and transmitting the bitstream.
Citation Information
Patent Citations
Method and apparatus for accelerating inverse transform, and method and apparatus for decoding video stream
CN105745929A
Video coding and decoding method and device, computer readable medium and electronic equipment
CN116095329A
Coefficient encoding and decoding method, encoder, decoder and computer storage medium
CN116888965A
Context adaptive position and amplitude coding of coefficients for video compression
US20090086815A1
Encoding and decoding method, encoder, decoder, and storage medium
WO2024148573A1