Video coding method and apparatus, video decoding method and apparatus, electronic device, and computer storage medium
By utilizing the syntax element identifiers of the current block and the reference block during video encoding and decoding, the context model is determined, thus solving the problem of inaccurate context model selection in CABAC and improving encoding performance and flexibility.
Patent Information
- Application Number
- PCT/CN2025/070093
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-05
- Filing Date
- 2025-01-02
- Publication Date
- 2026-01-08
AI Technical Summary
How to improve the accuracy of CABAC context model selection in video encoding and decoding to enhance encoding performance.
By obtaining the syntax element identifiers of the current block and the reference block, the context model of the current block is determined by spatial correlation, including reference blocks in the same frame or different frames. Multiple methods are used to determine the context index and initialization type, thereby improving the accuracy of context model selection.
It improves the coding performance of CABAC, enhances the accuracy and flexibility of context model selection, and reduces complexity.
Smart Images

Figure CN2025070093_08012026_PF_FP_ABST
Abstract
Description
Video coding method and device, electronic equipment and computer storage medium
[0001] The present disclosure claims priority to the patent application filed on July 5, 2024, with the application number 202410902128.0, the content of the prior application is incorporated herein by reference. TECHNICAL FIELD
[0002] The present disclosure relates to the technical field of video coding, in particular to a video coding method and device, electronic equipment and computer storage medium. BACKGROUND
[0003] Context-based adaptive binary arithmetic coding (CABAC) is an entropy coding scheme in video coding standards, which is a coding method combining context model and adaptive binary arithmetic coding.
[0004] An indispensable step in CABAC coding is the selection of context model, and the accuracy of context model selection will affect the performance of CABAC coding. Therefore, how to improve the accuracy of context model selection in CABAC is a problem to be solved at present. SUMMARY
[0005] Some embodiments of the present disclosure disclose a video coding method and device, electronic equipment and computer storage medium for improving the accuracy of context model selection in CABAC.
[0006] In a first aspect, some embodiments of the present disclosure provide a video decoding method, in which, in the case of decoding a current block in any image frame of video data, first, a first identifier of a reference block corresponding to the current block is obtained, the first identifier being an identifier of any syntax element, the current block and the reference block can be understood as units for decoding in the image frame, and the position of the reference block in the corresponding image frame is associated with the position of the current block in the corresponding image frame; then, according to the first identifier, a context model of the current block is determined, and the context model is used for entropy decoding of the syntax element identifier.
[0007] In the above technical solution, by introducing the identifier of the syntax element of the reference block associated with the position of the current block into the selection of the context model of the current block, the spatial correlation of the identifier of the syntax element can be used to improve the accuracy of the selection of the context model, and thus the coding performance of CABAC can be improved.
[0008] In some possible implementation manners, the reference block can include at least one of the following cases:
[0009] The reference block is a block in the same image frame as the current block, and the reference block is adjacent to the current block.
[0010] The reference block is a block in a different image frame from the current block, and the reference block is located at the same position in the corresponding image frame as the position of the current block in the corresponding image frame.
[0011] In the above technical solution, the selection of the reference block can be various, thereby improving the flexibility of the solution.
[0012] In some possible implementation manners, the reference block is a block in the image frame adjacent to the current block and located on the left side of the current block; and / or, the reference block is a block in the image frame adjacent to the current block and located on the top side of the current block.
[0013] In the above technical solution, since the reference block participates in the selection of the context model of the current block, a block with better relevance to the current block can be selected as the reference block, for example, if the reference block is a block in the same image frame as the current block, a block adjacent to the current block and having been encoded before the current block in the image frame can be selected as the reference block, for example, a block on the left side of the current block or a block on the top side of the current block, thereby further improving the accuracy of the selection of the context model.
[0014] In some possible implementation manners, the manner of determining the context model of the current block according to the first identifier can include but is not limited to the following two manners:
[0015] According to the first identifier, a context index is determined; and according to the context index, the context model of the current block is determined. Or,
[0016] According to the initialization type of the second identifier and the first identifier, the context model of the current block is determined, and the second identifier is an identifier of the syntax element of the current block.
[0017] In the above technical solution, a plurality of manners of determining the context model according to the first identifier are provided, thereby increasing the flexibility of the solution.
[0018] In some possible implementation manners, the determination of the context model of the current block according to the initialization type of the second identifier and the first identifier can include:
[0019] According to the first identifier, a context index is determined;
[0020] According to the initialization type of the second identifier and the context index, the context model of the current block is determined.
[0021] In the technical solution, the first identifier is used to determine the context index for determining the context model, and then the context index and the initialization type of the syntax element of the current block are used to determine the context model of the current block, which is similar to the determination process of the context model in the related art, so that the accuracy of selecting the context model can be improved without changing the determination process of the context model in the related art, and the implementation is easy.
[0022] In some possible implementation, the number of the reference blocks is one, and the context index is the value of the first identifier; or the number of the reference blocks is at least two, and the context index is the operation result of the values of the at least two first identifiers.
[0023] In the technical solution, different ways of determining the context index can be selected according to the number of the reference blocks, so that the flexibility of the solution can be improved.
[0024] In some possible implementation, the preset operation can include summation operation, and the context index is the sum of the values of the at least two first identifiers.
[0025] In the technical solution, in the case where the number of the reference blocks is at least two, the operation on the at least two first identifiers can have multiple ways, and as an example, the operation can be summation operation, so that the spatial correlation of the syntax element can be used to further improve the accuracy of determining the context model.
[0026] In some possible implementation, the determining the context model of the current block according to the context index can include:
[0027] The context model of the current block is determined from the preset context models according to the context index.
[0028] In the technical solution, the way of determining the context model through the context index is relatively simple, so that the complexity of the solution can be reduced.
[0029] In some possible implementation, one initialization type corresponds to multiple preset context models, and the context model of the current block is one of the multiple preset context models corresponding to the initialization type of the second identifier.
[0030] In the technical solution, the number of the context models corresponding to each initialization type is increased, so that the selection range of the context model is larger, and the flexibility of the solution can be further improved.
[0031] In some possible implementation manners, the initial value of the parameter of each preset context model is preset, or is determined according to the parameter of the context model after encoding of sample video data, or is determined according to the parameter of the context model after encoding of a previous frame of image.
[0032] In the technical solution, various setting manners of the initial value of the parameter of the preset context model are provided, and a person skilled in the art can set according to actual use requirements.
[0033] In some possible implementation manners, after the context index is determined, the context model offset can be updated according to the context index.
[0034] In the technical solution, the modified context index is assigned to the corresponding syntax element, for example, the context model offset, to save the context index.
[0035] In some possible implementation manners, the syntax element can include an inter-component linear prediction mode.
[0036] The above method can be used to select the context model corresponding to each syntax element involved in the entropy decoding process, wherein each syntax element can include an inter-component linear prediction mode, and the syntax element can also be other modes, for example, any one of a skip mode, an affine mode, a sub-block merge mode, and the like, which is not limited herein.
[0037] In a second aspect, some embodiments of the present disclosure provide a video encoding method, in which, first, a first identifier is obtained, the first identifier being an identifier of any syntax element of a reference block corresponding to a current block, the current block being a unit for encoding in any frame of image of video data, and the reference block being associated with the current block in the corresponding frame of image in position; then, a context model of the current block is determined according to the first identifier, the context model being used for entropy encoding of the identifier of the syntax element.
[0038] The video encoding method of the second aspect is corresponding to the video decoding method of the first aspect, and therefore, the related content is introduced with reference to the first aspect, which is not repeated herein.
[0039] In a third aspect, some embodiments of the present disclosure provide a video encoding apparatus, which can include an obtaining unit and a determining unit. The obtaining unit is configured to obtain a first identifier, the first identifier being an identifier of any syntax element of a reference block corresponding to a current block, the current block being a unit for encoding in any image frame of video data, and the reference block being associated with the current block in the corresponding image frame in terms of position. The determining unit is configured to determine a context model of the current block according to the first identifier, the context model being used for entropy encoding of the identifier of the syntax element.
[0040] In a fourth aspect, some embodiments of the present disclosure provide a video decoding apparatus, which can include an obtaining unit and a determining unit. The obtaining unit is configured to obtain a first identifier, the first identifier being an identifier of any syntax element of a reference block corresponding to a current block, the current block being a unit for decoding in any image frame of video data, and the reference block being associated with the current block in the corresponding image frame in terms of position. The determining unit is configured to determine a context model of the current block according to the first identifier, the context model being used for entropy decoding of the identifier of the syntax element.
[0041] In a fifth aspect, some embodiments of the present disclosure provide an electronic device, which can include a memory and a processor. The memory stores a computer program. The processor is configured to implement the method of any one of the first aspect or the second aspect when executing the computer program.
[0042] In a sixth aspect, some embodiments of the present disclosure provide a chip, which can include a processor and a memory. The processor can be a logic circuit, an integrated circuit, or a general-purpose processor, etc. The memory stores instructions. The processor can implement the method of the first aspect or the second aspect by reading the software code stored in the memory. The memory can be integrated in the processor or exist independently outside the processor.
[0043] In a seventh aspect, some embodiments of the present disclosure provide a computer program product, which can include a computer program (also referred to as code or instructions). When the computer program is executed, the computer program causes a computer to perform the method of the first aspect or the second aspect.
[0044] In an eighth aspect, some embodiments of the present disclosure provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the computer program implements the method of any one of the first aspect or the second aspect. BRIEF DESCRIPTION OF DRAWINGS
[0045] FIG. 1A is an architecture diagram of an example of a video data transmission system according to some embodiments of the present disclosure;
[0046] FIG. 1B is an architecture diagram of another example of a video data transmission system according to some embodiments of the present disclosure;
[0047] FIG. 2 is a functional block diagram of an example of a decoder according to some embodiments of the present disclosure;
[0048] FIG. 3 is a functional block diagram of an example of an encoder according to some embodiments of the present disclosure;
[0049] FIG. 4 is a schematic diagram of an encoding process of CABAC according to some embodiments of the present disclosure;
[0050] FIG. 5 is a schematic diagram of a decoding process of CABAC according to some embodiments of the present disclosure;
[0051] FIG. 6 is a flowchart of an example of a video decoding method according to some embodiments of the present disclosure;
[0052] FIG. 7A, FIG. 7B, and FIG. 7C are schematic diagrams of an example of a reference block according to some embodiments of the present disclosure;
[0053] FIG. 8A, FIG. 8B, and FIG. 8C are schematic diagrams of another example of a reference block according to some embodiments of the present disclosure;
[0054] FIG. 9A, FIG. 9B, and FIG. 9C are schematic diagrams of another example of a reference block according to some embodiments of the present disclosure;
[0055] FIG. 10 is a schematic diagram of another example of a reference block according to some embodiments of the present disclosure;
[0056] FIG. 11 is a schematic diagram of an example of calculating a context index according to some embodiments of the present disclosure;
[0057] FIG. 12 is a schematic diagram of an encoding process of CABAC according to some embodiments of the present disclosure;
[0058] FIG. 13 is a flowchart of an example of a video encoding method according to some embodiments of the present disclosure;
[0059] FIG. 14 is a schematic diagram of a decoding process of CABAC according to some embodiments of the present disclosure;
[0060] FIG. 15 is a flowchart of an example of a video decoding method according to some embodiments of the present disclosure;
[0061] FIG. 16 is a flowchart of another example of a video decoding method according to some embodiments of the present disclosure;
[0062] FIG. 17 is a flowchart of another example of a video encoding method according to some embodiments of the present disclosure;
[0063] FIG. 18 is an exemplary block diagram of a video encoding apparatus according to some embodiments of the present disclosure;
[0064] FIG. 19 is an exemplary block diagram of a video decoding apparatus according to some embodiments of the present disclosure;
[0065] FIG. 20 is an exemplary structural schematic diagram of a video encoding apparatus according to some embodiments of the present disclosure;
[0066] FIG. 21 is an exemplary structural schematic diagram of a video decoding apparatus according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0067] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure.
[0068] In the embodiments of the present disclosure, the same items or similar items with basically the same functions and effects are distinguished by using "first", "second", and the like. For example, the first instruction and the second instruction are used to distinguish different user instructions, and the order is not limited. Those skilled in the art can understand that "first", "second", and the like do not limit the number and execution order, and "first", "second", and the like do not necessarily mean different.
[0069] It should be noted that in the present disclosure, "exemplarily" or "for example" and the like are used to represent as an example, illustration or description. Any embodiment or design scheme described as "exemplarily" or "for example" in the present disclosure should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the use of "exemplarily" or "for example" and the like is intended to present the relevant concept in a specific manner.
[0070] The terms "include", "may include" and "have" and any variations thereof in the embodiments of the present disclosure and the drawings are intended to cover the non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally further include steps or units not listed, or can optionally further include other steps or units inherent to the process, method, product or device.
[0071] “at least one” means one or more, “multiple” means two or more. “And / or” describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character “ / ” generally represents that the associated objects before and after are in an “or” relationship. “At least one of the following” or similar expressions means any combination of these items, which can include any combination of single or multiple items. For example, at least one of a, b and c, which can represent: a, or b, or c, or a and b, or a and c, or b and c, or a, b and c, where a, b and c can be single or multiple.
[0072] To facilitate understanding of the technical solutions of the embodiments of the present disclosure, first, the application scenarios of the present disclosure are described.
[0073] Please refer to FIG. 1A, which is an architecture diagram of an example of a video data transmission system provided by some embodiments of the present disclosure. As shown in FIG. 1A, the video data transmission system can include a first terminal device 101, a second terminal device 102 and a communication network 103, wherein the first terminal device 101 and the second terminal device 102 can transmit video data through the communication network 103. For example, the first terminal device 101 can send local video data to the second terminal device 102, the second terminal device 102 can also send local video data to the first terminal device 101, or the first terminal device 101 and the second terminal device 102 can also send their respective local video data to each other. The communication network 103 can be a wired network, a wireless network or a satellite network, etc., which is not limited here.
[0074] It should be noted that the local video data can be encoded video data (encoded data of the video), and the terminal device receiving the video data can decode the encoded video data to obtain decoded video data.
[0075] Please refer to FIG. 1B, which is an architecture diagram of another example of a video data transmission system provided by some embodiments of the present disclosure. Different from FIG. 1A, the video data transmission system in FIG. 1B can include a collection subsystem 104 connected with the server 105, for collecting real-time video data, encoding the collected real-time video data to obtain encoded video data, and then transmitting the encoded video data to the server 105 for storage. The video data transmission system can also include a third terminal device 106 and a fourth terminal device 107, which are respectively connected with the server 105. The third terminal device 106 and the fourth terminal device 107 can obtain any one or more of the stored encoded video data from the server 105, and decode the encoded video data to obtain decoded video data after obtaining the encoded video data.
[0076] It should be noted that the collection subsystem 104 can include a camera or a camera and other collection components, which can collect real-time video data through the collection components.
[0077] It should also be understood that a terminal device can also be referred to as a user equipment, an access terminal, a user unit, a user station, a mobile station, a mobile station, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, a wireless communication device, a user agent or a user apparatus. It can be a mobile phone, a pad, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical treatment, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc.
[0078] In addition, in the video data transmission system as shown in FIG. 1A or FIG. 1B, the encoder for encoding the video data and the decoder for decoding the encoded video data can be used. In the following, the encoder and the decoder involved in the present disclosure are described respectively.
[0079] Please refer to FIG. 2, which is a functional block diagram of an example of a decoder provided by some embodiments of the present disclosure.
[0080] As shown in FIG. 2, in some embodiments, the decoder can include a first input module 201, through which the encoded video data is obtained, which can be an encoded video sequence, which can be a sequence encoded according to a video coding technology or standard, which can be Versatile Video Coding (H.266 / VVC) or High Efficiency Video Coding (H.265 / HEVC), without limitation.
[0081] In some embodiments, the decoder can include an entropy decoding module 202, which is connected to the first input module 201, and the first input module 201 sends the obtained encoded video sequence to the entropy decoding module 202. After obtaining the encoded video sequence, the entropy decoding module 202 performs an entropy decoding operation on the encoded video sequence, obtains bit symbols corresponding to the encoded video data, and extracts information from the bit symbols corresponding to the encoded video sequence. The extracted information can include transform coefficients, quantizer parameters (QP), and motion vectors, etc.
[0082] In some embodiments, the decoder can include an inverse transform module 203. After the entropy decoding module 202 extracts the information such as transform coefficients, quantizer parameters, and motion vectors, etc. in the encoded video sequence, it sends the transform coefficients to the inverse transform module 203, which performs inverse transform processing on the bit symbols according to the received transform coefficients.
[0083] In some embodiments, the decoder can include an inverse quantization module 204. The entropy decoding module 202 sends the quantizer parameters to the inverse quantization module 204, which performs inverse quantization processing on the bit symbols according to the quantizer parameters, thereby outputting each block and sample values of each block, which can be each coding unit (CU).
[0084] In some embodiments, the decoder can include a motion compensation module 205. In some cases, after being processed by the inverse transform module 203 and the inverse quantization module 204, the output block can be an inter-coded and potentially motion-compensated block. In this case, the entropy decoding module 202 can send the motion vector to the motion compensation module 205, which performs motion compensation processing on the bit symbols according to the motion vector.
[0085] In some embodiments, the decoder can include an intra prediction module 206. In some cases, the output block after being processed by the inverse transform module 203 and the inverse quantization module 204 can be an intra coded block; i.e., a block that is coded without using predictive information from previously reconstructed pictures, but can use predictive information from previously reconstructed parts of the current picture. The predictive information can be provided by the intra prediction module 206. The intra prediction module 206 can generate a block of the same size and shape as the block being reconstructed from surrounding already reconstructed information extracted from the picture that has already been completed reconstructed.
[0086] In some embodiments, the decoder can include an aggregation module 210. The aggregation module 210 can add the predictive information generated by the intra prediction module 206 to the output block after being processed by the inverse transform module 203 and the inverse quantization module 204. Also, the aggregation module 210 can add the samples used for prediction after being motion compensated to the output block after being processed by the inverse transform module 203 and the inverse quantization module 204.
[0087] In some embodiments, the decoder can include a loop filtering module 207. The output block of the aggregation module 210 can be subjected to various loop filtering techniques in the loop filtering module 207. The output of the loop filtering module 207 can be a stream of samples or sample data.
[0088] In some embodiments, the decoder can include a decoded picture buffer module 208. The motion compensation module 205 can obtain reference pictures from the decoded picture buffer module 208 to extract samples used for prediction and to motion compensate the extracted samples used for prediction. The sample data output by the loop filtering module 207 can also be stored in the decoded picture buffer module 208 for subsequent inter prediction and motion compensation.
[0089] In some embodiments, the decoder can include a first output module 209. The first output module 209 can be used for output of sample data, e.g., to a display device.
[0090] In FIG. 2, various modules that a decoder can include are shown. In some embodiments, a decoder can include fewer or additional modules than the ones described above, which is not limited herein.
[0091] FIG. 3 is a functional block diagram of an example of an encoder, according to some embodiments of the present disclosure.
[0092] As shown in FIG. 3, in some embodiments, the encoder can include a second input module 301, through which the video data to be encoded is obtained, which can be a video sequence, which can be an arbitrary-ary sequence, for example, an 8-ary sequence or a 16-ary sequence, etc., which can be a video sequence of the locally stored video data as shown in FIG. 1A or a video sequence corresponding to the captured real-time video data as shown in FIG. IB, which is not limited herein.
[0093] A mode selection and encoding end control module 302 is configured to set the relevant parameters of encoding, such as the code rate and the quantizer parameter (QP), etc., which can be used for the encoding of the video data to be encoded by the entropy encoding module 306.
[0094] In some embodiments, the encoder can include a prediction module 303, which can perform the prediction encoding, and can include obtaining the residual between the block in the input frame and the block of the reference frame according to one or more previously encoded frames designated as the reference frames from the video sequence to be encoded.
[0095] In some embodiments, the encoder can include a transform module 304, which is configured to perform the transform processing on the residual output by the prediction module 303. In some embodiments, the Discrete Cosine Transform (DCT), the Discrete Sine Transform (DST) or the wavelet transform can be used.
[0096] In some embodiments, the encoder can include a quantization module 305, which is configured to perform the quantization processing on the data processed by the transform module 304, to obtain the processed bit symbols.
[0097] In some embodiments, the encoder can include an entropy encoding module 306, which is configured to perform the entropy encoding on the bit symbols processed by the quantization module 305 according to the relevant parameters of encoding set by the mode selection and encoding end control module 302, which can be the variable length encoding or the Huffman encoding, etc., so as to convert the bit symbols into the encoded video sequence.
[0098] In some embodiments, the encoder can include a second output module 307, through which the video sequence encoded by the entropy encoding module 306 is output.
[0099] In some embodiments, the encoder can include a decoder 308. The decoder 308 can decode and reconstruct the encoded video data of the reference frames based on the bit symbols processed by the quantization module 305, and the reconstructed reference frames are stored in a decoded picture buffer module of the decoder 308. The intra prediction module and the motion compensation module in the decoder 308 can search for sample data, such as blocks of reference frames, in the decoded picture buffer module as the reference frames for the input frame to be encoded.
[0100] In FIG. 3, the various modules that can be included in the encoder are shown. In some embodiments, the encoder can include fewer or additional modules than those shown in FIG. 3.
[0101] During encoding, the mode selection and encoding end control module 302 can assign a picture frame type to each encoded image frame, and the picture frame type can be any one of the following frame types:
[0102] An intra-coded (I) picture frame, referred to as an I frame, can be a picture frame that can be encoded and decoded without using any other frame in the sequence as a prediction source.
[0103] A predicted (P) picture frame, referred to as a P frame, can be a picture frame that can be encoded and decoded using intra prediction or inter prediction.
[0104] A bidirectional predicted (B) picture frame, referred to as a B frame, can be a picture frame that can be encoded and decoded using intra prediction or inter prediction.
[0105] The encoding methods for I frames, P frames, and B frames are similar to those in related video coding technologies or standards, and will not be described here.
[0106] The above is an introduction to the video data transmission system, the encoder, and the decoder involved in the present disclosure. Next, the entropy coding involved in the present disclosure will be described in combination with the above content.
[0107] CABAC is an entropy coding scheme in a video coding standard, which is a coding method that combines a context model with adaptive binary arithmetic coding.
[0108] Please refer to FIG. 4 for a coding flowchart of CABAC. As shown in FIG. 4, the coding flow of CABAC can include the following steps:
[0109] Step S401, binary processing is performed on the input syntax element.
[0110] The binarization method in CABAC can include fixed-length binarization (FL), truncated rice binarization (TR), and k-th order exp-golomb binarization (EGK), etc. In CABAC, different binarization methods can be selected according to different probability distribution characteristics of input syntax elements.
[0111] Fixed-length binarization is a binary algorithm that converts the value of a syntax element into a fixed-length binary symbol. When the probability of the value of a syntax element is uniformly distributed, the fixed-length coding binarization scheme can be selected. For example, the value of a given syntax element is x, and 0≤x≤cMax, then the fixed-length binary symbol string of x is obtained by directly using the decimal number conversion to binary number method, and the length of the fixed-length binary symbol string of x is wherein, l FL is the length of the binary symbol string, represents the upward rounding.
[0112] The truncated rice code binary code is formed by splicing a prefix string and a suffix string. The calculation formula of the prefix value P is: P=V>>R. If the P value is less than (cMax>>R), the prefix string is composed of P ones and a zero, and the length is P+1; if the P value is greater than or equal to (cMax>>R), the prefix string is composed of (cMax>>R) ones, and the length is (cMax>>R). The calculation formula of the suffix value S is: S=V-(P<<R), and the suffix string is the binary string of S, and the length is R. When the value V of the syntax element is greater than or equal to cMax, there is no suffix string. Wherein, cMax is the maximum value of the syntax element to be binarized, R is the rice parameter, and V is the value of the syntax element to be binarized.
[0113] The binary code of k-th order exp-golomb binarization is also formed by splicing a prefix and a suffix. Wherein, the prefix part is composed of the unary code corresponding to the value of The suffix part can be calculated by using the binary value of x+2k(1-2l(x)) with a length of k+l(x)k+l(x) bits.
[0114] For example, when the probability of the syntax element in the input image frame is uniformly distributed, the FL scheme can be selected.
[0115] Step S402, context modeling is performed on the syntax element.
[0116] The key of context modeling of the syntax element is to determine the variable of the context model. The context model involves two types of variables: the probability update rate and the predicted probability state. When the probability update rate and the predicted probability state of the context model corresponding to the syntax element are determined, the context modeling can be performed according to the probability update rate and the predicted probability state of the context model corresponding to the syntax element.
[0117] In step S403, the arithmetic coding is performed on the syntax element after the binarization according to the context model corresponding to the syntax element, so that the coded bitstream is obtained.
[0118] The coded bitstream is the coded sequence of the input image frame after the CABAC processing.
[0119] The arithmetic coding used in the CABAC coding can include regular coding and bypass coding. In the regular coding, the adaptive probability model, i.e., the context model, is used for coding. In the bypass coding, the equal probability is used for coding, and the probability state does not need to be updated.
[0120] Correspondingly, referring to FIG. 5, a decoding flowchart of the CABAC is shown. As shown in FIG. 5, the decoding flowchart of the CABAC can include the following steps:
[0121] In step S501, the context model of each bit symbol of the coded bitstream is obtained.
[0122] In this process, it is determined according to the standard syntax element structure which syntax element is currently decoded, and the context model is obtained according to the position of the bit symbol currently decoded in the coded bitstream and some reference information. The specific manner is similar to that in step S402, and thus will not be described herein.
[0123] In step S502, the arithmetic decoding is performed on each bit symbol according to the context model corresponding to the bit symbol, so that the decoded binary sequence is obtained.
[0124] This step corresponds to step S403. The decoding manner can include regular decoding and bypass decoding. In the regular decoding, the adaptive probability model is used for decoding. In the bypass decoding, the equal probability is used for decoding, and the probability state can not need to be updated. The regular decoding is similar to the regular coding in step S403, and the bypass decoding is similar to the bypass coding in step S403.
[0125] In step S503, the syntax element is reconstructed according to the decoded binary sequence.
[0126] The value corresponding to the syntax element can be reconstructed through the inverse process of the binarization of the decoded binary sequence. The specific manner is similar to that in step S401, and thus will not be described herein.
[0127] As can be seen from the CABAC encoding and decoding processes shown in FIG. 4 and FIG. 5, one step in CABAC encoding and decoding is the selection of a context model, and whether the selection is accurate will affect the performance of CABAC encoding and decoding. Therefore, how to improve the accuracy of the selection of the context model in CABAC is a problem that needs to be solved at present.
[0128] In view of this, some embodiments of the present disclosure provide a video decoding method, in which, in the case of decoding a current block in any image frame of video data, a first identifier of a reference block corresponding to the current block is first obtained, the first identifier is an identifier of any syntax element, the current block and the reference block can both be understood as a unit for encoding in an image frame, and the position of the reference block in the corresponding image frame is associated with the position of the current block in the corresponding image frame; then, according to the first identifier, a context model of the current block is determined, and the context model is used for entropy decoding of the syntax element identifier.
[0129] In the above technical solution, by introducing the identifier of the syntax element of the reference block associated with the position of the current block into the selection of the context model of the current block, the spatial correlation between the current block and the reference block is utilized to improve the accuracy of the selection of the context model, and thus the decoding performance of CABAC can be improved.
[0130] Next, the video decoding method of the embodiments of the present disclosure will be described in combination with the accompanying drawings.
[0131] Please refer to FIG. 6, which is a flowchart of an example of the video decoding method provided by some embodiments of the present disclosure. The method can be applied to the decoder shown in FIG. 2, and as shown in FIG. 6, the method can include the following steps:
[0132] Step S601, a first identifier is obtained.
[0133] In some embodiments of the present disclosure, the first identifier is an identifier of any syntax element of the reference block corresponding to the current block, or the first identifier can be understood as an identifier of any prediction mode of the reference block. The any syntax element is, for example, a skip flag (cu_skip_flag), an affine flag (inter_affine_flag), a subblock merge flag (merge_subblock_flag), a CU split flag (qt_split_cu_flag, mtt_split_cu_flag, mtt_split_cu_vertical_flag, mtt_split_cu_binary_flag), a motion vector precision IMV flag (amvr_mode), an inter-component linear prediction mode (cclm_mode_flag), and the like. The any prediction mode can be a mode corresponding to the foregoing syntax element, for example, a skip mode, an affine mode, a subblock merge mode, and the like, which are not enumerated one by one herein.
[0134] Specifically, the current block is a unit for decoding of the bitstream to be decoded. As an example, the unit for decoding can be the same as the unit for encoding, and thus, the unit for encoding is described below.
[0135] Video data can be understood as a sequence of images in the order of a time axis, and thus, video coding can be actual coding of each image in the sequence. Each image can be referred to as an image frame or a frame image, and is described below by taking the image frame as an example. In order to facilitate coding of the image frame, each image frame can be divided into a plurality of coding units (CUs) before coding, and coding of an image frame can be understood as coding of the plurality of CUs of the image frame. In some embodiments of the present disclosure, a block is a CU, and the current block is a CU to be decoded.
[0136] It should be noted that the division of the CU can be determined according to the input coding structure, or the coding structure can be preset, which is not limited herein.
[0137] The reference block is a block associated with the position of the current block. The position association can be understood as that the position of the reference block in a corresponding image frame is associated with the position of the current block in the corresponding image frame. In some embodiments of the present disclosure, the reference block and the current block can be blocks in the same image frame or blocks in different image frames, which is not limited herein. If the reference block and the current block are blocks in the same image frame, the reference block can be a block in the image frame where the current block is located and associated with the position of the current block. If the reference block and the current block are blocks in different image frames, the reference block can be a block in another image frame different from the image frame where the current block is located and associated with the position of the current block.
[0138] In addition, it can be understood that the reference block can be a block that has been decoded before the current block.
[0139] The two cases are described below.
[0140] In the first case, the reference block and the current block are blocks in the same image frame.
[0141] In this case, the reference block is a block in the image frame where the current block is located and at a preset distance from the current block. The preset distance can be a distance that makes the correlation of the syntax elements of the reference block and the current block greater than or equal to a preset correlation threshold. The preset correlation threshold can be preset, set according to actual use requirements, or obtained according to a limited number of experiments. The preset distance can be in units of pixel size of an image frame, for example, 4 pixel units or 3 pixel units, or in units of a block, for example, 3 blocks or 2 blocks, etc. In addition, the number of reference blocks can be one or multiple, which is not limited herein.
[0142] Referring to FIGS. 7A-7C, they are schematic diagrams of an example of the reference block in some embodiments of the present disclosure. In FIGS. 7A and 7B, taking the preset distance as 1 block and the number of reference blocks as 1 for example, as shown in FIG. 7A, the reference block 712 is located at the left side of the current block 711 and at a distance of 1 block from the current block; as shown in FIG. 7B, the reference block 722 is located at the upper side of the current block 721 and at a distance of 1 block from the current block. In FIG. 7C, taking the preset distance as 1 block and the number of reference blocks as 2 for example, as shown in FIG. 7C, the reference block 732 is located at the left side of the current block 731 and at a distance of 1 block from the current block, and the reference block 733 is located at the upper side of the current block 731 and at a distance of 1 block from the current block.
[0143] It can be understood that the closer the distance between the reference block and the current block, the greater the correlation between the syntax elements of the reference block and the current block. Therefore, please refer to FIGS. 8A-8C for another example of the reference block provided by some embodiments of the present disclosure. In FIGS. 8A-8C, the reference block is adjacent to the current block. As shown in FIG. 8A, the reference block 812 is a block adjacent to the current block 811 and located on the left side of the current block in the image frame; as shown in FIG. 8B, the reference block 822 is a block adjacent to the current block 821 and located on the top side of the current block in the image frame; as shown in FIG. 8C, the reference block 832 is a block adjacent to the current block 831 and located on the left side of the current block in the image frame, and the reference block 833 is a block adjacent to the current block 831 and located on the top side of the current block in the image frame.
[0144] In the second case, the reference block and the current block are blocks in different image frames.
[0145] In this case, the difference between the distance between the position of the reference block in the corresponding image frame and the position of the current block in the corresponding image frame and the projection position in the same image frame is less than or equal to a preset distance. The preset distance can be a distance that makes the correlation between the syntax elements of the reference block and the current block greater than or equal to a preset correlation threshold, which is similar to the first case and will not be repeated here. It should be noted that the corresponding image frame can be understood as the image frame in which the block is located, for example, the corresponding image frame of the reference block can be understood as the image frame in which the reference block is located, and the corresponding image frame of the current block can be understood as the image frame in which the current block is located.
[0146] Please refer to FIGS. 9A-9C for another example of the reference block in some embodiments of the present disclosure. In FIGS. 9A and 9B, taking the preset distance as 1 block and the number of reference blocks as 1 for example, as shown in FIG. 9A, the position of the reference block 912 in the first image frame is projected to the projection position 913 in the second image frame, which is located on the left side of the position of the current block 911 in the second image frame and is 1 block away from the current block; as shown in FIG. 9B, the position of the reference block 922 in the first image frame is projected to the projection position 923 in the second image frame, which is located on the top side of the position of the current block 921 in the second image frame and is 1 block away from the current block. In FIG. 9C, taking the preset distance as 1 block and the number of reference blocks as 2 for example, as shown in FIG. 9C, the position of the reference block 932 in the first image frame is projected to the projection position 933 in the second image frame, which is located on the left side of the position of the current block 931 in the second image frame and is 1 block away from the current block, and the position of the reference block 934 in the first image frame is projected to the projection position 935 in the second image frame, which is located on the top side of the position of the current block in the second image frame and is 1 block away from the current block.
[0147] It can be understood that, since the reference block and the current block are not in the same image frame, the closer the distance between the projection positions of the reference block and the current block in the same image frame, the greater the correlation of the syntax elements of the reference block and the current block. Therefore, please refer to FIG. 10 for a schematic diagram of another example of a reference block provided by some embodiments of the present disclosure. In FIG. 10, the position of the reference block 102 in the corresponding image frame is the same as the position of the current block 101 in the corresponding image frame, that is, the position of the reference block 102 in the first image frame is the same as the position of the current block in the second image frame 102.
[0148] The first image frame is any one image frame different from the image frame in which the current block is located. It can be understood that the first image frame can be an image frame that has been completed encoding.
[0149] Of course, the relative position relationship between the reference block and the current block can also have other cases, which are not exemplified one by one here. In the above manner, the probability of the identification of the syntax element can be more finely divided, so as to improve the accuracy of the determined context model.
[0150] Step S602: determining the context model of the current block according to the first identification.
[0151] In some embodiments of the present disclosure, the context model is used for entropy decoding of the identification of the syntax element.
[0152] Step S602 can include, but is not limited to, the following two execution manners:
[0153] The first execution manner:
[0154] determining a context index according to the first identification;
[0155] determining the context model of the current block according to the context index.
[0156] In some embodiments, a preset context model can be pre-stored in the decoder, and when the context index is determined according to the first identification, the context model of the current block is determined from the preset context model according to the context index.
[0157] The second execution manner:
[0158] determining the context model of the current block according to the initialization type of the second identification and the first identification. The second identification is the identification of the syntax element of the current block.
[0159] After the first identification is acquired, the context model of the current block is determined together with the initialization type of the identification of the syntax element of the current block. Of course, in the case of different syntax elements, the context model of the current block can also be determined together according to the first identification and other parameters of the identification of the syntax element of the current block. In some embodiments, the initialization type of the identification is only one of the cases.
[0160] It should be noted that the determination method of the initialization type of the identification of the syntax element can be preset, or agreed in the standard, or indicated by the indication information, which is not limited herein. As an example, taking the syntax element cclm_mode_flag as an example, the initialization type initType of cclm_mode_flag can be determined by the picture frame type (sh_slice_type) and sh_cabac_init_flag. Among them, the syntax element sh_cabac_init_flag represents the initialization list determination method used in the context variable initialization process.
[0161] Specifically, if (sh_slice_type==I)
[0162] Among them, sh_cabac_init_flag is determined by pps_cabac_init_present_flag. Among them, the syntax element pps_cabac_init_present_flag represents whether sh_cabac_init_flag exists in the frame slice (slice) or frame segment header.
[0163] pps_cabac_init_present_flag is equal to 1, indicating that sh_cabac_init_flag exists in the segment header referring to the picture parameter set (PPS), and the value of sh_cabac_init_flag is inferred to be 1.
[0164] pps_cabac_init_present_flag is equal to 0, indicating that sh_cabac_init_flag does not exist in the segment header referring to the PPS, and the value is inferred to be 0.
[0165] According to the judgment code, the value of the initialization type initType of the cclm_mode_flag can include 0, 1 and 2, the initialization type initType of the I type frame is 0, the initialization type initType of the P type frame is 2 or 1 according to the sh_cabac_init_flag, and the initialization type initType of the B type frame is 1 or 2 according to the sh_cabac_init_flag.
[0166] In some embodiments of the present disclosure, the context model of the current block is determined according to the initialization type of the second identification and the value of the first identification. It can be understood that the context model of the current block is determined according to the initialization type of the second identification and the value of the first identification.
[0167] The determination of the context model of the current block according to the initialization type of the second identification and the value of the first identification can include, but is not limited to, the following two implementation manners:
[0168] The first implementation manner:
[0169] In this case, the decoder can pre-store the mapping relationship between the initialization type of the identification of the syntax element, the preset value and the preset context model. As an example, the mapping relationship can be a table, such as Table 1, which is an example of the mapping relationship between the initialization type of the identification of the syntax element, the preset value and the preset context model.
[0170] As shown in Table 1, the preset context model can include multiple types, and the same initialization type corresponds to multiple different preset context models. In Table 1, the preset value can include 0, 1 and 2. In this way, after determining the initialization type of the identification of the syntax element and the value of the first identification of the current block, the context model of the current block can be determined by looking up Table 1.
[0171] Table 1
[0172] It should be noted that Table 1 is an example, and the initialization type and the preset value can be set to more values. The mapping relationship corresponding to different syntax elements can be the same or different, and examples are not listed here.
[0173] The second implementation manner:
[0174] In this case, the determination of the context model of the current block according to the initialization type of the second identification and the value of the first identification can specifically include:
[0175] determining a context index according to the first identification;
[0176] According to the initialization type of the second identifier and the context index, a context model of the current block is determined.
[0177] Specifically, after the first identifier is acquired, the context index can be determined according to the value of the first identifier, and then the context model of the current block is determined according to the initialization type of the syntax element of the current block and the context index.
[0178] It should be noted that in this case, the number of preset context models corresponding to the initialization type is changed. The number of preset context models corresponding to each initialization type is increased from 1 to at least 2. The preset context model can also be a candidate context model. For the sake of unity, the preset context model is used as an example in the following.
[0179] According to the number of reference blocks, the number of preset context models can be different. In the case of one reference block, the number of preset context models can be set to two or three or more; in the case of two reference blocks, the number of preset context models can be set to three or four or more. The more the number of reference blocks, the more the number of preset context models, thereby improving the flexibility of selecting the context model used for encoding the current block.
[0180] As an example, taking the syntax element cclm_mode_flag as an example, when the reference block is adjacent to the current block and located on the left side and on the top side of the current block, the preset context model of cclm_mode_flag can be set to three; when the position of the reference block in the corresponding image frame is the same as the position of the current block in the corresponding image frame, that is, the reference block is the homonym block of the current block, the preset context model of cclm_mode_flag can be set to two.
[0181] Please refer to Table 2 for an example of the value of the preset context model ID (abbreviated as ctxId) corresponding to different initialization types when the syntax element is cclm_mode_flag and the reference block is adjacent to the current block and located on the left side and on the top side of the current block.
[0182] Table 2
[0183] As shown in Table 2, when the initialization type initType = 0, the ctxIdx candidate set is [0, 1, 2]; when the initialization type initType = 1, the ctxIdx candidate set is [3, 4, 5]; when the initialization type initType = 2, the ctxIdx candidate set is [6, 7, 8], and ctxIdx is an integer.
[0184] In the technical solution, the number of context models corresponding to each initialization type is increased, so that the selection range of the context model is larger, and the flexibility of the solution can be further improved.
[0185] It should be noted that, since the first identifier is the identifier of the syntax element of the reference block, the number of reference blocks is different, and the way of determining the context index according to the first identifier is also different. According to the number of reference blocks, different ways of determining the context index can be selected, and the flexibility of the solution can be improved. In some embodiments of the present disclosure, the two determination methods can include but are not limited to the following two determination methods:
[0186] The first determination method is:
[0187] The number of reference blocks is one, and the context index is the value of the first identifier.
[0188] That is, if there is only one reference block, the value of the identifier of the syntax element of the reference block is the context index value. For example, if the cclm_mode_flag of the reference block is 0, the context index is 0; if the cclm_mode_flag of the reference block is 1, the context index is 1.
[0189] The second determination method is:
[0190] The number of reference blocks is at least two, and the context index is the operation result of the values of the at least two first identifiers.
[0191] In some embodiments of the present disclosure, the operation result is obtained by performing a preset operation on the values of the at least two first identifiers. The preset operation can be a sum operation, a difference operation, a product operation or a quotient operation, etc., which is not limited herein. That is, if the preset operation is a sum operation, the context index is the sum of the values of the at least two first identifiers. By using the spatial correlation of the identifier of the syntax element through the sum operation, the accuracy of determining the context model can be further improved.
[0192] Referring to FIG. 11, an example of calculating the context index of the present disclosure is shown. In FIG. 11, the first flag is cclm_mode_flag. As shown in FIG. 11, if the left CU 112 and the top CU 113 of the current block 111 exist, the cclm_mode_flag value of the left CU 112 and the cclm_mode_flag value of the top CU 113 are obtained, respectively, and the flag values of the two CUs are added, i.e. the Neibor_CclmFlag_Sum value, which can be 0, 1 or 2. The Neibor_CclmFlag_Sum value is determined as the context index. If the left CU of the current block or the top CU of the current block does not exist, the cclm_mode_flag value of the non-existing CU is 0 by default.
[0193] Table 3
[0194] In this determination manner, the input can be the initialization type initType and the sum value Neibor_CclmFlag_Sum of the cclm_mode_flag of the reference block, and the output can be the ID of the context model of the current block or the index number of the context model of the current block, i.e. ctxIdx. As shown in Table 3.
[0195] As shown in Table 3, the context model of the current block can be one of a plurality of preset context models corresponding to the initialization type of the second flag. In some embodiments, the value of the first flag used to determine the context index can be understood as the binary bit value.
[0196] In addition, it should be noted that, according to the foregoing, the number of preset context models in some embodiments of the present disclosure is increased, and the initial value of the parameters of each preset context model also has a greater impact on the performance of coding. In some embodiments of the present disclosure, the initial value of the parameters of each preset context model is preset. Alternatively, it is determined according to the parameters of the context model after the sample video data is encoded, that is, the context model probability situation after each frame of picture in each sequence is encoded can be recorded in a training manner, and the initial value of the context model is changed according to the training result. As an example, when the context model probability after a frame is encoded reaches a relatively stable situation, the current parameter of the context model is set as the initial value of the parameters of the context model. Alternatively, the initial value of the parameters of each preset context model is determined according to the parameters of the context model after the previous frame of image is encoded. As an example, the initial value of the cclm_mode_flag context model and the context probability situation after encoding are recorded when a frame of picture is encoded, respectively. If the interpolation of the two is less than a threshold, the cclm_mode_flag context model is not initialized in the next frame of encoding, and the cclm_mode_flag context model parameter after the previous frame of encoding is continued to be used.
[0197] Please refer to Table 4 for an example of the initial value of the parameters of the cclm_mode_flag context model in some embodiments of the present disclosure.
[0198] Table 4
[0199] As shown in Table 4, the initialization parameters of the cclm_mode_flag context model can include an initialization value initValue and a window index shiftIdx, initValue is used to calculate the initial probability state of the model, and shiftIdx is used to control the updating rate of the probability of the model. Of course, the parameters of the cclm_mode_flag context model can also include other parameters, which are not described one by one here.
[0200] In the above technical solution, the context index used to determine the context model is first determined according to the first identifier, and then the context model of the current block is determined according to the two factors of the context index and the initialization type of the identifier of the syntax element of the current block. The flow is similar to the determination flow of the context model in the related art, so that the accuracy of selecting the context model can be improved without changing the determination flow of the context model in the related art, and it is easy to implement.
[0201] Step S603: updating the context model offset according to the context index.
[0202] After determining the context index of the context model of the current block, the modified context index can be assigned to the corresponding syntax element, as an example, assigned to the context model offset binIdx to save the context index.
[0203] Please refer to Table 5 for an example of assigning the modified context index to binIdx in some embodiments of the present disclosure, with the syntax element being cclm_mode_flag.
[0204] Table 5
[0205] As shown in Table 5, taking the blocks located on the left and top sides of the current block as the reference blocks, when the value of binIdx is 0, it corresponds to the sum of the values of the first identifier, and when the value of binIdx is greater than or equal to 1, it corresponds to na (null).
[0206] It should be noted that step S603 is an optional step, i.e., it is not necessary to be executed, and in FIG. 6, step S603 is represented by a dashed line to indicate that it is an optional step.
[0207] After determining the context model of the current block, the syntax element identifier of the current block is entropy decoded according to the context model, for example, the cclm_mode_flag identifier is entropy decoded.
[0208] In the above technical solution, by introducing the identifier of the syntax element of the reference block associated with the position of the current block into the selection of the context model of the current block, the spatial correlation of the identifier of the syntax element is utilized to improve the selection accuracy of the context model, and thus the coding and decoding performance of CABAC can be improved.
[0209] The foregoing describes the video decoding method in some embodiments of the present disclosure, and the following describes the video encoding method in some embodiments of the present disclosure.
[0210] Please refer to FIG. 12 for a flowchart of an example of the video encoding method provided by some embodiments of the present disclosure. The method can be applied to the encoder shown in FIG. 3, and as shown in FIG. 12, the method can include the following steps:
[0211] Step S1201: obtaining a first identifier.
[0212] In some embodiments of the present disclosure, the first identifier is an identifier of any syntax element of a reference block corresponding to the current block, the current block being a unit for encoding in any frame image of the video data, and the reference block being associated with the position of the current block in the corresponding frame image.
[0213] It should be noted that the unit for decoding and the unit for encoding can be the same, for example, both are coding units (CUs).
[0214] Step S1202, determining a context model of the current block according to the first identifier.
[0215] In some embodiments of the present disclosure, the context model is used for entropy encoding of the syntax element identifier.
[0216] It should be noted that step S1201 is similar to step S601, step S1202 is similar to step S602, and since the video encoding method and the video decoding method are reciprocal processes, the content similar to the decoding method in the video encoding method is not described in some embodiments of the present disclosure.
[0217] In the above technical solution, by introducing the identifier of the syntax element of the reference block associated with the position of the current block into the selection of the context model of the current block, the spatial correlation of the identifier of the syntax element is used to improve the selection accuracy of the context model, and thus the coding and decoding performance of CABAC can be improved.
[0218] In the following, the video encoding method provided in some embodiments of the present disclosure is described in detail taking the CABAC encoding process as an example.
[0219] Please refer to FIG. 13, which is a schematic diagram of the encoding process of CABAC provided in the present disclosure. For convenience of description, in the following, taking the syntax element cclm_mode_flag for encoding as an example, as shown in FIG. 13, the encoding process of CABAC can include the following steps:
[0220] Step S1301, obtaining the binarized cclm_mode_flag.
[0221] First, the meaning of cclm_mode_flag is described.
[0222] The chroma intra prediction modes can be classified into non-cross-component prediction modes and cross-component prediction modes according to whether the luminance reconstructed pixels are used. Three cross-component prediction modes are defined according to different neighboring reference regions involved in the model construction: cross-component linear model chroma intra-prediction derived from the left and top template (INTRA LT CCLM), cross-component linear model chroma intra-prediction derived from the left template (INTRA L CCLM) and cross-component linear model chroma intra-prediction derived from the top template (INTRA T CCLM). For each of the three cross-component prediction modes, four reference pixels (may include four pairs of luminance and chroma reconstructed values) are selected from its corresponding reference region to construct a linear model, and the chroma prediction values are calculated according to the linear model with the luminance reconstructed values of the current coding unit (CU) as the independent variable.
[0223] The cclm mode flag indicates whether the current chroma intra prediction mode is a cross-component prediction mode, i.e. whether the current chroma intra prediction mode is one of INTRA LT CCLM, INTRA L CCLM and INTRA T CCLM. When the value of cclm mode flag is 1, it means that the current chroma intra prediction mode is one of INTRA LT CCLM, INTRA L CCLM and INTRA T CCLM. When the value of cclm mode flag is 0, it means that the current chroma intra prediction mode is not one of INTRA LT CCLM, INTRA L CCLM and INTRA T CCLM. In addition, when the value of cclm mode flag is 1, the cross-component prediction mode index cclm mode idx is further transmitted to indicate which one of INTRA LT CCLM, INTRA L CCLM and INTRA T CCLM is the current chroma intra prediction mode. When cclm mode flag is not present, the value of cclm mode idx is also 0.
[0224] In some embodiments of the present disclosure, the syntax element cclm_mode_flag is selected to be fixed length binarization. Please refer to Table 6, which shows the binarization method of the syntax element cclm_mode_flag and the input parameters.
[0225] Table 6
[0226] In Table 6, the maximum value of the input parameters of the syntax element cclm_mode_flag is cMax, and cMax = 1. Thus, the syntax element cclm_mode_flag is binarized according to the binarization method of fixed length binarization to obtain the binarized cclm_mode_flag. The specific binarization process is similar to that in the related art, and thus will not be described here.
[0227] Step S1302, determining the context model corresponding to cclm_mode_flag.
[0228] Step S1302 is similar to steps S601-S603 in FIG. 6, and thus will not be described here.
[0229] Step S1303, performing binary arithmetic coding on the binarized cclm_mode_flag according to the context model corresponding to cclm_mode_flag to obtain the encoded bitstream.
[0230] Binary arithmetic coding is based on the recursive interval division method, and the coding interval and the lower limit of the interval are saved in the recursive process. Binary arithmetic coding can include two coding methods: regular coding and bypass coding. The former uses adaptive probability model for coding; the latter uses equal probability for coding, and the probability state does not need to be updated.
[0231] The input of the regular encoder is two types of variables of the context model and the Bin value to be coded, wherein the two types of variables of the context model can include: shiftIdx which controls the probability updating rate of the model, and pStateIdx0 and pStateIdx1 which predict the probability state; and the Bin value to be coded is the binary symbol of the binarized cclm_mode_flag.
[0232] It should be noted that in some embodiments of the present disclosure, a double probability model is used to predict the least probable symbol (LPS) symbol probability of each context. p (t+1) = (p0 (t+1) +p1 (t+1) ) / 2 (1) p0 (t+1) =p0 (t) · (1-α0) +x (t) ·α0 (2) p1 (t+1) =p0 (t) · (1-α1) +x (t) ·α1 (3)
[0233] wherein α0, α1 are the probability update rates of the two models, p0 (t), p1 (t) are the prediction probabilities of the two models of the double probability model at the tth time, x (t) is the symbol of the double probability model at the tth time, p0 (t+1), p1 (t+1) are the prediction probabilities of the two models of the double probability model at the t+1th time.
[0234] In order to avoid multiplication operation, α is limited to α = 2 -β (β∈N + ), and the probability prediction model is q (t+1) =q (t) -(q (t) > >β+x (t) · ((2 b -1) > >β) (4)
[0235] wherein q (t) is the b-bit integerized representation of p (t).
[0236] In some embodiments, b0 = 10, b1 = 14, and the probability prediction model can be p (t) =q (t) · 2 b +2 b-1 (5)
[0237] wherein pStateIdx0 is the prediction probability state of the probability prediction model q0 (t), and pStateIdx1 is the prediction probability state of the probability prediction model q1 (t).
[0238] As an example, the input of the conventional encoder is the context model (shiftIdx, pStateIdx0, pStateIdx1) and the binary symbol to be encoded (Bin), and the state of the encoder is the current encoding interval width Range and the interval lower limit Low. The current encoding interval width Range and the interval starting point Low. The initial value of Range is 510, and the initial value of Low is 0. The specific encoding process is as follows:
[0239] 1) Calculate the interval width R LPS , R MPs of the LPS: qRangeIdx = Range > > 5 (6) pState = pStateIdx1 + 16 · pStateIdx0 (7) R LPS = (qRangeIdx • pState5) » 1 + 4 (9) R MPS = Range - R LPS (10)
[0240] where pState5 represents the prediction probability with 5-bit precision; when pState » 14 = 1, the prediction probability is greater than 0.5, pState is the prediction probability of MPS, and the exclusive OR operation keeps it as the prediction probability of LPS.
[0241] 2) Update Low and Range:
[0242] The update formula of Range and Low is shown as follows: MPS = pState » 14 (11)
[0243] If Bin = LPS, then Low = Low + R MPS , Range = R LPS ;
[0244] If Bin = MPS, then Low remains unchanged, and Range = R MPS .
[0245] 3) Encode the interval width Range and renormalize.
[0246] With the update of the encoding interval Range, the value of Range can be less than 256, and renormalization is needed at this time. The renormalization method is to perform a left shift operation on the interval lower limit Low and the encoding interval width Range at the same time, until the value of the encoding interval width Range is greater than or equal to 256, and the bits shifted out of the interval lower limit Low are the encoding output bits.
[0247] Step S1304, update the probability model corresponding to the candidate ctxIdx.
[0248] According to the value of the encoded binary symbol and the update rate shiftIdx, the probability state pStateIdx0 and pStateIdx1 are constantly updated by using the probability prediction model. Further, the prediction probability obtained by using the double probability model is averaged to obtain the prediction probability pState of LPS. pState = pStateIdx1 + 16 • pStateIdx0 (12) pStateIdx0 = pStateIdx0 - (pStateIdx0 » shift0) + ((2 10- 1) • Bin » shiftO) (13) pStateIdxl = pStateIdxl - (pStateIdxl » shiftl) + ((2 14 - 1) • Bin » shiftl) (14) shiftO = (shiftIdx » 2) + 2 (15) shiftl = (shiftIdx & 3) + 3 + shiftO (16)
[0249] where shiftO (aO in the probability prediction model) and shiftl (al in the probability prediction model) are the update rates of the prediction probabilities pStateIdxO and pStateIdxl, respectively; Bin is the already coded bin; and the prediction probabilities pStateIdxO and pStateIdxl are 10-bit and 14-bit integer representations, respectively.
[0250] From the foregoing, each context index has an initial value initValue and a shiftIdx. The initValue is used to calculate the model probability state, in particular as follows: slopeIdx = initValue » 3 (17) offsetIdx = initValue (18) m = slopeIdx - 4 (19) n = (offsetIdx « 18) + 1 (20) preCtxState = Clip3(1, 127, ((m « (Clip3(0, 63, SliceQPy) - 16)) » 1) + n) (21) pStateIdxO = preCtxState « 3 (22) pStateIdxl = preCtxState « 7 (23)
[0251] where SliceQPy is the quantization parameter of the luma signal.
[0252] In some embodiments, the pStateIdxO and pStateIdxl for the double probability model prediction are increased by a corresponding weight, while the shiftO and shiftl are set to fine-tune the mechanism according to the coded bin being 0 or 1.
[0253] The operation for deriving the resulting probability for binary arithmetic coding by introducing the weight is as follows: p = ((32 - ω) • pO + ω • pl) » 5 (24)
[0254] where ω is a weight selected from a predefined set ω∈{10, 12, 16, 20, 22}, and different frame types, such as an I frame picture type, a B frame picture type, and a P frame picture type, have three different weights predetermined for the context models of the respective frame types. The weight of the I frame picture type is allowed to be used only for an intra slice, and the weights of the B frame picture type and the P frame picture type are allowed to be switched for an inter slice based on a same syntax sh_cabac_init_flag at a slice level.
[0255] In some embodiments of the present disclosure, two probability states of CABAC are updated with a short window and a long window, respectively, and the sizes of the long window and the short window can be specified by a shiftIdx item in a standard document, the update rates of the long window and the short window are controlled, and the sizes of the long window and the short window are fine-tuned according to an initial value of shiftIdx based on the encoded symbols being 0 or 1, the update ranges of the long window and the short window are both -7 to 7, and the lower limit of the window size is set to 2.
[0256] By the technical solution, the coding can be performed by using the fine context modeling method of the syntax element cclm_mode_flag, and the coding performance can be improved.
[0257] Next, the video decoding method provided by some embodiments of the present disclosure is described in detail by taking a CABAC coding process as an example.
[0258] Please refer to FIG. 14, which is a decoding flowchart of CABAC provided by the present disclosure. For convenience of description, the syntax element to be decoded is taken as cclm_mode_flag in the following description, and as shown in FIG. 14, the decoding flowchart of CABAC can include the following steps.
[0259] Step S1401, a code stream to be decoded is acquired.
[0260] In some embodiments of the present disclosure, the code stream to be decoded is a code stream obtained by coding the syntax element cclm_mode_flag.
[0261] Step S1401, a context model corresponding to each CU in the code stream is determined.
[0262] It is determined according to the syntax element structure of the standard which syntax element is currently decoded, for example, cclm_mode_flag. The ctxIdx corresponding to the current symbol is acquired according to the position of bin in the bit stream, the corresponding binIdx, and some context reference information.
[0263] Step S1402 is similar to steps S601-S603 in FIG. 6, and will not be described here again.
[0264] Step S1403, according to the interval division condition and the value corresponding to the interval lower limit in the code stream, binary arithmetic decoding is performed to obtain each symbol of the code stream.
[0265] The value of each symbol can be 0 or 1.
[0266] As an example, the estimated probability of the symbol is obtained by the shiftIdx, pStateIdx0, pStateIdx1 and the weight of the corresponding context model.
[0267] In some embodiments of the present disclosure, the binary decoding mode can also include a regular decoding mode and a bypass decoding mode. The former uses an adaptive probability model for decoding; the latter uses an equal probability mode for decoding, and the probability state does not need to be updated. The specific decoding mode can be obtained by the bypassFlag of the current symbol, which distinguishes the two modes by the position of the bin in the bit stream, the corresponding binIdx and some context reference information in the code stream.
[0268] The input of the regular encoder is the shiftIdx, pStateIdx0, pStateIdx1 and the current coding interval length Range in the context model. As an example, the initial value of Range is 510, and the interval lower limit m_value is obtained by reading the byte from the code stream. The specific decoding process is as follows:
[0269] After obtaining the corresponding context model used for decoding the current symbol, the size of the interval corresponding to the LPS symbol is calculated according to the estimated probability and the current coding interval length m_Range, and the interval lower limit is compared with the size of the current coding interval to determine whether the value of the symbol is 1 or 0. It should be noted that the coding interval and the decoding interval can be the same.
[0270] 1) Calculate the interval width R corresponding to LPS LPS , R MPS : qRangeIdx = Range » 5 (25) pState = pStateIdx1 + 16 pStateIdx0 (26)
[0271] R LPS = (qRangeIdx pState5) » 1 + 4 (28) R MPS = Range - R LPS (29)
[0272] Wherein, pState5 represents the prediction probability of 5-bit precision; when pState>>14=1, the prediction probability is greater than 0.5, pState is the prediction probability of MPS, and XOR The operation makes the prediction probability of LPS remain.
[0273] 2) Compare the lower limit of the interval with the corresponding sub-interval positions of MPS and LPS:
[0274] If m_Value MPS , m_Value remains unchanged, Range=R MPS , and Bin=MPS.
[0275] If m_Value MPS , m_Value=m_Value-R MPS , Range=R LPS , and Bin=LPS.
[0276] 3) With the update of the coding interval Range, the value of Range may be less than 256, in which case renormalization needs to be performed, and at the same time, Low and Range are left shifted until the value of Range is greater than or equal to 256, and the bits left shifted out of Low are the value of the decoded output symbol.
[0277] Step S1404, the value of each symbol obtained by decoding is assigned to cclm_mode_flag, and the syntax element cclm_mode_flag is reconstructed.
[0278] The inverse process of binarization of the binary string obtained by decoding one syntax element can obtain the value corresponding to the reconstructed syntax element.
[0279] Step S1405, the probability model corresponding to the candidate ctxIdx is updated.
[0280] Step S1405 is similar to step S1204, and will not be described here.
[0281] Through the above technical solutions, the decoding can be performed by using the fine context modeling method of the syntax element cclm_mode_flag, and the decoding performance can be improved.
[0282] Some embodiments of the present disclosure provide a video decoding method, which can include the following steps with reference to FIG. 15:
[0283] S151, entropy decoding data corresponding to the first syntax element of the first chroma block is obtained.
[0284] The first syntax element is a syntax element identifying whether an intra prediction mode of the chroma block is an inter-component prediction mode. That is, the first syntax element is cclm_mode_flag.
[0285] In some embodiments, the entropy encoding data of the first syntax element of the first chroma block can be obtained from the video bitstream according to a standard syntax element structure.
[0286] S152, obtaining spatial information of the first chroma block according to neighboring decoded chroma blocks of the first chroma block.
[0287] In some embodiments, the neighboring decoded chroma blocks of the first chroma block can include chroma blocks that are adjacent to the first chroma block and have been decoded before the first chroma block is decoded.
[0288] In some embodiments, the neighboring decoded chroma blocks of the first chroma block can include chroma blocks above the first chroma block and / or chroma blocks to the left of the first chroma block; and the step S152 of obtaining the spatial information of the first chroma block according to the neighboring decoded chroma blocks of the first chroma block can include obtaining the spatial information of the first chroma block according to the chroma blocks above the first chroma block and / or the chroma blocks to the left of the first chroma block.
[0289] In some embodiments, the step S152 of obtaining the spatial information of the first chroma block according to the neighboring decoded chroma blocks of the first chroma block can include obtaining the spatial information of the first chroma block according to whether an intra prediction mode of the neighboring decoded chroma blocks is an inter-component prediction mode.
[0290] S153, determining a target context model according to the spatial information of the first chroma block.
[0291] In some embodiments, the step of determining the target context model according to the spatial information of the first chroma block can include selecting the target context model from a preset context model set according to the spatial information of the first chroma block.
[0292] In some embodiments, the step of determining the target context model according to the spatial information of the first chroma block can include determining the target context model according to the spatial information of the first chroma block and an initialization type (initType) of the first syntax element of the first chroma block.
[0293] S154, obtaining a value of the first syntax element of the first chroma block based on the target context model and the entropy encoding data of the first syntax element of the first chroma block.
[0294] The video decoding method provided in the above embodiments first acquires the entropy encoding data of a first syntax element of a first chroma block, and then acquires the spatial domain information of the first chroma block according to the neighboring decoded chroma blocks of the first chroma block, and then determines a target context model according to the spatial domain information of the first chroma block, and finally acquires the value of the first syntax element of the first chroma block based on the target context model and the entropy encoding data of the first syntax element of the first chroma block. Since the video decoding method provided in the embodiments of the present disclosure acquires the spatial domain information of the first chroma block according to the neighboring decoded chroma blocks, and determines the context model for entropy decoding the value of the first syntax element of the first chroma block according to the spatial domain information of the first chroma block, some embodiments of the present disclosure can determine the context model for entropy decoding the first syntax element by combining the spatial characteristics of the first syntax element, thereby improving the accuracy of the context model for entropy decoding the first syntax element.
[0295] As an extension and refinement of the above embodiments, some embodiments of the present disclosure provide another video decoding method, which can include the following steps, as shown in FIG. 16:
[0296] S1601, acquiring the entropy encoding data of a first syntax element of a first chroma block.
[0297] The first syntax element is a syntax element indicating whether the intra prediction mode of a chroma block is an inter-component prediction mode.
[0298] S1602, acquiring the value of the first syntax element of the neighboring decoded chroma block.
[0299] That is, acquiring the value of cclm_mode_flag of the neighboring decoded chroma block.
[0300] S1603, acquiring the spatial domain information of the first chroma block according to the value of the first syntax element of the neighboring decoded chroma block.
[0301] In some embodiments, the step S1603 (acquiring the spatial domain information of the first chroma block according to the value of the first syntax element of the neighboring decoded chroma block) can include summing the values of the first syntax elements of the neighboring decoded chroma blocks to acquire the spatial domain information of the first chroma block.
[0302] For example, the neighboring decoded chroma blocks can include a chroma block above the first chroma block and a chroma block to the left of the first chroma block, and the value of the first syntax element of the chroma block above the first chroma block is 1, and the value of the first syntax element of the chroma block to the left of the first chroma block is 1, then the spatial domain information of the first chroma block can be determined as 2.
[0303] For example, the adjacent decoded chroma blocks can include the chroma blocks located above the first chroma block and the chroma blocks located left to the first chroma block, and the value of the first syntax element of the chroma block located above the first chroma block is 0 and the value of the first syntax element of the chroma block located left to the first chroma block is 0, and the spatial domain information of the first chroma block can be determined as 0.
[0304] For example, the adjacent decoded chroma blocks can include the chroma blocks located above the first chroma block and the chroma blocks located left to the first chroma block, and the value of the first syntax element of the chroma block located above the first chroma block is 0 and the value of the first syntax element of the chroma block located left to the first chroma block is 0, and the spatial domain information of the first chroma block can be determined as 0.
[0305] S1604, obtaining an initialization type (initType) of the first syntax element of the first chroma block.
[0306] In some embodiments, the implementation of obtaining the initialization type of the first syntax element of the first chroma block can include: determining the value of the syntax element sh_cabac_init_flag according to the value of the syntax element pps_cabac_init_present_flag, and determining the initialization type of the first syntax element of the first chroma block according to the value of the syntax element sh_cabac_init_flag and the frame type (I type or B type or P type) corresponding to the first chroma block.
[0307] In some embodiments, the implementation of obtaining the initialization type of the first syntax element of the first chroma block can include:
[0308] If the frame type corresponding to the first chroma block is I type, the initialization type of the first syntax element of the first chroma block is determined as a first initialization type; if the frame type corresponding to the first chroma block is P type and the value of sh_cabac_init_flag is 1, the initialization type of the first syntax element of the first chroma block is determined as a second initialization type; if the frame type corresponding to the first chroma block is P type and the value of sh_cabac_init_flag is 0, the initialization type of the first syntax element of the first chroma block is determined as a third initialization type; if the frame type corresponding to the first chroma block is B type and the value of sh_cabac_init_flag is 1, the initialization type of the first syntax element of the first chroma block is determined as the third initialization type; if the frame type corresponding to the first chroma block is B type and the value of sh_cabac_init_flag is 0, the initialization type of the first syntax element of the first chroma block is determined as the second initialization type.
[0309] S1605, determine a target index set according to the initialization type.
[0310] In some embodiments, the initialization type is a first initialization type, a second initialization type, or a third initialization type, and the first initialization type, the second initialization type, and the third initialization type correspond to a context model index set respectively; and the step of determining the target index set according to the initialization type can include: when the initialization type is the first initialization type, determining the context model index set corresponding to the first initialization type as the target index set; when the initialization type is the second initialization type, determining the context model index set corresponding to the second initialization type as the target index set; and when the initialization type is the third initialization type, determining the context model index set corresponding to the third initialization type as the target index set.
[0311] As shown in Table 2 above, the context model index set corresponding to the first initialization type (initType=0) is {0, 1, 2}, the context model index set corresponding to the second initialization type (initType=1) is {3, 4, 5}, and the context model index set corresponding to the third initialization type (initType=2) is {6, 7, 8}, so when the initialization type is the first initialization type, the target index set is {0, 1, 2}, when the initialization type is the first initialization type, the target index set is {3, 4, 5}, and when the initialization type is the first initialization type, the target index set is {6, 7, 8}.
[0312] S1606, selecting a target model index from the target index set according to the spatial information of the first chroma block.
[0313] In some embodiments, the target index set can include three context model indexes (as shown in Table 2 above). The step S1606 of selecting a target model index from the target index set according to the spatial information of the first chroma block can include:
[0314] When the spatial information of the first chroma block is m, the (m+1)th context model index in the target index set is selected as the target model index; 0≥m≥2.
[0315] wherein 0≥m≥2 means that m≥0 and m≤2. Since the spatial information of the first chroma block is the value of a first syntax element or the sum of the values of two first syntax elements, m is an integer and m is 0, 1, or 2.
[0316] That is, when the spatial information of the first chroma block is 0, the first context model index in the target index set is selected as the target model index; when the spatial information of the first chroma block is 1, the second context model index in the target index set is selected as the target model index; and when the spatial information of the first chroma block is 2, the third context model index in the target index set is selected as the target model index.
[0317] S1607, determining the target context model according to the target model index.
[0318] In some embodiments, the step S1607 of determining the target context model according to the target model index can include the following steps 16071 and 16072.
[0319] Step 16071, obtaining the initial value and the update rate of the target context model according to the target model index.
[0320] In some embodiments, the initial value (initValue) of the target context model is 35, and the update rate (shiftIdx) of the target context model is 8.
[0321] Step 16072, constructing the target context model according to the initial value and the update rate of the target context model.
[0322] In some embodiments, the step of constructing the target context model according to the initial value and the update rate of the target context model can include calculating the model probability state (pStateIdx0 and pStateIdx1) of the target context model according to the initial value (initValue) of the target context model. The implementation of calculating the model probability state of the target context model according to the initial value of the target context model can refer to the above formulas (17) to (23), and will not be repeated here to avoid redundancy.
[0323] In some embodiments, the step of determining the target context model according to the target model index can include reading the target context model according to the target model index.
[0324] S1608, using the target context model as the context model of the CABAC algorithm to perform arithmetic decoding on the entropy encoded data of the first syntax element of the first chroma block by the CABAC algorithm to obtain the value of the first syntax element of the first chroma block.
[0325] The arithmetic decoding of the entropy-encoded data of the first syntax element of the first chroma block using the CABAC algorithm includes two decoding methods: regular decoding and bypass decoding. Regular decoding utilizes an adaptive probability model for decoding; bypass decoding is performed with equal probability, its probability state does not need to be updated, and the two decoding methods are distinguished by a bypass flag.
[0326] The encoder input for conventional decoding is the context model (update rate shiftIdx, probability state pStateIdx0, probability state pStateIdx1) and the current encoding interval width Range. The initial value of the encoding interval width Range is 510, and the lower limit m_value of the interval is obtained by reading bytes from the bitstream. The basic principle of conventional decoding is as follows: after obtaining the corresponding context used for decoding the current symbol, the size of the interval corresponding to the LPS symbol is calculated based on the estimated probability and the current encoding interval width m_Range. The lower limit of the interval is compared with the size of the current encoding interval to determine whether the symbol symbol is 1 or 0. The specific decoding process can be described as follows: steps ① to ④:
[0327] Step ①: Calculate the interval width R corresponding to LPS. LPS and R MPS .
[0328] Calculate the interval width R corresponding to LPS LPS and R MPS The formulas are shown in formulas (6) to (10) above. To avoid redundancy, they will not be explained in detail here.
[0329] Step 2: Compare the lower limit m_Value of the interval with the corresponding sub-interval positions of MPS and LPS.
[0330] If m_Value <R MPS Then m_Value remains unchanged, Range = R MPS The corresponding syntax element's binary value is Bin = MPS;
[0331] If m_Value>=R MPS Then m_Value = m_Value - R LPS Range = R LPS The corresponding syntax element's binary value is Bin = LPS.
[0332] Step 3: Renormalize the encoding range width (Range).
[0333] Similarly, as the coding interval width Range is updated, the value of the coding interval width Range can be less than 256 (the interval width is initialized to 510), and renormalization is needed. The renormalization is performed by left-shifting both the interval lower bound Low and the coding interval width Range until the value of the coding interval width Range is greater than or equal to 256, and the bits left-shifted out of the interval lower bound Low are the coding output bits.
[0334] Step 4, reconstructing the value of the syntax element
[0335] In some embodiments, reconstructing the value of the syntax element can include performing an inverse process of binarization on a binary string obtained by decoding the value of the syntax element, thereby reconstructing the value of the syntax element.
[0336] S1609, updating the target context model according to the value of the first syntax element of the first chroma block.
[0337] That is, the probability states pStateIdx0 and pStateIdx1 of the target context model used are updated according to the value of tu_cb_coded_flag of the first chroma block.
[0338] In some embodiments, the video decoding method provided by some embodiments of the present disclosure can further include: determining a context model corresponding to a second chroma block according to the target context model; and wherein the first chroma block and the second chroma block are chroma blocks corresponding to two chroma components of the same coding block.
[0339] That is, the video decoding method provided by some embodiments of the present disclosure obtains a target context model corresponding to a chroma block of a U component of a coding block, obtains a value of a first syntax element of the first chroma block based on the target context model and entropy coding data of the first syntax element of the first chroma block, and determines a context model corresponding to a chroma block of a V component of the coding block according to the target context model.
[0340] In some embodiments, the determining a context model corresponding to a second chroma block according to the target context model can include: determining the target context model as the context model corresponding to the second chroma block.
[0341] After determining a context model corresponding to a second chroma block according to the target context model, the video encoding method provided by some embodiments of the present disclosure can further include:
[0342] obtaining a value of a first syntax element of the second chroma block based on the context model corresponding to the second chroma block and entropy coding data of the first syntax element of the second chroma block.
[0343] Some embodiments of the present disclosure provide a video encoding method, which can include the following steps with reference to FIG. 17:
[0344] S171, obtaining the spatial domain information of the first chroma block according to the adjacent coded chroma blocks of the first chroma block.
[0345] The implementation of step S171 can refer to the implementation of step S152, except that step S152 is to obtain the spatial domain information of the first chroma block according to the adjacent decoded chroma blocks of the first chroma block, and step S171 is to obtain the spatial domain information of the first chroma block according to the adjacent coded chroma blocks of the first chroma block. In addition, based on the order of coding and decoding, the adjacent decoded chroma blocks of the first chroma block in the decoding process and the adjacent coded chroma blocks of the first chroma block in the encoding process are the same chroma blocks.
[0346] S172, determining a target context model according to the spatial domain information of the first chroma block.
[0347] The implementation of determining a target context model according to the spatial domain information of the first chroma block can refer to the implementation of determining a target context model according to the spatial domain information of the first chroma block in the video decoding method, and will not be described in detail here to avoid repetition.
[0348] S173, entropy encoding the value of the first syntax element of the first chroma block based on the target context model to obtain the entropy encoded data of the first syntax element of the first chroma block.
[0349] The first syntax element is a syntax element for identifying whether the intra prediction mode of the chroma block is the inter-component prediction mode.
[0350] The video encoding method provided by the above embodiments first obtains the spatial domain information of the first chroma block according to the adjacent coded chroma blocks of the first chroma block, then determines a target context model according to the spatial domain information of the first chroma block, and finally entropy encodes the value of the first syntax element of the first chroma block based on the target context model to obtain the entropy encoded data of the first syntax element of the first chroma block. Since the video encoding method provided by the above embodiments obtains the spatial domain information of the first chroma block according to the adjacent coded chroma blocks and determines the context model for entropy encoding the value of the first syntax element of the first chroma block according to the spatial domain information of the first chroma block, the above embodiments can determine the context model for entropy encoding the first syntax element by combining the spatial characteristics of the first syntax element, thereby improving the accuracy of the context model for entropy encoding the first syntax element.
[0351] The above describes the schemes provided by some embodiments of the present disclosure mainly from the implementation perspective of the method flow. It can be understood that, in order to implement the above functions, the encoder or the decoder can include hardware structures and / or software modules for performing respective functions. It should be easily realized by those skilled in the art that, in combination with the units, modules and algorithm steps of the examples described in the embodiments disclosed herein, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.
[0352] In the case of using integrated units (modules), FIG. 18 shows a possible exemplary block diagram of a video coding apparatus involved in some embodiments of the present disclosure, which can exist in the form of software. The video coding apparatus 1800 can include an obtaining unit 1801 and a determining unit 1802. The obtaining unit 1801 is configured to obtain information for determining a context model, such as a first identifier. The determining unit 1802 is configured to determine the context model according to the information obtained by the obtaining unit 1801. Optionally, the video coding apparatus 1800 can further include a storage unit for storing program codes and / or data of the video coding apparatus 1800, which is not shown in FIG. 18.
[0353] The obtaining unit 1801 can be a communication interface, a transceiver or a transceiving circuit, etc., wherein the communication interface is a general term, which can include multiple interfaces in specific implementation. The storage unit can be a memory. The determining unit 1802 can be a processor or a controller, which can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosed content of the embodiments of the present disclosure.
[0354] The video coding apparatus 1800 can be a decoder in any of the above embodiments, or can also be a chip disposed in the decoder. The obtaining unit 1801 can support the video coding apparatus 1800 to perform the obtaining actions in the above method examples. For example, for performing step S1201 in the embodiment shown in FIG. 12, and / or for performing step S1301 in the embodiment shown in FIG. 13. The determining unit 1802 can support the video coding apparatus 1800 to perform other actions in the above method examples, for example, for performing step S1202 in the embodiment shown in FIG. 6, and / or for performing steps S1302-S1304 in the embodiment shown in FIG. 13.
[0355] It should be noted that the division of units (modules) in some embodiments of the present disclosure is illustrative, and is only a logical function division. In actual implementation, another division manner can be used. The function modules in the embodiments of the present disclosure can be integrated in one processing module, or each module can be physically present individually, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software function module.
[0356] When the integrated module is realized in the form of a software function module and sold or used as an independent product, the integrated module can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present disclosure essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods described in the embodiments of the present disclosure. The foregoing storage medium can be a memory or various other media that can store program codes.
[0357] In the case of using the integrated unit (module), FIG. 19 shows a possible exemplary block diagram of a video decoding apparatus involved in some embodiments of the present disclosure, which can exist in the form of software. The video decoding apparatus 1900 can include an acquisition unit 1901 and a determination unit 1902. The acquisition unit 1901 is configured to acquire information for determining a context model, for example, a first identifier. The determination unit 1902 is configured to determine the context model according to the information acquired by the acquisition unit 1901. Optionally, the video decoding apparatus 1900 can further include a storage unit for storing the program code and / or data of the video decoding apparatus 1900, which is not shown in FIG. 19.
[0358] The acquisition unit 1901 can be a communication interface, a transceiver or a transceiving circuit, etc., wherein the communication interface is a general term, and in specific implementation, the communication interface can include a plurality of interfaces. The storage unit can be a memory. The determination unit 1902 can be a processor or a controller, which can realize or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure of the embodiments of the present disclosure.
[0359] The video decoding apparatus 1900 can be a decoder in any of the above-described embodiments, or can also be a chip disposed in the decoder. The obtaining unit 1901 can support the video decoding apparatus 1900 to perform the obtaining action in each of the above method examples. For example, for performing step S601 in the embodiment shown in FIG. 6, and / or for performing step S1401 in the embodiment shown in FIG. 14. The determining unit 1902 can support the video decoding apparatus 1900 to perform other actions in each of the above method examples, for example, for performing steps S602 and S603 in the embodiment shown in FIG. 13, and / or for performing steps S1402-S1405 in the embodiment shown in FIG. 14.
[0360] It should be noted that the division of the units (modules) in some embodiments of the present disclosure is illustrative, and is only a logical function division. When actually implemented, another division manner can be used. Each functional module in the embodiments of the present disclosure can be integrated in one processing module, or each module can exist physically independently, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software functional module.
[0361] When the integrated module is realized in the form of a software functional module and sold or used as an independent product, the software functional module can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present disclosure essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium can be a memory or various other media that can store program codes.
[0362] As shown in FIG. 20, the video encoding apparatus 2000 provided by some embodiments of the present disclosure, wherein the video encoding apparatus 2000 can be a terminal, and can realize the functions of the terminal in the methods provided by some embodiments of the present disclosure; the video encoding apparatus 2000 can also be an apparatus capable of supporting the terminal to realize the functions of the terminal in the methods provided by some embodiments of the present disclosure. The video encoding apparatus 2000 can be a chip system. In some embodiments of the present disclosure, the chip system can be composed of a chip, or can include a chip and other discrete devices.
[0363] In hardware implementation, the obtaining unit 1801 can be a transceiver integrated in the video encoding apparatus 2000 to form a communication interface 2010.
[0364] The video encoding device 2000 can include at least one processor 2020 configured to implement or support implementation of the encoder of the methods provided by some embodiments of the present disclosure. For example, the processor 2020 can determine a context model corresponding to a current block, as described in detail in the method examples, which will not be repeated here.
[0365] The video encoding device 2000 can further include at least one memory 2030 configured to store program instructions and / or data. The memory 2030 and the processor 2020 can be coupled. The coupling between the various devices, units, or modules in some embodiments of the present disclosure can be indirect coupling or communication connection between devices, units, or modules, which can be electrical, mechanical, or other forms, for information exchange between devices, units, or modules. The processor 2020 can operate in cooperation with the memory 2030. The processor 2020 can execute program instructions stored in the memory 2030. At least one of the at least one memory can be included in the processor.
[0366] The video encoding device 2000 can further include a communication interface 2010 configured to communicate with other devices through a transmission medium, so that the devices in the video encoding device 2000 can communicate with other devices. For example, the other device can be a decoder. The processor 2020 can use the communication interface 2010 to transceive data. The communication interface 2010 can be a transceiver.
[0367] The specific connection medium between the communication interface 2010, the processor 2020, and the memory 2030 in some embodiments of the present disclosure is not limited. In FIG. 20, the memory 2030, the processor 2020, and the communication interface 2010 are connected by a bus 2040, which is represented by a thick line in FIG. 20. The connection mode between other components is only schematically illustrated and is not limited. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, only one thick line is used in FIG. 20, but it does not mean that there is only one bus or only one type of bus.
[0368] In some embodiments of the present disclosure, the processor 2020 can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, and can implement or execute the disclosed methods, steps, and logic block diagrams in some embodiments of the present disclosure. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as hardware processor execution or executed by a combination of hardware and software modules in the processor.
[0369] In some embodiments of the present disclosure, the memory 2030 can be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), and the like, and can also be a volatile memory, such as a random-access memory (RAM). The memory can be any other medium capable of carrying or storing desired program codes in the form of instructions or data structures and capable of being accessed by a computer, but is not limited thereto. The memory in some embodiments of the present disclosure can also be a circuit or any other device capable of realizing a storage function, for storing program instructions and / or data.
[0370] As shown in FIG. 21, the video decoding apparatus 2100 can be a terminal, which can realize the functions of the terminal in the methods provided by some embodiments of the present disclosure; or the video decoding apparatus 2100 can be an apparatus capable of supporting the terminal to realize the functions of the terminal in the methods provided by some embodiments of the present disclosure. The video decoding apparatus 2100 can be a chip system. In some embodiments of the present disclosure, the chip system can be composed of a chip, or can include a chip and other discrete devices.
[0371] In hardware implementation, the obtaining unit 1901 can be a transceiver integrated in the video decoding apparatus 2100 to form a communication interface 2110.
[0372] The video decoding apparatus 2100 can include at least one processor 2120 for realizing or for supporting the video decoding apparatus 2100 to realize the functions of the decoder in the methods provided by some embodiments of the present disclosure. For example, the processor 2120 can determine a context model corresponding to a current block, and details are described in the method examples, which are not repeated here.
[0373] The video decoding apparatus 2100 can further include at least one memory 2130 for storing program instructions and / or data. The memory 2130 is coupled with the processor 2120. The coupling in some embodiments of the present disclosure is an indirect coupling or communication connection between devices, units or modules, which can be electrical, mechanical or other forms, for information interaction between devices, units or modules. The processor 2120 can operate in cooperation with the memory 2130. The processor 2120 can execute program instructions stored in the memory 2130. At least one of the at least one memory can be included in the processor.
[0374] The video decoding apparatus 2100 can further include a communication interface 2110 for communicating with other devices through a transmission medium, so that the devices in the video decoding apparatus 2100 can communicate with other devices. For example, the other device can be an encoder. The processor 2120 can transceive data by using the communication interface 2110. The communication interface 2110 can be a transceiver in particular.
[0375] The specific connection medium between the communication interface 2110, the processor 2120 and the memory 2130 is not limited in some embodiments of the present disclosure. In FIG. 21, the connection between the memory 2130, the processor 2120 and the communication interface 2110 is through a bus 2140, which is represented by a thick line in FIG. 21, and the connection mode between other components is only illustrative and is not limited. The bus can be divided into an address bus, a data bus, a control bus and the like. For convenience of representation, only one thick line is used in FIG. 21, but it does not mean that there is only one bus or only one type of bus.
[0376] In some embodiments of the present disclosure, the processor 2120 can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the disclosed methods, steps and logic block diagrams in some embodiments of the present disclosure. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as hardware processor execution or executed by a combination of hardware and software modules in the processor.
[0377] In some embodiments of the present disclosure, the memory 2130 can be a non-volatile memory such as a hard disk drive (HDD) or a solid-state drive (SSD), and can also be a volatile memory such as a random-access memory (RAM). The memory can be any other medium capable of carrying or storing desired program codes in the form of instructions or data structures and capable of being accessed by a computer, but is not limited to this. The memory in some embodiments of the present disclosure can also be a circuit or any other device capable of realizing a storage function, for storing program instructions and / or data.
[0378] The embodiments of the present disclosure further provide a computer readable storage medium, which stores computer instructions, when the computer instructions run on an electronic device, the electronic device performs the video encoding method or the video decoding method described above.
[0379] The computer readable storage medium can be implemented in any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this document, the computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0380] A computer readable signal medium can include a propagated data signal with computer readable program code embodied therein, for use by or in connection with an instruction execution system, apparatus, or device. The computer readable program code can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0381] Computer readable program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0382] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0383] The embodiments of the present disclosure further provide a computer program product, which, when executed on a computer, causes the computer to perform some or all of the steps of the above-mentioned video encoding method or video decoding method embodiments.
[0384] Some embodiments of the present disclosure provide a chip system, which can include a processor, and can further include a memory for implementing the functions of the video encoding device or video decoding device in the above-mentioned method. The chip system can be composed of a chip, or can include a chip and other discrete devices.
[0385] Some embodiments of the present disclosure provide a system, which can include the above-mentioned video encoding device and video decoding device.
[0386] From the above description of the embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0387] In several embodiments provided in the present disclosure, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the division of the apparatus embodiments is merely illustrative, and the division of the modules or units can not mean physical division, and each module or unit can be integrated in a physical entity, or one or more modules or units can be combined or integrated in a physical entity. In addition, the display or discussion of a relative coupling or direct coupling or communication connection between modules or units can be implemented by an indirect coupling or a communication connection between modules or units through some interfaces, devices or units, and can be in electric, mechanical or other forms.
[0388] The units described as separated components can or can not be physically separated, and the components displayed as units can be one physical unit or multiple physical units, i.e., can be located in one place, or can be distributed in multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0389] In addition, each functional unit in the various embodiments of the present disclosure can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.
[0390] The above is merely specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present disclosure, which should be covered in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method of video decoding, wherein, The method comprises: obtaining a first identifier, the first identifier being an identifier of any syntax element of a reference block corresponding to a current block, the current block being a unit for decoding in any one image frame of video data, and the reference block being associated with the position of the current block in the corresponding image frame; determining a context model of the current block according to the first identifier, the context model being used for entropy decoding of the identifier of the syntax element.
2. The method of claim 1, wherein, The reference block comprises at least one of the following cases: The reference block and the current block are blocks in the same image frame, and the reference block is adjacent to the current block. The reference block and the current block are blocks in different image frames, and the position of the reference block in the corresponding image frame is the same as the position of the current block in the corresponding image frame.
3. The method of claim 2, wherein, The reference block is a block adjacent to the current block and located on the left side of the current block in the image frame; and / or The reference block is a block adjacent to the current block and located on the top side of the current block in the image frame.
4. The method of claim 1, wherein, The determining of the context model of the current block according to the first identifier comprises: determining a context index according to the first identifier; determining the context model of the current block according to the context index.
5. The method of claim 1, wherein, The determining of the context model of the current block according to the first identifier comprises: determining the context model of the current block according to the initialization type of the second identifier and the first identifier, the second identifier being an identifier of the syntax element of the current block.
6. The method of claim 5, wherein, The determining of the context model of the current block according to the initialization type of the second identifier and the first identifier comprises: determining a context index according to the first identifier; determining the context model of the current block according to the initialization type of the second identifier and the context index.
7. The method of any one of claims 4-6, wherein, The number of the reference blocks is one, and the context index is a value of the first identifier.
8. The method of any one of claims 4-6, wherein, The number of the reference blocks is at least two, and the context index is an operation result of values of at least two first identifiers, the operation result being obtained by preset operation on the values of the at least two first identifiers.
9. The method of claim 8, wherein, The preset operation comprises summation operation, and the context index is a sum of the values of the at least two first identifiers.
10. The method of claim 5, wherein, The determining of the context model of the current block according to the context index comprises: determining the context model of the current block from preset context models according to the context index.
11. The method of claim 5, wherein, A plurality of preset context models correspond to the same initialization type, and the context model of the current block is one of the plurality of preset context models corresponding to the initialization type of the second identifier.
12. The method of claim 10 or 11, wherein, A parameter initial value of the preset context model is preset, or is determined according to a parameter of a context model after encoding of sample video data, or is determined according to a parameter of a context model after encoding of a previous image frame.
13. The method of claim 4, wherein, The method further comprises: updating a context model offset according to the context index.
14. The method of any one of claims 1-13, wherein, The syntax element comprises an inter-component linear prediction mode.
15. A method of video encoding, wherein, The method comprises: obtaining a first identifier, the first identifier being an identifier of any syntax element of a reference block corresponding to a current block, the current block being a unit for encoding in any frame image of video data, the reference block being associated with the current block in the corresponding frame image in position; determining a context model of the current block according to the first identifier, the context model being used for entropy encoding of the identifier of the syntax element.
16. A video decoding apparatus, comprising: comprising: obtaining a first identifier, the first identifier being an identifier of any syntax element of a reference block corresponding to a current block, the current block being a unit for decoding in any frame image of video data, the reference block being associated with the current block in the corresponding frame image in position; determining a context model of the current block according to the first identifier, the context model being used for entropy decoding of the identifier of the syntax element.
17. A video encoding apparatus, wherein, comprising: obtaining a first identifier, the first identifier being an identifier of any syntax element of a reference block corresponding to a current block, the current block being a unit for encoding in any frame image of video data, the reference block being associated with the current block in the corresponding frame image in position; determining a context model of the current block according to the first identifier, the context model being used for entropy encoding of the identifier of the syntax element.
18. An electronic device, comprising: comprising: a memory and a processor, the memory storing a computer program, the processor being configured to implement the method of any one of claims 1-14 or 15 when executing the computer program.
19. A computer readable storage medium having stored thereon a computer program, wherein, The computer program is executed by the processor to implement the method of any one of claims 1-14 or 15.
20. A computer program product, which, when executed on a computer, causes the computer to implement the method of any one of claims 1-14 or 15.
21. A chip, comprising a processor and a memory, the memory being configured to store a program or instructions executable on the processor, and the processor being configured to execute the program or instructions to cause the method of any one of claims 1-14 or 15 to be performed.
Citation Information
Patent Citations
Method and system of video coding with context decoding and reconstruction bypass
CN109565587A
METHOD AND APPARATUS FOR VIDEO CODING, computer equipment and storage medium
CN112135135A
Video decoding method and device and storage medium
CN112399180A
Context model selection method and device, equipment and storage medium
CN114257810A
Performance improvements for geometry point cloud compression (GPCC) planar modes using inter prediction
CN117121492A