Image decoding method, video decoder, and non-temporary readable storage medium
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2023-09-12
- Publication Date
- 2026-08-07
Smart Images

Figure 0007902355000016 
Figure 0007902355000017 
Figure 0007902355000018
Abstract
Description
Technical Field
[0001] This application claims the priority of Chinese Patent Application No. 202211146464.4 filed on September 20, 2022, and all of its contents are incorporated herein by reference.
[0002] This application relates to the field of video encoding and video decoding technologies, and particularly to an image encoding method, an image decoding method, an apparatus, and a storage medium.
Background Art
[0003] To improve the performance of an encoder, a technology called sub-stream parallelism has been proposed.
[0004] Sub-stream parallelism means using multiple entropy encoders to encode syntax elements of different channels to obtain multiple sub-streams, embedding the multiple sub-streams into corresponding sub-stream buffers respectively, and interleaving the sub-streams in the sub-stream buffers into a bitstream (also called a code stream) according to a preset interleaving rule.
[0005] However, considering the dependencies between sub-streams, the embedding speeds of sub-streams in different sub-stream buffers are different. At the same time, the number of sub-stream bits embedded in the sub-stream buffer with a faster embedding speed is larger than that in the sub-stream buffer with a slower embedding speed. Therefore, in order to ensure the integrity of the embedded data, it is necessary to set all sub-stream buffers to be large, which increases the hardware cost.
Summary of the Invention
[0006] Based on the above technical problems, this application provides an image encoding method and an image decoding method, as well as an apparatus and a storage medium, that can encode using multiple improved encoding modes, rationally arrange the substream buffer space, and reduce hardware costs.
[0007] In a first aspect, the present application provides an image coding method, the method being The process involves obtaining an encoding unit, wherein the encoding unit includes encoding blocks for multiple channels, Encoding a preset codeword in a target substream that satisfies a preset condition until the target substream no longer satisfies a preset condition, wherein the target substream is a substream among a plurality of substreams, and the plurality of substreams are code streams obtained by encoding the coding blocks of the plurality of channels.
[0008] In a second aspect, the present application provides an image coding method, the method being The process involves obtaining an encoding unit, wherein the encoding unit includes encoding blocks for multiple channels, The current expansion rate is to encode the coding block of each channel based on the preset expansion rate such that the current expansion rate is less than or equal to the preset expansion rate, wherein the preset expansion rate includes a first preset expansion rate, the value of the current expansion rate is derived from the quotient of the number of bits of the largest substream and the number of bits of the smallest substream, the largest substream is the substream with the largest number of bits among the multiple substreams obtained by encoding the coding blocks of the multiple channels, and the smallest substream is the substream with the smallest number of bits among the multiple substreams obtained by encoding the coding blocks of the multiple channels. Includes.
[0009] In a third aspect, the present application provides an image coding method, the method being The process involves obtaining an encoding unit, wherein the encoding unit includes encoding blocks for multiple channels, Encoding the coding block of at least one channel among the plurality of channels in IBC mode, which is intrablock copy mode, The method involves obtaining the block vector BV of a reference prediction block, wherein the BV of the reference prediction block indicates the position of the reference prediction block in the encoded image block, and the reference prediction block represents the predicted value of the encoded block encoded in IBC mode. The method includes encoding the BV of the reference prediction block in at least one substream obtained by encoding the encoding block of at least one channel in the IBC mode.
[0010] In a fourth aspect, the present application provides an image coding method, the method being The process involves obtaining an encoding unit, wherein the encoding unit is an image block in the image to be processed, and the encoding unit includes encoding blocks for multiple channels. The first total code length is determined, wherein the first total code length is the total code length of a first stream obtained by encoding each of the plurality of channel encoding blocks in the corresponding target encoding mode, the target encoding mode includes a first encoding mode, the first encoding mode is a mode that encodes sample values in an encoding block with a first fixed-length code, the code length of the first fixed-length code is less than or equal to the image bit width of the image to be processed, and the image bit width is for representing the number of bits required to store each sample in the image to be processed. If the total code length of the first mode is greater than or equal to the remaining amount in the code stream buffer, the encoding blocks of the plurality of channels are encoded in fallback mode, wherein the mode flags of the fallback mode are the same as those of the first encoding mode.
[0011] In a fifth aspect, the present application provides an image decoding method, the method being The process involves analyzing a code stream obtained by encoding an encoding unit, wherein the encoding unit includes encoding blocks for multiple channels. The method includes determining the number of codewords, wherein the number of codewords is intended to indicate the number of preset codewords encoded in a target substream that satisfies a predetermined condition, and the preset codewords are encoded in the target substream if such a target substream exists that satisfies a predetermined condition, and decoding the code stream based on the number of codewords.
[0012] In a sixth aspect, the present application provides an image decoding method, the method being The process involves analyzing a code stream obtained by encoding an encoding unit, wherein the encoding unit includes encoding blocks for multiple channels. To determine the current expansion rate, The number of codewords is determined based on the current expansion rate and a first preset expansion rate, wherein the number of codewords is intended to indicate the number of preset codewords encoded in a target substream that satisfies a preset condition, and the preset codewords are encoded in the target substream if such a target substream exists that satisfies the preset condition. The method includes decoding the code stream based on the number of codewords.
[0013] In a seventh aspect, the present application provides an image decoding method, the method being The method involves analyzing a code stream obtained by encoding an encoding unit, wherein the encoding unit includes encoding blocks for multiple channels, and the code stream includes multiple substreams, each corresponding one-to-one to the multiple channels, from which the encoding blocks for multiple channels have been encoded. The method involves determining the position of a reference prediction block in the plurality of substreams based on the block vector BV of the reference prediction block analyzed from at least one of the plurality of substreams, wherein the reference prediction block represents the predicted value of a decoded block decoded in IBC mode, which is an intrablock copy mode, and the BV of the reference prediction block indicates the position of the reference prediction block in the reconstructed image block. Based on the location information of the reference prediction block, the predicted value of the decoded block to be decoded in IBC mode is determined. This includes reconstructing the decrypted block to be decrypted in IBC mode based on the predicted value.
[0014] In an eighth aspect, the present application provides an image decoding method, the method being The process involves analyzing a code stream obtained by encoding an encoding unit, wherein the encoding unit includes encoding blocks for multiple channels. The mode flag is analyzed from the substream obtained by encoding the encoding blocks of the plurality of channels, and if the second total code length is greater than the remaining amount in the code stream buffer, it is determined that the target decoding mode of the substream is the fallback mode, wherein the mode flag indicates whether or not the encoding blocks of the plurality of channels are encoded using the fallback mode, the second total code length is the total code length of the code stream obtained when all of the encoding blocks of the plurality of channels are encoded in the first encoding mode, the first encoding mode is a mode in which the sample values in the encoding block are encoded with a first fixed-length code, the code length of the first fixed-length code is less than or equal to the image bit width of the image to be processed, and the image bit width represents the number of bits required to store each sample in the image to be processed. Analyzing the preset flag bits in the substream to determine a target fallback mode, where the target fallback mode is a type of the fallback mode, and the preset flag bits are for indicating the types of fallback modes used when encoding blocks of the plurality of channels, and the fallback mode includes a first fallback mode and a second fallback mode, including decoding the substream in the target fallback mode.
[0015] As a ninth aspect, the present application provides an image decoding method, and the method includes: obtaining a code stream obtained by encoding an encoding unit, including decoding a substream corresponding to a first channel in a first decoding mode.
[0016] As a tenth aspect, the present application provides an image decoding method, and the method includes: obtaining a decoding unit, where the decoding unit is an image block in a processing target image, and the decoding unit includes encoding blocks of a plurality of channels, determining a first total code length, where the first total code length is the total code length of a first code stream obtained by decoding decoding blocks in the plurality of channels in respective corresponding target decoding modes, when the first total code length is greater than or equal to the remaining amount of the code stream buffer, encoding the decoding blocks in the plurality of channels in a fallback mode.
[0017] As an eleventh aspect, the present application provides an image decoding method, and the method includes: obtaining a decoding unit, where the decoding unit includes encoding blocks of a plurality of channels, decoding decoding blocks in at least one channel of the plurality of channels in an IBC mode. Obtaining the BV of the reference prediction block, and decoding the BV of the reference prediction block in at least one substream obtained by encoding the decoding block in at least one channel in the IBC mode.
[0018] As a twelfth aspect, the present application provides an image encoding apparatus, the apparatus comprising: an acquisition module for acquiring an encoding unit, the encoding unit including encoding blocks of a plurality of channels; a processing module for encoding a preset codeword in a target substream that satisfies preset conditions until the target substream no longer satisfies the preset conditions, the target substream being one of a plurality of substreams, and the plurality of substreams being code streams obtained by encoding the encoding blocks of the plurality of channels.
[0019] As a thirteenth aspect, the present application provides an image encoding apparatus, the apparatus comprising: an acquisition module for acquiring an encoding unit, the encoding unit including encoding blocks of a plurality of channels; a processing module for encoding the encoding blocks of each channel based on the preset expansion rate so that the current expansion rate is less than or equal to the preset expansion rate.
[0020] As a fourteenth aspect, the present application provides an image encoding apparatus, the apparatus comprising: an acquisition module for acquiring an encoding unit, the encoding unit including encoding blocks of a plurality of channels; The system includes a processing module used for encoding an encoded block of at least one channel among the plurality of channels in IBC mode, which is an intrablock copy mode; obtaining a block vector BV of a reference prediction block, wherein the BV of the reference prediction block is for indicating the position of the reference prediction block in the encoded image block, and the reference prediction block is for representing the predicted value of the encoded block encoded in IBC mode; and encoding the BV of the reference prediction block in a substream obtained by encoding the encoded block in each of the plurality of channels in IBC mode.
[0021] In a fifteenth aspect, the present application provides an image coding apparatus, the apparatus is An acquisition module for acquiring an encoding unit, wherein the encoding unit is an image block in the image to be processed, and the encoding unit includes an encoding block for multiple channels, The processing module is used to determine a first total code length, wherein the first total code length is the total code length of a first stream obtained by encoding each of the plurality of channel coding blocks in their respective target coding modes, the target coding mode includes a first coding mode, the first coding mode is a mode that encodes sample values in a coding block with a first fixed-length code, the code length of the first fixed-length code is less than or equal to the image bit width of the image to be processed, and the image bit width is for representing the number of bits required to store each sample in the image to be processed, and if the first total code length is greater than or equal to the remaining amount of the code stream buffer, encode the plurality of channel coding blocks in a fallback mode, wherein the mode flag of the fallback mode is the same as that of the first coding mode.
[0022] In a sixteenth aspect, the present application provides an image decoding device, the device is The processing module is used to analyze a code stream obtained by encoding an encoding unit, wherein the encoding unit includes encoding blocks for multiple channels, and to determine the number of codewords, the number of codewords being used to indicate the number of preset codewords encoded in a target substream that satisfies a set of conditions, the preset codewords being encoded in the target substream if such a target substream exists, and to decode the code stream based on the number of codewords.
[0023] In a 17th aspect, the present application provides an image decoding device, the device is The processing module is used to analyze a code stream obtained by encoding an encoding unit, wherein the encoding unit includes encoding blocks of multiple channels, to determine the current expansion rate, to determine the number of codewords based on the current expansion rate and a first preset expansion rate, wherein the number of codewords is used to indicate the number of preset codewords encoded in a target substream that satisfies a preset condition, the preset codewords are encoded in the target substream if such a target substream exists, and to decode the code stream based on the number of codewords.
[0024] In its eighteenth aspect, the present application provides an image decoding device, the device is The processing module is used to analyze a code stream obtained by encoding an encoding unit, wherein the encoding unit includes encoding blocks for multiple channels, and the code stream includes multiple substreams that correspond one-to-one to the multiple channels, in which the encoding blocks for multiple channels are encoded; to determine the position of the reference prediction block in the multiple substreams based on the block vector BV of the reference prediction block analyzed from at least one of the multiple substreams, wherein the reference prediction block is for representing the predicted value of a decoded block to be decoded in IBC mode, which is an intrablock copy mode, and the BV of the reference prediction block is for indicating the position of the reference prediction block in the reconstructed image block; to determine the predicted value of the decoded block to be decoded in IBC mode based on the position information of the reference prediction block; and to reconstruct the decoded block to be decoded in IBC mode based on the predicted value.
[0025] In a 19th aspect, the present application provides an image decoding device, the device is The method involves analyzing a code stream obtained by encoding an encoding unit, wherein the encoding unit includes encoding blocks of multiple channels, and a mode flag is analyzed from a substream obtained by encoding the encoding blocks of multiple channels, and if the second total code length is greater than the remaining amount in the code stream buffer, it is determined that the target decoding mode of the substream is a fallback mode, wherein the mode flag indicates whether or not the decoding blocks in the multiple channels are encoded using the fallback mode, the second total code length is the total code length of the code stream obtained by encoding all of the encoding blocks of the multiple channels in the first encoding mode, and the first encoding mode is a first fixed-length code with respect to the sangence in the encoding block. A processing module is provided which is used to encode the pull value, wherein the code length of the first fixed-length code is less than or equal to the image bit width of the image to be processed, the image bit width is for representing the number of bits required to store each sample in the image to be processed, and to analyze the preset flag bits in the substream to determine the target fallback submode, wherein the target fallback mode is a type of the fallback mode, the preset flag bits are for indicating the type of fallback mode used when the encoding blocks of the plurality of channels are encoded, the fallback mode includes a first fallback mode and a second fallback mode, and to decode the substream in the target fallback mode.
[0026] In a 20th aspect, the present invention provides a video encoder comprising a processor and a memory, the memory storing instructions that can be executed by the processor, and the processor is configured such that when it executes an instruction, it causes the video encoder to implement the image encoding method described in any one of the first to fourth aspects.
[0027] In a 21st aspect, the present invention provides a video decoder comprising a processor and a memory, the memory storing instructions that can be executed by the processor, and the processor is configured such that when it executes an instruction, it causes the video decoder to implement the image decoding method described in any one of the 5th to 11th aspects.
[0028] In a 22nd aspect, the present invention provides a computer program product, which, when executed by an image encoding device, causes the image encoding device to implement the image encoding method described in any one of the first to fourth aspects described above.
[0029] In a 23rd aspect, the present invention provides a readable storage medium which includes a software instruction, which, when executed by an image encoding device, causes the image encoding device to implement the image encoding method described in the first to fourth aspects, and when executed by an image decoding device, causes the image decoding device to implement the image decoding method described in any one of the fifth to eleventh aspects.
[0030] In a 24th aspect, the present invention provides a chip comprising a processor and an interface, wherein the processor is coupled to memory via the interface, and when the processor executes a computer program in memory or an image encoding device executes an instruction, the method described in the first to fourth aspects is performed. [Brief explanation of the drawing]
[0031] To more clearly explain the technical means of the embodiments of this application, the necessary drawings for the embodiments are briefly described below. Note that the drawings described below represent only some embodiments of this application, and those skilled in the art can obtain other drawings based on these without requiring any creative effort.
[0032] [Figure 1] Figure 1 is a schematic diagram of the configuration of the substream parallel technology. [Figure 2]Figure 2 is a schematic diagram of the substream interleaving process on the encoding side. [Figure 3] Figure 3 is a schematic diagram of the substream interleaved unit format. [Figure 4] Figure 4 is a schematic diagram of the inverse substream interleaving process on the decoding side. [Figure 5] Figure 5 is another schematic diagram of the substream interleaving process on the encoding side. [Figure 6] Figure 6 is a schematic diagram of the configuration of a video encoding and decoding system provided by an embodiment of the present invention. [Figure 7] Figure 7 is a schematic diagram of the configuration of a video encoder provided by an embodiment of the present invention. [Figure 8] Figure 8 is a schematic diagram of the configuration of a video decoder provided by an embodiment of the present invention. [Figure 9] Figure 9 is a schematic flowchart of video encoding and decoding provided by an embodiment of the present invention. [Figure 10] Figure 10 is a schematic diagram of the configuration of an image encoding device and an image decoding device provided by an embodiment of the present invention. [Figure 11] Figure 11 is a schematic flowchart of the image coding method provided by an embodiment of the present invention. [Figure 12] Figure 12 is a schematic flowchart of the image decoding method provided by the embodiment of the present invention. [Figure 13] Figure 13 is a schematic flowchart of another image decoding method provided by an embodiment of the present invention. [Figure 14] Figure 14 is a schematic flowchart of another image decoding method provided by an embodiment of the present invention. [Figure 15] Figure 15 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present invention. [Figure 16] Figure 16 is a schematic flowchart of yet another image coding method provided by an embodiment of the present invention. [Figure 17]Figure 17 is a schematic flowchart of yet another image coding method provided by an embodiment of the present invention. [Figure 18] Figure 18 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present invention. [Figure 19] Figure 19 is a schematic flowchart of yet another image coding method provided by an embodiment of the present invention. [Figure 20] Figure 20 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present invention. [Figure 21] Figure 21 is a schematic flowchart of yet another image coding method provided by an embodiment of the present invention. [Figure 22] Figure 22 is a schematic flowchart of yet another image coding method provided by an embodiment of the present invention. [Figure 23] Figure 23 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present invention. [Figure 24] Figure 24 is a schematic flowchart of yet another image coding method provided by an embodiment of the present invention. [Figure 25] Figure 25 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present invention. [Figure 26] Figure 26 is a schematic flowchart of yet another image coding method provided by an embodiment of the present invention. [Figure 27] Figure 27 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present invention. [Figure 28] Figure 28 is a schematic diagram of the configuration of an image encoding device provided by an embodiment of the present invention. [Modes for carrying out the invention]
[0033] The technical concepts in the embodiments of this application will be described below clearly and completely with reference to the drawings of the embodiments. Clearly, the embodiments described are not all embodiments, but only a part of the embodiments of this application. All other embodiments that can be obtained by those skilled in the art without any creative effort based on the embodiments of this application are included in the scope of protection of this application.
[0034] In this description, unless otherwise specified, " / " means "or," for example, A / B means A or B. The term "and / or" in this description is solely for the purpose of describing the relationship between the associated objects, indicating that three types of relationships are possible. For example, A and / or B can indicate three situations: A existing alone, A and B existing simultaneously, or B existing alone. The term "at least one" means one or more, and "multiple" means two or more. Words such as "first," "second," etc., do not limit the quantity or order of execution, nor do words such as "first," "second," etc., necessarily limit them to being "different."
[0035] In this application, terms such as “exemplary” or “for example” are used to mean example, illustration, or explanation. Any embodiment or design described “exemplary” or “for example” in this application should not be construed as being preferable or advantageous to other embodiments or designs. More precisely, the use of terms such as “exemplary” or “for example” is intended to express the relevant concepts in a specific form.
[0036] To improve the performance of encoders, a technique called substream parallelism (also known as substream interleaving) has been proposed.
[0037] For the encoding side, substream parallelism means that the encoding side uses multiple entropy encoders to encode the syntactic elements of a coding block (CB) in different channels of a coding unit (CU) (e.g., the luminance channel, the first chromaticity channel, and the second chromaticity channel) to obtain multiple substreams, and then interleaves these multiple substreams into a bitstream in a fixed-size packet. Correspondingly, for the decoding side, substream parallelism means that the decoding side uses different entropy decoders to decode different substreams in parallel.
[0038] Exemplary, Figure 1 is a schematic diagram of the substream parallel technology configuration. As shown in Figure 1, taking the encoding side as an example, the specific application timing of the substream parallel technology is after the syntax elements (e.g., transformation coefficients and quantization coefficients) have been encoded. The encoding flow of the rest of Figure 1 can be found in the description of the video encoding and decoding system provided by the embodiment of this application below, and will not be described further here.
[0039] Exemplary, Figure 2 is a schematic diagram of the substream interleaving process on the encoding side. As shown in Figure 2, taking the example that the image block to be encoded contains three channels, the encoding module (e.g., prediction module, transformation module, and quantization module) outputs the syntactic elements and quantized transformation coefficients of the three channels. Then, the syntactic elements and quantized transformation coefficients of the three channels are encoded by entropy encoder 1, entropy encoder 2, and entropy encoder 3, respectively, to obtain substreams corresponding to each channel, and these substreams corresponding to each of the three channels can be placed into encoded substream buffer 1, encoded substream buffer 2, and encoded substream buffer 3. The substream interleaving module interleaves the substreams in encoded substream buffer 1, encoded substream buffer 2, and encoded substream buffer 3, and can finally output a bitstream (also called a code stream) with multiple interleaved substreams.
[0040] Exemplary, Figure 3 is a schematic diagram of the format of a substream interleaved unit. As shown in Figure 3, a substream may consist of substream interleaved units, which are sometimes called substream segments. A substream segment has a length of N bits and contains an M-bit data header and an NM-bit data body.
[0041] Here, the data header is used to indicate the substream to which the current substream belongs. N can be 512 and M can be 2.
[0042] Illustratively, Figure 4 is a schematic diagram of the reverse substream interleaving process on the decoding side. As shown in Figure 4, taking the example that the image block to be encoded contains three channels, when the bitstream output by the encoding side is input to the decoding side, the substream interleaving module on the decoding side first performs a reverse substream interleaving process on the bitstream, dividing the bitstream into substreams corresponding to each of the three channels, and then placing the substreams corresponding to each of the three channels into decoding substream buffer 1, decoding substream buffer 2, and decoding substream buffer 3. For example, taking the substream segment shown in Figure 3 above as an example, the decoding side can extract an N-bit packet from the bitstream each time. By analyzing the M-bit data header within that packet, the target substream to which the current substream segment belongs is obtained, and the data body remaining in the current substream segment is placed into the decoding substream buffer corresponding to the target substream.
[0043] Entropy decoder 1 decodes the substream in decoding substream buffer 1, obtaining the syntactic elements and quantized transformation coefficients for one channel. Entropy decoder 2 decodes the substream in decoding substream buffer 2, obtaining the syntactic elements and quantized transformation coefficients for another channel. Entropy decoder 3 decodes the substream in decoding substream buffer 3, obtaining the syntactic elements and quantized transformation coefficients for yet another channel. Finally, the syntactic elements and quantized transformation coefficients for each of these three channels are input to a subsequent decoding module for decoding, and the decoded image is obtained.
[0044] The following explains the substream interleaving process using the encoding side as an example.
[0045] Each substream segment in the encoded substream buffer contains encoded bits generated by encoding at least one image block. During substream interleaving, first, the image block corresponding to the first bit of the data body of each substream segment is sequentially marked, and in the substream interleaving process, different substreams can be interleaved using this sequential marking.
[0046] In one embodiment, the substream interleaving process can mark image blocks using a block count queue.
[0047] For example, a block count queue is implemented using first-in, first-out (FIFO) queues. The encoding side sets a block count for the currently encoded image blocks and sets up one block count queue [ss_idx] for each substream. When encoding of each slice begins, the block count is initialized to 0, each counter queue [ss_idx] is initialized to null, and then one 0 is placed in each counter queue [ss_idx].
[0048] After each image block (or coding unit, CU) has been encoded, the block count queue is updated. The update process is as follows:
[0049] Step 1: Increment the current encoded image block count by 1, i.e., set block count += 1.
[0050] Step 2: Select one substream ss_idx.
[0051] Step 3: Calculate the number of substream segments that can be constructed in the encoded substream buffer corresponding to this substream ss_idx, num_in_buffer[ss_idx]. Let buffer[ss_idx] be the encoded substream buffer corresponding to this substream ss_idx, and let buffer[ss_idx].fullness be the amount of data contained in buffer[ss_idx]. Similarly, taking as an example that the size of the substream segment shown in Figure 3 above is N bits and the substream segment contains an M-bit data header, num_in_buffer[ss_idx] can be calculated using the following formula (1).
[0052] JPEG0007902355000001.jpg8122 Here, " / " represents divisibility.
[0053] Step 4: Compare the current block count queue length num_in_queue[ss_idx] with the number of substream segments that can be constructed in the encoded substream buffer num_in_buffer[ss_idx]. If they are equal, put the count of the currently encoded image block into this block count queue, i.e., counter_queue[ss_idx].push(block_count).
[0054] Step 5: Return to Step 2 and process the next substream until all substreams have been processed.
[0055] Once the block count queue update is complete, the encoding side can interleave the substreams within each encoded substream buffer, and the interleaving process is as follows:
[0056] Step 1: Select one substream ss_idx.
[0057] Step 2: Determine whether the amount of data contained in the encoded substream buffer buffer[ss_idx] corresponding to this substream ss_idx, buffer[ss_idx].fullness, is greater than or equal to NM. If it is greater than or equal to NM, perform Step 3. If it is not greater than or equal to NM, perform Step 6.
[0058] Step 3: Determine whether the value of the first element in the block count queue for this substream ss_idx is the minimum value in the block count queue for all substreams. If it is the minimum value, perform Step 4. If it is not the minimum value, perform Step 6.
[0059] Step 4: Construct a substream segment using the data in the current encoded substream buffer. For example, take NM-length data from the encoded substream buffer buffer[ss_idx], add an M-bit data header, set the data header data as ss_idx, combine the M-bit data header and the extracted NM-bit data into an N-bit substream segment, and send the substream segment to the final output bitstream from the encoding side.
[0060] Step 5: Pop up (or remove) the top element of the block count queue for this substream ss_idx, i.e., counter_queue[ss_idx].pop().
[0061] Step 6: Return to Step 1 and process the next substream until all substreams have been processed.
[0062] In one embodiment, if the current image block is the last image block of a slice, the encoding side may also perform the following steps to package the data remaining in the encoded substream buffer after the interleaving process described above.
[0063] Step 1: Determine if there is at least one non-null in all currently encoded substream buffers. If so, proceed to Step 2. If not, exit.
[0064] Step 2: Select one substream ss_idx.
[0065] Step 3: Determine whether the value of the first element of the block count queue for this substream ss_idx is the minimum value of the block count queue for all substreams. If it is the minimum value, perform Step 4. If it is not the minimum value, perform Step 6.
[0066] Step 4: If the amount of data in the encoded substream buffer buffer[ss_idx] corresponding to this substream ss_idx is less than NM bits, fill the encoded substream buffer buffer[ss_idx] with 0s until the data in buffer[ss_idx] reaches NM bits. At the same time, pop up (or delete) the value of the first element of the block count queue for this substream, i.e., counter_queue[ss_idx].pop(), and insert MAX_INT, which represents the maximum value within the data range, i.e., counter_queue[ss_idx].push(MAX_INT).
[0067] Step 5: Construct one substream segment. Step 5 here can refer to Step 4 in the substream interleaving process described above, and will not be explained further here.
[0068] Step 6: Return to Step 2 and process the next substream. Once all substreams have been processed, return to Step 1.
[0069] For illustrative purposes, Figure 5 is another schematic diagram of the substream interleaving process on the encoding side. As shown in Figure 5, taking the example that the image block to be encoded has three channels, the encoding substream buffers can each include encoding substream buffer 1, encoding substream buffer 2, and encoding substream buffer 3.
[0070] The substream segments in encoded substream buffer 1 are 1_1, 1_2, 1_3, and 1_4 from left to right. The flags in block count queue 1 for substream 1 corresponding to encoded substream buffer 1 are 12, 13, 27, and 28, respectively. The substream segments in encoded substream buffer 2 are 2_1 and 2_2 from left to right. The flags in block count queue 2 for substream 2 corresponding to encoded substream buffer 2 are 5 and 71, respectively. The substream segments in encoded substream buffer 3 are 3_1, 3_2, and 3_3 from left to right. The flags in block count queue 3 for substream 3 corresponding to encoded substream buffer 3 are 6, 13, and 25, respectively. The substream interleave module can interleave the substream segments in encoded substream buffer 1, encoded substream buffer 2, and encoded substream buffer 3 in the order of their flags.
[0071] The order of substream segments in the interleaved code stream is as follows: substream segment 2_1 corresponding to minimum flag 5, substream segment 3_1 corresponding to flag 6, substream segment 1_1 corresponding to flag 12, substream segment 1_2 corresponding to flag 13, substream segment 3_2 corresponding to flag 13, substream segment 3_3 corresponding to flag 25, substream segment 1_3 corresponding to flag 27, substream segment 1_4 corresponding to flag 28, and substream segment 2_2 corresponding to flag 71.
[0072] However, multiple substreams have dependencies when interleaved. For example, on the encoding side, the embedding speed of substreams differs in different substream buffers. More substreams are embedded in a faster substream buffer at the same time than in a slower substream buffer. The faster substream buffer continues to embed data simultaneously while the slower substream buffer waits for one substream segment to be embedded. To ensure the integrity of the embedded data, all substream buffers must be set to a large size, increasing hardware costs.
[0073] In this case, embodiments of the present invention provide an image encoding method, an image decoding method, an apparatus, and a storage medium. By encoding with multiple improved encoding modes, the number of bits in the maximum encoding module and the number of bits in the minimum encoding module can be controlled. This reduces the theoretical expansion rate of the encoded target block and reduces the difference in embedding speed between different substream buffers, thereby reducing the size of the substream buffer's preset space and lowering hardware costs.
[0074] The following explanation will be given with reference to the diagrams.
[0075] Figure 6 is a schematic diagram of the configuration of a video encoding and decoding system provided by an embodiment of the present invention. As shown in Figure 6, the video encoding and decoding system comprises a source device 10 and a target device 11.
[0076] Source device 10 generates encoded video data and is also called the encoding side, video encoding side, video encoding device, or video encoding device. Target device 11 can decode the encoded video data generated by source device 10 and is also called the decoding side, video decoding side, video decoding device, or video decoding device. Source device 10 and / or target device 11 may comprise at least one processor and memory coupled to the at least one processor. This memory may include, but is not limited to, read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or any other medium usable to store the necessary program code in the form of computer-accessible instructions or data structures.
[0077] The source device 10 and target device 11 may include various devices, such as desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, mobile phones such as so-called "smartphones," televisions, cameras, display devices, digital media players, video game consoles, in-car computers, or other similar devices.
[0078] The target device 11 can receive encoded video data from the source device 10 via link 12. Link 12 may comprise one or more media and / or devices that can move the encoded video data from the source device 10 to the target device 11. In one example, link 12 may include one or more communication media that enable the source device 10 to directly transmit encoded video data to the target device 11 in real time. In this example, the source device 10 may modulate the encoded video data and transmit the modulated video data to the target device 11 based on a communication standard (e.g., a wireless communication protocol). The one or more communication media may include wireless and / or wired communication media such as a radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network such as a local area network, a wide area network, or a global network (e.g., the Internet). The one or more communication media may comprise routers, switches, base stations, or other devices that enable communication from the source device 10 to the target device 11.
[0079] In another example, source device 10 can output encoded video data to storage device 13 via output interface 103. Similarly, target device 11 can access encoded video data from storage device 13 via input interface 113. Storage device 13 can include various local access data storage media such as Blu-ray discs, high-density digital video discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, or other suitable digital storage media for storing encoded video data.
[0080] In another example, the storage device 13 may correspond to a file server or another intermediate storage device that stores encoded video data generated by the source device 10. In this example, the target device 11 can stream or download the video data stored from the storage device 13. The file server may be any type of server that stores encoded video data and can send the encoded video data to the server on the target device 11. For example, the file server may comprise a World Wide Web (Web) server (e.g., for a website), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, and a local disk drive.
[0081] The target device 11 can access the encoded video data via a data connection of any standard (e.g., an internet connection). Examples of data connections include a wireless channel, a wired connection (e.g., a cable modem), or a combination of both, suitable for accessing encoded video data stored on a file server. The method by which the encoded video data is transmitted from the file server may be streaming, downloading, or a combination of both.
[0082] Furthermore, the image encoding method and image decoding method provided by the embodiments of this application are not limited to wireless application scenarios.
[0083] Exemplary examples, the image encoding and decoding methods provided by embodiments of the present application can be applied to video encoding and decoding used in a variety of multimedia applications, such as aerial television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (e.g., over the Internet), encoding of video data stored on a data storage medium, decoding of video data stored on a data storage medium, or other applications. In some examples, the video encoding and decoding system may be configured to support one-way or two-way video transmission, such as video streaming, video playback, video broadcasting, and / or video phone calls.
[0084] The video encoding and decoding system shown in Figure 6 is merely an example of a video encoding and decoding system and does not limit the video encoding and decoding systems described herein. The image encoding method and image decoding method provided herein can also be applied to scenes where there is no data communication between the encoding device and the decoding device. In other embodiments, the video data to be encoded or the encoded video data may be retrieved from local memory or streamed over a network. The video encoding device can encode the video data to be encoded and store the encoded video data in memory. The video decoding device can retrieve the encoded video data from memory and decode this encoded video data.
[0085] In the embodiment shown in Figure 6, the source device 10 includes a video source 101, a video encoder 102, and an output interface 103. In some embodiments, the output interface 103 may include a modulator / demodulator (modem) and / or transmitter. The video source 101 may include a video capture device (e.g., a video camera), a video archive containing captured video data, a video input interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources of video data.
[0086] The video encoder 102 can encode video data from the video source 101. In some embodiments, the source device 10 transmits the encoded video data directly to the target device 11 via the output interface 103. In other embodiments, the encoded video data may be stored in a storage device 13, which the target device 11 can then access for decoding and / or playback.
[0087] In the embodiment shown in Figure 6, the target device 11 includes a display device 111, a video decoder 112, and an input interface 113. In some embodiments, the input interface 113 includes a receiver and / or modem. The input interface 113 can receive encoded video data via link 12 and / or from storage device 13. The display device 111 may be integrated with the target device 11 or may be external to the target device 11. Generally, the display device 111 displays the decoded video data. The display device 111 can include a variety of displays, such as a liquid crystal display, a plasma display, an organic light-emitting diode display, or other types of displays.
[0088] In one embodiment, the video encoder 102 and video decoder 112 can be integrated with an audio encoder and decoder, respectively, and may include a suitable multiplexer-multiplexer unit or other hardware and software for encoding both audio and video in a common data stream or independent data streams.
[0089] The video encoder 102 and video decoder 112 may include at least one microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. If the encoding method provided herein is implemented by software, the instructions used in the software can be stored in a suitable non-volatile computer-readable storage medium, and the instructions can be executed by at least one processor to implement the invention.
[0090] The video encoder 102 and video decoder 112 in this application may operate in accordance with a video compression standard (e.g., HEVC) or in accordance with other industry standards, and this application is not particularly limited thereto.
[0091] Figure 7 is a schematic diagram of the configuration of a video encoder 102 provided by an embodiment of the present invention. As shown in Figure 7, the video encoder 102 can perform the processes of prediction, transformation, quantization, entropy coding, and substream interleaving in each of the following modules: prediction module 21, transformation module 22, quantization module 23, entropy coding module 24, coding substream buffer 25, and substream interleaving module 26. Here, the prediction module 21, transformation module 22, and quantization module 23 are the coding modules in Figure 1 above. The video encoder 102 further includes a preprocessing module 20 and an adder 202, the preprocessing module 20 comprising a splitting module and a code rate control module. To reconstruct the video block, the video encoder 102 includes an inverse quantization module 27, an inverse transformation module 28, an adder 201, and a reference image memory 29.
[0092] As shown in Figure 7, the video encoder 102 receives video data, and the preprocessing module 20 receives input parameters for the video data. These input parameters include information such as the resolution of the video data, the sampling format of the image, the pixel depth (bits per pixel, BPP), and the bit width (also called the image bit width). Here, BPP refers to the number of bits occupied per pixel. Bit width refers to the number of bits occupied by one pixel channel within a unit pixel. For example, if one pixel is represented by the values of three YUV pixel channels, and each pixel channel occupies 8 bits, then the bit width of that pixel is 8, and the BPP of that pixel is 3 × 8 = 24 bits.
[0093] The splitting module in the preprocessing module 20 divides the image into original blocks (sometimes called coding units (CUs)). An original block (sometimes called a coding unit (CU)) may include coding blocks for multiple channels. For example, these multiple channels may be RGB channels or YUV channels, etc. Embodiments of the present application are not limited thereto. In one embodiment, this splitting may include dividing into slices, image blocks or other larger units, and video block splitting by a Largest Coding Unit (LCU) and a four-tree structure of CUs. Exemplaryly, the video encoder 102 encodes components of video blocks within the video slice to be encoded. Generally, a slice can be divided into multiple original blocks (it can be divided into a set of original blocks called image blocks). Typically, the sizes of CUs, PUs, and TUs are determined in the splitting module. The splitting module is also used to determine the size of the code rate control unit. This code rate control unit is a fundamental processing unit in the code rate control module. For example, the code rate control module calculates complexity information for the original block based on the code rate control unit, and then calculates the quantization parameters of the original block based on the complexity information. Here, the partitioning policy of the partitioning module may be pre-set or it may be adjusted sequentially based on the image during encoding. If the partitioning policy is a pre-set policy, the same partitioning policy is pre-set on the decoding side accordingly, thereby obtaining the same image processing unit. This image processing unit is one of the image blocks mentioned above, and it corresponds one-to-one with the encoding side. If the partitioning policy is adjusted sequentially based on the image during encoding, this partitioning policy can be directly or indirectly incorporated into the code stream, and accordingly the decoding side obtains the corresponding parameters from the code stream, obtains the same partitioning policy, and obtains the same image processing unit.
[0094] The code rate control module in the preprocessing module 20 is used to generate quantization parameters so that the quantization module 23 and the inverse quantization module 27 can perform correlation calculations. Here, the code rate control module can acquire and calculate image information of the original block, such as the input information described above, during the process of calculating the quantization parameters, or it can acquire and calculate the reconstructed value reconstructed by the adder 201, and the present invention is not limited to this.
[0095] The prediction module 21 provides the prediction block to the adder 202 to generate a residual block, and can also provide this prediction block to the adder 201 to be reconstructed and obtain a reconstructed block, which is used for the reference pixels to be predicted later. Here, the video encoder 102 forms a pixel difference value by subtracting the pixel value of the prediction block from the pixel value of the original block, and this pixel difference value is a residual block, and the data in this residual block may include luminance difference and chromaticity difference. The adder 201 represents one or more components that perform this subtraction. The prediction module 21 can also send the relevant syntactic elements to the entropy encoding module 24 for merging into the code stream.
[0096] The conversion module 22 can convert the residual block by dividing it into one or more TUs. The conversion module 22 can convert the residual block from the pixel domain to the conversion domain (e.g., the frequency domain). For example, the residual block can be converted into conversion coefficients using a discrete cosine transform (DCT) or a discrete sine transform (DST). The conversion module 22 can transmit the obtained conversion coefficients to the quantization module 23.
[0097] The quantization module 23 can perform quantization based on a quantization unit, where the quantization unit may be the same as the CU, TU, and PU described above, or it may be further divided in the partitioning module. The quantization module 23 quantizes the conversion coefficients to further reduce the code rate and obtain quantization coefficients. Here, the quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be changed by adjusting the quantization parameters. In some viable embodiments, the quantization module 23 can then perform a scan of a matrix containing the quantization conversion coefficients. The entropy coding module 24 can perform the scan instead.
[0098] After quantization, the entropy coding module 24 can entropy code the quantization coefficients. For example, the entropy coding module 24 can perform context-adaptive variable-length coding (CAVLC), context-based adaptive binary arithmetic coding (CABAC), syntax-based binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) decoding, or other entropy coding methods or techniques. Multiple substreams can be obtained by entropy coding with the entropy coding module 24, and these substreams are interleaved by the coding substream buffer 25 and the substream interleaving module 26 to obtain a code stream, which is then sent to the video decoder 112 or to the archive for later transmission or retrieval by the video decoder 112.
[0099] The substream interleaving process can be explained in Figures 1 to 5 above, and will not be explained further here.
[0100] The inverse quantization module 27 and the inverse transform module 28 apply inverse quantization and inverse transform, respectively. The adder 201 adds the residual block after the inverse transform with the predicted residual block to generate a reconstructed block, which is then used as a reference pixel to predict the original block. This reconstructed block is stored in the reference image memory 29.
[0101] Figure 8 is a schematic diagram of the configuration of a video decoder 112 provided by an embodiment of the present invention. As shown in Figure 8, the video decoder 112 comprises a substream interleaving module 30, a decoding substream buffer 31, an entropy decoding module 32, a prediction module 33, an inverse quantization module 34, an inverse transform module 35, an adder 301, and a reference image memory 36.
[0102] Here, the entropy decoding module 32 comprises an analysis module and a code rate control module. In some executable embodiments, the video decoder 112 can perform a decoding flow that is exemplary the reverse of the coding flow with respect to the video encoder 102 as shown in Figure 7.
[0103] In the decoding process, the video decoder 112 receives the encoded video code stream from the video encoder 102. The substream interleaving module 30 then performs reverse substream interleaving on this code stream to obtain multiple substreams. These substreams then pass through their respective decoding substream buffers and enter their respective entropy decoding modules 32. Using one substream as an example, the analysis module in the entropy decoding module 32 of the video decoder 112 entropy decodes the substream to generate quantization coefficients and syntactic elements. The entropy decoding module 32 then transfers the syntactic elements to the prediction module 33. The video decoder 112 can receive syntactic elements at the video slice level and / or video block level.
[0104] In the entropy decoding module 32, the code rate control module generates quantization parameters based on the information of the image to be decoded obtained by the analysis module, so that the inverse quantization module 34 can perform the relevant calculations. The code rate control module can also calculate quantization parameters based on the reconstructed block reconstructed by the adder 301.
[0105] The inverse quantization module 34 inversely quantizes (e.g., dequantizes) the quantization coefficients provided to the substream and decoded by the entropy decoding module 32, and the generated quantization parameters. The inverse quantization process may include determining the degree of quantization using the quantization parameters calculated for each video block in the video slice using the video encoder 102, and similarly determining the degree of inverse quantization applied. The inverse transform module 35 applies an inverse transform (e.g., a transform method such as DCT or DST) to the inversely quantized transform coefficients to generate residual blocks in which the inversely quantized transform coefficients are inversely transformed into pixel regions by the inverse transform unit. Here, the size of the inverse transform unit is the same as the size of the TU, and the inverse transform method and transform method can utilize the corresponding forward and inverse transforms in the same transform method; for example, the inverse transform of DCT or DST is inverse DCT, inverse DST, or a conceptually similar inverse transform process.
[0106] If the prediction module 33 generates a prediction block, the video decoder 112 adds the prediction block to the residual block after the inverse transformation from the inverse transformation module 35 to form a decoded video block. The adder 301 represents one or more components that perform this addition operation. If necessary, a deblocking filter can be applied to filter the decoded blocks to remove block effect artifacts. The decoded image blocks in a frame or image are stored in the reference image memory 36 as reference pixels to be predicted later.
[0107] Embodiments of the present invention provide a feasible method for video (image) encoding and decoding. Referring to Figure 9, Figure 9 is a schematic flowchart of the video encoding and decoding provided by embodiments of the present invention. This feasible method for video encoding and decoding includes processes 1 to 5, which can be performed by one or more of the source device 10, video encoder 102, target device 11, or video decoder 112.
[0108] Process 1: Divide a single frame of an image into one or more parallel coding units that do not overlap with each other. There are no dependencies between these one or more parallel coding units, and they can be coded and decoded completely in parallel / independently, for example, parallel coding unit 1 and parallel coding unit 2.
[0109] Process 2: Each parallel coding unit can be further divided into one or more independent coding units that do not overlap with each other. These independent coding units do not need to be dependent on each other, but they can share header information from several parallel coding units.
[0110] An independent coding unit may include three channels: luminance Y, first chromaticity Cb, and second chromaticity Cr, or three channels: RGB, or it may include only one of these channels. If an independent coding unit includes three channels, the sizes of these three channels may be exactly the same or different, specifically relating to the image input format. This independent coding unit can also be understood as one or more processing units formed by the N channels contained in each parallel coding unit. For example, the three channels Y, Cb, and Cr mentioned above are the three channels that constitute this parallel coding unit, each of which may be an independent coding unit, or Cb and Cr may be collectively called chromaticity channels. In this case, this parallel coding unit includes an independent coding unit consisting of luminance channels and an independent coding unit consisting of chromaticity channels.
[0111] Process 3: Each independent coding unit can be further divided into one or more non-overlapping coding units, and each coding unit within an independent coding unit can depend on the others. For example, multiple coding units can refer to each other and pre-code or decode.
[0112] If the coding unit is the same size as the independent coding unit (i.e., the independent coding unit is divided into only one coding unit), then its size may be any of the sizes described in process 2.
[0113] The encoding unit may have three channels (or three channels of RGB) including luminance Y, first chromaticity Cb, and second chromaticity Cr, or it may contain only one of these channels. If it contains three channels, the sizes of some of the channels may be exactly the same or different, specifically relating to the image input format.
[0114] Process 3 is an optional step in the video encoding / decoding method, in which the video encoder / decoder can encode / decode the residual coefficients (or residual values) of the independent encoding units obtained in Process 2.
[0115] Process 4: The encoded unit can be further divided into one or more non-overlapping prediction groups (PGs), which can also be abbreviated as groups. Each PG is encoded in a selected prediction mode to obtain the predicted values of the PGs, which constitute the predicted values for the entire encoded unit. Based on the predicted values and the original values of the encoded unit, the residual values of the encoded unit are obtained.
[0116] Process 5: Based on the residual values of the encoded units, the encoded units are divided into groups to obtain one or more non-overlapping residual blocks (RBs). The residual coefficients of each RB are encoded in a selected mode to form a residual coefficient stream. Specifically, the residual coefficients can be divided into those that are transformed and those that are not.
[0117] Here, the selection mode for the residual coefficient coding / decoding method in process 5 may include, but is not limited to, any one of the following: semi-fixed-length coding, exponential Columbus (Golomb) coding, Golomb-Rice coding, truncated unary coding, run-length coding, or directly coding the original residual values.
[0118] For example, a video encoder can directly encode the coefficients within the RB.
[0119] For example, a video encoder can also transform residual blocks and further encode the resulting coefficients, using transformation methods such as DCT, DST, and Hadamard transform.
[0120] As a preferred example, when the RB is small, the video encoder can directly quantize each coefficient in the RB as a whole and then binarize it. When the RB is large, it can be further divided into multiple coefficient groups (CGs), each CG can be uniformly quantized, and then binarized. In some embodiments of the present application, the coefficient groups (CGs) and quantization groups (QGs) may be the same.
[0121] The following describes encoding residual coefficients using a semi-fixed-length coding scheme. First, the maximum absolute value of residuals within a single RB block is defined as the modified maximum (mm). Next, the number of bits used to encode the residual coefficients within that RB block (the number of bits used to encode residual coefficients within the same RB block is the same) is determined. For example, if the critical limit (CL) of the current RB block is 2 and the current residual coefficient is 1, encoding a residual coefficient of 1 requires 2 bits, which are represented by 01. If the CL of the current RB block is 7, it means encoding an 8-bit residual coefficient and a 1-bit sign bit. Determining the CL involves finding the smallest value of M within the range [-2^(M-1), 2^(M-1)] for all residuals in the current subblock. If two boundary values, -2^(M-1) and 2^(M-1), exist simultaneously, M should increase by 1, meaning that all residuals in the current RB block must be encoded with M+1 bits. If only one of the two boundary values, -2^(M-1) and 2^(M-1), exists, one trailing bit needs to be encoded to determine whether this boundary value is 2^(M-1) or 2^(M-1). If neither -2^(M-1) nor 2^(M-1) exists in any of the residuals, this trailing bit does not need to be encoded.
[0122] Furthermore, in special cases, video encoders can directly encode the original values of the image rather than the residual values.
[0123] The video encoder 102 and video decoder 112 described above can also be implemented by other embodiments, such as a general-purpose digital processing system. Referring to Figure 10, Figure 10 is a schematic diagram of the configuration of an encoding / decoding device provided by an embodiment of the present invention, which may be some of the devices in the video encoder 102 or some of the devices in the video decoder 112. This encoding / decoding device is applicable to both the encoding side (or encoding end) and the decoding side (or decoding end). As shown in Figure 10, this encoding / decoding device comprises a processor 41 and a memory 42. The processor 41 is connected to the memory 42 (for example, connected to each other via a bus 43). In one embodiment, this encoding / decoding device further comprises a communication interface 44, which connects the processor 41 and the memory 42 for sending and receiving data.
[0124] The processor 41 executes instructions stored in the memory 42 to implement the image encoding method and image decoding method provided by the embodiments of the present invention described below. The processor 41 may be a central processing unit (CPU), a general-purpose processor, a network processor (NP), a digital signal processing (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 41 may also be a device having any other processing function, such as a circuit, device, or software module, and the embodiments of the present invention are not limited thereto. In one example, the processor 41 may include one or more CPUs, such as CPU0 and CPU1 in Figure 10. As an optional embodiment, the electronic device may have multiple processors, for example, a processor 45 (illustrated by a dashed line in Figure 10) in addition to the processor 41.
[0125] Memory 42 is used to store instructions. For example, the instructions may be computer programs. In one embodiment, memory 42 may be read-only memory (ROM) or other types of static storage devices capable of storing static information and / or instructions, random access memory (RAM) or other types of dynamic storage devices capable of storing information and / or instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc storage devices, optical disc storage devices (including compressed discs, laser discs, optical discs, digital multipurpose discs, Blu-ray discs, etc.), disk storage media, or other magnetic storage devices, but embodiments of the present application are not limited thereto.
[0126] The memory 42 may exist independently of the processor 41, or it may be integrated with the processor 41. The memory 42 may be located within the encoding / decoding device or outside the encoding / decoding device, and the embodiments of this application are not limited to this.
[0127] Bus 43 is used to transmit information between the various components of the encoding / decoding device. Bus 43 may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Bus 43 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 shows the bus with only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0128] The communication interface 44 is for communicating with other devices or other communication networks. These other communication networks may be Ethernet, radio access networks (RAN), wireless local area networks (WLAN), etc. The communication interface 44 may be a module, circuit, transceiver, or any communication-capable device. Embodiments of the present invention are not limited thereto.
[0129] Note that the configuration shown in Figure 10 is not intended to limit the encoding / decoding device. In addition to the configuration shown in Figure 10, the encoding / decoding device may include more or fewer components than those shown, or combinations of several components, or different components.
[0130] The entity that executes the image encoding method and image decoding method provided by the embodiments of this application may be the encoding device, an application (APP) that provides encoding functionality installed on the encoding device, the CPU in the encoding / decoding device, or a functional block in the encoding / decoding device for executing the image encoding method and image decoding method. The embodiments of this application are not limited to these. For the sake of explanation, the following will be described as the encoding side or the decoding side.
[0131] The image encoding method and image decoding method provided by the embodiments of this application will be described below with reference to the drawings.
[0132] As explained in the background technology and in Figures 1 to 5 above, the substream embedding speed differs in different substream buffers. In the same amount of time, more substream bits are embedded in the faster substream buffer than in the slower substream buffer. Therefore, to ensure the integrity of the embedded data, all substream buffers must be set to a large size, which increases hardware costs.
[0133] To rationally arrange the substream buffer, embodiments of the present invention propose reducing the expansion rate of the coding unit with a series of improved coding schemes (e.g., coding mode / prediction mode, transmission of complexity information, coefficient grouping, code stream arrangement, etc.), thereby reducing the speed differences when coding blocks of each channel of the coding unit are coded, rationally arranging the substream buffer space, and reducing hardware costs.
[0134] Here, the expansion rate of a coding unit can include the theoretical expansion rate and the current (CU) expansion rate (actual expansion rate). The theoretical expansion rate is obtained by theoretical derivation after it has been determined by coding and decoding, and its value is greater than 1. If a coding unit has channels in which the number of bits in the substream obtained by coding a coding block is 0, this case must be excluded when calculating the current expansion rate. Theoretical expansion rate = number of bits of the CB with the largest theoretical number of bits in the CU / number of bits of the CB with the smallest theoretical number of bits in the CU. Current (CU) expansion rate = number of bits of the CB with the largest actual number of bits in the current CU / number of bits of the CB with the smallest actual number of bits in the current CU.
[0135] In one embodiment, the relevant expansion ratio may also include the current substream expansion ratio. Current substream expansion ratio = number of bits of data in the substream buffer with the largest number of bits of data in the current substream buffers / number of bits of data in the substream buffer with the smallest number of bits of data in the current substream buffers.
[0136] Furthermore, when the encoding side selects an encoding mode (or prediction mode) for each channel's encoding block or entire encoding unit, it can do so based on several policies as described below. Each improved encoding mode in the following embodiments can be selected and determined based on several policies as described below.
[0137] Policy 1: Calculate the bit consumption cost and select the encoding mode with the lowest bit consumption cost.
[0138] Here, bit consumption cost refers to the number of bits required to encode / decode a CU. Bit consumption cost mainly includes the code length of the mode flag (or mode codeword), the code length of the encoding / decoding tool information codeword, and the code length of the residual codeword.
[0139] Policy 2: Calculate the encoded distortion and select the encoding mode with the least distortion.
[0140] Here, distortion is used to indicate the difference between the reconstructed value and the original value. Distortion can be calculated using one or more of the following calculations: sum of squared difference (SSD), mean squared error (MSE), sum of absolute difference (SAD) in the time domain, sum of absolute transformed difference (SATD) in the frequency domain, and peak signal to noise ratio (PSNR). Embodiments of this application are not limited to these.
[0141] Policy 3: Calculate the rate distortion cost and select the encoding mode with the lowest rate distortion cost.
[0142] Here, rate distortion cost refers to the weighted sum of bit consumption cost and distortion. When performing the weighting calculation, both the weighting coefficient for bit consumption cost and the weighting coefficient for encoded distortion can be pre-set on the encoding side. The embodiments of this application do not limit the specific numerical values of the weighting coefficients.
[0143] The improved encoding modes in the image encoding method and image decoding method provided by the embodiments of the present application will be described below using a series of embodiments.
[0144] 1. Change in expansion prevention mode. Possible implementation: The encoding side encodes the image bit width as a fixed-length code for each pixel value in the encoding block of multiple channels of the encoding unit.
[0145] Proposed improvement: The encoding side encodes each pixel value of the encoding block of multiple channels in the encoding unit with a fixed-length code less than or equal to the image bit width.
[0146] Example 1: As an example, Figure 11 is a schematic flowchart of an image encoding method provided by an embodiment of the present invention. As shown in Figure 11, the image encoding method includes S101 to S102.
[0147] S101: The encoding side acquires the encoding unit.
[0148] Here, the encoding side is the source device 10 in Figure 6, or the video encoder 102 in the source device 10, or the encoding device in Figure 10. Embodiments of the present invention are not limited to these. The encoding unit is an image block in the image to be processed (i.e., the original block). The encoding unit includes encoding blocks for multiple channels, each of which includes a first channel, and the first channel is one of the multiple channels. For example, the size of the encoding unit may be 16 × 2 × 3, in which case the size of the encoding block for the first channel of the encoding unit is 16 × 2.
[0149] S102: The encoding side encodes the encoding block of the first channel in the first encoding mode.
[0150] Here, the first encoding mode is a mode in which the sampled values of the encoded blocks of the first channel are encoded based on a first fixed-length code. The code length of the first fixed-length code is less than or equal to the image bit width of the image to be processed. The image bit width represents the number of bits required to store each sample in the image to be processed. The first fixed-length code may be pre-set on the encoding / decoding side, or it may be determined by the encoding side, written to the header information of the code stream, and transmitted to the decoding side.
[0151] The anti-expansion mode is a mode in which the original pixel values within the coding blocks of multiple channels are directly coded. Compared to other coding modes, the number of bits consumed to directly code the original pixel values is usually large. Therefore, it should be understood that the coding block with the largest theoretical number of bits in the theoretical expansion rate usually originates from the coding block coded according to the anti-expansion mode. Embodiments of the present invention reduce the theoretical expansion rate by reducing the coding length of the fixed-length code in the anti-expansion mode, thereby reducing the theoretical number of bits (i.e., the numerator of the fractional expression) of the coding block with the largest theoretical number of bits in the theoretical expansion rate. When the theoretical expansion rate is low, the difference in the speed at which each channel's coding block codes the substream is small. Therefore, when setting the substream buffer, it is not necessary to set the substream buffer to be large, nor is the substream buffer space wasted, thereby achieving a rational arrangement of the substream buffer.
[0152] In one embodiment, if the image to be processed is an RGB image, the encoding side can convert it to YUV format and encode it, or if the image to be processed is a YUV image, the encoding side can convert it to RGB format and encode it. The embodiments of the present application are not limited to this.
[0153] In response to the above, embodiments of the present invention further provide an image decoding method. Figure 12 is a schematic flowchart of the image decoding method provided by embodiments of the present invention. As shown in Figure 12, the image decoding method includes S201 to S202.
[0154] S201: The decoding side obtains the code stream obtained by encoding the encoding unit.
[0155] Here, the code stream obtained by encoding the encoding unit may include multiple substreams, each corresponding one-to-one to multiple channels, in which the encoding blocks of multiple channels are encoded. For example, these multiple substreams may include a substream corresponding to the first channel (i.e., a substream in which the encoding block of the first channel is encoded).
[0156] S202: The decoding side decodes the substream corresponding to the first channel in the first decoding mode.
[0157] Here, the first decoding mode is a mode in which sample values are analyzed from the substream corresponding to the first channel based on the first fixed-length code.
[0158] In one embodiment, step S202 further includes, specifically, if the first fixed-length code is equal to the image bit width, the decoding side decodes the substream corresponding to the first channel according to the first decoding mode, and if the first fixed-length code is smaller than the image bit width, the decoding side inversely quantizes the pixel values of the coded blocks of the analyzed first channel.
[0159] Here, the quantization step size is 1<<(bitdepth-fixed_length). bitdepth represents the image bit width. fixed_length represents the length of the first fixed-length code. 1<< indicates shifting one bit to the left.
[0160] 2. Changing the fallback mode. Option 1: If the current code stream buffer is not satisfied with all channels of the encoding unit using the anti-bloat mode, the anti-bloat mode is turned off, and the encoding side selects a mode other than the original value. In this case, even if encoding is done with other encoding modes, bloat will occur (the number of bits in the encoded substream of an encoded block of a channel is too large), and the anti-bloat mode cannot be used.
[0161] Improvement proposal 1: Ensure the anti-expansion mode is open.
[0162] Example 2: Figure 13 is a schematic flowchart of another image decoding method provided by an embodiment of the present invention. As shown in Figure 13, the image decoding method includes steps S301 to S303.
[0163] S301: The encoding side obtains the encoding unit.
[0164] Here, the encoding unit is an image block in the image to be processed, and the encoding unit includes encoding blocks for multiple channels.
[0165] S302: The encoding side determines the first total code length.
[0166] Here, the first total code length is the total code length of the first stream obtained by encoding each of the multiple channel encoding blocks according to their respective target encoding modes. The target encoding mode includes the first encoding mode, which is a mode that encodes the sample values in the encoding block with a first fixed-length code, the code length of the first fixed-length code being less than or equal to the image bit width of the image to be processed, and the image bit width is used to represent the number of bits required to store each sample in the image to be processed.
[0167] For example, the encoding side can determine the rate distortion cost of each encoding mode for each of the multiple channels according to policy 3 above, and determine the mode with the lowest rate distortion cost as the target encoding mode corresponding to the encoding block of the channel.
[0168] S303: If the first total code length is greater than or equal to the remaining amount in the code stream buffer, the encoding side encodes multiple channel coding blocks in fallback mode.
[0169] Here, the fallback mode includes a first fallback mode and a second fallback mode. The first fallback mode involves obtaining the block vector of the reference prediction block in IBC mode, then calculating and quantizing the residuals, with the quantized step size determined based on the remaining code stream buffer and the target pixel depth (bite per pixel, BPP). The second fallback mode involves directly quantizing the pixel points, with the quantized step size determined based on the remaining code stream buffer and the target BPP. The mode flags for the fallback modes are the same as those for the first coding mode.
[0170] Furthermore, when encoding using other encoding modes, the encoded residual may be too large, resulting in an excessively large number of bits in the encoded block using the other encoding mode. However, the anti-bloat mode uses fixed-length codes for encoding, and the total encoded code length is fixed. Therefore, it should be understood that the excessively large residual mentioned above can be avoided by using the anti-bloat mode. In this embodiment of the present invention, when the mode flag of the fallback mode is the same as that of the first encoding mode (anti-bloat mode), the encoding mode used is notified to the decoding side by determining the remaining amount of the code stream buffer. This ensures that the anti-bloat mode is always selectable, thereby reducing the theoretical expansion rate. Because the theoretical expansion rate is low and the difference in the speed at which each channel's encoded block encodes the substream is small, when setting the substream buffer, it is not necessary to set the substream buffer to be large or to waste space in the substream buffer, thereby achieving a rational arrangement of the substream buffer.
[0171] In one embodiment, the image decoding method further includes the encoding side encoding mode flags in multiple substreams obtained by encoding encoding blocks of multiple channels.
[0172] Here, the mode flag indicates the encoding mode used for each of the multiple channel encoding blocks, and the mode flag for the first encoding mode is the same as the fallback mode.
[0173] In one embodiment, the mode flag of a substream can be encoded in a substream in which the encoding blocks of multiple channels have been encoded. For example, if the multiple channels include a first channel and the first channel is one of the multiple channels, then encoding the mode flag in the multiple substreams obtained by encoding the encoding blocks of multiple channels includes encoding the submode flag in the substream obtained by encoding the encoding block of the first channel.
[0174] Here, the submode flag indicates the type of fallback mode used for the first channel's coding block, or the type of fallback mode used for the coding block of multiple channels. As mentioned above, the fallback mode can include a first fallback mode and a second fallback mode. This will not be explained further here.
[0175] In one embodiment, similarly using the first component as an example, the encoding side encodes the mode flag in the substream obtained by encoding the encoding blocks of multiple channels, which includes encoding the first flag, the second flag, and the third flag in the substream obtained by encoding the encoding block of the luminance channel.
[0176] Here, the first flag indicates that the multi-channel coding block will be coded using either the first coding mode or the fallback mode. The second flag indicates that the multi-channel coding block will be coded using the target mode, which is either the first coding mode or the fallback mode. If the second flag indicates that the target mode used for the multi-channel coding block is the fallback mode, the third flag is used to indicate the type of fallback mode used for the multi-channel coding block.
[0177] In response to the above, embodiments of the present invention provide two image decoding methods. Figure 14 is a schematic flowchart of another image decoding method provided by embodiments of the present invention. As shown in Figure 14, the image decoding method includes S401 to S404.
[0178] S401: The decoding side analyzes the code stream obtained by encoding the encoding unit.
[0179] Here, the coding unit includes a coding block of multiple channels. The multiple channels include a first channel, and the first channel is one of the multiple channels.
[0180] S402: The decoding side analyzes the mode flag from the substream obtained by encoding the coded blocks of multiple channels, and if the second total code length is greater than the remaining amount in the code stream buffer, it determines that the target decoding mode of the substream is the fallback mode.
[0181] Here, the second total code length is the total code length of the code stream obtained when the coding blocks of multiple channels are all coded according to the first coding mode or the fallback mode.
[0182] S403: The decoding side analyzes the preset flag bits in the above substream and determines the target fallback mode.
[0183] Here, the target fallback mode is a type of fallback mode. The submode flag indicates the type of fallback mode used when encoding a multi-channel coding block. The fallback mode includes a first fallback mode and a second fallback mode.
[0184] S404: The decryption side decrypts the substream in target fallback mode.
[0185] Figure 15 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present invention. As shown in Figure 15, the image decoding method includes steps S501 to S505.
[0186] S501: The decoding side analyzes the code stream obtained by encoding the encoding unit.
[0187] S502: The decoding side analyzes the first flag from the substream obtained by encoding the coded block of the first channel.
[0188] Here, the first channel is one of several channels. The first flag can refer to those described in the encoding method above and will not be explained further here.
[0189] S503: The decoding side analyzes the second flag from the substream obtained by encoding the coded block of the first channel.
[0190] Here, the second flag can be found in the encoding method described above, and will not be explained further here.
[0191] S504: If the second flag indicates that the target mode used for the coding block of multiple channels is the fallback mode, the decoding side analyzes the third flag from the substream obtained by coding the coding block of the first channel.
[0192] Here, the third flag can be found in the encoding method described above, and will not be explained further here.
[0193] S505: The decoding side determines the target decoding mode for multiple channels based on the type of fallback mode indicated by the third flag, and decodes the substream obtained by encoding the coded blocks of multiple channels in the target decoding mode.
[0194] For example, the decoding side can use the type of fallback mode indicated by the third flag as the target decoding mode for multiple channels.
[0195] In one embodiment, the method further includes: the decoding side determining a target decoded code length of an encoding unit based on the remaining code stream buffer and the target pixel depth BPP, wherein the target decoded code length indicates the code length required to decode the code stream of the encoding unit; the decoding side determining the assigned code lengths of a plurality of channels based on the decoded code length, wherein the assigned code lengths indicate the code length required to decode the residuals of the code stream of the encoding blocks of the plurality of channels; and the decoding side determining the decoded code length assigned to each of the plurality of channels based on the average value of the assigned code lengths in the plurality of channels.
[0196] Based on the above embodiment 2, and taking the example that the multiple channels include a luminance (Y) channel, a first chromaticity (U) channel, and a second chromaticity (V) channel, we will explain two proposals mainly included in improved fallback mode proposal 1.
[0197] Proposal 1: Encoding side: Encoding side control always ensures that the bit consumption cost of encoding is highest in the anti-bloat mode (original value mode), and the encoding side cannot select a mode with a bit consumption cost greater than the original value. Therefore, an encoding mode is required for all components, and even in fallback mode, an encoding mode is required not only for the substream corresponding to the Y channel (first substream) but also for the substreams corresponding to the U / V channels (second and third substreams). The codeword in fallback mode is maintained to be the same as the codeword in original value mode. The specific types of fallback modes (first fallback mode and second fallback mode) can be encoded based on a certain channel.
[0198] Decoder side: First, the encoding modes of the three channels are analyzed. If all three channels (Y / U / V) are selected based on rate distortion cost, the remaining code stream buffer determines whether the code stream buffer allows all three channels to decode in their respective target encoding modes. If it does not allow this, the current decoding mode is fallback mode. In fallback mode, one flag bit can be analyzed based on a certain component to indicate whether the current fallback mode is the first fallback mode or the second fallback mode (where all three channels use the same fallback mode). Therefore, if the decoding side has channels that have not selected an anti-bloat mode, the current CU will not select a fallback mode.
[0199] Proposal 2: Encoder: Encoder control always ensures that the bit consumption cost of encoding is highest in the anti-bloat mode (original value mode), and the encoder cannot select a mode with a bit consumption cost greater than the original value mode. Original value mode and fallback mode still use the same mode flag (i.e., the first flag above), but when encoding anti-bloat mode and fallback mode, one additional flag (i.e., the second flag above) is encoded to indicate that the current encoding mode is either anti-bloat mode or fallback mode. If this additional flag indicates that the current encoding mode is fallback mode, then one more flag (i.e., the third flag above) is encoded to indicate that the current fallback mode is either the first or second fallback mode. If multiple channels all select a fallback mode, the type of fallback mode for multiple channels remains consistent. If a channel selects anti-bloat mode, the mode flag is encoded in the substream corresponding to that channel, but it is not necessary to encode a flag indicating whether or not it is in fallback mode (i.e., the second flag above).
[0200] Decoder: The mode flag is analyzed from the code stream. If the mode flag is a common mode flag for both the anti-bloat mode and the fallback mode, one flag (i.e., the second flag mentioned above) is analyzed to indicate whether or not the current mode is fallback mode. If the mode is fallback mode, another flag (i.e., the third flag mentioned above) is analyzed to indicate whether the current fallback mode is either the first or second fallback mode. Once the type of fallback mode for one channel is analyzed, other channels using fallback modes will use the same type of fallback mode as that channel. Channels that do not use fallback modes will analyze their own mode.
[0201] As described above, Improvement 1, as in Embodiment 2 above, can reduce the upper limit of the numerator in the expansion rate formula. In one embodiment, in the fallback mode, the lower limit of the denominator in the expansion rate formula can also be increased. The following will be explained using the selectable Embodiment 2 and Improvement 2.
[0202] Option 2: If it is decided to encode using fallback mode, the encoder provides a single target BPP, and in fallback mode, the sum of the number of bits in the substreams corresponding to the encoded blocks of the three channels must be less than or equal to the target BPP. Therefore, when allocating the code rate, common information (block vector and mode flag) is encoded in the luminance channel, and after subtracting the code length occupied by the common information from the allocated code length, the remaining code length is used to encode the residuals of the multiple channels, so the remaining code length needs to be allocated. As an allocation method, the remaining code length is divided by the number of pixels of the multiple channels (i.e., the code length allocated to each channel must be an integer multiple of the number of pixels of the encoded unit), and if it cannot be divided, the code length of the undivided portion is allocated to the luminance channel. If the remaining code length is small, it cannot be divided by the number of pixels of the multiple channels, and the code length cannot be allocated to the substreams corresponding to the first and second chromaticity channels, resulting in a small number of bits in the corresponding substreams and a large expansion rate.
[0203] Improvement plan 2: Allocate the remaining code length evenly across multiple channels according to the number of bits, thereby improving allocation accuracy.
[0204] Example 3: Figure 16 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present invention. As shown in Figure 16, in addition to S301 to S303 described above, the image encoding method further includes S601 to S603.
[0205] S601: The encoding side determines the target coded length of the encoding unit based on the remaining amount of the code stream buffer and the target BPP.
[0206] Here, the target coding length indicates the length of code required to encode the coding unit. The target BPP can be obtained by the encoding side, for example, by receiving a target BPP entered by the user.
[0207] S602: Determine the assigned code lengths for multiple channels based on the target coded length.
[0208] Here, the allocated code length indicates the code length required to encode the residuals of the coded blocks for multiple channels.
[0209] For example, as described above, the encoding side subtracts the code length occupied by common information from the allocated code length to obtain the allocated code lengths for multiple channels.
[0210] S603: Determine the code length assigned to each of the multiple channels based on the average value of the assigned code lengths across the multiple channels.
[0211] In one embodiment, the encoding side can also encode the block vector and mode flag in substreams corresponding to multiple channels.
[0212] Here, the mode flag can be found in the above explanation and will not be explained further here. The block vector can be found in the IBC mode explanation below and will not be explained further here.
[0213] In one embodiment, when encoding residuals, if the bits assigned to a channel cannot divide the number of pixels in that channel, that is, if the residuals of each pixel cannot be encoded using a fixed-length code, the encoding side can divide the pixels of that channel into groups, and each group can encode only one residual value.
[0214] Furthermore, if the remaining code length is small, the current code length allocation mode, which uses the number of pixels in the encoding unit as the unit, may result in not allocating a code length to the chromaticity channel, leading to a smaller number of encoded bits in the encoding block of the chromaticity channel and potentially a larger expansion rate. The image encoding method provided by the embodiment of the present invention allows for allocation by converting the allocated unit from the number of pixels to the number of bits of the allocated code length, so that each channel can be allocated a code length and encoded, thereby reducing the theoretical expansion rate.
[0215] In response to the above, embodiments of the present invention further provide an image decoding method, adding to S401-S404 or S501-S505 above, the method further includes: determining a target decoded code length of an encoding unit based on the remaining amount of the code stream buffer and the target pixel depth BPP, wherein the target decoded code length is to indicate the code length required to decode the code stream of the encoding unit; determining the assigned code lengths of a plurality of channels based on the decoded code length, wherein the assigned code lengths are to indicate the code length required to decode the residuals of the code stream of the encoding blocks of the plurality of channels; and determining the decoded code length assigned to each of the plurality of channels based on the average value of the assigned code lengths in the plurality of channels.
[0216] 3. Changing the Intra Block Copy (IBC) mode. Possible implementation: The mode identifier is at the CU level, the BV (Block Vector) is also at the CU level, and the mode identifier and BV are transmitted via the luminance channel.
[0217] Proposed improvement: Change the mode identifier to CB level, while BV remains at CU level. Each channel transmits the mode identifier, and each channel transmits the BV.
[0218] Example 4: Figure 17 is a schematic flowchart of yet another image coding method provided by an embodiment of the present invention. As shown in Figure 17, this method includes S701 to S704.
[0219] S701: The encoding side acquires the encoding unit.
[0220] Here, the coding unit includes coding blocks for multiple channels.
[0221] S702: The encoding side encodes an encoding block for at least one channel among multiple channels in IBC mode.
[0222] In one embodiment, S702 specifically includes the encoding side determining the target encoding mode for this at least one channel using a rate distortion optimization policy.
[0223] Here, the target coding mode includes IBC mode. The rate distortion optimization policy can be found in the description of policy 3 above and will not be explained further here.
[0224] S703: The encoding side obtains the BV of the reference prediction block.
[0225] Here, the BV of the reference prediction block indicates the position of the reference prediction block within the encoded image block. The reference prediction block represents the predicted value of the encoded block encoded in IBC mode.
[0226] S704: The encoding side encodes the BV of the reference prediction block in at least one subnet obtained by encoding the encoding block of at least one channel in IBC mode.
[0227] In one embodiment, for the BV of the reference prediction block of the encoding block of each channel, all of these BVs can be encoded in the substream corresponding to the luminance channel.
[0228] In one embodiment, for the BV of the reference prediction block of the encoding block for each channel, all of these BVs can be encoded in a substream corresponding to a channel with chromaticity (e.g., a first chromaticity channel or a second chromaticity channel), or the number of BVs can be equally divided and encoded in substreams corresponding to two chromaticity channels. If they cannot be equally divided, the encoding side can encode the other BV in a substream corresponding to either of the chromaticity channels.
[0229] In one embodiment, the BV of the reference prediction block of the coding block for each channel can be coded by dividing it equally in the substreams corresponding to the coding blocks of different channels. If it cannot be divided equally, the coding side can allocate it in a predetermined ratio. For example, taking the number of BVs as 8, the predetermined ratio can be Y channel:U channel:V channel = 2:3:3, or Y channel:U channel:V channel = 4:2:2, etc. Embodiments of this application do not limit the specific numerical values of the predetermined ratio. In this case, S602 specifically includes coding the coding blocks of multiple channels in IBC mode. S704 specifically includes coding the BV of the reference prediction block in the substream obtained by coding the coding blocks in each of the multiple channels in IBC mode.
[0230] In one embodiment, the BV of the reference prediction block includes multiple BVs. Encoding the BV of the reference prediction block in a substream obtained by encoding the coded blocks in each of the multiple channels in IBC mode includes encoding the BV of the reference prediction block in a code stream obtained by encoding the coded blocks in each of the multiple channels in IBC mode based on a predetermined ratio.
[0231] It should be understood that the number of bits in the substream corresponding to the luminance channel is usually large, and the number of bits in the substream corresponding to the chromaticity channel is usually small. In the embodiment of the present invention, when encoding using IBC mode, the transmission data in the substream corresponding to the chromaticity channel with a small number of bits can be increased by transmitting the BV of the reference prediction block in the substream corresponding to the encoding block of at least one channel, thereby increasing the denominator in the formula for calculating the theoretical expansion rate and reducing the theoretical expansion rate.
[0232] In one embodiment, the CU generates header information during encoding, and this CU header information can also be assigned and encoded in the substream using the BV assignment method described above.
[0233] In response to the above, embodiments of the present invention further provide an image decoding method. Figure 18 is a schematic flowchart of yet another image decoding method provided by embodiments of the present invention. As shown in Figure 18, this method includes S801 to S804.
[0234] S801: The decoding side analyzes the code stream obtained by encoding the encoding unit.
[0235] Here, the coding unit includes coding blocks for multiple channels, and the code stream includes multiple substreams, each corresponding one-to-one to the multiple channels, from which the coding blocks for multiple channels have been coded.
[0236] S802: The decoding side determines the position of the reference prediction block in multiple substreams based on the block vector BV of the reference prediction block analyzed from at least one of the multiple substreams.
[0237] S803: The decryption side determines the predicted value of the decrypted block to be decrypted in IBC mode based on the position information of the reference prediction block.
[0238] S804: The decryption side reconstructs the decryption block to be decrypted in IBC mode based on the predicted values.
[0239] S803 and S804 can be found in the description of the video encoding and decoding system above, and will not be explained further here.
[0240] In one embodiment, the encoding side can analyze the BV of the reference prediction block of the encoding block for each channel in a substream that corresponds to the luminance channel.
[0241] In one embodiment, for the BV of the reference prediction block of the coding block for each channel, all of these BVs can be analyzed in a substream corresponding to a channel with chromaticity (e.g., the first chromaticity channel or the second chromaticity channel), or the number of BVs can be equally divided and analyzed in substreams corresponding to two chromaticity channels. If equal division is not possible, the coding side can analyze the other BV in a substream corresponding to either of the chromaticity channels.
[0242] In one embodiment, the BV of the reference prediction block of the coding block for each channel can be coded by dividing it equally in the substreams corresponding to coding blocks of different channels. If it cannot be divided equally, the decoding side can allocate it in a predetermined ratio. For example, taking the number of BVs as 8, the predetermined ratio can be Y channel:U channel:V channel = 2:3:3, or Y channel:U channel:V channel = 4:2:2, etc. Embodiments of this application do not limit the specific numerical values of the predetermined ratio.
[0243] In one embodiment, the coded block encoded in IBC mode includes at least two channel coded blocks, and the at least two channel coded blocks share the BV of the reference prediction block.
[0244] In one embodiment, the method further includes determining that the target decoding mode corresponding to the multiple substreams is IBC mode once an IBC mode identifier has been parsed from any one of the multiple substreams.
[0245] In one embodiment, the method further includes the decoding side analyzing IBC mode identifiers one by one from a plurality of substreams and determining that the target decoding mode corresponding to the substream of the IBC mode identifier analyzed from the plurality of substreams is IBC mode.
[0246] Based on the above embodiment IV, and taking the example that the multiple channels include a luminance (Y) channel, a first chromaticity (U) channel, and a second chromaticity (V) channel, we will explain two proposals mainly included in improved fallback mode proposal 1.
[0247] Proposal 1: Encoding side: Step 1: Obtain a reference prediction block based on the multichannel. That is, for one multichannel encoding unit, the multichannel at a certain position within the search area is designated as the reference prediction block for the current multichannel encoding unit, and this position is denoted as BV. In other words, the multichannel shares one BV. If the input image is YUV400, it has only one channel, but otherwise it has three channels. After obtaining the prediction block, the residuals for each channel are calculated, the residuals are (transformed) quantized, inverse quantized (inverse transformed), and then the reconstruction is completed.
[0248] Step 2: Encode auxiliary information (including information such as the complexity level of the coding block) and mode information in the luminance channel, and also encode the quantized coefficients of the luminance channel.
[0249] Step 3: If a chromaticity channel exists, encode auxiliary information (including information such as the coding block complexity) in the chromaticity channel, and encode the quantized coefficients of the chromaticity channel.
[0250] 3.1: For each coding unit, the BV of the reference prediction block can all be coded in the luminance channel.
[0251] 3.2: For each coding unit, the BV of the reference prediction block can be coded entirely on a channel with chromaticity, or it can be coded by dividing the BV equally across two chromaticity channels. If it cannot be divided equally, a portion of the BV is coded on one of the chromaticity channels.
[0252] 3.3: For the BV of the reference prediction block of each coding unit, the BV can be divided equally and coded into different components. If it cannot be divided equally, it will be allocated according to a predetermined ratio.
[0253] Decryption side: Step 1: Analyze the auxiliary information and mode information of the encoding block in the luminance channel, and analyze the quantized coefficients in the luminance channel.
[0254] Step 2: If there is a chrominance channel, analyze the auxiliary information of the encoding block in the U channel. If the prediction mode of the luminance channel is the IBC mode, there is no need to analyze the prediction mode of the current chrominance channel and directly use the IBC mode. Analyze the quantized coefficients in the current chrominance channel.
[0255] Step 3: The analysis of BV is the same on the encoding side. That is, as follows. 3.1: Regarding the BV of the reference prediction block of each encoding unit, it can all be analyzed in the luminance channel.
[0256] 3.2: Regarding the BV of the reference prediction block of each analysis unit, it can all be analyzed in the channel with chrominance, or the BV can be equally divided and analyzed in two chrominance channels. If it cannot be equally divided, encode a part of the BV in a certain chrominance channel.
[0257] 3.3: Regarding the BV of the reference prediction block of each analysis unit, the BV can be equally divided and encoded into different components. If it cannot be equally divided, assign it according to a preset ratio.
[0258] Step 4: Based on the BV shared by the three channels, obtain the predicted value of each encoding block in each channel, inverse quantize (inverse transform) the coefficients obtained by analyzing in each channel to obtain the residual value, and complete the reconstruction of each encoding block based on the residual value and the predicted value.
[0259] Proposal 2: Encoding side: Step 1: For the same coding unit, that is, for each IBC mode in three channels, it is necessary to train to obtain only one group of BV. This group of BV is obtained based on the search regions and original values of the three channels, or based on the search region and original value of one of the channels, or based on the search regions and original values of any two of the channels.
[0260] Step 2: For the coding block of each channel, use the rate-distortion cost to determine the target coding mode. Here, the BV of the three channels in the IBC mode uses the BV calculated in Step 1. After a certain component selects the IBC mode, other components cannot select the IBC mode.
[0261] Step 3: For the coding block of each channel, it is necessary to encode its own optimal mode. When the target mode of one or more channels selects the IBC mode, the BV can be encoded based on a certain channel that selects the IBC mode, the number of BV can be equally divided and encoded based on two certain channels that select IBC, and the number of BV can be equally divided and encoded based on all channels that select the IBC mode. When the number of BV cannot be evenly divided, it can be allocated at a preset ratio.
[0262] Decoder side: Each channel analyzes one target mode. When it is analyzed that the target mode of one or more channels selects the IBC mode (only the same IBC mode can be selected), the BV analysis method is the same as the allocation method for the encoder to encode the BV.
[0263] IV. Grouping of coefficients. Possible implementation options: As shown in the flowchart in Figure 1, a residual skip mode may exist in the encoding process. If the encoding side selects this skip mode during encoding, it is not necessary to encode the residuals during encoding, and it is sufficient to encode 1 bit of data to represent the residual skip.
[0264] Suggested improvement: Divide the processing coefficients (residual coefficients and / or transformation coefficients) into groups.
[0265] Example 5: Figure 19 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present invention. As shown in Figure 19, the image encoding method includes S901 to S902.
[0266] S901: The encoding side obtains the processing coefficient corresponding to the encoding unit.
[0267] Here, the processing coefficient includes one or more of the residual coefficient and the transformation coefficient.
[0268] S902: The encoding side divides the processing coefficients into multiple groups based on the number threshold.
[0269] Here, the count threshold can be set in advance on the coding side. The count threshold is related to the size of the coding block. For example, in the case of a 16x2 coding block, the count threshold can be set to 16. Of the processing coefficients of multiple groups, the number of processing coefficients in each group is less than or equal to the count threshold.
[0270] Compared to the residual skipping mode in the selectable implementations, the image coding method provided by the embodiment of the present invention allows the processing coefficients to be divided into groups during coding, and the processing coefficients of each group require the addition of header information for that group to indicate the details of the processing coefficients of that group during transmission, which increases the denominator in the formula for calculating the theoretical expansion rate and reduces the theoretical expansion rate compared to representing the current residual skipping with 1 bit of data.
[0271] In response to the above, embodiments of the present invention further provide an image decoding method. Figure 20 is a schematic flowchart of yet another image decoding method provided by embodiments of the present invention. As shown in Figure 20, the image decoding method includes S1001 to S1002.
[0272] S1001: The decoding side analyzes the code stream obtained by encoding the encoding unit and determines the processing coefficient corresponding to the encoding unit.
[0273] Here, the processing coefficient includes one or more of the residual coefficients and transformation coefficients. The processing coefficient includes multiple groups. The number of processing coefficients in each group is less than or equal to the number threshold.
[0274] S1002: The decoding side decodes the code stream based on the processing coefficient.
[0275] S1002 can be found in the description of the video encoding and decoding system above, and will not be explained further here.
[0276] 5. Complexity transmission. Optional implementation: The coding unit includes a coding block for the luminance channel, a coding block for the first chromaticity channel, and a coding block for the second chromaticity channel. The substream corresponding to the coding block for the luminance channel is the first substream, the substream corresponding to the coding block for the first chromaticity channel is the second substream, and the substream corresponding to the coding block for the second chromaticity channel is the third substream. The first substream transmits the complexity level of the luminance channel in 1 or 3 bits, the second substream transmits the average value of the two chromaticity channels in 1 or 3 bits, and the third substream does not transmit the complexity level.
[0277] For example, the method for calculating the complexity level of a CU is as follows: After searching for BiasInit from a table based on the complexity level ComplexityLevel[0] of the luminance channel and the complexity level ComplexityLevel[1] of the chrominance channels in JPEG0007902355000002.jpg with 51163, the quantization parameter Qp[0] of the luminance channel and the quantization parameters Qp[1], Qp[2] of the two chrominance channels were calculated.
[0278] Here, the searched table can refer to Table 1 below, which will not be further explained here.
[0279] Taking the first sub-stream as an example, the specific realization of the first sub-stream is as follows. JPEG0007902355000003.jpg6787
[0280] Here, complexity_level_flag[0] is the update flag for the complexity level of the luminance channel and is a binary variable. A value of "1" indicates that the luminance channel of the coding unit needs to update the complexity level, and a value of "0" indicates that the luminance channel of the coding unit does not need to update the complexity level. The value of ComplexityLevelFlag[0] is equal to the value of complexity_level_flag[0]. delta_level[0] is the amount of change in the complexity level of the luminance channel and is a 2-bit unsigned integer. Determine the amount of change in the luminance complexity level. The value of DeltaLevel[0] is equal to the value of delta_level[0]. If delta_level[0] does not exist in the code stream, the value of DeltaLevel[0] is equal to 0. PrevComplexityLevel represents the complexity level of the luminance channel of the previous coding unit, and ComplexityLevel[0] represents the complexity level of the luminance channel.
[0281] Improvement: Transmit complexity information in the third sub-stream.
[0282] Example Six Figure 21 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present invention. As shown in Figure 21, the image encoding method includes S1101 to S1103.
[0283] S1101: The encoding side acquires the encoding unit.
[0284] Here, the coding unit contains coding blocks of P channels, where P is an integer greater than or equal to 2.
[0285] S1102: The encoding side obtains complexity information for each of the P channels.
[0286] Here, complexity information is used to represent the degree of difference in pixel values of the encoded blocks for each channel. For example, taking as an example that P channels include a luminance channel, a first chromaticity channel, and a second chromaticity channel, the complexity information of the encoded blocks for the luminance channel is used to represent the degree of difference in pixel values of the encoded blocks for the luminance channel, the complexity information of the encoded blocks for the first chromaticity channel is used to represent the degree of difference in pixel values of the encoded blocks for the first chromaticity channel, and the complexity information of the encoded blocks for the second chromaticity channel is used to represent the degree of difference in pixel values of the encoded blocks for the second chromaticity channel.
[0287] S1103: The encoding side encodes complexity information for each channel's encoded block in the substream obtained by encoding the encoded blocks of P channels.
[0288] In one embodiment, S1103 specifically includes the encoding side encoding the complexity level of each channel's encoding block in a substream obtained by encoding P channel encoding blocks.
[0289] For example, taking the second subnet corresponding to the first chromaticity component as an example, the second subnet is specifically implemented as follows. JPEG0007902355000004.jpg7294
[0290] Here, complexity_level_flag[1] is a binary variable that indicates the update flag for the complexity level of the first chromaticity channel. A value of "1" indicates that the complexity level of the coded block for the first chromaticity channel of the coding unit matches the complexity level of the coded block for the luminance channel, and a value of "0" indicates that the complexity level of the coded block for the first chromaticity channel of the coding unit does not match the complexity level of the coded block for the luminance channel. The value of ComplexityLevelFlag[1] is equal to the value of complexity_level_flag[1]. delta_level[1] is a 2-bit unsigned integer that indicates the change in the complexity level of the first chromaticity channel. It determines the change in the complexity level of the coded block for the first chromaticity channel. The value of DeltaLevel[1] is equal to the value of delta_level[1]. If delta_level[1] does not exist in the code stream, the value of DeltaLevel[1] is equal to 0. ComplexityLevel[1] indicates the complexity level of the coded block for the first chromaticity channel.
[0291] For example, taking a third subnet corresponding to the second chromaticity component as an example, the third subnet is specifically implemented as follows. JPEG0007902355000005.jpg79105
[0292] Here, complexity_level_flag[2] is a binary variable that indicates the update flag for the complexity level of the second chromaticity channel. A value of "1" indicates that the complexity level of the coded block for the second chromaticity channel of the coding unit matches the complexity level of the coded block for the first chromaticity channel, and a value of "0" indicates that the complexity level of the coded block for the second chromaticity channel of the coding unit does not match the complexity level of the coded block for the first chromaticity channel. The value of ComplexityLevelFlag[2] is equal to the value of complexity_level_flag[2]. delta_level[2] is a 2-bit unsigned integer that indicates the change in the complexity level of the second chromaticity channel. It determines the change in the complexity level of the coded block for the second chromaticity channel. The value of DeltaLevel[2] is equal to the value of delta_level[2]. If delta_level[2] does not exist in the code stream, the value of DeltaLevel[2] is equal to 0. ComplexityLevel[2] indicates the complexity level of the coding block for the second chromaticity channel.
[0293] In one embodiment, complexity information includes a complexity level and a first reference coefficient, the first reference coefficient being used to represent the ratio relationship between complexity levels of coding blocks of different channels. Specifically, S1103 includes the coding side coding the complexity level of the Q channel coding blocks in each of the substreams obtained by coding the Q channel coding blocks out of P channels, where Q is an integer less than P, and the coding side coding the first reference coefficient in the substreams obtained by coding the PQ channel coding blocks out of P channels.
[0294] For example, taking as an example that P channels include a luminance channel, a first chromaticity channel, and a second chromaticity channel, the complexity information of the coding block for the luminance channel is the complexity level of the coding block for the luminance channel, the complexity information of the coding block for the first chromaticity channel is the complexity level of the coding block for the first chromaticity channel, and the complexity information of the coding block for the second chromaticity channel is the reference coefficient, which is used to represent the ratio relationship between the complexity level of the coding block for the first chromaticity channel and the complexity level of the coding block for the second chromaticity channel.
[0295] For example, the meaning of the variable complexity_level_flag[2] can be changed so that a value of "1" indicates that the complexity level of the coding block for the second chromaticity channel of the coding unit is greater than the complexity level of the coding block for the first chromaticity channel, and a value of "0" indicates that the complexity level of the coding block for the second chromaticity channel of the coding unit is less than or equal to the complexity level of the coding block for the first chromaticity channel. The value of ComplexityLevelFlag[2] is equal to the value of complexity_level_flag[2].
[0296] In one embodiment, complexity information includes a complexity level, a reference complexity level, and a second reference coefficient. The reference complexity level includes one of a first complexity level, a second complexity level, and a third complexity level. The first complexity level is the maximum complexity level of the coded blocks of PQ channels out of P channels, where Q is an integer less than P. The second complexity level is the minimum complexity level of the coded blocks of PQ channels out of P channels. The third complexity level is the average complexity level of the coded blocks of PQ channels out of P channels. The second reference coefficient is used to represent the relationship and / or ratio relationship between the complexity levels of the coded blocks of PQ channels out of P channels. In this case, S1103 specifically includes the encoding side encoding the complexity level of the Q channel encoding block in each of the substreams obtained by encoding the Q channel encoding block out of the P channels, and the encoding side encoding the reference complexity level and a second reference coefficient in the substreams obtained by encoding the PQ channel encoding block out of the P channels.
[0297] For example, taking as an example that P channels include a luminance channel, a first chromaticity channel, and a second chromaticity channel, the complexity information of the coding block for the luminance channel is the complexity level of the coding block for the luminance channel, and the complexity information of the coding block for the first chromaticity channel is the reference complexity level. The reference complexity level includes one of the first complexity level, the second complexity level, and the third complexity level. The first complexity level is the maximum value among the complexity level of the coding block for the first chromaticity channel and the complexity level of the coding block for the second chromaticity channel. The second complexity level is the minimum value among the complexity level of the coding block for the first chromaticity channel and the complexity level of the coding block for the second chromaticity channel. The third complexity level is the average value among the complexity level of the coding block for the first chromaticity channel and the complexity level of the coding block for the second chromaticity channel. The complexity information of the coding block for the second chromaticity channel is a reference coefficient, which is used to represent the relationship and / or ratio relationship between the complexity level of the coding block for the first chromaticity channel and the complexity level of the coding block for the second chromaticity channel.
[0298] It should be understood that in the selectable implementations, complexity information is not transmitted in the third substream corresponding to the encoding block of the second chromaticity channel. The image encoding method provided by the embodiment of the present invention increases the number of bits in the substream with fewer bits by adding the encoding of complexity information in the third substream, thereby increasing the denominator in the theoretical expansion rate calculation formula and reducing the theoretical expansion rate.
[0299] In response to the above, embodiments of the present invention provide an image decoding method. Figure 2 is a diagram showing a further image decoding method provided by embodiments of the present invention. As shown in Figure 2, the decoding method includes S1201 to S1204.
[0300] S1201: The decoding side analyzes the code stream obtained by encoding the encoding unit.
[0301] Here, the coding unit contains a coding block of P channels, where P is an integer greater than or equal to 2. The code stream contains multiple substreams, each corresponding one-to-one to the P channels, with the coding block of P channels being coded.
[0302] S1202: The decoding side analyzes the complexity information of each channel's encoded block in the substream obtained by encoding the encoded blocks of P channels.
[0303] In one embodiment, S1202 specifically includes the decoding side encoding the encoding blocks of P channels and analyzing the complexity level of each channel's encoding block in each of the resulting substreams.
[0304] S1203: The decoding side determines the quantization parameters for each channel's coded block based on the complexity information of each channel's coded block.
[0305] S1204: The decoding side decodes the code stream based on the quantization parameters of the coded block for each channel.
[0306] In one embodiment, if the complexity level of each channel is transmitted in any of the substreams corresponding to each channel, the complexity level of the CU level (CuComplexityLevel) can be calculated according to the following process. JPEG0007902355000006.jpg43170
[0307] Here, the definition of ComplexityDivide3Table is ComplexityDivide3Table={0,0,0,1,1,1,2,2,2,3,3,3,4}; and image_format represents the image format of the image to be processed in which the encoding unit exists.
[0308] In one embodiment, as described above, complexity information includes a complexity level and a first reference coefficient, the first reference coefficient being used to represent the ratio relationship between the complexity levels of the coded blocks of different channels. In this case, S1202 specifically involves the decoding side analyzing the complexity level of the coded blocks of Q channels in each of the substreams obtained by encoding the coded blocks of Q channels out of P channels, where Q is an integer less than P; the decoding side analyzing the first reference coefficient in the substreams obtained by encoding the coded blocks of PQ channels out of P channels; and the decoding side determining the complexity level of the coded blocks of PQ channels based on the first reference coefficient and the complexity level of the coded blocks of Q channels.
[0309] In one embodiment, if the complexity level is not transmitted in the substream (i.e., the magnitude relationship is transmitted to the third substream), the complexity level of the CU level (CuComplexityLevel) can be calculated according to the following process. JPEG0007902355000007.jpg49163
[0310] Here, ChromaComplexityLevel represents the complexity level of the chromaticity channel.
[0311] The ChromaComplexityLevel needs to be calculated individually on the encoding side (the decoding side obtains it directly from the code stream).
[0312] ChromaComplexityLevel=(ComplexityLevel[1]+ComplexityLevel[2])>>1, or ChromaComplexityLevel=max(ComplexityLevel[1],ComplexityLevel[2]), or ChromaComplexityLevel=min(ComplexityLevel[1],ComplexityLevel[2]), or ChromaComplexityLevel=ComplexityLevel[1], or ChromaComplexityLevel=ComplexityLevel[2].
[0313] In one embodiment, complexity information includes a complexity level, a reference complexity level, and a second reference coefficient. The reference complexity level includes one of a first complexity level, a second complexity level, and a third complexity level. The first complexity level is the maximum complexity level of the coded blocks of PQ channels out of P channels, where Q is an integer less than P. The second complexity level is the minimum complexity level of the coded blocks of PQ channels out of P channels. The third complexity level is the average complexity level of the coded blocks of PQ channels out of P channels. The second reference coefficient is used to represent the relationship and / or ratio relationship between the complexity levels of the coded blocks of PQ channels out of P channels. In this case, S1202 specifically includes: the decoding side analyzing the complexity level of the coded blocks of Q channels in each of the substreams obtained by coding the coded blocks of Q channels out of P channels; the decoding side coding the reference complexity level and a second reference coefficient in the substreams obtained by coding the coded blocks of PQ channels out of P channels; and the decoding side determining the complexity level of the coded blocks of PQ channels based on the complexity level of the coded blocks of Q channels and the second reference coefficient.
[0314] In one embodiment, if the complexity level of each channel is transmitted to the substream corresponding to each channel, the decoding side searches for BiasInit1 from Table 2 below based on the complexity level ComplexityLevel[0] of the coding block of the luminance channel and the complexity level ComplexityLevel[1] of the coding block of the first chromaticity channel, and searches for BiasInit2 from Table 2 below based on the complexity level ComplexityLevel[0] of the coding block of the luminance channel and the complexity level ComplexityLevel[2] of the coding block of the second chromaticity channel, and calculates the quantization parameter Qp[0] of the luminance channel, the quantization parameter Qp[1] of the first chromaticity component, and the quantization parameter Qp[2] of the second chromaticity channel.
[0315] For example, the decoding side can calculate the quantization parameters according to the following process. JPEG0007902355000008.jpg46110
[0316] [Table 1]
[0317] [Table 2]
[0318] In one embodiment, the decoding side retrieves BiasInit from Table 2 based on the complexity level ComplexityLevel[0] and the complexity level ChromaComplexityLevel of the coding block of the chromaticity channel, and calculates Qp[0], Qp[1], and Qp[2] according to the following process. JPEG0007902355000011.jpg49155
[0319] In one embodiment, as described above, the first complexity level and the second complexity level can be transmitted in the substream corresponding to the coding block of the first chromaticity channel. In this case, the decoding side searches for BiasInit from Table 2 based on the complexity level ComplexityLevel[0] of the coding block of the luminance channel and the complexity level ChromaComplexityLevel (taken as the first complexity level / the complexity level of the coding block of the first chromaticity channel), and calculates Qp[0], Qp[1], and Qp[2] according to the following process. JPEG0007902355000012.jpg72117
[0320] In one embodiment, as described above, a third complexity level can be transmitted in the substream corresponding to the coding block of the first chromaticity channel. In this case, the ChromaComplexityLevel is obtained by taking the third complexity level and calculating Qp[0], Qp[1], and Qp[2] according to the following process. JPEG0007902355000013.jpg72107
[0321] 6. The chromaticity channel shares a substream. Optional implementation: For an image to be processed in YUV420 / YUV422 format, a total of three substreams are transmitted: a first substream, a second substream, and a third substream. Here, the first substream contains the syntactic elements and transformation / residual coefficients of the luminance channel, the second substream contains the syntactic elements and transformation / residual coefficients of the first chromaticity channel, and the third substream contains the syntactic elements and transformation / residual coefficients of the second chromaticity channel.
[0322] Proposed improvement: For images processed in YUV420 / YUV422 format, transmit the syntax elements and conversion coefficients / residual coefficients of the first chromaticity channel, and the syntax elements and conversion coefficients / residual coefficients of the second chromaticity channel to a second substream, and cancel the third substream.
[0323] Example 7: Figure 22 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present invention. As shown in Figure 22, the image encoding method includes S1201 to S1202.
[0324] S1201: The encoding side obtains the encoding unit.
[0325] Here, the encoding unit is an image block in the image to be processed. The encoding unit includes encoding blocks for multiple channels.
[0326] S1202: If the image format of the image to be processed is a preset format, the encoding side encodes encoding blocks of at least two of the multiple preset channels and merges the resulting substreams into a single merged substream.
[0327] Here, the preset format can be set in advance on the encoding / decoding side, for example, the preset format may be YUV420 or YUV422, and the preset channels may be the first chromaticity channel and the second chromaticity channel. Exemplarily, the encoding block of the first substream is defined as follows: JPEG0007902355000014.jpg94170
[0328] For example, the encoding block of the second substream is defined as follows: JPEG0007902355000015.jpg168170
[0329] It should be understood that the number of encoded bits for the syntactic elements and quantized conversion coefficients of the two chromaticity channels is typically smaller than that of the luminance channel, and that by fusing the syntactic elements of the two chromaticity channels into a second substream and transmitting them, the difference in the number of bits between the first and second substreams can be reduced, thereby reducing the theoretical expansion rate.
[0330] In response to the above, embodiments of the present invention further provide an image decoding method, and Figure 23 is a schematic flowchart of yet another image decoding method provided by embodiments of the present invention. As shown in Figure 23, this method includes S1301 to S1303.
[0331] S1301: The decoding side analyzes the code stream obtained by encoding the encoding unit.
[0332] S1302: The decrypting side confirms the merged substream.
[0333] Here, the merged substream is obtained by merging substreams that are obtained by encoding the encoded blocks of at least two preset channels out of multiple channels, in the case where the image format of the image to be processed is a preset format.
[0334] S1303: The decryption side decrypts the code stream based on the merger target substream in which the above two substreams are merged.
[0335] 7. Substream embedding. Proposed improvement: Embed preset codewords in substreams whose bit count is below the bit threshold.
[0336] Example 8: Figure 24 is a schematic flowchart of yet another image coding method provided by an embodiment of the present invention. As shown in Figure 24, this method includes S1401 to S1402.
[0337] S1401: The encoding side acquires the encoding unit.
[0338] Here, the coding unit includes coding blocks for multiple channels.
[0339] S1402: The encoding side encodes preset codewords in target substreams that satisfy the preset conditions until the target substream no longer satisfies the preset conditions.
[0340] Here, the target substream is one of several substreams. The multiple substreams are code streams obtained by encoding multiple channel coding blocks. The preset codeword may be "0" or another codeword, etc. Embodiments of the present application are not limited thereto.
[0341] In one embodiment, the preset condition includes that the number of bits in the substream is less than a preset first bit threshold.
[0342] Here, the first bit threshold can be pre-set on the encoding / decoding side or transmitted in the stream by the encoding / decoding side. Embodiments of the present application are not limited thereto. The first bit threshold is used to indicate the minimum number of bits of the allowable CB in the CU.
[0343] In one embodiment, the pre-set condition includes that the code stream of the coding unit contains coded blocks whose number of bits is less than a pre-set second bit threshold.
[0344] Here, the second bit threshold can be pre-set on the encoding / decoding side or transmitted in the stream by the encoding / decoding side. Embodiments of the present invention are not limited thereto. The second bit threshold is used to indicate the minimum allowable stream bit count in a plurality of substreams.
[0345] It should be understood that embedding preset codewords in a low-bit substream directly increases the denominator of the expansion rate calculation formula, thereby reducing the actual expansion rate of the coding unit.
[0346] In response to the above, embodiments of the present invention further provide an image decoding method. Figure 25 is a schematic flowchart of yet another image decoding method provided by embodiments of the present invention. As shown in Figure 25, this image decoding method includes S1501 to S1503.
[0347] S1501: The decoding side analyzes the code stream obtained by encoding the encoding unit.
[0348] Here, the coding unit includes coding blocks for multiple channels.
[0349] S1502: The decoding side determines the number of codewords.
[0350] Here, the number of codewords is used to indicate the number of preset codewords encoded in the target substream that satisfies the pre-defined conditions. Preset codewords are encoded in the target substream if such a target substream exists that satisfies the pre-defined conditions.
[0351] S1503: The decrypting side decrypts the code stream based on the number of codewords.
[0352] For example, the decryption side can remove preset codewords based on the number of codewords and decrypt the code stream from which the preset codewords have been removed.
[0353] 8. Pre-set the expansion rate. Proposed improvement: Control the current actual expansion rate by pre-setting the expansion rate as a threshold.
[0354] Example 9: Figure 26 is a schematic flowchart of yet another image coding method provided by an embodiment of the present invention, which, as shown in Figure 26, includes S1601 to S1603 S1601: The encoding side acquires the encoding unit.
[0355] Here, the coding unit includes coding blocks for multiple channels.
[0356] S1602: The encoding side determines the target encoding mode corresponding to each channel encoding block among the multiple channel encoding blocks based on the preset expansion ratio.
[0357] Here, the preset expansion rate can be set in advance on the encoding and decoding sides. Alternatively, the preset expansion rate may be encoded into a substream by the encoding side and transmitted to the decoding side. Embodiments of the present application are not limited thereto.
[0358] S1603: The encoding side encodes the encoding block for each channel in the target encoding mode so that the current expansion rate is less than or equal to the preset expansion rate.
[0359] In one embodiment, the preset expansion rate, including the first preset expansion rate, is equal to the quotient of the number of bits of the largest substream and the number of bits of the smallest substream. The largest substream is the substream with the largest number of bits among the multiple substreams obtained by encoding multiple channel encoding blocks. The smallest substream is the substream with the smallest number of bits among the multiple substreams obtained by encoding multiple channel encoding blocks.
[0360] In one embodiment, as described above, the encoding side can preset a first preset expansion rate as the maximum threshold allowed within the encoding unit, and before encoding the encoding unit, it can acquire the state of all current substreams, acquire the state of the substream of the most transmitted fixed-length code stream and the substream with the largest total remaining amount in the current substream, and acquire the state of the substream of the least transmitted fixed-length code stream and the substream with the smallest total remaining amount in the current substream. If the ratio or difference between the maximum state and the minimum state is greater than a preset difference threshold, the encoding side can not use rate distortion optimization as a criterion for selecting the target mode, but can select a lower encoding mode for the code stream for the substream in the maximum state and a higher code rate mode for the substream in the minimum state.
[0361] For example, if we set a predetermined expansion rate as B_th and B_delta as an auxiliary threshold, and encode the coding blocks for the three channels Y / U / V to obtain three substreams, the encoding side can obtain the state of these three substreams when encoding one coding unit. If the code rate corresponding to the first substream (the substream corresponding to the Y channel) is set to maximum bit_stream1 and the code rate corresponding to the second substream (the substream corresponding to the U channel) is set to minimum bit_stream2, then if bit_stream1 / bit_stream2 > B_th - B_delta, the encoding side can select a mode with a larger code rate for the coding block of the Y channel and a mode with a smaller code rate for the coding block of the U channel.
[0362] In one embodiment, the preset expansion rate includes a second preset expansion rate. The current expansion rate is equal to the quotient between the number of bits of the encoding block with the largest number of bits and the number of bits of the encoding block with the smallest number of bits among the encoding blocks of the multiple encoded channels.
[0363] In one embodiment, as described above, the encoding side preset a second preset expansion rate as the maximum threshold allowed within the encoding unit, and can obtain the code rate cost in each mode before encoding a certain encoding block. If the encoding mode corresponding to the optimal code rate cost obtained based on the rate-distortion cost satisfies that the actual expansion rate for encoding each substream is less than the preset expansion rate, the encoding mode corresponding to the optimal code rate cost is set as the target encoding mode for the encoding block, and the target encoding mode is incorporated into the substream corresponding to the encoding block. If the encoding mode corresponding to the optimal code rate cost cannot satisfy that the actual expansion rate for encoding each substream is less than the preset expansion rate, the encoding side changes the mode with the largest code rate to a mode with a code rate smaller than the code rate of the mode with the largest code rate, or changes the mode with the smallest code rate to a mode with a code rate larger than the code rate of the mode with the smallest code rate, or changes the mode with the largest code rate to a mode with a code rate smaller than the code rate of the mode with the largest code rate and changes the mode with the smallest code rate to a mode with a code rate larger than the code rate of the mode with the smallest code rate.
[0364] Exemplarily, taking the case where the preset expansion rate is A_th as an example, encoding blocks of three channels of Y / U / V to obtain three substreams. Assume that the optimal code rate costs of the three channels of Y / U / V are rate_y, rate_u, and rate_v respectively, and the code rate of rate_y is the largest and the code rate of rate_u is the smallest. When rate_y / rate_u≧A_th, the encoding side can change the target encoding mode of the Y channel so that its code rate cost becomes smaller than rate_y, or change the target encoding mode of the U channel so that its code rate cost becomes larger than rate_u, so as to satisfy rate_y / rate_u<A_th and rate_y / rate_v<A_th.
[0365] The image coding method provided by the embodiment of the present invention intervenes with a target coding mode selected for each channel coding block by a preset expansion rate, so that when the coding side codes each channel in the target coding mode, the actual expansion rate of the coding unit is less than the preset expansion rate, and thus the actual expansion rate can be reduced.
[0366] In one embodiment, the image encoding method further includes the encoding side determining the current dilation rate, and, if the current dilation rate is greater than a first preset dilation rate, encoding a preset codeword in the minimum substream such that the current dilation rate becomes less than or equal to the first preset dilation rate. For example, the encoding side embeds the preset codeword at the end of the minimum substream.
[0367] In one embodiment, this image encoding method includes the encoding side determining the current expansion rate, and if the encoding side is greater than a second preset expansion rate, encoding a preset codeword in the encoding block with the smallest number of bits so that the current expansion rate becomes less than or equal to the second preset expansion rate.
[0368] In response to the above, embodiments of the present invention further provide an image decoding method. Figure 27 is a schematic flowchart of yet another image decoding method provided by embodiments of the present invention. As shown in Figure 27, this method includes S1701 to S1703.
[0369] S1701: The decoding side analyzes the code stream obtained by encoding the encoding unit.
[0370] S1702: The decoding side determines the number of preset codewords based on the code stream.
[0371] S1703: The decryption side decrypts the code stream based on the number of preset codewords.
[0372] Sections S1701 to S1703 can be found by referring to the explanations in S1501 to S1503 above, and will not be explained further here.
[0373] The above describes the image encoding method and image decoding method provided by the embodiments of the present application in a series of independent embodiments. In actual use, the above series of embodiments and the selectable embodiments within the embodiments can be used in combination with each other. The embodiments of the present application are not limited to specific combinations.
[0374] The above primarily describes the technical means provided by embodiments of the present application based on the method. To realize the above functions, the invention includes corresponding hardware structures and / or software modules for performing each function. The technical objectives of the art should be readily apparent, with reference to the means and algorithmic steps of each example described in the embodiments disclosed herein, that the invention can be realized in hardware or in combination of hardware and computer software. Whether a function is performed in hardware or in computer software to operate the hardware depends on the specific application and design constraints of the technical means. Specialized technical objectives may require different methods to realize the described functions for each specific application, but such realizations should not be considered beyond the scope of the invention.
[0375] In exemplary embodiments, the embodiments of the present application provide an image encoding apparatus, wherein any one of the above image encoding methods is performed by the image encoding apparatus. The image encoding apparatus provided by the embodiments of the present application may be the source device 10 or the video encoder 102.
[0376] Figure 28 is a schematic diagram of the configuration of an image encoding device provided by an embodiment of the present invention. As shown in Figure 28, this image encoding device comprises an acquisition module 2801 and a processing module 2802.
[0377] In one embodiment, the acquisition module 2801 is used to acquire an encoding unit, which is an image block in the image to be processed, and the encoding unit includes an encoding block for multiple channels, which includes a first channel, and the first channel is any one of the multiple channels. The processing module 2802 is used to encode the encoding block for the first channel in a first encoding mode, which is a mode in which the sample value in the encoding block for the first channel is encoded with a first fixed-length code, the code length of the first fixed-length code is less than or equal to the image bit width of the image to be processed, and the image bit width represents the number of bits required to store each sample in the image to be processed.
[0378] In one embodiment, the acquisition module 2801 is further used to acquire a code stream obtained by encoding an encoding unit, the encoding unit being an image block in the image to be processed, the encoding unit including an encoding block of multiple channels, the multiple channels including a first channel, the first channel being any one of the multiple channels, and the code stream including a multiple substreams that correspond one-to-one to the multiple channels, with the encoding blocks of the multiple channels being encoded. The processing module 2802 is further used to decode the substreams corresponding to the first channel in a first decoding mode, the first decoding mode being a mode that analyzes sample values from the substreams corresponding to the first channel with a first fixed-length code, the code length of the first fixed-length code being less than or equal to the image bit width of the image to be processed, and the image bit width representing the number of bits required to store each sample in the image to be processed.
[0379] In one embodiment, the acquisition module 2801 is further used to acquire an encoding unit, the encoding unit being an image block in the image to be processed, and the encoding unit including encoding blocks for multiple channels. The processing module 2802 determines a first total code length, and if the first total code length is greater than or equal to the remaining amount in the code stream buffer, it is used to encode the encoding blocks for multiple channels in fallback mode, the first total code length being the total code length of a first stream obtained by encoding each of the encoding blocks for multiple channels in their respective target encoding modes, the target encoding mode including a first encoding mode, the first encoding mode being a mode that encodes the sample values in the encoding block with a first fixed-length code, the code length of the first fixed-length code being less than or equal to the image bit width of the image to be processed, the image bit width representing the number of bits required to store each sample in the image to be processed, and the mode flag for fallback mode being the same as for the first encoding mode.
[0380] In one embodiment, the processing module 2802 is used to encode mode flags in multiple substreams obtained by encoding multiple channel encoding blocks, and the mode flags are used to indicate the encoding mode used for each of the multiple channel encoding blocks.
[0381] In one embodiment, the plurality of channels include a first channel, the first channel being any one of the plurality of channels, and the processing module 2802 is specifically used to encode a submode flag in the substream obtained by encoding the encoding block of the first channel, the submode flag indicating the type of fallback mode used in the encoding block of the plurality of channels.
[0382] In one embodiment, the plurality of channels include a first channel, the first channel being any one of the plurality of channels, and the processing module 2802 is specifically used to encode a first flag, a second flag, and a third flag in the substream obtained by encoding the encoded block of the first channel, the first flag indicating that the encoded block of the plurality of channels is encoded using a first encoding mode or a fallback mode, the second flag indicating that the encoded block of the plurality of channels is encoded using a target mode which is either the first encoding mode or a fallback mode, and the third flag indicating the type of fallback mode used for the encoded block of the plurality of channels.
[0383] In one embodiment, the processing module 2802 is further used to determine the target coded length of the coding unit based on the remaining amount of the code stream buffer and the target pixel depth BPP, to determine the assigned coded lengths of the multiple channels based on the target coded length, and to determine the coded length assigned to each of the multiple channels based on the average value of the assigned coded lengths in the multiple channels, wherein the target coded length indicates the coded length required to encode the coding unit, and the assigned coded length indicates the coded length required to encode the residuals of the coded blocks of the multiple channels.
[0384] In one embodiment, the processing module 2802 further analyzes the code stream obtained by encoding the encoding unit, and if the mode flag is analyzed from the substream obtained by encoding the encoding blocks of multiple channels and the first total code length is greater than the remaining amount in the code stream buffer, it determines that the target decoding mode of the substream is the fallback mode, analyzes the preset flag bits in the substream to determine the target fallback mode, and uses this to decode the substream in the target fallback mode, the encoding unit includes encoding blocks of multiple channels, the mode flag indicates whether the encoding blocks of multiple channels are encoded using the first encoding mode or the fallback mode, and the first total code length is greater than the remaining amount in the encoding block of multiple channels The lock is the total code length of the first code stream obtained by encoding each of the first encoding modes, the target mode includes the first encoding mode, the first encoding mode is a mode that encodes sample values in an encoding block with a first fixed-length code, the code length of the first fixed-length code is less than or equal to the image bit width of the image to be processed, the image bit width represents the number of bits required to store each sample in the image to be processed, the target fallback mode is a type of fallback mode, the preset flag bits indicate the position of the submode flags, the submode flags indicate the type of fallback mode used when encoding multiple channel encoding blocks, and the fallback mode includes a first fallback mode and a second fallback mode.
[0385] In one embodiment, the processing module 2802 is used to determine the target decoded code length of the coding unit based on the remaining code stream buffer and the target pixel depth BPP, to determine the assigned code lengths of multiple channels based on the target decoded code length, and to determine the decoded code length assigned to each of the multiple channels based on the average value of the assigned code lengths in the multiple channels. The target decoded code length indicates the code length required to decode the coding unit, and the assigned code length indicates the code length required to decode the residuals of the coding blocks of the multiple channels.
[0386] In one embodiment, the processing module 2802 is further used to analyze the code stream obtained by encoding the encoding unit, and analyzes a first flag in the substream obtained by encoding the encoding block of the first channel, and if a second flag is analyzed from the substream obtained by encoding the encoding block of the first channel and the second flag indicates that the target mode used for the encoding blocks of multiple channels is a fallback mode, then analyzes a third flag from the substream obtained by encoding the encoding block of the first channel, determines the target decoding mode of multiple channels based on the type of fallback mode indicated by the third flag, and decodes the substream obtained by encoding the encoding blocks of multiple channels in the target decoding mode. The encoding unit includes a encoding block of multiple channels, each of which includes a first channel, and the first channel is one of the multiple channels. The first flag indicates that the encoding block of multiple channels uses either a first encoding mode or a fallback mode. The first encoding mode is a mode in which the sample values in the encoding block are encoded with a first fixed-length code, the code length of the first fixed-length code being less than or equal to the image bit width of the image to be processed, and the image bit width representing the number of bits required to store each sample in the image to be processed. The second flag is used to indicate that the encoding block of multiple channels is encoded using a target mode which is either the first encoding mode or a fallback mode.
[0387] In one embodiment, the processing module 2802 is further used to determine the target decoded code length of the coding unit based on the remaining code stream buffer and the target pixel depth BPP, to determine the assigned code lengths of the multiple channels based on the target decoded code length, and to determine the decoded code length assigned to each of the multiple channels based on the average value of the assigned code lengths in the multiple channels, wherein the target decoded code length indicates the code length required to decode the coding unit, and the assigned code length indicates the code length required to decode the residuals of the coding blocks of the multiple channels.
[0388] In one embodiment, the acquisition module 2801 is further used to acquire an encoding unit, the encoding unit comprising an encoding block for multiple channels. The processing module 2802 is further used to encode the encoding block for at least one of the multiple channels in IBC mode, which is an intra-block copy mode, to acquire the block vector BV of the reference prediction block, and to encode the BV of the reference prediction block in at least one substream obtained by encoding the encoding block for at least one channel in IBC mode, the BV of the reference prediction block being used to indicate the position of the reference prediction block in the encoded image block, and the reference prediction block being used to represent the predicted value of the encoding block encoded in IBC mode.
[0389] In one embodiment, the processing module 2802 is specifically used to determine the target coding mode for this at least one channel using a rate distortion optimization policy, the target coding mode includes an IBC mode, and the coding block of at least one channel among a plurality of channels is coded in the target coding mode.
[0390] In one embodiment, the processing module 2802 is specifically used to encode the coding blocks of multiple channels in IBC mode, and to encode the BV of the reference prediction block in the substream obtained by encoding the coding block of each of the multiple channels in IBC mode.
[0391] In one embodiment, the processing module 2802 is specifically used to encode the BV of the reference prediction block in a substream obtained by encoding the coding block of each channel among a plurality of channels in IBC mode based on a preset ratio.
[0392] In one embodiment, the processing module 2802 further analyzes the code stream obtained by encoding the encoding unit, determines the position of the reference prediction block in the multiple substreams based on the block vector BV of the reference prediction block analyzed from at least one of the multiple substreams, determines the predicted value of the decoded block to be decoded in IBC mode based on the position information of the reference prediction block, and is used to encode the encoding block to be encoded in IBC mode based on the predicted value. The encoding unit includes encoding blocks for multiple channels, the code stream includes multiple substreams in which the encoding blocks for multiple channels are encoded, each corresponding one-to-one to the multiple channels, the reference prediction block is for representing the predicted value of the decoded block to be decoded in IBC mode, which is an intra-block copy mode, and the BV of the reference prediction block is for indicating the position of the reference prediction block in the reconstructed image block.
[0393] In one embodiment, the BV of the reference prediction block is encoded in multiple streams, and the BV of the reference prediction block is obtained by encoding the encoding blocks of multiple channels in IBC mode.
[0394] In one embodiment, the acquisition module 2801 is further used to acquire processing coefficients corresponding to coding units, the processing coefficients including one or more of residual coefficients and transformation coefficients. The processing module 2802 is further used to divide the processing coefficients into multiple groups based on a count threshold, the number of processing coefficients in each of the multiple groups being less than or equal to the count threshold.
[0395] In one embodiment, the processing module 2802 is further used to analyze the code stream obtained by encoding the encoding unit, determine the processing coefficients corresponding to the encoding unit, and decode the code stream based on the processing coefficients, wherein the processing coefficients include one or more of the residual coefficients and the transformation coefficients, and the processing coefficients include processing coefficients of multiple groups, the number of processing coefficients in each group being less than or equal to a threshold.
[0396] In one embodiment, the acquisition module 2801 is used to acquire an encoding unit and obtain complexity information for each of the P channels, where the encoding unit includes the encoding blocks of the P channels, P is an integer of 2 or more, and the complexity information is used to represent the difference in pixel values of the encoding blocks of each channel. The processing module 2802 is used to encode the complexity information of each channel's encoding block in the substream obtained by encoding the encoding blocks of the P channels.
[0397] In one embodiment, complexity information includes complexity levels. Specifically, the processing module 2802 is used to encode the complexity level of each channel's encoded block in a substream obtained by encoding P channel encoded blocks.
[0398] In one embodiment, complexity information includes a complexity level and a first reference coefficient, the first reference coefficient being used to represent the ratio relationship between complexity levels of coding blocks of different channels. Specifically, the processing module 2802 is used to encode the complexity level of the Q channel coding blocks in each of the substreams obtained by encoding the Q channel coding blocks out of P channels, and to encode the first reference coefficient in the substreams obtained by encoding the PQ channel coding blocks out of P channels, where Q is an integer less than P.
[0399] In one embodiment, complexity information includes a complexity level, a reference complexity level, and a second reference coefficient. The reference complexity level includes one of a first complexity level, a second complexity level, and a third complexity level. The first complexity level is the maximum complexity level of the coded blocks of PQ channels out of P channels, where Q is an integer less than P. The second complexity level is the minimum complexity level of the coded blocks of PQ channels out of P channels. The third complexity level is the average complexity level of the coded blocks of PQ channels out of P channels. The second reference coefficient is used to represent the relationship and / or ratio relationship between the complexity levels of the coded blocks of PQ channels out of P channels. Specifically, the processing module 2802 is used to encode the complexity level of the Q channel coding blocks in each of the substreams obtained by encoding the Q channel coding blocks out of P channels, and to encode the reference complexity level and a second reference coefficient in the substream obtained by encoding the PQ channel coding blocks out of P channels.
[0400] In one embodiment, the processing module 2802 further analyzes the code stream obtained by encoding the encoding unit, analyzes the complexity information of each channel's encoding block in the substream obtained by encoding the P channel encoding blocks, determines the quantization parameter of each channel's encoding block based on the complexity information of each channel's encoding block, and is used to decode the code stream based on the quantization parameter of each channel's encoding block. The encoding unit includes P channel encoding blocks, where P is an integer greater than or equal to 2, and the code stream includes a plurality of substreams, each corresponding one-to-one to the P channels, in which the P channel encoding blocks have been encoded. The complexity information is used to represent the difference in pixel values of each channel's encoding block.
[0401] In one embodiment, complexity information includes complexity levels. The processing module 2802 is further used to analyze the complexity level of each channel's encoded block in the substream obtained by encoding the encoded blocks of P channels.
[0402] In one embodiment, complexity information includes a complexity level and a first reference coefficient, the first reference coefficient being used to represent the ratio relationship between the complexity levels of the coded blocks of different channels. Specifically, the processing module 2802 analyzes the complexity level of the coded blocks of Q channels in each of the substreams obtained by coding the coded blocks of Q channels out of P channels, and analyzes the first reference coefficient in the substream obtained by coding the coded blocks of PQ channels out of P channels, and is used to determine the complexity level of the coded blocks of PQ channels based on the first reference coefficient and the complexity level of the coded blocks of Q channels, where Q is an integer less than P.
[0403] In one embodiment, complexity information includes a complexity level, a reference complexity level, and a second reference coefficient. The reference complexity level includes one of a first complexity level, a second complexity level, and a third complexity level. The first complexity level is the maximum complexity level of the coded blocks of PQ channels out of P channels, where Q is an integer less than P. The second complexity level is the minimum complexity level of the coded blocks of PQ channels out of P channels. The third complexity level is the average complexity level of the coded blocks of PQ channels out of P channels. The second reference coefficient is used to represent the relationship and / or ratio relationship between the complexity levels of the coded blocks of PQ channels out of P channels. Specifically, the processing module 2802 is used to determine the complexity level of the coded blocks of PQ channels in each of the substreams obtained by encoding the coded blocks of Q channels out of P channels on the encoding side, to determine the complexity level of the coded blocks of PQ channels in the substreams obtained by encoding the coded blocks of PQ channels out of P channels, and to determine the complexity level of the coded blocks of PQ channels based on the complexity level of the coded blocks of Q channels and the second reference coefficient.
[0404] In one embodiment, the acquisition module 2801 is further used to acquire an encoding unit, which comprises an encoding block of multiple channels, and the encoding unit is an image block in the image to be processed. The processing module 2802 is used, if the image format of the image to be processed is a preset format, to merge the substreams obtained by encoding the encoding blocks of at least two preset channels among the multiple channels into a single merged substream.
[0405] In one embodiment, the processing module 2802 is used to analyze the code stream obtained by encoding the encoding unit, determine the merged substream, decode the code stream based on the merged substream in which at least two substreams are merged, and obtain the encoding unit, which includes an encoding block of multiple channels, and the encoding unit is an image block in the image to be processed, and the merged substream is obtained by merging substreams obtained by encoding the encoding blocks of at least two preset channels out of the multiple channels, if the image format of the image to be processed is a preset format.
[0406] In one embodiment, the acquisition module 2801 is further used to acquire an encoding unit, the encoding unit comprising encoding blocks of multiple channels. The processing module 2802 is further used to encode a preset codeword in a target substream that satisfies a preset condition until the target substream no longer satisfies the preset condition, the target substream being one of multiple substreams, and the multiple substreams being a code stream obtained by encoding encoding blocks of multiple channels.
[0407] In one embodiment, the preset condition includes that the number of bits in the substream is less than a preset first bit threshold.
[0408] In one embodiment, the pre-set conditions include the presence of encoded coded blocks in the coded stream of the coding unit, where the number of bits is less than a pre-set second bit threshold.
[0409] In one embodiment, the processing module 2802 is further used to analyze the code stream obtained by encoding the encoding unit, determine the number of codewords, and decode the code stream based on the number of codewords. The encoding unit includes encoding blocks for multiple channels, and the number of codewords indicates the number of preset codewords encoded in a target substream that satisfies pre-set conditions. The preset codewords are encoded in the target substream if such a target substream exists that satisfies pre-set conditions.
[0410] In one embodiment, the acquisition module 2801 is further used to acquire an encoding unit, which includes encoding blocks for multiple channels. The processing module 2802 is further used to determine a target encoding mode corresponding to each of the encoding blocks for each channel among the encoding blocks for multiple channels, based on a preset expansion rate, and to encode the encoding blocks in each component in the target encoding mode so that the current expansion rate is less than or equal to the preset expansion rate.
[0411] In one embodiment, the preset expansion rate includes a first preset expansion rate, the current expansion rate is equal to the quotient of the number of bits of the maximum substream and the number of bits of the minimum substream, the maximum substream is the substream with the largest number of bits among the multiple substreams obtained by encoding multiple channel encoding blocks, and the minimum substream is the substream with the smallest number of bits among the multiple substreams obtained by encoding multiple channel encoding blocks.
[0412] In one embodiment, the preset expansion rate includes a second preset expansion rate. The current expansion rate is equal to the quotient between the number of bits of the encoding block with the largest number of bits and the number of bits of the encoding block with the smallest number of bits among the encoding blocks of the multiple encoded channels.
[0413] In one embodiment, the processing module 2802 is further used to determine the current expansion rate and, if the current expansion rate is greater than the first preset expansion rate, to encode a preset codeword in the minimum substream such that the current expansion rate becomes less than or equal to the first preset expansion rate.
[0414] In one embodiment, the processing module 2802 is further used to determine the current expansion rate and, if the current expansion rate is greater than a second preset expansion rate, to encode a preset codeword in the encoding block with the smallest number of bits so that the current expansion rate becomes less than or equal to the second preset expansion rate.
[0415] In one embodiment, the processing module 2802 further analyzes the code stream obtained by encoding an encoding unit including encoding blocks of multiple channels, determines the number of preset codewords based on the code stream, and is used to decode the code stream based on the number of preset codewords.
[0416] In Figure 28, the partitioning of modules is schematic and represents only logical function partitions; in actual implementation, partitioning can be done in a different way. For example, two or more functions can be integrated into a single processing module. The integrated module may be implemented in hardware form or in the form of a software function module.
[0417] In exemplary embodiments, embodiments of the present application provide a readable storage medium containing an execution instruction, which, when executed by an image encoding / decoding device, causes the image encoding / decoding device to implement any one of the methods provided by the above embodiments.
[0418] In an exemplary embodiment, an embodiment of the present application provides a computer program product including an execution instruction, which, when executed by an image encoding / decoding device, causes the image encoding / decoding device to implement one of the methods provided by the above embodiment.
[0419] In an exemplary embodiment, an embodiment of the present application provides a chip comprising a processor and an interface, wherein the processor is coupled to memory via the interface, and when the processor executes a computer program in memory, or when an image encoding / decoding device executes an instruction, one of the methods provided by the above embodiment is performed.
[0420] In the embodiments described above, all or part of the implementation can be achieved by software, hardware, firmware, or any combination thereof. When implemented by a software program, all or part of the implementation can be achieved in the form of a computer program product. This computer program product includes one or more computer execution instructions. When the computer execution instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable device. The computer execution instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer execution instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via a wired connection (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless connection (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer, or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0421] While this specification has described the present application with reference to various embodiments, a person skilled in the art can understand and implement other variations of the disclosed embodiments by referring to the accompanying drawings, disclosures, and claims. In the claims, the term “comprising” does not exclude other components or steps, and “one” or “one” does not exclude the case of multiple components. A single processor or other means may implement some of the functions enumerated in the claims. Although several methods are described in the dependent claims which differ from each other, it is not impossible that these methods may be combined to produce a good effect.
[0422] While the present application has been described by combining specific features and embodiments, it is clear that various modifications and combinations are possible without departing from the spirit and scope of the present application. Therefore, this specification and drawings are merely illustrative descriptions as defined by the scope of the appended claims and are deemed to cover any and all modifications, alterations, combinations, or equivalents within the scope of the present application. Clearly, a person skilled in the art can make various modifications and alterations of the present application without departing from the spirit and scope of the present application. Thus, if these modifications and alterations of the present application fall within the scope of the claims and the equivalent art, the present application is also intended to include these modifications and alterations.
[0423] The above is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto, and any changes or substitutions within the scope of the art disclosed in the present application should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be the same as the scope of protection of the claims.
Claims
1. The process involves analyzing a code stream obtained by encoding an encoding unit, wherein the encoding unit includes encoding blocks for multiple channels. The number of codewords is determined, and the number of codewords is intended to indicate the number of preset codewords encoded in a target substream that satisfies a predetermined condition, and the preset codewords are encoded in the target substream if such a target substream exists that satisfies a predetermined condition. The process includes decoding the code stream based on the number of codewords, Determining the number of the aforementioned codewords is Currently, the expansion rate needs to be determined, This includes determining the number of codewords based on the current expansion rate and a first preset expansion rate. Image decoding method.
2. Decoding the code stream based on the number of codewords is: The process includes removing preset codewords based on the number of codewords, and decoding the code stream from which the preset codewords have been removed. The image decoding method according to claim 1.
3. The current expansion rate is equal to the quotient of the number of bits of the largest substream and the number of bits of the smallest substream. The image decoding method according to claim 1.
4. A video decoder comprising a processor and memory, The memory stores instructions that can be executed by the processor. When the processor executes the instruction, the video decoder, The method is configured to perform an analysis on a code stream obtained by encoding a coding unit, wherein the coding unit includes coding blocks for multiple channels, the current expansion rate is determined, the number of codewords is determined based on the current expansion rate and a first preset expansion rate, the number of codewords is intended to indicate the number of preset codewords encoded in a target substream that satisfies a preset condition, the preset codewords are encoded in the target substream if such a target substream exists, and the code stream is decoded based on the number of codewords. Video decoder.
5. A non-temporary readable storage medium that includes software instructions, When the software instruction is executed by the image decoding device, the image decoding device implements the image decoding method described in any one of claims 1 to 3. A non-temporary readable storage medium.
Citation Information
Patent Citations
Intra prediction from prediction block
JP2017508345A
Substream multiplexing for display stream compression
JP2019522413A
Parallel Entropy Coding
JP2024510268A
Substream multiplexing for display stream compression
US20180343471A1
Method and apparatus for video coding
WO2022132251A1