Image decoding method, video decoder and non-transitory readable storage medium
By controlling bit allocation in sub-streams through multiple encoding modes, the method addresses unequal embedding speeds and reduces hardware costs in video encoding and decoding, ensuring efficient data integrity.
Patent Information
- Application Number
- JP2025517178
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-20
- Filing Date
- 2023-09-12
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2043-09-12
AI Technical Summary
Substream parallelism in video encoding leads to unequal embedding speeds and increased hardware costs due to the need for larger sub-stream buffers to ensure data integrity.
Implementing multiple encoding modes to control the number of bits in maximum and minimum sub-streams, reducing the difference in filling speeds and buffer sizes, thereby minimizing hardware costs.
Reduces hardware costs by optimizing sub-stream buffer sizes and embedding speeds, ensuring efficient data integrity in video encoding and decoding processes.
Smart Images

Figure 2025531374000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority from Chinese Patent Application No. 202211146464.4, filed on September 20, 2022, the entire contents of which are incorporated herein by reference.
[0002] The present application relates to the technical field of video encoding and video decoding, and in particular to an image encoding method, an image decoding method, an apparatus, and a storage medium. [Background technology]
[0003] To improve the performance of the encoder, a technique called substream parallelism has been proposed.
[0004] Substream parallelism means using multiple entropy coders to encode syntax elements of different channels to obtain multiple substreams, embedding the multiple substreams into corresponding substream buffers, and interleaving the substreams in the substream buffers into a bitstream (also called a codestream) according to a preset interleaving rule.
[0005] However, considering the dependency between sub-streams, the embedding speeds of sub-streams in different sub-stream buffers are different, and in the same time, the number of sub-stream bits embedded in a sub-stream buffer with a fast embedding speed is greater than that in a sub-stream buffer with a slow embedding speed. Therefore, to ensure the integrity of the embedded data, all sub-stream buffers need to be set large, which increases the hardware cost. Summary of the Invention
[0006] Based on the above technical problem, the present application provides an image encoding method and an image decoding method, as well as a device and a storage medium, that can encode in multiple improved encoding modes, rationally allocate sub-stream buffer space, and reduce hardware costs.
[0007] In a first aspect, the present application provides an image encoding method, the method comprising: obtaining a coding unit, the coding unit including coding blocks of a plurality of channels; encoding a preset codeword in a target substream that satisfies a predetermined condition until the target substream no longer satisfies the predetermined condition, the target substream being a substream among a plurality of substreams, the plurality of substreams being codestreams obtained by encoding coding blocks of the plurality of channels.
[0008] In a second aspect, the present application provides an image encoding method, the method comprising: obtaining a coding unit, the coding unit including coding blocks of a plurality of channels; encoding the coding blocks of the channels based on preset expansion rates so that the current expansion rate is equal to or less than a preset expansion rate, the preset expansion rate including a first preset expansion rate, the value of the current expansion rate being derived from a quotient of a number of bits of a maximum substream and a number of bits of a minimum substream, the maximum substream being the substream with the largest number of bits among a plurality of substreams obtained by encoding the coding blocks of the plurality of channels, and the minimum substream being the substream with the smallest number of bits among a plurality of substreams obtained by encoding the coding blocks of the plurality of channels; Includes:
[0009] In a third aspect, the present application provides an image encoding method, the method comprising: obtaining a coding unit, the coding unit including coding blocks of a plurality of channels; encoding a coding block of at least one channel among the plurality of channels in an IBC mode that is an intra block copy mode; Obtaining a block vector BV of a reference prediction block, where the BV of the reference prediction block is for indicating a position of the reference prediction block in a coded image block, and the reference prediction block is for representing a predicted value of the coding block coded in IBC mode; and encoding a BV of the reference prediction block in at least one sub-stream obtained by encoding the coding block of the at least one channel in the IBC mode.
[0010] In a fourth aspect, the present application provides an image encoding method, the method comprising: Obtaining a coding unit, the coding unit being an image block in a target image, the coding unit including coding blocks of multiple channels; determining a first total code length, the first total code length being a total code length of a first stream obtained by encoding all of the encoding blocks of the plurality of channels in a corresponding target encoding mode, the target encoding mode including a first encoding mode, the first encoding mode being a mode for encoding sample values in the encoding blocks with a first fixed-length code, the code length of the first fixed-length code being equal to or less than an image bit width of the image to be processed, and the image bit width being for representing the number of bits required to store each sample in the image to be processed; If the first total code length is equal to or greater than the remaining capacity of a code stream buffer, encoding the coding blocks of the multiple channels in a fallback mode, and a mode flag of the fallback mode is the same as that of the first coding mode.
[0011] In a fifth aspect, the present application provides an image decoding method, the method comprising: analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; determining a number of codewords, the number of codewords being for indicating the number of preset codewords encoded in a target substream that satisfies a preset condition, the preset codewords being encoded into the target substream if a target substream that satisfies the preset condition exists; and decoding the codestream based on the number of codewords.
[0012] In a sixth aspect, the present application provides an image decoding method, the method comprising: analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; determining the current expansion rate; determining a number of codewords based on the current expansion factor and a first preset expansion factor, the number of codewords being for indicating a number of preset codewords encoded in a target substream that satisfies a preset condition, and the preset codewords being encoded into the target substream when a target substream that satisfies the preset condition exists; and decoding the codestream based on the number of codewords.
[0013] In a seventh aspect, the present application provides an image decoding method, the method comprising: analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of a plurality of channels, the codestream including a plurality of substreams in one-to-one correspondence with the plurality of channels, into which the coding blocks of the plurality of channels are encoded; determining a position of the reference prediction block in the plurality of substreams based on a block vector BV of the reference prediction block analyzed from at least one substream among the plurality of substreams, the reference prediction block representing a predicted value of a decoded block decoded in an IBC mode, which is an intra block copy mode, and the BV of the reference prediction block representing a position of the reference prediction block in a reconstructed image block; determining a predicted value of the decoded block to be decoded in IBC mode based on position information of the reference predicted block; and reconstructing a decoded block to be decoded in the IBC mode based on the predicted value.
[0014] In an eighth aspect, the present application provides an image decoding method, the method comprising: analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; a mode flag is analyzed from a substream obtained by encoding the coding blocks of the plurality of channels, and if a second total code length is greater than the remaining capacity of a code stream buffer, a target decoding mode of the substream is determined to be a fallback mode, wherein the mode flag indicates whether the coding blocks of the plurality of channels are encoded using the fallback mode, the second total code length is a total code length of a codestream obtained by encoding all of the coding blocks of the plurality of channels in a first coding mode, the first coding mode is a mode for encoding sample values in the coding blocks with a first fixed-length code, the code length of the first fixed-length code is equal to or less than an image bit width of the image to be processed, and the image bit width is for indicating the number of bits required to store each sample in the image to be processed; Analyzing preset flag bits in the substream to determine a target fallback mode, where the target fallback mode is one of the fallback modes, the preset flag bits are for indicating types of fallback modes to be used when encoding the coding blocks of the multiple channels, and the fallback modes include a first fallback mode and a second fallback mode; and decoding the sub-stream in the target fallback mode.
[0015] In a ninth aspect, the present application provides an image decoding method, the method comprising: encoding the coding unit to obtain a codestream; decoding a substream corresponding to the first channel in a first decoding mode.
[0016] In a tenth aspect, the present application provides an image decoding method, the method comprising: Obtaining a decoding unit, the decoding unit being an image block in a current image, the decoding unit including coding blocks of multiple channels; determining a first total code length, the first total code length being a total code length of a first code stream obtained by decoding all of the decoded blocks in the plurality of channels in the corresponding target decoding modes; If the first total code length is equal to or greater than the remaining capacity of a codestream buffer, encoding the decoded blocks in the multiple channels in a fallback mode.
[0017] In an eleventh aspect, the present application provides an image decoding method, the method comprising: Obtaining a decoding unit, the decoding unit including coding blocks of multiple channels; decoding a decoding block in at least one channel of the plurality of channels in an IBC mode; Obtaining a BV of a reference predicted block; and decoding a BV of a reference prediction block in at least one substream obtained by encoding a decoded block in at least one channel in IBC mode.
[0018] In a twelfth aspect, the present application provides an image encoding apparatus, the apparatus comprising: an acquisition module for acquiring a coding unit, the coding unit including coding blocks of multiple channels; and a processing module that encodes preset codewords in a target substream that satisfies a preset condition until the target substream no longer satisfies the preset condition, the target substream being a substream among a plurality of substreams, the plurality of substreams being codestreams obtained by encoding coding blocks of the plurality of channels.
[0019] In a thirteenth aspect, the present application provides an image encoding apparatus, the apparatus comprising: an acquisition module for acquiring a coding unit, the coding unit including coding blocks of multiple channels; a processing module for encoding the coding block of each channel based on the preset expansion rate such that the current expansion rate is equal to or less than the preset expansion rate.
[0020] In a fourteenth aspect, the present application provides an image coding apparatus, the apparatus comprising: an acquisition module for acquiring a coding unit, the coding unit including coding blocks of multiple channels; and a processing module used for encoding a coding block of at least one channel among the plurality of channels in an IBC mode, which is an intra block copy mode; obtaining a block vector BV of a reference prediction block, wherein the BV of the reference prediction block indicates the position of the reference prediction block in a coded image block and represents a predicted value of a coding block coded in the IBC mode; and coding the BV of the reference prediction block in a substream obtained by coding the coding block in each of the plurality of channels in the IBC mode.
[0021] In a fifteenth aspect, the present application provides an image encoding apparatus, the apparatus comprising: an acquisition module for acquiring a coding unit, the coding unit being an image block in a target image, the coding unit including coding blocks of multiple channels; a processing module used for determining a first total code length, the first total code length being the total code length of a first stream obtained by encoding all of the encoding blocks of the plurality of channels in corresponding target encoding modes, the target encoding mode including a first encoding mode, the first encoding mode being a mode for encoding sample values in the encoding blocks with a first fixed-length code, the code length of the first fixed-length code being equal to or less than an image bit width of the image to be processed, the image bit width being intended to represent the number of bits required to store each sample in the image to be processed; and if the first total code length is equal to or greater than the remaining capacity of a code stream buffer, encoding the encoding blocks of the plurality of channels in a fallback mode, the mode flag of the fallback mode being the same as that of the first encoding mode.
[0022] In a sixteenth aspect, the present application provides an image decoding device, the device comprising: The present invention includes a processing module for analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; determining a number of codewords, the number of codewords indicating the number of preset codewords encoded in a target substream that satisfies a predetermined condition, the preset codewords being coded into the target substream if a target substream that satisfies the predetermined condition exists; and decoding the codestream based on the number of codewords.
[0023] In a seventeenth aspect, the present application provides an image decoding device, the device comprising: The present invention includes a processing module for analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; determining a current expansion factor; determining a number of codewords based on the current expansion factor and a first preset expansion factor, the number of codewords indicating the number of preset codewords encoded in a target substream that satisfies a preset condition, the preset codewords being coded into the target substream if a target substream that satisfies the preset condition exists; and decoding the codestream based on the number of codewords.
[0024] In an eighteenth aspect, the present application provides an image decoding device, the device comprising: The present invention includes a processing module for: analyzing a code stream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels, the code stream including multiple sub-streams into which the coding blocks of the multiple channels are encoded, the multiple sub-streams corresponding one-to-one to the multiple channels; determining a position of the reference prediction block in the multiple sub-streams based on a block vector BV of the reference prediction block analyzed from at least one sub-stream of the multiple sub-streams, the reference prediction block representing a predicted value of a decoded block decoded in an IBC mode, which is an intra block copy mode, and the BV of the reference prediction block indicating a position of the reference prediction block in a reconstructed image block; determining a predicted value of the decoded block decoded in the IBC mode based on position information of the reference prediction block; and reconstructing the decoded block decoded in the IBC mode based on the predicted value.
[0025] In a nineteenth aspect, the present application provides an image decoding device, the device comprising: analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; analyzing a mode flag from a substream obtained by encoding the coding blocks of the multiple channels, and determining that a target decoding mode of the substream is a fallback mode if a second total code length is greater than the remaining capacity of a codestream buffer, the mode flag indicating whether the decoding blocks of the multiple channels are encoded using a fallback mode, the second total code length being a total code length of a codestream obtained by encoding all of the coding blocks of the multiple channels using a first coding mode, and the first coding mode being a first fixed-length code; the code length of the first fixed-length code is equal to or less than an image bit width of a target image, the image bit width representing the number of bits required to store each sample in the target image; analyzing preset flag bits in the substream to determine a target fallback submode, the target fallback mode being one of the fallback modes, the preset flag bits indicating the type of fallback mode to be used when encoding the coding blocks of the multiple channels, the fallback modes including a first fallback mode and a second fallback mode; and decoding the substream in the target fallback mode.
[0026] In a twentieth aspect, the present application provides a video encoder comprising a processor and a memory, wherein the memory stores instructions executable by the processor, and the processor is configured to, when executing the instructions, cause the video encoder to realize an image encoding method described in any one of the first to fourth aspects above.
[0027] In a 21st aspect, the present application provides a video decoder, the video decoder comprising a processor and a memory, the memory storing instructions executable by the processor, the processor being configured to, when executing the instructions, cause the video decoder to realize an image decoding method according to any one of the above 5th to 11th aspects.
[0028] In a 22nd aspect, the present application provides a computer program product, which, when executed by an image encoding device, causes the image encoding device to realize the image encoding method described in any one of the first to fourth aspects above.
[0029] In a 23rd aspect, the present application provides a readable storage medium, the readable storage medium including software instructions, which, when executed by an image encoding device, cause the image encoding device to realize the image encoding method described in any one of the first to fourth aspects, and which, when executed by an image decoding device, cause the image decoding device to realize the image decoding method described in any one of the fifth to 11th aspects.
[0030] In a 24th aspect, the present application provides a chip comprising a processor and an interface, the processor being coupled to a memory via the interface, and when the processor executes a computer program in the memory or the image encoding device executes an instruction, the methods described in the first to fourth aspects are performed. [Brief explanation of the drawings]
[0031] In order to more clearly describe the technical solutions of the embodiments of the present application, the drawings required in the embodiments will be briefly described below. Note that the drawings described below are only some of the embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without any creative efforts.
[0032] [Figure 1] FIG. 1 is a schematic diagram of the configuration of the sub-stream parallel technology. [Figure 2]FIG. 2 is a schematic diagram of the sub-stream interleaving process at the encoding side. [Figure 3] FIG. 3 is a schematic diagram of the format of a sub-stream interleaving unit. [Figure 4] FIG. 4 is a schematic diagram of the de-substream interleaving process at the decoding side. [Figure 5] FIG. 5 is another schematic diagram of the sub-stream interleaving process at the encoding side. [Figure 6] FIG. 6 is a schematic diagram of the configuration of a video encoding / decoding system provided by an embodiment of the present application. [Figure 7] FIG. 7 is a structural schematic diagram of a video encoder provided by an embodiment of the present application. [Figure 8] FIG. 8 is a structural schematic diagram of a video decoder provided by an embodiment of the present application. [Figure 9] FIG. 9 is a schematic flowchart of video encoding and decoding provided by an embodiment of the present application. [Figure 10] FIG. 10 is a schematic diagram illustrating the configuration of an image encoding device and an image decoding device provided by an embodiment of the present application. [Figure 11] FIG. 11 is a schematic flow chart of an image encoding method provided by an embodiment of the present application. [Figure 12] FIG. 12 is a schematic flowchart of an image decoding method provided by an embodiment of the present application. [Figure 13] FIG. 13 is a schematic flowchart of another image decoding method provided by an embodiment of the present application. [Figure 14] FIG. 14 is a schematic flowchart of another image decoding method provided by an embodiment of the present application. [Figure 15] FIG. 15 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. [Figure 16] FIG. 16 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. [Figure 17]FIG. 17 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. [Figure 18] FIG. 18 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. [Figure 19] FIG. 19 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. [Figure 20] FIG. 20 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. [Figure 21] FIG. 21 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. [Figure 22] FIG. 22 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. [Figure 23] FIG. 23 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. [Figure 24] FIG. 24 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. [Figure 25] FIG. 25 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. [Figure 26] FIG. 26 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. [Figure 27] FIG. 27 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. [Figure 28] FIG. 28 is a schematic diagram illustrating the configuration of an image encoding device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION
[0033] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.
[0034] In the description of this application, unless otherwise specified, " / " means "or," for example, A / B means A or B. The term "and / or" in this application is only used to describe the relationship between related objects and indicates that three types of relationships can exist, for example, A and / or B can indicate three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc. do not limit the quantity or execution order, and the terms "first," "second," etc. do not necessarily limit "different."
[0035] It should be noted that, in this application, terms such as "exemplary" or "for example" are used as examples, illustrations, or explanations. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as preferred or advantageous over other embodiments or designs. Rather, use of terms such as "exemplary" or "for example" is intended to express the relevant concept in a concrete manner.
[0036] To improve the performance of the encoder, a technique called substream parallelism (also called substream interleaving) has been proposed.
[0037] For the encoding side, substream parallelism means that the encoding side encodes syntax elements of a coding block (CB) in different channels (e.g., luma channel, first chroma channel, and second chroma channel) of a coding unit (CU) using multiple entropy encoders to obtain multiple substreams, and then interleaves the multiple substreams into a bitstream in fixed-size packets. Correspondingly, for the decoding side, substream parallelism means that the decoding side decodes different substreams in parallel using different entropy decoders.
[0038] For example, Figure 1 is a schematic diagram of the configuration of the sub-stream parallel technology. As shown in Figure 1, taking the encoding side as an example, the specific application timing of the sub-stream parallel technology is after encoding syntax elements (e.g., transform coefficients, quantization coefficients, etc.). The encoding flow of other parts of Figure 1 can refer to the description of the video encoding / decoding system provided by the embodiment of the present application below, and will not be further described here.
[0039] Illustratively, Figure 2 is a schematic diagram of a sub-stream interleaving process on the encoding side. As shown in Figure 2, for example, assuming that the image block to be encoded includes three channels, the encoding modules (such as a prediction module, a transform module, and a quantization module) output syntax elements and quantized transform coefficients for the three channels. Then, the syntax elements and quantized transform coefficients for the three channels are encoded by entropy coder 1, entropy coder 2, and entropy coder 3, respectively, to obtain sub-streams corresponding to the three channels. The sub-streams corresponding to the three channels can then be stored in encoding sub-stream buffer 1, encoding sub-stream buffer 2, and encoding sub-stream buffer 3. The sub-stream interleaving module interleaves the sub-streams in encoding sub-stream buffer 1, encoding sub-stream buffer 2, and encoding sub-stream buffer 3, and finally outputs a bitstream (also referred to as a codestream) in which multiple sub-streams are interleaved.
[0040] 3 is a schematic diagram of a format of a sub-stream interleave unit. As shown in FIG. 3, a sub-stream may be composed of sub-stream interleave units, which may also be called a sub-stream segment. The length of a sub-stream segment is N bits, and the sub-stream segment includes an M-bit data header and an NM-bit data body.
[0041] Here, the data header is used to indicate the substream to which the current substream belongs. N can be 512 and M can be 2.
[0042] For example, Figure 4 is a schematic diagram of a de-substream interleaving process on the decoding side. As shown in Figure 4, similarly, taking the example of an image block to be coded containing three channels, when a bitstream output by the coding side is input to the decoding side, the sub-stream interleaving module on the decoding side first performs a de-substream interleaving process on the bitstream to divide the bitstream into sub-streams corresponding to the three channels, and the sub-streams corresponding to the three channels can be stored in decoding sub-stream buffer 1, decoding sub-stream buffer 2, and decoding sub-stream buffer 3. For example, taking the sub-stream segment shown in Figure 3 above as an example, the decoding side can extract an N-bit long packet from the bitstream each time. By analyzing the M-bit data header of the packet, the target sub-stream to which the current sub-stream segment belongs is obtained, and the data body remaining in the current sub-stream segment is stored in the decoding sub-stream buffer corresponding to the target sub-stream.
[0043] Entropy decoder 1 decodes the substream in decoded substream buffer 1 to obtain syntax elements and quantized transform coefficients for one channel. Entropy decoder 2 decodes the substream in decoded substream buffer 2 to obtain syntax elements and quantized transform coefficients for another channel. Entropy decoder 3 decodes the substream in decoded substream buffer 3 to obtain syntax elements and quantized transform coefficients for yet another channel. Finally, the syntax elements and quantized transform coefficients for each of the three channels are input to a subsequent decoding module for decoding processing to obtain a decoded image.
[0044] The process of sub-stream interleaving will be explained below by taking the encoding side as an example.
[0045] Each sub-stream segment in the coded sub-stream buffer includes coded bits generated by coding at least one image block. During sub-stream interleaving, an order marking is first performed on the image block corresponding to the first bit of the data body of each sub-stream segment, and different sub-streams can be interleaved using this order marking in the sub-stream interleaving process.
[0046] In one embodiment, the sub-stream interleaving process can mark image blocks by a block count queue.
[0047] For example, the block count queue is implemented as a first-in, first-out queue. The encoding side sets a count of the currently encoded image block (block count) and sets one block count queue (counter queue[ss_idx]) for each substream. When encoding each slice starts, the encoding side initializes the block count to 0, initializes each counter queue[ss_idx] to null, and puts one 0 into each counter queue[ss_idx].
[0048] After each image block (or coding unit (CU)) is coded, the block count queue is updated. The update process is as follows:
[0049] Step 1: Set the count of the current coded image block to +1, i.e., block count+=1.
[0050] Step 2: Select one sub-stream ss_idx.
[0051] Step 3: Calculate num_in_buffer[ss_idx], the number of sub-stream segments that can be constructed in the encoding sub-stream buffer corresponding to this sub-stream ss_idx. Let the encoding sub-stream buffer corresponding to this sub-stream ss_idx be buffer[ss_idx], and the amount of data contained in buffer[ss_idx] be buffer[ss_idx].fullness. Similarly, using the example shown in Figure 3 above, where the size of the sub-stream segment is N bits and the sub-stream segment contains an M-bit data header, num_in_buffer[ss_idx] can be calculated using the following formula (1):
[0052] JPEG2025531374000002.jpg8122Here, " / " indicates division.
[0053] Step 4: Compare the length of the current block count queue, num_in_queue[ss_idx], with the number of substream segments that can be constructed in the encoding substream buffer, num_in_buffer[ss_idx]. If they are equal, put the count of the current encoded image block into this block count queue, i.e., counter_queue[ss_idx].push(block_count).
[0054] Step 5: Return to step 2 and process the next substream until all substreams have been processed.
[0055] After the update of the block count queue is completed, the encoding side can interleave the sub-streams in each encoding sub-stream buffer, and the interleaving process is as follows:
[0056] Step 1: Select one sub-stream ss_idx.
[0057] Step 2: Determine whether the amount of data buffer[ss_idx].fullness contained in the encoded substream buffer buffer[ss_idx] corresponding to this substream ss_idx is equal to or greater than NM. If it is equal to or greater than NM, execute step 3. If it is not equal to or greater than NM, execute step 6.
[0058] Step 3: Determine whether the queue head element value in the block count queue of this substream ss_idx is the minimum value in the block count queues of all substreams. If it is the minimum value, execute step 4. If it is not the minimum value, execute step 6.
[0059] Step 4: Build one sub-stream segment using the data in the current encoding sub-stream buffer. For example, retrieve NM-bit data from the encoding sub-stream buffer buffer[ss_idx], add an M-bit data header, set the data in the data header to ss_idx, combine the M-bit data header and the retrieved NM-bit data into an N-bit sub-stream segment, and send the sub-stream segment to the final bitstream output from the encoding side.
[0060] Step 5: Pop up (or remove) the queue-top element of the block count queue for this substream ss_idx, i.e., counter_queue[ss_idx].pop().
[0061] Step 6: Return to step 1 and process the next substream until all substreams have been processed.
[0062] In one embodiment, if the current image block is the last image block of a slice, the encoding side can also perform the following steps to package the data remaining in the encoding sub-stream buffer after the above interleaving process.
[0063] Step 1: Determine whether there is currently at least one non-null in all encoding sub-stream buffers. If there is, execute step 2. If there is not, end.
[0064] Step 2: Select one sub-stream ss_idx.
[0065] Step 3: Determine whether the queue head element value of the block count queue of this substream ss_idx is the minimum value of the block count queues of all substreams. If it is the minimum value, execute step 4. If it is not the minimum value, execute step 6.
[0066] Step 4: If the amount of data in the encoding substream buffer buffer[ss_idx] corresponding to this substream ss_idx is less than NM bits, fill this encoding substream buffer buffer[ss_idx] with 0 until the data in this encoding substream buffer buffer[ss_idx] reaches NM bits. At the same time, pop up (or delete) the queue head element value of the block count queue for this substream, i.e., counter_queue[ss_idx].pop(), and insert MAX_INT, which represents the maximum value within the data range, i.e., counter_queue[ss_idx].push(MAX_INT).
[0067] Step 5: Construct one sub-stream segment. Step 5 here can refer to step 4 in the above sub-stream interleaving process, and will not be further described here.
[0068] Step 6: Return to step 2 to process the next substream. Once all substreams have been processed, return to step 1.
[0069] For example, Figure 5 is another schematic diagram of the sub-stream interleaving process on the encoding side. As shown in Figure 5, for example, the image block to be encoded has three channels, and the encoding sub-stream buffers can include encoding sub-stream buffer 1, encoding sub-stream buffer 2, and encoding sub-stream buffer 3, respectively.
[0070] The substream segments in encoding substream buffer 1 are 1_1, 1_2, 1_3, and 1_4, from the front. The flags in block count queue 1 of substream 1 corresponding to encoding substream buffer 1 are 12, 13, 27, and 28, respectively. The substream segments in encoding substream buffer 2 are 2_1 and 2_2, from the front. The flags in block count queue 2 of substream 2 corresponding to encoding substream buffer 2 are 5 and 71, respectively. The substream segments in encoding substream buffer 3 are 3_1, 3_2, and 3_3, from the front. The flags in block count queue 3 of substream 3 corresponding to encoding substream buffer 3 are 6, 13, and 25, respectively. The substream interleaving module can interleave the substream segments in encoding substream buffer 1, encoding substream buffer 2, and encoding substream buffer 3 in the order of the flags in encoding substream buffer 1, encoding substream buffer 2, and encoding substream buffer 3.
[0071] The order of the substream segments in the interleaved codestream is: substream segment 2_1 corresponding to minimum flag 5, substream segment 3_1 corresponding to flag 6, substream segment 1_1 corresponding to flag 12, substream segment 1_2 corresponding to flag 13, substream segment 3_2 corresponding to flag 13, substream segment 3_3 corresponding to flag 25, substream segment 1_3 corresponding to flag 27, substream segment 1_4 corresponding to flag 28, and substream segment 2_2 corresponding to flag 71.
[0072] However, multiple substreams have dependencies when interleaved. Take the encoding side as an example, the embedding speeds of substreams in different substream buffers are different, and the number of substreams embedded in a substream buffer with a fast embedding speed at the same time is greater than that of a substream buffer with a slow embedding speed. The substream buffer with a fast embedding speed continues to embed data simultaneously while waiting for the substream buffer with a slow embedding speed to embed one substream segment. In order to ensure the integrity of the embedded data, all substream buffers need to be set large, which increases hardware costs.
[0073] In this case, the embodiments of the present application provide an image encoding method, an image decoding method, an apparatus, and a storage medium, and by encoding using multiple improved encoding modes, the number of bits of the maximum encoding module and the number of bits of the minimum encoding module can be controlled, thereby reducing the theoretical expansion rate of the encoded block to be encoded and reducing the difference in filling speed between different sub-stream buffers, thereby reducing the size of the preset space of the sub-stream buffer and reducing hardware costs.
[0074] The following description will be given with reference to the drawings.
[0075] 6 is a schematic diagram of a configuration of a video encoding / decoding system provided by an embodiment of the present application. As shown in FIG. 6, the video encoding / decoding system includes a source device 10 and a target device 11.
[0076] The source device 10 generates encoded video data and is also referred to as an encoding side, video encoding side, video encoding apparatus, or video encoding device. The target device 11 can decode the encoded video data generated by the source device 10 and is also referred to as a decoding side, video decoding side, video decoding apparatus, or video decoding device. The source device 10 and / or the target device 11 may include at least one processor and memory coupled to the at least one processor. This memory may include, but is not limited to, read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or any other medium usable for storing necessary program code in the form of computer-accessible instructions or data structures.
[0077] The source device 10 and the target device 11 may include a variety of devices, such as a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a mobile phone such as a so-called "smartphone," a television, a camera, a display device, a digital media player, a video game console, an in-vehicle computer, or other similar devices.
[0078] The target device 11 can receive encoded video data from the source device 10 via link 12. Link 12 can include one or more media and / or devices capable of moving the encoded video data from the source device 10 to the target device 11. In one example, link 12 can include one or more communication media that enable the source device 10 to transmit encoded video data directly to the target device 11 in real time. In this example, the source device 10 can modulate the encoded video data based on a communication standard (e.g., a wireless communication protocol) and transmit the modulated video data to the target device 11. The one or more communication media can include wireless and / or wired communication media, such as a radio frequency (RF) spectrum, one or more physical transmission lines, etc. The one or more communication media can form part of a packet-based network, such as a local area network, a wide area network, or a global network (e.g., the Internet). The one or more communication media can include routers, switches, base stations, or other devices that facilitate communication from the source device 10 to the target device 11.
[0079] In another example, source device 10 may output encoded video data to storage device 13 via output interface 103. Similarly, target device 11 may access encoded video data from storage device 13 via input interface 113. Storage device 13 may include a variety of locally accessible data storage media, such as a Blu-ray disc, a high-density digital video disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, or other suitable digital storage media for storing encoded video data.
[0080] In another example, storage device 13 may correspond to a file server or another intermediate storage device that stores encoded video data generated by source device 10. In this example, target device 11 can stream or download the stored video data from storage device 13. The file server may be any type of server capable of storing encoded video data and transmitting the encoded video data to a server of target device 11. For example, the file server may comprise a Global Wide Area Network (World Wide Web, Web) server (e.g., for a website), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, and a local disk drive.
[0081] The target device 11 can access the encoded video data via any standard data connection (e.g., an Internet connection). Examples of data connections include a wireless channel, a wired connection (e.g., a cable modem), or a combination of both suitable for accessing encoded video data stored on a file server. The encoded video data can be transmitted from the file server via streaming, download, or a combination of both.
[0082] It should be noted that the image encoding method and image decoding method provided by the embodiments of the present application are not limited to wireless application scenes.
[0083] Illustratively, the image encoding and decoding methods provided by the embodiments of the present application can be applied to video encoding and decoding used in multiple multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (e.g., via the Internet), encoding video data stored in a data storage medium, decoding video data stored in a data storage medium, or other applications. In some examples, the video encoding and decoding system may be configured to support one-way or two-way video transmission, such as video streaming, video playback, video broadcasting, and / or video telephony.
[0084] Note that the video encoding / decoding system shown in FIG. 6 is merely an example of a video encoding / decoding system and does not limit the video encoding / decoding system in the present application. The image encoding method and image decoding method provided by the present application can also be applied to scenes where there is no data communication between the encoding device and the decoding device. In other embodiments, the video data to be encoded or the encoded video data can be retrieved from a local memory or streamed over a network. The video encoding device can encode the video data to be encoded and store the encoded video data in a memory. The video decoding device can retrieve the encoded video data from the memory and decode the encoded video data.
[0085] 6 embodiment, source device 10 includes video source 101, video encoder 102, and output interface 103. In some embodiments, output interface 103 may include a modulator / demodulator (modem) and / or a transmitter. Video source 101 may include a video capture device (e.g., a video camera), a video archive containing captured video data, a video input interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources of video data.
[0086] Video encoder 102 can encode video data from video source 101. In some embodiments, source device 10 transmits the encoded video data directly to target device 11 via output interface 103. In other embodiments, the encoded video data may be stored in storage device 13 for later access by target device 11 for decoding and / or playback.
[0087] 6, target device 11 includes display device 111, video decoder 112, and input interface 113. In some embodiments, input interface 113 includes a receiver and / or a modem. Input interface 113 can receive encoded video data via link 12 and / or from storage device 13. Display device 111 may be integrated with target device 11 or may be external to target device 11. Generally, display device 111 displays decoded video data. Display device 111 may include various display devices, such as a liquid crystal display, a plasma display, an organic light-emitting diode display, or other types of display devices.
[0088] In one embodiment, the video encoder 102 and the video decoder 112 may be integrated with an audio encoder and decoder, respectively, and may include a multiplexer-multiplexer unit or other hardware and software suitable for encoding both audio and video in a common data stream or independent data streams.
[0089] The video encoder 102 and the video decoder 112 may include at least one microprocessor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. When the encoding methods provided herein are implemented in software, the instructions used in the software can be stored in a suitable non-volatile computer-readable storage medium and executed by at least one processor to implement the present invention.
[0090] The video encoder 102 and the video decoder 112 in this application may operate in accordance with a video compression standard (eg, HEVC) or other industry standards, and this application is not particularly limited thereto.
[0091] FIG. 7 is a schematic diagram of a video encoder 102 provided by an embodiment of the present application. As shown in FIG. 7, the video encoder 102 includes a prediction module 21, a transform module 22, a quantization module 23, an entropy encoding module 24, an encoded sub-stream buffer 25, and a sub-stream interleaving module 26, which perform prediction, transform, quantization, entropy encoding, and sub-stream interleaving processes, respectively. Here, the prediction module 21, the transform module 22, and the quantization module 23 are the encoding modules in FIG. 1 above. The video encoder 102 further includes a pre-processing module 20 and an adder 202, where the pre-processing module 20 includes a segmentation module and a code rate control module. To reconstruct video blocks, the video encoder 102 includes an inverse quantization module 27, an inverse transform module 28, an adder 201, and a reference image memory 29.
[0092] As shown in FIG. 7, the video encoder 102 receives video data, and the pre-processing module 20 receives input parameters for the video data. The input parameters include information such as the image resolution, image sampling format, pixel depth (bits per pixel, BPP), and bit width (also referred to as image bit width) of the video data. Here, BPP refers to the number of bits occupied per pixel. Bit width refers to the number of bits occupied by one pixel channel within a unit pixel. For example, if one pixel is represented by values of three YUV pixel channels, each occupying 8 bits, the bit width of the pixel is 8, and the BPP of the pixel is 3×8=24 bits.
[0093] The partitioning module in the pre-processing module 20 partitions an image into original blocks (also referred to as coding units (CUs)). An original block (or coding unit (CU)) may include coding blocks of multiple channels. For example, these multiple channels may be RGB channels or YUV channels, etc. The embodiments of the present application are not limited thereto. In one embodiment, this partitioning may include partitioning into slices, image blocks, or other large units, and partitioning video blocks according to a 4-tree structure of Largest Coding Units (LCUs) and CUs. Illustratively, the video encoder 102 encodes components of video blocks within a video slice to be coded. Generally, a slice may be partitioned into multiple original blocks (groups of original blocks called image blocks). The partitioning module typically determines the sizes of CUs, PUs, and TUs. The partitioning module is also used to determine the size of a code rate control unit. The code rate control unit is a basic processing unit in the code rate control module. For example, the code rate control module calculates complexity information for an original block based on the code rate control unit, and then calculates a quantization parameter for the original block based on the complexity information. Here, the partitioning policy of the partitioning module may be preset or may be adjusted incrementally based on an image during encoding. If the partitioning policy is a preset policy, the same partitioning policy is also preset on the decoding side accordingly, thereby obtaining the same image processing unit. This image processing unit is any one of the image blocks mentioned above, which has a one-to-one correspondence with the encoding side. If the partitioning policy is adjusted incrementally based on an image during encoding, the partitioning policy can be directly or indirectly incorporated into the codestream, and the decoding side accordingly obtains corresponding parameters from the codestream, obtains the same partitioning policy, and obtains the same image processing unit.
[0094] The code rate control module in the pre-processing module 20 is used to generate a quantization parameter so that the quantization module 23 and the inverse quantization module 27 can perform correlation calculations. Here, in the process of calculating the quantization parameter, the code rate control module can obtain image information of the original block, such as the input information, and perform the calculation, or can obtain a reconstructed value reconstructed by the adder 201 and perform the calculation, but this application is not limited thereto.
[0095] Prediction module 21 provides a predictive block to adder 202 to generate a residual block, and can provide the predictive block to adder 201 to be reconstructed to obtain a reconstructed block, which is subsequently used to predict reference pixels. Here, video encoder 102 subtracts pixel values of the predictive block from pixel values of the original block to form pixel difference values, which are residual blocks, and the data in the residual blocks can include luma and chroma differences. Adder 201 represents one or more components that perform this subtraction. Prediction module 21 can also send associated syntax elements to entropy coding module 24 for incorporation into the codestream.
[0096] The transform module 22 may divide and transform the residual block into one or more TUs. The transform module 22 may transform the residual block from the pixel domain to a transform domain (e.g., the frequency domain). For example, the residual block is transformed into transform coefficients using a discrete cosine transform (DCT) or a discrete sine transform (DST). The transform module 22 may transmit the resulting transform coefficients to the quantization module 23.
[0097] The quantization module 23 can perform quantization based on quantization units, which may be similar to the CUs, TUs, and PUs described above, or may be further divided in a division module. The quantization module 23 quantizes the transform coefficients to further reduce the coding rate and obtain quantized coefficients. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be changed by adjusting a quantization parameter. In some possible embodiments, the quantization module 23 can then perform a scan of a matrix containing the quantized transform coefficients. The entropy coding module 24 can perform the scan instead.
[0098] After quantization, entropy coding module 24 may entropy code the quantized coefficients. For example, entropy coding module 24 may perform context-adaptive variable-length coding (CAVLC), context-based adaptive binary arithmetic coding (CABAC), syntax-based binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) decoding, or another entropy coding method or technique. Entropy coding may be performed by entropy coding module 24 to obtain substreams, which may be entropy coded by multiple entropy coding modules 24 to obtain multiple substreams, which may be substream interleaved by coded substream buffer 25 and substream interleaving module 26 to obtain codestreams, which may be transmitted to video decoder 112 or an archive for later transmission or retrieval by video decoder 112.
[0099] Here, the sub-stream interleaving process can be referred to the description of FIGS. 1 to 5 above, and will not be further described here.
[0100] The inverse quantization module 27 and the inverse transform module 28 apply inverse quantization and inverse transform, respectively, and the adder 201 adds the inverse transformed residual block and the predicted residual block to generate a reconstructed block, which is subsequently used as a reference pixel for predicting the original block, and which is stored in the reference image memory 29.
[0101] 8 is a schematic diagram of the configuration of a video decoder 112 provided by an embodiment of the present application. As shown in FIG. 8, the video decoder 112 includes a sub-stream interleaving module 30, a decoded sub-stream buffer 31, an entropy decoding module 32, a prediction module 33, a dequantization module 34, an inverse transform module 35, an adder 301, and a reference image memory 36.
[0102] Here, the entropy decoding module 32 includes an analysis module and a code rate control module. In some possible embodiments, the video decoder 112 may perform a decoding flow that is illustratively the reverse of the encoding flow for the video encoder 102 as shown in FIG.
[0103] During the decoding process, the video decoder 112 receives the coded video codestream from the video encoder 102. The codestream is then de-substream interleaved by a substream interleaving module 30 to obtain multiple substreams, which then pass through corresponding decoded substream buffers to corresponding entropy decoding modules 32. Taking one substream as an example, a parsing module in the entropy decoding module 32 of the video decoder 112 entropy decodes the substream to generate quantized coefficients and syntax elements. The entropy decoding module 32 then forwards the syntax elements to a prediction module 33. The video decoder 112 can receive syntax elements at the video slice level and / or the video block level.
[0104] The code rate control module in the entropy decoding module 32 generates a quantization parameter based on the information of the image to be decoded obtained by the analysis module, so that the inverse quantization module 34 performs related calculations. The code rate control module can also calculate the quantization parameter based on the reconstructed block reconstructed by the adder 301.
[0105] The inverse quantization module 34 inversely quantizes (e.g., dequantizes) the quantized coefficients provided to the substream and decoded by the entropy decoding module 32 and the generated quantization parameters. The inverse quantization process may include determining the degree of quantization using the quantization parameters calculated for each video block in the video slice using the video encoder 102, and similarly determining the degree of inverse quantization applied. The inverse transform module 35 applies an inverse transform (e.g., a transform method such as a DCT or DST) to the inverse quantized transform coefficients, generating residual blocks inversely transformed from the inverse quantized transform coefficients into the pixel domain by an inverse transform unit. Here, the size of the inverse transform unit is the same as the size of the TU, and the inverse transform method and the transform method may utilize corresponding forward and inverse transforms of the same transform method; for example, the inverse transform of a DCT or DST may be an inverse DCT, an inverse DST, or a conceptually similar inverse transform process.
[0106] If prediction module 33 generates a predicted block, video decoder 112 sums the predicted block with an inverse transformed residual block from inverse transform module 35 to form a decoded video block. Summer 301 represents one or more components that perform this summation operation. Optionally, a deblocking filter may be applied to filter the decoded block to remove block effect artifacts. Decoded image blocks within a frame or image are stored in reference image memory 36 as reference pixels for subsequent prediction.
[0107] The embodiment of the present application provides a possible method for video (image) encoding / decoding. Referring to Figure 9, Figure 9 is a schematic flowchart of the video encoding / decoding provided by the embodiment of the present application. The possible method for video encoding / decoding includes Process 1 to Process 5, which can be performed by any one or more of the source device 10, the video encoder 102, the target device 11, or the video decoder 112.
[0108] Process 1: Divide a frame image into one or more non-overlapping parallel coding units. These parallel coding units have no dependencies between them and can be coded and decoded completely in parallel / independently, e.g., parallel coding unit 1 and parallel coding unit 2.
[0109] Process 2: Each parallel coding unit can be further divided into one or more independent coding units that do not overlap each other, and the independent coding units do not need to depend on each other, but header information of some parallel coding units can be shared.
[0110] The independent coding unit may include three channels: a luma channel Y, a first chromaticity channel Cb, and a second chromaticity channel Cr; or three RGB channels, or only one of these channels. When the independent coding unit includes three channels, the sizes of these three channels may be identical or different, depending on the input format of the image. The independent coding unit may also be understood as one or more processing units formed by the N channels included in each parallel coding unit. For example, the three channels Y, Cb, and Cr are the three channels that make up this parallel coding unit, and each may be an independent coding unit, or Cb and Cr may be collectively referred to as chroma channels. In this case, the parallel coding unit includes an independent coding unit consisting of a luma channel and an independent coding unit consisting of chroma channels.
[0111] Process 3: Each independent coding unit can be further divided into one or multiple non-overlapping coding units, and each coding unit within the independent coding unit can be dependent on each other, for example, multiple coding units can be pre-coded or pre-decoded by referring to each other.
[0112] If the coding unit is the same size as the independent coding unit (i.e., the independent coding unit is split into only one coding unit), the size may be any size described in process 2.
[0113] A coding unit may have three channels (RGB) including luminance Y, first chromaticity Cb, and second chromaticity Cr, or it may have only one of these channels. If it has three channels, the sizes of some of the channels may be the same or different, depending on the input format of the image.
[0114] Note that process 3 is an optional step in the video encoding / decoding method, and the video encoder / decoder can encode / decode the residual coefficients (or residual values) of the independent coding units obtained in process 2.
[0115] Process 4: The coding unit can be further divided into one or more non-overlapping prediction groups (PG), where PG can also be abbreviated as group. Each PG is coded using a selected prediction mode to obtain a predicted value for the PG, which is then used to construct a predicted value for the entire coding unit. A residual value for the coding unit is then obtained based on the predicted value and the original value of the coding unit.
[0116] Process 5: Based on the residual values of the coding units, the coding units are divided into groups to obtain one or more non-overlapping residual blocks (RBs), and the residual coefficients of each RB are coded in a selected mode to form a residual coefficient stream. Specifically, the residual coefficients can be divided into those that are transformed and those that are not.
[0117] Here, the selection mode of the residual coefficient encoding / decoding method in process 5 can include, but is not limited to, any one of semi-fixed length encoding, exponential Columbus (Golomb) encoding, Golomb-Rice encoding, truncated unary encoding, run-length encoding, and direct encoding of original residual values.
[0118] For example, the video encoder may directly encode the coefficients within the RBs.
[0119] For example, the video encoder may also transform the residual block and then encode the transformed coefficients, where the transform may be DCT, DST, Hadamard transform, etc.
[0120] For example, if the RB is small, the video encoder can directly quantize each coefficient in the RB and then binarize it. If the RB is large, the video encoder can further divide the RB into multiple coefficient groups (CG), and each CG can be unified quantized and then binarized. In some embodiments of the present application, the coefficient group (CG) and the quantization group (QG) may be the same.
[0121] The following describes coding of residual coefficients using a semi-fixed-length coding method. First, the maximum value of the residual absolute value within one RB block is defined as the modified maximum (mm). Next, the number of coding bits for the residual coefficients within the RB block is determined (the number of coding bits for residual coefficients within the same RB block is the same). For example, if the critical limit (CL) of the current RB block is 2 and the current residual coefficient is 1, coding the residual coefficient 1 requires 2 bits, which is represented as 01. If the CL of the current RB block is 7, this represents coding an 8-bit residual coefficient and a 1-bit code bit. Determining CL involves finding the smallest M value such that all residuals of the current sub-block are within the range of [-2^(M-1), 2^(M-1)]. If the two boundary values of -2^(M-1) and 2^(M-1) exist simultaneously, M should be increased by 1, meaning that all residuals of the current RB block need to be coded using M+1 bits. If only one of the two boundary values, -2^(M-1) and 2^(M-1), exists, one trailing bit needs to be coded to determine whether this boundary value is 2^(M-1) or 2^(M-1). If neither -2^(M-1) nor 2^(M-1) exists in all residuals, this trailing bit does not need to be coded.
[0122] Also, in special cases, the video encoder may directly encode the original values of the image instead of the residual values.
[0123] The video encoder 102 and the video decoder 112 may also be implemented by other embodiments, such as a general-purpose digital processing system. Referring to FIG. 10, FIG. 10 is a schematic diagram of a coding / decoding device provided by an embodiment of the present application, which may be a part of the video encoder 102 or a part of the video decoder 112. This coding / decoding device may be applicable to both the encoding side (or encoding end) and the decoding side (or decoding end). As shown in FIG. 10, the coding / decoding device includes a processor 41 and a memory 42. The processor 41 is connected to the memory 42 (e.g., they are connected to each other via a bus 43). In one embodiment, the coding / decoding device further includes a communication interface 44, which connects the processor 41 and the memory 42 for data transmission and reception.
[0124] The processor 41 executes instructions stored in the memory 42 to implement the image encoding and decoding methods provided by the embodiments of the present application described below. The processor 41 may be a central processing unit (CPU), a general-purpose processor, a network processor (NP), a digital signal processing (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 41 may also be any other device having processing capabilities, such as a circuit, a device, or a software module, and the embodiments of the present application are not limited thereto. In one example, the processor 41 may include one or more CPUs, such as CPU0 and CPU1 in FIG. 10. As an alternative implementation, the electronic device may include multiple processors, such as processor 45 (illustrated by dashed lines in FIG. 10) in addition to processor 41.
[0125] The memory 42 is used to store instructions. For example, the instructions may be a computer program. In one embodiment, the memory 42 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and / or instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and / or instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disk storage device, optical disk storage device (including compressed disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), disk storage medium, or other magnetic storage device, although the embodiments of the present application are not limited thereto.
[0126] The memory 42 may exist independently of the processor 41, or may be integrated with the processor 41. The memory 42 may be located inside the encoding / decoding device or outside the encoding / decoding device, and the embodiment of the present application is not limited thereto.
[0127] The bus 43 is for transmitting information between the components included in the encoding / decoding device. The bus 43 may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus 43 can be divided into an address bus, a data bus, a control bus, and the like. For convenience of illustration, the bus is shown by only one thick line in FIG. 10, but this does not mean that there is only one bus or only one type of bus.
[0128] The communication interface 44 is for communicating with other devices or other communication networks. These other communication networks may be Ethernet, radio access networks (RAN), wireless local area networks (WLAN), etc. The communication interface 44 may be a module, a circuit, a transceiver, or any device capable of communication. The embodiment of the present application is not limited in this respect.
[0129] Note that the configuration shown in Figure 10 is not intended to limit the encoding / decoding device, and in addition to the configuration shown in Figure 10, the encoding / decoding device may include more or fewer components than shown, or a combination of several components, or different components.
[0130] The image encoding method and the image decoding method provided by the embodiments of the present application may be executed by the encoding device, an application (APP) that provides an encoding function installed in the encoding device, a CPU in the encoding / decoding device, or a functional block for executing the image encoding method and the image decoding method in the encoding / decoding device. The embodiments of the present application are not limited thereto. For convenience of explanation, the following description will be made on the encoding side or the decoding side.
[0131] Hereinafter, an image encoding method and an image decoding method provided by an embodiment of the present application will be described with reference to the drawings.
[0132] As explained in the background art and above in FIGS. 1 to 5, the embedding speeds of sub-streams in different sub-stream buffers are different, and in the same time, the number of sub-stream bits embedded in a sub-stream buffer with a fast embedding speed is greater than that in a sub-stream buffer with a slow embedding speed. Therefore, in order to ensure the integrity of the embedded data, all sub-stream buffers need to be set large, which increases the hardware cost.
[0133] In order to rationally allocate the sub-stream buffer, the embodiments of the present application propose a series of improved coding methods (e.g., coding mode / prediction mode, complexity information transmission, coefficient grouping, code stream allocation, etc.) to reduce the expansion rate of the coding unit, thereby reducing the speed difference when the coding blocks of each channel of the coding unit are coded, rationally allocating the space of the sub-stream buffer, and reducing hardware costs.
[0134] Here, the expansion rate of a coding unit can include a theoretical expansion rate and a current (CU) expansion rate (actual expansion rate). The theoretical expansion rate is determined by encoding / decoding and then theoretically derived, and its value is greater than 1. If a coding unit has a channel in which the number of bits of a substream obtained by encoding a coding block is 0, this case must be excluded when calculating the current expansion rate. Theoretical expansion rate = number of bits of the CB with the largest theoretical number of bits in the CU / number of bits of the CB with the smallest theoretical number of bits in the CU. Current (CU) expansion rate = number of bits of the CB with the largest actual number of bits in the current CU / number of bits of the CB with the smallest actual number of bits in the current CU.
[0135] In one embodiment, the associated expansion ratio may also include a current sub-stream expansion ratio: current sub-stream expansion ratio = number of bits of data in the sub-stream buffer with the largest number of bits of data in the current plurality of sub-stream buffers / number of bits of data in the sub-stream buffer with the smallest number of bits of data in the current plurality of sub-stream buffers.
[0136] When the encoding side selects the encoding mode (or prediction mode) for the entire encoding block or encoding unit of each channel, the selection can be made based on several policies such as the following: Each improved encoding mode in the following embodiments can be selected and determined based on the following policies.
[0137] Policy 1: Calculate the bit consumption cost and select the encoding mode with the smallest bit consumption cost.
[0138] Here, the bit consumption cost refers to the number of bits required to encode / decode a CU, and mainly includes the code length of the mode flag (or mode codeword), the code length of the encoding / decoding tool information codeword, and the code length of the residual codeword.
[0139] Policy 2: Calculate the encoded distortion and select the encoding mode with the smallest distortion.
[0140] Here, distortion refers to the difference between the reconstructed value and the original value. The distortion can be calculated using one or more of the following: sum of squared difference (SSD), mean squared error (MSE), sum of absolute difference (time domain) (SAD), sum of absolute transformed difference (SATD), and peak signal to noise ratio (PSNR). The embodiments of the present application are not limited thereto.
[0141] Policy 3: Calculate the rate-distortion cost and select the coding mode with the smallest rate-distortion cost.
[0142] Here, the rate-distortion cost refers to a weighted sum of bit consumption cost and distortion. When performing weighting calculation, the weighting coefficients of bit consumption cost and coded distortion can both be preset in the coding side. The embodiments of the present application do not limit the specific values of the weighting coefficients.
[0143] The improved coding modes in the image coding method and image decoding method provided by the embodiments of the present application will be described below using a series of embodiments.
[0144] 1. Change the anti-inflation mode. Optional implementation: The encoding side encodes the image bit width as a fixed-length code for each pixel value of the coding block of multiple channels of the coding unit.
[0145] Improvement: The encoding side encodes each pixel value of the encoding blocks of the multiple channels of the encoding unit with a fixed-length code that is equal to or less than the image bit width.
[0146] Example 1: Illustratively, Fig. 11 is a schematic flowchart of an image encoding method provided by an embodiment of the present application. As shown in Fig. 11, the image encoding method includes S101 to S102.
[0147] S101: The encoding side obtains an encoding unit.
[0148] Here, the encoding side may be the source device 10 in FIG. 6, the video encoder 102 in the source device 10, or the encoding device in FIG. 10. The embodiments of the present application are not limited thereto. The encoding unit is an image block in the image to be processed (i.e., the original block). The encoding unit includes encoding blocks of multiple channels, where the multiple channels include a first channel, and the first channel is any one of the multiple channels. For example, the size of the encoding unit may be 16×2×3, in which case the size of the encoding block of the first channel of the encoding unit is 16×2.
[0149] S102: The encoding side encodes the encoding block of the first channel in a first encoding mode.
[0150] Here, the first coding mode is a mode in which the sample values of the coding block of the first channel are coded based on a first fixed-length code. The code length of the first fixed-length code is equal to or less than the image bit width of the image to be processed. The image bit width indicates the number of bits required to store each sample in the image to be processed. The first fixed-length code may be preset on the coding side / decoding side, or may be determined by the coding side, written in header information of the codestream, and transmitted to the decoding side.
[0151] It should be noted that the expansion prevention mode is a mode in which original pixel values in coding blocks of multiple channels are directly coded. Since the number of bits consumed for directly coding original pixel values is typically larger than in other coding modes, it should be understood that the coding block with the largest number of theoretical bits in a theoretical expansion rate typically comes from a coding block coded in the expansion prevention mode. In an embodiment of the present application, the code length of the fixed-length code in the expansion prevention mode is reduced to reduce the number of theoretical bits (i.e., the numerator of the fractional formula) of the coding block with the largest number of theoretical bits in the theoretical expansion rate, thereby reducing the theoretical expansion rate. When the theoretical expansion rate is low, the difference in the speed at which the coding blocks of each channel code the sub-streams is small. Therefore, when setting up the sub-stream buffers, the sub-stream buffers do not need to be set large, and space in the sub-stream buffers is not wasted, thereby achieving a rational allocation of the sub-stream buffers.
[0152] In one embodiment, if the image to be processed is an RGB image, the encoding side can convert it to a YUV format and encode it, or if the image to be processed is a YUV image, the encoding side can convert it to an RGB format and encode it, although the embodiment of the present application is not limited thereto.
[0153] In response to the above, an embodiment of the present application further provides an image decoding method. Figure 12 is a schematic flowchart of the image decoding method provided by the embodiment of the present application. As shown in Figure 12, the image decoding method includes S201 to S202.
[0154] S201: The decoding side obtains a codestream obtained by encoding a coding unit.
[0155] Here, the codestream obtained by encoding the coding unit may include a plurality of substreams in which the coding blocks of the plurality of channels are encoded, and which correspond one-to-one to the plurality of channels. For example, the plurality of substreams may include a substream corresponding to the first channel (i.e., a substream in which the coding blocks of the first channel are encoded).
[0156] S202: The decoding side decodes the sub-stream corresponding to the first channel in a first decoding mode.
[0157] Here, the first decoding mode is a mode in which sample values are analyzed from a substream corresponding to a first channel based on a first fixed-length code.
[0158] In one embodiment, step S202 specifically includes: if the first fixed-length code is equal to the image bit width, the decoding side directly decodes the sub-stream corresponding to the first channel according to the first decoding mode; and if the first fixed-length code is smaller than the image bit width, the decoding side dequantizes the pixel values of the analyzed coded block of the first channel.
[0159] Here, the quantization step size is 1<<(bitdepth-fixed_length), where bitdepth represents the image bit width, fixed_length represents the length of the first fixed-length code, and 1<< represents a 1-bit shift to the left.
[0160] 2. Change the fallback mode. Option 1: If the current code stream buffer is not sufficient for all channels of the encoding unit to use anti-expansion mode, the anti-expansion mode will be turned off and the encoding side will select a mode other than the original value. In this case, encoding using other coding modes will cause expansion (the number of bits in the sub-stream in which the coding block of a channel is coded will be too large), and the anti-expansion mode cannot be used.
[0161] Improvement suggestion 1: Ensure the anti-inflation mode is open.
[0162] Example 2: 13 is a schematic flowchart of another image decoding method provided by an embodiment of the present application. As shown in FIG. 13, the image decoding method includes S301 to S303.
[0163] S301: The encoding side obtains an encoding unit.
[0164] Here, the coding unit is an image block in the image to be processed, and includes coding blocks of multiple channels.
[0165] S302: The encoding side determines the first total code length.
[0166] Here, the first total code length is the total code length of the first stream obtained by encoding all of the encoding blocks of the multiple channels according to the corresponding target encoding modes, where the target encoding modes include a first encoding mode, the first encoding mode being a mode for encoding sample values in the encoding blocks with a first fixed-length code, the code length of the first fixed-length code being equal to or less than the image bit width of the target image, and the image bit width representing the number of bits required to store each sample in the target image.
[0167] For example, the encoding side can determine the rate-distortion cost of each encoding mode for each of a plurality of channels according to Policy 3 above, and determine the mode with the lowest rate-distortion cost as the target encoding mode corresponding to the encoding block of the channel.
[0168] S303: If the first total code length is equal to or greater than the remaining capacity of the codestream buffer, the encoding side encodes the encoding blocks of the multiple channels in a fallback mode.
[0169] Here, the fallback modes include a first fallback mode and a second fallback mode. The first fallback mode obtains the block vector of the reference prediction block in IBC mode, and then calculates and quantizes the residual, where the quantized step size is determined based on the remaining amount of the codestream buffer and the target pixel depth (bite per pixel, BPP). The second fallback mode directly quantizes pixel points, where the quantized step size is determined based on the remaining amount of the codestream buffer and the target BPP. The mode flags of the fallback modes are the same as those of the first coding mode.
[0170] It should be understood that when encoding using another encoding mode, the encoded residual may be too large, resulting in an excessively large number of bits for the encoding block using the other encoding mode. However, the anti-expansion mode encodes using a fixed-length code, and the total encoded code length is fixed. Therefore, the anti-expansion mode can be used to avoid the above-mentioned excessively large residual. In an embodiment of the present application, when the mode flag of the fallback mode is the same as that of the first encoding mode (anti-expansion mode), the decoding side is notified of the used encoding mode by determining the remaining capacity of the code stream buffer. This ensures that the anti-expansion mode is always selectable, thereby reducing the theoretical expansion rate. Because the theoretical expansion rate is low and the difference in the speed at which the encoding blocks of each channel encode the sub-streams is small, the sub-stream buffer does not need to be set large and space in the sub-stream buffer does not need to be wasted. This realizes a rational allocation of the sub-stream buffer.
[0171] In one embodiment, the image decoding method further includes encoding mode flags in a plurality of substreams obtained by encoding the coding blocks of the plurality of channels on the encoding side.
[0172] Here, the mode flag indicates the coding mode used for each of the coding blocks of the multiple channels, and the mode flag of the first coding mode is the same as the fallback mode.
[0173] In one embodiment, the mode flags of the coding blocks of the multiple channels may be coded in the coded substreams. For example, the multiple channels may include a first channel, and the first channel may be any one of the multiple channels. In this case, the coding of the mode flags in the multiple substreams obtained by coding the coding blocks of the multiple channels by the coding side may include coding a submode flag in the substream obtained by coding the coding block of the first channel.
[0174] Here, the sub-mode flag indicates the type of fallback mode used for the coding block of the first channel, or indicates the type of fallback mode used for the coding block of multiple channels. As mentioned above, the fallback mode can include a first fallback mode and a second fallback mode. This will not be further described here.
[0175] In one embodiment, again taking the above-mentioned first component as an example, in this case, the encoding side encodes a mode flag in a substream obtained by encoding coding blocks of multiple channels, including encoding a first flag, a second flag, and a third flag in a substream obtained by encoding a coding block of a luminance channel.
[0176] Here, the first flag is for indicating that the coding blocks of the multiple channels are coded using a first coding mode or a fallback mode, the second flag is for indicating that the coding blocks of the multiple channels are coded using a target mode, and the target mode is one of the first coding mode and the fallback mode, and when the second flag indicates that the target mode used for the coding blocks of the multiple channels is the fallback mode, the third flag is used to indicate the type of fallback mode used for the coding blocks of the multiple channels.
[0177] In response to the above, the embodiment of the present application provides two image decoding methods. Figure 14 is a schematic flowchart of another image decoding method provided by the embodiment of the present application. As shown in Figure 14, the image decoding method includes S401 to S404.
[0178] S401: The decoding side analyzes the code stream obtained by encoding the coding unit.
[0179] Here, the coding unit includes coding blocks of a plurality of channels, where the plurality of channels includes a first channel, and the first channel is any one of the plurality of channels.
[0180] S402: The decoding side analyzes a mode flag from a substream obtained by encoding coding blocks of multiple channels, and if the second total code length is greater than the remaining capacity of the codestream buffer, determines that the target decoding mode of the substream is the fallback mode.
[0181] Here, the second total code length is the total code length of a codestream obtained by encoding all of the encoding blocks of a plurality of channels according to the first encoding mode or the fallback mode.
[0182] S403: The decoding side analyzes the preset flag bit in the substream and determines the target fallback mode.
[0183] Here, the target fallback mode is a type of fallback mode. The submode flag indicates the type of fallback mode used when coding blocks of multiple channels. The fallback modes include a first fallback mode and a second fallback mode.
[0184] S404: The decoding side decodes the substream in the target fallback mode.
[0185] 15 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. As shown in FIG. 15, the image decoding method includes S501 to S505.
[0186] S501: The decoding side analyzes the code stream obtained by encoding the coding unit.
[0187] S502: The decoding side analyzes the first flag from the substream obtained by encoding the encoding block of the first channel.
[0188] Here, the first channel is any one of the multiple channels. The first flag can refer to that described in the encoding method above, and will not be further described here.
[0189] S503: The decoding side analyzes the second flag from the substream obtained by encoding the encoding block of the first channel.
[0190] Here, the second flag can refer to that described in the encoding method above, and will not be further described here.
[0191] S504: If the second flag indicates that the target mode used for the coding blocks of the multiple channels is a fallback mode, the decoding side analyzes the third flag from the substream obtained by coding the coding blocks of the first channel.
[0192] Here, the third flag can refer to that described in the encoding method above, and will not be further described here.
[0193] S505: The decoding side determines a target decoding mode for the multiple channels based on the type of fallback mode indicated by the third flag, and decodes the substreams obtained by encoding the encoding blocks of the multiple channels in the target decoding mode.
[0194] For example, the decoding side can set the type of fallback mode indicated by the third flag as the target decoding mode for multiple channels.
[0195] In one embodiment, the method further includes: the decoding side determining a target decoding code length of the coding unit based on the remaining amount of the code stream buffer and the target pixel depth BPP, where the target decoding code length indicates the code length required to decode the code stream of the coding unit; the decoding side determining assigned code lengths of the multiple channels based on the decoding code length, where the assigned code length indicates the code length required to decode the residual of the code stream of the coding block of the multiple channels; and the decoding side determining the decoding code length assigned to each of the multiple channels based on an average value of the assigned code lengths in the multiple channels.
[0196] Based on the above Example 2, taking the example where the multiple channels include a luminance (Y) channel, a first chrominance (U) channel, and a second chrominance (V) channel, we will explain two proposals mainly included in Improvement Scheme 1 of the fallback mode.
[0197] Suggestion 1: Encoding side: The encoding side controls that the bit cost of encoding in expansion prevention mode (original value mode) is always the highest, and the encoding side cannot select a mode with a bit cost greater than the original value. Therefore, all components require a coding mode, and even in fallback mode, not only the substream corresponding to the Y channel (first substream) but also the substreams corresponding to the U / V channels (second and third substreams) require a coding mode. The codewords in fallback mode are kept the same as those in original value mode. The specific types of fallback modes (first fallback mode and second fallback mode) can be coded based on a certain channel.
[0198] Decoding side: First, analyze the coding modes of the three channels. If all three channels (Y, U, and V) are in the target coding mode selected based on the rate-distortion cost, determine whether the remaining space in the codestream buffer allows the three channels to be decoded in their respective target coding modes. If not, the current decoding mode is the fallback mode. In the fallback mode, analyze one flag bit based on a certain component to indicate whether the current fallback mode is the first fallback mode or the second fallback mode (all three channels use the same fallback mode). Therefore, if there is a channel on the decoding side that does not select the expansion prevention mode, the current CU will not select the fallback mode.
[0199] Suggestion 2: Encoding side: At any time, the encoding side control determines that the bit cost of encoding in the expansion prevention mode (original value mode) is the highest, and the encoding side cannot select a mode with a bit cost greater than the original value mode. The original value mode and the fallback mode still use the same mode flag (i.e., the first flag described above). However, when encoding the expansion prevention mode and the fallback mode, one additional flag (i.e., the second flag described above) is encoded to indicate that the current encoding mode is either the expansion prevention mode or the fallback mode. If this additional flag indicates that the current encoding mode is the fallback mode, one more flag (i.e., the third flag described above) is encoded to indicate that the current fallback mode is either the first fallback mode or the second fallback mode. If multiple channels all select the fallback mode, the types of fallback modes for the multiple channels remain consistent. If a channel selects the expansion prevention mode, a mode flag is encoded in the substream corresponding to that channel, but a flag indicating whether it is in the fallback mode (i.e., the second flag described above) does not need to be encoded.
[0200] Decoding side: The mode flag is parsed from the codestream. If the mode flag is a common mode flag for the anti-inflation mode and the fallback mode, one flag (i.e., the second flag above) is parsed to indicate whether the current fallback mode is the fallback mode. If the current fallback mode is the fallback mode, one more flag (i.e., the third flag above) is parsed to indicate whether the current fallback mode is the first fallback mode or the second fallback mode. Once the fallback mode type of one channel is parsed, other channels using fallback modes use the same fallback mode type as that channel. Channels not using fallback modes parse their own modes.
[0201] As mentioned above, Improved Option 1 can reduce the upper limit of the numerator in the expansion rate equation by using Example 2 above. In one embodiment, in the fallback mode, the lower limit of the denominator in the expansion rate equation can also be increased. Below, we will explain using Option 2 and Improved Option 2 as alternatives.
[0202] Option 2: If it is determined that the fallback mode is used for encoding, the encoding side provides one target BPP. In the fallback mode, the total number of bits of the substreams corresponding to the coding blocks of the three channels must be equal to or less than the target BPP. Therefore, when allocating the code rate, the common information (block vector and mode flag) is encoded in the luma channel, and after subtracting the code length occupied by the common information from the allocated code length, the remaining code length must be allocated because it is used to encode the residuals of the multiple channels. The allocation method is to divide the remaining code length by the number of pixels of the multiple channels (i.e., the code length allocated to each channel must be an integer multiple of the number of pixels in the coding unit). If the remaining code length is not divisible, the remaining code length is allocated to the luma channel. If the remaining code length is small, it cannot be divided by the number of pixels of the multiple channels, and code lengths cannot be allocated to the substreams corresponding to the first and second chroma channels. As a result, the number of bits of the corresponding substreams is small and the expansion rate is large.
[0203] Improvement 2: Allocate the remaining code length evenly to multiple channels according to the number of bits, improving the accuracy of allocation.
[0204] Example 3: 16 is a schematic flowchart of another image encoding method provided by an embodiment of the present application. As shown in FIG. 16, in addition to the above S301 to S303, the image encoding method further includes S601 to S603.
[0205] S601: The encoding side determines a target encoding code length of an encoding unit based on the remaining amount of the codestream buffer and the target BPP.
[0206] Here, the target encoding code length indicates the code length required to encode the encoding unit. The target BPP can be obtained by the encoding side, for example, by receiving a target BPP input by a user.
[0207] S602: Determine the allocated code lengths of the multiple channels based on the target encoding code length.
[0208] Here, the assigned code length indicates the code length required to encode the residuals of the coding blocks of multiple channels.
[0209] For example, as described above, the encoding side subtracts the code length occupied by the common information from the assigned code length to obtain the assigned code length for the multiple channels.
[0210] S603: Determine the code length assigned to each of the multiple channels based on the average value of the assigned code lengths for the multiple channels.
[0211] In one embodiment, the encoding side may also encode the block vectors and mode flags in substreams corresponding to multiple channels.
[0212] Here, the mode flags can be referred to above and will not be further explained here, and the block vectors can be referred to the description of the IBC mode below and will not be further explained here.
[0213] In one embodiment, when encoding the residual, if the allocated bits of a channel cannot divide the number of pixels in this channel evenly, i.e., the residual for each pixel cannot be coded using a fixed-length code, the coding side can divide the pixels in this channel into groups, and each group can code only one residual value.
[0214] It should be understood that when the remaining code length is small, the code length allocation mode using the number of pixels in the current coding unit as a unit may result in no code length being allocated to the chroma channel, which may result in a smaller number of bits coded for the coding block of the chroma channel and a larger expansion rate.The image coding method provided by the embodiment of the present application can allocate the allocation unit by converting the number of pixels into the number of bits of the allocated code length, and each channel can be coded by allocating a code length, thereby reducing the theoretical expansion rate.
[0215] In response to the above, an embodiment of the present application further provides an image decoding method, which adds the above S401 to S404 or S501 to S505, and further includes: determining a target decoding code length of a coding unit based on a remaining capacity of a code stream buffer and a target pixel depth BPP, where the target decoding code length indicates a code length required to decode the code stream of the coding unit; determining assigned code lengths of multiple channels based on the decoding code length, where the assigned code length indicates a code length required to decode residuals of the code streams of the coding blocks of the multiple channels; and determining the decoding code length assigned to each of the multiple channels based on an average value of the assigned code lengths in the multiple channels.
[0216] 3. Change the intra block copy (IBC) mode. Optional implementation: The mode identifier is at the CU level, and the BV (Block Vector) is also at the CU level, and the mode identifier and BV are transmitted in the luminance channel.
[0217] Improvement: The mode identifier is changed to the CB level, the BV is still at the CU level, and each channel carries the mode identifier and each channel carries the BV.
[0218] Example 4: 17 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. As shown in FIG. 17, the method includes S701 to S704.
[0219] S701: The encoding side obtains an encoding unit.
[0220] Here, the coding unit includes coding blocks of multiple channels.
[0221] S702: The encoding side encodes a coding block of at least one channel among the multiple channels in IBC mode.
[0222] In one embodiment, S702 specifically includes the encoding side determining a target encoding mode for the at least one channel using a rate-distortion optimization policy.
[0223] Here, the target coding mode includes IBC mode. The rate-distortion optimization policy can refer to the description of Policy 3 above, and will not be further described here.
[0224] S703: The encoding side obtains the BV of the reference prediction block.
[0225] Here, the BV of the reference prediction block indicates the position of the reference prediction block in the coded image block, and the reference prediction block represents the predicted value of the coding block coded in IBC mode.
[0226] S704: The encoding side encodes the BV of the reference prediction block in at least one substream obtained by encoding the coding block of at least one channel in IBC mode.
[0227] In one embodiment, for the BVs of the reference prediction blocks of the coding block of each channel, the BVs may all be coded in the sub-stream corresponding to the luminance channel.
[0228] In one embodiment, for the BVs of the reference prediction blocks of the coding block of each channel, all of the BVs can be coded in the substream corresponding to a chrominance channel (e.g., the first chrominance channel or the second chrominance channel), or the number of the BVs can be equally divided and coded in the substreams corresponding to the two chrominance channels. If the number of BVs cannot be equally divided, the coding side can code another BV in the substream corresponding to one of the chrominance channels.
[0229] In one embodiment, the BVs of the reference prediction blocks of the coding blocks of each channel may be equally divided and coded in substreams corresponding to the coding blocks of different channels. If equal division is not possible, the coding side may allocate them at a preset ratio. For example, assuming that there are eight BVs, the preset ratio may be Y channel:U channel:V channel=2:3:3 or Y channel:U channel:V channel=4:2:2. The embodiment of the present application is not limited to a specific value of the preset ratio. In this case, S602 specifically includes coding the coding blocks of the multiple channels in IBC mode. S704 specifically includes coding the BVs of the reference prediction blocks in the substreams obtained by coding the coding blocks of each of the multiple channels in IBC mode.
[0230] In one embodiment, the BV of the reference prediction block includes a plurality of BVs, and encoding the BVs of the reference prediction block in the substreams obtained by encoding the coding blocks of each of the plurality of channels in IBC mode includes encoding the BVs of the reference prediction block in the codestreams obtained by encoding the coding blocks of each of the plurality of channels in IBC mode based on a preset ratio.
[0231] It should be understood that the number of bits of the sub-stream corresponding to the luma channel is usually large, and the number of bits of the sub-stream corresponding to the chroma channel is usually small. In the embodiment of the present application, when encoding using the IBC mode, the BV of the reference prediction block is transmitted in the sub-stream corresponding to the encoding block of at least one channel, so that the transmission data can be increased in the sub-stream corresponding to the chroma channel with a small number of bits, and thus the theoretical expansion factor can be reduced by increasing the denominator in the formula for calculating the theoretical expansion factor.
[0232] In one embodiment, a CU generates header information during encoding, and the header information of the CU can be allocated and encoded in the substream using the BV allocation method described above.
[0233] In response to the above, an embodiment of the present application further provides an image decoding method. Figure 18 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. As shown in Figure 18, this method includes S801 to S804.
[0234] S801: The decoding side analyzes the codestream obtained by encoding the coding unit.
[0235] Here, the coding unit includes coding blocks of multiple channels, and the codestream includes multiple substreams in one-to-one correspondence with the multiple channels, into which the coding blocks of the multiple channels are coded.
[0236] S802: The decoding side determines the position of the reference prediction block in the multiple sub-streams based on the block vector BV of the reference prediction block analyzed from at least one sub-stream among the multiple sub-streams.
[0237] S803: The decoding side determines a predicted value of a decoded block to be decoded in IBC mode based on the position information of the reference prediction block.
[0238] S804: The decoding side reconstructs the decoded block to be decoded in IBC mode based on the predicted value.
[0239] S803 and S804 can refer to the description of the video encoding / decoding system above, and will not be further described here.
[0240] In one embodiment, the encoding side can analyze the BVs of the reference predicted blocks of the coding blocks of each channel in sub-streams that all correspond to the luminance channel.
[0241] In one embodiment, for the BVs of the reference prediction blocks of the coding blocks of each channel, all the BVs can be analyzed in the substream corresponding to a chrominance channel (e.g., the first chrominance channel or the second chrominance channel), or the number of BVs can be equally divided and analyzed in the substreams corresponding to the two chrominance channels. If the number of BVs cannot be equally divided, the encoding side can analyze another BV in the substream corresponding to one of the chrominance channels.
[0242] In one embodiment, the BVs of the reference prediction blocks of the coding blocks of each channel can be equally divided and coded in substreams corresponding to the coding blocks of different channels. If equal division is not possible, the decoding side can allocate them at a preset ratio. For example, assuming that the number of BVs is eight, the preset ratio can be Y channel:U channel:V channel=2:3:3 or Y channel:U channel:V channel=4:2:2. The embodiment of the present application does not limit the specific value of the preset ratio.
[0243] In one embodiment, the coding blocks coded in IBC mode include coding blocks of at least two channels, and the coding blocks of the at least two channels share a BV of a reference prediction block.
[0244] In one embodiment, the method further includes, when an IBC mode identifier is parsed from any one of the plurality of substreams, the decoding side determines that the target decoding mode corresponding to the plurality of substreams is IBC mode.
[0245] In one embodiment, the method further includes the decoding side analyzing the IBC mode identifiers one by one from the multiple substreams, and determining that the target decoding mode corresponding to the substream of the analyzed IBC mode identifier from the multiple substreams is IBC mode.
[0246] Based on the above Example 4, taking the example where the multiple channels include a luminance (Y) channel, a first chrominance (U) channel, and a second chrominance (V) channel, we will explain two proposals mainly included in the improvement proposal 1 of the fallback mode.
[0247] Suggestion 1: Encoding side: Step 1: Obtain a reference prediction block based on the multi-channels. That is, for one multi-channel coding unit, the multi-channel at a certain position in the search area is taken as the reference prediction block of the current multi-channel coding unit, and this position is denoted as BV. That is, the multi-channels share one BV. If the input image is YUV400, it has only one channel, but otherwise it has three channels. After obtaining the prediction block, calculate the residual for each channel, quantize the residual (transform), and then perform inverse quantization (inverse transform) to complete the reconstruction.
[0248] Step 2: In the luminance channel, the side information (including information such as the complexity level of the coding block) and mode information are coded, and the quantized coefficients of the luminance channel are coded.
[0249] Step 3: If a chroma channel exists, encode the auxiliary information (including information such as coding block complexity) in the chroma channel and encode the quantized coefficients of the chroma channel.
[0250] 3.1: The BVs of the reference prediction blocks of each coding unit can all be coded in the luminance channel.
[0251] 3.2: For the BV of the reference prediction block of each coding unit, all of it can be coded in one chrominance channel, or it can be equally divided and coded in two chrominance channels. If it is not possible to equally divide it, a part of the BV is coded in one chrominance channel.
[0252] 3.3: For the BV of the reference prediction block of each coding unit, the BV can be equally divided and coded into different components. If it cannot be equally divided, it will be allocated according to a preset ratio.
[0253] Decryptor: Step 1: Analyze the side information and mode information of the coding block of the luminance channel, and analyze the quantized coefficients in the luminance channel.
[0254] Step 2: If a chrominance channel exists, analyze the auxiliary information of the coding block of the U channel. If the prediction mode of the luma channel is IBC mode, there is no need to analyze the prediction mode of the current chrominance channel, and directly set it to IBC mode, and analyze the quantized coefficients of the current chrominance channel.
[0255] Step 3: The analysis of BV is similar on the encoding side, i.e., 3.1: The BVs of the reference prediction blocks of each coding unit can all be analyzed in the luminance channel.
[0256] 3.2: For the BV of the reference prediction block of each analysis unit, all BVs can be analyzed in one chrominance channel, or the BVs can be equally divided and analyzed in two chrominance channels. If they cannot be equally divided, a part of the BV is coded in one chrominance channel.
[0257] 3.3: For the BV of the reference prediction block of each analysis unit, the BV can be equally divided and coded into different components. If it cannot be equally divided, it is allocated according to a preset ratio.
[0258] Step 4: Based on the BV shared by the three channels, obtain the predicted value of each coding block in each channel, and perform inverse quantization (inverse transformation) on the coefficients obtained by analyzing each channel to obtain the residual value, and complete the reconstruction of each coding block based on the residual value and the predicted value.
[0259] Suggestion 2: Encoding side: Step 1: For each IBC mode in the same coding unit, i.e., three channels, only one group of BVs needs to be trained, which is obtained based on the search area and original values of the three channels, or based on the search area and original values of one of the channels, or based on the search area and original values of any two of the channels.
[0260] Step 2: For the coding blocks of each channel, determine the target coding mode using the rate-distortion cost, where the BVs of the three channels in IBC mode are calculated using the BVs obtained in step 1. After a component selects the IBC mode, other components may not select the IBC mode.
[0261] Step 3: For the coding block of each channel, its own optimal mode needs to be coded. If the target mode of one or more channels selects IBC mode, the BVs can be coded based on one channel that selects IBC mode, or the number of BVs can be equally divided and coded based on two channels that select IBC mode, or the number of BVs can be equally divided and coded based on all channels that select IBC mode. If the number of BVs cannot be divided, they can be allocated at a preset ratio.
[0262] Decryptor: Each channel analyzes one target mode, and when it is analyzed that the target mode of one or more channels selects IBC mode (only the same IBC mode can be selected), the BV analysis method is the same as the allocation method used by the encoding side to encode the BV.
[0263] 4. Grouping of coefficients. Optional implementation: As shown in the flowchart of Figure 1, there may be a residual skip mode in the encoding process. If the encoding side selects this skip mode during encoding, there is no need to encode the residual during encoding, and only one bit of data is required to represent the residual skip.
[0264] Improvement: Divide the processing coefficients (residual coefficients and / or transform coefficients) into groups.
[0265] Example 5: 19 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. As shown in FIG. 19, the image encoding method includes S901 to S902.
[0266] S901: The encoding side obtains processing coefficients corresponding to an encoding unit.
[0267] Here, the processing coefficients include one or more of residual coefficients and transform coefficients.
[0268] S902: The encoding side divides the processing coefficients into a plurality of groups based on a number threshold.
[0269] Here, the number threshold can be preset on the encoding side. The number threshold is related to the size of the encoding block. For example, for a 16×2 encoding block, the number threshold can be set to 16. Among the multiple groups of processing coefficients, the number of processing coefficients in each group is equal to or less than the number threshold.
[0270] Compared with the residual skip mode in the alternative implementation, the image encoding method provided by the embodiment of the present application can divide the processing coefficients into groups during encoding, and the processing coefficients of each group need to add header information for the processing coefficients of the group to indicate the details of the processing coefficients of the group during transmission. Compared with the current residual skip being represented by 1-bit data, this increases the denominator in the theoretical expansion rate calculation formula and reduces the theoretical expansion rate.
[0271] In response to the above, an embodiment of the present application further provides an image decoding method. Figure 20 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. As shown in Figure 20, the image decoding method includes S1001 to S1002.
[0272] S1001: The decoding side analyzes the code stream obtained by encoding the encoding unit, and determines the processing coefficients corresponding to the encoding unit.
[0273] Here, the processing coefficients include one or more of residual coefficients and transform coefficients, and the processing coefficients include a plurality of groups, and the number of processing coefficients in each group is equal to or less than a number threshold.
[0274] S1002: The decoding side decodes the code stream based on the processing coefficients.
[0275] S1002 can refer to the description of the video encoding / decoding system above, and will not be further described here.
[0276] 5. Complexity transmission. Optional implementation: the encoding unit includes a coding block for a luma channel, a coding block for a first chroma channel, and a coding block for a second chroma channel. The substream corresponding to the coding block for the luma channel is a first substream, the substream corresponding to the coding block for the first chroma channel is a second substream, and the substream corresponding to the coding block for the second chroma channel is a third substream. The first substream transmits a complexity level for the luma channel using 1 or 3 bits, the second substream transmits an average value of the two chroma channels using 1 or 3 bits, and the third substream does not transmit a complexity level.
[0277] For example, the calculation method of the CU level complexity is as follows: JPEG2025531374000003.jpg51163After searching BiasInit from the table based on the complexity level ComplexityLevel[0] of the luma channel and the complexity level ComplexityLevel[1] of the chroma channel, the quantization parameter Qp[0] of the luma channel and the quantization parameters Qp[1] and Qp[2] of the two chroma channels were calculated.
[0278] Here, the searched table can be referred to as Table 1 below, and will not be further described here.
[0279] Taking the first sub-stream as an example, the specific implementation of the first sub-stream is as follows: JPEG2025531374000004.jpg6787
[0280] Here, complexity_level_flag[0] is the complexity level update flag for the luma channel and is a binary variable. A value of "1" indicates that the luma channel of the coding unit needs to have its complexity level updated, and a value of "0" indicates that the luma channel of the coding unit does not need to have its complexity level updated. The value of ComplexityLevelFlag[0] is equal to the value of complexity_level_flag[0]. delta_level[0] is the complexity level change amount for the luma channel and is a 2-bit unsigned integer. It determines the luma complexity level change amount. The value of DeltaLevel[0] is equal to the value of delta_level[0]. If delta_level[0] does not exist in the codestream, the value of DeltaLevel[0] is equal to 0. PrevComplexityLevel represents the luma channel complexity level of the previous coding unit, and ComplexityLevel[0] represents the luma channel complexity level.
[0281] Improvement: Transmit complexity information in the third substream.
[0282] Example 6 21 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. As shown in FIG. 21, the image encoding method includes S1101 to S1103.
[0283] S1101: The encoding side obtains an encoding unit.
[0284] Here, the coding unit includes coding blocks of P channels, where P is an integer equal to or greater than 2.
[0285] S1102: The encoding side obtains complexity information of the encoding block in each of the P channels.
[0286] Here, the complexity information is used to represent the degree of difference in pixel values of the coding blocks of each channel. For example, assuming that the P channels include a luma channel, a first chroma channel, and a second chroma channel, the complexity information of the coding blocks of the luma channel is used to represent the degree of difference in pixel values of the coding blocks of the luma channel, the complexity information of the coding blocks of the first chroma channel is used to represent the degree of difference in pixel values of the coding blocks of the first chroma channel, and the complexity information of the coding blocks of the second chroma channel is used to represent the degree of difference in pixel values of the coding blocks of the second chroma channel.
[0287] S1103: The encoding side encodes complexity information of the encoding blocks of each channel in the substreams obtained by encoding the encoding blocks of the P channels.
[0288] In one embodiment, S1103 specifically includes: the encoding side encoding the complexity level of the coding blocks of each channel in the substreams obtained by encoding the coding blocks of the P channels.
[0289] Illustratively, taking the second sub-stream corresponding to the first chrominance component as an example, the second sub-stream is specifically realized as follows: JPEG2025531374000005.jpg7294
[0290] Here, complexity_level_flag[1] indicates the complexity level update flag for the first chroma channel and is a binary variable. A value of "1" indicates that the complexity level of the coding block of the first chroma channel and the complexity level of the coding block of the luma channel of the coding unit match, while a value of "0" indicates that the complexity level of the coding block of the first chroma channel and the complexity level of the coding block of the luma channel of the coding unit do not match. The value of ComplexityLevelFlag[1] is equal to the value of complexity_level_flag[1]. delta_level[1] is a 2-bit unsigned integer indicating the amount of change in the complexity level of the first chroma channel. It determines the amount of change in the complexity level of the coding block of the first chroma channel. The value of DeltaLevel[1] is equal to the value of delta_level[1]. If delta_level[1] is not present in the codestream, the value of DeltaLevel[1] is equal to 0. ComplexityLevel[1] indicates the complexity level of the coding block of the first chroma channel.
[0291] Illustratively, taking the third sub-stream corresponding to the second chrominance component as an example, the third sub-stream is specifically realized as follows: JPEG2025531374000006.jpg79105
[0292] Here, complexity_level_flag[2] indicates the complexity level update flag for the second chroma channel and is a binary variable. A value of "1" indicates that the complexity level of the coding block of the second chroma channel of the coding unit matches the complexity level of the coding block of the first chroma channel, and a value of "0" indicates that the complexity level of the coding block of the second chroma channel of the coding unit does not match the complexity level of the coding block of the first chroma channel. The value of ComplexityLevelFlag[2] is equal to the value of complexity_level_flag[2]. delta_level[2] is a 2-bit unsigned integer that indicates the amount of change in the complexity level of the second chroma channel. It determines the amount of change in the complexity level of the coding block of the second chroma channel. The value of DeltaLevel[2] is equal to the value of delta_level[2]. If delta_level[2] is not present in the codestream, the value of DeltaLevel[2] is equal to 0. ComplexityLevel[2] indicates the complexity level of the coding block for the second chrominance channel.
[0293] In one embodiment, the complexity information includes a complexity level and a first reference coefficient, and the first reference coefficient is used to represent a ratio relationship between the complexity levels of the coding blocks of different channels. S1103 specifically includes: an encoding side encoding complexity levels of the coding blocks of Q channels in each of the substreams obtained by encoding the coding blocks of Q channels among the P channels, where Q is an integer less than P; and an encoding side encoding a first reference coefficient in the substreams obtained by encoding the coding blocks of PQ channels among the P channels.
[0294] For example, assuming that the P channels include a luminance channel, a first chrominance channel, and a second chrominance channel, the complexity information of the coding block of the luminance channel is the complexity level of the coding block of the luminance channel, the complexity information of the coding block of the first chrominance channel is the complexity level of the coding block of the first chrominance channel, and the complexity information of the coding block of the second chrominance channel is a reference coefficient, which is used to represent the ratio relationship between the complexity level of the coding block of the first chrominance channel and the complexity level of the coding block of the second chrominance channel.
[0295] For example, the meaning of the above variable complexity_level_flag[2] is changed so that a value of "1" indicates that the complexity level of the coding block of the second chrominance channel of the coding unit is greater than that of the coding block of the first chrominance channel, and a value of "0" indicates that the complexity level of the coding block of the second chrominance channel of the coding unit is equal to or less than that of the coding block of the first chrominance channel. The value of ComplexityLevelFlag[2] is equal to the value of complexity_level_flag[2].
[0296] In one embodiment, the complexity information includes a complexity level, a reference complexity level, and a second reference coefficient. The reference complexity level includes one of a first complexity level, a second complexity level, and a third complexity level. The first complexity level is the maximum complexity level of the coding blocks of PQ channels out of P channels, where Q is an integer less than P. The second complexity level is the minimum complexity level of the coding blocks of PQ channels out of P channels. The third complexity level is the average complexity level of the coding blocks of PQ channels out of P channels. The second reference coefficient is used to represent a relationship and / or a ratio relationship between the complexity levels of the coding blocks of PQ channels out of P channels. In this case, the above S1103 specifically includes the encoding side encoding the complexity levels of the coding blocks of Q channels in each of the substreams obtained by encoding the coding blocks of Q channels out of the P channels, and the encoding side encoding the reference complexity levels and the second reference coefficients in the substreams obtained by encoding the coding blocks of PQ channels out of the P channels.
[0297] For example, assuming that the P channels include a luma channel, a first chroma channel, and a second chroma channel, the complexity information of the coding block of the luma channel is the complexity level of the coding block of the luma channel, and the complexity information of the coding block of the first chroma channel is the reference complexity level. The reference complexity level includes one of a first complexity level, a second complexity level, and a third complexity level. The first complexity level is the maximum value among the complexity level of the coding block of the first chroma channel and the complexity level of the coding block of the second chroma channel. The second complexity level is the minimum value among the complexity level of the coding block of the first chroma channel and the complexity level of the coding block of the second chroma channel. The third complexity level is the average value among the complexity level of the coding block of the first chroma channel and the complexity level of the coding block of the second chroma channel. The complexity information of the coding block of the second chrominance channel is a reference coefficient, and the reference coefficient is used to represent the relationship and / or ratio relationship between the complexity level of the coding block of the first chrominance channel and the complexity level of the coding block of the second chrominance channel.
[0298] It should be understood that in an alternative embodiment, complexity information is not transmitted in the third sub-stream corresponding to the coding block of the second chrominance channel. The image coding method provided by the embodiment of the present application adds coding of complexity information in the third sub-stream, thereby increasing the number of bits in the sub-stream with a smaller number of bits, thereby increasing the denominator in the above theoretical expansion factor calculation formula, and reducing the theoretical expansion factor.
[0299] In response to the above, an embodiment of the present application provides an image decoding method. Figure 2 is a diagram showing a further image decoding method provided by an embodiment of the present application. As shown in Figure 2, the decoding method includes S1201 to S1204.
[0300] S1201: The decoding side analyzes the code stream obtained by encoding the coding unit.
[0301] Here, a coding unit includes coding blocks for P channels, where P is an integer equal to or greater than 2. A codestream includes multiple substreams, each substream corresponding to the P channels, into which the coding blocks for the P channels are coded.
[0302] S1202: The decoding side analyzes complexity information of the coded blocks of each channel in the substreams obtained by coding the coded blocks of the P channels.
[0303] In one embodiment, S1202 specifically includes the decoding side analyzing the complexity level of the coding blocks of each channel in each of the sub-streams obtained by coding the coding blocks of the P channels.
[0304] S1203: The decoding side determines the quantization parameters of the coded blocks of each channel based on the complexity information of the coded blocks of each channel.
[0305] S1204: The decoding side decodes the codestream based on the quantization parameters of the coded blocks of each channel.
[0306] In one embodiment, when the complexity level of each channel is transmitted in any of the substreams corresponding to each channel, the complexity level of the CU level (CuComplexityLevel) can be calculated according to the following process. JPEG2025531374000007.jpg43170
[0307] Here, the definition of ComplexityDivide3Table is ComplexityDivide3Table={0,0,0,1,1,1,2,2,2,3,3,3,4}; image_format represents the image format of the image to be processed in which the coding unit exists.
[0308] In one embodiment, as described above, the complexity information includes a complexity level and a first reference coefficient, and the first reference coefficient is used to represent a ratio relationship between the complexity levels of the coding blocks of different channels. In this case, S1202 specifically includes the following steps: a decoding side analyzes the complexity levels of the coding blocks of Q channels in each of the substreams obtained by encoding the coding blocks of Q channels among the P channels, where Q is an integer less than P; a decoding side analyzes the first reference coefficient in the substreams obtained by encoding the coding blocks of PQ channels among the P channels; and a decoding side determines the complexity levels of the coding blocks of PQ channels based on the first reference coefficient and the complexity levels of the coding blocks of the Q channels.
[0309] In one embodiment, if the complexity level is not transmitted in the substream (i.e., if the magnitude relationship is transmitted in the third substream), the complexity level at the CU level (CuComplexityLevel) can be calculated according to the following process: JPEG2025531374000008.jpg49163
[0310] Here, ChromaComplexityLevel represents the complexity level of the chrominance channel.
[0311] ChromaComplexityLevel must be calculated separately on the encoding side (it is obtained directly from the codestream on the decoding side).
[0312] ChromaComplexityLevel=(ComplexityLevel[1]+ComplexityLevel[2])>>1, or ChromaComplexityLevel=max(ComplexityLevel[1], ComplexityLevel[2]), or ChromaComplexityLevel=min(ComplexityLevel[1], ComplexityLevel[2]), or ChromaComplexityLevel=ComplexityLevel[1], or ChromaComplexityLevel=ComplexityLevel[2].
[0313] In one embodiment, the complexity information includes a complexity level, a reference complexity level, and a second reference coefficient. The reference complexity level includes one of a first complexity level, a second complexity level, and a third complexity level. The first complexity level is the maximum complexity level of the coding blocks of PQ channels out of P channels, where Q is an integer less than P. The second complexity level is the minimum complexity level of the coding blocks of PQ channels out of P channels. The third complexity level is the average complexity level of the coding blocks of PQ channels out of P channels. The second reference coefficient is used to represent a relationship and / or a ratio relationship between the complexity levels of the coding blocks of PQ channels out of P channels. In this case, the above S1202 specifically includes: the decoding side analyzing the complexity levels of the coding blocks of the Q channels in each of the substreams obtained by encoding the coding blocks of the Q channels out of the P channels; the decoding side encoding a reference complexity level and a second reference coefficient in the substreams obtained by encoding the coding blocks of the PQ channels out of the P channels; and the decoding side determining the complexity levels of the coding blocks of the PQ channels based on the complexity levels of the coding blocks of the Q channels and the second reference coefficient.
[0314] In one embodiment, when the complexity level of each channel is transmitted in the substream corresponding to each channel, the decoding side searches for BiasInit1 in Table 2 below based on the complexity level ComplexityLevel[0] of the coding block of the luma channel and the complexity level ComplexityLevel[1] of the coding block of the first chroma channel, and searches for BiasInit2 in Table 2 below based on the complexity level ComplexityLevel[0] of the coding block of the luma channel and the complexity level ComplexityLevel[2] of the coding block of the second chroma channel to calculate the quantization parameter Qp[0] of the luma channel, the quantization parameter Qp[1] of the first chroma component, and the quantization parameter Qp[2] of the second chroma channel.
[0315] For example, the decoding side can calculate the quantization parameter according to the following process. JPEG2025531374000009.jpg46110
[0316] [Table 1]
[0317] [Table 2]
[0318] In one embodiment, the decoding side searches for BiasInit from Table 2 based on the complexity level ComplexityLevel[0] of the coding block of the chrominance channel and the complexity level ChromaComplexityLevel of the coding block of the chrominance channel, and calculates Qp[0], Qp[1], and Qp[2] according to the following process. JPEG2025531374000012.jpg49155
[0319] In one embodiment, as described above, a first complexity level and a second complexity level can be transmitted in the substream corresponding to the coding block of the first chrominance channel. In this case, the decoding side searches for BiasInit from Table 2 based on the complexity level ComplexityLevel[0] of the coding block of the luminance channel and the complexity level ChromaComplexityLevel of the coding block of the chrominance channel (taking the first complexity level / complexity level of the coding block of the first chrominance channel), and calculates Qp[0], Qp[1], and Qp[2] according to the following process: JPEG2025531374000013.jpg72117
[0320] In one embodiment, as described above, a third complexity level can be transmitted in the substream corresponding to the coding block of the first chrominance channel, in which case the ChromaComplexityLevel is obtained by taking the third complexity level and calculating Qp[0], Qp[1], and Qp[2] according to the following process: JPEG2025531374000014.jpg72107
[0321] 6. Chroma channels share a sub-stream. Optional implementation: for a processing target image in YUV420 / YUV422 format, a total of three substreams, including a first substream, a second substream, and a third substream, are transmitted, where the first substream includes syntax elements and transform coefficients / residual coefficients of a luma channel, the second substream includes syntax elements and transform coefficients / residual coefficients of a first chroma channel, and the third substream includes syntax elements and transform coefficients / residual coefficients of a second chroma channel.
[0322] Improvement: For images to be processed in YUV420 / YUV422 format, the syntax elements and transform coefficients / residual coefficients of the first chrominance channel and the syntax elements and transform coefficients / residual coefficients of the second chrominance channel are transmitted to the second sub-stream, and the third sub-stream is cancelled.
[0323] Example 7: 22 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. As shown in FIG. 22, the image encoding method includes S1201 to S1202.
[0324] S1201: The encoding side obtains an encoding unit.
[0325] Here, the coding unit is an image block in the image to be processed, and includes coding blocks of multiple channels.
[0326] S1202: If the image format of the image to be processed is a preset format, the encoding side merges substreams obtained by encoding the encoding blocks of at least two preset channels among the multiple channels into one merged substream.
[0327] Here, the preset format may be set in advance on the encoding side / decoding side, for example, the preset format may be YUV420 or YUV422, etc., and the preset channels may be the first chrominance channel and the second chrominance channel. Exemplarily, the coding block of the first sub-stream is defined as follows: JPEG2025531374000015.jpg94170
[0328] Illustratively, the coding blocks of the second sub-stream are defined as follows: JPEG2025531374000016.jpg168170
[0329] It should be understood that the number of coded bits for the syntax elements and quantized transform coefficients of the two chrominance channels is typically smaller than that of the luma channel, and that by fusing the syntax elements of the two chrominance channels into the second substream for transmission, the difference in the number of bits between the first substream and the second substream can be reduced, thereby reducing the theoretical expansion rate.
[0330] In response to the above, an embodiment of the present application further provides an image decoding method, and Figure 23 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. As shown in Figure 23, this method includes S1301 to S1303.
[0331] S1301: The decoding side analyzes the code stream obtained by encoding the coding unit.
[0332] S1302: The decoding side determines the merged substream.
[0333] Here, when the image format of the image to be processed is a preset format, the merged substream is obtained by merging substreams obtained by encoding the encoding blocks of at least two preset channels among the multiple channels.
[0334] S1303: The decoding side decodes the codestream based on the merge target substream into which the at least two substreams are merged.
[0335] 7. Substream embedding. Improvement: Embed preset codewords in substreams whose bit count is less than a bit count threshold.
[0336] Example 8: 24 is a schematic flowchart of yet another image encoding method provided by an embodiment of the present application. As shown in FIG. 24, the method includes S1401 to S1402.
[0337] S1401: The encoding side obtains an encoding unit.
[0338] Here, the coding unit includes coding blocks of multiple channels.
[0339] S1402: The encoding side encodes preset codewords in the target substream that satisfy the preset condition until the target substream no longer satisfies the preset condition.
[0340] Here, the target substream is a substream among the multiple substreams. The multiple substreams are codestreams obtained by encoding coding blocks of multiple channels. The preset codeword may be "0" or another codeword, etc. However, the embodiment of the present application is not limited thereto.
[0341] In one embodiment, the preset condition includes the number of bits of the substream being less than a preset first bit number threshold.
[0342] Here, the first bit number threshold may be preset in the encoding side / decoding side or may be transmitted in the stream by the encoding side / decoding side, although the embodiment of the present application is not limited thereto. The first bit number threshold is used to indicate the minimum number of bits of CB allowed in a CU.
[0343] In one embodiment, the preset condition includes that the codestream of the coding unit contains coded coding blocks whose number of bits is less than a second preset bit number threshold.
[0344] Here, the second bit number threshold may be preset in the encoding side / decoding side or may be transmitted in the stream by the encoding side / decoding side, although the embodiment of the present application is not limited thereto. The second bit number threshold is used to indicate the minimum stream bit number allowed among the multiple substreams.
[0345] It should be understood that embedding preset codewords in substreams with a smaller number of bits directly increases the denominator of the expansion factor calculation formula, thereby reducing the actual expansion factor of the coding unit.
[0346] In response to the above, an embodiment of the present application further provides an image decoding method. Figure 25 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. As shown in Figure 25, this image decoding method includes S1501 to S1503.
[0347] S1501: The decoding side analyzes the code stream obtained by encoding the coding unit.
[0348] Here, the coding unit includes coding blocks of multiple channels.
[0349] S1502: The decoding side determines the number of code words.
[0350] Here, the number of codewords is used to indicate the number of preset codewords encoded in a target substream that satisfies a preset condition. The preset codewords are encoded into a target substream if a target substream that satisfies the preset condition exists.
[0351] S1503: The decoding side decodes the codestream based on the number of codewords.
[0352] For example, the decoding side can delete preset codewords based on the number of codewords and decode the codestream from which the preset codewords have been deleted.
[0353] 8. Pre-set the expansion rate. Improvement: By presetting the expansion rate as a threshold, the current actual expansion rate is controlled.
[0354] Example 9: FIG. 26 is a schematic flowchart of yet another image coding method provided by an embodiment of the present application. As shown in FIG. 26, this image coding method includes S1601 to S1603. S1601: The encoding side obtains an encoding unit.
[0355] Here, the coding unit includes coding blocks of multiple channels.
[0356] S1602: The encoding side determines a target encoding mode corresponding to each of the encoding blocks of each channel among the encoding blocks of the multiple channels based on a preset expansion factor.
[0357] Here, the preset expansion factor may be preset in the encoding side and the decoding side, or the preset expansion factor may be encoded into a substream by the encoding side and transmitted to the decoding side, although the embodiment of the present application is not limited thereto.
[0358] S1603: The encoding side encodes the encoding block of each channel in the target encoding mode so that the current expansion rate is equal to or less than the preset expansion rate.
[0359] In one embodiment, the preset expansion ratios include a first preset expansion ratio, and the current expansion ratio is equal to the quotient of the number of bits of the largest substream and the number of bits of the smallest substream. The largest substream is the substream with the largest number of bits among the multiple substreams obtained by encoding the coding blocks of the multiple channels. The smallest substream is the substream with the smallest number of bits among the multiple substreams obtained by encoding the coding blocks of the multiple channels.
[0360] In one embodiment, as described above, the encoding side can preset a first preset expansion rate as the maximum threshold allowed in the encoding unit. Before encoding the encoding unit, the encoding side can obtain the states of all current substreams, obtain the states of the substream with the most transmitted fixed-length code streams and the substream with the largest total remaining amount in the current substream, and obtain the states of the substream with the least transmitted fixed-length code streams and the substream with the smallest total remaining amount in the current substream. If the ratio or difference between the maximum state and the minimum state is greater than the preset difference threshold, the encoding side can not use rate-distortion optimization as a target mode selection criterion, and can select a low coding mode for the substream with the maximum state and a high coding rate mode for the substream with the minimum state.
[0361] For example, a preset expansion rate is set as B_th, B_delta is set as an auxiliary threshold, and coding blocks of three channels Y / U / V are coded to obtain three sub-streams. When coding one coding unit, the coding side can obtain the states of the three sub-streams. If the coding rate corresponding to the first sub-stream (the sub-stream corresponding to the Y channel) is set to the maximum bit_stream1 and the coding rate corresponding to the second sub-stream (the sub-stream corresponding to the U channel) is set to the minimum bit_stream2, when bit_stream1 / bit_stream2>B_th-B_delta, the coding side can select a mode with a higher coding rate for the coding block of the Y channel and a mode with a lower coding rate for the coding block of the U channel.
[0362] In one embodiment, the preset expansion rate includes a second preset expansion rate, and the current expansion rate is equal to a quotient of the number of bits of a coding block with the largest number of bits and the number of bits of a coding block with the smallest number of bits among the coding blocks of the multiple channels that have been coded.
[0363] In one embodiment, as described above, the encoding side pre-sets a second preset expansion rate as the maximum threshold value allowed within the encoding unit, and can obtain the code rate cost in each mode before encoding a certain encoding block. When the encoding mode corresponding to the optimal code rate cost obtained based on the rate-distortion cost satisfies the condition that the actual expansion rate for encoding each sub-stream is less than the preset expansion rate, the encoding mode corresponding to the optimal code rate cost is set as the target encoding mode for the encoding block, and the target encoding mode is incorporated into the sub-stream corresponding to the encoding block. When the encoding mode corresponding to the optimal code rate cost cannot satisfy the condition that the actual expansion rate for encoding each sub-stream is less than the preset expansion rate, the encoding side changes the mode with the largest code rate to a mode with a code rate smaller than the code rate of the mode with the largest code rate, or changes the mode with the smallest code rate to a mode with a code rate larger than the code rate of the mode with the smallest code rate, or changes the mode with the largest code rate to a mode with a code rate smaller than the code rate of the mode with the largest code rate and, at the same time, changes the mode with the smallest code rate to a mode with a code rate larger than the code rate of the mode with the smallest code rate.
[0364] Exemplarily, taking the case where the preset expansion rate is A_th as an example, encoding blocks of three channels of Y / U / V are encoded to obtain three sub-streams. Assume that the optimal code rate costs of the three channels of Y / U / V are rate_y, rate_u, and rate_v respectively, and the code rate of rate_y is the largest and the code rate of rate_u is the smallest. When rate_y / rate_u ≧ A_th, the encoding side can change the target encoding mode of the Y channel so that its code rate cost becomes smaller than rate_y, or change the target encoding mode of the U channel so that its code rate cost becomes larger than rate_u, so as to satisfy rate_y / rate_u < A_th and rate_y / rate_v < A_th.
[0365] The image encoding method provided by the embodiment of the present application involves a target encoding mode selected for the encoding block of each channel according to a preset expansion rate, so that when the encoding side encodes each channel in the target encoding mode, the actual expansion rate of the encoding unit is less than the preset expansion rate, and therefore the actual expansion rate can be reduced.
[0366] In one embodiment, the image encoding method further includes: the encoding side determining a current expansion rate; and if the current expansion rate is greater than the first preset expansion rate, encoding a preset codeword in the minimum substream such that the current expansion rate is less than or equal to the first preset expansion rate. For example, the encoding side embeds the preset codeword at the end of the minimum substream.
[0367] In one embodiment, the image encoding method includes: an encoding side determining a current expansion rate; and, if the current expansion rate is greater than a second preset expansion rate, the encoding side encoding a preset codeword in the encoding block with the smallest number of bits so that the current expansion rate is less than or equal to the second preset expansion rate.
[0368] In response to the above, an embodiment of the present application further provides an image decoding method. Figure 27 is a schematic flowchart of yet another image decoding method provided by an embodiment of the present application. As shown in Figure 27, this method includes S1701 to S1703.
[0369] S1701: The decoding side analyzes the code stream obtained by encoding the coding unit.
[0370] S1702: The decoding side determines the number of preset codewords based on the codestream.
[0371] S1703: The decoding side decodes the codestream based on the number of preset codewords.
[0372] For S1701 to S1703, the explanation of S1501 to S1503 above can be referred to, and no further explanation will be given here.
[0373] It should be noted that the image encoding method and image decoding method provided by the embodiments of the present application have been described above as a series of independent embodiments. In actual use, the series of embodiments and selectable implementations of the embodiments may be combined with each other. The embodiments of the present application do not limit the specific combinations.
[0374] The above mainly describes the technical means provided by the embodiments of the present application based on methods. To realize the above functions, corresponding hardware structures and / or software modules are included to perform each function. With reference to the means and algorithm steps of each example described in the embodiments disclosed herein, it should be readily understood that the present application can be realized by hardware or a combination of hardware and computer software. Whether a function is implemented by hardware or by computer software operating on hardware depends on the specific application and design constraints of the technical means. While specialized technical objectives may use different methods to realize the described functions for each specific application, such realization should not be considered beyond the scope of the present application.
[0375] In an exemplary embodiment, the present embodiment provides an image encoding device, and any one of the image encoding methods is performed by the image encoding device. The image encoding device provided by the present embodiment may be the source device 10 or the video encoder 102.
[0376] 28 is a schematic diagram of the configuration of an image encoding device provided by an embodiment of the present application. As shown in FIG. 28, this image encoding device includes an acquisition module 2801 and a processing module 2802.
[0377] In one embodiment, the acquisition module 2801 is used to acquire a coding unit, where the coding unit is an image block in a target image, the coding unit includes coding blocks of multiple channels, the multiple channels including a first channel, and the first channel is any one of the multiple channels. The processing module 2802 is used to encode the coding block of the first channel in a first coding mode, where the first coding mode is a mode for encoding sample values in the coding block of the first channel with a first fixed-length code, where the code length of the first fixed-length code is less than or equal to an image bit width of the target image, and the image bit width represents the number of bits required to store each sample in the target image.
[0378] In one embodiment, the acquisition module 2801 is further configured to acquire a codestream obtained by encoding a coding unit, the coding unit being an image block in a target image, the coding unit including coding blocks of multiple channels, the multiple channels including a first channel, the first channel being any one of the multiple channels, and the codestream including multiple substreams into which the coding blocks of the multiple channels are encoded, each substream corresponding to the multiple channels one-to-one. The processing module 2802 is further configured to decode the substream corresponding to the first channel in a first decoding mode, the first decoding mode being a mode for analyzing sample values from the substream corresponding to the first channel with a first fixed-length code, the code length of the first fixed-length code being less than or equal to an image bitwidth of the target image, the image bitwidth representing the number of bits required to store each sample in the target image.
[0379] In one embodiment, the acquisition module 2801 is further used to acquire coding units, where the coding units are image blocks in the target image, and the coding units include coding blocks of multiple channels. The processing module 2802 determines a first total code length, and if the first total code length is equal to or greater than the remaining capacity of the code stream buffer, uses the code units to encode the coding blocks of the multiple channels in a fallback mode, where the first total code length is the total code length of a first stream obtained by encoding all of the coding blocks of the multiple channels in corresponding target coding modes, the target coding mode includes a first coding mode, the first coding mode is a mode for encoding sample values in the coding blocks with a first fixed-length code, the code length of the first fixed-length code is equal to or less than the image bit width of the target image, the image bit width represents the number of bits required to store each sample in the target image, and the mode flag of the fallback mode is the same as that of the first coding mode.
[0380] In one embodiment, the processing module 2802 is used to encode mode flags in multiple sub-streams obtained by encoding the coding blocks of the multiple channels, and the mode flags are used to indicate the coding mode used for each of the coding blocks of the multiple channels.
[0381] In one embodiment, the plurality of channels includes a first channel, the first channel being any one of the plurality of channels, and the processing module 2802 is specifically used to encode a sub-mode flag in a sub-stream obtained by encoding the encoding block of the first channel, the sub-mode flag being for indicating the type of fallback mode used for the encoding block of the plurality of channels.
[0382] In one embodiment, the multiple channels include a first channel, the first channel being any one of the multiple channels, and the processing module 2802 is specifically used to encode a first flag, a second flag, and a third flag in a sub-stream obtained by encoding a coding block of the first channel, wherein the first flag is for indicating that the coding block of the multiple channels is to be encoded using a first coding mode or a fallback mode, the second flag is for indicating that the coding block of the multiple channels is to be encoded using a target mode which is either the first coding mode or the fallback mode, and the third flag is for indicating the type of fallback mode used for the coding block of the multiple channels.
[0383] In one embodiment, the processing module 2802 is further used to determine a target encoding code length of the encoding unit based on the remaining amount of the code stream buffer and the target pixel depth BPP, determine assigned code lengths for the multiple channels based on the target encoding code length, and determine an assigned encoding code length for each of the multiple channels based on an average value of the assigned code lengths in the multiple channels, where the target encoding code length indicates the code length required to encode the encoding unit, and the assigned code length indicates the code length required to encode the residual of the encoding block of the multiple channels.
[0384] In one embodiment, the processing module 2802 further analyzes the codestream obtained by encoding the coding unit; analyzes a mode flag from a substream obtained by encoding the coding blocks of multiple channels, and if the first total code length is greater than the remaining space of the codestream buffer, determines that the target decoding mode of the substream is a fallback mode; analyzes a preset flag bit in the substream to determine a target fallback mode, which is used to decode the substream in the target fallback mode, where the coding unit includes coding blocks of multiple channels, the mode flag is for indicating whether the coding blocks of the multiple channels are encoded using the first coding mode or the fallback mode, and the first total code length is greater than the remaining space of the codestream buffer. the target mode includes the first encoding mode, the first encoding mode being a mode for encoding sample values in the encoding block with a first fixed-length code, the code length of the first fixed-length code being equal to or less than the image bit width of the image to be processed, the image bit width being for representing the number of bits required to store each sample in the image to be processed, the target fallback mode being a type of fallback mode, the preset flag bit being for indicating the position of a sub-mode flag, the sub-mode flag being for indicating the type of fallback mode to be used when encoding the encoding blocks of multiple channels, and the fallback modes include a first fallback mode and a second fallback mode.
[0385] In one embodiment, the processing module 2802 is used to determine a target decoding code length of the coding unit based on the remaining amount of the code stream buffer and the target pixel depth BPP, determine assigned code lengths for multiple channels based on the target decoding code length, and determine the decoding code length assigned to each of the multiple channels based on the average value of the assigned code lengths in the multiple channels, where the target decoding code length indicates the code length required to decode the coding unit, and the assigned code length indicates the code length required to decode the residual of the coding block of the multiple channels.
[0386] In one embodiment, the processing module 2802 is further used to parse the codestream obtained by encoding the coding unit, parse a first flag in the substream obtained by encoding the coding block of the first channel, parse a second flag from the substream obtained by encoding the coding block of the first channel, and if the second flag indicates that the target mode used for the coding block of the multiple channels is a fallback mode, parse a third flag from the substream obtained by encoding the coding block of the first channel, determine a target decoding mode for the multiple channels based on the type of fallback mode indicated by the third flag, and decode the substream obtained by encoding the coding block of the multiple channels in the target decoding mode. the encoding unit includes coding blocks of a plurality of channels, the plurality of channels including a first channel, the first channel being any one of the plurality of channels; the first flag is for indicating that the coding blocks of the plurality of channels use a first coding mode or a fallback mode, the first coding mode being a mode for encoding sample values in the coding blocks with a first fixed-length code, the code length of the first fixed-length code being equal to or less than an image bit width of a target image to be processed, the image bit width being for indicating the number of bits required to store each sample in the target image to be processed; and the second flag is for indicating that the coding blocks of the plurality of channels use a target mode that is either the first coding mode or the fallback mode.
[0387] In one embodiment, the processing module 2802 is further used to determine a target decoding code length of the coding unit based on the remaining amount of the code stream buffer and the target pixel depth BPP, determine assigned code lengths for the multiple channels based on the target decoding code length, and determine the decoding code length assigned to each of the multiple channels based on the average value of the assigned code lengths in the multiple channels, where the target decoding code length indicates the code length required to decode the coding unit, and the assigned code length indicates the code length required to decode the residual of the coding block of the multiple channels.
[0388] In one embodiment, the obtaining module 2801 is further used to obtain a coding unit, where the coding unit includes coding blocks of multiple channels. The processing module 2802 is further used to encode the coding blocks of at least one channel of the multiple channels in an IBC mode, which is an intra block copy mode, obtain a block vector BV of a reference prediction block, and encode the BV of the reference prediction block in at least one sub-stream obtained by encoding the coding blocks of the at least one channel in the IBC mode, where the BV of the reference prediction block indicates the position of the reference prediction block in the coded image block, and the reference prediction block represents a predicted value of the coding block coded in the IBC mode.
[0389] In one embodiment, the processing module 2802 is specifically used to determine a target coding mode for the at least one channel using a rate-distortion optimization policy, where the target coding mode includes an IBC mode, and encodes a coding block of the at least one channel among the plurality of channels in the target coding mode.
[0390] In one embodiment, the processing module 2802 is specifically used to encode the coding blocks of multiple channels in IBC mode and encode the BV of the reference prediction block in the substream obtained by encoding the coding blocks of each channel of the multiple channels in IBC mode.
[0391] In one embodiment, the processing module 2802 is specifically used to encode the BV of the reference prediction block in a substream obtained by encoding the encoding block of each channel of multiple channels in IBC mode based on a preset ratio.
[0392] In one embodiment, the processing module 2802 further analyzes the codestream obtained by encoding the coding unit, and determines a position of the reference prediction block in multiple substreams based on a block vector BV of the reference prediction block analyzed from at least one substream among the multiple substreams, and determines a predicted value of a decoded block to be decoded in IBC mode based on the position information of the reference prediction block, which is used to encode the coded block to be coded in IBC mode based on the predicted value, wherein the coding unit includes coded blocks of multiple channels, and the codestream includes multiple substreams in which the coded blocks of the multiple channels are coded, corresponding one-to-one to the multiple channels, and the reference prediction block is for representing the predicted value of a decoded block to be decoded in IBC mode, which is an intra block copy mode, and the BV of the reference prediction block is for indicating the position of the reference prediction block in a reconstructed image block.
[0393] In one embodiment, the BV of the reference prediction block is coded in multiple streams, and the BV of the reference prediction block is obtained by coding the coded blocks of multiple channels in IBC mode.
[0394] In one embodiment, the obtaining module 2801 is further used for obtaining processing coefficients corresponding to the coding unit, where the processing coefficients include one or more of residual coefficients and transform coefficients, and the processing module 2802 is further used for dividing the processing coefficients into a plurality of groups based on a number threshold, where the number of processing coefficients in each group of the plurality of groups is equal to or less than the number threshold.
[0395] In one embodiment, the processing module 2802 is further used to analyze the codestream obtained by encoding the coding unit, determine processing coefficients corresponding to the coding unit, and decode the codestream based on the processing coefficients, where the processing coefficients include one or more of residual coefficients and transform coefficients, and the processing coefficients include multiple groups of processing coefficients, and the number of processing coefficients in each group is less than or equal to a number threshold.
[0396] In one embodiment, the obtaining module 2801 is used to obtain a coding unit and obtain complexity information of the coding blocks in each of P channels, where the coding unit includes coding blocks of P channels, where P is an integer greater than or equal to 2, and the complexity information is used to represent the difference in pixel values of the coding blocks of each channel. The processing module 2802 is used to encode the complexity information of the coding blocks of each channel in sub-streams obtained by encoding the coding blocks of the P channels.
[0397] In one embodiment, the complexity information includes a complexity level, and the processing module 2802 is specifically used to encode the complexity level of the coding block of each channel in the sub-streams obtained by encoding the coding blocks of the P channels.
[0398] In one embodiment, the complexity information includes a complexity level and a first reference coefficient, and the first reference coefficient is used to represent a ratio relationship between the complexity levels of the coding blocks of different channels. The processing module 2802 is specifically used in the encoding side to encode the complexity levels of the coding blocks of Q channels in each of the substreams obtained by encoding the coding blocks of Q channels out of the P channels, and to encode the first reference coefficient in the substreams obtained by encoding the coding blocks of PQ channels out of the P channels, where Q is an integer less than P.
[0399] In one embodiment, the complexity information includes a complexity level, a reference complexity level, and a second reference coefficient. The reference complexity level includes one of a first complexity level, a second complexity level, and a third complexity level. The first complexity level is the maximum complexity level of the coding blocks of PQ channels out of P channels, where Q is an integer less than P. The second complexity level is the minimum complexity level of the coding blocks of PQ channels out of P channels. The third complexity level is the average complexity level of the coding blocks of PQ channels out of P channels. The second reference coefficient is used to represent a relationship and / or a ratio relationship between the complexity levels of the coding blocks of PQ channels out of P channels. Specifically, the processing module 2802 is used on the encoding side to encode the complexity levels of the coding blocks of Q channels in each of the substreams obtained by encoding the coding blocks of Q channels out of the P channels, and to encode the reference complexity levels and second reference coefficients in the substreams obtained by encoding the coding blocks of PQ channels out of the P channels.
[0400] In one embodiment, the processing module 2802 further analyzes the code stream obtained by encoding the coding unit, analyzes complexity information of the coding blocks of each channel in the sub-stream obtained by encoding the coding blocks of P channels, determines quantization parameters of the coding blocks of each channel based on the complexity information of the coding blocks of each channel, and decodes the code stream based on the quantization parameters of the coding blocks of each channel, where the coding unit includes coding blocks of P channels, where P is an integer greater than or equal to 2, and the code stream includes multiple sub-streams into which the coding blocks of the P channels are encoded, each sub-stream corresponding to the P channels one-to-one, and the complexity information is used to represent the difference in pixel values of the coding blocks of each channel.
[0401] In one embodiment, the complexity information includes a complexity level, and the processing module 2802 is further used to analyze the complexity level of the coding blocks of each channel in the sub-streams obtained by encoding the coding blocks of the P channels.
[0402] In one embodiment, the complexity information includes a complexity level and a first reference coefficient, where the first reference coefficient is used to represent a ratio relationship between the complexity levels of the coding blocks of different channels. The processing module 2802 is specifically used to: analyze the complexity levels of the coding blocks of Q channels in each substream obtained by encoding the coding blocks of Q channels among the P channels; analyze the first reference coefficient in the substream obtained by encoding the coding blocks of PQ channels among the P channels; and determine the complexity levels of the coding blocks of PQ channels based on the first reference coefficient and the complexity levels of the coding blocks of the Q channels, where Q is an integer less than P.
[0403] In one embodiment, the complexity information includes a complexity level, a reference complexity level, and a second reference coefficient. The reference complexity level includes one of a first complexity level, a second complexity level, and a third complexity level. The first complexity level is the maximum complexity level of the coding blocks of PQ channels out of P channels, where Q is an integer less than P. The second complexity level is the minimum complexity level of the coding blocks of PQ channels out of P channels. The third complexity level is the average complexity level of the coding blocks of PQ channels out of P channels. The second reference coefficient is used to represent a relationship and / or a ratio relationship between the complexity levels of the coding blocks of PQ channels out of P channels. Specifically, the processing module 2802 is used by the encoding side to analyze the complexity levels of the coding blocks of Q channels in each of the substreams obtained by encoding the coding blocks of Q channels out of the P channels, analyze the reference complexity levels and second reference coefficients in the substreams obtained by encoding the coding blocks of PQ channels out of the P channels, and determine the complexity levels of the coding blocks of PQ channels based on the complexity levels of the coding blocks of Q channels and the second reference coefficients.
[0404] In one embodiment, the acquisition module 2801 is further used to acquire coding units, where the coding units include coding blocks of multiple channels, and the coding units are image blocks in the image to be processed. If the image format of the image to be processed is a preset format, the processing module 2802 is used to merge substreams obtained by encoding the coding blocks of at least two preset channels of the multiple channels into one merged substream.
[0405] In one embodiment, the processing module 2802 further analyzes the codestream obtained by encoding the coding unit, determines a merged substream, and decodes the codestream based on a merged substream in which at least two substreams are merged to obtain a coding unit, where the coding unit includes coding blocks of multiple channels, the coding unit is an image block in the image to be processed, and the merged substream is obtained by merging substreams obtained by encoding coding blocks of at least two preset channels among the multiple channels when the image format of the image to be processed is a preset format.
[0406] In one embodiment, the obtaining module 2801 is further used to obtain a coding unit, the coding unit including coding blocks of multiple channels, and the processing module 2802 is further used to encode preset codewords in a target substream that satisfies the preset condition until the target substream no longer satisfies the preset condition, the target substream being a substream of the multiple substreams, and the multiple substreams being codestreams obtained by encoding the coding blocks of the multiple channels.
[0407] In one embodiment, the preset condition includes the number of bits of the substream being less than a preset first bit number threshold.
[0408] In one embodiment, the predetermined condition includes that the codestream of the coding unit contains a coding block that is coded with a bit number less than a second predetermined bit number threshold.
[0409] In one embodiment, the processing module 2802 is further used to analyze the codestream obtained by encoding the encoding unit, determine the number of codewords, and decode the codestream based on the number of codewords, where the encoding unit includes coding blocks of multiple channels, and the number of codewords is for indicating the number of preset codewords encoded in the target substream that satisfies the preset conditions, and the preset codewords are to be encoded into the target substream if a target substream that satisfies the preset conditions exists.
[0410] In one embodiment, the acquisition module 2801 is further used to acquire a coding unit, where the coding unit includes coding blocks of multiple channels. The processing module 2802 is further used to determine a target coding mode corresponding to each of the coding blocks of each channel among the coding blocks of the multiple channels based on a preset expansion rate, and encode the coding block of each component in the target coding mode, so that the current expansion rate is less than or equal to the preset expansion rate.
[0411] In one embodiment, the preset expansion rate includes a first preset expansion rate, and the value of the current expansion rate is equal to the quotient of the number of bits of the maximum substream and the number of bits of the minimum substream, where the maximum substream is the substream with the largest number of bits among the multiple substreams obtained by encoding the encoding blocks of the multiple channels, and the minimum substream is the substream with the smallest number of bits among the multiple substreams obtained by encoding the encoding blocks of the multiple channels.
[0412] In one embodiment, the preset expansion rate includes a second preset expansion rate, and the current expansion rate is equal to a quotient of the number of bits of a coding block with the largest number of bits and the number of bits of a coding block with the smallest number of bits among the coding blocks of the multiple channels that have been coded.
[0413] In one embodiment, the processing module 2802 is further used to determine a current expansion rate, and if the current expansion rate is greater than the first preset expansion rate, to encode the preset codeword in the smallest substream so that the current expansion rate is less than or equal to the first preset expansion rate.
[0414] In one embodiment, the processing module 2802 is further used to determine a current expansion rate, and if the current expansion rate is greater than the second preset expansion rate, to encode the preset codeword in the encoding block with the smallest number of bits so that the current expansion rate is less than or equal to the second preset expansion rate.
[0415] In one embodiment, the processing module 2802 is further used to analyze a codestream obtained by encoding a coding unit including coding blocks of multiple channels, determine the number of preset codewords based on the codestream, and decode the codestream based on the number of preset codewords.
[0416] In Figure 28, the division of modules is only a schematic and logical function division, and in actual implementation, division may be performed in a different manner. For example, two or more functions may be integrated into one processing module. The integrated module may be implemented in the form of hardware or in the form of a software function module.
[0417] In an exemplary embodiment, an embodiment of the present application provides a readable storage medium containing executable instructions that, when executed in an image encoding / decoding device, cause the image encoding / decoding device to implement any one of the methods provided by the above embodiments.
[0418] In an exemplary embodiment, an embodiment of the present application provides a computer program product including executable instructions that, when executed in an image encoding / decoding device, cause the image encoding / decoding device to implement any one of the methods provided by the above embodiments.
[0419] In an exemplary embodiment, an embodiment of the present application provides a chip, the chip comprising a processor and an interface, the processor being coupled to a memory via the interface, and when the processor executes a computer program in the memory or the image encoding / decoding device executes an instruction, any one of the methods provided by the above embodiments is performed.
[0420] The above embodiments may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented as a software program, all or a portion thereof may be implemented in the form of a computer program product. This computer program product includes one or more computer-executable instructions. When the computer-executable instructions are loaded and executed on a computer, all or a portion of the processes or functions described in the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer-executable instructions may be stored on a computer-readable storage medium or transmitted from a computer-readable storage medium to another computer-readable storage medium. For example, the computer-executable instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wire (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium accessible by a computer, or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid state drives (SSDs)).
[0421] Although the present application has been described herein with reference to various embodiments, those skilled in the art will understand and realize other variations of the disclosed embodiments when practicing the present application, by referring to the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and "one" or "an" does not exclude a plurality. A single processor or other means may implement several functions recited in the claims. The fact that several methods are recited in mutually different dependent claims does not necessarily mean that these methods cannot be combined to produce advantageous results.
[0422] Although the present application has been described by combining specific features and embodiments, it is clear that various modifications and combinations are possible without departing from the spirit and scope of the present application. Therefore, the specification and drawings are merely exemplary descriptions defined by the scope of the appended claims, and are to be considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims and their equivalents, the present application also intends to include these modifications and variations.
[0423] The above is only a specific embodiment of the present application, and the scope of protection of the present application is not limited thereto, and any changes or replacements within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be in accordance with the scope of protection of the claims.
Claims
1. obtaining a coding unit, the coding unit including coding blocks of a plurality of channels; encoding the coding block of each channel based on the preset expansion rate so that the current expansion rate is equal to or less than the preset expansion rate; the preset expansion rate includes a first preset expansion rate, and the value of the current expansion rate is derived from a quotient of the number of bits of a maximum substream and the number of bits of a minimum substream, the maximum substream being the substream with the largest number of bits among a plurality of substreams obtained by encoding the encoding blocks of the plurality of channels, and the minimum substream being the substream with the smallest number of bits among a plurality of substreams obtained by encoding the encoding blocks of the plurality of channels; Image encoding method.
2. The image encoding method includes: Obtaining a first substream and a second substream, the first substream being a substream with the largest total amount of residuals in the most transmitted fixed-length code streams, and the second substream being a substream with the smallest total amount of residuals in the least transmitted fixed-length code streams; and determining, when a ratio value or a difference value between the first sub-stream and the second sub-stream is greater than a predetermined difference threshold, that a target coding mode of the coding block corresponding to the first sub-stream is a low code rate coding mode and that a target coding mode of the coding block corresponding to the second sub-stream is a high code rate coding mode. The image encoding method according to claim 1 .
3. The image encoding method includes: determining the current expansion rate; If the current expansion rate is greater than the first preset expansion rate, encoding preset codewords in the smallest substream so that the current expansion rate is less than or equal to the first preset expansion rate. The image encoding method according to claim 1 .
4. The preset expansion ratio further includes a second preset expansion ratio, and the image encoding method further comprises: determining the current expansion rate; If the current expansion rate is greater than a second preset expansion rate, encoding the preset codeword in the coding block with the smallest number of bits so that the current expansion rate is equal to or less than the second preset expansion rate. The image encoding method according to claim 3 .
5. encoding a preset codeword in the minimum substream, embedding the preset codeword at the end of the minimal substream.
5. The image coding method according to claim 3 or 4.
6. obtaining a coding unit, the coding unit including coding blocks of a plurality of channels; encoding a coding block of at least one channel among the plurality of channels in an IBC mode that is an intra block copy mode; Obtaining a block vector BV of a reference prediction block, the BV of the reference prediction block being for indicating a position of the reference prediction block in a coded image block, and the reference prediction block being for representing a predicted value of the coding block coded in IBC mode; encoding a BV of the reference prediction block in a substream obtained by encoding the coding blocks in each of the plurality of channels in the IBC mode; Image encoding method.
7. In the substream obtained by encoding the coding blocks in each of the plurality of channels in the IBC mode, encoding the BV of the reference prediction block includes: encoding a BV of the reference prediction block in a codestream obtained by encoding the coding blocks in each of the plurality of channels in the IBC mode based on a preset ratio; The image encoding method according to claim 6.
8. The image encoding method includes: Obtaining residual coefficients corresponding to a coding unit; further comprising: dividing the residual coefficients into a plurality of groups based on a number threshold, wherein the number of residual coefficients in each group among the plurality of groups is equal to or less than the number threshold. The image encoding method according to claim 6.
9. Obtaining a coding unit, the coding unit being an image block in a target image, the coding unit including coding blocks of multiple channels; determining a first total code length, the first total code length being a total code length of a first stream obtained by encoding all of the encoding blocks of the plurality of channels in a corresponding target encoding mode, the target encoding mode including a first encoding mode, the first encoding mode being a mode for encoding sample values in the encoding blocks with a first fixed-length code, the code length of the first fixed-length code being equal to or less than an image bit width of the image to be processed, and the image bit width being for representing the number of bits required to store each sample in the image to be processed; encoding the coding blocks of the plurality of channels in a fallback mode when the first total code length is equal to or greater than a remaining capacity of a code stream buffer, wherein a mode flag of the fallback mode is the same as that of the first coding mode. Image encoding method.
10. The image encoding method includes: further comprising encoding mode flags in a plurality of substreams obtained by encoding the coding blocks of the plurality of channels, the mode flags being for indicating a coding mode used for each of the coding blocks of the plurality of channels. The image encoding method according to claim 9.
11. the plurality of channels includes a first channel, the first channel being any one of the plurality of channels, and encoding a mode flag in a substream obtained by encoding coding blocks of the plurality of channels includes: and encoding a sub-mode flag in a sub-stream obtained by encoding the encoding block of the first channel, the sub-mode flag indicating a type of fallback mode used for the encoding block of the plurality of channels, the fallback mode including a first fallback mode and a second fallback mode. The image coding method according to claim 10.
12. the plurality of channels includes a first channel, the first channel being any one of the plurality of channels, and encoding a mode flag in a substream obtained by encoding coding blocks of the plurality of channels includes: encoding a first flag and a second flag in a substream obtained by encoding the encoding block of the first channel, the first flag being for indicating that the encoding block of the plurality of channels is to be encoded using the first encoding mode or the fallback mode; The image coding method according to claim 10.
13. The image encoding method includes: determining a target encoding code length of the encoding unit based on the remaining capacity of the code stream buffer and a target pixel depth BPP, wherein the target encoding code length indicates a code length required to encode the encoding unit; determining assigned code lengths for a plurality of channels based on the target coding code lengths, the assigned code lengths indicating code lengths required to encode residuals of coding blocks for the plurality of channels; determining an encoding code length assigned to each of the plurality of channels based on an average value of the assigned code lengths for the plurality of channels; The image coding method according to any one of claims 10 to 12.
14. determining assigned code lengths for a plurality of channels based on the target encoding code length, and subtracting a code length occupied by common information from the target encoding code length to obtain allocated code lengths for a plurality of channels. The image coding method according to claim 13.
15. analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; determining a number of codewords, the number of codewords being for indicating a number of preset codewords encoded in a target substream that satisfies a preset condition, and the preset codewords being encoded into the target substream when a target substream that satisfies the preset condition exists; and decoding the codestream based on the number of codewords. Image decoding method.
16. Decoding the codestream based on the number of codewords comprises: removing preset codewords based on the number of codewords; and decoding the codestream from which the preset codewords have been removed. The image decoding method according to claim 15.
17. analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; determining the current expansion rate; determining a number of codewords based on the current expansion factor and a first preset expansion factor, the number of codewords being for indicating a number of preset codewords encoded in a target substream that satisfies a preset condition, and the preset codewords being encoded into the target substream when a target substream that satisfies the preset condition exists; and decoding the codestream based on the number of codewords. Image decoding method.
18. analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of a plurality of channels, the codestream including a plurality of substreams in one-to-one correspondence with the plurality of channels, into which the coding blocks of the plurality of channels are encoded; determining a position of the reference prediction block in the plurality of substreams based on a block vector BV of the reference prediction block analyzed from at least one substream among the plurality of substreams, the reference prediction block representing a predicted value of a decoded block decoded in an IBC mode, which is an intra block copy mode, and the BV of the reference prediction block representing a position of the reference prediction block in a reconstructed image block; determining a predicted value of the decoded block to be decoded in IBC mode based on the position information of the reference predicted block; and reconstructing a decoded block decoded in the IBC mode based on the predicted value. Image decoding method.
19. the decoded blocks decoded in the IBC mode include decoded blocks in at least two channels, and the decoded blocks in the at least two channels share a BV of the reference prediction block; 19. The image decoding method according to claim 18.
20. The image decoding method includes: and determining, when an IBC mode identifier is analyzed from any one of the plurality of sub-streams, that a target decoding mode corresponding to the plurality of sub-streams is the IBC mode.
19. The image decoding method according to claim 18.
21. The image decoding method includes: Determining residual coefficients corresponding to the decoding unit, the residual coefficients including multiple groups of residual coefficients, and the number of residual coefficients in each group is less than or equal to a number threshold; decoding a codestream based on the residual coefficients.
19. The image decoding method according to claim 18.
22. analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; mode flags are analyzed from the plurality of substreams, and if a second total code length is greater than the remaining capacity of a codestream buffer, a target decoding mode of the substream is determined to be a fallback mode, the mode flags are for indicating whether decoded blocks in the plurality of channels are coded using a fallback mode or whether a first channel uses a first decoding mode, the second total code length is a total code length of a codestream obtained by coding all of the coding blocks of the plurality of channels using a first coding mode, the first coding mode is a mode for coding sample values in coding blocks with a first fixed-length code, the code length of the first fixed-length code is less than or equal to an image bit width of the image to be processed, and the image bit width is for indicating the number of bits required to store each sample in the image to be processed; Analyzing preset flag bits in the substream to determine a target fallback mode, where the target fallback mode is one of the fallback modes, the preset flag bits are for indicating types of fallback modes to be used when encoding the coding blocks of the multiple channels, and the fallback modes include a first fallback mode and a second fallback mode; and decoding the sub-stream in the target fallback mode. Image decoding method.
23. The image decoding method includes: determining a target decoding code length of the coding unit based on the remaining capacity of the code stream buffer and a target pixel depth BPP, wherein the target decoding code length indicates a code length required to decode the code stream of the coding unit; determining assigned code lengths for a plurality of channels based on the decoding code lengths, the assigned code lengths indicating code lengths required to decode residuals of code streams of the coded blocks of the plurality of channels; determining a decoding code length assigned to each of the plurality of channels based on an average value of the assigned code lengths for the plurality of channels; 23. The image decoding method according to claim 22.
24. encoding the coding unit to obtain a codestream; decoding a substream corresponding to a first channel in a first decoding mode; Image decoding method.
25. Decoding the substream corresponding to the first channel in a first decoding mode includes: If the first fixed-length code is equal to the image bit width, decoding the sub-stream corresponding to the first channel in the first decoding mode as is; or If the first fixed-length code is smaller than the image bit width, dequantizing pixel values of the analyzed coded block of the first channel.
25. The image decoding method according to claim 24.
26. Obtaining a decoding unit, the decoding unit being an image block in a current image, the decoding unit including decoding blocks in multiple channels; determining a first total code length, the first total code length being the total code length of a first code stream obtained by decoding all of the decoded blocks in the plurality of channels in the corresponding target decoding modes; decoding the decoded blocks in the plurality of channels in a fallback mode when the first total code length is equal to or greater than a remaining capacity of a code stream buffer; Image decoding method.
27. The image decoding method includes: further comprising decoding mode flags in a plurality of substreams obtained by decoding the decoded blocks in the plurality of channels.
27. The image decoding method according to claim 26.
28. analyzing a codestream obtained by encoding a coding unit to determine processing coefficients corresponding to the coding unit; and decoding the codestream based on the processing coefficients. Image decoding method.
29. the processing coefficients include one or more of residual coefficients and transform coefficients, the processing coefficients include a plurality of groups of processing coefficients, and the number of processing coefficients in each group is less than or equal to a number threshold; 29. The image decoding method according to claim 28.
30. an acquisition module for acquiring a coding unit, the coding unit including coding blocks of multiple channels; a processing module that encodes preset codewords in a target substream that satisfies a preset condition until the target substream no longer satisfies the preset condition, the target substream being one of a plurality of substreams, the plurality of substreams being codestreams obtained by encoding coding blocks of the plurality of channels; Image encoding device.
31. an acquisition module for acquiring a coding unit, the coding unit including coding blocks of multiple channels; a processing module for encoding the coding block of each channel based on the preset expansion rate so that the current expansion rate is equal to or less than the preset expansion rate; Image encoding device.
32. an acquisition module for acquiring a coding unit, the coding unit including coding blocks of multiple channels; a processing module used for: encoding a coding block of at least one channel among the plurality of channels in an IBC mode, which is an intra block copy mode; obtaining a block vector BV of a reference prediction block, wherein the BV of the reference prediction block indicates a position of the reference prediction block in a coded image block and represents a predicted value of a coding block coded in the IBC mode; and coding the BV of the reference prediction block in a substream obtained by coding the coding block of each of the plurality of channels in the IBC mode. Image encoding device.
33. an acquisition module for acquiring a coding unit, the coding unit being an image block in a target image, the coding unit including coding blocks of multiple channels; a processing module configured to determine a first total code length, the first total code length being a total code length of a first stream obtained by encoding all of the encoding blocks of the plurality of channels in corresponding target encoding modes, the target encoding mode including a first encoding mode, the first encoding mode being a mode for encoding sample values in the encoding blocks with a first fixed-length code, the code length of the first fixed-length code being equal to or less than an image bit width of the image to be processed, the image bit width being intended to represent a number of bits required to store each sample in the image to be processed; and, if the first total code length is equal to or greater than a remaining capacity of a code stream buffer, to encode the encoding blocks of the plurality of channels in a fallback mode, the mode flag of the fallback mode being the same as that of the first encoding mode. Image encoding device.
34. a processing module for analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; determining a number of codewords, the number of codewords indicating a number of preset codewords encoded in a target substream that satisfies a preset condition, the preset codewords being coded into the target substream if a target substream that satisfies the preset condition exists; and decoding the codestream based on the number of codewords. Image decoding device.
35. a processing module for analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of multiple channels; determining a current expansion factor; determining a number of codewords based on the current expansion factor and a first preset expansion factor, the number of codewords indicating a number of preset codewords encoded in a target substream that satisfies a preset condition, the preset codewords being coded into the target substream if a target substream that satisfies the preset condition exists; and decoding the codestream based on the number of codewords. Image decoding device.
36. the code stream obtained by encoding a coding unit includes coding blocks of a plurality of channels, and the code stream includes a plurality of substreams into which the coding blocks of the plurality of channels are coded, the substream corresponding to the plurality of channels in one-to-one correspondence; determining a position of the reference prediction block in the plurality of substreams based on a block vector BV of a reference prediction block analyzed from at least one substream of the plurality of substreams, the reference prediction block representing a predicted value of a decoded block decoded in an IBC mode, which is an intra block copy mode, and the BV of the reference prediction block indicating a position of the reference prediction block in a reconstructed image block; determining the predicted value of the decoded block decoded in the IBC mode based on position information of the reference prediction block; and reconstructing the decoded block decoded in the IBC mode based on the predicted value. Image decoding device.
37. analyzing a codestream obtained by encoding a coding unit, the coding unit including coding blocks of a plurality of channels; analyzing a mode flag from a substream obtained by encoding the coding blocks of the plurality of channels, and determining a target decoding mode of the substream to be a fallback mode if a second total code length is greater than a remaining capacity of a codestream buffer, the mode flag indicating whether the decoding blocks in the plurality of channels are encoded using a fallback mode, the second total code length being a total code length of a codestream obtained by encoding all of the coding blocks of the plurality of channels in a first coding mode, and the first coding mode being a substream in a coding block with a first fixed-length code; a processing module for: encoding a sample value of a substream in a mode in which a code length of the first fixed-length code is equal to or less than an image bit width of a processing target image, the image bit width representing the number of bits required to store each sample in the processing target image; analyzing preset flag bits in the substream to determine a target fallback mode, the target fallback mode being one of the fallback modes, the preset flag bits representing types of fallback modes to be used when encoding the coding blocks of the plurality of channels, the fallback modes including a first fallback mode and a second fallback mode; and decoding the substream in the target fallback mode. Image decoding device.
38. 1. A video encoder comprising: a processor; and a memory; The memory stores instructions executable by the processor; The processor is configured, when executing the instructions, to cause the video encoder to implement the image encoding method of any one of claims 1 to 14. Video encoder.
39. 1. A video decoder comprising: a processor; and a memory; The memory stores instructions executable by the processor; The processor is configured, when executing the instructions, to cause the video decoder to implement the image decoding method of any one of claims 15 to 29. Video decoder.
40. a readable storage medium containing software instructions; When the software instructions are executed by an image encoding device and an image decoding device, the image encoding device realizes the image encoding method according to any one of claims 1 to 14, and the image decoding device realizes the image decoding method according to any one of claims 15 to 29. Readable storage medium.
Citation Information
Patent Citations
Intra prediction from prediction block
JP2017508345A
Substream multiplexing for display stream compression
JP2019522413A
Parallel Entropy Coding
JP2024510268A
Substream multiplexing for display stream compression
US20180343471A1
Method and apparatus for video coding
WO2022132251A1