Image coding method and apparatus

By determining the reference frame number and layer number using channel feedback information, the problem of screen distortion caused by channel changes in scalable video coding is solved, coding efficiency and consistency are improved, and the reference image at the encoding and decoding ends is consistent.

CN116033147BActive Publication Date: 2026-05-22HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2021-10-27
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing scalable video coding schemes are prone to screen flickering issues when channel conditions change, and the reference images at the encoding and decoding ends are inconsistent, affecting coding efficiency.

Method used

The reference frame number and layer number of the current image are determined by the channel feedback information to ensure that the encoding and decoding ends use the same reference image. The reference frame number and reference layer number set are obtained by using the channel feedback information, and the reconstructed image is extracted from the decoding image buffer as the reference image for encoding.

Benefits of technology

It improves encoding efficiency, avoids screen tearing, saves storage space in the decoding image buffer, and ensures the consistency of reference images between the encoding and decoding ends.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116033147B_ABST
    Figure CN116033147B_ABST
Patent Text Reader

Abstract

The application provides an image coding method and device. The image coding method comprises the following steps: determining a reference frame number of a current image according to channel feedback information, wherein the channel feedback information is used for indicating information of an image frame received by a decoding end; obtaining a first reference layer number set of a first image frame corresponding to the reference frame number, wherein the first reference layer number set comprises N1 layer numbers of sub-layers, 1≤N1 L1, and L1 represents a total number of sub-layers of the first image frame; determining a reference layer number of the current image according to the channel feedback information and the first reference layer number set; and performing video coding on the current image according to the reference frame number and the reference layer number to obtain a code stream. The application can fully consider the change of a channel, ensure that the reference images adopted by the coding end and the decoding end are consistent, improve coding efficiency, and avoid the situation of a ghost screen.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to video encoding and decoding technology, and more particularly to an image encoding and decoding method and apparatus. Background Technology

[0002] Scalable video coding, also known as scalable video coding, is an extension of current video coding standards. In scalable video coding, different bitstream layers are formed by spatial domain scaling (resolution scaling), temporal domain scaling, or quality scaling in the encoder, thus allowing video bitstreams with different resolutions, frame rates, or bitrates to be contained within the same bitstream.

[0003] The encoder can encode video frames into a base layer bitstream and enhancement layer bitstreams according to different encoding configurations. The base layer typically encodes the lowest spatial, temporal, or lowest quality bitstream; the enhancement layers use the base layer as a foundation, superimposing higher-level spatial, temporal, or higher-quality bitstreams. As the number of enhancement layers increases, the encoded spatial, temporal, or quality levels also increase. During transmission, the base layer bitstream is prioritized for transmission, and as network capacity allows, increasingly higher-level enhancement layer bitstreams are transmitted gradually. The decoder first receives and decodes the base layer bitstream, and then, based on the received enhancement layer bitstreams from low to high levels, progressively decodes increasingly higher-level spatial, temporal, or quality bitstreams. By superimposing higher-level information onto lower-level information, higher-resolution, higher-frame-rate, or higher-quality video reconstructed frames are obtained.

[0004] However, the scalable coding schemes in related technologies are affected by changes in channel conditions, leading to screen distortion. Summary of the Invention

[0005] This application provides an image encoding and decoding method and apparatus to fully consider channel variations, ensure that the reference image used by the encoding and decoding ends is consistent, improve encoding efficiency, and avoid screen distortion.

[0006] In a first aspect, this application provides an image encoding method, comprising: determining a reference frame number of a current image based on channel feedback information, wherein the channel feedback information is used to indicate information of an image frame received by a decoding end; obtaining a first reference layer number set of a first image frame corresponding to the reference frame number, wherein the first reference layer number set includes N1 layer numbers of layers, 1≤N1<L1, and L1 represents the total number of layers of the first image frame; determining the reference layer number of the current image based on the channel feedback information and the first reference layer number set; and performing video encoding on the current image based on the reference frame number and the reference layer number to obtain a bitstream.

[0007] Channel feedback information is used to indicate the information of the image frame received by the decoding end. For example, if the current image frame has a total of 4 layers, the encoding end encodes the current image frame to obtain the bitstream corresponding to the 4 layers. However, during transmission, the decoding end only receives the bitstream corresponding to the first 3 layers of the current image frame. When encoding the next frame, the encoding end uses the reconstructed image of the 4th layer of the current image as the reference image. Since the decoding end has not received the bitstream corresponding to the 4th layer, it cannot use the reconstructed image of the 4th layer of the current image as the reference image to decode the next frame, resulting in the next frame failing to decode correctly. Therefore, in this application, the encoding end first obtains the channel feedback information and determines the information of the image frame received by the decoding end based on this information. For example, the channel feedback information includes the frame number and layer number of the image frame received by the decoding end. Then, based on this, the reference frame for the next frame is determined, thereby avoiding the situation described in the example above and ensuring that the reference image used by the encoding end and the decoding end is consistent.

[0008] In this application, the maximum number of layers L in the image frames of the video during scalable coding can be preset. max For example, L max =6. This maximum number of layers can be a threshold, meaning that the number of layers in each image frame will not exceed this number. However, in actual encoding, different image frames may have different total number of layers. The total number of layers in the first image frame is denoted by L1, and L1 can be less than or equal to the aforementioned maximum number of layers L. max After the first image frame is divided into L1 layers, N1 layers can be used as reference images for subsequent image frames, where 1 ≤ N1 < L1. The layer numbers of these N1 layers form the first reference layer number set for the first image frame. That is, the first image frame corresponds to the first reference layer number set, and only the reconstructed images of layers whose layer numbers are in the first reference layer number set can be used as reference images for subsequent image frames. The same applies to other image frames in the video, which will not be elaborated here.

[0009] After determining the reference frame number and reference layer number of the current image, the reconstructed image corresponding to the reference frame number and reference layer number can be extracted from the decoded picture buffer (DPB) as the reference image of the current image. Based on this reference image, the current image can be scalably encoded to obtain the bitstream.

[0010] This application determines the reference frame number of the current image based on channel feedback information, and then determines the reference layer number of the current image based on the reference frame number and a pre-set set of reference layer numbers. The set of reference layer numbers includes the layer numbers of N layers of the image frame corresponding to the reference frame number. Then, the reference image of the current image is obtained based on the reference frame number and the reference layer number. The reference image obtained in this way fully considers the changes in the channel, ensures that the reference image used by the encoding end and the decoding end is consistent, improves encoding efficiency, and avoids screen distortion.

[0011] In one possible implementation, the step of performing video encoding on the current image according to the reference frame number and the reference layer number to obtain a bitstream includes: obtaining a reconstructed image corresponding to the reference frame number and the reference layer number from a decoded image buffer (DPB), wherein the DPB contains only the N1 layered reconstructed images for the first image frame; using the obtained reconstructed image corresponding to the reference frame number and the reference layer number as a reference image, and performing video encoding on the current image according to the reference image to obtain the bitstream.

[0012] Since the DPB only stores the reconstructed images of its N1 layers for the first image frame, instead of storing the reconstructed images of all L1 layers of the first image frame, this can save the storage space of the DPB and improve the coding efficiency.

[0013] In one possible implementation, when there is only one decoding end, determining the reference frame number of the current image based on the channel feedback information includes: acquiring multiple channels feedback information, wherein the channels feedback information is used to indicate the frame number of the image frame received by the decoding end; and determining the frame number of the multiple frame numbers indicated by the multiple channels feedback information that is closest to the current image as the reference frame number of the current image.

[0014] In one possible implementation, determining the reference layer number of the current image based on the channel feedback information and the first reference layer number set includes: determining the highest layer number indicated by the channel feedback information indicating the reference frame number as the target layer number; when the first reference layer number set includes the target layer number, determining the target layer number as the reference layer number of the current image; or, when the first reference layer number set does not include the target layer number, determining the layer number in the first reference layer number set that is less than and closest to the target layer number as the reference layer number of the current image.

[0015] In one possible implementation, when there are multiple decoding ends, determining the reference frame number of the current image based on channel feedback information includes: acquiring multiple sets of channel feedback information, the multiple sets of channel feedback information corresponding to the multiple decoding ends, each set of channel feedback information including multiple sets of channel feedback information, the channel feedback information being used to indicate the frame number of the image frame received by the corresponding decoding end; determining one or more common frame numbers based on the multiple sets of channel feedback information, the common frame number being a frame number indicated by at least one channel feedback information in each set of channel feedback information; and determining the reference frame number of the current image based on the one or more common frame numbers.

[0016] In this application, the highest layer number indicated by the channel feedback information indicating the common frame number can be determined as the target layer number; when the first reference layer number set includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, when the first reference layer number set does not include the target layer number, the layer number in the first reference layer number set that is less than and closest to the target layer number is determined as the reference layer number of the current image.

[0017] The reference frame number determined based on the channel feedback information can not only conform to the channel conditions, but also identify the received image frame that is closest to the current image as the reference frame, which can improve coding efficiency.

[0018] In one possible implementation, determining the reference layer number of the current image based on the channel feedback information and the first set of reference layer numbers includes: obtaining the highest layer number indicated by the channel feedback information indicating the reference frame number in each of the multiple sets of channel feedback information; determining the smallest of the multiple highest layer numbers as the target layer number; and determining the reference layer number of the current image based on the target layer number and the first set of reference layer numbers.

[0019] In one possible implementation, the channel feedback information comes from the corresponding decoding end and / or network devices on the transmission link.

[0020] In one possible implementation, the channel feedback information is generated based on the transmitted code stream.

[0021] In one possible implementation, determining the reference frame number of the current image based on channel feedback information includes: acquiring multiple channels feedback information, the channels feedback information indicating the frame number of the image frame received by the decoding end; determining the frame number of the multiple frame numbers indicated by the multiple channels feedback information that is closest to the current image as the target frame number; when the highest layer number indicated by the channels feedback information indicating the target frame number is greater than or equal to the highest layer number in the second reference layer number set, determining the target frame number as the reference frame number, the second reference layer number set being the reference layer number set of the second image frame corresponding to the target frame number.

[0022] In one possible implementation, the method further includes: when the highest layer number indicated by the channel feedback information indicating the target frame number is less than the highest layer number in the second set of reference layer numbers, determining a specified frame number among the multiple frame numbers indicated by the multiple channel feedback information as the reference frame number of the current image.

[0023] In one possible implementation, the method further includes: when the first reference layer number set does not include the target layer number, if the first reference layer number set does not include a layer number less than the target layer number, then the reference frame number of the previous frame of the current image is determined as the reference frame number of the current image, and the reference layer number of the previous frame is determined as the reference layer number of the current image.

[0024] In one possible implementation, the bitstream further includes the first set of reference layer numbers.

[0025] In one possible implementation, the bitstream also includes the reference frame number.

[0026] In one possible implementation, the bitstream further includes the reference frame number and the reference layer number.

[0027] In one possible implementation, when the current image is an image fragment, the channel feedback information includes the image fragment number of the image frame received by the decoding end and the layer number corresponding to the image fragment number; determining the reference layer number of the current image based on the channel feedback information and the first reference layer number set includes: if the image fragment number of the current image is the same as the image fragment number of the image frame received by the decoding end, determining the layer number corresponding to the image fragment number of the image frame received by the decoding end as the target layer number; when the first reference layer number set includes the target layer number, determining the target layer number as the reference layer number of the current image; or, when the first reference layer number set does not include the target layer number, determining the layer number in the first reference layer number set that is smaller than and closest to the target layer number as the reference layer number of the current image.

[0028] An image frame can be divided into multiple image slices for encoding and transmission. Therefore, when the current image is an image slice, in addition to the frame number of the image frame received by the decoding end, the channel feedback information also includes the image slice number of the image frame received by the decoding end and the layer number corresponding to that image slice number. Thus, the encoding end can first determine the reference frame number of the current image using the method described above, and then determine the reference layer number of the current image based on the image slice number of the current image and the first set of reference layer numbers. That is, it finds the image slice number that is the same as the current image among the multiple image slice numbers indicated by the channel feedback information indicating the reference frame number, and then determines the layer number corresponding to this identical image slice number as the target layer number. Based on the target layer number, the reference layer number of the current image is determined from the first set of reference layer numbers. For example, if the reference frame number of the current image is 1, and the channel feedback information indicating frame number 1 indicates multiple image slice numbers including 1, 2, 3, and 4, the layer number corresponding to image slice number 1 is 3, the layer number corresponding to image slice number 2 is 4, the layer number corresponding to image slice number 3 is 5, the layer number corresponding to image slice number 4 is 6, and the first reference layer number set Rx = {1, 3, 5}. If the image slice number of the current image is 1, then the reference layer number of the current image is obtained based on the aforementioned layer number 3 corresponding to image slice number 1 and the first reference layer number set, and its reference layer number is 3. Alternatively, if the image slice number of the current image is 2, then the reference layer number of the current image is obtained based on the aforementioned layer number 4 corresponding to image slice number 2 and the first reference layer number set, and its reference layer number is 3. Or, if the image slice number of the current image is 3, then the reference layer number of the current image is obtained based on the aforementioned layer number 5 corresponding to image slice number 3 and the first reference layer number set, and its reference layer number is 5. Alternatively, if the image slice number of the current image is 4, then the reference layer number of the current image is obtained based on the layer number 6 corresponding to the aforementioned image slice number 4 and the first reference layer number set, and its reference layer number is 5.

[0029] Secondly, this application provides an image decoding method, comprising: acquiring a bitstream; parsing the bitstream to acquire a reference frame number of a current image; acquiring a third reference layer number set of a third image frame corresponding to the reference frame number, the third reference layer number set including N2 layer numbers, 1≤N2<L2, where L2 represents the total number of layers of the third image frame; determining the reference layer number of the current image based on the third reference layer number set; and performing video decoding based on the reference frame number and the reference layer number to obtain a reconstructed image of the current image.

[0030] In this application, the decoding end can send channel feedback information when it parses the bitstream indicating the start of receiving the next frame (based on the frame number in the bitstream), carrying the frame number of the previous frame and the highest layer number received in the previous frame; alternatively, it can send channel feedback information when it parses the bitstream indicating the current image has been received (based on the highest layer number of the current image in the bitstream), carrying the frame number of the current image and the highest layer number received in the current image. By sending channel feedback information in both of these cases, the decoding end can ensure that the highest layer received in any frame obtained by the encoding end is consistent with the highest layer of the same frame actually received by the decoding end, thereby avoiding errors that can occur at the encoding end due to the use of different reference images during encoding and decoding.

[0031] In one possible implementation, the step of performing video decoding based on the reference frame number and the reference layer number to obtain the reconstructed image of the current image includes: obtaining the reconstructed image corresponding to the reference frame number and the reference layer number from the decoded image buffer (DPB); using the obtained reconstructed image corresponding to the reference frame number and the reference layer number as a reference image, and performing video decoding based on the reference image to obtain the reconstructed image of the current image.

[0032] In one possible implementation, the method further includes: storing the reconstructed images of the N3 layers of the current image into the DPB, wherein the fourth reference layer number set of the current image includes the layer numbers of M layers, the M layers include the N3 layers, 1≤M<L3, and L3 represents the total number of layers of the current image; or, storing the reconstructed image of the highest layer among the N3 layers into the DPB.

[0033] The fourth reference layer number set of the current image can be obtained by parsing the bitstream. However, when the decoding end is decoding, the highest layer number L4 obtained for the current image may be less than the total number of layers L3 of the current image. Therefore, when storing the reconstructed image of the current image into the DPB, if L4 is greater than or equal to the highest layer number among the above M layers, then N3 = M; and if L4 is less than the highest layer number among the above M layers, then N3 < M.

[0034] In one possible implementation, the method further includes: displaying the reconstructed image of the L4 layer of the current image, where L4 represents the layer number of the highest layer obtained by decoding the current image.

[0035] In one possible implementation, determining the reference layer number of the current image based on the third reference layer number set includes: determining the highest layer number among the layer numbers corresponding to the multiple reconstructed images of the decoded third image frame; when the third reference layer number set includes the highest layer number, determining the highest layer number as the reference layer number of the current image; or, when the reference layer number set does not include the highest layer number, determining the layer number in the third reference layer number set that is less than and closest to the highest layer number as the reference layer number of the current image.

[0036] In one possible implementation, the method further includes: when the third reference layer number set does not include the highest layer number, if the third reference layer number set does not include a layer number less than the highest layer number, then the reference frame number of the previous frame of the current image is determined as the reference frame number of the current image, and the reference layer number of the previous frame is determined as the reference layer number of the current image.

[0037] In one possible implementation, the method further includes: determining the frame number and layer number of the received image frame; and sending channel feedback information to the encoding end, the channel feedback information being used to indicate the frame number and layer number of the received image frame.

[0038] In one possible implementation, sending channel feedback information to the encoding end includes: when it is determined that parsing of the second frame has begun based on the frame number in the bitstream, sending the channel feedback information to the encoding end, wherein the channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the received first frame, and the first frame is the frame preceding the second frame; or, when it is determined that the first frame has been completely received based on the layer number of the received image frame, sending the channel feedback information to the encoding end, wherein the channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the received first frame.

[0039] In one possible implementation, when the current image is an image fragment, the method further includes: determining the image fragment number of the received image frame; correspondingly, the channel feedback information is also used to indicate the image fragment number.

[0040] An image frame is divided into multiple image slices for encoding and transmission. When the current image is an image slice, the decoding end can determine the received image slice number along with the received frame number and slice number. This information is then included in the channel feedback information along with the frame number, slice number, and corresponding slice number of the received image slice. The above image-based processing can be performed in the same way for image slices.

[0041] Thirdly, this application provides an image encoding apparatus, comprising: an inter-frame prediction module, configured to determine a reference frame number of a current image based on channel feedback information, wherein the channel feedback information is used to indicate information of an image frame received by a decoding end; to obtain a first reference layer number set of a first image frame corresponding to the reference frame number, wherein the first reference layer number set includes N1 layer numbers, 1≤N1<L1, and L1 represents the total number of layers of the first image frame; to determine the reference layer number of the current image based on the channel feedback information and the first reference layer number set; and an encoding module, configured to perform video encoding on the current image based on the reference frame number and the reference layer number to obtain a bitstream.

[0042] In one possible implementation, the encoding module is specifically used to obtain the reconstructed image corresponding to the reference frame number and the reference layer number from the decoded image buffer DPB, wherein the DPB contains only the N1 layered reconstructed images for the first image frame; the obtained reconstructed image corresponding to the reference frame number and the reference layer number is used as a reference image, and the current image is video encoded according to the reference image to obtain the bitstream.

[0043] In one possible implementation, when there is only one decoding end, the inter-frame prediction module is specifically used to acquire multiple channel feedback information, the channel feedback information being used to indicate the frame number of the image frame received by the decoding end; and to determine the frame number of the current image that is closest to the multiple frame numbers indicated by the multiple channel feedback information as the reference frame number of the current image.

[0044] In one possible implementation, the inter-frame prediction module is specifically configured to determine the highest layer number indicated by the channel feedback information indicating the reference frame number as the target layer number; when the first set of reference layer numbers includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, when the first set of reference layer numbers does not include the target layer number, the layer number in the first set of reference layer numbers that is less than and closest to the target layer number is determined as the reference layer number of the current image.

[0045] In one possible implementation, when there are multiple decoding ends, the inter-frame prediction module is specifically used to acquire multiple sets of channel feedback information, which correspond to the multiple decoding ends. Each set of channel feedback information includes multiple sets of channel feedback information, which are used to indicate the frame number of the image frame received by the corresponding decoding end; determine one or more common frame numbers based on the multiple sets of channel feedback information, where the common frame number refers to the frame number indicated by at least one channel feedback information in each set of channel feedback information; and determine the reference frame number of the current image based on the one or more common frame numbers.

[0046] In one possible implementation, the inter-frame prediction module is specifically used to obtain the highest layer number indicated by the channel feedback information indicating the reference frame number in each of the multiple sets of channel feedback information; determine the smallest of the multiple highest layer numbers as the target layer number; and determine the reference layer number of the current image based on the target layer number and the first set of reference layer numbers.

[0047] In one possible implementation, the channel feedback information comes from the corresponding decoding end and / or network devices on the transmission link.

[0048] In one possible implementation, the channel feedback information is generated based on the transmitted code stream.

[0049] In one possible implementation, the inter-frame prediction module is specifically used to acquire multiple channel feedback information, the channel feedback information being used to indicate the frame number of the image frame received by the decoding end; determine the frame number closest to the current image among the multiple frame numbers indicated by the multiple channel feedback information as the target frame number; when the highest layer number indicated by the channel feedback information indicating the target frame number is greater than or equal to the highest layer number in the second reference layer number set, determine the target frame number as the reference frame number, the second reference layer number set being the reference layer number set of the second image frame corresponding to the target frame number.

[0050] In one possible implementation, the inter-frame prediction module is further configured to determine a specified frame number among the multiple frame numbers indicated by the multiple channel feedback information as the reference frame number of the current image when the highest layer number indicated by the channel feedback information indicating the target frame number is less than the highest layer number in the second reference layer number set.

[0051] In one possible implementation, the inter-frame prediction module is further configured to, when the first reference layer number set does not include the target layer number, if the first reference layer number set does not include a layer number less than the target layer number, determine the reference frame number of the previous frame of the current image as the reference frame number of the current image, and determine the reference layer number of the previous frame as the reference layer number of the current image.

[0052] In one possible implementation, the bitstream further includes the first set of reference layer numbers.

[0053] In one possible implementation, the bitstream also includes the reference frame number.

[0054] In one possible implementation, the bitstream further includes the reference frame number and the reference layer number.

[0055] In one possible implementation, when the current image is an image slice, the inter-frame prediction module is specifically used to determine that the image slice number of the current image is the same as the image slice number of the image frame received by the decoding end, and to determine the layer number corresponding to the image slice number of the image frame received by the decoding end as the target layer number; when the first reference layer number set includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, when the first reference layer number set does not include the target layer number, the layer number in the first reference layer number set that is smaller than and closest to the target layer number is determined as the reference layer number of the current image.

[0056] Fourthly, this application provides an image decoding apparatus, comprising: an acquisition module for acquiring a bitstream; an inter-frame prediction module for parsing the bitstream to acquire a reference frame number of a current image; acquiring a set of third reference layer numbers of a third image frame corresponding to the reference frame number, the set of third reference layer numbers including N2 layer numbers, 1≤N2<L2, where L2 represents the total number of layers of the third image frame; determining the reference layer number of the current image based on the set of third reference layer numbers; and a decoding module for performing video decoding based on the reference frame number and the reference layer number to obtain a reconstructed image of the current image.

[0057] In one possible implementation, the decoding module is specifically used to obtain the reconstructed image corresponding to the reference frame number and the reference layer number from the decoded image buffer DPB; use the obtained reconstructed image corresponding to the reference frame number and the reference layer number as a reference image, and perform video decoding based on the reference image to obtain the reconstructed image of the current image.

[0058] In one possible implementation, the decoding module is further configured to store the reconstructed images of the N3 layers of the current image into the DPB, wherein the fourth reference layer number set of the current image includes the layer numbers of M layers, the M layers include the N3 layers, 1≤M<L3, and L3 represents the total number of layers of the current image; or, to store the reconstructed image of the highest layer among the N3 layers into the DPB.

[0059] In one possible implementation, it further includes: a display module, used to display the reconstructed image of the L4 layer of the current image, where L4 represents the layer number of the highest layer obtained by decoding the current image.

[0060] In one possible implementation, the inter-frame prediction module is specifically used to determine the highest layer number among the layer numbers corresponding to the multiple reconstructed images of the decoded third image frame; when the third reference layer number set includes the highest layer number, the highest layer number is determined as the reference layer number of the current image; or, when the reference layer number set does not include the highest layer number, the layer number in the third reference layer number set that is less than and closest to the highest layer number is determined as the reference layer number of the current image.

[0061] In one possible implementation, the inter-frame prediction module is further configured to, when the third reference layer number set does not include the highest layer number, if the third reference layer number set does not include a layer number less than the highest layer number, determine the reference frame number of the previous frame of the current image as the reference frame number of the current image, and determine the reference layer number of the previous frame as the reference layer number of the current image.

[0062] In one possible implementation, the module further includes: a sending module, configured to determine the frame number and layer number of the received image frame; and to send channel feedback information to the encoding end, wherein the channel feedback information is used to indicate the frame number and layer number of the received image frame.

[0063] In one possible implementation, the sending module is specifically configured to send the channel feedback information to the encoding end when it is determined, based on the frame number in the bitstream, to start parsing the second frame. The channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the received first frame, wherein the first frame is the frame preceding the second frame; or, when it is determined, based on the layer number of the received image frame, that the first frame has been completely received, the sending module sends the channel feedback information to the encoding end. The channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the received first frame.

[0064] In one possible implementation, when the current image is an image fragment, the sending module is further configured to determine the image fragment number of the received image frame; correspondingly, the channel feedback information is further configured to indicate the image fragment number.

[0065] Fifthly, this application provides an encoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processors and storing a program executed by the processors, wherein the program, when executed by the processors, causes the encoder to perform the method according to any one of the first aspects.

[0066] Sixthly, this application provides a decoder, comprising: one or more processors;

[0067] A non-transitory computer-readable storage medium coupled to the processor and storing a program executed by the processor, wherein the program, when executed by the processor, causes the decoder to perform the method according to any one of the second aspects.

[0068] In a seventh aspect, this application provides a non-transitory computer-readable storage medium including program code, which, when executed by a computer device, is used to perform the method described according to any one of the first or second aspects.

[0069] Eighthly, this application provides a non-transient storage medium comprising a bit stream encoded according to the method described in any one of the first or second aspects.

[0070] Ninthly, this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in either the first or second aspect. Attached Figure Description

[0071] Figure 1A This is an exemplary block diagram of the decoding system 10 according to an embodiment of this application;

[0072] Figure 1B This is an exemplary block diagram of the video decoding system 40 according to an embodiment of this application;

[0073] Figure 2 This is an exemplary block diagram of the video encoder 20 according to an embodiment of this application;

[0074] Figure 3 This is an exemplary block diagram of the video decoder 30 according to an embodiment of this application;

[0075] Figure 4 This is an exemplary block diagram of a video decoding device 400 according to an embodiment of this application;

[0076] Figure 5 This is an exemplary hierarchical diagram illustrating the scalable video coding of this application;

[0077] Figure 6 An exemplary flowchart of the encoding method for the enhancement layer of this application;

[0078] Figure 7 Here is an exemplary flowchart of the image encoding method of this application;

[0079] Figure 8 Here is an exemplary flowchart of the image decoding method of this application;

[0080] Figure 9 This is a schematic diagram of the structure of the encoding device 900 according to an embodiment of this application;

[0081] Figure 10This is a schematic diagram of the structure of the decoding device 1000 according to an embodiment of this application. Detailed Implementation

[0082] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0083] The terms "first," "second," etc., used in the specification, embodiments, claims, and drawings of this application are for distinguishing purposes only and should not be construed as indicating or implying relative importance or order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0084] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0085] Video coding generally refers to the processing of image sequences that form a video or video sequence. In the field of video coding, the terms "picture," "frame," or "image" can be used synonymously. Video coding (or commonly referred to as encoding) comprises two parts: video encoding and video decoding. Video encoding is performed on the source side and typically involves processing (e.g., compressing) the raw video image to reduce the amount of data required to represent that video image (thus enabling more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves inverse processing relative to the encoder to reconstruct the video image. The "encoding" of the video image (or commonly referred to as an image) mentioned in the embodiments should be understood as the "encoding" or "decoding" of the video image or video sequence. The encoding and decoding parts are also collectively referred to as codec (encoding and decoding, CODEC).

[0086] In lossless video coding, the original video image can be reconstructed, meaning the reconstructed video image has the same quality as the original (assuming no transmission loss or other data loss during storage or transmission). In lossy video coding, further compression is performed through quantization to reduce the amount of data required to represent the video image, and the decoder cannot completely reconstruct the video image, meaning the quality of the reconstructed video image is lower or worse than the quality of the original video image.

[0087] Several video coding standards fall under the category of "lossy hybrid video coding and decoding" (i.e., combining spatial and temporal prediction in the pixel domain with 2D transform coding in the transform domain for applying quantization). Each image in a video sequence is typically segmented into a set of non-overlapping blocks, which are usually encoded at the block level. In other words, the encoder typically processes the video at the block (video block) level, for example, generating prediction blocks through spatial (intra-frame) prediction and temporal (inter-frame) prediction; subtracting the prediction blocks from the current block (the block currently being processed / to be processed) to obtain residual blocks; transforming and quantizing the residual blocks in the transform domain to reduce the amount of data to be transmitted (compressed), while the decoder applies the inverse processing relative to the encoder to the encoded or compressed blocks to reconstruct the current block for representation. Additionally, the encoder needs to repeat the decoder's processing steps so that the encoder and decoder generate the same predictions (e.g., intra-frame and inter-frame predictions) and / or reconstruct pixels for processing, i.e., encoding subsequent blocks.

[0088] In the following embodiment of the decoding system 10, the encoder 20 and decoder 30 are based on Figures 1A to 3 Describe it.

[0089] Figure 1AThis is an exemplary block diagram of a decoding system 10 according to an embodiment of this application, such as a video decoding system 10 (or simply decoding system 10) that can utilize the technology of this application. The video encoder 20 (or simply encoder 20) and video decoder 30 (or simply decoder 30) in the video decoding system 10 represent devices, etc., that can be used to perform various technologies according to the various examples described in this application.

[0090] like Figure 1A As shown, the decoding system 10 includes a source device 12, which provides encoded image data 21, such as encoded images, to a destination device 14 for decoding the encoded image data 21.

[0091] The source device 12 includes an encoder 20, and optionally may include an image source 16, a preprocessor (or preprocessing unit) 18 such as an image preprocessor, and a communication interface (or communication unit) 22.

[0092] Image source 16 may include or may be any type of image capture device for capturing real-world images, and / or any type of image generation device, such as a computer graphics processor for generating computer animation images or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage device storing any of the images described above.

[0093] To distinguish the processing performed by the preprocessor (or preprocessing unit) 18, the image (or image data) 17 may also be referred to as the raw image (or raw image data) 17.

[0094] The preprocessor 18 receives the raw image data 17 and preprocesses it to obtain a preprocessed image (or preprocessed image data) 19. For example, the preprocessing performed by the preprocessor 18 may include cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or noise reduction. It is understood that the preprocessing unit 18 may be an optional component.

[0095] Video encoder (or encoder) 20 is used to receive preprocessed image data 19 and provide encoded image data 21 (hereinafter referred to as...) Figure 2 (and so on, for further description).

[0096] The communication interface 22 in the source device 12 can be used to: receive encoded image data 21 and send encoded image data 21 (or other arbitrarily processed version) to another device such as the destination device 14 or any other device via the communication channel 13 for storage or direct reconstruction.

[0097] The target device 14 includes a decoder 30, and optionally may include a communication interface (or communication unit) 28, a post-processor (or post-processing unit) 32 and a display device 34.

[0098] The communication interface 28 in the destination device 14 is used to receive encoded image data 21 (or other processed versions) directly from the source device 12 or from any other source device such as a storage device, for example, the storage device is an encoded image data storage device, and to provide the encoded image data 21 to the decoder 30.

[0099] Communication interfaces 22 and 28 can be used to send or receive encoded image data (or encoded data 21) through a direct communication link between source device 12 and destination device 14, such as a direct wired or wireless connection, or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any combination thereof.

[0100] For example, the communication interface 22 can be used to encapsulate the encoded image data 21 into a suitable format such as a message, and / or process the encoded image data using any type of transmission encoding or processing, so as to transmit it on a communication link or communication network.

[0101] Communication interface 28 corresponds to communication interface 22. For example, it can be used to receive transmitted data and process the transmitted data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain encoded image data 21.

[0102] Both communication interface 22 and communication interface 28 can be configured as follows: Figure 1A The arrow pointing from the source device 12 to the corresponding communication channel 13 of the destination device 14 indicates a one-way or two-way communication interface, which can be used to send and receive messages, establish connections, acknowledge and exchange any other information related to the communication link and / or data transmission, such as encoded image data transmission, etc.

[0103] Video decoder (or decoder) 30 is used to receive encoded image data 21 and provide decoded image data (or decoded image data) 31 (hereinafter referred to as...). Figure 3 (and so on, for further description).

[0104] The post-processor 32 is used to post-process the decoded image data 31 (also known as the reconstructed image data) to obtain post-processed image data 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color adjustment, trimming or resampling, or any other processing to generate the decoded image data 31 for display by the display device 34, etc.

[0105] Display device 34 is used to receive post-processed image data 33 to display the image to a user or viewer. Display device 34 can be or includes any type of display for representing the reconstructed image, such as an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display screen.

[0106] although Figure 1A The source device 12 and destination device 14 are shown as independent devices, but device embodiments may also include both source device 12 and destination device 14, or the functions of both source device 12 and destination device 14, that is, simultaneously including source device 12 or its corresponding functions and destination device 14 or its corresponding functions. In these embodiments, source device 12 or its corresponding functions and destination device 14 or its corresponding functions may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.

[0107] According to the description, Figure 1A The presence and (accurate) division of different units or functions in the source device 12 and / or destination device 14 shown may vary depending on the actual device and application, which is obvious to those skilled in the art.

[0108] Encoder 20 (e.g., video encoder 20) or decoder 30 (e.g., video decoder 30) or both can be transmitted via, for example, Figure 1BThe processing circuitry shown can be implemented as, for example, one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding dedicated processors, or any combination thereof. Encoder 20 can be implemented via processing circuitry 46 to include reference... Figure 2 Encoder 20 refers to various modules discussed herein and / or any other encoder system or subsystem described herein. Decoder 30 may be implemented via processing circuitry 46 to include references. Figure 3 Decoder 30 comprises various modules discussed herein and / or any other decoder system or subsystem described herein. The processing circuitry 46 can be used to perform various operations discussed below. Figure 5 As shown, if some of the technology is implemented in software, the device can store the software instructions in a suitable computer-readable storage medium and execute the instructions in hardware using one or more processors, thereby performing the technology of the present invention. One of the video encoder 20 and the video decoder 30 can be integrated into a single device as part of a combined codec (encoder / decoder, CODEC), such as... Figure 1B As shown.

[0109] Source device 12 and destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, such as a laptop or notebook computer, mobile phone, smartphone, tablet or tablet computer, camera, desktop computer, set-top box, television, display device, digital media player, video game console, video streaming device (e.g., content service server or content distribution server), broadcast receiving device, broadcast transmitting device, etc., and may or may not use an operating system of any type. In some cases, source device 12 and destination device 14 may be equipped with components for wireless communication. Therefore, source device 12 and destination device 14 may be wireless communication devices.

[0110] In some cases, Figure 1AThe video decoding system 10 shown is merely exemplary, and the technology provided in this application can be applied to video encoding setups (e.g., video encoding or video decoding) that do not necessarily include any data communication between the encoding and decoding devices. In other examples, data is retrieved from local memory, sent over a network, etc. A video encoding device may encode data and store it in memory, and / or a video decoding device may retrieve data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data into memory and / or retrieve and decode data from memory.

[0111] Figure 1B This is an exemplary block diagram of the video decoding system 40 according to an embodiment of this application, such as... Figure 1B As shown, the video decoding system 40 may include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video encoder / decoder implemented by processing circuitry 46), an antenna 42, one or more processors 43, one or more memory storage devices 44, and / or a display device 45.

[0112] like Figure 1B As shown, the imaging device 41, antenna 42, processing circuitry 46, video encoder 20, video decoder 30, processor 43, memory storage 44, and / or display device 45 are capable of communicating with each other. In different instances, the video decoding system 40 may contain only the video encoder 20 or only the video decoder 30.

[0113] In some instances, antenna 42 can be used to transmit or receive encoded bitstreams of video data. Additionally, in some instances, display device 45 can be used to present video data. Processing circuitry 46 can include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. Video decoding system 40 can also include an optional processor 43, which similarly can include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. Furthermore, memory storage 44 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, memory storage 44 can be implemented using high-speed cache memory. In other instances, processing circuitry 46 can include memory (e.g., cache, etc.) for implementing image buffers, etc.

[0114] In some instances, the video encoder 20 implemented via logic circuitry may include (e.g., implemented via processing circuitry 46 or memory storage 44) an image buffer and (e.g., implemented via processing circuitry 46) a graphics processing unit. The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video encoder 20 implemented via processing circuitry 46 to implement reference... Figure 2 And / or any other encoder system or subsystem described herein, and the various modules discussed herein. Logic circuits may be used to perform the various operations discussed herein.

[0115] In some instances, the video decoder 30 can be implemented in a similar manner via the processing circuitry 46 to implement the reference. Figure 3 The video decoder 30 and / or any other decoder system or subsystem described herein are various modules discussed. In some instances, the logic circuit-implemented video decoder 30 may include (implemented via processing circuitry 46 or memory storage 44) an image buffer and (e.g., implemented via processing circuitry 46) a graphics processing unit. The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video decoder 30 implemented via processing circuitry 46 to implement reference... Figure 3 And / or the various modules discussed in any other decoder system or subsystem described herein.

[0116] In some instances, antenna 42 can be used to receive encoded bitstreams of video data. As discussed herein, the encoded bitstream may contain data related to encoded video frames, indicators, index values, mode selection data, etc., such as data related to code segmentation (e.g., transform coefficients or quantized transform coefficients, optional indicators, and / or data defining code segmentation). Video decoding system 40 may also include a video decoder 30 coupled to antenna 42 for decoding the encoded bitstream. Display device 45 is used to display the video frames.

[0117] It should be understood that, for the examples described with reference to video encoder 20 in this application embodiment, video decoder 30 can be used to perform the reverse process. Regarding signaling syntax elements, video decoder 30 can be used to receive and parse such syntax elements, and accordingly decode the associated video data. In some examples, video encoder 20 can entropy-encode syntax elements into an encoded video bitstream. In such instances, video decoder 30 can parse such syntax elements and accordingly decode the associated video data.

[0118] For ease of description, embodiments of the present invention are described with reference to the Universal Video Coding (VVC) reference software or the High-Efficiency Video Coding (HEVC) developed by the ITU-T Video Coding Experts Group (VCEG) and the Joint Collaboration Team on Video Coding (JCT-VC) of the ISO / IEC Moving Picture Experts Group (MPEG). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC.

[0119] Encoders and Encoding Methods

[0120] Figure 2 This is an exemplary block diagram of the video encoder 20 according to an embodiment of this application. Figure 2As shown, the video encoder 20 includes an input terminal (or input interface) 201, a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output terminal (or output interface) 272. The mode selection unit 260 may include an inter-frame prediction unit 244, an intra-frame prediction unit 254, and a segmentation unit 262. The inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The video encoder 20 shown can also be called a hybrid video encoder or a video encoder based on a hybrid video codec.

[0121] The residual calculation unit 204, transform processing unit 206, quantization unit 208, and mode selection unit 260 constitute the forward signal path of encoder 20, while the inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, buffer 216, loop filter 220, decoded picture buffer (DPB) 230, inter-frame prediction unit 244, and intra-frame prediction unit 254 constitute the backward signal path of encoder 20. The backward signal path of encoder 20 corresponds to the signal path of decoder (see [link to decoder]). Figure 3 The decoder 30 in the video encoder 20 consists of an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded image buffer 230, an inter-frame prediction unit 244, and an intra-frame prediction unit 254.

[0122] Image and image segmentation (images and patches)

[0123] Encoder 20 can be used to receive images (or image data) 17 via input terminal 201, for example, images in an image sequence forming a video or video sequence. The received images or image data can also be pre-processed images (or pre-processed image data) 19. For simplicity, the following description uses image 17. Image 17 can also be referred to as the current image or the image to be encoded (especially in video encoding when distinguishing the current image from other images, such as those in the same video sequence, i.e., the video sequence that also includes the current image, previously encoded images, and / or decoded images).

[0124] A digital image is, or can be viewed as, a two-dimensional array or matrix of pixels with intensity values. Pixels in an array are also called pixels (short for image element). The number of pixels in the array or image along the horizontal and vertical directions (or axes) determines the image size and / or resolution. To represent color, three color components are typically used, meaning an image can be represented as or comprise an array of three pixels. In RBG format or color space, an image includes corresponding arrays of red, green, and blue pixels. However, in video coding, each pixel is typically represented in a luma / chroma format or color space, such as YCbCr, which includes the luma component indicated by Y (sometimes also represented by L) and two chroma components represented by Cb and Cr. The luma component Y represents the brightness or grayscale level intensity (e.g., both are the same in grayscale images), while the two chroma components Cb and Cr represent the chroma or color information components. Accordingly, a YCbCr format image consists of a luminance pixel array for the luminance pixel value (Y) and two chrominance pixel arrays for the chrominance values ​​(Cb and Cr). An RGB format image can be converted or transformed to YCbCr format, and vice versa; this process is also known as color conversion or transformation. If the image is black and white, it may only include the luminance pixel array. Accordingly, the image can be, for example, a monochrome format luminance pixel array or a 4:2:0, 4:2:2, and 4:4:4 color format luminance pixel array and two corresponding chrominance pixel arrays.

[0125] In one embodiment, the video encoder 20 may include an image segmentation unit ( Figure 2 (Not shown in the image) is used to segment image 17 into multiple (typically non-overlapping) image blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTBs) or coding tree units (CTUs) in the H.265 / HEVC and VVC standards. Segmentation units can be used to apply the same block size and a corresponding grid with defined block sizes to all images in a video sequence, or to vary the block size between images, subsets of images, or groups of images, segmenting each image into corresponding blocks.

[0126] In other embodiments, the video encoder may be used to directly receive blocks 203 of image 17, such as one, several, or all of the blocks that make up image 17. Image block 203 may also be referred to as the current image block or the image block to be encoded.

[0127] Similar to image 17, image block 203 is also a two-dimensional array or matrix composed of pixels with intensity values ​​(pixel values), but image block 203 is smaller than that of image 17. In other words, block 203 may include a pixel array (e.g., a luminance array in the case of monochrome image 17 or a luminance or chrominance array in the case of a color image) or a three-pixel array (e.g., a luminance array and two chrominance arrays in the case of color image 17) or any other number and / or type of array depending on the color format used. The number of pixels in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Accordingly, the block may be an M×N (M columns × N rows) pixel array, or an M×N transform coefficient array, etc.

[0128] In one embodiment, Figure 2 The video encoder 20 shown is used to encode the image 17 block by block, for example, to perform encoding and prediction for each block 203.

[0129] In one embodiment, Figure 2 The video encoder 20 shown can also be used to segment and / or encode images using slices (also called video slices), where images can be segmented or encoded using one or more slices (typically non-overlapping). Each slice may include one or more blocks (e.g., coding tree units, CTUs) or one or more groups of blocks (e.g., coded tiles in the H.265 / HEVC / VVC standard and bricks in the VVC standard).

[0130] In one embodiment, Figure 2 The video encoder 20 shown can also be used to segment and / or encode an image using slice / encoding block groups (also known as video encoding block groups) and / or encoding blocks (also known as video encoding blocks), wherein the image can be segmented or encoded using one or more slice / encoding block groups (typically non-overlapping), each slice / encoding block group may include one or more blocks (e.g., CTUs) or one or more encoding blocks, wherein each encoding block may be rectangular or the like, and may include one or more complete or partial blocks (e.g., CTUs).

[0131] Residual calculation

[0132] The residual calculation unit 204 is used to calculate the residual block 205 based on the image block 203 and the prediction block 265 in the following manner (the prediction block 265 is described in detail later): for example, the residual block 205 in the pixel domain is obtained by subtracting the pixel value of the prediction block 265 from the pixel value of the image block 203 pixel by pixel.

[0133] Transformation

[0134] The transformation processing unit 206 performs discrete cosine transform (DCT) or discrete sine transform (DST) on the pixel values ​​of the residual block 205 to obtain the transformation coefficients 207 in the transform domain. The transformation coefficients 207 can also be called transformation residual coefficients, representing the residual block 205 in the transform domain.

[0135] Transform processing unit 206 can be used to apply an integer approximation of DCT / DST, such as the transform specified for H.265 / HEVC. This integer approximation is typically scaled by a certain factor compared to the orthogonal DCT transform. To maintain the norm of the residual block after both the forward and inverse transforms, other scaling factors are used as part of the transform process. These scaling factors are typically selected based on certain constraints, such as powers of 2 used for shift operations, the bit depth of the transform coefficients, and a trade-off between accuracy and implementation cost. For example, a specific scaling factor can be specified on the encoder 20 side via inverse transform processing unit 212 (and on the decoder 30 side via, for example, inverse transform processing unit 312) for the inverse transform, and correspondingly, a corresponding scaling factor can be specified on the encoder 20 side via transform processing unit 206 for the forward transform.

[0136] In one embodiment, the video encoder 20 (correspondingly, the transform processing unit 206) can be used to output transform parameters such as the type of one or more transforms, for example, directly outputting them or outputting them after being encoded or compressed by the entropy encoding unit 270, for example, so that the video decoder 30 can receive and use the transform parameters for decoding.

[0137] Quantification

[0138] Quantization unit 208 is used to quantize the transform coefficients 207 by, for example, scalar quantization or vector quantization, to obtain quantized transform coefficients 209. Quantized transform coefficients 209 can also be called quantized residual coefficients 209.

[0139] The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, n-bit transform coefficients can be rounded down to m-bit transform coefficients during quantization, where n is greater than m. The degree of quantization can be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different scales can be applied to achieve finer or coarser quantization. Smaller quantization steps correspond to finer quantization, while larger quantization steps correspond to coarser quantization. The appropriate quantization step size can be indicated by the quantization parameter (QP). For example, the quantization parameter can be an index to a predefined set of appropriate quantization steps. For example, a smaller quantization parameter can correspond to fine quantization (smaller quantization step size), a larger quantization parameter can correspond to coarse quantization (larger quantization step size), and vice versa. Quantization may include division by the quantization step size, while corresponding or inverse dequantization performed by the dequantization unit 210, etc., may include multiplication by the quantization step size. Embodiments of some HEVC standards, for example, can be used to determine the quantization step size using the quantization parameter. In general, the quantization step size can be calculated using a fixed-point approximation of an equation involving division based on the quantization parameter. Additional scaling factors can be introduced for quantization and dequantization to recover the norm of the residual block, which may have been modified by the scaling used in the fixed-point approximation of the equations used for the quantization step size and quantization parameters. In one exemplary implementation, the scaling of the inverse transform and dequantization can be combined. Alternatively, a custom quantization table can be used and indicated from the encoder to the decoder in the bitstream, etc. Quantization is a lossy operation, where the loss increases with the quantization step size.

[0140] In one embodiment, the video encoder 20 (correspondingly, the quantization unit 208) can be used to output the quantization parameter (QP), for example, directly outputting it or outputting it after being encoded or compressed by the entropy encoding unit 270, for example, so that the video decoder 30 can receive it and use it for decoding.

[0141] Inverse Quantization

[0142] The dequantization unit 210 is used to perform dequantization on the quantization coefficients by the quantization unit 208 to obtain the dequantization coefficients 211. For example, it performs a dequantization scheme based on or using the same quantization step size as the quantization unit 208 to perform the quantization scheme performed by the quantization unit 208. The dequantization coefficients 211 can also be called dequantization residual coefficients 211, corresponding to the transform coefficients 207. However, due to the loss caused by quantization, the dequantization coefficients 211 are usually not exactly the same as the transform coefficients.

[0143] Inverse Transformation

[0144] The inverse transform processing unit 212 is used to perform the inverse transform of the transform performed by the transform processing unit 206, such as the inverse discrete cosine transform (DCT) or the inverse discrete sine transform (DST), to obtain the reconstructed residual block 213 (or the corresponding dequantization coefficients 213) in the pixel domain. The reconstructed residual block 213 may also be referred to as the transform block 213.

[0145] reconstruction

[0146] The reconstruction unit 214 (e.g., summer 214) is used to add the transform block 213 (i.e., the reconstruction residual block 213) to the prediction block 265 to obtain the reconstruction block 215 in the pixel domain, for example, by adding the pixel values ​​of the reconstruction residual block 213 and the pixel values ​​of the prediction block 265.

[0147] Filtering

[0148] Loop filter unit 220 (or simply "loop filter" 220) is used to filter the reconstructed block 215 to obtain the filtered block 221, or typically to filter the reconstructed pixels to obtain filtered pixel values. For example, the loop filter unit is used to smoothly perform pixel transformations or improve video quality. Loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. For example, loop filter unit 220 may include a deblocking filter, a SAO filter, and an ALF filter. The filtering process may be performed in the order of deblocking filter, SAO filter, and ALF filter. As another example, a process called luma mapping with chromascaling (LMCS) (i.e., an adaptive in-loop shaper) may be added. This process is performed before deblocking. For example, the deblocking filtering process can also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Although the loop filter unit 220 in... Figure 2 The loop filter is shown in the diagram, but in other configurations, the loop filter unit 220 can be implemented as a post-loop filter. The filter block 221 can also be called the filter reconstruction block 221.

[0149] In one embodiment, the video encoder 20 (correspondingly, the loop filter unit 220) can be used to output loop filter parameters (e.g., SAO filter parameters, ALF filter parameters, or LMCS parameters), for example, directly outputting or outputting after entropy encoding by the entropy encoding unit 270, for example, enabling the decoder 30 to receive and decode using the same or different loop filter parameters.

[0150] Decoding image buffer

[0151] The decoded picture buffer (DPB) 230 can be a reference picture memory that stores reference picture data for use by the video encoder 20 when encoding video data. The DPB 230 can be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer 230 can be used to store one or more filter blocks 221. The decoded picture buffer 230 can also be used to store other previous filter blocks of the same current image or different images, such as previously reconstructed blocks, such as previously reconstructed and filtered blocks 221, and can provide complete previously reconstructed i.e., decoded images (and corresponding reference blocks and pixels) and / or partially reconstructed current images (and corresponding reference blocks and pixels), for example, for inter-frame prediction. The decoded image buffer 230 can also be used to store one or more unfiltered reconstruction blocks 215, or generally store unfiltered reconstruction pixels, such as reconstruction blocks 215 that have not been filtered by the loop filter unit 220, or reconstruction blocks or reconstruction pixels that have not undergone any other processing.

[0152] Pattern selection (segmentation and prediction)

[0153] The mode selection unit 260 includes a segmentation unit 262, an inter-frame prediction unit 244, and an intra-frame prediction unit 254. It receives or obtains raw image data, such as raw block 203 (the current block 203 of the current image 17) and reconstructed block data, from the decoded image buffer 230 or other buffers (e.g., a column buffer, not shown in the figure). This raw image data includes, for example, filtered and / or unfiltered reconstructed pixels or reconstructed blocks from the same (current) image and / or one or more previously decoded images. The reconstructed block data is used as reference image data for predictions such as inter-frame prediction or intra-frame prediction to obtain prediction block 265 or prediction value 265.

[0154] The mode selection unit 260 can be used to determine or select a segmentation for the prediction mode (e.g., intra-frame or inter-frame prediction mode) of the current block (including non-segmentation) to generate the corresponding prediction block 265 for calculating the residual block 205 and reconstructing the reconstructed block 215.

[0155] In one embodiment, the mode selection unit 260 can be used to select a segmentation and prediction mode (e.g., from prediction modes supported or available by the mode selection unit 260), which provides the best match or minimum residual (minimum residual refers to better compression in transmission or storage), or provides minimum signaling overhead (minimum signaling overhead refers to better compression in transmission or storage), or considers or balances both. The mode selection unit 260 can be used to determine the segmentation and prediction mode based on rate distortion optimization (RDO), i.e., selecting the prediction mode that provides minimum RDO optimization. The terms "best," "lowest," and "optimal" in this document do not necessarily refer to "best," "lowest," or "optimal" overall, but can also refer to situations that meet termination or selection criteria. For example, values ​​exceeding or falling below a threshold or other limitations may lead to a "suboptimal choice," but reduce complexity and processing time.

[0156] In other words, segmentation unit 262 can be used to segment images in a video sequence into a sequence of coding tree units (CTUs), CTUs 203 can be further segmented into smaller block portions or sub-blocks (forming blocks again), for example, by iteratively using quad-tree partitioning (QT), binary-tree partitioning (BT), or triple-tree partitioning (TT), or any combination thereof, and is used to perform prediction, for example, on each of the block portions or sub-blocks, wherein mode selection includes selecting the tree structure of the segmented block 203 and selecting the prediction mode applied to each of the block portions or sub-blocks.

[0157] The segmentation (e.g., performed by segmentation unit 262) and prediction processing (e.g., performed by inter-frame prediction unit 244 and intra-frame prediction unit 254) performed by video encoder 20 will be described in detail below.

[0158] segmentation

[0159] Segmentation unit 262 can divide (or partition) a coding tree unit 203 into smaller parts, such as small square or rectangular blocks. For an image with a three-pixel array, a CTU consists of N×N luminance pixel blocks and two corresponding chrominance pixel blocks.

[0160] The H.265 / HEVC video coding standard divides a frame of image into non-overlapping CTUs. The size of a CTU can be set to 64×64 (the size of the CTU can also be set to other values, such as 128×128 or 256×256 in the JVET reference software JEM). A 64×64 CTU contains a rectangular pixel array consisting of 64 columns, each column containing 64 pixels. Each pixel contains a luminance component and / or a chrominance component.

[0161] H.265 uses a QT-based CTU partitioning method, treating the CTU as the root node of QT. Following the QT partitioning method, the CTU is recursively divided into several leaf nodes. Each node corresponds to an image region. If a node is not partitioned, it is called a leaf node, and its corresponding image region is a CU (Cu). If a node continues to be partitioned, its corresponding image region can be divided into four regions of equal size (each half the length and width of the partitioned region). Each region corresponds to a node, and it is necessary to determine whether these nodes will be further partitioned. Whether a node is partitioned is indicated by the split_cu_flag flag in the bitstream. A node A is partitioned once to obtain four nodes Bi, i = 0~3. Bi is called a child node of A, and A is called the parent node of Bi. The QT level (qtDepth) of the root node is 0, and the QT level of a node is the four QT levels of its parent node plus 1.

[0162] In the H.265 / HEVC standard, for YUV4:2:0 format images, a CTU contains one luma block and two chroma blocks. The luma and chroma blocks can be partitioned in the same way, called a luma-chroma joint coding tree. In VVC, if the current frame is an I-frame, when a CTU is a node of a preset size (e.g., 64×64) in an intra-coded frame (I-frame), the luma block contained in that node is partitioned into a set of coding units containing only luma blocks through the luma coding tree, and the chroma block contained in that node is partitioned into a set of coding units containing only chroma blocks through the chroma coding tree; the partitioning of the luma and chroma coding trees is independent of each other. This use of independent coding trees for luma and chroma blocks is called separate trees. In H.265, a CU contains luma pixels and chroma pixels; in standards such as H.266 and AVS3, in addition to CUs containing both luma and chroma pixels, there are also luma CUs containing only luma pixels and chroma CUs containing only chroma pixels.

[0163] As described above, the video encoder 20 is used to determine or select the best or optimal prediction mode from a (predetermined) set of prediction modes. The set of prediction modes may include, for example, intra-frame prediction modes and / or inter-frame prediction modes.

[0164] Intra-frame prediction

[0165] The intra-prediction mode set can include 35 different intra-prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined in HEVC, or it can include 67 different intra-prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined in VVC. For example, several conventional angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes for non-square blocks as defined in VVC. As another example, to avoid division operations in DC prediction, only the longer side is used to calculate the average value of non-square blocks. Furthermore, the intra-prediction results of planar mode can be modified using the position-dependent intra-prediction combination (PDPC) method.

[0166] Intra-prediction unit 254 is used to generate intra-prediction block 265 using reconstructed pixels of adjacent blocks of the same current image according to the intra-prediction mode in the intra-prediction mode set.

[0167] Intra-prediction unit 254 (or typically mode selection unit 260) is also used to output intra-prediction parameters (or typically information indicating the selected intra-prediction mode of the block) to entropy coding unit 270 in the form of syntax element 266 to be included in encoded image data 21, so that video decoder 30 can perform operations such as receiving and using the prediction parameters for decoding.

[0168] Inter-frame prediction

[0169] In a possible implementation, the set of inter-frame prediction modes depends on the available reference image (i.e., at least a portion of the previously decoded image stored in the DBP230 as described above) and other inter-frame prediction parameters, such as whether to use the entire reference image or only a portion of the reference image, such as a search window region near the current block, to search for the best matching reference block, and / or, for example, whether to perform pixel interpolation of half-pixel, quarter-pixel, and / or 1 / 16th interpolation.

[0170] In addition to the prediction modes mentioned above, skip mode and / or direct mode can also be used.

[0171] For example, extended merge prediction, this mode's merge candidate list consists of five candidate types in sequence: spatial MVP from spatially adjacent CUs, temporal MVP from co-located CUs, history-based MVP from a FIFO table, pairwise average MVP, and zero MV. Decoder-side motion vector refinement (DMVR) based on bilateral matching can be used to increase the accuracy of the merge mode's MV. Mergemode with MVD (MMVD) comes from merge modes with motion vector differences. The MMVD flag is sent immediately after the skip flag and merge flag to specify whether the CU uses MMVD mode. A CU-level adaptive motion vector resolution (AMVR) scheme can be used. AMVR supports encoding CU MVD with different precisions. The MVD of the current CU is adaptively selected based on the current CU's prediction mode. When the CU is encoding in merge mode, combined inter / intra prediction (CIIP) mode can be applied to the current CU. CIIP prediction is obtained by weighted averaging of inter-frame and intra-frame prediction signals. For affine motion compensation prediction, the affine motion field of the block is described by motion information from motion vectors of 2 control points (4 parameters) or 3 control points (6 parameters). Subblock-based temporal motion vector prediction (SbTMVP) is similar to temporal motion vector prediction (TMVP) in HEVC, but predicts the motion vectors of sub-CUs within the current CU. Bidirectional optical flow (BDOF), formerly known as BIO, is a simplified version that reduces computation, particularly in terms of the number of multiplications and the size of the multipliers. In the triangular partitioning mode, the CU is uniformly divided into two triangular parts using both diagonal and anti-diagonal partitioning. Furthermore, the bidirectional prediction mode extends the simple averaging to support weighted averaging of the two prediction signals.

[0172] Inter-frame prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both in... Figure 2(Not shown in the image). The motion estimation unit can be used to receive or acquire image block 203 (current image block 203 of current image 17) and decoded image 231, or at least one or more previously reconstructed blocks, such as one or more other / different previously decoded image blocks 231, to perform motion estimation. For example, the video sequence may include the current image and the previously decoded image 231, or in other words, the current image and the previously decoded image 231 may be part of or form the image sequence that forms the video sequence.

[0173] For example, encoder 20 can be used to select a reference block from multiple reference blocks of the same or different images in multiple other images, and provide the offset (spatial offset) between the position (x, y coordinates) of the reference image (or reference image index) and / or the reference block and the position of the current block as an inter-frame prediction parameter to the motion estimation unit. This offset is also called a motion vector (MV).

[0174] The motion compensation unit is used to acquire, for example, receive, inter-frame prediction parameters, and perform inter-frame prediction based on or using these parameters to obtain inter-frame prediction blocks 246. Motion compensation performed by the motion compensation unit may include extracting or generating prediction blocks based on motion / block vectors determined by motion estimation, and may also include performing interpolation with sub-pixel precision. Interpolation filtering can generate pixels of other pixels from pixels of known pixels, thereby potentially increasing the number of candidate prediction blocks available for encoding image blocks. Once the motion vector corresponding to the PU of the current image block is received, the motion compensation unit can locate the prediction block pointed to by the motion vector in one of the reference image lists.

[0175] The motion compensation unit can also generate syntax elements associated with blocks and video slices for use by the video decoder 30 when decoding image blocks of the video slices. Alternatively, or as an alternative to slices and corresponding syntax elements, coded block groups and / or coded blocks and their corresponding syntax elements can be generated or used.

[0176] Entropy coding

[0177] Entropy coding unit 270 is used to apply entropy coding algorithms or schemes (e.g., variable length coding (VLC), context adaptive VLC (CALVC), arithmetic coding schemes, binarization algorithms, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to quantization residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements, to obtain encoded image data 21 that can be output as an encoded bitstream 21 through output terminal 272, so that video decoder 30 and the like can receive and use the parameters for decoding. The encoded bitstream 21 can be transmitted to video decoder 30, or stored in memory for later transmission or retrieval by video decoder 30.

[0178] Other architectural variations of the video encoder 20 can be used to encode the video stream. For example, a non-transform-based encoder 20 can directly quantize the residual signal in certain blocks or frames without the transform processing unit 206. In another implementation, the encoder 20 may have a quantization unit 208 and an inverse quantization unit 210 combined into a single unit.

[0179] Decoder and Decoding Method

[0180] Figure 3 This is an exemplary block diagram of a video decoder 30 according to an embodiment of this application. The video decoder 30 is used to receive encoded image data 21 (e.g., encoded bitstream 21) encoded by encoder 20, for example, to obtain a decoded image 331. The encoded image data or bitstream includes information for decoding the encoded image data, such as data representing image blocks (and / or groups or blocks of encoded video segments) and associated syntax elements.

[0181] exist Figure 3In the example, decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., a summer 314), a loop filter 320, a decoded image buffer (DBP) 330, a mode application unit 360, an inter-frame prediction unit 344, and an intra-frame prediction unit 354. The inter-frame prediction unit 344 may be or include a motion compensation unit. In some examples, video decoder 30 may perform substantially the same functions as the referenced unit. Figure 2 The video encoder 100 describes the encoding process as the opposite of the decoding process.

[0182] As described in encoder 20, the inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded image buffer DPB 230, inter-frame prediction unit 344, and intra-frame prediction unit 354 also constitute the "built-in decoder" of video encoder 20. Correspondingly, inverse quantization unit 310 can be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 can be functionally identical to inverse transform processing unit 122, reconstruction unit 314 can be functionally identical to reconstruction unit 214, loop filter 320 can be functionally identical to loop filter 220, and decoded image buffer 330 can be functionally identical to decoded image buffer 230. Therefore, the explanation of the corresponding units and functions of video encoder 20 is correspondingly applicable to the corresponding units and functions of video decoder 30.

[0183] Entropy Decoding

[0184] Entropy decoding unit 304 is used to parse bitstream 21 (or generally encoded image data 21) and perform entropy decoding on encoded image data 21 to obtain quantization coefficients 309 and / or decoded encoded parameters. Figure 3 (Not shown in the image) Examples of parameters include inter-frame prediction parameters (e.g., reference image index and motion vector), intra-frame prediction parameters (e.g., intra-frame prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 can be used to apply the decoding algorithm or scheme corresponding to the encoding scheme of the entropy coding unit 270 of the encoder 20. The entropy decoding unit 304 can also be used to provide inter-frame prediction parameters, intra-frame prediction parameters, and / or other syntax elements to the mode application unit 360, and to provide other parameters to other units of the decoder 30. The video decoder 30 can receive syntax elements at the video slice and / or video block level. Furthermore, or as an alternative to slices and corresponding syntax elements, it can receive or use coded block groups and / or coded blocks and corresponding syntax elements.

[0185] Inverse Quantization

[0186] The dequantization unit 310 can be used to receive quantization parameters (QP) (or generally information related to dequantization) and quantization coefficients from encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and dequantize the decoded quantization coefficients 309 based on the quantization parameters to obtain dequantization coefficients 311, which may also be referred to as transform coefficients 311. The dequantization process may include using the quantization parameters calculated by the video encoder 20 for each video block in the video slice to determine the degree of quantization, and also to determine the degree of dequantization to be performed.

[0187] Inverse Transformation

[0188] The inverse transform processing unit 312 can be used to receive the dequantized coefficients 311, also known as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain the reconstructed residual block 213 in the pixel domain. The reconstructed residual block 213 can also be called transform block 313. The transform can be an inverse transform, such as inverse DCT, inverse DST, inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 can also be used to receive transform parameters or corresponding information from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304) to determine the transform applied to the dequantized coefficients 311.

[0189] reconstruction

[0190] The reconstruction unit 314 (e.g., summer 314) is used to add the reconstruction residual block 313 to the prediction block 365 to obtain the reconstruction block 315 in the pixel domain, for example, by adding the pixel values ​​of the reconstruction residual block 313 and the pixel values ​​of the prediction block 365.

[0191] Filtering

[0192] Loop filter unit 320 (in or after the encoding loop) is used to filter the reconstructed block 315 to obtain filtered block 321, thereby facilitating pixel transformation or improving video quality. Loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. For example, loop filter unit 320 may include a deblocking filter, a SAO filter, and an ALF filter. The filtering process may be performed in the order of deblocking filter, SAO filter, and ALF filter. As another example, a process called luma mapping with chromascaling (LMCS) (i.e., an adaptive in-loop shaper) may be added. This process is performed before deblocking. For example, the deblocking filtering process can also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Although the loop filter unit 320 in... Figure 3 The loop filter is shown in the diagram, but in other configurations, the loop filter unit 320 can be implemented as a post-loop filter.

[0193] Decoding image buffer

[0194] The decoded video block 321 in one image is then stored in the decoded image buffer 330, which stores the decoded image 331 as a reference image. The reference image is used for subsequent motion compensation for other images and / or output displays respectively.

[0195] The decoder 30 is used to output the decoded image 311 through the output terminal 312, etc., for display to the user or for the user to view.

[0196] predict

[0197] Inter-frame prediction unit 344 is functionally identical to inter-frame prediction unit 244 (especially motion compensation unit), and intra-frame prediction unit 354 is functionally identical to inter-frame prediction unit 254. It determines segmentation or partitioning and performs prediction based on segmentation and / or prediction parameters or corresponding information received from coded image data 21 (e.g., parsed and / or decoded by entropy decoding unit 304). Mode application unit 360 can be used to perform prediction (intra-frame or inter-frame prediction) for each block based on the reconstructed block, block, or corresponding pixel (filtered or unfiltered), resulting in prediction block 365.

[0198] When a video slice is encoded as an intra-coded (I) slice, the intra-prediction unit 354 in the mode application unit 360 generates a prediction block 365 for the current video slice based on the indicated intra-prediction mode and data from the previous decoded block of the current image. When a video image is encoded as an inter-coded (i.e., B or P) slice, the inter-prediction unit 344 (e.g., a motion compensation unit) in the mode application unit 360 generates a prediction block 365 for the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit 304. For inter-prediction, these prediction blocks can be generated from one of the reference images in one of the reference image lists. The video decoder 30 can construct reference frame lists 0 and 1 using the default construction technique based on the reference images stored in the DPB 330. In addition to slices (e.g., video slices) or as a substitute for slices, the same or similar processes can be applied to embodiments of coding block groups (e.g., video coding block groups) and / or coding blocks (e.g., video coding blocks), such as video can be encoded using I, P, or B coding block groups and / or coding blocks.

[0199] The pattern application unit 360 is used to determine prediction information for video blocks in the current video slice by parsing motion vectors and other syntax elements, and to generate prediction blocks for the current video slice being decoded using the prediction information. For example, the pattern application unit 360 uses some received syntax elements to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction), inter-frame prediction slice type (e.g., B-slice, P-slice, or GPB-slice), construction information for one or more reference image lists for the slice, motion vectors for each inter-frame coded video block in the slice, inter-frame prediction state for each inter-frame coded video block in the slice, and other information to decode video blocks within the current video slice. In addition to slices (e.g., video slices) or as alternatives to slices, the same or similar process can be applied to embodiments of coding block groups (e.g., video coding block groups) and / or coding blocks (e.g., video coding blocks), for example, where video can be encoded using I, P, or B coding block groups and / or coding blocks.

[0200] In one embodiment, Figure 3 The video encoder 30 shown can also be used to segment and / or decode images using slices (also called video slices), where images can be segmented or decoded using one or more slices (typically non-overlapping). Each slice may include one or more blocks (e.g., CTUs) or one or more groups of blocks (e.g., coded blocks in the H.265 / HEVC / VVC standard and bricks in the VVC standard).

[0201] In one embodiment, Figure 3The video decoder 30 shown can also be used to segment and / or decode an image using slice / coded block groups (also known as video coded block groups) and / or coded blocks (also known as video coded blocks), wherein the image can be segmented or decoded using one or more slice / coded block groups (typically non-overlapping), each slice / coded block group may include one or more blocks (e.g., CTUs) or one or more coded blocks, wherein each coded block may be rectangular or the like, and may include one or more complete or partial blocks (e.g., CTUs).

[0202] Other variations of the video decoder 30 can be used to decode the encoded image data 21. For example, the decoder 30 can generate an output video stream without the loop filter unit 320. For example, the non-transform-based decoder 30 can directly dequantize the residual signal in certain blocks or frames without the inverse transform processing unit 312. In another implementation, the video decoder 30 may have a dequantization unit 310 and an inverse transform processing unit 312 combined into a single unit.

[0203] It should be understood that in encoder 20 and decoder 30, the processing result of the current step can be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations can be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering, such as clipping or shifting operations.

[0204] It should be noted that further calculations can be performed on the derived motion vector of the current block (including but not limited to control point motion vectors in affine mode, affine, planar, sub-block motion vectors in ATMVP mode, time motion vectors, etc.). For example, the value of the motion vector can be restricted to a predefined range based on the representation bits of the motion vector. If the representation bits of the motion vector are bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" represents exponentiation. For example, if bitDepth is set to 16, the range is -32768 to 32767; if bitDepth is set to 18, the range is -131072 to 131071. For example, the value of the derived motion vector (e.g., the MV of four 4×4 sub-blocks in an 8×8 block) is restricted such that the maximum difference between the integer parts of the MV of the four 4×4 sub-blocks does not exceed N pixels, for example, not more than 1 pixel. Two methods for restricting motion vectors based on bitDepth are provided here.

[0205] Although the above embodiments primarily describe video encoding and decoding, it should be noted that embodiments of the decoding system 10, encoder 20, and decoder 30, as well as other embodiments described herein, can also be used for still image processing or encoding and decoding, i.e., the processing or encoding and decoding of a single image independent of any previous or consecutive images in video encoding and decoding. Generally, if image processing is limited to a single image 17, the inter-frame prediction unit 244 (encoder) and inter-frame prediction unit 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and video decoder 30 can also be used for still image processing, such as residual calculation 204 / 304, transform 206, quantization 208, inverse quantization 210 / 310, (inverse) transform 212 / 312, segmentation 262 / 362, intra-frame prediction 254 / 354 and / or loop filtering 220 / 320, entropy coding 270, and entropy decoding 304.

[0206] Figure 4 This is an exemplary block diagram of a video decoding device 400 according to an embodiment of this application. The video decoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video decoding device 400 may be a decoder, such as... Figure 1A The video decoder 30 in the text can also be an encoder, for example... Figure 1A The video encoder 20 in the middle.

[0207] The video decoding device 400 includes: an input port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit (CPU) 430 for processing data; a transmitter unit (Tx) 440 and an output port 450 (or output port 450) for transmitting data; and a memory 460 for storing data. The video decoding device 400 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the input port 410, receiver unit 420, transmitter unit 440, and output port 450 for the entry or exit of optical or electrical signals.

[0208] Processor 430 is implemented in both hardware and software. Processor 430 may be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 430 communicates with ingress port 410, receiver unit 420, transmitter unit 440, egress port 450, and memory 460. Processor 430 includes a decoding module 470. Decoding module 470 implements the embodiments disclosed above. For example, decoding module 470 performs, processes, prepares, or provides various encoding operations. Therefore, decoding module 470 provides a substantial improvement to the functionality of video decoding device 400 and affects the switching of video decoding device 400 to different states. Alternatively, decoding module 470 may be implemented with instructions stored in memory 460 and executed by processor 430.

[0209] Memory 460 includes one or more disks, tape drives, and solid-state drives, which can be used as overflow data storage devices to store such programs when an executable program is selected, and to store instructions and data read during program execution. Memory 460 can be volatile and / or non-volatile, and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0210] Scalable video coding, also known as scalable video coding, is an extension of current video coding standards (generally, it's an extension of Advanced Video Coding (AVC) (H.264) called Scalable Video Coding (SVC), or an extension of High Efficiency Video Coding (HEVC) (H.265) called Scalable High Efficiency Video Coding (SHVC)). Scalable video coding was primarily developed to address packet loss and latency issues caused by real-time changes in network bandwidth during real-time video transmission.

[0211] The basic structure of scalable video coding can be called a hierarchy. Scalable video coding technology can obtain bitstreams of different resolution levels by spatially scaling the original image patches (resolution scaling). Resolution can refer to the size of the image patch in pixels; lower-level patches have lower resolutions, while higher-level patches have resolutions no lower than lower-level patches. Alternatively, by temporally scaling the original image patches (frame rate scaling), different frame rate levels can be obtained. Frame rate can refer to the number of image frames contained in the video per unit time; lower-level patches have lower frame rates, while higher-level patches have frame rates no lower than lower-level patches. Alternatively, by quality-domain scaling the original image patches, different quality levels can be obtained. Code quality refers to the video quality; lower-level patches have higher levels of image distortion, while higher-level patches have no higher levels of image distortion than lower-level patches.

[0212] Typically, the layer called the base layer is the lowest level in scalable video coding. In spatial scalability, base layer image patches are encoded using the lowest resolution; in temporal scalability, they are encoded using the lowest frame rate; and in quality scalability, they are encoded using the highest QP or the lowest bitrate. In other words, the base layer is the lowest quality layer in scalable video coding. The layers called enhancement layers are those above the base layer in scalable video coding, and can be divided into multiple enhancement layers from low to high. The lowest enhancement layer encodes a merged bitstream based on the coding information obtained from the base layer, and its coding resolution, frame rate, or bitrate is higher than that of the base layer. Higher enhancement layers can encode higher quality image patches based on the coding information of lower enhancement layers.

[0213] For example, Figure 5 This is an exemplary hierarchical diagram illustrating the scalable video coding of this application, such as... Figure 5As shown, after the original image blocks are fed into the scalable encoder, they can be layered into basic layer image blocks B and enhancement layer image blocks (E1~En, n≥1) according to different encoding configurations. These are then encoded separately to obtain bitstreams containing both basic and enhancement layer bitstreams. The basic layer bitstream is generally obtained by using the lowest resolution, lowest frame rate, or lowest coding quality parameters for the image blocks. The enhancement layer bitstream is obtained by superimposing high resolution, high frame rate, or high coding quality parameters on the image blocks, based on the basic layer. As the number of enhancement layers increases, the spatial, temporal, or quality levels of the encoding also increase. When the encoder transmits the bitstream to the decoder, it prioritizes the transmission of the basic layer bitstream, and gradually transmits higher-level bitstreams as the network has capacity. The decoder first receives and decodes the base layer bitstream. Then, based on the received enhancement layer bitstream, it decodes the spatial, temporal, or quality-level bitstreams layer by layer in order from low to high levels. The decoded information of the higher levels is then superimposed on the reconstructed blocks of the lower levels to obtain reconstructed blocks with higher resolution, higher frame rate, or higher quality.

[0214] As mentioned above, each image in a video sequence is typically segmented into a set of non-overlapping blocks, usually encoded at the block level. In other words, the encoder typically processes and encodes video at the block (image block) level. For example, it generates prediction blocks through spatial (intra-frame) prediction and temporal (inter-frame) prediction; subtracts the prediction blocks from the image blocks (the currently processed / to-be-processed blocks) to obtain residual blocks; transforming and quantizing the residual blocks in the transform domain can reduce the amount of data to be transmitted (compressed). The encoder also needs to perform inverse quantization and inverse transform to obtain reconstructed residual blocks, and then add the pixel values ​​of the reconstructed residual blocks to the pixel values ​​of the prediction blocks to obtain the reconstructed blocks. The reconstructed blocks of the base layer refer to the reconstructed blocks obtained by performing the above operations on the base layer image blocks obtained by layering the original image blocks. For example, Figure 6 An exemplary flowchart of the encoding method for the enhancement layer of this application is shown below. Figure 6 As shown, the encoder obtains the prediction block of the base layer based on the original image block (e.g., LCU). Then, it calculates the difference between corresponding pixels in the original image block and the prediction block of the base layer to obtain the residual block of the base layer. After dividing the residual block of the base layer, it performs transformation and quantization, and together with the base layer encoded control information, prediction information, motion information, etc., it performs entropy coding to obtain the bitstream of the base layer. The encoder performs inverse quantization and inverse transformation on the quantized coefficients to obtain the reconstructed residual block of the base layer. Then, it sums the corresponding pixels in the prediction block and the reconstructed residual block of the base layer to obtain the reconstructed block of the base layer.

[0215] The image mentioned below (e.g., the current image, the previous frame, etc.) can refer to the largest coding unit (LCU) in the entire frame, or the entire frame, or the region of interest (ROI) in the entire frame, that is, a specific image region that needs to be processed in the image, or a slice of an image.

[0216] Based on the above description, this application provides an image encoding and decoding method to solve the problem that the scalable encoding and decoding technology is greatly affected by changes in channel conditions, which can lead to screen distortion.

[0217] Figure 7 This is an exemplary flowchart of the image encoding method of this application. Process 700 can be performed by video encoder 20 (or encoder). Process 700 is described as a series of steps or operations, and it should be understood that process 700 can be performed in various orders and / or occur simultaneously, and is not limited to... Figure 7 The execution order is shown. Process 700 includes the following steps:

[0218] Step 701: Determine the reference frame number of the current image based on the channel feedback information.

[0219] Optionally, the channel feedback information comes from the corresponding decoding end and / or network devices on the transmission link. The encoding end can send the bitstream to one or more decoding ends, enabling the receiving decoding end to parse the bitstream and reconstruct the image frame. The bitstream from the encoding end to the decoding end may include not only the transmitting (encoding end) and receiving (decoding end) devices, but also network devices on the transmission link between them, such as switches, repeaters, base stations, hubs, routers, firewalls, bridges, gateways, network interface cards (NICs), printers, modems, fiber optic transceivers, and optical cables. To inform the encoding end about the status of the transmission link, the decoding end can send channel feedback information to the encoding end. Similarly, this channel feedback information can also be sent to the encoding end by network devices on the transmission link. This application does not specifically limit the sender of the channel feedback information.

[0220] The decoding end and / or network devices on the transmission link can parse the received bitstream to determine the frame number and layer number corresponding to the bitstream, and then send channel feedback information to the encoding end, carrying the frame number and layer number corresponding to the aforementioned bitstream, to inform the encoding end which frame and which layer it has received.

[0221] In this application, the decoding end can periodically send channel feedback information to the encoding end, carrying the frame number and layer number corresponding to the latest received bitstream; alternatively, it can send channel feedback information when it parses the bitstream to start receiving the next frame (based on the frame number in the bitstream), carrying the frame number of the previous frame and the highest layer number received in the previous frame; or, it can send channel feedback information when it parses the bitstream to finish receiving the current image (based on the highest layer number of the current image in the bitstream), carrying the frame number of the current image and the highest layer number received in the current image. Furthermore, the decoding end can send channel feedback information in other ways, without specific limitations.

[0222] Optionally, channel feedback information is generated based on the transmitted bitstream. In a reliable transmission channel mode, the encoder does not need to wait for feedback from the decoder or network devices on the transmission link when transmitting the bitstream. It can know the channel condition based on the signal transmission and reception status of the communication interface. Therefore, the encoder can adjust the transmission of the bitstream according to the known channel condition, assuming that the transmitted bitstream will definitely be received by the decoder. Thus, the encoder can generate channel feedback information based on the frame number and layer number corresponding to the transmitted bitstream. Similarly, the encoder can generate channel feedback information periodically, or it can generate it after each transmitted image frame; there is no specific limitation on this.

[0223] Therefore, channel feedback information is used to indicate the information of the image frame received by the decoding end. For example, if the current image frame has a total of 4 layers, the encoding end encodes the current image frame to obtain the bitstream corresponding to the 4th layer. However, during transmission, the decoding end only receives the bitstream corresponding to the 3rd layer of the current image frame. When encoding the next frame, the encoding end uses the reconstructed image of the 4th layer of the current image as the reference image. Since the decoding end has not received the bitstream corresponding to the 4th layer, it cannot use the reconstructed image of the 4th layer of the current image as the reference image to decode the next frame, resulting in decoding errors in the next frame. Therefore, in this application, the encoding end first obtains the channel feedback information, determines the information of the image frame received by the decoding end based on this information, including the frame number and layer number of the image frame received by the decoding end, and then determines the reference frame for the next frame based on this information. This avoids the situation described in the example above and ensures that the reference image used by the encoding end and the decoding end is consistent.

[0224] In one possible implementation, when there is only one decoding end, the encoding end can first obtain multiple channel feedback information from the decoding end, and then determine the frame number that is closest to the current image among the multiple frame numbers indicated by the multiple channel feedback information as the reference frame number of the current image.

[0225] As described above, the channel feedback information indicates information about the image frames received by the decoder, including the frame number and layer number of the image frames received by the decoder. The encoder acquires multiple channel feedback information messages, which reflect the frame number and layer number of the bitstream received by the decoder at different times. Therefore, the frame number closest to the current image among the multiple frame numbers indicated by the aforementioned channel feedback information messages can be determined as the reference frame number of the current image. For example, if the encoder acquires three channel feedback information messages, one indicating frame number 1, another indicating frame number 2, and yet another indicating frame number 1, and the current image is frame 3, the encoder can determine that the reference frame number of the current image is 2.

[0226] In one possible implementation, when there are multiple decoding ends, the encoding end can first acquire multiple sets of channel feedback information, which correspond to multiple decoding ends. Each set of channel feedback information includes multiple channel feedback information. Then, based on the multiple sets of channel feedback information, one or more common frame numbers are determined. A common frame number is a frame number indicated by at least one channel feedback information in each set of channel feedback information. Finally, based on the one or more common frame numbers, the reference frame number of the current image is determined.

[0227] As described above, the channel feedback information indicates the information of the image frames received by the decoding end, including the frame number and layer number of the image frames received by the decoding end. For each decoding end, the encoding end obtains multiple channel feedback information corresponding to that decoding end. The encoding end can determine the common frame number based on the multiple sets of channel feedback information corresponding to multiple decoding ends. The common frame number is the frame number indicated by at least one channel feedback information in each set of channel feedback information; that is, the common frame number is the frame number indicated in a set of channel feedback information fed back by each decoding end. If there is only one common frame number, the encoding end can determine this common frame number as the reference frame number of the current image; if there are multiple common frame numbers, the largest of the multiple common frame numbers can be determined as the reference frame number of the current image. For example, decoding end A corresponds to 3 channel feedback information, indicating frame numbers 1, 2, 3, and 4 respectively; decoding end B corresponds to 3 channel feedback information, indicating frame numbers 2, 3, 4, and 5 respectively; decoding end C corresponds to 3 channel feedback information, indicating frame numbers 2, 3, 4, and 6 respectively. We can determine that there are three frame numbers: 2, 3, and 4, with the largest being 4. Therefore, the reference frame number for the current image is 4.

[0228] Step 702: Obtain the first reference layer number set of the first image frame corresponding to the reference frame number. The first reference layer number set includes N1 layer numbers of the layer.

[0229] The first image frame corresponding to the reference frame number is the image frame indicated by the reference frame number determined in step 701. In this application, the maximum number of layers L in the video image frames during scalable coding can be preset.max For example, 6. This maximum number of layers can be a threshold, meaning that the number of layers in each image frame does not exceed this limit. However, in actual encoding, different image frames may have different total number of layers. The total number of layers in the first image frame is denoted by L1, and L1 can be less than or equal to the aforementioned maximum number of layers L. max After the first image frame is divided into L1 layers, N1 layers can be used as reference images for subsequent image frames, where 1 ≤ N1 < L1. The layer numbers of these N1 layers form the first reference layer number set for the first image frame. That is, the first image frame corresponds to the first reference layer number set, and only the reconstructed images of layers whose layer numbers are in the first reference layer number set can be used as reference images for subsequent image frames. The same applies to other image frames in the video, which will not be elaborated here.

[0230] In this application, the N1 layers can be preset, for example, the first reference layer number set Rx = {1,4,6}, L1 = 6, that is, the encoding end divides the image frame into 6 layers when encoding the image frame, and the layer number of the image frame that can be used as the reference image for subsequent image frames is 1, 4, 6.

[0231] This application does not impose specific limitations on the value of N1. For example, N1 can be set according to chip capabilities or dynamically set according to real-time encoding and network feedback. For instance, N1 can not exceed half of the total number of layers, and the layer number interval can be 2. For example, if L1 = 6, then the first reference layer number set Rx = {1, 3, 5}; or, N1 layers can include the highest layer number. For example, if L1 = 6, the first reference layer number set Rx = {1, 3, 6}.

[0232] It should be understood that each frame in a video can have an independent and different set of reference layer numbers; or, the image frames in a video can be divided into multiple groups, and a group of image frames can have the same set of reference layer numbers; or, all image frames in a video can have the same set of reference layer numbers. This application does not impose any specific limitations on this.

[0233] Step 703: Determine the reference layer number of the current image based on the channel feedback information and the first reference layer number set.

[0234] In one possible implementation, when there is only one decoding end, the encoding end can determine the highest layer number indicated by the channel feedback information indicating the reference frame number as the target layer number.

[0235] As mentioned above, the reference frame number is determined based on the channel feedback information. The encoder can further determine the target layer number based on the reference layer number indicated by the channel feedback information that indicates the corresponding reference frame number. When there is only one channel feedback information indicating the reference frame number, the layer number indicated by that channel feedback information is the highest layer number and is directly determined as the target layer number. When there are multiple channel feedback information indicating the reference frame number, the largest of the layer numbers indicated by each of the multiple channel feedback information is determined as the target layer number. For example, if the encoder obtains three channel feedback information all indicating frame number 2, and 2 is the reference frame number of the current image, and one channel feedback information indicates layer number 3, another indicates layer number 4, and the third indicates layer number 5, then to improve image quality, layer number 5 can be determined as the target layer number.

[0236] When the reference layer number set includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, when the reference layer number set does not include the target layer number, the layer number in the reference layer number set that is less than and closest to the target layer number is determined as the reference layer number of the current image.

[0237] In step 701, the reference frame number of the current image is determined, which indicates that the reference image of the current image comes from the image frame corresponding to that reference frame number. In step 702, the set of reference layer numbers of the image frame corresponding to the reference frame number is determined, which indicates that the reference image of the current image is one of the reconstructed images corresponding to the N1 layers included in the set of reference layer numbers of the image frame corresponding to the reference frame number.

[0238] Therefore, after determining the target layer number, we can first check whether the target layer number belongs to the reference layer number set of the image frame corresponding to the reference frame number. If the reference layer number set includes the target layer number, then the target layer number can be determined as the reference layer number of the current image. If the reference layer number set does not include the target layer number, then we need to find the layer number in the reference layer number set that is smaller than the target layer number and closest to the target layer number, and determine that layer number as the reference layer number of the current image. For example, if the reference layer number set Rx = {1, 3, 5}, the target layer number is 4, and the layer number in the reference layer number set Rx that is smaller than and closest to the target layer number is 3, then 3 is the reference layer number of the current image.

[0239] In one possible implementation, when there are multiple decoding ends, the encoding end can obtain the highest layer number indicated by the channel feedback information indicating the reference frame number in each of the multiple sets of channel feedback information, and determine the smallest of the multiple highest layer numbers as the target layer number.

[0240] As mentioned above, the reference frame number is determined based on the channel feedback information. The encoder can further determine the target layer number based on the channel feedback information indicating the reference frame number. Since the reference frame number is first and foremost a common frame number indicated by multiple sets of channel feedback information corresponding to multiple decoders, it is possible to obtain at least one channel feedback information indicating the reference frame number corresponding to each decoder. The maximum layer number indicated by the channel feedback information indicating the reference frame number of each decoder can be determined, and then the minimum value is taken from the highest layer number corresponding to each decoder as the target layer number. For example, if the reference frame number is 2, and the layer number indicated by the channel feedback information of the decoding end A for reference frame number 2 includes 1, 3, and 4, then the highest layer number 4 is taken for reference frame number 2; if the layer number indicated by the channel feedback information of the decoding end B for reference frame number 2 includes 1, 3, and 6, then the highest layer number 6 is taken for reference frame number 2; if the layer number indicated by the channel feedback information of the decoding end C for reference frame number 2 includes 3, 4, and 6, then the highest layer number 6 is taken for reference frame number 2; then the minimum value of these highest layer numbers is taken, thus the target layer number can be determined to be 4.

[0241] Similarly, when the reference layer number set includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, when the reference layer number set does not include the target layer number, the layer number in the reference layer number set that is less than and closest to the target layer number is determined as the reference layer number of the current image.

[0242] In step 701, the reference frame number of the current image is determined, which indicates that the reference image of the current image comes from the image frame corresponding to that reference frame number. In step 702, the set of reference layer numbers of the image frame corresponding to the reference frame number is determined, which indicates that the reference image of the current image is one of the reconstructed images corresponding to the N1 layers included in the set of reference layer numbers of the image frame corresponding to the reference frame number.

[0243] Therefore, after determining the target layer number, we can first check whether the target layer number belongs to the reference layer number set of the image frame corresponding to the reference frame number. If the reference layer number set includes the target layer number, then the target layer number can be determined as the reference layer number of the current image. If the reference layer number set does not include the target layer number, then we need to find the layer number in the reference layer number set that is smaller than the target layer number and closest to the target layer number, and determine that layer number as the reference layer number of the current image. For example, if the reference layer number set Rx = {1, 3, 5}, the target layer number is 6, and the layer number in the reference layer number set Rx that is smaller than and closest to the target layer number is 5, then 5 is the reference layer number of the current image.

[0244] Step 704: Perform scalable video encoding on the current image based on the reference frame number and reference layer number to obtain the bitstream.

[0245] After determining the reference frame number and reference layer number of the current image, the reconstructed image corresponding to the reference frame number and reference layer number can be extracted from the decoded picture buffer (DPB) as the reference image of the current image. Based on the reference image, the current image can be scalably encoded to obtain the bitstream.

[0246] In this application, in addition to the bitstream obtained by the hierarchical encoding described above, the encoding end can also carry the reference layer number set of the image frames in the bitstream.

[0247] This application determines the reference frame number of the current image based on channel feedback information, and then determines the reference layer number of the current image based on the reference frame number and a pre-set set of reference layer numbers. The set of reference layer numbers includes the layer numbers of N layers of the image frame corresponding to the reference frame number. Then, the reference image of the current image is obtained based on the reference frame number and the reference layer number. The reference image obtained in this way fully considers the changes in the channel, ensures that the reference image used by the encoding end and the decoding end is consistent, improves encoding efficiency, and avoids screen distortion.

[0248] In one possible implementation, the DPB contains only N1 layered reconstructed images for the image frame corresponding to the reference frame number.

[0249] In related technologies, for an image frame corresponding to a reference frame number, after performing scalable encoding on the image frame, the encoding end needs to store the reconstructed images of all layers into the DPB. For example, if L1=6 for the image frame corresponding to the reference frame number, the encoding end needs to store 6 layers of reconstructed images in the DPB. However, in this application, only the reconstructed images corresponding to the N1 layers included in the reference layer set of the image frame corresponding to the reference frame number need to be stored. For example, if L1=6 for the image frame corresponding to the reference frame number, and its reference layer set Rx={1,3,5}, after encoding the image frame, the encoding end only needs to store the reconstructed images with layer numbers 1, 3, and 5 into the DPB. Compared with related technologies, this application reduces the number of reconstructed images stored in the DPB, lowers the write bandwidth, improves the encoding processing speed, and saves DPB space.

[0250] In one possible implementation, the encoder can first determine the frame number closest to the current image among multiple frame numbers indicated by multiple channel feedback information as the target frame number. It then determines whether the highest layer number indicated by the channel feedback information pointing to the target frame number is greater than or equal to the highest layer number in a second set of reference layer numbers, which is the set of reference layer numbers for the second image frame corresponding to the target frame number. When the aforementioned condition (i.e., greater than or equal to) is met, the target frame number is then determined as the reference frame number.

[0251] For example, the encoder determines the reference frame number of the 4th frame as 3 and the reference layer number as 4 based on the channel feedback information. However, the decoder actually receives the 6th layer of the image frame with frame number 3. When decoding the aforementioned 4th frame, the decoder may determine its reference frame number as 3 and its reference layer number as 6. In this case, the encoder and decoder use different reference images for the "4th frame", resulting in inconsistency between encoding and decoding, and thus decoding errors.

[0252] To address the aforementioned issues, this application provides the above-described solution. Instead of directly determining the frame number obtained based on the conditions as the reference frame number, the encoding end uses it as the target frame number. Based on the channel feedback information indicating the target frame number, it determines whether the decoding end has received an image layer with a layer number greater than or equal to the highest layer number in the second reference layer number set. If the highest layer number indicated by the channel feedback information already satisfies the condition of being greater than or equal to the highest layer number in the second reference layer number set, even if the decoding end receives a higher layer than the second image frame, according to the scheme for determining the reference layer number in step 703, the highest layer number in the second reference layer number set will still be selected as the reference layer number. Therefore, the target frame number can be directly used as the reference frame number, and the problem of inconsistent reference layer numbers selected by the encoding and decoding ends will not occur.

[0253] If the condition of being greater than or equal to is not met, the encoder can determine the specified frame number from the multiple frame numbers indicated by the multiple channel feedback messages as the reference frame number of the current image. For example, if the encoder receives three channel feedback messages, one indicating frame number 1, another indicating frame number 2, and yet another indicating frame number 1, and the current image is frame 3, then the encoder determines the target frame number as 2. However, if the channel feedback message indicates that the decoder has received two frames at layer 4, which is less than the highest layer number 6 in the reference layer number set Rx = {1, 4, 6} for two frames, then the encoder can use the specified frame number 1 as the reference frame number.

[0254] In this application, the specified frame number can be a frame number that is 2 frames earlier than the current image, or it can be a fixed frame number; there is no specific limitation on this.

[0255] In one possible implementation, a specified frame number and a specified layer number are pre-defined. The encoding end can determine the specified frame number as the reference frame number of the current image. When the specified layer number is included in the set of reference layer numbers, the specified layer number is determined as the reference layer number of the current image; or, when the specified layer number is not included in the set of reference layer numbers, the layer number in the set of reference layer numbers that is less than and closest to the specified layer number is determined as the reference layer number of the current image.

[0256] That is, the encoding end can directly specify the reference frame number and reference layer number of the current image, which can improve the efficiency of determining the reference image.

[0257] In one possible implementation, when the reference layer number set does not include the target layer number, if the reference layer number set does not include a layer number less than the target layer number, then the reference frame number of the previous frame of the current image is determined as the reference frame number of the current image, and the reference layer number of the previous frame is determined as the reference layer number of the current image.

[0258] For example, if the reference layer number set Rx = {3,5} and the target layer number is 2, and there is no layer number less than 2 in the reference layer number set Rx, then the reference frame number and reference layer number determined in the previous frame can be directly used for the current image.

[0259] Optionally, the encoder can set the upper layer number to 1 in each reference layer number set, so that there is no situation where the reference layer number set does not include a layer number less than the target layer number, thus determining that the reference layer number of the current image is 1.

[0260] In one possible implementation, when there is only one decoding end, the encoding end can carry the reference frame number determined in step 701 in the bitstream, and the decoding end can use the logic in step 703 to determine the reference layer number. Alternatively, the encoding end can also carry the reference frame number determined in step 701 and the reference layer number determined in step 703 in the bitstream, so that the decoding end can directly obtain the reference frame number and reference layer number by parsing the bitstream.

[0261] In one possible implementation, when there are multiple decoding ends, the encoding end can also carry the reference frame number determined in step 701 and the reference layer number determined in step 703 in the bitstream, so that the decoding end can directly obtain the reference frame number and the reference layer number by parsing the bitstream.

[0262] In one possible implementation, when the current image is an image fragment, the channel feedback information includes the image fragment number of the image frame received by the decoding end and the layer number corresponding to the image fragment number; determining the reference layer number of the current image based on the channel feedback information and the first reference layer number set includes: if the image fragment number of the current image is the same as the image fragment number of the image frame received by the decoding end, determining the layer number corresponding to the image fragment number of the image frame received by the decoding end as the target layer number; if the first reference layer number set includes the target layer number, determining the target layer number as the reference layer number of the current image; or, if the first reference layer number set does not include the target layer number, determining the layer number in the first reference layer number set that is less than and closest to the target layer number as the reference layer number of the current image.

[0263] An image frame can be divided into multiple image slices for encoding and transmission. Therefore, when the current image is an image slice, in addition to the frame number of the image frame received by the decoding end, the channel feedback information also includes the image slice number of the image frame received by the decoding end and the layer number corresponding to that image slice number. Thus, the encoding end can first determine the reference frame number of the current image using the method described above, and then determine the reference layer number of the current image based on the image slice number of the current image and the first set of reference layer numbers. That is, it finds the image slice number that is the same as the current image among the multiple image slice numbers indicated by the channel feedback information indicating the reference frame number, and then determines the layer number corresponding to this identical image slice number as the target layer number. Based on the target layer number, the reference layer number of the current image is determined from the first set of reference layer numbers. For example, if the reference frame number of the current image is 1, and the channel feedback information indicating frame number 1 indicates multiple image slice numbers including 1, 2, 3, and 4, the layer number corresponding to image slice number 1 is 3, the layer number corresponding to image slice number 2 is 4, the layer number corresponding to image slice number 3 is 5, the layer number corresponding to image slice number 4 is 6, and the first reference layer number set Rx = {1, 3, 5}. If the image slice number of the current image is 1, then the reference layer number of the current image is obtained based on the aforementioned layer number 3 corresponding to image slice number 1 and the first reference layer number set, and its reference layer number is 3. Alternatively, if the image slice number of the current image is 2, then the reference layer number of the current image is obtained based on the aforementioned layer number 4 corresponding to image slice number 2 and the first reference layer number set, and its reference layer number is 3. Or, if the image slice number of the current image is 3, then the reference layer number of the current image is obtained based on the aforementioned layer number 5 corresponding to image slice number 3 and the first reference layer number set, and its reference layer number is 5. Alternatively, if the image slice number of the current image is 4, then the reference layer number of the current image is obtained based on the layer number 6 corresponding to the aforementioned image slice number 4 and the first reference layer number set, and its reference layer number is 5.

[0264] Figure 8 This is an exemplary flowchart of the image decoding method of this application. Process 800 may be performed by video decoder 30 (or a decoder). Process 800 is described as a series of steps or operations, and it should be understood that process 800 may be performed in various orders and / or occur simultaneously, and is not limited to... Figure 8 The execution order is shown. Process 800 includes the following steps:

[0265] Step 801: Obtain the bitstream.

[0266] The decoding end can obtain the bit stream through the transmission link between it and the encoding end.

[0267] Step 802: Parse the bitstream to obtain the reference frame number of the current image.

[0268] refer to Figure 7In the embodiment shown, the encoding end carries the reference frame number of the image frame in the video in the bitstream, so the decoding end can determine the reference frame number of the current image by parsing the bitstream.

[0269] Step 803: Obtain the set of third reference layer numbers of the third image frame corresponding to the reference frame number.

[0270] The third reference layer number set includes N² layer numbers, where 1 ≤ N² < L², and L² represents the total number of layers in the third image frame. The decoding end can parse the bitstream to obtain the reference layer number set of the image frames in the video. For a description of the reference layer number set, please refer to [link to documentation / reference]. Figure 7 Step 702 of the illustrated embodiment will not be repeated here.

[0271] Step 804: Determine the reference layer number of the current image based on the third reference layer number set.

[0272] If the bitstream does not carry the reference layer number of the image frame, the decoder can determine the highest layer number among the layer numbers of the multiple reconstructed images of the decoded third image frame. If the set of third reference layer numbers for the third image frame includes the highest layer number, that highest layer number is determined as the reference layer number of the current image; or, if the set of third reference layer numbers for the third image frame does not include the highest layer number, the layer number in the set of third reference layer numbers that is less than and closest to the highest layer number is determined as the reference layer number of the current image. If the set of third reference layer numbers does not include a layer number less than the highest layer number, then the reference frame number of the previous frame is determined as the reference frame number of the current image, and the reference layer number of the previous frame is determined as the reference layer number of the current image.

[0273] This can be referenced. Figure 7 Step 703 of the illustrated embodiment will not be described again here.

[0274] Step 805: Decode the video based on the reference frame number and reference layer number to obtain the reconstructed image of the current image.

[0275] The decoding end can obtain the reconstructed image corresponding to the reference frame number and reference layer number from the DPB, and then use the obtained reconstructed image corresponding to the reference frame number and reference layer number as the reference image. Based on the reference image, video decoding is performed to obtain the reconstructed image of the current image.

[0276] In one possible implementation, after the decoding end performs hierarchical decoding to obtain the reconstructed image of the L3 layer of the current image, it can store the reconstructed images of the N3 layers of the current image into the DPB. The fourth reference layer number set of the current image includes the layer numbers of M layers, and the M layers include N3 layers, 1≤M<L3, where L3 represents the total number of layers of the current image; or, the reconstructed image of the highest layer among the N3 layers can be stored into the DPB.

[0277] The fourth reference layer number set of the current image can be obtained by parsing the bitstream. However, when the decoding end is decoding, the highest layer number L4 obtained for the current image may be less than the total number of layers L3 of the current image. Therefore, when storing the reconstructed image of the current image into the DPB, if L4 is greater than or equal to the highest layer number among the above M layers, then N3 = M; and if L4 is less than the highest layer number among the above M layers, then N3 < M.

[0278] Each time the decoder acquires a reconstructed image of a layer of the current image, it can determine whether the layer number belongs to the fourth reference layer number set of the current image. If it does, the reconstructed image of that layer can be stored in the DPB; otherwise, it does not need to be stored in the DPB. That is, each frame only needs to store N3 reconstructed images of each layer. For example, if the fourth reference layer number set of the current image is Rx = {1,3,5}, M = 3, and the highest layer number of the current image obtained by the decoder is L4 = 6 > 5, the decoder only needs to store the reconstructed images of layers 1, 3, and 5 of the current image in the DPB, N3 = 3 = M; or, for another example, if the fourth reference layer number set of the current image is Rx = {1,3,5}, M = 3, and the highest layer number of the current image obtained by the decoder is L4 = 4 < 5, the decoder only needs to store the reconstructed images of layers 1 and 3 of the current image in the DPB, N3 = 2 < M. Alternatively, after the decoding end performs hierarchical decoding to obtain the reconstructed images of each layer of the current image, it can save only the reconstructed image of the highest layer among the N3 layers in the DPB. Each time the decoding end obtains the reconstructed image of a layer of the current image, it can determine whether the layer number belongs to the reference layer number set of the current image. If it does, the reconstructed image of that layer can be stored in the DPB, directly overwriting the previously stored reconstructed image of the current image. If it does not belong, it does not need to be stored in the DPB. For example, if the fourth reference layer number set of the current image is Rx = {1,3,5}, M = 3, and the highest layer number L4 of the current image obtained by the decoding end is 6 > 5, the decoding end only needs to retain the reconstructed image of layer number 5 of the current image in the DPB; as another example, if the fourth reference layer number set of the current image is Rx = {1,3,5}, M = 3, and the highest layer number L4 of the current image obtained by the decoding end is 4, the decoding end only needs to retain the reconstructed image of layer number 3 of the current image in the DPB. Compared to related technologies, this application reduces the number of reconstructed images stored in the DPB, lowers the write bandwidth, improves decoding speed, and saves DPB space.

[0279] In one possible implementation, the decoding end can send the reconstructed image of the L4 layer of the current image for display.

[0280] As described above, after decoding the current image, the decoder stores either the reconstructed image whose layer number belongs to the fourth reference layer set of the current image, or the reconstructed image whose layer number belongs to the fourth reference layer set of the current image and is the highest layer among them. However, when displaying the decoded current image, the decoder can use the reconstructed image of the highest layer of the decoded current image. For example, if the decoder decodes the current image with L4=6, it will display the reconstructed image with layer number 6, while storing the reconstructed images with layer numbers 1, 3, and 5 in the DPB for reference in decoding subsequent image frames. Similarly, if the decoder decodes the current image with L4=4, it will display the reconstructed image with layer number 4, while storing the reconstructed images with layer numbers 1 and 3 in the DPB for reference in decoding subsequent image frames. This results in better image quality displayed by the decoder, ensuring a better viewing experience for the user, and saving DPB storage space.

[0281] In one possible implementation, the decoding end can determine the frame number and layer number of the received image frame, and then send channel feedback information to the encoding end, which is used to indicate the aforementioned frame number and layer number.

[0282] Optionally, when the second frame is determined to be parsed based on the frame number in the bitstream, the decoding end sends channel feedback information to the encoding end. This channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the received first frame. The first frame is the frame preceding the second frame.

[0283] Optionally, when it is determined that the first frame has been received based on the layer number of the received image frame, the decoding end sends channel feedback information to the encoding end. This channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the received first frame.

[0284] In this application, the decoding end can send channel feedback information when it parses the bitstream indicating the start of receiving the next frame (based on the frame number in the bitstream), carrying the frame number of the previous frame and the highest layer number received in the previous frame; alternatively, it can send channel feedback information when it parses the bitstream indicating the current image has been received (based on the highest layer number of the current image in the bitstream), carrying the frame number of the current image and the highest layer number received in the current image. By sending channel feedback information in both of these cases, the decoding end can ensure that the highest layer received in any frame obtained by the encoding end is consistent with the highest layer of the same frame actually received by the decoding end, thereby avoiding errors that can occur at the encoding end due to the use of different reference images during encoding and decoding.

[0285] In addition, the decoding end can periodically send channel feedback information to the encoding end, which carries the frame number and layer number corresponding to the latest received bitstream. In this application, the decoding end can also send channel feedback information in other ways, without specific limitations.

[0286] In one possible implementation, when the current image is an image fragment, the method further includes: determining the image fragment number of the received image frame; correspondingly, channel feedback information is also used to indicate the image fragment number.

[0287] An image frame is divided into multiple image slices for encoding and transmission. When the current image is an image slice, the decoding end can determine the received image slice number along with the received frame number and slice number. This information is then included in the channel feedback information along with the frame number, slice number, and corresponding slice number of the received image slice. The above image-based processing can be performed in the same way for image slices.

[0288] The following describes the solution of the above method embodiment through several specific examples.

[0289] Example 1

[0290] Encoding end:

[0291] 1. The encoding end determines the total number of layers for scalable video encoding, sets a set of reference layer numbers, and sends the total number of layers and the set of reference layer numbers to the decoding end.

[0292] In this step, the reference layer number set is the set of reference layer numbers selected by the encoder when encoding a frame of image for reference. It includes the layer numbers to be referenced. For example, the reference layer number set Rx = {1, 4, 6} indicates that the reconstructed images with layer numbers 1, 4, and 6 will be referenced by subsequent image frames and need to be retained in the DPB. The layer numbers in the reference layer number set will not exceed the total number of layers L that can be scalably coded. max When a layer number in the reference layer number set is greater than the total number of layers that can be hierarchically encoded, that layer number will be ignored by the encoder.

[0293] The reference layer number set is set by the encoder, and the number of layer numbers in the reference layer number set is less than or equal to the total number of layers L. max The setting method is not limited. For example, based on the chip's capabilities, the number of layer numbers in the reference layer number set can be set to not exceed half of the maximum number of layers, and the layer number interval can be 2, such as the total number of layers L. max =6, then the reference layer number set Rx = {1,3,5}; or the reference layer number set contains the total number of layers, such as the maximum number of layers L. max =6, with the reference layer number set Rx = {1,3,6}; or dynamically set based on real-time encoding and network feedback.

[0294] The total number of layers and the set of reference layer numbers can be sent to the decoding end in the form of a bitstream, or they can be determined through negotiation between the decoding end and the encoding end. This application does not impose any specific limitations on this.

[0295] Each frame can have an independent and different set of reference layer numbers, or a group / all image frames can have the same set of reference layer numbers at the same time; this is not limited here.

[0296] 2. The encoding end performs hierarchical encoding on the current image and saves the reconstructed image based on the reference layer number set.

[0297] The encoding end acquires a frame of image and encodes it according to the total number of layers in step 1, either according to resolution scalability or quality scalability, thus encoding the image into a multi-layer bitstream. When a layer number is in the reference layer number set, the encoded reconstructed image of that layer is stored in the DPB for reference in subsequent image frames. For example, if the reference layer number set is Rx = {1, 4, 6}, then the reconstructed images of the current image with layer numbers 1, 4, and 6 will be stored in the DPB.

[0298] 3. Obtain channel feedback information and encode subsequent image frames based on the channel feedback information and the reference layer number set.

[0299] The encoded hierarchical bitstream is transmitted over a network, which can be a transmission network with packet loss characteristics. The bitstream can be prioritized according to its different levels, and dropped from lowest to highest priority. For example, the basic layer bitstream has the highest priority and its successful transmission must be maximized; the higher the enhancement layer, the lower its priority, and the lower the priority for ensuring successful transmission. To ensure that higher-priority layers pass, these lower-priority layers can be actively dropped and not transmitted.

[0300] After transmission is complete, the decoding end sends back the received frame number and layer number information, indicating which layer of which frame was received. Here, "transmission complete" refers to the end time of transmission within a certain period. For example, it could mean the previous frame is complete before the next frame is encoded; or it could mean the previous frame is complete before the next frame's encoded bitstream is sent to the transmission module; or it could mean the previous frame is complete when all layers have been transmitted. There is no specific limitation here.

[0301] Before encoding the next frame, the encoder acquires channel feedback information. Based on the frame number and layer number received by the decoder, and combined with layer numbers from the reference layer number set, a reference image is obtained. Specifically, based on the frame number and layer number received by the decoder, the frame number closest to the current image is selected, and the layer number in the reference layer number set corresponding to that frame number that is smaller than that layer number and closest to it is selected. For example, if the decoder receives layer 5 of the previous frame, and the reference layer number set Rx = {1, 4, 6} for the previous frame, then the reconstructed image of layer number 4 from the previous frame is selected as the reference image for encoding. Alternatively, if the encoder determines that the received layer 5 of the previous frame is smaller than the highest layer 6 in the reference layer number set for the previous frame, it can determine a specific frame number from the multiple frame numbers indicated by the channel feedback information as the reference frame number. In this case, the selected reference frame number is the frame that the decoder has confirmed has been completely received. For example, when encoding the 3rd frame, if the frame number received by the decoder includes the 1st frame, the 1st frame can be specified as the reference.

[0302] 4. Repeat steps 2 and 3 until all images in the video have been encoded and transmitted.

[0303] Decoding end:

[0304] 1. The decoding end obtains the total number of layers in the scalable video encoding, as well as the set of reference layer numbers.

[0305] The decoding end can obtain the total number of layers and the set of reference layer numbers by parsing the bitstream, or it can negotiate with the encoding end to determine the total number of layers and the set of reference layer numbers. This application does not make specific limitations on this.

[0306] 2. The decoding end acquires the scalable video bitstream for decoding, and puts the decoded reconstructed image into the DPB based on the layer number of each received frame and the reference layer number set of that frame.

[0307] After receiving the bitstream, the decoding end directly sends it to the decoder for decoding. The decoding end parses the bitstream in order from low to high levels, decoding the base layer and enhancement layer of a frame of image. After decoding each layer of reconstructed image, it determines whether the reconstructed image of that layer needs to be sent to the DPB (Deep Layer Block). Specifically, if the layer number of a certain layer is in the reference layer number set of the frame, then the reconstructed image of that layer is placed in the DPB; if a higher layer of the frame is decoded, and the layer number of that layer is also in the reference layer number set of the frame, then the image of that layer is placed in the DPB, replacing the lower layer reconstructed image of the frame entering the DPB, and serving as the reference layer image of that frame, which will be referenced by subsequent frames. This replacement can be a rewriting and overwriting of image data, reusing a storage space; or it can be a labeling process where the layer is marked as a reference layer, the lower layer is marked as a non-reference layer, the storage space of the reference layer is reserved, and the storage space of the non-reference layer is released.

[0308] If the decoding end provides feedback on the received frame number and layer number, then after obtaining all the receivable data for that frame, the decoding end will provide feedback on the received frame number and layer number corresponding to that frame. Alternatively, the decoding end can provide feedback on the received frame number and layer number for each layer of data received for that frame, and provide feedback in ascending order of layer number.

[0309] 3. After the decoding end completes the decoding of a frame of data, it obtains the reconstructed image of the highest layer of that frame and sends the image to the display module for display.

[0310] Decoding a frame of data involves two scenarios. First, the decoding end receives the highest-level bitstream of the frame and decodes the reconstructed image of that highest-level layer. This image can then be directly sent to the display module. Here, determining whether it is the highest-level layer of the frame can be done by parsing the layer information in the bitstream and checking if the layer number matches the total number of layers. Second, the decoding end receives a non-highest-level bitstream of the frame. After decoding and obtaining the reconstructed image of that layer, it determines that the next bitstream to be decoded belongs to the basic layer of the next frame. In this case, the decoded reconstructed image of that layer is used as the highest-level reconstructed image of the frame and sent to the display module.

[0311] It should be noted that the highest-level reconstructed image obtained in this step is not necessarily the reconstructed image stored in the DPB. Only the reconstructed images of those layers whose layer numbers are located in the reference layer number set of that frame can be stored in the DPB.

[0312] 4. Obtain reference frames based on the decoded image buffer, decode subsequent images, and continue until all images are completed.

[0313] When decoding subsequent images, if an encoded frame is needed as a reference frame, the reconstructed image of that frame is directly obtained into the DPB for reference and decoding is completed.

[0314] Before decoding subsequent images, a set of reference layer numbers for the image can be obtained. This set is used to store the reconstructed images of the corresponding layers into the DPB queue after decoding the reconstructed images of different layers of the image.

[0315] The technical effects of this embodiment include:

[0316] (1) Since the reconstructed image with the specified layer number is stored in the DPB, the reconstructed image with the unspecified layer number will not be stored in the DPB at the encoding end, which reduces the storage cost of the DPB.

[0317] (2) Neither the encoding nor decoding end needs to perform a large number of write operations to write the reconstructed images of unnecessary layers into the DPB's data storage, which reduces the write bandwidth and improves the encoding and decoding processing speed.

[0318] (3) The end-to-end coding scheme is based on channel feedback information. Therefore, even if frames are lost, the decoding end will not be unable to find the reference image, which would lead to screen tearing or incorrect decoding, thus improving the subjective experience. In addition, the channel feedback information can effectively guide the coding end in selecting the reference image and using the received better reference layer for reference, thereby improving the coding compression efficiency.

[0319] (4) The reconstructed layer image sent to the DPB by the decoding end can be different from the image sent for display. The image sent for display can be a layer with a higher layer number than the reconstructed layer image sent to the DPB, and the image quality is better. Therefore, this scheme can send and display images with better image quality, ensuring the user's viewing experience.

[0320] Example 2

[0321] Encoding end:

[0322] 1. Same as step 1 of the encoding end in Example 1.

[0323] In this embodiment, the encoding end sends the bitstream to multiple decoding ends after encoding. Therefore, a multi-way connection is established between one encoding end and multiple decoding ends, and the bitstream transmitted in each path is the same bitstream.

[0324] 2. Same as step 2 of the encoding end in Example 1.

[0325] 3. The steps are the same as step 3 of the encoding end in Embodiment 1, but there are some differences in the way the channel feedback information is used.

[0326] In this step, after the transmission is completed, the encoding end needs to receive channel feedback information from all decoding ends. That is, it needs to obtain the reception status of the current image at all decoding ends, including the frame number and layer number that all decoding ends have received. The layer number of the highest layer of the same frame received by different decoding ends may be different.

[0327] Before encoding the next frame, based on the received frame number and layer number, the largest frame number among the common frame numbers received in the channel feedback information of all decoding ends is selected as the reference frame number, and the layer number in the reference layer number set of the frame that is smaller than the target layer number (see the method embodiment above for determination) and closest to the target layer number is selected as the reference layer number.

[0328] The reference frame number and reference layer number used by this frame are encoded into the bitstream.

[0329] 4. Same as step 4 of the encoding end in Example 1.

[0330] Decoding end:

[0331] 1. Same as step 1 of the decoding end in Example 1.

[0332] 2. Similar to step 2 of the decoding end in Embodiment 1, but when storing the reconstructed images of different layers of the image into the DPB, each decoding end needs to retain the reconstructed images of all layers in the reference layer number set of the frame, and cannot replace the low-layer reconstructed images of the frame entering the DPB with the obtained high-layer reconstructed images, because the decoding end cannot know which layer is used as the reference layer for subsequent image frames.

[0333] 3. Same as step 3 of the decoding end in Example 1.

[0334] 4. Same as step 4 of the decoding end in Example 1.

[0335] Example 3

[0336] This embodiment provides an example of the syntax and semantics of the encoding end bitstream based on a specified layer reference and channel feedback information.

[0337] In Embodiments 1 and 2, the encoding end needs to write the reference layer number set of each frame into the bitstream and transmit it to the decoding end so that the decoding end can obtain this information. The bitstream information in this embodiment is not limited to being added to bitstreams of standard protocols such as H.264 and H.265; it can also be added to non-standard bitstreams. This embodiment uses the H.265 standard bitstream as an example.

[0338] The first example is based on the encoding reference layer number set of the picture parameter set (PPS). Syntax elements indicating scalable encoding are added to the PPS, and syntax elements indicating the reference layer number set are also added to the PPS, as shown in the table below:

[0339]

[0340] The semantics of the newly added syntax elements are as follows:

[0341] The `pps_shortrange_multilayer_flag` parameter indicates whether to add a layered coding configuration parameter. A value of 1 indicates that the current image sequence uses a layered coding method, requiring the parsing of the layered coding method's syntax elements; a value of 0 indicates that it is not used.

[0342] `pps_candidate_reference_layer` indicates the set of reference layer numbers. After the decoder finishes decoding a frame, it needs to store the layer numbers of the reconstructed image, which is then used as a reference frame in the DPB. In the example, this syntax can be 8 bits, with each bit representing a layer number. For example, bit 0 can represent the base layer, and bits 1 to 7 can represent enhancement layers 1 to 7, respectively. The representation is not limited; for example, the number of bits can be adjusted based on the highest layer number. If the highest layer number is greater than 8, the syntax can be more than 8 bits.

[0343] The decoding and processing methods are as follows:

[0344] The decoder acquires the bitstream and parses the PPS information pic_parameter_set_rbsp. When the element pps_extension_present_flag = 1 is parsed, pps_shortrange_multilayer_flag will be parsed. If pps_shortrange_multilayer_flag = 1 is parsed, pps_candidate_reference_layer will be further parsed to obtain the value of this element. Based on the value of each bit of this element, the set of reference layer numbers for this image sequence can be constructed.

[0345] If the parsed element pps_extension_present_flag = 0, then no further parsing of pps_shortrange_multilayer_flag is performed; if the parsed element pps_shortrange_multilayer_flag = 0, then no further parsing of pps_candidate_reference_layer is performed. In this case, the decoder will not process the data according to the method in Example 1.

[0346] The set of reference layer numbers for the image sequence will not be updated until a new pps_candidate_reference_layer syntax element is parsed.

[0347] The second example encodes the reference layer number set based on the slice segment header (SSH). A syntax element indicating scalable encoding is added to the SSH, and a syntax element indicating the reference layer number set is added to the PPS, as shown in the table below:

[0348]

[0349] The semantics of the newly added syntax elements are as follows:

[0350] The `ssh_shortrange_multilayer_flag` flag is used to indicate that layered encoding configuration parameters are added to this slice. When the value is 1, it means that the current image slice uses a layered encoding method, and the syntax elements of the layered encoding method corresponding to this slice need to be parsed; when the value is 0, it means that it is not used.

[0351] `ssh_candidate_reference_layer` indicates the set of reference layer numbers for this slice. After the decoder finishes decoding the current slice, it needs to store the layer numbers of the reconstructed image, which is then used as a reference frame in the DPB. In the example, this syntax can be 8 bits, with each bit representing a layer number. For example, bit 0 can represent the base layer, and bits 1 through 7 can represent enhancement layers 1 through 7, respectively. The representation is not limited; for example, the number of bits can be adjusted based on the highest layer number. If the highest layer number is greater than 8, the syntax can be more than 8 bits.

[0352] The decoding and processing methods are similar to the first example, with the main difference being that the information parsed in this part corresponds to image slices.

[0353] Figure 9 This is a schematic diagram of the structure of the encoding device 900 according to an embodiment of this application. The encoding device 900 includes: an inter-frame prediction module 901 and an encoding module 902.

[0354] Inter-frame prediction module 901 is used to determine the reference frame number of the current image based on channel feedback information, wherein the channel feedback information is used to indicate the information of the image frame received by the decoding end; obtain a first reference layer number set of the first image frame corresponding to the reference frame number, wherein the first reference layer number set includes N1 layer numbers of the first image frame, 1≤N1<L1, and L1 represents the total number of layers of the first image frame; determine the reference layer number of the current image based on the channel feedback information and the first reference layer number set; encoding module 902 is used to perform video encoding on the current image based on the reference frame number and the reference layer number to obtain a bitstream.

[0355] In one possible implementation, the encoding module 902 is specifically used to obtain the reconstructed image corresponding to the reference frame number and the reference layer number from the decoded image buffer DPB, wherein the DPB contains only the N1 layered reconstructed images for the first image frame; the obtained reconstructed image corresponding to the reference frame number and the reference layer number is used as a reference image, and the current image is video encoded according to the reference image to obtain the bitstream.

[0356] In one possible implementation, when there is only one decoding end, the inter-frame prediction module 901 is specifically used to acquire multiple channel feedback information, the channel feedback information being used to indicate the frame number of the image frame received by the decoding end; and to determine the frame number of the current image that is closest to the multiple frame numbers indicated by the multiple channel feedback information as the reference frame number of the current image.

[0357] In one possible implementation, the inter-frame prediction module 901 is specifically used to determine the highest layer number indicated by the channel feedback information indicating the reference frame number as the target layer number; when the first set of reference layer numbers includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, when the first set of reference layer numbers does not include the target layer number, the layer number in the first set of reference layer numbers that is less than and closest to the target layer number is determined as the reference layer number of the current image.

[0358] In one possible implementation, when there are multiple decoding ends, the inter-frame prediction module 901 is specifically used to acquire multiple sets of channel feedback information, the multiple sets of channel feedback information corresponding to the multiple decoding ends, each set of channel feedback information including multiple sets of channel feedback information, the channel feedback information being used to indicate the frame number of the image frame received by the corresponding decoding end; determine one or more common frame numbers based on the multiple sets of channel feedback information, the common frame number being a frame number indicated by at least one channel feedback information in each set of channel feedback information; and determine the reference frame number of the current image based on the one or more common frame numbers.

[0359] In one possible implementation, the inter-frame prediction module 901 is specifically used to obtain the highest layer number indicated by the channel feedback information indicating the reference frame number in each of the multiple sets of channel feedback information; determine the smallest of the multiple highest layer numbers as the target layer number; and determine the reference layer number of the current image based on the target layer number and the first set of reference layer numbers.

[0360] In one possible implementation, the channel feedback information comes from the corresponding decoding end and / or network devices on the transmission link.

[0361] In one possible implementation, the channel feedback information is generated based on the transmitted code stream.

[0362] In one possible implementation, the inter-frame prediction module 901 is specifically used to acquire multiple channel feedback information, the channel feedback information being used to indicate the frame number of the image frame received by the decoding end; determine the frame number closest to the current image among the multiple frame numbers indicated by the multiple channel feedback information as the target frame number; when the highest layer number indicated by the channel feedback information indicating the target frame number is greater than or equal to the highest layer number in the second reference layer number set, determine the target frame number as the reference frame number, the second reference layer number set being the reference layer number set of the second image frame corresponding to the target frame number.

[0363] In one possible implementation, the inter-frame prediction module 901 is further configured to determine a specified frame number among the multiple frame numbers indicated by the multiple channel feedback information as the reference frame number of the current image when the highest layer number indicated by the channel feedback information indicating the target frame number is less than the highest layer number in the second reference layer number set.

[0364] In one possible implementation, the inter-frame prediction module 901 is further configured to, when the first reference layer number set does not include the target layer number, if the first reference layer number set does not include a layer number less than the target layer number, determine the reference frame number of the previous frame of the current image as the reference frame number of the current image, and determine the reference layer number of the previous frame as the reference layer number of the current image.

[0365] In one possible implementation, the bitstream further includes the first set of reference layer numbers.

[0366] In one possible implementation, the bitstream also includes the reference frame number.

[0367] In one possible implementation, the bitstream further includes the reference frame number and the reference layer number.

[0368] In one possible implementation, when the current image is an image slice, the inter-frame prediction module 901 is specifically used to determine that the image slice number of the current image is the same as the image slice number of the image frame received by the decoding end, and to determine the layer number corresponding to the image slice number of the image frame received by the decoding end as the target layer number; when the first reference layer number set includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, when the first reference layer number set does not include the target layer number, the layer number in the first reference layer number set that is less than and closest to the target layer number is determined as the reference layer number of the current image.

[0369] Figure 10 This is a schematic diagram of the structure of a decoding device 1000 according to an embodiment of this application. The decoding device 1000 includes: an acquisition module 1001, an inter-frame prediction module 1002, a decoding module 1003, a display module 1004, and a transmission module 1005.

[0370] The module 1001 is used to acquire the bitstream; the inter-frame prediction module 1002 is used to parse the bitstream to obtain the reference frame number of the current image; acquire the third reference layer number set of the third image frame corresponding to the reference frame number, the third reference layer number set includes N2 layer numbers, 1≤N2<L2, L2 represents the total number of layers of the third image frame; determine the reference layer number of the current image according to the third reference layer number set; and the decoding module 1003 is used to perform video decoding according to the reference frame number and the reference layer number to obtain the reconstructed image of the current image.

[0371] In one possible implementation, the decoding module 1003 is specifically used to obtain the reconstructed image corresponding to the reference frame number and the reference layer number from the decoded image buffer DPB; use the obtained reconstructed image corresponding to the reference frame number and the reference layer number as a reference image, and perform video decoding based on the reference image to obtain the reconstructed image of the current image.

[0372] In one possible implementation, the decoding module 1003 is further configured to store the reconstructed images of the N3 layers of the current image into the DPB, wherein the fourth reference layer number set of the current image includes the layer numbers of M layers, the M layers include the N3 layers, 1≤M<L3, and L3 represents the total number of layers of the current image; or, store the reconstructed image of the highest layer among the N3 layers into the DPB.

[0373] In one possible implementation, the display module 1004 is used to display the reconstructed image of the L4 layer of the current image, where L4 represents the layer number of the highest layer obtained by decoding the current image.

[0374] In one possible implementation, the inter-frame prediction module 1002 is specifically used to determine the highest layer number among the layer numbers corresponding to the multiple reconstructed images of the decoded third image frame; when the third reference layer number set includes the highest layer number, the highest layer number is determined as the reference layer number of the current image; or, when the reference layer number set does not include the highest layer number, the layer number in the third reference layer number set that is less than and closest to the highest layer number is determined as the reference layer number of the current image.

[0375] In one possible implementation, the inter-frame prediction module 1002 is further configured to, when the third reference layer number set does not include the highest layer number, if the third reference layer number set does not include a layer number less than the highest layer number, determine the reference frame number of the previous frame of the current image as the reference frame number of the current image, and determine the reference layer number of the previous frame as the reference layer number of the current image.

[0376] In one possible implementation, the sending module 1005 is used to determine the frame number and layer number of the received image frame; and to send channel feedback information to the encoding end, wherein the channel feedback information is used to indicate the frame number and layer number of the received image frame.

[0377] In one possible implementation, the sending module 1005 is specifically configured to send the channel feedback information to the encoding end when it is determined that parsing of the second frame has started based on the frame number in the bitstream. The channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the received first frame, wherein the first frame is the frame preceding the second frame; or, when it is determined that the first frame has been completely received based on the layer number of the received image frame, the sending module 1005 sends the channel feedback information to the encoding end. The channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the received first frame.

[0378] In one possible implementation, when the current image is an image fragment, the sending module 1005 is further configured to determine the image fragment number of the received image frame; correspondingly, the channel feedback information is further configured to indicate the image fragment number.

[0379] In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly implemented by a hardware encoding processor, or implemented by a combination of hardware and software modules in the encoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0380] The memory mentioned in the above embodiments can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0381] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0382] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0383] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0384] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0385] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0386] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0387] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image encoding method, characterized in that, include: The reference frame number of the current image is determined based on the channel feedback information, which is used to indicate the information of the image frame received by the decoding end; Obtain a first reference layer number set for the first image frame corresponding to the reference frame number. The first reference layer number set includes N1 layer numbers for each layer, where 1 ≤ N1 < L1, and L1 represents the total number of layers in the first image frame. The reference layer number of the current image is determined based on the channel feedback information and the first reference layer number set; The reconstructed image corresponding to the reference frame number and the reference layer number is obtained from the decoded image buffer DPB, wherein the DPB contains only the N1 layered reconstructed images for the first image frame; The reconstructed image corresponding to the reference frame number and the reference layer number is used as a reference image, and the current image is video encoded according to the reference image to obtain a bitstream.

2. The method according to claim 1, characterized in that, When there is only one decoding end, determining the reference frame number of the current image based on channel feedback information includes: Acquire multiple channel feedback information, wherein the channel feedback information is used to indicate the frame number of the image frame received by the decoding end; The frame number that is closest to the current image among the multiple frame numbers indicated by the multiple channel feedback information is determined as the reference frame number of the current image.

3. The method according to claim 2, characterized in that, Determining the reference layer number of the current image based on the channel feedback information and the first reference layer number set includes: The highest layer number indicated by the channel feedback information that indicates the reference frame number is determined as the target layer number; When the first set of reference layer numbers includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, When the target layer number is not included in the first set of reference layer numbers, the layer number in the first set of reference layer numbers that is less than and closest to the target layer number is determined as the reference layer number of the current image.

4. The method according to claim 1, characterized in that, When there are multiple decoding ends, determining the reference frame number of the current image based on channel feedback information includes: Multiple sets of channel feedback information are acquired, and the multiple sets of channel feedback information correspond to the multiple decoding ends. Each set of channel feedback information includes multiple sets of channel feedback information, and the channel feedback information is used to indicate the frame number of the image frame received by the corresponding decoding end. One or more common frame numbers are determined based on the multiple sets of channel feedback information, wherein the common frame number refers to the frame number indicated by at least one channel feedback information in each set of channel feedback information; The reference frame number of the current image is determined based on the one or more shared frame numbers.

5. The method according to claim 4, characterized in that, Determining the reference layer number of the current image based on the channel feedback information and the first reference layer number set includes: Obtain the highest layer number indicated by the channel feedback information that indicates the reference frame number in each of the multiple sets of channel feedback information; The smallest of the multiple highest-level numbers is determined as the target level number; The reference layer number of the current image is determined based on the target layer number and the first set of reference layer numbers.

6. The method according to any one of claims 1-5, characterized in that, The channel feedback information comes from the corresponding decoding end and / or network devices on the transmission link.

7. The method according to any one of claims 1-5, characterized in that, The channel feedback information is generated based on the transmitted code stream.

8. The method according to claim 1, characterized in that, Determining the reference frame number of the current image based on channel feedback information includes: Acquire multiple channel feedback information, wherein the channel feedback information is used to indicate the frame number of the image frame received by the decoding end; The frame number that is closest to the current image among the multiple frame numbers indicated by the multiple channel feedback information is determined as the target frame number; When the highest layer number indicated by the channel feedback information indicating the target frame number is greater than or equal to the highest layer number in the second reference layer number set, the target frame number is determined as the reference frame number, and the second reference layer number set is the reference layer number set of the second image frame corresponding to the target frame number.

9. The method according to claim 8, characterized in that, The method further includes: When the highest layer number indicated by the channel feedback information indicating the target frame number is less than the highest layer number in the second reference layer number set, the specified frame number among the multiple frame numbers indicated by the multiple channel feedback information is determined as the reference frame number of the current image.

10. The method according to claim 3 or 5, characterized in that, The method further includes: When the first reference layer number set does not include the target layer number, if the first reference layer number set does not include a layer number smaller than the target layer number, then the reference frame number of the previous frame of the current image is determined as the reference frame number of the current image, and the reference layer number of the previous frame is determined as the reference layer number of the current image.

11. The method according to any one of claims 1-10, characterized in that, The bitstream also includes the first set of reference layer numbers.

12. The method according to any one of claims 1-3, characterized in that, The bitstream also includes the reference frame number.

13. The method according to any one of claims 1-11, characterized in that, The bitstream also includes the reference frame number and the reference layer number.

14. The method according to claim 1 or 2, characterized in that, When the current image is an image fragment, the channel feedback information includes the image fragment number of the image frame received by the decoding end and the layer number corresponding to the image fragment number; Determining the reference layer number of the current image based on the channel feedback information and the first reference layer number set includes: If the image segment number of the current image is the same as the image segment number of the image frame received by the decoding end, the layer number corresponding to the image segment number of the image frame received by the decoding end is determined as the target layer number; When the first set of reference layer numbers includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, When the target layer number is not included in the first set of reference layer numbers, the layer number in the first set of reference layer numbers that is less than and closest to the target layer number is determined as the reference layer number of the current image.

15. An image decoding method, characterized in that, include: Obtain the bitstream; Parse the bitstream to obtain the reference frame number of the current image; Obtain the third reference layer number set of the third image frame corresponding to the reference frame number. The third reference layer number set includes N2 layer numbers of the layers, 1≤N2<L2, where L2 represents the total number of layers of the third image frame. The reference layer number of the current image is determined based on the third reference layer number set; Reconstruct the image corresponding to the reference frame number and the reference layer number from the decoded image buffer DPB; The reconstructed image corresponding to the reference frame number and the reference layer number is used as a reference image, and video decoding is performed based on the reference image to obtain the reconstructed image of the current image.

16. The method according to claim 15, characterized in that, The method further includes: The reconstructed images of the N3 layers of the current image are stored in the DPB. The fourth reference layer number set of the current image includes the layer numbers of M layers, and the M layers include the N3 layers, 1≤M<L3, where L3 represents the total number of layers of the current image; or, The reconstructed image of the highest layer among the N3 layers is stored in the DPB.

17. The method according to claim 16, characterized in that, The method further includes: The reconstructed image of the L4 layer of the current image is displayed, where L4 represents the layer number of the highest layer obtained by decoding the current image.

18. The method according to any one of claims 15-17, characterized in that, Determining the reference layer number of the current image based on the third reference layer number set includes: Determine the highest layer number among the multiple reconstructed images of the decoded third image frame; When the third set of reference layer numbers includes the highest layer number, the highest layer number is determined as the reference layer number of the current image; or, When the highest layer number is not included in the reference layer number set, the layer number in the third reference layer number set that is less than and closest to the highest layer number is determined as the reference layer number of the current image.

19. The method according to claim 18, characterized in that, The method further includes: When the third reference layer number set does not include the highest layer number, if the third reference layer number set does not include a layer number less than the highest layer number, then the reference frame number of the previous frame of the current image is determined as the reference frame number of the current image, and the reference layer number of the previous frame is determined as the reference layer number of the current image.

20. The method according to any one of claims 15-19, characterized in that, The method further includes: Determine the frame number and layer number of the received image frame; The channel feedback information is sent to the encoding end, and the channel feedback information is used to indicate the frame number and the layer number.

21. The method according to claim 20, characterized in that, Sending channel feedback information to the encoding end includes: When the parsing of the second frame is determined to begin based on the frame number in the bitstream, the channel feedback information is sent to the encoding end. This channel feedback information indicates the frame number of the first frame and the layer number of the highest layer of the received first frame, where the first frame is the frame preceding the second frame; or... When it is determined that the first frame has been completely received based on the layer number of the received image frame, the channel feedback information is sent to the encoding end. The channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the first frame received.

22. The method according to claim 20 or 21, characterized in that, When the current image is an image slice, the method further includes: Determine the image segment number of the received image frame; Correspondingly, the channel feedback information is also used to indicate the image slice number.

23. An image encoding device, characterized in that, include: The inter-frame prediction module is used to determine the reference frame number of the current image based on channel feedback information, which is used to indicate the information of the image frame received by the decoding end. Obtain a first reference layer number set for the first image frame corresponding to the reference frame number. The first reference layer number set includes N1 layer numbers, 1≤N1<L1, where L1 represents the total number of layers in the first image frame. Determine the reference layer number of the current image based on the channel feedback information and the first reference layer number set. The encoding module is used to obtain the reconstructed image corresponding to the reference frame number and the reference layer number from the decoded image buffer DPB, wherein the DPB contains only the N1 layered reconstructed images for the first image frame; the obtained reconstructed image corresponding to the reference frame number and the reference layer number is used as a reference image, and the current image is video encoded according to the reference image to obtain a bitstream.

24. The apparatus according to claim 23, characterized in that, When there is only one decoding end, the inter-frame prediction module is specifically used to acquire multiple channel feedback information, which are used to indicate the frame number of the image frame received by the decoding end; and to determine the frame number of the current image that is closest to the frame number indicated by the multiple channel feedback information as the reference frame number of the current image.

25. The apparatus according to claim 24, characterized in that, The inter-frame prediction module is specifically used to determine the highest layer number indicated by the channel feedback information indicating the reference frame number as the target layer number; when the first set of reference layer numbers includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, when the first set of reference layer numbers does not include the target layer number, the layer number in the first set of reference layer numbers that is less than and closest to the target layer number is determined as the reference layer number of the current image.

26. The apparatus according to claim 23, characterized in that, When there are multiple decoding ends, the inter-frame prediction module is specifically used to acquire multiple sets of channel feedback information, which correspond to the multiple decoding ends. Each set of channel feedback information includes multiple sets of channel feedback information, which are used to indicate the frame number of the image frame received by the corresponding decoding end. Based on the multiple sets of channel feedback information, one or more common frame numbers are determined. The common frame number refers to the frame number indicated by at least one channel feedback information in each set of channel feedback information. Based on the one or more common frame numbers, the reference frame number of the current image is determined.

27. The apparatus according to claim 26, characterized in that, The inter-frame prediction module is specifically used to obtain the highest layer number indicated by the channel feedback information indicating the reference frame number in each of the multiple sets of channel feedback information; determine the smallest of the multiple highest layer numbers as the target layer number; and determine the reference layer number of the current image based on the target layer number and the first set of reference layer numbers.

28. The apparatus according to any one of claims 23-27, characterized in that, The channel feedback information comes from the corresponding decoding end and / or network devices on the transmission link.

29. The apparatus according to any one of claims 23-27, characterized in that, The channel feedback information is generated based on the transmitted code stream.

30. The apparatus according to claim 23, characterized in that, The inter-frame prediction module is specifically used to acquire multiple channel feedback information, which are used to indicate the frame number of the image frame received by the decoding end; and to determine the frame number that is closest to the current image among the multiple frame numbers indicated by the multiple channel feedback information as the target frame number. When the highest layer number indicated by the channel feedback information indicating the target frame number is greater than or equal to the highest layer number in the second reference layer number set, the target frame number is determined as the reference frame number, and the second reference layer number set is the reference layer number set of the second image frame corresponding to the target frame number.

31. The apparatus according to claim 30, characterized in that, The inter-frame prediction module is further configured to determine a specified frame number among the multiple frame numbers indicated by the multiple channel feedback information as the reference frame number of the current image when the highest layer number indicated by the channel feedback information indicating the target frame number is less than the highest layer number in the second reference layer number set.

32. The apparatus according to claim 25 or 27, characterized in that, The inter-frame prediction module is further configured to, when the first reference layer number set does not include the target layer number, if the first reference layer number set does not include a layer number less than the target layer number, determine the reference frame number of the previous frame of the current image as the reference frame number of the current image, and determine the reference layer number of the previous frame as the reference layer number of the current image.

33. The apparatus according to any one of claims 23-32, characterized in that, The bitstream also includes the first set of reference layer numbers.

34. The apparatus according to any one of claims 23-25, characterized in that, The bitstream also includes the reference frame number.

35. The apparatus according to any one of claims 23-33, characterized in that, The bitstream also includes the reference frame number and the reference layer number.

36. The apparatus according to any one of claims 23-35, characterized in that, When the current image is an image fragment, the inter-frame prediction module is specifically used to determine that the image fragment number of the current image is the same as the image fragment number of the image frame received by the decoding end, and to determine the layer number corresponding to the image fragment number of the image frame received by the decoding end as the target layer number; When the first set of reference layer numbers includes the target layer number, the target layer number is determined as the reference layer number of the current image; or, when the first set of reference layer numbers does not include the target layer number, the layer number in the first set of reference layer numbers that is less than and closest to the target layer number is determined as the reference layer number of the current image.

37. An image decoding device, characterized in that, include: The acquisition module is used to acquire the bitstream; An inter-frame prediction module is used to parse the bitstream to obtain the reference frame number of the current image; Obtain the third reference layer number set of the third image frame corresponding to the reference frame number. The third reference layer number set includes N2 layer numbers of the layers, 1≤N2<L2, where L2 represents the total number of layers of the third image frame. The reference layer number of the current image is determined based on the third reference layer number set; The decoding module is used to obtain the reconstructed image corresponding to the reference frame number and the reference layer number from the decoded image buffer DPB; The reconstructed image corresponding to the reference frame number and the reference layer number is used as a reference image, and video decoding is performed based on the reference image to obtain the reconstructed image of the current image.

38. The apparatus according to claim 37, characterized in that, The decoding module is further configured to store the reconstructed images of the N3 layers of the current image into the DPB, wherein the fourth reference layer number set of the current image includes the layer numbers of M layers, the M layers include the N3 layers, 1≤M<L3, and L3 represents the total number of layers of the current image; or, store the reconstructed image of the highest layer among the N3 layers into the DPB.

39. The apparatus according to claim 38, characterized in that, Also includes: The display module is used to display the reconstructed image of the L4 layer of the current image, where L4 represents the layer number of the highest layer obtained by decoding the current image.

40. The apparatus according to any one of claims 37-39, characterized in that, The inter-frame prediction module is specifically used to determine the highest layer number among the layer numbers corresponding to the multiple reconstructed images of the decoded third image frame; when the third reference layer number set includes the highest layer number, the highest layer number is determined as the reference layer number of the current image; or, when the reference layer number set does not include the highest layer number, the layer number in the third reference layer number set that is less than and closest to the highest layer number is determined as the reference layer number of the current image.

41. The apparatus according to claim 40, characterized in that, The inter-frame prediction module is further configured to, when the third reference layer number set does not include the highest layer number, if the third reference layer number set does not include a layer number less than the highest layer number, determine the reference frame number of the previous frame of the current image as the reference frame number of the current image, and determine the reference layer number of the previous frame as the reference layer number of the current image.

42. The apparatus according to any one of claims 37-41, characterized in that, Also includes: The transmitting module is used to determine the frame number and layer number of the received image frame; and to send channel feedback information to the encoding end, wherein the channel feedback information is used to indicate the frame number and the layer number.

43. The apparatus according to claim 42, characterized in that, The sending module is specifically used to send the channel feedback information to the encoding end when the second frame is determined to be parsed based on the frame number in the bit stream. The channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the received first frame. The first frame is the frame preceding the second frame. Alternatively, when it is determined that the first frame has been completely received based on the layer number of the received image frame, the channel feedback information is sent to the encoding end. The channel feedback information is used to indicate the frame number of the first frame and the layer number of the highest layer of the first frame received.

44. The apparatus according to claim 42 or 43, characterized in that, When the current image is an image fragment, the sending module is further configured to determine the image fragment number of the received image frame; Correspondingly, the channel feedback information is also used to indicate the image slice number.

45. An encoder, characterized in that, include: One or more processors; A non-transitory computer-readable storage medium coupled to the processor and storing a program executed by the processor, wherein the program, when executed by the processor, causes the encoder to perform the method according to any one of claims 1-14.

46. ​​A decoder, characterized in that, include: One or more processors; A non-transitory computer-readable storage medium coupled to the processor and storing a program executed by the processor, wherein the program, when executed by the processor, causes the decoder to perform the method according to any one of claims 15-22.

47. A non-transitory computer-readable storage medium, characterized in that, Includes program code, which, when executed by a computer device, is used to perform the method according to any one of claims 1-22.

48. A non-transient storage medium, characterized in that, The system includes a computer program and an encoded bitstream, wherein the computer program, when executed by a processor, implements the image encoding method according to any one of claims 1-14 to generate the bitstream.

49. A computer program product comprising instructions that, when run on a computer, cause the computer to perform the method as described in any one of claims 1-22.