Image coding method and apparatus

By selecting the highest quality or resolution image layer determined by the feedback information from the decoding end as the reference frame in image coding, the problems of low quality and resolution caused by the selection of image layer reference frames are solved, higher coding quality and resolution are achieved, and the amount of calculation and error transmission are reduced.

CN115699745BActive Publication Date: 2025-10-17HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080101374.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-26
Publication Date
2025-10-17
Estimated Expiration
2040-05-26

AI Technical Summary

Technical Problem

In existing image coding technologies, the selection of reference frames at the image layer results in low coding quality and resolution, and there are problems of error propagation and high computational complexity.

Method used

By obtaining feedback information from the decoding end, the image layer with the highest quality or resolution is selected as the reference frame of the base layer, and inter-frame coding is performed during encoding to reduce error transmission and calculation complexity.

Benefits of technology

The coding quality and resolution of the base layer and enhancement layer are improved, the bit rate is reduced, the calculation amount and error transmission are reduced, and the image quality and resolution are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115699745B_ABST
    Figure CN115699745B_ABST
Patent Text Reader

Abstract

The application provides an image coding method and device. The image coding method comprises the following steps: obtaining an image to be coded, the image to be coded is divided into a basic layer and at least one enhancement layer; when receiving feedback information sent by a decoding end, determining a reconstructed image corresponding to a frame number and a layer number indicated in the feedback information as a first reference frame, and performing interframe coding on the basic layer according to the first reference frame to obtain a code stream of the basic layer; respectively coding the at least one enhancement layer to obtain code streams of the at least one enhancement layer; and sending the code stream of the basic layer and the code streams of the at least one enhancement layer to the decoding end, wherein the code stream of the basic layer carries coding reference information. The application can improve the quality or resolution of a current image frame.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to image coding technology, and in particular, to an image coding method and device. BACKGROUND

[0002] Wireless screen projection technology refers to a technology of encoding and compressing video data (for example, game pictures rendered by a graphics processing unit (GPU)) generated by a device with strong processing capability, and transmitting the video data to a device with weak processing capability but good display effect (for example, a television, a virtual reality (VR) headset, etc.) for display in a wireless transmission manner. Applications using wireless screen projection technology, such as game screen projection, VR glasses, etc., have the characteristic of interaction, and therefore require extremely low transmission delay. In order to avoid image quality problems caused by data loss, anti-interference is also an important requirement of such applications. In addition, since the larger the data volume is, the greater the transmission power consumption is, it is also important to improve video compression efficiency and reduce transmission power consumption.

[0003] A scalable video coding (SVC) protocol encodes image frames in a source video into multiple image layers corresponding to different qualities or resolutions, and the multiple image layers have reference relationships therebetween. When transmitted, relevant data is transmitted in order of a base layer, a lower quality / smaller resolution image layer to a higher quality / larger resolution image layer. The more image layer data a decoder receives for one frame of image, the better the quality of the reconstructed image is. This technology can more easily match the code rate of transmission to a changing bandwidth without switching code streams and avoids delay caused by switching code streams.

[0004] However, the above technology needs a large amount of calculation to determine the reference frame for encoding of each image layer, and at the same time, has the problem of quality degradation of the reconstructed image caused by loss of the image layer. SUMMARY

[0005] The present application provides an image coding method and device to improve the quality or resolution of a current image frame.

[0006] In a first aspect, the present application provides an image encoding method, comprising: obtaining an image to be encoded, the image to be encoded being divided into a base layer and at least one enhancement layer; when receiving feedback information sent by a decoding end, determining a reconstructed image corresponding to a frame number and a layer number indicated in the feedback information as a first reference frame, and performing inter-frame encoding on the base layer according to the first reference frame to obtain a code stream of the base layer; performing encoding on the at least one enhancement layer respectively to obtain code streams of the at least one enhancement layer; and sending the code stream of the base layer and the code streams of the at least one enhancement layer to the decoding end, wherein the code stream of the base layer carries encoding reference information, and the encoding reference information comprises the frame number and the layer number of the first reference frame.

[0007] In the existing scheme (for example, the SVC protocol, the scalable high-efficiency video coding (SHVC) protocol), the base layer can only refer to the reconstructed image corresponding to the base layer of the previous nth image, where n is a positive integer greater than or equal to 1, and it should be understood that the previous nth image refers to a certain image before the image to be encoded. However, the quality or resolution of the reconstructed image corresponding to the image layer (for example, any enhancement layer) higher than the base layer in the previous nth image is higher than that of the reconstructed image corresponding to the base layer. However, the reconstructed image corresponding to any enhancement layer cannot be used as a reference frame of the base layer, resulting in a lower quality of the code stream obtained by encoding the base layer, and a lower quality or resolution of the image reconstructed based on the code stream, and even a lower quality or resolution of the reconstructed image obtained by decoding the code stream at the decoding end. In the present application, the encoding end obtains the image layer of the image frame with the highest quality or resolution that can be obtained by the decoding end based on the feedback information from the decoding end, and uses the reconstructed image corresponding to the image layer as a reference frame of the base layer, that is, when encoding the base layer, the image referred to by the inter-frame encoding is the reconstructed image corresponding to the image layer with the highest quality or resolution in the previous nth image that is successfully decoded, successfully received or about to be decoded by the decoding end. The image layer is also the highest level image layer that meets the network transmission state and code rate requirements and is fed back by the decoding end. Therefore, the encoding layer uses the reconstructed image corresponding to such image layer as a reference frame to perform inter-frame encoding on the base layer, which can improve the quality of the code stream obtained by encoding the base layer, and can also improve the quality or resolution of the image reconstructed based on the code stream, and even can improve the quality or resolution of the reconstructed image obtained by decoding the code stream of the base layer at the decoding end, thereby improving the overall quality or resolution of the current image frame.

[0008] In addition, in the existing scheme (for example, the SVC protocol and the SHVC protocol), feedback is not required for each frame image or sub-image, which may cause image errors and error propagation problems, and a method of periodically inserting an intra-coded frame is required for periodic correction. In the present application, the decoding end can feed back each frame image or sub-image, thereby avoiding error propagation and improving image quality. In addition, the method of periodically inserting an intra-coded frame is avoided, thereby reducing the code rate.

[0009] In a possible implementation, the image to be encoded is a whole frame image or one of sub-images of the whole frame image.

[0010] In a possible implementation, when the image to be encoded is one of sub-images of the whole frame image, the feedback information further includes position information, and the position information is used to indicate the position of the sub-image to be encoded in the whole frame image.

[0011] In a possible implementation, the frame sequence number indicates the (n-1)th frame image before the image to be encoded, n is a positive integer; the layer sequence number corresponds to the image layer with the highest quality or resolution that is successfully decoded by the decoding end from the code stream of the (n-1)th frame image before the image to be encoded; or the layer sequence number corresponds to the image layer with the highest quality or resolution that is successfully received by the decoding end from the code stream of the (n-1)th frame image before the image to be encoded; or the layer sequence number corresponds to the image layer with the highest quality or resolution that is determined by the decoding end to be decoded from the code stream of the (n-1)th frame image before the image to be encoded.

[0012] In the existing scheme (for example, the SVC protocol, the scalable high-efficiency video coding (SHVC) protocol), the base layer can only refer to the reconstructed image corresponding to the base layer of the previous n-th image, n is a positive integer greater than or equal to 1, and it should be understood that the previous n-th image refers to a certain frame image before the image to be encoded. However, the quality or resolution of the reconstructed image corresponding to the image layer (for example, any enhancement layer) higher than the base layer in the previous n-th image is higher than that of the reconstructed image corresponding to the base layer, but the reconstructed image corresponding to any enhancement layer cannot be used as a reference frame of the base layer, resulting in a lower quality of the code stream obtained by encoding the base layer, and a lower quality or resolution of the image reconstructed based on the code stream, and even a lower quality or resolution of the reconstructed image obtained by decoding the code stream at the decoding end. In this application, the encoding end obtains the image layer of the image frame with the highest quality or resolution that the decoding end can obtain based on the feedback information from the decoding end, and uses the reconstructed image corresponding to the image layer as the reference frame of the base layer, that is, when encoding the base layer, the image referred to by the inter-frame encoding is the reconstructed image corresponding to the image layer with the highest quality or resolution in the previous n-th image that is successfully decoded, successfully received or about to be decoded by the decoding end. The feedback of the decoding end usually also reflects the network transmission state, that is, the current network state can meet the transmission requirements and code rate of which image layer. Therefore, the encoding layer uses the reconstructed image corresponding to such image layer as the reference frame for inter-frame encoding of the base layer, provides a good reference basis for the related region (for example, the static region) of the image to be encoded, and can improve the quality of the code stream obtained by encoding the base layer, and can also improve the quality or resolution of the image reconstructed based on the code stream, and even can improve the quality or resolution of the reconstructed image obtained by decoding the code stream of the base layer at the decoding end, and further improve the quality or resolution of the current image frame as a whole.

[0013] In a possible implementation, after obtaining the image to be encoded, the method further includes: when no feedback information is received or the feedback information includes identification information indicating a receiving failure or a decoding failure, performing inter-frame encoding on the base layer according to a third reference frame, the third reference frame being a reference frame of the base layer of the previous image of the image to be encoded.

[0014] In this application, since the change between adjacent image frames in the video is small, even if the latest feedback information cannot be received due to network factors, the previous image can be referred to, and the quality or resolution of the current image frame will not be greatly affected.

[0015] In a possible implementation, after the image to be encoded is acquired, the method further includes: when no feedback information is received or the feedback information includes identification information indicating a receiving failure or a decoding failure, performing intra-frame encoding on the base layer.

[0016] In a possible implementation, the encoding of the at least one enhancement layer respectively to obtain the code streams of the at least one enhancement layer includes: performing inter-frame encoding on a first enhancement layer according to a second reference frame to obtain a code stream of the first enhancement layer, the first enhancement layer being any one of the at least one enhancement layer, and the second reference frame being a reconstructed image corresponding to a first image layer, the quality or resolution of the first image layer being lower than that of the first enhancement layer.

[0017] In the existing scheme (for example, the SVC protocol and the SHVC protocol), an enhancement layer needs to simultaneously refer to a reconstructed image corresponding to a same-layer image layer of a previous n-th frame image and a reconstructed image corresponding to a low-layer image layer of a same-frame image, that is, to provide a better reference basis for a region (for example, a static region) to be encoded of any one of the enhancement layers, the reconstructed image corresponding to the same-layer image layer of the previous n-th frame image needs to be referred to; to provide a better reference for an occlusion region to be encoded, the reconstructed image corresponding to the low-layer image layer of the same-frame image needs to be referred to. The related processing of the two reference frames increases the calculation amount. Moreover, the reference frame of the enhancement layer can only be the reconstructed image corresponding to the same-layer image layer of the previous n-th frame image and the reconstructed image corresponding to the low-layer image layer of the same-frame image, which limits the quality or resolution of the enhancement layer. In this application, for any one of the enhancement layers, if the base layer is referred to, as described above, the image referred to by the base layer encoding is the image layer with the highest quality or resolution that is successfully decoded, successfully received or about to be decoded by the decoding end in the previous n-th frame image, which has improved the quality or resolution of the base layer, and further improved the quality of the code stream of the enhancement layer encoding referring to the base layer, and also improved the quality or resolution of the reconstructed image obtained based on the code stream, or even improved the quality or resolution of the reconstructed image obtained by decoding the code stream of the base layer by the decoding end. If the enhancement layer is referred to, the enhancement layer itself also directly or indirectly refers to the base layer, so the quality of the code stream of the enhancement layer encoding can also be improved, and the quality or resolution of the reconstructed image obtained based on the code stream can also be improved, or even the quality or resolution of the reconstructed image obtained by decoding the code stream of the base layer by the decoding end can also be improved. Therefore, on the basis of providing a very good reference basis for the region (for example, the static region) to be encoded in the base layer encoding, the high-layer image layer of the same-frame image uses the low-layer image layer as the reference frame, which further provides a reference for the occlusion region, and finally improves the image quality or resolution at a higher level. In addition, the enhancement layer only refers to the reconstructed image corresponding to the low-layer image layer of the same-frame image, which reduces the calculation amount.

[0018] In a possible implementation, the first image layer is an image layer one layer lower than the first enhancement layer; or, the first image layer is the base layer.

[0019] In a possible implementation, the base layer and the low-level enhancement layer use a low-rate MCS to enable user equipment with poor channel conditions to obtain basic video services, and the high-level enhancement layer uses a high-rate MCS to enable user equipment with good channel conditions to obtain higher-quality, higher-resolution video services.

[0020] In a possible implementation, in the process of encoding the at least one enhancement layer to obtain the code stream of the at least one enhancement layer, the method further includes: buffering the reconstructed images corresponding to the base layer and the at least one enhancement layer respectively.

[0021] In a possible implementation, before determining the reconstructed image corresponding to the frame sequence number and the layer sequence number indicated in the feedback information as the first reference frame when the feedback information sent by the decoding end is received, the method further includes: monitoring the feedback information within a set time length; and if the feedback information is received within the set time length, determining that the feedback information is received.

[0022] In this application, if the encoding end does not receive the feedback information within the set time length, it is considered that no feedback information is received, at this time the encoding end will not continue to monitor, on the one hand, unnecessary waiting is avoided, consumption is reduced, on the other hand, the received invalid feedback information can be avoided as useful information processing, so as to cause the encoding end to make wrong judgment on the reference frame.

[0023] In a possible implementation, the base layer and the low-level enhancement layer use a low-rate MCS to enable user equipment with poor channel conditions to obtain basic video services, and the high-level enhancement layer uses a high-rate MCS to enable user equipment with good channel conditions to obtain higher-quality, higher-resolution video services.

[0024] In a possible implementation, the to-be-decoded image is an entire image or one sub-image of an entire image.

[0025] In a possible implementation, when the to-be-decoded picture is one of the sub-pictures of the whole picture, the feedback information further comprises position information, which is used to indicate the position of the to-be-decoded picture in the whole picture.

[0026] In a possible implementation, the second layer sequence number corresponds to the picture layer with the highest quality or resolution in the base layer and the at least one enhancement layer of the to-be-decoded picture, and specifically comprises: the second layer sequence number corresponds to the picture layer with the highest quality or resolution successfully decoded from the code stream of the base layer and the code stream of the at least one enhancement layer of the to-be-decoded picture; or, the second layer sequence number corresponds to the picture layer with the highest quality or resolution successfully received from the code stream of the base layer and the code stream of the at least one enhancement layer of the to-be-decoded picture; or, the second layer sequence number corresponds to the picture layer with the highest quality or resolution currently determined to be decoded from the code stream of the base layer and the code stream of the at least one enhancement layer of the to-be-decoded picture.

[0027] In a possible implementation, the feedback information comprises identification information used to indicate the reception failure when the code stream of the base layer and the code stream of the at least one enhancement layer are both failed to be received; or, the feedback information comprises identification information used to indicate the decoding failure when the code stream of the base layer and / or the code stream of the at least one enhancement layer is failed to be decoded.

[0028] In a possible implementation, after the feedback information is sent to the encoding end, the method further comprises: obtaining the to-be-decoded picture according to the reconstructed picture corresponding to the base layer and the reconstructed picture corresponding to the at least one enhancement layer.

[0029] In a possible implementation, the code stream of the at least one enhancement layer is decoded to obtain the reconstructed picture corresponding to each of the at least one enhancement layer, comprising: performing inter-frame decoding on the code stream of the first enhancement layer according to a second reference frame to obtain the reconstructed picture corresponding to the first enhancement layer, the first enhancement layer being any one of the at least one enhancement layer, and the second reference frame being the reconstructed picture corresponding to a first picture layer, the quality or resolution of the first picture layer being lower than that of the first enhancement layer.

[0030] In a possible implementation, the first picture layer is a picture layer one layer lower than the first enhancement layer; or, the first picture layer is the base layer.

[0031] In a possible implementation, when the feedback information comprises frame sequence numbers and layer sequence numbers of all image layers that are successfully decoded, are about to be decoded, or are successfully received, the reconstructed images corresponding to all image layers are buffered; or when the feedback information comprises frame sequence numbers and layer sequence numbers of an image layer with the highest quality or resolution that is successfully decoded, is about to be decoded, or is successfully received, the reconstructed image corresponding to the image layer with the highest quality or resolution that is successfully decoded, is about to be decoded, or is successfully received is buffered.

[0032] In a possible implementation, after the code stream of the base layer and the code stream of the at least one enhancement layer of the to-be-decoded image are received from the encoding end, the method further includes: when the code stream of the base layer and / or the code stream of the at least one enhancement layer comprises encoding mode indication information, decoding the corresponding image layer in a mode indicated by the encoding mode indication information, the mode indicated by the encoding mode indication information comprising intra decoding or inter decoding.

[0033] In a third aspect, the present application provides an encoding apparatus, comprising: a receiving module configured to acquire a to-be-encoded image, the to-be-encoded image being divided into a base layer and at least one enhancement layer; an encoding module configured to, when receiving feedback information sent by a decoding end, determine a reconstructed image corresponding to frame sequence numbers and layer sequence numbers indicated in the feedback information as a first reference frame, and perform inter-frame encoding on the base layer according to the first reference frame to obtain a code stream of the base layer; and perform encoding on the at least one enhancement layer respectively to obtain code streams of the at least one enhancement layer; and a sending module configured to send the code stream of the base layer and the code streams of the at least one enhancement layer to the decoding end, the code stream of the base layer carrying encoding reference information, the encoding reference information comprising the frame sequence numbers and the layer sequence numbers of the first reference frame.

[0034] In a possible implementation, the to-be-encoded image is an entire frame image or one sub-image of an entire frame image.

[0035] In a possible implementation, when the to-be-encoded image is one sub-image of the entire frame image, the feedback information further comprises position information, the position information being used to indicate a position of the to-be-encoded sub-image in the entire frame image.

[0036] In a possible implementation, the frame sequence number indicates a first n-th frame image of the image to be encoded, n being a positive integer; the layer sequence number corresponds to an image layer with the highest quality or resolution that is successfully decoded by the decoding end from the code stream of the first n-th frame image of the image to be encoded; or, the layer sequence number corresponds to an image layer with the highest quality or resolution that is successfully received by the decoding end from the code stream of the first n-th frame image of the image to be encoded; or, the layer sequence number corresponds to an image layer with the highest quality or resolution that is determined by the decoding end to be decoded from the code stream of the first n-th frame image of the image to be encoded.

[0037] In a possible implementation, the processing module is further configured to inter-frame encode the base layer according to a third reference frame when the feedback information is not received or the feedback information includes identification information indicating a receiving failure or a decoding failure, the third reference frame being a reference frame of a base layer of a previous frame image of the image to be encoded.

[0038] In a possible implementation, the processing module is further configured to intra-frame encode the base layer when the feedback information is not received or the feedback information includes identification information indicating a receiving failure or a decoding failure.

[0039] In a possible implementation, the encoding module is specifically configured to inter-frame encode a first enhancement layer according to a second reference frame to obtain a code stream of the first enhancement layer, the first enhancement layer being any one of the at least one enhancement layer, and the second reference frame being a reconstructed image corresponding to a first image layer, the quality or resolution of the first image layer being lower than that of the first enhancement layer.

[0040] In a possible implementation, the first image layer is an image layer one layer lower than the first enhancement layer; or, the first image layer is the base layer.

[0041] In a possible implementation, the apparatus further includes a processing module configured to buffer reconstructed images corresponding to the base layer and the at least one enhancement layer respectively.

[0042] In a possible implementation, the processing module is further configured to monitor the feedback information within a set time length, and determine that the feedback information is received if the feedback information is received within the set time length.

[0043] In a fourth aspect, the present application provides a decoding apparatus, comprising: a receiving module, configured to receive a code stream of a base layer and code streams of at least one enhancement layer of a to-be-decoded image from an encoding end, wherein the code stream of the base layer carries encoding reference information, and the encoding reference information comprises a first frame sequence number and a first layer sequence number; a decoding module, configured to determine a first reference frame according to the first frame sequence number and the first layer sequence number, and perform inter-frame decoding on the code stream of the base layer according to the first reference frame to obtain a reconstructed image corresponding to the base layer; decode the code streams of the at least one enhancement layer respectively to obtain reconstructed images respectively corresponding to the at least one enhancement layer; and a sending module, configured to send feedback information to the encoding end, wherein the feedback information comprises a second frame sequence number and a second layer sequence number, the second frame sequence number corresponds to the to-be-decoded image, and the second layer sequence number corresponds to an image layer with the highest quality or resolution among the base layer and the at least one enhancement layer of the to-be-decoded image.

[0044] In a possible implementation, the to-be-decoded image is an entire frame image or one sub-image of the entire frame image.

[0045] In a possible implementation, when the to-be-decoded image is one sub-image of the entire frame image, the feedback information further comprises position information, and the position information is used to indicate a position of the to-be-decoded image in the entire frame image.

[0046] In a possible implementation, the second layer sequence number corresponds to an image layer with the highest quality or resolution among the base layer and the at least one enhancement layer of the to-be-decoded image, and specifically comprises: the second layer sequence number corresponds to an image layer with the highest quality or resolution successfully decoded from the code stream of the base layer and the code streams of the at least one enhancement layer of the to-be-decoded image; or, the second layer sequence number corresponds to an image layer with the highest quality or resolution successfully received from the code stream of the base layer and the code streams of the at least one enhancement layer of the to-be-decoded image; or, the second layer sequence number corresponds to an image layer with the highest quality or resolution to be decoded from the code stream of the base layer and the code streams of the at least one enhancement layer of the to-be-decoded image at present.

[0047] In a possible implementation, when the code stream of the base layer and the code streams of the at least one enhancement layer are both failed to be received, the feedback information comprises identification information used to indicate the reception failure; or, when the code stream of the base layer and / or the code streams of the at least one enhancement layer are failed to be decoded, the feedback information comprises identification information used to indicate the decoding failure.

[0048] In a possible implementation, the decoding module is further configured to obtain the to-be-decoded image according to the reconstructed image corresponding to the base layer and the reconstructed images corresponding to the at least one enhancement layer.

[0049] In a possible implementation, the decoding module is specifically configured to perform inter-frame decoding on the code stream of the image layer to obtain a reconstructed image corresponding to the first enhancement layer, the first enhancement layer being any one of the at least one enhancement layer, and the second reference frame being a reconstructed image corresponding to the first image layer, the first image layer having a quality or resolution lower than that of the image layer.

[0050] In a possible implementation, the first image layer is an image layer one layer lower than the first enhancement layer; or the first image layer is the base layer.

[0051] In a possible implementation, the method further includes: when the feedback information includes frame sequence numbers and layer sequence numbers of all image layers that are successfully decoded, are about to be decoded, or are successfully received, buffering the reconstructed images corresponding to all the image layers; or when the feedback information includes frame sequence numbers and layer sequence numbers of an image layer having the highest quality or resolution that is successfully decoded, is about to be decoded, or is successfully received, buffering the reconstructed image corresponding to the image layer having the highest quality or resolution that is successfully decoded, is about to be decoded, or is successfully received.

[0052] In a possible implementation, the decoding module is further configured to, when the code stream of the base layer and / or the code stream of the at least one enhancement layer includes encoding mode indication information, decode the corresponding image layer in a mode indicated by the encoding mode indication information, the mode indicated by the encoding mode indication information including intra-frame decoding or inter-frame decoding.

[0053] In a fifth aspect, the present application provides an encoder, including: a processor and a transmission interface;

[0054] The processor is configured to invoke program instructions stored in the memory to implement the method in any one of the above first aspect.

[0055] In a sixth aspect, the present application provides a decoder, including: a processor and a transmission interface;

[0056] The processor is configured to invoke program instructions stored in the memory to implement the method in any one of the above second aspect.

[0057] In a seventh aspect, the present application provides a computer readable storage medium, including a computer program, when the computer program is executed on a computer or a processor, the computer or the processor executes the method in any one of the above first to second aspects.

[0058] In an eighth aspect, the present application provides a computer program product, which comprises computer program codes, and when the computer program codes are run on a computer or a processor, the computer or the processor executes the method of any one of the first to seventh aspects. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1A is a block diagram of a video coding system 10 for implementing embodiments of the present application;

[0060] Figure 1A is a block diagram of a video coding system 40 for implementing embodiments of the present application;

[0061] Figure 2 is a flowchart of an image encoding method embodiment of the present application;

[0062] Figure 3 is a flowchart of an image decoding method embodiment of the present application;

[0063] Figure 4 shows an exemplary schematic diagram of an image encoding and decoding process;

[0064] Figure 5 shows an exemplary schematic diagram of image hierarchical encoding and decoding;

[0065] Figure 6 shows an exemplary schematic diagram of an encoding process at an encoding end;

[0066] Figure 7 shows an exemplary schematic diagram of a decoding process at a decoding end;

[0067] Figure 8 shows an exemplary schematic diagram of an image encoding method of the present application;

[0068] Figure 9 is a structural schematic diagram of an encoding device embodiment of the present application;

[0069] Figure 10 is a structural schematic diagram of a decoding device embodiment of the present application. DETAILED DESCRIPTION

[0070] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0071] The terms "first", "second", etc. in the description of the embodiments of the present application and claims and drawings are only used for the purpose of distinguishing description, and cannot be understood as indicating or implying relative importance, nor can it be understood as indicating or implying sequence. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, comprising a series of steps or units. The method, system, product or device is not necessarily limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0072] It should be understood that in the present application, "at least one" refers to one or more, and "multiple" refers to two or more. "And / or" is used to describe the association between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0073] The technical solutions related to the embodiments of the present application can not only be applied to existing video coding standards (such as H.264 / advanced video coding (AVC), H.265 / high efficiency video coding (HEVC) standards), but also can be applied to future video coding standards (such as H.266 / versatile video coding (VVC) standard). The terms used in the implementation part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application. The following will first introduce some concepts that may be involved in the embodiments of the present application.

[0074] In the field of video coding, the terms "picture," "frame," or "image" can be used synonymously. Video coding is performed at a source side, typically including processing (e.g., by compression) of original video pictures to reduce the amount of data needed to represent that video picture, for more efficient storage and / or transmission. Video decoding is performed at a destination side, typically including inverse processing relative to the encoder, to reconstruct video pictures. Embodiments involving video picture "encoding" are to be understood as involving "encoding" or "decoding" of a video sequence. The combination of an encoding part and a decoding part is also referred to as coding.

[0075] A system architecture to which embodiments of the present application apply is described below. Referring to Figure 1A , Figure 1A is a block diagram of an example of a video encoding and decoding system 10 for implementing embodiments of the present application. As shown in Figure 1A , the video encoding and decoding system 10 can include a source device 12 and a destination device 14. The source device 12 generates encoded video data and, as such, the source device 12 can be referred to as a video encoding apparatus. The destination device 14 can decode the encoded video data generated by the source device 12 and, as such, the destination device 14 can be referred to as a video decoding apparatus. The various embodiments of the source device 12 or the destination device 14 can include one or more processors and a memory coupled to the one or more processors. The memory can include, but is not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer, as described herein. The source device 12 and the destination device 14 can comprise various apparatuses, including a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone, a television, a camera, a display device, a digital media player, a video gaming console, an automotive computer, a wireless communication device, or similar device.

[0076] Although Figure 1A the source device 12 and the destination device 14 are illustrated as separate devices, device embodiments can also include both the source device 12 and the destination device 14 or functionality of both, i.e., the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality, simultaneously. In such embodiments, the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality can be implemented using the same hardware and / or software, or separate hardware and / or software, or any combination thereof.

[0077] Source device 12 and destination device 14 can communicate information through the use of a communication channel 13. Destination device 14 can receive encoded video data from source device 12 via communication channel 13. Communication channel 13 can include one or more media or devices that enable capture of encoded video data from source device 12 and communication of the encoded video data to destination device 14. In one example, communication channel 13 can include one or more communication media that enable source device 12 to transmit encoded video data directly to destination device 14 in real-time. In this example, source device 12 can modulate encoded video data signals

[0078] Source device 12 includes an encoder 20, and can optionally include a picture source 16, a picture preprocessor 18, and a communication interface 22. In various implementations, the encoder 20, picture source 16, picture preprocessor 18, and communication interface 22 can be hardware components of source device 12, or can be software programs stored in memory of source device 12. Each is described separately as follows:

[0079] Picture source 16 can include or be any kind of picture capturing device, e.g., for capturing real world pictures, and / or any kind of picture or comment (for screen content coding, some text on the screen is also considered as part of the picture or image to be coded) generating device, e.g., a computer graphics processor for generating computer animated pictures, or any kind of device for acquiring and / or providing real world pictures, computer animated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). Picture source 16 can be a camera for capturing pictures or a memory for storing pictures, and picture source 16 can also include any kind of (internal or external) interface for storing previously captured or generated pictures and / or for acquiring or receiving pictures. When picture source 16 is a camera, picture source 16 can be, e.g., a local or integrated camera integrated in the source device; when picture source 16 is a memory, picture source 16 can be a local or integrated memory, e.g., integrated in the source device. When picture source 16 includes an interface, the interface can be, e.g., an external interface for receiving pictures from an external video source, e.g., an external picture capturing device such as a camera, an external memory or an external picture generating device, e.g., an external computer graphics processor, a computer or a server. The interface can be any kind of interface according to any proprietary or standardized interface protocol, e.g., a wired or wireless interface, an optical interface.

[0080] Picture pre-processor 18 is configured to receive raw picture data 17 and perform pre-processing on raw picture data 17 to obtain pre-processed picture 19 or pre-processed picture data 19. For example, the pre-processing performed by picture pre-processor 18 can include trimming, color format conversion, color adjustment or de-noising. It is to be noted that performing pre-processing on picture data 17 is not a mandatory process in the present application, and the present application does not limit the pre-processing.

[0081] Encoder 20 (or video encoder 20) is configured to receive pre-processed picture data 19, process pre-processed picture data 19 using a relevant prediction mode (e.g., the prediction modes in various embodiments described herein) to provide encoded picture data 21. In some embodiments, encoder 20 can be configured to perform various embodiments described later to implement the application of the picture encoding method described in the present application at the encoding side.

[0082] The communication interface 22 can be configured to receive the encoded picture data 21 and transmit the encoded picture data 21 to the destination device 14 or any other device (e.g. a storage) via the link 13 for storage or direct reconstruction, which can be any device for decoding or storage. The communication interface 22 can be configured to, for example, encapsulate the encoded picture data 21 into a suitable format, e.g. packets, for transmission over the link 13.

[0083] The destination device 14 comprises a decoder 30, and optionally, the destination device 14 can further comprise a communication interface 28, a picture post-processor 32 and a display device 34, which are described as follows.

[0084] The communication interface 28 can be configured to receive the encoded picture data 21 from the source device 12 or any other source, e.g. a storage device, e.g. an encoded picture data storage device. The communication interface 28 can be configured to transmit or receive the encoded picture data 21 via the link 13 between the source device 12 and the destination device 14, or via any kind of network, e.g. a wired or wireless network or any combination thereof, or any kind of private and public network or any combination thereof. The communication interface 28 can be configured to, for example, de-encapsulate the packets transmitted by the communication interface 22 to obtain the encoded picture data 21.

[0085] Both the communication interface 28 and the communication interface 22 can be configured as unidirectional or bidirectional communication interfaces, and can be configured to, for example, send and receive messages to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission, e.g. transmission of the encoded picture data.

[0086] The decoder 30 (or referred to as the decoder 30) can be configured to receive the encoded picture data 21 and provide decoded picture data 31 or decoded pictures 31. In some embodiments, the decoder 30 can be configured to perform various embodiments described hereinafter to implement the application of the image decoding method described in the present application at the decoding side.

[0087] The picture post-processor 32 can be configured to perform post-processing on the decoded picture data 31 (or referred to as reconstructed picture data) to obtain post-processed picture data 33. The post-processing performed by the picture post-processor 32 can include color format conversion, toning, retouching or resampling, or any other processing, and can be further configured to transmit the post-processed picture data 33 to the display device 34. It is noted that the post-processing on the decoded picture data 31 (or referred to as reconstructed picture data) is not a mandatory process in the present application, which is not limited in this regard.

[0088] A display device 34 for receiving the post-processed picture data 33 to display the picture to, for example, a user or viewer. The display device 34 can be or can comprise any kind of display for presenting a reconstructed picture, for example, an integrated or external display or monitor. For example, the display can comprise a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any kind of other display.

[0089] Although, Figure 1A The source device 12 and the destination device 14 are illustrated as separate devices, but a device embodiment can also comprise both the source device 12 and the destination device 14 or the functionality of both, i.e., the source device 12 or the corresponding functionality and the destination device 14 or the corresponding functionality. In such embodiments, the source device 12 or the corresponding functionality and the destination device 14 or the corresponding functionality can be implemented using the same hardware and / or software or using separate hardware and / or software or any combination thereof.

[0090] It is apparent to a person skilled in the art on the basis of the description that the functionality of different units or Figure 1A The presence and (precise) division of the functionality of the illustrated source device 12 and / or destination device 14 can differ depending on the actual device and application. The source device 12 and the destination device 14 can comprise any of a variety of devices, including any kind of handheld or stationary device, for example, a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camcorder, a desktop computer, a set-top box, a television, a camera, a car device, a display device, a digital media player, a video game console, a video streaming device, for example, a content service server or a content distribution server, a broadcast receiver device, a broadcast transmitter device, etc., and can not use or use any kind of operating system.

[0091] The encoder 20 and the decoder 30 can each be implemented as any of a variety of suitable circuitry, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combinations thereof. If the techniques are implemented partially in software, a device can store instructions for the software in any suitable non-transitory computer-readable storage medium, and can execute the instructions with one or more processors to perform the techniques of this disclosure. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) can be considered to be one or more processors.

[0092] In some cases, Figure 1A The video encoding and decoding system 10 shown in FIG. 1 is merely one example. The techniques of this disclosure can be applied in a video encoding setting (e.g., video encoding or video decoding) that does not necessarily involve any data communication between an encoding and decoding device. In other examples, data can be retrieved from local storage, streamed over a network, etc. A video encoding device can encode data and store the data to memory, and / or a video decoding device can retrieve data from memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other, but rather only encode data to memory and / or retrieve data from memory and decode the data.

[0093] Referring to Figure 1B , Figure 1B is a block diagram of an example of a video coding system 40 for implementing embodiments of the present disclosure. The video coding system 40 can implement a combination of various techniques of embodiments of the present disclosure. In the illustrated implementation, the video coding system 40 can include an imaging device 41, the encoder 20, the decoder 30 (and / or a video encoder / decoder implemented by logic circuitry 47 of a processing unit 46), an antenna 42, one or more processors 43, one or more memories 44, and / or a display device 45.

[0094] As Figure 1B illustrated, the imaging device 41, the antenna 42, the processing unit 46, the logic circuitry 47, the encoder 20, the decoder 30, the processor(s) 43, the memory(ies) 44, and / or the display device 45 can be in communication with each other. As discussed, while the video coding system 40 is illustrated with the encoder 20 and the decoder 30, in different examples, the video coding system 40 can include only the encoder 20 or only the decoder 30.

[0095] In some examples, the antenna 42 can be used to transmit or receive an encoded bitstream of video data. Additionally, in some examples, the display device 45 can be used to present video data. In some examples, the processing unit 46 can include application-specific integrated circuit (ASIC) logic, a graphics processor, a general purpose processor, etc. The video coding system 40 can also include an optional processor 43, which similarly can include ASIC logic, a graphics processor, a general purpose processor, etc. In some examples, the processing unit 46 can be implemented in hardware, such as video encoding specific hardware, etc., and the processor 43 can be implemented in general software, operating systems, etc. Additionally, the memory 44 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.), etc. In non-limiting examples, the memory 44 can be implemented by cache memory. In some examples, the logic circuit 47 can access the memory 44 (e.g., for implementing an image buffer). In other examples, the logic circuit 47 and / or the processing unit 46 can include memory (e.g., cache, etc.) for implementing an image buffer, etc.

[0096] In some examples, an encoder 20 implemented by a logic circuit can include an image buffer (implemented by the processing unit 46 or the memory 44) and a graphics processing unit (implemented by the processing unit 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit can include the encoder 20 implemented by the logic circuit 47 to implement various modules discussed by any other encoder system or subsystem described herein. The logic circuit can be used to perform various operations discussed herein.

[0097] In some examples, a decoder 30 can be implemented by a logic circuit 47 in a similar manner to implement various modules discussed by any other decoder system or subsystem described herein. In some examples, a logic circuit implemented decoder 30 can include an image buffer (implemented by the processing unit 2820 or the memory 44) and a graphics processing unit (implemented by the processing unit 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit can include the decoder 30 implemented by the logic circuit 47 to implement various modules discussed by any other decoder system or subsystem described herein.

[0098] In some examples, the antenna 42 can be used to receive an encoded bitstream of video data. As discussed, the encoded bitstream can include data, indicators, index values, mode selection data, etc. discussed herein related to encoding video frames, e.g., data related to encoding partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators (as discussed), and / or data defining encoding partitions). The video coding system 40 can also include a decoder 30 coupled to the antenna 42 and used to decode the encoded bitstream. The display device 45 is used to present video frames.

[0099] It should be understood that the decoder 30 can be used to perform the reverse process for the examples described in the present application embodiments with respect to the reference encoder 20. With respect to the signaling syntax elements, the decoder 30 can be used to receive and parse such syntax elements and decode the related video data accordingly. In some examples, the encoder 20 can entropy encode the syntax elements into the encoded video bitstream. In such examples, the decoder 30 can parse such syntax elements and decode the related video data accordingly.

[0100] It should be noted that the encoder 20 and the decoder 30 in the present application embodiments can be the corresponding encoder / decoder of a video standard protocol, e.g., H.263, H.264, HEVC, moving picture experts group (MPEG)-2, MPEG-4, VP8, VP9, etc., or a next generation video standard protocol, e.g., H.266, etc.

[0101] The scheme of the present application embodiments is described in detail as follows:

[0102] Figure 2 A flowchart of the image encoding method embodiments of the present application is shown. The process 200 can be performed by an encoder of a source device. The process 200 is described as a series of steps or operations, and it should be understood that the process 200 can be performed in various orders and / or simultaneously, not limited to the execution order shown. As shown, the method of the present embodiments can include: Figure 2 Figure 2

[0103] Step 201, obtaining an image to be encoded.

[0104] The image to be encoded is a whole frame image or one of the sub-images of the whole frame image, which can be referred to the above description of the image frame, and will not be described here. In the present application, the image to be encoded is divided into a base layer and at least one enhancement layer, and the at least one enhancement layer is arranged in order from low to high in quality or resolution.

[0105] ​​The Scalable Video Coding (SVC) protocol can be used to refer to image layering. The SVC protocol divides video frames into a base layer and multiple enhancement layers based on user needs. The base layer provides users with the most basic image quality, frame rate, and resolution, while the enhancement layers refine image quality by providing more information such as image resolution, grayscale, and pixel values. The more image layers, the higher the resulting image quality. When an SVC-encoded bitstream is transmitted over a communication network, different modulation and coding schemes (MCSs) can be used for different image layers. For example, using a low-rate MCS for the base layer and lower-level enhancement layers can provide basic video services to user devices in poor channel conditions, while using a high-rate MCS for higher-level enhancement layers can provide higher-quality, higher-resolution video services to user devices in good channel conditions.

[0106] Step 202: When feedback information sent by the decoding end is received, the reconstructed image corresponding to the frame number and layer number indicated in the feedback information is determined as the first reference frame, and the base layer is inter-coded according to the first reference frame to obtain a code stream of the base layer.

[0107] Feedback information is provided by the decoder to the encoder based on the status of the codestream received or decoded during the codestream reception process. Due to factors such as network transmission latency and the decoder's processing power, while the encoder is processing the current image (i.e., the image to be encoded), the decoder may be processing the nth frame (frame number mn) before the current image (frame number m). If n is 1, the decoder may be processing the previous frame (frame number m-1) before the current image. If n is 2, the decoder may be processing the two previous frames (frame number m-2) before the current image, and so on. To ensure that the encoder is updated with the decoder's latest processing status, the decoder may include information about the nth frame, including its frame number (mn) and layer number, in the feedback information it sends to the encoder.

[0108] In one possible implementation, the encoder and decoder determine, by a pre-agreed or pre-set method, that the layer number carried in the feedback information is based on successful decoding. In this case, the layer number corresponds to the image layer with the highest quality or resolution successfully decoded by the decoder from the bitstream of the previous n-th frame image (frame number is mn).

[0109] In a possible implementation, the encoding end and the decoding end determine, by pre-agreement or presetting, that the layer number carried in the feedback information is based on successful reception, and at this time the layer number corresponds to the image layer with the highest quality or resolution that the decoding end successfully receives from the code stream of the (m-n)th image in the past.

[0110] In a possible implementation, the decoding end can determine the image layer that it will decode according to the size of the received code stream, and determine the image layer that it will decode according to the determination result, and at this time the layer number corresponds to the image layer with the highest quality or resolution that the decoding end determines to decode from the code stream of the (m-n)th image in the past. That is, the image layer to be decoded means the image layer that the decoding end can successfully decode within a predetermined time but has not decoded yet; that is, after receiving the code stream, the decoding end will make a judgment combining the size of the code stream and its decoding capability, and when it determines that it can successfully decode the code stream within a predetermined time, it can send feedback information to the encoding end, without waiting for successful decoding to send feedback information.

[0111] In a possible implementation, when the image to be encoded is one of the sub-images of the whole image, the feedback information further includes position information, which is used to indicate the position of the image to be encoded in the whole image. For example, the pixels of the whole image are 64x64, which are divided into four 32x32 non-overlapping sub-images, and the positions of the four sub-images are located at the upper left, upper right, lower left or lower right of the whole image, and the position information is used to indicate which of the four sub-images is the image to be encoded.

[0112] In a possible implementation, when the image to be encoded is one of the sub-images of the whole image, the feedback information further includes information reflecting the position of the image layer fed back by the decoding end in the whole image, such as the starting position of the slice (when the sub-image is a slice), the serial number of the sub-image (the size of the sub-image is pre-agreed), the width or height of the sub-image, and the like.

[0113] In this application, the encoding end can monitor the feedback information within a set time length, and if the feedback information is received within the set time length, it is determined that the feedback information is received. That is, the encoding end can set a time length, and start timing after sending the code stream of an image, and if the feedback information is received within the time length, it is considered that the feedback information is received, and if the feedback information is not received within the time length, it is considered that the feedback information is not received.

[0114] After the encoding end encodes the base layer and at least one enhancement layer of the image to be encoded, it decodes according to the corresponding method of each layer encoding to obtain the corresponding reconstructed images of each layer. These reconstructed images will be cached as reference frames for subsequent images.

[0115] In existing schemes (e.g., the SVC protocol and the scalable high-efficiency video coding (SHVC) protocol), the base layer can only refer to the reconstructed image corresponding to the base layer of the previous n-th frame image, where n is a positive integer greater than or equal to 1. It should be understood that the previous n-th frame image represents a frame image before the image to be encoded. However, the quality or resolution of the reconstructed image corresponding to the image layer higher than the base layer in the previous n-th frame image (e.g., any enhancement layer) is higher than the base layer. However, the reconstructed image corresponding to any enhancement layer cannot be used as a reference frame for the base layer, resulting in lower quality of the code stream obtained by encoding the base layer, lower quality or resolution of the image reconstructed based on the code stream, and even lower quality or resolution of the reconstructed image obtained by decoding at the decoding end. In the present application, the encoding end obtains the image layer of the image frame with the highest quality or resolution that the decoding end can obtain based on the feedback information from the decoding end, and uses the reconstructed image corresponding to the image layer as the reference frame of the base layer, that is, when the encoding end encodes the base layer, the image referenced by the inter-frame coding is the reconstructed image corresponding to the image layer with the highest quality or resolution that has been successfully decoded, successfully received, or is about to be decoded by the decoding end in the previous n-th frame. The feedback from the decoding end usually also reflects the network transmission status, that is, which image layer's transmission requirements and bit rate can be met by the current network status. Therefore, the encoding layer uses the reconstructed image corresponding to such an image layer as a reference frame to perform inter-frame coding on the base layer, providing a good reference basis for the relevant areas of the image to be encoded (such as static areas), which can improve the quality of the code stream obtained by encoding the base layer, and can also improve the quality or resolution of the image reconstructed based on the code stream, and can even improve the quality or resolution of the reconstructed image obtained by decoding the code stream of the base layer by the decoding end, thereby improving the quality or resolution of the current image frame as a whole.

[0116] Step 203: Perform inter-frame coding on the first enhancement layer according to the second reference frame to obtain a code stream of the first enhancement layer, where the first enhancement layer is any one of the at least one enhancement layer.

[0117] The first enhancement layer is any one of the at least one enhancement layer, the first image layer is one of the base layer and the at least one enhancement layer, and the quality or resolution of the first image layer is lower than that of the first enhancement layer. In a frame of the image to be encoded, the image layer of a higher layer can refer to the reconstructed image of a lower layer during encoding. For example, the image to be encoded includes a base layer and three enhancement layers, the layer number of the base layer is 0, and the layer numbers of the enhancement layers from low to high are 1, 2 and 3 according to the quality or resolution. The reference frame of the enhancement layer 1 during encoding is the reconstructed image of the base layer 0, the reference frame of the enhancement layer 2 during encoding is the reconstructed image of the enhancement layer 1 or the reconstructed image of the base layer 0, and the reference frame of the enhancement layer 3 during encoding is the reconstructed image of the enhancement layer 2 or the reconstructed image of the enhancement layer 1 or the reconstructed image of the base layer 0. As long as the condition that the image layer of a higher layer can refer to the reconstructed image of a lower layer during encoding is met, the application does not make specific limitations on the specific reference of the reconstructed image of which image layer corresponding to the same frame of image the enhancement layer refers to.

[0118] In the existing scheme (for example, the SVC protocol, the SHVC protocol), the enhancement layer needs to simultaneously refer to the reconstructed image corresponding to the same layer image layer of the previous n-th frame image and the reconstructed image corresponding to the low layer image layer of the same frame image, that is, for any one enhancement layer, to provide a better reference basis for the to-be-encoded related region (for example, a static region), the reconstructed image corresponding to the same layer image layer of the previous n-th frame image needs to be referred to; to provide a better reference for the to-be-encoded occluded region, the reconstructed image corresponding to the low layer image layer of the same frame image needs to be referred to. The related processing of the two reference frames increases the calculation amount. Moreover, the reference frame of the enhancement layer can only be the reconstructed image corresponding to the same layer image layer of the previous n-th frame image and the reconstructed image corresponding to the low layer image layer of the same frame image, which limits the quality or resolution of the enhancement layer. In the present application, for any one enhancement layer, if the basic layer is referred to, as described above, the image referred to by the basic layer encoding is the image layer with the highest quality or resolution successfully decoded, successfully received or about to be decoded by the decoding end in the previous n-th frame image of the basic layer, which has improved the quality or resolution of the basic layer, and further improved the quality of the code stream encoded by the enhancement layer referring to the basic layer, and also can improve the quality or resolution of the image reconstructed based on the code stream, and even can improve the quality or resolution of the reconstructed image decoded by the decoding end based on the code stream of the basic layer. If the enhancement layer is referred to, the enhancement layer itself also directly or indirectly refers to the basic layer, so the quality of the code stream encoded by the enhancement layer can also be improved, and the quality or resolution of the image reconstructed based on the code stream can also be improved, and even the quality or resolution of the reconstructed image decoded by the decoding end based on the code stream of the basic layer can also be improved. Therefore, on the basis of providing a good reference basis for the to-be-encoded related region (for example, a static region) in the basic layer encoding, the high layer image layer of the same frame image uses the low layer image layer as the reference frame, which further provides a reference for the occluded region, and finally improves the image quality or resolution at a higher level.

[0119] Step 204, sending the code stream of the basic layer and the code stream of at least one enhancement layer to the decoding end.

[0120] The code stream of the basic layer carries the coding reference information, and the coding reference information includes the frame sequence number and the layer sequence number of the first reference frame. The coding end can package the code stream of the basic layer and the code stream of at least one enhancement layer together and send them to the decoding end, or can package the code stream of the basic layer and the code stream of at least one enhancement layer separately according to the image layer and send them to the decoding end in turn, which is not limited in the present application. The coding end sends the frame sequence number and the layer sequence number of the reference frame used when encoding the basic layer to the decoding end, and the decoding end can directly obtain the reconstructed image of the corresponding image layer as the reference image when performing inter-frame decoding.

[0121] After the encoding end sends the code stream, a timer is started, and feedback information from the decoding end is monitored within a set time period, so as to determine the reference frame of the base layer of a subsequent image frame during encoding.

[0122] In the existing scheme (for example, the SVC protocol and the SHVC protocol), feedback is not required for each image frame or sub-image, which may cause image errors and error propagation problems, and a method of periodically inserting an intra-coded frame is required for periodic correction. However, the present application can feed back each image frame or sub-image, avoid error propagation, and improve image quality. In addition, the present application avoids the periodic insertion of an intra-coded frame, thereby reducing the code rate.

[0123] Therefore, in the image encoding method provided by the present application, the encoding end obtains the image layer of the image frame with the highest quality or resolution that can be obtained by the decoding end based on the feedback information from the decoding end. The image layer is the most consistent with the network transmission state and the code rate requirement, and therefore the quality or resolution of the base layer can be improved. In addition, the enhancement layer of the same image frame is encoded by referring to the reconstructed image of the lower layer, and the quality or resolution of the current image frame can be improved as a whole.

[0124] In a possible implementation, when no feedback information is received or the feedback information includes identification information indicating a reception failure or a decoding failure, the base layer is inter-coded according to a third reference frame, which is a reference frame of the base layer of a previous image frame of the image to be encoded. Before the step 202, if the encoding end does not receive feedback information from the decoding end within a set time period during monitoring of the feedback information, the base layer of the current image frame can be encoded by referring to the reference frame of the base layer of the previous image frame. Since the change between adjacent image frames in a video is small, even if the latest feedback information cannot be received due to network factors, the previous image frame can be referred to, and the quality or resolution of the current image frame will not be greatly affected.

[0125] In a possible implementation, when no feedback information is received or the feedback information includes identification information indicating a reception failure or a decoding failure, the base layer is intra-coded. Similarly, before the step 202, if the encoding end does not receive feedback information from the decoding end within a set time period during monitoring of the feedback information, the base layer of the current image frame can also be encoded by using an intra-coding manner. In this way, the intra-coding manner does not affect the quality or resolution of the base layer, and the quality or resolution of the current image frame is ensured.

[0126] Figure 3A flowchart of an embodiment of the image decoding method of the present application is shown. The process 300 can be performed by a decoder of a destination device. The process 300 is described as a series of steps or operations, which should be understood as not necessarily being limited to the order shown, and / or occurring at the same time. For example, Figure 3 As shown, the method of the present embodiment can include: Figure 3

[0127] Step 301, receiving a code stream of a base layer and a code stream of at least one enhancement layer of a to-be-decoded image from an encoding end.

[0128] Corresponding to step 204 of the above method embodiment, the decoding end receives a code stream of a base layer or a code stream of a base layer and at least one enhancement layer of a to-be-decoded image from an encoding end, and the code stream of the base layer carries encoding reference information, which includes frame sequence numbers and layer sequence numbers of reference frames used by the encoding end when encoding a base layer of an image corresponding to the to-be-decoded image. The to-be-decoded image can be an entire image or one sub-image of an entire image. Optionally, when the to-be-decoded image is one sub-image of an entire image, the encoding reference information further includes position information, which is used to indicate a position of a reference frame used by the encoding end when encoding a base layer of an image corresponding to the to-be-decoded image in the entire image.

[0129] Step 302, determining a first reference frame according to the frame sequence numbers and the layer sequence numbers, and performing inter-frame decoding on the code stream of the base layer according to the first reference frame to obtain a reconstructed image corresponding to the base layer.

[0130] The decoding end can directly obtain the reference frame of the base layer based on the information carried in the code stream, and perform inter-frame decoding on the base layer based on the reference frame.

[0131] Step 303, performing inter-frame decoding on the code stream of a first enhancement layer according to a second reference frame to obtain a reconstructed image corresponding to the first enhancement layer, the first enhancement layer being any one of the at least one enhancement layer.

[0132] The first enhancement layer is any one of the at least one enhancement layer, the second reference frame is a reconstructed image corresponding to a first image layer, the first image layer being one of the base layer and the at least one enhancement layer, and the quality or resolution of the first image layer being lower than that of the first enhancement layer. The present application uses a decoder corresponding to an encoder to decode layer by layer starting from the base layer, and the reconstructed image of a lower layer is used as a reference frame of a higher image layer. It should be noted that the reference frame of the higher image layer can be the reconstructed image of a lower layer, the reconstructed image of the base layer, or the reconstructed images of several lower layers, which are not limited in the present application.

[0133] ​In a possible implementation, when the code stream of the base layer and / or the code stream of the at least one enhancement layer comprises the coding mode indication information, the decoding end can decode the corresponding image layer in the mode indicated by the coding mode indication information, which includes intra decoding or inter decoding. Corresponding to the encoding end, if the encoding end adopts intra coding when encoding a certain image layer, the decoding end also needs to adopt intra decoding when decoding the image layer; if the encoding end adopts inter coding based on a certain reference frame when encoding a certain image layer, the decoding end also needs to adopt inter decoding based on the reference frame when decoding the image layer.

[0134] In the present application, the decoding end can obtain the to-be-decoded image according to the reconstructed image corresponding to the base layer and the reconstructed image corresponding to the at least one enhancement layer.

[0135] Step 304: sending feedback information to the encoding end.

[0136] The feedback information comprises a second frame number and a second layer number, the second frame number corresponding to the to-be-decoded image, and the second layer number corresponding to the image layer with the highest quality or resolution in the base layer and the at least one enhancement layer of the to-be-decoded image. In the process of processing the to-be-decoded image, the decoding end can send feedback information related to the to-be-decoded image to the encoding end, as described in the above embodiments, at this time, the feedback information is used to let the encoding end determine the reference frame when encoding the base layer of the subsequent image frame.

[0137] The frame number in the above feedback information corresponds to the frame number of the to-be-decoded image. The layer number corresponds to the image layer with the highest quality or resolution successfully decoded from the code stream of the base layer and the code stream of the at least one enhancement layer of the to-be-decoded image; or, the layer number corresponds to the image layer with the highest quality or resolution successfully received from the code stream of the base layer and the code stream of the at least one enhancement layer of the to-be-decoded image; or, the layer number corresponds to the image layer with the highest quality or resolution currently determined to be decoded from the code stream of the base layer and the code stream of the at least one enhancement layer of the to-be-decoded image. Similar to the description in step 202, the layer number corresponds to one of successful decoding, successful receiving or decoding to be performed, which is related to the priority agreement between the encoding end and the decoding end or the mode set in advance, or related to the processing capability of the decoding end, which will not be described herein.

[0138] In a possible implementation, when the code stream of the base layer and the code stream of the at least one enhancement layer are both failed to be received, the decoding end can carry identification information indicating the reception failure in the feedback information; or, when the code stream of the base layer and / or the code stream of the at least one enhancement layer are failed to be decoded, the decoding end can carry identification information indicating the decoding failure in the feedback information.

[0139] In a possible implementation, when the feedback information comprises frame sequence numbers and layer sequence numbers of all image layers that are successfully decoded, are about to be decoded, or are successfully received, the decoding end can buffer reconstructed images corresponding to all image layers of the image to be decoded; or when the feedback information comprises frame sequence numbers and layer sequence numbers of an image layer with the highest quality or resolution that is successfully decoded, is about to be decoded, or is successfully received, the decoding end can only buffer a reconstructed image corresponding to the image layer with the highest quality or resolution that is successfully decoded, is about to be decoded, or is successfully received in the image to be decoded.

[0140] Based on the technical solutions of the foregoing method embodiments, the following specific embodiments are used for detailed description.

[0141] Figure 4 An exemplary schematic diagram of an image coding process is shown in FIG. 1, which includes encoding end reference frame establishment, encoding, and code stream sending at the encoding end, and code stream receiving and feedback, decoding end reference frame establishment, and decoding at the decoding end. Figure 4 The image coding method provided in the present application mainly involves encoding end / decoding end reference frame establishment, encoding / decoding, and feedback.

[0142] Figure 5 An exemplary schematic diagram of image layered coding is shown in FIG. 2, in which a source image is divided into a base layer and at least one enhancement layer (for example, enhancement layer 1 and enhancement layer 2), and the image layers are respectively encoded to generate multiple code streams (including a code stream of the base layer, a code stream of enhancement layer 1, and a code stream of enhancement layer 2), which are transmitted to the decoding end through a network. Figure 5 The decoding end decodes the code stream of the base layer, the code stream of enhancement layer 1, and the code stream of enhancement layer 2 layer by layer to obtain a reconstructed image corresponding to the base layer, a reconstructed image corresponding to enhancement layer 1, and a reconstructed image corresponding to enhancement layer 2.

[0143] Figure 6 An exemplary schematic diagram of an encoding process at the encoding end is shown in FIG. 3, which includes source image input, image layer division, image layer encoding, and code stream output. Figure 6As shown, the base layer of the source image is encoded by the base layer encoder to obtain the base layer code stream, and its inter-frame coding reference frame is the optimal reference frame. The determination of the optimal reference frame is related to the feedback information received by the transceiver from the decoding end. The base layer encoder can also reconstruct the reconstructed image of the base layer. The enhancement layer 1 of the source image is encoded by the enhancement layer 1 encoder to obtain the enhancement layer 1 code stream, and its inter-frame coding reference frame is the reconstructed image of the base layer. The enhancement layer 1 encoder can also reconstruct the reconstructed image of the enhancement layer 1. The enhancement layer 2 of the source image is encoded by the enhancement layer 2 encoder to obtain the enhancement layer 2 code stream, and its inter-frame coding reference frame is the reconstructed image of the enhancement layer 1. The enhancement layer 2 encoder can also reconstruct the reconstructed image of the enhancement layer 2. And so on. The base layer code stream, the enhancement layer 1 code stream, and the enhancement layer 2 code stream are sent out by the transceiver.

[0144] Figure 7 An exemplary schematic diagram of the decoding process at the decoding end is shown, as shown in FIG. Figure 7 As shown, the transceiver at the decoding end receives the base layer code stream, the enhancement layer 1 code stream, and the enhancement layer 2 code stream from the encoding end. The base layer decoder performs inter-frame decoding on the base layer code stream to obtain a reconstructed image of the base layer, whose reference frame is determined based on the information carried in the base layer code stream. The enhancement layer 1 decoder performs inter-frame decoding on the enhancement layer 1 code stream to obtain a reconstructed image of the enhancement layer 1, whose reference frame is the reconstructed image of the base layer. The enhancement layer 2 decoder performs inter-frame decoding on the enhancement layer 2 code stream to obtain a reconstructed image of the enhancement layer 2, whose reference frame is the reconstructed image of the enhancement layer 1. And so on. The decoding end can store the reconstructed image of the base layer, the reconstructed image of the enhancement layer 1, and the reconstructed image of the enhancement layer 2.

[0145] Figure 8 An exemplary schematic diagram of the image encoding method of the present application is shown as follows: Figure 8 As shown, a frame image is divided into three sub-images (Slice0, Slice1 and Slice2), and each sub-image is divided into a base layer (BL) and multiple enhancement layers (EL0, EL1, ...) for encoding respectively.

[0146] During the encoding and decoding process, the optimal reference frame of the base layer is updated on a slice-by-slice basis based on an update signal. On the encoder side, the update signal is a new feedback signal, indicating the highest quality or resolution image layer successfully decoded, successfully received, or about to be decoded by the decoder. On the decoder side, the update signal is the coding reference information carried in the base layer bitstream, indicating the image layers of the image frame used by the encoder during encoding. If all image layers of an image frame are not received or successfully decoded by the decoder, the optimal reference frame for that image frame is not updated.

[0147] Encoding side:

[0148] 1. After encoding image frame 1, reconstructed images corresponding to all image layers of all sub-images of image frame 1 are cached, that is, Slice0 BL, Slice0 EL0, Slice0 EL1, ..., Slice1 BL, Slice1 EL0, Slice1 EL1, ..., Slice2BL, Slice2 EL0, Slice2 EL1.

[0149] 2. Transmit the code stream of each image layer of each sub-image of image frame 1 and obtain a feedback signal from the decoding end. The feedback signal includes the layer sequence number of the image layer with the highest quality or resolution that the decoding end successfully decodes, successfully receives, or is about to decode.

[0150] 3. Update the reconstructed image indicated by the layer number of each Slice to the optimal reference frame of the corresponding Slice, that is, Figure 8 The black image layer corresponding to image frame 1: Slice0 EL1, Slice1 EL0, Slice2 BL.

[0151] 4. The updated optimal reference frames are used as reference frames of the base layer of each sub-image of image frame 2 for inter-frame coding of the base layer of each sub-image of image frame 2.

[0152] 5. After encoding image frame 2, reconstructed images corresponding to all image layers of all sub-images of image frame 2 are cached, that is, Slice0 BL, Slice0 EL0, Slice0 EL1, ..., Slice1 BL, Slice1 EL0, Slice1 EL1, ..., Slice2BL, Slice2 EL0, Slice2 EL1.

[0153] 6. Transmit the code stream of each image layer of each sub-image of image frame 2 and obtain a feedback signal from the decoding end. The feedback signal includes the layer sequence number of the image layer with the highest quality or resolution that the decoding end successfully decodes, successfully receives, or is about to decode.

[0154] 7. Update the reconstructed image indicated by the layer number of each Slice to the optimal reference frame of the corresponding Slice, that is, Figure 8 The black image layer corresponding to image frame 2: Slice0 EL1, Slice1 EL1. Due to transmission loss, the optimal reference frame of all layers of Slice2 is not updated, and the reference frame of the base layer of Slice2 is still the reference frame Slice2 BL of the base layer of Slice2 of image frame 1.

[0155] 8. The updated optimal reference frames are used as reference frames of the base layer of each sub-image of the image frame 3, for inter-frame coding of the base layer of each sub-image of the image frame 3.

[0156] 9、Image frame 3 is encoded, and the reconstructed images corresponding to all image layers of all sub-images of image frame 3 are cached, namely Slice0 BL, Slice0 EL0, Slice0 EL1, …, Slice1 BL, Slice1 EL0, Slice1 EL1, …, Slice2 BL, Slice2 EL0, Slice2 EL1.

[0157] 10、The code streams of the image layers of the sub-images of image frame 3 are transmitted, and a feedback signal of the decoding end is obtained, which includes the layer sequence number of the image layer with the highest quality or resolution that is successfully decoded or successfully received or is about to be decoded by the decoding end.

[0158] 11、The reconstructed images indicated by the layer sequence numbers corresponding to each Slice are updated into the optimal reference frames corresponding to the Slices, namely Figure 8 The black image layers corresponding to image frame 3 in the above-mentioned method are Slice0 EL1 and Slice2 EL1. All layers of Slice1 are not updated into the optimal reference frames due to the transmission loss, and the reference frame of the base layer of Slice1 is still Slice1 EL1, which is the reference frame of the base layer of Slice1 of image frame 2.

[0159] 12、The updated optimal reference frames are respectively used as the reference frames of the base layers of the corresponding sub-images of image frame 4, and are used for the inter-frame encoding of the base layers of the sub-images of image frame 4.

[0160] In this way, the method is iterated.

[0161] Decoding end:

[0162] 1、The code stream of image frame 1 is received and decoded.

[0163] 2、After image frame 1 is decoded, case 1: if a feedback signal is sent for each layer of image frame 1, that is, a feedback signal is sent every time a code stream of an image layer is successfully received, or a feedback signal is sent every time a code stream of an image layer is successfully decoded, and the like, the reconstructed images corresponding to all image layers of all sub-images of image frame 1 are cached, namely Slice0 BL, Slice0 EL0, Slice0 EL1, …, Slice1 BL, Slice1 EL0, Slice1 EL1, …, Slice2 BL, Slice2 EL0, Slice2 EL1; case 2: if only one feedback signal is sent for image frame 1, only the reconstructed image corresponding to the image layer with the highest quality or resolution of image frame 1 is stored, namely Slice0 EL1, Slice1 EL0, Slice2 BL.

[0164] 3. Update the reference frames of the base layer of each slice of picture 1 to the corresponding optimal reference frames according to the coding reference information in the bitstream of the base layer of picture 1, for example, SliceO EL1, Slice1 EL0, Slice2 BL.

[0165] 4. The updated optimal reference frames are respectively used as the reference frames of the base layer of the corresponding sub-pictures of picture 2 for the inter-frame decoding of each base layer of picture 2.

[0166] 5. Receive and decode the bitstream of picture 2.

[0167] 6. After picture 2 is decoded, case 1: if a feedback signal is sent for each layer of picture 2, i.e. a feedback signal is sent for each successfully received bitstream of an image layer, or a feedback signal is sent for each successfully decoded bitstream of an image layer, etc., then the reconstructed images corresponding to all image layers of all sub-pictures of picture 2 are buffered, i.e. SliceO BL, SliceO EL0, SliceO EL1, …, Slice1 BL, Slice1 EL0, Slice1 EL1, …, Slice2 BL, Slice2 EL0, Slice2 EL1; case 2: if only one feedback signal is sent for picture 1, then only the reconstructed image corresponding to the image layer with the highest quality or resolution of picture 2 is stored, i.e. SliceO EL1, Slice1 EL1. In this example, all the bitstreams of Slice2 are lost.

[0168] 7. Update the reference frames of the base layer of each slice of picture 2 to the corresponding optimal reference frames according to the coding reference information in the bitstream of the base layer of picture 2, for example, SliceO EL1, Slice1 EL1, all the bitstreams of Slice2 are lost, which has been informed to the encoding end through the feedback signal, thus the optimal reference frames of the encoding end are not updated for Slice2, and the decoding end is informed through the bitstream, thus the optimal reference frames of the decoding end are not updated for Slice2 either.

[0169] 8. The updated optimal reference frames are respectively used as the reference frames of the base layer of the corresponding sub-pictures of picture 3 for the inter-frame decoding of each base layer of picture 3.

[0170] 9. Receive and decode the bitstream of picture 3.

[0171] 10. After decoding image frame 3, case 1: If a feedback signal is sent for each layer of image frame 3, that is, a feedback signal is sent for each successful reception of a layer's codestream, or a feedback signal is sent for each successful decoding of a layer's codestream, and so on, then the reconstructed images corresponding to all layers of all sub-images of image frame 3 are cached, that is, Slice0 BL, Slice0 EL0, Slice0 EL1, ..., Slice1 BL, Slice1 EL0, Slice1 EL1, ..., Slice2 BL, Slice2 EL0, Slice2 EL1. Case 2: If only one feedback signal is sent for image frame 3, then only the reconstructed images corresponding to the highest quality or resolution layer of image frame 3, Slice0 EL1 and Slice2 EL1, are stored. In this case, all codestreams of Slice 1 are lost.

[0172] 11. According to the coding reference information in the code stream of the base layer of image frame 3, the reference frame of the base layer of each slice of image frame 3 is updated to the corresponding optimal reference frame. For example, all the code streams of Slice0 EL1, Slice2 EL1, and Slice1 are lost, and the encoder has been informed through the feedback signal. Therefore, the optimal reference frame of the encoder does not update Slice1, and the decoder is informed through the code stream. At this time, the optimal reference frame of the decoder does not update Slice1 either.

[0173] 12. The updated optimal reference frames are used as reference frames of the base layers of the sub-images corresponding to the image frame 4, for inter-frame decoding of the base layers of the image frame 4.

[0174] And so on.

[0175] Figure 9 This is a schematic diagram of the structure of an embodiment of the encoding device of the present application, as shown in FIG. Figure 9 As shown, the apparatus of this embodiment may include: a receiving module 901, an encoding module 902, a processing module 903, and a sending module 904. The apparatus of this embodiment may be an encoding apparatus or an encoder used at an encoding end.

[0176] The receiving module 901 is configured to acquire a to-be-encoded image, the to-be-encoded image being divided into a base layer and at least one enhancement layer; the encoding module 902 is configured to, when receiving feedback information sent by a decoding end, determine a reconstructed image corresponding to a frame sequence number and a layer sequence number indicated in the feedback information as a first reference frame, perform inter-frame encoding on the base layer according to the first reference frame to obtain a code stream of the base layer, and perform encoding on the at least one enhancement layer respectively to obtain code streams of the at least one enhancement layer; and the sending module 903 is configured to send the code stream of the base layer and the code streams of the at least one enhancement layer to the decoding end, wherein the code stream of the base layer carries encoding reference information, and the encoding reference information includes the frame sequence number and the layer sequence number of the first reference frame.

[0177] In a possible implementation, the to-be-encoded image is an entire frame image or one sub-image of an entire frame image.

[0178] In a possible implementation, when the to-be-encoded image is one sub-image of the entire frame image, the feedback information further includes position information, and the position information is used to indicate a position of the to-be-encoded sub-image in the entire frame image.

[0179] In a possible implementation, the frame sequence number indicates an nth frame image of the to-be-encoded image, n is a positive integer; the layer sequence number corresponds to an image layer with the highest quality or resolution that is successfully decoded by the decoding end from a code stream of the nth frame image of the to-be-encoded image; or the layer sequence number corresponds to an image layer with the highest quality or resolution that is successfully received by the decoding end from the code stream of the nth frame image of the to-be-encoded image; or the layer sequence number corresponds to an image layer with the highest quality or resolution that is determined by the decoding end to be decoded from the code stream of the nth frame image of the to-be-encoded image.

[0180] In a possible implementation, the processing module 902 is further configured to, when no feedback information is received or the feedback information includes identification information indicating a receiving failure or a decoding failure, perform inter-frame encoding on the base layer according to a third reference frame, the third reference frame being a reference frame of a base layer of a previous frame image of the to-be-encoded image.

[0181] In a possible implementation, the processing module 902 is further configured to, when no feedback information is received or the feedback information includes identification information indicating a receiving failure or a decoding failure, perform intra-frame encoding on the base layer.

[0182] In a possible implementation, the encoding module 902 is specifically configured to perform inter-frame encoding on the first enhancement layer according to a second reference frame to obtain a code stream of the first enhancement layer, the first enhancement layer being any one of the at least one enhancement layer, and the second reference frame being a reconstructed image corresponding to a first image layer, the first image layer having a quality or resolution lower than that of the any one of the image layers.

[0183] In a possible implementation, the first image layer is an image layer one layer lower than the first enhancement layer, or the first image layer is the base layer.

[0184] In a possible implementation, the processing module 903 is configured to buffer the reconstructed images corresponding to the base layer and the at least one enhancement layer, respectively.

[0185] In a possible implementation, the processing module 903 is further configured to monitor the feedback information within a set time length, and determine that the feedback information is received if the feedback information is received within the set time length.

[0186] The apparatus of the embodiment can be used to execute the technical solutions of the method embodiments shown in FIGS. 8-10, and achieve similar implementation principles and technical effects, which will not be described herein again. Figure 2 4 The apparatus of the embodiment can be used to execute the technical solutions of the method embodiments shown in FIGS. 8-10, and achieve similar implementation principles and technical effects, which will not be described herein again.

[0187] Figure 10 FIG. 11 shows a structural schematic diagram of a decoding apparatus embodiment of the present application, and the apparatus of the embodiment can include a receiving module 1001, a decoding module 1002, a processing module 1003, and a sending module 1004. The apparatus of the embodiment can be a decoding apparatus or a decoder for a decoding end. Figure 10

[0188] The receiving module 1001 is configured to receive, from a coding end, a code stream of a base layer and code streams of at least one enhancement layer of a to-be-decoded image, the code stream of the base layer carrying encoding reference information, the encoding reference information including a first frame sequence number and a first layer sequence number; the decoding module 1002 is configured to determine a first reference frame according to the first frame sequence number and the first layer sequence number, and perform inter-frame decoding on the code stream of the base layer according to the first reference frame to obtain a reconstructed image corresponding to the base layer; and decode the code streams of the at least one enhancement layer to obtain reconstructed images corresponding to the at least one enhancement layer, respectively; and the sending module 1004 is configured to send feedback information to the coding end, the feedback information including a second frame sequence number and a second layer sequence number, the second frame sequence number corresponding to the to-be-decoded image, and the second layer sequence number corresponding to an image layer having the highest quality or resolution.

[0189] ​​In a possible implementation, the image to be decoded is a whole frame image or a sub-image of the whole frame image.

[0190] In a possible implementation, when the image to be decoded is a sub-image of the whole frame image, the feedback information further includes position information, which is used to indicate the position of the image to be decoded in the whole frame image.

[0191] In a possible implementation, the second layer sequence number corresponds to an image layer with the highest quality or resolution in the base layer and the at least one enhancement layer of the image to be decoded, and specifically includes: the second layer sequence number corresponds to an image layer with the highest quality or resolution successfully decoded from the code stream of the base layer and the code stream of the at least one enhancement layer of the image to be decoded; or, the second layer sequence number corresponds to an image layer with the highest quality or resolution successfully received from the code stream of the base layer and the code stream of the at least one enhancement layer of the image to be decoded; or, the second layer sequence number corresponds to an image layer with the highest quality or resolution currently determined to be decoded from the code stream of the base layer and the code stream of the at least one enhancement layer of the image to be decoded.

[0192] In a possible implementation, when the code stream of the base layer and the code stream of the at least one enhancement layer are both failed to be received, the feedback information includes identification information used to indicate the reception failure; or, when the code stream of the base layer and / or the code stream of the at least one enhancement layer is failed to be decoded, the feedback information includes identification information used to indicate the decoding failure.

[0193] In a possible implementation, the decoding module 1002 is further configured to obtain the image to be decoded according to the reconstructed image corresponding to the base layer and the reconstructed image corresponding to the at least one enhancement layer.

[0194] In a possible implementation, the decoding module 1002 is specifically configured to obtain the reconstructed image corresponding to the first enhancement layer according to inter-frame decoding of the code stream of the first enhancement layer, the first enhancement layer being any one of the at least one enhancement layer, and the second reference frame being the reconstructed image corresponding to the first image layer, the quality or resolution of the first image layer being lower than that of the first enhancement layer.

[0195] In a possible implementation, the first image layer is an image layer one layer lower than the first enhancement layer; or, the first image layer is the base layer.

[0196] In a possible implementation, the processing module 1003 is configured to cache the reconstructed images corresponding to all image layers when the feedback information comprises the frame sequence numbers and layer sequence numbers of all image layers that are successfully decoded, are about to be decoded, or are successfully received; or cache the reconstructed image corresponding to the image layer with the highest quality or resolution that is successfully decoded, is about to be decoded, or is successfully received when the feedback information comprises the frame sequence number and layer sequence number of the image layer with the highest quality or resolution that is successfully decoded, is about to be decoded, or is successfully received.

[0197] In a possible implementation, the decoding module 1002 is further configured to decode the corresponding image layer in a manner indicated by the encoding manner indication information when the code stream of the base layer and / or the code stream of the at least one enhancement layer comprises the encoding manner indication information, where the manner indicated by the encoding manner indication information comprises intra decoding or inter decoding.

[0198] The apparatus of the embodiment can be used to execute the method of the embodiment. Figures 3-8 The technical solutions of the method embodiment shown in the drawings have similar implementation principles and technical effects, which will not be described herein.

[0199] In the implementation process, each step of the method embodiment can be completed by integrated logic circuits of hardware in the processor or instructions in the form of software. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiment of the application can be directly embodied as hardware coding executed by the processor, or executed by a combination of hardware and software modules in the coding processor. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the storage memory, and the processor reads the information in the storage memory and combines the hardware to complete the steps of the above method.

[0200] The memory mentioned in each of the above embodiments can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct Rambus RAM (DR RAM). It should be noted that the memory of the system and method described herein is intended to include, but not be limited to, these and any other suitable types of memory.

[0201] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or in a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0202] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0203] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. The division of the units is merely logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0204] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0205] In addition, each functional unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit.

[0206] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts of the prior art that make contributions or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various program codes that can be stored in the medium.

[0207] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image coding method, characterized in that: include: Acquire a to-be-encoded image, where the to-be-encoded image is divided into a base layer and at least one enhancement layer; When feedback information sent by the decoding end is received, determining the reconstructed image corresponding to the frame sequence number and layer sequence number indicated in the feedback information as a first reference frame, and performing inter-frame coding on the base layer according to the first reference frame to obtain a code stream of the base layer; Encoding the at least one enhancement layer separately to obtain a code stream of the at least one enhancement layer; Sending a code stream of the base layer and a code stream of the at least one enhancement layer to the decoding end, wherein the code stream of the base layer carries coding reference information, and the coding reference information includes a frame sequence number and a layer sequence number of the first reference frame; The frame number indicates the nth frame before the image to be encoded, where n is a positive integer; The layer sequence number corresponds to the image layer with the highest quality or resolution successfully decoded by the decoding end from the code stream of the nth frame image before the image to be encoded; or, the layer sequence number corresponds to the image layer with the highest quality or resolution successfully received by the decoding end from the code stream of the nth frame image before the image to be encoded; or, the layer sequence number corresponds to the image layer with the highest quality or resolution to be decoded from the code stream of the nth frame image before the image to be encoded determined by the decoding end.

2. The method according to claim 1, characterized in that The image to be encoded is a whole frame image or a sub-image of the whole frame image.

3. The method according to claim 2, characterized in that When the image to be encoded is a sub-image of the entire frame image, the feedback information further includes position information, where the position information is used to indicate the position of the sub-image to be encoded in the entire frame image.

4. The method according to any one of claims 1 to 3, characterized in that After obtaining the image to be encoded, the method further includes: When the feedback information is not received or the feedback information includes identification information for indicating reception failure or decoding failure, the base layer is inter-frame encoded according to a third reference frame, and the third reference frame is a reference frame of the base layer of the previous frame image of the image to be encoded.

5. The method according to any one of claims 1 to 3, characterized in that After obtaining the image to be encoded, the method further includes: When the feedback information is not received or the feedback information includes identification information indicating reception failure or decoding failure, intra-frame encoding is performed on the base layer.

6. The method according to any one of claims 1 to 5, characterized in that The encoding of the at least one enhancement layer to obtain a code stream of the at least one enhancement layer includes: The first enhancement layer is inter-frame encoded according to the second reference frame to obtain the code stream of the first enhancement layer, where the first enhancement layer is any one of the at least one enhancement layer, the second reference frame is the reconstructed image corresponding to the first image layer, and the quality or resolution of the first image layer is lower than the quality or resolution of the first enhancement layer.

7. The method according to claim 6, characterized in that The first image layer is an image layer one layer lower than the first enhancement layer; or, the first image layer is the base layer.

8. The method according to claim 6 or 7, characterized in that In the process of respectively encoding the at least one enhancement layer to obtain the bitstream of the at least one enhancement layer, the method further includes: The reconstructed images corresponding to the base layer and the at least one enhancement layer are cached.

9. The method according to any one of claims 1 to 8, characterized in that The method further includes: before determining the reconstructed image corresponding to the frame number and layer number indicated in the feedback information as the first reference frame upon receiving the feedback information sent by the decoding end, the method further includes: Monitoring the feedback information within a set period of time; If the feedback information is received within the set time period, it is determined that the feedback information is received.

10. An image decoding method, characterized in that: include: receiving a code stream of a base layer and a code stream of at least one enhancement layer of an image to be decoded from an encoding end, wherein the code stream of the base layer carries coding reference information, and the coding reference information includes a first frame sequence number and a first layer sequence number; Determining a first reference frame according to the first frame sequence number and the first layer sequence number, and performing inter-frame decoding on the code stream of the base layer according to the first reference frame to obtain a reconstructed image corresponding to the base layer; Decoding the bitstream of the at least one enhancement layer to obtain reconstructed images corresponding to the at least one enhancement layer; Sending feedback information to the encoder, the feedback information including a second frame sequence number and a second layer sequence number, the second frame sequence number corresponding to the image to be decoded, and the second layer sequence number corresponding to an image layer with the highest quality or resolution among a base layer and at least one enhancement layer of the image to be decoded; The second layer sequence number corresponds to the image layer with the highest quality or resolution among the base layer and at least one enhancement layer of the image to be decoded, specifically including: The second layer sequence number corresponds to the image layer with the highest quality or resolution successfully decoded from the code stream of the base layer and the code stream of at least one enhancement layer of the image to be decoded; or The second layer sequence number corresponds to the image layer with the highest quality or resolution successfully received from the code stream of the base layer and the code stream of at least one enhancement layer of the image to be decoded; or The second layer sequence number corresponds to the image layer with the highest quality or resolution to be decoded from the code stream of the base layer and the code stream of at least one enhancement layer of the image to be decoded, which is currently determined.

11. The method according to claim 10, characterized in that The image to be decoded is a whole frame image or a sub-image of the whole frame image.

12. The method according to claim 11, characterized in that When the image to be decoded is a sub-image of the entire frame image, the feedback information further includes position information, where the position information is used to indicate the position of the image to be decoded in the entire frame image.

13. The method according to any one of claims 10 to 12, characterized in that Also includes: When both the code stream of the base layer and the code stream of the at least one enhancement layer fail to be received, the feedback information includes identification information for indicating the reception failure; or, When decoding of the code stream of the base layer and / or the code stream of the at least one enhancement layer fails, the feedback information includes identification information for indicating decoding failure.

14. The method according to any one of claims 10 to 13, characterized in that After sending the feedback information to the encoding end, the method further includes: The image to be decoded is obtained according to the reconstructed image corresponding to the base layer and the reconstructed image corresponding to the at least one enhancement layer.

15. The method according to any one of claims 10 to 14, characterized in that The decoding of the code stream of the at least one enhancement layer to obtain the reconstructed images corresponding to the at least one enhancement layer includes: The code stream of the first enhancement layer is inter-frame decoded according to the second reference frame to obtain a reconstructed image corresponding to the first enhancement layer, where the first enhancement layer is any one of the at least one enhancement layer, and the second reference frame is the reconstructed image corresponding to the first image layer, and the quality or resolution of the first image layer is lower than the quality or resolution of the first enhancement layer.

16. The method according to claim 15, characterized in that The first image layer is an image layer one layer lower than the first enhancement layer; or, the first image layer is the base layer.

17. The method according to any one of claims 10 to 12, characterized in that When the feedback information includes the frame sequence numbers and layer sequence numbers of all image layers that are successfully decoded, to be decoded, or successfully received, cache the reconstructed images corresponding to all image layers; or When the feedback information includes the frame number and layer number of the image layer with the highest quality or resolution that is successfully decoded, about to be decoded, or successfully received, the reconstructed image corresponding to the image layer with the highest quality or resolution that is successfully decoded, about to be decoded, or successfully received is cached.

18. The method according to any one of claims 10 to 17, characterized in that After receiving the code stream of the base layer and the code stream of at least one enhancement layer of the image to be decoded from the encoding end, the method further includes: When the code stream of the base layer and / or the code stream of the at least one enhancement layer includes coding mode indication information, the corresponding image layer is decoded using the mode indicated by the coding mode indication information, and the mode indicated by the coding mode indication information includes intra-frame decoding or inter-frame decoding.

19. An encoding device, characterized in that: include: A receiving module, configured to obtain an image to be encoded, where the image to be encoded is divided into a base layer and at least one enhancement layer; an encoding module configured to, upon receiving feedback information sent by a decoding end, determine the reconstructed image corresponding to the frame sequence number and layer sequence number indicated in the feedback information as a first reference frame, and perform inter-frame encoding on the base layer based on the first reference frame to obtain a code stream of the base layer; Encoding the at least one enhancement layer separately to obtain a code stream of the at least one enhancement layer; a sending module, configured to send a code stream of the base layer and a code stream of the at least one enhancement layer to the decoding end, wherein the code stream of the base layer carries coding reference information, and the coding reference information includes a frame sequence number and a layer sequence number of the first reference frame; The frame sequence number indicates the nth preceding frame image of the image to be encoded, where n is a positive integer; the layer sequence number corresponds to the image layer with the highest quality or resolution successfully decoded by the decoding end from the code stream of the nth preceding frame image of the image to be encoded; or, the layer sequence number corresponds to the image layer with the highest quality or resolution successfully received by the decoding end from the code stream of the nth preceding frame image of the image to be encoded; or, the layer sequence number corresponds to the image layer with the highest quality or resolution to be decoded from the code stream of the nth preceding frame image of the image to be encoded determined by the decoding end.

20. The device according to claim 19, characterized in that The image to be encoded is a whole frame image or a sub-image of the whole frame image.

21. The device according to claim 20, characterized in that When the image to be encoded is a sub-image of the entire frame image, the feedback information further includes position information, where the position information is used to indicate the position of the sub-image to be encoded in the entire frame image.

22. The device according to any one of claims 19 to 21, characterized in that The encoding module is also used to perform inter-frame encoding on the base layer according to a third reference frame when the feedback information is not received or the feedback information includes identification information for indicating reception failure or decoding failure, and the third reference frame is a reference frame of the base layer of the previous frame image of the image to be encoded.

23. The device according to any one of claims 19 to 21, characterized in that The encoding module is further configured to perform intra-frame encoding on the base layer when the feedback information is not received or the feedback information includes identification information indicating a reception failure or a decoding failure.

24. The device according to any one of claims 19 to 23, characterized in that The encoding module is specifically used to inter-frame encode the first enhancement layer according to the second reference frame to obtain the code stream of the first enhancement layer, where the first enhancement layer is any one of the at least one enhancement layer, the second reference frame is the reconstructed image corresponding to the first image layer, and the quality or resolution of the first image layer is lower than the quality or resolution of the first enhancement layer.

25. The device according to claim 24, characterized in that The first image layer is an image layer one layer lower than the first enhancement layer; or, the first image layer is the base layer.

26. The device according to claim 24 or 25, characterized in that Also includes: A processing module is used to cache the reconstructed images corresponding to the base layer and the at least one enhancement layer respectively.

27. The device according to claim 26, characterized in that The processing module is further configured to monitor the feedback information within a set time period; if the feedback information is received within the set time period, it is determined that the feedback information is received.

28. A decoding device, characterized in that: include: a receiving module, configured to receive a code stream of a base layer and a code stream of at least one enhancement layer of an image to be decoded from an encoding end, wherein the code stream of the base layer carries coding reference information, and the coding reference information includes a first frame sequence number and a first layer sequence number; a decoding module, configured to determine a first reference frame according to the first frame sequence number and the first layer sequence number, and perform inter-frame decoding on the code stream of the base layer according to the first reference frame to obtain a reconstructed image corresponding to the base layer; and respectively decode the code stream of the at least one enhancement layer to obtain reconstructed images corresponding to the at least one enhancement layer; a sending module, configured to send feedback information to the encoding end, the feedback information including a second frame sequence number and a second layer sequence number, the second frame sequence number corresponding to the image to be decoded, and the second layer sequence number corresponding to the image layer with the highest quality or resolution among the base layer and at least one enhancement layer of the image to be decoded; The second layer sequence number corresponds to the image layer with the highest quality or resolution among the base layer and at least one enhancement layer of the image to be decoded, specifically including: The second layer sequence number corresponds to the image layer with the highest quality or resolution successfully decoded from the code stream of the base layer and the code stream of at least one enhancement layer of the image to be decoded; or The second layer sequence number corresponds to the image layer with the highest quality or resolution successfully received from the code stream of the base layer and the code stream of at least one enhancement layer of the image to be decoded; or The second layer sequence number corresponds to the image layer with the highest quality or resolution to be decoded from the code stream of the base layer and the code stream of at least one enhancement layer of the image to be decoded, which is currently determined.

29. The device according to claim 28, characterized in that The image to be decoded is a whole frame image or a sub-image of the whole frame image.

30. The device according to claim 29, characterized in that When the image to be decoded is a sub-image of the entire frame image, the feedback information further includes position information, where the position information is used to indicate the position of the image to be decoded in the entire frame image.

31. The device according to any one of claims 28 to 30, characterized in that When both the code stream of the base layer and the code stream of the at least one enhancement layer fail to be received, the feedback information includes identification information for indicating the reception failure; or when the code stream of the base layer and / or the code stream of the at least one enhancement layer fail to be decoded, the feedback information includes identification information for indicating the decoding failure.

32. The device according to any one of claims 28 to 31, characterized in that The decoding module is further configured to obtain the image to be decoded according to the reconstructed image corresponding to the base layer and the reconstructed image corresponding to the at least one enhancement layer.

33. The device according to any one of claims 28 to 32, characterized in that The decoding module is specifically used to perform inter-frame decoding on the code stream of the first enhancement layer according to the second reference frame to obtain a reconstructed image corresponding to the first enhancement layer, where the first enhancement layer is any one of the at least one enhancement layer, the second reference frame is the reconstructed image corresponding to the first image layer, and the quality or resolution of the first image layer is lower than the quality or resolution of the first enhancement layer.

34. The device according to claim 33, characterized in that The first image layer is an image layer one layer lower than the first enhancement layer; or, the first image layer is the base layer.

35. The device according to any one of claims 28 to 30, characterized in that Also includes: A processing module is used to cache the reconstructed images corresponding to all image layers when the feedback information includes the frame sequence numbers and layer sequence numbers of all image layers that are successfully decoded, about to be decoded or successfully received; or, when the feedback information includes the frame sequence number and layer sequence number of the image layer with the highest quality or resolution that is successfully decoded, about to be decoded or successfully received, cache the reconstructed images corresponding to the image layer with the highest quality or resolution that is successfully decoded, about to be decoded or successfully received.

36. The device according to any one of claims 28 to 35, characterized in that The decoding module is further configured to, when the code stream of the base layer and / or the code stream of at least one enhancement layer includes coding mode indication information, decode the corresponding image layer using the method indicated by the coding mode indication information, where the method indicated by the coding mode indication information includes intra-frame decoding or inter-frame decoding.

37. An encoder, characterized in that include: processor and transmission interface; The processor is configured to call program instructions stored in the memory to implement the method according to any one of claims 1 to 9.

38. A decoder, characterized in that include: processor and transmission interface; The processor is configured to call program instructions stored in the memory to implement the method according to any one of claims 10 to 18.

39. A computer-readable storage medium, characterized in that The method comprises a computer program which, when executed on a computer or a processor, causes the computer or the processor to perform the method according to any one of claims 1 to 9 or 10 to 18.

Citation Information

Patent Citations

  • Dynamic insertion of synchronization predicted video frames

    CN104137543A