Decoding methods, encoding methods, and related device
Patent Information
- Application Number
- PCT/CN2025/080241
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-06
- Filing Date
- 2025-03-03
- Publication Date
- 2025-10-02
AI Technical Summary
In the prior art, an encoding end or a decoding end uses a unified resolution conversion method for the entire reconstructed image sequence, which cannot guarantee the resolution conversion effect of the reconstructed image or reconstructed block.
The decoding end decodes the code stream to obtain a first identifier to indicate that a specific resolution conversion method is used for the reconstructed image or reconstructed block. The encoding end reconstructs the original image and encodes the corresponding identifier to indicate the resolution conversion method.
The flexibility and effect of resolution conversion are improved, ensuring the resolution conversion effect of reconstructed images or reconstructed blocks.
Smart Images

Figure CN2025080241_02102025_PF_FP_ABST
Abstract
Description
Decoding method, encoding method and related equipment
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on March 6, 2024, with application number 202410256484.X and invention name “Decoding method, encoding method and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of coding and decoding, and more specifically, to a decoding method, an encoding method and related devices. Background Art
[0004] In related art, when performing super-resolution reconstruction, an encoder or decoder uses a unified resolution conversion method for the entire reconstructed image sequence.
[0005] However, using a unified resolution conversion method for the entire visual reconstructed image sequence may not guarantee the resolution conversion effect of some reconstructed images in the reconstructed image sequence or some reconstructed blocks in the reconstructed images. Summary of the Invention
[0006] The embodiments of the present application provide a decoding method, an encoding method and related devices, which can not only improve the flexibility of resolution conversion for a reconstructed image or a reconstructed block in a reconstructed image, but also help improve the resolution conversion effect.
[0007] In a first aspect, a decoding method is provided, comprising:
[0008] The decoding end obtains a first reconstructed image by decoding the code stream;
[0009] The decoding end obtains a first identifier by decoding the code stream;
[0010] The first identifier indicates that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method;
[0011] The decoding end performs resolution conversion based on the first resolution conversion method.
[0012] In a second aspect, an encoding method is provided, comprising:
[0013] The encoding end reconstructs the first original image to obtain a first reconstructed image;
[0014] The encoding end encodes a first identifier, where the first identifier indicates that the first reconstructed image or a first reconstructed block in the first reconstructed image uses a first resolution conversion method.
[0015] A third aspect provides a decoding method, comprising:
[0016] The decoding end obtains a first reconstructed image by decoding the code stream;
[0017] The decoding end determines a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0018] The decoding end performs resolution conversion based on the first resolution conversion method.
[0019] In a fourth aspect, an encoding method is provided, comprising:
[0020] The encoding end reconstructs the first original image to obtain a first reconstructed image;
[0021] The encoder determines, based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image, a first resolution conversion mode:
[0022] The encoding end performs resolution conversion based on the first resolution conversion method.
[0023] In a fifth aspect, a decoding device is provided, comprising:
[0024] Decoding unit, used to:
[0025] Obtaining a first reconstructed image by decoding the code stream;
[0026] Obtaining a first identifier by decoding the code stream;
[0027] The first identifier indicates that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method;
[0028] A conversion unit is configured to perform resolution conversion based on the first resolution conversion method.
[0029] In a sixth aspect, an encoding device is provided, comprising:
[0030] a reconstruction unit, configured to reconstruct the first original image to obtain a first reconstructed image;
[0031] The encoding unit is configured to encode a first identifier, where the first identifier indicates that the first reconstructed image or a first reconstructed block in the first reconstructed image uses a first resolution conversion method.
[0032] In a seventh aspect, a decoding device is provided, comprising:
[0033] A decoding unit, configured to obtain a first reconstructed image by decoding the code stream;
[0034] a determining unit, configured to determine a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0035] A conversion unit is configured to perform resolution conversion based on the first resolution conversion method.
[0036] In an eighth aspect, an encoding device is provided, comprising:
[0037] a reconstruction unit, configured to reconstruct the first original image to obtain a first reconstructed image;
[0038] a determining unit, configured to determine a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0039] A conversion unit is configured to perform resolution conversion based on the first resolution conversion method.
[0040] In the ninth aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented, or the steps of the method described in the third aspect are implemented, or the steps of the method described in the fourth aspect are implemented.
[0041] In a tenth aspect, an electronic device is provided, comprising a processor and a communication interface, wherein the processor is configured to:
[0042] Obtaining a first reconstructed image by decoding the code stream;
[0043] Obtaining a first identifier by decoding the code stream;
[0044] The first identifier indicates that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method;
[0045] Resolution conversion is performed based on the first resolution conversion method.
[0046] In an eleventh aspect, an electronic device is provided, comprising a processor and a communication interface, wherein the processor is configured to:
[0047] Reconstructing the first original image to obtain a first reconstructed image;
[0048] A first identifier is encoded, where the first identifier indicates that the first reconstructed image or a first reconstructed block in the first reconstructed image uses a first resolution conversion method.
[0049] In a twelfth aspect, an electronic device is provided, including a processor and a communication interface, wherein the processor is configured to:
[0050] Obtaining a first reconstructed image by decoding the code stream;
[0051] Determining a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0052] Resolution conversion is performed based on the first resolution conversion method.
[0053] In a thirteenth aspect, an electronic device is provided, including a processor and a communication interface, wherein the processor is configured to:
[0054] Reconstructing the first original image to obtain a first reconstructed image;
[0055] Determining a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0056] Resolution conversion is performed based on the first resolution conversion method.
[0057] In the fourteenth aspect, an electronic device is provided, comprising: a memory configured to store video data, and a processing circuit configured to implement the steps of the method described in the first aspect, or the steps of the method described in the second aspect, or the steps of the method described in the third aspect, or the steps of the method described in the fourth aspect.
[0058] In the fifteenth aspect, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented, or the steps of the method described in the third aspect are implemented, or the steps of the method described in the fourth aspect are implemented.
[0059] In the sixteenth aspect, a coding and decoding system is provided, including: an encoding end device and a decoding end device, the encoding end device can be used to execute the steps of the method described in the first aspect, and the decoding end device can be used to execute the steps of the method described in the second aspect, or implement the steps of the method described in the third aspect, or implement the steps of the method described in the fourth aspect.
[0060] In the seventeenth aspect, a chip is provided, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method described in the first aspect, or the steps of the method described in the second aspect, or the steps of the method described in the third aspect, or the steps of the method described in the fourth aspect.
[0061] In aspect 18, a computer program / program product is provided, which is stored in a storage medium, and is executed by at least one processor to implement the steps of the method described in aspect 1, or the steps of the method described in aspect 2, or the steps of the method described in aspect 3, or the steps of the method described in aspect 4.
[0062] In an embodiment of the present application, a first identifier is used to indicate that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method, so that the decoding end can perform resolution conversion based on the first resolution conversion method, which not only improves the flexibility of resolution conversion, but also helps to improve the resolution conversion effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0064] FIG1 is a schematic diagram of a coding and decoding system provided in an embodiment of the present application.
[0065] FIG2 is a schematic diagram of the structure of the encoder provided in an embodiment of the present application.
[0066] FIG3 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application.
[0067] FIG4 is a schematic diagram illustrating the principle of RPR provided in an embodiment of the present application.
[0068] FIG5 is a schematic diagram of the principle of NNSR provided in an embodiment of the present application.
[0069] FIG6 is a schematic diagram of the grid structure used by the NNSR provided in an embodiment of the present application.
[0070] FIG7 is a schematic diagram of the principle of the pixel shuffling layer provided in an embodiment of the present application.
[0071] FIG8 is a schematic flowchart of a decoding method provided in an embodiment of the present application.
[0072] FIG9 is a schematic flowchart of an encoding method provided in an embodiment of the present application.
[0073] FIG10 is a schematic flowchart of another decoding method provided in an embodiment of the present application.
[0074] FIG11 is a schematic flowchart of another encoding method provided in an embodiment of the present application.
[0075] FIG12 is a schematic block diagram of a decoding device provided in an embodiment of the present application.
[0076] FIG13 is a schematic block diagram of an encoding device provided in an embodiment of the present application.
[0077] FIG14 is a schematic block diagram of another decoding device provided in an embodiment of the present application.
[0078] FIG15 is a schematic block diagram of another encoding device provided in an embodiment of the present application.
[0079] FIG16 is a schematic block diagram of an electronic device provided in an embodiment of the present application.
[0080] FIG17 is a schematic diagram of the hardware structure of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION
[0081] The following will be combined with the accompanying drawings in the embodiments of this application to clearly describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0082] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in this application represents at least one of the connected objects. For example, "A or B" covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0083] FIG1 is a schematic diagram of a codec system 10 provided in an embodiment of the present application. The technical solution of the embodiment of the present application relates to encoding and decoding (CODEC) (including encoding or decoding) of video data. The video data includes original unencoded video, encoded video, decoded (e.g., reconstructed) video, or syntax elements.
[0084] As shown in FIG1 , a codec system 10 includes a source device 100 that provides encoded video data to be decoded and displayed by a destination device 110. Specifically, the source device 100 provides the video data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a mobile phone, a wearable device (e.g., a smartwatch or a wearable camera), a television, a camera, a display device, an in-vehicle device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a digital media player, a video game console, a video conferencing device, a video streaming device, a broadcast receiver device, a broadcast transmitter device, a spacecraft, an aircraft, a robot, a satellite, and the like.
[0085] In the example of Figure 1, the source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. The destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. The source device 100 represents an example of a video encoding device, while the destination device 110 represents an example of a video decoding device. In other examples, the source device 100 and the destination device 110 may not include some of the components in Figure 1, or may also include other components outside of Figure 1. For example, the source device 100 can receive video data from an external data source (such as an external camera). Similarly, the destination device 110 can be connected to an external display device interface without including an integrated display device. For another example, the memory 102 and the memory 113 can be external memories.
[0086] Although FIG1 illustrates source device 100 and destination device 110 as separate devices, in some examples, the two may be integrated into a single device. In such embodiments, the functions corresponding to source device 100 and the functions corresponding to destination device 110 may be implemented using the same hardware or software, or using separate hardware or software, or any combination thereof.
[0087] In some examples, source device 100 and destination device 110 can perform one-way video transmission or two-way video transmission. If it is two-way video transmission, source device 100 and destination device 110 can operate in a substantially symmetrical manner, that is, each of source device 100 and destination device 110 includes an encoder and a decoder.
[0088] Data source 101 represents a source of video data (i.e., raw, unencoded video data) and provides a series of pictures containing the video data to encoder 200, which encodes the picture data. Data source 101 of source device 100 may include a video capture device (such as a video camera), a video archive containing previously captured raw video, or a video feed interface for receiving video from a video content provider. Alternatively, data source 101 may generate computer graphics-based data as the source video, or combine real-time video, archived video, and computer-generated video. In these cases, encoder 200 encodes the captured, pre-captured, or computer-generated video data. Encoder 200 may rearrange the pictures from the order in which they were received (sometimes referred to as "display order") into an encoding order. Encoder 200 may generate a bitstream comprising the encoded video data. Source device 100 may then output the encoded video data to communication medium 120 via output interface 104 for receipt or retrieval, for example, by input interface 111 of destination device 110.
[0089] Memory 102 of source device 100 and memory 113 of destination device 110 represent general purpose memories. In some examples, memory 102 may store raw video data from data source 101, and memory 113 may store decoded video data from decoder 300. Additionally or alternatively, memories 102 and 113 may store software instructions executable by, for example, encoder 200 and decoder 300, respectively. Although memory 102 and memory 113 are shown separately from encoder 200 and decoder 300 in this example, it should be understood that encoder 200 and decoder 300 may also include internal memory for functionally similar or equivalent purposes. If encoder 200 and decoder 300 are deployed on the same hardware device, memory 102 and memory 113 may be the same memory. Furthermore, memories 102 and 113 may store, for example, encoded video data output from encoder 200 and input to decoder 300. In some examples, portions of memory 102 , 113 may be allocated as one or more video buffers, eg, for storing raw, decoded, or encoded video data.
[0090] In some examples, source device 100 can output the encoded data from output interface 104 to memory 113. Similarly, destination device 110 can access the encoded data from memory 113 via input interface 111. Memory 113 or storage 102 can include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a Digital Versatile Disc (DVD), a Compact Disc Read-Only Memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0091] Output interface 104 may include any type of medium or device capable of transmitting encoded video data from source device 100 to destination device 110. For example, output interface 104 may include a transmitter or transceiver, such as an antenna, configured to transmit encoded video data directly from source device 100 to destination device 110 in real time. The encoded video data may be modulated according to a communication standard of a wireless communication protocol and transmitted to destination device 110.
[0092] The communication medium 120 may include a transient medium such as a wireless broadcast or a wired network transmission. For example, the communication medium 120 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 120 may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium) such as a hard disk, a flash drive, a compact disk, a digital video disk, a Blu-ray disc, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0093] In some embodiments, the communication medium 120 may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device 100 to the destination device 110. For example, a server (not shown) can receive encoded video from the source device 100 and provide the encoded video data to the destination device 110, for example, via a network transmission. The server may include a web server (e.g., for a website), a server configured to provide a file transfer protocol service (such as the File Transfer Protocol (FTP) or the File Delivery Over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or an evolved Multimedia Broadcast Multicast Service (eMBMS) server, a Network-attached storage (NAS) device, and the like. The server can implement one or more HTTP streaming protocols, such as MPEG Media Transport (MMT) protocol, Dynamic Adaptive Streaming over HTTP (DASH) protocol, HTTP Live Streaming (HLS) protocol or Real Time Streaming Protocol (RTSP).
[0094] Destination device 110 can access the encoded video data from the server, for example, via a wireless channel (e.g., a Wi-Fi connection) or a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.) for accessing the encoded video data stored on the server.
[0095] The output interface 104 and the input interface 111 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to the IEEE 802.11 standard or the IEEE 802.15 standard (e.g., ZigBee™), the Bluetooth standard, or other physical components. In examples where the output interface 104 and the input interface 111 include wireless components, the output interface 104 and the input interface 111 may be configured to communicate data, such as encoded video data, according to WIFI, Ethernet, a cellular network (such as 4G, LTE (Long Term Evolution), LTE-Advanced, 5G, 6G, etc.).
[0096] The technology provided in the embodiments of the present application can be applied to support video encoding and decoding in one or more multimedia applications such as: video conferencing, over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission, digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0097] Input interface 111 of destination device 110 receives an encoded video bitstream from communication medium 120. The encoded video bitstream may include syntax elements and encoded data units (e.g., sequences, groups of pictures, pictures, slices, blocks, etc.), wherein the syntax elements are used to decode the encoded data units to obtain decoded video data. Display device 114 displays the decoded video data to a user. Display device 114 may include a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0098] The encoder 200 and the decoder 300 may be implemented as one or more of a variety of processing circuits, which may include a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. When the technology is implemented in whole or in part in software, the device may store instructions for the software in an appropriate non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided in the embodiments of the present application.
[0099] The encoder 200 and the decoder 300 can be processed based on the following video coding and decoding standards: H.263, H.264, H.265 (also known as High Efficiency Video Coding (HEVC)), H.266 (also known as Versatile Video Coding (VVC), the second generation Moving Picture Experts Group 2 (MPEG-2), MPEG-4, VP8, VP9, the first generation of open media alliance video (Alliance for Open Media Video 1, AV1), the first generation audio and video coding standard (Audio Video coding Standard 1, AVS1), AVS2, AVS3 or the next generation video standard protocol, which is not specifically limited in the embodiments of the present application.
[0100] Generally, the encoder 200 and decoder 300 can perform block-based encoding and decoding of pictures. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used in the encoding or decoding process). For example, a block can include a two-dimensional matrix of samples of luma or chroma data. For example, the encoder 200 and decoder 300 can encode and decode video data represented in YUV format.
[0101] 2 , which is a schematic diagram of the structure of an encoder 200 provided in an embodiment of the present application, and the encoder 200 may be the encoder 200 in FIG1 . In the example of FIG2 , the encoder 200 includes a memory 201, a coding parameter determination unit 210, a residual generation unit 202, a transform processing unit 203, a quantization unit 204, an inverse quantization unit 205, an inverse transform processing unit 206, a reconstruction unit 207, a filter unit 208, a decoded picture buffer (DPB) 209, and an entropy coding unit 220.
[0102] The memory 201 can store video data to be encoded. For example, the encoder 200 can receive and store video data from the data source 104 shown in FIG1 . In some examples, the memory 201 can be on the same chip as other components of the encoder 200 (as shown in FIG2 ), or can be independent of the chip where these components are located.
[0103] The coding parameter determination unit 210 includes a mode selection unit 211, an inter-frame prediction unit 212, and an intra-frame prediction unit 213. The inter-frame prediction unit 212 is used to obtain a first prediction block of the current block using the inter-frame prediction mode, and the intra-frame prediction unit 213 is used to obtain a second prediction block of the current block using the intra-frame prediction mode. The mode selection unit 211 is used to obtain a target prediction block based on the first prediction block and the second prediction block, and to determine a final prediction mode. In addition, the coding parameter determination unit 210 may also include other functional units, such as a functional unit for determining the division method of the coding unit (CU), a functional unit for determining the transform type of the residual data of the CU or the quantization parameter of the residual data of the CU, etc.
[0104] For ease of description and understanding, the embodiment of the present application refers to the CU to be processed in the current image as the current CU, and the image block to be processed in the current CU as the current block or the image block to be processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.
[0105] The inter prediction unit 212 may include a motion estimation unit and a motion compensation unit. For inter prediction of the current block, the motion estimation unit may perform a motion search to identify one or more matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in the DPB 209).
[0106] The motion estimation unit may form one or more motion vectors (MVs) that represent the position of a reference block in a reference picture relative to the position of a current block in a current picture, and the motion compensation unit may obtain a prediction value of the accuracy indicated by the motion vectors through interpolation.
[0107] The encoding parameter determination unit 210 may provide the target prediction block to the residual generation unit 202. The residual generation unit 202 receives the original unencoded video data of the current block from the memory 201 and calculates the residual between the current block and the target prediction block to obtain a residual block. In some examples, the functionality of the residual generation unit 202 may be implemented using one or more subtractor circuits that perform binary subtraction.
[0108] As an example, the coding parameter determination unit 210 may provide a syntax element representing the coding parameter to the entropy coding unit 220 for encoding. The coding parameter includes one or more of the CU partitioning method, the final prediction mode, the transform type of the CU residual data, or the quantization parameter of the CU residual data.
[0109] The transform processing unit 203 transforms the residual block output by the residual generating unit 202 to obtain a transform coefficient block. The transform may include discrete cosine transform (DCT), integer transform, directional transform, or Karhunen-Loeve transform. In some examples, the encoder 200 may not include the transform processing unit 203.
[0110] The quantization unit 204 may quantize the transform coefficients in the transform coefficient block according to a quantization parameter (QP) value associated with the current block to generate a quantized transform coefficient block.
[0111] The inverse quantization unit 205 and the inverse transform processing unit 206 can respectively perform inverse quantization and inverse transform on the transform coefficient block to obtain a reconstructed residual block. The reconstruction unit 207 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the target prediction block generated by the coding parameter determination unit 210.
[0112] The filter unit 208 may perform one or more filter operations on the reconstructed block. For example, the filter unit 208 may be a deblocking filter (DBF), an adaptive loop filter (ALF), a sample adaptive offset (SAO) filter, etc. In some examples, the encoder 200 may not include the filter unit 208.
[0113] The encoder 200 stores the reconstructed picture obtained from the reconstructed block in the DPB 209. For example, in examples where the operation of the filter unit 208 is not required, the reconstruction unit 207 may store the reconstructed block in the DPB 209. In examples where the operation of the filter unit 208 is required, the filter unit 208 may store the filtered reconstructed block in the DPB 209. The inter-frame prediction unit 212 obtains the reconstructed picture from the DPB 209 to perform inter-frame prediction on blocks of subsequent pictures to be encoded. In some examples, the DPB 209 may be replaced with other types of memory.
[0114] The entropy coding unit 220 may entropy encode syntax elements of other components in the encoder 200 and output encoded video data. For example, the entropy coding unit 220 may entropy encode the quantized transform coefficient blocks from the quantization unit 204. As another example, the entropy coding unit 220 may entropy encode syntax elements from the coding parameter determination unit 210 (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction).
[0115] It is understandable that the composition of the encoder 200 shown in FIG2 is merely an illustration and does not constitute a limitation to the embodiments of the present application.
[0116] FIG3 is a schematic diagram of the structure of a decoder 300 provided in an embodiment of the present application, and the decoder 300 may be the decoder 300 described in FIG1 . In the example of FIG3 , the decoder 300 includes a coded picture buffer (CPB) 301, an entropy decoding unit 302, a prediction processing unit 310, an inverse quantization unit 303, an inverse transform processing unit 304, a reconstruction unit 305, a filter unit 306, and a DPB 307.
[0117] The entropy decoding unit 302 can receive the encoded video data from the CPB 301 and perform entropy decoding on the video data to obtain syntax elements, where the syntax elements indicate encoding parameters, and the encoding parameters include one or more of the CU partitioning method, the final prediction mode, the transform type of the CU's residual data, or the quantization parameter of the CU's residual data.
[0118] When the syntax element includes the final prediction mode, the prediction processing unit 310 obtains the final prediction mode. If the final prediction mode is an inter-frame prediction mode, the prediction block of the current CU can be obtained by the inter-frame prediction unit 311 of the prediction processing unit 310; if the final prediction mode is an intra-frame prediction mode, the prediction block of the current CU can be obtained by the intra-frame prediction unit 312 of the prediction processing unit 310. In some examples, the prediction processing unit 310 may also include a unit for performing prediction functions according to other prediction modes.
[0119] The CPB 301 can obtain and store encoded video data from the communication medium 120 shown in Figure 1. The DPB 307 is used to store decoded pictures. Optionally, the CPB 301 and DPB 307 can also be replaced with other types of memory, which is not specifically limited in this application. In some examples, the CPB 301 can be on the same chip as other components of the decoder 300 (as shown in the figure), or it can be independent of the chip where these components are located.
[0120] The decoder 300 can perform reconstruction operations on each block separately. The entropy decoding unit 302 can entropy decode the syntax elements of the quantized transform coefficients and transform information (such as QP or transform mode indication) to obtain quantized transform coefficients. The quantized transform coefficients are dequantized by the dequantization unit 303 to obtain a transform coefficient block including the transform coefficients. The transform coefficient block is inversely transformed by the inverse transform processing unit 304 to generate a residual block corresponding to the current block. This inverse transform is the inverse operation of the above-mentioned transform.
[0121] The reconstruction unit 305 may reconstruct the current block according to the prediction block and the residual block. For example, the reconstruction unit 305 may add samples of the residual block to corresponding samples of the prediction block to reconstruct the current block.
[0122] The filter unit 306 may perform one or more filter operations on the reconstructed block. For example, the type of the filter unit 306 may refer to the type of the filter unit 208 and will not be described in detail here. In some examples, the operation of the filter unit 306 may be skipped.
[0123] The decoder 300 may store the reconstructed picture obtained from the reconstructed block in the DPB 307. For example, in an example in which the operation of the filter unit 306 is not performed, the reconstruction unit 305 may store the reconstructed block in the DPB 307. In an example in which the operation of the filter unit 306 is performed, the filter unit 306 may store the filtered reconstructed block in the DPB 307. The decoder 300 may output a decoded picture (e.g., decoded video) from the DPB 307 for subsequent presentation to a display device (such as the display device 114 of FIG. 1 ).
[0124] In order to facilitate a better understanding of the embodiments of the present application, the technologies related to the present application are explained.
[0125] (1) Reference Picture Resampling (RPR)
[0126] The increasing use of high-resolution video poses a significant challenge to video transmission in bandwidth-constrained environments. VVC addresses this issue by employing a resampling encoding method. RPR adaptively adjusts resolution based on network conditions. When network bandwidth is low, it encodes downsampled low-resolution (LR) frames. The decoder then upsamples the reconstructed low-resolution frames to obtain original-resolution reconstructed frames. When network bandwidth improves, it can then encode high-resolution (HR) original frames.
[0127] FIG4 is a schematic diagram illustrating the principle of RPR provided in an embodiment of the present application.
[0128] As shown in FIG4 , the encoding end downsamples the original image, and the decoding end upsamples the reconstructed image.
[0129] Downsampling principle: For an image with a size of M×N, the downsampling multiple is s, that is, a resolution image with a size of (M / s)×(N / s) is obtained. If the image is in matrix form, the image in the s×s window of the original image is converted into a pixel. The interpolation method can be used, that is, based on the original image pixels, a suitable interpolation algorithm is used to calculate the surrounding pixels to obtain new pixels.
[0130] Upsampling principle: Image enlargement can be done by using the interpolation method, that is, inserting new pixels between pixels based on the original image pixels using a suitable interpolation algorithm.
[0131] (2) Neural network super resolution (NNSR).
[0132] NNSR refers to the use of neural networks to perform super-resolution reconstruction of images, enhancing clarity and detail. This method achieves super-resolution reconstruction by training neural networks to learn image features. Common neural networks include convolutional neural networks and recurrent neural networks. In image super-resolution schemes, neural networks can be used for super-resolution reconstruction of single images and for joint super-resolution reconstruction of multiple images. This method has been widely applied in image processing, medical image analysis, security monitoring, and other fields, and continues to develop and improve.
[0133] FIG5 is a schematic diagram of the principle of NNSR provided in an embodiment of the present application.
[0134] As shown in Figure 5, in addition to the low-resolution reconstructed image You can also use low-resolution prediction images The frame-level quantization parameter (QP) and base QP are also input into NNSR. The dashed box represents the auxiliary input. The reconstructed image after RPR serves as the input to NNSR, ultimately resulting in a high-resolution reconstructed image after resolution conversion.
[0135] FIG6 is a schematic diagram of the grid structure used by the NNSR provided in an embodiment of the present application.
[0136] As shown in Figure 6, the network structure may include a convolutional layer, a parameterized ReLU (Parametric Rectified Linear Unit, PReLU), N residual blocks, and a pixel shuffle layer (PixelShuffle). Specifically, the convolutional layer, PReLU, and residual block can be used to extract effective features from the input, and the pixel shuffle layer outputs a high-resolution residual based on the extracted effective features. Finally, the reconstructed image can be obtained based on the high-resolution residual output by the pixel shuffle layer and the RPR. To generate high-resolution reconstructed images
[0137] The block size in the input image can be 3×3, and the convolution kernel size of the convolution layer can be k×k.
[0138] Residual connections are a type of skip connection commonly used in deep neural networks. Their main idea is to add skip connections to the network, allowing for more direct information transfer. Specifically, residual connections can directly add the output of one layer to the input of a subsequent layer, allowing for more direct information transfer within the network. This prevents information loss and distortion within the network, helps address issues like vanishing gradients, and improves network training efficiency and performance.
[0139] The pixel shuffling layer is a neural network layer used for super-resolution image reconstruction. It converts a low-resolution input image into a high-resolution image while maintaining the perceived quality. As shown in Figure 7, the main idea of the pixel shuffling layer is to increase image resolution through pixel reorganization. This is achieved by dividing the input image's channels into several small groups, then recombining the pixels in each group into a larger image block according to specific rules. This method effectively reduces the number of parameters and improves model training efficiency.
[0140] The following describes the encoding method and decoding method provided by the embodiments of the present application in conjunction with the accompanying drawings. The encoding method provided by the embodiments of the present application can be performed by an encoding end, such as the encoder 200 shown in Figure 1 or Figure 2. The decoding method provided by the embodiments of the present application can be performed by a decoding end, such as the decoder 300 described in Figure 1 or Figure 3. The encoding end and the decoding end can be implemented by software, hardware, or a combination thereof. When implemented by hardware, the encoding end can be referred to as an encoding end device or a video encoding device, and the decoding end can be referred to as a decoding end device or a decoding device.
[0141] FIG8 is a schematic flowchart of a decoding method 410 according to an embodiment of the present application.
[0142] As shown in FIG8 , the decoding method 410 may include at least part of the following:
[0143] S411: The decoding end decodes the code stream to obtain a first reconstructed image.
[0144] S412: The decoding end obtains a first identifier by decoding the code stream; wherein the first identifier indicates that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method.
[0145] Exemplarily, the first identifier may also be referred to as super-resolution reconstruction method information or super-resolution reconstruction method indication information.
[0146] Exemplarily, the first identifier indicates that the first reconstructed image uses a first resolution conversion method, which means that the first identifier is a picture-level (or frame-level) identifier, i.e., the decoding end may apply the first identifier to all reconstructed blocks in the first reconstructed image. The first identifier indicates that the first reconstructed block uses a first resolution conversion method, which means that the first identifier is a block-level identifier, i.e., the decoding end will only apply the first identifier to the first reconstructed block.
[0147] S413: The decoding end performs resolution conversion based on the first resolution conversion method.
[0148] Exemplarily, the first identifier indicates that the first reconstructed image uses a first resolution conversion method, and the decoding end performs resolution conversion on the first reconstructed image based on the first resolution conversion method to obtain a second reconstructed image.
[0149] Exemplarily, the first identifier indicates that the first reconstructed block uses a first resolution conversion method, and the decoding end performs resolution conversion on the first reconstructed block based on the first resolution conversion method to obtain a second reconstructed block.
[0150] In an embodiment of the present application, a first identifier is used to indicate that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method, so that the decoding end can perform resolution conversion based on the first resolution conversion method, which not only improves the flexibility of resolution conversion, but also helps to improve the resolution conversion effect.
[0151] In some embodiments, the S412 includes:
[0152] The decoding end obtains a second identifier by decoding the code stream; when the second identifier indicates that the first reconstructed image is allowed to undergo resolution conversion, the decoding end obtains the first identifier by decoding the code stream, and the first identifier indicates that the first reconstructed block uses the first resolution conversion method.
[0153] Exemplarily, the decoding end obtains a second identifier by decoding the code stream; when the second identifier indicates that the first reconstructed image is allowed to undergo resolution conversion, the decoding end obtains the first identifier by decoding the code stream, and the first identifier indicates that the first reconstructed block uses the first resolution conversion method; when the second identifier indicates that the first reconstructed image is not allowed to undergo resolution conversion, the decoding end does not perform resolution conversion on the reconstructed block in the first reconstructed image, or the decoding end may directly use the first reconstructed image as the decoded image.
[0154] In this embodiment, by introducing the second identifier, when all reconstructed blocks in the first reconstructed image do not need to undergo resolution conversion, the decoding end can avoid decoding the identifier at the image block level for indicating the resolution conversion method, and the encoding end can avoid encoding the identifier at the image block level for indicating the resolution conversion method, thereby improving the encoding and decoding performance.
[0155] In some embodiments, the decoding end obtains a third identifier by decoding the code stream; when the third identifier indicates that resolution conversion of the reconstructed image sequence to which the first reconstructed image belongs is allowed, the decoding end obtains the second identifier by decoding the code stream.
[0156] Exemplarily, the decoding end obtains a third identifier by decoding the code stream; when the third identifier indicates that resolution conversion is allowed for the reconstructed image sequence to which the first reconstructed image belongs, the decoding end obtains the second identifier by decoding the code stream; when the third identifier indicates that resolution conversion is not allowed for the reconstructed image sequence to which the first reconstructed image belongs, the decoding end does not perform resolution conversion on the reconstructed image in the reconstructed image sequence, or the decoding end may directly use the reconstructed image sequence as the decoded image sequence.
[0157] In this embodiment, by introducing the third identifier, when the reconstructed image sequence does not require resolution conversion, the decoding end can avoid decoding the image-level identifier for indicating whether the resolution conversion mode is allowed to be converted, and the encoding end can avoid encoding the image-level identifier for indicating whether the resolution conversion mode is allowed to be converted, thereby improving the encoding and decoding performance.
[0158] In some embodiments, the S412 includes:
[0159] The decoding end obtains a third identifier by decoding the code stream; when the third identifier indicates that the reconstructed image sequence to which the first reconstructed image belongs is allowed to perform resolution conversion, the decoding end obtains the first identifier by decoding the code stream, and the first identifier indicates that the first reconstructed image uses the first resolution conversion method.
[0160] Exemplarily, the decoding end obtains a third identifier by decoding the code stream; when the third identifier indicates that the reconstructed image sequence to which the first reconstructed image belongs is allowed to perform resolution conversion, the decoding end obtains the first identifier by decoding the code stream, and the first identifier indicates that the first reconstructed image uses the first resolution conversion method; when the third identifier indicates that the reconstructed image sequence to which the first reconstructed image belongs is not allowed to perform resolution conversion, the decoding end does not perform resolution conversion on the reconstructed image in the reconstructed image sequence, or the decoding end can directly use the reconstructed image sequence as the decoded image sequence.
[0161] In this embodiment, by introducing the second identifier, when all reconstructed blocks in the first reconstructed image do not need to undergo resolution conversion, the decoding end can avoid decoding the identifier at the image block level for indicating the resolution conversion method, and the encoding end can avoid encoding the identifier at the image block level for indicating the resolution conversion method, thereby improving the encoding and decoding performance.
[0162] In some embodiments, the S412 includes:
[0163] When at least one of the following conditions is met, the decoding end obtains the first identifier by decoding the code stream:
[0164] The size of the first reconstructed image or the size of the first reconstructed block is a predefined size;
[0165] The image type of the first reconstructed image or the sequence type of the reconstructed image sequence to which the first reconstructed image belongs is a predefined type.
[0166] Exemplarily, when the size of the first reconstructed image is a predefined size, the decoding end obtains the first identifier by decoding the code stream, and the first identifier indicates that the first reconstructed image uses the first resolution conversion method; when the size of the first reconstructed image is not a predefined size, the decoding end does not perform resolution conversion on the reconstructed block in the first reconstructed image, or the decoding end uses the first reconstructed image as the decoded image. Similarly, when the size of the first reconstructed block is a predefined size, the decoding end obtains the first identifier by decoding the code stream, and the first identifier indicates that the first reconstructed block uses the first resolution conversion method; when the size of the first reconstructed block is not a predefined size, the decoding end does not perform resolution conversion on the first reconstructed block, or the decoding end can directly use the first reconstructed block as the decoded image block.
[0167] Illustratively, the predefined size may include one or more sizes.
[0168] Exemplarily, when the image type of the first reconstructed image or the sequence type of the reconstructed image sequence to which the first reconstructed image belongs is a predefined type, the decoding end obtains the first identifier by decoding the code stream; when the image type of the first reconstructed image or the sequence type of the reconstructed image sequence to which the first reconstructed image belongs is not a predefined type, the decoding end does not perform resolution conversion on the first reconstructed image or the first reconstructed block, or the decoding end can directly use the first reconstructed image as the decoded image, or the decoding end can directly use the first reconstructed block as the decoded image block.
[0169] Exemplarily, the predefined type includes but is not limited to at least one of the following: application type, scenario type, prediction type used in the prediction process, etc.
[0170] In this embodiment, when at least one of the following conditions is met, the decoding end obtains the first identifier by decoding the code stream, thereby avoiding the decoding end from decoding the identifier at the image level or sequence level for indicating whether the resolution conversion mode is allowed to be converted, and avoiding the encoding end from encoding the identifier at the image level or sequence level for indicating whether the resolution conversion mode is allowed to be converted, thereby improving the encoding and decoding performance.
[0171] In some embodiments, the resolution of the first reconstructed block is smaller than the resolution of the second reconstructed block obtained after resolution conversion of the first reconstructed block, or the resolution of the first reconstructed image is smaller than the resolution of the second reconstructed image obtained after resolution conversion of the first reconstructed image.
[0172] Exemplarily, the decoding end performs super-resolution conversion on the first reconstructed block based on the first resolution conversion method to obtain a second reconstructed block. The resolution of the first reconstructed block is lower than the resolution of the second reconstructed block. Super-resolution conversion is an image processing technology designed to improve image detail and clarity.
[0173] Exemplarily, the decoding end performs super-resolution conversion on the first reconstructed image based on the first resolution conversion method to obtain a second reconstructed image. The resolution of the first reconstructed image is lower than the resolution of the second reconstructed image. Super-resolution conversion is an image processing technology designed to improve image detail and clarity.
[0174] Of course, in other alternative embodiments, the resolution of the first reconstructed block may be greater than the resolution of the second reconstructed block, or the resolution of the first reconstructed image may be greater than the resolution of the second reconstructed image obtained after resolution conversion of the first reconstructed image. This application does not impose any specific limitations on this.
[0175] In some embodiments, the first resolution conversion method includes at least one of the following: reference image resampling RPR, neural network-based super-resolution NNSR.
[0176] Exemplarily, the first resolution conversion method may be RPR. In this case, the decoding end performs resolution conversion on the first reconstructed image or the first reconstructed block based on RPR. For example, the decoding end may obtain the second reconstructed image by upsampling the first reconstructed image. For another example, the decoding end may obtain the second reconstructed block by upsampling the first reconstructed block.
[0177] Exemplarily, the first resolution conversion method may be NNSR. In this case, the decoder performs resolution conversion on the first reconstructed image or the first reconstructed block based on the NNSR. For example, the decoder may use the NNSR to predict a first residual image having a resolution greater than that of the first reconstructed image, and then determine the second reconstructed image based on the first residual image. For example, the decoder may add the first residual image to a reconstructed image obtained by upsampling the first reconstructed image using RPR to obtain the second reconstructed image. Optionally, when using the NNSR to predict the first residual image, the decoder may also input at least one of the predicted image corresponding to the first reconstructed image, the frame-level (also known as slice-level) QP, and the base QP corresponding to the first reconstructed image as auxiliary information into the NNSR. For another example, the decoder may use the NNSR to predict a first residual block having a resolution greater than that of the first reconstructed block, and then determine the second reconstructed block based on the first residual block. For example, the decoder may add the first residual block to a reconstructed block obtained by upsampling the first reconstructed block using RPR to obtain the second reconstructed block. Optionally, when the decoding end uses the NNSR to predict the first residual block, it can also input at least one of the prediction block corresponding to the first reconstructed block, the frame-level (also called slice-level) QP and the basic QP corresponding to the first reconstructed block as auxiliary information into the NNSR.
[0178] FIG9 is a schematic flowchart of an encoding method 420 according to an embodiment of the present application.
[0179] As shown in FIG9 , the encoding method 420 may include at least part of the following:
[0180] S421: The encoding end reconstructs the first original image to obtain a first reconstructed image.
[0181] Exemplarily, the encoder downsamples the first original image to obtain a low-resolution image having a resolution lower than that of the first original image, and then reconstructs the obtained low-resolution image to obtain the first reconstructed image. For example, the encoder encodes and then decodes the first-resolution image to obtain the first reconstructed image.
[0182] S422: The encoding end encodes a first identifier, where the first identifier indicates that the first reconstructed image or a first reconstructed block in the first reconstructed image uses a first resolution conversion method.
[0183] Exemplarily, the encoding end downsamples the first original image to obtain a low-resolution image having a resolution smaller than that of the first original image, and after encoding and decoding the obtained low-resolution image to obtain a first reconstructed image, a reconstructed block in the first reconstructed image corresponding to the first original block in the first original image can be used as the first reconstructed block.
[0184] Exemplarily, the encoding end downsamples the first original block in the first original image to obtain a low-resolution block having a resolution lower than that of the first original block, and then encodes and decodes the obtained low-resolution block to obtain the first reconstructed block.
[0185] In some embodiments, before S422, the method 420 further includes:
[0186] The encoder performs resolution conversion on the first reconstructed block using multiple resolution conversion modes to obtain multiple candidate resampled blocks corresponding to the multiple resolution conversion modes; the encoder determines the loss of each of the multiple candidate resampled blocks based on the multiple candidate resampled blocks and a first original block in the first original image corresponding to the first reconstructed block; and the encoder determines, based on the loss of each candidate resampled block, the resolution conversion mode corresponding to the candidate resampled block with the smallest loss as the first resolution conversion mode.
[0187] Exemplarily, when the encoding end uses multiple resolution conversion methods to perform resolution conversion on the first reconstructed block, the multiple resolution conversion methods can be used to convert the first reconstructed block from low resolution to high resolution to obtain the multiple candidate resampling blocks.
[0188] Exemplarily, the loss of each candidate resampled block is used to represent the quality difference between the first original block and each candidate resampled block.
[0189] Exemplarily, the loss of each candidate resampling block includes but is not limited to: Sum of Absolute Differences (SAD), Sum of Squared Differences (SSD), Structural Similarity Index (SSIM), etc.
[0190] In some embodiments, the S422 includes:
[0191] The encoding end encodes the second identifier; when the second identifier indicates that the first reconstructed image is allowed to perform resolution conversion, the encoding end encodes the first identifier, and the first identifier indicates that the first reconstructed block uses the first resolution conversion method.
[0192] In some embodiments, the encoding end encodes a third identifier; and when the third identifier indicates that resolution conversion is allowed for a reconstructed image sequence to which the first reconstructed image belongs, the encoding end encodes the second identifier.
[0193] In some embodiments, before S422, the method 420 further includes:
[0194] The encoder performs resolution conversion on the first reconstructed image using a plurality of resolution conversion modes to obtain a plurality of candidate resampled images corresponding to the plurality of resolution conversion modes; the encoder determines the loss of each of the plurality of candidate resampled images based on the plurality of candidate resampled images and the first original image; and the encoder determines, based on the loss of each candidate resampled image, the resolution conversion mode corresponding to the candidate resampled image with the smallest loss as the first resolution conversion mode.
[0195] Exemplarily, when the encoding end uses multiple resolution conversion methods to convert the resolution of the first reconstructed image, the multiple resolution conversion methods can be used to convert the first reconstructed image from low resolution to high resolution to obtain the multiple candidate resampled images.
[0196] Exemplarily, the loss of each candidate resampled image is used to represent the quality difference between the first original image and each candidate resampled image.
[0197] Exemplarily, the loss of each candidate resampled image includes but is not limited to: SAD, SSD, SSIM, etc.
[0198] In some embodiments, the S422 includes:
[0199] The encoding end encodes the third identifier; when the third identifier indicates that the reconstructed image sequence to which the first reconstructed image belongs is allowed to perform resolution conversion, the encoding end encodes the first identifier, and the first identifier indicates that the first reconstructed image uses the first resolution conversion method.
[0200] In some embodiments, the S422 includes:
[0201] The encoding end encodes the first identifier when at least one of the following conditions is met:
[0202] The size of the first reconstructed image or the size of the first reconstructed block is a predefined size;
[0203] The image type of the first reconstructed image or the sequence type of the reconstructed image sequence to which the first reconstructed image belongs is a predefined type.
[0204] In some embodiments, the method 420 further includes:
[0205] The encoding end converts the resolution of the first reconstructed block based on the first resolution conversion mode to obtain a second reconstructed block.
[0206] In some embodiments, the resolution of the first reconstructed block is smaller than the resolution of the second reconstructed block obtained after resolution conversion of the first reconstructed block, or the resolution of the first reconstructed image is smaller than the resolution of the second reconstructed image obtained after resolution conversion of the first reconstructed image.
[0207] In some embodiments, the first resolution conversion method includes at least one of the following: reference image resampling RPR, neural network-based super-resolution NNSR.
[0208] It should be understood that the encoding method can be understood as the inverse process of the decoding method (or called the reverse process). Therefore, the specific scheme of the encoding method 420 can refer to the relevant content of the decoding method 410. For the sake of ease of description, this application will not go into details.
[0209] The specific embodiments provided in this application are described below.
[0210] Example 1:
[0211] In this embodiment, the decoding end obtains a first reconstructed image by decoding the code stream; the decoding end obtains a first identifier by decoding the code stream; wherein the first identifier indicates that the first reconstructed image uses a first resolution conversion method; the decoding end performs resolution conversion on the first reconstructed image based on the first resolution conversion method to obtain a second reconstructed image.
[0212] Taking the first resolution conversion mode as RPR or NNSR as an example, it may include the following process:
[0213] Step 1:
[0214] The decoding end decodes the bitstream to obtain a first reconstructed image, and obtains a first flag (e.g., flag 1) for the first reconstructed image by decoding the bitstream, the first flag indicating whether to use RPR or NNSR. For example, if flag 1 is 0, RPR is used, and if flag 1 is 1, NNSR is used.
[0215] Step 2:
[0216] The decoding end processes the first reconstructed image based on the first identifier to obtain a high-resolution reconstructed image (i.e., a second reconstructed image). For example, if the value of the first identifier is 0, the first reconstructed image is processed using RPR to obtain a high-resolution reconstructed image; if the value of the first identifier is 1, the first reconstructed image is processed using NNSR to obtain a high-resolution reconstructed image.
[0217] In this embodiment, a first identifier is used to indicate that the first reconstructed image uses a first resolution conversion method, so that the decoding end can perform resolution conversion on the first reconstructed image based on the first resolution conversion method to obtain a high-resolution second reconstructed image, which not only improves the flexibility of resolution conversion, but also helps to improve the resolution conversion effect.
[0218] Example 2:
[0219] Step 1:
[0220] In this embodiment, a decoding end obtains a first reconstructed image by decoding a code stream; the decoding end obtains a first identifier by decoding the code stream; wherein the first identifier indicates that a first reconstructed block in the first reconstructed image uses a first resolution conversion method; the decoding end performs resolution conversion on the first reconstructed block based on the first resolution conversion method to obtain a second reconstructed block.
[0221] Taking the first resolution conversion mode as RPR or NNSR as an example, it may include the following process:
[0222] The decoding end decodes the bitstream to obtain a first reconstructed image, and obtains a first flag (e.g., flag 1) for a first reconstructed block in the first reconstructed image by decoding the bitstream. The first flag indicates whether to use RPR or NNSR. For example, if flag 1 is 0, RPR is used, and if flag 1 is 1, NNSR is used.
[0223] Step 2:
[0224] The decoding end processes the first reconstructed block based on the first identifier to obtain a high-resolution reconstructed block (i.e., a second reconstructed block). For example, if the value of the first identifier is 0, the first reconstructed block is processed using RPR to obtain a high-resolution reconstructed block; if the value of the first identifier is 1, the first reconstructed block is processed using NNSR to obtain a high-resolution reconstructed block.
[0225] Similarly, the decoding end can process each block in the low-resolution reconstructed image according to the flag indicating the resolution conversion mode of each block in the low-resolution reconstructed image to obtain a high-resolution reconstructed image. For example, if the value of the flag indicating the resolution conversion mode of the current block is 0, the current block is processed using RPR to obtain a high-resolution reconstructed block; if the value of the flag indicating the resolution conversion mode of the current block is 1, the current block is processed using NNSR to obtain a high-resolution reconstructed block.
[0226] In this embodiment, a first identifier is used to indicate that the first reconstructed block uses a first resolution conversion method, so that the decoding end can perform resolution conversion on the first reconstructed block based on the first resolution conversion method to obtain a high-resolution second reconstructed block, which not only improves the flexibility of resolution conversion, but also helps to improve the resolution conversion effect.
[0227] Example 3:
[0228] In this embodiment, an encoder downsamples a first original image to obtain a low-resolution image having a resolution lower than that of the first original image, encodes and decodes the obtained low-resolution image to obtain a first reconstructed image, and then converts the first reconstructed image from the low resolution to a high resolution using multiple resolution conversion modes to obtain multiple candidate resampled images corresponding to the multiple resolution conversion modes. The encoder determines a loss for each of the multiple candidate resampled images based on the multiple candidate resampled images and the first original image. Based on the loss of each candidate resampled image, the encoder determines the resolution conversion mode corresponding to the candidate resampled image with the smallest loss as the first resolution conversion mode, and writes a first flag into a bitstream, the first flag indicating that the first reconstructed image uses the first resolution conversion mode.
[0229] The following uses multiple resolution conversion methods including RPR and NNSR as an example, which may include the following process:
[0230] Step 1:
[0231] The encoding end downsamples the first original image to obtain a low-resolution image with a resolution lower than that of the first original image, encodes and decodes the obtained low-resolution image to obtain a first reconstructed image, and then uses RPR to convert the first reconstructed image from low resolution to high resolution to obtain a first candidate resampled image; the encoding end determines the loss of the first candidate resampled image based on the first candidate resampled image and the first original image; the loss includes but is not limited to: SAD, SSD, SSIM, etc. of each pixel point in the two images, which is not limited in this application.
[0232] Step 2:
[0233] The encoding end uses NNSR to convert the first reconstructed image from low resolution to high resolution to obtain a second candidate resampled image; the encoding end determines the loss of the second candidate resampled image based on the second candidate resampled image and the first original image; the loss includes but is not limited to: SAD, SSD, SSIM, etc. of each pixel point in the two images, which is not limited in this application.
[0234] Step 3:
[0235] The encoder selects a resolution conversion mode adopted by the candidate resampled image with less loss from the first candidate resampled image and the second candidate resampled image as a first resolution conversion mode used for the first reconstructed image, and writes a first identifier into the bitstream, the first identifier indicating that the first reconstructed image uses the first resolution conversion mode.
[0236] In this embodiment, for the first reconstructed image, the encoding end can select the first resolution conversion method with the best resolution conversion effect through loss comparison, and indicate that the first reconstructed image uses the first resolution conversion method through a first identifier, so that the decoding end can perform resolution conversion on the first reconstructed image based on the first resolution conversion method to obtain a high-resolution second reconstructed image, which not only improves the flexibility of resolution conversion, but also helps to improve the resolution conversion effect.
[0237] It is worth noting that Example 3 corresponds to Example 1, that is, the encoder can adopt the solution of Example 3 and select the resolution conversion method used for the low-resolution reconstructed image. Correspondingly, the decoder can adopt Example 1 to perform resolution conversion on the low-resolution reconstructed image to obtain a high-resolution reconstructed image.
[0238] Example 4:
[0239] In this embodiment, the encoder downsamples the first original image to obtain a low-resolution image having a resolution lower than that of the first original image, and encodes and decodes the obtained low-resolution image to obtain a first reconstructed image. Then, a first reconstructed block in the first reconstructed image corresponding to the first original block in the first original image is converted from low resolution to high resolution using multiple resolution conversion methods to obtain multiple candidate resampled blocks corresponding to the multiple resolution conversion methods (or, the encoder downsamples the first original block in the first original image to obtain a low-resolution block having a resolution lower than that of the first original block, and encodes and decodes the obtained low-resolution block). After decoding to obtain a first reconstructed block, the first reconstructed block is converted from a low resolution to a high resolution using multiple resolution conversion modes to obtain multiple candidate resampled blocks corresponding to the multiple resolution conversion modes; the encoder determines, based on the multiple candidate resampled blocks and the first original block, a loss for each of the multiple candidate resampled blocks; the encoder determines, based on the loss of each candidate resampled block, the resolution conversion mode corresponding to the candidate resampled block with the smallest loss as the first resolution conversion mode, and writes a first identifier into a bitstream, the first identifier indicating that the first reconstructed block uses the first resolution conversion mode.
[0240] The following uses multiple resolution conversion methods including RPR and NNSR as an example, which may include the following process:
[0241] Step 1:
[0242] The encoding end downsamples the first original block to obtain a low-resolution block having a resolution smaller than that of the first original block, and encodes and decodes the obtained low-resolution block to obtain a first reconstructed block (or the encoding end downsamples the first original image to obtain a low-resolution image having a resolution smaller than that of the first original image, and encodes and decodes the obtained low-resolution image to obtain a first reconstructed block in the first reconstructed image corresponding to the first original block in the first original image), and then uses RPR to convert the first reconstructed block from low resolution to high resolution to obtain a first candidate resampled block; the encoding end determines the loss of the first candidate resampled block based on the first candidate resampled block and the first original block; the loss includes but is not limited to: SAD, SSD, SSIM, etc. of each pixel point in the two blocks, which is not limited in this application.
[0243] Step 2:
[0244] The encoding end uses NNSR to convert the first reconstructed block from low resolution to high resolution to obtain a second candidate resampled block; the encoding end determines the loss of the second candidate resampled block based on the second candidate resampled block and the first original block; the loss includes but is not limited to: SAD, SSD, SSIM, etc. of each pixel point in the two blocks, which is not limited in this application.
[0245] Step 3:
[0246] The encoder selects a resolution conversion mode adopted by the candidate resampling block with less loss from among the first candidate resampling blocks and the second candidate resampling blocks, and uses the selected resolution conversion mode as a first resolution conversion mode used for a first reconstructed block. The encoder also writes a first identifier into a bitstream, the first identifier indicating that the first reconstructed block uses the first resolution conversion mode.
[0247] In this embodiment, for the first reconstructed block, the encoding end can select the first resolution conversion method with the best resolution conversion effect through loss comparison, and instruct the first reconstructed block to use the first resolution conversion method through a first identifier, so that the decoding end can perform resolution conversion on the first reconstructed block based on the first resolution conversion method to obtain a high-resolution second reconstructed block, which not only improves the flexibility of resolution conversion, but also helps to improve the resolution conversion effect.
[0248] It is worth noting that Example 4 corresponds to Example 2, that is, the encoder can adopt the solution of Example 4 and select the resolution conversion method used for the low-resolution reconstructed block. Correspondingly, the decoder can adopt Example 2 to perform resolution conversion on the low-resolution reconstructed block to obtain a high-resolution reconstructed block.
[0249] FIG10 is a schematic flowchart of a decoding method 510 according to an embodiment of the present application.
[0250] As shown in FIG10 , the decoding method 510 may include at least part of the following:
[0251] S511: The decoding end decodes the bit stream to obtain a first reconstructed image;
[0252] S512: The decoding end determines a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0253] Exemplarily, the decoding end determines the first resolution conversion method based on the information of the first reconstructed image or the information of the first reconstructed block in the first reconstructed image, which is intended to illustrate that the first resolution conversion method can be determined based on the information of the first reconstructed image or the information of the first reconstructed block in the first reconstructed image in a manner derived by the decoding end.
[0254] S513: The decoding end performs resolution conversion based on the first resolution conversion method.
[0255] Exemplarily, the decoding end determines a first resolution conversion method based on information of the first reconstructed image, and the decoding end performs resolution conversion on the first reconstructed image based on the first resolution conversion method to obtain a second reconstructed image.
[0256] Exemplarily, the decoding end determines a first resolution conversion mode based on information of the first reconstructed block, and the decoding end performs resolution conversion on the first reconstructed block based on the first resolution conversion mode to obtain a second reconstructed block.
[0257] In an embodiment of the present application, the decoding end determines a first resolution conversion method based on information of the first reconstructed image or information of the first reconstructed block in the first reconstructed image, so that the decoding end can perform resolution conversion based on the first resolution conversion method, which not only improves the flexibility of resolution conversion, but also helps to improve the resolution conversion effect.
[0258] In some embodiments, the S512 includes:
[0259] The decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed image.
[0260] In some embodiments, the decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed image, including at least one of the following:
[0261] When the size of the first reconstructed image is less than or equal to a preset threshold, the decoding end determines the first type of resolution conversion mode as the first resolution conversion mode;
[0262] When the size of the first reconstructed image is larger than a preset threshold, the decoding end determines the second type of resolution conversion mode as the first resolution conversion mode;
[0263] The decoding end determines the resolution conversion mode corresponding to the size of the first reconstructed image as the first resolution conversion mode.
[0264] For example, the decoding end may determine the first resolution conversion mode based on the size of the first reconstructed image. In one implementation, when the size of the first reconstructed image is less than or equal to a preset threshold, the decoding end determines the first type of resolution conversion mode as the first resolution conversion mode. When the size of the first reconstructed image is greater than the preset threshold, the decoding end determines the second type of resolution conversion mode as the first resolution conversion mode. In another implementation, the decoding end may determine the first resolution conversion mode as the resolution conversion mode corresponding to the size of the first reconstructed image.
[0265] In some embodiments, the decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed image, including:
[0266] The decoding end determines the resolution conversion mode corresponding to the type of the first reconstructed image as the first resolution conversion mode.
[0267] For example, the decoding end may determine the first resolution conversion mode based on the type of the first reconstructed image. In one implementation, when the type of the first reconstructed image is a first preset type, the decoding end determines the first type of resolution conversion mode as the first resolution conversion mode, and when the type of the first reconstructed image is a second preset type, the decoding end determines the second type of resolution conversion mode as the first resolution conversion mode. In another implementation, the decoding end may determine the resolution conversion mode corresponding to the type of the first reconstructed image as the first resolution conversion mode. The type of the first reconstructed image includes but is not limited to at least one of the following: application type, scene type, prediction type used in the prediction process, etc. The prediction type includes but is not limited to intra-frame prediction type and inter-frame prediction type. Alternatively, the intra-frame prediction type may be represented as a specific prediction mode, such as the number or index of the intra-frame prediction mode.
[0268] In some embodiments, the S512 includes:
[0269] The decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed block.
[0270] In some embodiments, the decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed block, including at least one of the following:
[0271] When the size of the first reconstructed block is less than or equal to a preset threshold, the decoding end determines the first type of resolution conversion mode as the first resolution conversion mode;
[0272] When the size of the first reconstructed block is larger than a preset threshold, the decoding end determines the second type of resolution conversion mode as the first resolution conversion mode;
[0273] The decoding end determines the resolution conversion mode corresponding to the size of the first reconstructed block as the first resolution conversion mode.
[0274] For example, the decoding end may determine the first resolution conversion mode based on the size of the first reconstructed block. In one implementation, when the size of the first reconstructed block is less than or equal to a preset threshold, the decoding end determines the first type of resolution conversion mode as the first resolution conversion mode. When the size of the first reconstructed block is greater than the preset threshold, the decoding end determines the second type of resolution conversion mode as the first resolution conversion mode. In another implementation, the decoding end may determine the first resolution conversion mode as the resolution conversion mode corresponding to the size of the first reconstructed block.
[0275] In some embodiments, the decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed block, including:
[0276] The decoding end determines the resolution conversion mode corresponding to the type of the first reconstructed block as the first resolution conversion mode.
[0277] For example, the decoding end may determine the first resolution conversion mode based on the type of the first reconstructed block. In one implementation, when the type of the first reconstructed block is a first preset type, the decoding end determines the first type of resolution conversion mode as the first resolution conversion mode, and when the type of the first reconstructed block is a second preset type, the decoding end determines the second type of resolution conversion mode as the first resolution conversion mode. In another implementation, the decoding end may determine the resolution conversion mode corresponding to the type of the first reconstructed block as the first resolution conversion mode. The type of the first reconstructed block includes but is not limited to: application type, scene type, prediction type used in the prediction process, etc. The prediction type includes but is not limited to intra-frame prediction type and inter-frame prediction type. Alternatively, the intra-frame prediction type may be represented as a specific prediction mode, such as the number or index of the intra-frame prediction mode.
[0278] In some embodiments, the resolution of the first reconstructed block is smaller than the resolution of the second reconstructed block obtained after resolution conversion of the first reconstructed block, or the resolution of the first reconstructed image is smaller than the resolution of the second reconstructed image obtained after resolution conversion of the first reconstructed image.
[0279] Exemplarily, the decoding end performs super-resolution conversion on the first reconstructed block based on the first resolution conversion method to obtain a second reconstructed block. Super-resolution conversion is an image processing technology that aims to improve image detail and clarity. In this case, the resolution of the first reconstructed block is smaller than the resolution of the second reconstructed block obtained after the resolution conversion of the first reconstructed block, or the resolution of the first reconstructed image is smaller than the resolution of the second reconstructed image obtained after the resolution conversion of the first reconstructed image.
[0280] Of course, in other alternative embodiments, the resolution of the first reconstructed block may be greater than the resolution of the second reconstructed block, or the resolution of the first reconstructed image may be greater than the resolution of the second reconstructed image obtained after resolution conversion of the first reconstructed image. This application does not impose any specific limitations on this.
[0281] In some embodiments, the first resolution conversion method includes at least one of the following: reference image resampling RPR, neural network-based super-resolution NNSR.
[0282] Exemplarily, the first resolution conversion method may be RPR. In this case, the decoding end performs resolution conversion on the first reconstructed image or the first reconstructed block based on RPR. For example, the decoding end may obtain the second reconstructed image by upsampling the first reconstructed image. For another example, the decoding end may obtain the second reconstructed block by upsampling the first reconstructed block.
[0283] Exemplarily, the first resolution conversion method may be NNSR. In this case, the decoder performs resolution conversion on the first reconstructed image or the first reconstructed block based on the NNSR. For example, the decoder may use the NNSR to predict a first residual image having a resolution greater than that of the first reconstructed image, and then determine the second reconstructed image based on the first residual image. For example, the decoder may add the first residual image to a reconstructed image obtained by upsampling the first reconstructed image using RPR to obtain the second reconstructed image. Optionally, when using the NNSR to predict the first residual image, the decoder may also input at least one of the predicted image corresponding to the first reconstructed image, the frame-level (also known as slice-level) QP, and the base QP corresponding to the first reconstructed image as auxiliary information into the NNSR. For another example, the decoder may use the NNSR to predict a first residual block having a resolution greater than that of the first reconstructed block, and then determine the second reconstructed block based on the first residual block. For example, the decoder may add the first residual block to a reconstructed block obtained by upsampling the first reconstructed block using RPR to obtain the second reconstructed block. Optionally, when the decoding end uses the NNSR to predict the first residual block, it can also input at least one of the prediction block corresponding to the first reconstructed block, the frame-level (also called slice-level) QP and the basic QP corresponding to the first reconstructed block as auxiliary information into the NNSR.
[0284] FIG11 is a schematic flowchart of an encoding method 520 according to an embodiment of the present application.
[0285] As shown in FIG11 , the encoding method 520 may include at least part of the following:
[0286] S521: The encoder reconstructs the first original image to obtain a first reconstructed image.
[0287] Exemplarily, the encoder downsamples the first original image to obtain a low-resolution image having a resolution lower than that of the first original image, and then reconstructs the obtained low-resolution image to obtain the first reconstructed image. For example, the encoder encodes and then decodes the first-resolution image to obtain the first reconstructed image.
[0288] S522: The encoder determines a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0289] Exemplarily, the encoding end downsamples the first original image to obtain a low-resolution image having a resolution smaller than that of the first original image, and after encoding and decoding the obtained low-resolution image to obtain a first reconstructed image, a reconstructed block in the first reconstructed image corresponding to the first original block in the first original image can be used as the first reconstructed block.
[0290] Exemplarily, the encoding end downsamples the first original block in the first original image to obtain a low-resolution block having a resolution lower than that of the first original block, and then encodes and decodes the obtained low-resolution block to obtain the first reconstructed block.
[0291] S523: The encoding end performs resolution conversion based on the first resolution conversion method.
[0292] In some embodiments, the S522 includes:
[0293] The encoding end determines the first resolution conversion mode based on the size or type of the first reconstructed image.
[0294] In some embodiments, the encoder determines the first resolution conversion mode based on the size or type of the first reconstructed image, including at least one of the following:
[0295] When the size of the first reconstructed image is smaller than or equal to a preset threshold, the encoder determines the first type of resolution conversion mode as the first resolution conversion mode;
[0296] When the size of the first reconstructed image is larger than a preset threshold, the encoder determines the second type of resolution conversion mode as the first resolution conversion mode;
[0297] The encoding end determines the resolution conversion mode corresponding to the size of the first reconstructed image as the first resolution conversion mode.
[0298] In some embodiments, the encoder determines the first resolution conversion mode based on the size or type of the first reconstructed image, including:
[0299] The encoding end determines the resolution conversion mode corresponding to the type of the first reconstructed image as the first resolution conversion mode.
[0300] In some embodiments, the S522 includes:
[0301] The encoding end determines the first resolution conversion mode based on the size or type of the first reconstructed block.
[0302] In some embodiments, the encoder determines the first resolution conversion mode based on the size or type of the first reconstructed block, including at least one of the following:
[0303] When the size of the first reconstructed block is less than or equal to a preset threshold, the encoder determines the first type of resolution conversion mode as the first resolution conversion mode;
[0304] When the size of the first reconstructed block is larger than a preset threshold, the encoder determines the second type of resolution conversion mode as the first resolution conversion mode;
[0305] The encoding end determines the resolution conversion mode corresponding to the size of the first reconstructed block as the first resolution conversion mode.
[0306] In some embodiments, the encoder determines the first resolution conversion mode based on the size or type of the first reconstructed block, including:
[0307] The encoding end determines the resolution conversion mode corresponding to the type of the first reconstructed block as the first resolution conversion mode.
[0308] It should be understood that the encoding method can be understood as the inverse process of the decoding method (or called the reverse process). Therefore, the specific scheme of the encoding method 520 can refer to the relevant content of the decoding method 510. For the sake of ease of description, this application will not go into details.
[0309] The decoding method provided in the embodiments of the present application may be executed by a decoding device. In the embodiments of the present application, a decoding device executing the decoding method is used as an example to illustrate the decoding device provided in the embodiments of the present application. The encoding method provided in the embodiments of the present application may be executed by an encoding device. In the embodiments of the present application, an encoding device executing the encoding method is used as an example to illustrate the encoding device provided in the embodiments of the present application.
[0310] FIG12 is a schematic block diagram of a decoding device 610 provided according to an embodiment of the present application.
[0311] As shown in FIG12 , the decoding device 610 includes:
[0312] The decoding unit 611 is configured to:
[0313] Obtaining a first reconstructed image by decoding the code stream;
[0314] Obtaining a first identifier by decoding the code stream;
[0315] The first identifier indicates that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method;
[0316] The conversion unit 612 is configured to perform resolution conversion based on the first resolution conversion method.
[0317] In some embodiments, the decoding unit 611 is specifically configured to:
[0318] Obtaining a second identifier by decoding the code stream;
[0319] In a case where the second flag indicates that resolution conversion is allowed for the first reconstructed image, the first flag is obtained by decoding the code stream, and the first flag indicates that the first reconstructed block uses the first resolution conversion method.
[0320] In some embodiments, the decoding unit 611 is specifically configured to:
[0321] Obtaining a third identifier by decoding the code stream;
[0322] In a case where the third identifier indicates that resolution conversion is allowed for the reconstructed image sequence to which the first reconstructed image belongs, the second identifier is obtained by decoding the code stream.
[0323] In some embodiments, the decoding unit 611 is specifically configured to:
[0324] Obtaining a third identifier by decoding the code stream;
[0325] When the third identifier indicates that resolution conversion is allowed for the reconstructed image sequence to which the first reconstructed image belongs, the first identifier is obtained by decoding the code stream, and the first identifier indicates that the first reconstructed image uses the first resolution conversion method.
[0326] In some embodiments, the decoding unit 611 is specifically configured to:
[0327] The first identifier is obtained by decoding the code stream when at least one of the following conditions is met:
[0328] The size of the first reconstructed image or the size of the first reconstructed block is a predefined size;
[0329] The image type of the first reconstructed image or the sequence type of the reconstructed image sequence to which the first reconstructed image belongs is a predefined type.
[0330] In some embodiments, the resolution of the first reconstructed block is smaller than the resolution of the second reconstructed block obtained after resolution conversion of the first reconstructed block, or the resolution of the first reconstructed image is smaller than the resolution of the second reconstructed image obtained after resolution conversion of the first reconstructed image.
[0331] In some embodiments, the first resolution conversion method includes at least one of the following: reference image resampling RPR, neural network-based super-resolution NNSR.
[0332] It should be understood that the decoding device 610 provided in the embodiment of the present application may correspond to the execution subject in the method embodiment of the present application, and the various units in the decoding device 610 are respectively for implementing the corresponding processes of the decoding method 410 shown in Figure 8. For the sake of brevity, they will not be repeated here.
[0333] The decoding device provided in the embodiment of the present application can implement the various processes implemented in the method embodiment of Figure 8 and achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0334] FIG13 is a schematic block diagram of an encoding device 620 provided according to an embodiment of the present application.
[0335] As shown in FIG13 , the encoding device 620 includes:
[0336] The reconstruction unit 621 is configured to reconstruct the first original image to obtain a first reconstructed image;
[0337] The encoding unit 622 is configured to encode a first identifier, where the first identifier indicates that the first reconstructed image or a first reconstructed block in the first reconstructed image uses a first resolution conversion method.
[0338] In some embodiments, before encoding the first identifier, the encoding unit 622 is further configured to:
[0339] Performing resolution conversion on the first reconstructed block using multiple resolution conversion methods to obtain multiple candidate resampled blocks corresponding to the multiple resolution conversion methods;
[0340] determining a loss of each of the plurality of candidate resampled blocks based on the plurality of candidate resampled blocks and a first original block in the first original image corresponding to the first reconstructed block;
[0341] Based on the loss of each candidate resampling block, the resolution conversion mode corresponding to the candidate resampling block with the smallest loss is determined as the first resolution conversion mode.
[0342] In some embodiments, the encoding unit 622 is specifically configured to:
[0343] encoding the second identifier;
[0344] In a case where the second flag indicates that the first reconstructed image is allowed to undergo resolution conversion, the first flag is encoded, and the first flag indicates that the first reconstructed block uses the first resolution conversion method.
[0345] In some embodiments, the encoding unit 622 is specifically configured to:
[0346] encoding the third identifier;
[0347] In a case where the third flag indicates that resolution conversion is allowed for the reconstructed image sequence to which the first reconstructed image belongs, the second flag is encoded.
[0348] In some embodiments, before encoding the first identifier, the encoding unit 622 is further configured to:
[0349] Performing resolution conversion on the first reconstructed image using multiple resolution conversion methods to obtain multiple candidate resampled images corresponding to the multiple resolution conversion methods;
[0350] determining a loss for each of the plurality of candidate resampled images based on the plurality of candidate resampled images and the first original image;
[0351] Based on the loss of each candidate resampled image, a resolution conversion mode corresponding to the candidate resampled image with the smallest loss is determined as the first resolution conversion mode.
[0352] In some embodiments, the encoding unit 622 is specifically configured to:
[0353] encoding the third identifier;
[0354] In a case where the third flag indicates that the reconstructed image sequence to which the first reconstructed image belongs is allowed to perform resolution conversion, the first flag is encoded, and the first flag indicates that the first reconstructed image uses the first resolution conversion mode.
[0355] In some embodiments, the encoding unit 622 is specifically configured to:
[0356] The first identifier is encoded when at least one of the following conditions is met:
[0357] The size of the first reconstructed image or the size of the first reconstructed block is a predefined size;
[0358] The image type of the first reconstructed image or the sequence type of the reconstructed image sequence to which the first reconstructed image belongs is a predefined type.
[0359] In some embodiments, the encoding unit 622 is further configured to:
[0360] Based on the first resolution conversion method, the resolution of the first reconstructed block is converted to obtain a second reconstructed block.
[0361] In some embodiments, the resolution of the first reconstructed block is smaller than the resolution of the second reconstructed block obtained after resolution conversion of the first reconstructed block, or the resolution of the first reconstructed image is smaller than the resolution of the second reconstructed image obtained after resolution conversion of the first reconstructed image.
[0362] In some embodiments, the first resolution conversion method includes at least one of the following: reference image resampling RPR, neural network-based super-resolution NNSR.
[0363] It should be understood that the encoding device 620 provided in the embodiment of the present application may correspond to the execution entity in the method embodiment of the present application, and the various units in the encoding device 620 are respectively for implementing the corresponding processes of the encoding method 420 shown in Figure 9. For the sake of brevity, they will not be repeated here.
[0364] The encoding device provided in the embodiment of the present application can implement the various processes implemented in the method embodiment of Figure 9 and achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0365] FIG14 is a schematic block diagram of a decoding device 710 provided according to an embodiment of the present application.
[0366] As shown in FIG14 , the decoding device 710 includes:
[0367] The decoding unit 711 is configured to obtain a first reconstructed image by decoding the code stream;
[0368] The determining unit 712 is configured to determine a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0369] The conversion unit 713 is configured to perform resolution conversion based on the first resolution conversion method.
[0370] In some embodiments, the determining unit 712 is specifically configured to:
[0371] The first resolution conversion mode is determined based on the size or type of the first reconstructed image.
[0372] In some embodiments, the determining unit 712 is specifically configured to perform at least one of the following:
[0373] When the size of the first reconstructed image is less than or equal to a preset threshold, determining the first type of resolution conversion mode as the first resolution conversion mode;
[0374] When the size of the first reconstructed image is greater than a preset threshold, determining the second type of resolution conversion mode as the first resolution conversion mode;
[0375] The resolution conversion mode corresponding to the size of the first reconstructed image is determined as the first resolution conversion mode.
[0376] In some embodiments, the determining unit 712 is specifically configured to:
[0377] A resolution conversion mode corresponding to the type of the first reconstructed image is determined as the first resolution conversion mode.
[0378] In some embodiments, the determining unit 712 is specifically configured to:
[0379] The first resolution conversion mode is determined based on the size or type of the first reconstructed block.
[0380] In some embodiments, the determining unit 712 is specifically configured to perform at least one of the following:
[0381] When the size of the first reconstructed block is smaller than or equal to a preset threshold, determining the first type of resolution conversion mode as the first resolution conversion mode;
[0382] When the size of the first reconstructed block is greater than a preset threshold, determining the second type of resolution conversion mode as the first resolution conversion mode;
[0383] A resolution conversion mode corresponding to the size of the first reconstructed block is determined as the first resolution conversion mode.
[0384] In some embodiments, the determining unit 712 is specifically configured to:
[0385] A resolution conversion mode corresponding to the type of the first reconstructed block is determined as the first resolution conversion mode.
[0386] It should be understood that the decoding device 710 provided in the embodiment of the present application may correspond to the execution subject in the method embodiment of the present application, and the various units in the decoding device 710 are respectively for implementing the corresponding processes of the decoding method 510 shown in Figure 10. For the sake of brevity, they will not be repeated here.
[0387] The decoding device provided in the embodiment of the present application can implement the various processes implemented in the method embodiment of Figure 10 and achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0388] FIG15 is a schematic block diagram of an encoding device 720 provided according to an embodiment of the present application.
[0389] As shown in FIG15 , the encoding device 720 includes:
[0390] The reconstruction unit 721 is configured to reconstruct the first original image to obtain a first reconstructed image;
[0391] The determining unit 722 is configured to determine a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0392] The conversion unit 723 is configured to perform resolution conversion based on the first resolution conversion method.
[0393] In some embodiments, the determining unit 722 is specifically configured to:
[0394] The first resolution conversion mode is determined based on the size or type of the first reconstructed image.
[0395] In some embodiments, the determining unit 722 is specifically configured to perform at least one of the following:
[0396] When the size of the first reconstructed image is less than or equal to a preset threshold, determining the first type of resolution conversion mode as the first resolution conversion mode;
[0397] When the size of the first reconstructed image is greater than a preset threshold, determining the second type of resolution conversion mode as the first resolution conversion mode;
[0398] The resolution conversion mode corresponding to the size of the first reconstructed image is determined as the first resolution conversion mode.
[0399] In some embodiments, the determining unit 722 is specifically configured to:
[0400] A resolution conversion mode corresponding to the type of the first reconstructed image is determined as the first resolution conversion mode.
[0401] In some embodiments, the determining unit 722 is specifically configured to:
[0402] The first resolution conversion mode is determined based on the size or type of the first reconstructed block.
[0403] In some embodiments, the determining unit 722 is specifically configured to perform at least one of the following:
[0404] When the size of the first reconstructed block is smaller than or equal to a preset threshold, determining the first type of resolution conversion mode as the first resolution conversion mode;
[0405] When the size of the first reconstructed block is greater than a preset threshold, determining the second type of resolution conversion mode as the first resolution conversion mode;
[0406] A resolution conversion mode corresponding to the size of the first reconstructed block is determined as the first resolution conversion mode.
[0407] In some embodiments, the determining unit 722 is specifically configured to:
[0408] A resolution conversion mode corresponding to the type of the first reconstructed block is determined as the first resolution conversion mode.
[0409] It should be understood that the encoding device 720 provided in the embodiment of the present application may correspond to the execution entity in the method embodiment of the present application, and the various units in the encoding device 720 are respectively for implementing the corresponding processes of the encoding method 520 shown in Figure 11. For the sake of brevity, they will not be repeated here.
[0410] The encoding device provided in the embodiment of the present application can implement the various processes implemented in the method embodiment of Figure 11 and achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0411] An embodiment of the present application further provides an electronic device 800 , as shown in FIG16 , including a processor 801 and a memory 802 , where the memory 802 stores programs or instructions that can be run on the processor 801 .
[0412] For example, when the electronic device 800 is a decoding end, the program or instruction is executed by the processor 801 to implement the various steps of the above-mentioned decoding method embodiment, and can achieve the same technical effect. When the electronic device 800 is an encoding end, the program or instruction is executed by the processor 801 to implement the various steps of the above-mentioned encoding method embodiment, and can achieve the same technical effect. To avoid repetition, it is not repeated here. Optionally, the memory 802 can be the memory 102 or the memory 113 in the embodiment shown in Figure 1, and the processor 801 can implement the functions of the encoder 200 or the decoder 300 in the embodiment shown in Figures 1-3.
[0413] The present application also provides an electronic device including: a memory configured to store video data; and a processing circuit configured to implement the steps of the above-described decoding method embodiment or encoding method embodiment. Optionally, the memory may be memory 102 or memory 113 in the embodiment shown in FIG1 , and the processing circuit may implement the functions of encoder 200 or decoder 300 in the embodiments shown in FIG1-3 .
[0414] The present application also provides an electronic device including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to execute a program or instruction to implement the steps of the above-described decoding method embodiment or encoding method embodiment. This device embodiment corresponds to the above-described method embodiment, and each implementation process and implementation method of the above-described method embodiment is applicable to this terminal embodiment and can achieve the same technical effects.
[0415] The electronic device may be a terminal, or may be other devices other than a terminal, such as a server, a network attached storage (NAS), etc.
[0416] Among them, the terminal can be a mobile phone, tablet personal computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile Internet device (MID), augmented reality (AR), virtual reality (VR) equipment, mixed reality (MR) equipment, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipborne equipment, pedestrian user equipment (PUE), smart home (home appliances with wireless communication function, such as refrigerator, TV, washing machine or furniture, etc.), game console, personal computer (PC), ATM or self-service machine and other terminal-side devices. Wearable devices include: smart watches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart bracelets, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among them, vehicle-mounted devices can also be called vehicle-mounted terminals, vehicle-mounted controllers, vehicle-mounted modules, vehicle-mounted components, vehicle-mounted chips, or vehicle-mounted units, etc. It should be noted that the specific type of terminal is not limited in the embodiments of this application.
[0417] The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.
[0418] For example, the electronic device may include but is not limited to the source device 100 or the destination device 110 shown in FIG. 1 .
[0419] Taking an electronic device as a terminal as an example, FIG17 is a schematic diagram of the hardware structure of a terminal 900 for implementing an embodiment of the present application.
[0420] As shown in Figure 17, the terminal 900 includes but is not limited to: a radio frequency unit 901, a network module 902, an audio output unit 903, an input unit 904, a sensor 905, a display unit 906, a user input unit 907, an interface unit 908, a memory 909 and at least some of the components of the processor 910.
[0421] Those skilled in the art will appreciate that the terminal 900 may further include a power supply (such as a battery) for powering various components. The power supply may be logically connected to the processor 910 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The terminal structure shown in FIG17 does not limit the terminal. The terminal may include more or fewer components than shown, or may combine certain components, or have different component arrangements, which will not be described in detail here.
[0422] It should be understood that in an embodiment of the present application, the input unit 904 may include a graphics processing unit (GPU) 9041 and a microphone 9042. The graphics processor 9041 processes image data of a static picture or video obtained by an image acquisition device (such as a camera) in a video acquisition mode or an image acquisition mode, or may process the obtained point cloud data. The display unit 906 may include a display panel 9061, which may be configured in the form of a liquid crystal display, an organic light emitting diode, or the like. The user input unit 907 includes a touch panel 9071 and at least one of other input devices 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 may include two parts: a touch detection device and a touch controller. Other input devices 9072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.
[0423] In the embodiment of the present application, after receiving data from a peer, the RF unit 901 may transmit the data to the processor 910 for processing. Furthermore, the RF unit 901 may send data to the peer. Typically, the RF unit 901 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and the like.
[0424] The memory 909 can be used to store software programs or instructions and various data. The memory 909 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 909 may include a volatile memory or a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 909 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0425] Processor 910 may include one or more processing units. Optionally, processor 910 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 910.
[0426] As an implementation manner, the terminal 900 is a decoding end, and the processor 910 is configured to:
[0427] Obtaining a first reconstructed image by decoding the code stream;
[0428] Obtaining a first identifier by decoding the code stream;
[0429] The first identifier indicates that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method;
[0430] Resolution conversion is performed based on the first resolution conversion method.
[0431] As another implementation, the terminal 900 is an encoding end, and the processor 910 is configured to:
[0432] Reconstructing the first original image to obtain a first reconstructed image;
[0433] A first identifier is encoded, where the first identifier indicates that the first reconstructed image or a first reconstructed block in the first reconstructed image uses a first resolution conversion method.
[0434] In an embodiment of the present application, a first identifier is used to indicate that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method, so that the decoding end can perform resolution conversion based on the first resolution conversion method, which not only improves the flexibility of resolution conversion, but also helps to improve the resolution conversion effect.
[0435] As an implementation manner, the terminal 900 is a decoding end, and the processor 910 is configured to:
[0436] Obtaining a first reconstructed image by decoding the code stream;
[0437] Determining a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0438] Resolution conversion is performed based on the first resolution conversion method.
[0439] As another implementation, the terminal 900 is an encoding end, and the processor 910 is configured to:
[0440] Reconstructing the first original image to obtain a first reconstructed image;
[0441] Determining a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image:
[0442] Resolution conversion is performed based on the first resolution conversion method.
[0443] In an embodiment of the present application, the decoding end determines a first resolution conversion method based on information of the first reconstructed image or information of the first reconstructed block in the first reconstructed image, so that the decoding end can perform resolution conversion based on the first resolution conversion method, which not only improves the flexibility of resolution conversion, but also helps to improve the resolution conversion effect.
[0444] It can be understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the method embodiment and achieve the same or corresponding technical effects. To avoid repetition, it will not be described here.
[0445] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned decoding method embodiment or the above-mentioned encoding method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0446] The processor is the processor in the terminal described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as ROM, RAM, a magnetic disk, or an optical disk. In some examples, the readable storage medium may be a non-transitory readable storage medium.
[0447] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned decoding method embodiment or the above-mentioned encoding method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0448] It should be understood that the chip mentioned in the embodiments of the present application may include a system-level chip (also known as a system chip, a chip system or a system-on-chip chip), and may also include an independent display chip, etc.
[0449] An embodiment of the present application further provides a computer program / program product, which is stored in a storage medium. The computer program / program product is executed by at least one processor to implement the various processes of the above-mentioned decoding method embodiment or the above-mentioned encoding method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0450] An embodiment of the present application also provides a coding and decoding system, including: an encoding end and a decoding end, wherein the encoding end can be used to execute the steps of the encoding method described above, such as executing the steps of method 420 or method 520 described above, and the decoding end can be used to execute the steps of the decoding method described above, such as executing the steps of method 410 or method 510 described above.
[0451] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0452] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of a computer software product plus the necessary general hardware platform, or of course, by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes a number of instructions for causing the decoding end to execute the decoding method described in each embodiment of the present application or the encoding end to execute the encoding method described in each embodiment of the present application.
[0453] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms of implementation methods without departing from the purpose of this application and the scope of protection of the claims. These implementation methods are all within the protection of this application.
Claims
1. A decoding method, wherein: include: The decoding end obtains a first reconstructed image by decoding the code stream; The decoding end obtains a first identifier by decoding the code stream; The first identifier indicates that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method; The decoding end performs resolution conversion based on the first resolution conversion method.
2. The method according to claim 1, wherein The decoding end obtains a first identifier by decoding the code stream, including: The decoding end obtains a second identifier by decoding the code stream; When the second flag indicates that the first reconstructed image is allowed to undergo resolution conversion, the decoding end obtains the first flag by decoding the code stream, where the first flag indicates that the first reconstructed block uses the first resolution conversion method.
3. The method according to claim 2, wherein: The decoding end obtains a second identifier by decoding the code stream, including: The decoding end obtains a third identifier by decoding the code stream; In a case where the third identifier indicates that resolution conversion is allowed for the reconstructed image sequence to which the first reconstructed image belongs, the decoding end obtains the second identifier by decoding the code stream.
4. The method according to claim 1, wherein The decoding end obtains a first identifier by decoding the code stream, including: The decoding end obtains a third identifier by decoding the code stream; When the third identifier indicates that resolution conversion is allowed for the reconstructed image sequence to which the first reconstructed image belongs, the decoding end obtains the first identifier by decoding the code stream, and the first identifier indicates that the first reconstructed image uses the first resolution conversion method.
5. The method according to any one of claims 1 to 4, wherein The decoding end obtains a first identifier by decoding the code stream, including: When at least one of the following conditions is met, the decoding end obtains the first identifier by decoding the code stream: The size of the first reconstructed image or the size of the first reconstructed block is a predefined size; The image type of the first reconstructed image or the sequence type of the reconstructed image sequence to which the first reconstructed image belongs is a predefined type.
6. The method according to any one of claims 1 to 5, wherein The resolution of the first reconstructed block is smaller than the resolution of a second reconstructed block obtained after resolution conversion of the first reconstructed block, or the resolution of the first reconstructed image is smaller than the resolution of a second reconstructed image obtained after resolution conversion of the first reconstructed image.
7. The method according to any one of claims 1 to 6, wherein The first resolution conversion method includes at least one of the following: reference image resampling RPR, neural network-based super-resolution NNSR.
8. A coding method, wherein: include: The encoding end reconstructs the first original image to obtain a first reconstructed image; The encoding end encodes a first identifier, where the first identifier indicates that the first reconstructed image or a first reconstructed block in the first reconstructed image uses a first resolution conversion method.
9. The method according to claim 8, wherein Before the encoding end encodes the first identifier, the method further includes: The encoder performs resolution conversion on the first reconstructed block using multiple resolution conversion modes to obtain multiple candidate resampled blocks corresponding to the multiple resolution conversion modes; Determining, by the encoder, a loss of each of the plurality of candidate resampled blocks based on the plurality of candidate resampled blocks and a first original block in the first original image corresponding to the first reconstructed block; The encoder determines, based on the loss of each candidate resampling block, a resolution conversion mode corresponding to the candidate resampling block with the smallest loss as the first resolution conversion mode.
10. The method according to claim 8 or 9, wherein: The encoding end encodes the first identifier, including: The encoding end encodes the second identifier; In a case where the second flag indicates that the first reconstructed image is allowed to perform resolution conversion, the encoding end encodes the first flag, where the first flag indicates that the first reconstructed block uses the first resolution conversion mode.
11. The method according to claim 10, wherein: The encoding end encodes the second identifier, including: The encoding end encodes the third identifier; When the third identifier indicates that resolution conversion is allowed for the reconstructed image sequence to which the first reconstructed image belongs, the encoding end encodes the second identifier.
12. The method according to claim 8, wherein Before the encoding end encodes the first identifier, the method further includes: The encoding end performs resolution conversion on the first reconstructed image using multiple resolution conversion methods to obtain multiple candidate resampled images corresponding to the multiple resolution conversion methods; determining, by the encoder, a loss of each of the plurality of candidate resampled images based on the plurality of candidate resampled images and the first original image; The encoder determines, based on the loss of each candidate resampled image, a resolution conversion mode corresponding to the candidate resampled image with the smallest loss as the first resolution conversion mode.
13. The method according to claim 8 or 12, wherein: The encoding end encodes the first identifier, including: The encoding end encodes the third identifier; When the third flag indicates that resolution conversion is allowed for the reconstructed image sequence to which the first reconstructed image belongs, the encoding end encodes the first flag, where the first flag indicates that the first reconstructed image uses the first resolution conversion mode.
14. The method according to any one of claims 8 to 13, wherein The encoding end encodes the first identifier, including: The encoding end encodes the first identifier when at least one of the following conditions is met: The size of the first reconstructed image or the size of the first reconstructed block is a predefined size; The image type of the first reconstructed image or the sequence type of the reconstructed image sequence to which the first reconstructed image belongs is a predefined type.
15. The method according to any one of claims 8 to 14, wherein The method further comprises: The encoding end converts the resolution of the first reconstructed image based on the first resolution conversion method to obtain a second reconstructed image, or the encoding end converts the resolution of the first reconstructed block based on the first resolution conversion method to obtain a second reconstructed block.
16. The method according to any one of claims 8 to 15, wherein The resolution of the first reconstructed block is smaller than the resolution of a second reconstructed block obtained after resolution conversion of the first reconstructed block, or the resolution of the first reconstructed image is smaller than the resolution of a second reconstructed image obtained after resolution conversion of the first reconstructed image.
17. The method according to any one of claims 8 to 16, wherein The first resolution conversion method includes at least one of the following: reference image resampling RPR, neural network-based super-resolution NNSR.
18. A decoding method, wherein: include: The decoding end obtains a first reconstructed image by decoding the code stream; The decoding end determines a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image: The decoding end performs resolution conversion based on the first resolution conversion method.
19. The method according to claim 18, wherein The decoding end determines a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image, including: The decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed image.
20. The method according to claim 19, wherein The decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed image, including at least one of the following: When the size of the first reconstructed image is less than or equal to a preset threshold, the decoding end determines the first type of resolution conversion mode as the first resolution conversion mode; When the size of the first reconstructed image is larger than a preset threshold, the decoding end determines the second type of resolution conversion mode as the first resolution conversion mode; The decoding end determines the resolution conversion mode corresponding to the size of the first reconstructed image as the first resolution conversion mode.
21. The method according to claim 19 or 20, wherein The decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed image, including: The decoding end determines the resolution conversion mode corresponding to the type of the first reconstructed image as the first resolution conversion mode.
22. The method according to claim 18, wherein The decoding end determines a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image, including: The decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed block.
23. The method according to claim 22, wherein The decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed block, including at least one of the following: When the size of the first reconstructed block is less than or equal to a preset threshold, the decoding end determines the first type of resolution conversion mode as the first resolution conversion mode; When the size of the first reconstructed block is larger than a preset threshold, the decoding end determines the second type of resolution conversion mode as the first resolution conversion mode; The decoding end determines the resolution conversion mode corresponding to the size of the first reconstructed block as the first resolution conversion mode.
24. The method according to claim 22 or 23, wherein The decoding end determines the first resolution conversion mode based on the size or type of the first reconstructed block, including: The decoding end determines the resolution conversion mode corresponding to the type of the first reconstructed block as the first resolution conversion mode.
25. A coding method, wherein: include: The encoding end reconstructs the first original image to obtain a first reconstructed image; The encoder determines, based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image, a first resolution conversion mode: The encoding end performs resolution conversion based on the first resolution conversion method.
26. The method according to claim 25, wherein The encoder determines, based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image, a first resolution conversion mode, including: The encoding end determines the first resolution conversion mode based on the size or type of the first reconstructed image.
27. The method according to claim 26, wherein The encoder determines, based on the size or type of the first reconstructed image, the first resolution conversion mode, including at least one of the following: When the size of the first reconstructed image is smaller than or equal to a preset threshold, the encoder determines the first type of resolution conversion mode as the first resolution conversion mode; When the size of the first reconstructed image is larger than a preset threshold, the encoder determines the second type of resolution conversion mode as the first resolution conversion mode; The encoding end determines the resolution conversion mode corresponding to the size of the first reconstructed image as the first resolution conversion mode.
28. The method according to claim 26 or 27, wherein The encoder determines the first resolution conversion mode based on the size or type of the first reconstructed image, including: The encoding end determines the resolution conversion mode corresponding to the type of the first reconstructed image as the first resolution conversion mode.
29. The method according to claim 25, wherein The encoder determines, based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image, a first resolution conversion mode, including: The encoding end determines the first resolution conversion mode based on the size or type of the first reconstructed block.
30. The method according to claim 29, wherein The encoder determines, based on the size or type of the first reconstructed block, the first resolution conversion mode, including at least one of the following: When the size of the first reconstructed block is less than or equal to a preset threshold, the encoder determines the first type of resolution conversion mode as the first resolution conversion mode; When the size of the first reconstructed block is larger than a preset threshold, the encoder determines the second type of resolution conversion mode as the first resolution conversion mode; The encoding end determines the resolution conversion mode corresponding to the size of the first reconstructed block as the first resolution conversion mode.
31. The method according to claim 29 or 30, wherein The encoder determines, based on the size or type of the first reconstructed block, the first resolution conversion mode, including: The encoding end determines the resolution conversion mode corresponding to the type of the first reconstructed block as the first resolution conversion mode.
32. A decoding device, wherein: include: Decoding unit, used to: Obtaining a first reconstructed image by decoding the code stream; Obtaining a first identifier by decoding the code stream; The first identifier indicates that the first reconstructed image or the first reconstructed block in the first reconstructed image uses a first resolution conversion method; A conversion unit is configured to perform resolution conversion based on the first resolution conversion method.
33. The apparatus according to claim 32, wherein The decoding unit is specifically configured to: Obtaining a second identifier by decoding the code stream; In a case where the second flag indicates that resolution conversion is allowed for the first reconstructed image, the first flag is obtained by decoding the code stream, and the first flag indicates that the first reconstructed block uses the first resolution conversion method.
34. The apparatus of claim 32, wherein: The decoding unit is specifically configured to: Obtaining a third identifier by decoding the code stream; When the third identifier indicates that resolution conversion is allowed for the reconstructed image sequence to which the first reconstructed image belongs, the first identifier is obtained by decoding the code stream, and the first identifier indicates that the first reconstructed image uses the first resolution conversion method.
35. An encoding device, wherein: include: a reconstruction unit, configured to reconstruct the first original image to obtain a first reconstructed image; The encoding unit is configured to encode a first identifier, where the first identifier indicates that the first reconstructed image or a first reconstructed block in the first reconstructed image uses a first resolution conversion method.
36. The apparatus of claim 35, wherein: Before encoding the first identifier, the encoding unit is further configured to: Performing resolution conversion on the first reconstructed block using multiple resolution conversion methods to obtain multiple candidate resampled blocks corresponding to the multiple resolution conversion methods; determining a loss of each of the plurality of candidate resampled blocks based on the plurality of candidate resampled blocks and a first original block in the first original image corresponding to the first reconstructed block; Based on the loss of each candidate resampling block, the resolution conversion mode corresponding to the candidate resampling block with the smallest loss is determined as the first resolution conversion mode.
37. The apparatus of claim 35, wherein: Before encoding the first identifier, the encoding unit is further configured to: Performing resolution conversion on the first reconstructed image using multiple resolution conversion methods to obtain multiple candidate resampled images corresponding to the multiple resolution conversion methods; determining a loss for each of the plurality of candidate resampled images based on the plurality of candidate resampled images and the first original image; Based on the loss of each candidate resampled image, a resolution conversion mode corresponding to the candidate resampled image with the smallest loss is determined as the first resolution conversion mode.
38. A decoding device, wherein: include: A decoding unit, configured to obtain a first reconstructed image by decoding the code stream; a determining unit, configured to determine a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image: A conversion unit is configured to perform resolution conversion based on the first resolution conversion method.
39. The apparatus according to claim 38, wherein The determining unit is specifically configured to: The first resolution conversion mode is determined based on the size or type of the first reconstructed image.
40. The apparatus of claim 38, wherein The determining unit is specifically configured to: The first resolution conversion mode is determined based on the size or type of the first reconstructed block.
41. An encoding device, wherein: include: a reconstruction unit, configured to reconstruct the first original image to obtain a first reconstructed image; a determining unit, configured to determine a first resolution conversion mode based on information of the first reconstructed image or information of a first reconstructed block in the first reconstructed image: A conversion unit is configured to perform resolution conversion based on the first resolution conversion method.
42. The apparatus according to claim 41, wherein The determining unit is specifically configured to: The encoding end determines the first resolution conversion mode based on the size or type of the first reconstructed image.
43. The apparatus according to claim 41, wherein The determining unit is specifically configured to: The first resolution conversion mode is determined based on the size or type of the first reconstructed block.
44. An electronic device, wherein: The invention comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the program or instruction implements the steps of the decoding method according to any one of claims 1 to 7, or implements the steps of the encoding method according to any one of claims 8 to 17, or implements the steps of the decoding method according to any one of claims 18 to 24, or implements the steps of the encoding method according to any one of claims 25 to 31.
45. A readable storage medium, wherein: The readable storage medium stores a program or instruction, which, when executed by a processor, implements the steps of the decoding method according to any one of claims 1 to 7, or implements the steps of the encoding method according to any one of claims 8 to 17, or implements the steps of the decoding method according to any one of claims 18 to 24, or implements the steps of the encoding method according to any one of claims 25 to 31.
46. A chip, wherein The chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run a program or instructions to implement the steps of the decoding method according to any one of claims 1 to 7, or implement the steps of the encoding method according to any one of claims 8 to 17, or implement the steps of the decoding method according to any one of claims 18 to 24, or implement the steps of the encoding method according to any one of claims 25 to 31.