Image coding method, image decoding method andapparatus
Patent Information
- Application Number
- HK62026127345
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-07-28
- Filing Date
- 2026-08-11
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-04-22
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(12) International application published under the Patent Cooperation Treaty (19) International Bureau of WIPO (43) International Publication Date: 31 October 2024 (31.10.2024) WIPO I PCT (51) International Patent Classification: H04N19 / 184 (2014.01) G06N 3 / 02 (2006.01) (21) International Application Number: (22) International Application Date: (25) Application Language: (26) Publication Language: (30) Priority: 202310476967.6 202310956879.6 (71) Applicant: Huawei PCT / CN2024 / 089342 23 April 2024 (23.04.2024) Chinese 24 April 2023 (24.04.2023) CN July 28, 2023 (28.07.2023) CN Technology Co., Ltd. (HUAWEI = _-_======== = = = == === TECHNOLOGIES CO., LTD.) [CN / CN]; Huawei Headquarters Office Building, Bantian, Longgang District, Shenzhen, Guangdong 518129 (CN)0 (72) Inventors: Yu Dequan (YU, Dequan); Huawei Headquarters Office Building, Bantian, Longgang District, Shenzhen, Guangdong 518129 (CN)0 Zhao Yin (ZHAO, Yin); Huawei Headquarters Office Building, Bantian, Longgang District, Shenzhen, Guangdong 518129 (CN)0 (10) International Publication No.: WO 2024 / 222681 Al Shenzhen, Longgang District, Bantian, Huawei Headquarters Office Building, Guangdong 518129 (CN)0 Alshina Elena Alexandrovna (ALSHINA, Elena) Alexandrovna); Huawei Headquarters Office Building, Bantian, Longgang District, Shenzhen, Guangdong 518129, China (CN) 0 (74) Agent: Beijing Tongda Xinheng Intellectual Property Agency Co., Ltd. (TDIP & PARTNERS); Room 2002, Block A, Beihuan Center, No. 18 Yumin Road, Xicheng District, Beijing 100029, China (CN) 0 (81) Designated Country (unless otherwise specified, each requiring available national protection): AE, AG, AL, AM, AO, AT, AU, AZ, BA, BB, BG, BH, BN, BR, BW, BY, BZ, CA, CH, CL, CN, CO, CR, CU, CV, CZ, DE, DJ, DK, DM, DO, DZ, EC, EE, EG,ES, FI, GB, GD, GE, GH, GM, GT, HN, HR, HU, ID, IL, IN, IQ, IR, IS, IT, JM, JO, JP, KE, KG, KH, KN, KP, KR, KW, KZ, LA, LC, LK, LR, LS, LU, LY, MA, MD, MG, MK, MN, MU, MW, MX, MY, MZ, NA, NG, NI, NO, NZ, OM, PA, PE, PG, PH, PL, PT, QA, RO, RS, RU, RW, SA, SC, SD, (54) Title: IMAGE CODING METHOD, IMAGE DECODING METHOD AND APPARATUS (54) Invention Title: An Image Coding and Decoding Method and Apparatus 501 / 501 Encode into a code stream identifier information used for indicating a decoding network to be used 503 Decode from a received code stream the identifier information used for indicating the decoding network to be used 504 When the identifier information is a first value, use a first decoding network to decode from the code stream an image to be processed; or, when the identifier information is a second value, use a second decoding network to decode from the code stream said image AA Coder BB Decoder (57) Abstract: An image coding method, an image decoding method and an apparatus, relatingThis invention relates to the fields of artificial intelligence and image compression, and provides a coding and decoding solution to meet the requirements of different application scenarios. The coding and decoding methods provided by this application can determine the coding and decoding networks to be used based on profile information (or identifier information). Specifically, a coder and a decoder can select corresponding profiles in formation according to the capability of a decoding device, thus selecting or indicating different coding and decoding networks. Therefore, the application can adapt to end-sides requiring low computational power, and also adapt to end-sides requiring higher computational power. (57) Abstract: This invention relates to the fields of artificial intelligence and image compression, and provides a coding and decoding solution to meet the requirements of different application scenarios. The coding and decoding methods provided by this application can determine the coding and decoding networks used based on profile information (or identifier information). In other words, the codec can select the appropriate level of information based on the capabilities of the decoding device, thereby selecting or indicating different codec networks. This allows it to adapt to both low-computing-power devices and devices requiring higher computing power. [See continued page] WO 2024 / 222681Al IIIIIIIIIIIIIIIIIIIIIIIIIIIIM SE, SG, SK, SL, ST, SY SY TH, TJ, TM, TN, TR, TT, TZ, UA, UG, US, UZ, VC, VN, WS, ZA, ZM, ZWO (84) Designated countries (unless otherwise specified, each of the available regional protections is required): ARIPO (BW, CV, GH, GM, KE, LR, LS, MW, MZ, NA, RW, SC, SD, SL, ST, SZ, TZ, UG, ZM, ZW), Eurasia (AM, AZ, BY, KG, KZ, RU, TJ, TM), Europe (AL, AT, BE, BG, CH, CY, CZ, DE, DK, EE, ES, FI, FR, GB, GR, HR, HU, IE, IS, IT, LT, LU, LV, MC, ME, MK, MT, NL, NO, PL, PT, RO, RS, SE, SI, SK, SM, TR), OAPI (BF, BJ, CF, CG, CI, CM, GA, GN, GQ, GW, KM, ML, MR, NE, SN, TD,TG). This international publication includes: International search report (Article 21(3) of the Treaty WO 2024 / 222681 PCT / CN2024 / 089342 5 10 15 20 25 30 35 40 Cross-reference to related applications of an image encoding and decoding method and apparatus This application claims priority to Chinese Patent Application No. 202310476967.6, filed on April 24, 2023, entitled "An image encoding and decoding method and apparatus", the entire contents of which are incorporated herein by reference; This application claims priority to Chinese Patent Application No. 202310956879.6, filed on July 28, 2023, entitled "An image encoding and decoding method and apparatus", the entire contents of which are incorporated herein by reference. Technical Field This application relates to the fields of image compression technology and artificial intelligence technology, and particularly to an image encoding and decoding method and apparatus. Background Technology Many consumer applications (such as news, social networking, and shopping internet applications) require image decoding to be completed on edge devices with relatively low computing power (such as mobile phones, personal PCs, and televisions); other industrial applications allow image decoding to be completed on edge devices with higher computing power (such as GPU workstations equipped with dedicated graphics cards), and also have higher requirements for image compression rates.Current neural network-based image encoding and decoding schemes typically have fixed network structures, failing to meet the needs of diverse application scenarios. This application provides an image encoding and decoding method and apparatus to offer an encoding and decoding scheme that meets the requirements of different application scenarios. In a first aspect, this application provides an image encoding method, comprising: encoding identification information indicating the decoding network used into a bitstream; wherein the identification information is a first value, used to indicate that the decoding network used to decode the image to be processed from the bitstream is a first decoding network; or, the identification information is a second value, used to indicate that the decoding network used to decode the image to be processed from the bitstream is a second decoding network; the first decoding network requires higher processing resources than the second decoding network; and transmitting the bitstream. The identification information can also be called profile ID. This application, through the above scheme, instructs the receiving end to adopt a network structure via the transmitting end. Different network structures can achieve different decoding performances, improving the flexibility of the decoding end. Users can adjust the computing power of their encoding and decoding network according to their own scenarios to flexibly balance latency and compression performance. In one possible implementation, the first decoding network and the second decoding network are completely different decoding networks; or, the first decoding network and the second decoding network share some sub-networks; or, the second decoding network is a sub-network of the first decoding network. If the second decoding network is a sub-network of the first decoding network, it can be understood that when the identification information is a second value, some network layers in the first decoding network are skipped, i.e., the decoding process is implemented using the second decoding network. In one possible implementation, the method further includes: obtaining the identification information; when the identification information is a first value, encoding the residual information obtained by encoding the image to be processed using the first encoding network into the bitstream; or, when the identification information is a second value, encoding the bitstream using the residual information obtained by encoding the image to be processed using the second encoding network; the processing resources required by the first encoding network are higher than those required by the second encoding network. In the above scheme, for bitstreams generated using different AI encoding networks, the decoding end can select different decoding network structures based on the bitstream content to achieve decoding. This provides greater flexibility to the encoding and decoding end, allowing users to adjust the computing power of their encoding and decoding networks according to their own scenarios to flexibly balance latency and compression performance. In one possible implementation, the first encoding network and the second encoding network are two different encoding networks; or, the first encoding network and the second encoding network share some sub-networks; or, the second encoding network is a subset of the first encoding network.The sub-network has a 30 35 40 45 sub-network configuration. In one possible implementation, the first encoding network includes a first feature extraction network, an autoregressive network, a side information extraction network, and a probability estimation network. The residual information obtained by encoding the image to be processed using the first encoding network includes: extracting a three-dimensional feature map of the image to be processed through the first feature extraction network, the three-dimensional feature map including multiple feature elements; extracting side information of the feature elements to be encoded from the three-dimensional feature map through the side information extraction network; estimating the first probability distribution mean of the feature elements to be encoded based on the side information through the probability estimation network; inputting the encoded feature elements and the first probability distribution mean into the autoregressive network to obtain a second probability distribution mean of the feature elements to be encoded; and obtaining the residual information of the feature elements to be encoded based on the feature elements to be encoded and the second probability distribution mean of the feature elements to be encoded. In one possible implementation, the first encoding network includes a second feature extraction network, a side information extraction network, and a probability estimation network. The residual information obtained by encoding the image to be processed using the first encoding network includes: extracting a three-dimensional feature map of the image to be processed through the second feature extraction network, the three-dimensional feature map including multiple feature elements; extracting side information of the feature elements to be encoded from the three-dimensional feature map through the side information extraction network; estimating the mean probability distribution of the feature elements to be encoded based on the side information through the probability estimation network; and obtaining the residual information of the feature elements to be encoded based on the feature elements to be encoded and the mean probability distribution. In one possible implementation, the second feature extraction network is a sub-network of the first feature extraction network, or the second feature extraction network and the first feature extraction network are two completely different sub-networks. In one possible implementation, the method further includes: encoding the side information into the bitstream. In one possible implementation, the identification information is located in the bitstream header. Secondly, embodiments of this application provide an image decoding method, comprising: receiving a bitstream; decoding identification information from the bitstream indicating the decoding network used; when the identification information is a first value, decoding an image to be processed from the bitstream using a first decoding network; or, when the identification information is a second value, decoding the image to be processed from the bitstream using a second decoding network; wherein the processing resources required by the first decoding network are higher than those required by the second decoding network. In one possible implementation, the first decoding network and the second decoding network are completely different decoding networks; or...The first decoding network and the second decoding network share some sub-networks; or, the second decoding network is a sub-network of the first decoding network. In one possible implementation, the first decoding network includes an entropy decoding network, a probability estimation network, an autoregressive network, and a first image restoration network; decoding the image to be processed from the bitstream using the first decoding network includes: decoding the edge information of a three-dimensional feature map of the image to be processed from the bitstream using the entropy decoding network; the three-dimensional feature map includes multiple feature elements; estimating the first probability distribution mean of the feature elements to be decoded based on the edge information using the probability estimation network; determining the second probability distribution mean of the feature elements to be decoded based on the first probability distribution mean and the already decoded feature elements using the autoregressive network; decoding the difference information of the feature elements to be decoded from the bitstream using the entropy decoding network based on the second probability distribution mean, and obtaining the feature elements to be decoded based on the residual information and the second probability distribution mean; and restoring the image to be processed using the first image restoration network based on the decoded three-dimensional feature map. In one possible implementation, the second decoding network includes the entropy decoding network, the probability estimation network, and the second image restoration network; decoding the image to be processed from the bitstream using the second decoding network includes: decoding the edge information of the three-dimensional feature map of the image to be processed from the bitstream using the entropy decoding network; the three-dimensional feature map includes multiple feature elements; estimating the mean of a first probability distribution of the feature elements to be decoded based on the edge information using the probability estimation network; decoding the residual information of the feature elements to be decoded from the bitstream based on the mean of the first probability distribution using the entropy decoding network, and obtaining the feature elements to be decoded based on the residual information and the mean of the first probability distribution; and restoring the image to be processed based on the decoded three-dimensional feature map using the second image restoration network. In one possible implementation, the second image restoration network is a sub-network of the first image restoration network, or the image restoration network shares a common sub-network with the first image restoration network, or the second image restoration network and the first image restoration network are two different networks. In a third aspect, embodiments of this application provide an image encoding apparatus, including a memory and a video encoder, wherein: the memory is used to store video data, the video data including an image to be processed; the video encoder is used to encode identification information indicating the decoding network used into a bitstream; wherein the identification information is a first value, indicating that the decoding network used to decode the image to be processed from the bitstream is a first decoding network; or...The identification information is a second value, indicating that the decoding network used to decode the image to be processed from the bitstream is a second decoding network; the processing resources required by the first decoding network are higher than those required by the second decoding network. Fourthly, embodiments of this application provide an image decoding apparatus, including a memory and a video decoder, wherein the memory is used to store video data in the form of a bitstream, the video data including an image to be processed; the video decoder is used to decode identification information from the bitstream indicating the decoding network used; when the identification information is a first value, the image to be processed is decoded from the bitstream using a first decoding network; or, when the identification information is a second value, the image to be processed is decoded from the bitstream using a second decoding network; the processing resources required by the first decoding network are higher than those required by the second decoding network. Fifthly, embodiments of this application provide a video decoding device, including: a non-volatile memory and a processor coupled to each other, the processor calling program code stored in the memory to execute the method described in any implementation of the second aspect. Sixthly, embodiments of this application provide a video encoding device, including: a non-volatile memory and a processor coupled to each other, wherein the processor invokes program code stored in the memory to execute the method described in any implementation of the first aspect or the seventeenth aspect. Seventhly, embodiments of this application provide a computer-readable storage medium storing program code that, when the computer program is run on a computer, causes the computer to execute the method described in any implementation of the second aspect. Eighthly, embodiments of this application provide a computer-readable storage medium storing program code that, when the computer program is run on a computer, causes the computer to execute the method described in any implementation of the first aspect or the seventeenth aspect. Ninthly, embodiments of this application provide a computer-readable storage medium storing a video stream decoded by the method of any implementation of the second aspect executed by one or more processors. Tenthly, embodiments of this application provide a computer-readable storage medium storing a video stream encoded by the method of any implementation of the first aspect or the seventeenth aspect executed by one or more processors. Eleventhly, embodiments of this application provide a computer-readable storage medium storing a bitstream, the bitstream including identification information; wherein the identification information is a first value indicating that the decoding network used to decode the image to be processed from the bitstream is a first decoding network; or, the identification information is a second value indicating that the decoding network used to decode the image to be processed from the bitstream is a second decoding network; the firstThe processing resources required by the first decoding network are higher than those required by the second decoding network. In a twelfth aspect, embodiments of this application provide an encoded bitstream, the encoded bitstream including a plurality of syntax elements, the plurality of syntax elements including identification information for indicating the decoding network used to decode an image to be processed from the bitstream. In a thirteenth aspect, embodiments of this application provide a video encoder for encoding an image to be processed. Exemplarily, the video encoder can implement the method described in the first aspect or the seventeenth aspect. In a fourteenth aspect, embodiments of this application provide a video decoder for decoding an image to be processed from a bitstream. Exemplarily, the video encoder can implement the method described in the second aspect. In a fifteenth aspect, embodiments of this application provide an encoding network, comprising: a first feature extraction network, a second feature extraction network, a quantization network, an autoregressive network, an edge information extraction network, and a probability estimation network; when the identifier information used to indicate the encoding network being used is a first value, a three-dimensional feature map of the image to be processed is extracted through the first feature extraction network; when the identifier information is a second value, a three-dimensional feature map of the image to be processed is extracted through the first feature extraction network; edge information of the image to be processed is extracted from the edges of the three-dimensional feature map through the edge information extraction network; and a first probability distribution mean of the feature elements to be encoded is estimated through the probability estimation network based on the edge information. When the identifier information used to indicate the encoding network is a first value, the encoded feature elements and the mean of the first probability distribution are input into the autoregressive network to obtain the mean of the second probability distribution of the feature elements to be encoded; based on the feature elements to be encoded and the mean of the second probability distribution of the feature elements to be encoded, the residual information of the feature elements to be encoded is obtained; when the identifier information used to indicate the encoding network is a second value, based on the feature elements to be encoded and the mean of the first probability distribution, the residual information of the feature elements to be encoded is obtained. In one possible implementation, the second feature extraction network is a sub-network of the first feature extraction network, or the second feature extraction network and the first feature extraction network are two completely different sub-networks. In a sixteenth aspect, embodiments of this application provide a decoding network, including: a width decoding network, a probability estimation network, an autoregressive network, a first image restoration network, and a second image restoration network; the entropy decoding network decodes the edge information and identifier information of the three-dimensional feature map of the image to be processed from the bitstream; the three-dimensional feature map includes multiple feature elements; the probability estimation network estimates the mean of the first probability distribution of the feature elements to be decoded based on the edge information;When the identification information is a first value, the autoregressive network determines the second probability distribution mean of the feature element to be decoded based on the first probability distribution mean and the already decoded feature elements. The entropy decoding network decodes the residual information of the feature element to be decoded from the bitstream based on the second probability distribution mean, and obtains the feature element to be decoded based on the residual information and the second probability distribution mean. The first image restoration network restores the image to be processed based on the decoded 3D feature map. Alternatively, when the identification information is a second value, the autoregressive network decodes the residual information of the feature element to be decoded from the bitstream based on the first probability distribution mean, and obtains the feature element to be decoded based on the residual information and the first probability distribution mean. The second image restoration network restores the image to be processed based on the decoded 3D feature map. In one possible implementation, the second image restoration network is a sub-network of the first image restoration network, or the second image restoration network shares a common sub-network with the first image restoration network, or the second image restoration network and the first image restoration network are two different networks. In a seventeenth aspect, embodiments of this application provide an image encoding method, comprising: acquiring the identification information; when the identification information is a first value, encoding residual information obtained by encoding the image to be processed based on (or using) a first encoding network into the bitstream; or, when the identification information is a second value, encoding the bitstream with residual information obtained by encoding the image to be processed based on (or using) a second encoding network; wherein the processing resources required by the first encoding network are higher than those required by the second encoding network. In one possible implementation, the method further comprises: encoding the identification information into the bitstream. In one possible implementation, the identification information is further used to indicate the decoding network used to decode the image to be processed from the bitstream; wherein, the identification information is a first value, used to indicate that the decoding network used to decode the image to be processed from the bitstream is a first decoding network; or, the identification information is a second value, used to indicate that the decoding network used to decode the image to be processed from the bitstream is a second decoding network; wherein the processing resources required by the first decoding network are higher than those required by the second decoding network. The identification information may also be referred to as profile ID. In one possible implementation, the first decoding network and the second decoding network are completely different decoding networks; or, the first decoding network and the second decoding network share some sub-networks; or, the second decoding network is a sub-network of the first decoding network. In one possible implementation, the first encoding network and the second encoding network are two different encoding networks; or...The first encoding network and the second encoding network share some sub-networks; or, the second encoding network is a sub-network of the first encoding network. In one possible implementation, the first encoding network includes a first feature extraction network, an autoregressive network, a side information extraction network, and a probability estimation network. The residual information obtained by encoding the image to be processed using the first encoding network includes: extracting a three-dimensional feature map of the image to be processed through the first feature extraction network, the three-dimensional feature map including multiple feature elements; extracting side information of the feature elements to be encoded from the three-dimensional feature map through the side information extraction network; estimating the first probability distribution mean of the feature elements to be encoded based on the side information through the probability estimation network; inputting the encoded feature elements and the first probability distribution mean into the autoregressive network to obtain a second probability distribution mean of the feature elements to be encoded; and obtaining the residual information of the feature elements to be encoded based on the feature elements to be encoded and the second probability distribution mean of the feature elements to be encoded. In one possible implementation, the first encoding network includes a second feature extraction network, a side information extraction network, and a probability estimation network. The residual information obtained by encoding the image to be processed using the first encoding network includes: extracting a three-dimensional feature map of the image to be processed through the second feature extraction network, the three-dimensional feature map including multiple feature elements; extracting side information of the feature elements to be encoded from the three-dimensional feature map through the side information extraction network; estimating the mean probability distribution of the feature elements to be encoded based on the side information through the probability estimation network; and obtaining the residual information of the feature elements to be encoded based on the feature elements to be encoded and the mean probability distribution. In one possible implementation, the second feature extraction network is a sub-network of the first feature extraction network, or the second feature extraction network and the first feature extraction network are two completely different sub-networks. In one possible implementation, the method further includes: encoding the side information into the bitstream. Based on the implementations provided in the above aspects, this application can be further combined to provide more implementations. Figure 1 is an exemplary block diagram of a decoding system provided in an embodiment of this application; Figure 2 is a schematic diagram of a convolutional neural network structure provided in an embodiment of this application; Figure 3 is a schematic diagram of a deep learning-based video encoding and decoding network provided in an embodiment of this application; Figure 4 is a schematic diagram of an end-to-end deep learning-based video encoding and decoding network structure provided in an embodiment of this application.Figure 5 is a flowchart illustrating an encoding / decoding method according to an embodiment of this application; Figure 6A is a network structure diagram of a first encoding network according to an embodiment of this application; Figure 6B is a network structure diagram of a second encoding network according to an embodiment of this application; Figure 7A is a flowchart illustrating an encoding process according to an embodiment of this application; Figure 7B is a flowchart illustrating another encoding process according to an embodiment of this application; Figure 8 is a structural diagram illustrating a decoding network according to an embodiment of this application; Figure 9A is a possible decoding process using a first decoding network according to an embodiment of this application; Figure 9B is a possible decoding process using a second decoding network according to an embodiment of this application; Figure 10A is a flowchart illustrating the execution process of an encoding network according to an embodiment of this application; Figure 10B is a flowchart illustrating the execution process of a decoding network according to an embodiment of this application; Figure 11 is a structural diagram illustrating an encoding network according to an embodiment of this application; Figure 12 is a network structure diagram illustrating ResAU 3N3 no tanh according to an embodiment of this application; Figure 13 is a structural diagram illustrating RNAB according to an embodiment of this application; Figure 14 is a structural diagram illustrating a residual module layer according to an embodiment of this application; Figure 15 is a network structure diagram illustrating a super-decoding network according to an embodiment of this application. Figure 16 is a schematic diagram of the network structure of a large-scale decoding network provided in an embodiment of this application; Figure 17 is a schematic diagram of the execution flow of a decoding network provided in an embodiment of this application; Figure 18 is a schematic diagram of the network structure of a lightweight residual module (LightResBlock) provided in an embodiment of this application; Figure 19 is a schematic diagram of the decoding network structure provided in Example 2 of an embodiment of this application; Figure 20A is a schematic diagram of the encoding network execution flow provided in Example 3 of an embodiment of this application; Figure 20B is a schematic diagram of the encoding and decoding network structure provided in Example 3 of an embodiment of this application; Figure 21 is a schematic diagram of the encoding network structure provided in Example 3 of an embodiment of this application; Figure 22 is a schematic diagram of the decoding network structure provided in Example 3 of an embodiment of this application. Detailed Description The embodiments of this application will now be described with reference to the accompanying drawings. In the following description, reference is made to the accompanying drawings, which form part of this disclosure and illustrate specific aspects of the embodiments of this application or which may be used to illustrate specific aspects of the embodiments of this application. It should be understood that embodiments of this application may be used in other aspects and may include structural or logical variations not depicted in the drawings. Therefore, the following detailed description should not be construed in a limiting sense, and the scope of this application is defined by the appended claims. For example, it should be understood that the disclosure in conjunction with the described methods can be equally applied to corresponding devices or systems used to perform the methods, and vice versa. For example, if one or more specific...The method steps and corresponding devices may include one or more units, such as functional units, to perform the described one or more method steps (e.g., one unit performs one or more steps, or multiple units, each performing one or more of the multiple steps), even if such one or more units are not explicitly described or illustrated in the accompanying drawings. On the other hand, for example, if a specific device is described based on one or more units, such as functional units, the corresponding method may include a step to perform the functionality of one or more units (e.g., one step performs the functionality of one or more units, or multiple steps, each performing the functionality of one or more of the multiple units), even if such one or more steps are not explicitly described or illustrated in the accompanying drawings. Furthermore, it should be understood that, unless otherwise explicitly stated, the features of the various exemplary embodiments and / or aspects described herein can be combined with each other. The technical solutions involved in the embodiments of this application may be applied not only to existing video coding standards (such as H.264, HEVC, etc.) but also to future video coding standards (such as the h.266 standard). The terminology used in the implementation section of this application is only used to explain the specific embodiments of this application and is not intended to limit this application. A brief introduction to some concepts that may be involved in the embodiments of this application is given below. The image decoding and encoding methods provided in this application can be applied to the fields of video encoding and image encoding. Specifically, these methods can be applied in scenarios such as photo album management, human-computer interaction, video compression or transmission, and image compression or transmission. Taking the application of the encoding and decoding methods in an end-to-end video image encoding and decoding system as an example, this system includes two parts: image encoding and image decoding. Image encoding is determined at the source end and typically involves processing (e.g., compressing) the original video image to reduce the amount of data required to represent the video image (thus enabling more efficient storage and / or transmission). Image decoding is determined at the destination end and typically involves performing inverse processing relative to the encoder to reconstruct the image. Current neural network-based image encoding and decoding schemes typically have fixed network structures, such as the encoding and decoding model in JPEG AI VM1.0. If this network structure is adapted to the capabilities of low-computing-power devices, the compression efficiency of the encoding scheme will decrease to some extent; if the network structure is adapted to the computing power of high-computing-power devices, then this network cannot run on low-computing-power devices. In an end-to-end video image encoding and decoding system, the encoding and decoding method provided in this application can determine the encoding and decoding network used based on the quality information. Quality information can also be called identification information or network identification, or other names may be used. This application embodiment does not specifically limit this; quality information is used to indicate the decoding network used. That is, the codec can select the corresponding quality information according to the capabilities of the decoding device to select or indicate different encoding and decoding networks. This not only has the capability to adapt to low-computing-power devices, but also...The adaptation requires higher computing power on the edge. Video encoding and decoding generally refers to processing image sequences that form videos or video sequences. In the field of video encoding and decoding, the terms "picture," "frame," or "image" can be used as synonyms. Figure 1 is an exemplary block diagram of a decoding system provided by an embodiment of this application, such as a video decoding system 10 (or simply decoding system 10) that can utilize the technology of this application. The video encoder 20 (or simply encoder 20) and video decoder 30 (or simply decoder 30) in the video decoding system 10 represent devices, etc., that can be used to perform various techniques according to the various examples described in this application. As shown in Figure 1, the decoding system 10 includes a source device 12, which provides encoded image data 21, such as encoded images, to a destination device 14 for decoding the encoded image data 21. The source device 12 includes the encoder 20, and optionally may include an image source 16, a preprocessor (or preprocessing unit) 18 such as an image preprocessor, and a communication interface (or communication unit) 22. Image source 16 may include or may be any type of image capture device for capturing real-world images, and / or any type of image generation device, such as a computer graphics processor for generating computer animation images or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof) (e.g., augmented reality (AR) images). The image source may be any type of memory or storage device storing any of the aforementioned images. To distinguish the processing performed by the preprocessor (or preprocessing unit) 18, image (or image data) 17 may also be referred to as raw image (or raw image data) 17. The preprocessor 18 receives the raw image data 17 and preprocesses it to obtain a preprocessed image (or preprocessed image data) 19. For example, the preprocessing performed by the preprocessor 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise reduction. It is understood that the preprocessing unit 18 may be an optional component. The video encoder (or encoder) 20 receives the preprocessed image data 19 and provides encoded image data 21. The communication interface 22 in the source device 12 can be used to receive the encoded image data 21 and transmit it (or any other processed version) to another device, such as the destination device 14, or any other device via the communication channel 13 for storage or direct reconstruction.Source device 12 may also include a memory (not shown in FIG1) for storing at least one of the following data: raw image data 17, preprocessed image (or preprocessed image data) 19, and encoded image data 21. Destination device 14 includes a decoder 30 and, optionally, may include a communication interface (or communication unit) 28, a post-processor (or post-processing unit) 32, and a display device 34. Communication interface 28 in destination device 14 is used to receive encoded image data 21 (or other any processed version) directly from source device 12 or from any other source device such as a storage device (e.g., the storage device is an encoded image data storage device) and to provide encoded image data 21 to decoder 30. Communication interfaces 22 and 28 can be used to send or receive encoded image data (or encoded data) 21 via a direct communication link between source device 12 and destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired network, wireless network, or any combination thereof, any type of private network and public network, or any combination thereof. For example, communication interface 22 can be used to encapsulate the encoded image data 21 into a suitable format such as a message, and / or process the encoded image data using any type of transmission encoding or processing, so as to transmit it on a communication link or communication network. Communication interface 28 corresponds to communication interface 22, and for example, can be used to receive transmitted data and process the transmitted data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain the encoded image data 21. Both communication interface 22 and communication interface 28 can be configured as unidirectional or bidirectional communication interfaces as indicated by the arrow pointing from source device 12 to destination device 14 in FIG. 1, and can be used to send and receive messages, etc., to establish connections, acknowledge and exchange any other information related to communication links and / or data transmissions such as encoded image data transmissions, etc. Video decoder (or decoder) 30 is used to receive the encoded image data 21 and provide decoded image data (or decoded image data) 31. Post-processor 32 is used to post-process the decoded image data 31 (also called reconstructed image data) to obtain post-processed image data 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, or resampling, or any other processing to generate decoded image data 31 for display by the display device 34, etc. The display device 34 receives the post-processed image data 33 to display the image to a user or viewer, etc. The display device 34 may be or include any type of display for representing the reconstructed image, such as an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD).The target device 14 may also include a memory (not shown in FIG1) for storing at least one of the following data: encoded image data 21, decoded image data 31, and post-processed image data 33. The decoding system 10 further includes a training engine 25 for training the encoder 20 to process the input image or image region or image block to obtain a feature map of the input image or image region or image block, and to obtain an estimated probability distribution of the feature map, and to encode the feature map according to the estimated probability distribution. The training engine 25 is also used to train the decoder 30 to obtain an estimated probability distribution of the bitstream, to decode the bitstream according to the estimated probability distribution to obtain a feature map, and to reconstruct the feature map to obtain a reconstructed image. Although Figure 1 shows source device 12 and destination device 14 as independent devices, device embodiments may also include source device 12 and destination device 14 simultaneously, or include the functions of both source device 12 and destination device 14 simultaneously, i.e., simultaneously including source device 12 or corresponding functions and destination device 14 or corresponding functions. In these embodiments, source device 12 or corresponding functions and destination device 14 or corresponding functions may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof. 7 WO 2024 / 222681 PCT / CN2024 / 089342 5 10 15 20 25 30 35 40 45 As described, the presence and (precise) division of different units or functions in source device 12 and / or destination device 14 shown in Figure 1 may vary depending on the actual device and application, which is obvious to those skilled in the art. In recent years, the application of deep learning to the field of video encoding and decoding has gradually become a trend. Deep learning refers to learning at multiple levels at different abstraction levels through machine learning algorithms. Video codecs based on deep learning can also be called AI video codecs or video codecs based on neural networks. Since the embodiments of this application involve the application of neural networks, for ease of understanding, some nouns or terms used in the embodiments of this application will be explained below, and these nouns or terms are also part of the content of the invention. (1) Artificial neural network (ANN): An artificial neural network, also known as a neural network (NN), is a dynamic system artificially established with a directed graph as its topology. It learns through...Artificial neural networks (CNNs) are information processing systems that mimic the structure and function of the human brain by processing continuous or discontinuous inputs as state responses. After decades of development, CNNs have been widely used in many fields such as pattern recognition, automatic control, signal processing, decision support, artificial intelligence, and scientific computing, and have achieved widespread success. A network usually consists of an input layer, a hidden layer, and an output layer. The neural networks involved in this application can include various types, such as deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), residual networks, neural networks using the transformer model, or other neural networks. Some neural networks are introduced below as examples. (2) Convolutional Neural Networks: Convolutional neural networks (CNNs) are a type of deep neural network with a convolutional structure. They are a deep learning architecture, which refers to learning at multiple levels at different abstraction levels through machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network in which each neuron processes the input data. As shown in Figure 2, the Convolutional Neural Network (CNN) 100 may include an input layer 110, convolutional / pooling layers 120 (pooling layers are optional), and a neural network layer 130. As shown in Figure 2, the convolutional / pooling layer 120 may include layers 121-126 as in examples. In one implementation, layer 121 is a convolutional layer, layer 122 is a pooling layer, layer 123 is a convolutional layer, layer 124 is a pooling layer, layer 125 is a convolutional layer, and layer 126 is a pooling layer; in another implementation, layers 121 and 122 are convolutional layers, layer 123 is a pooling layer, layers 124 and 125 are convolutional layers, and layer 126 is a pooling layer. That is, the output of the convolutional layer can be used as the input of a subsequent pooling layer, or as the input of another convolutional layer to continue the convolution operation. Taking convolutional layer 121 as an example, convolutional layer 121 may include multiple convolution operators, also called convolution kernels. A convolution operator can essentially be a weight matrix, which is usually predefined. Taking image processing as an example, different weight matrices extract different features from the image. For instance, one weight matrix might be used to extract edge information, another to extract specific colors, and yet another to blur unwanted noise. The weight values in these matrices need to be obtained through extensive training in practical applications. The various weights formed by these trained weights are then used to define the different features.Each weight matrix can extract information from the input data, thereby helping the convolutional neural network 100 to make correct predictions. When the convolutional neural network 100 has multiple convolutional layers, the initial convolutional layers (e.g., 121) tend to extract more general features, which can also be called low-level features. As the depth of the convolutional neural network 100 increases, the features extracted by later convolutional layers (e.g., 126) become increasingly complex, such as high-level semantic features. Features with higher semantic levels are more suitable for the problem to be solved. Pooling layers: Since it is often necessary to reduce the number of training parameters, pooling layers are often introduced periodically after convolutional layers, i.e., layers 121-126 as shown in Figure 2 (120 example). It can be a convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. In image processing, the sole purpose of pooling layers is to reduce the spatial size of the image. Pooling layers can include average pooling operators and / or max pooling operators to sample the input image to obtain a smaller image size. The average pooling operator calculates the average value of pixels in an image within a specific range. The max pooling operator takes the pixel with the largest value within that range as the result of max pooling. Furthermore, just as the size of the weight matrix in a convolutional layer should be related to the image size, the operators in a pooling layer should also be related to the image size. The size of the output image after pooling can be smaller than the size of the input image of the pooling layer. Each pixel in the output image represents the average or maximum value of the corresponding sub-region of the input image of the pooling layer. After processing by the convolutional / pooling layer 120, the convolutional neural network 100 is still insufficient to output the required output information. As mentioned earlier, the convolutional / pooling layer 120 only extracts features and reduces the parameters introduced by the input image. However, to generate the final output information (the required class information or other relevant information), the convolutional neural network 100 needs to utilize the neural network layer 130 to generate one or a set of required class numbers of output. Therefore, neural network layer 130 may include multiple hidden layers (as shown in Figure 2, 13L·132 8 WO 2024 / 222681 PCT / CN2024 / 089342 5 10 15 20 25 30 35 40 45 to 13n) and an output layer 140. The parameters contained in these multiple hidden layers can be pre-trained based on relevant training data for specific task types, such as image recognition, image classification, image super-resolution reconstruction, etc. After the multiple hidden layers in neural network layer 130, that is, the last layer of the entire convolutional neural network 100, is the output layer 140. This output layer 140 has a loss function similar to classification cross-entropy, specifically used to calculate the prediction error. Once the previous layers of the entire convolutional neural network 100 are complete, the output layer 140 will be used to calculate the prediction error.After forward propagation (as shown in Figure 2, propagation from 110 to 140 is forward propagation), backward propagation (as shown in Figure 2, propagation from 140 to 110 is backward propagation) begins to update the weight values and biases of the aforementioned layers to reduce the loss of the convolutional neural network 100 and the error between the output of the convolutional neural network 100 through the output layer and the ideal result. It should be noted that the convolutional neural network 100 shown in Figure 2 is only an example of a convolutional neural network. In specific applications, convolutional neural networks can also exist in the form of other network models, for example, multiple convolutional / pooling layers in parallel, with the extracted features input to neural network layer 130 for processing. (3) During the training of a neural network, the goal is to make the network's output as close as possible to the desired predicted value. Therefore, the weight vector of each layer is updated based on the difference between the network's predicted value and the target value. (Of course, there is usually an initialization process before the first update, i.e., pre-configuring parameters for each layer in the neural network). For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict lower, and this adjustment continues until the neural network can predict the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value," which is the loss function or objective function. These are important equations used to measure the difference between the predicted value and the target value. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, so training the neural network becomes a process of minimizing this loss as much as possible. (4) Linear Operations: Linearity refers to a proportional, linear relationship between quantities. Mathematically, it can be understood as a function whose first derivative is a constant. Linear operations can include, but are not limited to, addition, empty operations, identity operations, convolution operations, layer normalization (LN) operations, and pooling operations. Linear operations can also be called linear mappings. Linear mappings need to satisfy two conditions: homogeneity and additivity. If either condition is not satisfied, it is nonlinear. Homogeneity means f(ax) = af(x); additivity means f(x+y) = f(x) + f(y). For example, f(x) = ax is linear. It should be noted that X, a, and f(x) here are not necessarily scalars; they can be vectors or matrices, forming a linear space of arbitrary dimensions. If X and f(X) are n-dimensional vectors, when a is a constant, it is equivalent to satisfying homogeneity; when a is a matrix, it is equivalent to satisfying additivity. In contrast, functions whose graphs are straight lines do not necessarily conform to linear mappings. For example, f(X) = ax + b does not satisfy either homogeneity or additivity, and therefore belongs to nonlinear mappings.In this embodiment, the composition of multiple linear operations can be called a linear operation, and each linear operation included in the linear operation can also be called a sub-linear operation. (5) Attention model. The attention model is a neural network that applies the attention mechanism. In deep learning, the attention mechanism can be broadly defined as a weight vector describing importance: through this weight vector, an element can be predicted or inferred. For example, for a pixel in an image or a word in a sentence, the correlation between the target element and other elements can be quantitatively estimated using the attention vector, and the weighted sum of the attention vectors can be used as an approximation of the target. The attention mechanism in deep learning simulates the attention mechanism of the human brain. For example, when a human views a painting, although the human eye can see the whole picture, when the human observes it carefully, the eye actually focuses on only a part of the pattern in the whole picture, and at this time the human brain mainly focuses on this small part of the pattern. That is to say, when a human carefully observes an image, the human brain's attention to the whole image is not balanced, but has a certain weight distinction, which is the core idea of the attention mechanism. In simple terms, the human visual processing system often selectively focuses on certain parts of an image while ignoring other irrelevant information, thus aiding the brain's perception. Similarly, in deep learning's attention mechanisms, in problems involving language, speech, or vision, certain parts of the input may be more relevant than others. Therefore, through the attention mechanism in the attention model, the attention model can dynamically focus only on the parts of the input that are helpful in effectively performing the task at hand. (6) Self-attention network. A self-attention network is a neural network that applies a self-attention mechanism. The self-attention mechanism is an extension of the attention mechanism. The self-attention mechanism is actually an attention mechanism that associates different positions of a single sequence to compute the representation of the same sequence. The self-attention mechanism can play a key role in machine reading, abstract summarization, or image description generation. Taking the application of self-attention networks in natural language processing as an example, self-attention networks process input data of arbitrary length and generate new feature representations of the input data, and then convert the feature representations into target words. The self-attention network layer in a self-attention network utilizes an attention mechanism to obtain the relationships between all other words, thereby generating new feature representations for each word. The advantage of self-attention networks is that the attention mechanism can directly capture the relationships between all words in a sentence without considering word position. Figure 3 is a schematic diagram of a deep learning-based video encoding / decoding network (or system) provided in an embodiment of this application. Figure 2 shows the entropy...The encoding and decoding process is illustrated using an example. This network includes a feature extraction module, a feature quantization module, an entropy coding module, an entropy decoding module, a feature dequantization module, and a feature decoding (or image reconstruction) module. At the encoding end, the original image (or the image to be compressed) is input to the feature extraction module. The feature extraction module, through stacked convolutional layers and combined with a nonlinear mapping activation function, outputs the extracted 3D feature map of the original image. The feature quantization module quantizes the floating-point feature values in the 3D feature map to obtain the quantized feature map. The quantized 3D feature map undergoes entropy coding to obtain the bitstream. At the decoding end, the entropy decoding module parses the bitstream to obtain the quantized 3D feature map. The feature dequantization module dequantizes the integer feature values in the quantized feature map to obtain the dequantized feature map. The dequantized feature map is reconstructed by the feature decoding module to obtain the reconstructed image. Entropy coding is encoding that follows the entropy principle without losing any information during the encoding process. Entropy coding is used to apply entropy coding algorithms or schemes to quantization coefficients and other syntax elements to obtain encoded data that can be output as an encoded bitstream, allowing decoders to receive and use the parameters for decoding. The encoded bitstream can be transmitted to the decoder or stored in memory for later transmission or retrieval by the decoder. The encoding algorithms or schemes include, but are not limited to: variable length coding (VLC), context adaptive VLC (CALVC), arithmetic coding schemes, binarization algorithms, context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other encoding methods or techniques. The network may also exclude feature quantization and dequantization modules. In this case, the network can directly perform a series of processes on the feature map, which is a floating-point number. Alternatively, the network can perform integer processing so that the feature values in the feature map output by the feature extraction module are all integers. After the image to be processed (or the image to be compressed) passes through the feature extraction module and the feature quantization module, a three-dimensional feature quantization map is obtained. When the encoding module processes each feature value in the three-dimensional feature quantization map, it can use the processed feature values in the neighborhood as context to estimate the probability distribution of the feature value, obtain the probability distribution of the feature value, and perform subsequent encoding based on the probability distribution to obtain the encoded bitstream.Figure 4 is a schematic diagram of an end-to-end video encoding and decoding network structure based on deep learning provided in an embodiment of this application. Figure 4 uses entropy encoding and decoding as an example for illustration. The neural network includes a feature extraction module, a quantization module, a side information extraction module, a fine encoding module, a fine decoding module, a probability estimation module, and a reconstruction module. Entropy encoding can be an autoencoder (AE), and entropy decoding can be an autodecoder (AD). At the encoding end, the original image X is input to the feature extraction module, which outputs a feature map y of the original image. On one hand, the feature map y is input to the quantization module, which outputs a quantized feature map, and the quantized feature map is input to the fine encoding module. On the other hand, the feature map y is input to the side information extraction module, which outputs side information Z. The side information Z is input to the quantization module, which outputs quantized side information. The quantized side information is processed by the fine encoding module to obtain the bitstream of the side information, and then processed by the fine decoding module to obtain the decoded side information. The decoded side information is input to the probability estimation module, which outputs the probability distribution of each feature element in the quantized feature map, and the probability distribution of each feature element is input to the entropy encoding module. The entropy encoding module performs entropy encoding on each input feature element based on the probability distribution of each feature element to obtain a super-prior bitstream. Here, side information z is a type of feature information, represented as a three-dimensional feature map, containing fewer feature elements than feature map y. At the decoding end, the entropy decoding module parses the bitstream of side information to obtain side information, inputs the side information into the probability estimation module, and the probability estimation module outputs the probability distribution of each feature element [x][y][i] in the symbol to be decoded. The probability distribution of each feature element [x][y][i] is input into the entropy decoding module. The entropy decoding module performs entropy decoding on each feature element based on its probability distribution to obtain the decoded feature map. The decoded feature map is input into the reconstruction module, and the reconstruction module outputs the reconstructed image. Furthermore, some variational autoencoders (VAEs) utilize previously encoded or decoded feature elements around the current feature element in their probability estimation modules to more accurately estimate the probability distribution of the current feature element. It should be noted that the network structures shown in Figures 3 and 4 are merely illustrative examples, and the embodiments of this application do not limit the modules included in the network or their structures. 10 WO 2024 / 222681 PCT / CN2024 / 089342 5 10 15 20 25 30 35 40 45 In some possible scenarios, in order to further improve the accuracy of the mean, an autoregressive module can be added. The autoregressive module can further obtain the probability distribution value for obtaining the residual based on the mean output by the probability distribution module and the quantized feature map.The encoding and decoding method provided in the embodiments of this application will be described in detail below. Referring to Figure 5, it is a schematic flowchart of an encoding and decoding method provided in an embodiment of this application. This method can be executed by two electronic devices or by one electronic device. For example, when the method is executed by two electronic devices, one electronic device includes an encoder for indicating encoding operations, and the other electronic device includes a decoder for performing decoding operations. When the method is executed by one electronic device, the electronic device may include both an encoder / decoder and a decoder. The method can be executed by the electronic device by invoking a neural network model. This method is described as a series of operations; it should be understood that the method can be executed in various orders and / or occur simultaneously, and is not limited to the execution order shown in Figure 5. 501, The encoder encodes identification information indicating the decoding network used into the bitstream. Wherein, the identification information is a first value, used to indicate that the decoding network used to decode the image to be processed from the bitstream is a first decoding network; or, the identification information is a second value, used to indicate that the decoding network used to decode the image to be processed from the bitstream is a second decoding network. This can also be understood as follows: if the identifier information is a first value, the encoder uses the first encoding network corresponding to the first decoding network to perform encoding operations on the image to be processed; or, if the identifier information is a second value, the encoder uses the second encoding network corresponding to the second decoding network to perform encoding operations on the image to be processed. This identifier information can also be called profile information (Profile ID), network information, network identifier, or other names; this application embodiment does not limit this. In other words, this identifier information is used to indicate the processing that the decoder needs to support, such as the `general_profile_idc` syntax element in the H.265 standard. As an example, the first value can be 0, the second value can be 1, or the first value can be 1 and the second value can be 0. The first and second values can also be other values. The processing resources (or computing power) required by the first decoding network are higher than the processing resources (computing power) required by the second decoding network. Processing resources (computing power) may include memory resources, processor resources, and other resources. In some embodiments, it can also be understood that the decoding rate (or decompression efficiency) of the first decoding network and the second decoding network are different, for example, the decoding rate of the first decoding network is higher than that of the second decoding network; or, the image quality recovered by the first decoding network is different from that recovered by the second decoding network, for example, the image quality recovered by the first decoding network is higher than that recovered by the second decoding network. 502, the encoder sends the bitstream. 503, the decoder decodes the identification information used to indicate the decoding network employed from the received bitstream. 504, when the identification information is a first value, the first decoding network is used to decode the image to be processed from the bitstream; or, when...When the identification information is a second value, the image to be processed is decoded from the bitstream using a second decoding network. In one possible implementation, the first decoding network and the second decoding network are completely different decoding networks; or, the first decoding network and the second decoding network share some sub-networks; or, the second decoding network is a sub-network of the first decoding network. In some embodiments, the identification information may also include other values to indicate different decoding networks. It can be understood that multiple different values are used to indicate multiple different decoding networks. For example, if the identification information is a third value, it indicates that the decoding network used is a third decoding network. The decoding rate of the third decoding network is different from that of the first decoding network (or the second decoding network). In some embodiments, the decoding rate of the third decoding network is higher than that of the second decoding network, and the decoding rate of the second decoding network is higher than that of the first decoding network. In other embodiments, the decoding speed of the third decoding network is between the decoding rates of the first and second decoding networks. Here, only three decoding networks are used as an example, and the number of decoding networks is not specifically limited in this application embodiment. A higher image decoding rate can be understood as a smaller image decoding latency. For example, the image quality recovered by the third decoding network may differ from that recovered by the first (or second) decoding network. The image quality recovered by the third decoding network is higher than that recovered by the second decoding network, and the image quality recovered by the second decoding network is higher than that recovered by the first decoding network. In other embodiments, the image quality recovered by the third decoding network falls between that recovered by the first and second decoding networks. In some scenarios, where a third decoding network is also included, it differs from the first (and second) decoding networks. For instance, the third decoding network, the second decoding network, and the first decoding network are three different decoding networks; or, the third decoding network shares some sub-networks with the second (or first) decoding network; or, the third decoding network is a sub-network of the second (or first) decoding network. The first and second decoding networks sharing some sub-networks can be understood as the first decoding network reusing some sub-networks of the second decoding network. For example, the first decoding network includes networks A, B, and C, and the second decoding network includes networks D, B, and C, with both decoding networks sharing network B. Therefore, when using the first decoding network, data is input to network A, the output of network A is input to network B, and the output of network B is input to network C. If the second decoding network is used, [the process is as follows: 11 WO 2024 / 222681 PCT / CN2024 / 089342 5 10 15 20 25 30 35 40 45]This can be understood as follows: data is input into network D, the output of network D is input into network B, and the output of network B is input into network C. For example, the first decoding network is a sub-network of the second decoding network. For instance, the first decoding network includes networks A1, A2, and A3. The second decoding network includes networks A1 and A3. When using the first decoding network, data is input into network A1, the output of network A1 is input into network A2, and the output of network A2 is input into network A3. When using the second decoding network, it can be understood that when data is input into network A1, the output of network A1 is no longer input into network A2, but is skipped and input into network A3. In another possible implementation, the encoder can also use different encoding networks based on different values of the identification information during encoding. Alternatively, after the encoder encodes the image to be processed into a bitstream using encoding networks, it can encode the identification information of the decoding network corresponding to the used encoding network into the bitstream. This identification information can be understood as indicating both the decoding network and the encoding network used. When the identification information is a first value, the residual information obtained by encoding the image to be processed using the first encoding network is incorporated into the bitstream; or, when the identification information is a second value, the bitstream is encoded using the residual information obtained by encoding the image to be processed using the second encoding network; the processing resources (or computing power) required by the first encoding network are higher than those required by the second encoding network. It should be noted that the first encoding network and the first decoding network can be a pair of networks, where encoding is done using the first encoding network and then decoding using the first decoding network; the second encoding network and the second decoding network are a pair of networks, where encoding is done using the second encoding network and then decoding using the second decoding network. In some embodiments, the first encoding network and the second encoding network are two different encoding networks; or, the first encoding network and the second encoding network share some sub-networks; or, the first encoding network is a sub-network of the second encoding network. The identification information can also have other values, with different values indicating the use of different encoding networks. For example, if the identification information is a third value, the encoding network used is the third encoding network, and the decoding network used is the third decoding network. It should be noted that the third encoding network and the third decoding network are a pair of networks; encoding is done using the third encoding network, and decoding is done using the third decoding network. The encoding rate of the third encoding network differs from that of the first encoding network (or the second encoding network). In some embodiments, the encoding rate of the third encoding network is higher than that of the second encoding network, and the encoding rate of the second encoding network is higher than that of the first encoding network. In other embodiments, the encoding rate of the third encoding network is between that of the first and second encoding networks. A higher image encoding rate can be understood as a lower image encoding latency.For example, the image quality recovered by the third encoding network differs from that recovered by the first encoding network (or the second encoding network). The image quality recovered by the third encoding network is higher than that recovered by the second encoding network, and the image quality recovered by the second encoding network is higher than that recovered by the first encoding network. In other embodiments, the image quality recovered by the third encoding network is between that recovered by the first and second encoding networks. In some scenarios, when a third encoding network is also included, it differs from the first encoding network (and the second encoding network). For example, the third encoding network, the second encoding network, and the first encoding network are three different encoding networks; or, the third encoding network shares some sub-networks with the second encoding network (or the first encoding network); or, the third encoding network is a sub-network of the second encoding network (or the first encoding network). Exemplarily, the first encoding network includes a feature extraction module, a quantization module, a side information extraction module, an entropy encoding module, and a probability estimation module. The second encoding network also includes a feature extraction module, a quantization module, a side information extraction module, an entropy encoding module, and a probability estimation module. In one approach, at least one module in the first encoding network and the second encoding network uses a different network structure. For example, the probability estimation module in the first encoding network has a different network structure than the probability estimation module in the second encoding network. Another example is that the feature extraction module in the first encoding network has a different network structure than the feature extraction module in the second encoding network. For instance, the feature extraction module in the first encoding network can be referred to as the first feature extraction module, and the feature extraction module in the second encoding network can be referred to as the second feature extraction module. It should be noted that the modules belonging to a neural network can also be referred to as "networks". For example, the feature extraction module can be called a feature extraction network, and the quantization module can also be called a quantization network. The network structures of the feature extraction modules in the first encoding network and the second encoding network may differ. This could mean that the second feature extraction network is a subnetwork of the first feature extraction network, or that the second feature extraction network and the first feature extraction network are two completely different subnetworks. Alternatively, the second feature extraction network and the first feature extraction network may share one or more subnetworks. Similarly, the network structures of the probability estimation network in the first encoding network and the second encoding network may differ. This could mean that the second feature extraction network is a subnetwork of the first feature extraction network, or that the second feature extraction network and the first feature extraction network are two completely different subnetworks. Alternatively, the second feature extraction network and the first feature extraction network may share one or more subnetworks.As an example, referring to Figure 6A, the network structure of the first encoding network is as follows: the first encoding network includes a first feature extraction network 610, a quantization network 620, an autoregressive network 630, a side information extraction network 640, and a probability estimation network 650. Further, the residual information obtained by encoding the image to be processed using the first encoding network can be implemented in the following way, as shown in Figure 7A, which is a schematic diagram of a possible process for encoding residual information. 701a, Extract the three-dimensional feature map of the image to be processed through the first feature extraction network 610; 702a, Quantize the three-dimensional image features through the quantization network 620 to obtain a quantized three-dimensional feature map; 703a, Extract the edge information of the image to be processed from the edges in the three-dimensional feature map through the edge information extraction network 640; 704a, Estimate the mean of the first probability distribution of the image to be processed based on the edge information through the probability estimation network 650; 705a, Input the quantized three-dimensional feature map and the first probability distribution information into the autoregressive network 630 to obtain the mean of the second probability distribution; 706a, Obtain the residual information based on the third-dimensional feature map of the image to be processed and the mean of the second probability distribution. The first encoding network may also include an entropy encoding network 660. The entropy encoding network 660 can encode the residual information and the edge information into the bitstream. In some scenarios, this can be understood as encoding the edge information into bitstream 1, encoding the residual information into bitstream 2, and then merging bitstream 1 and bitstream 2 into one bitstream. In another scenario, edge information and residual information can be encoded into a single bitstream. As another example, referring to Figure 6B, the network structure of the second encoding network is as follows: the second encoding network includes a second feature extraction network 611, an edge information extraction network 640, and a probability estimation network 650. Further, the residual information obtained by encoding the image to be processed using the second encoding network can be implemented as follows, as shown in Figure 7B, which is a possible flowchart for encoding residual information. 701b, Extract the three-dimensional feature map of the image to be processed using the second feature extraction network 611. 702b, Extract edge information from the three-dimensional feature map using the edge information extraction network 640. 703b, Estimate the mean probability distribution of the image to be processed based on the edge information using the probability estimation network 650. 704b, Obtain the residual information based on the three-dimensional feature map of the image to be processed and the mean probability distribution. The second encoding network may also include a supplementary encoding network 660. The supplementary encoding network 660 can encode the residual information and edge information into the bitstream. In some scenarios, this can be understood as encoding the edge information into bitstream 1, encoding the residual information into bitstream 2, and then merging bitstream 1 and bitstream 2 into a single bitstream. In another scenario, the edge information and residual information can be encoded into a single bitstream.The identification information in this embodiment can be located in the stream header. In some scenarios, the identification information can also be added to the file extension of the stream file. For example, different identification information corresponds to different file extensions. For example, the header contains information such as image width and height, image format, profileID, etc. This information should be stored in a pre-defined order, and this application does not specifically limit the specific storage order. For example, the stream header can contain one or more of the following parameter information. The parameter information includes: profile ID, image height h, width W, position and size of tiles in the latent space, control flags for each tool, scaling factors of primary and secondary components, model index (model IDx) - learnable model index, and rate control parameter βv, etc. Rate control parameters include rate control parameter βy for primary components and rate control parameter βuν for secondary components, etc. For example, the parameters of the stream header can be encoded using a fixed bit length. The following explains each parameter information. W represents the width of the input image. For example, the range of W can be from 1 pixel to 8192 pixels. h represents the height of the input image. For example, h can range from 1 pixel to 8192 pixels. format represents the data format of the input image, such as YUV420, YUV444, sRGB, etc. bit_depth represents the bit depth of the input image, such as 8 and 10. 0 is a parameter used to indicate the quality level for variable rate. The primary and secondary components may have different betas, so the primary component is represented as beta luma (py), and the secondary component is represented as beta chroma (Puv). The value of βy is between 0 and 1 and can be represented as a 16-bit fixed-point number. γ represents luminance (or luma). UV represents chroma. Color transform information, by default, is encoded in YUV Bt.709 (full range).However, custom color transformation is also supported. In this case, 12 coefficients (conversion matrix and offset) can be used, encoded as fixed-point numbers with 8-bit resolution. (By default, the coded representation of the signal is YUV Bt.709 (full range), but custom color transformation is also supported. In that case, 12 coefficients may be sent (conversion matrix and offset), 13 WO 2024 / 222681 PCT / CN2024 / 089342 5 10 15 20 25 30 35 40 45 encoded as a fixed-point number with 8-bit resolution) « Tile information, representing the tile size and overlap of the luminance decoder, and the tile size and overlap of the chrominance decoder. Inter-Channel Correlation Information (ICCI) filter: Tile size and overlap. Generally, the tile shape is square. However, because the right or bottom edge of the image may be smaller, the shape of the tiles on the right or edge does not have to be square. (decoder tile size and overlap for Luma, decoder tile size and overlap for Chroma. ICCI tile size and overlap. Tiles have square shape except at the right or bottom picture boundary, where they may be smaller and non-rectangular). The SkipMode_enable_flag indicates whether SkipMode is used for image encoding. The RVS_enable_flag indicates whether Residual and Variance Scale (RVS) is used for each image encoding. The LSBS_enable_flag...The enable flag indicates whether latent scale before synthesis (LSBS) is used for encoding each image. The ICIC enable flag (ICIC_enable_flag) indicates whether inter-channel correlation information (ICCI) is used for encoding each image. numThreads is a 16-bit unsigned integer used to indicate the number of samples processed in parallel. It should be noted that in some scenarios, different values of the identifier information (grade information) may only indicate the use of different decoding networks; in other scenarios, different encoding networks are used only based on different values of the identifier information; and in still other scenarios, different values of the identifier information indicate the use of different encoding and decoding networks. In one possible example, as shown in Figure 8, the first decoding network (or the second decoding network) may include a dragonfly decoding network 810, a probability estimation network 820, and an image restoration network. The image restoration network may also be called a reconstruction network, or other names may be used; this embodiment does not specifically limit this. The decoding network at the decoding end decodes the edge information and residual information of the image to be processed from the bitstream. At least one of the first and second decoding networks has a different network structure; for example, the image restoration network structure may differ, or the probability estimation network structure may differ. For instance, the image restoration network in the first and second decoding networks may have different network structures. For ease of distinction, the image restoration network in the first decoding network is referred to as the first image restoration network 831, and the image restoration network in the second decoding network is referred to as the second image restoration network 832. The first image restoration network 831 and the second image restoration network 832 are different; for example, the second image restoration network 832 may be a sub-network of the first image restoration network 831, or it may share a common sub-network with the first image restoration network 831, or the second image restoration network 832 and the first image restoration network 831 may be two different networks. The first decoding network also includes an autoregressive network 840. The decoding process is explained using the above examples of the structures of the first and second decoding networks. See Figure 9A, which is a schematic diagram of a possible decoding process using the first decoding network. 9 01 a, The edge information of the three-dimensional feature map of the image to be processed is decoded from the bitstream by the docking decoding network 810; the three-dimensional feature map includes multiple feature elements.902a, The mean of the first probability distribution of the feature elements to be decoded is estimated based on the edge information by the probability estimation network 820. 903a, The mean of the second probability distribution of the feature elements to be decoded is determined by the autoregressive network 840 based on the mean of the first probability distribution and the already decoded feature elements. 904a, The residual information of the feature elements to be decoded is decoded from the bitstream by the IDM network 810 based on the mean of the second probability distribution, and the feature elements to be decoded are obtained based on the residual information and the mean of the second probability distribution. 905a, The image to be processed is recovered by the first image recovery network 831 based on the decoded three-dimensional feature map. See Figure 9B, which is a schematic diagram of a possible decoding process using the second decoding network. 901b, The edge information of the three-dimensional feature map of the image to be processed is decoded from the bitstream by the entropy decoding network 810; the three-dimensional feature map includes multiple feature elements. 902b, The mean of the first probability distribution of the feature elements to be decoded is estimated based on the edge information by the probability estimation network 820. 903b, The residual information of the feature element to be decoded is decoded from the bitstream by the decoding network 810 according to the mean of the first probability distribution, and the feature element to be decoded is obtained based on the residual information and the mean of the first probability distribution. 904b, The image to be processed is recovered by the second image recovery network 832 according to the three-dimensional feature map obtained by decoding. In some possible implementations, the probability estimation network in the encoding network (including the first encoding network and the second encoding network) and the probability estimation network used in the decoding network can be the same. The following describes the scheme of the embodiment of this application with specific examples. The following example uses the end-to-end image encoding and decoding process as an example for illustration. Example 1: Referring to Figures 10A and 10B, it is a schematic diagram of the encoding and decoding network execution process provided in the embodiment of this application. The encoding and decoding network used is dynamically adjusted according to the Profile ID. Combining the network structures of Figures 6A and 6B, the encoding network, specifically the first feature extraction network (module) 610, includes encoding network submodules 1-3. Encoding network submodules 1-3 extract features from the image to be processed, gradually transforming the image from the pixel domain to the feature domain, making it easier to compress. The second feature extraction module in the second encoding network includes encoding network submodules 1 and 3. Correspondingly, the decoding network combines the network structure shown in Figure 8. Decoding network submodules 1-2-13 or 1-2-14 gradually reconstruct the three-dimensional feature map into an image. The difference between decoding network submodules 3 and 4 lies in their structure. For example, one possible difference is that decoding network submodule 3 is a lightweight module adapted to low-computing-power devices on the edge, offering faster operation but less image recovery compared to decoding network submodule 4.Unlike decoding network submodule 3, which has slightly lower image quality, decoding network submodule 4 is designed for high-performance computing devices and, while slower, offers higher image quality. Figure 10B illustrates this, using the first image restoration network of the first decoding network (including decoding network submodules 1-2 and 4) as an example, and the second image restoration network of the second decoding network (including decoding network submodules 1-3) as an example. See Figure 10A for a schematic diagram of the encoding process provided in Example 1. Specifically, the encoding process is as follows: The first step is to calculate the output image features y using the feature extraction module. During the calculation, certain network submodules are selected for execution or skipping based on the Profile ID. When the Profile ID is 0, encoding network submodule 2 is executed, i.e., the first encoding network is used for encoding. When the Profile ID is 1, encoding network submodule 2 is skipped, i.e., the second encoding network is used for encoding. In some scenarios, encoding network submodule 2 can be skipped when the Profile ID is 1, and executed when the Profile ID is 0. For example, image features y can also be called feature map y, or three-dimensional feature map y. After the feature extraction module extracts features from the image to be encoded to obtain image features y, the image features y can be quantized. This can be understood as processing the floating-point feature values (e.g., rounding) to obtain integer feature values, thus obtaining the quantized feature map y. The second step inputs the image features y obtained in the first step into the edge information extraction network (module) to extract edge information z. z is then quantized to obtain z2, which is compressed into bitstream 1. It should be noted that the edge information extraction module is not mandatory. In some possible scenarios, the original image, after feature extraction, is directly quantized and compressed (or encoded) into a bitstream. In some embodiments, the quantized feature map z can also be input into the edge information extraction network to output the quantized edge information z2. The edge information extraction module can be implemented using a neural network; specific neural network structures will be illustrated later and will not be elaborated here. Edge information z2 can be understood as feature map z2 obtained by further feature extraction from the quantized feature map z2. The number of feature elements in z2 is less than the number of feature elements in feature map z2. In some scenarios, the encoding network (first encoding network, second encoding network) may also include a quantization network to perform quantization operations on the image feature y. In other scenarios, the side information extraction network may also have the function of performing quantization operations, thus performing quantization operations on the image feature y by the side information extraction network. The third step is to obtain the probability distribution of the image feature y from the side information. When the Profile ID is 1, the side information is input into the probability estimation network. The probability estimation network (also called the probability estimation module) includes feature probability distribution modules A and B to predict the mean of the image feature y.The feature probability distribution modules A and B can also be called feature map probability distribution estimation modules A and B, or other names, which are not limited in this embodiment. When Profile ID is 0, the edge information is input into feature probability distribution modules A and B. Feature probability distribution module B outputs the variance information of image feature y, and the output of feature probability distribution module A and the quantized image feature y are sent to the autoregressive module to generate the mean of image feature y. In some scenarios, feature probability distribution modules A and B can also be merged into one module, that is, the function is performed by one module. For example, the probability estimation network can use a Gaussian single model (GSM), an asymmetric Gaussian model, a Gaussian mixture model (GMM), or a Laplace distribution model, etc. The probability estimation network can also use deep learning networks, such as recurrent neural networks (RNN) and convolutional neural networks (CNN), etc., which are not limited here. The fourth step involves combining the probability distribution information (mean and variance) of the image feature y obtained in the third step to calculate the residual information r = y - mean relative to the mean of the image feature y. The quantized residual information f is then encoded to obtain compressed bitstream 2. The residual information r can also be called the residual feature map r, and the quantized residual information f can also be called the quantized residual feature map f, or simply the quantized residual feature map f. The fifth step involves merging bitstreams 1 and 2 into a single bitstream and writing the Profile ID into the bitstream, for example, by writing it into the bitstream header. It should be noted that the encoding steps in steps two, four, and five at the encoding end can be merged. In step two, instead of encoding the edge information 2 and writing it into the bitstream, in step four, after obtaining the quantized residual information f, the quantized residual information f and edge information 2 are encoded together (e.g., entropy encoding) and written into the bitstream. See Figure 10B, a schematic diagram of the decoding process provided in Example 1. Specifically, the decoding process is as follows: First, the bitstream is parsed through the entropy decoding network (module) to obtain the Profile ID information. For example, the Profile ID is obtained from the bitstream header. The Profile ID is profile information in the bitstream, used to indicate the processing that the decoder needs to support, such as in the H.265 standard.The `general_profile_idc` syntax element. The Profile ID can be an integer (though not necessarily an integer; this application does not impose a specific limitation). Profile information indicates the processing the decoder needs to support; it can also be understood as indicating the different networks the decoder needs to use. The second step involves decoding the bitstream (e.g., bitstream 1) using an entropy decoding network to obtain side information. For example, quantized side information 2 can be obtained from bitstream 1 using an asymmetric numeral system (ANS) / arithmetic decoding. The third step involves obtaining the probability distribution of the feature map y[x][y][i] from the side information 2 using a probability estimation network (module). When profile ID=1, the probability estimation module (or probability estimation network) for side information 2 performs probability estimation on each feature element y[x][y][i] in the feature map to be decoded, obtaining the probability distribution of the feature element y[x][y][i]. Assume that the feature elements y[x][y][i] satisfy a Gaussian distribution with mean u[x][y][i] and variance o[x][y][i], where the mean mx][y][i] can be used as the predicted value of the feature element x][y][i]. When profile ID=0, the edge information 2 is input into the probability estimation module (or probability estimation network) to perform probability estimation on each feature element x][y][i] in the feature map to be decoded, obtaining the probability distribution of the feature element kx]. Then, based on the autoregressive network, using the information of the already decoded feature elements and the mean output of the probability estimation network, the predicted value of the current feature element to be decoded is obtained. Here, the parameters X, y, i in the feature element x][y][i] are all positive integers, and the coordinates (x, y, i) represent the position of the current feature element to be decoded. Specifically, the coordinates (X, y, i) represent the position of the current feature element to be decoded relative to the feature element at the top left vertex in the current 3D feature map. This step can be specifically implemented by the probability estimation module. The probability estimation method used at the decoding end can be the same as that at the encoding end; that is, the structure of the probability estimation module at the decoding end can be the same as that at the encoding end, which will not be elaborated here. Bitstream 2 can be understood as a bitstream converted from multiple matrices y. Decoding is the process of recovering multiple matrices y from this bitstream. Recovery of y is done element-wise in an orderly manner. For example, for a 10x10 matrix, the elements are recovered one by one from left to right and top to bottom. When the element in the 7th row and 8th column is recovered, the elements preceding it (i.e., all elements with x-coordinate < 7 and y-coordinate < 8) can be called the context of this element, which can be understood as the decoded context information.The fourth step involves using the Gaussian distribution mean *n* and variance *0* of each feature element in the quantized feature map obtained in the third step to continue parsing the bitstream through the decoding network, obtaining the quantized residual feature map *i*. Then, based on *f* and *μ*, the quantized feature map *y* = *i* + *μ*0 is obtained. As an example, a possible implementation of parsing feature map 2 from the bitstream is as follows: Based on the probability distribution (e.g., a Gaussian distribution with mean 0 and variance 6), the probability P(k) of the feature element *nx][y][i* taking the value *k* is obtained. Based on P(k), the feature element is obtained from the bitstream using ANS decoding / arithmic decoding. Here, *k* can be any integer, such as 0, 1, 2, 3, etc. The fifth step involves reconstructing the image from the quantized *T* through the image restoration network. During the execution of the image restoration network (or reconstruction network), after executing decoding network submodule 1 and decoding network submodule 2, the image is selected from decoding network submodules 3 and 4 based on the value of the Profile ID. If Profile ID is 1, then decoding network submodule 3 is executed, i.e., the second decoding network is used; if Profile ID is 0, then decoding network submodule 4 is executed, i.e., the first decoding network is used. The structure of each subnetwork of the above encoding network (including the first and second encoding networks) is illustrated below with specific examples. See Figure 11, which is a schematic diagram of the encoding network execution flow. It should be noted that Figure 11 is only an example and does not constitute a limitation on the specific structure of the encoding network. Referring to Figure 11, the encoding network submodule 1 includes multiple layers, namely padding layer 1_1, convolutional layer 1_11, residual activation function layer 1_21, padding layer 1_2, convolutional layer 1_12, residual activation function layer 1_22, and padding layer 1_30. In this embodiment, a convolution with a size of KxK, an output channel number of M, and a stride of N can be represented as ConvMxKxKSNo. In Figure 11, taking the convolutional layer 1_11 and convolutional layer 1_12 as an example using Convl2 28x3x3 S2. For example, the padding layer can use zeros padding (default 0 padding), reflect padding, replicated padding, or circular padding. For example, Padding1, Padding1_2, and Padding1_3 all use the Replicate property.Padding is used to fill the length and width of the input tensor to even numbers using replicated padding (filling with the nearest element). For example, if the length and width of the input tensor are 5 and 6 respectively, the padding layer will pad one element in the length direction to make the length even number 6, while the width of the input tensor is 6, which is also even, so no padding is done in the width direction. Residual activation function layers 1_21 and 1_22 are mainly used as activation functions and can also provide attention mechanisms. As an example, residual activation function layers 1_21 and 1_22 can adopt a ResAU3x3 no tanh network. The ResAU 3x3 no tanh network can adopt the structure shown in Figure 12. Residual activation function layers 1_21 (and 1_22) include: activation function layer 2_1 and convolutional layer 2_2_1. In Figure 12, o and ㊉ are stepwise multiplication and stepwise addition operations. For example, convolutional layer 2" uses a CX3x3 kernel, with a kernel size of 3x3 and an output channel number of Co. In Figure 12, W represents the width of the input image patch (or the number of rows in the vector matrix), h represents the height of the input image patch (or the number of columns in the vector matrix), and n represents the stride. For example, activation function layer 2" can use the Leaky ReLU function. The Leaky ReLU function is used to assign a non-zero slope to all negative values. Encoding network submodule 2 can use a residual non-local attention block (RNAB) to provide an attention mechanism, such as providing spatial global or local attention information. For example, see Figure 13 for a possible RNAB structure diagram. RNAB uses a network structure including multiple residual block (RB) layers, multiple convolutional layers, deconvolutional layers, and activation function layers (e.g., using the sigmoid function). Spatial global or local attention information is extracted through RNAB. As an example, the residual block layer can use the network structure shown in Figure 14. The RB layer may include a convolutional layer 4_1, an activation function layer 4_11, and a convolutional layer 4_2. In some embodiments, the activation function layer 4_11 may use the LeakyReLU function. The main line of the residual module layer in Figure 14 passes the input features through a 3x3 convolutional layer 4_1 to obtain the feature matrix, then outputs it through an activation function, and then adds the result of a 3x3 convolutional layer 4_2 to the input features. Referring to Figure 11, the encoding network submodule 3 includes a convolutional layer 5_1, a residual activation function layer 5_1, a padding layer 5_2, a convolutional layer 5_2, and a convolutional layer 5_3. In Figure 11, convolutional layers 5_1 and 5_2 may use a conv 128x3x3 array.S2. Convolutional layer 5_3 uses a conv 128x1x1 Slo residual activation function layer 5” can use a ResAU 3x3 no Tanh, such as the network structure shown in Figure 12. The quantization network can include a Round layer 611 for performing quantization operations, or rounding operations. It returns the rounded value of a floating-point number. In some embodiments, referring to Figure 11, the quantization network can also include a Gunit layer 6” and an invGunit layer 6_21. The Gunit layer 6” and invGunit layer 6_21 are used to perform rate matching, enabling the encoding network to adjust the rate. As an example, the Gunit layer 6” and invGunit layer 6_21 can adopt the Gain and Inverse Gain structure from the paper "G-VAE: A CONTINUOUSLY VARIABLE RATE DEEP IMAGE COMPRESSION FRAMEWORK" by Ze Cui, Jing Wang et al. In Figure 11, the edge information extraction module can include a Hyper Encoder Net and a round layer. Hyper Encoder Net, also known as a hyper-encoder network, extracts side information z from the quantized image features y. In Figure 11, the autoregressive network can include Context Model Net and Prediction Fusion Net. Context Model Net is an autoregressive process. It uses the already encoded element yjl (i' Wi, j' ej — i — j') information combined with Prediction Fusion Net to predict the expected value of the element to be encoded. For example, mask convolution can be used in its implementation. Prediction Fusion Net receives the already encoded element information extracted by Context Model Net and the side information extracted by Hyper Decoder to predict the expected value (or predicted value) of the element to be encoded. The expected value of the element to be encoded can be understood as the mean of the predicted probability distribution. In some possible implementation scenarios, feature probability estimation module A can use Hyper Decoder Net, and feature probability estimation module B can use Hyper Scale Decoder Net. Of course, other network structures can also be used; any network capable of probabilistic estimation of edge information is applicable to this application. As an example, feature probability estimation module A can...The network structure of the super-decoding network shown in Figure 15 is adopted. The feature probability estimation module A includes a convolutional layer 7_1, a deconvolutional layer 7_21, a cropping layer 7_31, a convolutional layer 7_2, a deconvolutional layer 7_22, a cropping layer 7_22, an activation function layer 7_32, a convolutional layer 7_3, and an activation function layer 7_33. Crop layers 7_21 and 7_22 are used to perform crop operations on the input tensor. The cropping operation can be represented as Crop{Hout, Wout, d, sd), where xout and wout are the width and height of the final output image of the decoding network (or can be understood as the size of the input image of the encoding network, which can be obtained from the bitstream header), and sd is the stride information of the deconvolution operation. As an example, sd = 2. d represents the depth of the deconvolutional layer. The Crop layer takes a tensor of size [C, 5, Sd, Wd] as input and outputs a tensor of size [&l-1, y / -1], where φ = ce〃(%-i / Sd); wd = ceiKw^ / si), h0 = Hout, w0 = Wout. As an example, Figure 11 shows activation function layers 7_31, 7_32, and 7_33 using the LeakyReLU function. Figure 11 shows an example with convolutional layer 7_n using conv 128X1X1 S1, deconvolutional layer 7_n using DConv 128X4X4 S2, convolutional layer 7_2 using conv 128X3X3 S1, deconvolutional layer 7_12 using DConv 128X4X4 S2, and convolutional layer 7_3 using conv 128X3X3 S1. As another example, the feature probability estimation module B can adopt the network structure of the ultra-large-scale decoding network shown in Figure 16. The feature probability estimation module B includes a deconvolutional layer 713, a cropping layer 723, an activation function layer 734, a convolutional layer 74, an activation function layer 734, a deconvolutional layer 714, a cropping layer 724, an activation function layer 735, a convolutional layer 75, and a Gunit layer 76. In Figure 11, the deconvolutional layer 713 uses a DConv 128X4X4 S2, the activation function layer 734 uses the LeakyReLU function, the convolutional layer 74 uses a conv 128X3X3 S1, the deconvolutional layer 714 uses a DConv 128X4X4 S2, the activation function layer 735 uses the LeakyReLU function, and the convolutional layer 75 uses a conv 128X3X3 S1.Taking S1 as an example. In Figure 11, the entropy encoding network uses a lossless encoder, whose function is to convert the features to be encoded into a bitstream. The structure of each sub-network of the above decoding network (including the first and second decoding networks) is illustrated below with specific examples. See Figure 17, which is a schematic diagram of the decoding network's execution flow. It should be noted that Figure 17 is only an example and does not limit the specific structure of the decoding network. In Figure 17, the entropy decoding network uses a lossless decoder, whose function is to restore the bitstream to be decoded into features. The probability estimation network in the decoding network can adopt the same structure as the encoding network, as shown in Figure 11. Decoding network sub-module 1 includes an invGunit layer, a lightweight residual module (LightResBlock), a deconvolution layer 8_1, a crop layer 8_11, and a residual activation function layer 8_21. Decoding network sub-module 2 includes a deconvolution layer 8_2, a crop layer 8_12, and a residual activation function layer 8_22. Residual activation function layers 8-21 and 22 can adopt a ResAU structure. Deconvolution layer 81 can adopt a Dconv 96X4X4 S2 structure, and deconvolution layer 8_2 can adopt a Dconv 64X4X4 S2 structure. Decoding network submodule 3 can include convolutional layer 8_31, residual activation function layer 8_23, convolutional layer 8_32, PxlShuffleS4, and pruning layer 8_13. Decoding network submodule 4 includes deconvolutional layer 8_3, RNAB, pruning layer 8_14, residual activation function layer 8_24, deconvolutional layer 8_4, and pruning layer 8_15. For example, RNAB can adopt the network structure shown in Figure 13. As an example, the network structure of the lightweight residual module (LightResBlock) can be seen in Figure 18. PxlShuffleS4: represents a pixel shuffle operation with a 4x upsampling. Example 2: Example 2 uses the same encoding network as Example 1, and the execution flow is similar, so it will not be repeated here. The decoding network in Example 2 is different from that in Example 1. In Example 2, the second decoding network is a subnetwork of the first decoding network, or in other words, the second image restoration network of the second decoding network is a subnetwork of the structure of the first image restoration network in the first decoding network. As shown in Figure 19, the second image restoration network includes decoding network submodule 5 and decoding network submodule 7, and the first image restoration network includes decoding network submodules 5-7. Unlike Example 1, in Example 2, at the decoding end, when the Profile ID is different, the selection is not between decoding network submodule 3 and decoding network submodule 4, but rather whether to skip a certain decoding network submodule. When Profile ID=1, decoding network submodule 6 is skipped.When Profile ID=0, the decoding network submodule 6 is executed. Example 3: Example 3 uses two different feature extraction networks for the two encoding networks and two different image restoration networks for the two decoding networks. Referring to Figure 20A, the encoding network includes a first feature extraction network, a second feature extraction network, a quantization network, an autoregressive network, a side information extraction network, a probability estimation network, and an entropy encoding network. Correspondingly, referring to Figure 20B, the decoding network includes an autoregressive network, a side information extraction network, a probability estimation network, a first image restoration network, a second image restoration network, and a decryption network. As shown in Figure 20A, the implementation process at the encoding end differs from Example 1 in the first step. During the calculation of the input image feature y, different feature extraction networks are selected based on the Profile ID. When Profile ID=0, the first feature extraction network is selected; when Profile ID=1, the second feature extraction network is selected. Similarly, the implementation process at the decoding end differs from Example 1 in the fifth step, where the image is restored from y through the image restoration network. During the operation of the decoding network, image restoration networks with different structures are selected based on the Profile ID. When Profile ID=0, the first image restoration network is selected; when Profile ID=1, the second image restoration network is selected. The structure of each subnetwork of the above encoding network (including the first and second encoding networks) is illustrated below with specific examples. See Figure 21 for a schematic diagram of the encoding network's execution flow. It should be noted that Figure 21 is only an example and does not limit the specific structure of the encoding network. In Figure 21, the first feature extraction network includes a padding layer 1_11, a convolutional layer 1_11, a residual activation function layer 1_21, a padding layer 12, a convolutional layer 1_12, a residual activation function layer 1_22, a padding layer 13, a convolutional layer 5_1, a residual activation function layer 5_1, a padding layer 5_2, a convolutional layer 5_3. The second feature extraction network includes a padding layer 1_1, a convolutional layer 1_11, a residual activation function layer (ResAU) 1_21, padding 12, a convolutional layer 1_12, a residual activation function layer 1_22, padding 1_3, a convolutional layer 5_1, a residual activation function layer 5_1, a padding layer 5_12, a convolutional layer 5_2, and a convolutional layer 5_3. For a description of each of these layers, please refer to the relevant description in the embodiment corresponding to Figure 11, which will not be repeated here. The structures of the other networks in Figure 21 are as described in Example 1, and will not be repeated here.Referring to Figure 22, which is a schematic diagram of the execution flow of the decoding network, it should be noted that Figure 22 is only an example and does not limit the specific structure of the encoding network. The first image restoration network includes the invGunit layer, the lightweight residual module (LightResBlock), deconvolution layer 8_1, crop layer 8_11, residual activation function layer 8_21, deconvolution layer 8_2, crop layer 8_12, residual activation function layer 8_22, deconvolution layer 8_3, RNAB, crop layer 8_14, residual activation function layer 8_24, deconvolution layer 4, and crop layer 8_15. The second image restoration network includes an invGunit layer, a lightweight residual module (LightResBlock), a deconvolution layer 8_11, a crop layer 8_21, a residual activation function layer 8_2, a deconvolution layer 8_2, a crop layer 12, a residual activation function layer 8_22, a convolutional layer 8_31, a residual activation function layer 8_23, a convolutional layer 8_32, a PxlShuffle layer 8_4, and a crop layer 8_13. For a description of each of the above layers, please refer to the relevant description of the embodiment corresponding to Figure 11, which will not be repeated here. The structures of the other networks in Figure 22 are as described in Example 1, and will not be repeated here. It should be noted that the structures of the above networks are only examples and do not specifically limit the specific network structure. Any network structure that can achieve the corresponding function is applicable to this application. Furthermore, it should be noted that the above adjustments to the encoding / decoding network at the sub-module level and the overall network level based on the Profile ID can be flexibly combined. For example, one possible scenario is that the encoding end executes a dynamic computation graph on the cloud side using frameworks such as PyTorch, adjusting the encoding network sub-modules based on the Profile ID, while the decoding end executes a static computation graph on the edge side, switching the entire decoding network based on the Profile information. This application embodiment transmits network structure information through the encoding end, allowing the decoding end to adjust the decoding network based on the bitstream content. This scheme has the following advantages: 1. For bitstreams generated using different AI encoding networks, the decoding end can select different decoding network structures based on the bitstream content to achieve decoding. This provides greater flexibility to the encoding and decoding ends, allowing users to adjust the computing power of their encoding and decoding networks according to their own scenarios to flexibly balance latency and compression performance. 2. Users can dynamically select and adjust parts of the decoding network based on the Profile ID, or switch between different decoding networks, according to their own usage scenarios. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the communication system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. One embodiment of this application provides...A computer-readable medium is provided for storing a computer program including instructions for performing method steps in the method embodiment corresponding to FIG5. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and limitations. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowcharts and / or one or more block diagrams. It is obvious that those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations. 19 5 10 15 20 25 30 35 40 45 Claims WO 2024 / 222681 PCT / CN2024 / 089342 1. An image encoding method, characterized in that it includes: encoding identification information for indicating the decoding network used into a bitstream; wherein the identification information is a first value, used to indicate that the decoding network used to decode the image to be processed from the bitstream is a first decoding network; or, the identification information is a second value, used to indicate that the decoding network used to decode the image to be processed from the bitstream is a second decoding network; the processing resources required by the first decoding network are higher than the processing resources required by the second decoding network; and transmitting the bitstream. 2. The method of claim 1, characterized in that the first decoding network and the second decoding network are completely different decoding networks; or, the first decoding network and the second decoding network share some sub-networks; or, the second decoding network is a sub-network of the first decoding network. 3. The method as described in claim 1 or 2, characterized in that the method further comprises: acquiring the identification information; when the identification information is a first value, encoding residual information obtained by encoding the image to be processed based on the first coding network into the bitstream; or,When the identification information is a second value, the bitstream is encoded based on the residual information obtained by encoding the image to be processed using the second encoding network; the processing resources required by the first encoding network are higher than those required by the second encoding network. 4. The method as described in claim 3, wherein the first encoding network and the second encoding network are two different encoding networks; or, the first encoding network and the second encoding network share some sub-networks; or, the second encoding network is a sub-network of the first encoding network. 5. The method as described in claim 4, wherein the first encoding network comprises a first feature extraction network, an autoregressive network, an edge information extraction network, and a probability estimation network; the residual information obtained by encoding the image to be processed using the first encoding network comprises: extracting a three-dimensional feature map of the image to be processed through the first feature extraction network, the three-dimensional feature map comprising multiple feature elements; extracting edge information of the feature elements to be encoded from the three-dimensional feature map through the edge information extraction network; estimating the first probability distribution mean of the feature elements to be encoded based on the edge information through the probability estimation network; inputting the encoded feature elements and the first probability distribution mean into the autoregressive network to obtain a second probability distribution mean of the feature elements to be encoded; and obtaining the residual information of the feature elements to be encoded based on the feature elements to be encoded and the second probability distribution mean of the feature elements to be encoded. 6. The method of claim 5, wherein the second encoding network comprises a second feature extraction network, a side information extraction network, and a probability estimation network; the residual information obtained by encoding the image to be processed using the second encoding network includes: extracting a three-dimensional feature map of the image to be processed through the second feature extraction network, the three-dimensional feature map comprising multiple feature elements; extracting side information of the feature elements to be encoded from the three-dimensional feature map through the side information extraction network; estimating the mean probability distribution of the feature elements to be encoded based on the side information through the probability estimation network; and obtaining the difference information of the feature elements to be encoded based on the feature elements to be encoded and the mean probability distribution. 7. The method of claim 6, wherein the second feature extraction network is a sub-network of the first feature extraction network, or the second feature extraction network and the first feature extraction network are two completely different sub-networks. 20 WO 2024 / 222681 PCT / CN2024 / 089342 5 10 15 20 25 30 35 40 45 8. The method according to any one of claims 5-7, characterized in that the method further comprises: encoding the side information into the bitstream. 910. The method according to any one of claims 1-8, wherein the identification information is located in the bitstream header. A method for image decoding, comprising: receiving a bitstream; decoding identification information from the bitstream indicating the decoding network used; when the identification information is a first value, decoding an image to be processed from the bitstream using a first decoding network; or, when the identification information is a second value, decoding the image to be processed from the bitstream using a second decoding network; wherein the processing resources required by the first decoding network are higher than those required by the second decoding network. 11. The method according to claim 10, wherein the first decoding network and the second decoding network are completely different decoding networks; or, the first decoding network and the second decoding network share some sub-networks; or, the second decoding network is a sub-network of the first decoding network. 12. The method as described in claim 10 or 11, wherein the first decoding network comprises an entropy decoding network, a probability estimation network, an autoregressive network, and a first image restoration network; decoding the image to be processed from the bitstream using the first decoding network comprises: decoding the edge information of a three-dimensional feature map of the image to be processed from the bitstream using the entropy decoding network; the three-dimensional feature map comprises multiple feature elements; estimating the mean of a first probability distribution of the feature elements to be decoded based on the edge information using the probability estimation network; determining the mean of a second probability distribution of the feature elements to be decoded based on the mean of the first probability distribution and the already decoded feature elements using the autoregressive network; decoding the residual information of the feature elements to be decoded from the bitstream based on the mean of the second probability distribution using the entropy decoding network, and obtaining the feature elements to be decoded based on the residual information and the mean of the second probability distribution; and restoring the image to be processed based on the decoded three-dimensional feature map using the first image restoration network. 13. The method of claim 12, wherein the second decoding network comprises the fine decoding network, the probability estimation network, and the second image restoration network; decoding the image to be processed from the bitstream using the second decoding network comprises: decoding the edge information of a three-dimensional feature map of the image to be processed from the bitstream using the entropy decoding network; the three-dimensional feature map comprises multiple feature elements; estimating the mean of a first probability distribution of the feature elements to be decoded based on the edge information using the probability estimation network; decoding the residual information of the feature elements to be decoded from the bitstream based on the mean of the first probability distribution using the entropy decoding network, and obtaining the feature elements to be decoded based on the residual information and the mean of the first probability distribution; and restoring the image to be processed based on the decoded three-dimensional feature map using the second image restoration network. 14The method as described in claim 13, wherein the second image restoration network is a sub-network of the first image restoration network, or the image restoration network and the first image restoration network share a common sub-network, or the second image restoration network and the first image restoration network are two different networks. 15. An image encoding apparatus, comprising a memory and a video encoder, wherein the memory is used to store video data, the video data including an image to be processed; the video encoder is used to encode identification information indicating the decoding network used into a bitstream; wherein the identification information is a first value, indicating that the decoding network used to decode the image to be processed from the bitstream is a first decoding network; or, the identification information is a second value, indicating that the decoding network used to decode the image to be processed from the bitstream is a second decoding network; the processing resources required by the first decoding network are higher than the processing resources required by the second decoding network. 16. An image decoding apparatus, characterized in that it comprises a memory and a video decoder, wherein the memory is used to store video data in the form of a bitstream, the video data including an image to be processed; the video decoder is used to decode identification information from the bitstream indicating the decoding network used; when the identification information is a first value, a first decoding network is used to decode the image to be processed from the bitstream; or, when the identification information is a second value, a second decoding network is used to decode the image to be processed from the bitstream; the processing resources required by the first decoding network are higher than those required by the second decoding network. 17. A video decoding device, characterized in that it comprises: a non-volatile memory and a processor coupled to each other, the processor calling program code stored in the memory to perform the method as described in any one of claims 10-14. 18. A video encoding device, characterized in that it comprises: a non-volatile memory and a processor coupled to each other, the processor calling program code stored in the memory to perform the method as described in any one of claims 1-14. 19. A computer-readable storage medium, characterized in that the computer-readable storage medium stores program code that, when the computer program is run on a computer, causes the computer to perform the method as described in any one of claims 10 to 14. 20. A computer-readable storage medium, characterized in that the computer-readable storage medium stores program code that, when the computer program is run on a computer, causes the computer to perform the method as described in any one of claims 1 to 9. 21. A computer-readable storage medium, characterized in that the computer-readable storage medium stores information processed by one or more processing methods.The device executes the method of any one of claims 10 to 14 to decode the video bitstream. 22. A computer-readable storage medium, characterized in that the computer-readable storage medium stores a video bitstream encoded by the method of any one of claims 1 to 9, executed by one or more processors. 22 WO 2024 / 222681 PCT / CN2024 / 089342 Figure 1 1 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Convolutional Neural Network (CNN) 100 Figure 2 Data to be processed Figure 3 2 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Figure 4 Figure 5 3 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 First coding network diagram 6A Second coding network diagram 6B 4 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 First coding network diagram 7A Second coding network diagram 7B 5 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 I---------- I Second Decoding Network First Decoding Network Two-Image Complex Network 2 Figure 8 Figure 8 First Decoding Network 903a, 902a, Figure 9A 6 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Second Decoding Network 902b, Figure 9B 7 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Figure 10A 8 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Figure 10B 9 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Figure 11 Input [N,C,H,W] Res AU 3x3 no Tanh Figure 12 10 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Figure 13 11 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Feature Probability Estimation Module B Gunit Layer 7_6 Convolutional Layer 7_5 Activation Function Layer 7_35 7_24 Deconvolutional Layer 7_14 Activation Function Layer 7_34 Convolutional Layer 7_4 Activation Function Layer 7_34 Triple Layer 7_23 Deconvolutional Layer 7_13 Figure 16 12 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 - SPOJ 9 PSS - SS 0 7 Flowing es Infant 1Through Context Model Net Prediction Fusion Net *z κ ^ Double 8 NCQ 0 V kI 8 & 13 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Figure 18 Code Stream 2 Figure 19 14 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Figure 20A 15 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Figure 20B 16 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 I Z ® I Z| r- ( J O P O Q U g S°SOJ ) A w I r- ( X U B L I P A ) V * η * " Z 1| S 3 」 Z | 口 3 」 1 1 S N u £ s n 工 u o ap ald S N Ep o n * Inv-Unit layer 6_ 21 g 9 — 口 Z ii C Z 8 X V 」 *3 口 K」 「 1 17 / 18 WO 2024 / 222681 PCT / CN2024 / 089342 Context Model Net Prediction Fusion Net - i ■ I g p o ' J O P S-SOJ * Γ-· Γ- rJ IΓ- 7 rJ I Γ- : *k 9 CN (N I ε ® HS 3 f 8 2 Z 1| QO kq co & q d3 «nr EX 18 / 18 INTERNATIONAL SEARCHREPORT International application No. PCT / CN2024 / 089342 A. CLASSIFICATION OF SUBJECT MATTER H04N19 / 184(2014.01)i; G06N3 / 02(2006.01)i According to International Patent Classification (IPC) or to both national classification and IPC B. FIELDS SEARCHED Minimum documentation searched (classification system followed by classification symbols) IPC: H04N, G06N Documentation searched other than minimum documentation to the extent that such documents are included in the fields searched Electronic database consulted during the international search (name of database and, where practicable, search terms used) CNABS, CNTXT, DWPI, ENTXTC, WPABSC, CNKI, IEEE, JVET: coding, compression, decoding, marking, identification, marking, grade, flag, residual, resource, first, second, neural network, probability, code stream, feature map, indication, code, encode, compress, decode, identification, profile, profile, flag, residual, resource, first, second, neural network, probability, bitstream, feature map, indicate C. DOCUMENTS CONSIDERED TO BE RELEVANT Category* Citation of document, with indication, where appropriate, of therelevant passages Relevant to claim No. X WO 2022257134 Al (GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.) 15 December 2022 (2022-12-15) description, page 3, line 39-page 16, line 20 1-4, 9-11, 15-22 Y WO 2022257134 Al (GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.) 15 December 2022 (2022-12-15) description, page 3, line 39-page 16, line 20 5-8, 12-14 X CN 104853209 A (TONGJI UNIVERSITY et al.) 19 August 2015 (2015-08-19) description, paragraphs 84-263 1-4, 9-11, 15-22 Y CN 104853209 A (TONGJI UNIVERSITY et al.) 19 August 2015 (2015-08-19) description, paragraphs 84-263 5-8, 12-14 Y CN 111641832 A (HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO., LTD.) 08 September 2020 (2020-09-08) description, paragraphs 155-251 5-8, 12-14 | | Further documents are listed in the continuation of Box C. | / 1 See patent family annex. * Special categories of cited documents: “T" later document published after the international filing date or priority“A" document defining the general state of theart which is not considered date and not in conflict with the application but cited to understand theto be of paiticulai' relevance principle or theory underlying the invention"D” document cited by the applicant in the international application “χ” document of particular relevance; the claimed invention cannot be“E” eailier application or patent but published on or after the international considered novel or cannot be considered to involve an inventive stepfiling date when the document is taken alone“L" document which may thi'ow doubts on priority claim(s) or which is “Y" document of paiticulai' relevance; the claimed invention cannot be cited to establish the publication date of another citation or other considered to involve an inventive step when the document isspecial reason (as specified) combined with one or more other such documents, such combination"O” document refen'ing to an oral disclosure, use, exhibition or other being obvious to a person skilled in the artmeans documentmember of the same patent family"P” document published prior to the international filing date but later thanthe priority date claimed Date of the actual completion of the international search 18 July 2024 Date of mailing of the international search report 18 July 2024 Name and mailing address of the ISA / CN China National Intellectual Property Administration (ISA / CN) China No. 6, Xitucheng Road, Jimenqiao, Haidian District, Beijing 100088 Authorized officer Telephone No. Form PCT / ISA / 210 (second sheet) (July 2022) INTERNATIONAL SEARCH REPORT PCT / CN2024 / 089342 International application No. C. DOCUMENTS CONSIDERED TO BE RELEVANT Category* Citation of document, with indication, where appropriate, of the relevant passages Relevant to claim No. A US 2022166976 Al (ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE) 26 May 2022 (2022-05-26) entire document 1-22 Form PCT / ISA / 210 (second sheet) (July 2022) INTERNATIONAL SEARCH REPORT Information on patent family members Internationalapplication No. PCT / CN2024 / 089342 Patent document cited in search report Publication date (day / month / year) Patent family member(s) Publication date (day / month / year) WO 2022257134 Al 15 December 2022 None CN 104853209 A 19 August 2015 JP 2017509279 A 30 March 2017 JP 6659586 B2 04 March 2020 EP 3107289 Al 21 December 2016 EP 3107289 A4 15 February 2017 KR 20160124190 A 26 October 2016 KR 102071764 Bl 30 January 2020 WO 2015120818 Al 20 August 2015 US 2017054988 A1 23 February 2017 CN 111641832 A 08 September 2020 None US 2022166976 A1 26 May 2022 US 12003719 B2 04 June 2024 Form PCT / ISA / 210 (patent family annex) (July 2022) International Search Report International Application Number PCT / CN2024 / 089342 A. Subject Classification H04N19 / 184(2014.01)i; G06N3 / 02(2006.01)i According to the International Patent Classification (IPC) or both national classification and IPC classification B. Search Field Minimum Documents Searched (Indicate classification system and classification number) IPC: H04N, G06N Electronic databases consulted during international searches, excluding the minimum required literature, included in the search domain (database names and search terms used, such as those for CNABS, CNTXT, DWPLENTXTC, WPABSC, CNKI, IEEE JVET): encoding, compression, decoding, tagging, identifier, flag, residual, resource, first, second, neural network probability, bitstream, feature map, indicator, code, encode, compress, decode, identification, profile.Profile, flag, residual, resource, first, second, neural network, probability, bitstream, feature map, indicateC. Related File Types* Referenced Documents, specifying relevant paragraphs where necessary. Relevant Claims X WO 2022257134 A1 (GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD) December 15, 2022 (2022-12-15) Specification, Page 3, Line 39 - Page 16, Line 20, 1-4, 9-11, 15-22 Y WO 2022257134 A1 (GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD) December 15, 2022 (2022-12-15) Specification, Page 3, Line 39 - Page 16, Line 20, 5-8, 12-14 X CN 104853209 A (Tongji University, etc.) August 19, 2015 (2015-08-19) Instruction Manual, Sections 84-263, Paragraphs 1-4, 9-11, 15-22 Y CN 104853209 A (Tongji University, etc.) August 19, 2015 (2015-08-19) Instruction Manual, Sections 84-263, Paragraphs 5-8, 12-14 Y CN 111641832 A (Hangzhou Hikvision Digital Technology Co., Ltd.) September 8, 2020 (2020-09-08) Instruction Manual, Sections 155-251, Paragraphs 5-8, 12-14 A US 2022166976 A1 (ELECTRONICS & TELECOMMUNICATIONS RES INST) May 26, 2022 (2022-05-26) Full text 1-22 *Specific types of cited documents: “A” Documents deemed not particularly relevant that represent the general state of the prior art; “D” Documents cited by the applicant in an international application; “E” Earlier applications or patents published on or after the international filing date; “L” Documents that may cast doubt on the priority claim, or documents cited to determine the publication date of another cited document, or documents cited for other specific reasons (as specifically stated); Documents involving □head disclosure, use, exhibition, or other forms of disclosure; “P” Documents published before the international filing date but later than the claimed priority date; □ Other documents are listed on the next page in column C. See the patent family appendix."1" is published after the claim date or priority date, does not conflict with the application, but is a document that is particularly different from the later document "X" for understanding the inventive theory or principle. Considering only that document, the claimed invention is not novel or lacks inventiveness. "Y" is a document particularly related to the claim, and when that document is not combined with any of the other documents of the same class and such combination is obviously difficult for those skilled in the art to understand, the claimed invention lacks inventiveness. International search of patent family documents: Date of actual completion: July 18, 2024; International search report mailing date: July 18, 2024; Name and mailing address of ISA / CN: China National Intellectual Property Administration, No. 6, Tucheng Road, Xijimenqiao, Haidian District, Beijing, 100088, China; Authorized Officer: He Meiling; Telephone: (+86) 010-62412254; PCT / ISA / 210 Form (Page 2) (July 2022) International Search Report Information on Patent Families International Application No. PCT / CN2024 / 089342 Publication Dates of Patent Documents Cited in the Search Report (Year / Month / Day) Publication Dates of Patent Families (Year / Month / Day) WO 2022257134 A1 December 15, 2022 None CN 104853209 A August 19, 2015 JP 2017509279 A March 30, 2017 JP 6659586 B2 March 4, 2020 EP 3107289 A1 December 21, 2016 EP 3107289 A4 February 15, 2017 KR 20160124190 A October 26, 2016 KR 102071764 B1 January 30, 2020 WO 2015120818 Al August 20, 2015 US 2017054988 Al February 23, 2017 CN 111641832 A September 8, 2020 None US 2022166976 Al May 26, 2022 US 12003719 B2 June 4, 2024 PCT / ISA / 210 Form (Appendix to Patent Family) (July 2022) (19) *EP004694127A1* (11) EP 4 694 127 A1 (12) EUROPEAN PATENT APPLICATION published in accordance with Art. 153(4) EPC (43) Date of publication: 11.02.2026 Bulletin 2026 / 07 (21) Application number: 24796061.0 (22) Date of filing:23.04.2024 (51) International Patent Classification (IPC): H04N 19 / 184 (2014.01) G06N 3 / 02 (2006.01) (52) Cooperative Patent Classification (CPC): G06N 3 / 02; G06T 9 / 00; H04N 19 / 124; H04N 19 / 184; H04N 19 / 85 (86) International application number: PCT / CN2024 / 089342 (87) International publication number: WO 2024 / 222681 (31.10.2024 Gazette 2024 / 44) (84) Designated Contracting States: AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR Designated Extension States: BA Designated Validation States: GE KH MA MD TN (30) Priority: 24.04.2023 CN 202310476967 28.07.2023 CN 202310956879 (71) Applicant: Huawei Technologies Co., Ltd. Shenzhen, Guangdong 518129 (CN) (72) Inventors: • YU, Dequan Shenzhen, Guangdong 518129 (CN) • ZHAO, Yin Shenzhen, Guangdong 518129 (CN) • ALSHINA, Elena Alexandrovna Shenzhen, Guangdong 518129 (CN) (74) Representative: Eisenführ Speiser Patentanwälte Rechtsanwälte PartGmbB Johannes-Brahms-Platz 1 20355Hamburg (DE) (54) IMAGE CODING METHOD, IMAGE DECODING METHOD AND APPARATUS (57) A picture encoding and decoding method and apparatus are provided, and relate to the artificial intelli- gence field and the picture compression field, to provide an encoding and decoding scheme, thereby meeting requirements of different application scenarios. Accord- ing to the encoding and decoding method provided in this application, a used encoder and decoder network may be determined based on profile information (or identification information). That is, a codec may select corresponding profile information based on a capability of a decoding device, to select or indicate different encoder and deco- der networks. In this way, the network may not only have a capability of adapting to a terminal side with low comput- ing power, but also have a capability of adapting to a terminal side with higher computing power. EP 4 69 4 12 7 A 1 Processed by Luminess, 75001 PARIS (FR) 2 1 EP 4 694 127 A1 2 DescriptionCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priorities to Chinese Patent Application No. 202310476967.6, filed with the China National Intellectual Property Administration on April 24, 2023 and entitled "PICTURE ENCODING AND DECODING METHOD AND APPARATUS", and to Chinese Patent Application No. 202310956879.6, filed with the China National Intellectual Property Administra- tion on July 28, 2023 and entitled "PICTURE ENCODING AND DECODING METHOD AND APPARATUS", both of which are incorporated herein by reference in their entire- ties. TECHNICAL FIELD
[0002] This application relates to the field of picture compression technologies and the field of artificial intelli- gence technologies, and in particular, to a picture encod- ing and decoding method and apparatus. BACKGROUND
[0003] Many consumer applications (such as news, social, and shopping networking applications) require that picture decoding be completed on terminal-side devices with low computing power(such as mobile phones, personal PCs, and televisions). In some other industrial applications, picture decoding is allowed to be completed on terminal-side devices with higher comput- ing power (such as GPU workstations equipped with independent graphics cards), and higher requirements are posed on picture compression rates.
[0004] A current neural network-based picture encod- ing and decoding scheme usually has a fixed network structure, and cannot meet requirements of different application scenarios. SUMMARY
[0005] Embodiments of this application provide a pic- ture encoding and decoding method and apparatus, to provide an encoding and decoding scheme, thereby meeting requirements of different application scenarios.
[0006] According to a first aspect, an embodiment of this application provides a picture encoding method, including: encoding identification information indicating a used decoder network into a bitstream, where the identification information is a first value, indicat- ingthat the decoder network used to decode the bitstream to obtain a to-be-processed picture is a first decoder network; or the identification information is a second value, in- dicating that the decoder network used to decode the bitstream to obtain a to-be-processed picture is a second decoder network, where a processing re- source required by the first decoder network is higher than a processing resource required by the second decoder network; and sending the bitstream.
[0007] The identification information may also be re- ferred to as profile information (Profile ID).
[0008] According to the foregoing solution in this em- bodiment of this application, a transmit end indicates a receive end to use a network structure, so that different network structures can implement different decoding performance, thereby improving flexibility of a decoder side. A user may adjust encoder and decoder network computing power of the user based on a scenario of the user, to flexibly balance a delay andcompression per- formance.
[0009] In a possible implementation, the first decoder network and the second decoder network are completely different decoder networks, or the first decoder network and the second decoder network share a part of subnet, or the second decoder network is a subnet of the first decoder network.
[0010] If the second decoder network is a subnet of the first decoder network, it may be understood that, when the identification information is the second value, some network layers in the first decoder network are skipped, that is, a process of performing decoding by using the second decoder network is implemented.
[0011] In a possible implementation, the method further includes: obtaining the identification information; and when the identification information is the first value, encoding, into the bitstream, residual information obtained by encoding the to-be-processed picture by using a first encoder network; or when the identification information is the second value,encoding, into the bitstream, residual informa- tion obtained by encoding the to-be-processed pic- ture by using a second encoder network, where a processing resource required by the first encoder network is higher than a processing resource re- quired by the second encoder network.
[0012] In the foregoing solution, for bitstreams gener- ated by using different AI encoder networks, a decoder side may select different decoder network structures by using bitstream content, to implement decoding. This brings high flexibility to a codec side. A user may adjust encoder and decoder network computing power of the user based on a scenario of the user, to flexibly balance a delay and compression performance.
[0013] In a possible implementation, the first encoder network and the second encoder network are two differ- ent encoder networks, or the first encoder network and the second encoder network share a part of subnet, or the 5 10 15 20 25 30 35 40 45 50 55 3 3 EP 4 694 127 A1 4 second encodernetwork is a subnet of the first encoder network.
[0014] In a possible implementation, the first encoder network includes a first feature extraction network, an autoregressive network, a side information extraction network, and a probability estimation network; and the residual information obtained by encoding the to-be- processed picture by using the first encoder network includes: the residual information obtained by encoding the to-be- processed picture by using the first encoder network includes: extracting a three-dimensional feature map of the to- be-processed picture by using the first feature ex- traction network, where the three-dimensional fea- ture map includes a plurality of feature elements; extracting side information of a to-be-encoded fea- ture element from the three-dimensional featuremap by using the side information extraction network; estimating a first probability distribution mean of the to-be-encoded feature element by using the prob- ability estimation networkbased on the side informa- tion; inputting an encoded feature element and the first probability distribution mean into the autoregressive network to obtain a second probability distribution mean of the to-be-encoded feature element; and obtaining residual information of the to-be-encoded feature element based on the to-be-encoded feature element and the second probability distribution mean of the to-be-encoded feature element.
[0015] In a possible implementation, the first encoder network includes a second feature extraction network, a side information extraction network, and a probability estimation network; and the residual information obtained by encoding the to-be- processed picture by using the first encoder network includes: extracting a three-dimensional feature map of the to- be-processed picture by using the second feature extraction network, where the three-dimensional feature map includes a plurality of feature elements; extracting side information of a to-be-encoded fea-ture element from the three-dimensional featuremap by using the side information extraction network; estimating a probability distribution mean of the to- be-encoded feature element by using the probability estimation network based on the side information; and obtaining residual information of the to-be-encoded feature element based on the to-be-encoded feature element and the probability distribution mean.
[0016] In a possible implementation, the second fea- ture extraction network is a subnet of the first feature extraction network, or the second feature extraction net- work and the first feature extraction network are two completely different subnets.
[0017] In a possible implementation, the method further includes: encoding the side information into the bitstream.
[0018] In a possible implementation, the identification information is located in a header of the bitstream.
[0019] According to a second aspect, an embodiment of this application provides a picture decoding method,including: receiving a bitstream; decoding the bitstream to obtain identification infor- mation indicating a used decoder network; and when the identification information is a first value, decoding the bitstream to obtain a to-be-processed picture by using a first decoder network; or when the identification information is a second value, decoding the bitstream to obtain a to-be-processed picture by using a second decoder network, where a processing resource required by the first decoder network is higher than a processing resource re- quired by the second decoder network.
[0020] In a possible implementation, the first decoder network and the second decoder network are completely different decoder networks, or the first decoder network and the second decoder network share a part of subnet, or the second decoder network is a subnet of the first decoder network.
[0021] In a possible implementation, the first decoder network includes an entropy decoder network, a prob- ability estimationnetwork, an autoregressive network, and a first picture restoration network; and decoding the bitstream to obtain the to-be-processed picture by using the first decoder network includes: decoding the bitstream to obtain side information of a three-dimensional feature map of the to-be-pro- cessed picture by using the entropy decoder net- work, where the three-dimensional feature map in- cludes a plurality of feature elements; estimating a first probability distribution mean of a to- be-decoded feature element by using the probability estimation network based on the side information; determining a second probability distribution mean of the to-be-decoded feature element by using the autoregressive network based on the first probability distribution mean and a decoded feature element; decoding the bitstream to obtain residual information of the to-be-decoded feature element by using the entropy decoder network based on the second prob- ability distribution mean, and obtaining the to-be-decoded feature element based on the residual in- formation and the second probability distribution mean; and restoring the to-be-processed picture by using the 5 10 15 20 25 30 35 40 45 50 55 4 5 EP 4 694 127 A1 6 first picture restoration network based on the three- dimensional feature map obtained through decod- ing.
[0022] In a possible implementation, the second deco- der network includes the entropy decoder network, the probability estimation network, and the second picture restoration network; and decoding the bitstream to obtain the to-be-processed picture by using the second decoder network includes: decoding the bitstream to obtain side information of a three-dimensional feature map of the to-be-pro- cessed picture by using the entropy decoder net- work, where the three-dimensional feature map in- cludes a plurality of feature elements; estimating a first probability distribution mean of a to- be-decoded feature element by using the probability estimation network based on theside information; decoding the bitstream to obtain residual information of the to-be-decoded feature element by using the entropy decoder network based on the first probabil- ity distribution mean, and obtaining the to-be-de- coded feature element based on the residual infor- mation and the first probability distribution mean; and restoring the to-be-processed picture by using the second picture restoration network based on the three-dimensional feature map obtained through decoding.
[0023] In a possible implementation, the second pic- ture restoration network is a subnet of the first picture restoration network, or the picture restoration network and the first picture restoration network share a part of subnet, or the second picture restoration network and the first picture restoration network are two different net- works.
[0024] According to a third aspect, an embodiment of this application provides a picture encoding apparatus, including a memory and a video encoder, where thememory is configured to store video data, where the video data includes a to-be-processed picture; and the video encoder is configured to encode identifica- tion information indicating a used decoder network into a bitstream, where the identification information is a first value, indicat- ing that the decoder network used to decode the bitstream to obtain a to-be-processed picture is a first decoder network; or the identification information is a second value, in- dicating that the decoder network used to decode the bitstream to obtain a to-be-processed picture is a second decoder network, where a processing re- source required by the first decoder network is higher than a processing resource required by the second decoder network.
[0025] According to a fourth aspect, an embodiment of this application provides a picture decoding apparatus, including a memory and a video decoder, where the memory is configured to store video data in a bitstream form, where the video data includes a to-be-processed picture; and the video decoder is configured to: decode the bit- stream to obtain identification information indicating a used decoder network; and when the identification information is a first value, decode the bitstream to obtain a to-be-processed picture by using a first decoder network; or when the identification information is a second value, decode the bitstream to obtain a to-be-processed picture by using a second decoder network, where a processing resource required by the first decoder network is higher than a processing resource re- quired by the second decoder network.
[0026] According to a fifth aspect, an embodiment of this application provides a video decoding device, includ- ing a nonvolatile memory and a processor that are coupled to each other, where the processor invokes program code stored in the memory to perform the meth- od described in any implementation of the second as- pect.
[0027] According to a sixth aspect, an embodiment of this applicationprovides a video encoding device, includ- ing a nonvolatile memory and a processor that are coupled to each other, where the processor invokes program code stored in the memory to perform the meth- od described in any implementation of the first aspect or the seventeenth aspect.
[0028] According to a seventh aspect, an embodiment of this application provides a computer-readable storage medium, where the computer-readable storage medium stores program code, and when the computer program is run on a computer, the computer is enabled to perform the method according to any implementation of the sec- ond aspect.
[0029] According to an eighth aspect, an embodiment of this application provides a computer-readable storage medium, where the computer-readable storage medium stores program code, and when the computer program is run on a computer, the computer is enabled to perform the method according to any implementation of the first aspect or the seventeenth aspect.
[0030] According to a ninthaspect, an embodiment of this application provides a computer-readable storage medium, where the computer-readable storage medium stores a video bitstream decoded by one or more pro- cessors according to the method according to any im- plementation of the second aspect.
[0031] According to a tenth aspect, an embodiment of this application provides a computer-readable storage medium, where the computer-readable storage medium stores a video bitstream obtained through encoding by 5 10 15 20 25 30 35 40 45 50 55 5 7 EP 4 694 127 A1 8 one or more processors according to the method accord- ing to any implementation of the first aspect or the se- venteenth aspect.
[0032] According to an eleventh aspect, an embodi- ment of this application provides a computer-readable storage medium, where the computer-readable storage medium stores a bitstream, and the bitstream includes identification information, where the identification information is a first value, indicat- ing that a decoder networkused to decode the bit- stream to obtain a to-be-processed picture is a first decoder network; or the identification information is a second value, in- dicating that a decoder network used to decode the bitstream to obtain a to-be-processed picture is a second decoder network, where a processing re- source required by the first decoder network is higher than a processing resource required by the second decoder network.
[0033] According to a twelfth aspect, an embodiment of this application provides an encoded bitstream, where the encoded bitstream includes a plurality of syntax ele- ments, and the plurality of syntax elements include iden- tification information indicating a decoder network used to decode the bitstream to obtain a to-be-processed picture.
[0034] According to a thirteenth aspect, an embodi- ment of this application provides a video encoder, con- figured to encode a to-be-processed picture. For exam- ple, the video encoder may implement the method ac- cording to thefirst aspect or the seventeenth aspect.
[0035] According to a fourteenth aspect, an embodi- ment of this application provides a video decoder, con- figured to decode a bitstream to obtain a to-be-processed picture. For example, the video encoder may implement the method according to the second aspect.
[0036] According to a fifteenth aspect, an embodiment of this application provides an encoder network, includ- ing: a first feature extraction network, a second feature extraction network, a quantization network, an auto- regressive network, a side information extraction network, and a probability estimation network, where when identification information indicating a used encoder network is a first value, the first feature extraction network extracts a three-dimensional fea- ture map of a to-be-processed picture; or when identification information is a second value, the first feature extraction network extracts a three-dimen- sional feature map of a to-be-processed picture; the sideinformation extraction network extracts side information of the to-be-processed picture from an edge in the three-dimensional feature map; the probability estimation network estimates a first probability distribution mean of a to-be-encoded feature element based on the side information; and when the identification information indicating the used encoder network is the first value, an encoded feature element and the first probability distribution mean are input into the autoregressive network to obtain a second probability distribution mean of the to-be-encoded feature element; and residual infor- mation of the to-be-encoded feature element is ob- tained based on the to-be-encoded feature element and the second probability distribution mean of the to-be-encoded feature element; or when the identification information indicating the used encoder network is the second value, residual information of the to-be-encoded feature element is obtained based on the to-be-encoded feature ele- mentand the first probability distribution mean.
[0037] In a possible implementation, the second fea- ture extraction network is a subnet of the first feature extraction network, or the second feature extraction net- work and the first feature extraction network are two completely different subnets.
[0038] According to a sixteenth aspect, an embodi- ment of this application provides a decoder network, including: an entropy decoder network, a probability estimation network, an autoregressive network, a first picture restoration network, and asecond picture restoration network, where the entropy decoder network decodes a bitstream to obtain side information of a three-dimensional fea- ture map of a to-be-processed picture and identifica- tion information, where the three-dimensional fea- ture map includes a plurality of feature elements; the probability estimation network estimates a first probability distribution mean of a to-be-decoded feature element based on the side information; andwhen the identification information is a first value, the autoregressive network determines a second prob- ability distribution mean of the to-be-decoded fea- ture element based on the first probability distribution mean and a decoded feature element, and the en- tropy decoder network decodes the bitstream to obtain residual information of the to-be-decoded feature element based on the second probability distribution mean, and obtain the to-be-decoded feature element based on the residual information and the second probability distribution mean; and the first picture restoration network restores the to-be- processed picture based on the three-dimensional feature map obtained through decoding; or when the identification information is a second value, the entropy decoder network decodes the bitstream to obtain residual information of the to-be-decoded feature element based on the first probability distri- bution mean, and obtain the to-be-decoded feature element based on the residualinformation and the 5 10 15 20 25 30 35 40 45 50 55 6 9 EP 4 694 127 A1 10 first probability distribution mean; and the second picture restoration network restores the to-be-pro- cessed picture based on the three-dimensional fea- ture map obtained through decoding.
[0039] In a possible implementation, the second pic- ture restoration network is a subnet of the first picture restoration network, or the second picture restoration network and the first picture restoration network share a part of subnet, or the second picture restoration network and the first picture restoration network are two different networks.
[0040] According to a seventeenth aspect, an embodi- ment of this application provides a picture encoding method, including: obtaining identification information; and when the identification information is a first value, encoding, into a bitstream, residual information ob- tained by encoding a to-be-processed picture based on (or by using) a first encoder network; or whenidentification information is a second value, encoding, into a bitstream, residual information ob- tained by encoding a to-be-processed picture based on (or by using) a second encoder network, where a processing resource required by the first encoder network is higher than a processing resource re- quired by the second encoder network.
[0041] In a possible implementation, the method further includes: encoding the identification information into the bitstream.
[0042] In a possible implementation, the identification information further indicates a decoder network used to decode the bitstream to obtain the to-be-processed pic- ture, where the identification information is a first value, indicat- ing that the decoder network used to decode the bitstream to obtain the to-be-processed picture is a first decoder network; or the identification information is a second value, in- dicating that the decoder network used to decode the bitstream to obtain the to-be-processed picture is a seconddecoder network, where a processing re- source required by the first decoder network is higher than a processing resource required by the second decoder network.
[0043] The identification information may also be re- ferred to as profile information (Profile ID).
[0044] In a possible implementation, the first decoder network and the second decoder network are completely different decoder networks, or the first decoder network and the second decoder network share a part of subnet, or the second decoder network is a subnet of the first decoder network.
[0045] In a possible implementation, the first encoder network and the second encoder network are two differ- ent encoder networks, or the first encoder network and the second encoder network share a part of subnet, or the second encoder network is a subnet of the first encoder network.
[0046] In a possible implementation, the first encoder network includes a first feature extraction network, an autoregressive network, a side informationextraction network, and a probability estimation network; and the residual information obtained by encoding the to-be- processed picture by using the first encoder network includes: the residual information obtained by encoding the to-be- processed picture by using the first encoder network includes: extracting a three-dimensional feature map of the to- be-processed picture by using the first feature ex- traction network, where the three-dimensional fea- ture map includes a plurality of feature elements; extracting side information of a to-be-encoded fea- ture element from the three-dimensional featuremap by using the side information extraction network; estimating a first probability distribution mean of the to-be-encoded feature element by using the prob- ability estimation network based on the side informa- tion; inputting an encoded feature element and the first probability distribution mean into the autoregressive network to obtain a second probability distribution mean of theto-be-encoded feature element; and obtaining residual information of the to-be-encoded feature element based on the to-be-encoded feature element and the second probability distribution mean of the to-be-encoded feature element.
[0047] In a possible implementation, the first encoder network includes a second feature extraction network, a side information extraction network, and a probability estimation network; and the residual information obtained by encoding the to-be- processed picture by using the first encoder network includes: extracting a three-dimensional feature map of the to- be-processed picture by using the second feature extraction network, where the three-dimensional feature map includes a plurality of feature elements; extracting side information of a to-be-encoded fea- ture element from the three-dimensional featuremap by using the side information extraction network; estimating a probability distribution mean of the to- be-encoded feature element by using theprobability estimation network based on the side information; and obtaining residual information of the to-be-encoded feature element based on the to-be-encoded feature element and the probability distribution mean. 5 10 15 20 25 30 35 40 45 50 55 7 11 EP 4 694 127 A1 12
[0048] In a possible implementation, the second fea- ture extraction network is a subnet of the first feature extraction network, or the second feature extraction net- work and the first feature extraction network are two completely different subnets.
[0049] In a possible implementation, the method further includes: encoding the side information into the bitstream.
[0050] In this application, based on the implementa- tions provided in the foregoing aspects, the implementa- tions may be further combined to provide more imple- mentations. BRIEF DESCRIPTION OF DRAWINGS
[0051] FIG. 1 is an example block diagram of a coding system according to an embodiment of this applica- tion; FIG. 2 is a diagram of a structure of aconvolutional neural network according to an embodiment of this application; FIG. 3 is a diagram of a deep learning-based video encoder and decoder network according to an em- bodiment of this application; FIG. 4 is a diagram of a structure of a deep learning- based end-to-end video encoder and decoder net- work according to an embodiment of this application; FIG. 5 is a schematic flowchart of an encoding and decoding method according to an embodiment of this application; FIG. 6A is a diagram of a structure of a first encoder network according to an embodiment of this applica- tion; FIG. 6B is a diagram of a structure of a second encoder network according to an embodiment of this application; FIG. 7A is a diagram of an encoding process accord- ing to an embodiment of this application; FIG. 7B is a diagram of another encoding process according to an embodiment of this application; FIG. 8 is a diagram of a structure of a decoder net- work according to an embodiment of this application;FIG. 9A is a diagram of a possible decoding process using a first decoder network according to an embo- diment of this application; FIG. 9B is a diagram of a possible decoding process using a second decoder network according to an embodiment of this application; FIG. 10A is a diagram of an execution process of an encoder network according to an embodiment of this application; FIG. 10B is a diagram of an execution process of a decoder network according to an embodiment of this application; FIG. 11A and FIG. 11B are a diagram of a structure of an encoder network according to an embodiment of this application; FIG. 12 is a diagram of a structure of a ResAU 3×3 no tanh network according to an embodiment of this application; FIG. 13 is a diagram of an RNAB structure according to an embodiment of this application; FIG. 14 is a diagram of a structure of a residual block layer according to an embodiment of this application; FIG. 15 is a diagram of a network structure of a hyper decoder networkaccording to an embodiment of this application; FIG. 16 is a diagram of a network structure of a hyper scale decoder network according to an embodiment of this application; FIG. 17A and FIG. 17B are a diagram of an execution process of a decoder network according to an em- bodiment of this application; FIG. 18 is a diagram of a network structure of a light residual block (LightResBlock) according to an em- bodiment of this application; FIG. 19 is a diagram of a structure of a decoder network according to Example 2 according to an embodiment of this application; FIG. 20A is a diagram of an execution process of an encoder network according to Example 3 according to an embodiment of this application; FIG. 20B is a diagram of a structure of an encoder and decoder network according to Example 3 ac- cording to an embodiment of this application; FIG. 21A and FIG. 21B are a diagram of a structure of an encoder network according to Example 3 accord- ing to an embodiment of this application; andFIG. 22A and FIG. 22B are a diagram of a structure of a decoder network according to Example 3 accord- ing to an embodiment of this application. DESCRIPTION OF EMBODIMENTS
[0052] The following describes embodiments of this application with reference to the accompanying drawings in embodiments of this application. In the following de- scription, reference is made to the accompanying draw- ings, which form a part of the present disclosure and show, by way of illustration, specific aspects of embodi- ments of this application or specific aspects in which embodiments of this application may be used. It should be understood that embodiments of this application may be used in other aspects, and may include structural or logical changes not depicted in the accompanying draw- ings. Therefore, the following detailed descriptions shall not be construed in a limitative sense, and the scope of this application is defined by the appended claims. For example, it should be understood that thedisclosed content with reference to the described method may also be applied to a corresponding device or system for per- forming the method, and vice versa. For example, if one or more specific method steps are described, a corre- sponding device may include one or more units such as 5 10 15 20 25 30 35 40 45 50 55 8 13 EP 4 694 127 A1 14 functional units for performing the described one or more method steps (for example, one unit performs the one or more steps; or a plurality of units, each of which performs one or more of the plurality of steps), even if such one or more units are not explicitly described or illustrated in the accompanying drawings. In addition, for example, if a specific apparatus is described based on one or more units such as a functional unit, a corresponding method may include one step for implementing functionality of one or more units (for example, one step for implement- ing functionality of one or more units; or a plurality of steps, each of which is forimplementing functionality of one or more units in a plurality of units), even if such one or more of steps are not explicitly described or illustrated in the accompanying drawings. Further, it should be understood that features of example embodiments an- d / or aspects described in this specification may be com- bined with each other, unless expressly stated otherwise.
[0053] The technical solutions in embodiments of this application may not only be applied to existing video coding standards (for example, standards such as H.264 and HEVC), but also be applied to future video coding standards (for example, the H.266 standard). Terms used in implementations of this application are merely used to explain specific embodiments of this application, but are not intended to limit this application. The following first briefly describes some related con- cepts in embodiments of this application.
[0054] A picture decoding and encoding method pro- vided in embodiments of this application can beapplied to the video encoding field and the picture encoding field. Specifically, the decoding and encoding method may be applied to album management, human-computer inter- action, video compression or transmission, and picture compression or transmission scenarios.
[0055] An example in which the encoding and decod- ing method is applied to an end-to-end video picture encoding and decoding system is used. The end-to- end video picture encoding and decoding system in- cludes two parts: picture encoding and picture decoding. Picture encoding is determined at a source, and usually includes processing (for example, compressing) an ori- ginal video picture to reduce an amount of data required for representing the video picture (for more efficient storage and / or transmission). Picture decoding is deter- mined at a destination, and usually includes inverse processing relative to an encoder, to reconstruct a pic- ture. A current neural network-based picture encoding and decoding scheme usuallyhas a fixed network struc- ture, for example, an encoding and decoding model in JPEG AI VM1.0. If the network structure is adapted to a capability of a terminal side with low computing power, compression efficiency of the encoding scheme is re- duced to some extent. If the network structure is adapted to computing power of a device with high computing power, the network cannot run on a device with low computing power. In the end-to-end video picture encod- ing and decoding system, by using the encoding and decoding method provided in this application, a used encoder and decoder network may be determined based on profile information. The profile information may also be referred to as identification information or a network identifier, and may have another name. This is not spe- cifically limited in embodiments of this application. The profile information indicates the used decoder network. That is, a codec may select corresponding profile infor- mation based on a capability of adecoding device, to select or indicate different encoder and decoder net- works. In this way, the network may not only have a capability of adapting to a terminal side with low comput- ing power, but also have a capability of adapting to a terminal side with higher computing power.
[0056] Video encoding and decoding generally refers to processing a picture sequence that forms a video or a video sequence. In the video encoding and decoding field, terms "picture (picture)", "frame (frame)", and "im- age (image)" may be used as synonyms.
[0057] FIG. 1 is an example block diagram of a coding system according to an embodiment of this application, for example, a video coding system 10 (or a coding system 10 for short) that may utilize technologies of this application. A video encoder 20 (or an encoder 20 for short) and a video decoder 30 (or a decoder 30 for short) of the video coding system 10 represent examples of devices that may be configured to perform technologies based on variousexamples described in this application.
[0058] As shown in FIG. 1, the coding system 10 includes a source device 12. The source device 12 is configured to provide encoded picture data 21 such as an encoded picture to a destination device 14 configured to decode the encoded picture data 21.
[0059] The source device 12 includes the encoder 20, and optionally, may include a picture source 16, a pre- processor (or preprocessing unit) 18 such as a picture preprocessor, and a communication interface (or com- munication unit) 22.
[0060] The picture source 16 may include or may be any type of picture capturing device configured to capture a real-world picture, and / or any type of picture generation device, for example, a computer graphics processing unit configured to generate a computer-animated picture, or any type of device configured to obtain and / or provide a real-world picture, a computer-generated picture (for example, screen content, a virtual reality (virtual reality, VR) picture,and / or any combination thereof (for exam- ple, an augmented reality (augmented reality, AR) pic- ture). The picture source may be any type of memory or storage that stores any of the foregoing pictures.
[0061] To distinguish processing performed by the pre- processor (or preprocessing unit) 18, a picture (or picture data) 17 may also be referred to as an original picture (or original picture data) 17.
[0062] The preprocessor 18 is configured to receive the original picture data 17, and preprocess the original picture data 17, to obtain a preprocessed picture (or preprocessed picture data) 19. For example, the prepro- 5 10 15 20 25 30 35 40 45 50 55 9 15 EP 4 694 127 A1 16 cessing performed by the preprocessor 18 may include cropping, color format conversion (for example, from RGB to YCbCr), color correction, or denoising. It may be understood that the preprocessing unit 18 may be an optional component.
[0063] The video encoder (or encoder) 20 is configured to receive the preprocessedpicture data 19 and provide the encoded picture data 21.
[0064] The communication interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and send the encoded picture data 21 (or any other processed version) over a communication channel 13 to another device such as the destination device 14 or any other device, for storage or direct reconstruction.
[0065] The source device 12 may further include a memory (not shown in FIG. 1). The memory may be configured to store at least one of the following data: the original picture data 17, the preprocessed picture (or preprocessed picture data) 19, and the encoded picture data 21.
[0066] The destination device 14 includes a decoder 30, and optionally, may include a communication inter- face (or communication unit) 28, a post-processor (or post-processing unit) 32, and a display device 34.
[0067] The communication interface 28 of the destina- tion device 14 is configured to directly receive the en- codedpicture data 21 (or any other processed version) from the source device 12 or any other source device such as a storage device, for example, an encoded picture data storage device, and provide the encoded picture data 21 to the decoder 30.
[0068] The communication interface 22 and the com- munication interface 28 may be configured to send or receive the encoded picture data (or encoded data) 21 via a direct communication link between the source device 12 and the destination device 14, for example, a direct wired or wireless connection, or via any type of network, for example, a wired network, a wireless network, or any combination thereof, or any type of private network, public network, or any combination thereof.
[0069] For example, the communication interface 22 may be configured to package the encoded picture data 21 into an appropriate format such as a packet, and / or process the encoded picture data by using any type of transmission encoding or processing, for transmission over acommunication link or a communication network.
[0070] The communication interface 28 corresponds to the communication interface 22, and may be, for ex- ample, configured to receive transmitted data and pro- cess the transmitted data by using any type of corre- sponding transmission decoding or processing and / or de-packaging, to obtain the encoded picture data 21.
[0071] The communication interface 22 and commu- nication interface 28 each may be configured as a uni- directional communication interface indicated by an ar- row of the communication channel 13 pointing from the source device 12 to the destination device 14 in FIG. 1, or a bidirectional communication interface; and may be configured to send and receive a message and the like, to establish a connection, confirm and exchange any other information related to the communication link an- d / or data transmission such as transmission of the en- coded picture data.
[0072] The video decoder (or decoder) 30 is configured to receive theencoded picture data 21 and provide decoded picture data (or decoded picture data) 31.
[0073] The post-processor 32 is configured to post- process the decoded picture data 31 (also referred to as reconstructed picture data) such as a decoded picture, to obtain post-processed picture data 33 such as a post- processed picture. For example, the post-processing performed by the post-processing unit 32 may include color format conversion (for example, from YCbCr to RGB), color correction, cropping, or re-sampling, or any other processing for generating the decoded picture data 31 for display by display device 34 or the like.
[0074] The display device 34 is configured to receive the post-processed picture data 33, to display the picture to a user, a viewer, or the like. The display device 34 may be or may include any type of display for representing the reconstructed picture, for example, an integrated or ex- ternal display or monitor. For example, the display may include a liquid crystaldisplay (liquid crystal display, LCD), an organic light-emitting diode (organic light-emit- ting diode, OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (liquid crystal on silicon, LCoS), a digital light processor (digital light processor, DLP), or any type of other display.
[0075] The destination device 14 may further include a memory (not shown in FIG. 1). The memory may be configured to store at least one of the following data: the encoded picture data 21, the decoded picture data 31, and the post-processed picture data 33.
[0076] The coding system 10 further includes a training engine 25. The training engine 25 is configured to train the encoder 20 to process an input picture, picture region, or picture block, to obtain a feature map of the input picture, picture region, or picture block, obtain an esti- mated probability distribution of the feature map, and encode the feature map based on the estimated prob- ability distribution.
[0077] The training engine 25 is further configured to train the decoder 30, to obtain an estimated probability distribution of a bitstream, decode the bitstream based on the estimated probability distribution to obtain a feature map, and reconstruct the feature map to obtain a recon- structed picture.
[0078] As shown in FIG. 1, the source device 12 and the destination device 14 are separate devices. However, device embodiments may include both the source device 12 and the destination device 14, or include functions of both the source device 12 and the destination device 14, that is, include both the source device 12 or a correspond- ing function thereof and the destination device 14 or a corresponding function thereof. In these embodiments, 5 10 15 20 25 30 35 40 45 50 55 10 17 EP 4 694 127 A1 18 the source device 12 or the corresponding function there- of and the destination device 14 or the corresponding function thereof may be implemented by same hardware and / or software or byseparate hardware and / or software or any combination thereof.
[0079] Based on the description, it is clear for a skilled person that existence and (accurate) division of different units or functions of the source device 12 and / or the destination device 14 shown in FIG. 1 may vary depend- ing on an actual device and application.
[0080] In recent years, applying deep learning (deep learning) to the video encoding and decoding field gra- dually becomes a trend. The deep learning is multi-layer learning at different abstraction layers by using a ma- chine learning algorithm. Deep learning-based video encoding and decoding may also be referred to as AI video encoding and decoding or neural network-based video encoding and decoding. Embodiments of this ap- plication relate to application of a neural network. For ease of understanding, the following first explains some nouns or terms used in embodiments of this application. The nouns or terms are also used as a part of content of the presentinvention. (1) Artificial neural network (artificial neural network, ANN):
[0081] The artificial neural network is also referred to as a neural network (NN), and is a dynamic system that is established manually and uses a directed graph as a topology structure. The artificial neural network pro- cesses information by using a continuous or discontin- uous input as a status response, and is an information processing system that simulates a human brain struc- ture and its functions. After decades of development, the artificial neural network has been widely used in many fields, such as pattern recognition, automatic control, signal processing, decision-making assistance, artificial intelligence, and scientific computing, and has achieved extensive success. Generally, one network includes an input layer, a hidden layer, and an output layer. The neural network in this application may include a plurality of types, for example, a deep neural network (deep neural network, DNN), a convolutionalneural network (convolu- tional neural network, CNN), a recurrent neural network (recurrent neural network, RNN), a residual network, a neural network using a transformer model, or another neural network. Some neural networks are described by way of example below. (2) Convolutional neural network:
[0082] The convolutional neural network (convolu- tional neuron network, CNN) is a deep neural network with a convolutional structure, and is a deep learning (deep learning) architecture. The deep learning architec- ture means that multi-layer learning is performed at dif- ferent abstraction lays by using a machine learning algo- rithm. As a deep learning architecture, the CNN is a feed- forward (feed-forward) artificial neural network, and each neuron in the feed-forward artificial neural network pro- cesses data input into the neuron.
[0083] As shown in FIG. 2, a convolutional neural net- work (CNN) 100 may include an input layer 110, a con- volutional layer / pooling layer 120, where thepooling layer is optional, and a neural network layer 130. As shown in FIG. 2, the convolutional layer / pooling layer 120 may include, for example, layers 121 to 126. In an implementation, the layer 121 is a convolutional layer, the layer 122 is a pooling layer, the layer 123 is a convolu- tional layer, the layer 124 is a pooling layer, the layer 125 is a convolutional layer, and the layer 126 is a pooling layer. In another implementation, the layers 121 and 122 are convolutional layers, the layer 123 is a pooling layer, the layers 124 and 125 are convolutional layers, and the layer 126 is a pooling layer. That is, an output of a convolutional layer may be used as an input of a sub- sequent pooling layer, or may be used as an input of another convolutional layer to continue to perform a convolutional operation. The convolutional layer 121 is used as an example. The convolutional layer 121 may include a plurality of convolutional operators, and the convolutional operators are alsoreferred to as convolu- tional kernels. The convolutional operator may be essen- tially a weight matrix, and the weight matrix is usually predefined. Picture processing is used as an example. Different weight matrices are used to extract different features in a picture. For example, one weight matrix is used to extract edge information of the picture, another weight matrix is used to extract a specific color of the picture, and still another weight matrix is used to blur unnecessary noise in the picture.
[0084] Weight values in these weight matrices need to be obtained through a large amount of training in actual application. Each weight matrix formed by the weight values obtained through training may be used to extract information from input data, to help the convolutional neural network 100 perform correct prediction.
[0085] When the convolutional neural network 100 has a plurality of convolutional layers, a large quantity of general features are usually extracted at an initial con-volutional layer (for example, 121). The general feature may also be referred to as a low-level feature. As a depth of the convolutional neural network 100 increases, a feature extracted at a later convolutional layer (for ex- ample, 126) is more complex, for example, a higher-level semantic feature. A higher semantic feature is more applicable to a to-be-resolved problem. Pooling layer:
[0086] Because a quantity of training parameters often needs to be reduced, a pooling layer often needs to be periodically introduced after a convolutional layer. For the layers 121 to 126 shown in 120 in FIG. 2, one convolu- tional layer may be followed by one pooling layer, or a 5 10 15 20 25 30 35 40 45 50 55 11 19 EP 4 694 127 A1 20 plurality of convolutional layers may be followed by one or more pooling layers. During picture processing, the pool- ing layer is only used to reduce a space size of a picture. The pooling layer may include an average pooling op- erator and / or a maximum poolingoperator, to perform sampling on the input picture to obtain a picture with a small size. The average pooling operator may calculate a pixel value in the picture in a specific range, to generate an average value. The maximum pooling operator may use a maximum pixel in a specific range as a maximum pooling result. In addition, similar to the size of the weight matrix that needs to be related to the size of the picture at the convolutional layer, an operator also needs to be related to a size of a picture at the pooling layer. A size of a processed picture output from the pooling layer may be less than a size of a picture input to the pooling layer. Each pixel in the picture output from the pooling layer represents an average value or a maximum value of a corresponding sub-region of the picture input into the pooling layer.
[0087] After processing is performed at the convolu- tional layer / pooling layer 120, the convolutional neural network 100 still cannot output required output informa-tion. This is because the convolutional layer / pooling layer 120 only extracts a feature and reduces a parameter brought by the input picture, as described above. How- ever, to generate final output information (required type information or other related information), the convolu- tional neural network 100 needs to use the neural net- work layer 130 to generate an output of one required type or a group of required types. Therefore, the neural net- work layer 130 may include a plurality of hidden layers (131, 132, ..., and 13n shown in FIG. 2) and an output layer 140. Parameters included in the plurality of hidden layers may be obtained through pre-training based on related training data of a specific task type. For example, the task type may include picture recognition, picture classification, and super-resolution picture reconstruc- tion.
[0088] The plurality of hidden layers in the neural net- work layer 130 are followed by the output layer 140, namely, the last layer of the entireconvolutional neural network 100. The output layer 140 has a loss function similar to classification cross entropy, and the loss func- tion is specifically used to calculate a prediction error. Once forward propagation of the entire convolutional neural network 100 (for example, propagation from layers 110 to 140 in FIG. 2 is forward propagation) is completed, back propagation (for example, propagation from layers 140 to 110 in FIG. 2 is back propagation) is started to update weight values and deviations of the layers mentioned above, to reduce a loss of the convolu- tional neural network 100 and an error between an ideal result and a result output by the convolutional neural network 100 through the output layer.
[0089] It should be noted that the convolutional neural network 100 shown in FIG. 2 is merely used as an example of a convolutional neural network. In specific application, the convolutional neural network may alter- natively exist in a form of another network model, forexample, a plurality of parallel convolutional layers / pool- ing layers, and extracted features are all input into the neural network layer 130 for processing. (3) Loss function:
[0090] In a process of training a neural network, be- cause it is expected that an output of the neural network is as close as possible to a value that actually needs to be predicted, a current predicted value of the network and an actually expected target value may be compared, and then a weight vector of each layer of the neural network is updated based on a difference between the current pre- dicted value and the target value (certainly, before a first update, there is usually an initialization process, that is, preconfiguring a parameter for each layer of the neural network). For example, if the predicted value of the net- work is large, the weight vector is adjusted to decrease the predicted value, and adjustment is continuously per- formed, until the neural network can predict the actually expected targetvalue. Therefore, "how to obtain, through comparison, a difference between the predicted value and the target value" needs to be predefined. This is a loss function (loss function) or an objective function (ob- jective function). The loss function and the objective function are important equations that measure the differ- ence between the predicted value and the target value. The loss function is used as an example. A higher output value (loss) of the loss function indicates a larger differ- ence. Therefore, training of the neural network is a pro- cess of minimizing the loss as much as possible. (4) Linear operation:
[0091] Linearity refers to a proportional and straight- line relationship between quantities, and may be math- ematically understood as a function whose first-order derivative is a constant. The linear operation may be but is not limited to an addition operation, a null operation, an identity operation, a convolutional operation, a layer normalization (layernormalization, LN) operation, and a pooling operation. The linear operation may also be referred to as linear mapping. The linear mapping needs to meet two conditions: homogeneity and additivity. If either condition is not met, non-linear mapping occurs.
[0092] Homogeneity means that f(ax)=af(x), and ad- ditivity means that f(x+y)=f(x)+f(y). For example, f(x)=ax is linear. It should be noted that x, a, and f(x) herein are not necessarily scalars, and may be vectors or matrices, forming linear space of any dimension. If x and f(x) are n-dimensional vectors, when a is a constant, it is equiva- lent that homogeneity is met; or when a is a matrix, it is equivalent that additivity is met. Relatively, a function graph that is a straight line does not necessarily comply with linear mapping. For example, f(x)=ax+b does not meet homogeneity or additivity, and therefore belongs to 5 10 15 20 25 30 35 40 45 50 55 12 21 EP 4 694 127 A1 22 non-linear mapping.
[0093] In embodiments of thisapplication, a combina- tion of a plurality of linear operations may be referred to as a linear operation, and each linear operation included in the linear operation may also be referred to as a sub- linear operation. (5) Attention model:
[0094] The attention model is a neural network that uses an attention mechanism. In deep learning, the attention mechanism may be defined in a broad sense as a weight vector that describes importance: to predict or infer an element by using the weight vector. For example, for a pixel in a picture or a word in a sentence, a correla- tion between a target element and another element may be quantitatively estimated by using an attention vector, and a weighted sum of the attention vector is used as an approximate value of a target.
[0095] The attention mechanism in deep learning si- mulates an attention mechanism of a human brain. For example, when a man views a picture, although the hu- man eyes can see the whole picture, when the man observes thepicture in depth, the eyes focus only on a part of the picture, and at this time, the human brain focuses on this small pattern. In other words, when the man observes a picture carefully, attention of the human brain to the entire picture is not balanced, and is distin- guished by a specific weight. This is a core idea of the attention mechanism.
[0096] Simply, a human visual processing system usually selectively focuses on some parts of a picture and ignores other irrelevant information, thereby facilitat- ing perception of the human brain. Similarly, in the atten- tion mechanism of deep learning, some parts of an input may be more relevant than others in some issues invol- ving language, speech, or vision. Therefore, by using the attention mechanism in the attention model, the attention model can dynamically focus only on a part of input that helps effectively execute a task at hand. (6) Self-attention network:
[0097] The self-attention network is a neural network that uses aself-attention mechanism. The self-attention mechanism is an extension of the attention mechanism. The self-attention mechanism is actually an attention mechanism that associates different locations of a single sequence to calculate a representation of a same se- quence. The self-attention mechanism can play a key role in machine reading, abstract summarization, or pic- ture description generation. For example, the self-atten- tion network is applied to natural language processing. The self-attention network processes input data of any length, generates a new feature representation of the input data, and then converts the feature expression into a target word. A self-attention network layer in the self- attention network uses the attention mechanism to obtain a relationship between all other words, thereby generat- ing a new feature representation of each word. An ad- vantage of the self-attention network is that the attention mechanism can directly capture a relationship between allwords in a sentence without considering a word posi- tion.
[0098] FIG. 3 is a diagram of a deep learning-based video encoder and decoder network (or system) accord- ing to an embodiment of this application. FIG. 2 is de- scribed by using entropy encoding and decoding as an example. The network includes a feature extraction mod- ule, a feature quantization module, an entropy encoding module, an entropy decoding module, a feature dequan- tization module, and a feature decoding (or picture re- construction) module.
[0099] At an encoder side, an original picture (or a to- be-compressed picture) is input into the feature extrac- tion module, and the feature extraction module outputs an extracted three-dimensional feature map of the origi- nal picture by stacking a plurality of convolutional layers with reference to a nonlinear mapping activation function. The feature quantization module quantizes a feature value of a floating-point number in the three-dimensional feature map, to obtain aquantized feature map. Entropy encoding is performed on the quantized three-dimen- sional feature map to obtain a bitstream.
[0100] At a decoder side, the entropy decoding module parses a bitstream to obtain a quantized three-dimen- sional feature map. The feature dequantization module dequantizes a feature value of an integer in the quantized feature map, to obtain a dequantized feature map. After the dequantized feature map is reconstructed by the feature decoding module, a reconstructed picture is ob- tained.
[0101] Entropy encoding is encoding that no informa- tion is lost according to an entropy principle in an encod- ing process. Entropy encoding is used to apply an en- tropy encoding algorithm or scheme to a quantized coef- ficient and another syntax element, to obtain encoded data that can be output by an output end in a form of an encoded bitstream or the like, so that a decoder or the like can receive and use a parameter used for decoding. The encoded bitstream may betransmitted to the decoder, or stored in a memory for later transmission or retrieval by the decoder. The entropy encoding algorithm or scheme includes but is not limited to: a variable length coding (variable length coding, VLC) scheme, a context-adap- tive VLC scheme (context-adaptive VLC, CALVC), an arithmetic coding scheme, a binarization algorithm, con- text-adaptive binary arithmetic coding (context adaptive binary arithmetic coding, CABAC), syntax-based con- text-adaptive binary arithmetic coding (syntax-based context-adaptive binary arithmetic coding, SBAC), prob- ability interval partitioning entropy (probability interval partitioning entropy, PIPE) coding, or another entropy coding method or technology.
[0102] Alternatively, the network may not include a feature quantization module and a feature dequantiza- 5 10 15 20 25 30 35 40 45 50 55 13 23 EP 4 694 127 A1 24 tion module. In this case, the network may directly per- form a series of processing on a feature map whosefeature map is a floating-point number. Alternatively, integerization processing may be performed on the net- work, so that all feature values in a feature map output by the feature extraction module are integers.
[0103] After the to-be-processed picture (or the to-be- compressed picture) passes through the feature extrac- tion module and the feature quantization module, the quantized three-dimensional feature map is obtained. When processing each feature value in the quantized three-dimensional feature map, the entropy encoding module may estimate a probability distribution of the feature value by using a processed feature value in a neighborhood as a context, to obtain a probability dis- tribution of the feature value, and perform subsequent encoding based on the probability distribution, to obtain an encoded bitstream.
[0104] FIG. 4 is a diagram of a structure of a deep learning-based end-to-end video encoder and decoder network according to an embodiment of this application. FIG. 4is described by using entropy encoding and de- coding as an example. The neural network includes a feature extraction module, a quantization module, a side information extraction module, an entropy encoding module, an entropy decoding module, a probability esti- mation module, and a reconstruction module. Entropy encoding may be an autoencoder (Auto Encoder, AE), and entropy decoding may be an autodecoder (Auto Decoder, AD).
[0105] At an encoder side, an original picture x is input into the feature extraction module, and the feature ex- traction module outputs a feature map y of the original picture. The feature map y is input into the quantization module, the quantization module outputs a quantized feature map, and the quantized feature map is input into the entropy encoding module. In addition, the feature map y is input into the side information extraction module, and the side information extraction module outputs side information z. The side information z is input into thequantization module, and the quantization module out- puts quantized side information. The quantized side in- formation passes through the entropy encoding module to obtain a bitstream of the side information, and then passes through the entropy decoding module to obtain decoded side information. The decoded side information is input into the probability estimation module. The prob- ability estimation module outputs a probability distribu- tion of each feature element [x] [y] [i] in the quantized feature map, and inputs the probability distribution of each feature element into the entropy encoding module. The entropy encoding module performs entropy encod- ing on each input feature element based on the prob- ability distribution of each feature element, to obtain a hyperprior bitstream.
[0106] The side information z is feature information, which is represented as a three-dimensional feature map. A quantity of feature elements included in the three-dimensional feature map is less than aquantity of feature elements included in the feature map y.
[0107] At a decoder side, the entropy decoding module parses a bitstream of side information to obtain the side information, and inputs the side information into the probability estimation module. The probability estimation module outputs a probability distribution of each feature element [x][y][i] in a to-be-decoded symbol. The prob- ability distribution of each feature element [x] [y] [i] is input into the entropy decoding module. The entropy decoding module performs entropy decoding on each feature ele- ment based on the probability distribution of each feature element, to obtain a decoded feature map. The decoded feature map is input into the reconstruction module, and the reconstruction module outputs a reconstructed pic- ture.
[0108] In addition, in probability estimation modules of some variational autoencoders (Variational Auto Enco- der, VAE), an encoded or decoded feature element around a current feature element isfurther used to esti- mate a probability distribution of the current feature ele- ment more accurately.
[0109] It should be noted that the network structures shown in FIG. 3 and FIG. 4 are merely examples for description. Modules included in the network and struc- tures of the modules are not limited in embodiments of this application.
[0110] In some possible scenarios, to further improve accuracy of a mean, an autoregressive module may be added. The autoregressive module may further obtain, based on a mean output by the probability distribution module and the quantized feature map, a probability distribution used to obtain a residual.
[0111] The following describes in detail an encoding and decoding method provided in embodiments of this application. FIG. 5 is a schematic flowchart of an encod- ing and decoding method according to an embodiment of this application. The method process may be performed by two electronic devices or by one electronic device. For example, when the methodprocess is performed by two electronic devices, one electronic device includes an encoder, configured to indicate an encoding operation, and the other electronic device includes a decoder, con- figured to perform a decoding operation. When the meth- od process is performed by one electronic device, the electronic device may include an encoder and a decoder. The method may be performed by the electronic device by invoking a neural network model. The method process is described as a series of operations. It should be under- stood that the method process may be performed in various sequences and / or simultaneously, and is not limited to an execution sequence shown in FIG. 5.
[0112] 501: The encoder encodes identification infor- mation indicating a used decoder network into a bit- stream.
[0113] The identification information is a first value, indicating that the decoder network used to decode the bitstream to obtain a to-be-processed picture is a first 5 10 15 20 25 30 35 40 45 50 55 14 25EP 4 694 127 A1 26 decoder network; or the identification information is a second value, indicating that the decoder network used to decode the bitstream to obtain a to-be-processed picture is a second decoder network. It may also be understood that the identification information is the first value, and the encoder performs an encoding operation on the to-be-processed picture by using a first encoder network corresponding to the first decoder network; or the identification information is the second value, and the encoder performs an encoding operation on the to-be- processed picture by using a second encoder network corresponding to the second decoder network. The iden- tification information may also be referred to as profile information (Profile ID), or may also be referred to as network information, a network identifier, or another name. This is not limited in this embodiment of this application. In other words, the identification information indicates processing that needs to besupported by the decoder, for example, a general_profile_idc syntax ele- ment in the H.265 standard. In an example, the first value may be 0, and the second value may be 1; or the first value may be 1, and the second value may be 0. The first value and the second value may alternatively be other values. A processing resource (or computing power) required by the first decoder network is higher than a processing resource (or computing power) required by the second decoder network. The processing resource (computing power) may include a memory resource, a processor resource, or the like. In some embodiments, it may also be understood that decoding rates (or decom- pression efficiency) of the first decoder network and the second decoder network are different. For example, the decoding rate of the first decoder network is higher than the decoding rate of the second decoder network; or quality of a picture restored by the first decoder network is different from quality of a picture restoredby the second decoder network. For example, the quality of the picture restored by the first decoder network is higher than the quality of the picture restored by the second decoder network.
[0114] 502: The encoder sends the bitstream.
[0115] 503: The decoder decodes the received bit- stream to obtain the identification information indicating the used decoder network.
[0116] 504: When the identification information is the first value, decode the bitstream to obtain the to-be- processed picture by using the first decoder network; or when the identification information is the second value, decode the bitstream to obtain the to-be-processed pic- ture by using the second decoder network.
[0117] In a possible implementation, the first decoder network and the second decoder network are completely different decoder networks, or the first decoder network and the second decoder network share a part of subnet, or the second decoder network is a subnet of the first decoder network.
[0118] In someembodiments, the identification infor- mation may further include other values indicating differ- ent decoder networks. It may be understood that a plur- ality of different decoder networks are indicated by a plurality of different values. For example, the identifica- tion information is a third value, indicating that the used decoder network is a third decoder network. A decoding rate of the third decoder network is different from the decoding rate of the first decoder network (or the second decoder network). In some embodiments, the decoding rate of the third decoder network is higher than the decoding rate of the second decoder network, and the decoding rate of the second decoder network is higher than the decoding rate of the first decoder network. In some other embodiments, the decoding rate of the third decoder network is between the decoding rate of the first decoder network and the decoding rate of the second decoder network. Herein, only three decoder networks are used as anexample. A quantity of decoder networks is not specifically limited in this embodiment of this ap- plication.
[0119] It may be understood that a higher picture de- coding rate indicates a shorter picture decoding delay.
[0120] For another example, quality of a picture re- stored by the third decoder network is different from the quality of the picture restored by the first decoder network (or the second decoder network).
[0121] The quality of the picture restored by the third decoder network is higher than the quality of the picture restored by the second decoder network, and the quality of the picture restored by the second decoder network is higher than the quality of the picture restored by the first decoder network. In some other embodiments, the qual- ity of the picture restored by the third decoder network is between the quality of the picture restored by the first decoder network and the quality of the picture restored by the second decoder network.
[0122] In some scenarios,when the third decoder net- work is further included, the third decoder network is different from the first decoder network (and the second decoder network). For example, the third decoder net- work, the second decoder network, and the first decoder network are three different decoder networks; or the third decoder network and the second decoder network (or the first decoder network) share a part of subnet; or the third decoder network is a subnet of the second decoder net- work (or the first decoder network).
[0123] That the first decoder network and the second decoder network share a part of subnet may be under- stood as that the first decoder network reuses a part of subnet of the second decoder network. For example, the first decoder network includes a network A, a network B, and a network C, the second decoder network includes a network D, the network B, and the network C, and the two decoder networks share the network B. Therefore, when the first decoder network is used, afterdata is input into the network A, an output result of the network A is input into the network B, and an output result of the network B is input into the network C. When the second decoder net- work is used, it may be understood that data is input into 5 10 15 20 25 30 35 40 45 50 55 15 27 EP 4 694 127 A1 28 the network D, an output result of the network D is input into the network B, and an output result of the network B is input into the network C.
[0124] For another example, the first decoder network is a subnet of the second decoder network. For example, the first decoder network includes a network A1, a net- work A2, and a network A3. The second decoder network includes the network A1 and the network A3. When the first decoder network is used, data is input into the net- work A1, an output result of the network A1 is input into the network A2, and an output of the network A2 is input into the network A3. When the second decoder network is used, it may be understood that when datais input into the network A1, an output result of the network A1 is not input into the network A2, but skips the network A2 and is input into the network A3.
[0125] In another possible implementation, when per- forming encoding, the encoder may use different encoder networks based on different values of the identification information. Alternatively, after encoding the to-be-pro- cessed picture into the bitstream by using an encoder network, the encoder may encode, into the bitstream, identification information of a decoder network corre- sponding to the used encoder network. It may be under- stood that the identification information indicates both the used decoder network and the used encoder network. When the identification information is the first value, residual information obtained by encoding the to-be-pro- cessed picture by using the first encoder network is encoded into the bitstream; or when the identification information is the second value, residual information obtained byencoding the to-be-processed picture by using the second encoder network is encoded into the bitstream, where a processing resource (or computing power) required by the first encoder network is higher than a processing resource (or computing power) re- quired by the second encoder network. It should be noted that the first encoder network and the first decoder net- work may be a pair of networks, and after the first encoder network is used for encoding, the first decoder network is used for decoding; and the second encoder network and the second decoder network are a pair of networks, and after the second encoder network is used for encoding, the second decoder network is used for decoding.
[0126] In some embodiments, the first encoder net- work and the second encoder network are two different encoder networks, or the first encoder network and the second encoder network share a part of subnet, or the first encoder network is a subnet of the second encoder network.
[0127] Theidentification information may further have other values, and different values indicate different used encoder networks. For example, the identification infor- mation is a third value, the used encoder network is a third encoder network, and the used decoder network is a third decoder network. It should be noted that the third encoder network and the third decoder network are a pair of networks, and after the third encoder network is used for encoding, the third decoder network is used for de- coding. An encoding rate of the third encoder network is different from an encoding rate of the first encoder net- work (or the second encoder network). In some embodi- ments, the encoding rate of the third encoder network is higher than the encoding rate of the second encoder network, and the encoding rate of the second encoder network is higher than the encoding rate of the first encoder network. In some other embodiments, the en- coding rate of the third encoder network is between the encodingrate of the first encoder network and the en- coding rate of the second encoder network.
[0128] It may be understood that a higher picture en- coding rate indicates a shorter picture encoding delay.
[0129] For another example, quality of a picture re- stored by the third encoder network is different from quality of a picture restored by the first encoder network (or the second encoder network).
[0130] The quality of the picture restored by the third encoder network is higher than the quality of the picture restored by the second encoder network, and the quality of the picture restored by the second encoder network is higher than the quality of the picture restored by the first encoder network. In some other embodiments, the qual- ity of the picture restored by the third encoder network is between the quality of the picture restored by the first encoder network and the quality of the picture restored by the second encoder network.
[0131] In some scenarios, when the third encoder net-work is further included, the third encoder network is different from the first encoder network (and the second encoder network). For example, the third encoder net- work, the second encoder network, and the first encoder network are three different encoder networks; or the third encoder network and the second encoder network (or the first encoder network) share a part of subnet; or the third encoder network is a subnet of the second encoder net- work (or the first encoder network).
[0132] For example, the first encoder network includes a feature extraction module, a quantization module, a side information extraction module, an entropy encoding module, and a probability estimation module. The second encoder network also includes a feature extraction mod- ule, a quantization module, a side information extraction module, an entropy encoding module, and a probability estimation module. In one manner, a used network struc- ture of at least one module in the first encoder network isdifferent from that of at least one module in the second encoder network. For example, a network structure of the probability estimation module in the first encoder network is different from that of the probability estimation module in the second encoder network. For another example, a network structure of the feature extraction module in the first encoder network is different from that of the feature extraction module in the second encoder network.
[0133] For example, the feature extraction module in the first encoder network is referred to as a first feature extraction module, and the feature extraction module in 5 10 15 20 25 30 35 40 45 50 55 16 29 EP 4 694 127 A1 30 the second encoder network is referred to as a second feature extraction module. It should be noted that, for each module that belongs to a neural network, the "mod- ule" may also be referred to as a "network". For example, the feature extraction module may be referred to as a feature extraction network. For anotherexample, the quantization module may also be referred to as a quanti- zation network. That the network structure of the feature extraction module in the first encoder network is different from that of the feature extraction module in the second encoder network may be that the second feature extrac- tion network is a subnet of the first feature extraction network, or the second feature extraction network and the first feature extraction network are two completely different subnets, or the second feature extraction net- work and the first feature extraction network share one or more subnets. That the network structure of the prob- ability estimation module in the first encoder network is different from that of the probability estimation module in the second encoder network may be that the second feature extraction network is a subnet of the first feature extraction network, or the second feature extraction net- work and the first feature extraction network are two completely differentsubnets, or the second feature ex- traction network and the first feature extraction network share one or more subnets.
[0134] In an example, refer to FIG. 6A. A network structure of the first encoder network is as follows: The first encoder network includes a first feature extraction network 610, a quantization network 620, an autoregres- sive network 630, a side information extraction network 640, and a probability estimation network 650.
[0135] Further, the residual information obtained by encoding the to-be-processed picture by using the first encoder network may be implemented in the following implementation. FIG. 7A is a diagram of a possible pro- cess of encoding the residual information.
[0136] 701a: Extract a three-dimensional feature map of the to-be-processed picture by using the first feature extraction network 610.
[0137] 702a: Quantize the three-dimensional picture feature by using the quantization network 620 to obtain a quantized three-dimensional feature map.
[0138] 703a: Extract side information of the to-be-pro- cessed picture from an edge in the three-dimensional feature map by using the side information extraction network 640.
[0139] 704a: Estimate a first probability distribution mean of the to-be-processed picture by using the prob- ability estimation network 650 based on the side informa- tion.
[0140] 705a: Input the quantized three-dimensional feature map and the first probability distribution informa- tion into the autoregressive network 630 to obtain a second probability distribution mean.
[0141] 706a: Obtain the residual information based on the third-dimensional feature map of the to-be-processed picture and the second probability distribution mean.
[0142] The first encoder network may further include an entropy encoder network 660. The entropy encoder network 660 may encode the residual information and the side information into the bitstream.
[0143] In some scenarios, it may be understood that the side information is encoded into abitstream 1, the residual information is encoded into a bitstream 2, and then the bitstream 1 and the bitstream are combined into one bitstream. In another scenario, the side information and the residual information may be encoded into one bitstream.
[0144] In another example, refer to FIG. 6B. A network structure of the second encoder network is as follows: The second encoder network includes a second feature extraction network 611, the side information extraction network 640, and the probability estimation network 650.
[0145] Further, the residual information obtained by encoding the to-be-processed picture by using the sec- ond encoder network may be implemented in the follow- ing implementation. FIG. 7B is a diagram of a possible process of encoding the residual information.
[0146] 701b: Extract a three-dimensional feature map of the to-be-processed picture by using the second fea- ture extraction network 611.
[0147] 702b: Extract side information from the three- dimensionalfeature map by using the side information extraction network 640.
[0148] 703b: Estimate a probability distribution mean of the to-be-processed picture by using the probability estimation network 650 based on the side information.
[0149] 704b: Obtain the residual information based on the third-dimensional feature map of the to-be-processed picture and the probability distribution mean.
[0150] The second encoder network may further in- clude an entropy encoder network 660. The entropy encoder network 660 may encode the residual informa- tion and the side information into a bitstream. In some scenarios, it may be understood that the side information is encoded into a bitstream 1, the residual information is encoded into a bitstream 2, and then the bitstream 1 and the bitstream are combined into one bitstream. In another scenario, the side information and the residual informa- tion may be encoded into one bitstream.
[0151] The identification information in this embodi- ment of thisapplication may be located in a header (header) of the bitstream. In some scenarios, the identi- fication information may alternatively be added to a suffix of a bitstream file. For example, different identification information corresponds to different suffixes. For exam- ple, the header includes information such as a picture length and width, a picture format, and a profile ID. The information needs to be stored in an agreed sequence. A specific storage sequence is not specifically limited in this application.
[0152] For example, the header of the bitstream may include one or more of the following parameter informa- tion: The parameter information includes profile informa- tion (profile ID), a picture height (H) and width (W), a 5 10 15 20 25 30 35 40 45 50 55 17 31 EP 4 694 127 A1 32 position and a size of a tile (Tiles) in latent space, a control flag of each tool, scaling factors of primary and secondary components, a model index (model Idx): a learnable model index, and a bit ratecontrol parameter βv. The rate control parameter includes a rate control parameter βY of the primary component, a rate control parameter of the secondary component βUV, and the like.
[0153] For example, the parameters of the header of the bitstream may be encoded by using a fixed bit length.
[0154] The following describes the parameter informa- tion.
[0155] W represents a width of the input picture. For example, W may range from 1 pixel to 8192 pixels.
[0156] H represents a height of the input picture. For example, H may range from 1 pixel to 8192 pixels.
[0157] format represents a data format of the input picture, for example, YUV420, YUV444, or sRGB.
[0158] bit_depth represents a bit depth of the input picture, for example, 8 and 10.
[0159] β is a parameter representing a quality level of a variable rate. The primary component and the secondary component may have different β. Therefore, the primary component is represented as beta_luma (BY), and the secondary component is representedas beta_chroma (βUV). A value of βY is between 0 and 1, and may be represented in a form of a 16-bit fixed-point number. Y represents luminance (Luminance or Luma). UV repre- sents chrominance (chroma). (parameter indicating quality level for variable rate. Primary and secondary component might have different betas, so for primary (beta_luma) and secondary (beta_chroma) are signaled. The value βv lays between 0 and 1, and signaled as a 16 bit fixed point numbers).
[0160] Color transform information (color_transfor- m_info): By default, a coded representation of a signal is YUV Bt.709 (full range). However, custom color trans- form is also supported. In this case, 12 coefficients (con- version matrix and offset) may be used and encoded as a fixed-point number with 8-bit resolution. (by default coded representation of the signal is YUV Bt.709 (full range), but also custom color transform is supported. In that case, 12 coefficients may be sent (conversion matrix and offset), encoded as afixed point numbers with 8 bit resolution).
[0161] Tile information (tiles info): represents a deco- der tile size and overlap for luminance and a decoder tile size and overlap for chrominance. Inter-channel correla- tion information filter (Inter Channel Correlation Informa- tion filter, ICCI) tile size and overlap. Generally, the tile has a square shape. However, because the tile may be smaller at the right or bottom picture boundary, the tile at the right or boundary may be non-rectangular. (decoder tile size and overlap for luma, decoder tile size and over- lap for chroma. ICCI tiles size and overlap. Tiles have square shape except at right or bottom picture boundary, where they may be smaller and non-rectangular).
[0162] Skip mode enable flag (SkipMode_ena- ble_flag): indicates whether a skip mode (SkipMode) is used for picture encoding. (signaled per image, indicates if SkipMode is used).
[0163] RVS enable flag (RVS_enable_flag): indicates whether to use the residual and variancescale (Residual and Variance Scale, RVS) for encoding each picture.
[0164] LSBS enable flag (LSBS_enable_flag): indi- cates whether to use the decoder-side latent scale before synthesis (Latent Scale Before Synthesis, LSBS) for encoding each picture.
[0165] ICIC enable flag (ICIC_enable_flag): indicates whether to use the inter-channel correlation information filter (Inter Channel Correlation Information filter, ICCI) for encoding each picture.
[0166] numThreads: is a 16-bit unsigned integer, and specifies a number of samples processed in parallel. (16 bit unsigned integer. Specifies number of samples pro- cessed in parallel).
[0167] It should be noted that, in some scenarios, different values of the identification information (profile information) may indicate only different used decoder networks. In some other scenarios, different encoder networks are used based only on different values of the identification information. In some other scenarios, different values of the identificationinformation indicate different used encoder networks and decoder networks.
[0168] In a possible example, refer to FIG. 8. The first decoder network (or the second decoder network) may include an entropy decoder network 810, a probability estimation network 820, and a picture restoration net- work. The picture restoration network may also be re- ferred to as a reconstruction network, or may have an- other name. This is not specifically limited in this embodi- ment of this application. At a decoder side, the entropy decoder network decodes a bitstream to obtain side information and residual information of a to-be-pro- cessed picture. A network structure of at least one net- work in the first decoder network is different from that of at least one network in the second decoder network, for example, the picture restoration network, or the probabil- ity estimation network. For example, the network struc- ture of the picture restoration network in the first decoder network is different fromthat of the picture restoration network in the second decoder network. For ease of distinguishing, the picture restoration network in the first decoder network is referred to as a first picture restoration network 831, and the picture restoration network in the second decoder network is referred to as a second picture restoration network 832. That the first picture restoration network 831 is different from the second picture restoration network 832 may be, for example, that the second picture restoration network 832 is a subnet of the first picture restoration network 831, or the second picture restoration network 832 and the first picture restoration network 831 share a part of subnet, or the second picture restoration network 832 and the first picture restoration network 831 are two different net- works. The first decoder network further includes an autoregressive network 840. 5 10 15 20 25 30 35 40 45 50 55 18 33 EP 4 694 127 A1 34
[0169] A decoding process is described withreference to the structure examples of the first decoder network and the second decoder network.
[0170] FIG. 9A is a diagram of a possible decoding process using the first decoder network.
[0171] 901a: Decode the bitstream to obtain side in- formation of a three-dimensional feature map of the to- be-processed picture by using the entropy decoder net- work 810, where the three-dimensional feature map in- cludes a plurality of feature elements.
[0172] 902a: Estimate a first probability distribution mean of a to-be-decoded feature element by using the probability estimation network 820 based on the side information.
[0173] 903a: Determine a second probability distribu- tion mean of the to-be-decoded feature element by using the autoregressive network 840 based on the first prob- ability distribution mean and a decoded feature element.
[0174] 904a: Decode the bitstream to obtain residual information of the to-be-decoded feature element by using the entropy decoder network 810 based on thesecond probability distribution mean, and obtain the to- be-decoded feature element based on the residual in- formation and the second probability distribution mean.
[0175] 905a: Restore the to-be-processed picture by using the first picture restoration network 831 based on the three-dimensional feature map obtained through de- coding.
[0176] FIG. 9B is a diagram of a possible decoding process using the second decoder network.
[0177] 901b: Decode the bitstream to obtain side in- formation of a three-dimensional feature map of the to- be-processed picture by using the entropy decoder net- work 810, where the three-dimensional feature map in- cludes a plurality of feature elements.
[0178] 902b: Estimate a first probability distribution mean of a to-be-decoded feature element by using the probability estimation network 820 based on the side information.
[0179] 903b: Decode the bitstream to obtain residual information of the to-be-decoded feature element by using the entropy decodernetwork 810 based on the first probability distribution mean, and obtain the to-be- decoded feature element based on the residual informa- tion and the first probability distribution mean.
[0180] 904b: Restore the to-be-processed picture by using the second picture restoration network 832 based on the three-dimensional feature map obtained through decoding.
[0181] In some possible implementations, the prob- ability estimation network in the encoder network (includ- ing the first encoder network and the second encoder network) may be the same as the probability estimation network used in the decoder network.
[0182] The following describes the solutions in embo- diments of this application with reference to specific examples. The following examples are described by using an end-to-end picture encoding and decoding pro- cess as an example. Example 1:
[0183] FIG. 10A and FIG. 10B are diagrams of execu- tion processes of an encoder network and that of a decoder network according toembodiments of this ap- plication. The encoder and decoder network is dynami- cally adjusted based on a profile ID. The encoder network is described with reference to the foregoing network structures in FIG. 6A and FIG. 6B. The first feature extraction network (module) 610 includes encoder net- work submodules 1 to 3. The encoder network submo- dules 1 to 3 extract features from a to-be-processed picture, and gradually convert the picture from a pixel domain to a feature domain, so that the picture is more easily compressed. The second feature extraction mod- ule in the second encoder network includes an encoder network submodule 1 and an encoder network submo- dule 3.
[0184] Correspondingly, the decoder network is de- scribed with reference to the network structure shown in FIG. 8. Decoder network submodules 1→2→3 or 1→ 2→4 gradually restore a three-dimensional feature map to a picture. A difference between the decoder network submodule 3 and the decoder network submodule 4 lies in astructure. For example, a possible difference lies in that the decoder network submodule 3 is a light module that adapts to a terminal-side device with low computing power, and is characterized by faster running but poorer picture restoration quality than the decoder network sub- module 4, while the decoder network submodule 4 is a module that adapts to a device with high computing power, and is characterized by slower running but better picture restoration quality than the decoder network sub- module 3. In FIG. 10B, an example in which the first picture restoration network of the first decoder network includes the decoder network submodules 1, 2 and 4 is used, and an example in which the second picture re- storation network of the second decoder network in- cludes the decoder network submodules 1 to 3 is used.
[0185] FIG. 10A is a diagram of an encoding process according to Example 1. Specifically, an implementation process of an encoder side is as follows:
[0186] Step 1: Calculateand output a picture feature y by using the feature extraction module. During the cal- culation, whether to execute or skip some network sub- modules is selected based on the profile ID. When the profile ID is 0, the encoder network submodule 2 is executed, that is, encoding is performed by using the first encoder network. When the profile ID is 1, the en- coder network submodule 2 is skipped, that is, encoding is performed by using the second encoder network. In some scenarios, the encoder network submodule 2 may be skipped when the profile ID is 1, or the encoder net- work submodule 2 may be executed when the profile ID is 0. For example, the picture featureymay also be referred to as a feature map y, or may be referred to as a three- 5 10 15 20 25 30 35 40 45 50 55 19 35 EP 4 694 127 A1 36 dimensional feature map y. After feature extraction is performed on the to-be-encoded picture by using feature extraction module to obtain the picture feature y, the picture feature y may befurther quantized, which may be understood as processing (for example, rounding off) a feature value of a floating-point number to obtain an integer feature value, so as to obtain a quantized feature map ŷ.
[0187] Step 2: Input the picture feature y calculated in step 1 into the side information extraction network (mod- ule), to extract side information z;and quantize z to obtain ẑ, and compress ẑ into a bitstream 1. It should be noted that the side information extraction module is not man- datory. In some possible scenarios, after feature extrac- tion is performed on an original picture, quantization compression (or encoding) is directly performed to gen- erate a bitstream.
[0188] In some embodiments, the quantized feature map ŷ may be input into the side information extraction network, to output quantized side information ẑ. The side information extraction module may be implemented by a neural network. A specific neural network structure is described by using an example subsequently,and details are not described herein. The side information ẑ may be understood as a feature map ẑ obtained by performing further feature extraction on the quantized feature map ŷ, and a quantity of feature elements included in ẑ is less than a quantity of feature elements included in the feature map ŷ.
[0189] In some scenarios, the encoder network (the first encoder network and the second encoder network) may further include a quantization network, configured to perform a quantization operation on the picture feature y. In some other scenarios, the side information extraction network may have a function of performing a quantization operation, so that the side information extraction network performs a quantization operation on the picture feature y.
[0190] Step 3: Obtain a probability distribution of the picture feature y from the side information. When the profile ID is 1, the side information is input into the prob- ability estimation network. The probability estimation net- work(which may also be referred to as a probability estimation module) includes feature probability distribu- tion modules A and B that predict a mean (mean) and variance information of the picture feature y. The feature probability distribution modules A and B may also be referred to as feature map probability distribution estima- tion modules A and B, or may be referred to as other names. This is not limited in this embodiment of this application. When the profile ID is 0, the side information is input into the feature probability distribution modules A and B. The feature probability distribution module B out- puts the variance information of the picture feature y.The output of the feature probability distribution module A and the quantized picture feature y need to be sent to the autoregressive module to generate the mean (mean) of the picture feature y. In some scenarios, the feature probability distribution modules A and B may be com- bined into one module, that is, functions thereofare performed by one module.
[0191] For example, the probability estimation network may use a Gaussian single model (Gaussian single model, GSM), an asymmetric Gaussian model, a Gaus- sian mixture model (Gaussian mixture model, GMM), or a Laplace distribution (Laplace distribution) model. The probability estimation network may alternatively be a deep learning-based network, for example, a recurrent neural network (recurrent neural network, RNN) or a convolutional neural network (convolutional neural net- work, CNN). This is not limited herein.
[0192] Step 4: Calculate residual information of the picture feature y relative to the mean r=y-mean with reference to the probability distribution information (mean mean and variance variance) of the picture feature y obtained in step 3, and perform entropy encoding on quantized residual information r̂ to obtain a compressed bitstream 2. The residual information r may also be referred to as a residual feature map r. Therefore, the quantizedresidual information r̂ may also be referred to as a quantized residual feature map r̂, or may be briefly referred to as a quantized residual feature r̂.
[0193] Step 5: Combine the bitstream 1 and the bit- stream 2 into one bitstream, and write the profile ID into the bitstream, for example, into header information (header) of the bitstream.
[0194] It should be noted that the encoding steps in step 2, step 4, and step 5 of the encoder side may be combined. In step 2, the side information ẑ is not encoded and written into the bitstream. Instead, after the quan- tized residual information r̂ is obtained in step 4, the quantized residual information r̂ and the side information ẑ are encoded (for example, entropy encoded) and writ- ten into the bitstream.
[0195] FIG. 10B is a diagram of a decoding process according to Example 1. Specifically, an implementation process of a decoder side is as follows:
[0196] Step 1: Parse a bitstream to obtain profile ID information by using the entropydecoder network (mod- ule), for example, obtain the profile ID from a header of the bitstream. The profile ID is profile information in the bitstream, and indicates processing that needs to be supported by the decoder, for example, a general_pro- file_idc syntax element in the H.265 standard. The profile ID may be an integer (certainly, the profile ID may not be an integer, and this is not specifically limited in this application). The profile information indicates processing that needs to be supported by the decoder, or may be understood as that the profile information indicates dif- ferent networks that need to be used by the decoder.
[0197] Step 2: Decode the bitstream (for example, a bitstream 1) to obtain side information by using the en- tropy decoder network, for example, may decode the bitstream 1 to obtain quantized side information ẑ through asymmetric numeral system (asymmetric numeral sys- tem, ANS) / arithmetic decoding. 5 10 15 20 25 30 35 40 45 50 55 20 37 EP 4 694 127 A138
[0198] Step 3: Obtain a probability distribution of a feature map ŷ from the side information ẑ by using the probability estimation network (module). When profile ID=1, the side information ẑ is input into the probability estimation module (or referred to as a probability estima- tion network), and probability estimation is performed on each feature element ŷ[x][y][i] in the to-be-decoded fea- ture map ŷ, to obtain a probability distribution of the feature element ŷ[x][y][i]. It is assumed that the feature elementŷ[x][y][i] meets a Gaussian distribution of a mean µ[x][y][i] and a variance σ[x][y][i], where the mean µ[x][y] [i] may be used as a predicted value of the feature element ŷ[x][y][i]. When profile ID=0, the side information ẑ is input into the probability estimation module (or re- ferred to as a probability estimation network), and prob- ability estimation is performed on each feature element ŷ[x][y][i] in the to-be-decoded feature map ŷ, to obtain a probabilitydistribution of the feature element ŷ[x][y][i]. Then, a predicted value of the current to-be-decoded feature element is obtained based on the autoregressive network by using information of a decoded feature ele- ment and a mean output by the probability estimation network.
[0199] Parameters x, y, and i in the feature elementŷ[x] [y][i] are all positive integers, and coordinates (x, y, i) represent a position of the current to-be-decoded feature element. Specifically, the coordinates (x, y, i) represent a position of the current to-be-decoded feature element relative to a feature element of an upper left vertex in a current three-dimensional feature map. This step may be specifically implemented by the probability estimation module. The probability estimation method used at the decoder side may be correspondingly the same as the probability estimation method used at the encoder side, that is, the structure of the probability estimation module of the decoder side may be the same as thestructure of the probability estimation module of the encoder side, and details are not described herein.
[0200] The bitstream 2 may be understood as a bit- stream converted from a plurality of matrices y, and decoding is a process of restoring the plurality of matrices y from the bitstream. Restoring y is sequentially restoring an element and then an element. For example, in a 10x10 matrix, elements of the matrix are restored one by one in order from left to right and from top to bottom. When the element in the seventh row and the eighth column is restored, elements (that is, all elements whose horizontal coordinates are less than 7 and vertical coordinates are less than 8) before the element may be referred to as context of the element, that is, may be understood as decoded context information.
[0201] Step 4: Continue to parse the bitstream to obtain a quantized residual feature map r̂ by using the entropy decoder network by using the Gaussian distribution of the mean µ and thevariance σ of each feature element in the quantized feature map ŷ obtained in step 3, and further obtain the quantized feature map ŷ = r̂ + µ based on r̂ and µ.
[0202] In an example, a possible implementation of parsing the bitstream to obtain the feature map r̂ is as follows:
[0203] A probability with a value k P(k) of the to-be- decoded feature element r̂[x|[y][i] is obtained based on the probability distribution (for example, the Gaussian distribution of the mean value 0 and the variance σ), and the bitstream is parsed to obtain the feature element r̂[x] [y][i] through the ANS decoding / arithmetic decoding based on P(k). k may be any integer, for example, 0, 1, 2, or 3.
[0204] Step 5: Restore the picture from the quantized ŷ by using the picture restoration network. In a process of running the picture restoration network (or reconstruction network), after the decoder network submodule 1 and the decoder network submodule 2 are executed, the decoder network submodule 3 and the decodernetwork submo- dule 4are selected based on a value of the profile ID. If the profile ID is 1, the decoder network submodule 3 is selected, that is, the second decoder network is used; or if the profile ID is 0, the decoder network submodule 4 is selected, that is, the first decoder network is used.
[0205] The following describes, with reference to spe- cific examples, a structure of each sub-network of the foregoing encoder network (including the first encoder network and the second encoder network). FIG. 11A and FIG. 11B are a diagram of an execution process of an encoder network. It should be noted that FIG. 11A and FIG. 11B are merely an example, and does not constitute a limitation on a specific structure of the encoder network.
[0206] Refer to FIG. 11A and FIG. 11B. The encoder network submodule 1 includes a plurality of layers, which are respectively a padding (padding) layer 1_1, a con- volutional (Conv) layer 1_11, a residual activation func- tion (ResAU) layer 1_21, padding1_2, a convolutional layer 1_12, a residual activation function layer 1_22, and padding 1_3. In this embodiment of this application, a convolution with a convolution size of K×K, a quantity of output channels of M, and a stride (Stride) of N may be represented as Conv M×K×K SN. In FIG. 11A and FIG. 11B, an example in which the convolutional (Conv) layer 1_11 and the convolutional layer 1_12 use Conv12 28×3×3 S2 is used.
[0207] For example, the padding layer may use zeros padding (constant padding) (padding 0 by default), reflect padding (reflect padding), replicated padding (replicated padding), and circular padding (circular padding). For example, padding 1_1, padding 1_2, and padding 1_3 all use the replicate padding, to make a length and a width of an input tensor to an even number through padding by using the replicated padding (padding with a nearest element). For example, if the length and the width of the input tensor are 5 and 6 respectively, the padding layer pads an elementin a length direction to change the length to an even number 6. However, the width 6 of the input tensor is an even number. Therefore, a padding operation is not performed in a width direction.
[0208] The residual activation function layer 1_21 and 5 10 15 20 25 30 35 40 45 50 55 21 39 EP 4 694 127 A1 40 the residual activation function layer 1_22 are mainly used as activation functions, and may further provide an attention mechanism. In an example, the residual activation function layer 1_21 and the residual activation function layer 1_22 may use a ResAU 3×3 no tanh network. The ResAU 3×3 no tanh network may use a structure shown in FIG. 12. The residual activation func- tion layer 1_21 (and the residual activation function layer 1_22) includes an activation function layer 2_1 and a convolutional layer 2_1. In FIG. 12, ⊙ and ⊕ are an element-wise multiplication operation and an element- wise addition operation. For example, the convolutional layer 2_1 uses C×3×3, where a size of aconvolutional kernel is 3×3, and a number of output channels is C. In FIG. 12, W represents a width of an input picture block (or a number of rows of a vector matrix), H represents a height of an input picture block (or a number of columns of a vector matrix), and N represents a stride. For example, the activation function layer 2_1 may use a Leaky ReLU function. The Leaky ReLU function is used to assign a non-zero slope to all negative values.
[0209] The encoder network submodule 2 may use a residual non-local attention block (residual non-local at- tention block, RNAB), configured to provide an attention mechanism, for example, may be configured to provide global or local attention information in space. For exam- ple, FIG. 13 is a possible diagram of a structure of an RNAB. The RNAB uses a network structure including a plurality of residual block (residual block, RB) layers, a plurality of convolutional layers, a deconvolutional layer, and an activation function layer (for example, asigmoid function is used). Global or local attention information in space is extracted by using the RNAB.
[0210] In an example, the residual block layer may use a network structure shown in FIG. 14. The RB layer may include a convolutional layer 4_1, an activation function layer 4_11, and a convolutional layer 4_2. In some em- bodiments, the activation function layer 4_11 may use a Leaky ReLU function. A main line of the residual block layer in FIG. 14 inputs features into the 3×3 convolutional layer 4_1 to obtain a feature matrix, then outputs the feature matrix by using an activation function, and next perform an operation of adding a result obtained by using the 3×3 convolutional layer 4_2 to the input feature.
[0211] Refer to FIG. 11A and FIG. 11B. The encoder network submodule 3 includes a convolutional layer 5_1, a residual activation function layer 5_11, a padding layer 5_12, a convolutional layer 5_2, and a convolutional layer 5_3. In FIG. 11A and FIG. 11B, the convolutionallayer 5_1 and the convolutional layer 5_2 may use conv 128×3×3 S2. The convolutional layer 5_3 uses conv 128×1×1 S1. The residual activation function layer 5_1 may use ResAU 3×3 no Tanh, for example, a network structure shown in FIG. 12.
[0212] The quantization network may include a round layer 6_11, configured to: perform a quantization opera- tion, which may also be referred to as a rounding opera- tion; and return a rounded value of a floating-point num- ber. In some embodiments, refer to FIG. 11A and FIG. 11B. The quantization network may further include a Gunit layer 6_1 and an invGunit layer 6_21. The Gunit layer 6_1 and the invGunit layer 6_21 are configured to perform bit rate matching, so that the encoder network has a bit rate adjustment capability. In an example, the Gunit layer 6_1 and the invGunit layer 6_21 may use structures of Gain and Inverse Gain in the paper "G-VAE: A CONTINUOUSLY VARIABLE RATE DEEP IMAGE COMPRESSION FRAMEWORK" by Ze Cui, Jing Wang et al.
[0213] InFIG. 11A and FIG. 11B, the side information extraction module may include a hyper encoder net and a round layer. The hyper encoder net may also be referred to as a hyper encoder network. A function of the hyper encoder net is to extract side information z by using an input quantized picture feature y. In FIG. 11A and FIG. 11B, the autoregressive network may include a context model net and a prediction fusion net. The context model net is an autoregressive process. The context model net predicts an expected value of a to-be-encoded element ŷ[:,i,j] by using information about an encoded element ŷ[:,i’,j’] (i’ ≤ i,j’ < j - i - j’) with reference to the prediction fusion net. For example, in an implementation process, mask convolution (mask conv) may be used for imple- mentation. The prediction fusion net is configured to receive the information about the encoded element ex- tracted by the context model net and the side information extracted by the hyper decoder, to predict the expectedvalue (or a predicted value) of the to-be-encoded ele- ment. The expected value of the to-be-encoded element may be understood as a predicted probability distribution mean.
[0214] In some possible implementation scenarios, the feature probability estimation module A may use a super decoder network (Hyper Decoder Net). The feature prob- ability estimation module B may use a hyper scale deco- der network (Hyper Scale Decoder Net). Certainly, an- other network structure may alternatively be used, and a network that can implement probability estimation on the side information is applicable to this application. In an example, the feature probability estimation module A may use a network structure of a hyper decoder network shown in FIG. 15. The feature probability estimation module A includes a convolutional layer 7_1, a decon- volutional layer 7_11, a crop (Crop) layer 7_21, an activa- tion function layer 7_31, a convolutional layer 7_2, a deconvolutional layer 7_12, a crop (Crop) layer7_22, an activation function layer 7_32, a convolutional layer 7_3, and an activation function layer 7_33.
[0215] The crop layer 7_21 and the crop layer 7_22 are configured to perform a crop operation on an input tensor. The crop operation may be represented as Crop(Hout, Wout, d, sd), where Hout, Wout is a length and a width of a picture finally output by the decoder network (or may be understood as a size of a picture input by the encoder network, and the size information may be obtained from a header of the bitstream), and sd is stride (stride) informa- 5 10 15 20 25 30 35 40 45 50 55 22 41 EP 4 694 127 A1 42 tion of a deconvolutional operation. In an example, sd = 2. d represents a depth of the deconvolutional layer. The crop layer inputs a tensor with a size of [C, sdhd, sdwd], and outputs a tensor with a size of [C,hd-1, wd-1]. hd = ceil (hd-1 / sd); wd = ceil(wd-1 / sd), h0 = Hout, w0 = Wout. In an example, in FIG. 11A and FIG. 11B, an example in which the activation function layer7_31, the activation function layer 7_32, and the activation function layer 7_33 use a LeakyRelu function is used. In FIG. 11A and FIG. 11B, an example in which the convolutional layer 7_1 uses conv 128×1×1 S1, the deconvolutional layer 7_11 uses DConv 128×4×4 S2, the convolutional layer 7_2 uses conv 128×3×3 S1, the deconvolutional layer 7_12 uses DConv 128×4×4 S2, and the convolutional layer 7_3 uses conv 128×3×3 S1 is used.
[0216] In another example, the feature probability es- timation module B may use a network structure of a hyper scale decoder network shown in FIG. 16. The feature probability estimation module B includes a deconvolu- tional layer 7_13, a crop (Crop) layer 7_23, an activation function layer 7_34, a convolutional layer 7_4, an activa- tion function layer 7_34, a deconvolutional layer 7_14, a crop (Crop) layer 7_24, an activation function layer 7_35, a convolutional layer 7_5, and a Gunit layer 7_6. In FIG. 11A and FIG. 11B, an example in which the deconvolu-tional layer 7_13 uses DConv 128×4×4 S2, the activa- tion function layer 7_34 uses a LeakyRelu function, the convolutional layer 7_4 uses conv 128×3×3 S1, the deconvolutional layer 7_14 uses DConv 128×4x4 S2, the activation function layer 7_35 uses a LeakyRelu function, and the convolutional layer 7_5 uses conv 128×3×3 S1 is used.
[0217] In FIG. 11A and FIG. 11B, a lossless encoder (lossless encoder) is used in an entropy encoder net- work, and a function of the lossless encoder is to convert a to-be-encoded feature into a bitstream.
[0218] The following describes, with reference to spe- cific examples, a structure of each sub-network of the foregoing decoder network (including the first decoder network and the second decoder network). FIG. 17A and FIG. 17B are a diagram of an execution process of a decoder network. It should be noted that FIG. 17A and FIG. 17B are merely an example, and does not constitute a limitation on a specific structure of the decoder network. In FIG. 17A andFIG. 17B, a lossless decoder is used in an entropy decoder network, and a function of the loss- less decoder is to restore a to-be-decoded bitstream to a feature. The probability estimation network in the deco- der network may use a same structure as the encoder network. For details, refer to FIG. 11A and FIG. 11B. The decoder network submodule 1 includes an invGunit layer, a light residual block (LightResBlock), a deconvolutional layer 8_1, a crop (crop) layer 8_11, and a residual activa- tion function layer 8_21. The decoder network submo- dule 2 includes a deconvolutional layer 8_2, a crop (crop) layer 8_12, and a residual activation function layer 8_22. The residual activation function layer 8_21 and the re- sidual activation function layer 8_22 may use a ResAU structure. The deconvolutional layer 8_1 may use Dconv 96×4×4 S2, and the deconvolutional layer 8_2 may use Dconv 64×4×4 S2.
[0219] The decoder network submodule 3 may include a convolutional layer 8_31, a residualactivation function layer 8_23, a convolutional layer 8_32, PxlShuffleS4, and a crop layer 8_13. The decoder network submodule 4 includes a deconvolutional layer 8_3, an RNAB, a crop layer 8_14, a residual activation function layer 8_24, a deconvolutional layer 8_4, and a crop layer 8_15. For example, the RNAB may use the network structure shown in FIG. 13.
[0220] In an example, for a network structure of the light residual block (LightResBlock), refer to FIG. 18. PxlShuffleS4: represents a pixel shuffle operation of 4x upsampling. Example 2:
[0221] An encoder network used in Example 2 is the same as that used in Example 1, and an execution process is also similar. Details are not described herein. The decoder network in Example 2 is different from the decoder network in Example 1. In Example 2, the second decoder network in the decoder network is a subnet of the first decoder network, or the second picture restoration network in the second decoder network is a subnet of the structureof the first picture restoration network in the first decoder network. Refer to FIG. 19. The second picture restoration network includes a decoder network submo- dule 5 and a decoder network submodule 7, and the first picture restoration network includes decoder network submodules 5 to 7.
[0222] Different from Example 1, in Example 2, at the decoder side, when profile IDs are different, selection is not performed in the decoder network submodule 3 and the decoder network submodule 4, but whether to skip a decoder network submodule is selected. When profile ID=1, the decoder network submodule 6 is skipped. When profile ID=0, the decoder network submodule 6 is executed. Example 3:
[0223] In Example 3, an example in which the feature extraction networks of the two encoder networks are two different networks, and the picture restoration networks of the two decoder networks are two different networks is used. Refer to FIG. 20A. The encoder network includes a first feature extractionnetwork, a second feature extrac- tion network, a quantization network, an autoregressive network, a side information extraction network, a prob- ability estimation network, and an entropy encoder net- work. Correspondingly, refer to FIG. 20B. The decoder network includes an autoregressive network, a side in- formation extraction network, a probability estimation network, a first picture restoration network, a second picture restoration network, and an entropy decoder net- 5 10 15 20 25 30 35 40 45 50 55 23 43 EP 4 694 127 A1 44 work.
[0224] As shown in FIG. 20A, a difference between an implementation process of the encoder side and Exam- ple 1 lies in step 1. In a process of calculating the input picture feature y, different feature extraction networks are selected based on profile IDs. When profile ID=0, the first featureextraction network is selected. When profile ID=1, the second feature extraction network is selected.
[0225] Similarly, a difference between an implementa- tionprocess of the decoder side and that of Example 1 lies in step 5, that is, the picture is restored from ŷ by using the picture restoration network. When the decoder net- work is running, picture restoration networks of different structures are selected based on profile IDs. When profile ID=0, the first picture restoration network is selected. When profile ID=1, the second picture restoration net- work is selected.
[0226] The following describes, with reference to spe- cific examples, a structure of each sub-network of the foregoing encoder network (including the first encoder network and the second encoder network). FIG. 21A and FIG. 21B are a diagram of an execution process of an encoder network. It should be noted that FIG. 21A and FIG. 21B are merely an example, and does not constitute a limitation on a specific structure of the encoder network. In FIG. 21A and FIG. 21B, the first feature extraction network includes a padding (padding) layer 1_1, a con- volutional (Conv) layer 1_11,a residual activation func- tion (ResAU) layer 1_21, padding 1_2, a convolutional layer 1_12, a residual activation function layer 1_22, padding 1_3, an RNAB, a convolutional layer 5_1, a residual activation function layer 5_11, a padding layer 5_12, a convolutional layer 5_2, and a convolutional layer 5_3. The second feature extraction network includes a padding (padding) layer 1_1, a convolutional (Conv) layer 1_11, a residual activation function (ResAU) layer 1_21, padding 1_2, a convolutional layer 1_12, a residual activation function layer 1_22, padding 1_3, a convolu- tional layer 5_1, a residual activation function layer 5_11, a padding layer 5_12, a convolutional layer 5_2, and a convolutional layer 5_3. For descriptions of the foregoing layers, refer to the related descriptions in the embodi- ment corresponding toFIG. 11A and FIG. 11B. Details are not described herein. For structures of other networks in FIG. 21A and FIG. 21B, refer to the descriptions in Ex- ample 1. Detailsare not described herein.
[0227] FIG. 22A and FIG. 22B are a diagram of an execution process of a decoder network. It should be noted that FIG. 22A and FIG. 22B are merely an example, and does not constitute a limitation on a specific structure of the encoder network. The first picture restoration net- work includes an invGunit layer, a light residual block (LightResBlock), a deconvolutional layer 8_1, a crop (crop) layer 8_11, a residual activation function layer 8_21, a deconvolutional layer 8_2, a crop (crop) layer 8_12, a residual activation function layer 8_22, a decon- volutional layer 8_3, an RNAB, a crop layer 8_14, a residual activation function layer 8_24, a deconvolutional layer 8_4, and a crop layer 8_15. The second picture restoration network includes an invGunit layer, a light residual block (LightResBlock), a deconvolutional layer 8_1, a crop (crop) layer 8_11, a residual activation func- tion layer 8_21, a deconvolutional layer 8_2, a crop (crop) layer 8_12, a residualactivation function layer 8_22, a convolutional layer 8_31, a residual activation function layer 8_23, a convolutional layer 8_32, PxlShuffleS4, and a crop layer 8_13. For descriptions of the foregoing layers, refer to the related descriptions in the embodi- ment corresponding toFIG. 11A and FIG. 11B. Details are not described herein. For structures of other networks in FIG. 22A and FIG. 22B, refer to the descriptions in Ex- ample 1. Details are not described herein.
[0228] It should be noted that the structures of the foregoing networks are merely used as examples, and a specific network structure is not specifically limited. A network structure that can implement a corresponding function is applicable to this application.
[0229] In addition, it should be noted herein that the foregoing submodule-level adjustment and the foregoing entire network-level adjustment performed on the enco- der and decoder network based on the profile ID may be flexibly combined. For example, in a possiblecase, the encoder side executes a dynamic computation graph on a cloud side by using a framework such as PyTorch, and adjusts an encoder network submodule based on a pro- file ID; and the decoder side executes a static computa- tion graph on a device side, and switches the entire decoder network based on profile information.
[0230] In embodiments of this application, the encoder side transmits network structure information, so that the decoder side can adjust the decoder network by using bitstream content. The solution has the following advan- tages: 1. For bitstreams generated by using different AI encoder networks, a decoder side may select different decoder network structures by using bitstream content, to implement decoding. This brings high flexibility to a codec side. A user may adjust encoder and decoder network computing power of the user based on a scenar- io of the user, to flexibly balance a delay and compression performance. 2. A user can dynamically select and adjust somedecoder network modules based on a profile ID or switch between different decoder networks.
[0231] It may be clearly understood by a person skilled in the art that, for the purpose of convenient and brief description, for a specific working process of the com- munication system described above, refer to a corre- sponding process in the foregoing method embodiments. Details are not described herein again.
[0232] An embodiment of this application provides a computer-readable medium, configured to store a com- puter program. The computer program includes instruc- tions used to perform the method steps in the method embodiment corresponding to FIG. 5.
[0233] A person skilled in the art should understand that embodiments of this application may be provided as a method, a system, or a computer program product. 5 10 15 20 25 30 35 40 45 50 55 24 45 EP 4 694 127 A1 46 Therefore, this application may use a form of hardware only embodiments, software only embodiments, or em- bodiments with acombination of software and hardware. Moreover, this application may use a form of a computer program product that is implemented on one or more computer-usable storage media (including but not limited to a disk memory, an optical memory, and the like) that include computer-usable program code.
[0234] This application is described with reference to the flowcharts and / or block diagrams of the method, the device (system), and the computer program product according to the embodiments of this application. It should be understood that computer program instruc- tions may be used to implement each process and / or each block in the flowcharts and / or the block diagrams and a combination of a process and / or a block in the flowcharts and / or the block diagrams. These computer program instructions may be provided for a general- purpose computer, a dedicated computer, an embedded processor, or a processor of any other programmable data processing device to generate a machine, so that the instructionsexecuted by a computer or a processor of any other programmable data processing device gener- ate an apparatus for implementing a specific function in one or more processes in the flowcharts and / or in one or more blocks in the block diagrams.
[0235] It is clear that a person skilled in the art can make various modifications and variations to this appli- cation without departing from the scope of this applica- tion. This application is intended to cover these modifica- tions and variations of this application provided that they fall within the scope of protection defined by the following claims and their equivalent technologies. Claims 1. A picture encoding method, comprising: encoding identification information indicating a used decoder network into a bitstream, wherein the identification information is a first value, in- dicating that the decoder network used to de- code the bitstream to obtain a to-be-processed picture is a first decoder network; or the identification information isa second value, indicating that the decoder network used to de- code the bitstream to obtain a to-be-processed picture is a second decoder network, wherein a processing resource required by the first deco- der network is higher than a processing re- source required by the second decoder network; and sending the bitstream. 2. The method according to claim 1, wherein the first decoder network and the second decoder network are completely different decoder networks, or the first decoder network and the second decoder net- work share a part of subnet, or the second decoder network is a subnet of the first decoder network. 3. The method according to claim 1 or 2, wherein the method further comprises: obtaining the identification information; and when the identification information is the first value, encoding, into the bitstream, residual in- formation obtained by encoding the to-be-pro- cessed picture based on a first encoder network; or when the identification information is the secondvalue, encoding, into the bitstream, residual in- formation obtained by encoding the to-be-pro- cessed picture based on a second encoder net- work, wherein a processing resource required by the first en- coder network is higher than a processing re- source required by the second encoder network. 4. The method according to claim 3, wherein the first encoder network and the second encoder network are two different encoder networks, or the first en- coder network and the second encoder network share a part of subnet, or the second encoder net- work is a subnet of the first encoder network. 5. The method according to claim 4, wherein the first encoder network comprises a first feature extraction network, an autoregressive network, a side informa- tion extraction network, and a probability estimation network; and the residual information obtained by encoding the to- be-processed picture by using the first encoder net- work comprises: extracting a three-dimensional feature map of theto-be-processed picture by using the first feature extraction network, wherein the three- dimensional feature map comprises aplurality of feature elements; extracting side information of a to-be-encoded feature element from the three-dimensional fea- ture map by using the side information extraction network; estimating a first probability distribution mean of the to-be-encoded feature element by using the probability estimation network based on the side information; inputting an encoded feature element and the first probability distribution mean into the auto- regressive network to obtain a second probabil- ity distribution mean of the to-be-encoded fea- ture element; and obtaining residual information of the to-be-en- coded feature element based on the to-be-en- 5 10 15 20 25 30 35 40 45 50 55 25 47 EP 4 694 127 A1 48 coded feature element and the second probabil- ity distribution mean of the to-be-encoded fea- ture element. 6. The method according to claim 5, wherein the sec- ondencoder network comprises a second feature extraction network, a side information extraction net- work, and a probability estimation network; and the residual information obtained by encoding the to- be-processed picture by using the second encoder network comprises: extracting a three-dimensional feature map of the to-be-processed picture by using the second feature extraction network, wherein the three- dimensional feature map comprises aplurality of feature elements; extracting side information of a to-be-encoded feature element from the three-dimensional fea- ture map by using the side information extraction network; estimating a probability distribution mean of the to-be-encoded feature element by using the probability estimation network based on the side information; and obtaining residual information of the to-be-en- coded feature element based on the to-be-en- coded feature element and the probability dis- tribution mean. 7. The method according to claim 6, wherein the sec- ondfeature extraction network is a subnet of the first feature extraction network, or the second feature extraction network and the first feature extraction network are two completely different subnets. 8. The method according to any one of claims 5 to 7, wherein the method further comprises: encoding the side information into the bitstream. 9. The method according to any one of claims 1 to 8, wherein the identification information is located in a header of the bitstream. 10. A picture decoding method, comprising: receiving a bitstream; decoding the bitstream to obtain identification information indicating a used decoder network; and when the identification information is a first va- lue, decoding the bitstream to obtain a to-be- processed picture by using a first decoder net- work; or when the identification information is a second value, decoding the bitstream to obtain a to-be- processed picture by using a second decoder network, wherein a processing resource re- quired by the firstdecoder network is higher than a processing resource required by the sec- ond decoder network. 11. The method according to claim 10, wherein the first decoder network and the second decoder network are completely different decoder networks, or the first decoder network and the second decoder net- work share a part of subnet, or the second decoder network is a subnet of the first decoder network. 12. The method according to claim 10 or 11, wherein the first decoder network comprises an entropy decoder network, a probability estimation network, an auto- regressive network, and a first picture restoration network; and decoding the bitstream to obtain the to-be-pro- cessed picture by using the first decoder network comprises: decoding the bitstream to obtain side informa- tion of a three-dimensional feature map of the to- be-processed picture by using the entropy de- coder network, wherein the three-dimensional feature map comprises a plurality of feature elements; estimating a firstprobability distribution mean of a to-be-decoded feature element by using the probability estimation network based on the side information; determining a second probability distribution mean of the to-be-decoded feature element by using the autoregressive network based on the first probability distribution mean and a decoded feature element; decoding the bitstream to obtain residual infor- mation of the to-be-decoded feature element by using the entropy decoder network based on the second probability distribution mean, and ob- taining the to-be-decoded feature element based on the residual information and the sec- ond probability distribution mean; and restoring the to-be-processed picture by using the first picture restoration network based on the three-dimensional feature map obtained through decoding. 13. The method according to claim 12, wherein the second decoder network comprises the entropy de- coder network, the probability estimation network, and the second picture restorationnetwork; and decoding the bitstream to obtain the to-be-pro- cessed picture by using the second decoder network comprises: decoding the bitstream to obtain side informa- tion of a three-dimensional feature map of the to- be-processed picture by using the entropy de- 5 10 15 20 25 30 35 40 45 50 55 26 49 EP 4 694 127 A1 50 coder network, wherein the three-dimensional feature map comprises a plurality of feature elements; estimating a first probability distribution mean of a to-be-decoded feature element by using the probability estimation network based on the side information; decoding the bitstream to obtain residual infor- mation of the to-be-decoded feature element by using the entropy decoder network based on the first probability distribution mean, and obtaining the to-be-decoded feature element based on the residual information and the first probability dis- tribution mean; and restoring the to-be-processed picture by using the second picture restoration network based on thethree-dimensional feature map obtained through decoding. 14. The method according to claim 13, wherein the second picture restoration network is a subnet of the first picture restoration network, or the picture restoration network and the first picture restoration network share a part of subnet, or the second picture restoration network and the first picture restoration network are two different networks. 15. A picture encoding apparatus, comprising a memory and a video encoder, wherein the memory is configured to store video data, wherein the video data comprises a to-be-pro- cessed picture; and the video encoder is configured to encode iden- tification information indicating a used decoder network into a bitstream, wherein the identification information is a first value, in- dicating that the decoder network used to de- code the bitstream to obtain a to-be-processed picture is a first decoder network; or the identification information is a second value, indicating that the decodernetwork used to de- code the bitstream to obtain a to-be-processed picture is a second decoder network, wherein a processing resource required by the first deco- der network is higher than a processing re- source required by the second decoder network. 16. A picture decoding apparatus, comprising a memory and a video decoder, wherein the memory is configured to store video data in a bitstream form, wherein the video data com- prises a to-be-processed picture; and the video decoder is configured to: decode the bitstream to obtain identification information in- dicating a used decoder network; and when the identification information is a first va- lue, decode the bitstream to obtain a to-be-pro- cessed picture by using a first decoder network; or when the identification information is a second value, decode the bitstream to obtain a to-be- processed picture by using a second decoder network, wherein a processing resource re- quired by the first decoder network is higher than a processingresource required by the sec- ond decoder network. 17. A video decoding device, comprising a nonvolatile memory and a processor that are coupled to each other, wherein the processor invokes program code stored in the memory to perform the method accord- ing to any one of claims 10 to 14. 18. A video encoding device, comprising a nonvolatile memory and a processor that are coupled to each other, wherein the processor invokes program code stored in the memory to perform the method accord- ing to any one of claims 1 to 14. 19. A computer-readable storage medium, wherein the computer-readable storage medium stores program code, and when the computer program is run on a computer, the computer is enabled to perform the method according to any one of claims 10 to 14. 20. A computer-readable storage medium, wherein the computer-readable storage medium stores program code, and when the computer program is run on a computer, the computer is enabled to perform the method according to any one ofclaims 1 to 9. 21. A computer-readable storage medium, wherein the computer-readable storage medium stores a video bitstream decoded by one or more processors ac- cording to the method according to any one of claims 10 to 14. 22. A computer-readable storage medium, wherein the computer-readable storage medium stores a video bitstream obtained through encoding by one or more processors according to the method according to any one of claims 1 to 9. 5 10 15 20 25 30 35 40 45 50 55 27 EP 4 694 127 A1 28 EP 4 694 127 A1 29 EP 4 694 127 A1 30 EP 4 694 127 A1 31 EP 4 694 127 A1 32 EP 4 694 127 A1 33 EP 4 694 127 A1 34 EP 4 694 127 A1 35 EP 4 694 127 A1 36 EP 4 694 127 A1 37 EP 4 694 127 A1 38 EP 4 694 127 A1 39 EP 4 694 127 A1 40 EP 4 694 127 A1 41 EP 4 694 127 A1 42 EP 4 694 127 A1 43 EP 4 694 127 A1 44 EP 4 694 127 A1 45 EP 4 694 127 A1 46 EP 4 694 127 A1 47 EP 4 694 127 A1 48 EP 4 694 127 A1 49 EP 4 694 127 A1 50 EP 4 694 127 A1 51 EP 4 694 127 A1 52 EP 4 694 127 A1 53 EP 4 694 127 A1 54EP 4 694 127 A1 55 EP 4 694 127 A1 56 EP 4 694 127 A1 57 EP 4 694 127 A1 58 EP 4 694 127 A1 5 10 15 20 25 30 35 40 45 50 55 59 EP 4 694 127 A1 5 10 15 20 25 30 35 40 45 50 55 60 EP 4 694 127 A1 5 10 15 20 25 30 35 40 45 50 55 61 EP 4 694 127 A1 REFERENCES CITED IN THE DESCRIPTION This list of references cited by the applicant is for the reader’s convenience only. It does not form part of the European patent document. Even though great care has been taken in compiling the references, errors or omissions cannot be excluded and the EPO disclaims all liability in this regard. Patent documents cited in the description • CN 202310476967
[0001] • CN 202310956879
[0001]
Claims
1. A picture encoding method, comprising: encoding identification information indicating a used decoder network into a bitstream, wherein the identification information is a first value, indicating that the decoder network used to decode the bitstream to obtain a to-be-processed picture is a first decoder network; or the identification information is a second value, indicating that the decoder network used to decode the bitstream to obtain a to-be-processed picture is a second decoder network, wherein a processing resource required by the first decoder network is higher than a processing resource required by the second decoder network; and sending the bitstream.
2. The method according to claim 1, wherein the first decoder network and the second decoder network are completely different decoder networks, or the first decoder network and the second decoder network share a part of subnet, or the second decoder network is a subnet of the first decoder network.
3. The method according to claim 1 or 2, wherein the method further comprises: obtaining the identification information; and when the identification information is the first value, encoding, into the bitstream, residual information obtained by encoding the to-be-processed picture based on a first encoder network; or when the identification information is the second value, encoding, into the bitstream, residual information obtained by encoding the to-be-processed picture based on a second encoder network, wherein a processing resource required by the first encoder network is higher than a processing resource required by the second encoder network.
4. The method according to claim 3, wherein the first encoder network and the second encoder network are two different encoder networks, or the first encoder network and the second encoder network share a part of subnet, or the second encoder network is a subnet of the first encoder network.
5. The method according to claim 4, wherein the first encoder network comprises a first feature extraction network, an autoregressive network, a side information extraction network, and a probability estimation network; and the residual information obtained by encoding the to-be-processed picture by using the first encoder network comprises: extracting a three-dimensional feature map of the to-be-processed picture by using the first feature extraction network, wherein the three-dimensional feature map comprises a plurality of feature elements; extracting side information of a to-be-encoded feature element from the three-dimensional feature map by using the side information extraction network; estimating a first probability distribution mean of the to-be-encoded feature element by using the probability estimation network based on the side information; inputting an encoded feature element and the first probability distribution mean into the autoregressive network to obtain a second probability distribution mean of the to-be-encoded feature element; and obtaining residual information of the to-be-encoded feature element based on the to-be-encoded feature element and the second probability distribution mean of the to-be-encoded feature element.
6. The method according to claim 5, wherein the second encoder network comprises a second feature extraction network, a side information extraction network, and a probability estimation network; and the residual information obtained by encoding the to-be-processed picture by using the second encoder network comprises: extracting a three-dimensional feature map of the to-be-processed picture by using the second feature extraction network, wherein the three-dimensional feature map comprises a plurality of feature elements; extracting side information of a to-be-encoded feature element from the three-dimensional feature map by using the side information extraction network; estimating a probability distribution mean of the to-be-encoded feature element by using the probability estimation network based on the side information; and obtaining residual information of the to-be-encoded feature element based on the to-be-encoded feature element and the probability distribution mean.
7. The method according to claim 6, wherein the second feature extraction network is a subnet of the first feature extraction network, or the second feature extraction network and the first feature extraction network are two completely different subnets.
8. The method according to any one of claims 5 to 7, wherein the method further comprises: encoding the side information into the bitstream.
9. The method according to any one of claims 1 to 8, wherein the identification information is located in a header of the bitstream.
10. A picture decoding method, comprising: receiving a bitstream; decoding the bitstream to obtain identification information indicating a used decoder network; and when the identification information is a first value, decoding the bitstream to obtain a to-be-processed picture by using a first decoder network; or when the identification information is a second value, decoding the bitstream to obtain a to-be-processed picture by using a second decoder network, wherein a processing resource required by the first decoder network is higher than a processing resource required by the second decoder network.
11. The method according to claim 10, wherein the first decoder network and the second decoder network are completely different decoder networks, or the first decoder network and the second decoder network share a part of subnet, or the second decoder network is a subnet of the first decoder network.
12. The method according to claim 10 or 11, wherein the first decoder network comprises an entropy decoder network, a probability estimation network, an autoregressive network, and a first picture restoration network; and decoding the bitstream to obtain the to-be-processed picture by using the first decoder network comprises: decoding the bitstream to obtain side information of a three-dimensional feature map of the to-be-processed picture by using the entropy decoder network, wherein the three-dimensional feature map comprises a plurality of feature elements; estimating a first probability distribution mean of a to-be-decoded feature element by using the probability estimation network based on the side information; determining a second probability distribution mean of the to-be-decoded feature element by using the autoregressive network based on the first probability distribution mean and a decoded feature element; decoding the bitstream to obtain residual information of the to-be-decoded feature element by using the entropy decoder network based on the second probability distribution mean, and obtaining the to-be-decoded feature element based on the residual information and the second probability distribution mean; and restoring the to-be-processed picture by using the first picture restoration network based on the three-dimensional feature map obtained through decoding.
13. The method according to claim 12, wherein the second decoder network comprises the entropy decoder network, the probability estimation network, and the second picture restoration network; and decoding the bitstream to obtain the to-be-processed picture by using the second decoder network comprises: decoding the bitstream to obtain side information of a three-dimensional feature map of the to-be-processed picture by using the entropy decoder network, wherein the three-dimensional feature map comprises a plurality of feature elements; estimating a first probability distribution mean of a to-be-decoded feature element by using the probability estimation network based on the side information; decoding the bitstream to obtain residual information of the to-be-decoded feature element by using the entropy decoder network based on the first probability distribution mean, and obtaining the to-be-decoded feature element based on the residual information and the first probability distribution mean; and restoring the to-be-processed picture by using the second picture restoration network based on the three-dimensional feature map obtained through decoding.
14. The method according to claim 13, wherein the second picture restoration network is a subnet of the first picture restoration network, or the picture restoration network and the first picture restoration network share a part of subnet, or the second picture restoration network and the first picture restoration network are two different networks.
15. A picture encoding apparatus, comprising a memory and a video encoder, wherein the memory is configured to store video data, wherein the video data comprises a to-be-processed picture; and the video encoder is configured to encode identification information indicating a used decoder network into a bitstream, wherein the identification information is a first value, indicating that the decoder network used to decode the bitstream to obtain a to-be-processed picture is a first decoder network; or the identification information is a second value, indicating that the decoder network used to decode the bitstream to obtain a to-be-processed picture is a second decoder network, wherein a processing resource required by the first decoder network is higher than a processing resource required by the second decoder network.
16. A picture decoding apparatus, comprising a memory and a video decoder, wherein the memory is configured to store video data in a bitstream form, wherein the video data comprises a to-be-processed picture; and the video decoder is configured to: decode the bitstream to obtain identification information indicating a used decoder network; and when the identification information is a first value, decode the bitstream to obtain a to-be-processed picture by using a first decoder network; or when the identification information is a second value, decode the bitstream to obtain a to-be-processed picture by using a second decoder network, wherein a processing resource required by the first decoder network is higher than a processing resource required by the second decoder network.
17. A video decoding device, comprising a nonvolatile memory and a processor that are coupled to each other, wherein the processor invokes program code stored in the memory to perform the method according to any one of claims 10 to 14.
18. A video encoding device, comprising a nonvolatile memory and a processor that are coupled to each other, wherein the processor invokes program code stored in the memory to perform the method according to any one of claims 1 to 14.
19. A computer-readable storage medium, wherein the computer-readable storage medium stores program code, and when the computer program is run on a computer, the computer is enabled to perform the method according to any one of claims 10 to 14.
20. A computer-readable storage medium, wherein the computer-readable storage medium stores program code, and when the computer program is run on a computer, the computer is enabled to perform the method according to any one of claims 1 to 9.
21. A computer-readable storage medium, wherein the computer-readable storage medium stores a video bitstream decoded by one or more processors according to the method according to any one of claims 10 to 14.
22. A computer-readable storage medium, wherein the computer-readable storage medium stores a video bitstream obtained through encoding by one or more processors according to the method according to any one of claims 1 to 9.