Pretrained adapters for decoder-side neural networks

Pretrained adapters with adaptation parameters enhance neural network flexibility and efficiency by allowing integration into specific network positions, addressing the challenge of adapting to diverse data types without extensive retraining.

WO2026009114A1PCT designated stage Publication Date: 2026-01-08NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/056570
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-03
Filing Date
2025-06-27
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing neural networks face challenges in efficiently adapting to diverse data processing tasks without requiring extensive retraining, particularly in decoder-side operations for media data encoding and decoding.

Method used

The use of pretrained adapters with adaptation parameters that can be inserted or integrated into specific positions within neural networks, allowing for efficient adaptation based on selection criteria such as similarity metrics and decoder-side indicators, enabling flexible and efficient processing of various media data types.

Benefits of technology

This approach allows for adaptive neural networks that can process diverse media data efficiently, reducing the need for extensive retraining and improving performance across different data types and processing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000077_0000
    Figure 00000077_0000
  • Figure 00000077_0001
    Figure 00000077_0001
  • Figure 00000078_0000
    Figure 00000078_0000
Patent Text Reader

Abstract

Various embodiments provide methods, apparatuses, and computer program products. An example apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform; selecting at least one adapter from one or more adapters based at least on a selection criterion to obtain at least one selected adapter; wherein each of the one or more adapters comprises one or more adaptation parameters; using the at least one selected adapter to adapt a neural network, based at least on an adaptation process, to obtain an adapted neural network; and wherein the adapted neural network is intended to be used to process a data item that is input to the neural network.
Need to check novelty before this filing date? Find Prior Art

Description

PRETRAINED ADAPTERS FOR DECODER-SIDE NEURAL NETWORKSTECHNICAL FIELD

[0001] The examples and non-limiting embodiments relate generally to neural networks and, more particularly to, using pretrained adapters for decoder-side neural networks.BACKGROUND

[0002] It is known to use neural networks for encoding or decoding of media data.SUMMARY

[0003] Example 1: An apparatus comprising; at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: selecting at least one adapter from one or more adapters based at least on a selection criterion to obtain at least one selected adapter; wherein each of the one or more adapters comprises one or more adaptation parameters; using the at least one selected adapter to adapt a neural network, based at least on an adaptation process, to obtain an adapted neural network; and wherein the adapted neural network is intended to be used to process a data item that is input to the neural network.

[0004] Example 2: The apparatus of example 1, wherein the neural network and / or the adapted neural network are part of an encoding process at an encoder side, a decoding process at a decoder side, and / or a post-processing process at a decoder side.

[0005] Example 3: The apparatus of any of examples 1 or 2, wherein the one or more adaptation parameters of each of the one or more adapters are associated with respective one or more positions within the neural network.

[0006] Example 4: The apparatus of any of the previous examples, wherein the one or more adaptation parameters of each of the one or more adapters are grouped into one or more adaptation layers, wherein the one or more adaptation layers are associated with respective one or more layer positions within the neural network.

[0007] Example 5: The apparatus of any of the previous examples, wherein the apparatus is further caused to perform; using the one or more adaptation parameters as parameters or weights of the neural network; using the one or more adaptation parameters as bias parameters of the neural network; using the one or more adaptation parameters as multiplier parameters of the neural network; using the one ormore adaptation parameters as parameters of one or more convolutional layers; using the one or more adaptation parameters as an external or internal input of the neural network; using the one or more adaptation parameters as an input to the neural network; or using the one or more adaptation parameters as an input to one or more layers of the neural network.

[0008] Example 6: The apparatus of any of the previous examples, wherein the adaptation process comprises; inserting the one or more adaptation parameters of the at least one selected adapter into the respective one or more associated positions within the neural network; replacing values of at least one parameter of the neural network with values of the one or more adaptation parameters of the at least one selected adapter; modifying the values of at least one parameter of the neural network based on the values of the one or more adaptation parameters of the at least one selected adapter; and inserting the one or more adaptation layers of the at least one selected adapter into respective one or more associated layer positions within the neural network.

[0009] Example 7: The apparatus of any of the previous examples, wherein the one or more adapters are associated with respective one or more identifiers or respective one or more nominal or numerical values of a feature or characteristic of at least one of: the data item, the neural network, the encoding process, or the decoding process.

[0010] Example 8: The apparatus of any of the previous examples, wherein the neural network is adapted on one or more of: a video sequence basis, a coded layer video sequence basis, a random-access segment basis, an intra-frame period basis, a frame or picture basis, or a block basis.

[0011] Example 9: The apparatus of any of the previous examples, wherein the apparatus is further caused to perform: grouping two or more adapters of the one or more adapters into two or more categories, wherein a category of the two or more categories comprises one or more adapters that are similar or substantially similar based on a similarity metric or a feature, and wherein one or more first adapters in a first category of the two or more categories are different from one or more second adapters in a second category of the two or more categories based on the similarity metric or the feature.

[0012] Example 10: The apparatus of any of the previous examples, wherein the apparatus is further caused to perform: combining two or more adapters, by using a combination operation, to obtain a combined adapter, wherein the combined adapter is used to adapt the neural network.

[0013] Example 11: The apparatus of example 10, wherein the combination operation comprises averaging, weighted averaging, or using the two or more adapters by cascading adaptation parameters of different adapters in the two or more adapters.

[0014] Example 12: The apparatus of example 11, wherein when the combination operation comprises weighted averaging, weights or coefficients used in the weighted averaging are determined based on; a similarity between a data item to be processed and the two or more adapters, in terms of a similarity metric or a similarity criterion; or a similarity between a characteristic of the data item to be processed and characteristics of the two or more adapters, in terms of the similarity metric or similarity criterion.

[0015] Example 13: The apparatus of any of the examples 10 to 12, wherein the two or more adapters that are combined are selected based on; a similarity between the data item to be processed and the two or more adapters, in terms of the similarity metric or similarity criterion; or a similarity between a characteristic of a data item to be processed and characteristics of the two or more adapters, in terms of the similarity metric or similarity criterion.

[0016] Example 14: The apparatus of example 13, wherein the similarity criterion comprises a similarity between a value of a particular feature of the data item to be processed and values of particular features that are associated with the two or more adapters to be combined.

[0017] Example 15: The apparatus of any of the previous examples, wherein the selection criterion comprises; selecting the at least one selected adapter based on an indication received from an encoder; or selecting the at least one selected adapter based on one or more parameters or indicators that are available at a decoder side and that are used for some other operations of the decoder in addition to being used for selecting the at least one selected adapter.

[0018] Example 16: The apparatus of example 15, wherein the indication is associated with information about a scope of the at least one selected adapter.

[0019] Example 17: The apparatus of example 16, wherein the indication is indicative of a position of the neural network within a processing pipeline; the indication is indicative of a content type, wherein the content type is associated to the at least one selected adapter; or the indication is indicative of a loss function that is associated to the at least one selected adapter.

[0020] Example 18: The apparatus of example 15, wherein: the one or more parameters or indicators comprise one or more of the following; a quantization parameter associated with a picture to be processed by using the neural network; a picture type or block type of a picture or a block to be processed by using the neural network; a resolution information of one or more video frames to be processed by using the neural network; a temporal layer identifier of a temporal layer of an inter-frame prediction hierarchy structure; a position of the neural network within a coding pipeline; or a type of content to be processed by using the neural network.

[0021] Example 19: The apparatus of example 18, wherein the apparatus is further caused to perform: signaling or receiving a correction or adjustment for the one or more parameters or indicators that are available at decoder side and that are used for some other operations of the decoder in addition to being used for selecting the at least one selected adapter, wherein the one or more parameters or indicators are corrected or adjusted based on the correction or adjustment to optimize selection of the at least one selected adapter.

[0022] Example 20: The apparatus of example 19, wherein the correction or adjustment are intended to be used by the decoder to correct or adjust the one or more parameters or indicators to obtain respective one or more corrected or adjusted parameters or indicators that are suitable or optimal to be used for selection of the at least one adapter.

[0023] Example 21: The apparatus of any of the previous examples, wherein the apparatus is further caused to perform: adapting the at least one selected adapter, based on an adaptation signal, to obtain an updated adapter, wherein the updated adapter is used to adapt the neural network.

[0024] Example 22: The apparatus of any of the previous examples, wherein the one or more adaptation parameters are associated with one or more adaptation layers of the at least one selected adapter, and wherein the one or more adaptation parameters of the one or more adaptation layers of the at least one selected adapter are integrated into already present parameters of the neural network.

[0025] Example 23: The apparatus of any of the previous examples, wherein the apparatus is further caused to perform: initializing the one or more adaptation parameters based on: a predetermined value, values sampled from a probability distribution, or a decomposition operation that is performed on parameters of the neural network associated with the one or more adaptation parameters.

[0026] Example 24: The apparatus of any of the previous examples, wherein the neural network comprises a decoder side neural network.

[0027] Example 25 : The apparatus of example 2, wherein the decoder side comprises N adapters that are associated with N quantization parameters (QPs) or with N QP ranges, wherein QP refers to the quantization parameter used to derive the data item, wherein N is a natural number greater than or equal to 1.

[0028] Example 26: The apparatus of example 25, wherein the QP comprises one of the following: a sequence QP, a picture QP, a slice QP, or a block QP.

[0029] Example 27 : The apparatus of any of examples 25 or 26, wherein when the data item is codedwith a particular QP, the particular QP information is available at decoder side, and wherein the decoder selects an adapter associated with a QP that is equal to the particular QP, an adapter that is associated with a QP that is most similar to the particular QP, or an adapter that is associated with a QP range that comprises the particular QP.

[0030] Example 28: The apparatus of example 2, wherein the decoder side comprises N adapters that are associated with N possible positions of the neural network within a pipeline or cascade of processing steps, wherein N is a natural number greater than or equal to 1.

[0031] Example 29: The apparatus of example 28, wherein the apparatus is further caused to perform: receiving an indication of the position of the neural network to be used for processing the data item, and wherein the indication of the position is further used for selecting the adapter that is associated with that position.

[0032] Example 30: The apparatus of example 2, wherein the decoder side comprises N adapters that are associated with N types of content, wherein N is a natural number greater than or equal to 1.

[0033] Example 31: The apparatus of example 30, wherein a type of content comprises a particular set of characteristics, and different types of content comprise different sets of characteristics, and wherein the content or data in a certain content type share same or substantially same characteristics, and wherein the content or data in different content types do not share same or substantially same characteristics.

[0034] Example 32: The apparatus of example 30, wherein the apparatus is further caused to perform; receiving an indication of the content type associated with the data item; selecting the adapter that is associated with the same or substantially same content type as indicated by an encoder; using the selected adapter for adapting the neural network; and using the neural network for processing the data item.

[0035] Example 33: The apparatus of example 2, wherein the decoder side comprises N adapters that are associated with N picture types, wherein N is a natural number greater than or equal to 1.

[0036] Example 34: The apparatus of example 33, wherein a picture type comprises an intra-coded picture, or an inter-coded picture.

[0037] Example 35: The apparatus of example 33, wherein the decoder comprises information about a picture type associated to a picture to be decoded, and wherein information about the picture type associated to the picture to be decoded is received from an encoder or is derived at the decoder side.

[0038] Example 36: The apparatus of example 35, wherein when the decoder decodes the picture of the particular picture type, the decoder is caused to perform: selecting the adapter that is associated with that particular picture type, or that is associated with a picture type that is most similar to that particular picture type.

[0039] Example 37 : The apparatus of example 2, wherein the decoder side comprises N adapters that are associated with N resolutions or N resolution-ranges, wherein N is a natural number greater than or equal to 1.

[0040] Example 38: The apparatus of example 37, wherein when the data item has been coded with a particular resolution, information about the particular resolution is available at the decoder side, and wherein the information about that particular resolution is received from an encoder, or may be derived at decoder side based on other information.

[0041] Example 39: The apparatus of example 37, wherein the decoder is caused to perform: selecting an adapter that is associated with a resolution that is equal to the particular resolution, or an adapter that is associated with a resolution that is most similar to the particular resolution, or an adapter that is associated with a resolution range that comprises the particular resolution.

[0042] Example 40: The apparatus of example 2, wherein the decoder side comprises N adapters that are associated with N temporal layers of a predetermined hierarchy for inter-frame prediction patterns, wherein N is a natural number greater than or equal to 1.

[0043] Example 41: The apparatus of example 40, wherein when the data item has been coded with a particular temporal layer, information about the particular temporal layer is available at the decoder side, and wherein the information about the particular temporal layer is received from an encoder, or may be derived at decoder side based on other information.

[0044] Example 42: The apparatus of example 41 , wherein the decoder is caused to perform: selecting an adapter that is associated with a temporal layer that is equal to the particular temporal layer, or an adapter that is associated with a resolution that is most similar to the particular temporal layer according to a predefined metric or evaluation process.

[0045] Example 43 : The apparatus of example 2, wherein the decoder side comprises N adapters that are associated with loss functions that were used to train the N adapters.

[0046] Example 44: The apparatus of example 43, wherein the apparatus is further caused to perform: receiving information about which adapter to use for processing the data item.

[0047] Example 45: The apparatus of example 44, wherein the information explicitly indicates a loss function used to train the adapter to be used for processing the data item.

[0048] Example 46: The apparatus of example 43, wherein the apparatus is further caused to perform; selecting an adapter that is associated with a loss function that is equal to the loss function; or selecting an adapter that is associated with a loss function that is most similar to the loss function according to a predefined metric or evaluation process.

[0049] Example 47: The apparatus of example 44, wherein the information indicates which adapter to use for processing the data item without an explicit reference to how the adapter is trained.

[0050] Example 48: A method comprising; selecting at least one adapter from one or more adapters based at least on a selection criterion to obtain at least one selected adapter; wherein each of the one or more adapters comprises one or more adaptation parameters; using the at least one selected adapter to adapt a neural network, based at least on an adaptation process, to obtain an adapted neural network; and wherein the adapted neural network is intended to be used to process a data item that is input to the neural network.

[0051] Example 49: The method of example 48, wherein the neural network and / or the adapted neural network are part of an encoding process at an encoder side, a decoding process at a decoder side, and / or a post-processing process at a decoder side.

[0052] Example 50: The method of any of examples 48 or 49, wherein the one or more adaptation parameters of each of the one or more adapters are associated with respective one or more positions within the neural network.

[0053] Example 51 : The method of any of the examples 48 to 50, wherein the one or more adaptation parameters of each of the one or more adapters are grouped into one or more adaptation layers, wherein the one or more adaptation layers are associated with respective one or more layer positions within the neural network.

[0054] Example 52: The method of any of the examples 48 to 51 further comprising; using the one or more adaptation parameters as parameters or weights of the neural network; using the one or more adaptation parameters as bias parameters of the neural network; using the one or more adaptation parameters as multiplier parameters of the neural network; using the one or more adaptation parameters as parameters of one or more convolutional layers; using the one or more adaptation parameters as an external or internal input of the neural network; using the one or more adaptation parameters as an input to the neural network; or using the one or more adaptation parameters as an input to one or more layersof the neural network.

[0055] Example 53: The method of any of the examples 48 to 52, wherein the adaptation process comprises; inserting the one or more adaptation parameters of the at least one selected adapter into the respective one or more associated positions within the neural network; replacing values of at least one parameter of the neural network with values of the one or more adaptation parameters of the at least one selected adapter; modifying the values of at least one parameter of the neural network based on the values of the one or more adaptation parameters of the at least one selected adapter; and inserting the one or more adaptation layers of the at least one selected adapter into respective one or more associated layer positions within the neural network.

[0056] Example 54: The method of any of the examples 48 to 53, wherein the one or more adapters are associated with respective one or more identifiers or respective one or more nominal or numerical values of a feature or characteristic of at least one of: the data item, the neural network, the encoding process, or the decoding process.

[0057] Example 55: The method of any of the examples 48 to 54, wherein the neural network is adapted on one or more of: a video sequence basis, a coded layer video sequence basis, a random-access segment basis, an intra-frame period basis, a frame or picture basis, or a block basis.

[0058] Example 56: The method of any of the examples 48 to 55 further comprising: grouping two or more adapters of the one or more adapters into two or more categories, wherein a category of the two or more categories comprises one or more adapters that are similar or substantially similar based on a similarity metric or a feature, and wherein one or more first adapters in a first category of the two or more categories are different from one or more second adapters in a second category of the two or more categories based on the similarity metric or the feature.

[0059] Example 57: The method of any of the examples 48 to 56 further comprising: combining two or more adapters, by using a combination operation, to obtain a combined adapter, wherein the combined adapter is used to adapt the neural network.

[0060] Example 58: The method of example 57, wherein the combination operation comprises averaging, weighted averaging, or using the two or more adapters by cascading adaptation parameters of different adapters in the two or more adapters.

[0061] Example 59: The method of example 58, wherein when the combination operation comprises weighted averaging, weights or coefficients used in the weighted averaging are determined based on; a similarity between a data item to be processed and the two or more adapters, in terms of a similaritymetric or a similarity criterion; or a similarity between a characteristic of the data item to be processed and characteristics of the two or more adapters, in terms of the similarity metric or similarity criterion.

[0062] Example 60: The method of any of the examples 57 to 59, wherein the two or more adapters that are combined are selected based on; a similarity between the data item to be processed and the two or more adapters, in terms of the similarity metric or similarity criterion; or a similarity between a characteristic of a data item to be processed and characteristics of the two or more adapters, in terms of the similarity metric or similarity criterion.

[0063] Example 61: The method of example 60, wherein the similarity criterion comprises a similarity between a value of a particular feature of the data item to be processed and values of particular features that are associated with the two or more adapters to be combined.

[0064] Example 62: The method of any of the examples 48 to 61, wherein the selection criterion comprises; selecting the at least one selected adapter based on an indication received from an encoder; or selecting the at least one selected adapter based on one or more parameters or indicators that are available at a decoder side and that are used for some other operations of the decoder in addition to being used for selecting the at least one selected adapter.

[0065] Example 63: The method of example 62, wherein the indication is associated with information about a scope of the at least one selected adapter.

[0066] Example 64: The method of example 63, wherein the indication is indicative of a position of the neural network within a processing pipeline; the indication is indicative of a content type, wherein the content type is associated to the at least one selected adapter; or the indication is indicative of a loss function that is associated to the at least one selected adapter.

[0067] Example 65: The method of example 62, wherein: the one or more parameters or indicators comprise one or more of the following; a quantization parameter associated with a picture to be processed by using the neural network; a picture type or block type of a picture or a block to be processed by using the neural network; a resolution information of one or more video frames to be processed by using the neural network; a temporal layer identifier of a temporal layer of an inter-frame prediction hierarchy structure; a position of the neural network within a coding pipeline; or a type of content to be processed by using the neural network.

[0068] Example 66: The method of example 65 further comprising: signaling or receiving a correction or adjustment for the one or more parameters or indicators that are available at decoder side and that are used for some other operations of the decoder in addition to being used for selecting the atleast one selected adapter, wherein the one or more parameters or indicators are corrected or adjusted based on the correction or adjustment to optimize selection of the at least one selected adapter.

[0069] Example 67 : The method of example 66, wherein the correction or adjustment are intended to be used by the decoder to correct or adjust the one or more parameters or indicators to obtain respective one or more corrected or adjusted parameters or indicators that are suitable or optimal to be used for selection of the at least one adapter.

[0070] Example 68: The method of any of the examples 48 to 67 further comprising: adapting the at least one selected adapter, based on an adaptation signal, to obtain an updated adapter, wherein the updated adapter is used to adapt the neural network.

[0071] Example 69: The method of any of the examples 48 to 68, wherein the one or more adaptation parameters are associated with one or more adaptation layers of the at least one selected adapter, and wherein the one or more adaptation parameters of the one or more adaptation layers of the at least one selected adapter are integrated into already present parameters of the neural network.

[0072] Example 70: The method of any of the examples 48 to 69 further comprising: initializing the one or more adaptation parameters based on: a predetermined value, values sampled from a probability distribution, or a decomposition operation that is performed on parameters of the neural network associated with the one or more adaptation parameters.

[0073] Example 71: The method of any of the examples 48 to 70, wherein the neural network comprises a decoder side neural network.

[0074] Example 72: The method of example 49, wherein the decoder side comprises N adapters that are associated with N quantization parameters (QPs) or with N QP ranges, wherein QP refers to the quantization parameter used to derive the data item, wherein N is a natural number greater than or equal to 1.

[0075] Example 73: The method of example 72, wherein the QP comprises one of the following: a sequence QP, a picture QP, a slice QP, or a block QP.

[0076] Example 74: The method of any of examples 73 or 73, wherein when the data item is coded with a particular QP, the particular QP information is available at decoder side, and wherein the decoder selects an adapter associated with a QP that is equal to the particular QP, an adapter that is associated with a QP that is most similar to the particular QP, or an adapter that is associated with a QP range that comprises the particular QP.

[0077] Example 75: The method of example 49, wherein the decoder side comprises N adapters that are associated with N possible positions of the neural network within a pipeline or cascade of processing steps, wherein N is a natural number greater than or equal to 1.

[0078] Example 76: The method of example 75 further comprising: receiving an indication of the position of the neural network to be used for processing the data item, and wherein the indication of the position is further used for selecting the adapter that is associated with that position.

[0079] Example 77: The method of example 49, wherein the decoder side comprises N adapters that are associated with N types of content, wherein N is a natural number greater than or equal to 1.

[0080] Example 78: The method of example 77, wherein a type of content comprises a particular set of characteristics, and different types of content comprise different sets of characteristics, and wherein the content or data in a certain content type share same or substantially same characteristics, and wherein the content or data in different content types do not share same or substantially same characteristics.

[0081] Example 79: The method of example 77 further comprising; receiving an indication of the content type associated with the data item; selecting the adapter that is associated with the same or substantially same content type as indicated by an encoder; using the selected adapter for adapting the neural network; and using the neural network for processing the data item.

[0082] Example 80: The method of example 49, wherein the decoder side comprises N adapters that are associated with N picture types, wherein N is a natural number greater than or equal to 1.

[0083] Example 81: The method of example 80, wherein a picture type comprises an intra-coded picture, or an inter-coded picture.

[0084] Example 82: The method of example 80, wherein the decoder comprises information about a picture type associated to a picture to be decoded, and wherein information about the picture type associated to the picture to be decoded is received from an encoder or is derived at the decoder side.

[0085] Example 83: The method of example 82, wherein when the decoder decodes the picture of the particular picture type, the decoder is caused to perform: selecting the adapter that is associated with that particular picture type, or that is associated with a picture type that is most similar to that particular picture type.

[0086] Example 84: The method of example 49, wherein the decoder side comprises N adapters that are associated with N resolutions or N resolution-ranges, wherein N is a natural number greater than or equal to 1.

[0087] Example 85: The method of example 84, wherein when the data item has been coded with a particular resolution, information about the particular resolution is available at the decoder side, and wherein the information about that particular resolution is received from an encoder, or may be derived at decoder side based on other information.

[0088] Example 86: The method of example 84, wherein the decoder is caused to perform: selecting an adapter that is associated with a resolution that is equal to the particular resolution, or an adapter that is associated with a resolution that is most similar to the particular resolution, or an adapter that is associated with a resolution range that comprises the particular resolution.

[0089] Example 87 : The method of example 49, wherein the decoder side comprises N adapters that are associated with N temporal layers of a predetermined hierarchy for inter-frame prediction patterns, wherein N is a natural number greater than or equal to 1.

[0090] Example 88: The method of example 87, wherein when the data item has been coded with a particular temporal layer, information about the particular temporal layer is available at the decoder side, and wherein the information about the particular temporal layer is received from an encoder, or may be derived at decoder side based on other information.

[0091] Example 89: The method of example 88, wherein the decoder is caused to perform: selecting an adapter that is associated with a temporal layer that is equal to the particular temporal layer, or an adapter that is associated with a resolution that is most similar to the particular temporal layer according to a predefined metric or evaluation process.

[0092] Example 90: The method of example 49, wherein the decoder side comprises N adapters that are associated with loss functions that were used to train the N adapters.

[0093] Example 91: The method of example 90 further comprising: receiving information about which adapter to use for processing the data item.

[0094] Example 92: The method of example 91, wherein the information explicitly indicates a loss function used to train the adapter to be used for processing the data item.

[0095] Example 93: The method of example 92 further comprising; selecting an adapter that is associated with a loss function that is equal to the loss function; or selecting an adapter that is associated with a loss function that is most similar to the loss function according to a predefined metric or evaluation process.

[0096] Example 94: The method of example 93, wherein the information indicates which adapter touse for processing the data item without an explicit reference to how the adapter is trained.

[0097] Example 95: An apparatus comprising means for performing methods as described in any of the examples 48 to 94.

[0098] Example 96: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as described in any of the examples 48 to 94.

[0099] Example 97 : The computer readable medium clam 96, wherein the computer readable medium comprises a non-transitory computer readable medium.BRIEF DESCRIPTION OF THE DRAWINGS

[0100] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:

[0101] FIG. 1 shows schematically an apparatus employing embodiments of the examples described herein.

[0102] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.

[0103] FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.

[0104] FIG. 4 is a block diagram illustrating a system in accordance with an example.

[0105] FIG. 5 illustrates an example of modified video coding pipeline based on neural networks.

[0106] FIG. 6 illustrates an example neural network-based end-to-end learned codec, such as an end- to-end learned video codec.

[0107] FIG. 7 illustrates an example pipeline performing video coding for machines.

[0108] FIG. 8 illustrates an example of encoder-side operations for overfitting a neural network based filter.

[0109] FIG. 9 illustrates an example of decoder or receiver side operations for updating a neural network based filter.

[0110] FIG. 10 describes an illustration of example embodiments.

[0111] FIG. 11 is an example apparatus, which may be implemented in hardware, and is caused to, implement examples described herein.

[0112] FIG. 12 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.

[0113] FIG. 13 is an example method performed with an encoder or a decoder, based on the examples described herein.DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS

[0114] The following acronyms and abbreviations that may be found in the specification and / or the drawing figures are defined as follows (the abbreviations may be appended with each other or with other characters using e.g. a hyphen or dash (-), and may be case insensitive):4CC four character code5G fifth generation cellular network technology5GC 5G core network a.k.a. also known asAVC advanced video codingCU coding unitDSP digital signal processorDU distributed unit eNB (or eNodeB) evolved Node B (for example, an LTE base station)EN-DC E-UTRA-NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN-DCE-UTRA evolved universal terrestrial radio access, for example, theLTE radio access technologyFl or Fl-C interface between CU and DU control interfacegNB (or gNodeB) base station for 5G / NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GCIEC International Electrotechnical Commission loT internet of thingsISO International Organization for StandardizationISOBMFF ISO base media file formatJPEG joint photographic experts groupLTE long-term evolution mdat MediaDataBoxMIME Multipurpose Internet Mail ExtensionMME mobility management entity moov MovieBoxMP4 file format for MPEG-4 Part 14 filesMPEG moving picture experts groupMPEG-2 H.222 / H.262 as defined by the ITUMPEG-4 audio and video coding standard for ISO / IEC 14496 ng or NG new generation ng-eNB or NG-eNB new generation eNBNR new radio (5G radio)N / W or NW networkPDCP packet data convergence protocolPHY physical layerPNG portable network graphicsRAN radio access networkRFC request for commentsRLC radio link controlRRC radio resource controlRRH remote radio headRU radio unitRx receiverSDAP service data adaptation protocolSGW serving gatewaySMF session management functionSPS sequence parameter setSVC scalable video codingSI interface between eNodeBs and the EPC trak TrackBoxTx transmitterUE user equipmentUICC Universal Integrated Circuit CardUPF user plane functionURL uniform resource locatorX2 interconnecting interface between two eNodeBs in LIE networkXn interface between two NG-RAN nodes

[0115] Some embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments may be shown. Indeed, various embodiments of the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals refer to like elements throughout. As used herein, the terms ‘data,’ ‘content,’ ‘information,’ and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance with embodiments of the present invention. Thus, use of any such terms should not be taken to limit the spiritand scope of embodiments.

[0116] Described herein is a method and apparatus for determining and using pretrained adapters for decoder-side neural networks.

[0117] The following describes in detail a suitable apparatus and possible method for determining and using pretrained adapters for decoder-side neural networks according to embodiments. In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an electronic device or apparatus 100. The apparatus 100 may be an Internet of Things (loT) apparatus configured to perform various functions, such as for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 are explained next.

[0118] The apparatus 100 may for example be a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or other lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.

[0119] The apparatus 100 may comprise a housing 101 for incorporating and protecting the device. The apparatus 100 further may comprise a display 102 in the form of a liquid crystal display. In other embodiments of the examples described herein the display may be any suitable display technology suitable to display an image or video. The apparatus 100 may further comprise a keypad 104. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.

[0120] The apparatus may comprise a microphone 106 or any suitable audio input which may be a digital or analog signal input. The apparatus 100 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 108, speaker, or an analog audio or digital audio output connection. The apparatus 100 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus 100 may further comprise a camera 109 capable of recording or capturing images and / or video. The apparatus 100 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 100 may further comprise any suitable short range communication solutionsuch as for example a Bluetooth wireless connection or a USB / firewire wired connection.

[0121] The apparatus 100 may comprise a controller 110, processor or processor circuitry for controlling the apparatus 100. The controller 110 may be connected to memory 112 which in embodiments of the examples described herein may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 110. The controller 110 may further be connected to codec circuitry 114 suitable for carrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.

[0122] The apparatus 100 may further comprise a card reader 118 and a smart card 116, for example a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.

[0123] The apparatus 100 may comprise radio interface circuitry 120 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 100 may further comprise an antenna 122 connected to the radio interface circuitry 120 for transmitting radio frequency signals generated at the radio interface circuitry 120 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).

[0124] The apparatus 100 may comprise a camera capable of recording or detecting individual frames which are then passed to the codec circuitry 114 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and / or storage. The apparatus 100 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 100 described above represent examples of means for performing a corresponding function.

[0125] With respect to FIG. 3, an example of a system within which embodiments of the examples described herein can be utilized is shown. The system 300 comprises multiple communication devices which can communicate through one or more networks. The system 300 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network, etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.

[0126] The system 300 may include both wired and wireless communication devices and / or apparatus 100 suitable for implementing embodiments of the examples described herein.

[0127] For example, the system shown in FIG. 3 shows a mobile telephone network 301 and a representation of the internet 302. Connectivity to the internet 302 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.

[0128] The example communication devices shown in the system 300 may include, but are not limited to, an electronic device or apparatus 100, a combination of a personal digital assistant (PDA) and a mobile telephone 304, a PDA 306, an integrated messaging device (IMD) 308, a desktop computer 310, a notebook computer 312, or a head-mounted apparatus. The head-mounted apparatus may be a head-mounted display (HMD), or glasses having a device such as a camera configured to encode and / or decode images and / or video. The apparatus 100 may be stationary or mobile when carried by an individual who is moving. The apparatus 100 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.

[0129] The embodiments may also be implemented in a set-top box; e.g., a digital TV receiver, gaming consoles, e.g., a steam deck which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.

[0130] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 314 to a base station 316. The base station 316 may be connected to a network server 318 that allows communication between the mobile telephone network 301 and the internet 302. The system may include additional communication devices and communication devices of various types.

[0131] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocolinternet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.

[0132] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.

[0133] The embodiments may also be implemented in so-called loT devices. The Internet of Things (loT) may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home / building automation, etc. to be included in the Internet of Things (loT). In order to utilize the Internet loT devices are provided with an IP address as a unique identifier. loT devices may be provided with a radio transmitter, such as a WLAN or Bluetooth transmitter or a RFID tag. Alternatively, loT devices may have access to an IP -based network via a wired network, such as an Ethernet-based network or a power-line connection (PLC).

[0134] FIG. 4 is a block diagram illustrating a system or apparatus 400 in accordance with several examples. In an example, the encoder 402 is used to encode an image or video from the scene 404, and the encoder 402 is implemented in a transmitting apparatus 406. The encoder 402 produces a bitstream 408 comprising signaling that is received by the receiving apparatus 410, which implements a decoder 412. The encoder 402 sends the bitstream 408 that comprises the herein described signaling. The decoder 412 forms the image or video for the scene 404-1, and the receiving apparatus 410 would present this to the user, e.g., via a smartphone, television, or projector among many other options.

[0135] In some examples, the transmitting apparatus 406 and the receiving apparatus 410 are at least partially within a common apparatus, and for example, are located within a common housing 414. In other examples the transmitting apparatus 406 and the receiving apparatus 410 are at least partially not within a common apparatus and have at least partially different housings. Therefore in some examples, the encoder 402 and the decoder 412 are at least partially within a common apparatus, and for example are located within a common housing 414. For example, the common apparatus comprising the encoder 402 and decoder 412 implements a codec. In other examples, the encoder 402 and the decoder 412 are at least partially not within a common apparatus and have at least partially different housings, but when together still implement a codec.

[0136] In some examples, 3D media from the capture (e.g., volumetric capture) at a viewpoint 416 of the scene 404, which includes a person 418) is converted via projection to a series of 2D representations with texture, occupancy, geometry, attributes and / or displacements. Additional atlasinformation is also included in the bitstream to enable inverse reconstruction. For decoding, the received bitstream 408 is separated into its components with atlas information; texture, occupancy, geometry, displacement, and attribute 2D representations. A 3D reconstruction is performed to reconstruct the scene 404-1 created looking at the viewpoint 416-1 with a “reconstructed” person 418-1. The “-1” are used to indicate that these are reconstructions of the original. As indicated at 420, the decoder 412 performs an operation(s) or action(s) based on the received signaling.

[0137] Encoding 422 generates bitstreams of different bitrates, based on the example described herein. In an example, a bitrate of a bitstream depends on an adaptation signal. Decoding 424 generates outputs of different qualities, based on the examples described herein. In an example, a quality of an output depends on an adaptation signal.

[0138] Having thus introduced a suitable but non-limiting technical context for the practice of the example embodiments of the present disclosure, example embodiments will now be described in detail.

[0139] Fundamentals of neural networks

[0140] A neural network (NN) may be described as a computation graph comprising several layers of computation. Each layer may include one or more units, where each unit performs an elementary computation. A unit is connected to one or more other units, and the connection may be associated with a weight. The weight may be used for scaling the signal passing through the associated connection. Weights are learnable parameters, e.g., values which can be learned from training data. There may be other learnable parameters, such as those of batch-normalization layers.

[0141] In some neural networks, such as convolutional neural networks for image classification, initial layers (those close to the input data) extract semantically low-level features such as edges and textures in images, whereas intermediate layers extract more high-level features. After the feature extraction layers there may be one or more layers performing a certain task, such as classification, semantic segmentation, object detection, denoising, style transfer, super-resolution, and the like.

[0142] Neural networks are being utilized in an ever-increasing number of applications for many different types of devices, such as mobile phones. Examples include image and video analysis and processing, social media data analysis, device usage data analysis, and the like.

[0143] One property of neural nets (and other machine learning tools) is that they are able to learn properties from input data, e.g., in a supervised way or in unsupervised way. Such learning is a result of a training algorithm, or of a meta-level neural network providing the training signal.

[0144] In general, the training algorithm includes changing some properties of the neural network so that its output is as close as possible to a desired output. For example, in the case of classification of objects in images, the output of the neural network can be used to derive a class or category index which indicates the class or category that the object in the input image belongs to. Training usually happens by minimizing or decreasing the output’s error, also referred to as the loss or loss function. Examples of losses are mean squared error, cross-entropy, etc. In recent deep learning techniques, training is an iterative process, where at each iteration the algorithm modifies the weights of the neural net to make a gradual improvement of the network’s output, e.g., to gradually decrease the loss, by means of gradient descent technique. In one example, at each training iteration, gradients of the loss function with respect to one or more weights or parameters of the NN are computed, for example by backpropagation technique; the computed gradients are then used by an optimization routine, such as Adam or Stochastic Gradient Descent (SGD) to obtain an update to the one or more weights or parameters.

[0145] In various embodiment, the terms “model”, “neural network”, “neural net” and “network” may be used interchangeably, and also the weights of neural networks are sometimes referred to as learnable parameters or simply as parameters.

[0146] Training a neural network is an optimization process, but the final goal may be different from the typical goal of optimization. In optimization, the only goal is to minimize a function. In machine learning, the goal of the optimization or training process is to make the model learn the properties of the data distribution from a limited training dataset. In other words, the goal is to learn to use a limited training dataset in order to learn to generalize to previously unseen data, i.e., data which was not used for training the model. This is usually referred to as generalization. In practice, data is usually split into at least two sets, the training set and the validation set. The training set is used for training the network, e.g., to modify its learnable parameters in order to minimize the loss. The validation set is used for checking the performance of the network on data, which was not used to minimize the loss, as an indication of the final performance of the model. In particular, the errors on the training set and on the validation set are monitored during the training process to understand the following things:

[0147] when the network is learning at all - in this case, the training set error should decrease, otherwise the model is in the regime of underfitting.

[0148] when the network is learning to generalize - in this case, also the validation set error needs to decrease and to be not too much higher than the training set error. When the training set error is low, but the validation set error is much higher than the training set error, or it does not decrease, or it even increases, the model may be in the regime of overfitting. This means that the model has just memorized the training set’ s properties and performs well only on that set, but performs poorly on a set not usedfor tuning its parameters.

[0149] Fundamentals of video / image coding

[0150] Video codec includes an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically, an encoder discards some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).

[0151] Typical hybrid video codecs, for example ITU-T H.263 and H.264, encode the video information in two phases. Firstly pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, e.g., the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).

[0152] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, exploits temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures (a.k.a. reference pictures).

[0153] In temporal inter prediction, the sources of prediction are previously decoded pictures in the same scalable layer. In intra block copy (IBC; a.k.a. intra-block-copy prediction), prediction may be applied similarly to temporal inter prediction but the reference picture is the current picture and only previously decoded samples may be referred in the prediction process. Inter-layer or inter-view prediction may be applied similarly to temporal inter prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal inter prediction only, while in other cases inter prediction may refer collectively to temporal inter prediction and any of intra block copy, inter-layer prediction, and inter- view prediction provided that they are performed with the same or similar process than temporal prediction. Inter prediction, temporal inter prediction, or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.

[0154] Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to becorrelated. Intra prediction can be performed in spatial or transform domain, e.g., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.

[0155] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.

[0156] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.

[0157] In typical video codecs the motion information is indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side) and the prediction source block in one of the previously coded or decoded pictures. In order to represent motion vectors efficiently those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or or co-located blocks in temporal reference picture. Moreover, typical high efficiency video codecs employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available referencepicture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures and the used motion field information is signaled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.

[0158] In typical video codecs the prediction residual after motion compensation is first transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.

[0159] Typical video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g., the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor X to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:

[0160] C = D + R

[0161] where C is the Lagrangian cost to be minimized, D is the image distortion (e.g. Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).

[0162] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI NAL units, and some video coding specifications contain both prefix SEI NAL units and suffix SEI NAL units, where the former type can start a picture unit or alike and the latter type can end a picture unit or alike. An SEI NAL unit contains one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may contain the syntax and semantics for the specified SEI messages but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret thesupplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient can be specified.

[0163] Information on neural network based image / video coding

[0164] Recently, neural networks (NNs) have been used in the context of image and video compression, by following mainly two approaches.

[0165] In an example approach, NNs are used to replace one or more of the components of a traditional codec such as a VVC / H.266-compliant codec. Here, by “traditional” we mean those codecs whose components and their parameters are typically not learned from data by means of machine learning techniques. Some examples of components that may be implemented as neural networks are, but not limited to:An in-loop filter, for example a NN that works as an additional in-loop filter with respect to the traditional loop filters, or a NN that works as the only additional in-loop filter, thus replacing any other in-loop filter.Intra-frame prediction.Inter-frame prediction.Transform and / or inverse transform.Probability model for lossless coding.

[0166] In another example approach, commonly referred to as “end-to-end learned compression” (or end-to-end learned codec), NNs are used as the main components of the image / video codecs. However, the codec may still comprise components which are not based on machine learning techniques. In this approach, two example design options are as follows:

[0167] Option 1: re-use the traditional video coding pipeline, but replace most or all the components with NNs, as shown in FIG. 5.

[0168] Referring to FIG. 5, it illustrates an example of modified video coding pipeline based on neural networks. An example of neural network may include, but is not limited, a compressed representation of a neural network. FIG. 5 is shown to include following components:

[0169] A neural transform block or circuit 502: this block or circuit transforms the output of a summation / subtraction operation 503 to a new representation of that data, which may have lower entropy and thus be more compressible.

[0170] A quantization block or circuit 504: this block or circuit quantizes an input data 501 to a smaller set of possible values.

[0171] An inverse transform and inverse quantization blocks or circuits 506. These blocks or circuits perform the inverse or approximately inverse operation of the transform and the quantization, respectively.

[0172] An encoder parameter control block or circuit 508. This block or circuit may control and optimize some or all the parameters of the encoding process, such as parameters of one or more of the encoding blocks or circuits.

[0173] An entropy coding block or circuit 510. This block or circuit may perform lossless coding, for example, based on entropy. One popular entropy coding technique is arithmetic coding.

[0174] A neural intra-codec block or circuit 512. This block or circuit may be an image compression and decompression block or circuit, which may be used to encode and decode an intra frame. An encoder 514 may be an encoder block or circuit, such as the neural encoder part of an auto-encoder neural network. A decoder 516 may be a decoder block or circuit, such as the neural decoder part of an auto-encoder neural network. An intra-coding block or circuit 518 may be a block or circuit performing some intermediate steps between encoder and decoder, such as quantization, entropy encoding, entropy decoding, and / or inverse quantization.

[0175] A deep loop filter block or circuit 520. This block or circuit performs filtering of reconstructed data, in order to enhance it.

[0176] A decode picture buffer block or circuit 522. This block or circuit is a memory buffer, keeping the decoded frame, for example, reconstructed frames 524 and enhanced reference frames 526 to be used for inter prediction.

[0177] An inter -prediction block or circuit 528. This block or circuit performs inter-frame prediction, for example, predicts from frames, for example, frames 432, which are temporally nearby. An ME / MC 430 performs motion estimation and / or motion compensation, which are two key operations to be performed when performing inter-frame prediction. ME / MC stands for motion estimation / motion compensation.

[0178] In this example (Option 1), the forward and inverse transforms were replaced with two neural networks. Also, the loop filter is a neural network.

[0179] Option 2 (also referred to as end-to-end learned coding): re-design the whole pipeline as aneural network auto-encoder with a quantization and lossless coding in the middle part, as follows:

[0180] Encoder NN (also referred to as neural network based encoder, or NN encoder): performs a non-linear transformation of the input. The output is typically referred to as latent tensor.

[0181] Quantization and lossless encoding of the encoder NN’s output.

[0182] Lossless decoding and dequantization.

[0183] Decoder NN (also referred to as neural network based decoder, or NN decoder): performs a non-linear inverse transformation from dequantized latent tensor to a reconstructed input.

[0184] It is to be understood that even in end-to-end learned approaches, there may be components which are not learned from data, such as the arithmetic codec.

[0185] More information on option 2 is provided in the following section.

[0186] Further information on neural network-based end-to-end learned coding

[0187] FIG. 6 illustrates an example neural network-based end-to-end learned codec, such as an end- to-end learned video codec. Even though some examples are provided with respect to coding images or videos, it is to be understood that other types of data may be coded in a similar way, such as audio, speech, text, features, and the like. As shown in FIG. 6, a typical neural network-based end-to-end learned coding system 600 includes an encoder 602 and a decoder 604.

[0188] The encoder comprises an encoder NN 606, a quantizer or quantization 608, a probability model 610, a lossless encoder 612 (for example, an arithmetic encoder). The decoder 604 comprises a lossless decoder 614 (for example, an arithmetic decoder), a probability model 616, a dequantizer or dequantization 618, a decoder NN 620.

[0189] It is to be noted that the probability model 610 present at encoder side and the probability model 616 present at decoder side may be same or substantially the same. For example, they may be two copies of the same probability model.

[0190] The lossless encoder 612 and the lossless decoder 614 form a lossless codec 622. The lossless codec 622 may be an entropy-based lossless codec. An example of lossless codec is an arithmetic codec, such as, a context-adaptive binary arithmetic coding (CABAC).

[0191] The encoder NN 606 and decoder NN 620 are typically two neural networks, or mainly comprise neural network components.

[0192] The probability models 610, 616 may also be neural networks and / or comprise mainly neural network components, and may be referred to as neural network based probability models or learned probability models.

[0193] Sometimes, the term lossless codec may refer to a system that comprise also the probability model, in addition to, for example, an arithmetic encoder and an arithmetic decoder.

[0194] The quantizer or quantization 608, the dequantizer or dequantization 618, and the lossless codec 622 are typically not based on neural network components, but they may also potentially comprise neural network components.

[0195] The encoder NN 606 takes an input x. which may comprise, for example, an image to be compressed. The encoder NN 606 outputs a latent tensor z. In one example, the latent tensor may be a 3D tensor, where the three dimensions of such tensor represent a channel dimension, a vertical dimension (also sometimes referred to as height dimension) and a horizontal dimension (also sometimes referred to as width dimension). In another example, the latent tensor may be a 4D tensor, where the four dimensions of such tensor represent sample dimension (also sometimes referred to as batch dimension, which is the dimension along which different samples of data may be placed), a channel dimension, a vertical dimension (also sometimes referred to as height dimension) and a horizontal dimension (also sometimes referred to as width dimension). The latent tensor is input to the quantizer or quantization 608 or a quantization operation, obtaining a quantized latent tensor zq. The quantized latent tensor is lossless-encoded into a bitstream b by the lossless encoder 612, based also on the output of the probability model 610. In particular, the probability model 610 takes as input at least part of the quantized latent tensor and outputs an estimate of a probability, an estimate of a probability distribution, or an estimate of one or more parameters of a probability distribution for one or more elements of the quantized latent tensor. The bitstream represents an encoded or compressed version of the input x.

[0196] The bitstream is lossless-decoded by the lossless decoder 614 also based on the output of the probability model 616 present at the decoder side, obtaining a quantized latent tensor zq. The quantized latent tensor is dequantized by the dequantizer or dequantization 618, obtaining a reconstructed latent tensor z. The reconstructed latent tensor is input to the decoder NN 620, obtaining a reconstructed input x, e.g., a reconstructed version of the input x. The reconstructed input may also be referred to as reconstructed data, reconstruction, decoded data, decoded input, decoded output, and the like.

[0197] The coding system 600 is a simplified description of an end-to-end learned codec, and it is to be understood that more sophisticated designs or variations of this design are possible.

[0198] The neural network components, or a subset of the neural network components, of an end-to- end learned codec may be trained by minimizing a rate-distortion loss function:

[0199] L=D+ / .R,

[0200] where D is a distortion loss term, R is a rate loss term, and X is a weight that controls the balance between the two losses.

[0201] The distortion loss term may be referred to also as reconstruction loss term, or simply reconstruction loss.

[0202] The rate loss term may be referred to simply as rate loss.

[0203] The distortion loss term measures the quality of the reconstructed or decoded output, and may comprise (but may not be limited to) one or more of the following:- Mean square error (MSE)- Structure similarity (SSIM)-MS-SSIM-Losses derived from the use of a pretrained neural network. For example, error(fl, f2), where fl and f2 are the features extracted by a pretrained neural network for the input data and the decoded data, respectively, and errorQ is an error or distance function, such as LI norm or L2 norm.- Losses derived from the use of a neural network that is trained simultaneously with the end-to- end learned codec. For example, adversarial loss can be used, which is the loss provided by a discriminator neural network that is trained adversarially with respect to the codec, following the settings proposed in the context of Generative Adversarial Networks (GANs) and their variants.- Loss that is related to a performance of one or more machine analysis tasks or to an estimated performance of one or more machine analysis tasks, where the one or more machine analysis tasks may comprise classification, object detection, image segmentation, instance segmentation, etc. In one example, the estimated performance of one or more machine analysis tasks may comprise a distortion computed based at least on a first set of features extracted from an output of the decoder and a second set of features extracted from a respective ground truth data, where the first set of features and the second set of features are output by one or more layers of a pretrained feature-extraction neural network.

[0204] Multiple distortion losses may be used and integrated into D, such as a weighted sum of MSE and SSIM.

[0205] The rate loss term may be used to train the encoder NN to output a low-entropy latent tensor, or a latent tensor such that the quantized latent tensor has low entropy, or a latent tensor such that theprobability distribution of the quantized latent tensor can be better estimated or predicted by the probability model.

[0206] The rate loss term may be used to train the probability model to better estimate or predict the probability distribution of the quantized latent tensor.

[0207] Examples of the rate loss terms are the following:- In one example, the rate loss term is derived from the output of the probability model, and it represents the estimated entropy of the quantized latent representation, which indicates the number of bits necessary to represent the quantized latent tensor.- A sparsification loss, e.g., a loss that encourages the quantized latent tensor to comprise many zeros. Examples are L0 norm, LI norm, LI norm divided by L2 norm.

[0208] In order to train the neural network components, or a subset of the neural network components, of an end-to-end learned codec, one or more of reconstruction losses may be used, and one or more rate losses may be used. In one example the one or more reconstruction losses and / or one or more rate losses are combined by means of a weighted sum. Typically, the different loss terms are weighted using different weights, and these weights determine how the final system performs in terms of rate-distortion performance. For example, when more weight is given to the reconstruction losses with respect to the rate losses, the system may learn to compress less but to reconstruct with higher accuracy (as measured by a metric that correlates with the reconstruction losses). These weights are usually considered to be hyper-parameters of the training process, and may be set manually by the person designing the training process, or automatically for example by grid search or by using additional neural networks.

[0209] In an example, the training process may be performed jointly with respect to the distortion loss D and the rate loss R. In another case, the training process may be performed in two alternating phases, where in a first phase only the distortion loss D may be used, and in a second phase only the rate loss R may be used.

[0210] For lossless video / image compression, the system may comprise only the probability model and lossless encoder and lossless decoder. The loss function would comprise only the rate loss, since the distortion loss is always zero (e.g., no loss of information).

[0211] In various embodiments, an inference phase, an inference stage, an inference time, or a test time, refers to a phase when a neural network or a codec is used for its purpose, such as encoding and decoding an input image.

[0212] Information on Video Coding for Machines (VCM)

[0213] Reducing the distortion in image and video compression is often intended to increase human perceptual quality, as humans are considered to be the end users, e.g., consuming / watching the decoded images or videos. Recently, with the advent of machine learning, especially deep learning, there is a rising number of machines (e.g., autonomous agents) that analyze data independently from humans and that may even take decisions based on the analysis results without human intervention. Examples of such analysis are object detection, scene classification, semantic segmentation, video event detection, anomaly detection, pedestrian tracking, and the like. For example, such analysis tasks may be performed by neural networks.

[0214] It is likely that the device where the analysis takes place has multiple “machines” or neural networks (NNs). These multiple machines may be used in a certain combination which, for example, is determined by an orchestrator sub-system. The multiple machines may be used, for example, in succession, based on the output of the previously used machine, and / or in parallel. For example, a video may be analyzed by one machine (NN) for detecting pedestrians, by another machine (another NN) for detecting cars, and by another machine (another NN) for estimating the depth of all the pixels in the frames.

[0215] Example use cases and applications are self-driving cars, video surveillance cameras and public safety, smart sensor networks, smart TV and smart advertisement, person re-identification, smart traffic monitoring, drones, etc. In addition to image and video data, automatic analysis and processing is increasingly being performed for other types of data, such as audio, speech, text.

[0216] Compressing (and decompressing) data where the end user comprises machines (e.g., neural networks) is commonly referred to as compression or coding for machines. In the case of video data, it is referred to as video compression or coding for machines (VCM).

[0217] Compressing for machines may differ from compressing for humans, for example, with respect to the algorithms and technology used in the codec, or the training losses used to train any neural network components of the codec, or the evaluation methodology of codecs.

[0218] It is to be understood that, when considering the case of coding for machines, the term “receiver- side” or “decoder-side” may refer to a physical entity, an abstract entity, or a device which includes one or more machines, and runs these one or more machines on some encoded and eventually decoded video representation which is encoded by another physical or abstract entity or device, the “encoder-side device”.

[0219] FIG. 7 illustrates an example pipeline 700 performing video coding for machines. A VCM encoder 704 encodes the input video 703 into a bitstream 706. A bitrate 710 may be computed 709 fromthe bitstream 706 in order to evaluate the size of the bitstream 706. A VCM decoder 712 decodes the bitstream 706 by the VCM encoder 704. The output 714 of the VCM decoder 712 is referred in FIG. 7 as “Decoded data for machines”. This data 714 may be considered as the decoded or reconstructed video. However, in some implementations of this pipeline 700, this data 714 may not have the same or similar characteristics as the original or input video 703 which was input to the VCM encoder 704. For example, this data 714 may not be easily understandable by a human by simply rendering the data onto a screen. The output 714 of VCM decoder 712 is then input to one or more task neural networks (716, 718, 720, 722). In FIG. 7, for the sake of illustrating that there may be any number of task-NNs, there are three example task-NNs, namely a task-NN 716 for object detection, a task-NN 718 for object segmentation, a task-NN 3 for object tracking, and a non-specified one (Task-NN X 722). The goal of VCM is to obtain a low bitrate while guaranteeing that the task-NNs (716, 718, 720, 722) still perform well in terms of the evaluation metric associated to each task.

[0220] As shown in FIG. 7, a performance (732) of the first task (e.g., object detection) is evaluated (724) and a performance (734) of the second task (e.g., object segmentation) is evaluated (726), a performance (736) of the third task (e.g., object tracking) is evaluated (728), and a performance (738) of the unspecified task is evaluated (730). The evaluated performances (732, 734, 736, 738) are collectively given as 740.

[0221] It is to be understood that, in some cases, the VCM decoder may not be present. In an example, the machines are run directly on the bitstream. In another example, the VCM decoder may comprise only a lossless decoding stage, and the lossless decoded data is provided as input to the machines. In yet some other cases, the VCM decoder may comprise a lossless decoding stage following by a dequantization operation, and the loss-decoded and dequantized data is provided as input to the machines.

[0222] When a conventional video encoder, such as a H.266 / V VC encoder, is used as a VCM encoder, one or more of the following approaches may be used to adapt the encoding to be suitable to machine analysis tasks:- One or more regions of interest (ROIs) may be detected. An ROI detection method may be used. For example, ROI detection may be performed using a task NN, such as an object detection NN. In some cases, ROI boundaries of a group of pictures or an intra period may be spatially overlaid and rectangular areas may be formed to cover the ROI boundaries. The detected ROIs (or rectangular areas, likewise) may be used in one or more of the following ways: o The quantization parameter (QP) may be adjusted spatially in a manner that ROIs are encoded using finer quantization step size(s) than other regions.For example, QP may be adjusted CTU-wise. o The video is preprocessed to contain only the ROIs, while the other areas are replaced by one or more constant values or removed. o The video is preprocessed so that the areas outside the ROIs are blurred or filtered. o A grid is formed in a manner that a single grid cell covers a ROI. Grid rows or grid columns that contain no ROIs are downsampled as preprocessing to encoding.-Quantization parameter of the highest temporal sublayer(s) is increased (e.g., a coarser quantization is used) when compared to practices for human watchable video.- The original video is temporally downsampled as preprocessing prior to encoding. A frame rate upsampling method may be used as postprocessing subsequent to decoding when machine analysis at the original frame rate is desired.- A filter is used to preprocess the input to the conventional encoder. The filter may be a machine learning based filter, such as a convolutional neural network.

[0223] It is to be understood that, in the context of video coding for machines, the terms “machine vision”, “machine vision task”, “machine task”, “machine analysis”, “machine analysis task”, “computer vision”, “computer vision task”, "task network" and “task” may be used interchangeably.

[0224] Also, it is to be understood that, in the context of video coding for machines, the terms “machine consumption” and “machine analysis” may be used interchangeably.

[0225] Information on neural network based filtering

[0226] A neural network may be used for filtering or processing input data. Such a neural network may be referred to as a neural network based filter, a NN filter, a filter. A NN filter may comprise one or more neural networks, and / or one or more components that may not be categorized as neural networks.

[0227] The purpose of a NN filter may comprise, (but may not be limited to, visual enhancement, colorization, upsampling, super-resolution, inpainting, temporal extrapolation, generating content, and the like.

[0228] In some video codecs, a neural network may be used as filter in the encoding and decoding loop (also referred to simply as coding loop), and it may be referred to as neural network loop filter, or neural network in-loop filter. The NN loop filter may replace all other loop filters of an existing videocodec, or may represent an additional loop filter with respect to the already present loop filters in an existing video codec.

[0229] In one example, a codec is a modified VVC / H.266 compliant codec (e.g., a VVC / H.266 compliant codec that has been modified and thus it may not be compliant to the VVC / H.266) that comprises one or more NN loop filters. An input to the one or more NN loop filters may comprise at least a reconstructed block or frames (simply referred to as reconstruction) or data derived from a reconstructed block or frame (e.g., the output of a conventional loop filter). The reconstruction may be obtained based on predicting a block or frame (e.g., by means of intra-frame prediction or inter-frame prediction) and performing residual compensation. The one or more NN loop filters may enhance the quality of at least one of their input, so that a rate-distortion loss is decreased. The rate may indicate a bitrate (estimate or real) of the encoded video. The distortion may indicate a pixel fidelity distortion such as the following:- Mean-squared error (MSE)- Mean absolute error (MAE)-Mean Average Precision (mAP) computed based on the output of a task NN (such as an object detection NN) when the input is the output of the post-processing NN.- Other machine task-related metric, for tasks such as object tracking, video activity classification, video anomaly detection, etc.

[0230] The enhancement may result into a coding gain, which may be expressed for example in terms of BD-rate or BD-PSNR.

[0231] A neural network filter may be used as post-processing filter for a codec, e.g., may be applied to an output of an image or video decoder in order to remove or reduce coding artifacts. In an example, the NN filter is used as a post-processing filter where the input comprises data that is output by or is derived from an output of a traditional decoder, such as a decoder that is compliant with the VVC / H.266 standard. In another example, the NN filter is used as a post-processing filter where the input comprises data that is output by or is derived from an output of a decoder of an end-to-end learned decoder.

[0232] Input to a NN filter

[0233] In the case of filtering images, a filter may take as input at least one or more first images to be filtered and may output at least one or more second images, where the one or more second images are the filtered version of the one or more first images. In an example, the filter takes as input one image and outputs one image. In another example, the filter takes as input more than one images and outputs one image. In another example, the filter takes as input more than one images and outputs more thanone images.

[0234] It is to be understood that a filter may also take other data as an input (also referred to as auxiliary data, or extra data) than the data that is to be filtered, for example, data that may aid the filter to perform a better filtering than when no auxiliary data is provided as input. In an example, the auxiliary data comprises information about prediction data, and / or information about the picture type, and / or information about the slice type, and / or information about a Quantization Parameter (QP) used for encoding, and / or information about boundary strength, and the like. In an example, the filter takes as input one image and other data associated to that image, such as information about the quantization parameter (QP) used for quantizing and / or dequantizing that image, and outputs one image.

[0235] Information on overfitting a neural network filter

[0236] A NN filter can be adapted at test time based at least on part of the data to be encoded and / or decoded and / or post-processed.

[0237] Although, for simplicity, the case of a NN filter is being considered herein, similar adaptation may be performed for other coding tools and / or post-processing tools that are based on neural network technology. For example, a neural network based intra-frame prediction, or a neural network based inter-frame prediction, and the like.

[0238] Such operation may be referred to, for example, with one of the following terms, when their meaning is clear from the context: adaptation, content adaptation, overfitting, finetuning, optimization, specialization, and the like.

[0239] The NN filter that results from the adaptation process may be referred to, for example, with one of the following terms: adapted filter, content-adapted filter, overfitted filter, finetuned filter, optimized filter, specialized filter, and the like.

[0240] The overfitting process may be performed at encoder side based on a training process. The resulting overfitted filter is then used to derive an overfitting signal, or adaptation signal. The adaptation signal may be compressed and then signaled from encoder to decoder, in or along a bitstream that represents encoded data, such as an encoded image or video. FIG. 8 illustrates an example of such encoder-side operations.

[0241] Referring to FIG. 8, x represents an input to the NN filter 802, x represents an output of the NN filter 802, x represents a ground-truth data associated with x, compute loss circuit / module 804 computes a training loss I in order to overfit the NN filter 802, overfit circuit / module 806 uses I tooverfit the NN filter. As a result of the overfitting process 801, an overfitted NN filter 808 is obtained, which is used by the derive overfitting circuit / module 810, together with the NN filter 802 (prior to being overfitted), to derive an adaptation signal. The adaptation signal is compressed 812 and signaled 814 to a decoder or receiver.

[0242] At decoder side, the compressed adaptation signal is received and is decompressed 902 to generate a decompressed adaptation signal. The decompressed adaptation signal, overfitting signal, or a signal derived from the overfitting signal, is used to update 904 the NN filter 906. The updated NN filter 908 is then used to filter one or more pictures, or one or more blocks. FIG. 9 illustrates an example of such decoder or receiver side operations.

[0243] The overfitted NN filter 808 that is obtained from the overfitting process at encoder side may be different from the updated NN filter 908 that is obtained from the updating process at decoder side. For example, one reason may be that the adaptation signal may be compressed in a lossy way. Thus, the former NN filter may be referred to as overfitted filter or adapted filter (or other similar terms, see above), and the latter NN filter may be referred to as updated filter.

[0244] Overfitting process performed at encoder side

[0245] The adaptation process starts with an initial NN filter.

[0246] In an example, the initial NN filter is a pretrained NN filter, which was pretrained during an offline stage on a sufficiently large dataset.

[0247] In another example, the initial NN filter is a randomly initialized NN filter.

[0248] In the adaptation, one or more parameters of the NN filter may be adapted. Examples of such parameters may include (but may not be limited to) the following:- The bias terms of a convolutional neural network.- Multiplier parameters, that multiply one or more tensors produced by the NN filter, such as one or more feature tensors that are output by respective one or more layers of the NN filter.- Parameters of the kernels of a convolutional neural network.- Parameters of an adapter layer.- One or more arrays or tensors that are used as input to respective one or more layers of the NN filter.

[0249] The adaptation may be performed by means of a training process, e.g., by minimizing a loss function until a stopping criterion is met. The data used for this training process may comprise one ormore pictures or blocks of input to the NN filter and associated respective one or more pictures or blocks of ground-truth data. In an example, where the filter is an in-loop filter, the input to the NN filter is reconstruction data, after prediction and residual compensation; the ground-truth data is the uncompressed data that is given as input to the encoder. In another example, where the filter is a postprocessing filter, the input to the NN filter is decoded data (e.g., the output of a video decoder); the ground-truth data is the uncompressed data that is given as input to the encoder.

[0250] The loss function used during the training process may comprise one or more distortion loss functions (also referred to as reconstruction loss functions) and zero or more rate loss functions. A rate loss function may measure, for example, the cost in terms of bitrate of signaling any adaptation signal, such as updates to the parameters of the NN filter. A distortion loss function may comprise one of MSE, MS-SSIM, VMAF, and the like.

[0251] Deriving the adaptation signal

[0252] The adaptation signal may be derived based on the adapted NN filter and on the original NN filter (e.g., the NN filter before the overfitting process).

[0253] In an example, the adaptation signal comprises an update to one or more parameters of the NN filter. Such an update may also be referred to as weight update, or parameter update. Such update may be computed, for example, by subtracting the values of the adapted parameters (e.g., the parameters of the adapted NN filter) from the corresponding values of the original parameters (e.g., the parameters of the original NN filter).

[0254] In another example, the adaptation signal comprises the parameters (of the NN filter) that were adapted, also referred to as updated parameters, or adapted parameters, or adapted weights, or overfitted parameters, and the like.

[0255] Compression of adaptation signal

[0256] In order to keep the size of the adaptation signal low, the adaptation signal may go through one or more compression steps, such as sparsification, quantization and lossless coding.

[0257] In one example, an encoder that compresses the adaptation signal into a bitstream that is compliant with a neural network compression standard, such as MPEG NNC, may be used.

[0258] Signaling

[0259] The compressed adaptation signal may be signaled from encoder to decoder in or along abitstream that represents encoded image or video data.

[0260] In one example, the compressed adaptation signal is signaled in an Adaptation Parameter Set (APS) syntax structure of a video coding bitstream.

[0261] In another example, the compressed adaptation signal is signaled in a Supplemental Enhancement Information (SEI) message of a video coding bitstream.

[0262] Signaling may comprise also other information which is associated with the adaptation signal and that may be required for correctly parsing and / or decompressing and / or using the adaptation signal, such as any quantization parameters.

[0263] Decoder or receiver side operations

[0264] At decoder side, the signaled compressed adaptation signal is received and decompressed. The decompressed adaptation signal may then be used to update the NN filter.

[0265] In one example, where the adaptation signal comprises a weight update, where the weight update comprises one or more updates to respective one or more parameters of the NN filter, the one or more updates are added to the one or more parameters.

[0266] In another example, where the adaptation signal comprises one or more updated or adapted parameters, the one or more updated or adapted parameters are used to replace respective one or more parameters of the NN filter.

[0267] Once the NN filter has been updated based on the adaptation signal, the updated NN filter may be used for its purpose. For example, for filtering an input picture or an input block.

[0268] Terminology

[0269] In various embodiments, terms frame, picture, and image may be used interchangeably.

[0270] For example, the input and output to an end-to-end learned codec may be pictures. The input and output of a NN filter may be pictures.

[0271] It is to be understood that also the term block, when it refers to a portion of a picture, may be simply referred to as a frame, or a picture, or an image. In other words, at least some of the embodiments herein, even when described as applied to a picture, may be applicable also to a block, e.g., to a portion of a picture.

[0272] A neural network that is present at decoder side, e.g., a decoder-side neural network (DSNN), is usually trained on a large corpus of data. However, at test time, data to be encoded and / or decoded may have peculiar characteristics that may be slightly or significantly different from the main or average characteristics of data used for training. Consequently, the DSNN may not perform optimally.

[0273] General information

[0274] In various embodiments, a neural network that is present at decoder side is referred to as decoder-side neural network (DSNN). The neural network performs a processing of an input data item (also referred to simply as data item), where the input data item (or a portion of the input data item) is mapped to an output data item. The neural network may be used as part of a decoder process and / or as part of a post-processing process that is performed on data derived from an output of the decoder process. In other words, the neural network may be comprised in a decoder process and / or in a postprocessing process. The decoder process and / or the post-processing process may be considered to be comprised of or to be part of the decoder side.

[0275] In an example, the DSNN may be a neural network (NN) loop filter.

[0276] In another example, the DSNN may be a NN post -processing filter.

[0277] In yet another example, the DSNN may be a NN that is used as part of an intra-frame prediction process.

[0278] In yet another example, the DSNN may be a NN that is used as part of an inter-frame prediction process.

[0279] In yet another example, the DSNN may be a NN that is used as part of a transformation process, such as a transform of a prediction residual signal.

[0280] In yet another example, the DSNN may be a NN that is used as part of an end-to-end learned codec, such as a NN that takes a lossless-decoded latent tensor and outputs reconstructed or decoded data.

[0281] In yet another example, the DSNN may be a NN that performs spatial upsampling.

[0282] In yet another example, the DSNN may be a NN that performs super-resolution.

[0283] In yet another example, the DSNN may be a NN that performs temporal upsampling.

[0284] For the sake of simplicity, at least some embodiments or examples are described as applied toa NN filter. However, it is to be understood that the at least some embodiments or examples may be valid or applicable to other DSNNs than a NN filter.

[0285] Also, for the sake of simplicity, at least some embodiments or examples are described as applied to a DSNN that is used for decoding a data item or is part of an encoding or decoding process that encodes or decodes a data item. However, it is to be understood that the at least some embodiments or examples may be valid or applicable to a DSNN that is used for post-processing a decoded data item or data derived from a decoded data item.

[0286] While at least some embodiments are described such that the input and output data are in the form of images or (video) frames or pictures, those embodiments may be applicable also to other types of data, such as audio frames. Furthermore, while at least some embodiments are described by considering a full image, those embodiments may be applicable also to one or more blocks or portions of an image.

[0287] It is to be noted that a DSNN may be available also at encoder side. A DSNN at encoder side may be the same or substantially the same as a DSNN at decoder side. For example, a DSNN at encoder side may have same architecture and values of its parameters as a DSNN at decoder side. Thus, a DSNN may be a neural network that is part of an encoding process and / or a decoding process and / or a postprocessing process. In other words, the DSNN may be comprised in an encoding process and / or in a decoding process and / or in a post-processing process.

[0288] It is to be noted that in some cases an encoder comprises a decoder or a subset of components of a decoder. An encoder may be comprised in an encoder side. The encoder side may comprise one or more parts or processes that are same or substantially same as respective one or more parts or processes of a decoder side, such as the decoder process and / or a subset of the decoder process and / or a postprocessing process.

[0289] It is to be noted that some embodiments may use the term “data item” or “data unit” to refer to data that is input to the DSNN, or to data from which an input to the DSNN is derived (such as data that is input to an encoder or to a decoder), or to data that is derived from an output of the DSNN, and its meaning may be understood from the context.

[0290] Example embodiment

[0291] In an embodiment, a decoder or receiver comprises a neural network that is referred to also as decoder-side neural network (DSNN). For example, the DSNN is a NN loop filter, or NN filter for short. The DSNN is assumed to have been trained during an offline or development stage, for example, whenthe codec or decoder is developed. The decoder may comprise one or more adapters, where each of the one or more adapters comprises one or more adaptation parameters. The one or more adaptation parameters of each of the one or more adapters may have been pretrained during an offline or development stage. The one or more adaptation parameters of each of the one or more adapters may be associated with respective one or more positions within the DSNN. The one or more adaptation parameters of each of the one or more adapters may be grouped into one or more adaptation layers, where the one or more adaptation layers may be associated with respective one or more layer positions within the DSNN. At decoding time, e.g., when a data item such as a block or an image or a video frame is to be decoded, the decoder may select at least one adapter based at least on a selection criterion. The selected at least one adapter (also referred to simply as the at least one adapter) may be used to adapt the DSNN based at least on an adaptation process. The adaptation process may comprise modifying the DSNN or one or more external or internal input of the DSNN by using the at least one adapter or the adaptation parameters in the at least one adapter.

[0292] In an embodiment, the adaptation parameters are used as parameters or weights of the DSNN, e.g., similar to parameters of a convolutional layer. In an example, the adaptation parameters are used as bias parameters. In another example, the adaptation parameters are used as multiplier parameters. In yet another example, the adaptation parameters are used as parameters of one or more convolutional layers.

[0293] In an alternative embodiment, the adaptation parameters are used as an external or internal input of the DSNN. In an example, the adaptation parameters are used as an input to the DSNN. In another example, the adaptation parameters are used as an input to one or more layers of the DSNN. When the adaptation parameters are used as an external or internal input of the DSNN, the adaptation process may need information about how the adaptation parameters are input, such as by concatenation with other one or more inputs, by summation, by multiplication, and the like. Such information may be signaled from an encoder or may be predetermined and known at decoder side.

[0294] In an embodiment, the adaptation process comprises inserting the one or more adaptation parameters of the at least one adapter into the respective one or more associated positions within the DSNN.

[0295] In an embodiment, the adaptation process comprises replacing the values of some parameters of the DSNN with the values of the adaptation parameters of the at least one adapter.

[0296] In an embodiment, the adaptation process comprises modifying (e.g., by addition or by multiplication) the values of some parameters of the DSNN based on the values of the adaptationparameters of the at least one adapter.

[0297] In an embodiment, the adaptation process comprises inserting the one or more adaptation layers of the at least one adapter into the respective one or more associated layer positions within the DSNN.

[0298] In an embodiment, the one or more adapters are associated with respective one or more identifiers or with respective one or more nominal or numerical values of a feature or characteristic, such as a characteristic of an input to the DSNN or a characteristic of the DSNN or a characteristic of the encoding or decoding process. In an example, the one or more identifiers are indexes of a look-up table. In another example, the one or more identifiers are derived from positions of the DSNN within a certain pipeline or cascade of operations. In yet another example, the one or more identifiers are derived from content types. In yet another example, the one or more identifiers are derived from quantization parameters (QPs). In yet another example, the one or more identifiers are derived from picture types. In yet another example, the one or more identifiers are derived from resolutions or resolution ranges. In yet another example, the one or more identifiers are derived from temporal layer identifiers. In yet another example, the one or more identifiers are derived from loss functions used to train the one or more adapters. In yet another example, the one or more identifiers are derived from types of scanning of a transformed quantized residual signal (e.g., z-scan, zig-zag scan, diagonal scan, horizontal scan, vertical scan, and the like). In yet another example, the one or more identifiers are derived from frame rates.

[0299] FIG. 10 describes an illustration of example embodiments. In FIG. 10, an NN filter is adapted based on a selected adapter. A decoder comprises the NN filter and N adapters (e.g., adapter 1004a, adapter 1004b, ..., adapter 1004n). One of the adapters (for example, adapter 1004b) is selected by means of a select process 1006 and a selection criterion. The selected adapter (for example, adapter 1004b) is used to adapt the NN filter by using an adapt process 1008, obtaining an adapted NN filter 1010. The adapted NN filter 1010 is used to filter an input x, obtaining an output x.

[0300] Example

[0301] In this example, the decoder comprises a DSNN and 4 pretrained adapters. The DSNN is a NN loop filter that comprises 10 convolutional layers. Each adapter comprises 10 multiplier layers associated to the 10 convolutional layers of the DSNN, where the layer position associated to a multiplier layer is after a convolutional layer, and where a i-th multiplier layer comprises C[i] multiplier parameters, where C[i] is the number of channels of the output tensor from the convolutional layer associated to the i-th multiplier layer. Adapting the DSNN comprises selecting one of the 4 pretrainedadapters, then inserting the 10 multiplier layers comprised in the selected adapter into their respective layer positions (e.g., each multiplier layer is inserted after its associated convolutional layer), obtaining an adapted DSNN. The operation performed by a convolutional layer followed by a multiplier layer may be described mathematically as follows:

[0302] out = (conv(input) + biases) * m

[0303] where input stands for an input to a convolutional layer, conv stands for a convolution operation performed by a convolutional layer (excluding any addition of biases), biases stand for an array of bias parameters, m stands for an array of multiplier parameters, out stands for the output of the operation of the convolutional layer followed by the multiplier layer.

[0304] The adapted DSNN is used as a NN loop filter (e.g., within a chain of loop filters) to filter a reconstructed block or frame.

[0305] Embodiments on frequency of adaptation

[0306] In an embodiment, the DSNN may be adapted on a video sequence basis, on a coded layer video sequence (CLVS) basis, on a random-access (RA) segment basis, on an intra-frame period basis, on a frame or picture basis, and / or on a block basis (e.g., on a CTU basis).

[0307] In an example, the DSNN is adapted for each random-access segment of a video sequence. E.g., for each RA segment, a certain adapter from the one or more adapters is selected and used in the inference of the DSNN when decoding that RA segment.

[0308] In an example, there may be multiple (at least two) sets of adapters labeled with identifiers and an identifier can be signaled for each data unit (depending on the frequency of adaptation, a data unit may be a CTU or a slice or data associated to a temporal layer id, or data associated to a picture type, or the like) to indicate which set of adapters is associated with the current data unit. Alternatively, a map may be signaled for a particular data unit (e.g., for RA segment) where the map comprises two or more identifiers of respective two or more adapters to be used for decoding respective two or more sub-units (e.g., two or more temporal layers within the RA segment).

[0309] Embodiments on combination of adapters

[0310] In an embodiment, two or more adapters are grouped into two or more categories, where a category may comprise one or more adapters that are similar with respect to a similarity metric or with respect to a certain feature. One or more first adapters in any first category of the two or more categories are different from one or more second adapters in any second category of the two or more categorieswith respect to a similarity metric or with respect to a certain feature. In an example, a first category comprises one or more first adapters that are associated with respective one or more quantization parameters, and a second category comprises one or more second adapters that are associated with respective one or more content types.

[0311] In an embodiment, two or more adapters may be combined by means of a combination operation for obtaining a combined adapter that is then used for processing a data item. Any two of the two or more adapters may belong to a same category or to different categories.

[0312] In an embodiment, the two or more adapters that are combined are selected based on a similarity between a data item to be processed and the two or more adapters, in terms of a similarity metric or similarity criterion.

[0313] In an embodiment, the two or more adapters that are combined are selected based on a similarity between a characteristic of a data item to be processed and characteristics of the two or more adapters, in terms of a similarity metric or similarity criterion.

[0314] In an example, a similarity criterion comprises a similarity between a value of a particular feature of the data item to be processed (e.g., a value of a quantization parameter used to encode or decode the data item) and values of that particular feature that are associated with the two or more adapters to be combined (e.g., values of a quantization parameter that are associated with the two or more adapters to be combined).

[0315] In an embodiment, the combination operation comprises averaging.

[0316] In another embodiment, the combination operation comprises weighted averaging.

[0317] In the embodiment, where the combination operation comprises weighted averaging, weights or coefficients used in the weighted averaging may be determined based on a similarity between a data item to be processed and the two or more adapters, in terms of a similarity metric or similarity criterion (e.g., the similarity criterion described above), or based on a similarity between a characteristic of a data item to be processed and characteristics of the two or more adapters, in terms of a similarity metric or similarity criterion (e.g., the similarity criterion described above).

[0318] In an example, a first adapter and a second adapter belong to a category that comprises adapters that are associated with quantization parameters. The first adapter and the second adapter are selected based on the similarity between the quantization parameter (QP) of the data item to be processed and the QPs associated to the adapters in that category. For example, the first adapter and thesecond adapter are the adapters associated with QPs that are most similar to the QP of the data item to be processed, among all the adapters belonging to the same category of the first adapter and second adapter. In this example, the similarity metric may comprise or be derived from an algebraic difference between two QP values, where the similarity is higher when the difference is smaller. One or more adaptation parameters in the first adapter are combined with respective one or more adaptation parameters in the second adapter by means of a weighted averaging operation. Weights of the weighted averaging operation are selected based on the QP of the data item to be processed and the QPs associated with the first adapter and second adapter. For example, when a quantization parameter of a data item to be processed is 39, two adapters that are associated to quantization parameters 37 and 42 may be combined. The weights are determined based on the difference between the associated QPs and the QP of the data item. For example, a bigger weight is determined for the adapter associated with QP 37 and a smaller weight is determined for the adapter associated with QP 42.

[0319] In an embodiment, the combination operation may comprise using the two or more adapters by cascading adaptation parameters of different adapters in the two or more adapters. In an example, where each of the two or more adapters comprise a multiplier layer associated to a particular layer of a DSNN, a multiplier layer of a first adapter is cascaded to a multiplier layer of a second adapter. In this example, the first adapter and the second adapter are comprised in the two or more adapters.

[0320] Embodiments on selection criterion

[0321] In an embodiment, the selection criterion (or the select process that is based on the selection criterion) comprises selecting one or more adapters based on a received indication of which of the one or more adapters is to be selected or used. The indication may be sent from an encoder to the decoder, in or along a bitstream representing an encoded data, for example, in an adaptation parameter set (APS) or in a supplemental enhancement information (SEI) message.

[0322] In an embodiment, the indication may be associated with information about the scope of the indicated adapter. The scope may comprise, for example, one or more data items to be processed by a neural network that is adapted based on the indicated adapter.

[0323] In an example, an encoder may include an indicator in an APS, where the indicator indicates which of the one or more adapters is to be used for decoding one or more frames that are in the scope of the APS. The scope of the APS may be the frame associated with the APS, or a scope that is explicitly indicated within an APS. The decoder would then select the adapter indicated by the received indicator for decoding the one or more frames in the scope of the APS.

[0324] In an embodiment, the indication is indicative of a position of the DSNN within a pipeline,such as the position of a NN loop filter within a pipeline or chain or cascade of loop filters, where the position of the DSNN is also associated to a particular adapter of the one or more adapters.

[0325] In an embodiment, the indication is indicative of a content type, where a content type is associated to a particular adapter of the one or more adapters.

[0326] In an embodiment, the indication is indicative of a loss function that is associated to a particular adapter of the one or more adapters.

[0327] In another embodiment, the selection criterion (or the select process that is based on the selection criterion) comprises selecting one or more adapters based on one or more parameters or indicators that are available at decoder side and that are also used for some other operations of the decoder, in addition to being used for selecting an adapter. For example, in this embodiment, the one or more parameters or indicators may not have been explicitly signaled for the purpose of selecting one or more adapters from the one or more adapters, but they can be reused for such a purpose.

[0328] In an embodiment, the one or more parameters or indicators comprise a quantization parameter (QP) associated to a picture to be processed by means of the DSNN.

[0329] In an embodiment, the one or more parameters or indicators comprise a picture type or block type (e.g., intra-predicted picture or block, or inter-predicted picture or block) of the picture or block to be processed by means of the DSNN.

[0330] In an embodiment, the one or more parameters or indicators comprise a resolution information, e.g., the spatial resolution of the video frame(s) to be processed by means of the DSNN.

[0331] In an embodiment, the one or more parameters or indicators comprise a temporal layer identifier, e.g., an identifier of the temporal layer of an inter-frame prediction hierarchy structure.

[0332] In an embodiment, the one or more parameters or indicators comprise a position of the DSNN within a pipeline, such as a position of a NN loop filter within the pipeline of loop filters in a video decoder.

[0333] In an embodiment, the one or more parameters or indicators comprise a type of content to be decoded by means of the DSNN.

[0334] In an embodiment, an encoder may signal a correction or adjustment for the one or more parameters or indicators that are available at decoder side and that are used for some other operations of the decoder. The signaled correction or adjustment may be used by the decoder for correcting oradjusting the one or more parameters or indicators, obtaining respective one or more corrected or adjusted parameters or indicators which may be more suitable or optimal to be used for selecting respective one or more adapters. For example, when a particular adapter to be used to decode a particular picture is selected based on the QP (e.g., 42) used to code that particular picture, an encoder may signal an adjustment to that QP (e.g., -2), which would result into an adjusted QP (e.g., 42-2=40) that is more optimal for selecting the adapter.

[0335] Embodiments on overfitting

[0336] In an embodiment, one or more selected adapters may be adapted based at least on an adaptation signal that is signaled from an encoder to the decoder, obtaining respective one or more updated adapters.

[0337] In an example, an encoder may select a particular adapter and overfit one or more adaptation parameters comprised in the selected adapter, where the selected adapter is assumed to have been trained or pretrained during an offline or development phase. An adaptation signal may be derived based on the overfitted adapter, e.g., the adaptation signal may be derived from a difference between the parameters of the overfitted adapter and the parameters of the selected pretrained adapter, where such a difference may be referred to as weight-update. The weight-update may be compressed (in a lossy and / or lossless way) and signaled to the decoder, together with an indication of which adapter it refers to and with an indication of which data it is associated with. At decoder side, the compressed weightupdate is decompressed and used to update or adapt the indicated adapter, obtaining an updated adapter. The updated adapter is then used in the DSNN for decoding data associated to the weight-update.

[0338] Embodiments on integrating the adapters into the DSNN

[0339] In an embodiment, the one or more adaptation parameters of the one or more adaptation layers of the one or more selected adapters are integrated into the already present parameters of the DSNN, based at least on an integration method.

[0340] In an example, a particular adapter is selected by a selection criterion. The particular adapter comprises one or more multiplier layers. In an example, the one or more multiplier layers comprises one or more multiplier parameters, the one or more multiplier layers are associated to respective one or more convolutional layers of the DSNN, and the one or more multiplier parameters of a multiplier layer are associated to respective one or more kernels of a convolutional layer associated with the multiplier layer. One or more multiplier parameters of an i-th multiplier layer are integrated into the parameters of respective one or more convolutional kernels of the i-th convolutional layer, based at least on an integration method.

[0341] For example, an integration method may comprise multiplying a j-th multiplier parameter of the i-th multiplier layer with all parameters of a j-th kernel of the i-th convolutional layer.

[0342] Embodiments on initializing the adaptation parameters

[0343] In an embodiment, at the beginning of training the adapters, the adaptation parameters comprised in the adapters are initialized based on a predetermined value. In an example, when the adaptation parameters are multiplier parameters, they may be initialized to a value equal to 1. In another example, when the adaptation parameters are additive parameters (e.g., bias parameters), they may be initialized to a value equal to 0.

[0344] In another embodiment, the adaptation parameters comprised in the adapters are initialized based on values that are sampled from a probability distribution, for example, by using a random or pseudo-random number generator.

[0345] In another embodiment, the adaptation parameters comprised in the adapters are initialized based at least on a decomposition operation that is performed on the DSNN’s parameters associated with the adaptation parameters. The decomposition operation may result into two components, where a first component is used to initialize the adaptation parameters and a second component is used as the new or updated values of the DSNN’s parameters associated with the adaptation parameters.

[0346] In an example, an adapter comprises one or more multiplier layers. In an example, the one or more multiplier layers comprises one or more multiplier parameters, the one or more multiplier layers are associated to respective one or more convolutional layers of a DSNN, and the one or more multiplier parameters in a i-th multiplier layer are associated to respective one or more kernels of a i-th convolutional layer of the DSNN. For example, a j-th multiplier parameter in the i-th multiplier layer is associated to a j-th kernel of the i-th convolutional layer. A matrix of parameters of the j-th kernel of the i-th convolutional layer is decomposed into two components that may be referred to a j-th magnitude component and a j-th directional component. In an example, the j-th magnitude component may comprise a j-th magnitude value and the j-th directional component may comprise a j-th matrix of directional values. The j-th multiplier parameter in the i-th multiplier layer is initialized to a value equal to or derived from the j-th magnitude value of the j-th kernel of the i-th convolutional layer. The magnitude component and the directional component may be computed, for example, as follows: the magnitude component is computed as the value of a norm of the kernel’s matrix, such as the Frobenius norm; the directional component may be computed by dividing the kernel’s matrix by the value of the norm. The matrix representing the directional component may be used as a new or updated kernel’s matrix, and the initialized multiplier parameters will be trained by using the DSNN comprising the newor updated kernel’s matrices multiplied by the multiplier parameters.

[0347] Embodiments on QP-wise adapters - Inference

[0348] In an embodiment, a decoder comprises N adapters that are associated with N Quantization Parameters (QPs) or with N QP ranges, where QP refers to the quantization parameter used to encode a data item such as a video sequence or a frame or a block. Such adapters may be referred to as QP- wise adapters or as QP-based adapters.

[0349] It is to be understood that QP may refer to one of the following: a sequence QP (e.g., a QP used to code a CLVS), a picture QP (e.g., a QP used to code a picture), a slice QP (e.g., a QP used to code a slice of a picture), or a block QP (e.g., a QP used to code a CTU).

[0350] When a data item (e.g., a CLV S, or a picture) has been coded with a particular QP, information about the particular QP is available at decoder side. The decoder may select the adapter that is associated with a QP that is equal to the particular QP, that is associated with a QP that is most similar to the particular QP, or that is associated with a QP range that comprises the particular QP.

[0351] In an example, a decoder comprises 4 adapters that are associated with the following QP ranges:

[0352] Adapter 1: QPs in the range [22, 26].

[0353] Adapter 2: QPs in the range [27, 31].

[0354] Adapter 3: QPs in the range [32, 36].

[0355] Adapter 4: QPs in the range [37, 42] .

[0356] When a CLVS has been coded with a QP equal to 33, the decoder selects the Adapter 3 for adapting the DSNN that will be used for decoding the CLVS.

[0357] Embodiments on QP-wise adapters - Training

[0358] During training of QP-based adapters, a number N of adapters is chosen, such as N=4. Then, the training dataset may be divided into N portions or subsets, where each portion or subset comprises data associated (e.g., coded) with a particular QP or QP range. A i-th adapter is trained by using training data in the i-th portion or subset.

[0359] Embodiments on position-wise adapters - Inference

[0360] In an embodiment, a decoder comprises N adapters that are associated with N possible positions of the DSNN within a pipeline or cascade of processing steps. Such adapters may be referred to as position-wise adapters or position-based adapters. For example, when the DSNN is a NN loop filter, the N adapters are associated with N possible positions of the NN loop filter within the loop filter chain or pipeline. In an example, the loop filter chain or pipeline may comprise a deblocking filter (DBF), a sample-adaptive offset (SAO) filter, an adaptive loop filter (ALF); so N may be equal to 4 because of the following 4 possible positions of the NN loop filter: before DBF, between DBF and SAO, between SAO and ALF, after ALF. It is to be understood that other possible positions may comprise positions that are in parallel with one or more processing steps in the pipeline or cascade of processing steps, where an output of the DSNN may be combined with one or more outputs of the one or more processing steps. In one example, a position of a DSNN is in parallel with a DBF, and an output of the DSNN is combined with an output of the DBF by means of weighted average.

[0361] An encoder may signal an indication of the position of the DSNN to be used when decoding a particular data item. In an example, the indication of the position may be used also for selecting the adapter that is associated with that position.

[0362] In an example, a decoder comprises 4 adapters for a NN loop filter, where the 4 adapters are associated with respective 4 positions within the chain of loop filters. An encoder signals to the decoder, for example, as part of an APS or as part of a picture header, an indication of a position of the NN loop filter to be used for decoding a picture associated with the APS or the picture header. The decoder selects the adapter that is associated with the position indicated by the encoder, uses the selected adapter for adapting the NN loop filter and uses the adapted NN loop filter in the position indicated by the encoder for decoding the picture associated with the APS or the picture header.

[0363] Embodiments on position-wise adapters - Training

[0364] In an embodiment, N position-based adapters are trained based on N versions of a training dataset, e.g., each position-based adapter is trained by using a different version of the training dataset. A training dataset usually includes a number of pairs of data items, where a pair comprises an input data item and a ground-truth data item. The input data item is meant to be input to the DSNN and the groundtruth data item is meant to be used as ground-truth for computing a loss function. In the case of video data, the input data items are usually generated by encoding a set of videos with a video codec and saving the data in the position where the DSNN is expected to be used, e.g., saving the data that is output by a previous processing step with respect to the DSNN, or saving the data at the position of the input of the DSNN. The N versions of the training dataset are different with respect to the position where the input data items are generated or saved. For example, a first version of the training datasetcomprises input data items that are saved at the position before the DBF (e.g., data items that are input to the DBF); a second version of the training dataset comprises input data items that are saved at the position between the DBF and the SAO filter (e.g., data items that are output by the DBF); a third version of the training dataset comprises input data items that are saved at the position between the SAO filter and the ALF (e.g., data items that are output by the SAO); a fourth version of the training dataset comprises input data items that are saved at the position after the ALF (e.g., data items that are output by the ALF).

[0365] Embodiments on content- wise adapters - Inference

[0366] In an embodiment, the decoder comprises N adapters which are associated with N types of content. Such adapters may be referred to as content-wise adapters or content-based adapters. Herein, a type of content (or content type) may comprise a particular set of characteristics, and different types of content (or content types) may comprise different sets of characteristics; content or data in a certain content type may share same or similar characteristics, whereas content or data in different content types may not share same or similar characteristics.

[0367] An encoder may signal to the decoder an indication of the content type associated with a particular data unit such as a video sequence, a CLVS, a RA segment, a picture, or a block. The decoder may select the adapter that is associated with the same content type as indicated by the encoder, use that adapter for adapting the DSNN, and use the DSNN for decoding the particular data unit.

[0368] In an embodiment, at inference time, an encoder may perform two or more inferences of a DSNN by using respective two or more of the N adapters. In an example, the encoder may perform N inferences of the DSNN by using respective N adapters. Based on an evaluation of the performance or quality of the output of the two or more inferences (e.g., based on a comparison of two or more PSNR gains, or based on a comparison of two or more rate-distortion costs), the encoder may determine an optimal adapter. The encoder may signal to the decoder an indication of the optimal adapter, for example, in the form of an index, in association with an indication of one or more data items to be processed by using a neural network that is adapted based on the indicated optimal adapter. The decoder may use the received indication of the optimal adapter to select and use the indicated optimal adapter for decoding the indicated one or more data items.

[0369] Embodiments on content-wise adapters - Training

[0370] Training N content-based adapters based on grouping the adapters

[0371] In an embodiment, N content-based adapters are obtained as follows. A training dataset isdivided into a set of small data items. In an example, the definition of “small” may depend on the nature and modality of the data. In the case of visual data (image, video), a data item is a patch or block of an image or video frame; a small data item may be a patch or block of size less or equal to 144x144 pixels. It is to be understood that a training dataset usually comprises pairs of input data item and associated ground-truth data item. Accordingly, a small data item may refer either to an input data item or to a pair of input data item and associated ground-truth data item, depending on the context. For each small data item (also referred to as data item for simplicity), a temporary adapter is trained or overfitted by using that small data item as training data. As a result, there may be as many temporary adapters as there are small data items in the training dataset. Subsequently, the temporary adapters may be grouped into a smaller number N of adapters, for example, by means of clustering. The N adapters obtained from such grouping represent the final N adapters that are included into the decoder.

[0372] Training N content-based adapters based on grouping data items in the training dataset

[0373] In another embodiment, N content-based adapters are obtained as follows. A training dataset is divided into a set of small data items. The small data items are grouped into N groups based on a grouping algorithm such as clustering or classification. The N groups represent N subsets of training data. For each subset of training data, a separate adapter is trained on the data items included in that subset, obtaining N adapters that are included into the decoder.

[0374] Training N content-based adapters based on weighted loss functions

[0375] In another embodiment, N content-based adapters are obtained via the following training process. N neural networks are trained jointly or substantially jointly. At the beginning of the training process, each of the N neural networks comprises the pretrained parameters and an adapter. The pretrained parameters of the N neural networks are the same or substantially the same (e.g., they have same values across different neural networks). An adapter in one of the N neural networks may be initialized to be same or substantially same as an adapter in another of the N neural networks (e.g., all N adapters have adaptation parameters with value equal to 1), or may be initialized to be different than an adapter in another of the N neural networks (e.g., the adapters or the adaptation parameters are initialized by means of a random sampling process). One or more iterations may be performed as part of the training process.

[0376] At each iteration of the training process, an input sample (or input data item) is provided as input to N neural networks. The N neural networks are run or executed (e.g., a forward pass is performed), obtaining respective N outputs. The N outputs and associated N ground-truth data items are used to compute respective N initial loss functions. The N initial loss functions are used to computeN weights or coefficients, where each of the N weights or coefficients may be computed by using two or more of the N initial loss functions (in one example, each of the N weights or coefficients are computed by using all the N initial loss functions). The N weights or coefficients are used to compute respective N final loss functions. The N final loss functions are used to compute respective N sets of gradients of the N final loss functions with respect to the parameters of the respective N neural networks. The N sets of gradients are used for updating the adapters or the adaptation parameters of the respective N neural networks.

[0377] As a result of this training process, N adapters are trained.

[0378] Embodiments on picture-type-wise adapters - Inference

[0379] In an embodiment, a decoder comprises N adapters that are associated with respective N picture types. Such adapters may be referred to as picture- type- wise adapters or picture- type-based adapters. A picture type may comprise, for example, an intra-coded picture, or an inter-coded picture.

[0380] It is to be understood that, while some embodiments are described in terms of picture type, the same embodiments may be valid or applicable also to block type, or anyway type of another data structure or data unit than a picture, such as a slice, a CTU, a group of pictures, and the like.

[0381] A decoder may comprise information about a picture type associated to a picture to be decoded. Such information may be signaled from an encoder, or may be derived at decoder side.

[0382] When the decoder decodes a picture of a particular picture type, the decoder may select the adapter that is associated with that particular picture type, or that is associated with a picture type that is most similar to that particular picture type (e.g., according to a predefined similarity metric or evaluation process).

[0383] Embodiments on picture-type-wise adapters - Training

[0384] N picture-type based adapters are trained by means of the following training process.

[0385] At the beginning of training, a number N of adapters is chosen, for example, based on an expected number of picture types. In an example, N is chosen to be equal to 2, and the two picture types are intra-coded picture type and inter-coded picture type. Then, the training dataset may be divided into N portions or subsets, where each portion or subset comprises data associated (e.g., coded) with a particular picture type. In an example, the training dataset is divided into two subsets, where a first subset comprises intra-coded pictures and a second subset comprises inter-coded pictures. Thereafter, N training sessions are performed, where respective N adapters are trained on respective N subsets oftraining data. In a i-th training session, a DSNN comprising a i-th adapter (e.g., adaptation parameters comprised in a i-th adapter) is used for training the i-th adapter based on the i-th subset of training data.

[0386] Embodiments on resolution- wise adapters - Inference

[0387] In an embodiment, a decoder comprises N adapters that are associated with respective N resolutions or N resolution-ranges. Such adapters may be referred to as resolution- wise adapters or resolution-based adapters.

[0388] When a data item (e.g., a CLVS, or a picture) has been coded with a particular resolution, information about that particular resolution may be available at decoder side. For example, the information about the particular resolution may be signaled from an encoder to the decoder, or may be derived at decoder side based on other information. The decoder may select the adapter that is associated with a resolution that is equal to the particular resolution, that is associated with a resolution that is most similar to the particular resolution, or that is associated with a resolution range that comprises the particular resolution. The selected adapter is then used to adapt the DSNN, obtaining an adapted DSNN that may be used for decoding the data item.

[0389] Embodiments on resolution-wise adapters - Training

[0390] During training of resolution-based adapters, a number N of adapters is chosen, such as N=4. Then, the training dataset may be divided into N portions or subsets, where each portion or subset comprises data associated (e.g., coded) with a particular resolution or resolution range. A i-th adapter is trained by using training data comprised in the i-th portion or subset.

[0391] Embodiments on temporal-layer- wise adapters - Inference

[0392] In one embodiment, a decoder comprises N adapters that are associated with respective N temporal layers of a predetermined hierarchy for inter-frame prediction patterns. Such adapters may be referred to as temporal-layer-wise adapters or temporal-layer-based adapters.

[0393] When a data item (e.g., a CLVS, or a picture) has been coded at a particular temporal layer, information about the particular temporal layer may be available at decoder side. For example, the information about the particular temporal layer may be signaled from an encoder to the decoder, or may be derived at decoder side based on other information. The decoder may select the adapter that is associated with a temporal layer that is equal to the particular temporal layer, or that is associated with a temporal layer that is most similar to the particular temporal layer according to a predefined metric or evaluation process. The selected adapter is then used to adapt the DSNN, obtaining an adapted DSNNthat may be used for decoding the data item.

[0394] Embodiments on temporal-layer-wise adapters - Training

[0395] During training of temporal-layer -based adapters, a number N of adapters is chosen, for example, based on the number of layers in a hierarchy for inter-frame prediction patterns. Then, the training dataset may be divided into N portions or subsets, where each portion or subset comprises data associated (e.g., coded) with a particular temporal layer. A i-th adapter is trained by using training data in the i-th portion or subset.

[0396] Embodiments on loss-function-wise adapters - Inference

[0397] In an embodiment, a decoder comprises N adapters that are associated with respective N loss functions that were used to train the N adapters. Such adapters may be referred to as loss-function-wise adapters or loss-function-based adapters. The decoder may or may not comprise information about the actual loss functions that were used to train the N adapters. In an example, the decoder has information that a first adapter was trained by means of LI loss function, and a second adapter was trained by means of L2 loss function. In another example, the decoder comprises two adapters that are associated with respective two indexes, such as 0 and 1.

[0398] An encoder may signal to the decoder information about which adapter to use for decoding a certain data item. In one example, the encoder signals information that explicitly indicates the loss function used to train the adapter to be used for decoding a certain data item. In another example, the encoder signals information that indicates the adapter to be used for decoding a certain data item without an explicit reference to how that adapter was trained, such as by using an index.

[0399] When the encoder signals to the decoder information indicative of a particular loss function used to train the adapter to be used for decoding a certain data item, the decoder may select the adapter that is associated with a loss function that is equal to the particular loss function, or that is associated with a loss function that is most similar to the particular loss function according to a predefined metric or evaluation process.

[0400] It is to be noted that, while some embodiments describe the case of adapters associated with loss functions used to train the adapters, the same embodiments may be valid or applicable to the other hyper-parameters or characteristics of a training process used to train the adapters. Other examples include, but are not limited to, size of training dataset, content type of training dataset, learning rate, optimization routine, data augmentation method, preprocessing method, post-processing method, and the like.

[0401] Embodiments on loss-function-wise adapters - Training

[0402] At the beginning of the training of N loss-function-wise adapters, a number N of adapters is chosen, for example, based on the number of available loss functions. N training sessions are performed based at least on respective N loss functions, obtaining respective N adapters. At each i-th training session, a i-th adapter is trained by using a i-th loss function.

[0403] FIG. 11 is an example apparatus 1100, which may be implemented in hardware, configured to implement the examples described herein. The apparatus 1100 comprises at least one processor 1102 (e.g., an FPGA and / or CPU), at least one memory 1104 including computer program code 1105, the computer program code 1105 having instructions to carry out the methods described herein, wherein the at least one memory 1104 and the computer program code 1105 are configured to, with the at least one processor 1102, cause the apparatus 1100 to implement circuitry, a process, component, module, or function (implemented with control module 1106) to implement the examples described herein, including determining and using pretrained adapters for decoder-side neural networks. Optionally included encoder 1108 of the control module 1106 implements encoding based on the examples described herein, and optionally included decoder 1011 implements decoding based on the examples described herein. The at least one memory 1104 may be a non-transitory memory, a transitory memory, a volatile memory (e.g. RAM), or a non-volatile memory (e.g., ROM).

[0404] The apparatus 1100 includes a display and / or I / O interface 1112, which includes user interface (UI) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one or more sensors, etc. The apparatus 1100 includes one or more communication e.g., network (N / W) interfaces (I / F(s)) 1114. The communication I / F(s) 1114 may be wired and / or wireless and communicate over the Internet / other network(s) via any communication technique including via one or more links 1116. The communication I / F(s) 1114 may comprise one or more transmitters or one or more receivers.

[0405] The transceiver 1118 comprises one or more transmitters 1120 and one or more receivers 1122. The transceiver 1118 and / or communication I / F(s) 1114 may comprise standard well-known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 1124 used for communication over wireless link 1126.

[0406] The control module 1106 of the apparatus 1100 comprises one of or both parts 1106-1 and / or1106-2, which may be implemented in a number of ways. The control module 1106 may be implemented in hardware as control module 1106-1, such as being implemented as part of the at least one processor 1102. The control module 1106-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 1106 may be implemented as control module 1106-2, which is implemented as computer program code (having corresponding instructions) 1105 and is executed by the at least one processor 1102. For instance, the at least one memory 1104 store instructions that, when executed by the at least one processor 1102, cause the apparatus 1100 to perform one or more of the operations as described herein. Furthermore, the at least one processor 1102, the at least one memory 1104, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.

[0407] The apparatus 1100 to implement the functionality of control module 1106 may correspond to any of the apparatuses depicted herein. Alternatively, apparatus 1100 and its elements may not correspond to any of the other apparatuses depicted herein, as apparatus 1100 may be part of a self- organizing / optimizing network (SON) node or other node, such as a node in a cloud.

[0408] The apparatus 1100 may also be distributed throughout the network including within and between apparatus 1100 and any network element (such as a base station and / or terminal device and / or user equipment).

[0409] Interface 1128 enables data communication and signaling between the various items of apparatus 1100, as shown in FIG. 11. For example, the interface 1128 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g. instructions) 1105, including control module 1106 may comprise object-oriented software configured to pass data or messages between objects within computer program code 1105. The apparatus 1100 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 1100 may at least partially reside in a housing 1130, or a subset of the various components of apparatus 1100 may at least partially be located in different housings, which different housings may include housing 1130.

[0410] FIG. 12 shows a schematic representation of non-volatile memory media 1200a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 1200b (e.g. universal serial bus (USB) memory stick) and 1200c (e.g. cloud storage for downloading instructions and / or parameters 1202 or receiving emailed instructions and / or parameters 1202) storing instructions and / or parameters 1202 which when executed by a processor allows the processor to perform one or more of the operations ofthe methods described herein. Instructions and / or parameters 1202 may represent or correspond to a non-transitory computer readable medium.

[0411] FIG. 13 is an example method 1300 performed with an encoder or a decoder, based on the examples described herein. At 1302, the method 1300 includes selecting at least one adapter from one or more adapters based at least on a selection criterion to obtain at least one selected adapter. At 1304, the method 1300 includes, wherein each of the one or more adapters comprises one or more adaptation parameters. At 1306, the method 1300 includes using the at least one selected adapter to adapt a neural network, based at least on an adaptation process, to obtain an adapted neural network. At 1308, the method 1300 includes, wherein the adapted neural network is intended to be used to process a data item that is input to the neural network.

[0412] The method 1300 may be performed with an apparatus, such as the apparatus 100, 1100, or the apparatuses depicted in FIG. 3 and FIG. 4, for example, the transmitting apparatus 406 with the encoder 402, or the apparatus 400 with the encoder 402; or the receiving apparatus 410 with the decoder 412, or the apparatus 400 with the decoder 412.

[0413] As described above, FIG. 13 include flowcharts of an apparatus (e.g. 100, 1100, or any other apparatuses described herein), method, and computer program product according to certain example embodiments. It will be understood that each block of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other devices associated with execution of software including one or more computer program instructions. For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instructions which embody the procedures described above may be stored by a memory (e.g., 112 or 1104) of an apparatus employing an embodiment of the present invention and executed by processing circuitry (e.g., 110 or 1102) of the apparatus. As will be appreciated, any such computer program instructions may be loaded onto a computer or other programmable apparatus (e.g., hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture, the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or otherprogrammable apparatus provide operations for implementing the functions specified in the flowchart blocks.

[0414] A computer program product is therefore defined in those instances in which the computer program instructions, such as computer-readable program code portions, are stored by at least one non- transitory computer-readable storage medium with the computer program instructions, such as the computer-readable program code portions, being configured, upon execution, to perform the functions described above, such as in conjunction with the flowchart(s) of FIG. 13. In other embodiments, the computer program instructions, such as the computer-readable program code portions, need not be stored or otherwise embodied by a non-transitory computer-readable storage medium, but may, instead, be embodied by a transitory medium with the computer program instructions, such as the computer- readable program code portions, still being configured, upon execution, to perform the functions described above.

[0415] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions for performing the specified functions. It will also be understood that one or more blocks of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by special purpose hardware-based computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.

[0416] In some embodiments, certain ones of the operations above may be modified or further amplified. Furthermore, in some embodiments, additional optional operations may be included. Modifications, additions, or amplifications to the operations above may be performed in any order and in any combination.

[0417] Some embodiments have been described in relation to one or more neural networks performing visual temporal extrapolation. It is to be understood that embodiments can be realized with any generative modelling neural networks.

[0418] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.

[0419] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needsto be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.

[0420] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, although the foregoing descriptions and the associated drawings describe example embodiments in the context of certain example combinations of elements and / or functions, it should be appreciated that different combinations of elements and / or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, different combinations of elements and / or functions than those explicitly described above are also contemplated as may be set forth in some of the appended claims. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0421] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.

[0422] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not only computers having different architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device, and the like.

[0423] As used herein, the term ‘circuitry’ may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of circuits and software (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii)portions of processor(s) / software including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even when the software or firmware is not physically present. This description of ‘circuitry’ applies to uses of this term in this application. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.

[0424] Circuitry or Circuit: As used in this application, the term ‘circuitry’ or ‘circuit’ may refer to one or more or all of the following:(a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); and(b) combinations of hardware circuits and software, such as (as applicable):(i) a combination of analog and / or digital hardware circuit(s) with software / firmware; and(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and(c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.

[0425] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example, and when applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

Claims

CLAIMSWhat is claimed is:

1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: selecting at least one adapter from one or more adapters based at least on a selection criterion to obtain at least one selected adapter; wherein each of the one or more adapters comprises one or more adaptation parameters; using the at least one selected adapter to adapt a neural network, based at least on an adaptation process, to obtain an adapted neural network; and wherein the adapted neural network is intended to be used to process a data item that is input to the neural network.

2. The apparatus of claim 1, wherein the neural network and / or the adapted neural network are part of an encoding process at an encoder side, a decoding process at a decoder side, and / or a post-processing process at a decoder side.

3. The apparatus of any of claims 1 or 2, wherein the one or more adaptation parameters of each of the one or more adapters are associated with respective one or more positions within the neural network.

4. The apparatus of any of the previous claims, wherein the one or more adaptation parameters of each of the one or more adapters are grouped into one or more adaptation layers, wherein the one or more adaptation layers are associated with respective one or more layer positions within the neural network.

5. The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: using the one or more adaptation parameters as parameters or weights of the neural network; using the one or more adaptation parameters as bias parameters of the neural network; using the one or more adaptation parameters as multiplier parameters of the neural network; using the one or more adaptation parameters as parameters of one or more convolutional layers;using the one or more adaptation parameters as an external or internal input of the neural network; using the one or more adaptation parameters as an input to the neural network; or using the one or more adaptation parameters as an input to one or more layers of the neural network.

6. The apparatus of any of the previous claims, wherein the adaptation process comprises: inserting the one or more adaptation parameters of the at least one selected adapter into the respective one or more associated positions within the neural network; replacing values of at least one parameter of the neural network with values of the one or more adaptation parameters of the at least one selected adapter; modifying the values of at least one parameter of the neural network based on the values of the one or more adaptation parameters of the at least one selected adapter; and inserting the one or more adaptation layers of the at least one selected adapter into respective one or more associated layer positions within the neural network.

7. The apparatus of any of the previous claims, wherein the one or more adapters are associated with respective one or more identifiers or respective one or more nominal or numerical values of a feature or characteristic of at least one of: the data item, the neural network, the encoding process, or the decoding process.

8. The apparatus of any of the previous claims, wherein the neural network is adapted on one or more of: a video sequence basis, a coded layer video sequence basis, a random-access segment basis, an intra-frame period basis, a frame or picture basis, or a block basis.

9. The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: grouping two or more adapters of the one or more adapters into two or more categories, wherein a category of the two or more categories comprises one or more adapters that are similar or substantially similar based on a similarity metric or a feature, and wherein one or more first adapters in a first category of the two or more categories are different from one or more second adapters in a second category of the two or more categories based on the similarity metric or the feature.

10. The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: combining two or more adapters, by using a combination operation, to obtain a combined adapter, wherein the combined adapter is used to adapt the neural network.

11. The apparatus of claim 10, wherein the combination operation comprises averaging,weighted averaging, or using the two or more adapters by cascading adaptation parameters of different adapters in the two or more adapters.

12. The apparatus of claim 11, wherein when the combination operation comprises weighted averaging, weights or coefficients used in the weighted averaging are determined based on: a similarity between a data item to be processed and the two or more adapters, in terms of a similarity metric or a similarity criterion; or a similarity between a characteristic of the data item to be processed and characteristics of the two or more adapters, in terms of the similarity metric or similarity criterion.

13. The apparatus of any of the claims 10 to 12, wherein the two or more adapters that are combined are selected based on: a similarity between the data item to be processed and the two or more adapters, in terms of the similarity metric or similarity criterion; or a similarity between a characteristic of a data item to be processed and characteristics of the two or more adapters, in terms of the similarity metric or similarity criterion.

14. The apparatus of claim 13, wherein the similarity criterion comprises a similarity between a value of a particular feature of the data item to be processed and values of particular features that are associated with the two or more adapters to be combined.

15. The apparatus of any of the previous claims, wherein the selection criterion comprises: selecting the at least one selected adapter based on an indication received from an encoder; or selecting the at least one selected adapter based on one or more parameters or indicators that are available at a decoder side and that are used for some other operations of the decoder in addition to being used for selecting the at least one selected adapter.

16. The apparatus of claim 15, wherein the indication is associated with information about a scope of the at least one selected adapter.

17. The apparatus of claim 16, wherein the indication is indicative of a position of the neural network within a processing pipeline; the indication is indicative of a content type, wherein the content type is associated to the at least one selected adapter; or the indication is indicative of a loss function that is associated to the at least one selected adapter.

18. The apparatus of claim 15, wherein: the one or more parameters or indicators comprise one or more of the following:a quantization parameter associated with a picture to be processed by using the neural network; a picture type or block type of a picture or a block to be processed by using the neural network; a resolution information of one or more video frames to be processed by using the neural network; a temporal layer identifier of a temporal layer of an inter-frame prediction hierarchy structure; a position of the neural network within a coding pipeline; or a type of content to be processed by using the neural network.

19. The apparatus of claim 18, wherein the apparatus is further caused to perform: signaling or receiving a correction or adjustment for the one or more parameters or indicators that are available at decoder side and that are used for some other operations of the decoder in addition to being used for selecting the at least one selected adapter, wherein the one or more parameters or indicators are corrected or adjusted based on the correction or adjustment to optimize selection of the at least one selected adapter.

20. The apparatus of claim 19, wherein the correction or adjustment are intended to be used by the decoder to correct or adjust the one or more parameters or indicators to obtain respective one or more corrected or adjusted parameters or indicators that are suitable or optimal to be used for selection of the at least one adapter.

21. The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: adapting the at least one selected adapter, based on an adaptation signal, to obtain an updated adapter, wherein the updated adapter is used to adapt the neural network.

22. The apparatus of any of the previous claims, wherein the one or more adaptation parameters are associated with one or more adaptation layers of the at least one selected adapter, and wherein the one or more adaptation parameters of the one or more adaptation layers of the at least one selected adapter are integrated into already present parameters of the neural network.

23. The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: initializing the one or more adaptation parameters based on: a predetermined value, values sampled from a probability distribution, or a decomposition operation that is performed on parameters of the neural network associated with the one or more adaptation parameters.

24. The apparatus of any of the previous claims, wherein the neural network comprises a decoder side neural network.

25. The apparatus of claim 2, wherein the decoder side comprises N adapters that are associated with N quantization parameters (QPs) or with N QP ranges, wherein QP refers to the quantization parameter used to derive the data item, wherein N is a natural number greater than or equal to 1.

26. The apparatus of claim 25, wherein the QP comprises one of the following: a sequence QP, a picture QP, a slice QP, or a block QP.

27. The apparatus of any of claims 25 or 26, wherein when the data item is coded with a particular QP, the particular QP information is available at decoder side, and wherein the decoder selects an adapter associated with a QP that is equal to the particular QP, an adapter that is associated with a QP that is most similar to the particular QP, or an adapter that is associated with a QP range that comprises the particular QP.

28. The apparatus of claim 2, wherein the decoder side comprises N adapters that are associated with N possible positions of the neural network within a pipeline or cascade of processing steps, wherein N is a natural number greater than or equal to 1.

29. The apparatus of claim 28, wherein the apparatus is further caused to perform: receiving an indication of the position of the neural network to be used for processing the data item, and wherein the indication of the position is further used for selecting the adapter that is associated with that position.

30. The apparatus of claim 2, wherein the decoder side comprises N adapters that are associated with N types of content, wherein N is a natural number greater than or equal to 1.

31. The apparatus of claim 30, wherein a type of content comprises a particular set of characteristics, and different types of content comprise different sets of characteristics, and wherein the content or data in a certain content type share same or substantially same characteristics, and wherein the content or data in different content types do not share same or substantially same characteristics.

32. The apparatus of claim 30, wherein the apparatus is further caused to perform: receiving an indication of the content type associated with the data item; selecting the adapter that is associated with the same or substantially same content type asindicated by an encoder; using the selected adapter for adapting the neural network; and using the neural network for processing the data item.

33. The apparatus of claim 2, wherein the decoder side comprises N adapters that are associated with N picture types, wherein N is a natural number greater than or equal to 1.

34. The apparatus of claim 33, wherein a picture type comprises an intra-coded picture, or an inter-coded picture.

35. The apparatus of claim 33, wherein the decoder comprises information about a picture type associated to a picture to be decoded, and wherein information about the picture type associated to the picture to be decoded is received from an encoder or is derived at the decoder side.

36. The apparatus of claim 35, wherein when the decoder decodes the picture of the particular picture type, the decoder is caused to perform: selecting the adapter that is associated with that particular picture type, or that is associated with a picture type that is most similar to that particular picture type.

37. The apparatus of claim 2, wherein the decoder side comprises N adapters that are associated with N resolutions or N resolution-ranges, wherein N is a natural number greater than or equal to 1.

38. The apparatus of claim 37, wherein when the data item has been coded with a particular resolution, information about the particular resolution is available at the decoder side, and wherein the information about that particular resolution is received from an encoder, or may be derived at decoder side based on other information.

39. The apparatus of claim 37, wherein the decoder is caused to perform: selecting an adapter that is associated with a resolution that is equal to the particular resolution, or an adapter that is associated with a resolution that is most similar to the particular resolution, or an adapter that is associated with a resolution range that comprises the particular resolution.

40. The apparatus of claim 2, wherein the decoder side comprises N adapters that are associated with N temporal layers of a predetermined hierarchy for inter-frame prediction patterns, wherein N is a natural number greater than or equal to 1.

41. The apparatus of claim 40, wherein when the data item has been coded with a particular temporal layer, information about the particular temporal layer is available at the decoder side,and wherein the information about the particular temporal layer is received from an encoder, or may be derived at decoder side based on other information.

42. The apparatus of claim 41, wherein the decoder is caused to perform: selecting an adapter that is associated with a temporal layer that is equal to the particular temporal layer, or an adapter that is associated with a resolution that is most similar to the particular temporal layer according to a predefined metric or evaluation process.

43. The apparatus of claim 2, wherein the decoder side comprises N adapters that are associated with loss functions that were used to train the N adapters.

44. The apparatus of claim 43, wherein the apparatus is further caused to perform: receiving information about which adapter to use for processing the data item.

45. The apparatus of claim 44, wherein the information explicitly indicates a loss function used to train the adapter to be used for processing the data item.

46. The apparatus of claim 43, wherein the apparatus is further caused to perform: selecting an adapter that is associated with a loss function that is equal to the loss function; or selecting an adapter that is associated with a loss function that is most similar to the loss function according to a predefined metric or evaluation process.

47. The apparatus of claim 44, wherein the information indicates which adapter to use for processing the data item without an explicit reference to how the adapter is trained.

48. A method comprising: selecting at least one adapter from one or more adapters based at least on a selection criterion to obtain at least one selected adapter; wherein each of the one or more adapters comprises one or more adaptation parameters; using the at least one selected adapter to adapt a neural network, based at least on an adaptation process, to obtain an adapted neural network; and wherein the adapted neural network is intended to be used to process a data item that is input to the neural network.

49. The method of claim 48, wherein the neural network and / or the adapted neural network are part of an encoding process at an encoder side, a decoding process at a decoder side, and / or a post-processing process at a decoder side.

50. The method of any of claims 48 or 49, wherein the one or more adaptation parameters of each of the one or more adapters are associated with respective one or more positions within the neural network.

51. The method of any of the claims 48 to 50, wherein the one or more adaptation parameters of each of the one or more adapters are grouped into one or more adaptation layers, wherein the one or more adaptation layers are associated with respective one or more layer positions within the neural network.

52. The method of any of the claims 48 to 51 further comprising: using the one or more adaptation parameters as parameters or weights of the neural network; using the one or more adaptation parameters as bias parameters of the neural network; using the one or more adaptation parameters as multiplier parameters of the neural network; using the one or more adaptation parameters as parameters of one or more convolutional layers; using the one or more adaptation parameters as an external or internal input of the neural network; using the one or more adaptation parameters as an input to the neural network; or using the one or more adaptation parameters as an input to one or more layers of the neural network.

53. The method of any of the claims 48 to 52, wherein the adaptation process comprises: inserting the one or more adaptation parameters of the at least one selected adapter into the respective one or more associated positions within the neural network; replacing values of at least one parameter of the neural network with values of the one or more adaptation parameters of the at least one selected adapter; modifying the values of at least one parameter of the neural network based on the values of the one or more adaptation parameters of the at least one selected adapter; and inserting the one or more adaptation layers of the at least one selected adapter into respective one or more associated layer positions within the neural network.

54. The method of any of the claims 48 to 53, wherein the one or more adapters are associated with respective one or more identifiers or respective one or more nominal or numerical values of a feature or characteristic of at least one of: the data item, the neural network, the encoding process, or the decoding process.

55. The method of any of the claims 48 to 54, wherein the neural network is adapted on one or more of: a video sequence basis, a coded layer video sequence basis, a random-access segment basis, an intra-frame period basis, a frame or picture basis, or a block basis.

56. The method of any of the claims 48 to 55 further comprising: grouping two or more adapters of the one or more adapters into two or more categories, wherein a category of the two or more categories comprises one or more adapters that are similar or substantially similar based on a similarity metric or a feature, and wherein one or more first adapters in a first category of the two or more categories are different from one or more second adapters in a second category of the two or more categories based on the similarity metric or the feature.

57. The method of any of the claims 48 to 56 further comprising: combining two or more adapters, by using a combination operation, to obtain a combined adapter, wherein the combined adapter is used to adapt the neural network.

58. The method of claim 57, wherein the combination operation comprises averaging, weighted averaging, or using the two or more adapters by cascading adaptation parameters of different adapters in the two or more adapters.

59. The method of claim 58, wherein when the combination operation comprises weighted averaging, weights or coefficients used in the weighted averaging are determined based on: a similarity between a data item to be processed and the two or more adapters, in terms of a similarity metric or a similarity criterion; or a similarity between a characteristic of the data item to be processed and characteristics of the two or more adapters, in terms of the similarity metric or similarity criterion.

60. The method of any of the claims 57 to 59, wherein the two or more adapters that are combined are selected based on: a similarity between the data item to be processed and the two or more adapters, in terms of the similarity metric or similarity criterion; or a similarity between a characteristic of a data item to be processed and characteristics of the two or more adapters, in terms of the similarity metric or similarity criterion.

61. The method of claim 60, wherein the similarity criterion comprises a similarity between a value of a particular feature of the data item to be processed and values of particular features that are associated with the two or more adapters to be combined.

62. The method of any of the claims 48 to 61, wherein the selection criterion comprises:selecting the at least one selected adapter based on an indication received from an encoder; or selecting the at least one selected adapter based on one or more parameters or indicators that are available at a decoder side and that are used for some other operations of the decoder in addition to being used for selecting the at least one selected adapter.

63. The method of claim 62, wherein the indication is associated with information about a scope of the at least one selected adapter.

64. The method of claim 63, wherein the indication is indicative of a position of the neural network within a processing pipeline; the indication is indicative of a content type, wherein the content type is associated to the at least one selected adapter; or the indication is indicative of a loss function that is associated to the at least one selected adapter.

65. The method of claim 62, wherein: the one or more parameters or indicators comprise one or more of the following: a quantization parameter associated with a picture to be processed by using the neural network; a picture type or block type of a picture or a block to be processed by using the neural network; a resolution information of one or more video frames to be processed by using the neural network; a temporal layer identifier of a temporal layer of an inter-frame prediction hierarchy structure; a position of the neural network within a coding pipeline; or a type of content to be processed by using the neural network.

66. The method of claim 65 further comprising: signaling or receiving a correction or adjustment for the one or more parameters or indicators that are available at decoder side and that are used for some other operations of the decoder in addition to being used for selecting the at least one selected adapter, wherein the one or more parameters or indicators are corrected or adjusted based on the correction or adjustment to optimize selection of the at least one selected adapter.

67. The method of claim 66, wherein the correction or adjustment are intended to be used by the decoder to correct or adjust the one or more parameters or indicators to obtain respective one or more corrected or adjusted parameters or indicators that are suitable or optimal to beused for selection of the at least one adapter.

68. The method of any of the claims 48 to 67 further comprising: adapting the at least one selected adapter, based on an adaptation signal, to obtain an updated adapter, wherein the updated adapter is used to adapt the neural network.

69. The method of any of the claims 48 to 68, wherein the one or more adaptation parameters are associated with one or more adaptation layers of the at least one selected adapter, and wherein the one or more adaptation parameters of the one or more adaptation layers of the at least one selected adapter are integrated into already present parameters of the neural network.

70. The method of any of the claims 48 to 69 further comprising: initializing the one or more adaptation parameters based on: a predetermined value, values sampled from a probability distribution, or a decomposition operation that is performed on parameters of the neural network associated with the one or more adaptation parameters.

71. The method of any of the claims 48 to 70, wherein the neural network comprises a decoder side neural network.

72. The method of claim 49, wherein the decoder side comprises N adapters that are associated with N quantization parameters (QPs) or with N QP ranges, wherein QP refers to the quantization parameter used to derive the data item, wherein N is a natural number greater than or equal to 1.

73. The method of claim 72, wherein the QP comprises one of the following: a sequence QP, a picture QP, a slice QP, or a block QP.

74. The method of any of claims 73 or 73, wherein when the data item is coded with a particular QP, the particular QP information is available at decoder side, and wherein the decoder selects an adapter associated with a QP that is equal to the particular QP, an adapter that is associated with a QP that is most similar to the particular QP, or an adapter that is associated with a QP range that comprises the particular QP.

75. The method of claim 49, wherein the decoder side comprises N adapters that are associated with N possible positions of the neural network within a pipeline or cascade of processing steps, wherein N is a natural number greater than or equal to 1.

76. The method of claim 75 further comprising: receiving an indication of the position of the neural network to be used for processing the data item, and wherein the indication of theposition is further used for selecting the adapter that is associated with that position.

77. The method of claim 49, wherein the decoder side comprises N adapters that are associated with N types of content, wherein N is a natural number greater than or equal to 1.

78. The method of claim 77, wherein a type of content comprises a particular set of characteristics, and different types of content comprise different sets of characteristics, and wherein the content or data in a certain content type share same or substantially same characteristics, and wherein the content or data in different content types do not share same or substantially same characteristics.

79. The method of claim 77 further comprising: receiving an indication of the content type associated with the data item; and. selecting the adapter that is associated with the same or substantially same content type as indicated by an encoder; using the selected adapter for adapting the neural network; and using the neural network for processing the data item.

80. The method of claim 49, wherein the decoder side comprises N adapters that are associated with N picture types, wherein N is a natural number greater than or equal to 1.

81. The method of claim 80, wherein a picture type comprises an intra-coded picture, or an inter-coded picture.

82. The method of claim 80, wherein the decoder comprises information about a picture type associated to a picture to be decoded, and wherein information about the picture type associated to the picture to be decoded is received from an encoder or is derived at the decoder side.

83. The method of claim 82, wherein when the decoder decodes the picture of the particular picture type, the decoder is caused to perform: selecting the adapter that is associated with that particular picture type, or that is associated with a picture type that is most similar to that particular picture type.

84. The method of claim 49, wherein the decoder side comprises N adapters that are associated with N resolutions or N resolution-ranges, wherein N is a natural number greater than or equal to 1.

85. The method of claim 84, wherein when the data item has been coded with a particular resolution, information about the particular resolution is available at the decoder side, andwherein the information about that particular resolution is received from an encoder, or may be derived at decoder side based on other information.

86. The method of claim 84, wherein the decoder is caused to perform: selecting an adapter that is associated with a resolution that is equal to the particular resolution, or an adapter that is associated with a resolution that is most similar to the particular resolution, or an adapter that is associated with a resolution range that comprises the particular resolution.

87. The method of claim 49, wherein the decoder side comprises N adapters that are associated with N temporal layers of a predetermined hierarchy for inter-frame prediction patterns, wherein N is a natural number greater than or equal to 1.

88. The method of claim 87, wherein when the data item has been coded with a particular temporal layer, information about the particular temporal layer is available at the decoder side, and wherein the information about the particular temporal layer is received from an encoder, or may be derived at decoder side based on other information.

89. The method of claim 88, wherein the decoder is caused to perform: selecting an adapter that is associated with a temporal layer that is equal to the particular temporal layer, or an adapter that is associated with a resolution that is most similar to the particular temporal layer according to a predefined metric or evaluation process.

90. The method of claim 49, wherein the decoder side comprises N adapters that are associated with loss functions that were used to train the N adapters.

91. The method of claim 90 further comprising: receiving information about which adapter to use for processing the data item.

92. The method of claim 91, wherein the information explicitly indicates a loss function used to train the adapter to be used for processing the data item.

93. The method of claim 92 further comprising: selecting an adapter that is associated with a loss function that is equal to the loss function; or selecting an adapter that is associated with a loss function that is most similar to the loss function according to a predefined metric or evaluation process.

94. The method of claim 93, wherein the information indicates which adapter to use for processing the data item without an explicit reference to how the adapter is trained.

95. An apparatus comprising means for performing methods as claimed in any of the claims 48 to 94.

96. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as claimed in any of the claims 48 to 94.

97. The computer readable medium clam 96, wherein the computer readable medium comprises a non-transitory computer readable medium.