Single-bit overfitting
The apparatus addresses the challenge of single-bit overfitting in multimedia systems by updating decoder-side neural networks based on overfitting signals, resulting in improved performance in multimedia tasks.
Patent Information
- Application Number
- PCT/IB2024/061373
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-15
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-22
AI Technical Summary
Existing multimedia systems face challenges in efficiently addressing single-bit overfitting in decoder-side neural networks, which can lead to suboptimal performance in tasks like filtering, object detection, and object tracking.
The proposed solution involves an apparatus that receives an overfitting signal, updates a decoder-side neural network by overfitting one or more bits of its weights, and uses the updated network for specific tasks. This process can involve signaling information related to the overfitted bits, such as their value or position, to the decoder.
This approach allows for improved performance in multimedia tasks by specifically targeting and optimizing the overfitted bits within the neural network, leading to enhanced filtering, detection, and tracking capabilities.
Smart Images

Figure IB2024061373_22052025_PF_FP_ABST
Abstract
Description
SINGLE-BIT OVERFITTINGSTATEMENT OF GOVERNMENT SUPPORT
[0001] The project leading to this application has received funding from the ECSEL Joint Undertaking (JU) under grant agreement No 876019. The JU receives support from the European Union’s Horizon 2020 research and innovation programme and Germany, Netherlands, Austria, Romania, France, Sweden, Cyprus, Greece, Lithuania, Portugal, Italy, Finland, Turkey.TECHNICAL FIELD
[0002] The examples and non-limiting embodiments relate generally to multimedia transport and, more particularly, to a single-bit overfitting.BACKGROUND
[0003] It is known to perform data compression and decoding in a multimedia system.SUMMARY
[0004] Example 1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving an overfitting signal; updating a decoder-side neural network, based at least on the overfitting signal, to obtain an updated decoder-side neural network, wherein for one or more weights of the decoder-side neural network, at least one bit of a binary word representing a weight is overfitted; and using the updated decoder-side neural network for at least one task.
[0005] Example 2. The apparatus of example 1, wherein the at least one task comprises one or more of filtering input data, objection detection, object segmentation, or object tracking.
[0006] Example 3. The apparatus of any of examples 1 to 2, wherein: all weights of the decoderside neural network are overfitted to obtain the overfitting signal; a subset of the weights of the decoder-side neural network are overfitted to obtain the overfitting signal; or bias weights of the decoder-side neural network are overfitted to obtain the overfitting signal.
[0007] Example 4. The apparatus of any of examples 1 to 3, wherein the apparatus is further caused at least to perform: receiving an indication of a value of the at least one bit that is overfitted; wherein when the at least one bit that is overfitted comprises an original value of 0, the overfitted value comprises 1 and the indication comprises a bit with the value 1; and wherein when the at least one bit that is overfitted comprises the original value of 1, the overfitted value comprises 0, and the indication comprises the bit with the value 0.
[0008] Example 5. The apparatus of any of examples 1 to 4, wherein a position of the at least one bit that is overfitted is predetermined and is known to a decoder.
[0009] Example 6. The apparatus of example 5, wherein the position of the at least one bit that is overfitted is predetermined based on statistics collected during an offline phase where overfitting is performed on a dataset.
[0010] Example 7. The apparatus of any of examples 1 to 4, wherein the apparatus is further caused at least to perform: receiving, from an encoder, an indication of a position of the at least one bit that is overfitted.
[0011] Example 8. The apparatus of example 7, wherein the at least one bit that is overfitted is determined at the encoder.
[0012] Example 9. The apparatus of any of the previous examples, wherein the indication comprises a change indication bit, wherein when the change-indication bit is 0, an original bit of a weight is left unmodified, wherein when the change-indication bit is 1, the original bit of the weight is flipped.
[0013] Example 10. The apparatus of example 9, wherein a bitwise Boolean operation is applied to the original bit of the weight based on the change indication bit.
[0014] Example 11. The apparatus of example 10, wherein the bitwise Boolean operation is predetermined and is known at the decoder.
[0015] Example 12. The apparatus of example 10, wherein the apparatus is further caused to receive an indication of the bitwise Boolean operation, and wherein the bitwise Boolean operation is determined at the encoder.
[0016] Example 13. The apparatus of any of examples 1 to 12, wherein all the weights that are comprised in a layer or a group of layers and that are overfitted share the same predetermined position of the bit that is overfitted.
[0017] Example 14. The apparatus of any of examples 1 to 12, wherein for each weight or for a group of weights, a number of bits and positions of the bits are predetermined, and wherein the apparatus is further caused to receive an indication of values of the bits, or an indication of whether the bits are to be flipped.
[0018] Example 15. The apparatus of any of the previous examples, wherein the apparatus is further caused to perform: receiving, for each overfitted weight, a lossless coded binary wordrepresenting the weight update.
[0019] Example 16. The apparatus of example 1, wherein a set of candidate weight updates are predetermined for the weight that is overfitted.
[0020] Example 17. The apparatus of example 16, wherein predetermination of the set of candidate weight updates is based on determining one or more weight-updates based on a dataset.
[0021] Example 18. The apparatus of any of examples 16 or 17, wherein the apparatus is further caused to perform: receiving an indication of whether an indicated candidate weight update needs to be added to or subtracted from an original weight.
[0022] Example 19. The apparatus of example 18, wherein a sign of the weight update which determines whether the weight update is added to or subtracted from the original weight is predetermined.
[0023] Example 20. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: for one or more weights of a decoder-side neural network, overfitting at least one bit of a binary word representing a weight that is overfitted to obtain an overfitted weight; and signaling information related to the at least one bit of the binary word representing the weight that is overfitted to a decoder.
[0024] Example 21. The apparatus of example 20, wherein: all weights of the decoder-side neural network are overfitted; a subset of weights of the decoder-side neural network are overfitted; or bias weights of the decoder-side neural network are overfitted.
[0025] Example 22. The apparatus of any of examples 20 or 21, wherein to signal the information related to the at least one bit of the binary word representing the weight that is overfitted, the apparatus is further caused at least to perform: signaling an indication of a value of the at least one bit that was overfitted; wherein when the at least one bit that is overfitted comprises an original value of 0, an overfitted value comprises 1 and the indication comprises a bit with the value 1; and wherein when the at least one bit that is overfitted comprises the original value of 1, the overfitted value comprises 0, and the indication comprises a bit with the value 0.
[0026] Example 23. The apparatus of any of examples 20 to 22, wherein the apparatus is further caused to perform: determining a position of the at least one bit that is overfitted.
[0027] Example 24. The apparatus of example 23, wherein the position of the at least one bit that is overfitted is predetermined based on statistics collected during an offline phase whereoverfitting is performed on a dataset.
[0028] Example 25. The apparatus of any of examples 20 to 23, wherein the apparatus is further caused at least to perform: signaling an indication of the position of the at least one bit that is overfitted to the decoder.
[0029] Example 26. The apparatus of any of the examples 20 to 25, wherein the indication comprise a change-indication bit, wherein when the change-indication bit is 0, an original bit of a weight is left unmodified, wherein when the change-indication bit is 1, the original bit of the weight is flipped.
[0030] Example 27. The apparatus of example 26, wherein a change indication bit comprises an indicator of whether to apply a bitwise Boolean operation to the original bit of the weight, and wherein the apparatus is further caused to perform: determining the bitwise Boolean operation; and signaling, to the decoder, the indication whether to apply the bitwise Boolean operation.Example 28. The apparatus of any of examples 20 to 27, wherein all the weights that are comprised in a layer or a group of layers and that are overfitted share the same predetermined position of the bit that is overfitted.
[0031] Example 29. The apparatus of any of examples 20 to 27, wherein for each weight or for a group of weights, a number of bits and positions of the bits are predetermined, and wherein the apparatus is further caused to receive an indication of values of the bits, or an indication of whether the bits are to be flipped.
[0032] Example 30. The apparatus of any of the examples 20 to 29, wherein the apparatus is further caused to perform: for each overfitted weight, lossless coding the binary word representing the weight update; and for each overfitted weight, signaling the lossless coded the binary word representing the weight update to the decoder.
[0033] Example 31. The apparatus of example 20, wherein the apparatus is further caused to perform: predetermining a set of candidate weight updates for the weight that is overfitted.
[0034] Example 32. The apparatus of example 31, wherein predetermination of the set of candidate weight updates is based on determining one or more weight-updates based on a dataset.
[0035] Example 33. The apparatus of any of examples 31 or 32, wherein the apparatus is further caused to perform: signaling an indication of whether an indicated candidate weight update needs to be added to or subtracted from the weight that is overfitted.
[0036] Example 34. The apparatus of example 33, wherein the apparatus is further caused to perform predetermining a sign of the weight update which determines whether the weight update is added to or subtracted from the weight that is overfitted.
[0037] Example 35. A method comprising: receiving an overfitting signal; updating a decoderside neural network, based at least on the overfitting signal, to obtain an updated decoder-side neural network, wherein for one or more weights of the decoder-side neural network, at least one bit of a binary word representing a weight is overfitted; and using the updated decoder-side neural network for at least one task.
[0038] Example 36. The method of example 35, wherein the at least one task comprises one or more of filtering input data, objection detection, object segmentation, or object tracking.
[0039] Example 37. The method of any of examples 35 to 36, wherein: all weights of the decoder-side neural network are overfitted to obtain the overfitting signal; a subset of the weights of the decoder-side neural network are overfitted to obtain the overfitting signal; or bias weights of the decoder-side neural network are overfitted to obtain the overfitting signal.
[0040] Example 38. The method of any of examples 35 to 37 further caused at comprising: receiving an indication of a value of the at least one bit that is overfitted; wherein when the at least one bit that is overfitted comprises an original value of 0, the overfitted value comprises 1 and the indication comprises a bit with the value 1; and wherein when the at least one bit that is overfitted comprises the original value of 1, the overfitted value comprises 0, and the indication comprises the bit with the value 0.
[0041] Example 39. The method of any of examples 35 to 38, wherein a position of the at least one bit that is overfitted is predetermined and is known to a decoder.
[0042] Example 40. The method of example 39, wherein the position of the at least one bit that is overfitted is predetermined based on statistics collected during an offline phase where overfitting is performed on a dataset.
[0043] Example 41. The method of any of examples 35 to 38 further comprising: receiving, from an encoder, an indication of a position of the at least one bit that is overfitted.
[0044] Example 42. The method of example 41, wherein the at least one bit that is overfitted is determined at the encoder.
[0045] Example 43. The method of any of the examples 35 to 42, wherein the indication comprises a change-indication bit, wherein when the change-indication bit is 0, an original bit of aweight is left unmodified, wherein when the change-indication bit is 1, the original bit of the weight is flipped.
[0046] Example 44. The method of example 43, wherein a bitwise Boolean operation is applied to the original bit of the weight based on the change indication bit.
[0047] Example 45. The method of example 44, wherein the bitwise Boolean operation is predetermined and is known at the decoder.
[0048] Example 46. The method of example 44 further comprising receiving an indication of the bitwise Boolean operation, and wherein the bitwise Boolean operation is determined at the encoder.
[0049] Example 47. The method of any of examples 35 to 46, wherein all the weights that are comprised in a layer or a group of layers and that are overfitted share the same predetermined position of the bit that is overfitted.
[0050] Example 48. The method of any of examples 35 to 46, wherein for each weight or for a group of weights, a number of bits and positions of the bits are predetermined, and wherein the method further comprising receiving an indication of values of the bits, or an indication of whether the bits are to be flipped.
[0051] Example 49. The method of any of the examples 35 to 48 further comprising receiving, for each overfitted weight, a lossless coded binary word representing the weight update.
[0052] Example 50. The method of example 35, wherein a set of candidate weight updates are predetermined for the weight that is overfitted.
[0053] Example 51. The method of example 50, wherein predetermination of the set of candidate weight updates is based on determining one or more weight-updates based on a dataset.
[0054] Example 52. The method of any of examples 50 or 51 further comprising receiving an indication of whether an indicated candidate weight update needs to be added to or subtracted from an original weight.
[0055] Example 53. The method of example 52, wherein a sign of the weight update which determines whether the weight update is added to or subtracted from the original weight is predetermined.
[0056] Example 54. An method comprising: for one or more weights of a decoder-side neural network, overfitting at least one bit of a binary word representing a weight that is overfitted to obtainan overfitted weight; and signaling information related to the at least one bit of the binary word representing the weight that is overfitted to a decoder.
[0057] Example 55. The method of example 54, wherein: all weights of the decoder-side neural network are overfitted; a subset of weights of the decoder-side neural network are overfitted; or bias weights of the decoder-side neural network are overfitted.
[0058] Example 56. The method of any of examples 54 or 55, wherein to signal the information related to the at least one bit of the binary word representing the weight that is overfitted, the method is further caused at least to perform: signaling an indication of a value of the at least one bit that was overfitted; wherein when the at least one bit that is overfitted comprises an original value of 0, an overfitted value comprises 1 and the indication comprises a bit with the value 1; and wherein when the at least one bit that is overfitted comprises the original value of 1, the overfitted value comprises 0, and the indication comprises a bit with the value 0.
[0059] Example 57. The method of any of examples 54 to 56 further comprising determining a position of the at least one bit that is overfitted.
[0060] Example 58. The method of example 57, wherein the position of the at least one bit that is overfitted is predetermined based on statistics collected during an offline phase where overfitting is performed on a dataset.
[0061] Example 59. The method of any of examples 54 to 57 further comprising signaling an indication of the position of the at least one bit that is overfitted to the decoder.
[0062] Example 60. The method of any of the examples 54 to 59, wherein the indication comprise a change-indication bit, wherein when the change-indication bit is 0, an original bit of a weight is left unmodified, wherein when the change-indication bit is 1, the original bit of the weight is flipped.
[0063] Example 61. The method of example 60, wherein the change indication bit comprises an indicator of whether to apply a bitwise Boolean operation to the original bit of the weight, and wherein the method is further caused to perform: determining the bitwise Boolean operation; and signaling, to the decoder, the indication whether to apply the bitwise Boolean operation.
[0064] Example 62. The method of any of examples 54 to 61, wherein all the weights that are comprised in a layer or a group of layers and that are overfitted share the same predetermined position of the bit that is overfitted.
[0065] Example 63. The method of any of examples 54 to 61, wherein for each weight or for agroup of weights, a number of bits and positions of the bits are predetermined, and wherein the method further comprising receiving an indication of values of the bits, or an indication of whether the bits are to be flipped.
[0066] Example 64. The method of any of the examples 54 to 63 further comprising: for each overfitted weight, lossless coding the binary word representing the weight update; and for each overfitted weight, signaling the lossless coded the binary word representing the weight update to the decoder.
[0067] Example 65. The method of example 54 further comprising predetermining a set of candidate weight updates for the weight that is overfitted.
[0068] Example 66. The method of example 65, wherein predetermination of the set of candidate weight updates is based on determining one or more weight-updates based on a dataset.
[0069] Example 67. The method of any of examples 65 or 66 further comprising signaling an indication of whether an indicated candidate weight update needs to be added to or subtracted from the weight that is overfitted.
[0070] Example 68. The method of example 67 further comprising predetermining a sign of the weight update which determines whether the weight update is added to or subtracted from the weight that is overfitted.
[0071] Example 69. An apparatus comprising: means for receiving an overfitting signal; means for updating a decoder-side neural network, based at least on the overfitting signal, to obtain an updated decoder-side neural network, wherein for one or more weights of the decoder-side neural network, at least one bit of a binary word representing a weight is overfitted; and means for using the updated decoder-side neural network for at least one task.
[0072] Example 70. The apparatus of example 69, wherein the apparatus is further caused to perform the methods as described in any of the examples 36 to53.
[0073] Example 71. Apparatus comprising: means for overfitting, for one or more weights of a decoder-side neural network, at least one bit of a binary word representing a weight that is overfitted to obtain an overfitted weight; and means for signaling information related to the at least one bit of the binary word representing the weight that is overfitted to a decoder.
[0074] Example 72. The apparatus of example 69, wherein the apparatus is further caused to perform the methods as described in any of the examples 55 to 68.
[0075] Example 73. A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform at least the following: receiving an overfitting signal; updating a decoder-side neural network, based at least on the overfitting signal, to obtain an updated decoder-side neural network, wherein for one or more weights of the decoderside neural network, at least one bit of a binary word representing a weight is overfitted; and using the updated decoder-side neural network for at least one task.
[0076] Example 74. The computer readable medium of example 73, wherein the computer readable medium comprises a non-transitory computer readable medium.
[0077] Example 75. The computer readable medium of any of the examples 73 or 74, wherein the computer readable medium causes the apparatus to further perform the methods as described in any of the examples 36 to 53.
[0078] Example 76. A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform at least the following: for one or more weights of a decoder-side neural network, overfitting at least one bit of a binary word representing a weight that is overfitted to obtain an overfitted weight; and signaling information related to the at least one bit of the binary word representing the weight that is overfitted to a decoder.
[0079] Example 77. The computer readable medium of example 76, wherein the computer readable medium comprises a non-transitory computer readable medium.
[0080] Example 78. The computer readable medium of any of the examples 76 or 77, wherein the computer readable medium causes the apparatus to further perform the methods as described in any of the examples 55 to 68.BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:
[0082] FIG. 1 shows schematically an electronic device employing embodiments of the examples described herein.
[0083] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.
[0084] FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.
[0085] FIG. 4 shows schematically a block chart of an encoder used for data compression on a general level.
[0086] FIG. 5 shows a system pipeline for video coding for machines (VCM).
[0087] FIG. 6 illustrates where a decoder-side neural network (DSNN) is overfitted based at least on an overfitting signal, obtaining an overfitted DSNN.
[0088] FIG. 7 illustrates that the overfitted DSNN may be used for its purpose, such as for filtering input data.
[0089] FIG. 8 is an example apparatus configured to implement the examples described herein.
[0090] FIG. 9 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.
[0091] FIG. 10 is an example method, based on the examples described herein.
[0092] FIG. 11 is another example method, based on the examples described herein.DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS
[0093] Described herein is a method and apparatus for implementing and using overfitting signal.
[0094] The following describes in detail a suitable apparatus and possible mechanisms for implementing and using overfitting signal according to embodiments. In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an apparatus 50. The apparatus may be an Internet of Things (loT) apparatus configured to perform various functions, such as for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 are explained next.
[0095] The electronic device 50 may for example be a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or other lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.
[0096] The apparatus 50 may comprise a housing 30 for incorporating and protecting the device.The apparatus 50 further may comprise a display 32 in the form of a liquid crystal display. In other embodiments of the examples described herein the display may be any suitable display technology suitable to display an image or video. The apparatus 50 may further comprise a keypad 34. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0097] The apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analog signal input. The apparatus 50 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 38, speaker, or an analog audio or digital audio output connection. The apparatus 50 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera 42 capable of recording or capturing images and / or video. The apparatus 50 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 50 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.
[0098] The apparatus 50 may comprise a controller 56, processor or processor circuitry for controlling the apparatus 50. The controller 56 may be connected to memory 58 which in embodiments of the examples described herein may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.
[0099] The apparatus 50 may further comprise a card reader 48 and a smart card 46, for example a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.
[0100] The apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).
[0101] The apparatus 50 may comprise a camera capable of recording or detecting individual frames which are then passed to the codec 54 or the controller for processing. The apparatus mayreceive the video image data for processing from another device prior to transmission and / or storage. The apparatus 50 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 50 described above represent examples of means for performing a corresponding function.
[0102] With respect to FIG. 3, an example of a system within which embodiments of the examples described herein can be utilized is shown. The system 10 comprises multiple communication devices which can communicate through one or more networks. The system 10 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.
[0103] The system 10 may include both wired and wireless communication devices and / or apparatus 50 suitable for implementing embodiments of the examples described herein.
[0104] For example, the system shown in FIG. 3 shows a mobile telephone network 11 and a representation of the internet 28. Connectivity to the internet 28 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.
[0105] The example communication devices shown in the system 10 may include, but are not limited to, an electronic device or apparatus 50, a combination of a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22, or a head-mounted apparatus 21, which head-mounted apparatus 21 may be a head-mounted display (HMD), or glasses having a camera or other device used for processing images and / or video. The apparatus 50 may be stationary or mobile when carried by an individual who is moving. The apparatus 50 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.
[0106] The embodiments may also be implemented in a set-top box; e.g. a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.
[0107] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that allows communication between the mobile telephone network 11 and the internet 28. The system may include additional communication devices and communication devices of various types.
[0108] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-internet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0109] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.
[0110] The embodiments may also be implemented in so-called loT devices. The Internet of Things (loT) may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home / building automation, etc. to be included in the Internet of Things (loT). In order to utilize the Internet loT devices are provided with an IP address as a unique identifier. loT devices may be provided with a radio transmitter, such as a WLAN or Bluetooth transmitter or a RFID tag. Alternatively, loT devices may have access to an IP-based network via a wired network, such as an Ethernet-based network or a power-line connection (PLC).
[0111] An MPEG-2 transport stream (TS), specified in ISO / IEC 13818-1 or equivalently in ITU- T Recommendation H.222.0, is a format for carrying audio, video, and other media as well as program metadata or other metadata, in a multiplexed stream. A packet identifier (PID) is used to identify an elementary stream (a.k.a. packetized elementary stream) within the TS. Hence, a logicalchannel within an MPEG-2 TS may be considered to correspond to a specific PID value.
[0112] Available media file format standards include ISO base media file format (ISO / IEC 14496-12, which may be abbreviated ISOBMFF) and file format for NAE unit structured video (ISO / IEC 14496-15), which derives from the ISOBMFF.
[0113] FIG. 4 shows a block diagram of a general structure of a video encoder. FIG. 4 presents an encoder for two layers, but it would be appreciated that presented encoder could be similarly extended to encode more than two layers. FIG. 4 illustrates a video encoder comprising a first encoder section 500 for a base layer and a second encoder section 502 for an enhancement layer. Each of the first encoder section 500 and the second encoder section 502 may comprise similar elements for encoding incoming pictures. The encoder sections 500, 502 may comprise a pixel predictor 302, 402, prediction error encoder 303, 403 and prediction error decoder 304, 404. FIG. 4 also shows an embodiment of the pixel predictor 302, 402 as comprising an inter-predictor 306, 406 (Pinter), an intra-predictor 308, 408 (Pintra), a mode selector 310, 410, a filter 316, 416 (F), and a reference frame memory 318, 418 (RFM). The pixel predictor 302 of the first encoder section 500 receives base layer images 300 (Io,n) of a video stream to be encoded at both the inter-predictor 306 (which determines the difference between the image and a motion compensated reference frame 318) and the intra-predictor 308 (which determines a prediction for an image block based only on the already processed parts of the current frame or picture). The output of both the inter-predictor and the intra-predictor are passed to the mode selector 310. The intra-predictor 308 may have more than one intra-prediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 310. The mode selector 310 also receives a copy of the base layer picture 300. Correspondingly, the pixel predictor 402 of the second encoder section 502 receives 400 enhancement layer images (Ii,n) of a video stream to be encoded at both the interpredictor 406 (which determines the difference between the image and a motion compensated reference frame 418) and the intra-predictor 408 (which determines a prediction for an image block based only on the already processed parts of the current frame or picture). The output of both the inter-predictor and the intra-predictor are passed to the mode selector 410. The intra-predictor 408 may have more than one intra-prediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 410. The mode selector 410 also receives a copy of the enhancement layer picture 400.
[0114] Depending on which encoding mode is selected to encode the current block, the output of the inter-predictor 306, 406 or the output of one of the optional intra-predictor modes or the output of a surface encoder within the mode selector is passed to the output of the mode selector 310, 410. The output of the mode selector is passed to a first summing device 321, 421. The first summing device may subtract the output of the pixel predictor 302, 402 from the base layer picture300 / enhancement layer picture 400 to produce a first prediction error signal 320, 420 (Dn) which is input to the prediction error encoder 303, 403.
[0115] The pixel predictor 302, 402 further receives from a preliminary reconstructor 339, 439 the combination of the prediction representation of the image block 312, 412 (P’n) and the output 338, 438 (D’n) of the prediction error decoder 304, 404. The preliminary reconstructed image 314, 414 (I’n) may be passed to the intra-predictor 308, 408 and to the filter 316, 416. The filter 316, 416 receiving the preliminary representation may filter the preliminary representation and output a final reconstructed image 340, 440 (R’n) which may be saved in a reference frame memory 318, 418. The reference frame memory 318 may be connected to the inter-predictor 306 to be used as the reference image against which a future base layer picture 300 is compared in inter-prediction operations. Subject to the base layer being selected and indicated to be the source for inter-layer sample prediction and / or inter-layer motion information prediction of the enhancement layer according to some embodiments, the reference frame memory 318 may also be connected to the inter-predictor 406 to be used as the reference image against which a future enhancement layer picture 400 is compared in inter-prediction operations. Moreover, the reference frame memory 418 may be connected to the inter-predictor 406 to be used as the reference image against which a future enhancement layer picture 400 is compared in inter-prediction operations.
[0116] Filtering parameters from the filter 316 of the first encoder section 500 may be provided to the second encoder section 502 subject to the base layer being selected and indicated to be the source for predicting the filtering parameters of the enhancement layer according to some embodiments.
[0117] The prediction error encoder 303, 403 comprises a transform unit 342, 442 (T) and a quantizer 344, 444 (Q). The transform unit 342, 442 transforms the first prediction error signal 320, 420 to a transform domain. The transform is, for example, the DCT transform. The quantizer 344, 444 quantizes the transform domain signal, e.g. the DCT coefficients, to form quantized coefficients.
[0118] The prediction error decoder 304, 404 receives the output from the prediction error encoder 303, 403 and performs the opposite processes of the prediction error encoder 303, 403 to produce a decoded prediction error signal 338, 438 which, when combined with the prediction representation of the image block 312, 412 at the second summing device 339, 439, produces the preliminary reconstructed image 314, 414. The prediction error decoder 304, 404 may be considered to comprise a dequantizer 346, 446 (Q1), which dequantizes the quantized coefficient values, e.g. DCT coefficients, to reconstruct the transform signal and an inverse transformation unit 348, 448 (T-1), which performs the inverse transformation to the reconstructed transform signal wherein the output of the inverse transformation unit 348, 448 includes reconstructed block(s). The predictionerror decoder may also comprise a block filter which may filter the reconstructed block(s) according to further decoded information and filter parameters.
[0119] The entropy encoder 330, 430 (E) receives the output of the prediction error encoder 303, 403 and may perform a suitable entropy encoding / variable length encoding on the signal to provide error detection and correction capability. The outputs of the entropy encoders 330, 430 may be inserted into a bitstream e.g. by a multiplexer 508 (M).
[0120] In some embodiments, the terms “picture”, “image”, and “frame” may be used interchangeably.
[0121] Fundamentals of neural networks
[0122] A neural network (NN) is a computation graph consisting of several layers of computation. Each layer consists of one or more units, where each unit performs an elementary computation. A unit is connected to one or more other units, and the connection may have associated with a weight. The weight may be used for scaling the signal passing through the associated connection. Weights are learnable parameters, e.g., values which can be learned from training data. There may be other learnable parameters, such as those of batch-normalization layers.
[0123] Two of the most widely used architectures for neural networks are feed-forward and recurrent architectures. Feed-forward neural networks are such that there is no feedback loop: each layer takes input from one or more of the layers before and provides its output as the input for one or more of the subsequent layers. Also, units inside a certain layer take input from units in one or more of preceding layers, and provide output to one or more of following layers.
[0124] Initial layers (those close to the input data) extract semantically low-level features such as edges and textures in images, and intermediate and final layers extract more high-level features. After the feature extraction layers there may be one or more layers performing a certain task, such as classification, semantic segmentation, object detection, denoising, style transfer, super-resolution, etc. In recurrent neural nets, there is a feedback loop, so that the network becomes stateful, e.g., it is able to memorize information or a state.
[0125] It is to be understood that, as described herein, the terms “machine vision”, “machine vision task”, “machine task”, “machine analysis”, “machine analysis task”, “computer vision”, “computer vision task”, "task network" and “task” may be used interchangeably.
[0126] Neural networks are being utilized in an ever-increasing number of applications for many different types of device, such as mobile phones. Examples include image and video analysis and processing, social media data analysis, device usage data analysis, etc.
[0127] The most important property of neural nets (and other machine learning tools) is that they are able to learn properties from input data, either in supervised way or in unsupervised way. Such learning is a result of a training algorithm, or of a meta-level neural network providing the training signal.
[0128] In general, the training algorithm consists of changing some properties of the neural network so that its output is as close as possible to a desired output. For example, in the case of classification of objects in images, the output of the neural network can be used to derive a class or category index which indicates the class or category that the object in the input image belongs to. Training usually happens by minimizing or decreasing the output’s error, also referred to as the loss. Examples of losses are mean squared error, cross-entropy, etc. In recent deep learning techniques, training is an iterative process, where at each iteration the algorithm modifies the weights of the neural net to make a gradual improvement of the network’s output, e.g., to gradually decrease the loss.
[0129] Herein, the terms “model”, “neural network”, “neural net” and “network” are used and described interchangeably, and also the weights of neural networks are sometimes referred to as learnable parameters or simply as parameters.
[0130] Training a neural network is an optimization process, but the final goal is different from the typical goal of optimization. In optimization, the only goal is to minimize a function. In machine learning, the goal of the optimization or training process is to make the model learn the properties of the data distribution from a limited training dataset. In other words, the goal is to learn to use a limited training dataset in order to learn to generalize to previously unseen data, e.g., data which was not used for training the model. This is usually referred to as generalization. In practice, data is usually split into at least two sets, the training set and the validation set. The training set is used for training the network, e.g., to modify its learnable parameters in order to minimize the loss. The validation set is used for checking the performance of the network on data which was not used to minimize the loss, as an indication of the final performance of the model. In particular, the errors on the training set and on the validation set are monitored during the training process to understand the following things (1-2):
[0131] 1. When the network is learning at all - in this case, the training set error should decrease, otherwise the model is in the regime of underfitting.
[0132] 2. When the network is learning to generalize - in this case, also the validation set error needs to decrease and to be not too much higher than the training set error. When the training set error is low, but the validation set error is much higher than the training set error, or it does notdecrease, or it even increases, the model is in the regime of overfitting. This means that the model has just memorized the training set’s properties and performs well only on that set, but performs poorly on a set not used for tuning its parameters.
[0133] Lately, neural networks have been used for compressing and de-compressing data such as images, e.g., in an image codec. The most widely used architecture for realizing one component of an image codec is the auto-encoder, which is a neural network consisting of two parts: a neural encoder and a neural decoder (these are referred to simply as encoder and decoder in this description, even though algorithms which are learned from data instead of being tuned by hand may be referred to herein). The encoder takes as input an image and produces a code which requires less bits than the input image. This code may be obtained by applying a binarization or quantization process to the output of the encoder. The decoder takes in this code and reconstructs the image which was input to the encoder.
[0134] Such encoder and decoder are usually trained to minimize a combination of bitrate and distortion, where the distortion may be based on one or more of the following metrics: Mean Squared Error (MSE), Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), or similar. These metrics are meant to be correlated to the human visual perception quality, so that minimizing or maximizing one or more of these metrics results into improving the visual quality of the decoded image as perceived by humans.
[0135] Fundamentals of video / image coding
[0136] Video codec consists of an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically encoder discards some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).
[0137] Typical hybrid video codecs, for example ITU-T H.263 and H.264, encode the video information in two phases. Firstly pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, e.g., the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder cancontrol the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).
[0138] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, exploits temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures (a.k.a. reference pictures).
[0139] In temporal inter prediction, the sources of prediction are previously decoded pictures in the same scalable layer. In intra block copy (IBC; a.k.a. intra-block-copy prediction), prediction may be applied similarly to temporal inter prediction but the reference picture is the current picture and only previously decoded samples can be referred in the prediction process. Inter-layer or inter-view prediction may be applied similarly to temporal inter prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal inter prediction only, while in other cases inter prediction may refer collectively to temporal inter prediction and any of intra block copy, inter-layer prediction, and interview prediction provided that they are performed with the same or similar process than temporal prediction. Inter prediction, temporal inter prediction, or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0140] Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, e.g., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0141] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently when they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0142] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the qualityof the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.
[0143] In typical video codecs the motion information is indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side) and the prediction source block in one of the previously coded or decoded pictures. In order to represent motion vectors efficiently those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or or co-located blocks in temporal reference picture. Moreover, typical high efficiency video codecs employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures and the used motion field information is signaled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.
[0144] In typical video codecs the prediction residual after motion compensation is first transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.
[0145] Typical video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g. the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor X to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + R where C is the Lagrangian cost to be minimized, D is the image distortion (e.g. Mean Squared Error)with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).
[0146] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI NAL units, and some video coding specifications include both prefix SEI NAL units and suffix SEI NAL units, where the former type can start a picture unit or alike and the latter type can end a picture unit or alike. An SEI NAL unit includes one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may include the syntax and semantics for the specified SEI messages but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient can be specified.
[0147] Scalable video coding
[0148] A scalable bitstream may include a "base layer" providing the lowest quality video available and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layers. In order to improve coding efficiency for the enhancement layers, the coded representation of that layer may depend on the lower layers. E.g., the motion and mode information of the enhancement layer can be predicted from lower layers. Similarly, the pixel data of the lower layers can be used to create prediction for the enhancement layer.
[0149] A scalable video codec for quality scalability (also known as SignaLto-Noise or SNR) and / or spatial scalability may be implemented as follows. For a base layer, a conventional non- scalable video encoder and decoder is used. The reconstructed / decoded pictures of the base layer are included in the reference picture buffer for an enhancement layer. In H.264 / AVC, HEVC, and similar codecs using reference picture list(s) for inter prediction, the base layer decoded pictures may be inserted into a reference picture list(s) for coding / decoding of an enhancement layer picturesimilarly to the decoded reference pictures of the enhancement layer. Consequently, the encoder may choose a base-layer reference picture as inter prediction reference and indicate its use e.g., with a reference picture index in the coded bitstream. The decoder decodes from the bitstream, for example from a reference picture index, that a base-layer picture is used as inter prediction reference for the enhancement layer. When a decoded base-layer picture is used as prediction reference for an enhancement layer, it is referred to as an inter-layer reference picture.
[0150] Scalability modes or scalability dimensions may include but are not limited to the following (1-9):
[0151] 1. Quality scalability: Base layer pictures are coded at a lower quality than enhancement layer pictures, which may be achieved for example using a greater quantization parameter value (e.g., a greater quantization step size for transform coefficient quantization) in the base layer than in the enhancement layer.
[0152] 2. Spatial scalability: Base layer pictures are coded at a lower resolution (e.g., have fewer samples) than enhancement layer pictures. Spatial scalability and quality scalability may sometimes be considered the same type of scalability.
[0153] 3. Bit-depth scalability: Base layer pictures are coded at lower bit-depth (e.g., 8 bits) than enhancement layer pictures (e.g., 10 or 12 bits).
[0154] 4. Dynamic range scalability: Scalable layers represent a different dynamic range and / or images obtained using a different tone mapping function and / or a different optical transfer function.
[0155] 5. Chroma format scalability: Base layer pictures provide lower spatial resolution in chroma sample arrays (e.g., coded in 4:2:0 chroma format) than enhancement layer pictures (e.g., 4:4:4 format).
[0156] 6. Color gamut scalability: enhancement layer pictures have a richer / broader color representation range than that of the base layer pictures - for example the enhancement layer may have UHDTV (ITU-R BT.2020) color gamut and the base layer may have the ITU-R BT.709 color gamut.
[0157] 7. Region-of-interest (ROI) scalability: An enhancement layer represents a spatial subset of the base layer. ROI scalability may be used together with other types of scalabilities, e.g., quality or spatial scalability so that the enhancement layer provides higher subjective quality for the spatial subset.
[0158] 8. View scalability, which may also be referred to as multiview coding. The base layerrepresents a first set of views, whereas an enhancement layer represents a second set of views.
[0159] 9. Depth scalability, which may also be referred to as depth-enhanced coding. A layer or some layers of a bitstream may represent texture view(s), while other layer or layers may represent depth view(s).
[0160] In all of the above scalability cases, base layer information could be used to code enhancement layer to minimize the additional bitrate overhead.
[0161] Scalability can be enabled in two basic ways. Either by introducing new coding modes for performing prediction of pixel values or syntax from lower layers of the scalable representation or by placing the lower layer pictures to the reference picture buffer (decoded picture buffer, DPB) of the higher layer. The first approach is more flexible and thus can provide better coding efficiency in most cases. However, the second, reference frame -based scalability, approach can be implemented very efficiently with minimal changes to single layer codecs while still achieving majority of the coding efficiency gains available. Essentially a reference frame -based scalability codec can be implemented by utilizing the same hardware or software implementation for all the layers, just taking care of the DPB management by external means.
[0162] In ROI scalability, spatial correspondence of an ROI enhancement layer in relation to its reference layer(s) is indicated. In VVC, scaling windows can be used to indicate this spatial correspondence.
[0163] It has been proposed, e.g. in JVET-O1150 (https: / / www.jvet- experts.org / doc_end_user / documents / 15_Gothenburg / wgl 1 / JVET-O1150-v2.zip), that temporal sublayers could be used for any type of scalability. A mapping of scalability dimensions to sublayer identifiers could be provided, e.g., in a VPS or in an SEI message.
[0164] Neural-network post-filter characteristics (NNPFC) and neural-network post-filter activation (NNPFA) SEI messages
[0165] The neural-network post-filter characteristics (NNPFC) SEI message and the neural- network post-filter activation (NNPFA) SEI message have been described in document JVET- AE2006.
[0166] The NNPFC SEI message comprises the nnpfc_id syntax element, which includes an identifying number that may be used to identify a post-processing filter. A base post-processing filter is the filter that is included in or identified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within a coded layer video sequence (CLVS). The NNPFC SEI message defining a base post-processing filter has nnpfc_base_flag equal to 1. When an NNPFCSEI message is neither the first NNPFC SEI message, in decoding order, in the current CLVS nor a repetition of the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within the current CLVS, the NNPFC SEI message defines an update relative to the base postprocessing filter, and the update relative to the base post-processing filter is applied to obtain a postprocessing filter associated with the nnpfc_id value. The NNPFC SEI message defining an update relative to a base post-processing filter has nnpfc_base_flag equal to 0. The update may be obtained by decoding the coded neural network bitstream in the second NNPFC SEI message (when nnpfc_mode_idc is equal to 0) or through the Uniform Resource Identifier defining the update (when nnpfc_mode_idc is equal to 1). Otherwise (e.g., when there is no update defined by an NNPFC SEI message), the post-processing filter associated with the nnpfc_id value is assigned to be the same as the base post-processing filter.
[0167] The NNPFC SEI message comprises nnpfc_mode_idc syntax element, the semantics of which may be defined as follows (1-2):
[0168] 1. nnpfc_mode_idc equal to 1 specifies that the base post-processing filter or the update relative to the base post-processing filter associated with the nnpfc_id value is a neural network identified by the Uniform Resource Identifier (URI) nnpfc_uri with the format identified by the tag URI nnpfc_tag_uri.
[0169] 2. nnpfc_mode_idc equal to 0 indicates that this SEI message includes an ISO / IEC 15938- 17 bitstream that specifies the base post-processing filter or updates relative to the base postprocessing filter with the same nnpfc_id value.
[0170] The NNPFC SEI message may also comprise (1-4):
[0171] 1. Purpose of the post-processing filter, which may comprise one or more of the following:Visual quality improvement, Chroma upsampling from the 4:2:0 chroma format to the 4:2:2 or 4:4:4 chroma format, or from the 4:2:2 chroma format to the 4:4:4 chroma format, Resolution resampling, Frame rate upsampling, Bit depth upsampling, Colorization.
[0172] 2. Formatting of the input tensors that are given as input to the neural network inference.
[0173] 3. Formatting of the output tensors that are resulting from the neural network inference.
[0174] 4. Characterization of the complexity of the neural network
[0175] The NNPFA SEI message specifies the neural-network post-processing filter that may be used for post-processing filtering for the current picture, or for post-processing filtering for the current picture and one or more other pictures. The NNPFA SEI message comprises thennpfa_target_id syntax element, which indicates that the neural-network post-processing filter with nnpfc_id equal to nnfpa_target_id may be used for post-processing filtering for the indicated persistence. The indicated persistence may be the current picture only (nnpfa_persistence_flag equal to 0), or until the end of the current CLVS or the next picture, in output order, in the current layer associated with a NNPFA SEI message with the same nnpfa_target_id as the current SEI message (nnpfa_persistence_flag equal to 1).
[0176] It is to be understood that, as described herein, the terms “post-filter”, “post-processing filter” and “postprocessing filter” may be used interchangeably.
[0177] Background information on Video Coding for Machines (VCM)
[0178] Reducing the distortion in image and video compression is often intended to increase human perceptual quality, as humans are considered to be the end users, e.g. consuming / watching the decoded image. Recently, with the advent of machine learning, especially deep learning, there is a rising number of machines (e.g., autonomous agents) that analyze data independently from humans and that may even take decisions based on the analysis results without human intervention. Examples of such analysis are object detection, scene classification, semantic segmentation, video event detection, anomaly detection, pedestrian tracking, etc. Example use cases and applications are selfdriving cars, video surveillance cameras and public safety, smart sensor networks, smart TV and smart advertisement, person re-identification, smart traffic monitoring, drones, etc. This may raise the following question: when decoded data is consumed by machines, shouldn’t the goal be to aim at a different quality metric -other than human perceptual quality- when considering media compression in inter-machine communications? Also, dedicated algorithms for compressing and decompressing data for machine consumption are likely to be different than those for compressing and decompressing data for human consumption. The set of tools and concepts for compressing and decompressing data for machine consumption is referred to here as Video Coding for Machines.
[0179] It is to be understood that, as described herein, the terms “machine consumption” and “machine analysis” may be used interchangeably.
[0180] It is likely that the receiver-side device has multiple “machines” or neural networks (NNs). These multiple machines may be used in a certain combination which is for example determined by an orchestrator sub-system. The multiple machines may be used for example in succession, based on the output of the previously used machine, and / or in parallel. For example, a video which was compressed and then decompressed may be analyzed by one machine (NN) for detecting pedestrians, by another machine (another NN) for detecting cars, and by another machine (another NN) for estimating the depth of all the pixels in the frames.
[0181] Also, please notice that, when describing aspects or embodiments related to video coding for machines, the term “receiver-side” or “decoder-side” may be used to refer to the physical or abstract entity or device which includes one or more machines, and runs these one or more machines on some encoded and eventually decoded video representation which is encoded by another physical or abstract entity or device, the “encoder-side device”.
[0182] The encoded video data may be stored into a memory device, for example as a file. The stored file may later be provided to another device.
[0183] Alternatively, the encoded video data may be streamed from one device to another.
[0184] FIG. 5 is a general illustration of the pipeline 501 of Video Coding for Machines. A VCM encoder 504 encodes the input video 503 into a bitstream 506. A bitrate 510 may be computed 509 from the bitstream 506 in order to evaluate the size of the bitstream 506. A VCM decoder 512 decodes the bitstream output 506 by the VCM encoder 504. The output 514 of the VCM decoder 512 is referred in FIG. 5 as “Decoded data for machines”. This data 514 may be considered as the decoded or reconstructed video. However, in some implementations of this pipeline 501, this data 514 may not have the same or similar characteristics as the original video 503 which was input to the VCM encoder 504. For example, this data 514 may not be easily understandable by a human by simply rendering the data onto a screen. The output 514 of VCM decoder 512 is then input to one or more task neural networks (516, 518, 520, 522). In FIG. 5, for the sake of illustrating that there may be any number of task-NNs, there are three example task-NNs, namely a task -NN 516 for object detection, a task-NN 518 for object segmentation, a task-NN 3 for object tracking, and a nonspecified one (Task-NN X 522). The goal of VCM is to obtain a low bitrate while guaranteeing that the task-NNs (516, 518, 520, 522) still perform well in terms of the evaluation metric associated to each task.
[0185] As shown in FIG. 5, a performance (532) of the first task (e.g., object detection) is evaluated (524) and, a performance (534) of the second task (e.g., object segmentation) is evaluated (526), a performance (536) of the third task (e.g., object tracking) is evaluated (528), and a performance (538) of the unspecified task is evaluated (530). The evaluated performances (532, 534, 536, 538) are collectively given as 540.
[0186] When a conventional video encoder, such as a H.266 / VVC encoder, is used as a VCM encoder, one or more of the following approaches may be used to adapt the encoding to be suitable to machine analysis tasks (1-4):
[0187] 1. One or more regions of interest (ROIs) may be detected. Any ROI detection method may be used. For example, ROI detection may be performed using a task NN, such as an objectdetection NN. In some cases, ROI boundaries of a group of pictures or an intra period may be spatially overlaid and rectangular areas may be formed to cover the ROI boundaries. The detected ROIs (or rectangular areas, likewise) may be used in one or more of the following ways (i-iii): i) The quantization parameter (QP) may be adjusted spatially in a manner that ROIs are encoded using finer quantization step size(s) than other regions. For example, QP may be adjusted CTU-wise, ii) The video is preprocessed to include only the ROIs, while the other areas are replaced by one or more constant values or removed, iii) A grid is formed in a manner that a single grid cell covers a ROI. Grid rows or grid columns that include no ROIs are downsampled as preprocessing to encoding.
[0188] 2. Quantization parameter of the highest temporal sublayer(s) is increased (e.g. coarser quantization is used) when compared to practices for human watchable video.
[0189] 3. The original video is temporally downsampled as preprocessing prior to encoding. A frame rate upsampling method may be used as postprocessing subsequent to decoding, when machine analysis at the original frame rate is desired.
[0190] 4. A filter is used to preprocess the input to the conventional encoder. The filter may be a machine learning based filter, such as a convolutional neural network.
[0191] Neural network based filtering
[0192] In some video codecs, a neural network may be used as filter in the decoding loop, and it may be referred to as neural network loop filter, or neural network in-loop filter. The NN loop filter may replace all other loop filters of an existing video codec, or may represent an additional loop filter with respect to the already present loop filters in an existing video codec.
[0193] In the context of image and video enhancement, a neural network may be used as postprocessing filter, for example applied to the output of an image or video decoder in order to remove or reduce coding artifacts.
[0194] The proposed embodiments described herein concern the architecture of neural networks used as part of the decoding operations (such as a NN loop filter, or an intra-frame prediction NN, or an inter-frame prediction NN) or as part of post-processing operations (a NN post-processing filter). Also, it concerns the signaling of information related to those NNs, where the information is signaled from an encoder to a decoder.
[0195] The following example system will be used in several embodiments to illustrate or describe the idea. The example system comprises a codec that comprises one or more NN loop filters. For example, the codec could comprise a modified VVC / H.266 compliant codec (e.g., a VVC / H.266compliant codec that has been modified so that it would comprise one or more NN loop filters). The input to the one or more NN loop filters may comprise at least a reconstructed block or frames (simply referred to as reconstruction) or data derived from a reconstructed block or frame (e.g., the output of a conventional loop filter). The reconstruction may be obtained based on predicting a block or frame (e.g., by means of intra-frame prediction or inter-frame prediction) and performing residual compensation. The one or more NN loop filters (may be referred to simply as NN filters in some of the embodiments) may enhance the quality of at least one of their input, so that a rate-distortion loss is decreased. The rate may indicate a bitrate (estimate or real) of the encoded video. The distortion may indicate a pixel fidelity distortion such as the following:- mean-squared error (MSE);- mean absolute error (MAE);- mean average precision (mAP) computed based on the output of a task NN (such as an object detection NN) when the input is the output of the post-processing NN; and- Other machine task-related metric, for tasks such as object tracking, video activity classification, video anomaly detection, and the like.
[0196] The enhancement may result into a coding gain, which can be expressed for example in terms of BD-rate, BD-PSNR, or BD-mAP.
[0197] However, at least some of the embodiments described herein are applicable to a NN filter which is not a loop filter of a codec. For example, the NN filter may be a NN post-processing filter, whose input may comprise one or more outputs of a video codec. In this case, the filter may be used only for increasing a quality metric for at least one of its inputs, where the quality metric may be, for example, peak signal-to-noise ratio (PSNR), mAP for object detection, MOTA for object tracking, etc.
[0198] Information on overfitting a decoder-side neural network
[0199] Various embodiments consider the case where a data unit is encoded by an encoder and decoded by a decoder.
[0200] Some examples of data units are:- One or more video sequences;- One or more frames of one or more video sequences.; or- One or more blocks or regions of one or more frames.
[0201] The decoder or the receiver comprising the decoder is assumed to comprise at least a neural network, referred to as decoder- side neural network (DSNN). Various embodiments consider the case where the DSNN is optimized for a certain data unit, with respect to one or more metrics.
[0202] Some examples of metrics are:- Peak signal-to-noise ratio (PSNR);- Mean- squared error (MSE); and- Mean average precision (mAP).
[0203] Herein, the process of optimizing a NN on a certain data unit with respect to one or more metrics may be referred to as overfitting, specializing, or finetuning.
[0204] The NN which is optimized may have been pretrained (e.g., may have been previously trained on a training dataset).
[0205] Overfitting of a DSNN may be performed by computing a weight-update at encoder side by means of one or more training iterations, e.g., based on backpropagation. Then, the obtained weight update is compressed by a neural network encoder and provided to a neural network decoder. The neural network encoder may be part of the encoder of data units, such as a video encoder; the neural network decoder may be part of the decoder of data units, such as a video decoder. The decoder decompresses the compressed weight-update, uses the decompressed weight-update for updating the DSNN, and uses the updated DSNN for its purpose, such as for decoding a data unit or for post-processing a decoded data unit; for example, the updated DSNN may be used for decoding a video frame or for post-processing a decoded video frame or data derived from the decoded video frame.
[0206] In some embodiments, the term ’overfitting’ may be used to refer to the operation performed by the decoder side or receiver side for updating the DSNN based on the decompressed weight-update.
[0207] Sending a weight update from encoder to decoder may cause a bitrate increase, or bitrate overhead, with respect to the bitrate required for sending an encoded data unit (e.g., an encoded video).
[0208] An example problem addressed by at least some embodiments herein is how to reduce such bitrate overhead or, in other words, how to generate a low-bitrate bitstream representing a weight update.
[0209] It is assumed that one or more weights or parameters of a DSNN are updated or overfitted. In an example, all the parameters of the DSNN are overfitted. In another example, a subset of parameters of the DSNN are overfitted. In yet another example, the bias parameters of the DSNN are overfitted. In yet another example, a set of multipliers or scaling parameters are overfitted, where the set of multipliers or scaling parameters multiply or scale at least some of the outputs of the layers of the DSNN.
[0210] In an embodiment, for each overfitted weight, only one bit of the binary word representing the weight is overfitted.
[0211] In an additional embodiment, the encoder may send an indication of the value of the bit that was overfitted. In one example, the bit that was overfitted had an original value of 0 (before overfitting), the overfitted value is 1, thus the indication comprises a bit with value 1. In another example, the bit that was overfitted had an original value of 1, the overfitted value is 0, thus the indication comprises a bit with value 0.
[0212] In an additional embodiment, the position of the bit (within the binary word of the weight) that is overfitted may be predetermined, and thus already known at decoder side. For example, it may be predetermined based on statistics collected during an offline phase where overfitting is performed on a dataset.
[0213] In an alternative additional embodiment, the position of the bit that is overfitted may not be predetermined, and thus an indication of the position of the bit within the binary word may be signaled from encoder to decoder.
[0214] In a yet alternative embodiment, the position of the bit that is overfitted may be predetermined, and the encoder may signal an indication of the position of the bit, such as an adjustment or modification of the position of the bit.
[0215] In an additional embodiment, the encoder may send an indication that the bit that was overfitted was changed (e.g., from 0 to 1, or from 1 to 0). In an example, the indication may comprise one bit and is referred to as change-indication bit: when the change-indication bit is 0, the original bit of the weight is left unmodified, whereas when the change-indication bit is 1, the bit of the weight is flipped (e.g., from 0 to 1, or from 1 to 0). In this example, the indication bit represents an indication of whether to apply a bitwise NOT operation to the bit of the weight.
[0216] In an additional embodiment, the change-indication bit that is sent from encoder to decoder may be used as part of another bitwise operation than a bitwise NOT operation. For example, AND, OR, XOR, and the like, may be applied to the change-indication bit and the bit of the binary word representing the weight. In an example, the bitwise operation is predetermined. In another example, the bitwise operation to be used for a weight or for a group of weights (e.g., for the weights in a layer) is signaled from encoder to decoder. In one example, for each overfitted weight or for a group of overfitted weights, the encoder signals one or more of the following: a change-indication bit, the position of the bit that was overfitted, an indication of which bitwise operation to apply to the change-indication bit and the bit of the weight in order to obtain a resulting value for the bit of the weight.
[0217] In an embodiment, the operation that is performed on the binary word representing the weight, based on the change-indication bit, may be applied to all bits of the binary word.
[0218] In an embodiment, the signaled change-indication information, which may be a single bit, is converted at decoder side into a change-indication binary word where the element or bit for which the position is signaled is set to the signaled change-indication bit and all other elements or bits are set to zero. The obtained change-indication binary word is then used when applying a certain operation (such as bit-wise NOT, or bit-wise XOR) to the binary word representing the weight.
[0219] In an embodiment, all the parameters that are comprised in a layer or a group of layers and that are overfitted share the same predetermined position of the bit that is overfitted. In an example, for a first layer LI, only the set of bits in position 2 (assuming a certain order, such as from least important bit to most important bit) in all the bias parameters is overfitted, whereas for a second layer L2, only the set of bits in position 5 in all the bias parameters is overfitted.
[0220] In an embodiment, for each weight or for a group of weights (such as for all weights in a layer), a number B of bits and their positions are predetermined, for example as the bits that change most commonly when overfitting on a dataset. The encoder may signal an indication of the values of those bits, or an indication of whether those bits are to be flipped.
[0221] In an embodiment, for each overfitted weight, the binary word representing the weight update is lossless coded and sent from encoder to decoder, and a bitwise operation may be used to obtain the overfitted weight. In one example, the run-length coding algorithm with fixed word-length may be utilized as lossless coding tool. In another example, the run-length coding algorithm may utilize different word-lengths for different overfitted weights and this word-length may be signaled from encoder to decoder.
[0222] An example advantage of at least some of the embodiments above is that, at decoder side, there is no need to perform arithmetic operations (e.g., adding a weight update to the pretrained weight). Instead, only bitwise operations are performed.
[0223] Candidate weight-updates
[0224] In an embodiment, a set of candidate weight updates are predetermined for each weight (or each layer) that is overfitted, and is known at both encoder side and decoder side. The encoder may signal an indication of which of the candidate weight updates in the set applies to a certain weight or to the weights in a certain layer. The set may comprise one or more candidate weight updates. When the set comprises only one candidate weight update, the encoder may signal to the decoder an indication of whether the candidate weight update is to be used for updating the original weight(s). The selected weight update may be then added to the original weight(s).
[0225] The predetermination of the candidate weight updates may be based on determining one or more weight-updates based on a dataset. In one example, the predetermination includes overfitting a DSNN on the data samples in the dataset, where for each data sample an overfitted DSNN is obtained, and thus for each data sample and each weight a weight update is obtained; for the weights in a certain layer, one or more most common weight-updates are determined and included into the set of candidate weight updates for that layer; a “common weight update” may be determined as a weight update that is similar to or same as one or more of the weight updates obtained based on respective one or more data samples; alternatively, a “common weight update” may be determined as the average of all the weight updates obtained based on the respective data samples.
[0226] In additional embodiment, the encoder may signal to the decoder an indication of whether an indicated candidate weight update should be added to or subtracted from the original weight.
[0227] In an alternative additional embodiment, the sign of the weight update (which determines whether the weight update is added to or subtracted from the original weight) is predetermined based for example on a dataset.
[0228] General information
[0229] For the sake of simplicity, at least some embodiments are described herein as applied to a filter. A filter may take as input at least one or more first images to be filtered and may output at least one or more second images, where the one or more second images may be the filtered versionof the one or more first images. In an example, the filter takes as input one image and outputs one image. In another example, the filter takes as input more than one image and outputs one image. In yet another example, the filter takes as input more than one image and outputs more than one image.
[0230] It is to be understood that a filter may take as input also other data (also referred to as auxiliary data) than the data that is to be filtered, such as data that can aid the filter to perform a better filtering than when no auxiliary data was provided as input. In an example, the auxiliary data comprises information about prediction data, and / or information about the picture type, and / or information about the slice type, and / or information about a Quantization Parameter (QP) used for encoding, and / or information about boundary strength, and the like. In an example, the filter takes as input one image and other data associated to that image, such as information about the quantization parameter (QP) used for quantizing and / or dequantizing that image, and outputs one image.
[0231] A filter may be a neural network based filter, or may be another type of filter. However, several embodiments describe training features which may be applicable to machine learning based filters.
[0232] A filter may be, for example, an in-loop filter that is used in the decoding loop of a codec, or a post-processing filter applied on the data decoded by the codec.
[0233] Even when at least some of the embodiments are described with reference to a filter, those embodiments may be applied also to other operations than just filters, such as an operation performing intra-frame prediction, or an operation performing inter-frame prediction, or an operation performing frame -rate upsampling, or an operation performing encoding and / or decoding (e.g., an end-to-end learned codec). The described embodiments may also be applied to components in an end-to-end learned image / video codec, for example, a decoder network, an optical flow estimation network, or a probability model neural network.
[0234] While at least some embodiments are described such that the input and output data are in the form of images or (video) frames or pictures, those embodiments may be applicable also to other types of data, such as audio frames. Furthermore, while at least some embodiments are described by considering a full image, those embodiments may be applicable also when considering one or more blocks or portions of an image.
[0235] It is to be noticed that, for simplicity, at least some of the figures do not include visual representations of other components that may be present at encoder side and / or at decoder side. In an example, in case the DSNN is a decoder neural network which is part of an end-to-end learnedimage codec, a figure may not include information about some of the encoder-side operations performed by the encoder of the end-to-end learned image codec, such as an encoder neural network to process an input image into a latent tensor or a lossless encoder to encode a latent tensor into a bitstream, and / or may not include information about some of the decoder-side operations performed by the decoder of the end-to-end learned image codec, such as a lossless decoder that decodes a bitstream representing an encoded image. In another example, in case the DSNN is a neural network based loop filter that is part of video decoding loop, a figure may not depict some of the video encoding operations and / or some of the video decoding operations.
[0236] Several embodiments are described herein, where a decoder-side neural network (DSNN) is overfitted based at least on an overfitting signal, obtaining an overfitted DSNN. The aim of the overfitting process is that the overfitted DSNN performs better than the DSNN with respect to a predefined metric, on at least some input data. See the following illustration.
[0237] Referring to FIG. 6, several embodiments are described herein, where a decoder-side neural network (DSNN) 602 is overfitted 604 based at least on an overfitting signal 606, obtaining an overfitted DSNN 608. An example objective of the overfitting process is that the overfitted DSNN performs better than the DSNN with respect to a predefined metric, on at least some input data.
[0238] Referring to FIG. 7, after a DSNN has been overfitted, the overfitted DSNN may be used (704) for its purpose, such as for filtering input data 702 to obtain filtered data 706.
[0239] A DSNN is a neural network which is used in a decoder or as a post -processing operation. Some examples of a DSNN used in a decoder are the following (1-3): 1. An NN loop filter that is comprised in a conventional decoder or in an end-to-end learned codec, 2. An NN post-filter that follows, in processing order, an inner decoder, such as a conventional decoder or an end-to-end learned codec, 3. An NN decoder, which is comprised in an end-to-end learned codec.
[0240] Several embodiments herein are described with reference to a DSNN that performs filtering of at least a portion of the input data. In some of the embodiments and examples, the at least portion of input data to be filtered by the DSNN is referred to as “data to be filtered”. However, the DSNN may take other inputs than the data to be filtered, and those other inputs may not be mentioned in the description of some embodiments or examples for the sake of simplicity. Even though several embodiments herein are described with reference to a DSNN that performs filtering, the same embodiments may be applicable to or valid for a DSNN that performs other purposes or tasks, such as a neural network decoder that is part of an end-to-end learned codec.
[0241] The DSNN may be pretrained, e.g., may have been previously trained on a training dataset, or may be initialized in some other suitable way, such as by random or pseudo-randominitialization.
[0242] In some of the embodiments described herein, the DSNN, or a copy of the DSNN, or a neural network which is same or substantially the same as the DSNN, is available at encoder side, and it may be referred to simply as DSNN even when meaning a DSNN present at encoder side. In other words, in at least some embodiments, both the encoder and the decoder may comprise the same DSNN or two copies of the same DSNN.
[0243] It is to be understood that at least some of the embodiments may use the term decoder to refer to an entity which performs decoding and some post-processing operations, or to a player, or to a receiver.
[0244] Overfitting a DSNN refers to optimizing the DSNN for a certain data unit, with respect to one or more metrics. Examples of data units are: one or more video sequences, one or more frames of one or more video sequences, and / or one or more blocks or regions of one or more frames.
[0245] Examples of metrics are: Peak signal-to-noise ratio (PSNR), Mean-squared error (MSE), Mean average precision (mAP).
[0246] Herein, the process of optimizing a DSNN on a certain data unit with respect to one or more metrics may be referred to as overfitting, or specializing, or finetuning. The DSNN which is optimized may be pretrained (e.g., may have been previously trained on a training dataset).
[0247] It is to be noted that a DSNN that has been overfitted or optimized on a certain data unit may be then used for processing (e.g., filtering, decoding, post-processing) one or more of the following data: the same data unit used for overfitting, and / or other data unit than the data unit used for overfitting.
[0248] In an embodiment, the overfitting signal may comprise updated weights. The updated weights may comprise one or more updated values associated to respective one or more weights or parameters of the DSNN. Overfitting a DSNN based on updated weights may comprise replacing one or more weights or parameters of the DSNN with respective one or more updated values comprised in the updated weights.
[0249] In an example, the updated weights comprise one or more updated values associated to respective one or more multiplying parameters, where the one or more multiplying parameters are comprised in a DSNN, and where each of the one or more multiplying parameters multiply an output of a layer of the DSNN.
[0250] In another example, the updated weights comprise one or more updated values associatedto respective one or more bias parameters, where the one or more bias parameters are comprised in a DSNN, and where each of the one or more bias parameters are added to an output of a layer of the DSNN.
[0251] In another embodiment, the overfitting signal may comprise a weight-update. The weightupdate may comprise one or more update values associated to respective one or more weights or parameters of the DSNN, where each update value represents an update or change to a respective weight or parameter of the DSNN. Overfitting a DSNN based on a weight-update may comprise adding or subtracting (or other suitable operation) the weight-update to one or more weights or parameters of the DSNN.
[0252] In an example, the weight-update comprises one or more updates associated to respective one or more multiplying parameters, where the one or more multiplying parameters are comprised in a DSNN, and where each of the one or more multiplying parameters multiply an output of a layer of the DSNN.
[0253] In another example, the weight-update comprises one or more updates associated to respective one or more bias parameters, where the one or more bias parameters are comprised in a DSNN, and where each of the one or more bias parameters are added to an output of a layer of the DSNN.
[0254] In another embodiment, the overfitting signal may comprise a modulating signal. The modulating signal may comprise one or more modulating values, or one or more sets of modulating values, associated to respective one or more outputs of one or more layers of a DSNN. Overfitting a DSNN based on a modulating signal may comprise modifying the one or more outputs based at least on the one or more modulating values, or the one or more sets of modulating values, and a modulation operation.
[0255] In an example, a modulating signal comprises one or more sets of modulating values associated to respective one or more channels of tensors output by one or more convolutional layers of a DSNN, where each set of modulating values comprises a multiplicative value and an additive value, and where the modulation operation comprises multiplying the multiplicative value with the respective channel and adding the additive value to the result of the multiplication for that channel.
[0256] In some embodiments, the term ’overfitting’ may be used to refer to the operation performed by the decoder side or receiver side for updating the DSNN based on an overfitting signal or data derived from an overfitting signal such as a decompressed overfitting signal.
[0257] Example Embodiments
[0258] It is assumed that one or more weights or parameters of a DSNN are updated or overfitted. In one example, all the parameters of the DSNN are overfitted. In another example, only the bias parameters of the DSNN are overfitted.
[0259] In one embodiment, for each overfitted weight, only one bit of the binary word representing the weight is overfitted.
[0260] In an additional embodiment, the encoder may send an indication of the value of the bit that was overfitted. In one example, the bit that was overfitted had an original value of 0 (before overfitting), the overfitted value is 1, thus the indication comprises a bit with value 1. In another example, the bit that was overfitted had an original value of 1, the overfitted value is 0, thus the indication comprises a bit with value 0.
[0261] In one additional embodiment, the position of the bit (within the binary word of the weight) that is overfitted may be predetermined and thus already known at decoder side. For example, it may be predetermined based on statistics collected during an offline phase where overfitting is performed on a dataset.
[0262] In an alternative additional embodiment, the position of the bit that is overfitted may not be predetermined, and thus an indication of the position of the bit within the binary word may be signaled from encoder to decoder.
[0263] In a yet alternative embodiment, the position of the bit that is overfitted may be predetermined, and the encoder may signal an indication of the position of the bit, such as an adjustment or modification of the position of the bit.
[0264] In an additional embodiment, the encoder may send an indication that the bit that was overfitted was changed (e.g., from 0 to 1, or from 1 to 0). In one example, the indication may comprise one bit and is referred to as change-indication bit: when the change-indication bit is 0, the original bit of the weight is left unmodified, whereas when the change-indication bit is 1, the bit of the weight is flipped (i.e., from 0 to 1, or from 1 to 0). In this example, the indication bit represents an indication of whether to apply a bitwise NOT operation to the bit of the weight.
[0265] In an additional embodiment, the change-indication bit that is sent from encoder to decoder may be used as part of another bitwise operation than a bitwise NOT operation. For example,AND, OR, XOR, etc., could be applied to the change-indication bit and the bit of the binary word representing the weight. In another additional embodiment, the bitwise operation is predetermined. In yet another additional embodiment, the bitwise operation to be used for a weight or for a group of weights (e.g., for the weights in a layer) is signaled from encoder to decoder. In one example, for each overfitted weight or for a group of overfitted weights, the encoder signals one or more of the following: a change-indication bit, the position of the bit that was overfitted, an indication of which bitwise operation to apply to the change-indication bit and the bit of the weight in order to obtain a resulting value for the bit of the weight.
[0266] In one embodiment, all the parameters that are comprised in a layer or a group of layers and that are overfitted share the same predetermined position of the bit that is overfitted. In one example, for a first layer LI, only the set of bits in position 2 (assuming a certain order, such as from least important bit to most important bit) in all the bias parameters is overfitted, whereas for a second layer L2, only the set of bits in position 5 in all the bias parameters is overfitted.
[0267] In one embodiment, for each weight or for a group of weights (such as for all weights in a layer), a number B of bits and their positions are predetermined, for example as the bits that change most commonly when overfitting on a dataset. The encoder may signal an indication of the values of those bits, or an indication of whether those bits are to be flipped.
[0268] In one embodiment, for each overfitted weight, the binary word representing the weight update is lossless coded and sent from encoder to decoder, and a bitwise operation may be used to obtain the overfitted weight. In one example, the run-length coding algorithm with fixed word-length could be utilized as lossless coding tool. In another example, the run-length coding algorithm could utilize different word-lengths for different overfitted weights and this word-length could be signaled from encoder to decoder.
[0269] One claimed advantage of at least some of the embodiments above is that, at decoder side, there is no need to perform arithmetic operations (e.g., adding a weight update to the pretrained weight). Instead, only bitwise operations are performed.
[0270] Candidate weight-updates
[0271] In one embodiment, a set of candidate weight updates are predetermined for each weight (or each layer) that is overfitted, and is known at both encoder side and decoder side. The encoder may signal an indication of which of the candidate weight updates in the set applies to a certain weight or to the weights in a certain layer. The set may comprise one or more candidate weightupdates. When the set comprises only one candidate weight update, the encoder may signal to the decoder an indication of whether the candidate weight update is to be used for updating the original weight(s). The selected weight update may be then added to the original weight(s).
[0272] The predetermination of the candidate weight updates may be based on determining one or more weight-updates based on a dataset. In one example, the predetermination consists of overfitting a DSNN on the data samples in the dataset, where for each data sample an overfitted DSNN is obtained, and thus for each data sample and each weight a weight update is obtained; for the weights in a certain layer, one or more most common weight-updates are determined and included into the set of candidate weight updates for that layer; a “common weight update” may be determined as a weight update that is similar to or same as one or more of the weight updates obtained based on respective one or more data samples; alternatively, a “common weight update” may be determined as the average of all the weight updates obtained based on the respective data samples.
[0273] In one additional embodiment, the encoder may signal to the decoder an indication of whether an indicated candidate weight update should be added to or subtracted from the original weight.
[0274] In one alternative additional embodiment, the sign of the weight update (which determines whether the weight update is added to or subtracted from the original weight) is predetermined based for example on a dataset.
[0275] In some embodiments, the terms “picture”, “image”, and “frame” may be used interchangeably.
[0276] It is to be understood that, as described herein, the terms “machine vision”, “machine vision task”, “machine task”, “machine analysis”, “machine analysis task”, “computer vision”, “computer vision task”, "task network" and “task” may be used interchangeably.
[0277] It is to be understood that, as described herein, the terms “machine consumption” and “machine analysis” may be used interchangeably.
[0278] It is to be understood that, as described herein, the terms “post-filter”, “post-processing filter” and “postprocessing filter” may be used interchangeably.
[0279] FIG. 8 is an example apparatus 800, which may be implemented in hardware, configuredto implement the examples described herein. The apparatus 800 comprises at least one processor 802 (e.g., an FPGA and / or CPU), one or more memories 804 including computer program code 805, the computer program code 805 having instructions to carry out the methods described herein, wherein the at least one memory 804 and the computer program code 805 are configured to, with the at least one processor 802, cause the apparatus 800 to implement circuitry, a process, component, module, or function (implemented with control module 806) to implement the examples described herein, including a learned overfitting signal. Optionally included encoder 830 of the control module 806 performs encoding, and optionally included decoder 840 implements decoding. The memory 804 may be a non-transitory memory, a transitory memory, a volatile memory (e.g. RAM), or a nonvolatile memory (e.g., ROM).
[0280] The apparatus 800 includes a display and / or I / O interface 808, which includes user interface (UI) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one or more sensors, etc. The apparatus 800 includes one or more communication e.g. network (N / W) interfaces (I / F(s)) 810. The communication I / F(s) 810 may be wired and / or wireless and communicate over the Intemet / other network(s) via any communication technique including via one or more links 824. The communication I / F(s) 810 may comprise one or more transmitters or one or more receivers.
[0281] The transceiver 816 comprises one or more transmitters 818 and one or more receivers 820. The transceiver 816 and / or communication I / F(s) 810 may comprise standard well-known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 814 used for communication over wireless link 88.
[0282] The control module 806 of the apparatus 800 comprises one of or both parts 806-1 and / or 806-2, which may be implemented in a number of ways. The control module 806 may be implemented in hardware as control module 806-1, such as being implemented as part of the one or more processors 802. The control module 806-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 806 may be implemented as control module 806-2, which is implemented as computer program code (having corresponding instructions) 805 and is executed by the one or more processors 802. For instance, the one or more memories 804 store instructions that, when executed by the one or more processors 802, cause the apparatus 800 to perform one or more of the operations as described herein. Furthermore, the one or more processors 802, one or more memories 804, andexample algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.
[0283] The apparatus 800 to implement the functionality of control 806 may correspond to any of the apparatuses depicted herein. Alternatively, apparatus 800 and its elements may not correspond to any of the other apparatuses depicted herein, as apparatus 800 may be part of a self- organizing / optimizing network (SON) node or other node, such as a node in a cloud.
[0284] The apparatus 800 may also be distributed throughout the network (e.g. internet 28) including within and between apparatus 800 and any network element (such as a base station 24 and / or apparatus 50).
[0285] Interface 812 enables data communication and signaling between the various items of apparatus 800, as shown in FIG. 8. For example, the interface 812 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g. instructions) 805, including control 806 may comprise object-oriented software configured to pass data or messages between objects within computer program code 805. The apparatus 800 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 800 may at least partially reside in a common housing 828, or a subset of the various components of apparatus 800 may at least partially be located in different housings, which different housings may include housing 828.
[0286] FIG. 9 shows a schematic representation of non-volatile memory media 900a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 900b (e.g. universal serial bus (USB) memory stick) and 900c (e.g. cloud storage for downloading instructions and / or parameters 902 or receiving emailed instructions and / or parameters 902) storing instructions and / or parameters 902 which when executed by a processor allows the processor to perform one or more of the operations of the methods described herein.
[0287] FIG. 10 is an example method 1000, based on the example embodiments described herein. At 1002, the method 1000 includes receiving an overfitting signal. At 1004, the method 1000 includes updating a decoder- side neural network, based at least on the overfitting signal, to obtain an updated decoder-side neural network, wherein for one or more weights of the decoder-side neural network, at least one bit of a binary word representing a weight is overfitted. At 1006 the method 1000 includes, using the updated decoder-side neural network for at least one task. Some examples of the at least one task include, but are not limited to, filtering input data, objection detection, objectsegmentation, or object tracking.
[0288] In one or more embodiments, only one bit of the binary word representing the weight is overfitted.
[0289] FIG. 11 is an example method 1100, based on the example embodiments described herein. At 1102, the method 1100 includes overfitting, for one or more weights of a decoder-side neural network, at least one bit of a binary word representing a weight that is overfitted to obtain an overfitted weight. At 1104, the method 1100 includes signaling information related to the at least one bit of the binary word representing the weight that is overfitted to a decoder.
[0290] In one or more embodiment, the at least one bit of the binary word comprises only one bit.
[0291] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.
[0292] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.
[0293] In the above, some embodiments have been described with reference to specific SEI messages, such as NNPFC SEI message(s) and / or NNPFA SEI message(s). It needs to be understood that embodiments may similarly be realized with any SEI messages of similar nature. For example, some embodiments may be realized with post-filter characteristics and / or activation SEI message(s) where post-filters are not based on neural networks.
[0294] In the above, some example embodiments have been described with reference to an SEI message or an SEI NAL unit. It needs to be understood, however, that embodiments may similarly be realized with any similar structures or data units, such as metadata OBUs. Where example embodiments have been described with SEI messages included in a structure, any independently parsable structures could likewise be used in embodiments. Specific SEI NAL unit and SEI message syntax structures have been presented in example embodiments, but it needs to be understood that embodiments generally apply to any syntax structures with a similar intent as SEI NAL units and / or SEI messages.
[0295] In the above, some example embodiments have been described with reference to postfilters. It is to be understood that embodiments may similarly be realized with in-loop filters. Furthermore, it is to be understood that some embodiments have been described with reference to syntax structures, such as SEI messages, that are suitable for post-filters. It is to be understood that embodiments for in-loop filters may be similarly realized with other syntax structures, such as parameter sets (e.g., video, sequence, picture, and / or adaptation parameter sets) and / or headers of different structural level(s) (e.g., sequence, group of pictures, picture, and / or slice header).
[0296] In the above, some embodiments have been described with reference to neural-network filters. It is to be understood that embodiments may similarly be realized with any filters.
[0297] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not only computers having different architectures such as single / multi-processor architectures and sequential / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGAs), application specific circuits (ASICs), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed- function device, gate array or programmable logic device, etc.
[0298] As used herein, the term ‘circuitry’, ‘circuit’ and variants may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of circuits and software (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s) / software including digital signal processor(s), software, and one or more memories that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even when the software or firmware is not physically present. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device. Circuitry or circuit may also be used to mean a function or a process used to execute a method.
[0299] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitablecombination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.
[0300] The following acronyms and abbreviations that may be found in the specification and / or the drawing figures are defined as follows (the abbreviations may be appended with each other or with other characters using e.g. a hyphen or dash (-), and may be case insensitive):2D two-dimensional3D three-dimensional3GPP 3rd generation partnership project4G fourth generation of broadband cellular network technology5G fifth generation cellular network technology802.x family of IEEE standards dealing with local area networks and metropolitan area networksASIC application specific integrated circuitAVC advanced video codingBD Bjontegaard delta (e.g. BD-rate)BT.2020 set of specifications covering various aspects of video broadcastingBT.709 standard developed by ITU-R for image encoding and signal characteristics of high-definition televisionCDMA code-division multiple accessCL VS coded layer video sequenceCPU central processing unitCTU coding tree unitDCT discrete cosine transformDPB decoded picture bufferDSNN decoder-side neural networkDSP digital signal processorFDMA frequency division multiple accessFPGA field programmable gate arrayGSM global system for mobile communicationsH.222.0 MPEG-2 systems, standard for the generic coding of moving pictures and associated audio informationH.2xx family of video coding standards in the domain of the ITU-T (e.g. H.263,H.264, H.266)HE VC high efficiency video codingHMD head-mounted displayIBC intra block copyIEC International Electrotechnical CommissionIEEE Institute of Electrical and Electronics EngineersLE interfaceIMD integrated messaging deviceIMS instant messaging serviceVO input / output loT internet of thingsIP internet protocolISO International Organization for StandardizationISOBMFF ISO base media file formatITU International Telecommunication UnionITU-R ITU Radiocommunication SectorITU-T ITU Telecommunication Standardization SectorJVET Joint Video Experts TeamJVET-O1150 scalable coding proposalLTE long-term evolutionMAE mean absolute error mAP mean average precisionMMS multimedia messaging serviceMOTA multi-object tracking accuracyMPEG-2 moving picture experts group, H.222 / H.262 as defined by the ITUMSE mean squared errorNAL network abstraction layerNN neural networkNNPF neural-network post-processing filterNNPFA neural-network post-filter activationNNPFC neural-network post-filter characteristicsN / W networkPC personal computerPDA personal digital assistantPID packet identifierPLC power line communicationPSNR peak signal-to-noise ratioQP quantization parameterRAM random access memoryRFID radio frequency identificationRFM reference frame memoryROI region of interestSEI supplemental enhancement informationSMS short messaging serviceSNR signal to noise ratioSON self-organizing / optimizing networkSSIM structural similarity index measureTCP-IP transmission control protocol-internet protocolTDMA time divisional multiple accessTS transport streamTV televisionUHDTV ultra-high -definition televisionUI user interfaceUICC universal integrated circuit cardUMTS universal mobile telecommunications systemURI uniform resource identifierURL uniform resource locatorUSB universal serial busVCM video coding for machinesVPS video parameter setVSEI versatile supplemental enhancement informationVVC versatile video coding
Claims
CLAIMSWhat is claimed is:
1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving an overfitting signal; updating a decoder-side neural network, based at least on the overfitting signal, to obtain an updated decoder-side neural network, wherein for one or more weights of the decoderside neural network, at least one bit of a binary word representing a weight is overfitted; and using the updated decoder-side neural network for at least one task.
2. The apparatus of claim 1, wherein the at least one task comprises one or more of filtering input data, objection detection, object segmentation, or object tracking.
3. The apparatus of any of claims 1 to 2, wherein: all weights of the decoder-side neural network are overfitted to obtain the overfitting signal; a subset of the weights of the decoder-side neural network are overfitted to obtain the overfitting signal; or bias weights of the decoder-side neural network are overfitted to obtain the overfitting signal.
4. The apparatus of any of claims 1 to 3, wherein the apparatus is further caused at least to perform: receiving an indication of a value of the at least one bit that is overfitted; wherein when the at least one bit that is overfitted comprises an original value of 0, the overfitted value comprises 1 and the indication comprises a bit with the value 1 ; and wherein when the at least one bit that is overfitted comprises the original value of 1, the overfitted value comprises 0, and the indication comprises the bit with the value 0.
5. The apparatus of any of claims 1 to 4, wherein a position of the at least one bit that is overfitted is predetermined and is known to a decoder.
6. The apparatus of claim 5, wherein the position of the at least one bit that is overfitted is predetermined based on statistics collected during an offline phase where overfitting is performed on a dataset.
7. The apparatus of any of claims 1 to 4, wherein the apparatus is further caused at least to perform: receiving, from an encoder, an indication of a position of the at least one bit that is overfitted.
8. The apparatus of claim 7, wherein the at least one bit that is overfitted is determined at the encoder.
9. The apparatus of any of the previous claims, wherein the indication comprises a change indication bit, wherein when the change indication bit is 0, an original bit of the weight is left unmodified, wherein when the change indication bit is 1, the original bit of the weight is flipped.
10. The apparatus of claim 9, wherein a bitwise Boolean operation is applied to the original bit of the weight based on the change indication bit.
11. The apparatus of claim 10, wherein the bitwise Boolean operation is predetermined and is known at the decoder.
12. The apparatus of claim 10, wherein the apparatus is further caused to receive an indication of the bitwise Boolean operation, and wherein the bitwise Boolean operation is determined at an encoder.
13. The apparatus of any of claims 1 to 12, wherein all the weights that are comprised in a layer or a group of layers and that are overfitted share the same predetermined position of the bit that is overfitted.
14. The apparatus of any of claims 1 to 12, wherein for each weight or for a group of weights, a number of bits and positions of the bits are predetermined, and wherein the apparatus is further caused to receive an indication of values of the bits, or an indication of whether the bits are to be flipped.
15. The apparatus of any of the previous claims, wherein the apparatus is further caused toperform: receiving, for each overfitted weight, a lossless coded binary word representing the weight update.
16. The apparatus of claim 1, wherein a set of candidate weight updates are predetermined for the weight that is overfitted.
17. The apparatus of claim 16, wherein predetermination of the set of candidate weight updates is based on determining one or more weight -updates based on a dataset.
18. The apparatus of any of claims 16 or 17, wherein the apparatus is further caused to perform: receiving an indication of whether an indicated candidate weight update needs to be added to or subtracted from an original weight.
19. The apparatus of claim 18, wherein a sign of the weight update which determines whether the weight update is added to or subtracted from the original weight is predetermined.
20. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: for one or more weights of a decoder-side neural network, overfitting at least one bit of a binary word representing a weight that is overfitted to obtain an overfitted weight; and signaling information related to the at least one bit of the binary word representing the weight that is overfitted to a decoder.
21. The apparatus of claim 20, wherein: all weights of the decoder-side neural network are overfitted; a subset of weights of the decoder-side neural network are overfitted; or bias weights of the decoder-side neural network are overfitted.
22. The apparatus of any of claims 20 or 21, wherein to signal the information related to the at least one bit of the binary word representing the weight that is overfitted, the apparatus is further caused at least to perform: signaling an indication of a value of the at least one bit that was overfitted; wherein when the at least one bit that is overfitted comprises an original value of 0, an overfitted value comprises 1 and the indication comprises a bit with the value 1; andwherein when the at least one bit that is overfitted comprises the original value of 1, the overfitted value comprises 0, and the indication comprises a bit with the value 0.
23. The apparatus of any of claims 20 to 22, wherein the apparatus is further caused to perform: determining a position of the at least one bit that is overfitted.
24. The apparatus of claim 23, wherein the position of the at least one bit that is overfitted is predetermined based on statistics collected during an offline phase where overfitting is performed on a dataset.
25. The apparatus of any of claims 20 to 23, wherein the apparatus is further caused at least to perform: signaling an indication of the position of the at least one bit that is overfitted to the decoder.
26. The apparatus of any of the claims 20 to 25, wherein the indication comprise a change indication bit, wherein when the change indication bit is 0, an original bit of a weight is left unmodified, wherein when the change indication bit is 1, the original bit of the weight is flipped.
27. The apparatus of claim 26, wherein the change indication bit comprises an indicator of whether to apply a bitwise Boolean operation to the original bit of the weight, and wherein the apparatus is further caused to perform: determining the bitwise Boolean operation; and signaling, to the decoder, the indication whether to apply the bitwise Boolean operation.
28. The apparatus of any of claims 20 to 27, wherein all the weights that are comprised in a layer or a group of layers and that are overfitted share the same predetermined position of the bit that is overfitted.
29. The apparatus of any of claims 20 to 27, wherein for each weight or for a group of weights, a number of bits and positions of the bits are predetermined, and wherein the apparatus is further caused to receive an indication of values of the bits, or an indication of whether the bits are to be flipped.
30. The apparatus of any of the claims 20 to 29, wherein the apparatus is further caused to perform:for each overfitted weight, lossless coding the binary word representing the weight update; and for each overfitted weight, signaling the lossless coded the binary word representing the weight update to the decoder.
31. The apparatus of claim 20, wherein the apparatus is further caused to perform: predetermining a set of candidate weight updates for the weight that is overfitted.
32. The apparatus of claim 31 , wherein predetermination of the set of candidate weight updates is based on determining one or more weight -updates based on a dataset.
33. The apparatus of any of claims 31 or 32, wherein the apparatus is further caused to perform: signaling an indication of whether an indicated candidate weight update needs to be added to or subtracted from the weight that is overfitted.
34. The apparatus of claim 33, wherein the apparatus is further caused to perform predetermining a sign of the weight update which determines whether the weight update is added to or subtracted from the weight that is overfitted.
35. A method comprising: receiving an overfitting signal; updating a decoder-side neural network, based at least on the overfitting signal, to obtain an updated decoder-side neural network, wherein for one or more weights of the decoderside neural network, at least one bit of a binary word representing a weight is overfitted; and using the updated decoder-side neural network for at least one task.
36. The method of claim 35, wherein the at least one task comprises one or more of filtering input data, objection detection, object segmentation, or object tracking.
37. The method of any of claims 35 to 36, wherein: all weights of the decoder-side neural network are overfitted to obtain the overfitting signal; a subset of the weights of the decoder-side neural network are overfitted to obtain the overfitting signal; or bias weights of the decoder-side neural network are overfitted to obtain the overfitting signal.
38. The method of any of claims 35 to 37 further caused at comprising: receiving an indication of a value of the at least one bit that is overfitted; wherein when the at least one bit that is overfitted comprises an original value of 0, the overfitted value comprises 1 and the indication comprises a bit with the value 1 ; and wherein when the at least one bit that is overfitted comprises the original value of 1, the overfitted value comprises 0, and the indication comprises the bit with the value 0.
39. The method of any of claims 35 to 38, wherein a position of the at least one bit that is overfitted is predetermined and is known to a decoder.
40. The method of claim 39, wherein the position of the at least one bit that is overfitted is predetermined based on statistics collected during an offline phase where overfitting is performed on a dataset.
41. The method of any of claims 35 to 38 further comprising: receiving, from an encoder, an indication of a position of the at least one bit that is overfitted.
42. The method of claim 41, wherein the at least one bit that is overfitted is determined at the encoder.
43. The method of any of the claims 35 to 42, wherein the indication comprises a change indication bit, wherein when the change indication bit is 0, an original bit of a weight is left unmodified, wherein when the change indication bit is 1, the original bit of the weight is flipped.
44. The method of claim 43, wherein a bitwise Boolean operation is applied to the original bit of the weight based on the change indication bit.
45. The method of claim 44, wherein the bitwise Boolean operation is predetermined and is known at the decoder.
46. The method of claim 44 further comprising receiving an indication of the bitwise Boolean operation, and wherein the bitwise Boolean operation is determined at the encoder.
47. The method of any of claims 35 to 46, wherein all the weights that are comprised in a layeror a group of layers and that are overfitted share the same predetermined position of the bit that is overfitted.
48. The method of any of claims 35 to 46, wherein for each weight or for a group of weights, a number of bits and positions of the bits are predetermined, and wherein the method further comprising receiving an indication of values of the bits, or an indication of whether the bits are to be flipped.
49. The method of any of the claims 35 to 48 further comprising receiving, for each overfitted weight, a lossless coded binary word representing the weight update.
50. The method of claim 35, wherein a set of candidate weight updates are predetermined for the weight that is overfitted.
51. The method of claim 50, wherein predetermination of the set of candidate weight updates is based on determining one or more weight -updates based on a dataset.
52. The method of any of claims 50 or 51 further comprising receiving an indication of whether an indicated candidate weight update needs to be added to or subtracted from an original weight.
53. The method of claim 52, wherein a sign of the weight update which determines whether the weight update is added to or subtracted from the original weight is predetermined.
54. An method comprising: for one or more weights of a decoder-side neural network, overfitting at least one bit of a binary word representing the weight that is overfitted to obtain an overfitted weight; and signaling information related to the at least one bit of the binary word representing a weight that is overfitted to a decoder.
55. The method of claim 54, wherein: all weights of the decoder-side neural network are overfitted; a subset of weights of the decoder-side neural network are overfitted; or bias weights of the decoder-side neural network are overfitted.
56. The method of any of claims 54 or 55, wherein to signal the information related to the at least one bit of the binary word representing the weight that is overfitted, the method is further caused at least to perform: signaling an indication of a value of the at least one bit that was overfitted; wherein when the at least one bit that is overfitted comprises an original value of 0, an overfitted value comprises 1 and the indication comprises a bit with the value 1; and wherein when the at least one bit that is overfitted comprises the original value of 1, the overfitted value comprises 0, and the indication comprises a bit with the value 0.
57. The method of any of claims 54 to 56 further comprising determining a position of the at least one bit that is overfitted.
58. The method of claim 57, wherein the position of the at least one bit that is overfitted is predetermined based on statistics collected during an offline phase where overfitting is performed on a dataset.
59. The method of any of claims 54 to 57 further comprising signaling an indication of the position of the at least one bit that is overfitted to the decoder.
60. The method of any of the claims 54 to 59, wherein the indication comprise a change indication bit, wherein when the change indication bit is 0, an original bit of a weight is left unmodified, wherein when the change indication bit is 1, the original bit of the weight is flipped.
61. The method of claim 60, wherein the change indication bit comprises an indicator of whether to apply a bitwise Boolean operation to the original bit of the weight, and wherein the method is further caused to perform: determining the bitwise Boolean operation; and signaling, to the decoder, the indication whether to apply the bitwise Boolean operation.
62. The method of any of claims 54 to 61, wherein all the weights that are comprised in a layer or a group of layers and that are overfitted share the same predetermined position of the bit that is overfitted.
63. The method of any of claims 54 to 61, wherein for each weight or for a group of weights, a number of bits and positions of the bits are predetermined, and wherein the method furthercomprising receiving an indication of values of the bits, or an indication of whether the bits are to be flipped.
64. The method of any of the claims 54 to 63 further comprising: for each overfitted weight, lossless coding the binary word representing the weight update; and for each overfitted weight, signaling the lossless coded the binary word representing the weight update to the decoder.
65. The method of claim 54 further comprising predetermining a set of candidate weight updates for the weight that is overfitted.
66. The method of claim 65, wherein predetermination of the set of candidate weight updates is based on determining one or more weight -updates based on a dataset.
67. The method of any of claims 65 or 66 further comprising signaling an indication of whether an indicated candidate weight update needs to be added to or subtracted from the weight that is overfitted.
68. The method of claim 67 further comprising predetermining a sign of the weight update which determines whether the weight update is added to or subtracted from the weight that is overfitted.
69. An apparatus comprising: means for receiving an overfitting signal;; means for updating a decoder-side neural network, based at least on the overfitting signal, to obtain an updated decoder-side neural network, wherein for one or more weights of the decoder-side neural network, at least one bit of a binary word representing a weight is overfitted; and means for using the updated decoder-side neural network for at least one task.
70. The apparatus of claim 69, wherein the apparatus is further caused to perform the methods as claimed in any of the claims 36 to53.
71. Apparatus comprising: means for overfitting, one or more weights of a decoder-side neural network, at leastone bit of a binary word representing the weight that is overfitted to obtain an overfitted weight; and means for signaling information related to the at least one bit of the binary word representing a weight that is overfitted to a decoder.
72. The apparatus of claim 69, wherein the apparatus is further caused to perform the methods as claimed in any of the claims 55 to 68.
73. A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform at least the following: receiving an overfitting signal; updating a decoder-side neural network, based at least on the overfitting signal, to obtain an updated decoder-side neural network, wherein for one or more weights of the decoderside neural network, at least one bit of a binary word representing a weight is overfitted; and using the updated decoder-side neural network for at least one task.
74. The computer readable medium of claim 73, wherein the computer readable medium comprises a non-transitory computer readable medium.
75. The computer readable medium of any of the claims 73 or 74, wherein the computer readable medium causes the apparatus to further perform the methods as claimed in any of the claims 36 to 53.
76. A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform at least the following: one or more weights of a decoder-side neural network, overfitting at least one bit of a binary word representing the weight that is overfitted to obtain an overfitted weight; and signaling information related to the at least one bit of the binary word representing a weight that is overfitted to a decoder.
77. The computer readable medium of claim 76, wherein the computer readable medium comprises a non-transitory computer readable medium.
78. The computer readable medium of any of the claims 76 or 77, wherein the computer readable medium causes the apparatus to further perform the methods as claimed in any of the claims 55 to 68.
Citation Information
Patent Citations
Iterative overfitting and freezing of decoder-side neural networks
US20230196072A1
Cited By
Internet of vehicles sensing data transmission method and device
CN120935533A