Method and data processing system for lossy image or video encoding, transmission and decoding
By training neural networks for lossy image and video encoding, transmission, and decoding, the method addresses the challenge of efficiently compressing high-resolution content, achieving reduced data transmission and energy use while maintaining image quality.
Patent Information
- Application Number
- PCT/EP2024/083641
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-02
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-12
AI Technical Summary
Current lossy image and video compression techniques face challenges in efficiently reducing data quantity while maintaining acceptable visual quality, especially with increasing demand for higher resolution content, which strains communications networks and increases energy use.
A method involving the training of one or more neural networks for lossy image or video encoding, transmission, and decoding, where an input image is downscaled, encoded, decoded, and upscaled, with the differences evaluated to update the neural network parameters, optimizing the compression process.
This approach enables efficient compression and decompression of images and videos, reducing data transmission requirements while maintaining image quality, thus alleviating the strain on communications networks and lowering energy consumption.
Smart Images

Figure EP2024083641_12062025_PF_FP_ABST
Abstract
Description
[0001] Method and data processing system for lossy image or videoencoding, transmission and decodingBACKGROUNDThis invention relates to a method and system for lossy image or video encoding, transmissionand decoding, a method, apparatus, computer program and computer readable storage mediumfor lossy image or video encoding and transmission, and a method, apparatus, computerprogram and computer readable storage medium for lossy image or video receipt and decoding.There is increasing demand from users of communications networks for images and videocontent. Demand is increasing not just for the number of images viewed, and for the playingtime of video; demand is also increasing for higher resolution content. This places increasingdemand on communications networks and increases their energy use because of the largeramount of data being transmitted.To reduce the impact of these issues, image and video content is compressed for transmissionacross the network. The compression of image and video content can be lossless or lossycompression. In lossless compression, the image or video is compressed such that all of theoriginal information in the content can be recovered on decompression. However, when usinglossless compression there is a limit to the reduction in data quantity that can be achieved. Inlossy compression, some information is lost from the image or video during the compressionprocess. Known compression techniques attempt to minimise the apparent loss of informationby the removal of information that results in changes to the decompressed image or video thatis not particularly noticeable to the human visual system. JPEG, JPEG2000, AVC, HEVC andAVI are examples of compression processes for image and / or video files.In general terms, known lossy image compression techniques use the spatial correlationsbetween pixels in images to remove redundant information during compression. For example,in an image of a blue sky, if a given pixel is blue, there is a high likelihood that the neighbouringpixels, and their neighbouring pixels, and so on, are also blue. There is accordingly no need toretain all the raw pixel data. Instead, we can retain only a subset of the pixels which take upfewer bits and infer the pixel values of the other pixels using information derived from spatialcorrelations.A similar approach is applied in known lossy video compression techniques. That is, spatialcorrelations between pixels allow the removal of redundant information during compression.However, in video compression, there is further information redundancy in the form of temporalcorrelations. For example, in a video of an aircraft flying across a blue-sky background, mostof the pixels of the blue sky do not change at all between frames of the video. The mostof the blue sky pixel data for the frame at position t = 0 in the video is identical to that atposition t = 10. Storing this identical, temporally correlated, information is inefficient. Instead,only the blue sky pixel data for a subset of the frames is stored and the rest are inferred frominformation derived from temporal correlations.In the realm of lossy video compression in particular, the removal of redundant temporallycorrelated information in a video sequence is as known inter-frame redundancy.One technique using inter-frame redundancy that is widely used in standard video compressionalgorithms involves the categorization of video frames into three types: I-frames, P-frames, andB-frames. Each frame type carries distinct properties concerning their encoding and decodingprocess, playing different roles in achieving high compression ratios while maintainingacceptable visual quality.I-frames, or intra-coded frames, serve as the foundation of the video sequence. These framesare self-contained, each one encoding a complete image without reference to any other frame.In terms of compression, I-frames are least compressed among all frame types, thus carryingthe most data. However, their independence provides several benefits, including being thestarting point for decompression and enabling random access, crucial for functionalities likefast-forwarding or rewinding the video.P-frames, or predictive frames, utilize temporal redundancy in video sequences to achievegreater compression. Instead of encoding an entire image like an I-frame, a P-frame representsthe difference between itself and the closest preceding I- or P-frame. The process, known asmotion compensation, identifies and encodes only the changes that have occurred, therebysignificantly reducing the amount of data transmitted. Nonetheless, P-frames are dependent onprevious frames for decoding. Consequently, any error during the encoding or transmissionprocess may propagate to subsequent frames, impacting the overall video quality.B-frames, or bidirectionally predictive frames, represent the highest level of compression.Unlike P-frames, B-frames use both the preceding and following frames as references in theirencoding process. By predicting motion both forwards and backwards in time, B-framesencode only the differences that cannot be accurately anticipated from the previous and nextframes, leading to substantial data reduction. Although this bidirectional prediction makesB-frames more complex to generate and decode, it does not propagate decoding errors sincethey are not used as references for other frames.Artificial intelligence (AI) based compression techniques achieve compression and decom-pression of images and videos through the use of trained neural networks in the compressionand decompression process. Typically, during training of the neutral networks, the differencebetween the original image and video and the compressed and decompressed image and videois analyzed and the parameters of the neural networks are modified to reduce this differencewhile minimizing the data required to transmit the content. However, AI based compressionmethods may achieve poor compression results in terms of the appearance of the compressedimage or video or the amount of information required to be transmitted.An example of an AI based image compression process comprising a hyper-network is describedin Ballé, Johannes, et al. “Variational image compression with a scale hyperprior.” arXivpreprint arXiv:1802.01436 (2018), which is hereby incorporated by reference.An example of an AI based video compression approach is shown in Agustsson, E., Minnen, D.,Johnston, N., Balle, J., Hwang, S. J., and Toderici, G. (2020), Scale-space flow for end-to-endoptimized video compression. In Proceedings of the IEEE / CVF Conference on ComputerVision and Pattern Recognition (pp. 8503-8512), which is hereby incorporated by reference.A further example of an AI based video compression approach is shown in Mentzer, F.,Agustsson, E., Ballé, J., Minnen, D., Johnston, N., and Toderici, G. (2022, November). Neuralvideo compression using gans for detail synthesis and propagation. In Computer Vision–ECCV2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, PartXXVI (pp. 562-578), which is hereby incorporated by reference. Figure 3 of which shows anarchitecture that calculates optical flow with a flow model, UFlow, and encodes the calculatedoptical flow with a flow encoder, Eflow.SUMMARYAccording to an aspect of the present disclosure, there is provided a method of training one ormore neural networks for use in lossy image or video encoding, transmission and decoding.The method comprises receiving an input image at a first computer system; downsamplingthe input image with a downsampler to produce a downsampled input image; encoding thedownsampled input image using a first neural network to produce a latent representation;decoding the latent representation using a second neural network to produce an output image,wherein the output image is an approximation of the input image; upsampling the output imagewith an upsampler to produce an upsampled output image; evaluating a function based on adifference between one or more of: the output image and the input image, the output imageand the downsampled input image, the upsampled output image and the input image, and / orthe upsampled output image and the downsampled input image; updating the parameters ofthe first neural network and the second neural network based on the evaluated function; andrepeating the above steps using a first set of input images to produce a first trained neuralnetwork and a second trained neural network.Optionally, the method described above may comprise a third neural network for upsampling,and wherein the method may include updating the parameters of said third neural networkbased on the evaluated function.Optionally, the method described above may comprise a downsampler configured for eitherbilinear or bicubic downsampling.Optionally, the method described above may comprise a Gaussian blur filter in the downsampler.Optionally, method according to any one as described above may comprise (i) updating theparameters of the first neural network and the second neural network based on the evaluatedfunction for a first number of steps to produce the first and second trained neural networkswithout performing said upsampling and downsampling and without updating the parametersof the third neural network, (ii) freezing the parameters of the first and second neural networksafter said first number of steps, and performing said upsampling and downsampling, and saidupdating of the parameters of the third neural network for a second number of said steps.Optionally, the method described above may comprise a fourth neural network in the down-sampler, and may further include updating the parameters of said fourth neural network basedon the evaluated function.Optionally, the method described above may comprise an upsampler configured for eitherbilinear or bicubic upsampling.Optionally, method as described above may comprise (i) updating the parameters of the firstneural network and the second neural network based on the evaluated function for a first numberof said steps to produce the first and second trained neural networks without performing saidupsampling and downsampling and without updating the parameters of the fourth neuralnetwork, (ii) freezing the parameters of the first and second neural networks after said numberfirst of steps, and performing said downsampling and said updating of the parameters of thefourth neural network for a second number of said steps.Optionally, the method described above may comprise entropy encoding the latent representationinto a bitstream having a length, wherein the function is further based on said bitstream length,and wherein said updating the parameters of the third or fourth neural network is based on theevaluated function based on the bitstream length.Optionally, the method described above may comprise determining the difference between oneor more of the output image and the input image, the output image and the downsampled inputimage, the upsampled output image and the input image, and / or the upsampled output imageand the downsampled input image based on the output of a fifth neural network acting as adiscriminator.Optionally, the method of any as described above may comprise calculating the differencebetween one or more of: the output image and the input image, the output image and thedownsampled input image, the upsampled output image and the input image, and / or theupsampled output image and the downsampled input image. The difference is expressed interms of a mean squared error (MSE) and / or a structural similarity index measure (SSIM).Optionally, the method described above may comprise a term defining a visual perceptualmetric.Optionally, the method described above may comprise a visual perceptual metric, wherein theterm defining the metric comprises a MS-SIM metric.According to an aspect of the present disclosure, there is provided a method of training oneor more neural networks, the one or more neural networks being for use in lossy image orvideo encoding, transmission and decoding. The method comprises receiving an input imageat a first computer system, encoding the input image using a first neural network to producea latent representation, decoding the latent representation using a second neural network toproduce an output image, wherein the output image is an approximation of the input image,upsampling the output image with an upsampler to produce an upsampled output image, theupsampler comprising a third neural network, evaluating a function based on a differencebetween one or both of: the output image and the input image, and / or the upsampled outputimage and the input image, updating the parameters of the third neural network based on theevaluated function, and repeating the above steps using a first set of input images to produce afirst trained neural network and a second trained neural network.According to an aspect of the present disclosure, there is provided a method of training one ormore neural networks for use in lossy image or video encoding, transmission and decoding.The method comprises receiving an input image at a first computer system, downsampling theinput image with a downsampler comprising a fourth neural network to produce a downsampledinput image, encoding the downsampled input image using a first neural network to producea latent representation, decoding the latent representation using a second neural network toproduce an output image approximating the input image. Further steps involve evaluating afunction based on differences between various images and updating the parameters of thefourth neural network based on the evaluated function. This process is repeated using a first setof input images to produce a first trained neural network and a second trained neural network.Optionally, the method described above may comprise producing the previously downsampledinput image by performing bilinear or bicubic downsampling on the input image.According to an aspect of the present disclosure, there is provided a method for lossy image orvideo encoding, transmission and decoding. The method comprises the steps of receiving aninput image at a first computer system, downsampling the input image with a downsampler,encoding the downsampled input image using a first trained neural network to produce a latentrepresentation, transmitting the latent representation to a second computer system, decodingthe latent representation using a second trained neural network to produce an output image,wherein the output image is an approximation of the input image, and upsampling the outputimage with an upsampler to produce an upsampled output image.According to an aspect of the present disclosure, there is provided a method for lossy image orvideo encoding, transmission and decoding, the method comprising the steps of receiving aninput image at a first computer system; encoding the input image using a first trained neuralnetwork to produce a latent representation; transmitting the latent representation to a secondcomputer system; and decoding the latent representation using a second trained neural networkto produce an output image, wherein the output image is an approximation of the input image;wherein one or more of the above steps comprises performing a downsampling or upsamplingoperation, and wherein the downsampling or upsampling operation comprises performing oneor convolution operations without performing a space-to-depth or depth-to-space operation.Optionally, the method described above may comprise performing the downsampling orupsampling operation on a CPU without performing a space-to-depth or depth-to-spaceoperation, and wherein said downsampling or upsampling is configured to be performed inreal-time or near real-time.The method as described above may optionally comprise performing the downsamplingor upsampling operation on a neural accelerator without performing a space-to-depth ordepth-to-space operation, and wherein said downsampling or upsampling is configured to beperformed in real-time or near real-time.The method as described above may optionally comprise a downsampling operation thatincludes applying one or more convolutional layers with a kernel size based on a downsamplingfactor. These convolutional layers are configured to sequentially reduce the spatial dimensionsof an input while increasing the depth or channel dimension of the input.Optionally, the method described above may comprise an input image.Optionally, the method may comprise a tensor representation of the input image.Optionally, the method described above may comprise a downsampling operation performed byapplying one or more convolutional layers configured with a stride equal to the downsamplingfactor. The number of filters in each convolutional layer is based on the original number ofchannels in the input tensor and on the downsampling factor.Optionally, the method may include performing a first convolution and second deconvolution.This may further involve performing additional upsampling steps and utilizing additional layerssuch as Maxpool and Relu.Optionally, the method described above may comprise the input including a latent representation.Optionally, the method may comprise a tensor representation of the latent representation or theoutput image as part of its input.Optionally, the method as described may include upsampling layers having strides determinedby an upsampling factor.Optionally, the method described above may further comprise applying an activation functionafter each convolutional layer in the upsampling operation.Optionally, the method described above may include the upsampling layers being selectedfrom a group consisting of nearest neighbor upsampling, bilinear upsampling, and bicubicupsampling, alternated with the convolutional layers.According to an aspect of the present disclosure, there is provided a method for lossy imageor video encoding, transmission and decoding, comprising the steps of receiving an inputimage and a second image at a first computer system; estimating optical flow informationusing the second image and the input image using a first neural network, wherein the opticalflow information is indicative of a difference between a representation of the second imageand a representation of the input image; transmitting the optical flow information to a secondcomputer system; decoding the optical flow information using a second neural network; andusing the second image and the decoded optical flow information to produce an output image,wherein the output image is an approximation of the input image. The estimating of optical flowinformation further comprises estimating differences between the input image and the secondimage by applying a first convolution operation on a one or more pixels of a representation ofthe input image and / or on one or more pixels of a representation of the second image, whereinthe convolution operation comprises applying one or more filters comprising weights havingvalues randomly distributed between a minimum and a maximum value.Optionally, the method comprises estimating a compressively encoded cost volume indicativeof said differences by applying said first convolution operation.Optionally, the first convolution substantially preserves a norm of a distribution of the respectivepixels of the representation of the input image and / or respective pixels of the representation ofthe second image.Optionally, a distribution of values of pixels of the representation of the input image and / or thedistribution of values of pixels of the representation of the second image are sparse distributionsin a spatial domain of the representation of the input image and / or the second image.Optionally, a distribution of values of pixels of the representation of the input image and / orthe distribution of values of pixels of the representation of the second image comprise sparsedistributions in a spatial domain.Optionally, the method described above may comprise assigning weights with values distributedaccording to a sub-Gaussian distribution.Optionally, the method described above may comprise determining the minimum value and / ormaximum value based on the number of channels of the input image and / or second image,kernel size of the first convolution operation, and / or pixel radius across which said differencesare estimated.Optionally, the method described above may comprise performing a second convolutionoperation on an output of the first convolution operation, wherein the second convolutionoperation substantially preserves a norm of a distribution of said output of the first convolutionoperation.Optionally, the method described above may comprise estimating a difference between anoutput of the second convolution operation and an output of the first convolution operation.Optionally, the difference may comprise an absolute difference.Optionally, the difference defines a cost volume.Optionally, the method described above may comprise using the optical flow information towarp a representation of the second image.Optionally, the method may involve estimating a difference between the warped second imageand the input image in order to create a residual representation of the input image relative tothe warped second image.Optionally, the method described above may comprise: (i) using a third neural network toencode the residual representation of the input image; (ii) transmitting the encoded residualrepresentation of the input image to the second computer system; (iii) using a fourth neuralnetwork to decode the residual representation of the input image; and (iv) using decoded theresidual representation of the input image to produce said output image.Optionally, the method described above may comprise applying a third convolution operationto an output of the first convolution operation and / or to an output of the second convolutionoperation.Optionally, a kernel size of the second convolution operation is greater than a kernel size ofthe first convolution operation.Optionally, the first convolution operation is defined by a 1x1 kernel.Optionally, the second convolution operation is defined by a 3x3 kernel.Optionally, the third convolution operation is defined by a 1x1 kernel.Optionally, the method described above may comprise performing the second convolutionoperation to entangle information associated with respective pixels of the representation ofthe input image with information associated with pixels adjacent corresponding pixels in therepresentation of the second image.Optionally, the method described above may comprise a first, second, and where present thirdconvolution operation, wherein these operations are performed without group convolutions.Optionally, one or more outputs of the first, second and / or third convolution operation arestored in contiguous memory blocks, and wherein estimating a difference comprises retrievingsaid stored outputs from said contiguous memory blocks.Optionally, a distribution of pixel values of the input image and of the second image are sparseand incoherent in a spatial domain and / or a transform of a spatial domain.According to an aspect of the present disclosure, there is provided a system configured toperform any of the above methods.According to an aspect of the present disclosure, there is provided a method for lossy imageor video encoding and transmission. The method includes receiving an input image and asecond image at a first computer system; estimating optical flow information using the secondimage and the input image using a first neural network, wherein the optical flow information isindicative of a difference between a representation of the second image and a representation ofthe input image; transmitting the optical flow information to a second computer system. In thismethod, estimating the optical flow information comprises estimating differences between theinput image and the second image by by applying a first convolution operation on a one or morepixels of a representation of the input image and / or on one or more pixels of a representationof the second image, wherein the convolution operation comprises applying one or more filterscomprising weights having values randomly distributed between a minimum and a maximumvalue.According to an aspect of the present disclosure, there is provided a method for lossy image orvideo decoding, comprising receiving an input image and a second image at a first computersystem; estimating optical flow information using the second image and the input imageusing a first neural network, wherein the optical flow information is indicative of a differencebetween a representation of the second image and a representation of the input image basedon a compressively encoded cost volume; receiving optical flow information at a secondcomputer system, wherein the optical flow information is indicative of a difference between arepresentation of a second image and a representation of an input image; decoding the opticalflow information using a second neural network; and using the second image and the decodedoptical flow information to produce an output image, which approximates the input image.According to an aspect of the present disclosure, there is provided an apparatus configured toperform any of the above methods.According to an aspect of the present disclosure, there is provided a method for estimating adifference between a first image and a second image. The method comprises performing afirst convolution operation on respective pixels of a representation of the first image and onrespective pixels of a representation of the second image; and estimating a difference betweenthe first image and the second image based on one or more outputs of the first convolutionoperation on the first and second images by estimating a compressively encoded cost volumeindicative of said differences.Optionally, the method may include performing a second convolution operation on an outputof the first convolution operation, and estimating a difference between an output of the secondconvolution operation and the first convolution operation.Optionally, the method described above may comprise performing a second convolutionoperation that entangles information associated with respective pixels of the representationof the first with information associated with pixels adjacent corresponding pixels in therepresentation of the second image.Optionally, the method described above may include a first convolution operation where one ormore filters are applied with weights having values randomly distributed between a minimumvalue and a maximum value.Optionally, the method described above may comprise determining the minimum value and / ormaximum value based on the number of channels of the input image and / or second image,kernel size of the first convolution operation, and pixel radius across which said differences areestimated.Optionally, the method described above may comprise a difference comprising an absolutedifference.Optionally, the method described above may comprise defining a cost volume based on thedifference.Optionally, the method described above may comprise applying a third convolution operationto an output of the first convolution operation and / or to an output of the second convolutionoperation.Optionally, the method described may involve adjusting a kernel size of the second convolutionoperation to be larger than that of the first.Optionally, the method described above may comprise a first convolution operation defined bya 1x1 kernel.Optionally, the method described above may include the step whereby the second convolutionoperation is defined by a 3x3 kernel.Optionally, the method described above may comprise the third convolution operation definedby a 1x1 kernel.Optionally, the method described above may comprise storing a plurality of respective outputsof the first, second and / or third convolution operations in contiguous memory blocks, andwherein estimating a difference comprises retrieving said stored outputs from said contiguousmemory blocks.Optionally, the method described above may comprise using said difference to identify one ormore pixel patches in the second image as movement-containing pixel patches, and generatinga bounding box around one or more of said movement-containing pixel patches.According to an aspect of the present disclosure, there is provided a data processing apparatusconfigured to perform any of the above described methods.According to a method of the present disclosure, there is provided a method of training one ormore neural networks, the one or more neural networks being for use in lossy image or videoencoding, transmission and decoding, the method comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of optical flow information,the optical flow information being indicative of a difference between the first image and thesecond image;with a second neural network, decoding the latent representation of optical flowinformation to produce an approximation of the optical flow information;with a third neural network, producing an output image using the optical flow informationand using a representation of the first image, wherein the output image is an approximation ofthe first image;evaluating a function based on a difference between the first image and the second image,the function comprising a Jacobian penalty term,updating the parameters of the first, second and / or third neural networks based on theevaluated function; andrepeating the above steps using a first set of input images to produce first, second and / orthird trained neural networks.Optionally, the Jacobian penalty term is based on a rate of change of one or more first variableswith respect to one or more second variables, the first variables and second variables selectedfrom inputs and / or outputs associated with the one or more neural networks.Optionally, at least input and / or output associated with the one or more neural networks is botha first variable and a second variable.Optionally, the method comprises producing the second variables from the first variables bymapping the first variables to the second variables.Optionally, the mapping is defined by an auxiliary function.Optionally, the first variables are inputs to the auxiliary function and the second variables areoutputs of the auxiliary function.Optionally, at least one input of said inputs to the auxiliary function is also an output of theauxiliary function.Optionally, the inputs of said mapping are defined in an input space, and the outputs of saidmapping are defined in an output space, and wherein the auxiliary function maps the inputspace to the output space.Optionally, the input space matches the output space.Optionally, the auxiliary function is based on the third neural network.Optionally, the third neural network comprises a residual decoder neural network.Optionally, the at least one input to the auxiliary function that is also an output of the auxiliaryfunction comprises said latent representation of the first image.Optionally, the method comprises weighting the Jacobian penalty term.Optionally, said weighting is based on a difference between the first image and the secondimage.Optionally, said weighting is defined by a weighted norm based on a matrix associated withsaid rate of change.Optionally, the method comprises estimating the Jacobian penalty term by approximating anorm of a matrix associated with said rate of change.Optionally, approximating the norm of the matrix comprises making a single sample approxi-mation.Optionally, the method comprises introducing the Jacobian penalty term into said functionafter a first number of said repeated steps.Optionally, said first number of said repeated steps is based on a GOP-size of one or moreframe sequences in said first set of input images.According to an aspect of the present disclosure, there is provided, a method of performinglossy image or video encoding, transmission and decoding, the method comprising the stepsof: receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of optical flow information,the optical flow information being indicative of a difference between the first image and thesecond image;transmitting the latent representation of optical flow information to a second computersystem;with a second neural network, decoding the latent representation of optical flowinformation to produce an approximation of the optical flow information;with a third neural network, producing an output image using the optical flow informationand using a representation of the first image, wherein the output image is an approximation ofthe first image;wherein the first neural network, the second neural network, and the third neural networkare produced according to any of the methods described above.According to an aspect of the present disclosure, there is provided, a method of performinglossy image or video encoding, transmission, the method comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of optical flow information,the optical flow information being indicative of a difference between the first image and thesecond image;transmitting the latent representation of optical flow information to a second computersystem; wherein the first neural network,is produced according to any of the methods describedabove.According to an aspect of the present disclosure, there is provided, a method of performinglossy image decoding, the method comprising the steps of:receiving a latent representation of optical flow information at a second computer system;,the optical flow information being indicative of a difference between a first image and a second image; with a second neural network, decoding the latent representation of optical flowinformation to produce an approximation of the optical flow information;with a third neural network, producing an output image using the optical flow informationand using a representation of the first image, wherein the output image is an approximation ofthe first image;wherein the second neural network and the third neural network are produced accordingto any of the methods described above.According to an aspect of the present disclosure, there is provided, a method of performinglossy image or video decoding, the method comprising the steps of:with a second neural network, at a second computer system, decoding a latent represen-tation to produce a first output image, wherein the first output image is an approximation ofone image of an image pair of a first sequence of input images;repeating the above step to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the second neural network is produced according to any of the methods describedabove.According to an aspect of the present disclosure, there is provided, a data processing apparatusconfigured to perform any of the above methods.According to an aspect of the present disclosure, there is provided, a computer programcomprising instructions which, when the program is executed by a computer, cause thecomputer to carry out the method of any of the above methods.According to an aspect of the present disclosure, there is provided, a computer-readable storagemedium comprising instructions which, when executed by a computer, cause the computercarry out the method of any of the above methods.BRIEF DESCRIPTION OF THE DRAWINGSAspects of the invention will now be described by way of examples, with reference to thefollowing figures in which:Figure 1 illustrates an example of an image or video compression, transmission and decom-pression pipeline.Figure 2 illustrates a further example of an image or video compression, transmission anddecompression pipeline including a hyper-network.Figure 3 illustrates an example of a video compression, transmission and decompressionpipeline.Figure 4 illustrates an example of a video compression, transmission and decompressionsystem.Figure 5 illustrates an example of an image or video compression, transmission and decom-pression pipeline.Figure 6 illustrates an example of an image or video compression, transmission and decom-pression pipeline.Figure 7 illustrates an example of an image or video compression, transmission and decom-pression pipeline.Figure 8 illustrates an example of an image or video compression, transmission and decom-pression pipeline.Figure 9 illustratively shows an example sequence of layers of an image or video compression,transmission and decompression pipeline.Figure 10 illustratively shows an example sequence of layers of an image or video compression,transmission and decompression pipeline.Figure 11 illustrates an example of how optical flow information may be calculated betweentwo images.Figure 12a illustrates steps of a MAD cost volume calculation.Figure 12b illustrates steps of a MAD cost volume calculation.Figure 12c illustrates steps of a MAD cost volume calculation.Figure 13 illustrates steps of an RKADe cost volume calculation.Figure 14 illustrates an example of an image or video compression, transmission and decom-pression pipeline.DETAILED DESCRIPTION OF THE DRAWINGSCompression processes may be applied to any form of information to reduce the amountof data, or file size, required to store that information. Image and video information is anexample of information that may be compressed. The file size required to store the information,particularly during a compression process when referring to the compressed file, may bereferred to as the rate. In general, compression can be lossless or lossy. In both forms ofcompression, the file size is reduced. However, in lossless compression, no information is lostwhen the information is compressed and subsequently decompressed. This means that theoriginal file storing the information is fully reconstructed during the decompression process.In contrast to this, in lossy compression information may be lost in the compression anddecompression process and the reconstructed file may differ from the original file. Image andvideo files containing image and video data are common targets for compression.In a compression process involving an image, the input image may be represented as ^^. Thedata representing the image may be stored in a tensor of dimensions ^^ × ^^ × ^^, where ^^represents the height of the image, ^^ represents the width of the image and ^^ represents thenumber of channels of the image. Each ^^ × ^^ data point of the image represents a pixel valueof the image at the corresponding location. Each channel ^^ of the image represents a differentcomponent of the image for each pixel which are combined when the image file is displayed bya device. For example, an image file may have 3 channels with the channels representing thered, green and blue component of the image respectively. In this case, the image informationis stored in the RGB colour space, which may also be referred to as a model or a format.Other examples of colour spaces or formats include the CMKY and the YCbCr colour models.However, the channels of an image file are not limited to storing colour information and otherinformation may be represented in the channels. As a video may be considered a series ofimages in sequence, any compression process that may be applied to an image may also beapplied to a video. Each image making up a video may be referred to as a frame of the video.The output image may differ from the input image and may be represented by ^^. The differencebetween the input image and the output image may be referred to as distortion or a differencein image quality. The distortion can be measured using any distortion function which receivesthe input image and the output image and provides an output which represents the differencebetween input image and the output image in a numerical way. An example of such a methodis using the mean square error (MSE) between the pixels of the input image and the outputimage, but there are many other ways of measuring distortion, as will be known to the personskilled in the art. The distortion function may comprise a trained neural network.Typically, the rate and distortion of a lossy compression process are related. An increase inthe rate may result in a decrease in the distortion, and a decrease in the rate may result in anincrease in the distortion. Changes to the distortion may affect the rate in a correspondingmanner. A relation between these quantities for a given compression technique may be definedby a rate-distortion equation.AI based compression processes may involve the use of neural networks. A neural network isan operation that can be performed on an input to produce an output. A neural network maybe made up of a plurality of layers. The first layer of the network receives the input. One ormore operations may be performed on the input by the layer to produce an output of the firstlayer. The output of the first layer is then passed to the next layer of the network which mayperform one or more operations in a similar way. The output of the final layer is the output ofthe neural network.Each layer of the neural network may be divided into nodes. Each node may receive at leastpart of the input from the previous layer and provide an output to one or more nodes in asubsequent layer. Each node of a layer may perform the one or more operations of the layer onat least part of the input to the layer. For example, a node may receive an input from one ormore nodes of the previous layer. The one or more operations may include a convolution, aweight, a bias and an activation function. Convolution operations are used in convolutionalneural networks. When a convolution operation is present, the convolution may be performedacross the entire input to a layer. Alternatively, the convolution may be performed on at leastpart of the input to the layer.Each of the one or more operations is defined by one or more parameters that are associatedwith each operation. For example, the weight operation may be defined by a weight matrixdefining the weight to be applied to each input from each node in the previous layer to eachnode in the present layer. In this example, each of the values in the weight matrix is a parameterof the neural network. The convolution may be defined by a convolution matrix, also knownas a kernel. In this example, one or more of the values in the convolution matrix may be aparameter of the neural network. The activation function may also be defined by values whichmay be parameters of the neural network. The parameters of the network may be varied duringtraining of the network.Other features of the neural network may be predetermined and therefore not varied duringtraining of the network. For example, the number of layers of the network, the number ofnodes of the network, the one or more operations performed in each layer and the connectionsbetween the layers may be predetermined and therefore fixed before the training process takesplace. These features that are predetermined may be referred to as the hyperparameters of thenetwork. These features are sometimes referred to as the architecture of the network.To train the neural network, a training set of inputs may be used for which the expected output,sometimes referred to as the ground truth, is known. The initial parameters of the neuralnetwork are randomized and the first training input is provided to the network. The output ofthe network is compared to the expected output, and based on a difference between the outputand the expected output the parameters of the network are varied such that the differencebetween the output of the network and the expected output is reduced. This process is thenrepeated for a plurality of training inputs to train the network. The difference between theoutput of the network and the expected output may be defined by a loss function. The result ofthe loss function may be calculated using the difference between the output of the networkand the expected output to determine the gradient of the loss function. Back-propagation ofthe gradient descent of the loss function may be used to update the parameters of the neuralnetwork using the gradients ^^^^ / ^^^^ of the loss function. A plurality of neural networks in asystem may be trained simultaneously through back-propagation of the gradient of the lossfunction to each network.In the context of image or video compression, this type of system, where simultaneous trainingwith back-propagation through each element or the whole network architecture may be referredto as end-to-end, learned image or video compression. Unlike in traditional compressionalgorithms that use primarily handcrafted, manually constructed steps, an end-to-end learnedsystem learns itself during training what combination of parameters best achieves the goal ofminimising the loss function. This approach is advantageous compared to systems that are notend-to-end learned because an end-to-end system has a greater flexibility to learn weights andparameters that might be counter-intuitive to someone handcrafting features.It will be appreciated that the term "training" or "learning" as used herein means the processof optimizing an artificial intelligence or machine learning model, based on a given set of data.This involves iteratively adjusting the parameters of the model to minimize the discrepancybetween the model’s predictions and the actual data, represented by the above-describedrate-distortion loss function.The training process may comprise multiple epochs. An epoch refers to one complete passof the entire training dataset through the machine learning algorithm. During an epoch, themodel’s parameters are updated in an effort to minimize the loss function. It is envisaged thatmultiple epochs may be used to train a model, with the exact number depending on variousfactors including the complexity of the model and the diversity of the training data.Within each epoch, the training data may be divided into smaller subsets known as batches.The size of a batch, referred to as the batch size, may influence the training process. A smallerbatch size can lead to more frequent updates to the model’s parameters, potentially leading tofaster convergence to the optimal solution, but at the cost of increased computational resources.Conversely, a larger batch size involves fewer updates, which can be more computationallyefficient but might converge slower or even fail to converge to the optimal solution.The learnable parameters are updated by a specified amount each time, determined by thelearning rate. The learning rate is a hyperparameter that decides how much the parametersare adjusted during the training process. A smaller learning rate implies smaller steps in theparameter space and a potentially more accurate solution, but it may require more epochs toreach that solution. On the other hand, a larger learning rate can expedite the training processbut may risk overshooting the optimal solution or causing the training process to diverge.The training described herein may involve use of a validation set, which is a portion of thedata not used in the initial training, which is used to evaluate the model’s performance and toprevent overfitting. Overfitting occurs when a model learns the training data too well, to thepoint that it fails to generalize to unseen data. Regularization techniques, such as dropout orL1 / L2 regularization, can also be used to mitigate overfitting.It will be appreciated that training a machine learning model is an iterative process thatmay comprise selection and tuning of various parameters and hyperparameters. As will beappreciated, the specific details, such as hyper parameters and so on, of the training processmay vary and it is envisaged that producing a trained model in this way may achieved in anumber of different ways with different epochs, batch sizes, learning rates, regularisations,and so on, the details of which are not essential to enabling the advantages and effects of thepresent disclosure, except where stated otherwise. The point at which an “untrained” neuralnetwork is considered be “trained” is envisaged to be case specific and depend on, for example,on a number of epochs, a plateauing of any further learning, or some other metric and is notconsidered to be essential in achieving the advantages described herein.More details of an end-to-end, learned compression process will now be described. It will beappreciated that in some cases, end-to-end, learned compression processes may be combinedwith one or more components that are handcrafted or trained separately.In the case of AI based image or video compression, the loss function may be defined by therate distortion equation. The rate distortion equation may be represented by ^^^^^^^^ = ^^ + ^^ ∗ ^^,where ^^ is the distortion function, ^^ is a weighting factor, and ^^ is the rate loss. ^^ may bereferred to as a lagrange multiplier. The langrange multiplier provides as weight for a particularterm of the loss function in relation to each other term and can be used to control which termsof the loss function are favoured when training the network.In the case of AI based image or video compression, a training set of input images maybe used. An example training set of input images is the KODAK image set (for exampleat www.cs.albany.edu / xypan / research / snr / Kodak.html). An example training set of inputimages is the IMAX image set. An example training set of input images is the Imagenetdataset (for example at www.image-net.org / download). An example training set of inputimages is the CLIC Training Dataset P (“professional”) and M (“mobile”) (for example athttp: / / challenge.compression.cc / tasks / ).An example of an AI based compression, transmission and decompression process 100 isshown in Figure 1. As a first step in the AI based compression process, an input image 5 isprovided. The input image 5 is provided to a trained neural network 110 characterized by afunction ^^^^ acting as an encoder. The encoder neural network 110 produces an output basedon the input image. This output is referred to as a latent representation of the input image 5. Ina second step, the latent representation is quantised in a quantisation process 140 characterisedby the operation ^^, resulting in a quantized latent. The quantisation process transforms thecontinuous latent representation into a discrete quantized latent. An example of a quantizationprocess is a rounding function.In a third step, the quantized latent is entropy encoded in an entropy encoding process 150 toproduce a bitstream 130. The entropy encoding process may be for example, range or arithmeticencoding. In a fourth step, the bitstream 130 may be transmitted across a communicationnetwork.In a fifth step, the bitstream is entropy decoded in an entropy decoding process 160. Thequantized latent is provided to another trained neural network 120 characterized by a function^^^^ acting as a decoder, which decodes the quantized latent. The trained neural network 120produces an output based on the quantized latent. The output may be the output image of theAI based compression process 100. The encoder-decoder system may be referred to as anautoencoder.Entropy encoding processes such as range or arithmetic encoding are typically able to losslesslycompress given input data up to close to the fundamental entropy limit of that data, as determinedby the total entropy of the distribution of that data. Accordingly, one way in which end-to-end,learned compression can minimise the rate loss term of the rate-distortion loss function andthereby increase compression effectiveness is to learn autoencoder parameter values thatproduce low entropy latent representation distributions. Producing latent representationsdistributed with as low an entropy as possible allows entropy encoding to compress the latentdistributions as close to or to the fundamental entropy limit for that distribution. The lowerthe entropy of the distribution, the more entropy encoding can losslessly compress it and thelower the amount of data in the corresponding bitstream. In some cases where the latentrepresentation is distributed according to a gaussian or Laplacian distribution, this learningmay comprise learning optimal location and scale parameters of the gaussian or Laplaciandistributions, in other cases, it allows the learning of more flexible latent representationdistributions which can further help to achieve the minimising of the rate-distortion lossfunction in ways that are not intuitive or possible to do with handcrafted features. Examples ofthese and other advantages are described in WO2021 / 220008A1, which is incorporated in itsentirety by reference.Something which is closely linked to the entropy encoding of the latent distribution and whichaccordingly also has an effect on the effectiveness of compression of end-to-end learnedapproaches is the quantisation step. During inference, a rounding function may be used toquantise a latent representation distribution into bins of given sizes, a rounding function isnot differentiable everywhere. Rather, a rounding function is effectively one or more stepfunctions whose gradient is either zero (at the top of the steps) or infinity (at the boundarybetween steps). Back propagating a gradient of a loss function through a rounding functionis challenging. Instead, during training, quantisation by rounding function is replaced byone or more other approaches. For example, the functions of a noise quantisation model aredifferentiable everywhere and accordingly do allow backpropagation of the gradient of theloss function through the quantisation parts of the end-to-end, learned system. Alternatively, astraight-through estimator (STE) quantisation model or one other quantisation models may beused. It is also envisaged that different quantisation models may be used for during evaluationof different term of the loss function. For example, noise quantisation may used to evaluate therate or entropy loss term of the rate-distortion loss function while STE quantisation may beused to evaluate the distortion term.In a similar manner to how learning parameters to produce certain distributions of the latentrepresentation facilitates achieving better rate loss term minimisation, end-to-end learning ofthe quantisation process achieves a similar effect. That is, learnable quantisation parametersprovide the architecture with a further degree of freedom to achieve the goal of minimising theloss function. For example, parameters corresponding to quantisation bin sizes may be learnedwhich is likely to result in an improved rate-distortion loss outcome compared to approachesusing hand-crafted quantisation bin sizes.Further, as the rate-distortion loss function constantly has to balance a rate loss term against adistortion loss term, it has been found that the more degrees of freedom the system has duringtraining, the better the architecture is at achieving optimal rate and distortion trade off.Returning to the compression pipeline more generally, the systems described above may bedistributed across multiple locations and / or devices. For example, the encoder 110 may belocated on a device such as a laptop computer, desktop computer, smart phone or server. Thedecoder 120 may be located on a separate device which may be referred to as a recipient device.The system used to encode, transmit and decode the input image 5 to obtain the output image 6may be referred to as a compression pipeline.As described above in the context of quantisation, the AI based compression process mayfurther comprise a hyper-network 105 for the transmission of meta-information that improvesthe compression process. The hyper-network 105 comprises a trained neural network 115acting as a hyper-encoder ^^ ℎ ℎ^^and a trained neural network 125 acting as a hyper-decoder ^^^^.An example of such a is shown in Figure 2. Components of the system not furtherdiscussed may be assumed to be the same as discussed above. The neural network 115 acting asa hyper-decoder receives the latent that is the output of the encoder 110. The hyper-encoder 115produces an output based on the latent representation that may be referred to as a hyper-latentrepresentation. The hyper-latent is then quantized in a quantization process 145 characterisedby ^^ℎ to produce a quantized hyper-latent. The quantization process 145 characterised by ^^ℎmay be the same as the quantisation process 140 characterised by ^^ discussed above.In a similar manner as discussed above for the quantized latent, the quantized hyper-latent isthen entropy encoded in an entropy encoding process 155 to produce a bitstream 135. Thebitstream 135 may be entropy decoded in an entropy decoding process 165 to retrieve thequantized hyper-latent. The quantized hyper-latent is then used as an input to trained neuralnetwork 125 acting as a hyper-decoder. However, in contrast to the compression pipeline 100,the output of the hyper-decoder may not be an approximation of the input to the hyper-decoder115. Instead, the output of the hyper-decoder is used to provide parameters for use in theentropy encoding process 150 and entropy decoding process 160 in the main compressionprocess 100. For example, the output of the hyper-decoder 125 can include one or more ofthe mean, standard deviation, variance or any other parameter used to describe a probabilitymodel for the entropy encoding process 150 and entropy decoding process 160 of the latentrepresentation. In the example shown in Figure 2, only a single entropy decoding process 165and hyper-decoder 125 is shown for simplicity. However, in practice, as the decompressionprocess usually takes place on a separate device, duplicates of these processes will be presenton the device used for encoding to provide the parameters to be used in the entropy encodingprocess 150.Further transformations may be applied to at least one of the latent and the hyper-latent at anystage in the AI based compression process 100. For example, at least one of the latent and thehyper latent may be converted to a residual value before the entropy encoding process 150,155is performed. The residual value may be determined by subtracting the mean value of thedistribution of latents or hyper-latents from each latent or hyper latent. The residual valuesmay also be normalised.To perform training of the AI based compression process described above, a training set ofinput images may be used as described above. During the training process, the parameters ofboth the encoder 110 and the decoder 120 may be simultaneously updated in each trainingstep. If a hyper-network 105 is also present, the parameters of both the hyper-encoder 115and the hyper-decoder 125 may additionally be simultaneously updated in each training step.The training process may further include a generative adversarial network (GAN). Whenapplied to an AI based compression process, in addition to the compression pipeline describedabove, an additional neutral network acting as a discriminator is included in the system. Thediscriminator receives an input and outputs a score based on the input providing an indicationof whether the discriminator considers the input to be ground truth or fake. For example, theindicator may be a score, with a high score associated with a ground truth input and a lowscore associated with a fake input. For training of a discriminator, a loss function is used thatmaximizes the difference in the output indication between an input ground truth and input fake.When a GAN is incorporated into the training of the compression process, the output image 6may be provided to the discriminator. The output of the discriminator may then be used in theloss function of the compression process as a measure of the distortion of the compressionprocess. Alternatively, the discriminator may receive both the input image 5 and the outputimage 6 and the difference in output indication may then be used in the loss function of thecompression process as a measure of the distortion of the compression process. Training ofthe neural network acting as a discriminator and the other neutral networks in the compressionprocess may be performed simultaneously. During use of the trained compression pipelinefor the compression and transmission of images or video, the discriminator neural network isremoved from the system and the output of the compression pipeline is the output image 6.Incorporation of a GAN into the training process may cause the decoder 120 to performhallucination. Hallucination is the process of adding information in the output image 6 thatwas not present in the input image 5. In an example, hallucination may add fine detail tothe output image 6 that was not present in the input image 5 or received by the decoder 120.The hallucination performed may be based on information in the quantized latent received bydecoder 120.Details of a video compression process will now be described. As discussed above, a video ismade up of a series of images arranged in sequential order. AI based compression process100 described above may be applied multiple times to perform compression, transmissionand decompression of a video. For example, each frame of the video may be compressed,transmitted and decompressed individually. The received frames may then be grouped toobtain the original video.The frames in a video may be labelled based on the information from other frames that is usedto decode the frame in a video compression, transmission and decompression process. Asdescribed above, frames which are decoded using no information from other frames may bereferred to as I-frames. Frames which are decoded using information from past frames may bereferred to as P-frames. Frames which are decoded using information from past frames andfuture frames may be referred to as B-frames. Frames may not be encoded and / or decoded inthe order that they appear in the video. For example, a frame at a later time step in the videomay be decoded before a frame at an earlier time.The images represented by each frame of a video may be related. For example, a number offrames in a video may show the same scene. In this case, a number of different parts of thescene may be shown in more than one of the frames. For example, objects or people in a scenemay be shown in more than one of the frames. The background of the scene may also beshown in more than one of the frames. If an object or the perspective is in motion in the video,the position of the object or background in one frame may change relative to the position ofthe object or background in another frame. The transformation of a part of the image froma first position in a first frame to a second position in a second frame may be referred to asflow, warping or motion compensation. The flow may be represented by a vector. One or moreflows that represent the transformation of at least part of one frame to another frame may bereferred to as a flow map.An example AI based video compression, transmission, and decompression process 200 isshown in Figure 3. The process 200 shown in Figure 3 is divided into an I-frame part 201for decompressing I-frames, and a P-frame part 202 for decompressing P-frames. It will beunderstood that these divisions into different parts are arbitrary and the process 200 may bealso be considered as a single, end-to-end pipeline.As described above, I-frames do not rely on information from other frames so the I-frame part201 corresponds to the compression, transmission, and decompression process illustrated inFigures 1 or 2. The specific details will not be repeated here but, in summary, an input image^^0 is passed into an encoder neural network 203 producing a latent representation which isquantised and entropy encoded into a bitstream 204. The subscript 0 in ^^0 indicates the inputimage corresponds to a frame of a video stream at position t = 0. This may be the first frame ofan entire video stream or the first frame of a chunk of a video stream made up of, for example,an I-frame and a plurality of subsequent P-frames and / or B-frames. The bitstream 204 is thenentropy decoded and passed into a decoder neural network 205 to reproduce a reconstructedimage ^^0 which in this case is an I-frame. The decoding step may be performed both locallyat the same location as where the input image compression occurs as well as at the locationwhere the decompression occurs. This allows the reconstructed image ^^0 to be available forlater use by components of both the encoding and decoding sides of the pipeline.In contrast to I-frames, P-frames (and B-frames) do rely on information from other frames.Accordingly, the P-frame part 202 at the encoding side of the pipeline takes as input not onlythe input image ^^^^ that is to be compressed (corresponding to a frame of a video stream atposition t), but also one or more previously reconstructed images ^^^^−1 from an earlier framet-1. As described above, the previously reconstructed ^^^^−1 is available at both the encodeand decode side of the pipeline and can accordingly be used for various purposes at both theencode and decode sides.At the encode side, previously reconstructed images may be used for generating a flow mapscontaining information indicative of inter-frame movement of pixels between frames. In theexample of Figure 3, both the image being compressed ^^^^ and the previously reconstructedimage from an earlier frame ^^^^−1 are passed into a flow module part 206 of the pipeline. Theflow module part 206 comprises an autoencoder such as that of the autoencoder systems ofFigures 1 and 2 but where the encoder neural network 207 has been trained to produce alatent representation of a flow map from inputs ^^^^−1 and ^^^^ , which is indicative of inter-framemovement of pixels or pixel groups between ^^^^−1 and ^^^^ . The latent representation of the flowmap is quantised and entropy encoded to compress it and then transmitted as a bitstream 208.On the decode side, the bitstream is entropy decoded and passed to a decoder neural network209 to produce a reconstructed flow map ^^ .The reconstructed flow map ^^ is applied to the previously reconstructed image ^^^^−1 to generatea warped image ^^^^−1,^^. It is envisaged that any suitable warping technique may be used, forexample bi-linear or tri-linear warping, as is described in Agustsson, E., Minnen, D., Johnston,N., Balle, J., Hwang, S. J., and Toderici, G. (2020), Scale-space flow for end-to-end optimizedvideo compression. In Proceedings of the IEEE / CVF Conference on Computer Vision andPattern Recognition (pp. 8503-8512), which is hereby incorporated by reference. It is furtherenvisaged that a scale-space flow approach as described in the above paper may also optionallybe used. The warped image ^^^^−1,^^ is a prediction of how the previously reconstructed image^^^^−1 might have changed between frame positions t-1 and t, based on the output flow mapproduced by the flow module part 206 autoencoder system from the inputs of ^^^^ and ^^^^−1.As with the I-frame, the reconstructed flow map ^^ and corresponding warped image ^^^^−1,^^may be produced both on the encode side and the decode side of the pipeline so they areavailable for use by other components of the pipeline on both the encode and decode sides.In the example of Figure 3, both the image being compressed ^^^^ and the ^^^^−1,^^ are passedinto a residual module part 210 of the pipeline. The residual module part 210 comprises anautoencoder system such as that of the autoencoder systems of Figures 1 and 2 but where theencoder neural network 211 has been trained to produce a latent representation of a residualmap indicative of differences between the input mage ^^^^ and the warped image ^^^^−1,^^. Thelatent representation of the residual map is then quantised and entropy encoded into a bitstream212 and transmitted. The bitstream 212 is then entropy decoded and passed into a decoderneural network 213 which reconstructs a residual map ^^ from the decoded latent representation.Alternatively, a residual map may first be pre-calculated between ^^^^ and the ^^^^−1,^^ and thepre-calculated residual map may be passed into an autoencoder for compression only. Thishand-crafted residual map approach is computationally simpler, but reduces the degrees offreedom with which the architecture may learn weights and parameters to achieve its goalduring training of minimising the rate-distortion loss function.Finally, on the decode side, the residual map ^^ is applied (e.g. combined by addition, subtractionor a different operation) to the warped image to produce a reconstructed image ^^^^ which is areconstruction of image ^^^^ and accordingly corresponds to a P-frame at position t in a sequenceof frames of a video stream. It will be appreciated that the reconstructed image ^^^^ can then beused to process the next frame. That is, it can be used to compress, transmit and decompress^^^^+1, and so on until an entire video stream or chunk of a video stream has been processed.Thus, for a block of video frames comprising an I-frame and ^^ subsequent P-frames, thebitstream may contain (i) a quantised, entropy encoded latent representation of the I-frameimage, and (ii) a quantised, entropy encoded latent representation of a flow map and residualmap of each P-frame image. For completeness, whilst not illustrated in Figure 3, any of theautoencoder systems of Figure 3 may comprise hyper and hyper-hyper networks such as thosedescribed in connection with Figure 2. Accordingly, the bitstream may also contain hyper andhyper-hyper parameters, their latent quantised, entropy encoded latent representations and soon, of those networks as applicable.Finally, the above approach may generally also be extended to B-frames, for example as isdescribed in Pourreza, R., and Cohen, T. (2021). Extending neural p-frame codecs for b-framecoding. In Proceedings of the IEEE / CVF International Conference on Computer Vision (pp.6680-6689).The above-described flow and residual based approach is highly effective at reducing theamount of data that needs to be transmitted because, as long as at least one reconstructed frame(e.g. I-frame ^^^^−1) is available, the encode side only needs to compress and transmit a flowmap and a residual map (and any hyper or hyper-hyper parameter information, as applicable)to reconstruct a subsequent frame.Figure 4 shows an example of an AI image or video compression process such as that describedabove in connection with Figures 1-3 implemented in a video streaming system 400. Thesystem 400 comprises a first device 401 and a second device 402. The first and second devices401, 402 may be user devices such as smartphones, tablets, AR / VR headsets or other portabledevices. In contrast to known systems which primarily perform inference on GPUs such asNvidia A100, Geforce 3090, Gefore 4090 GPU cards, the system 400 of Figure 4 performsinference on a CPU (or NPU or TPU as applicable) of the first and second devices respectively.That is, compute for performing both encoding and decoding are performed by the respectiveCPUs of the first and second devices 401, 402. This places very different power usage, memoryand runtime constraints on the implementation of the above methods than when implementingAI-based compression methods on GPUs. In one example, the CPU of first and second devices401, 402 may comprise, for example, a Qualcomm Snapdragon CPU.The first device 401 comprises a media capture device 403, such as a camera, arranged tocapture a plurality of images, referred to hereafter as a video stream 404, of a scene 404. Thevideo stream 404 is passed to a pre-processing module 406 which splits the video stream intoblocks of frames, various frames of which will be designated as I-frames, P-frames, and / orB-frames. The blocks of frames are then compressed by an AI-compression module 407comprising the encode side of the AI-based video compression pipeline of Figure 3. Theoutput of the AI-compression module is accordingly a bitstream 408a which is transmittedfrom the first device 401, for example via a communications channel, for example over oneor more of a WiFi, 3G, 4G or 5G channel, which may comprise internet or cloud-based 409communications.The second device 402 receives the communicated bitstream 408b which is passed to anAI-decompression module 410 comprising the decode side of the AI-based video compressionpipeline of Figure 3. The output of the AI-decompression module 402 is the reconstructedI-frames, P-frames and / or B-frames which are passed to a post-processing module 411 wherethey can prepared, for example passed into a buffer, in preparation for streaming 412 to andrendering on a display device 413 of the second device 402.It is envisaged that the system 400 of Figure 4 may be used for live video streaming at 30fpsof a 1080p video stream, which means a cumulative latency of both the encode and decodeside is below substantially 50ms, for example substantially 30ms or less. Achieving thislevel of runtime performance with only CPU (or NPU or TPU) compute on user devicespresents challenges which are not addressed by known methods and systems or in the widerAI-compression literature.For example, execution of different parts of the compression pipeline during inferencemay be optimized by adjusting the order in which operations are performed using one ormore known CPU scheduling methods. Efficient scheduling can allow for operations to beperformed in parallel, thereby reducing the total execution time. It is also envisaged thatefficient management of memory resources may be implemented, including optimising cachingmethods such as storing frequently-accessed data in faster memory locations, and memoryreuse, which minimizes memory allocation and deallocation operations.A number of concepts related to the AI compression processes and / or their implementationin a hardware system discussed above will now be described. Although each concept isdescribed separately, one or more of the concepts described below may be applied in an AIbased compression process as described above.Concept 1: Training super resolution with learned image and video compressionFigure 5 illustrates an example of an image or video compression, transmission and decompres-sion pipeline 500. The pipeline illustrates a method of the present disclosure that correspondsto that described in relation to Figures 1 to 4. Like-numbered features correspond to those inFigures 1 to 4. However in Figure 5, pipeline is further wrapped in a super-resolution wrapper.That is, the encoder ^^ is preceded by a downsampler 501, and decoder ^^ is followed by anupsampler 502.We first introduce the super-resolution wrapper 501, 502 around the pipeline 500 duringtraining which comprises making the evaluated loss function be based on one or more termsfrom the list of: a difference between the output image ^^ and the input image ^^ , a differencebetween the output image ^^ and the downsampled input image ^^, a difference between theupsampled output image ^̂^ and the input image ^^ , and / or a difference between the upsampledoutput image ^̂^ and the downsampled input image ^^. Modifying the loss function in this wayallows the pipeline to be "super-resolution" aware. That is, the loss function comprises a termthat is not just based on differences between the input to the encoder and the output of thedecoder, but also or alternatively on the input ^^ and the output ^^ of the downsampler 501and / or the input ^^ and output ^^ of the upsampler 502. In this way, the weights of the neuralnetworks of the pipeline 500 (e.g. the encoder ^^ , the decoder ^^, and / or the correspondinghyper encoder and hyper decoders) will be optimised to allow the encoder ^^ to produce alatent representation that has a distribution that is optimally entropy encodable to hit low targetbit rates while at the same time allow the decoder ^^ to output images that have distributionsthat can be optimally upsampled by the upsampler 502 to produce upsampled output images ^̂^that are as close to the original input images ^^ as possible.More generally, making the neural compression pipeline 500 super-resolution aware in thisway during training results in trained networks of the pipeline 500 that, when wrapped in thesuper-resolution down- and / or up-samplers 501, 502 during inference, produces output images^̂^ that are closer to the the input images ^^ than a network or networks that were not trained ina super-resolution aware manner.Accordingly, the method comprises receiving an input image ^^ at a first computer system,downsampling the input image with a downsampler 501 to produce a downsampled inputimage ^^, encoding the downsampled input image ^^ using a first neural network ^^ to producea latent representation, decoding the latent representation using a second neural network ^^to produce an output image ^^, wherein the output image ^^ is an approximation of the inputimage (e.g. of the downsampled input image ^^ which in turn is an approximation of theoriginal input image ^^), upsampling the output image ^^ with an upsampler 502 to produce anupsampled output image ^̂^ , evaluating a function (i.e. a loss function) based on a differencebetween one or more of: the output image ^^ and the input image ^^ , the output image ^^ and thedownsampled input image ^^, the upsampled output image ^̂^ and the input image ^^ , and / or theupsampled output image ^̂^ and the downsampled input image ^^, updating the parameters ofthe first neural network ^^ and the second neural network ^^ based on the evaluated function,and repeating the above steps using a first set of input images to produce a first trained neuralnetwork and a second trained neural network.In the above-described method, it is envisaged that the upsampler 502 may comprise a thirdneural network, and the method further comprises updating the parameters of the third neuralnetwork based on the evaluated function.This is illustrated in Figure 6 showing a pipeline 600 corresponding to that of Figure 5 butwhere the upsampler 602 is shown as a neural network ^^ . The upsampler neural network ^^ isresponsible for upscaling the output image obtained from the second neural network back toits original size or resolution.The upsampler neural network ^^ may comprise a convolutional neural network architecture. Forexample, the upsampler may be implemented using transposed convolutions or deconvolutionlayers. These layers perform the inverse operation of regular convolutions and can be used toincrease the spatial resolution of an image.The upsampling process can be further enhanced using various techniques such as skipconnections or residual connections. Skip connections allow for direct transmission ofinformation from earlier layers in the network to later layers, bypassing some of the intermediatelayers and thereby allowing the model to leverage detailed information present in the initialstages of processing. Residual connections add the output of a layer directly to the input ofanother layer, effectively performing addition or subtraction operations within the network.These techniques can improve the accuracy and stability of the neural network-based upsamplerby allowing it to better capture fine details in the image.More specifically, a neural network based upsampler such as ^^ can be trained together, e.g. inan end-to-end manner with the trainable neural networks of the neural compression pipeline600, making the entire pipeline super-resolution aware. This is an important distinctioncompared to simply applying a super resolution wrapper to a neural or traditional compressionpipeline. This is because it affords the freedom of the networks of the pipeline to learn toproduce latent representations that compress optimally that get decoded into output images ^^that may not be visually pleasing or indeed look to the human visual system like an accuratereconstruction of the original input images ^^ or their downsampled versions ^^, but whichnonetheless have pixel value distributions that the upsampler neural network ^^ can take betteradvantage of to produce upsampled output images ^̂^ that are more accurate approximations ofthe original input images ^^ than an upsampler that is acting on an output of a compressionpipeline that is not super-resolution aware (in this case upsampler aware)In the example of Figure 6, the downsampler may be a traditional downsampler, for example itis envisaged that the downsampler may comprise a bilinear or bicubic downsampler.Bilinear and bicubic downsampling are exemplary methods used for image resizing. Theyinvolve reducing the resolution of an input image, for example by a factor of 2x2 (e.g., from100x100 to 50x50). Further exemplary details of bilinear and bicubic downsampling areprovided below.Bilinear Downsampling: In bilinear downsampling, the algorithm approximates the originalpixel values based on the average intensity values of its surrounding pixels in the scaled image.This method assumes that the pixel intensities are uniformly distributed across the image. Thebasic implementation for a 2x2 downsampling operation is as follows:1. Read the four input pixels (A, B, C, and D).2. Set the output pixel value to be an average of the input pixels.3. Replace the output pixel positions with their respective calculated values from step 2.Bicubic Downsampling: In bicubic downsampling, the algorithm uses a third-order polynomialfunction to approximate the original pixel values based on the intensities of a neighborhoodsurrounding the current pixel. This method takes into account more details than bilinearinterpolation but requires more computational resources. The basic implementation for a 2x2downsampling operation is as follows:1. Read the four input pixels (A, B, C, and D).2. Calculate the coefficients of the third-order polynomial function.3. Calculate the output pixel values: x and y, using the coefficients obtained in step 2.In the context of the the compression pipeline 600 shown in Figure 6, either bilinear andbicubic downsampling methods can be used. The choice between these two methods dependson the desired tradeoff between computation complexity and visual quality of the output imageswhereby bilinear is envisaged to be quicker and accordingly contributes to faster run timeson the encode side of the pipeline while retaining the neural network based upsampler ^^ toemphasise reconstruction accuracy.Optionally, the evaluated function (i.e. the loss function) may further be based on saiddifferences comprising a structural similarity index measure (SSIM). The SSIM is a qualitymetric that compares two images in terms of their structure and contrast. For example, whenused here it evaluates the similarity between an output image and its corresponding input imageor downsampled input image, upsampled output image, and / or any combination thereof. Byusing the SSIM as the evaluation function, the method aims to optimize the neural networks forpreserving the structural information in the images during the encoding and decoding process,thus improving the overall quality of the generated output images. As the human visual systemis often able to perceive these higher level structural information, it allows the nerwork tolearn to optimise for this type of difference over a simpler mean square error (MSE) loss.Alternatively, MSE may be used as it is quicker and simpler to calculate and can accordinglyspeed up training times.In the above-described method, the downsampling can be performed using a Gaussian blurfilter. That is, it is envisaged that downsampling can be achieved through an implementationwherein the input image ^^ is filtered with a Gaussian blur. The Gaussian blur filter is atype of low-pass filter that smoothes out the image by reducing high-frequency details whilepreserving lower frequency information. This helps to reduce the complexity of the inputimage and makes it easier for the first, second and third neural networks to learn the underlyingpatterns in the data and may help the loss to converge during training. In this implementation,the downsampled input image ^^ produced is a smoother representation of the original inputimage, which can be used as an input for encoding using the first neural network ^^ .It is envisaged that end-to-end training is preferable as it makes the pipeline fully super-resolution aware. However, it can be challenging to get the loss to converge during trainingand / or for training to be stable when all neural networks are being optimised simultaneously.This training instability and slow convergence can be mitigated by splitting the training intomultiple phases. For example, the method may comprise (i) updating the parameters of thefirst neural network ^^ and the second neural network ^^ based on the evaluated function fora first number of said steps to produce the first and second trained neural networks withoutperforming said upsampling and downsampling and without updating the parameters of thethird neural network ^^ , (ii) freezing the parameters of the first and second neural networks ^^ , ^^after said number first number of steps, and performing said upsampling and downsampling,and said updating of the parameters of the third neural network ^^ for a second number of saidsteps.That is, the underlying compression pipeline 600 is trained first. After this initial trainingphase, the parameters of the first and second neural networks ^^ , ^^ are frozen, and then thesystem proceeds with a secondary training phase where it updates the third neural network’s ^^parameters.As described above, this split training approach can mitigate training instability.Moving now to Figure 7 which shows a compression pipeline 700 similar to that of Figure 5,except now the downsampler comprises a neural network 701. The features that correspondto those of Figure 5 are not repeated here for brevity. More specifically, it is envisaged thatthe downsampler 701 may consist of a fourth neural network ^^. Further the training methodalso comprises updating the parameters of this fourth neural network ^^ based on the evaluatedfunction.The fourth neural network ^^, referred to as the downsampler 701, is responsible for reducingthe spatial resolution of the input image while preserving important features and details inorder to process them efficiently during encoding by the first neural network ^^ . In a similarmanner as described in connection with the upsampler in Figure 6, the downsampler neuralnetwork ^^ may be trained in an end-to-end manner with the first and second neural networks ^^ ,^^ of the compression pipeline 700. This approach provides the end-to-end system with anextra degree of freedom to produce downsampled input images ^^ that may not be in any wayvisually pleasing or accurate representations of ^^ but which can be optimally encoded by ^^into a latent representation that is distributed in a way that is efficiently entropy encodable andwhich can be decoded and subsequently upsampled into more accurate reconstructions of ^^than would otherwise be possible with a pipeline that is not super-resolution aware (in thiscase down-sampler aware).That is, unlike in traditional super resolution approaches where the downsampling attemptsto produce as accurate or visually pleasing representations of the input image, the outputof the neural network downsampler ^^ can produce whatever output is needed by the neuralcompression pipeline to help achieve a target bitrate while maintaining final output imageaccuracy.One examplary downsampler having a neural network architecture is a network comprisinga plurality of convolutional layers with decreasing filter sizes and increasing strides, as thisapproach can effectively reduce the spatial dimensions of the input image while maintainingits overall structure.In more detail, let’s consider a simple example of a downsampling process using a convolutionalneural network (CNN). The CNN architecture typically consists of multiple layers, each layerbeing composed of several filters applied across the spatial dimensions of the input image.These filters are learnable parameters that enable the network to extract various features fromthe image and recognize complex patterns or shapes within it.In the case of a downsampler, we can start with an initial convolutional layer having a largefilter size (e.g., 7x7) and small strides (e.g., 2x2). This combination results in a significantreduction in spatial dimensions while still allowing the network to capture essential informationfrom the input image. Subsequent layers can then use smaller filters (e.g., 3x3, 5x5) with largerstrides (e.g., 2x2, 4x4), further reducing the size of the feature maps while also encouragingmore localized receptive fields within the network. One drawback of large kernel sizes isthat they are more computationally expensive than smaller kernels, even if they are strictlyspeaking more expressive.The choice of downsampler, and filter sizes and strides in a downsampler controls the balancebetween preserving important image details and efficiently processing the data. In practice,the downsampler may comprise multiple convolutional layers with decreasing filter sizes andincreasing strides, followed by one or more max-pooling layers to further reduce the spatialdimensions of the input image. Reducing the number of layers and using small filters or kernelshelps to speed up run time.A further illustrative downsampler may be comprise a network architecture with a numberof layers with a stride greater than 1. Every such layer will downsample by a factor of thestride. While the filter or kernel sizes can affect the spatial dimensions of the input (i.e. biggerkernel resulting in greater downsampling) zero padding may be applied in such a way thatthe output remains the same size as the input if stride equals 1. This type of downsamplerstructure is typically fast and accordingly works well in the context of real time or near realtime compression.As alluded to above, in Figure 7, it is envisaged that the upsampler may comprise a bilinear orbicubic upsampler.Bilinear upsampling (or interpolation) is a method for upsampling an image. It involvesestimating pixel values by linearly interpolating between the neighboring pixels in the originaland downsampled images.An example implementation algorithm may be as follows:1. For each pixel in the output image, find its corresponding location or pixel-coordinate in theinput image.2. Find four nearby pixel coordinates around this central coordinate of the input image. Theseare typically referred to as northeast (NE), southeast (SE), northwest (NW), and southwest(SW).3. Compute a weighted average of these four pixels, where the weights depend on theirdistances from the desired output pixel location. The weights are usually determined by abilinear function.4. Repeat steps 1-3 for every pixel in the output image.Bicubic upsampling (or interpolation) is technique to upscale an image. It is similar to bilinearinterpolation but uses a bicubic function instead of a linear one. The algorithm is slightlymore complex, and the resulting images tend to have smoother edges than those produced bybilinear interpolation.An example implementaiton algorithm may be as follows:1. For each pixel in the output image, find its corresponding pixel in the input image.2. Find 16 nearby pixel coordinates around this central coordinate of the input image. Theseare typically referred to as northeast (NE), north-northeast (NNE), northwest (NW), southwest(SW), southeast (SE), south-southeast (SSE), south (SO), and south-southwest (SSW).3. Fit a bicubic function with coefficients that comprise weighted sums of the values of theinput pixels.4. Repeat steps 1-3 for every pixel in the output image.Both bilinear and bicubic interpolation can produce reasonably good results when upsamplingimages, but the choice between them will depend on the specific use case and desired level ofdetail preservation. In the present case, bilinear is envisaged to be preferred as it is faster andcan reduce runtime while working in a pipeline 700 with a (typically slower) neural networkbased downsampler such as ^^ on the encode side.As was the case with the neural network upsampler ^^ , training in a fully end-to-end mannermay introduce training instability and slow convergence. To address this, the training maybe split into phases. For example, the method may comprise (i) updating the parameters ofthe first neural network and the second neural network based on the evaluated function fora first number of said steps to produce the first and second trained neural networks withoutperforming said upsampling and downsampling and without updating the parameters of thefourth neural network, (ii) freezing the parameters of the first and second neural networks aftersaid number of steps, and performing said downsampling and said updating of the parametersof the fourth neural network for a second number of said steps.More generally, training a neural network based downsampler ^^ in an end-to-end manner isfurther complicated by it not being straightforward as to what the downsampler’s trainingobjective should be i.e. what its loss ought to be based on. For example, should the loss termsthat include the output of the downsampler ^^ compare the input image ^^ and the immediateoutput of the downsampler ^^ i.e. ^^, or on a downsampled version of ^^ that was produced by apreviously downsampled image (e.g. created by a traditional downsampling method) to tryto teach the downsampler to mimic a traditional downsampler, or on some other difference.The present inventors have found that as long as the loss function includes a term based oncomparing the output of ^^ with something that is not just the original input image ^^ but alsosome other output of the pipeline and / or a downsampled image produced by a traditionalmethod of downsampling, the loss converges more quickly indicating the neural networkdownsampler is learning to become super resolution aware.It is also envisaged that both the upsampler and downsampler may be neural networks. This isshown in Figure 8 which corresponds to Figure 5 but where the upsamplers and downsamplerscomprise neural networks. Like-numbered references indicate like features which are notrepeated here for brevity. Specifically Figure 8 illustratively shows a pipeline 800 comprisinga neural network downsampler ^^ 801 and a neural network upsampler ^^ 802 wrapped around aneural compression pipeline. It is envisaged that these may be trained in an end to end fashion.More specifically, by making the loss function be based on comparisons between not just theinput image ^^ and the final upsampled output image ^̂^ , but also between the outputs of thedownsampler ^^, and the various other outputs of the pipeline, as well as optionally a previouslydownsampled image, the network learns to become super-resolution aware and outperformsnetworks where the training of the neural compression pipeline is not connected in any wayeither through training or through the calculated comparisons of the terms of the loss function.Considering all of the above approaches more generally now, it is envisaged that the methodsdescribed above may comprise entropy encoding the latent representation into a bitstreamwith a specified length, wherein the function used in the method is also dependent on saidbitstream length. That is, the loss function further comprises a rate term. Including the rateterm in the loss function allows the networks to learn to optimise for bit rates (e.g. in bits perpixel) simultaneously with image reconstruction accuracy. Some or all of the loss terms maybe scaled or weighted with respect to each other to focus the learning on any of the objectivesas defined by the different loss terms.Additionally and / or alternatively, the difference between one or more of: the output imageand the input image, the output image and the downsampled input image, the upsampledoutput image and the input image, and / or the upsampled output image and the downsampledinput image may be determined based on the output of a fifth neural network acting as adiscriminator.It will be understood that the output of the discriminator may be the differentiation between aground truth image and a "fake" (i.e. compressed) reconstructed image during training. Thediscriminator loss term used in the training of the encoder / decoders of the AI compressionpipeline are only a function of the compressed image. In essence, the training tries to encouragethe neural networks to change in such a way that its output will be more realistic and like theground truth image. Faithfulness to the ground truth image is taken care of by the distortionloss term (e.g. mean squared error) or other loss. The evaluation of the function based onthese differences guides the process of updating the parameters of the first neural network(encoder) and the second neural network (decoder), leading to improved performance andbetter results over time. This approach can help to improve the overall quality of the generatedoutput images.It is accordingly also envisaged that the difference between one or more of: the output imageand the input image, the output image and the downsampled input image, the upsampled outputimage and the input image, and / or the upsampled output image and the downsampled inputimage comprises a mean squared error (MSE) and / or a structural similarity index measure(SSIM). These correspond to the distortion term of the loss function.MSE measures the average squared difference between two images, while SSIM computesstructural similarity based on luminance, contrast, and structure. By using these metrics, themethod is able to optimize the parameters of the neural networks to better approximate theinput image in subsequent iterations, resulting in improved image reconstruction performanceover multiple runs with different sets of input images.In the method as described above, the function may further comprise a term defining a visualperceptual metric that models how a human visual system may perceive differences. In theabove-described method, it is envisaged that the term defining a visual perceptual metric maycomprise a MS-SIM loss. This loss function serves to gauge how effectively the network isapproximating the input image with the output image. By iteratively minimizing this lossfunction through parameter updates in the neural networks, the trained neural network improvesits ability to generate an output image that closely resembles the input image.It is further noted that the above described methods may be used in the context of anypre-trained neural compression network, and accordingly the present disclosure envisages amethod where only the weights of the upsampler and / or downsampler are updated duringtraining. Accordingly, such a method comprises receiving an input image at a first computersystem, encoding the input image using a first neural network to produce a latent representation,decoding the latent representation using a second neural network to produce an output image,wherein the output image is an approximation of the input image. Next, the output image isupsampled with an upsampler comprising a third neural network. Thereafter, the differencebetween one or both of: the output image and the input image, and / or the upsampled outputimage and the input image is evaluated using a function, parameters of the third neural networkare updated based on the evaluated function. These steps are then repeated using a first set ofinput images to produce a first trained neural network and a second trained neural network.Or in the case where only the downsampler weights are trained, the present disclosure envisagesa method comprising receiving an input image at a first computer system, downsampling theinput image with a downsampler comprising a fourth neural network to produce a downsampledinput image, encoding the downsampled input image using a first neural network to producea latent representation, decoding the latent representation using a second neural network toproduce an output image that approximates the input image, evaluating a function based ondifferences between the output image and other images, updating the parameters of the fourthneural network based on the evaluated function, and repeating these steps with a first set ofinput images to create trained versions of the first and second neural networks.As described above, it is envisaged that producing the previously downsampled input imagemay be performed by either bilinear or bicubic downsampling techniques.Finally the present disclosure also proposes using a network trained in accordance with theabove-described methods. That is: receiving an input image at a first computer system,downsampling the input image with a downsampler, encoding the downsampled input imageusing a first trained neural network to produce a latent representation, transmitting the latentrepresentation to a second computer system, decoding the latent representation using asecond trained neural network to produce an output image, wherein the output image is anapproximation of the input image, and upsampling the output image with an upsampler toproduce an upsampled output image.A specific example use cases of the above described super resolution approaches will now bedescribed: more flexible bitrate ladders.A bitrate ladder refers to a set of predefined bitrates applied within an encoded file. In thecontext of video compression, it refers to a series of different bitrates that can be chosen toachieve the desired trade-off between video quality and file size.In general, when encoding video files, a balance is struck between two conflicting goals:achieving high visual quality while minimizing the file size. The process involves convertingthe raw data into a compressed format that requires less storage space.To accomplish this task in traditional compression, a number of known algorithms are used,often dictated by compression codec standards. One such standard is H.264 / AVC (AdvancedVideo Coding), which is widely adopted due to its balance between encoding complexityand image quality. The H.264 standard includes various profiles and levels that define themaximum bitrate and other parameters for a given video stream. Implementers of the standardtypically target these profiles and levels to ensure their implementation is standard compliant.More specifically, When using these standards, implementers can choose from different presetbitrates. These predefined bitrates are often referred to as a "ladder" because they represent aseries of steps or options available when choosing the optimal encoding settings for a givenvideo file. Typically, bitrate ladders make use of the idea that the encode resolution can bevaried a priori whereby the streaming provider has already pre-encoded its content at a pluralityof different resolutions, which in turn facilitates giving an end user a choice of what qualitysetting to apply given some particular bitrate budget.The bitrate ladder approach works by progressively decreasing the bitrate from one level toanother, allowing a given balance between quality and file size to be found. For example,starting at a higher-than-desired bitrate, you can gradually reduce it until you reach an acceptablelevel of visual degradation without sacrificing too much detail or clarity.In a neural compression pipeline, the neural networks are typically trained to perform optimallyat a given bitrate. Accordingly, covering all the predefined bitrates of a given bitrate ladder mayrequire training separate neural networks for each predefined bitrate which can be burdensomeand may result in the final codec library memory footprint being potentially very large.To overcome this issue, the above-described approaches may be used to make the base,neural compression neural networks super-resolution aware and allows a given base, neuralcompression pipeline to be used not only for compression to its targeted bitrate, but also toother bitrates in the bitrate ladder by applying the super-resolution wrapper around the basemodels when desired. This in turn makes it more viable to use neural compression pipelinesin the implementation of a bitrate ladder to comply with a given codec standard, or to aimfor an entirely different bitrate ladder that provides even better rate-quality trade offs thantraditional codec bitrate ladders, given how significantly better neural compression pipelinescan compress images and video compared to known codecs.Concept 2: Lightweight convolutional downsampling and upsamplingAs has been explained above in the general section, there is a wide gap between the ideationof high level ideas, the research-stage implementations of super resolution architectures,and the production level implementation of such architectures. This is particularly the casewhere the implementation is intended to be performed in real time or near real time onresource-constrained hardware, such as edge devices. For example, a research-stage approachthat works well and fast on a GPU such as an NVIDIA 4090, A10 or A100 card is very unlikelyto achieve the same performance on resource-constrained mobile device platforms such aslaptops, tablets and smartphones.One area of AI-based video compression where this is particularly problematic is in theimplementation of downsampling and upsampling algorithms.More particularly, one common component of such upsampling and downsampling algorithmsis a process known as PixelShuffle and PixelUnshuffle. Both operations manipulate thearrangement of data in tensors (multi-dimensional arrays) that represent images.PixelShuffle is often used in super-resolution models. That is, in general terms, PixelShuffleincreases the resolution of an input image by rearranging the elements of a tensor.The pseudocode below illustrates a PixelShuffle operation:Input Tensor: shape [^^^^^^^^ℎ^^^^^^^^, ^^ ∗ ^^2, ^^, ^^], where: C is the number of channels (e.g., 3for an RGB image), r is the upscale factor, H and W are the height and width of the tensor.Rearrangement of Data: PixelShuffle rearranges elements in this tensor to form a new tensorof shape [^^^^^^^^ℎ^^^^^^^^, ^^, ^^ ∗ ^^, ^^ ∗ ^^]. Essentially, it "shuffles" the data from the channeldimension into the spatial dimensions (height and width).Upscaling: This operation effectively upscales the image by a factor of r, increasing theresolution. For example, if r = 2, each pixel in the original image is rearranged to form a 2x2block in the output image.Application in Super-Resolution: In super-resolution models like EDSR or SRGAN, PixelShuf-fle is used in the latter stages to upscale the low-resolution input to a high-resolution output.It’s a part of the sub-pixel convolution technique where the model first increases the number ofchannels with additional convolutions and then uses PixelShuffle to upscale the image spatially.PixelUnshuffle is the reverse operation of PixelShuffle. It’s used to decrease the spatialresolution of an image while increasing the number of channels.The pseudocode below illustrates a PixelUnshuffle operation:Input Tensor: shape [^^^^^^^^ℎ^^^^^^^^, ^^, ^^, ^^].Rearrangement of Data: PixelUnshuffle rearranges elements to form a new tensor of shape[^^^^^^^^ℎ^^^^^^^^, ^^ ∗ ^^2, ^^ / ^^, ^^ / ^^]. It does so by taking spatial blocks of size r x r and stackingthem depth-wise in the channel dimension.Downscaling: This process effectively downscales the image by a factor of r, reducing itsspatial dimensions. For example, if r = 2, a 2x2 block of pixels in the input image is rearrangedinto a single pixel in the output, with the depth (channels) increased by a factor of 4.Application: PixelUnshuffle can be used in tasks like feature extraction, where reducing spatialresolution while retaining information in the channel dimension might be beneficial. It’s alsouseful in certain generative models or autoencoders where manipulating spatial resolution atdifferent stages of the network is required.More generally, PixelShuffle is used for upscaling an image by rearranging the channel datainto the spatial dimensions, whereas PixelUnshuffle does the opposite, downscaling an imageby rearranging spatial data into the channel dimension.PixelShuffle and PixelUnshuffle are specific implementations of "depth-to-space" and "space-to-depth" operations. These can be explained in their generalised form as follows.The pseudocode below illustrates a depth-to-space operation:Input Tensor: The operation takes an input tensor of shape [^^^^^^^^ℎ^^^^^^^^, ^^ ∗ ^^2, ^^, ^^], whereC is the number of channels, r is the upscale factor, and H and W are the height and width.Rearrangement of Data: It rearranges the elements of this tensor to form a new tensor of shape[^^^^^^^^ℎ^^^^^^^^, ^^, ^^ ∗ ^^, ^^ ∗ ^^]. This rearrangement involves redistributing the elements fromthe depth (channels) into the spatial dimensions (height and width).Upscaling Effect: The result is an upscaling of the image or feature map by a factor of r, witheach pixel in the original tensor contributing to a block of pixels in the output tensor.The pseudocode below illustrates a space-to-depth operation:Input Tensor: the input of a space-to-depth operation is a tensor of shape [^^^^^^^^ℎ^^^^^^^^, ^^, ^^, ^^].Rearrangement of Data: Space-to-Depth rearranges elements to produce a new tensor ofshape [^^^^^^^^ℎ^^^^^^^^, ^^ ∗ ^^2, ^^ / ^^, ^^ / ^^]. It does this by taking blocks of pixels from the spatialdimensions and stacking them in the channel dimension.Downscaling Effect: This leads to a reduction in the spatial resolution by a factor of r, whileincreasing the depth (channels) of the tensor.A problem with all of the above approaches is that they involve a very large number of smallparallelisable operations. These can be very efficiently performed in parallel on GPUs but causebottlenecks and dramatic decreases in performance on CPUs or other more resource-constrainedhardware.The present inventors have realised that both depth-to-space and space-to-depth operationsused in upsampling and downsampling can be replaced wholly with convolutional operations.This is made possible due by virtue of the realisation that the change in dimensions fromthe key "rearrangement of data" steps of depth-to-space and space-to-depth can be achievedby one or more convolutional operations, for example performed by corresponding one ormore convolutional layers in a neural network. Convolutional operations are typically alreadyoptimised at the chip level in commonly available commercial hardware chips on laptops,tablets and smartphones (such as the M2 and M3 Apple chips, and the Qualcomm snapdragonchips, and the Meteor Lake Intel chips, and others). Accordingly, replacing depth-to-space andspace-to-depth with convolutions allows upsampling and downsampling to be performed farmore efficiently than traditional depth-to-space and space-to-depth operations.Whilst the above improvement is envisaged to be used in the context of, for example super-resolution such as that described in concept 1 such as in the upsamplers and / or downsamplersin Figures 5 to 8. It is generally applicable to any instances where upsampling or downsamplingmight be performed in a neural compression pipeline. For example, one or more layers in thefirst or second neural networks of Figures 1, 2, 11 or 14 may be configured to downsampleor upsample an input. Or there may be an intermediate layer within these networks, ormodules (not shown in the Figures) positioned throughout the pipeline that may performdownsampling or upsampling. In each of these cases it is envisaged that these downsampling orupsampling operations may be performed without depth-to-space or space-to-depth approaches,but with convolutional operations instead. Alternatively, there may be some combination ofdepth-to-space or space-to-depth in some instances but convolutional operations in others.For example, the flow module 206 in Figure 3 may comprise one or more layers or modulesconfigured to downsample the input. This is one way to speed up runtime as the flow often neednot be estimated in as high a resolution as the input image as the output quality of reconstructedimages created with high resolution flow can be similar to that of the quality of reconstructedimages created with a low resolution flow. This downsampling, if performed using traditionalspace-to-depth would be a bottleneck. However, replacing space-to-depth with one or moreconvolutional operations removes this bottleneck due to the convolutional operations beingfaster and better optimised for hardware-constrained platforms. The corresponding inverseupsampling may then be applied at the output of the flow module 206. again using convolutionaloperations rather than depth-to-space.A corresponding set of operations may be performed with the residual module of Figure 3, inany of the modules of the hypernetwork in Figure 2, and / or in any of the compression Pipelineof Figure 1.An exemplary implementation of mimicking space-to-depth (i.e. downsampling) may be asfollows.CONVOLUTIONAL LAYER SETUP:Kernel Size: The kernel size is envisaged to match the block size that is to be mimicked.For example, if the block size for space-to-depth is r (say, 2 for a 2x2 block), the correspondingconvolution kernel size would be r x r (2x2 in the above example).Stride: The stride size is envisaged to equal the block size (r). This ensures that theconvolutional filters move across the image in steps equal to the block size, effectively capturingthe spatial blocks.Number of Filters: It is envisaged that the number of filters is set to ^^ ∗ ^^2, where C is theoriginal number of channels. This ensures that each filter produces an output that correspondsto one depth level in the Space-to-Depth transformationSEQUENTIAL CONVOLUTION LAYERS:To fully replicate Space-to-Depth, it is envisaged that a number of convolutions maybe used sequentially, e.g. by using a series of convolutional layers. This allows for thehandling of where the channel increase (^^^^^^^^ ∗ ^^2) is significant. Each layer progressivelyaccumulates more spatial information into the depth dimension. For completeness, it is alsopossible to replicate space-to-depth with a single strided convolution. Splitting it into multipleconvolutions with activations between them provides additional expressive power, but isn’tneeded to just replicate the functionality of space-to-depth.ACTIVATION FUNCTIONS:It is also envisaged that one or more activation functions may be applied betweensequential convolution layers. These help in spreading out the information spatially. Anexemplary activation function may be ReLU, which introduces a non-linearity and helps inlearning spatial patterns.CHANNEL REARRANGEMENT:Finally, after these convolutional operations, the output channels may optionally berearranged to match the order that a space-to-depth operation produces. Alternatively, this canbe done at any point in the process, and need not be done "on the fly".An exemplary implemenation of mimicking depth-to-space (i.e. upsampling) may be asfollows: CONVOLUTIONAL LAYER SETUP:Kernel Size: A kernel size is envisaged that aligns with the desired spatial expansion. Forexample, if the upscale factor is r, a larger kernel size (like 3x3 or larger) can be more effectivein spreading out the information across a larger spatial area. Alternatively, the depth-to-spaceoperation can be implemented with a single transposed convolution with stride equal to theupsampling factor. It is also possible to make a more expressive process by using larger kernelsizes and / or splitting a more extensive upsample into multiple stages.Stride: It is envisaged that the stride may be set to 1, ensuring a uniform spread ofinformation. Or the stride may be equal to the upsampling factor, as described above.Number of Filters: This may be less than the original number of channels to reduce thechannel dimension gradually. The exact number can vary depending on the architecture anddesired output.SEQUENTIAL CONVOLUTION LAYERS:Multiple convolutional layers may be advantageous, especially if the change from depthto spatial dimensions is significant. Each layer can gradually increase the spatial dimensionsand reduce the depth.ACTIVATION FUNCTIONS:It is also envisaged that one or more activation functions may be applied betweensequential convolution layers. These help in spreading out the information spatially. Anexemplary activation function may be ReLU, which can introduce non-linearity and help inlearning spatial patterns.UPSAMPLING LAYERS:Alongside the convolutional layers, upsampling layers (like nearest neighbor or bilinearupsampling) can be used to increase the spatial dimensions. These can be alternated withconvolutional layers to progressively achieve the desired spatial expansion.The above exemplary implementation is illustrative only and is not intended to be limiting. Forexample, any suitable kernel size, stride and filter numbers are envisaged, as are the number ofoptional sequential convolution layers, activation functions and other steps.By way of illustration, Figure 9 shows an example sequence of layers of a neural networkwhich takes an input image and downsamples it. The sequence of layers comprises a 3x3Conv layer, a ReLU activation function, a space-to-depth (2x) operation, a 1x1 conv layer,another ReLU activation layer and finally a depth-to-space (3x) operation. This approachuses space-to-depth and depth-to-space. Implementing this sequence of layers in resourceconstrained environments e.g. on a CPU results in bottlenecks.In contrast, Figure 10 illustratively shows an example sequence of layers of a neural networkor compression pipeline, for example one or more of the neural networks or compressionpipelines shown in any of Figures 1 to 8 but where the space-to-depth and depth-to-spaceoperations have been replaced by convolution operations. For example, the depth-to-spacereplacement comprises a strided transposed convolution and the space-to-depth replacementcomprises a strided convolution. This implementation substantially reduces the bottleneckswhen running in resource-constrained environments such as on a CPU.Concept 3: Regional Kernel Absolute Deviation (RKADe) for FlowIn image processing, Mean Absolute Difference (MAD) is a technique for detecting andnumerically estimating differences between pixels and / or pixel patches, that is, differencesbetween the values of a one or more pixels in one image and the values one or more pixels inanother image.In the context of AI-based compression pipelines, MAD may be used in the estimation of costvolumes in order to estimate flow as part of a flow-residual compression pipeline such as thatshown in Figures 3, 11 and 14.Figure 3 has already been described above. Figure 11 illustrates an example of a flow module,in this case a network 1100, configured to estimate information indicative of a differencebetween an image ^^^^−1 and an image ^^^^ , e.g. flow information. Figure 11 is provided as anon-limiting example of how such flow information may be calculated between two images.The flow module may be used in or together with the flow module part 206 (Figure 3) of thecompression pipeline. An alternative approach to estimating flow is described in Mentzer, F.,Agustsson, E., Ballé, J., Minnen, D., Johnston, N., and Toderici, G. (2022, November). Neuralvideo compression using gans for detail synthesis and propagation. In Computer Vision–ECCV2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, PartXXVI (pp. 562-578), which is hereby incorporated by reference.The network 1100 in Figure 5 comprises a set of layers 501a, 501b respectively for an image^^^^−1 and an image ^^^^ from respective times or positions in a sequence ^^ − 1 and ^^ of a sequenceof frames. The set of layers 1101a, 1101b may define one or more convolution operationsand / or nonlinear activations (for example as described in concept 2 above) to sequentiallydownsample the input images to produce a pyramid of feature maps for different levels ofcoarseness or spatial resolution. This may comprise performing ℎ / 2 ^^ / 2 downsamplingin a first layer, ℎ / 4 ^^ / 4 downsampling in a second layer ℎ / 8 ^^ / 8 downsampling in a thirdlayer, ℎ / 16 ^^ / 16 downsampling in a fourth layer, ℎ / 32 ^^ / 32 downsampling in a fifth layer,ℎ / 64 ^^ / 64 downsampling in a sixth layer, and so on. It will of course be appreciated thatthese downsampling operations and levels of coarseness or spatial resolution of a pyramidfeature map are exemplary only and other layers, operations and levels are also envisaged. Forexample, other operations, not only those from concept 2 above, may be used.With the downsampling operations performed and the corresponding pyramid of feature mapsgenerated, a first cost volume 1102 is calculated at the most coarse level between the featuremap pixels of the first image ^^^^−1 and the corresponding feature map pixels of the second image^^^^ . Cost volumes define the matchmaking cost of matching the pixels in one image with thepixels in a second image (which may be later in time, or earlier in time, for example due to theorder in which B-frame processing typically occurs which is not necessarily the chronologicalorder of the frames). That is, the closeness of each pixel, or a subset of all pixels, in the initialimage to one or more pixels in the later image is determined with a measure of similaritysuch as a vector or dot product, a cosine similarity, a mean absolute difference, or some othermeasure of similarity. This metric may be calculated against all pixels in the later image, oronly for pixels in a predetermined search radius such as a 1-10 pixel radius (preferably a 1, 2,3, or 4 pixel radius), or some other radius as described in connection with concept 4 below,around the pixel coordinate corresponding to the pixel against which the closeness metric isbeing calculated. This process is computationally expensive in floating points space but whichcan be implemented efficiently in integer or fixed point space.Once the first cost volume 1102 at the most coarse level is estimated, a first flow 1103 can beestimated from the first cost volume 1102. This may be achieved using, for example a flowextractor network which may comprise a convolutional neural network comprising a pluralityof layers trained to output a tensor defining a flow map from the input cost volumes. Othermethods of calculating flow information from cost volumes will also be known to the skilledperson.The same process is then repeated for the other levels of feature map coarseness to calculatea second cost volume 1104 and second flow 1105, and so on for the cost volumes and flowsassociated with each of the levels of coarseness until they have all been calculated, up to thefinal cost volume 1106 and flow 1107.The weights and / or biases of any activation layers in network 1100 (e.g. optionally in thedownsampling convolution layers and / or in a flow extractor network that produces flow mapsfrom the cost volumes) are trainable parameters and can accordingly be updated during trainingeither alone, or in an end to end manner with the rest of the compression pipeline. The trainablenature of these parameters provides the network 1100 with flexibility to produce feature mapsat each level of spatial resolution (i.e. pyramid feature maps) and / or at the flow outputsthat are forced into a distribution that best allows the network to meet its training objective(e.g. better compression, better reconstruction accuracy, more accurate reconstruction of flow,and so on). For example, it allows the network 1100 to produce feature maps that, whencost volumes and / or flows are calculated therefrom, produce cost volumes or flows that aredistributed roughly matching the latent representation distribution that would previously havebeen expected to be output by a dedicated flow encoder module. This effectively allows adedicated flow encoder to be omitted entirely from the flow compression part of the pipeline.Optionally, for each level of coarseness or resolution, the flow of the previous level or levelsof coarseness or resolution may be used to warp 1108, 1109, the feature maps before thecost volume is calculated. This has the effect of artificially reducing the amount of relativemovement between the pixels of the t and t - 1 images or feature maps when calculating thecost volumes, reducing flow errors for high movement details. Optionally, removing warpingentirely or in some levels of coarseness or resolution can substantially reduce run-time of flowcalculation while maintaining good levels of flow accuracy.As the warping process uses inputs from different levels of coarseness or spatial resolution,the flow estimation output may be upsampled 1110, 1111 (for example using the methods ofconcept 2, or using any other upsampling method) first to match the coarseness resolution ofthe feature map to which the flow is being applied in the warping process.The outputs of the flow module may accordingly be one or more cost volumes or somerepresentation thereof, and / or one or more flows or some representation thereof).The flow or representation thereof may then be transmitted in a bitstream and decoded by aflow decoder, the output of which may in turn be used to warp a previously decoded image foruse in the residual encoder / decoder 1110 arrangement as shown in e.g. Figure 3.Returning now to the estimation of cost volumes, the cost volumes may be used to compute local(translational) alignment through patch-wise comparisons. For example, let ^^, ^^ ∈ R^^×^^×^^be two tensors each with ^^ ∈ N channels and spatial dimensions ^^ × ^^ ∈ N2. A standardcost-volume to compute isCV(^^, ^^)^^, ^^ := ^^^(:,^^) , ^^ (:,^^+ ^^)^,where the subscript denotes , ^^ ∈ [^^] × [^^] and ^^ = ( ^^1, ^^2) ∈{−^^, . .. , ^^}2 with ^^ ∈ N being the radius of the cost volume.The function above ^·, ·^ is the inner product and means that the cost volume is constructed bycomputing the channel-wise correlation.MAD may be used in the calculation of cost volumes as follows.Let P : R^^×^^×^^ → R^^×(2^^+1)2×^^×^^ be a patch operator that associates to each pixel^^ ∈ {^^} centered at pixel ^^ for an integer ^^ ≥ 0.Then the MAD cost volume is defined by∑^^^ cv (^^, ^^) := ∥P(^^) − P( 2mad ^^, ^^ ^^,^^ ^^)^^,^^+ ^^ ∥1, ^^ ∈ {^^} × {^^}, ^^ ∈ {−^^, .. . , ^^} . Exemplary r values that may be used include r = 1 (giving comparisons of 3x3 patches), r = 2(giving comparisons of 5x5 patches), r = 3 (giving comparisons of 7x7 patches), and more.The MAD-based cost volume estimation is more computationally efficient than other knowncost volume estimation methods and accordingly synergistically helps to reduce run times ofthe flow estimation of an AI-based compression pipeline. However, the present inventors havefound that implementing the operations used to perform MAD calculations at the machinelevel has a number of downsides, particularly when the MAD calculations are implementedusing convolutions.Specifically, an input tensor of arbitrary ^^^^^^^^^^ dimensions is typically stored in non-contiguous blocks of memory. For example, in one example, the values of the elements of^^^^ and ^^^^ for one channel ^^^^ of an input tensor may be stored in a first block of memorywhile the values of the elements ^^^^ and ^^^^ for another channel ^^^^ may be stored in a secondblock of memory and so on. When a MAD value is estimated that involves this multi-channelinput tensor, the values of the elements of ^^^^ and ^^^^ are accessed from one of the memoryblocks, the values of the elements of ^^^^ and ^^^^ are accessed from another of the memoryblocks, and so on. This means the number of memory access operations can be very higheven for relatively simple operations. Consider for example the following pseudo code forimplementing the above-described MAD approach may be as follows:1) Receive input tensor input_1 representing first image, and input tensor input_2 represent-ing second image.2) Apply repeated interleaving of input_1 to obtain tensor x of desired shape:x = repeat_interleave(input_1).3) Apply flattening and unfolding of input_2 to obtain tensor y of desired shape:y = flatten_unfold(input_2).4) Calculate absolute difference between tensors x and y:absolute_differences = absolute_difference_calculation(x, y).5) Sum the absolute differences with a strided sum operation:mad_output = strided_sum(absolute_differences).In the above example, the strided sum operation comprises depth or channel-wise groupedconvolutions (i.e. convolutions applied in depth-wise groups, each group taking as inputs thevalues of spatial dimensions ^^ and ^^ stored in non-contiguous memory blocks). That is, thestride is in the depth (i.e. channel) dimension necessitating the access and retrieval of thevalues of the elements of ^^ and ^^ stored in separate memory blocks. Alternatively, if thegrouping of the convolutions is in the spatial dimensions rather than the depth dimension, thenon-contiguous memory block problem still arises but now in the depth dimension. In otherwords, the use of grouped convolutions results in a convolution-based MAD approach that hasa memory access bottleneck during run time.Accordingly, implementing MAD-based, naive, patch-wise comparisons using convolutionaloperations on most commercial hardware CPUs, GPUs and NPUs (i.e. neural accelerators)is slow due to the interleaving of data in memory that results from the order in which theoperations of the convolutions are performed.A schematic of this strided sum based implementation of a MAD cost volume calculation isshown in Figures 12a, 12b and 12c. Illustrated in Figure 12a is a toy representation of inputtensors 1200a, 1200b respectively associated with a first image and a second image. Eachinput tensor has 3 channels: channel 1 (Ch1), channel 2 (Ch2) and channel 3 (Ch3). In a firststep, repeat interleaving 1201 is performed on the first image input tensor 1200a to produce afirst intermediate output. In a second step, an unfold convolution operation 1202 is performedon the second image input tensor 1200b to produce a second intermediate output. The unfoldconvolution operations are performed as group convolutions and accordingly each group isassigned its own memory block. In a third step 1203, an absolute difference is estimatedbetween the intermediate outputs of the first step and the second step to produce a thirdintermediate output. As is shown in Figure 12b, the elements of the third intermediate step arestored in said respective, different memory blocks - in this case memory blocks 1, 2 and 3(Mb1, Mb2, Mb3) to match the number of outputs Finally, a strided sum 1204 is performedon the estimated absolute differences stored in the respective memory blocks 1, 2, and 3 toproduce the MAD output cost volume tensor. Again by virtue of the group convolutions ofthe strided sum 1204 and by virtue of the memory block locations, this operation requiresnon-contiguous memory blocks to be accessed for each of the convolutions of the strided sum1204, as is illustrated in Figure 12c. Whilst it would in principle be possible to introduce ashuffling of the layers of the input tensor earlier in the flow, for example before the unfoldconvolutions and / or before the strided sum convolutions, this introduces an additional shufflingoperation which is also slow given the number of memory read and write operations requiredto complete a full interleaving or de-interleaving process, and can also further complicate anyother parts of the AI compression pipeline that rely on this information in unshuffled form.The additional shuffling operation to resolve the downstream issue is accordingly undesirableand does not result in appreciable run-time improvements.To address these and other problems, it is helpful to first consider in more detail the purpose ofthe cost volume in the presently described AI-compression pipeline. Specifically, the goal ofthe cost volume is to construct a spatial comparison operator that encodes some notion of howtwo patches are related. For a given pixel ^^, the cost volume thus has a measure of comparisonbetween ^^ and ^^ for a collection of local offsets ^^ . The present inventors have realised that anyeffective encoding of this information can suffice in AI-compression pipelines because theneural networks that make up the flow encoder and / or decoder and residual encoder and / ordecoder, and indeed any other neural networks that make up the AI-compression pipelines ofFigures 1-14, are able to learn to accommodate the encoding of this information, regardlessof how it is represented. It cannot be understated how significant of an advantage this is forend-to-end AI-based compression pipelines over traditional compression methods.In particular, this facilitates the use of approaches to estimating cost volumes that simplyaren’t viable in traditional compression pipelines. Presently described concept 3 is directedto such approaches, which are described herein as compressive approaches. That is, the useof a compressive encoding of the measure of comparison between ^^ and ^^ for a collection oflocal offsets ^^ , is made possible and in particular a compressive cost volume estimation (i.e. acompressively encoded cost volume) is made possible.In traditional, local cost volume approaches, a pixel is compared to pixels in a given referenceimage that lie within a given radius. In the context of flow estimation, this may be thecomparison of a pixel in a first image to pixels in a radius around a corresponding pixelcoordinate in a second image. For example, in a radius ^^ local cost volume, one must comparea pixel to (2^^ + 1)2 reference pixels (e.g., for radius 3 there are 49 reference pixels in a 7x7block). In classical, local cost volume approaches a version of this mapping might be:^^^^^( ^^) =^ ^^^^^ − ^^^^+ ^^^ .This approach effectively comprises a naive patch-wise comparison of the two images.An inductive bias for structured data suggests that the above map is approximately sparse in abasis, meaning that the data can be equivalently represented in a low-dimensional subspaceapproximately logarithmic in the dimension of the ambient space.Thus, a compressive encoding of the cost volume may look like:^^ · ^^^^( ^^),where ^^ is an appropriately chosen random matrix with ^^ ∈ R^^×^^ satisfying ^^ ∼ O(log ^^). From this, the present inventors have realised that it is possible to build a learned mappingthat replaces the above-described classical way, naive patch-wise comparison approach ofestimating local cost volumes that require large number of operations. The learned mappingeffectively bypasses the direct computation of the local cost volume and instead computes thelower dimensional compressively encoded cost volume ^^ · ^^^^( ^^) directly. This compressivelyencoded cost volume, or a representation thereof produced by a post-processing step, containssubstantially the same information as a traditional cost volume but this information is providedin a lower dimensional representation that can still be passed to any subsequent, downstreamcomponents of the AI-compression pipeline that relies on cost volumes in the usual way.Examples of downstream components may include layers in the flow modules, the finalestimation of flow, the warping in or after the flow module, and so on. Given that thesedownstream components are agnostic as to how they receive the cost volume information asend to end training allows them to adapt to whatever form the information is received in, thecompressively encoded cost volume facilitates the estimation of image differences far moreefficiently than traditional cost volume calculation methods.This approach is described hereinafter as Regional Kernel Absolute Deviation (RKADe) andallows for any MAD operations in the AI-compression pipeline to be substituted by RKADeoperations which mitigate the memory interleaving issue described above in connection withusing MAD and naive patch-wise comparisons for cost volume estimation, and generallyprovide a way to more efficiently estimate differences between images. The inventors havefound that run-time speed ups with very little additional optimisation were observed whenRKADe was tested across a number of widely available commercial CPUs, such as the MacM2 chip, and the Qualcomm Snapdragon chip. For example, in some instances, RKADeresults in a greater than 50% runtime reduction on widely-used standard Intel processors, ascompared to a purpose-built custom MAD cost volume implementation.At a general level, a realisation of the present inventors is that, where a mapping (e.g. aninput, an operation applied to that input, and an output) is sparse in a basis (e.g. in the caseof a representation of an image such as a flow or residual between two images representedby mostly zeros), the mapping can be almost identically represented in fewer dimensions.This facilitates the implementation of that mapping in a simplified manner. In the case offlow estimation in image sequences, there is often very little movement so cost volumes aretypically sparse in a transform of the spatial domain. This means that the calculation of costvolumes representing the difference between two images can be substantially simplified. Forexample, if a single pixel in one image is being compared to a pixel patch in a pixel radius of 3around in a second image, this would entail 49 pixel comparison operations to obtain the costvolume associated with that pixel using traditional methods. It turns out that the vast majorityof these pixel operations are related to redundant information when the input and / or outputare sparse as there is little or no difference. In such circumstances, the cost volume can moreefficiently be estimated by applying a fixed or learned map to the input to produce an identicalor almost identical cost volume. Applying a map, e.g. a linear map through a number ofconvolution operations, is computationally efficient and fast and provides a lower-dimensionalrepresentation of otherwise the same information that would have been provided in a costvolume estimated using traditional methods.Further, the compressive cost volumes that are computed using the RKADe approach comewith an additional, substantial run-time saving, because any downstream tensor operations(TOPs) that take place are in the lower dimensions of the compressive cost volume compared tothe higher dimensions of the traditional cost volume; and none of the steps comprising RKADerequire grouped convolutions, slow memory access or array operations like de-interleaving.An examplary pseudocode implementation of RKADe is provided below:1) Receive input tensor x_t representing first image, and input tensor x_ref representingsecond image.2) Generate a feature map of x_t:feature_map_x_t = feature_map(x_t)3) Generate a feature map of x_ref:feature_map_x_ref = feature_map(x_ref)4) Generate a region map of x_ref from the feature map of x_ref:region_map_x_ref = region_map(feature_map_x_ref)5) Generate transform maps of x_t and region_map_x_ref to ensure same shape to facilitatedirect comparison:transform_map_x_t = transform_map(x_t)transform_map_region = transform_map(region_map_x_ref)6) Calculate absolute difference:RKADe_output = absolute_difference(transform_map_x_t, transform_map_region).Figure 13 illustratively shows the above steps of RKADe. As shown in Figure 13, an illustrativeRKADe workflow on a toy example comprises four elements: a feature map F, a region map R,a transform map T, and an absolute difference Δ.The feature map F, defined by one or more layers comprising one or more filters definedby a plurality of weights, operates as a linear embedding of the input tensor into a featurespace. It is a local embedding in the sense that it may comprise a 1x1 convolution. Efficientimplementations of 1x1 convolutions on a wide variety of commercial CPUs, GPUs and NPUsare known to the skilled person. However, unique in the case of RKADe is that the featuremap operation filters may comprise random weights, and does not need to be trained (althoughit is envisaged that it may be trained in some circumstances). The feature map operationweights may be randomly distributed weights, and the operation relies on the favourableproperties of high-dimensional random embeddings to preserve local geometry. Accordingly,the feature map F operation filters are instantiated using random weights with a normalisationthat makes it an isometry in expectation. That is, the feature map operation filters compriseweights that apply a transformation that, on average, preserves geometrical distances of thedistribution it is applied to. In other words, the feature map operation preserves a norm ofthe inputs in expectation. Thus, if a convolution defining a feature map is an isometry inexpectation, the norm of the input to which the convolution is applied is preserved in the output(in a probabilistic sense). These fixed, random weights of F and the isometry in expectationproperty of F effectively mean that, during inference, estimating differences between twoimages comprises applying a convolution with the random weights that correspond to therandom weights the convolution was initialised with, rather than weights modified in someway during subsequent training. An example of a suitable random distribution of weights of Fis any suitably initialised sub-gaussian distribution. Here suitably initialised means initialisedbased on the shape of the input and / or output tensors.The region map R operation comprises a composition of 3x3 convolutions, optionally withintermediate non-linearities such as one or more ReLU maps or activations. The purpose of theregion map R is to entangle, in an output pixel, the information present in a local patch aboutthe same pixel in the input (reference) image. Here entangling information means combineinformation. The "radius" of the local patch is determined by the number of convolutions in R(e.g., 3 convolutions gives a patch radius of 7). Because R is comprised of 3x3 convolutionsand optionally simple non-linearities, it is possible to use known, efficient 3x3 convolutionimplementations to run efficiently on a wide variety of commercial CPUs, GPUs and NPUs.As above with the 1x1 convolutions, R is instantiated using a weight normalisation that makesit an isometry in expectation.The transform map T operation serves as a post-embedding or post-processing of the featureembedding and the entangled patch information, permitting the effective comparison of thetwo. This transform map T operation can be a simple 1x1 convolution, thereby being local,linear, fast, and efficient to implement across a wide variety of commercially available CPUsGPUs and NPUs. As above, T may be instantiated using a weight normalisation that makesit an isometry in expectation. In some implementations, the weights of the transform mapT may correspond to those of the feature map F. Further, the number of input channels forF can naïvely be set to any positive integer. The number of output channels of F may matchthat of R and hence the number of input and output channels of R must be the same. Thenumber of input channels of T do not need to match its number of output channels. Notefor completeness that the number of input channels of F can be set naïvely because if themathematical relationships of RKADe are to hold in practice, it is envisaged that there issufficient dimensional relationship between the pixel radius (encoded by the number of layersof R) and the number of output channels of F. This ensures the shape of the objects beingcompared match when the absolute difference is subsequently calculated.As shown in Figure 13, the feature map F is applied to a first input tensor representation1300a of a first image, for example a current frame ^^^^ of a sequence of images, and a secondinput tensor representation 1300b of a second image for example a previous frame ^^ref of thesequence which may contain movement relative to the first image. Each input pixel of the firstinput tensor representation 1300a, a comparison will be made with pixels at coordinates arounda predetermined radius of that coordinate in the second input tensor 1300b. An illustrativetoy example of a 1 pixel radius around a center pixel coordinate is indicated with the dottedborders in Figure 7. The feature map convolution operation F is applied to the pixels of thefirst input tensor 1300a and the associated patches of the predetermined radius in the secondinput tensor 1300b. The region map R convolution operation is then applied to the output ofthe feature map convolution operation on the second input tensor 1300b. A transform map Tconvolution operation is then optionally applied to the intermediate outputs, before an absolutedifference is estimated, resulting in the RKADe cost volume tensor. It will be appreciated fromFigure 13 that none of the feature map F convolution, region map R convolution or transformmap T operation require grouped convolutions. Accordingly, the intermediate outputs maybe easily stored in contiguous memory blocks without requiring a large number of memoryread and write operations to interleave or de-interleave the data in memory. As a result, costvolume estimations in the RKADe approach are substantially sped up compared to traditional,naive patch-wise comparison approaches.It is also noted that F, R and / or T may be kept fixed (e.g. F’s weights may be fixed andrandomly distributed between minimum and maximum values, as described above), or maybe trained. Keeping the maps fixed facilitates straightforward deployment by substitutingany naive patch-wise comparisons in AI-compression pipelines, (e.g. a large number ofMAD-based operations). However, when fixed, the expressivity of the maps may be reducedand thus overall accuracy of this approach may be reduced. Conversely, training F, R and / or Tincreases expressivity of RKADe but introduces greater complexity in training and inferencepipelines and can adversely affect training stability of an end-to-end trained AI compressionpipeline by virtue of the introduction of a further trainable element. The choice of using fixedor trained F, R and / or T may accordingly depend on a complexity-accuracy-run-time trade offfor a given application. In exemplary embodiments, it is envisaged that the weights of R and Tare trained or learned, whereas the weights of F are random and fixed.Combining these components together to get the absolute difference Δ that defines RKADe,let ^^, ^^ ∈ R^^in×^^×^^ and let F and T be convolutions with kernel sizes [^^out, ^^in, 1, 1]and comprising R have kernels with size[^^out, ^^out, 3, 3]. Mathematically, one may write RKADe asRKADe(^^, ^^) := |T (F^^ − R(F^^)) | .Above, note that F is a linear map operating pixel-wise (i.e. performing the same operation oneach ^^in-dimensional pixel); that R also acts pixel-wise, but insofar as each pixel is representedby a patch of a given radius that has been induced by the number of convolutions used toconstruct R; and that T also acts linearly and pixel-wise on each ^^out-dimensional pixel. Byvirtue of its construction using elementary convolutional components, RKADe runs on hardwarewithout the requirement of grouped convolutions, and thus solves the slow memory accessor array operations like de-interleaving of cost volume approaches such as naive patch-wiseMAD.As alluded to above, F, R and / or T may be individually or separately trainable, for exam-ple by setting a "requires_grad_" PyTorch or similar flag in their respective code-levelimplementations to permit backpropagation through them. However, it is envisaged that Fis generally kept fixed while R and / or T are trainable. In this case, they may be trained inan end-to-end manner with the rest of the pipeline, whereby the weight values are simplyadditional parameters, on top of the other parameters of the pipeline, that may be updatedduring back-propagation. More specifically, it has been found that end-to-end training requiresno special auxiliary loss terms to guarantee stability during training. Indeed, advantageously,the F, R and T maps are "plug-and-play" onto the rest of the AI compression pipeline duringtraining and subsequent inference.Optionally, student-teacher training of weights of the (optionally F), R and T maps is alsoeffective and achieves good training stability out of the box without difficulty. Multipleapproaches to student-teacher training are envisaged. For example, the teacher may be set upto push the teaching towards a fine-grained level that represents similar features to classicalcost volumes, or at a less granular level where the teacher comprises a flow network with MADcost volumes, and the student comprises a flow network with RKADe cost volumes, with theloss based on the difference between these two. Alternatively, the teacher may be set up topush the training towards some other objective and may be set up accordingly.More generally, if there are multiple training stages of the compression pipeline (e.g. pre-training and main training), the weights of F, R and / or T may be frozen at different timesduring these stages by setting the "requires_grad_" flag appropriately at different trainingsteps.By way of illustrative example, consider a scenario where F is fixed, and R and T are learned,R and T may be frozen with initialisation weights during the initial pre-training phase of thecompression pipeline before being unfrozen and trained during the main training phase. Thisapproach ensures that the weights of R and T are being updated based on a stronger trainingsignal from the rest of the neural networks of the compression pipeline in order to decreaseoverall convergence time, thereby speeding up training.As described above, the distribution (i.e. values) of the initialisation weights of F, R and / orT may be random, based on some predetermined distribution, or based on prior knowledgeobtained from experimentation to provide a warm-started signal from the rest of the model atthe point when one or more of F, R and / or T unfreeze and become trainable. In one illustrativeexample, the initialisation weights may be initialised using any appropriately normalisedsub-gaussian distribution producing a map that is an isometry in expectation (for example, aGaussian distribution, a truncated Weibull distribution, or a uniform distribution). In someembodiments, it is envisaged that, during training, the property of being an isometry inexpectation is maintained as the weights of F, R and / or T as applicable are adjusted. Thisproperty may be enforced using Jacobian regularisation, such as that described in concept 5below. Alternatively, the isometry preserving property may only be retained upon initialisation,and training will gradually eliminate that property as the weights converge to some final values.Finally, it is also envisaged that optionally some of the weights of F, R, and / or T may be keptfixed, while others may be trained, for example to avoid significant departure from the isometrypreserving property during training.In an illustrative toy example, the weights of F, R and / or T may be initialised by generatingrandomly distributed values uniformly distributed between a minimum value and a maximumvalue, whereby the minimum and maximum value are based on a number of input or outputchannels on which F, R and / or T are applied, based on a kernel size of F, R and / or T and / or ona pixel radius across which the RKADe cost volume is to be calculated (e.g. a 3 pixel radiusresulting in a comparison against a 7x7 pixel patch centered on a given pixel coordinate). Inother words, the minimum and maximum values are based on the dimensions of the inputtensor. For example, an initialisation weight distribution can be defined as:weight_distribution = uniform_noise(−^^, ^^),where −^^ is the minimum value and ^^ > 0 is the maximum value of the uniform randomweight distribution defined by:√^ ^^ =input_channels output_channels · kernel_dimension_one · kernel_dimension_two .A filter comprising weights distributed as above preserves norm in expectation of the input towhich it is applied. As described above, optionally, Jacobian regularisation may be appliedduring training to ensure the isometry in expectation is preserved even as the weights areupdated. Alernatively, the weights may be free to lose this property if the training results inthem doing so. Of course, if any of F, R, and / or T are fixed, then the property will be preservednaturally as the initalised weights do not change. It is particularly envisaged that F may befixed in this way, while the weights of R and / or T may be trained.Optionally, the training regime may be implemented with one or more switches that specifyexactly when during training to freeze or unfreeze the trainable parameters of F, R and / or Tbased on some predetermined one or more conditions (such as a number of iterations or a lossthreshold, or other conditions).Where one or more of the F, R and / or T maps are fixed and, for example, there are nonon-linearities such as ReLUs in any of F, R and / or T, the weights are initialised as describedabove e.g. with random weights for R and / or experimentally determined weights for F and T,and there are no special loss or training considerations that need to be taken into account asthe maps are simply linear mappings.Whilst concept 3 has been described in the context of using spatial comparisons to estimateflow between two images in an image sequence as part of a flow-residual-based AI compressionpipeline, it is envisaged that it may be used anywhere where spatial comparisons are calculated,both in and outside of the neural image and video compression domains, as is described inmore detail below in concept 4.Concept 4: Regional Kernel Absolute Deviation (RKADe) for General Computer Visionand Image Processing TasksSpatially comparing two images is widely performed across various technical domains.Specifically, naive patch-wise comparisons (using any measure of similarity) are slow ontypical commercial hardware. For example, if naive patch-wise comparisons using MADoperations are used and implemented as convolutions on existing commercial CPUs, GPUsand NPUs, the patch-wise comparisons introduce a bottleneck to run time. This is particularlyproblematic for technical domains where low-latency and fast run times are critical tofunctionality. Accordingly run-time advantages are realised for any use case that is presentlyimplemented with patch-wise comparisons (using any measure of similarity), for example anyuse cases currently implemented as a MAD-based patch-wise comparison.In contrast, RKADe is faster than naive patch-wise comparison operations, in any computervision and / or image processing task because it is a compressive approximation of such apatch-wise comparison.A first example use case where run-time improvements may be realised with RKADe is in thegeneration of bounding boxes around image patches where movement is to be detected.Consider a first image at a first time and a second image temporally separated from the firstimage. The objects in the second image have moved relative to their positions in the firstimage. In computer vision tasks such as surveillance, satellite image comparisons, dronenavigation, and others, detecting such movement is a common task as it facilitates tracking ofobjects across different views and across time. One approach to such detection is to generate abounding box around objects whose pixels differ between the first and second images.One approach to generating such bounding boxes is to divide the first and second images intogrids, and to compare the pixel values in the first image to individual pixels or groups ofpixels in the second image. One way to make this comparison is to use MAD implemented asconvolutions. If the MAD for a given pixel or patch of pixels exceeds a threshold, that pixel orpatch of pixels may be identified as a movement-containing patch (whereas any that don’t exceedthe threshold may be identified as static patches). The boundaries of one or more boundingboxes may then be generated that encompass some or all of these movement-containing patchesand used to identify the moving object across frames.As described above, a naive patch-wise comparison implemented as convolutions are run-timebottlenecks. Accordingly, applying RKADE approach from concept 3 to the bounding boxtask: RKADe(^^, ^^) := |T (F^^ − R(F^^)) | .where x is the first image, y is the second image, and T, F and R are as described in connectionwith concept 3.This approach to detecting moving objects facilitates the running of object detection locallyand in real-time on end devices such as onboard of a camera-based surveillance system, adrone such as a small quadcopter, and other hardware platforms that typically do not haveaccess to powerful processor resources, or where power draw is at a premium due to hardwareconstraints such as battery life.More generally, the bounding box generation approach may also be used in image and videocompression pipelines including both traditional and AI-based compression pipelines. In thiscase, it may be used to facilitate partial frame-skipping to further reduce the amount of datathat needs to be sent to reconstruct an image or image sequence.For example, if objects in two images have hardly changed except for a small number ofpixels or pixel patches, the above-described bounding box generation approach may be used toidentify and extract those movement-containing patches. Only these movement-containingpatches then need to be compressed and transmitted in order for the full image sequence tobe accurately reconstructed. That is, on the decode side, the original image sequence can beconstructed efficiently by stitching together previously decoded static image patches with thenewly received movement-containing patches.As described above, applying RKADe to estimating differences between pixel patches insteadof a naive patch-wise comparison facilitates a substantial run-time improvement.Other non-limiting use cases where applying RKADe instead of naive patch-wise comparisonresults in run time improvement include:Image Registration: In medical imaging or remote sensing, images taken at different times orfrom different sensors need to be aligned or "registered" to each other. Here the alignment orregistering of images to each other with naive patch-wise comparison (for example where MADis used as a similarity metric to align these images accurately by finding the transformationthat minimizes the average absolute intensity differences between them) can be replaced by thepresent RKADe approach for run-time improvements.Stereo Vision and Depth Estimation: When calculating depth from stereo images, naivepatch-wise comparisons using MAD are typically used to compare corresponding patches inthe left and right images. The disparity (difference in horizontal position) that minimizes theMAD is often chosen as the correct match, which is then used to compute depth information.Here, the replacement of the naive, MAD-based patch-wise comparison results in run-timeimprovements in calculating depth.Template Matching: In object detection and computer vision, template matching involvessliding a template image over a target image to find the region that best matches the template.A naive, patch-wise comparison, MAD-based is typically used as a measure to find the locationwhere the template and the target image have the least absolute difference, indicating a potentialmatch. Again, replacement of this approach with RKADe results in faster matching times.Noise Reduction: In image denoising, MAD can be used to compare the local neighborhoodof pixels. Filters like the median filter or adaptive filters use MAD to determine the level ofnoise in a local patch and to adjust the filtering strength accordingly to reduce noise whilepreserving details. This initial determination of noise levels in a local patch can be achievedfaster by applying the RKADe approach.Quality Assessment: For quality control in manufacturing, naive patch-wise comparisons arebe used to compare images of a product against a standard reference image. Differences beyonda certain threshold can indicate defects or deviations from the desired quality. Again, this mayinstead be implemented with RKADe to provide run time speed ups and the facilitation ofrunning quality control algorithms on edge devices that do not have significant computingpower.Change Detection: In satellite imagery analysis, naive patch-wise comparison can be used todetect changes over time by comparing pixel intensities of the same location across differentdates. This is useful in monitoring urban development, deforestation, or the effects of naturaldisasters. The use of RKADe instead of a naive patch-wise comparison faciliates the runningof such change detection algorithms more quickly and thus allowing real-time change detectionon resource-constrained devices.Photogrammetry: In reconstructing 3D models from 2D images, naive patch-wise comparisonscan be used to ensure that the matching of pixels across multiple images is accurate, which iscrucial for generating a reliable 3D representation. Again, the use of RKADe to replace thenaive patch-wise comparisons results in faster run times, allowing phogrammetry systems toreconstruct 3D models in real time on smaller, resource constrained devices.In each of these cases, the granular detail that naive patch-wise comparisons provide aboutthe differences between pixel intensities can be obtained more efficiently, faster, and onresource-constrained devices by using an RKADe-based approach instead.Concept 5: Jacobian penalty for temporal sequence modellingNeural network training stability refers to how well convergence of a neural network’s learningprocess progresses during training. Training instability manifests as large fluctuations inlearning performance, where, during training, the model’s loss and / or validation curvessignificantly vary, or even fail to improve at all, despite using the same data and trainingparameters.Some non-limiting factors that influence training stability include the choice of the optimizationalgorithm (e.g. stochastic gradient descent, momentum, adagrad, adam, and so on), thelearning rate, the initialization of network weights, the network architecture, the quality andpre-processing of the input data and so on.Traditionally, to improve training stability, techniques such as batch normalization, gradientclipping, regularisation, and hyper-parameter tuning are applied. Finding an optimum approachusing these techniques requires burdensome experimentation, hyper-parameter sweeps andablation studies because, in many cases, an optimum approach to achieving training stabilityfor one type of model, architecture, and data set may only work on that model, architectureand data set.A known regularisation technique is to introduce a Jacobian regularisation or penalty term tothe loss function. A Jacobian regularisation term or penalty in the context of training neuralnetworks is a method used to control or influence the behavior of the model by regularizingits sensitivity to input changes. The Jacobian matrix represents the partial derivatives of themodel’s outputs with respect to its inputs, effectively capturing how changes in the input affectchanges in the output. In matrix form, these partial derivatives are the network’s Jacobianmatrix. In order to control the behaviour of the loss function, a norm of the Jacobian matrix iscalculated and added to the loss function. Thus if the input causes the network output to changesignificantly, the partial derivatives will have high values and the norm of the Jacobian matrixwill be high, thus the penalty added to the overall loss will be high. As training progressesand the network learns to minimise the overall loss, it learns weights which avoid producing aJacobian matrix which, when the norm is calculated, produces a high value.In machine learning generally, temporal sequence modelling is a challenging problem. Inthe context of AI-based image and video compression, this challenge manifests itself in thecompression of image or frame sequences of a video. In low-motion sequences where there issignificant information redundancy across frames, the networks of an AI-based compressionpipeline such as that of Figure 3 and Figure 14, would learn to produce latent representationsof inputs that contain only minimal information and are highly compressible and then re-useinformation form previously decoded frames when reconstructing the input. Conversely, inhigh-motion sequences the latent representations contain more information. It turns out thatteaching these behaviours is extremely challenging.For example, the networks of a pipeline that performs well on high-motion frame sequenceswill not perform well on low-motion frame sequences and vice versa. In general terms, thiscan be understood as the networks not having learned that when two images ^^^^−1 and ^^^^ of asequence are the same or similar, they can produce a substantially empty latent representationto encode into the bit stream when compressing ^^^^ because all the information from ^^^^−1 can bere-used. Indeed it turns out that flow-residual architectures such as that shown in Figure 3 andFigure 14 struggle to perform well on low-motion frame sequences and this can be attributedto the challenge of modelling the temporal sequence of frames.Present concept 5 is directed to solving this problem by introducing a special type of Jacobianpenalty term into the loss function.In more detail, let ^^ ≥ 1 be an integer. Denote by S^^−1 the (^^ − 1)-sphere. Denote by [^^] theordered set {1, 2, .. . , ^^}.Let ^^ := ( ^^ (1) , .. . , ^^ (^^)) : R^^ → R^^ be an almost-everywhere continuously differentiable X ⊆ R^^ is compact. The directional derivative of anycomponent function ^^ (^^) in the (unit) direction ^^ ∈ S^^−1 at a point ^^ ∈ R^^ is given by^^ ^^ (^^) (^^) (^^ + ^^^^) − ^^ (^^)^^ [ ^^ (^^)] (^^) := lim^^↘0 ^^When ^^ is a standard basis vector ^^( ^^) we call ^^^^ [ ^^ (^^)] (^^) the partial derivative of ^^ (^^) (in thedirection ^^( ^^)) (at the point ^^) ^^ ^ (^^)^^ [ ^^ (^^) (^^) ( ^^) ^^^ ] (^^) = ^∇ ^^ (^^), ^^ ^ =^^^^.The Jacobian matrix of ^^ whose (^^, ^^)-th element is^^ ^^ ^^) := ^^ (^^)^^ [ ^^ ] ( ^^, (^^, ^^) ∈ [^^] × [^^] .^^^^ Next, define the fixed point set of a mapping ^^ restricted to a (sub-)space X byFix ^^ := {^^ ∈ X : ^^ (^^) = ^^}.Define a (discrete, Markovian) temporal sequence model as a sequence of tuples (^^^^ , ^^^^ , ^^^^)^^where ^^ ≥ 0 is an integer and where ^^^^ := ^^^^ (^^^^ , ^^^^−1, ^^^^−1). We accommodate the case ^^ = 0 as^^0 := ^^0(^^0). Applying this approach to image and video compression, ^^^^ may be a frame attime ^^, ^^^^−1 may be a frame at time ^^ − 1, ^^^^−1 may be a previously reconstructed frame fromtime ^^ − 1, and ^^ may be the function corresponding to the neural networks of the AI-basedcompression pipeline that we are training. More specifically, the special ^^ = 0 case maycorrespond to encoding / decoding an I-frame with no dependency on a frame at a differenttime, whereas all other times correspond encoding / decoding a P-frame or B-frame from time ^^that have some dependency on a frame from a different time ^^ − 1.Now, assume that ^^1 = ^^ ^^ for all ^^ ≥ 2, and define I := ^^0 and P := ^^1. Take for example thetemporal sequence(^^0, ^^^^ , ^^^^)^^ = (^^0,I, ^^0), (^^0, P, ^^1), (^^0, P, ^^2), .. . ,Observe that if P were a function of ^^^^−1 only, then P would (favourably) exhibit perfectrecovery if we had ^^^^−1 = ^^0 and it held thatP(^^^^−1) = P(^^0) = ^^^^ = ^^0.In other words, if P were a function of ^^^^−1 only, then we would desire that the ground truthdata ^^0 be a fixed point of P. Applying this specifically to image and video compression, if Pwere a function of ^^^^−1 only (i.e. there is no new information in the current frame ^^^^ comparedto the previously decoded frame ^^^^−1), then we would want the networks of the pipeline toperfectly reproduce the previously decoded frame ^^^^−1 as its output when encoding / decodinginput ^^^^ .Our wish accordingly is to invoke known results on fixed point existence theorems that rely onsmoothness characterisations of the relevant mapping. However, in the present context themap P is generally non-smooth and is possibly a composite representation of mappings, someof which have discrete domain, destroying applicability of such smoothness characterisations.Concept 5 is directed to overcoming this problem by introducing the assumption that P is acomposition of two maps: D ◦E such that D is almost-everywhere continuously differentiableand whose domain contains a compact Borel set of positive measure. One could think ofE as an “encoder” that maps a signal to its compressed or latent representation; and D as a“decoder” that maps a compressed or latent representation to a corresponding signal in theoutput domain.With this in mind, define the auxiliary mapQ : ( ^̂^^^ , ^^^^−1) ↦→ ( ^̂^^^+1, ^^^^) := (E( ^̂^^^ , ^^^^), D( ^̂^^^+1, ^^^^−1)).In the setting of the exemplary temporal sequence, one has (^^^^ , ^^^^−1, ^^^^−1) = (^^0, ^^0, ^^^^−1) ∈FixQ if ^^^^−1 = ^^^^ . Morevoer, ^^ exhibits favourable recovery if ^^^^−1 = ^^^^ ≈ ^^0. In other words,if we apply the defined auxiliary map to a given input, the following behaviour is observed:if a current frame ^^^^ has no motion or new information relative to a previous frame ^^^^−1, thereconstructed current frame ^^^^ correctly corresponds exactly to the reconstructed previousframe ^^^^−1, satisfying our requirement that the networks should just re-use previous informationwhere there is no motion between frames to perfectly reconstruct the previous information.Thus we desire a method of promoting that Q admit such a fixed point and in particular to have(^^, ^^, ^^) ∈ FixQ for all ^^ ∼ D where D is some data distribution of interest that is supportedon X.Under suitable conditions it is guaranteed that if the Lipschitz constant of a function issufficiently small, that function admits a fixed point on a given domain (e.g., the Banach fixedpoint theorem and Browder fixed point theorem are both admitted by settings relevant to thecurrent continuous-space framing of the question under present consideration). We therebymake the following Ansatz: an appropriately regularised and sufficiently over-parametrisedfunction P^^ := E^^ ◦ D^^ with weights ^^ admits a configuration ^^0, auxiliary mappingQ := and sufficiently small parameter ^^ ≥ 0 (that may depend on parameter spacecomplexity or other complexity factors) such that LipQ ≤ 1 and, therebyE ^^∼D^^ ((^^, ^^, ^^), ^^^^^^Q) < ^^,where dist(^^, A) is the distance (e.g., Euclidean distance) of the point ^^ to the set A. Observethat ^^ encodes how well the “discovered” fixed point manifold Fix Q is adapted to the (unknownor incompletely characterised) data distribution D. In other words, P exists and can be learnedsuch that its auxiliary mapping has a small Lipschitz constant and the fixed point set of theauxiliary mapping, induced thereby, significantly coincides with the points from the datadistribution of interest. Thus, the training objective becomes the task of learning weightsthat produce a fixed point set for mappings that act on low-motion input frames, effectivelyproducing a network that acts like an identity operator on a previously decoded frame whenthe current frame is substantially the same as or similar to the previously decoded frame.Applying all of the above to temporal sequence modelling, such as that which occurs in (forexample) AI-based image or video compression pipelines, our subsequent Ansatz is that thisparameter ^^0 can be approximated with a suitable training regime. That is, the desired fixedpoint manifold can be well approximated by a deep-learning based temporal sequence modelpermitting stable long-term temporal dependence modelling with recurrent implementations.In general terms, there exists a set of weights for which the network acts as an identity operatorwhen two input frames ^^^^ and ^^^^−1 are identical.Recall that an ^^-Lipschitz function ^^ that is differentiable at a point ^^0 satisfies, for a unitdirection vector ^^:^^^^ [ ^^ ] (^^0) ≤ ^^.In particular, for simplicity assume that the domain X is compact and that ^^ is everywherecontinuous and almost everywhere continuously differentiable on X. Then, there exists apoint ^^ ∈ X such that ^^ = sup^^∈S^^−1 ^^^^ [ ^^ ] (^^) = ∥^^ [ ^^ ] (^^)∥ (where ∥ · ∥ denotes the operatornorm). Accordingly, when ^^ is sufficiently well behaved, controlling the operator norm of theJacobian of ^^ controls its Lipschitz constant, in turn permitting the opportunity for obtaining afunction ^^ with a fixed point manifold, and allowing us to train the network of the AI-basedcompression pipeline corresponding to this function ^^ by including a Jacobian penalty in theloss function based on the auxiliary map defined above.That is, a Jacobian penalty calculated from an auxiliary function that is based on one or moreof the components of the pipeline but modified so that (1) the input space matches the outputspace (that is, the input variables to the auxiliary function have the same form, shape anddimensions as the outputs) and (2) the input latent is passed through and returned by thefunction without amendment.To illustrate this at a practical level, consider the residual decoder of the compression pipelineof Figure 14. Our goal is to encourage the residual decoder to learn weights that perfectlyreconstruct a previously decoded frame when the current frame is substantially identical (i.e.no motion between the frames of the sequence). The residual decoder of the actual pipelinereceives as input a latent tensor, an optical flow information tensor (e.g. in the form of awarped current frame), and outputs a reconstructed image. In this case, the input space doesnot match the output space (i.e. the number of variables, the forms, shapes and dimensions ofthe inputs and outputs are different because the output is only the reconstructed frame andnot the latent tensor). Thus, if we calculate a Jacobian penalty based on this (i.e. based onor approximated from a matrix of partial derivatives of how output changes with respect tochanges in the input), and add this Jacobian penalty to the loss function, the weights willnot generally converge to values that produce our goal of a residual decoder that perfectlyreconstruct the previously decoded frame when the current frame is substantially identical.Instead, in this illustrative example, we construct an auxiliary function that: (i) is based onthe residual decoder i.e. a function that operates identically to the residual decoder during aforward pass and accordingly similarly receives as inputs a latent tensor and an optical flowinformation tensor in the form of a warped input image, but (ii) now not only returns thereconstructed image, but also returns as output the original input latent tensor. Thus we havea function where the input space (latent tensor and warped input image) matches the outputspace (latent tensor and reconstructed image). When we subsequently calculate the Jacobianpenalty from this function, the latent tensor will appear as both a variable considered an inputand a variable considered an output when the partial derivatives of the Jacobian matrix arebeing estimated or approximated. In layman’s terms, the effect of this is the moderation ofsignificant changes to the latent tensor as sequences progress. In more formal terms, theabove-described mathematical relationships apply and the weights converge to values duringtraining that, during inference, exhibit the behaviour of perfectly reconstructing a previouslydecoded frame when the current frame is substantially identical.Accordingly, training an AI compression pipeline by using a Jacobian regularisation termbased on an auxiliary function such as this, in which the input space matches the outputspace, encourages the learning of weights in which the networks are better able to reconstructsequences of frames in a temporally consistent way.This illustrative example is given in the pseudocode below:
[0002] Algorithm 1 AI-based compression network training with residual decoder auxiliary mapping Jacobian penaltyInputs: Training dataset X, learning rate ^^, regularization parameter ^^, number of epochs ^^ ,network architecture ^^^^Initialize network parameters ^^for epoch = 1 to ^^ dofor each batch (^^^^−1, ^^^^ ) in X doForward pass to compute predictions ^^^^ = ^^^^ (^^^^−1, ^^^^ ), including ^̂^^^Forward pass through auxiliary function to compute Jacobian penalty ^^aux ( ^̂^^^ , ^^^^−1 ↦→ ^̂^^^ , ^^^^ )Compute loss Loss = ^^ + ^^^^ + ^^auxBackward pass to compute gradients ∇^^LossUpdate parameters with optimizer O: ^^ ← O(^^, ∇^^Loss, ^^)end forOptionally evaluate on validation setend forThat is, a training data set ^^. A learning rate ^^, regularisation parameter ^^, and a number oftraining steps or epochs ^^ is selected. The network architecture of ^^^^ is defined, for exampleas shown in Figure 3 and Figure 14. The network parameters ^^ are randomly initialisedand then the training loop is started. For each batch (^^^^−1, ^^^^) in the training data ^^, theforward pass is computed for the network being trained ^^^^ , as well as the auxiliary function.A Jacobian penalty ^^^^^^^^ ( ^̂^^^ , ^^^^−1 ↦→ ^̂^^^ , ^^^^) is then estimated as described above, and the total^^^^^^^^ is calculated by combining a distortion term ^^, a rate term ^^, and the estimated Jacobianpenalty. The backwards pass is then performed to compute gradients based on the loss, andthe parameters ^^ are optimised using the optimiser, such as stochastic gradient descent SGD,or some other known optimiser. Optionally, a validation loss can be calculated and the trainingloop is repeated until the predetermined number of steps or epochs ^^ have been calculated, orsome other criteria have been reached. The learning rate, batch size, and or number of epochsmay be optimised during training, for example using a learning rate scheduler or some otherhyperparameter optimisation method. More generally, the hyperparameters may be optimisedexperimentally.As described above, the Jacobian penalty ^^aux( ^̂^^^ , ^^^^−1 ↦→ ^̂^^^ , ^^^^) is based on an auxiliarymapping function that includes the latent tensor ^̂^^^ on both sides of the mapping so that theinput space matches the output space, thus resulting in convergence to a set of weights thatexhibit the desired temporally stability for frame sequences.Note that, whilst the Jacobian penalty term is shown to be based on the mapping ( ^̂^^^ , ^^^^−1 ↦→^̂^^^ , ^^^^), it is envisaged that, each of these variables may be processed in some way withoutthe general applicability of the above-described general mathematical properties ofthe input space matching the output space. For example, where the architecture of ^^^^ includesa flow module, and a flow estimation and / or warped previously decoded image is used in theplace of ^^^^−1, then the mapping may be based on this warped image e.g. ^^^^−1,warped, and so on.Finally, it will be appreciated that, whilst the above example has been described in the contextof the residual decoder in that the auxiliary function from which the Jacobian penalty iscalculated is based on the residual decoder, the above-described principles can be extended tobe more generally applicable to any of the networks in an AI-based compression pipeline inwhich encouraging temporal consistency is challenging.For example, the algorithm described above can be extended to introduce a Jacobian penaltyterm based on any input to output mappings where one of the terms is included as both an inputvariable and output variable. This generalisation is illustrated in the pseudocode providedbelow:
[0003] Algorithm 2 AI-based compression network training with general auxiliary mapping Jacobian penaltyInputs: Training dataset X, learning rate ^^, regularization parameter ^^, number of epochs ^^ ,network architecture ^^^^Initialize network parameters ^^for epoch = 1 to ^^ dofor each batch (^^^^−1, ^^^^ ) in X doForward pass to compute predictions ^^^^ = ^^^^ (^^^^−1, ^^^^ ), including ^̂^^^Forward pass through auxiliary function to compute Jacobian penalty ^^aux (^^,^^ ↦→ ^^, ^^)Compute loss Loss = ^^ + ^^1^^ + ^^2^^auxBackward pass to compute gradients ∇^^LossUpdate parameters with optimizer O: ^^ ← O(^^, ∇^^Loss, ^^)end forOptionally evaluate on validation setend forSimilarly to the specific residual decoder example, a training data set ^^ is provided. A learningrate ^^, regularisation parameter ^^, and a number of training steps or epochs ^^ is selected. Thenetwork architecture of ^^^^ is defined, for example as shown in Figure 3 and Figure 14. Thenetwork parameters ^^ are randomly initialised and then the training loop is started. For eachbatch (^^^^−1, ^^^^) in the training data ^^, the forward pass is computed for network being trained^^^^ , as well as the auxiliary function. A Jacobian penalty ^^^^^^^^ (^^,^^ ↦→ ^^, ^^) is then estimatedas described above, and the total ^^^^^^^^ is calculated by combining a distortion term ^^, a rateterm ^^, and the estimated Jacobian penalty. Some or all of the terms may be regularised bysome constant ^^1 and / or ^^2. The backwards pass is then performed to compute gradientsbased on the loss, and the parameters ^^ are optimised using the optimiser, such as stochasticgradient descent SGD, or some other known optimiser. A validation loss then optionally becalculated and the training loop is repeated until the predetermined number of steps or epochs^^ have been calculated, or some other criteria has been reached. The learning rate, batchsize, and or number of epochs may be optimised during training, for example using a learningrate scheduler or some other hyper parameter optimisation method. More generally, the hyperparameters may be optimised experimentally.In this general case, ^^ , ^^ and ^^ can be any variables or set of variables from an AI-basedcompression pipeline including, but not limited to one or more of: a current input image: ^^^^ ,a reference input image: ^^^^−1 or ^^^^+1, a reference previously reconstructed image (warpedor unwarped): ^^^^−1, ^^^^ or ^^^^+1, ^^^^−1,warped, ^^^^,warped or ^^^^+1,warped, a flow: ^^^^ , and / or any latent on, including in using any ofthese variables in any combination in upsampled and downsampled spaces as applicable.For example, for low-motion type frame sequences, such as static nature scenes or CCTV feedswhere it is typically expected that flow will not vary significantly between frames and wherenetworks are failing to converge to a set of weights where the desired behaviour to produceflow information that varies little between frames, we can construct an auxiliary mapping fromwhich to estimate a Jacobian penalty by using, for example:^^^^^^^^ ( ^̂^ ^^^^ , ^^^^−1 ↦→ ^̂^^^ ^^ , ^^^^) That is, the reconstructed flow latent is provided on both sides of the auxiliary mappingfrom which the Jacobian penalty is calculated. When the loss function then incorporatesthis Jacobian penalty term, the network weights converge to a set of values where flow isencouraged to be reconstructed in a way that varies little between across a sequence of frames.Effectively the networks are learning the set of weights that have a fixed point for the desiredauxiliary mapping behaviour.A similar approach can be taken for any and all of the example variables referred to aboveto encourage the networks of the AI-based compression pipeline to learn how to behave in atemporally consistent way for frame sequences.It is also envisaged that multiple different Jacobian penalties based on such mappings may beintroduced into the loss term, for example ^^^^^^^^1 , ^^^^^^^^2 , ..., ^^^^^^^^^^ as applicable to encourage thelearning of temporally consistent behaviour by any number of components of the AI-basedcompression pipeline. Further, any weighted norm of any such Jacobian penalty or penaltiesmay be incorporated into the AI-based compression pipeline. Specifically, a weighted normof the Jacobian multiplies the components of the matrix using weights, computed in somemanner relevant to the task at hand, such that the resulting "weight and norm" operation stillsatisfies the mathematical definition of being a (quasi) norm. Such weights may be computedfor example as a function of the amount of motion information that is present in the image; oraccording to metrics that define the presence of occlusion between two frames. The motioninformation may be captured indirectly in the flow information estimated by the flow moduleof the pipeline, or by some other measure such as, for example, a direct pixel differencecalculation between the pixels of two or more images.Moving onto practical considerations, given the many different variables and mappings fromwhich the Jacobian penalty described above may be estimated, the above-described approachcan increase network training times significantly, for example 30% or more. To help addressthis, it is envisaged that the presently described Jacobian penalties may be estimated using thefollowing approach.It will be understood that the Frobenius norm of a square matrix ^^ ∈ R^^×^^ also controlsits operator norm. Further recall that Hutchison’s trace estimator (where ^^ ∈ R^^ with^^i^^ ∼idN(0, 1)) is given by[ ] [ ∑ℓ^ tr(^^) = ^^E^^⊤^^= E^^⊤ ] 1 ^^^^ ≈ (^̂^( ^^))⊤^^^̂^( ^^)^^ ^^=1 where ^̂^( ^^) , ^^ = 1, .. . , ℓ are ℓ independent realisations of ^^.Now suppose that ^^ := ^^⊤ [ ^^ ] (^^)^^ [ ^^ ] (^^) and observe that the above equation provides anℓ-sample estimate of the Jacobian-vector product of a function ^^ .Thus, in the standard set-up for minimising a loss function using stochastic gradient descent,suppose on iteration ^^ that one has batch ^^ (^^) and a function ^^ (^^; ^^ (^^)) with parameters ^^ (^^) .Set ℓ = 1 to take a 1-sample approximation for the trace estimate, i.e.,tr(^^) = tr(^^⊤^^) = ∥^^∥2 ≈ ^^⊤^^⊤ [ ^^ (^^) (^^) (^^) (^^)^^ (·;^^ )] (^^ )^^ [ ^^ (·; ^^ )] (^^ )^^Now (^^)^^ [ ^^ (·; ^^ ))] (^^ (^^))^^ = ^^ [ ^^ ( (^^) (^^) ^^ (^^ + ℎ (^^) (^^) (^^)(^^ ^^; ^^ ) − ^^ (^^ ; ^^ )^^ ·;^^ )] (^^ ) ≈for step sizeestimation strategy are here omitted for brevity.Applying the above formal explanation to the specific example of calculating Jacobian penaltiesfor an AI-based compression pipeline loss function, recall that the Jacobian penalty may becalculated by estimating the partial derivatives of the network’s outputs with respect to its inputs,which together form a Jacobian matrix. The penalty may be estimated by calculating the normof this matrix (or the trace of a related matrix - as is known in the art). However, calculatingthe norm (or trace, as applicable) of the Jacobian matrix is computationally very expensive.Instead we use an approach based on Hutchison’s trace estimator and finite differences. Inthe specific case of AI-based compression pipelines, it turns out that we can get a good traceestimate by making a 1-sample approximation. Even though the sample size is only 1, thepresent inventors have found that the estimate is nevertheless accurate enough to estimate aJacobian penalty that encourages the learning of weights with the desired behaviours describedabove.Accordingly, the 1-sample approximation using finite differences facilitates the estimation of aJacobian penalty in a way that is significantly faster than traditional methods and facilitatestraining with multiple Jacobian penalties without significantly increasing training times. Thatis, the time attributed to estimating the Jacobian penalty is an insignificant fraction of theoverall training time.This in turn enables training with Jacobian penalties applied to multiple components ofthe pipeline to encourage temporally stable behaviour while keeping overall training timessubstantially the same thereby reducing overall cost per training run.Another practical consideration that is now discussed is how to ensure that the estimatedJacobian penalty(ies) are not over-penalising the loss when training frame sequences whereits introduction is not conducive to learning a suitable set of weights. For example, consideragain the Jacobian penalty described above that is based on an auxiliary function based onthe residual decoder. As described above, the Jacobian penalty in this case encourages thelearning of a set of weights that allow the near-perfect reconstruction of a previously decodedimage when there is little or no movement between frames.However, when this approach is applied to very high-motion frame sequences, the behaviourwe are encouraging with the Jacobian penalty results in a drop in accuracy.To overcome this issue, the present inventors have realised that the Jacobian penalty describedabove can itself be regularised based on a property of the frames of the sequence being trainedon. That is, the Jacobian penalty may be made "motion-aware" by weighting it according to aproperty of the frames of the sequence being trained on such as how much movement there isbetween frames. This movement may be captured indirectly in the Jacobian matrix from whichthe Jacobian penalty is calculated and accordingly the penalty may comprise a weighted norm^^(^^aux) where the mapping ^^ may comprise e.g. Frobenius norm, a spectral norm, or anyother norm. Whereby, for example, the the weighting may scale the Jacobian penalty based onthe combined strength of all the partial derivatives it contains, which will be higher in highmotion frame sequences.Alternatively, the movement may be captured directly and the mapping that encodes theweighting ^^ of the Jacobian may be based on e.g. pixel differences, MSE, or some othermeasure of differences between the frames at time ^^ and some other time ^^ − 1.At a general level, the idea is that if the motion between two frames ^^^^ and ^^^^−1 is large, thenwe want the Jacobian penalty described above to be dampened to a lower value. Conversely,if the motion between two frames, two corresponding pixel regions within a frame, or anyother information between which relative motion may exist, is small, then we don’t want todampen the Jacobian penalty associated with those frames / regions / pixels and so on at all andinstead allow it to have the effects described above to learn the desired fixed point weights forlow-motion frames. This regularisation of the Jacobian penalty term may comprise a weightednorm or some other weighting and is illustrated in the pseudocode below:
[0004] Algorithm 3 AI-based compression network training with motion-aware auxiliary mapping Jacobian penaltyInputs: Training dataset X, learning rate ^^, regularization parameter ^^, number of epochs ^^ ,network architecture ^^^^Initialize network parameters ^^for epoch = 1 to ^^ dofor each batch (^^^^−1, ^^^^ ) in X doForward pass to compute predictions ^^^^ = ^^^^ (^^^^−1, ^^^^ ), including ^̂^^^Forward pass through auxiliary function to compute Jacobian penalty ^^^^^^^^ (^^,^^ ↦→ ^^, ^^)Compute loss Loss = ^^ + ^^1^^ + ^^2^^(^^aux)Backward pass to compute gradients ∇^^LossUpdate parameters with optimizer O: ^^ ← O(^^, ∇^^ ^^^^^^^^, ^^)end forOptionally evaluate on validation setend forSimilarly to the specific residual decoder example, a training data set ^^ is provided. A learningrate ^^, regularisation parameter ^^, and a number of training steps or epochs ^^ is selected.The network architecture of ^^^^ is defined, for example as shown in Figure 3 and Figure 14.The network parameters ^^ are randomly initialised and then the training loop is started. Foreach batch (^^^^−1, ^^^^) in the training data ^^, the forward pass is computed for the network beingtrained ^^^^ , as well as the auxiliary function. A Jacobian penalty ^^^^^^^^ (^^,^^ ↦→ ^^, ^^) is thenestimated as described above, and the total Loss is calculated by combining a distortion term^^, a rate term ^^, and the estimated Jacobian penalty that in this case is weighted by ^^. Asabove, the loss term may be regularised using e.g. ^^1 and / or ^^2 The backwards pass is thenperformed to compute gradients based on the loss, and the parameters ^^ are optimised using theoptimiser, such as stochastic gradient descent SGD, or some other known optimiser. Optionally,a validation loss may be calculated and the training loop is repeated until the predeterminednumber of steps or epochs ^^ have been calculated, or some other criteria have been reached.The learning rate, batch size, and or number of epochs may be optimised during training, forexample using a learning rate scheduler or some other hyper parameter optimisation method.More generally, the hyper parameters may be optimised experimentally.As explained above, the Jacobian penalty here is weighted by a function ^^, where ^^ is definedto act element-wise on the object returned by a Jacobian-vector product between the Jacobianmatrix and some random vector (in practice a tensor) such that the weighted norm induced by^^ satisfies the mathematical definitions of a quasi-norm or pseudo-norm.A further difficulty that arises when training using the above-described Jacobian penalty is itsinteractions with different phases of training and different training schedules. For example,the present description has been framed in the context of training on sequences of frames,the specific number of frames in a given sequence or "group of pictures" (GOP) used duringtraining can vary. Note that a GOP is typically considered to be an I-frame and ^^ P- orB-frames. Consider for example a training schedule that starts training on short GOPs of5-6 frames, and then after a predetermined number of steps switches to training on longerGOPs of 7 or more frames e.g. 8, 9, 10, 20, 30, 40, 50 frames or more. The Jacobian penaltydescribed herein is ideally suited for encouraging the networks to be performant on said longerGOPs but can in some cases act as noise when performing the initial training on the shorterGOPs. This may be, for example, because the temporal consistency is less of a problem forshorter GOPs and so the Jacobian penalty actually weakens the strength of the training signal.More generally, Jacobian regularisation serves as an inductive bias to improve the network’sgeneralisation performance to sequences with GOP sizes that are significantly greater thanthat seen in training. Indeed, if the network is only evaluated on sequences with the sameGOP seen in training, for example short GOP sequences, then temporal stability can be lessproblematic. However, it follows that, for temporal stability arising solely from training ondifferent GOP sizes to be present in the wild, the training data necessary to achieve this wouldhave to contain equal numbers of samples of all GOP sizes distributed equally across all videosequences and so on - something which is burdensome to obtain in real world settings. Thepresently described approach accordingly facilitates obtaining the same effect but withoutneeding as complete a training data set. Indeed, it becomes computationally prohibitive to trainon very large GOP sizes so even if a complete training data set comprising large samples of allGOP sizes encountered in the wild is available, it may not be possible to practically train on allof the data. Accordingly, the present disclosure faclitates the generalisation of performance inthe sense of temporal stability for large GOP sizes — specifically those that are significantlylarger than what’s seen during training, irrespective of what GOP sizes are in the training data.To overcome this problem, it is envisaged that the presently described Jacobian penalty termmay be introduced only at a predetermined time or times (e.g. in terms of number of trainingsteps or consequential to changing training frame sequence length) during a training schedule.Further, the Jacobian penalty term may also be removed at a predetermined time or times forsimilar reason or reasons as may be applicable for a given training schedule. This adaptiveapproach thus enables the advantages of the above-described Jacobian penalty term to berealised in a flexible way to fit in with any existing training schedules.Figure 14 illustrates a further example of a flow-residual compression pipeline, such as that ofFigure 3, whereby the representation of flow information that the residual decoder receivesas input comprises a warped previously decoded image. As this architecture correspondsgenerally to the flow-residual compression pipeline shown in Figure 3 it accordingly usesthe same reference numbers for corresponding features. The flow module 1400 is shown tocomprise a flow encoder 1401 that produces a latent representation of optical flow information^̂^ ^^ ^^ which is decoded by a flow decoder 1402 into a reconstructed flow ^^^^ which in turn isused to warp a previously reconstructed image ^^^^−1 to produce ^^^^−1,^^ which in turn is fedinto the residual decoder 1413 as the representation of optical flow information. Thus theresidual decoder neural network 1413 uses a latent representation ^^^^ of the current frame^^^^ , and the warped previously reconstructed image ^^^^−1 to produce the reconstructed currentframe ^^^^ . Thus, based on this architecture, the Jacobian penalty term described above may beimplemented by constructing an auxiliary function that copies the operations of the residualencoder 1411 and / or residual decoder 1413 of the residual part 1410 including using the sameinputs as the residual decoder 1413, but the auxiliary function also returns as an output the latentrepresentation ^^^^ that was one of its inputs, effectively simply passing the latent representation^^^^ 1412 directly through the function. The Jacobian penalty based on this auxiliary functionwith the latent being both an input and an output has the desired mathematical properties toencourage convergence to a set of weights that produce a network that behaves in a temporallystable manner for sequences of frames.While this specification contains many specific implementation details, these should beconstrued as descriptions of features that may be specific to particular examples of particularinventions. Certain features that are described in this specification in the context of separateexamples can also be implemented in combination in a single example. Conversely, variousfeatures that are described in the context of a single example can also be implemented in 25multiple examples separately or in any suitable sub-combination.Similarly, while operations are depicted in the drawings in a particular order, this should notbe understood as requiring that such operations be performed in the particular order shownor in sequential order, or that all illustrated operations be performed, to achieve desirableresults. In certain circumstances, multitasking and parallel processing may be advantageous.Moreover, the separation of various system modules and components in the examples describedabove should not be understood as requiring such separation in all examples, and it should beunderstood that the described program components and systems can generally be integratedtogether in a single software product or packaged into multiple software products.The subject matter and the functional operations described in this specification can beimplemented in digital electronic circuitry, in tangibly-embodied computer software orfirmware, in computer hardware, including the structures disclosed in this specification andtheir structural equivalents, or in combinations of one or more of them. The subject matterdescribed in this specification can be implemented as one or more computer programs, i.e.,one or more modules of computer program instructions encoded on a tangible non transitoryprogram carrier for execution by, or to control the operation of, data processing apparatus.Alternatively or in addition, the program instructions can be encoded on an artificially generatedpropagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, thatis generated to encode information for transmission to suitable receiver apparatus for executionby a data processing apparatus. The computer storage medium can be a machine-readablestorage device, a machine-readable storage substrate, a random or serial access memory device,or a combination of one or more of them. The computer storage medium is not, however, apropagated signal.The term “data processing apparatus” encompasses all kinds of apparatus, devices, andmachines for processing data, including by way of example a programmable processor, acomputer, or multiple processors or computers. The apparatus can include special purposelogic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specificintegrated circuit). The apparatus can also include, in addition to hardware, code that createsan execution environment for the computer program in question, e.g., code that constitutesprocessor firmware, a protocol stack, a database management system, an operating system, ora combination of one or more of them.A computer program (which may also be referred to or described as a program, software, asoftware application, a module, a software module, a script, or code) can be written in anyform of programming language, including compiled or interpreted languages, or declarative orprocedural languages, and it can be deployed in any form, including as a stand alone program oras a module, component, subroutine, or other unit suitable for use in a computing environment.A computer program may, but need not, correspond to a file in a file system. A program can bestored in a portion of a file that holds other programs or data, e.g., one or more scripts storedin a markup language document, in a single file dedicated to the program in question, or inmultiple coordinated files, e.g., files that store one or more modules, sub programs, or portionsof code. A computer program can be deployed to be executed on one computer or on multiplecomputers that are located at one site or distributed across multiple sites and interconnected bya communication network.The processes and logic flows described in this specification can be performed by one or moreprogrammable computers executing one or more computer programs to perform functionsby operating on input data and generating output. The processes and logic flows can also beperformed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g.,an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).Computers suitable for the execution of a computer program include, by way of example,can be based on general or special purpose microprocessors or both, or any other kind ofcentral processing unit. Generally, a central processing unit will receive instructions and datafrom a read only memory or a random access memory or both. The essential elements ofa computer are a central processing unit for performing or executing instructions and oneor more memory devices for storing instructions and data. Generally, a computer will alsoinclude, or be operatively coupled to receive data from or transfer data to, or both, one or moremass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks.However, a computer need not have such devices. Moreover, a computer can be embedded inanother device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio orvideo player, a VR headset, a game console, a Global Positioning System (GPS) receiver, aserver, a mobile phones, a tablet computer, a notebook computer, a music player, an e-bookreader, a laptop or desktop computer, a PDAs, a smart phone, or other stationary or portabledevices, that includes one or more processors and computer readable media, or a portablestorage device, e.g., a universal serial bus (USB) flash drive, to name just a few.Computer readable media suitable for storing computer program instructions and data includeall forms of non-volatile memory, media and memory devices, including by way of examplesemiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magneticdisks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM andDVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in,special purpose logic circuitry.The subject matter described in this specification can be implemented in a computing systemthat includes a back end component, e.g., as a data server, or that includes a middlewarecomponent, e.g., an application server, or that includes a front end component, e.g., a clientcomputer having a graphical user interface or a Web browser through which a user can interactwith an implementation of the subject matter described in this specification, or any combinationof one or more such back end, middleware, or front end components. The components of thesystem can be interconnected by any form or medium of digital data communication, e.g., acommunication network. Examples of communication networks include a local area network(“LAN”) and a wide area network (“WAN”), e.g., the Internet.The computing system can include clients and servers. A client and server are generally remotefrom each other and typically interact through a communication network. The relationship ofclient and server arises by virtue of computer programs running on the respective computersand having a client-server relationship to each other.While this specification contains many specific implementation details, these should beconstrued as descriptions of features that may be specific to particular examples of particularinventions. Certain features that are described in this specification in the context of separateexamples can also be implemented in combination in a single example. Conversely, variousfeatures that are described in the context of a single example can also be implemented inmultiple examples separately or in any suitable subcombination.Similarly, while operations are depicted in the drawings in a particular order, this should notbe understood as requiring that such operations be performed in the particular order shownor in sequential order, or that all illustrated operations be performed, to achieve desirableresults. In certain circumstances, multitasking and parallel processing may be advantageous.Moreover, the separation of various system modules and components in the examples describedabove should not be understood as requiring such separation in all examples, and it should beunderstood that the described program components and systems can generally be integratedtogether in a single software product or packaged into multiple software products.
Claims
CLAIMS1. A method of training one or more neural networks, the one or more neural networks beingfor use in lossy image or video encoding, transmission and decoding, the method comprisingthe steps of:receiving an input image at a first computer system;downsampling the input image with a downsampler to produce a downsampled inputimage; encoding the downsampled input image using a first neural network to produce a latentrepresentation; decoding the latent representation using a second neural network to produce an outputimage, wherein the output image is an approximation of the input image;upsampling the output image with an upsampler to produce an upsampled output image;evaluating a function based on a difference between one or more of: the output imageand the input image, the output image and the downsampled input image, the upsampled outputimage and the input image, and / or the upsampled output image and the downsampled inputimage; updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps using a first set of input images to produce a first trained neuralnetwork and a second trained neural network.
2. The method of claim 1, wherein the upsampler comprises a third neural network, andwherein the method comprises updating the parameters of the third neural network based onthe evaluated function.
3. The method of claim 2, wherein the downsampler comprises a bilinear or bicubicdownsampler.
4. The method of claim 2 or 3, wherein the downsampler comprises a Gaussian blur filter.
5. The method of any of claims 2 to 4, comprising (i) updating the parameters of the first neuralnetwork and the second neural network based on the evaluated function for a first numberof said steps to produce the first and second trained neural networks without performingsaid upsampling and downsampling and without updating the parameters of the third neuralnetwork, (ii) freezing the parameters of the first and second neural networks after said numberfirst number of steps, and performing said upsampling and downsampling, and said updatingof the parameters of the third neural network for a second number of said steps.
6. The method of claim 1, wherein the downsampler comprises a fourth neural network, andwherein the method comprises updating the parameters of the fourth neural network based onthe evaluated function.
7. The method of claim 5, wherein the upsampler comprises a bilinear or bicubic upsampler.
8. The method of claim 6 or 7, comprising (i) updating the parameters of the first neuralnetwork and the second neural network based on the evaluated function for a first number ofsaid steps to produce the first and second trained neural networks without performing saidupsampling and downsampling and without updating the parameters of the fourth neuralnetwork, (ii) freezing the parameters of the first and second neural networks after said numberfirst number of steps, and performing said downsampling and said updating of the parametersof the fourth neural network for a second number of said steps.
9. The method of any of claims 2 to 8, wherein the method comprises entropy encoding thelatent representation into a bitstream having a length, wherein the function is further based onsaid bitstream length, and wherein said updating the parameters of the third or fourth neuralnetwork is based on the evaluated function based on the bitstream length.
10. The method of any of claims 1 to 9, wherein the difference between one or more of: theoutput image and the input image, the output image and the downsampled input image, theupsampled output image and the input image, and / or the upsampled output image and thedownsampled input image is determined based on the output of a fifth neural network actingas a discriminator.
11. The method of any of claims 1 to 10, wherein the said difference between one or more of:the output image and the input image, the output image and the downsampled input image,the upsampled output image and the input image, and / or the upsampled output image andthe downsampled input image, comprises a mean squared error (MSE) and / or a structuralsimilarity index measure (SSIM).
12. The method of any of claims 1 to 11, wherein the function further comprises a termdefining a visual perceptual metric.
13. The method of claim 12, wherein the term defining a visual perceptual metric comprises aMS-SIM metric.
14. A method of training one or more neural networks, the one or more neural networks beingfor use in lossy image or video encoding, transmission and decoding, the method comprisingthe steps of:receiving an input image at a first computer system;encoding the input image using a first neural network to produce a latent representation;decoding the latent representation using a second neural network to produce an outputimage, wherein the output image is an approximation of the input image;upsampling the output image with an upsampler to produce an upsampled output image,the upsampler comprising a third neural network;evaluating a function based on a difference between one or both of: the output imageand the input image, and / or the upsampled output image and the input image;updating the parameters of the third neural network based on the evaluated function; andrepeating the above steps using a first set of input images to produce a first trained neuralnetwork and a second trained neural network.
15. A method of training one or more neural networks, the one or more neural networks beingfor use in lossy image or video encoding, transmission and decoding, the method comprisingthe steps of:receiving an input image at a first computer system;downsampling the input image with a downsampler to produce a downsampled inputimage, the downsampler comprising a fourth neural network;encoding the downsampled input image using a first neural network to produce a latentrepresentation; decoding the latent representation using a second neural network to produce an outputimage, wherein the output image is an approximation of the input image;evaluating a function based on a difference between one or more of: the output imageand the input image, the output image and the downsampled input image and / or the inputimage and a previously downsampled input image;updating the parameters of the fourth neural network based on the evaluated function;and repeating the above steps using a first set of input images to produce a first trained neuralnetwork and a second trained neural network.
16. The method of claim 15, comprising producing the previously downsampled input imageby performing bilinear or bicubic downsampling on the input image.
17. A method for lossy image or video encoding, transmission and decoding, the methodcomprising the steps of:receiving an input image at a first computer system;downsampling the input image with a downsampler;encoding the downsampled input image using a first trained neural network to produce alatent representation;transmitting the latent representation to a second computer system;decoding the latent representation using a second trained neural network to produce anoutput image, wherein the output image is an approximation of the input image; andupsampling the output image with an upsampler to produce an upsampled output image.
18. A method for lossy image or video encoding, transmission and decoding, the methodcomprising the steps of:receiving an input image at a first computer system;encoding the input input image using a first trained neural network to produce a latentrepresentation; transmitting the latent representation to a second computer system; anddecoding the latent representation using a second trained neural network to produce anoutput image, wherein the output image is an approximation of the input image;wherein one or more of the above steps comprises performing a downsampling orupsampling operation, andwherein the downsampling or upsampling operation comprises performing one orconvolution operations without performing a space-to-depth or depth-to-space operation.
19. The method of claim 18, comprising performing the downsampling or upsamplingoperation on a CPU without performing a space-to-depth or depth-to-space operation, andwherein said downsampling or upsampling is configured to be performed in real-time or nearreal-time.
20. The method of claim 18, comprising performing the downsampling or upsamplingoperation on a neural accelerator without performing a space-to-depth or depth-to-spaceoperation, andwherein said downsampling or upsampling is configured to be performed in real-time or nearreal-time.
21. The method of claim 19 or 20, wherein the downsampling operation comprises applyingone or more convolutional layers with a kernel size based on a downsampling factor, andwherein the convolutional layers are configured to sequentially reduce the spatial dimensions ofan input to the series of convolutional layers while increasing the depth or channel dimensionof the input.
22. The method of claim 21, wherein the input comprises said input image.
23. The method of claim 21, wherein the input comprises a tensor representation of the inputimage.
24. The method of claim 19 or 20, wherein the downsampling operation comprises applyingone or more convolutional layers configured with a stride equal to the downsampling factor,and wherein the number of filters in each convolutional layer is based on the original numberof channels in the input tensor and on the downsampling factor.
25. The method of claim 19 or 20, wherein the upsampling operation comprises applyingsequential convolutional layers and upsampling layers, wherein the convolutional layersare configured to progressively increase the spatial dimensions and decrease the channeldimensions of an input to said layers.
26. The method of claim 25, wherein the input comprises the latent representation.
27. The method of claim 25, wherein the input comprises a tensor representation of the latentrepresentation or the output image.
28. The method of claim 25, wherein the convolutional layers for upsampling are configuredwith a stride based on an upsampling factor.
29. The method of claim 25, further comprising applying an activation function after eachconvolutional layer in the upsampling operation.
30. The method of claim 25, wherein the upsampling layers are selected from a groupconsisting of nearest neighbor upsampling, bilinear upsampling, and bicubic upsampling, andare alternated with the convolutional layers.
31. A method for lossy image or video encoding, transmission and decoding, the methodcomprising the steps of:receiving an input image and a second image at a first computer system;estimating optical flow information using the second image and the input image using afirst neural network, the optical flow information being indicative of a difference between arepresentation of the second image and a representation of the input image;transmitting the optical flow information to a second computer system;decoding the optical flow information using a second neural network;andusing the second image and the decoded optical flow information to produce an outputimage, wherein the output image is an approximation of the input image;wherein estimating the optical flow information comprises estimating differences betweenthe input image and second image by:applying a first convolution operation on respective pixels of a representation of theinput image and / or on respective pixels of a representation of the second image, whereinthe convolution operation comprises applying one or more filters comprising weights havingvalues randomly distributed between a minimum and a maximum value.
32. The method of claim 31, comprising estimating a compressively encoded cost volumeindicative of said differences by applying said first convolution operation.
33. The method of claim 31, wherein the first convolution substantially preserves a norm of adistribution of the respective pixels of the representation of the input image and / or respectivepixels of the representation of the second image.
34. The method of any of claims 31 to 33, wherein a distribution of values of pixels of therepresentation of the input image and / or the distribution of values of pixels of the representationof the second image are sparse distributions in a spatial domain of the representation of theinput image and / or the second image.
35. The method of claim 34, wherein the weights have values distributed according to asub-Gaussian distribution.
36. The method of claim 34 or 35, wherein the minimum value and / or the maximum valueare based on a number of channels of the input image and / or second image, on a kernel sizeof the first convolution operation, and / or on a pixel radius across which said differences areestimated.
37. The method of any of claims 31 to 36, comprising performing a second convolutionoperation on an output of the first convolution operation, wherein the second convolutionoperation substantially preserves a norm of a distribution of said output of the first convolutionoperation.
38. The method of any of claims 31 to 37, comprising estimating a difference between anoutput of the second convolution operation and an output of the first convolution operation.
39. The method of claim 38, wherein said difference comprises an absolute difference.
40. The method of claim 39, wherein said difference defines a cost volume.
41. The method of any of claims 31 to 40, comprising using the optical flow information towarp a representation of the second image.
42. The method of claim 41, comprising estimating a difference between the warped secondimage and the input image to generate a residual representation of the input image with respectto the warped second image.
43. The method of claim 42, comprising: (i) using a third neural network to encode the residualrepresentation of the input image, (ii) transmitting the encoded residual representation of theinput image to the second computer system, (iii) using a fourth neural network to decode theresidual representation of the input image, and (iv) using decoded the residual representationof the input image to produce said output image.
44. The method of any of claims 37 to 43, comprising applying a third convolution operationto an output of the first convolution operation and / or to an output of the second convolutionoperation.
45. The method of any of claims 31 to 44, wherein a kernel size of the second convolutionoperation is greater than a kernel size of the first convolution operation.
46. The method of any of claims 31 to 45, wherein the first convolution operation is defined bya 1x1 kernel.
47. The method of any of claims 31 to 46, wherein the second convolution operation is definedby a 3x3 kernel.
48. The method of any of claims 44 to to 47, wherein the third convolution operation is definedby a 1x1 kernel.
49. The method of any of claims 31 to 48, wherein performing the second convolutionoperation entangles information associated with respective pixels of the representation ofthe input image with information associated with pixels adjacent corresponding pixels in therepresentation of the second image.
50. The method of any of claims 31 to 49, wherein the first, second, and where present thirdconvolution operation are performed without grouped convolutions.
51. The method of claim 50, wherein one or more outputs of the first, second and / or thirdconvolution operation are stored in contiguous memory blocks, and wherein said estimating adifference comprises retrieving said stored outputs from said contiguous memory blocks.
52. The method of any of claims 31 to 51, wherein a distribution of pixel values of the inputimage and of the second image is comprises a sparse distribution, and wherein the sparsedistribution is incoherent in a spatial domain.
53. A data processing system configured to perform the method of any one of claims 31 to 52.
54. A method for lossy image or video encoding and transmission, the method comprising thesteps of:receiving an input image and a second image at a first computer system;estimating optical flow information using the second image and the input image using afirst neural network, the optical flow information being indicative of a difference between arepresentation of the second image and a representation of the input image;transmitting the optical flow information to a second computer system;wherein estimating the optical flow information comprises estimating differences betweenthe input image and second image by:applying a first convolution operation on respective pixels of a representation of theinput image and / or on respective pixels of a representation of the second image, whereinthe convolution operation comprises applying one or more filters comprising weights havingvalues randomly distributed between a minimum and a maximum value.
55. A method for lossy image or video decoding, the method comprising the steps of:receiving an input image and a second image at a first computer system;estimating optical flow information using the second image and the input image using afirst neural network, the optical flow information being indicative of a difference between arepresentation of the second image and a representation of the input image and based on acompressively encoded cost volume;at a second computer system, receiving optical flow information, the optical flowinformation being indicative of a difference between a representation of a second image and arepresentation of an input image;decoding the optical flow information using a second neural network; andusing the second image and the decoded optical flow information to produce an outputimage, wherein the output image is an approximation of the input image.
56. A data processing apparatus configured to perform the method of claims 54 or 55.
57. A method of estimating a difference between a first image and a second image, the methodcomprising: performing a first convolution operation on respective pixels of a representation of thefirst image and on respective pixels of a representation of the second image; andestimating a difference between the first image and the second image based on one ormore outputs of the first convolution operation on the first and second images by:estimating a compressively encoded cost volume indicative of said differences.
58. The method of claim 57, comprising performing a second convolution operation on anoutput of the first convolution operation; and estimating a difference between an output of thesecond convolution operation and the first convolution operation.
59. The method of claim 58, wherein performing the second convolution operation entanglesinformation associated with respective pixels of the representation of the first with informationassociated with pixels adjacent corresponding pixels in the representation of the second image.
60. The method of any of claims 57 to 59, wherein the first convolution operation comprisesapplying one or more filters comprising weights having values randomly distributed between aminimum value and a maximum value.
61. The method of claim 60, wherein the minimum value and / or the maximum value are basedon a number of channels of the input image and / or second image, on a kernel size of the firstconvolution operation, and / or on a pixel radius across which said differences are estimated.
62. The method of any of claims 57 to 61, wherein said difference comprises an absolutedifference.
63. The method of any of claims 57 to 62, wherein said difference defines a cost volume.
64. The method of any of claims 57 to 63, comprising applying a third convolution operationto an output of the first convolution operation and / or to an output of the second convolutionoperation.
65. The method of any of claims 57 to 64, wherein a kernel size of the second convolutionoperation is greater than a kernel size of the first convolution operation.
66. The method of any of claims 57 to 65, wherein the first convolution operation is defined bya 1x1 kernel.
67. The method of any of claims 57 to 66, wherein the second convolution operation is definedby a 3x3 kernel.
68. The method of any of claims 57 to 67, wherein the third convolution operation is definedby a 1x1 kernel.
69. The method of any of claims 57 to 68, comprising storing a plurality of said respectiveoutputs of the first, second and / or third convolution operations in contiguous memory blocks,and wherein said estimating a difference comprises retrieving said stored outputs from saidcontiguous memory blocks.
70. The method of any of claims 57 to 69, comprising using said difference to identify one ormore pixel patches in the second image as movement-containing pixel patches, and generatinga bounding box around one or more of said movement-containing pixel patches.
71. data processing apparatus configured to perform the method of claims 57 to 70.
72. A method of training one or more neural networks, the one or more neural networks beingfor use in lossy image or video encoding, transmission and decoding, the method comprisingthe steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of optical flow information,the optical flow information being indicative of a difference between the first image and thesecond image;with a second neural network, decoding the latent representation of optical flowinformation to produce an approximation of the optical flow information;with a third neural network, producing an output image using the optical flow informationand using a representation of the first image, wherein the output image is an approximation ofthe first image;evaluating a function based on a difference between the first image and the second image,the function comprising a Jacobian penalty term,updating the parameters of the first, second and / or third neural networks based on theevaluated function; andrepeating the above steps using a first set of input images to produce first, second and / orthird trained neural networks.
73. The method of claim 72, wherein the Jacobian penalty term is based on a rate of change ofone or more first variables with respect to one or more second variables, the first variables andsecond variables selected from inputs and / or outputs associated with the one or more neuralnetworks.
74. The method of claim 73, wherein at least the input and / or output associated with the oneor more neural networks is both a first variable and a second variable.
75. The method of any of claims 73 to 74, comprising producing the second variables fromthe first variables by mapping the first variables to the second variables.
76. The method of claim 75 wherein the mapping is defined by an auxiliary function.
77. The method of claim 76, wherein the first variables are inputs to the auxiliary function andthe second variables are outputs of the auxiliary function.
78. The method of claim 77, wherein at least one input of said inputs to the auxiliary functionis also an output of the auxiliary function.
79. The method of claim 78, wherein the inputs of said mapping are defined in an input space,and the outputs of said mapping are defined in an output space, and wherein the auxiliaryfunction maps the input space to the output space.
80. The method of claim 79, wherein the input space matches the output space.
81. The method of any of claims 76 to 80, wherein the auxiliary function is based on the thirdneural network.
82. The method of claim 81, wherein the third neural network comprises a residual decoderneural network.
83. The method of any of claims 78 to 82, wherein the at least one input to the auxiliaryfunction that is an output of the auxiliary function comprises said latent representation of thefirst image.
84. The method of any of claims 72 to 83, comprising weighting the Jacobian penalty term.
85. The method of claim 84, wherein said weighting is based on a difference between the firstimage and the second image.
86. The method of claim 84 when dependent on claim 73, wherein said weighting is definedby a weighted norm based on a matrix associated with said rate of change.
87. The method of any of claims 73 to 86 when dependent on claim 83, comprising estimatingthe Jacobian penalty term by approximating a norm of a matrix associated with said rate ofchange.
88. The method of claim 87, wherein approximating the norm of the matrix comprises makinga single sample approximation.
89. The method of any of claims 72 to 88, wherein the method comprises introducing theJacobian penalty term into said function after a first number of said repeated steps.
90. The method of claim 89, wherein said first number of said repeated steps is based on aGOP-size of one or more frame sequences in said first set of input images.
91. A method of performing lossy image or video encoding, transmission and decoding, themethod comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of optical flow information,the optical flow information being indicative of a difference between the first image and thesecond image;transmitting the latent representation of optical flow information to a second computersystem; with a second neural network, decoding the latent representation of optical flowinformation to produce an approximation of the optical flow information;with a third neural network, producing an output image using the optical flow informationand using a representation of the first image, wherein the output image is an approximation ofthe first image;wherein the first neural network, the second neural network, and the third neural networkare produced according to the method of any of claims 72 to 90.
92. A method of performing lossy image or video encoding, transmission, the methodcomprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of optical flow information,the optical flow information being indicative of a difference between the first image and thesecond image;transmitting the latent representation of optical flow information to a second computersystem; wherein the first neural network,is produced according to the method of any of claims72 to 90.
93. A method of performing lossy image decoding, the method comprising the steps of:receiving a latent representation of optical flow information at a second computer system;,the optical flow information being indicative of a difference between a first image and a secondimage; with a second neural network, decoding the latent representation of optical flowinformation to produce an approximation of the optical flow information;with a third neural network, producing an output image using the optical flow informationand using a representation of the first image, wherein the output image is an approximation ofthe first image;wherein the second neural network and the third neural network are produced accordingto the method of any of claims 72 to 90.
94. A method of performing lossy image or video decoding, the method comprising the stepsof: with a second neural network, at a second computer system, decoding a latent represen-tation to produce a first output image, wherein the first output image is an approximation ofone image of an image pair of a first sequence of input images;repeating the above step to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the second neural network is produced according to the method of any of claims72 to 90.
95. A data processing apparatus configured to perform the method of any of claims 72 to 94.
96. A computer program comprising instructions which, when the program is executed by acomputer, cause the computer to carry out the method of any of claims 72 to 94.
97. A computer-readable storage medium comprising instructions which, when executed by acomputer, cause the computer carry out the method of any of claims 72 to 94.
Citation Information
Patent Citations
Image compression and decoding, video compression and decoding: methods and systems
WO2021220008A1
Method for chroma subsampled formats handling in machine-learning-based picture coding
WO2022106013A1