Method and data processing system for lossy image or video encoding, transmission and decoding

By training neural networks to generate latent representations of image differences and iteratively updating parameters, the method addresses inefficiencies in existing lossy compression techniques, enhancing compression efficiency and visual quality in image and video encoding.

WO2025210218A1PCT designated stage Publication Date: 2025-10-09DEEP RENDER LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/059249
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-16
Filing Date
2025-04-04
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing lossy image and video compression techniques, including AI-based methods, struggle to balance compression efficiency and visual quality, particularly in handling spatial and temporal redundancies, leading to potential errors and inefficiencies in data transmission.

Method used

A method involving training neural networks to produce latent representations of image differences, using masking and decoding processes to approximate original images, with iterative parameter updates based on evaluation functions, enhances compression efficiency and quality.

Benefits of technology

The method improves the accuracy and efficiency of lossy image and video encoding and decoding, reducing data transmission requirements while maintaining acceptable visual quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025059249_09102025_PF_FP_ABST
    Figure EP2025059249_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A method of training one or more neural networks, the one or more neural networks being for use in lossy image or video encoding, transmission and decoding, the method comprising the steps of: receiving a first sequence of input images at a first computer system; producing a masked sequence of input images by masking a portion of one or more of the input images; with a first neural network, for a pair of input images of the masked sequence, producing a latent representation of a difference between the images of the pair; with a second neural network, decoding the latent representation to produce a first output image, wherein the first output image is an approximation of one of the images of the pair; repeating the above steps for a plurality of further sequences of input images to produce first and second trained neural networks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method and data processing system for lossy image or videoencoding, transmission and decodingBACKGROUNDThis invention relates to a method and system for lossy image or video encoding, transmissionand decoding, a method, apparatus, computer program and computer readable storage mediumfor lossy image or video encoding and transmission, and a method, apparatus, computerprogram and computer readable storage medium for lossy image or video receipt and decoding.There is increasing demand from users of communications networks for images and videocontent. Demand is increasing not just for the number of images viewed, and for the playingtime of video; demand is also increasing for higher resolution content. This places increasingdemand on communications networks and increases their energy use because of the largeramount of data being transmitted.To reduce the impact of these issues, image and video content is compressed for transmissionacross the network. The compression of image and video content can be lossless or lossycompression. In lossless compression, the image or video is compressed such that all of theoriginal information in the content can be recovered on decompression. However, when usinglossless compression there is a limit to the reduction in data quantity that can be achieved. Inlossy compression, some information is lost from the image or video during the compressionprocess. Known compression techniques attempt to minimise the apparent loss of informationby the removal of information that results in changes to the decompressed image or video thatis not particularly noticeable to the human visual system. JPEG, JPEG2000, AVC, HEVC andAVI are examples of compression processes for image and / or video files.In general terms, known lossy image compression techniques use the spatial correlationsbetween pixels in images to remove redundant information during compression. For example,in an image of a blue sky, if a given pixel is blue, there is a high likelihood that the neighbouringpixels, and their neighbouring pixels, and so on, are also blue. There is accordingly no need toretain all the raw pixel data. Instead, we can retain only a subset of the pixels which take upfewer bits and infer the pixel values of the other pixels using information derived from spatialcorrelations.A similar approach is applied in known lossy video compression techniques. That is, spatialcorrelations between pixels allow the removal of redundant information during compression.However, in video compression, there is further information redundancy in the form of temporalcorrelations. For example, in a video of an aircraft flying across a blue-sky background, mostof the pixels of the blue sky do not change at all between frames of the video. The mostof the blue sky pixel data for the frame at position t = 0 in the video is identical to that atposition t = 10. Storing this identical, temporally correlated, information is inefficient. Instead,only the blue sky pixel data for a subset of the frames is stored and the rest are inferred frominformation derived from temporal correlations.In the realm of lossy video compression in particular, the removal of redundant temporallycorrelated information in a video sequence is known inter-frame redundancy.One technique using inter-frame redundancy that is widely used in standard video compressionalgorithms involves the categorization of video frames into three types: I-frames, P-frames, andB-frames. Each frame type carries distinct properties concerning their encoding and decodingprocess, playing different roles in achieving high compression ratios while maintainingacceptable visual quality.I-frames, or intra-coded frames, serve as the foundation of the video sequence. These framesare self-contained, each one encoding a complete image without reference to any other frame.In terms of compression, I-frames are least compressed among all frame types, thus carryingthe most data. However, their independence provides several benefits, including being thestarting point for decompression and enabling random access, crucial for functionalities likefast-forwarding or rewinding the video.P-frames, or predictive frames, utilize temporal redundancy in video sequences to achievegreater compression. Instead of encoding an entire image like an I-frame, a P-frame representsthe difference between itself and the closest preceding I- or P-frame. The process, known asmotion compensation, identifies and encodes only the changes that have occurred, therebysignificantly reducing the amount of data transmitted. Nonetheless, P-frames are dependent onprevious frames for decoding. Consequently, any error during the encoding or transmissionprocess may propagate to subsequent frames, impacting the overall video quality.B-frames, or bidirectionally predictive frames, represent the highest level of compression.Unlike P-frames, B-frames use both the preceding and following frames as references in theirencoding process. By predicting motion both forwards and backwards in time, B-framesencode only the differences that cannot be accurately anticipated from the previous and nextframes, leading to substantial data reduction. Although this bidirectional prediction makesB-frames more complex to generate and decode, it does not propagate decoding errors sincethey are not used as references for other frames. Artificial intelligence (AI) based compressiontechniques achieve compression and decompression of images and videos through the use oftrained neural networks in the compression and decompression process. Typically, duringtraining of the neutral networks, the difference between the original image and video and thecompressed and decompressed image and video is analyzed and the parameters of the neuralnetworks are modified to reduce this difference while minimizing the data required to transmitthe content. However, AI based compression methods may achieve poor compression resultsin terms of the appearance of the compressed image or video or the amount of informationrequired to be transmitted.An example of an AI based image compression process comprising a hyper-network is describedin Ballé, Johannes, et al. “Variational image compression with a scale hyperprior.” arXivpreprint arXiv:1802.01436 (2018), which is hereby incorporated by reference.An example of an AI based video compression approach is shown in Agustsson, E., Minnen, D.,Johnston, N., Balle, J., Hwang, S. J., and Toderici, G. (2020), Scale-space flow for end-to-endoptimized video compression. In Proceedings of the IEEE / CVF Conference on ComputerVision and Pattern Recognition (pp. 8503-8512), which is hereby incorporated by reference.A further example of an AI based video compression approach is shown in Mentzer, F.,Agustsson, E., Ballé, J., Minnen, D., Johnston, N., and Toderici, G. (2022, November). Neuralvideo compression using gans for detail synthesis and propagation. In Computer Vision–ECCV2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, PartXXVI (pp. 562-578), which is hereby incorporated by reference.SUMMARYAccording to an aspect of the present disclosure, there is provided a method of training one ormore neural networks, the one or more neural networks being for use in lossy image or videoencoding, transmission and decoding, the method comprising the steps of:receiving a first sequence of input images at a first computer system;producing a masked sequence of input images by masking a portion of one or more ofthe input images;with a first neural network, for a pair of input images of the masked sequence, producinga latent representation of a difference between the images of the pair;with a second neural network, decoding the latent representation to produce a first outputimage, wherein the first output image is an approximation of one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;evaluating a function based on a difference between the first sequence of output imagesand the first sequence of input images;updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps for a plurality of further sequences of input images to producefirst and second trained neural networks.Optionally, the masking comprises selecting a portion of a first image of the first sequence,and replacing pixels in the other images of the first sequence using the selected portion.Optionally, an unselected portion of the first image of the first sequence remains unmasked inthe first sequence of input images.Optionally, a position and size of the portion for first sequence is different to the position andsize of the position for one or more of the the further sequences.Optionally, the position and size of the portions for the first sequence and the further sequencesare distributed according to a predetermined distribution.Optionally, the distribution is a uniform distribution.Optionally, the distribution is a Gaussian distribution with a mean at a center coordinate of theinput images.Optionally, the distribution is based on a type of the input images.Optionally, the type comprises at least one of: webcam video frames, static scene frames,video game frames, high motion video frames, low motion video frames.Optionally, the function is based on a difference between the first sequence of output imagesand the masked sequence of input images.According to an aspect of the present disclosure, there is provided a method of performinglossy image or video encoding, transmission and decoding, the method comprising the stepsof: receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce a first output image, wherein the first output image is an approximationof one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the first neural network and the second neural network are produced accordingto any of the above methods.According to an aspect of the present disclosure, there is provided a method of performinglossy image or video encoding and transmission, the method comprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;wherein the first neural network is produced according to any of the above methods.According to an aspect of the present disclosure, there is provided a method of performinglossy image or video decoding, the method comprising the steps of:with a second neural network, at a second computer system, decoding a latent represen-tation to produce a first output image, wherein the first output image is an approximation ofone image of an image pair of a first sequence of input images;repeating the above step to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the second neural network is produced according to any of the above describedmethods.According to an aspect of the present disclosure, there is provided a data processing apparatusconfigured to perform any of the above methods.According to an aspect of the present disclosure, there is provided a computer programcomprising instructions which, when the program is executed by a computer, cause thecomputer to carry out the method of any of the above methods.According to an aspect of the present disclosure, there is provided a computer-readable storagemedium comprising instructions which, when executed by a computer, cause the computercarry out the method of any of the above methods.According to a further aspect, there is provided a method of training one or more neuralnetworks, the one or more neural networks being for use in lossy image or video encoding,transmission and decoding, the method comprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;with a second neural network, decoding the latent representation to produce a first outputimage, wherein the first output image is an approximation of one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;evaluating a function based on a difference between the first sequence of output imagesand the first sequence of input images;repeating the above steps for one or more further sequences of input images;accumulating gradients associated with said evaluating of the function for the firstsequence and for the one or more further sequences;updating the parameters of the first neural network and the second neural network basedon the accumulated gradients; andrepeating the above steps for a plurality of said sequences of input images to producefirst and second trained neural networks.Optionally, the first sequence of input images and the one or more further sequences of inputimages comprise different images.Optionally, a final hidden state of the first and second neural networks after evaluating thefunction for the first sequence is used as a starting hidden state of the first and second neuralnetworks when repeating the steps for one of the one or more further sequences.Optionally, the first sequence comprises 10 or more input images.Optionally, the first sequence comprises 30 or more input images.Optionally, the first sequence comprises 60 or more input images.Optionally, a number of input images in the first sequence of input images is different to anumber of frames in at least one of the one or more further sequences of input images.Optionally, the number of input images in the one or more further sequences of input images isbased on a difference between the first output image and the one of the images of the pair.Optionally, the number of input images in the one or more further sequences is based on adifference between the first sequence of output images and the first sequence of input images.Optionally, the number of input images in the one or more further sequences is based on anoutput of a classifier configured to detect an artefact in the first output image.Optionally, the first sequence and the further sequences comprise a batch, and repeating saidsteps for a plurality of batches.Optionally, said updating the parameters comprises applying an optimiser to weights of thefirst and second neural networks based on said accumulated gradients.Optionally, the accumulated gradients are associated with a greater number of input imagesthan a number of input images in the first sequence or a number input images in any one of theone or more further sequences.Optionally, the method comprises repeating said steps a first number of times using sequenceshaving a first number of input images, and repeating said steps a second number of times usingone or more further sequences having a second, greater number of input images.According to an aspect of the present disclosure, there is provided a method of performinglossy image or video encoding, transmission and decoding, the method comprising the stepsof: receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce a first output image, wherein the first output image is an approximationof one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the first neural network and the second neural network are produced accordingto any of the above methods.According to an aspect of the present disclosure, there is provided a method of performinglossy image or video encoding and transmission, the method comprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;wherein the first neural network is produced according to any of the above methods.According to an aspect of the present disclosure, there is provided a method of performinglossy image or video decoding, the method comprising the steps of:with a second neural network, at a second computer system, decoding a latent represen-tation to produce a first output image, wherein the first output image is an approximation ofone image of an image pair of a first sequence of input images;repeating the above step to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the second neural network is produced according to any of the above methods.According to an aspect of the present disclosure, there is provided a data processing apparatusconfigured to perform any of the above methods.According to an aspect of the present disclosure, there is provided a computer programcomprising instructions which, when the program is executed by a computer, cause thecomputer to carry out the method of any of the above methods.According to an aspect of the present disclosure, there is provided a computer-readable storagemedium comprising instructions which, when executed by a computer, cause the computercarry out the method of any of the above methods.According to an aspect of the present disclosure, there is provided a method of training one ormore neural networks, the one or more neural networks being for use in lossy image or videoencoding, transmission and decoding, the method comprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;with a second neural network, decoding the latent representation to produce a first outputimage, wherein the first output image is an approximation of one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;evaluating a function based on a difference between the first sequence of output imagesand the first sequence of input images;updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps for a plurality of further sequences of input images to producefirst and second trained neural networks, wherein at least some further sequences comprise agreater number of input images than the first sequence of input images.Optionally, the first sequence of input images comprises fewer than 10 input images.Optionally, the at least some further sequences each comprise 10 or more input images.Optionally, the at least some further sequences each comprise 30 or more input images.Optionally, the at least some further sequences each comprise 60 or more input images.According to an aspect of the present disclosure, there is provided a method of performinglossy image or video encoding, transmission and decoding, the method comprising the stepsof: receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce a first output image, wherein the first output image is an approximationof one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the first neural network and the second neural network are produced accordingto any of the above methods.According to an aspect of the present disclosure, there is provided a method of performinglossy image or video encoding and transmission, the method comprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;wherein the first neural network is produced according to any of the above methods.According to an aspect of the present disclosure, there is provided a method of performinglossy image or video decoding, the method comprising the steps of:with a second neural network, at a second computer system, decoding a latent represen-tation to produce a first output image, wherein the first output image is an approximation ofone image of an image pair of a first sequence of input images;repeating the above step to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the second neural network is produced according to any of the above methods.According to an aspect of the present disclosure, there is provided a data processing apparatusconfigured to perform any of the above methods.According to an aspect of the present disclosure, there is provided a computer programcomprising instructions which, when the program is executed by a computer, cause thecomputer to carry out any of the above methods.According to an aspect of the present disclosure, there is provided a computer-readable storagemedium comprising instructions which, when executed by a computer, cause the computercarry out any of the above methods.According to an aspect, there is provided a method of training one or more neural networks,the one or more neural networks being for use in lossy image or video encoding, transmissionand decoding, the method comprising the steps of:receiving a first image and a second image at a first computer system;modifying the pixels of the first image and the second image at respective first andsecond coordinates to introduce a pixel representation of an object into the first image and thesecond image;with a first neural network, producing a latent representation of a difference between thefirst image and the second image;with a second neural network, decoding the latent representation to produce an outputimage, wherein the output image is an approximation of the first image;evaluating a function based on a difference between the output image image and the firstimage; updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps for one or more sequences of images to produce first andsecond trained neural networks.Optionally, the first and second coordinates are based on a predetermined motion path of theobject between the first image and the second image.Optionally, the one or more sequences of images comprise a predetermined balance of imagesequences having optical flows above a threshold and image sequences having optical flowsbelow the threshold.Optionally, said modifying the pixels is performed on one or more of said image sequenceshaving optical flows below the threshold.Optionally, pixels of the object define one or more alphanumeric characters.Optionally, pixels of the object define a non-uniform texture.Optionally, pixels of the object define a solid colour block.Optionally, pixels of the object are based on an image patch from an image sequence of thesequences of image different to the image sequence comprising the first image and the secondimage.Optionally, pixels of the object define a box shape.Optionally, the pixels of the object define a circular shape.Optionally, the pixels of the object define a polygonal shape.Optionally, the first and second coordinates for a first image sequence of the one or moresequences of images are different to the first and second coordinates for a second imagesequence of the one or more sequences of images, and wherein said modifying is based on thethreshold.According to a further aspect, there is provided a method of performing lossy image or videoencoding, transmission and decoding, the method comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation ofthe first image;wherein the first neural network and the second neural network are produced accordingto any of the above methods.According to a further aspect, there is provided a method of performing lossy image or videoencoding and transmission, the method comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image; andtransmitting the latent representation to a second computer system;wherein the first neural network is produced according to any of the above methodsAccording to a further aspect, there is provided a method of performing lossy image or videodecoding, the method comprising the steps of:with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation of afirst image;wherein the first neural network and the second neural network are produced accordingto any of the above methods.According to a further aspect, there is provided a data processing apparatus configured toperform any of the above methods.According to a further aspect, there is provided a computer program comprising instructionswhich, when the program is executed by a computer, cause the computer to carry out any ofthe above methods.According to a further aspect, there is provided a computer-readable storage medium comprisinginstructions which, when executed by a computer, cause the computer carry out any of theabove methods.According to a further aspect, there is provided a method of training one or more neuralnetworks, the one or more neural networks being for use in lossy image or video encoding,transmission and decoding, the method comprising the steps of:receiving a first image and a plurality of second images at a first computer system;selecting one of the plurality of second images as a reference image;with a first neural network, producing a latent representation of a difference between thefirst image and the reference image;with a second neural network, decoding the latent representation to produce an outputimage, wherein the output image is an approximation of the first image;evaluating a function based on a difference between the output image image and the firstimage; updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps for one or more sequences of images to produce first andsecond trained neural networks.Optionally, said selecting comprises sampling an image from the plurality of second images.Optionally, said selecting comprises selecting one of the plurality of second images from theimage sequence based on one or more predetermined rules.Optionally, the first image and the plurality of second images comprise consecutive images ofan image sequence.Optionally, the selected one of the plurality of second images is a positioned in the sequencebefore the first image.Optionally, the selected one of the plurality of second images is a positioned in the sequenceafter the first image .Optionally, the method comprises biasing said selecting towards a beginning of the imagesequence.Optionally, the method comprises biasing said selecting towards an end of the image sequence.Optionally, said biasing is based on a property of the first image and / or the plurality of secondimages.Optionally, said property is a scene type of the first image and / or the plurality of second images.Optionally, the one or more sequences of images comprise a predetermined balance of imagesequences having optical flows above a threshold and image sequences having optical flowsbelow the threshold, and wherein said selecting is based on the threshold.According to a further aspect, there is provided a method of performing lossy image or videoencoding, transmission and decoding, the method comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation ofthe first image;wherein the first neural network and the second neural network are produced accordingto any of the above methods.According to a further aspect, there is provided a method of performing lossy image or videoencoding and transmission, the method comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image; andtransmitting the latent representation to a second computer system;wherein the first neural network is produced according to any of the above methods.According to a further aspect, there is provided a method of performing lossy image or videodecoding, the method comprising the steps of:with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation of afirst image;wherein the first neural network and the second neural network are produced accordingto any of the above methods.According to a further aspect, there is provided a data processing apparatus configured toperform any of the above methods.According to a further aspect, there is provided a computer program comprising instructionswhich, when the program is executed by a computer, cause the computer to carry out any ofthe above methods.According to a further aspect, there is provided a computer-readable storage medium comprisinginstructions which, when executed by a computer, cause the computer carry out any of theabove methods.According to a further aspect, there is provided a method of training one or more neuralnetworks, the one or more neural networks being for use in lossy image or video encoding,transmission and decoding, the method comprising the steps of:receiving a first image and a second image at a first computer system;estimating first optical flow information and second optical flow information, each beingindicative of a difference between the first image and the second image;producing combined optical flow information from the first optical flow information andsecond optical flow information;with a first neural network, producing third optical flow information indicative of adifference between the first image and the second image;with a second neural network, decoding the third optical flow information to produce anoutput image, wherein the output image is an approximation of the first image;evaluating a function based on a difference between the combined optical flow informationand the third optical flow information;updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps for one or more sequences of images to produce first andsecond trained neural networks.Optionally, producing the combined optical flow information comprises producing a parame-terised combination of the first optical flow information and the second optical flow informationusing one or more combination parameters.Optionally, the method comprises updating the combination parameters based on one or moreterms of the evaluated function.Optionally, the parameterised combination comprises a weighted sum of the first optical flowinformation and the second optical flow information and the combination parameters compriseone or more weights of the weighted sum.Optionally, the parameterised combination comprises a pixel-wise parameterised combinationof the first optical flow information and the second optical flow information.Optionally, the function is further based on a rate term and a distortion term.Optionally, the difference between the combined optical flow information and the third opticalflow information comprises a mean squared error.Optionally, the method comprises estimating the first optical flow information using a firstoptical flow algorithm, and estimating the second optical flow information using a secondoptical flow algorithm.Optionally, the first optical flow algorithm and the second optical flow algorithm comprise anensemble of optical flow algorithms.Optionally, the method comprises estimating further optical flow information using a furtheroptical flow algorithm, and producing the combined optical flow information from the firstoptical flow information, the second optical flow information and the further optical flowinformation.Optionally, the method comprises estimating the difference between the combined optical flowinformation and the third optical flow information using a discriminator function.According to a further aspect, there is provided a method of performing lossy image or videoencoding, transmission and decoding, the method comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation ofthe first image;wherein the first neural network and the second neural network are produced accordingto any of the above methods.According to a further aspect, there is provided a method of performing lossy image or videoencoding and transmission, the method comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image; andtransmitting the latent representation to a second computer system;wherein the first neural network is produced according to any of the above methods.According to a further aspect, there is provided a method of performing lossy image or videodecoding, the method comprising the steps of:with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation of afirst image;wherein the first neural network and the second neural network are produced accordingto any of the above methods.According to a further aspect, there is provided a data processing apparatus configured toperform any of the above methods.According to a further aspect, there is provided a computer program comprising instructionswhich, when the program is executed by a computer, cause the computer to carry out any ofthe above methods.According to a further aspect, there is provided a computer-readable storage medium comprisinginstructions which, when executed by a computer, cause the computer carry out any of theabove methods.BRIEF DESCRIPTION OF THE DRAWINGSAspects of the invention will now be described by way of examples, with reference to thefollowing figures in which:Figure 1 illustrates an example of an image or video compression, transmission and decom-pression pipeline.Figure 2 illustrates a further example of an image or video compression, transmission anddecompression pipeline including a hyper-network.Figure 3 illustrates an example of a video compression, transmission and decompressionpipeline.Figure 4 illustrates an example of a video compression, transmission and decompressionsystem.Figure 5a illustrates an example plot of a distortion score of a frame sequence against framenumber.Figure 5b illustrates an example plot of a distortion score of of a frame sequence against framenumber.Figure 6 illustratively shows steps of a method of modifying training data of a video compression,transmission and decompression pipeline.Figure 7a illustrates an example plot of a distortion score of a frame sequence against framenumber.Figure 7b illustrates an example plot of a distortion score of a frame sequence against framenumber.Figure 8 illustrates a high motion frame sequence.Figure 9 illustrates a training data augmentation method for use with a video compression,transmission, and decompression method.Figure 10 illustrates a training data augmentation method for use with a video compression,transmission, and decompression method.Figure 11 illustrates method for use with training a video compression, transmission, anddecompression method.Figure 12 illustrates method for use with training a video compression, transmission, anddecompression method.DETAILED DESCRIPTION OF THE DRAWINGSCompression processes may be applied to any form of information to reduce the amountof data, or file size, required to store that information. Image and video information is anexample of information that may be compressed. The file size required to store the information,particularly during a compression process when referring to the compressed file, may bereferred to as the rate. In general, compression can be lossless or lossy. In both forms ofcompression, the file size is reduced. However, in lossless compression, no information is lostwhen the information is compressed and subsequently decompressed. This means that theoriginal file storing the information is fully reconstructed during the decompression process.In contrast to this, in lossy compression information may be lost in the compression anddecompression process and the reconstructed file may differ from the original file. Image andvideo files containing image and video data are common targets for compression.In a compression process involving an image, the input image may be represented as ^^. Thedata representing the image may be stored in a tensor of dimensions ^^ × ^^ × ^^, where ^^represents the height of the image, ^^ represents the width of the image and ^^ represents thenumber of channels of the image. Each ^^ × ^^ data point of the image represents a pixel valueof the image at the corresponding location. Each channel ^^ of the image represents a differentcomponent of the image for each pixel which are combined when the image file is displayed bya device. For example, an image file may have 3 channels with the channels representing thered, green and blue component of the image respectively. In this case, the image informationis stored in the RGB colour space, which may also be referred to as a model or a format.Other examples of colour spaces or formats include the CMKY and the YCbCr colour models.However, the channels of an image file are not limited to storing colour information and otherinformation may be represented in the channels. As a video may be considered a series ofimages in sequence, any compression process that may be applied to an image may also beapplied to a video. Each image making up a video may be referred to as a frame of the video.The output image may differ from the input image and may be represented by ^^. The differencebetween the input image and the output image may be referred to as distortion or a differencein image quality. The distortion can be measured using any distortion function which receivesthe input image and the output image and provides an output which represents the differencebetween input image and the output image in a numerical way. An example of such a methodis using the mean square error (MSE) between the pixels of the input image and the outputimage, but there are many other ways of measuring distortion, as will be known to the personskilled in the art. The distortion function may comprise a trained neural network.Typically, the rate and distortion of a lossy compression process are related. An increase inthe rate may result in a decrease in the distortion, and a decrease in the rate may result in anincrease in the distortion. Changes to the distortion may affect the rate in a correspondingmanner. A relation between these quantities for a given compression technique may be definedby a rate-distortion equation.AI based compression processes may involve the use of neural networks. A neural network isan operation that can be performed on an input to produce an output. A neural network maybe made up of a plurality of layers. The first layer of the network receives the input. One ormore operations may be performed on the input by the layer to produce an output of the firstlayer. The output of the first layer is then passed to the next layer of the network which mayperform one or more operations in a similar way. The output of the final layer is the output ofthe neural network.Each layer of the neural network may be divided into nodes. Each node may receive at leastpart of the input from the previous layer and provide an output to one or more nodes in asubsequent layer. Each node of a layer may perform the one or more operations of the layer onat least part of the input to the layer. For example, a node may receive an input from one ormore nodes of the previous layer. The one or more operations may include a convolution, aweight, a bias and an activation function. Convolution operations are used in convolutionalneural networks. When a convolution operation is present, the convolution may be performedacross the entire input to a layer. Alternatively, the convolution may be performed on at leastpart of the input to the layer.Each of the one or more operations is defined by one or more parameters that are associatedwith each operation. For example, the weight operation may be defined by a weight matrixdefining the weight to be applied to each input from each node in the previous layer to eachnode in the present layer. In this example, each of the values in the weight matrix is a parameterof the neural network. The convolution may be defined by a convolution matrix, also knownas a kernel. In this example, one or more of the values in the convolution matrix may be aparameter of the neural network. The activation function may also be defined by values whichmay be parameters of the neural network. The parameters of the network may be varied duringtraining of the network.Other features of the neural network may be predetermined and therefore not varied duringtraining of the network. For example, the number of layers of the network, the number ofnodes of the network, the one or more operations performed in each layer and the connectionsbetween the layers may be predetermined and therefore fixed before the training process takesplace. These features that are predetermined may be referred to as the hyperparameters of thenetwork. These features are sometimes referred to as the architecture of the network.To train the neural network, a training set of inputs may be used for which the expected output,sometimes referred to as the ground truth, is known. The initial parameters of the neuralnetwork are randomized and the first training input is provided to the network. The output ofthe network is compared to the expected output, and based on a difference between the outputand the expected output the parameters of the network are varied such that the differencebetween the output of the network and the expected output is reduced. This process is thenrepeated for a plurality of training inputs to train the network. The difference between theoutput of the network and the expected output may be defined by a loss function. The result ofthe loss function may be calculated using the difference between the output of the networkand the expected output to determine the gradient of the loss function. Back-propagation ofthe gradient descent of the loss function may be used to update the parameters of the neuralnetwork using the gradients ^^^^ / ^^^^ of the loss function. A plurality of neural networks in asystem may be trained simultaneously through back-propagation of the gradient of the lossfunction to each network.In the context of image or video compression, this type of system, where simultaneous trainingwith back-propagation through each element or the whole network architecture may be referredto as end-to-end, learned image or video compression. Unlike in traditional compressionalgorithms that use primarily handcrafted, manually constructed steps, an end-to-end learnedsystem learns itself during training what combination of parameters best achieves the goal ofminimising the loss function. This approach is advantageous compared to systems that are notend-to-end learned because an end-to-end system has a greater flexibility to learn weights andparameters that might be counter-intuitive to someone handcrafting features.It will be appreciated that the term "training" or "learning" as used herein means the processof optimizing an artificial intelligence or machine learning model, based on a given set of data.This involves iteratively adjusting the parameters of the model to minimize the discrepancybetween the model’s predictions and the actual data, represented by the above-describedrate-distortion loss function.The training process may comprise multiple epochs. An epoch refers to one complete passof the entire training dataset through the machine learning algorithm. During an epoch, themodel’s parameters are updated in an effort to minimize the loss function. It is envisaged thatmultiple epochs may be used to train a model, with the exact number depending on variousfactors including the complexity of the model and the diversity of the training data.Within each epoch, the training data may be divided into smaller subsets known as batches.The size of a batch, referred to as the batch size, may influence the training process. A smallerbatch size can lead to more frequent updates to the model’s parameters, potentially leading tofaster convergence to the optimal solution, but at the cost of increased computational resources.Conversely, a larger batch size involves fewer updates, which can be more computationallyefficient but might converge slower or even fail to converge to the optimal solution.The learnable parameters are updated by a specified amount each time, determined by thelearning rate. The learning rate is a hyperparameter that decides how much the parametersare adjusted during the training process. A smaller learning rate implies smaller steps in theparameter space and a potentially more accurate solution, but it may require more epochs toreach that solution. On the other hand, a larger learning rate can expedite the training processbut may risk overshooting the optimal solution or causing the training process to diverge.The training described herein may involve use of a validation set, which is a portion of thedata not used in the initial training, which is used to evaluate the model’s performance and toprevent overfitting. Overfitting occurs when a model learns the training data too well, to thepoint that it fails to generalize to unseen data. Regularization techniques, such as dropout orL1 / L2 regularization, can also be used to mitigate overfitting.It will be appreciated that training a machine learning model is an iterative process thatmay comprise selection and tuning of various parameters and hyperparameters. As will beappreciated, the specific details, such as hyper parameters and so on, of the training processmay vary and it is envisaged that producing a trained model in this way may achieved in anumber of different ways with different epochs, batch sizes, learning rates, regularisations,and so on, the details of which are not essential to enabling the advantages and effects of thepresent disclosure, except where stated otherwise. The point at which an “untrained” neuralnetwork is considered be “trained” is envisaged to be case specific and depend on, for example,on a number of epochs, a plateauing of any further learning, or some other metric and is notconsidered to be essential in achieving the advantages described herein.More details of an end-to-end, learned compression process will now be described. It will beappreciated that in some cases, end-to-end, learned compression processes may be combinedwith one or more components that are handcrafted or trained separately.In the case of AI based image or video compression, the loss function may be defined by therate distortion equation. The rate distortion equation may be represented by ^^^^^^^^ = ^^ + ^^ ∗ ^^,where ^^ is the distortion function, ^^ is a weighting factor, and ^^ is the rate loss. ^^ may bereferred to as a lagrange multiplier. The langrange multiplier provides as weight for a particularterm of the loss function in relation to each other term and can be used to control which termsof the loss function are favoured when training the network.In the case of AI based image or video compression, a training set of input images maybe used. An example training set of input images is the KODAK image set (for exampleat www.cs.albany.edu / xypan / research / snr / Kodak.html). An example training set of inputimages is the IMAX image set. An example training set of input images is the Imagenetdataset (for example at www.image-net.org / download). An example training set of inputimages is the CLIC Training Dataset P (“professional”) and M (“mobile”) (for example athttp: / / challenge.compression.cc / tasks / ).An example of an AI based compression, transmission and decompression process 100 isshown in Figure 1. As a first step in the AI based compression process, an input image 5 isprovided. The input image 5 is provided to a trained neural network 110 characterized by afunction ^^^^ acting as an encoder. The encoder neural network 110 produces an output basedon the input image. This output is referred to as a latent representation of the input image 5. Ina second step, the latent representation is quantised in a quantisation process 140 characterisedby the operation ^^, resulting in a quantized latent. The quantisation process transforms thecontinuous latent representation into a discrete quantized latent. An example of a quantizationprocess is a rounding function.In a third step, the quantized latent is entropy encoded in an entropy encoding process 150 toproduce a bitstream 130. The entropy encoding process may be for example, range or arithmeticencoding. In a fourth step, the bitstream 130 may be transmitted across a communicationnetwork.In a fifth step, the bitstream is entropy decoded in an entropy decoding process 160. Thequantized latent is provided to another trained neural network 120 characterized by a function^^^^ acting as a decoder, which decodes the quantized latent. The trained neural network 120produces an output based on the quantized latent. The output may be the output image of theAI based compression process 100. The encoder-decoder system may be referred to as anautoencoder.Entropy encoding processes such as range or arithmetic encoding are typically able to losslesslycompress given input data up to close to the fundamental entropy limit of that data, as determinedby the total entropy of the distribution of that data. Accordingly, one way in which end-to-end,learned compression can minimise the rate loss term of the rate-distortion loss function andthereby increase compression effectiveness is to learn autoencoder parameter values thatproduce low entropy latent representation distributions. Producing latent representationsdistributed with as low an entropy as possible allows entropy encoding to compress the latentdistributions as close to or to the fundamental entropy limit for that distribution. The lowerthe entropy of the distribution, the more entropy encoding can losslessly compress it and thelower the amount of data in the corresponding bitstream. In some cases where the latentrepresentation is distributed according to a gaussian or Laplacian distribution, this learningmay comprise learning optimal location and scale parameters of the gaussian or Laplaciandistributions, in other cases, it allows the learning of more flexible latent representationdistributions which can further help to achieve the minimising of the rate-distortion lossfunction in ways that are not intuitive or possible to do with handcrafted features. Examples ofthese and other advantages are described in WO2021 / 220008A1, which is incorporated in itsentirety by reference.Something which is closely linked to the entropy encoding of the latent distribution and whichaccordingly also has an effect on the effectiveness of compression of end-to-end learnedapproaches is the quantisation step. During inference, a rounding function may be used toquantise a latent representation distribution into bins of given sizes, a rounding function isnot differentiable everywhere. Rather, a rounding function is effectively one or more stepfunctions whose gradient is either zero (at the top of the steps) or infinity (at the boundarybetween steps). Back propagating a gradient of a loss function through a rounding functionis challenging. Instead, during training, quantisation by rounding function is replaced byone or more other approaches. For example, the functions of a noise quantisation model aredifferentiable everywhere and accordingly do allow backpropagation of the gradient of theloss function through the quantisation parts of the end-to-end, learned system. Alternatively, astraight-through estimator (STE) quantisation model or one other quantisation models may beused. It is also envisaged that different quantisation models may be used for during evaluationof different term of the loss function. For example, noise quantisation may used to evaluate therate or entropy loss term of the rate-distortion loss function while STE quantisation may beused to evaluate the distortion term.In a similar manner to how learning parameters top produce certain distributions of the latentrepresentation facilitates achieving better rate loss term minimisation, end-to-end learning ofthe quantisation process achieves a similar effect. That is, learnable quantisation parametersprovide the architecture with a further degree of freedom to achieve the goal of minimising theloss function. For example, parameters corresponding to quantisation bin sizes may be learnedwhich is likely to result in an improved rate-distortion loss outcome compared to approachesusing hand-crafted quantisation bin sizes.Further, as the rate-distortion loss function constantly has to balance a rate loss term against adistortion loss term, it has been found that the more degrees of freedom the system has duringtraining, the better the architecture is at achieving optimal rate and distortion trade off.The system described above may be distributed across multiple locations and / or devices. Forexample, the encoder 110 may be located on a device such as a laptop computer, desktopcomputer, smart phone or server. The decoder 120 may be located on a separate device whichmay be referred to as a recipient device. The system used to encode, transmit and decode theinput image 5 to obtain the output image 6 may be referred to as a compression pipeline.The AI based compression process may further comprise a hyper-network 105 for thetransmission of meta-information that improves the compression process. The hyper-network105 comprises a trained neural network 115 acting as a hyper-encoder ^^ ℎ^^and a trained neuralnetwork 125 acting as a hyper-decoder ^^ℎ^^. An example of such a is shown in Figure 2.Components of the system not further discussed may be assumed to be the same as discussedabove. The neural network 115 acting as a hyper-decoder receives the latent that is the output ofthe encoder 110. The hyper-encoder 115 produces an output based on the latent representationthat may be referred to as a hyper-latent representation. The hyper-latent is then quantizedin a quantization process 145 characterised by ^^ℎ to produce a quantized hyper-latent. Thequantization process 145 characterised by ^^ℎ may be the same as the quantisation process 140characterised by ^^ discussed above.In a similar manner as discussed above for the quantized latent, the quantized hyper-latent isthen entropy encoded in an entropy encoding process 155 to produce a bitstream 135. Thebitstream 135 may be entropy decoded in an entropy decoding process 165 to retrieve thequantized hyper-latent. The quantized hyper-latent is then used as an input to trained neuralnetwork 125 acting as a hyper-decoder. However, in contrast to the compression pipeline 100,the output of the hyper-decoder may not be an approximation of the input to the hyper-decoder115. Instead, the output of the hyper-decoder is used to provide parameters for use in theentropy encoding process 150 and entropy decoding process 160 in the main compressionprocess 100. For example, the output of the hyper-decoder 125 can include one or more ofthe mean, standard deviation, variance or any other parameter used to describe a probabilitymodel for the entropy encoding process 150 and entropy decoding process 160 of the latentrepresentation. In the example shown in Figure 2, only a single entropy decoding process 165and hyper-decoder 125 is shown for simplicity. However, in practice, as the decompressionprocess usually takes place on a separate device, duplicates of these processes will be presenton the device used for encoding to provide the parameters to be used in the entropy encodingprocess 150.Further transformations may be applied to at least one of the latent and the hyper-latent at anystage in the AI based compression process 100. For example, at least one of the latent and thehyper latent may be converted to a residual value before the entropy encoding process 150,155is performed. The residual value may be determined by subtracting the mean value of thedistribution of latents or hyper-latents from each latent or hyper latent. The residual valuesmay also be normalised.To perform training of the AI based compression process described above, a training set ofinput images may be used as described above. During the training process, the parameters ofboth the encoder 110 and the decoder 120 may be simultaneously updated in each trainingstep. If a hyper-network 105 is also present, the parameters of both the hyper-encoder 115and the hyper-decoder 125 may additionally be simultaneously updated in each training step.The training process may further include a generative adversarial network (GAN). Whenapplied to an AI based compression process, in addition to the compression pipeline describedabove, an additional neutral network acting as a discriminator is included in the system. Thediscriminator receives an input and outputs a score based on the input providing an indicationof whether the discriminator considers the input to be ground truth or fake. For example, theindicator may be a score, with a high score associated with a ground truth input and a lowscore associated with a fake input. For training of a discriminator, a loss function is used thatmaximizes the difference in the output indication between an input ground truth and input fake.When a GAN is incorporated into the training of the compression process, the output image 6may be provided to the discriminator. The output of the discriminator may then be used in theloss function of the compression process as a measure of the distortion of the compressionprocess. Alternatively, the discriminator may receive both the input image 5 and the outputimage 6 and the difference in output indication may then be used in the loss function of thecompression process as a measure of the distortion of the compression process. Training ofthe neural network acting as a discriminator and the other neutral networks in the compressionprocess may be performed simultaneously. During use of the trained compression pipelinefor the compression and transmission of images or video, the discriminator neural network isremoved from the system and the output of the compression pipeline is the output image 6.Incorporation of a GAN into the training process may cause the decoder 120 to performhallucination. Hallucination is the process of adding information in the output image 6 thatwas not present in the input image 5. In an example, hallucination may add fine detail tothe output image 6 that was not present in the input image 5 or received by the decoder 120.The hallucination performed may be based on information in the quantized latent received bydecoder 120.Details of a video compression process will now be described. As discussed above, a video ismade up of a series of images arranged in sequential order. AI based compression process100 described above may be applied multiple times to perform compression, transmissionand decompression of a video. For example, each frame of the video may be compressed,transmitted and decompressed individually. The received frames may then be grouped toobtain the original video.The frames in a video may be labelled based on the information from other frames that is usedto decode the frame in a video compression, transmission and decompression process. Asdescribed above, frames which are decoded using no information from other frames may bereferred to as I-frames. Frames which are decoded using information from past frames may bereferred to as P-frames. Frames which are decoded using information from past frames andfuture frames may be referred to as B-frames. Frames may not be encoded and / or decoded inthe order that they appear in the video. For example, a frame at a later time step in the videomay be decoded before a frame at an earlier time.The images represented by each frame of a video may be related. For example, a number offrames in a video may show the same scene. In this case, a number of different parts of thescene may be shown in more than one of the frames. For example, objects or people in a scenemay be shown in more than one of the frames. The background of the scene may also beshown in more than one of the frames. If an object or the perspective is in motion in the video,the position of the object or background in one frame may change relative to the position ofthe object or background in another frame. The transformation of a part of the image froma first position in a first frame to a second position in a second frame may be referred to asflow, warping or motion compensation. The flow may be represented by a vector. One or moreflows that represent the transformation of at least part of one frame to another frame may bereferred to as a flow map.An example AI based video compression, transmission, and decompression process 200 isshown in Figure 3. The process 200 shown in Figure 3 is divided into an I-frame part 201for decompressing I-frames, and a P-frame part 202 for decompressing P-frames. It will beunderstood that these divisions into different parts are arbitrary and the process 200 may bealso be considered as a single, end-to-end pipeline.As described above, I-frames do not rely on information from other frames so the I-frame part201 corresponds to the compression, transmission, and decompression process illustrated inFigures 1 or 2. The specific details will not be repeated here but, in summary, an input image^^0 is passed into an encoder neural network 203 producing a latent representation which isquantised and entropy encoded into a bitstream 204. The subscript 0 in ^^0 indicates the inputimage corresponds to a frame of a video stream at position t = 0. This may be the first frame ofan entire video stream or the first frame of a chunk of a video stream made up of, for example,an I-frame and a plurality of subsequent P-frames and / or B-frames. The bitstream 204 is thenentropy decoded and passed into a decoder neural network 205 to reproduce a reconstructedimage ^^0 which in this case is an I-frame. The decoding step may be performed both locallyat the same location as where the input image compression occurs as well as at the locationwhere the decompression occurs. This allows the reconstructed image ^^0 to be available forlater use by components of both the encoding and decoding sides of the pipeline.In contrast to I-frames, P-frames (and B-frames) do rely on information from other frames.Accordingly, the P-frame part 202 at the encoding side of the pipeline takes as input not onlythe input image ^^^^ that is to be compressed (corresponding to a frame of a video stream atposition t), but also one or more previously reconstructed images ^^^^−1 from an earlier framet-1. As described above, the previously reconstructed ^^^^−1 is available at both the encodeand decode side of the pipeline and can accordingly be used for various purposes at both theencode and decode sides.At the encode side, previously reconstructed images may be used for generating a flow mapscontaining information indicative of inter-frame movement of pixels between frames. In theexample of Figure 3, both the image being compressed ^^^^ and the previously reconstructedimage from an earlier frame ^^^^−1 are passed into a flow module part 206 of the pipeline. Theflow module part 206 comprises an autoencoder such as that of the autoencoder systems ofFigures 1 and 2 but where the encoder neural network 207 has been trained to produce alatent representation of a flow map from inputs ^^^^−1 and ^^^^ , which is indicative of inter-framemovement of pixels or pixel groups between ^^^^−1 and ^^^^ . The latent representation of the flowmap is quantised and entropy encoded to compress it and then transmitted as a bitstream 208.On the decode side, the bitstream is entropy decoded and passed to a decoder neural network209 to produce a reconstructed flow map ^^ .The reconstructed flow map ^^ is applied to the previously reconstructed image ^^^^−1 to generatea warped image ^^^^−1,^^. It is envisaged that any suitable warping technique may be used, forexample bi-linear or tri-linear warping, as is described in Agustsson, E., Minnen, D., Johnston,N., Balle, J., Hwang, S. J., and Toderici, G. (2020), Scale-space flow for end-to-end optimizedvideo compression. In Proceedings of the IEEE / CVF Conference on Computer Vision andPattern Recognition (pp. 8503-8512), which is hereby incorporated by reference. It is furtherenvisaged that a scale-space flow approach as described in the above paper may also optionallybe used. The warped image ^^^^−1,^^ is a prediction of how the previously reconstructed image^^^^−1 might have changed between frame positions t-1 and t, based on the output flow mapproduced by the flow module part 206 autoencoder system from the inputs of ^^^^ and ^^^^−1.As with the I-frame, the reconstructed flow map ^^ and corresponding warped image ^^^^−1,^^may be produced both on the encode side and the decode side of the pipeline so they areavailable for use by other components of the pipeline on both the encode and decode sides.In the example of Figure 3, both the image being compressed ^^^^ and the ^^^^−1,^^ are passedinto a residual module part 210 of the pipeline. The residual module part 210 comprises anautoencoder system such as that of the autoencoder systems of Figures 1 and 2 but where theencoder neural network 211 has been trained to produce a latent representation of a residualmap indicative of differences between the input mage ^^^^ and the warped image ^^^^−1,^^. Thelatent representation of the residual map is then quantised and entropy encoded into a bitstream212 and transmitted. The bitstream 212 is then entropy decoded and passed into a decoderneural network 213 which reconstructs a residual map ^^ from the decoded latent representation.Alternatively, a residual map may first be pre-calculated between ^^^^ and the ^^^^−1,^^ and thepre-calculated residual map may be passed into an autoencoder for compression only. Thishand-crafted residual map approach is computationally simpler, but reduces the degrees offreedom with which the architecture may learn weights and parameters to achieve its goalduring training of minimising the rate-distortion loss function.Finally, on the decode side, the residual map ^^ is applied (e.g. combined by addition, subtractionor a different operation) to the warped image to produce a reconstructed image ^^^^ which is areconstruction of image ^^^^ and accordingly corresponds to a P-frame at position t in a sequenceof frames of a video stream. It will be appreciated that the reconstructed image ^^^^ can then beused to process the next frame. That is, it can be used to compress, transmit and decompress^^^^+1, and so on until an entire video stream or chunk of a video stream has been processed.Alternatively, the residual autoencoder may be trained to reconstruct the frame ^^^^ directlyfrom the entropy decoded bitstream by removing the connection between ^^^^−1,^^ and the outputof the residual block 210, thereby eliminating any direct combination step with the warpedpreviously decoded image to speed up inference. In this case, the flow information is intuitivelyunderstood to be indirectly captured within the residual information, which the residual decoderis able to learn to use to directly reconstruct the output image ^^^^ .Alternatively, the residual autoencoder may be trained to reconstruct the frame ^^^^ directly fromthe entropy decoded bitstream in combination with some representation of flow injected intoone or more layers of the residual decoder. In this case, the flow information is intuitivelyunderstood to be indirectly captured within the injected information, which the residual decoderis able to learn to use while decoding the latent representation of flow information to directlyreconstruct the output image ^^^^ .Thus, for a block of video frames comprising an I-frame and ^^ subsequent P-frames, thebitstream may contain (i) a quantised, entropy encoded latent representation of the I-frameimage, and (ii) a quantised, entropy encoded latent representation of a flow map and residualmap of each P-frame image. For completeness, whilst not illustrated in Figure 3, any of theautoencoder systems of Figure 3 may comprise hyper and hyper-hyper networks such as thosedescribed in connection with Figure 2. Accordingly, the bitstream may also contain hyper andhyper-hyper parameters, their latent quantised, entropy encoded latent representations and soon, of those networks as applicable.Finally, the above approach may generally also be extended to B-frames, for example as isdescribed in Pourreza, R., and Cohen, T. (2021). Extending neural p-frame codecs for b-framecoding. In Proceedings of the IEEE / CVF International Conference on Computer Vision (pp.6680-6689).The above-described flow and residual based approach is highly effective at reducing theamount of data that needs to be transmitted because, as long as at least one reconstructed frame(e.g. I-frame ^^^^−1) is available, the encode side only needs to compress and transmit a flowmap and a residual map (and any hyper or hyper-hyper parameter information, as applicable)to reconstruct a subsequent frame.Figure 4 shows an example of an AI image or video compression process such as that describedabove in connection with Figures 1-3 implemented in a video streaming system 400. Thesystem 400 comprises a first device 401 and a second device 402. The first and seconddevices 401, 402 may be user devices such as smartphones, tablets, AR / VR headsets or otherportable devices. In contrast to known systems which primarily perform inference on GPUssuch as Nvidia A100, Geforce 3090, Gefore 4090 GPU cards, the system 400 of Figure 4performs inference on a CPU of the first and second devices respectively. That is, computefor performing both encoding and decoding are performed by the respective CPUs of the firstand second devices 401, 402. This places very different power usage, memory and runtimeconstraints on the implementation of the above methods than when implementing AI-basedcompression methods on GPUs. In one example, the CPU of first and second devices 401, 402may comprise a Qualcomm Snapdragon CPU.The first device 401 comprises a media capture device 403, such as a camera, arranged tocapture a plurality of images, referred to hereafter as a video stream 404, of a scene 404. Thevideo stream 404 is passed to a pre-processing module 406 which splits the video stream intoblocks of frames, various frames of which will be designated as I-frames, P-frames, and / orB-frames. The blocks of frames are then compressed by an AI-compression module 407comprising the encode side of the AI-based video compression pipeline of Figure 3. Theoutput of the AI-compression module is accordingly a bitstream 408a which is transmittedfrom the first device 401, for example via a communications channel, for example over oneor more of a WiFi, 3G, 4G or 5G channel, which may comprise internet or cloud-based 409communications.The second device 402 receives the communicated bitstream 408b which is passed to anAI-decompression module 410 comprising the decode side of the AI-based video compressionpipeline of Figure 3. The output of the AI-decompression module 402 is the reconstructedI-frames, P-frames and / or B-frames which are passed to a post-processing module 411 wherethey can prepared, for example passed into a buffer, in preparation for streaming 412 to andrendering on a display device 413 of the second device 402.It is envisaged that the system 400 of Figure 4 may be used for live video streaming at 30fps ofa 1080p video stream, which means a cumulative latency of both the encode and decode sideis below substantially 50ms, for example substantially 30ms or less. Achieving this level ofruntime performance with only CPU compute on user devices presents challenges which arenot addressed by known methods and systems or in the wider AI-compression literature.For example, execution of different parts of the compression pipeline during inferencemay be optimized by adjusting the order in which operations are performed using one ormore known CPU scheduling methods. Efficient scheduling can allow for operations to beperformed in parallel, thereby reducing the total execution time. It is also envisaged thatefficient management of memory resources may be implemented, including optimising cachingmethods such as storing frequently-accessed data in faster memory locations, and memoryreuse, which minimizes memory allocation and deallocation operations.A number of concepts related to the AI compression processes and / or their implementationin a hardware system discussed above will now be described. Although each concept isdescribed separately, one or more of the concepts described below may be applied in an AIbased compression process as described above.Concept 1: Frozen-patch long GOP trainingA compression artefact type that appears specifically in AI-based image and video compression,but not in traditional compression, is known as "unrolling". Unrolling is an artefact thattypically appears in image sequences where there is little to no movement between frames andmanifests itself in the reconstructed images as halo-like rings around edges and objects in theimages.These artefacts typically arise where small, barely noticeable changes in pixel values betweenframes such as tiny lighting changes, block artefacts from traditional compression (if the inputto the AI compression pipeline came from a traditionally compressed source), noise, and so onbecome amplified over time by the AI based compression pipeline.This results in the networks of the compression pipeline treating what is effectively imagenoise as real movement and then exaggerating this as the sequence of frames progresses. Atthe end of the sequence of frames, the artefacts can be quite severe and detrimental to theimage reconstruction quality.In long group of picture (GOP) sequences, the effect can be particularly pronounced as errorspropagate much further over the course of long P- and / or B-frame sequences before the nextI-frame provides "fresh" information without artefacts for the next set of P- and / or B-frames tobe generated from.Unrolling artefacts occur across mostly all types of image sequences but are particularlyapparent when AI based flow-residual compression models are used to encode and decodeimage sequences of internet video calls (whereby a foreground user who moves very little issurrounded by a largely static background which also does not move much), screen sharingimage sequences, security camera feeds and others. As described above, if there is no movementin the scene after a period of time, objects and edges start to exhibit halo-like ringing aroundtheir edges. The artefacts also result in the compression pipeline assigning a disproportionatenumber of bits to handle the hallucinated movement of the unrolling artefacts even thoughthe corresponding parts of the actual images aren’t changing or are hardly changing betweenframes so ought not to have many bits assigned to them.One advantage of an AI-based compression pipeline over traditional compression is that neuralnetworks are significantly more expressive than the hand-crafted techniques of traditionalcompression and can adapt very well to different training data. At a very general level, if anAI-based compression pipeline is not good at compressing certain types of sequences, biasingthe training data the networks of the pipeline have been trained on by adding more examplesof the problematic sequences is a generally successful approach to improving that pipeline’sability to handle the problematic sequences.This can be achieved for example by doing a base model training run on a balanced data set andthen finetuning the (trained) base model on the problematic sequences, or it can be achieved bysimply biasing the base model training data from the beginning.However, even when taking these approach, there are some types of sequences where biasingthe training data does not result in the expected improvement. For Long GOP sequences are anexample of such a type of sequence. The term "long GOP" as used herein means a sequenceof frames of more than substantially 10 frames. For example, a sequence of one I-frame andnine P- and / or B frames, or some other combination thereof. The advantages of the presentdisclosure become particularly apparent in respect of long GOP sequences of at least 100frames, for example 120 or more frames (for example one I-frame and 119 P- and / or B-frames).It turns out that even when the networks of the architecture are trained on training data biasedwith more long GOP sequences, there is only a very small performance improvement (in termsof bit rate and reconstruction accuracy e.g. as measured by PSNR), and the above-describedunrolling artefacts are still present.Accordingly, the present disclosure proposes a number of different approaches to help improveperformance on long GOP sequences.The first of these approaches is referred to herein as frozen-patch long GOP training.Frozen-patch long GOP training modifies one or more GOPs in the training data by selectingonly a patch of the frames of the GOP to to show to the network, and masking the rest of theframe using a frozen frame from the GOP. For example, we take the first I-frame of a GOP,select a patch (e.g. by pixel coordinates) I-frame, and modify the rest of the frames in the GOPby masking everything outside of those pixel coordinates in each frame by copying over thepixels from the I-frame that were not in the patch. The selection of the patch may be random,or based on some other criteria such as an image statistics property associated with the frame,and the patch size can be different across the training data (e.g. some with a patch size of 0where we mask the entire frame sequence with a duplicated I-frame. some with a patch sizecorresponding to the entire image resolution where we mask none of the pixels and keep theframe sequence unmodified, and some patch sizes in between these two extremes).In this way, the networks during training are exposed to static portions (the masked portionthat is duplicated from the I-frame), as well as movement containing portions (the selectedpatch in which the scene is played in the usual way across the sequence of frames).Whilst this modification of the training data results in what, at first glance, looks like artefact-ridden, blocky training data, it turns out, counter-intuitively, that an AI-based compressionpipeline trained on this type of training data is substantially better at handling longer GOPsequences in terms of bit rate (e.g. bits per pixel) and reconstruction accuracy (e.g. PSNR,MSE or some other quantified difference). It is also noted that this improvement is seen acrossall types of long GOP sequence scene types including video conference frame sequences,video game sequences, nature scene sequences, and so on. This above-described approach alsohelps to mitigate unrolling artefacts. This is in part understood to be due to the networks beinggiven the (noiseless) frozen parts of the GOP in each frame and so learning that any smalldeviations that can result in unrolling artefacts should not be propagated across the sequence.Figure 5a is a plot 500 of reconstruction accuracy (quantified in mean squared error, MSE,between the input frame and the output frame) against frame number of a sequence of one I-frame (frame index 1) and 149 P-frames (frame index 2 to 150) produced by an AI-compressionpipeline trained on training data where frozen-patch long GOP training was not applied. It canbe seen from the gradually increasing MSE as the frames progress that the image reconstructionaccuracy is decreasing, indicating that unrolling artefacts start to dominate the further awaythe last I-frame was. Frames beyond number 150 are not shown as the MSE just continues toincrease until the artefacts are so significant the image becomes unrecognisable.Figure 5b is a similar plot 501 but now the frames are produced by an AI-compressionpipeline that has been trained using the above-described frozen-patch long GOP training. Itis immediately apparent that the MSE does not substantially grow as the frames progressindicating that artefacts are not propagating through time across the sequence of frames. Indeed,data up to frame number 300 is plotted in the plot of Figure 5b and there is no substantialincrease in MSE compared to the early frame numbers. This can be contrasted with the plot500 in Figure 5a which shows substantial MSE increases even after only 150 frames.Figure 6 illustratively shows steps 600 of a method of modifying training data using theabove-described frozen-patch long GOP training. First, a GOP 601 is provided. Whilst only asmall number of frames are shown for brevity, in practice a full GOP of some predeterminedlength is provided. Each of the frames is part of a sequence of a scene in which a character ismoving his arms up and down and then walking off.A patch 602 of the first frame of the GOP, for example an I-frame 603, is selected by pixelcoordinates and the pixels of the first frame outside of those coordinates are copied as a mask604. The mask 604, made up of the non-patch first frame pixels is are applied 605 across allthe other frames of the GOP 601, leaving only those patches at the non-mask pixel coordinatesof the other frames untouched.In the example of Figure 6, as the character walks off, this will only be apparent from thepixels within the patch, but the mask frame remains unchanged, containing the original I-framepixels that include the characters arm, body and part of the head. Effectively, the I-frame isfrozen and applied to the whole GOP except for the pixels within the selected patch.The modified GOP is then stored as a sequence of frames of training data, and the same processis repeated for another GOP, and so on to produce the full training data set.It is envisaged that the patch sizes and positions may be chosen randomly according tosome distribution, for example a uniform distribution, or Gaussian distribution or some otherdistribution. It has been found that a uniform distribution of different patch sizes and positionsproduces a training data set that, when trained on, is most robust to different GOP lengths (bothlong and short) during inference. Alternatively, patch sizes and positions may be selectedbased on some property of the data the method is being applied to (e.g. selecting a patch sizeof around 128x128 pixels or 256x256 pixels or more pixels and positioned around the centerpixels of webcam video data around where the person in the webcam video data is positioned).Envisaged data types where the patch size and position may be predetermined include but arenot limited to webcam video data, static scene data, video game data, high motion video data,low motion CCTV data. In each case, the patch size and position is typically based aroundtypical motion or non-motion patterns and where detail is to be preserved as much as possible,such as at or around edges, at or around typical movement paths and patterns, and so on.Further, it is envisaged that the distribution of patch sizes and positions can be used to controlthe performance of the model trained on it for different needs. For example, if a given modelis only intended to be used on long GOP sequences where the center of the frame is expectedto contain detail and movement (e.g. webcam video conference scenes), the distribution ofpatch sizes and positions can be focused around the center of the frames when producing thefrozen-patch long GOP training data (e.g. a Gaussian distribution with a mean at the centerpixel of the frames). At a high level, this means the frozen-patch long GOP training approachcan advantageously be used as a way to specialise the networks of an AI-based compressionpipeline for a given purpose, such as compressing and reconstructing webcam video conferenceframes.For completeness, it is noted that the frozen-patch approach can also be applied to short GOPsequences (e.g. fewer than 10 frames). In this case, whilst short GOP sequences are alreadytypically less problematic, the technique can still be used to control the performance of themodel at by, for example, focusing the patch distribution around different positions within theframe to teach the networks to improve performance around the positions of the patches in theframe.Below is provided pseudocode for training one or more networks of an AI-based compressionpipeline using the above-described frozen-patch GOP training approach.

[0002] Algorithm 1 Neural Network Training frozen-patch GOPsRequire: Training dataset D, learning rate ^^, regularization parameter ^^, number of epochs ^^Select patches and frozen pixels of GOPs of DGenerate frozen-patch training data D^^^^^^ using selected patches and frozen pixelsInitialize network parameters ^^ randomlyfor ^^^^^^^^ℎ = 1 to ^^ dofor each batch (^^, ^^) in D^^^^^^ doForward pass to compute predictions ^^ = ^^^^ (^^)Compute ^^^^^^^^Backward pass to compute gradients ∇^^ ^^^^^^^^Update parameters with optimizer O: ^^ ← O(^^, ∇^^ ^^^^^^^^, ^^)end forOptionally evaluate on validation setend forA further technique that can be used, either together with the above-described frozen-patchGOP training, or in a standalone manner is long GOP training with truncated back propagationin time, as will be described below in concept 2.Concept 2: Long GOP training with truncated back propagation through timeOne problem with training on long GOP sequences is that training can be unstable. This can,in part, be attributed to the unrolling artefacts (and other artefacts) manifesting themselves inthe reconstructed frames of the long GOPs. This in turn increases the distortion loss whentraining on long GOP sequences and can prevent the loss from converging well. In the casewhere unrolling artefacts are particularly prominent this can cause the loss to increase andbehave unpredictably.The inventors have realised that one way to address this problem is by using TBPTT.Truncated Backpropagation Through Time (TBPTT) is a technique used to train neuralnetworks by limiting the number of time steps over which gradients are propagated backwardthrough the network. This approach helps to reduce computational costs and mitigate thevanishing or exploding gradient problem, enhancing the stability and efficiency of trainingneural networks.At a very general level, known TBPTT techniques comprise the following steps:1. Segment the Input Sequence: Divide the input sequence into overlapping smallerchunks of a fixed length, ^^ .2. Forward Pass: For each chunk, perform a forward pass through the neural network for^^ time steps, computing the hidden states and outputs.3. Backward Pass: Backpropagate the error through only the last ^^ time steps, insteadof the entire sequence.4. Update Parameters: Update the network parameters based on the gradients computedin the backward pass.5. Move to Next Chunk: Repeat the process for the next chunk of the sequence, usingthe final hidden state of the previous chunk as the initial state for the next chunk.By way of worked example, consider a sequence of length 12 and a TBPTT chunk size of 4 fora sequence modeling task using an neural network.• Sequence: [^^1, ^^2, ..., ^^12]• Chunk Size: 4The following steps use a known TBPTT technique with overlapping chunks whereby theoverlap in this example is 1:Step 1: Divide the sequence into overlapping chunks: [^^1, ^^2, ^^3, ^^4], [^^2, ^^3, ^^4, ^^5],[^^3, ^^4, ^^5, ^^6], ...,[^^9, ^^10, ^^11, ^^12]. Steps 2: For a first chunk:1. Perform a forward pass for 4 time steps.2. Compute the loss at the end of the chunk.3. Backpropagate the error through these 4 time steps.4. Apply optimiser to update weights.Step 3: Move to the next chunk and repeat the process, using the final hidden state of theprevious chunk as the initial state for the next chunk.This approach can facilitate the efficient learning from sequences by reducing the computationalcost and preventing issues with vanishing or exploding gradients.However, one challenge in the field of image and video compression is that sequences are long.For example in a video sequence running at 30fps, large parts of an image may be unchangedfor 1, 2, 3 minutes or more at a time, which means temporal information and correlations canstretch across sequences of thousands of frames or more. For example, if a background of ascene remains unchanged for 3 minutes, that information is redundant and in an ideal worldwould be compressed away and not transmitted in a bit stream. Increasing the GOP lengthallows for more of this temporal information to be harnessed to facilitate better compressionrates by removal of redundant information. For the networks to learn how to use temporalinformation across long GOP sizes, one approach is to increase GOP sizes in training data.However, this results in a very significant compute overhead and training time increase thatis expensive and / or not practical, even when using truncated backpropagation through timewith a sliding window of overlapping chunks. For example, consider a GOP frame sequenceof 60 frames long [^^0, ..., ^^59], with a batch size of 1 (i.e. the single 60 frame sequence is asingle batch) and chunk size 10. Applying an overlapping, sliding window of truncated backpropagation through time, we have to perform gradient calculations for all of the 51 followingchunks of the single batch:[^^0, ..., ^^9], [^^1, ..., ^^10], [^^2, ..., ^^11], [^^50, ..., ^^59]This can be contrasted with short GOP lengths, where for example, a GOP of 4 frames longwould only have a single chunk. Accordingly, it can be seen that increasing GOP size duringtraining and using truncated back propagation through time on its own can result in orders ofmagnitude increase in compute overhead and training time.The present disclosure proposes a solution to this problem by counter-intuitively removing theoverlap of the chunks while performing truncated back propagation through time. For example,considering the same GOP length of 60 frames and batch size 1. By removing the overlap, wenow only have to perform gradient calculations for the 6 following chunks of the single batch:[^^0, ..., ^^9], [^^10, ..., ^^19], [^^20, ..., ^^29], [^^50, ..., ^^59] This approach is counter-intuitive because one advantage typically associated with truncatedback propagation through time is that the overlap in the sliding window helps to capturetemporal relationships across all time steps. The present inventors have realised that, inthe context of image and video compression, removing this overlap does not result in anysignificant drops of performance that might intuitively be expected by less temporal overlap.Accordingly, training performed using this approach is 5-10x faster than using traditionaltruncated back propagation through time, while at the same time being just as stable andresulting in comparable final loss levels. As a result, the present disclosure facilitates theuse of long GOPs during training, which contributes to overall reduction in artefacts such asunrolling, as is described above.More generally, the above approach may be illustrated with the following pseudocode:

[0003] Algorithm 2 Truncated back propagation through time without overlapRequire: Training dataset D, learning rate ^^, regularization parameter ^^, number of epochs ^^Initialize network parameters ^^ randomlyfor ^^^^^^^^ℎ = 1 to ^^ dofor each batch in D^^^^^^ doChunk batch into non-overlapping chunks (^^) of image pairs ^^^^ and ^^^^−1for each chunk ^^ in ^^ doForward pass to compute predictions ^^^^,^^ = ^^^^ (^^^^,^^ , ^^^^,^^−1)Compute ^^^^^^^^^^ Backward pass to compute gradients ∇^^ ^^^^^^^^^^Accumulate gradients in ∇^^ ^^^^^^^^^^end forWith accumulated gradients, update parameters with optimizer O: ^^ ← O(^^, ∇^^ ^^^^^^^^^^ , ^^)end forOptionally evaluate on validation setend forThat is, initially, the network parameters ^^ are set to random values. Within each epoch, thetraining dataset, D, is divided into batches. These batches are processed one at a time. Thedivision into batches helps manage memory efficiently and can speed up the training process.Each batch from the dataset is further divided into non-overlapping chunks ^^ that togethercover a number of time steps ^^, each containing a plurality of image pairs ^^^^ and ^^^^−1, forexample images from a first point in time and a second point in time across the total number oftime steps ^^, whereby when compressing and reconstructing one of those images during theforward pass, the temporal redundancy between the two images is used (e.g. by applying aflow-residual approach as described above in the introductory section). For each chunk, ^^, aforward pass is performed to reconstruct the images of the sequence, ^^^^, based on the currentparameters, ^^, and the input images of the sequence, e.g. using all the pairs ^^^^,^^ and ^^^^,^^−1 inthe chunk ^^. The loss, ^^^^^^^^^^, for the chunk is computed. This loss measures between the predicted values and the actual target values in the chunk and may be, for examplean MSE loss or any other loss. A backward pass is then executed to calculate the gradients ofchunk ^^ of the loss with respect to the network parameters, ∇^^^^^^^^^^^^. The calculated gradientsare accumulated into an overall accumulated gradient ∇^^^^^^^^^^^^ for the batch. Only once allchunks ^^ in the batch have been processed and the gradients accumulated are the networkparameters are updated. This update may be performed using an optimizer, O, which adjuststhe parameters, ^^, based on the accumulated gradients, the learning rate, ^^, and optionallyother factors, for example regularization term. Finally, optionally, at the end of each epoch,validation may be performed using a validation dataset. For completeness, the final hiddenstate of the first chunk once completed is stored and used as the initial hidden state from whichthe forward pass of the second chunk is started. The same process is repeated for the secondchunk, the third chunk, and the fourth chunk to complete the batch.In a further generalisation of the above approach, the window size need not be static across abatch. Instead it can vary across a sequence of frames of a GOP. Consider for example, thesame GOP length of 60 frames and batch size 1 e.g. [^^0, ..., ^^59]. Instead of setting staticwindow or chunk length as described above, the chunk sizes can be dynamically adjusted asthe forward passes progress.For example, consider a 60 frame GOP with a starting chunk size of 4. That is:^^ℎ^^^^^^1 = [^^0, ..., ^^3]The forward pass is performed for ^^ℎ^^^^^^1, the distortion scores (e.g. MSE) are calculatedfor each of the frames of ^^ℎ^^^^^^1. If the distortion scores are indicative of the appearanceof artefacts (such as unrolling artefacts) in the reconstructed frames, the chunk size for thenext chunk, ^^ℎ^^^^^^2, can be reduced or kept the same. Conversely, if the distortion scoresare not indicative of the appearance of artefacts in the reconstructed frames, the chunk sizecan be increased for the next chunk. Checking distortion scores for signs of artefacts maybe performed by applying a threshold to the distortion score. For example, if the distortionis above the threshold it is indicative that the reconstructed image is too different to theground truth image and thus that artefacts are present. Conversely, if they are at or belowthe threshold, the reconstructed image is sufficiently similar to the ground truth image andno or only minor artefacts are present. More complex methods of detecting artefacts mayadditionally or alternatively be applied. For example, by comparing image statistics of thereconstructed image with the ground truth image, by applying a classifier, and so on.In the above example a chunk size of 4 is relatively small so it is unlikely that there will beunrolling artefacts. Accordingly, the chunk size of ^^ℎ^^^^^^2 is increased to e.g. 6. That is:^^ℎ^^^^^^2 = [^^4, ..., ^^9]The forward pass is then performed for ^^ℎ^^^^^^2, the distortion scores are calculated for eachof the frames of ^^ℎ^^^^^^2 and checked for signs of artefacts. If there are no signs of artefacts,the chunk size for ^^ℎ^^^^^^3 may be increased again. For the sake of illustration, we increase itto e.g. 27. That is:^^ℎ^^^^^^3 = [^^10, ..., ^^36]As a chunk size of 27 is relatively large, unrolling artefacts are likely to appear and this will beapparent from the increase in distortion scores in the reconstructed frames as the forward passprogresses through the sequence of 30 frames. In this case, the chunk size for the next chunk^^ℎ^^^^^^4 can be decreased, for example to 23. That is:^^ℎ^^^^^^4 = [^^37, ..., ^^59]A chunk size of 23 is also relatively large, but smaller than 27, so it is expected that thedistortion scores will be less than for the chunk size of 27. If the decrease is enough, thedistortion scores may fall below the threshold or meet whatever criteria of the image statisticsand / or classifier being used to detect distortions. Thus, for the single GOP of 60 frames, anoptimum chunk size has been identified at which temporal data from longer GOPs is optimallycaptured at just the threshold chunk length at which artefacts such as unrolling artefactsare unlikely appear. This adaptive chunk size approach accordingly facilitates the use oftruncated backpropagation through time in a way that is optimal for AI-based image and videocompression.The above example is generalised in the following pseudocode:

[0004] Algorithm 3 Truncated back propagation through time with adaptive chunk sizeRequire: Training dataset D, learning rate ^^, regularization parameter ^^, number of epochs ^^Initialize network parameters ^^ randomlyfor ^^^^^^^^ℎ = 1 to ^^ dofor each batch in D^^^^^^ doChunk batch into non-overlapping chunks (^^) of size ^^ of image pairs ^^^^ and ^^^^−1for each chunk ^^ in ^^ doForward pass to compute predictions ^^^^,^^ = ^^^^ (^^^^,^^ , ^^^^,^^−1)Compute ^^^^^^^^^^ Backward pass to compute gradients ∇^^ ^^^^^^^^^^Accumulate gradients in ∇^^ ^^^^^^^^^^Identify presence of artefacts in ^^^^,^^ and update ^^Update ^^ end forWith accumulated gradients, update parameters with optimizer O: ^^ ← O(^^, ∇^^ ^^^^^^^^^^ , ^^)end forOptionally evaluate on validation setend forIn an additional or alternative approach to adaptive chunk sizes is to introduce a relativedistortion check at some arbitrary point or points in the sequence (e.g. predetermined points orpoints determined dynamically). That is, we take some reference frame ^^^^ , for example anI-frame or a frame at the start of a chunk, we measure the distortion loss (e.g. MSE) at thatpoint, and then when we get to frame ^^^^+^^ we measure the distortion loss again and compare itto the previously measured distortion for the reference frame ^^^^ . If the new distortion is someamount greater than the previously measured distortion (e.g. some percentage, for example1-25%, such as 5%, 10%, or 20%), then it is indicative that the distortion is going up too fastand accordingly that we should cut the chunk short there and proceed to backpropagate, ratherthan waiting until the end of all the frames in that chunk.At a general level, this approach effectively immediately cuts short any chunks when we seethe distortion increasing too much, and adapts the next chunk size accordingly based on wherethe current chunk was cut short, rather than waiting for the end of the chunk to adapt the sizeof the next chunk. That is, by measuring a relative distortion of frames of the chunk comparedto a reference frame, it is possible to track changes in the distortion at an arbitrary level ofgranularity (e.g. frame to frame, every 5 frames, every 10 frames, and so on), thus allowingfor the dynamic updating of chunk sizes during training.The above approaches approaches also open up training (and finetuning) on substantiallylonger GOP size data than has been possible up to now and on long GOP data of all types (e.g.webcam call scenes, high motion scenes, low motion scenes etc) without a priori knowledgeof that data. This is because the adaptive window facilitates the automatic finding optimalchunk sizes for truncated back propogation through time, unique for each GOP and data type -something which is burdensome using handcrafted chunk sizes or static chunk sizes across anentire GOP.As alluded to above, the presently described truncated backpropagation through time approachesmay be used with concept 1, whereby the frozen patch approach may be applied to a GOPbefore that GOP is chunked into non-overlapping chunks as described above. Finally, bothof the above approaches may be jointly or individually combined with the concept describedbelow.Concept 3: Long GOP finetuningA further problem of training on long GOP training data is that the training signal strength andtraining stability can decrease as GOP length increases. This can be attributed, at least in part,due to errors and artefacts propagating temporally across the sequence as each subsequentframe relies on information from earlier frames that can include those errors and artefacts,resulting in worse MSE losses and unpredictable loss curve behaviours.The present inventors have realised that this problem can be partially mitigated by performingbase training on short GOP lengths which are associated with stronger training signal strengthand stability and, once base training is completed (e.g. once loss is below a predeterminedthreshold or has plateaued or after a predetermined number of training steps), finetuning onlonger GOP length training data. This approach allows the networks of the AI-compressionpipeline to learn the easier task of short GOP compression first without the issues of long GOPtraining, resulting in the overall loss moving roughly in the direction of a global minimum.Then, after base training is complete, the finetuning on long GOP training data allows thealready-performant networks to learn the more difficult task of long GOP compression, whichmoves the loss even closer to the global minimum for long GOP compression.As above, the term short GOP may be understood to be GOP lengths of below substanially 10frames, while 10 frames or above may be understood as a long GOP.The above approach may be illustrated using the following pseudocode:Algorithm 4 Neural Network Training with Long GOP Fine-tuningRequire: Base training dataset D^^^^^^^^, fine-tuning dataset D ^^ ^^^^^^, learning rate ^^, regularization parameter ^^,number of epochs ^^^^^^^^^^, number of fine-tuning epochs ^^ ^^ ^^^^^^Initialize network parameters ^^ randomly⊲ Base Training Loopfor ^^^^^^^^ℎ = 1 to ^^^^^^^^^^ dofor each batch (^^, ^^) in D^^^^^^^^ doForward pass to compute predictions ^^ = ^^^^ (^^)Compute base loss ^^^^^^^^^^^^^^^^Backward pass to compute gradients ∇^^ ^^^^^^^^^^^^^^^^Update parameters with optimizer O^^^^^^^^: ^^ ← O^^^^^^^^ (^^, ∇^^ ^^^^^^^^^^^^^^^^, ^^)end forOptionally evaluate on a base validation setend for⊲ Fine-tuning Loopfor ^^^^^^^^ℎ = 1 to ^^ ^^ ^^^^^^ dofor each batch (^^, ^^) in D ^^ ^^^^^^ doForward pass to compute predictions ^^ = ^^^^ (^^)Compute fine-tuning loss ^^^^^^^^ ^^ ^^^^^^Backward pass to compute gradients ∇^^ ^^^^^^^^ ^^ ^^^^^^Update parameters with optimizer O ^^ ^^^^^^: ^^ ← O ^^ ^^^^^^ (^^, ∇^^ ^^^^^^^^ ^^ ^^^^^^, ^^)end forOptionally evaluate on a fine-tuning validation setend forThat is, for the base training phase, the network’s parameters ^^ are initialised randomly and thebase, short GOP dataset, D^^^^^^^^ is provided. For each epoch, up to a predetermined number ofepochs ^^^^^^^^^^, the network processes batches of data from D^^^^^^^^. In each batch, the networkperforms a forward pass to compute predictions based on its current parameters. It thencalculates a loss, ^^^^^^^^^^^^^^^^. A backward pass is performed to compute gradients of the losswith respect to the parameters. The parameters are updated using an optimization algorithm,O^^^^^^^^, which applies these adjustments in a direction expected to reduce the loss, using thelearning rate ^^ as a step size. After each epoch, the network can be optionally evaluated on avalidation set to monitor its performance on unseen data. After base training on short GOPdata, the networks of the compression pipeline ^^^^ are able to compress short GOP sequenceswell but will produce errors and artefacts, such as unrolling, on long GOP sequences.To mitigate this, the network undergoes fine-tuning on long GOP data. This phase adjusts thenetwork to perform well on a specific long GOP compression task using long GOP trainingdata, D ^^ ^^^^^^, The fine-tuning process also proceeds in epochs, for a number of epochs ^^ ^^ ^^^^^^.The network processes batches from D ^^ ^^^^^^ in a manner similar to the base training phase,performing forward passes, computing a fine-tuning loss ^^^^^^^^ ^^ ^^^^^^, and then performingbackward passes to adjust the parameters. The parameters are updated, either using the sameoptimiser as the base training or optionally using a different optimizer, O ^^ ^^^^^^. Finally, similarto the base training phase, the network can be evaluated on a separate validation set specific tothe fine-tuning dataset. The number of training steps may be based on epochs or may insteadbe set as a predetermined number of steps or based on some stopping criteria such as losscurve plateau or some other stop criteria.Counter-intuitively, the present inventors have found that finetuning on long GOP data doesnot result in a significant drop in performance on short GOP sequences during validationeven though the performance on long GOP sequences significantly increases (e.g. in terms ofrate and distortion scores in e.g. bits per pixel and / or MSE scores). It is understood that thisphenomena is likely due to the tasks of short GOP and long GOP compression being similar innature. Thus, finetuning does not result in the base training model suffering from catastrophicforgetting as might be the case if the finetuning task was significantly different to the basetraining task.Figure 7a illustrates a plot 700 of distortion (MSE) score of compressed and reconstructedframes compared to corresponding ground truth frames of a sequence that has been compressedand reconstructed by an AI-based compression pipeline that has been base-trained on shortGOP data. The figure shows a long GOP data series 701 made up of a single sequence of thesame 150 frames as a single GOP. It is apparent that the distortion increases as the sequenceprogresses on the long GOP series 701 after the initial I frame. It is also noted that I-framescannot take advantage of temporal redundancies so use up more bits than P-frames to compressresulting in the initial MSE spike at frame 0 (i.e. the I frame). It can be seen that fromthe initial MSE spike at frame 0 of around 40, the MSE hits 40 again around frame 50 andcontinues to go extend higher and higher reaching MSE scores of around 55-60 at frame 150.Figure 7b illustrates the same plot 700 but now after finetuning on long GOP data. The longGOP data series 701 in Figure 7b has the same initial MSE spike to around 40 for the I framebefore dropping as the P-frames progress. However, this time the MSE scores only hit 40 againby frame 150 (contrasting to Figure 7a where the MSE scores already hit 40 by 50 frame).This indicates an improvement of temporal stability in sequences of roughly 3x the length.Accordingly, Figure 7b illustrates that finetuning on long GOP data after base training on shortGOP data produces networks of an AI-compression pipeline that achieve better bit rates whilemaintaining low distortion across longer GOP data during inference.For completeness, the finetuning of the networks of the pipeline associated with Figure 7b maybe performed using GOPs of whatever GOP length the finetuned networks are intended toused on. For example, the inference toy example of Figure 7b is performed at GOP length 150,accordingly it is envisaged that the finetune data comprises GOP lengths of 150. Extendingthis principle generally, including a distribution of different GOP lengths in the finetune datacan be used to produce a pipeline that is able to perform well across said variety of GOPlengths. More specifically, it has been found that performing the base training on the easy taskof short GOPs produces a base model that is particularly suited to subsequent finetuning onlonger GOPs and results in a pipeline that outperforms a model where the base training datacomprises a full, complete distribution of GOP sizes.Concept 4: Flying-box data augmentationAs described above in connection with concept 1, AI-based compression pipelines frequentlystruggle to handle high motion scenes well. That is, they struggle to achieve compression ratescompetitive with traditional codecs and assign a disproportionately large number of bits to thecompression of high motion scenes, and / or they struggle to achieve reconstruct the framesof the scene with high accuracy, resulting in poor distortion scores. Concept 4 is directedto a data augmentation technique that has been found to significantly improve the ability ofAI-based compression pipelines to learn how to handle high motion scenes, resulting in bettercompression rates and distortion scores. It is noted for completeness that improving the abilityto learn may be understood as meaning that, for a given AI-based compression pipeline, therate and distortion losses after training converge to lower values than that pipeline is able toachieve without using the methods of concept 4.Figure 8 illustrates a sequence of frames 800 of an image sequence that may be part of atraining data set for training an AI-based compression pipeline, for example the pipeline ofFigure 3. The sequence of frames 800 contains high motion features: a figure 801 movesfrom one side of the scene to the other side of the scene within two frames. This kind ofhigh motion feature is difficult for AI-based compression pipelines to learn to handle well.Whilst it is in principle possible to bias a training data set with a higher proportion of this kindof high motion scene data, the inventors have found that an AI-based compression pipelinethat is the result of training on a high-motion biased training data set will improve its ratedistortion scores on such high-motion scenes but at the expense of very poor performance onstatic scenes that were previously not a challenge.The inventors have instead realised that taking a balanced training data sets (i.g. a data set thatis not specifically biased towards a higher proportion of high motion scenes) and modifyingsome or all of the scenes to include a synthetically introduced box or patch that moves acrossthe scene, for example gradually at some predetermined speed in some predetermined directionor path, or for example appearing and disappearing at different places in the frames, has theeffect of improving high motion scene rate distortion scores when the data set is used fortraining, without negatively impacting the low motion scene rate distortion scores in the sameway that biasing the balanced training data set towards a higher proportion of natively highmotion scenes. The term box is used for ease of understanding but it will be appreciated thatany shape of any size may be used.In general terms, it is understood that the synthetic enhancement of the training data thatintroduces "flying" boxes into the scenes but leaves the scenes otherwise unmodified, facilitatesthe learning of high motion features (i.e. the flying box is an "easy-to-learn", high motionfeature) in scenes which are otherwise low motion. For example, an almost perfectly staticnature panorama scene is very low motion. A "flying" box that starts in one corner of the sceneand moves to another corner at some predetermined speed as the frame sequence progressesis a very simple, but high motion feature effectively super imposed onto the static naturescene. The resulting scene thus simultaneously comprises easy-to-learn low motion features(the pixel distributions of an almost unchanging static nature scene), and easy-to-learn highmotion features (the pixel distributions of a simple, synthetic "flying" box). This syntheticallyenhanced scene can be contrasted with native, high motion training data scenes (e.g. anaction scene from a movie with explosions and debris) where the high motion features aresignificantly more complex and difficult to learn, and do not frequently appear at same time asthe kind of pixel distributions that are present in low motion scenes. That is, the distributionsof pixels (and any associated optical flow information) in natively high motion scene aredistinct from distributions of pixels and optical flow information in natively low motion scenes.Biasing a training data set towards the natively high motion scenes results in the distributionsof the low motion scenes becoming under-represented, resulting in reduced performance onsuch scenes. This problem is at least partially solved through the introduction of a synthetic"flying" box across the training data set because, unlike biasing the data set with natively highmotion scenes, the "flying" box does not significantly alter the underlying distribution of thelow motion scenes onto which the "flying" box is super-imposed, thus ensuring that these typesof distributions do not become under-represented in the training data. Thus, the propertiesof the "flying" boxes (e.g. their starting positions, movement speed and so on) may be basedbe based on the balance of motion in the underlying training data set. For example, if theunderlying training data set has few high motion image sequences (for example, optical flowvalues across the data set are in aggregate below some threshold), then it may be advantageousto increase the movement speed of the "flying" boxes, and vice versa.When the neural networks of an AI-based compression pipeline are then trained on such a"flying" box enhanced training data set, the inventors have found that the trained network hasimproved rate distortion scores on high motion scenes (compared to training on a non-enhanceddata set) without significantly negatively impacting rate distortion scores on low motion scenes.This can be contrasted with simply biasing a training data set with more natively high motionscenes which the inventors have found in many cases catastrophically reduces performance onlow motion scenes. Thus training data augmentation using "flying" boxes allows the networksto better learn how to simultaneously handle both high and low motion scenes.It will be appreciated that the "flying" box enhancement may be introduced into some orall of the image sequences of a training data set. For example, it is envisaged that it maybe introduced only on scenes where some motion requirement is considered to be met. Forexample, where motion is below a threshold motion amount based on optical flow or someother pixel difference metric. Or it may be introduced based on some other requirement(e.g. predetermined scene type such as screen capture content, action content, or some othercontent), or indeed it may be introduced on in all scenes in the training data set.Figure 9 illustrates an image sequence 900 that may be used in a training data set showing astatic nature scene into which a "flying" box has been introduced.It will be appreciated that Figures 8 and 9, for illustrative purposes, show only three framesof the respective sequences but that in practice each image sequence may GOP sizes of anynumber and the "flying" boxes (or indeed other shapes) may be introduced onto some or all ofthe frames of a given GOP.An illustrative method of adding "flying" boxes, in this case black pixel rectangles, to asequence of images is set out in the pseudocode below:

[0005] Algorithm 5 Enhance training image sequence with flying boxInputs: sequence, rectangle_width, rectangle_height, start_x, start_y, speed_x, speed_yOutputs: output_sequence Variables: current_x, current_y, new_framecurrent_x← start_x current_y← start_y for each frame in sequence donew_frame ← copy(frame) for y from current_y to (current_y + rectangle_height) dofor x from current_x to (current_x + rectangle_width) doif x is within frame width and y is within frame height thenset_pixel(new_frame, x, y, color.BLACK)end ifend forend forAdd new_frame to output_sequencecurrent_x ← current_x + speed_xcurrent_y ← current_y + speed_yend forReturn: output_sequenceThat is, an input image sequence is provided, and the flying box width, height, start x and ypixel coordinate position and movement speed in each direction is specified. Then, for eachframe in the input image sequence, the frame is copied and the pixels at the position where theflying box is set to be are set to black, or some other colour, to introduce the flying box intothe frame. The position of where the flying box will be in the next frame is then determinedby considering the current positions and speeds, and then the process is repeated for all theframes of the sequence.A further optional enhancement of the above methods is the introduction of a non-uniformtexture or pattern into the pixels of the flying box. The inventors have unexpectedly found thatintroducing image features that AI-based compression pipelines in many cases have difficultywith into the pixels of the flying boxes results in the same effect of improving the rate distortionscores of the pipeline on frames with those features importantly without negatively impactingrate distortion scores associated with other types of scene features. One specific example ofthis effect is the introduction of text (e.g. alpha numeric or other characters, typed or handwritten) into the flying boxes. In a similar manner as described above, it is understood thatsynthetically introducing such features onto scenes where such features would not typicallybe expected to be found has the effect of allowing the networks to learn such to effectivelycompress and reconstruct such features without negatively impacting the networks’ abilitiesto effectively compress and reconstruct the features of the rest of the scene onto which theflying box is introduced. That is, in general terms the training data set substantially retainswhatever balance of pixel distributions it originally had while the the introduction of thetextured or patterned flying box becomes an easy-to-learn feature superimposed on those pixeldistributions.Figure 10 illustrates an example image sequence 1000 onto which a flying box 1001 containingalpha numeric characters has been introduced onto a scene, in this case a high motion scene ofa rocket during lift off as it moves across the scene. As explained above, the pixel distributionof the majority of the scene remains consistent with the natively captured footage of the scenewhile super imposing the "flying" box with text onto the scene introduces an easy-to-learnhigh motion feature, as well as an easy to learn alpha-numeric feature. Training an AI-basedcompression pipeline on such data results in a set of networks that have improved rate distortionscores on high motion scenes, low motion scenes, and scenes that include alphanumeric textcompared to networks trained on the same data set but without such enhancements.The alpha-numeric character example has been found to be particularly effective at improvingthe rate distortion performance of AI-based compression pipelines on screen capture imagesequences such as screen sharing by participants in a video call, and such as video gamestreamers who may quickly switch between browser screens and have various overlays ontheir streams, many of which contain text, without significantly reducing the rate distortionperformance of the pipeline on other types of content such as movies and so on.In very general terms, the inventors have found that applying the "flying" box methodologyto training data enhancement may be seen as a method of blending specific training datatypes (e.g. certain pixel distributions associated with certain scene and content types, and ormotions associated therewith) into an underlying, balanced training dataset without negativelyeffecting the balance of that underlying training dataset. Thus the networks trained thereon,retain the good across-the-board performance of the balanced training data while boosting theperformance on the specific data types blended into it using the "flying" boxes.In the specific case of video compression, this effect is particularly advantageous becauseit allows a single model (i.e. the weights and biases and so on of the neural networks ofthe pipeline) to be performant across a much higher number of different use cases, therebyreducing the memory requirements of the pipeline when it is deployed as a codec on edgedevices. This is a significantly different approach to the paradigm of using multiple differentmodels (i.e. sets of weights) and selecting that model which is finetuned specifically for goodperformance on a given image sequence type.In a further optional enhancement, it is envisaged that the texture of the "flying" box may besourced from other image sequences in the underlying training data set. That is, rather thansynthetically creating the pixel values of the "flying" box, they may instead be sampled, forexample randomly, from image patches of the training data set into which the "flying" boxesare being introduced. That is, for a given image sequence A, a patch of pixels from one ormore frames of image sequence B may be sampled and used as the texture of the "flying" boxadded to image sequence A. This approach increases the variance of the pixel distribution ofthe "flying" box, which can help to prevent overfitting that may occur when solid colour block"flying" boxes are used.In a further optional enhancement, it is envisaged that the motion of the "flying" box acrossthe frames of the sequence may be different between sequences. That is, it is envisaged that"flying" boxes together having a distribution of different motion amounts, directions andspeeds may be introduced into the sequences to increase the variance of the pixel distributionsattributed to the "flying" boxes. As above, this can help to prevent overfitting that may occurwhen all "flying" boxes across all sequences have the same or similar motion.In general terms, training the neural networks of an AI-based compression pipeline may bedefined by the following pseudocode:Algorithm 6 Training AI-based compression pipeline using flying box data augmentationInputs: Training dataset X, learning rate ^^, regularization parameter ^^, number of epochs ^^ ,network architecture ^^^^Initialize network parameters ^^for epoch = 1 to ^^ dofor each batch in X doAdd flying boxes to (^^^^−1, ^^^^ ) and perform forward pass to compute predictions ^^^^ = ^^^^ (^^^^−1, ^^^^ )Compute loss Loss = ^^ + ^^^^Backward pass to compute gradients ∇^^LossUpdate parameters with optimizer O: ^^ ← O(^^, ∇^^Loss, ^^)end forOptionally evaluate on validation setend forThat is, a training data set X, a learning rate ^^, regularisation parameter ^^, and a number oftraining steps or epochs ^^ is selected. The network architecture of ^^^^ is defined, for exampleas shown in Figure 3. The network parameters ^^ are randomly initialised and then the trainingloop is started. For each batch in the training data X, flying boxes are introduced as describedabove, and the forward pass is computed for the network being trained ^^^^ . The total ^^^^^^^^ iscalculated by combining a distortion term ^^ and a rate term ^^, and any other loss terms (notshown). The backwards pass is then performed to compute gradients based on the loss, andthe parameters ^^ are optimised using the optimiser, such as stochastic gradient descent SGD,or some other known optimiser. Optionally, a validation loss can be calculated and the trainingloop is repeated until the predetermined number of steps or epochs ^^ have been calculated,or some other criteria have been reached. The learning rate, batch size, and or number ofepochs may be optimised during training, for example using a learning rate scheduler or someother hyperparameter optimisation method. More generally, the hyperparameters may beoptimised experimentally. It will be appreciated that the flying boxes may be added "on thefly" or pre-computed for a given sequence and added to an input training dataset X beforestarting training so that the training dataset X already contains the flying boxes when trainingcommences, in which case the "add flying boxes" step in the pseudocode above does not needto be performed again.It will be appreciated that the presently decscribed concept may be used in a standalone manner,and / or together with any of the other concepts described herein.Concept 5: Temporal frame sub-samplingAs described above, AI-based compression pipelines in many cases struggle to effectivelycompress high motion scenes and training data augmentations, such as those of concept 4,may be used to improve the ability of the networks of the pipeline to learn to better compresssuch scenes.An alternative or additional approach is to increase motion variance in the image sequencesof the training data during training by varying how far back (or forward) in terms of frameposition in the image sequence the reference frame is taken from. That is, how far separatedfrom the current frame the reference frame for estimating the motion against is. For example,when estimating optical flow, the optical flow is not necessarily always estimated between^^^^ and ^^^^−1, but instead, for each current frame ^^^^ , the optical flow may instead be estimatedby comparing ^^^^ with ^^^^−^^ where ^^ may be sampled from ^^ ∈ [1, 2, 3, ...]. The effect of thisapproach is that the networks of the pipeline learn better how to predict larger optical flows asthe pixel differences between a frame at ^^ and for example ^^ − 5 are typically greater than thosebetween a frame at ^^ and ^^ − 1. The value of ^^ may be randomly sampled from Z using anysuitable method, for example using a random or pseudo-random number generator, where suchsampling may be drawn from a parameterised distribution to allow for some control over thevalue of ^^, and / or ^^ may be sampled according to some predetermined heuristics (e.g. one ormore predetermined rules). For example, certain scene types may result in lower loss valueswhen ^^ is biased to lower values compared to higher values, and so on and accordingly thesampling of ^^ may take this into account by being biased towards lower ^^ values (for example byparameterising a distribution of values from which ^^ may be selected using a mean, varianceand / or other parameters), or vice versa. Example scene types include high motion actionscenes of movies, low motion nature scenes from documentaries, and so on. In any event,the result is that there is a distribution of different frame positions from which the referenceframe is selected, and thus a greater distribution of motion amounts in the training data whileotherwise leaving the pixel distributions of the underlying data unchanged. When adoptingthis approach, the inventors have found that, for a given AI-based compression pipeline, therate and distortion losses after training converge to lower values than that pipeline is able toachieve without using the methods of concept 5.More generally, the higher the frame gap (i.e. where ^^ is large), the higher the pixel differencesbetween ^^^^ and ^^^^−^^ are and thus the higher the motion that the network has to try to learnduring training to minimise the rate distortion for. As a result, this approach has the effectof synthetically introducing motion into the training data set without otherwise changing theunderlying pixel distributions of the training data. That is, it facilitates the learning of highmotion without otherwise biasing an already balanced training data set with an increase ofadditional natively high motion image sequences, which in some cases has the effect of reducingperformance on low motion scenes. The result is accordingly a network that has improvedperformance (better rate distortion scores) on low and high motion scenes simultaneously. It isalso envisaged that, the sampling of ^^ (e.g. its biasing towards higher or lower ^^ values) maybe based on the balance of motion in the underlying training data set. For example, if theunderlying training data set has few high motion image sequences (for example, optical flowvalues across the data set are in aggregate below some threshold), then it may be advantageousto bias the sampling of ^^ towards higher values, or vice versa.Figure 11 illustratively shows this approach with reference to a sequence of frames 1100 withframe positions ^^, ^^ − 1, ^^ − 2,...,^^ − ^^. Frame ^^^^ is the current frame of a sequence beingprocessed while ^^^^−^^ is the frame which is being used as a reference frame against whichoptical flow information is being determined (e.g. in a P-frame module of Figure 3).For a current frame of the image sequence being fed into the pipeline e.g. ^^^^ , a random integer^^ is generated and ^^ is set to the value of ^^ and used to select the reference frame ^^^^−^^ againstwhich ^^^^ will be compared.For example, if ^^ = 1, then the reference frame input into the P-frame module of the pipelinewill be ^^^^−1 whereby, if the operation of the networks of the optical flow part of the pipelineare defined as a function ^^, then:^^ ^^^^^^ = ^^(^^^^ , ^^^^−1)Or, for example, if ^^ = 4, then:^^ ^^^^^^ = ^^(^^^^ , ^^^^−4)and so on up to the general case of:^^ ^^^^^^ = ^^(^^^^ , ^^^^−^^)where ^^ ∈ Z. In some implementations ^^ may be randomly generated from a subset of Z, forexample ^^ ∈ [1, 2, 3, 4, 5, ..., 60]. Whilst not shown in Figure 11, it is also envisaged that,when using B-frames, ^^ may take on negative values so that the reference frame is selectedfrom e.g. ^^ + 1, ^^ + 2, ... ^^ + ^^.Once the processing of the current frame ^^^^ is complete, ^^^^+1 becomes the current frame andthe process is repeated for ^^^^+1 and so on until the entire sequence of frames (which may makeup a GOP) has been processed.An example training implementation is illustrated in the pseudocode provided below.Algorithm 7 Training AI-based compression pipeline using temporal frame sub-samplingInputs: Training dataset X, learning rate ^^, regularization parameter ^^, number of epochs ^^ ,network architecture ^^^^Initialize network parameters ^^for epoch = 1 to ^^ dofor each batch in X doGenerate ^^ ∈ Z and perform forward pass to compute predictions ^^^^ = ^^^^ (^^^^−^^ , ^^^^ )Compute loss Loss = ^^ + ^^^^Backward pass to compute gradients ∇^^LossUpdate parameters with optimizer O: ^^ ← O(^^, ∇^^Loss, ^^)end forOptionally evaluate on validation setend forThat is, a training data set X, a learning rate ^^, regularisation parameter ^^, and a number oftraining steps or epochs ^^ is selected. The network architecture of ^^^^ is defined, for exampleas shown in Figure 3. The network parameters ^^ are randomly initialised and then the trainingloop is started. For each batch in the training data X, an ^^ value is generated and used to selectthe reference frame ^^^^−^^, and the forward pass is computed for the network being trained ^^^^ .The total ^^^^^^^^ is calculated by combining a distortion term ^^ and a rate term ^^, and anyother loss terms (not shown). The backwards pass is then performed to compute gradientsbased on the loss, and the parameters ^^ are optimised using the optimiser, such as stochasticgradient descent SGD, or some other known optimiser. Optionally, a validation loss canbe calculated and the training loop is repeated until the predetermined number of steps orepochs ^^ have been calculated, or some other criteria have been reached. The learning rate,batch size, and or number of epochs may be optimised during training, for example using alearning rate scheduler or some other hyperparameter optimisation method. More generally,the hyperparameters may be optimised experimentally.It will be appreciated that the presently decscribed concept may be used in a standalone manner,and / or together with any of the other concepts described herein.Concept 6: Flow teacher ensembleAs has been described generally above, successfully training the networks of an AI-basedcompression pipeline, such as that of Figure 3, to perform well simultaneously on high andlow motion image sequences can be challenging. One area of focus of such training is thetraining of the flow part of the P-frame module of such a pipeline. By way of illustrativeexample, this may refer to the training of one or more of the networks of the flow module206 in Figure 3. As described generally above, the distortion term(s) of a rate distortion lossfunction Loss = ^^ +^^^^ are typically based on a difference between a ground truth input frame,for example ^^^^ and the reconstructed version of that frame ^^^^ , that is:Loss = ^^ (^^^^ , ^^^^) + ^^^^When considering the training of the networks of the flow module 206 specifically, the lossfunction may incorporate one or more additional terms, that are based on a difference betweensome ground truth optical flow information ^^^^^^ , and the optical flow information ^^ that a givenforward pass of the networks produce during a training step. For example:Loss = ^^ (^^^^ , ^^^^) + ^^( ^^^^^^ (^^^^ , ^^^^−1), ^^ (^^^^ , ^^^^−1)) + ^^^^Whilst not shown for ^^ and ^^, each of the loss terms may be regularised relative to each otherusing a regularisation parameter such as ^^.Unlike in the case of the basic distortion loss term where the ground truth may simply be theinput image ^^^^ , ground truth optical flow information ^^^^^^ is not readily available and must beestimated first from the input image ^^^^ and a reference frame ^^^^−1. However, there are manydifferent methods of estimating optical flow, each with their own strengths and weaknesses,and each producing different optical flow maps. There is accordingly not necessarily any wayto determine a single ground truth of optical flow information in the same way that there isalways a ground truth input image ^^^^ that can be accessed. There are instead only estimates ofground truth optical flow information that are highly dependent on the estimation method oralgorithm used. Some exemplary optical flow estimation techniques and their strengths andweaknesses are set out below.Lucas-Kanade algorithm - this algorithm relies on the assumption that the optical flow is constantwithin a small neighborhood of each pixel. The method solves the optical flow equations byperforming a least-squares fit of the local image velocities, making it computationally efficientand relatively simple to implement. It performs well in regions with high-texture contentand small movements. However, it struggles with large displacements between images and issensitive to image noise due to its reliance on intensity gradients. Additionally, it does nothandle occlusions well, where pixels may appear or disappear between frames.Horn-Schunck algorithm - this algorithm takes a global approach by computing a dense flowfield over the entire image. It minimizes an energy function that includes both a data fidelityterm and a smoothness constraint, resulting in a smooth flow field. The algorithm estimatesmotion for every pixel, providing a comprehensive motion map. While it is mathematicallyrigorous and offers global consistency, it may over-smooth motion boundaries, blurring finedetails. It also requires careful tuning of the smoothness parameter and is more computationallyintensive than local methods like Lucas-Kanade.Farnebäck’s algorithm - this algorithm models the neighborhood of each pixel using quadraticpolynomials, allowing it to calculate flow in terms of displacement fields. It provides denseflow fields and is relatively fast. However, the polynomial approximation may not capturecomplex motions accurately, and the method requires careful selection of window sizes andother parameters to function optimally.TV-L1 optical flow algorithm - this algorithm introduces robustness to noise and illuminationchanges by using a total variation regularization term and an L1 norm for the data fidelity term.This variational approach preserves sharp motion boundaries and is less sensitive to outliers,making it effective in handling abrupt motion changes and varying lighting conditions.Brox’s optical flow algorithm - incorporates additional constancy assumptions, such asgradient constancy, to improve performance in challenging scenarios. It is better at capturingsignificant displacements and is more invariant to illumination changes due to its use ofgradient information. While it offers high-quality flow estimation in complex scenes, it issensitive to the choice of regularization parameters.SimpleFlow algorithm - this is a non-parametric algorithm that relies on color matching andspatial consistency rather than image gradients. It can handle large, non-linear motions andis less sensitive to textureless regions. However, it can be slow due to exhaustive searchmechanisms and may miss subtle motions, making it less suitable for applications requiringfine detail.FlowNet - this is a convolutional neural network architecture designed for flow estimation. Itis capable of modeling intricate motions learned from data but may not generalize well to datasignificantly different from its training data.PWC-Net - this is based on FlowNet but is an improvement over FlowNet in that it incorporatesa feature pyramid and warping layers, effectively capturing large motions with better efficiencyand accuracy than FlowNet.SIFT Flow algorithm - this algorithm aligns images through dense matching of Scale-InvariantFeature Transform (SIFT) descriptors. This method is robust to scale, rotation, and illuminationchanges, making it useful for aligning images with similar scenes but different content.Frequency domain methods - these types of methods estimate optical flow by analyzing phaseinformation using Fourier transforms. These methods can effectively capture periodic motionsand are resistant to certain types of noise. However, they may not localize motion accuratelyin spatial terms and struggle with non-linear or non-periodic movements.This list is not intended to be exhaustive and is provided primarily to illustrate that the thereare many different ways to estimate optical flow, each of which may produce different opticalflow information (e.g. a flow map).The inventors have realised that the performance on different scene types of the networks ofAI-based compression pipelines that estimate flow, such as the flow module 206 in Figure 3, iscorrelated with the performance on those scene types of the algorithm used to estimate theground truth flow ^^^^^^ in the loss function. For example, if ^^^^^^ is estimated using a Lucas-Kanaealgorithm implementation which performs well on small local motion but not on large globalmotion then the networks of the flow module 206 after training also perform well on small localmotion but not on large global motion. Conversely, if ^^^^^^ is estimated using a Horn-Schunckalgorithm implementation which performs well on global motion, then the networks of the flowmodule 206 after training will also perform better on global motion but not on local motion.Concept 6 is directed to harnessing these correlations by providing a learnable or criterion-based ensemble flow estimation to use as ^^^^^^ that comprises flow information (e.g. a flow map)produced by one or more of a plurality of different flow algorithms based on whichever oneor combination minimises the loss function or a specific loss term, or based on some otherminimisation criteria. For example, the resulting "ground truth" flow map may be stitchedtogether on a pixel-wise basis from pixels of whichever flow algorithm produced an optimalflow value for each pixel. Where optimal here refers to meeting some criteria such as a smallestflow pixel difference between the ^^ and ^^^^^^ , or minimises the overall loss, or one or more otherspecific loss terms, and so on.For example, consider a set of optical flow algorithms F = {^^1, ^^2, . .. , ^^^^} where n is thenumber of flow estimation algorithms included in the set. During a forward pass of thenetworks of the flow module, an ensemble flow may be calculated by producing a respectiveindividual flow maps using each of the optical flow algorithms in F and combining these, forexample using a weighted sum, or some any other combination By way of illustrativeexample, consider estimation of optical flow information ^^ such as a flow map for a given inputframe ^^^^ and ^^^^−1. We can define an ensemble ^^^^^^^^ of optical flow information estimationswith: ^^ (^^ , ^^∑^^^^^^^^ ^^ ^^−1) = ^^=1 ^^^^^^^^ (^^^^ , ^^^^−1) where ^^^^ is a learnable or criterion-based parameter or parameters, for example in tensor form,indicative of how much each of the different flow estimations produced by the respectiveoptical flow algorithms in F contribute to the values that make up the ensemble flow map ^^^^^^^^.This approach also allows different optical flow algorithms to contribute different amounts todifferent parts of the flow map. For example, if one optical flow algorithm produces opticalflow information for certain pixels or patches of pixels that results in a very low loss during atraining step, the networks may update the weights when applying the optimiser to increase thecontribution of this optical flow algorithm to the optical flow information for such patchesby updating the values of ^^^^ accordingly. As training progresses, the networks learn byupdating ^^^^ which optical flow algorithms are best applied to different pixels of different inputsto minimise the rate distortion loss. This approach effectively results in a learned teacherflow ensemble that forces the networks of the AI-compression pipeline to learn to emulatethe good-performance abilities of the various optical flow algorithms of the ensemble fordifferent scenarios on a pixel-wise basis, while simultaneously learning to not emulate thebad-performance abilities of those algorithms (because that would result in a higher loss andthus training pushes the values of ^^^^ towards minimising any contribution of bad flow valuesto the ensemble flow). Alternatively, ^^^^ may be criterion-based where, for each training stepof the AI-based compression pipeline, the ^^^^^^ is stitched together from flow pixels selectedbased on whichever algorithm produced a flow value for that pixel which met some criterion(e.g. lowest difference between ^^ and ^^^^^^ , and so on). For example, consider a toy examplewhere we consider what flow value to use for pixel (0,0) of a flow map ^^^^^^ , a first optical flowalgorithm produces a flow value for that pixel which is very different from ^^ whereas a secondoptical flow algorithm produces a flow value for that pixel which results in a smaller difference.In this case, the flow value selected for pixel (0,0) will be that produced by the second opticalflow algorithm rather than the first. This is of course only a toy example and in practice, manymore optical flow algorithms may be used and the resulting final flow map may be stitchedtogether from flow values taken from many of such optical flow algorithms.We can thus define our new loss function as:Loss = ^^ (^^^^ , ^^^^) + ^^( ^^^^^^^^ (^^^^ , ^^^^−1), ^^ (^^^^ , ^^^^−1)) + ^^^^ and:Loss = ^^ (^^^^ , ^^^^) + ^^(∑^^^^=1 ^^^^^^^^ (^^^^ , ^^^^−1), ^^ (^^^^ , ^^^^−1)) + ^^^^ where ^^^^ is a learnable or criterion-based parameter(s) that defines the contribution of each ofthe optical flow algorithms ^^^^ in F to the ensemble flow ^^^^^^^^ which acts as a "ground truth" or"teacher" optical flow against which the estimated flow ^^ may be compared when calculatingthe flow-specific loss term ^^ during a training step of the networks of the flow module of anAI-compression pipeline. As above, whilst not shown, each of the loss terms may be regularisedrelative to each other using a regularisation parameter such as ^^. Whilst the example of aweighted sum is used here, it is envisaged that any combination may be used, for example anyparameterised combination based on one or more combination parameters may be used suchas a parameterised linear or non-linear combination, a parameterised power law combination,a parameterised polynomial combination, a parameterised logarithmic combination and so on.This approach is also advantageous from an implementation perspective because it makes itstraightforward to include any number (for example 2, 3, 4, 5 or more) and any type of opticalflow algorithms (for example any of the optical flow algorithms listed above and / or any othersnot listed above) in F without burdensome, manual experimentation around which worksbest because the learnable parameter ^^^^ during training automatically learns how much ofeach optical flow algorithm should contribute to each pixel of a flow map for a given input ^^^^and ^^^^−1 to achieve a given training objective (e.g. minimised rate distortion loss) for a giventraining data set. In very general term, this approach allows the networks to decide on their ownwhich behaviours of which optical flow algorithms the networks of the flow module shouldemulate to achieve the desired minimisation of the rate distortion loss. It should be notedin contrast that a criterion-based approach may require additional work in terms of manualhand-crafting of the criterion (e.g. selecting the pixel flow values that are least different to^^ ), the criterion-based approach is advantageous is that it does not typically training stabilitywhen training the rest of the pipeline - something which may occur when using a learnable ^^^^.An illustrative training implementation is set out in the pseudocode below:

[0006] Algorithm 8 Training AI-based compression pipeline using ensemble flow teacherInputs: Training dataset X, learning rate ^^, regularization parameter ^^, number of epochs ^^ ,network architecture ^^^^ , flow algorithm ensemble FInitialize network parameters ^^for epoch = 1 to ^^ dofor each batch in X doCalculate ^^∑ ^^^^^^ (^^^^ , ^^^^−1) = ^^^^=1 ^^^^^^^^ (^^^^ , ^^^^−1) and perform forward pass to compute predictions^^^^ = ^^^^ (^^^^ , ^^^^−1) and ^^ (^^^^ , ^^^^−1)Compute loss Loss = ^^ (^^^^ , ^^^^−1) + ^^( ^^^^^^^^ , ^^ ) + ^^^^Backward pass to compute gradients ∇^^LossUpdate parameters with optimizer O: ^^, ^^ ← O(^^, ∇^^Loss, ^^)end forOptionally evaluate on validation setend forThat is, a training data set X, a learning rate ^^, regularisation parameter ^^, a number oftraining steps or epochs ^^ , and a set of optical flow estimation algorithms (the flow algorithmensemble) F is selected. The network architecture of ^^^^ is defined, for example as shown inFigure 3. The network parameters ^^ are randomly initialised and then the training loop isstarted. For each batch in the training data X the forward pass is computed for the networkbeing trained ^^^^ . Additionally, the flow ensemble produces an ensemble flow estimation^^^^^^^^ (^^^^ , ^^^^−1) =∑^^^^=1 ^^^^^^^^ (^^^^ , ^^^^−1) which will later be used as a "ground truth" flow whenestimating the loss function. Note that ^^ is a learnable parameter specifying how much of eachflow algorithm contributes to the values of the ensemble flow map and will accordingly beupdated by the optimiser when it is applied. The total ^^^^^^^^ is then calculated by combininga distortion term ^^ and a rate term ^^, and the flow loss term ^^( ^^^^^^^^, ^^ ) which is indiciativeof a difference between the actual flow ^^ estimated by the networks in the current trainingstep and the ensemble flow ^^^^^^^^. The backwards pass is then performed to compute gradientsbased on the loss, and the parameters ^^ and ^^ are optimised using the optimiser, such asstochastic gradient descent SGD, or some other known optimiser. Optionally, a validation losscan be calculated and the training loop is repeated until the predetermined number of steps orepochs ^^ have been calculated, or some other criteria have been reached. The learning rate,batch size, and or number of epochs may be optimised during training, for example using alearning rate scheduler or some other hyperparameter optimisation method. More generally,the hyperparameters may be optimised experimentally.As training progresses, ^^ will converge to a set of values defining the optimal composition offlow algorithms, either in aggregate or on a pixel-wise basis, that achieve a minimised ratedistortion loss. In this way, an optimal flow "teacher" is learned which results in the networksof the flow module 206 better emulating the most desired behaviours from the flow algorithmsin F while avoiding the least desired behaviours. It is envisaged that the optimiser may beapplied at different times to both of ^^ and ^^, and to only one of them. For example, duringearly stages of a training schedule (e.g. for a predetermined number of steps), ^^ may be frozenin its initialisation state, for example, all algorithms contribute equally to the ensemble flow,or only one does. This may increase the initial stability of the training of ^^. Then, after anext predetermined number of training steps, ^^ may be unfrozen, allowing the compositionof the flow ensemble to vary, and thus allowing ^^ and ^^ to be trained together to achieve thetraining objective of minimising the loss function. Other training schedules with freezing andunfreezing of ^^ and ^^ after predetermined number of steps are also envisaged. The specificdetails of the training schedule may be determined empirically during ablation runs.Alternatively, training can be simplified by using a criterion-based selection of the pixels usedto produce ^^^^^^^^, in which case there is no need to update ^^ with the optimiser as this is insteada hard coded criterion-based parameter.Figure 12 illustratively part 1200 of the above method. A plurality of optical flow estimationalgorithms ^^^^ 1201, 1202, 1203 are provided and used to calculate an output optical flowbetween a first frame ^^^^ and a second frame ^^^^−1. Their respective output flow maps areweighted or otherwise emphasised or de-emphasised by learned parameters ^^^^ and combinedinto an ensemble flow map or other representation of optical flow information ^^^^^^^^. Thisprovides a reference "ground truth" flow for use in a flow loss term ^^ of the loss function. Theoptical flow part 206 of the P-frame module, comprising at least an encoder neural network1204 and decoder neural network 1205 also performs a forward pass to produce its ownestimate of optical flow ^^ from the first frame ^^^^ and the second frame ^^^^−1. As this occursduring training, the optical flow ^^ produced by the (initially untrained) flow part 206 of theP-frame module may not be particularly accurate and / or may not be entropy encodable into aparticularly small bitstream size - this is something that is learned during training. The flowloss term ^^ is estimated from the the optical flow ^^ from the forward pass of the untrainednetworks of the flow part 206 and from the ensemble flow ^^^^^^^^. As described above, the flowloss term ^^ may be a simple difference metric such as a mean sqaured error, or other differencemetric and / or it may comprise a discriminator loss estimate, whereby the generator is the flowpart 206 performing its forward pass and the discriminator may comprise a neural networktrained to distinguish ^^ from ^^^^^^^^. The rate term ^^ and the distortion term ^^ may then becalculated in the usual way as described earlier herein to produce the overall loss ^^, on whichthe updates to the weights of the neural networks of the AI-based compression pipeline maybe based.Whilst the above described example describes a learned combination of flow algorithms, thecriterion-based approach is also envisaged as a more easily implemented approach given thetraining instability that a learned ^^ can introduce. That is, rather than a learned parameter thatdetermines how much of each flow algorithm contributes to the flow maps, instead each flowalgorithm computes a flow and then the choice of which teacher to use is made (pixel-wise) byapplying a criterion. For example, whichever teacher minimises one or more loss terms of theloss function or the loss function overall is the one that is selected. In this way, the resultingoptical flow information is still "combined" in the sense that the flow algorithm choice is stilla pixel-wise choice resulting in an overall optical flow map which is the aggregation of allthe individual flow pixels from a chosen algorithm for that pixel. In very general terms thefinal flow map may be said to be stitched together from flow pixels taken from whicheveralgorithm resulted in most closely matching a given criterion (e.g. minimising a specificloss term such as a rate term or distortion term or some other criterion). Thus, the final flowcomprises a weighted combination where one weight is 1 and the others are 0 without anylearned parameters. This implementation is substantially simpler and avoids any traininginstability that a learned combination can in some cases introduce when training the pipeline.It will be appreciated that the presently described concept may be used in a standalone manner,and / or together with any of the other concepts described herein.The subject matter and the functional operations described in this specification can beimplemented in digital electronic circuitry, in tangibly-embodied computer software orfirmware, in computer hardware, including the structures disclosed in this specification andtheir structural equivalents, or in combinations of one or more of them. The subject matterdescribed in this specification can be implemented as one or more computer programs, i.e.,one or more modules of computer program instructions encoded on a tangible non transitoryprogram carrier for execution by, or to control the operation of, data processing apparatus.Alternatively or in addition, the program instructions can be encoded on an artificially generatedpropagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, thatis generated to encode information for transmission to suitable receiver apparatus for executionby a data processing apparatus. The computer storage medium can be a machine-readablestorage device, a machine-readable storage substrate, a random or serial access memory device,or a combination of one or more of them. The computer storage medium is not, however, apropagated signal.The term “data processing apparatus” encompasses all kinds of apparatus, devices, andmachines for processing data, including by way of example a programmable processor, acomputer, or multiple processors or computers. The apparatus can include special purposelogic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specificintegrated circuit). The apparatus can also include, in addition to hardware, code that createsan execution environment for the computer program in question, e.g., code that constitutesprocessor firmware, a protocol stack, a database management system, an operating system, ora combination of one or more of them.A computer program (which may also be referred to or described as a program, software, asoftware application, a module, a software module, a script, or code) can be written in anyform of programming language, including compiled or interpreted languages, or declarative orprocedural languages, and it can be deployed in any form, including as a stand alone program oras a module, component, subroutine, or other unit suitable for use in a computing environment.A computer program may, but need not, correspond to a file in a file system. A program can bestored in a portion of a file that holds other programs or data, e.g., one or more scripts storedin a markup language document, in a single file dedicated to the program in question, or inmultiple coordinated files, e.g., files that store one or more modules, sub programs, or portionsof code. A computer program can be deployed to be executed on one computer or on multiplecomputers that are located at one site or distributed across multiple sites and interconnected bya communication network.The processes and logic flows described in this specification can be performed by one or moreprogrammable computers executing one or more computer programs to perform functionsby operating on input data and generating output. The processes and logic flows can also beperformed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g.,an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).Computers suitable for the execution of a computer program include, by way of example,can be based on general or special purpose microprocessors or both, or any other kind ofcentral processing unit. Generally, a central processing unit will receive instructions and datafrom a read only memory or a random access memory or both. The essential elements ofa computer are a central processing unit for performing or executing instructions and oneor more memory devices for storing instructions and data. Generally, a computer will alsoinclude, or be operatively coupled to receive data from or transfer data to, or both, one or moremass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks.However, a computer need not have such devices. Moreover, a computer can be embedded inanother device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio orvideo player, a VR headset, a game console, a Global Positioning System (GPS) receiver, aserver, a mobile phones, a tablet computer, a notebook computer, a music player, an e-bookreader, a laptop or desktop computer, a PDAs, a smart phone, or other stationary or portabledevices, that includes one or more processors and computer readable media, or a portablestorage device, e.g., a universal serial bus (USB) flash drive, to name just a few.Computer readable media suitable for storing computer program instructions and data includeall forms of non-volatile memory, media and memory devices, including by way of examplesemiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magneticdisks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM andDVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in,special purpose logic circuitry.The subject matter described in this specification can be implemented in a computing systemthat includes a back end component, e.g., as a data server, or that includes a middlewarecomponent, e.g., an application server, or that includes a front end component, e.g., a clientcomputer having a graphical user interface or a Web browser through which a user can interactwith an implementation of the subject matter described in this specification, or any combinationof one or more such back end, middleware, or front end components. The components of thesystem can be interconnected by any form or medium of digital data communication, e.g., acommunication network. Examples of communication networks include a local area network(“LAN”) and a wide area network (“WAN”), e.g., the Internet.The computing system can include clients and servers. A client and server are generally remotefrom each other and typically interact through a communication network. The relationship ofclient and server arises by virtue of computer programs running on the respective computersand having a client-server relationship to each other.While this specification contains many specific implementation details, these should beconstrued as descriptions of features that may be specific to particular examples of particularinventions. Certain features that are described in this specification in the context of separateexamples can also be implemented in combination in a single example. Conversely, variousfeatures that are described in the context of a single example can also be implemented inmultiple examples separately or in any suitable subcombination.Similarly, while operations are depicted in the drawings in a particular order, this should notbe understood as requiring that such operations be performed in the particular order shownor in sequential order, or that all illustrated operations be performed, to achieve desirableresults. In certain circumstances, multitasking and parallel processing may be advantageous.Moreover, the separation of various system modules and components in the examples describedabove should not be understood as requiring such separation in all examples, and it should beunderstood that the described program components and systems can generally be integratedtogether in a single software product or packaged into multiple software products.Finally, for completeness, whilst a short GOP is described herein as a GOP of length under 10frames and a long GOP is a GOP of 10 frames or greater. The advantages of each conceptbecomes significantly more apparent and visible as GOP length increases. For example, at 30frames, 60 frames, 90 frames, 120 frames and so on, training time and stability is significantlyworse without using the presently described concepts. Further, artefacts and errors produced bynetworks trained on long GOPs without using the presently described concepts are significantlymore apparent at GOP lengths of 30, 60, 90, 120, and greater numbers of frames. Accordingly,whilst the present disclosure is presented in the context of the boundary between long andshort GOP terminology is 10 frames, it is also envisaged that the boundary between long andshort GOP terminology may be considered at 30, 60, 90, 120 frames or other frames of theseorders of magnitude or greater.

Claims

CLAIMS1. A method of training one or more neural networks, the one or more neural networks beingfor use in lossy image or video encoding, transmission and decoding, the method comprisingthe steps of:receiving a first sequence of input images at a first computer system;producing a masked sequence of input images by masking a portion of one or more ofthe input images;with a first neural network, for a pair of input images of the masked sequence, producinga latent representation of a difference between the images of the pair;with a second neural network, decoding the latent representation to produce a first outputimage, wherein the first output image is an approximation of one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;evaluating a function based on a difference between the first sequence of output imagesand the first sequence of input images;updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps for a plurality of further sequences of input images to producefirst and second trained neural networks.

2. The method of claim 1, wherein masking comprises selecting a portion of a first imageof the first sequence, and replacing pixels in the other images of the first sequence using theselected portion.

3. The method of claim 2, wherein an unselected portion of the first image of the first sequenceremains unmasked in the first sequence of input images.

4. The method of claim 3, wherein a position and size of the portion for first sequence isdifferent to the position and size of the position for one or more of the the further sequences.

5. The method of claim 4, wherein the position and size of the portions for the first sequenceand the further sequences are distributed according to a predetermined distribution.

6. The method of claim 5, wherein the distribution is a uniform distribution.

7. The method of claim 5, wherein the distribution is a Gaussian distribution with a mean at acenter coordinate of the input images.

8. The method of claim 5, wherein the distribution is based on a type of the input images.

9. The method of claim 8, wherein the type comprises at least one of: webcam video frames,static scene frames, video game frames, high motion video frames, low motion video frames.

10. The method of any of claims 1 to 9, wherein the function is based on a difference betweenthe first sequence of output images and the masked sequence of input images.

11. A method of performing lossy image or video encoding, transmission and decoding, themethod comprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce a first output image, wherein the first output image is an approximationof one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the first neural network and the second neural network are produced accordingto the method of any of claims 1 to 10.

12. A method of performing lossy image or video encoding and transmission, the methodcomprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;wherein the first neural network is produced according to the method of any of claims 1to 10.

13. A method of performing lossy image or video decoding, the method comprising the stepsof: with a second neural network, at a second computer system, decoding a latent represen-tation to produce a first output image, wherein the first output image is an approximation ofone image of an image pair of a first sequence of input images;repeating the above step to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the second neural network is produced according to the method of any of claims1 to 10.

14. A data processing apparatus configured to perform the method of any of claims 1 to 13.

15. A computer program comprising instructions which, when the program is executed by acomputer, cause the computer to carry out the method of any of claims 1 to 13.

16. A computer-readable storage medium comprising instructions which, when executed by acomputer, cause the computer carry out the method of any of claims 1 to 13.

17. A method of training one or more neural networks, the one or more neural networks beingfor use in lossy image or video encoding, transmission and decoding, the method comprisingthe steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;with a second neural network, decoding the latent representation to produce a first outputimage, wherein the first output image is an approximation of one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;evaluating a function based on a difference between the first sequence of output imagesand the first sequence of input images;repeating the above steps for one or more further sequences of input images;accumulating gradients associated with said evaluating of the function for the firstsequence and for the one or more further sequences;updating the parameters of the first neural network and the second neural network basedon the accumulated gradients; andrepeating the above steps for a plurality of said sequences of input images to producefirst and second trained neural networks.

18. The method of claim 17, wherein the first sequence of input images and the one or morefurther sequences of input images comprise different images.

19. The method of claim 17 or 18, wherein a final hidden state of the first and second neuralnetworks after evaluating the function for the first sequence is used as a starting hidden stateof the first and second neural networks when repeating the steps for one of the one or morefurther sequences.

20. The method of any of claims 17 to 19, wherein the first sequence comprises 10 or moreinput images.

21. The method of any of claims 17 to 20, wherein the first sequence comprises 30 or moreinput images.

22. The method of any of claims 17 to 21, wherein the first sequence comprises 60 or moreinput images.

23. The method of any of claims 17 to 22, wherein a number of input images in the firstsequence of input images is different to a number of frames in at least one of the one or morefurther sequences of input images.

24. The method of claim 23, wherein the number of input images in the one or more furthersequences of input images is based on a difference between the first output image and the oneof the images of the pair.

25. The method of claim 23, wherein the number of input images in the one or more furthersequences is based on a difference between the first sequence of output images and the firstsequence of input images.

26. The method of claim 23, wherein the number of input images in the one or more furthersequences is based on an output of a classifier configured to detect an artefact in the first outputimage.

27. The method of any of claims 17 to 26, wherein the first sequence and the further sequencescomprise a batch, and repeating said steps for a plurality of batches.

28. The method of any of claims 17 to 27, wherein said updating the parameters comprisesapplying an optimiser to weights of the first and second neural networks based on saidaccumulated gradients.

29. The method of any of claims 17 to 28, wherein the accumulated gradients are associatedwith a greater number of input images than a number of input images in the first sequence or anumber input images in any one of the one or more further sequences.

30. The method of claim 17 to 29, comprising repeating said steps a first number of times usingsequences having a first number of input images, and repeating said steps a second number oftimes using one or more further sequences having a second, greater number of input images.

31. A method of performing lossy image or video encoding, transmission and decoding, themethod comprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce a first output image, wherein the first output image is an approximationof one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the first neural network and the second neural network are produced accordingto the method of any of claims 17 to 30.

32. A method of performing lossy image or video encoding and transmission, the methodcomprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;wherein the first neural network is produced according to the method of any of claims 17to 30.

33. A method of performing lossy image or video decoding, the method comprising the stepsof: with a second neural network, at a second computer system, decoding a latent represen-tation to produce a first output image, wherein the first output image is an approximation ofone image of an image pair of a first sequence of input images;repeating the above step to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the second neural network is produced according to the method of any of claims17 to 30.

34. A data processing apparatus configured to perform the method of any of claims 17 to 33.

35. A computer program comprising instructions which, when the program is executed by acomputer, cause the computer to carry out the method of any of claims 17 to 33.

36. A computer-readable storage medium comprising instructions which, when executed by acomputer, cause the computer carry out the method of any of claims 17 to 33.

37. A method of training one or more neural networks, the one or more neural networks beingfor use in lossy image or video encoding, transmission and decoding, the method comprisingthe steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;with a second neural network, decoding the latent representation to produce a first outputimage, wherein the first output image is an approximation of one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;evaluating a function based on a difference between the first sequence of output imagesand the first sequence of input images;updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps for a plurality of further sequences of input images to producefirst and second trained neural networks, wherein at least some further sequences comprise agreater number of input images than the first sequence of input images.

38. The method of claim 37, wherein the first sequence of input images comprises fewer than10 input images.

39. The method of claim 37 or 38, wherein the at least some further sequences each comprise10 or more input images.

40. The method of claim 37 or 38, wherein the at least some further sequences each comprise30 or more input images.

41. The method of claim 37 or 38, wherein the at least some further sequences each comprise60 or more input images.

42. A method of performing lossy image or video encoding, transmission and decoding, themethod comprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce a first output image, wherein the first output image is an approximationof one of the images of the pair;repeating the above steps to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the first neural network and the second neural network are produced accordingto the method of any of claims 37 to 41.

43. A method of performing lossy image or video encoding and transmission, the methodcomprising the steps of:receiving a first sequence of input images at a first computer system;with a first neural network, for a pair of input images of the first sequence, producing alatent representation of a difference between the images of the pair;transmitting the latent representation to a second computer system;wherein the first neural network is produced according to the method of any of claims 37to 41.

44. A method of performing lossy image or video decoding, the method comprising the stepsof: with a second neural network, at a second computer system, decoding a latent represen-tation to produce a first output image, wherein the first output image is an approximation ofone image of an image pair of a first sequence of input images;repeating the above step to produce a first sequence of output images, the first sequenceof output images being an approximation of the first sequence of input images;wherein the second neural network is produced according to the method of any of claims37 to 41.

45. A data processing apparatus configured to perform the method of any of claims 37 to 44.

46. A computer program comprising instructions which, when the program is executed by acomputer, cause the computer to carry out the method of any of claims 37 to 44.

47. A computer-readable storage medium comprising instructions which, when executed by acomputer, cause the computer carry out the method of any of claims 37 to 44.

48. A method of training one or more neural networks, the one or more neural networks beingfor use in lossy image or video encoding, transmission and decoding, the method comprisingthe steps of:receiving a first image and a second image at a first computer system;modifying the pixels of the first image and the second image at respective first andsecond coordinates to introduce a pixel representation of an object into the first image and thesecond image;with a first neural network, producing a latent representation of a difference between thefirst image and the second image;with a second neural network, decoding the latent representation to produce an outputimage, wherein the output image is an approximation of the first image;evaluating a function based on a difference between the output image image and the firstimage; updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps for one or more sequences of images to produce first andsecond trained neural networks.

49. The method of claim 48, wherein the first and second coordinates are based on apredetermined motion path of the object between the first image and the second image.

50. The method of any of claims 48 to 49, wherein the one or more sequences of imagescomprise a predetermined balance of image sequences having optical flows above a thresholdand image sequences having optical flows below the threshold.

51. The method of claim 50, wherein said modifying the pixels is performed on one or moreof said image sequences having optical flows below the threshold.

52. The method of any of claims 48 or 51, wherein pixels of the object define one or morealphanumeric characters.

53. The method of any of claims 48 to 51, wherein pixels of the object define a non-uniformtexture.

54. The method of any of claims 48 to 51, wherein pixels of the object define a solid colourblock.

55. The method of any of claims 48 to 51, wherein pixels of the object are based on an imagepatch from an image sequence of the sequences of image different to the image sequencecomprising the first image and the second image.

56. The method of any of claims 48 to 55, wherein pixels of the object define a box shape.

57. The method of any of claims 48 to 55, wherein the pixels of the object define a circularshape.

58. The method of any of claims 48 to 55, wherein the pixels of the object define a polygonalshape.

59. The method of any of claims 48 to 58, wherein the first and second coordinates for a firstimage sequence of the one or more sequences of images are different to the first and secondcoordinates for a second image sequence of the one or more sequences of images, and whereinsaid modifying is based on the threshold.

60. A method of performing lossy image or video encoding, transmission and decoding, themethod comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation ofthe first image;wherein the first neural network and the second neural network are produced accordingto the method of any of claims 48 to 59.

61. A method of performing lossy image or video encoding and transmission, the methodcomprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image; andtransmitting the latent representation to a second computer system;wherein the first neural network is produced according to the method of any of claims48 to 59.

62. A method of performing lossy image or video decoding, the method comprising the stepsof: with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation of afirst image;wherein the first neural network and the second neural network are produced accordingto the method of any of claims 48 to 59.

63. A data processing apparatus configured to perform the method of any of claims 48 to 62.

64. A computer program comprising instructions which, when the program is executed by acomputer, cause the computer to carry out the method of any of claims 48 to 62.

65. A computer-readable storage medium comprising instructions which, when executed by acomputer, cause the computer carry out the method of any of claims 48 to 62.

66. A method of training one or more neural networks, the one or more neural networks beingfor use in lossy image or video encoding, transmission and decoding, the method comprisingthe steps of:receiving a first image and a plurality of second images at a first computer system;selecting one of the plurality of second images as a reference image;with a first neural network, producing a latent representation of a difference between thefirst image and the reference image;with a second neural network, decoding the latent representation to produce an outputimage, wherein the output image is an approximation of the first image;evaluating a function based on a difference between the output image image and the firstimage; updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps for one or more sequences of images to produce first andsecond trained neural networks.

67. The method of claim 66, wherein said selecting comprises sampling an image from theplurality of second images.

68. The method of claim 66, wherein said selecting comprises selecting one of the plurality ofsecond images from the image sequence based on one or more predetermined rules.

69. The method of any of claims 66 to 68, wherein the first image and the plurality of secondimages comprise consecutive images of an image sequence.

70. The method of claim 69, wherein the selected one of the plurality of second images is apositioned in the sequence before the first image.

71. The method of claim 69, wherein the selected one of the plurality of second images is apositioned in the sequence after the first image .

72. The method of any of claims 66 to 71, comprising biasing said selecting towards abeginning of the image sequence.

73. The method of any of claims 66 to 71, comprising biasing said selecting towards an end ofthe image sequence.

74. The method of claim 72 or 73, wherein said biasing is based on a property of the firstimage and / or the plurality of second images.

75. The method of claim 74, wherein said property is a scene type of the first image and / or theplurality of second images.

76. The method of any of claims 66 to 75, wherein the one or more sequences of imagescomprise a predetermined balance of image sequences having optical flows above a thresholdand image sequences having optical flows below the threshold, and wherein said selecting isbased on the threshold.

77. A method of performing lossy image or video encoding, transmission and decoding, themethod comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation ofthe first image;wherein the first neural network and the second neural network are produced accordingto the method of any of claims 66 to 76.

78. A method of performing lossy image or video encoding and transmission, the methodcomprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image; andtransmitting the latent representation to a second computer system;wherein the first neural network is produced according to the method of any of claims66 to 76.

79. A method of performing lossy image or video decoding, the method comprising the stepsof: with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation of afirst image;wherein the first neural network and the second neural network are produced accordingto the method of any of claims 66 to 76.

80. A data processing apparatus configured to perform the method of any of claims 66 to 79.

81. A computer program comprising instructions which, when the program is executed by acomputer, cause the computer to carry out the method of any of claims 66 to 79.

82. A computer-readable storage medium comprising instructions which, when executed by acomputer, cause the computer carry out the method of any of claims 66 to 79.

83. A method of training one or more neural networks, the one or more neural networks beingfor use in lossy image or video encoding, transmission and decoding, the method comprisingthe steps of:receiving a first image and a second image at a first computer system;estimating first optical flow information and second optical flow information, each beingindicative of a difference between the first image and the second image;producing combined optical flow information from the first optical flow information andsecond optical flow information;with a first neural network, producing third optical flow information indicative of adifference between the first image and the second image;with a second neural network, decoding the third optical flow information to produce anoutput image, wherein the output image is an approximation of the first image;evaluating a function based on a difference between the combined optical flow informationand the third optical flow information;updating the parameters of the first neural network and the second neural network basedon the evaluated function; andrepeating the above steps for one or more sequences of images to produce first andsecond trained neural networks.

84. The method of claim 83, wherein producing the combined optical flow informationcomprises producing a parameterised combination of the first optical flow information and thesecond optical flow information using one or more combination parameters.

85. The method of claim 84, comprising updating the combination parameters based on oneor more terms of the evaluated function.

86. The method of claim 85, wherein the parameterised combination comprises a weightedsum of the first optical flow information and the second optical flow information and thecombination parameters comprise one or more weights of the weighted sum.

87. The method of any of claims 84 to 86, wherein the parameterised combination comprisesa pixel-wise parameterised combination of the first optical flow information and the secondoptical flow information.

88. The method of any of claims 83 to 87, wherein the function is further based on a rate termand a distortion term.

89. The method of any of claims 83 to 88, wherein the difference between the combined opticalflow information and the third optical flow information comprises a mean squared error.

90. The method of any of claims 83 to 89, comprising estimating the first optical flowinformation using a first optical flow algorithm, and estimating the second optical flowinformation using a second optical flow algorithm.

91. The method of claim 90, wherein the first optical flow algorithm and the second opticalflow algorithm comprise an ensemble of optical flow algorithms.

92. The method of claim 90 or 91, comprising estimating further optical flow informationusing a further optical flow algorithm, and producing the combined optical flow informationfrom the first optical flow information, the second optical flow information and the furtheroptical flow information.

93. The method of any of claims 83 to 92, comprising estimating the difference between thecombined optical flow information and the third optical flow information using a discriminatorfunction.

94. A method of performing lossy image or video encoding, transmission and decoding, themethod comprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image;transmitting the latent representation to a second computer system;with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation ofthe first image;wherein the first neural network and the second neural network are produced accordingto the method of any of claims 83 to 93.

95. A method of performing lossy image or video encoding and transmission, the methodcomprising the steps of:receiving a first image and a second image at a first computer system;with a first neural network, producing a latent representation of a difference between thefirst image and the second image; andtransmitting the latent representation to a second computer system;wherein the first neural network is produced according to the method of any of claims83 to 93.

96. A method of performing lossy image or video decoding, the method comprising the stepsof: with a second neural network, at the second computer system, decoding the latentrepresentation to produce an output image, wherein the output image is an approximation of afirst image;wherein the first neural network and the second neural network are produced accordingto the method of any of claims 83 to 93.

97. A data processing apparatus configured to perform the method of any of claims 83 to 96.

98. A computer program comprising instructions which, when the program is executed by acomputer, cause the computer to carry out the method of any of claims 83 to 96.

99. A computer-readable storage medium comprising instructions which, when executed by acomputer, cause the computer carry out the method of any of claims 83 to 96.

Citation Information

Patent Citations

  • Image compression and decoding, video compression and decoding: methods and systems

    WO2021220008A1