Frame rendering based on previous frames conditioned by quality prediction from machine-learning model

WO2026202744A1PCT designated stage Publication Date: 2026-10-01SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2026/052851
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-24
Publication Date
2026-10-01

Smart Images

  • Figure IB2026052851_01102026_PF_FP_ABST
    Figure IB2026052851_01102026_PF_FP_ABST
Patent Text Reader

Abstract

There is provided an image processing method. The method comprises: inputting data relating to one or more first image frames for content to a machine learning model; determining, by the machine learning model, a prediction of quality of at least part of a second image frame for the content if the second image frame was generated based on the one or more first image frames; and generating, in dependence on the prediction of quality, the second image frame at least partly based on the one or more first image frames.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No. 116457-1550661

[0002] Client Ref. No.: SYP356387WO01

[0003] IMAGE PROCESSING METHOD AND SYSTEM CROSS-REFERENCES TO RELATED APPLICATIONS This application claims the benefit of and priority to United Kingdom (GB) Application No.

[0004] 2504406.6, filed on March 26, 2025, the entire disclosure of which is hereby incorporated by reference in its entirety for all purposes.

[0005] BACKGROUND OF THE INVENTION

[0006] Field of the Invention

[0007] The present invention relates to an image processing method and system.

[0008] Description of the Prior Art

[0009] The speed and realism with which a scene can be rendered is a key consideration in the field of computer graphics processing. In particular, a high frame rate is desirable for many applications, such as multi-player games. As virtual systems become more complex with increasingly complex and feature-rich virtual environments, existing rendering systems (such as graphics engines, graphics drivers, and / or graphics cards) can at times struggle to render content at a target frame rate, while maintaining a sufficient quality of the frames. This can result in reduced realism and immersiveness of the content (e.g. videogame) for the user.

[0010] The present invention seeks to mitigate or alleviate these problems.

[0011] SUMMARY OF THE INVENTION

[0012] One general aspect includes an image processing method. The image processing method can also include inputting data relating to one or more first image frames for content to a machine learning model. The method may also include determining, by the machine learning model, a prediction of quality of at least part of a second image frame for the content if the second image frame was generated based on the one or more first image frames. The method may also include generating, in dependence on the prediction of quality, the second image frame at least partly based on the one or more first image frames. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0013] Implementations may include one or more of the following features. The image processing method may also include generating the second image frame based on the one or more first image frames may include at least one of interpolating and extrapolating the second image frame using the oneAttorney Docket No. 116457-1550661

[0014] Client Ref. No.: SYP356387WO01 or more first image frames. Generating the second image frame in dependence on the prediction of quality may include modifying a process for generating the second image frame in dependence on the prediction of quality. Generating the second image frame in dependence on the prediction of quality may include generating at least part of the second image frame based on the one or more first image frames if the prediction of quality is above a predetermined threshold. If the prediction of quality is below the predetermined threshold, generating the second image frame may include rendering the second image frame. Determining the prediction of quality of the second image frame may include determining a prediction of quality for a plurality of portions of the second image frame if the portions of the second image frame were generated based on the one or more first image frames. Generating the second image frame in dependence on the prediction of quality may include: for a portion of the second image frame for which the prediction of quality is above a predetermined threshold, generating the portion of the second image frame based on the one or more first image frames; and for a portion of the second image frame for which the prediction of quality is below the predetermined threshold, rendering the portion of the second image frame. The image processing method may include modifying the predetermined threshold for at least one portion of the second image frame in dependence on gaze data indicative of a location of gaze of a user of the second image frame. Generating the second image frame may include accessing the data buffer and modifying a process for generating the plurality of portions of the second image frame in dependence on the predictions of quality stored in the data buffer. Generating the second image frame in dependence on the prediction of quality may include selecting, in dependence on the prediction of quality of the second image frame, one of a plurality of image generation techniques for generating the second image frame based on the one or more first image frames. Determining the prediction of quality of the second image frame may include determining a prediction of quality of the second image frame if the second image frame was generated based on the one or more first image frames using a first image generation technique; and where generating the second image frame may include selecting a different, second, image generation technique for generating the second image frame based on the one or more first image frames. Generating the second image frame in dependence on the prediction of quality may include generating the second image frame based on the one or more first image frames, and performing one or more postprocessing operations on the second image frame in dependence on the prediction of quality. The data relating to one or more first image frames may include one or more selected from the list may include of: image data, motion data, temporal data, and context data. The machine learning model may be trained using training data may include pairs of: data relating to one or more third images frames, and indicators of quality of fourth image frames generated based on the one or more thirdAttorney Docket No. 116457-1550661

[0015] Client Ref. No.: SYP356387WO01 image frames. The indicators of quality of the fourth image frames are obtained by evaluating the fourth image frames against ground truth image frames corresponding to the fourth image frames. A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.

[0016] Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0017] BRIEF DESCRIPTION OF THE DRAWINGS

[0018] A more complete appreciation of the disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings, wherein:

[0019] Figure 1 schematically illustrates an entertainment device;

[0020] Figure 2 schematically illustrates an image processing system;

[0021] Figure 3 is a schematic flowchart illustrating an image processing method using a machine learning model;

[0022] Figure 4 is a schematic flowchart illustrating a further image processing method;

[0023] Figure 5 schematically illustrates a sequence of image frames for content;

[0024] Figure 6 is a schematic flowchart illustrating a yet further image processing method; and Figure 7 schematically illustrates an image frame comprising a plurality of portions.

[0025] DESCRIPTION OF THE EMBODIMENTS

[0026] An image processing method and system are disclosed. In the following description, a number of specific details are presented in order to provide a thorough understanding of the embodiments of the present invention. It will be apparent, however, to a person skilled in the art that these specific details need not be employed to practice the present invention. Conversely, specific details known to the person skilled in the art are omitted for the purposes of clarity where appropriate.Attorney Docket No. 116457-1550661

[0027] Client Ref. No.: SYP356387WO01 In an example embodiment of the present invention, a suitable system and / or platform for implementing the methods and techniques herein may be an entertainment device.

[0028] Referring now to the drawings, wherein like reference numerals designate identical or corresponding parts, Figure 1 shows an example of an entertainment device 10 which may be a computer or video game console, for example.

[0029] The entertainment device 10 comprises a central processor 20. The central processor 20 may be a single or multi core processor. The entertainment device also comprises a graphical processing unit or GPU 30. The GPU can be physically separate to the CPU, or integrated with the CPU as a system on a chip (SoC).

[0030] The GPU, optionally in conjunction with the CPU, may process data and generate video images (image data) and optionally audio for output via an AV output. Optionally, the audio may be generated in conjunction with or instead by an audio processor (not shown).

[0031] The video and optionally the audio may be presented to a television or other similar device. Where supported by the television, the video may be stereoscopic. The audio may be presented to a home cinema system in one of a number of formats such as stereo, 5.1 surround sound or 7.1 surround sound. Video and audio may likewise be presented to a head mounted display unit 120 worn by a user 1.

[0032] The entertainment device also comprises RAM 40, and may have separate RAM for each of the CPU and GPU, and / or may have shared RAM. The or each RAM can be physically separate, or integrated as part of an SoC. Further storage is provided by a disk 50, either as an external or internal hard drive, or as an external solid state drive, or an internal solid state drive.

[0033] The entertainment device may transmit or receive data via one or more data ports 60, such as a USB port, Ethernet® port, Wi-Fi® port, Bluetooth® port or similar, as appropriate. It may also optionally receive data via an optical drive 70.

[0034] Audio / visual outputs from the entertainment device are typically provided through one or more A / V ports 90, or through one or more of the wired or wireless data ports 60.

[0035] An example of a device for displaying images output by the entertainment device is the head mounted display ‘HMD’ 120 worn by the user 1. The images output by the entertainment device may be displayed using various other devices - e.g. using a conventional television display connected to A / V ports 90.Attorney Docket No. 116457-1550661

[0036] Client Ref. No.: SYP356387WO01 Where components are not integrated, they may be connected as appropriate either by a dedicated data link or via a bus 100.

[0037] Interaction with the device is typically provided using one or more handheld controllers 130, 130A and / or one or more VR controllers 130A-L,R in the case of the HMD. The user typically interacts with the system, and any content displayed by, or virtual environment rendered by the system, by providing inputs via the handheld controllers 130, 130A. For example, when playing a game, the user may navigate around the game virtual environment by providing inputs using the handheld controllers 130, 130A.

[0038] In embodiments of the present disclosure, the entertainment device 10 generates a plurality of image frames for content. The image frames may be output for display (e.g. via a television or the HMD 120).

[0039] Figure 1 therefore provides an example of a data processing apparatus suitable for executing an application such as a video game and generating images for the video game for display. Images may be output via a display device such as a television or other similar monitor and / or an HMD (e.g. HMD 120). More generally, user inputs can be received by the data processing apparatus and an instance of a video game can be executed accordingly with images being rendered for display to the user.

[0040] Embodiments of the present disclosure relate to use of a trained machine learning (ML) model. The machine learning model may be trained using various techniques, such as supervised learning and / or unsupervised learning.

[0041] In one or more example embodiments of the present disclosure, the machine learning model may be trained using supervised learning. Such a machine learning model may be referred to as a supervised (machine) learning model.

[0042] The supervised learning model is trained using labelled training data to learn a function that maps inputs (typically provided as feature vectors) to outputs (i.e. labels). The labelled training data comprises pairs of inputs and corresponding output labels. The output labels are typically provided by an operator to indicate the desired output for each input. The supervised learning model processes the training data to produce an inferred function that can be used to map new (i.e. unseen) inputs to a label.

[0043] The input data (during training and / or inference) may comprise various types of data, such as numerical values, images, video, text, or audio. Raw input data may be pre-processed to obtain anAttorney Docket No. 116457-1550661

[0044] Client Ref. No.: SYP356387WO01 appropriate feature vector used as input to the model - for example, features of an image or audio input may be extracted to obtain a corresponding feature vector. It will be appreciated that the type of input data and techniques for pre-processing of the data (if required) may be selected based on the specific task the supervised learning model is used for.

[0045] Once prepared, the labelled training data set is used to train the supervised learning model. During training the model adjusts its internal parameters (e.g. weights) so as to optimize (e.g. minimize) an error function, aiming to minimize the discrepancy between the model’s predicted outputs and the labels provided as part of the training data. In some cases, the error function may include a regularization penalty to reduce overfitting of the model to the training data set.

[0046] The supervised learning model may use one or more machine learning algorithms in order to learn a mapping between its inputs and outputs. Example suitable learning algorithms include linear regression, logistic regression, artificial neural networks, decision trees, support vector machines (SVM), random forests, and the K-nearest neighbour algorithm.

[0047] Once trained, the supervised learning model may be used for inference - i.e. for predicting outputs for previously unseen input data. The supervised learning model may perform classification and / or regression tasks. In a classification task, the supervised learning model predicts discrete class labels for input data, and / or assigns the input data into predetermined categories. In a regression task, the supervised learning model predicts labels that are continuous values.

[0048] In some cases, limited amounts of labelled data may be available for training of the model (e.g. because labelling of the data is expensive or impractical). In such cases, the supervised learning model may be extended to further use unlabelled data and / or to generate labelled data.

[0049] Considering using unlabelled data, the training data may comprise both labelled and unlabelled training data, and semi-supervised learning may be used to learn a mapping between the model’s inputs and outputs. For example, a graph-based method such as Laplacian regularization may be used to extend a SVM algorithm to Laplacian SVM in order to perform semi-supervised learning on the partially labelled training data.

[0050] Considering generating labelled data, an active learning model may be used in which the model actively queries an information source (such as a user, or operator) to label data points with the desired outputs. Labels are typically requested for only a subset of the training data set thus reducing the amount of labelling required as compared to fully supervised learning. The model may choose the examples for which labels are requested - for example, the model may request labels for data points that would most change the current model, or that would most reduce theAttorney Docket No. 116457-1550661

[0051] Client Ref. No.: SYP356387WO01 model's generalization error. Semi-supervised learning algorithms may then be used to train the model based on the partially labelled data set.

[0052] As discussed herein, existing rendering systems (such as graphics engines, graphics drivers, and / or graphics cards) can at times struggle to render content at a target frame rate, while maintaining a sufficient quality of the frames. An approach to counteract this is to generate part of the frames from existing frames, as opposed to rendering all frames from scratch. For instance, a portion of the frames can be interpolated from existing frames. However, generating frames in this way often results in artefacts, or distortions in the generated frames. These artefacts in turn cause reduced realism and immersivness of the content (e.g. videogame) for the user.

[0053] Embodiments of the present disclosure seek to address this issue, and relate to using a trained machine learning model to predict a quality of a new image frame if it was generated based on one or more existing frames (e.g. if it the new image frame was interpolated from the existing image frames). The machine learning model receives, as input, data relating to the existing / first image frames for content (e.g. a videogame), which data may for example comprise image (e.g. pixel) data for the existing image frames. The new / second image frame is then generated in dependence on the prediction of quality output by the machine learning model. For example, the new / second image frame may be generated (e.g. interpolated or extrapolated) based on the one or more existing / first image frames if the prediction of quality is above a predetermined threshold, and rendered (or omitted entirely) if the prediction of quality is below the predetermined threshold. This approach provides an improved balance between efficiency and quality, in particular allowing achieving a target frame rate while maintaining high quality of the output image frames. By determining a predicted quality of new frames generated from existing frames, the present approach allows making more efficient use of the available computational resources by generating new frames from existing frames (thus reducing computational costs) where possible (e.g. where the predicted quality is above a threshold), while ensuring that frames are rendered where needed to maintain quality.

[0054] As used herein in relation to the new / second image frame, the term “quality” preferably connotes a degree of realism (and / or accuracy) of the second image frame if it was generated from existing / first frames (e.g. how accurately and effectively the second image frame represents the intended motion and visual content), where the quality increases with increasing realism (and / or accuracy). The prediction of quality of the second image frame may for example comprise one or more of: a prediction of a degree of artefacts and / or distortions in the second image frame, and aAttorney Docket No. 116457-1550661

[0055] Client Ref. No.: SYP356387WO01 prediction of a difference (or similarity) between the second image frame and a corresponding ground-truth (e.g. rendered) second image frame.

[0056] It will be appreciated that the first image frame(s) (also abbreviated as the “first frame(s)”) may relate to any frame of a content (e.g. a videogame, or movie), with the term “first” merely indicating a distinction from a “second” frame. It will also be appreciated that the temporal relationship between the first and second image frames will depend on the image generation technique to be used to generate the second image frame based on the first image frames. In some cases, the second frame may precede or follow the first image frames (e.g. in cases where the second image frame is extrapolated from the first image frames). Alternatively, the second frame may be disposed temporally in between the first image frames (e.g. in cases where the second image frame is interpolated between the first image frames). Where multiple first image frames are used, these first image frames may be consecutive frames, or may be separated by one or more intervening frames.

[0057] Figure 2 shows an example of an image processing system 200 in accordance with one or more embodiments of the present disclosure.

[0058] The image processing system 200 may be provided as part of a user device (such as the entertainment device 10 of Figure 1) and / or as part of a server device. The image processing system 200 may be implemented in a distributed manner using two or more respective processing devices that communicate via a wired and / or wireless communications link. The image processing system 200 may be implemented as a special purpose hardware device or a general purpose hardware device operating under suitable software instruction. The image processing system 200 may be implemented using any suitable combination of hardware and software.

[0059] The image processing system comprises an input processor 210, a machine learning (ML) model 220, and an image generation processor 230. The operations discussed in relation to the input processor 210, ML model 220, and image generation processor 230 may be implemented using the CPU 20 and / or GPU 30, for example. For instance, the ML model 220 may be deployed on the GPU 30.

[0060] The input processor 210 inputs data relating to one or more first image frames for content (e.g. a given scene in a videogame) to the ML model 220. Based on this input data, the ML model 220 determines a prediction of quality (e.g. a numerical value indicative of the predicted quality) of a second image frame for the content if the second image frame was generated (e.g. interpolated) based on the one or more first image frames. The image generation processor 230 then generates,Attorney Docket No. 116457-1550661

[0061] Client Ref. No.: SYP356387WO01 in dependence on the prediction of quality of the second image frame output by the ML model 220, the second image frame at least partly based on the one or more first image frames. For example, the image generation processor 230 may interpolate the second image frame based on the first image frames if the prediction of quality of the second image frame if it were interpolated is above a predetermined threshold.

[0062] Figure 3 shows an example of an image processing method 300 in accordance with one or more embodiments of the present disclosure.

[0063] A step 310 comprises inputting data relating to one or more first image frames for content to a machine learning (ML) model.

[0064] The first image frames (i.e. ‘first frames’) relate to the same content. The content may for example be a videogame, or any other appropriate content (e.g. a film). The content may be pre-generated or generated in real-time. In the case of pre-generated content, the present approach allows improving the efficiency of streaming the pre-generated content by allowing the streaming system to determine whether to stream a given frame, or part thereof (which increases the communication load), or whether the frame can be generated by the receiving device based on previously streamed frames. In a similar manner, for content generated (e.g. rendered in-real time), as described herein, the present approach allows improving efficiency by determining whether a given frame, or part thereof, needs to be rendered (which increases computational cost), or whether existing frames can be used to generate the frame (thus reducing computational cost).

[0065] The first frames comprise one or more image frames. For example, the first frames may comprise one, two, three, five, or ten image frames. The number of frames that comprise the first frames may depend on an image generation technique used to generate the second image frame based on the first image frames at step 330. For instance, in a case where the ML model predicts the quality of a second image frame if it were interpolated between two first image frames, data relating to two first frames may be input to the ML model at step 310. Alternatively, in a case where the ML model predicts the quality of a second image frame if it were extrapolated from four first image frames, data relating to four first frames may be input to the ML model at step 310.

[0066] Alternatively, in some cases, the number of frames comprising the first frames may be different to (e.g. fewer) than the number of frames used to generate the second image frame at step 330. For example, the prediction of quality may be made in dependence on a single first image frame, independent of how many frames are used in the image generation step 330. For instance, where the image generation process comprises interpolation between two frames, data relating to one ofAttorney Docket No. 116457-1550661

[0067] Client Ref. No.: SYP356387WO01 the two frames (to be interpolated from) may be input to the ML model at step 310. Inputting data relating to fewer frames to the ML model at step 310 can reduce the computational cost of inference using the ML model at step 320. Alternatively, the number of first frames may be greater (e.g. 3 or 5) than the number of first frames actually used to generate the second frame, which may improve the accuracy of the quality prediction as the ML model has more data to learn from. The data relating to the one or more first image frames may for example comprise one or more of: image data, motion data, temporal data, context data, and / or gaze data. In some cases, the data relating to the first frames may comprise at least the image data, and optionally one or more of motion data, temporal data, context data, and / or gaze data. The image data may comprise the one or more first image frames (e.g. pixel data for the first frames). The image data may for example comprise matrices including RGB values for one or more pixels in the first frames. Raw pixel data may be input to the ML model as image data. Alternatively, or in addition, the image data may comprise data computed based on the pixel data, such as one or more image properties of the first frames. For example, the image data may comprise the level of detail (LOD) in the first frames. The motion data may comprise data relating to motion of one or more objects (and / or image portions, e.g. pixels or groups thereof) in the one or more first image frames. The motion data may for example be obtained from motion vectors and / or velocity buffers associated with the first frames. The temporal data may comprise data relating to a time interval between the second frame and one or more of the first frames, and / or a time interval between two or more of the first frames. The temporal data may for example comprise frame numbers (and / or timestamps) associated with the first frames, and / or the second frame, in a sequence of frames for the content.

[0068] The context data may comprise contextual data relating to the scene depicted in the first frames, such as depth data or object property data. For example, the context data may comprise data relating to depth of one or more objects in first frames. This depth data may for example be retrieved from depth buffers associated with the first frames. Alternatively, or in addition, the context data may comprise data relating to one or more properties associated with one or more objects in the first frames. The object properties may for example comprise identifiers for the objects, saliency indicators for the objects (which may e.g. indicate the saliency and / or importance of the objects to the visual scene), and / or game state data for videogame content. Alternatively, or in addition, the context data may comprise segmentation data defining one or more objects in the one or more first image frames, for example in the form of segmentation masks that delimit the boundaries of different objects in the first frames. Alternatively, or in addition, the context data may comprise data relating to one or more parameters associated with a camera used to captureAttorney Docket No. 116457-1550661

[0069] Client Ref. No.: SYP356387WO01 the one or more first image frames. The camera parameters may for example comprise focal length, and / or magnification level (i.e. zoom).

[0070] In some cases, the data relating to the first frames may comprise gaze data relating to a location of a user’s gaze with respect to the first frames. As described elsewhere herein, the gaze data may for example be used to determine a threshold quality used at step 230 to determine whether to generate a new second frame based on the existing first frames.

[0071] Referring back to Figure 3, a step 320 comprises determining, by the ML model, a prediction of quality of at least part of a second image frame for the content if the second image frame was generated based on the one or more first image frames.

[0072] In other words, the ML model outputs a predicted quality of the second frame assuming the second frame would be generated based on the first frames, e.g. assuming the second frame would be interpolated or extrapolated using the first frames. Using the ML model in this way allows preemptively assessing the quality of a second frame predicted / generated from first frames, without requiring the second frame to actually be generated. Accordingly, the predicted quality of the second frame can be used to guide the image generation process at step 330, e.g. to modify the image generation technique used at step 330 or to determine whether to perform the image generation at all. This allows improving the quality of the second frame generated at step 330. While the inference using the ML model at step 320 to predict quality inevitably introduces a small degree of latency into the image generation process, this latency is offset by the advantages associated with the present approach. It will be appreciated that the quality prediction at step 320 has a lower computational cost than the generation of second frames. The prediction of quality provides an efficient way to evaluate the quality of second frames without requiring them to be generated, and helps guide the image generation process at step 330.

[0073] The prediction of quality output by the ML model at step 320 may relate to one or more of: predicting a degree of realism (and / or accuracy) of the second frame if it was generated based on the one or more first frames, predicting a degree of artefacts and / or distortions in the second frame, and predicting a difference between the second frame and a corresponding ground-truth (e.g. rendered) second frame.

[0074] The ML model makes the quality prediction at step 320 based on the input data provided to the ML model at step 310. The second image frame relates to the same content as the first image frames. For example, the second frame and first frames may correspond to frames for a given videogame. The quality prediction output at step 320 may be output in any suitable format, suchAttorney Docket No. 116457-1550661

[0075] Client Ref. No.: SYP356387WO01 as binary (e.g. an indication of good vs bad quality), or a numerical score in a predetermined range (e.g. between 0 and 100, 0 indicating worst quality and 100 indicating best quality).

[0076] Step 320 may comprise predicting the quality of at least part of one or more second frames. For example, the quality of a plurality of second frames to be extrapolated from first frames may be predicted at once at step 320. The quality prediction in this case may comprise individual predictions for each second frame, and / or an aggregate quality of the plurality of second frames, such as the average, median, or minimum quality amongst quality predictions for the individual second frames. This may for example be useful in cases where multiple second frames are interpolated in between the first frames.

[0077] The quality of the entire second frame (or frames) may be predicted. Alternatively, the quality of only part of the second frame may be predicted. Predicting the quality of only part of the second frame may improve efficiency by reducing the amount of data input to the ML model at step 310 (e.g. reducing the number of pixels for which data is input to the ML model at step 310) and thus reducing the computational cost of the quality prediction. The part of the second frame for which quality is predicted may be selected in dependence on the data relating to the one or more first frames input at step 310. The part of the second frame for which quality is predicted may comprise a part of the second frame where artefacts and / or distortions are most likely expected to occur, for example the part of the second frame generated from parts of first frames that include the most moving objects (e.g. the parts of first frame for which motion vectors have the highest average magnitudes).

[0078] In some cases, the quality of a plurality of portions of a second frame may be predicted. The plurality of portions may make up the entire second frame, or only part of the second image frame. This example is described in further detail with reference to Figures 6 and 7.

[0079] The ML model is trained to predict a quality of at least part of a new second frame for the content if the second frame was generated based on the first frames, using the data relating to the first frames input at step 310.

[0080] The ML model learns to predict, based on the input data relating to the first frames, how realistic or accurate the second image frame would be if it were generated based on the first frames. The ML model may learn to infer the second frame quality based on the various examples of data relating to first frames described herein. For example, the ML model may make quality predictions for the second frame based on image data for the first frames, the degree of motion in the first frames, depth data for the first frames, and the time interval between the first frames and the secondAttorney Docket No. 116457-1550661

[0081] Client Ref. No.: SYP356387WO01 frame. For instance, the ML model may learn that the quality of the second frame decreases for particular arrangements of pixels in the first frames (e.g. interpolation resulting in more artefacts for such pixel arrangements), and / or with increasing motion at specific depths in the first frames. In an example, the ML model is trained using supervised learning. The ML model may be trained using training data comprising pairs of inputs and corresponding output labels. The training data may comprise, as inputs, data relating to one or more ‘source’ image frames (e.g. image and / or motion data, as described above); and, as output labels, one or more indicators of quality of ‘output’ image frames generated based on the ‘source’ image frames. It will be appreciated that a plurality of such pairs of inputs and output labels is used for training the ML model.

[0082] The inputs for the training data may be obtained from the source images, and / or associated metadata (which may e.g. provide depth or motion data). The input training data may be in the same format as the data input at inference, as described elsewhere herein. The output labels for the training data may be obtained by generating output frames using the source frames, and determining a quality of the output frames. The output frames may be generating using any suitable image generation technique, such as frame interpolation or extrapolation from the source frames. The quality of the output frames may be determined with reference to ground truth frames corresponding to the output frames. This can provide more accurate assessment of the quality of the output frames, by allowing evaluating the realism of the output frames against corresponding ground truth frames. A ground truth frame corresponds to the output frame had it not been generated based on other frames. For example, the ground truth frame may correspond to the original output frame (e.g. original intermediate frames in slow-motion video corresponding to the generated output frame), and / or to the output frame if it were natively rendered, as opposed to being interpolated based on other frames. The quality indicators for the output frames may be determined by comparing the output frames to the ground truth frames. Any suitable image comparison algorithm may be used to compare the output and ground truth frames, such as Peak Signal to Noise Ratio (PSNR), Structural Similarity Index (SSIM), or Video Multi-Method Assessment Fusion (VMAF). The comparison of the output frames to ground truth frames may comprise comparing image data of the frames, and / or any other data associated with the frames, such as the data described herein with reference to data input to the ML model at inference. For instance, motion vectors for the output frames may be compared to motion vectors for the ground truth frames to determine a motion compensated error (MCE) between the frames.Attorney Docket No. 116457-1550661

[0083] Client Ref. No.: SYP356387WO01 Alternatively, or in addition, the quality of the output frames may be determined by evaluating distortions and / or artefacts in the output frames. The evaluation of distortions and artefacts may be performed without reference to ground truth images. For example, the quality of the output frames may be determined by determining the realism of the output frames using techniques such as Naturalness Image Quality Evaluator (NIQE), or Blind / Referenceless Image Spatial Quality Evaluator (BRISQUE). Alternatively, or in addition, the level (e.g. quantity and severity) of temporal and / or spatial artefacts may be determined by detecting and measuring flicker, edge artefacts, ghosting, or blurring.

[0084] The indicators of quality (i.e. quality indicators) for the output frames may be stored in any suitable format, such as binary (e.g. good vs bad quality), or a numerical score in a predetermined range (e.g. between 0 and 100, 0 indicating worst quality and 100 indicating best quality).

[0085] In some cases, a plurality of ML models for determining a prediction of quality of a frame generated based on other frames may be trained. Each of the plurality of ML models may be specifically trained for different scenarios. This more granular approach can allow improved prediction of second frame quality by using ML models that are optimised for the particular scenarios. The ML model relevant for the current scenario may then be selected at inference at step 320.

[0086] For instance, a plurality of ML models may be trained for different image generation techniques the quality of which is being evaluated. For example, a first ML model may be trained to predict the quality of interpolation from first frames, and a second ML model may be trained to predict the quality of extrapolation from first frames. In a similar manner, different ML model may be used for different interpolation techniques, such as for bilinear interpolation and for optical flow-based interpolation.

[0087] Alternatively, or in addition, a plurality of ML models may be trained for different numbers of first frames to be used for the image generation (e.g. the number of frames used in extrapolation). For example, a first ML model may be trained for one input first frame, and a second ML model may be trained for three input first frames. Alternatively, or in addition, a plurality of ML models may be trained for different properties of the first frames to be used for the image generation. For example, a first, low motion, ML model may be trained for first frames in which the degree of motion (e.g. as determined using motion vectors for the first frames) is below a predetermined threshold, and a second, high motion, ML model may be trained for first frames in which the degree of motion is above the predetermined threshold.Attorney Docket No. 116457-1550661

[0088] Client Ref. No.: SYP356387WO01 An example of a supervised learning process for the ML model has been described, however, it will be appreciated that the ML model may be trained in any other appropriate manner. For example, the ML model may be trained using imitation learning to imitate human operators rating the quality of output images generated from other source images, so as to learn a mapping between source image properties and the quality of the output images.

[0089] Referring back to Figure 3, a step 330 comprises generating, in dependence on the prediction of quality output by the ML model at step 320, the second image frame at least partly based on the one or more first image frames.

[0090] Generating the second frame in dependence on the predicted quality of the second frame allows improved control over the image generation process, and generating image frames of improved and more consistent quality, thus providing a more realistic and immersive virtual environment. Generating the second frame based on the one or more first frames may comprise interpolating and / or extrapolating at least part of the second frame based on the first frames. Frame inter / extra-polation provides an efficient way to increase the frame rate of content. However, inter / extra-polation can in some cases result in artefacts in the generated images, which artefacts may be particularly noticeable to the user due to the jump (i.e. reduction) in quality between the first frames and the inter / extra-polated second frame. The present approach at least partially mitigates this disadvantage by allowing frame inter / extra-polation to continue being used for some second frames or parts thereof, while making the second frame generation process dependent on the predicted quality to help prevent generating or outputting second frames of noticeably inferior quality.

[0091] Any suitable image interpolation and extrapolation techniques may be used to generate the second frame. For example, interpolating of the second frame using the first frames may be performed using one or more of: linear interpolation, motion-compensated frame interpolation (MCFI), optical flow-based interpolation, and / or deep learning-based interpolation (e.g. using RIFE (Real-Time Intermediate Flow Estimation). In turn, extrapolating of the second frame using the first frames may for example be performed using one or more of: linear motion extrapolation, optical-flow based extrapolation, block-based motion prediction, and / or deep learning-based extrapolation (e.g. using Recurrent Neural Networks (RNNs) or Generative Adversarial Networks (GANs)). It will be appreciated the terms “interpolation” and “extrapolation” preferably connote the relative order of a generated frame and source frames. Frame interpolation relates to generating a new frame in-between source frames; where the new frame may directly neighbour one or more of theAttorney Docket No. 116457-1550661

[0092] Client Ref. No.: SYP356387WO01 source frames or be separated from the source frames by one or more other frames. Frame extrapolation relates to generating a new frame that follows or precedes the source frames.

[0093] Alternatively, or in addition to frame inter / extra-polation, the second frame may be generated based on the first frames using one or more alternative techniques. For example, one or more GANs may be used to generate new second frames for the content based on the first frames. The image generation at step 330 may comprise generating one or more second frames using one or more of the image generation techniques described herein. For instance, one or both of image interpolation and extrapolation may be used to generate the second frame. For example, a first portion of the second frame may be interpolated between two first frames A and B, and a second portion of the second frame may be extrapolated from first frame A. Alternatively, both the first and second portions of the second frame may be interpolated between first frames A and B. It will be appreciated that, in some cases, multiple second frames may be generated at step 330. For example, multiple intermediate second frames may be interpolated between a pair of first frames.

[0094] Generating the second image frame in dependence on the prediction of quality output by the ML model at step 320 may comprise modifying a process for generating the second image frame in dependence on the prediction of quality. Examples of modifying the second frame image generation process are illustrated in Figures 4 to 7. A further example of modifying the second frame image generation process relating to selecting one or more image generation technique for generating the second frame in dependence on the predicted quality is described herein. It will be appreciated that each of these examples may be used alone or in any combination to modify the process for generating the second image frame in dependence on the prediction of quality output at step 320.

[0095] Referring to Figures 4 and 5, these figures illustrate an example in which the generation of the second frame based on the first frames is conditional on the prediction of quality of the second frame output at step 320. Figure 4 illustrates a schematic flowchart of an example method for generating a second frame based on the first frames, in dependence on the prediction of quality of the second frame.

[0096] Step 420 comprises determining, using the ML model, a prediction of quality of at least part of the second frame if it was generated (e.g. interpolated) based on the one or more first frames. Step 420 may be performed using the same techniques and the same inputs as described herein with reference to step 320 of the method of Figure 3.Attorney Docket No. 116457-1550661

[0097] Client Ref. No.: SYP356387WO01 Steps 432 to 436 correspond to an example of steps that may be performed as part of step 330 of Figure 3.

[0098] Step 432 comprises evaluating whether the predicted quality is above a predetermined threshold. The threshold is used to determine whether the second frame can be generated using the first frames (thus saving computational and / or network resources as compared to e.g. rendering and / or streaming the second frame) while maintaining a minimum desired quality of the second frame. As discussed in relation to steps 434 and 436, the threshold may define whether a predicted quality of the second frame is acceptable such that the second frame can be generated based on the first frames (see e.g. step 434); or alternatively that the predicted quality of the second frame is not acceptable such that the second frame cannot be generated based on the first frames (see e.g. step 436).

[0099] The threshold used at step 432 may be determined empirically. For example, a plurality of second frames generated based on the first frames and their respective quality predictions output by the ML model may be analysed (manually or automatically, e.g. using an image quality grading algorithm) to determine a cut-off threshold below which the quality of the second frames is considered not acceptable. In some cases, a default (e.g. empirically determined) threshold may be set, which default threshold may be modified in dependence on properties of the first frames, e.g. as included in the data relating to the first frames described with reference to step 310. For instance, the threshold may be increased with increasing LOD in the first frames, to account for the fact that a higher quality may be required to avoid distortions in regions of the second frame corresponding to the high LOD regions of the first frames.

[0100] Determining whether the predicted quality is above the threshold may for example comprise comparing the numerical value assigned to the predicted quality (e.g. 55 out of 100) against a threshold numerical value (e.g. 50 out of 100).

[0101] Step 434 comprises generating at least part of the second frame based on the first frames. The method of Figure 4 proceeds to step 434 if it is determined at step 432 that the quality prediction output at step 420 is above or equal to (or above) the predetermined threshold. The second frame may be generated based on the first frames using one or more of the techniques described with reference to step 330 of Figure 3. Generating the second frame based on the first frames may for example comprise interpolating the second frame between the first frames. It will be appreciated that all or part of the second frame may be generated based on the first frames at step 434.Attorney Docket No. 116457-1550661

[0102] Client Ref. No.: SYP356387WO01 Step 436 comprises generating the second image frame, without using the first frames. The method of Figure 4 proceeds to step 436 if it is determined at step 432 that the quality prediction output at step 420 is below (or below or equal to) the predetermined threshold. Generating the second frame at step 436 may for example comprise rendering the second frame. The second frame may for example be rendered using a rendering pipeline associated with the content.

[0103] By making the generation of the second frame based on the first frames conditional on the predicted quality of the second frame generated in this way being above a predetermined threshold, the approach described with reference to Figure 4 provides improved control of output frame quality, whilst keep computational (and / or networking) costs low. It will be appreciated that generating the second frame using existing first frames at step 434 (e.g. via interpolation) is typically computationally cheaper than generating the second frame without the use of first frames 436 (e.g. via rendering); and e.g. provides an efficient way to increase frame rate. Thus, to minimise computational costs, one approach would be to always generate the second frame using the first frames. However, this can result in second frames of poor quality (e.g. when first frames are particularly difficult to interpolate from accurately). The present approach helps prevent this issue by pre-emptively predicting the expected quality of a second frame generated based on first frames, and generating the second frame based on the first frames only if the predicted quality meets predefined requirements (e.g. is above a predetermined threshold).

[0104] This approach therefore allows further improving the balance between efficiency and quality, e.g. by allowing using interpolation to the extent possible to reduce computational costs but without sacrificing quality.

[0105] Figure 5 shows an example sequence of frames generated using the method of Figure 4. Figure 5 illustrates how the techniques of Figure 4 may be applied to the generation and output of a sequence of image frames for content, such as a videogame. In the example of Figure 5, the prediction of quality output by the ML model is a numerical value between 0 and 100 (100 indicating highest quality), and the predetermined quality threshold is 70.

[0106] In a first cycle of the method, the first frames comprise frames 54 and 56, with second frame 55 to be interpolated between frames 54 and 56. The predicted quality, output by the ML model, of frame 55 interpolated based on frames 54 and 55 is 88, and so above the predetermined threshold. Accordingly, frame 55 is interpolated based on frames 54 and 56.

[0107] In a second cycle of the method, the first frames comprise frames 56 and 58, with second frame 57 to be interpolated between frames 56 and 58. The predicted quality, output by the ML model,Attorney Docket No. 116457-1550661

[0108] Client Ref. No.: SYP356387WO01 of frame 57 interpolated based on frames 56 and 58 is 77, and so above the predetermined threshold. Accordingly, frame 57 is interpolated based on frames 56 and 58.

[0109] In a third cycle of the method, the first frames comprise frames 58 and 60, with second frame 59 to be interpolated between frames 58 and 60. The predicted quality, output by the ML model, of frame 59 interpolated based on frames 58 and 60 is 44, and so below the predetermined threshold. Accordingly, rather than being interpolated from frames 58 and 60, frame 59 is rendered instead. As shown in Figure 5, the present approach therefore allows controlling the image interpolation process such that it can be used to achieve higher frame rate at lower computational cost, but not over-used to the detriment of output image frames for content.

[0110] Referring to Figures 6 and 7, these figures illustrate a further example in which the generation of the second frame based on the first frames is conditional on the prediction of quality of the second frame output at step 320. In the example of Figures 6 and 7, quality predictions are made for multiple portions of the second frame such that different portions of the second frame may be generated in a different manner from one another.

[0111] Figure 6 illustrates a schematic flowchart of an example method for generating a second frame based on the first frames, in dependence on predictions of quality of the second frame.

[0112] Step 620 comprises determining, using the ML model, predictions of quality for a plurality of portions of the second frame if the portions were generated based on the first frames.

[0113] Each portion of the second frame may comprise a portion corresponding to one or more pixels of the second frame. The plurality of portions may each have the same or variable shapes. In one example, the second frame is divided into a plurality of portions using a rectangular (e.g. square) grid, as shown in Figure 7. It will be appreciated that the second frame may be divided into portions of any other shape. For instance, arbitrary boundaries may be defined in the second image frame (e.g. boundaries around objects depicted in the second image frame) to define the portions of the second image frame.

[0114] A separate quality prediction may be determined for each portion of the second frame. Alternatively, quality predictions may be determined only for a subset of portions of the second frame.

[0115] It will be appreciated that while the second frame has not yet been generated at step 620, it can nonetheless be divided into portions by dividing the shape of (e.g. pixel dimensions) of the second frame into portions, such as square portions of 2 by 2, 10 by 10, 50 by 50, or 100 by 100 pixels.Attorney Docket No. 116457-1550661

[0116] Client Ref. No.: SYP356387WO01 The quality prediction for each portion (e.g. portions 1 to n) at step 620 may be determined using similar techniques to those described with reference to step 320 of Figure 3, and / or step 420 of Figure 4. The ML model may learn to determine which parts of a second frame can be generated using first frames to a higher quality than other parts, based on different properties of the first frames in different portions of the first frames. For example, the ML model may learn that the quality of a portion of a second frame located, in the pixel space, adjacent portions of first frames with large magnitude motion vectors and / or high LOD if the second frame portion was interpolated using the first frames may be lower than for second frame portions adjacent other portions of the first frames. Likewise, the ML may learn which portions of the first frames are likely to be used to generate a given portion of the second frame (e.g. based on motion vectors for the first frames), and the ML model may predict the quality of a portion of a second frame based on data relating to the portions of the first frames that are likely to be used to generate the portion of the second frame. The quality predictions for portions of the second frame may for example be output by the ML model in the form of a matrix with each entry in the matrix corresponding to a quality prediction for a given portion of the second frame. An example of such a rectangular matrix 800 is shown in Figure 7; each element 810 of the matrix 800 includes a prediction of quality for the corresponding portion of the second frame.

[0117] Steps 632 to 636 are performed for one or more portions of the second frame for which a quality prediction was determined at step 620. The description of steps 632 to 636 below relates to steps performed in respect of a given portion of the second frame. Steps 632 to 636 correspond respectively to steps 432 to 436 of the method of Figure 4, and can be performed using corresponding techniques.

[0118] Step 632 comprises determining, for the given portion of the second frame, whether its predicted quality is above a predetermined threshold.

[0119] Considering steps 620 and 632 together, these steps may in effect output a mask (see e.g. matrix 800) that can be imposed over the second image frame, which mask dictates which portions of the second frame have a predicted quality above a threshold (and so e.g. can be interpolated while maintaining requisite quality) and which portions of the second frame have a predicted quality below the threshold (and so e.g. need to be rendered to maintain high frame quality). This is further illustrated in Figure 7.

[0120] In some cases, the threshold (used at step 632) may be modified for at least one portion of the second image frame in dependence on gaze data indicative of a location of gaze of a user of theAttorney Docket No. 116457-1550661

[0121] Client Ref. No.: SYP356387WO01 second image frame. For example, a higher threshold (e.g. 90) may be used for frame portions closer to the gaze location, and a lower threshold (e.g. 70) may be used for frame portions further away from the gaze location. This allows further improving the balance between the quality of the second frames and the efficiency of generating the second frames. For example, lowering the threshold away from the gaze location allows saving computational costs by performing frame generation based on existing frames (e.g. using interpolation) more freely further away from the gaze location even if quality of the resulting frames is lower by lowering the threshold. Alternatively, or in addition, increasing the threshold in the vicinity of the gaze location allows ensuring that the frame portions the user is looking at are generated to a high quality. In some cases, three or more different thresholds may be used for different portions depending on their position relative to the gaze location.

[0122] Step 634 comprises generating the portion of the second frame based on the first frames. The method of Figure 6 proceeds to step 634 if it is determined at step 632 that the quality prediction output at step 620 is above or equal to (or above) the predetermined threshold. The portion of the second frame may be generated based on the first frames using one or more of the techniques described with reference to step 330 of Figure 3. Generating the portion of the second frame based on the first frames may for example comprise interpolating the portion between the first frames. Step 636 comprises generating the portion of the second image frame, without using the first frames. The method of Figure 6 proceeds to step 636 if it is determined at step 632 that the quality prediction output at step 620 is below (or below or equal to) the predetermined threshold. Generating the portion of the second frame at step 636 may for example comprise rendering the portion of the second frame using a rendering pipeline.

[0123] Determining separate predictions of quality for multiple portions of the second frame to then modify how the portions are generated (e.g. whether to interpolate or render a given portion) advantageously provides a more granular approach that further improves the balance between efficiency and image quality, as the most appropriate image generation process can be used for different portions of the second frame.

[0124] In some cases, the predictions of quality for portions of the second frame may be stored in a data buffer associated with the second frame. For example, the predictions of quality may be stored in a depth buffer and / or stencil buffer associated with the second frame. For instance, the predictions of quality may be stored on a per pixel basis in a depth buffer for the second frame, in dependence on the predictions of quality for different portions of the second frame. When generating theAttorney Docket No. 116457-1550661

[0125] Client Ref. No.: SYP356387WO01 second frame, the predictions of quality may be obtained by accessing the data buffer, and a process for generating the portions of the second frame may be modified in dependence on the predictions of quality stored in the data buffer (e.g. a given portion may be rendered or interpolated depending on the prediction of quality for that portion).

[0126] Storing the predictions of quality in the data buffer in this way allows further improving the efficiency of generating the second frame. For example, by storing the quality predictions in a depth or stencil buffer, the determination whether to render a given portion or generate (e.g. interpolate) the portion based on the first frames can take place at hardware level earlier in the rendering pipeline than using arbitrary image-based masking, thereby facilitating more efficient generation of the second frame.

[0127] Figure 7 illustrates an example second frame 700 generated using the method of Figure 6. The second frame depicts an example virtual scene for videogame content.

[0128] The quality predictions for portions of the second frame 700 are shown using a predicted quality matrix 800. As shown in Figure 7, the second frame is divided into portions 710, with each portion 710 assigned a quality prediction in the form of a numerical value 810. In the example of Figure 7, the predetermined threshold used at step 632 is 70. The quality predictions for portions 710-1 are below the predetermined threshold, and portions 710-1 are rendered using a rendering pipeline for the videogame. The quality predictions for the remaining portions are above the predetermined threshold, and these portions are generated (e.g. extrapolated) using the first frames.

[0129] As illustrated in Figure 7, the present techniques allow more efficiently allocating computing resources by only rendering portions of a second frame for which computationally cheaper approaches (e.g. interpolation) is predicted to result in poor image quality, while using the computationally cheaper approaches for other portions of the second frame.

[0130] It will be appreciated that a grid of relatively large portions 710 is shown in Figure 7 for illustrative purposes. In practice, the portions may be much smaller, and e.g. each comprise one or more pixels of the second image frame.

[0131] In one or more examples, modifying a process for generating the second frame in dependence on the prediction of quality comprises selecting one or more image generation techniques for generating the second frame. For example, generating the second frame in dependence on the prediction of quality output at step 320 may comprise selecting, in dependence on the prediction of quality, one or more of a plurality of image generation techniques for generating the secondAttorney Docket No. 116457-1550661

[0132] Client Ref. No.: SYP356387WO01 frame based on the first frames. This provides further improved control over the image generation process, thus allowing further improving the efficiency and the quality of output image frames. For instance, at step 320, a prediction of quality of the second frame may be determined assuming that a first image generation technique (e.g. MCFI) was used to generate the second frame using the first frames. Depending on this prediction of quality, a different, second, image generation technique (e.g. linear interpolation, or optical flow-based interpolation) may then be selected for generating the second frame at step 330.

[0133] For example, if the predicted quality of the second frame (e.g. 65) is below a first predetermined threshold (e.g. 80), a second image generation technique associated with higher quality outputs may be selected to generate the second frame. For example, more computationally expensive, but higher quality, optical flow-based interpolation may be used in place of MCFI. This helps to ensure that the second frames are generated to an acceptable quality.

[0134] Alternatively, or in addition, if the predicted quality of the second frame (e.g. 98) is above a second predetermined threshold (e.g. 90), which may the same or different as the first threshold, a second image generation technique (e.g. linear interpolation) associated with lower computational costs may be selected to generate the second frame. For example, in view of the high predicted quality of the second frame, computationally cheaper linear interpolation may be used in place of MCFI to improve efficiency while still expecting a sufficiently high quality of the output frame.

[0135] Predictions of quality of the second frame assuming different image generation techniques may be obtained by using different ML models, each ML model being trained to predict the quality of a second frame generated based on first frames using a given image generation technique. It will be appreciated that the first and second image generation techniques may relate to techniques of the same type (e.g. different interpolation techniques), or of different types (e.g. interpolation and extrapolation techniques).

[0136] In some cases, different image generation techniques may be used for different portions of the second image frame. Use of different techniques for different portions may for example be implemented as part of the method of Figure 6. For example, a more computationally complex, higher quality, generation technique may be used for portions of a second frame with lower predicted qualities; and a computationally cheaper, lower quality, generation technique may be used for portions of a second frame with higher predicted qualities.

[0137] Alternatively, or in addition, to selecting an image generation technique for initially generating the second frame (e.g. selecting an appropriate interpolation technique), one or more image post-Attorney Docket No. 116457-1550661

[0138] Client Ref. No.: SYP356387WO01 processing operations may be selected for performing on the second frame in dependence on the prediction of quality output at step 320. This provides improved quality control over the second frames and more efficient allocation of resources by using appropriate post-processing only when needed to.

[0139] The post-processing operations may for example include one or more of: upscaling (e.g. using bicubic interpolation, Lanczos resampling, and / or super-resolution techniques), noise reduction, and / or sharpening.

[0140] For example, considering upscaling as an example of a post-processing operation, the second frame may be generated based on the first frames, optionally even if the predicted quality of the second frame is below the predetermined threshold. If the predicted quality of the second frame is below a threshold (which may be the same or different to the predetermined threshold), the second frame may be upscaled to improve its quality. If the predicted quality of the second frame is above the threshold (which may be the same or different to the predetermined threshold), upscaling operations may be omitted to improve efficiency.

[0141] Referring back to Figure 3, in a summary embodiment of the present invention an image processing method comprises the following steps.

[0142] A step 310 comprises inputting data relating to one or more first image frames for content to a machine learning model, as described elsewhere herein.

[0143] A step 320 comprises determining, by the machine learning model, a prediction of quality of at least part of a second image frame for the content if (i.e. assuming that) the second image frame was generated based on the one or more first image frames, as described elsewhere herein.

[0144] A step 330 comprises generating, in dependence on the prediction of quality, the second image frame at least partly based on the one or more first image frames, as described elsewhere herein. It will be apparent to a person skilled in the art that variations in the above method corresponding to operation of the various embodiments of the method and / or apparatus as described and claimed herein are considered within the scope of the present disclosure, including but not limited to that:

[0145] generating 330 the second image frame based on the one or more first image frames comprises at least one of interpolating and extrapolating the second image frame using the one or more first image frames, as described elsewhere herein;Attorney Docket No. 116457-1550661

[0146] Client Ref. No.: SYP356387WO01 in this case, optionally generating 330 the second image frame based on the one or more first image frames comprises one or more selected from the list consisting of: interpolating at least part of the second image frame based on the one or more first image frames, and extrapolating at least part of the second image frame based on the one or more first image frames, as described elsewhere herein;

[0147] generating 330 the second image frame in dependence on the prediction of quality comprises modifying a process for generating the second image frame in dependence on the prediction of quality, as described elsewhere herein;

[0148] generating 330 the second image frame in dependence on the prediction of quality comprises generating at least part of the second image frame based on the one or more first image frames if the prediction of quality is above a predetermined threshold, as described elsewhere herein;

[0149] in this case, optionally if the prediction of quality is below the predetermined threshold, generating the second image frame comprises rendering the second image frame, as described elsewhere herein;

[0150] in this case, optionally if the prediction of quality is below the predetermined threshold, generating the second image frame comprises generating the second frame without using the first frames, or omitting the second frame, as described elsewhere herein;

[0151] determining 320 the prediction of quality of the second image frame comprises determining a prediction of quality for a plurality of portions of the second image frame if the portions of the second image frame were generated based on the one or more first image frames, as described elsewhere herein;

[0152] in this case, optionally generating 330 the second image frame in dependence on the prediction of quality comprises: for a portion of the second image frame for which the prediction of quality is above a predetermined threshold, generating the portion of the second image frame based on the one or more first image frames; and for a portion of the second image frame for which the prediction of quality is below the predetermined threshold, rendering the portion of the second image frame, as described elsewhere herein;

[0153] in this case, optionally further comprising modifying the predetermined threshold for at least one portion of the second image frame in dependence on gaze data indicative of a location of gaze of a user of the second image frame, as described elsewhere herein;Attorney Docket No. 116457-1550661

[0154] Client Ref. No.: SYP356387WO01 □ where, optionally modifying the predetermined threshold in dependence on gaze data comprises reducing the threshold with increasing distance from the gaze location, as described elsewhere herein;

[0155] in this case, optionally further comprising storing the prediction of quality for each of the plurality of portions of the second image frame in a data buffer associated with the second image frame; where generating 330 the second image frame comprises accessing the data buffer and modifying a process for generating the plurality of portions of the second image frame in dependence on the predictions of quality stored in the data buffer, as described elsewhere herein; □ where, optionally the data buffer is a depth buffer and / or a stencil buffer, as described elsewhere herein;

[0156] further comprising determining the threshold in dependence on the data relating to the one or more first image frames, as described elsewhere herein;

[0157] generating 330 the second image frame in dependence on the prediction of quality comprises selecting, in dependence on the prediction of quality of the second image frame, one of a plurality of image generation techniques for generating the second image frame based on the one or more first image frames, as described elsewhere herein;

[0158] in this case, optionally determining 320 the prediction of quality of the second image frame comprises determining a prediction of quality of the second image frame if the second image frame was generated based on the one or more first image frames using a first image generation technique, as described elsewhere herein;

[0159] in this case, optionally generating 330 the second image frame comprises selecting a different, second, image generation technique for generating the second image frame based on the one or more first image frames, as described elsewhere herein;

[0160] □ where, optionally if the predicted quality is below a first threshold (optionally the same as the predetermined threshold), a second image generation technique associated with higher quality outputs is selected, as described elsewhere herein;

[0161] □ where, optionally if the predicted quality is above a second threshold (optionally the same as the predetermined threshold), a second image generation technique associated with lower computational cost is selected, as described elsewhere herein;Attorney Docket No. 116457-1550661

[0162] Client Ref. No.: SYP356387WO01 generating 330 the second image frame in dependence on the prediction of quality comprises generating the second image frame based on the one or more first image frames, and performing one or more post-processing operations on the second image frame in dependence on the prediction of quality, as described elsewhere herein;

[0163] the data relating to one or more first image frames comprises one or more selected from the list consisting of: image data, motion data, temporal data, and context data, as described elsewhere herein;

[0164] the data relating to one or more first image frames comprises one or more selected from the list consisting of: the one or more first image frames, data relating to motion of one or more objects in the one or more first image frames, a time interval between the second image frame and the one or more first image frames, a time interval between the one or more first image frames, data relating to depth of one or more objects in the one or more first image frames, data relating to one or more parameters associated with a camera used to capture the one or more first image frames, data relating to one or more properties associated with one or more objects in the one or more first image frames, and segmentation data defining one or more objects in the one or more first image frames, as described elsewhere herein;

[0165] the machine learning model is trained using training data comprising pairs of: data relating to one or more third images frames, and indicators of quality of fourth image frames generated based on the one or more third image frames, as described elsewhere herein;

[0166] in this case, optionally the indicators of quality of the fourth image frames are obtained by evaluating the fourth image frames against ground truth image frames corresponding to the fourth image frames, as described elsewhere herein;

[0167] □ where, optionally ground truth image frames are rendered image frames, as described elsewhere herein;

[0168] predicting the quality of the second image frame comprises predicting a degree of realism (and / or accuracy) of the second image frame if it was generated based on the one or more first image frames, as described elsewhere herein;

[0169] predicting the quality of the second image frame comprises one or more of: predicting a degree of artefacts and / or distortions in the second image frame, and predicting a difference between the second image frame and a corresponding ground-truth (e.g. rendered) second image frame, as described elsewhere herein;Attorney Docket No. 116457-1550661

[0170] Client Ref. No.: SYP356387WO01 further comprising selecting the machine learning model amongst a plurality of candidate machine learning models in dependence on one or more selected from the listed consisting of: one or more image generation techniques used to generate the second image frame at step 330, the number of image frames in the first image frames, and the data relating to the one or more first image frames (e.g. motion data for the first image frames), as described elsewhere herein;

[0171] determining the prediction of quality of the second image frame comprises: selecting a part of the second image frame in dependence on the data relating to the one or more first frames; and determining, by the machine learning model, a prediction of quality of the selected part of the second image frame, as described elsewhere herein;

[0172] the content is videogame content, as described elsewhere herein;

[0173] the step 320 of determining the prediction of quality is performed by a first computing device (e.g. server device), the step 330 of generating the second image frame is performed by a second computing device (e.g. client device), and the method further comprises transmitting, by the first computing device to the second computing device, the prediction of quality, as described elsewhere herein;

[0174] the steps of determining 320 the prediction of quality by the machine learning model and of generating 330 the second image frame are performed using the same computing device, as described elsewhere herein;

[0175] the machine learning model is trained using supervised learning, as described elsewhere herein;

[0176] the machine learning model is trained using imitation learning, as described elsewhere herein; and

[0177] further comprising outputting the second image frame for display, as described elsewhere herein.

[0178] In another summary embodiment of the present invention, there is provided a method of training a machine learning model for use in image processing, as described elsewhere herein. In another summary embodiment of the present invention, there is provided a trained machine learning model for use in image processing, as described elsewhere herein.Attorney Docket No. 116457-1550661

[0179] Client Ref. No.: SYP356387WO01 It will be appreciated that the above methods may be carried out on conventional hardware suitably adapted as applicable by software instruction or by the inclusion or substitution of dedicated hardware.

[0180] Thus the required adaptation to existing parts of a conventional equivalent device may be implemented in the form of a computer program product comprising processor implementable instructions stored on a non-transitory machine-readable medium such as a floppy disk, optical disk, hard disk, solid state disk, PROM, RAM, flash memory or any combination of these or other storage media, or realised in hardware as an ASIC (application specific integrated circuit) or an FPGA (field programmable gate array) or other configurable circuit suitable to use in adapting the conventional equivalent device. Separately, such a computer program may be transmitted via data signals on a network such as an Ethernet, a wireless network, the Internet, or any combination of these or other networks.

[0181] Referring back to Figure 2, in a summary embodiment of the present invention, an image processing system 200 may comprise the following:

[0182] An input processor 210 configured (for example by suitable software instruction) to input data relating to one or more first image frames for content to a machine learning model, as described elsewhere herein.

[0183] A machine learning model 220 configured (for example by suitable software instruction) and / or trained to determine a prediction of quality of at least part of a second image frame for the content if the second image frame was generated based on the one or more first image frames, as described elsewhere herein.

[0184] An image generation processor 230 configured (for example by suitable software instruction) to generate, in dependence on the prediction of quality, the second image frame at least partly based on the one or more first image frames, as described elsewhere herein.

[0185] It will be appreciated that the above system 200, operating under suitable software instruction, may implement the methods and techniques described herein.

[0186] Of course, the functionality of these processors may be realised by any suitable number of processors located at any suitable number of devices and any suitable number of devices as appropriate rather than requiring a one-to-one mapping between the functionality and a device or processor.Attorney Docket No. 116457-1550661

[0187] Client Ref. No.: SYP356387WO01 The foregoing discussion discloses and describes merely exemplary embodiments of the present invention. As will be understood by those skilled in the art, the present invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. Accordingly, the disclosure of the present invention is intended to be illustrative, but not limiting of the scope of the invention, as well as other claims. The disclosure, including any readily discernible variants of the teachings herein, defines, in part, the scope of the foregoing claim terminology such that no inventive subject matter is dedicated to the public.

Claims

Attorney Docket No. 116457-1550661Client Ref. No.: SYP356387WO01 CLAIMSWHAT IS CLAIMED IS:

1. An image processing method comprising:inputting data relating to one or more first image frames for content to a machine learning model;determining, by the machine learning model, a prediction of quality of at least part of a second image frame for the content if the second image frame was generated based on the one or more first image frames; andgenerating, in dependence on the prediction of quality, the second image frame at least partly based on the one or more first image frames.

2. The image processing method of claim 1, wherein generating the second image frame based on the one or more first image frames comprises at least one of interpolating and extrapolating the second image frame using the one or more first image frames.

3. The image processing method of claim 1 or 2, wherein generating the second image frame in dependence on the prediction of quality comprises modifying a process for generating the second image frame in dependence on the prediction of quality.

4. The image processing method of any preceding claim, wherein generating the second image frame in dependence on the prediction of quality comprises generating at least part of the second image frame based on the one or more first image frames if the prediction of quality is above a predetermined threshold.

5. The image processing method of claim 4, wherein, if the prediction of quality is below the predetermined threshold, generating the second image frame comprises rendering the second image frame.

6. The image processing method of any preceding claim, wherein determining the prediction of quality of the second image frame comprises determining a prediction of quality for a plurality of portions of the second image frame if the portions of the second image frame were generated based on the one or more first image frames.

7. The image processing method of claim 6, wherein generating the second image frame in dependence on the prediction of quality comprises:Attorney Docket No. 116457-1550661Client Ref. No.: SYP356387WO01 for a portion of the second image frame for which the prediction of quality is above a predetermined threshold, generating the portion of the second image frame based on the one or more first image frames; andfor a portion of the second image frame for which the prediction of quality is below the predetermined threshold, rendering the portion of the second image frame.

8. The image processing method of claim 6 or 7, further comprising modifying the predetermined threshold for at least one portion of the second image frame in dependence on gaze data indicative of a location of gaze of a user of the second image frame.

9. The image processing method of any of claims 6 to 8, further comprising storing the prediction of quality for each of the plurality of portions of the second image frame in a data buffer associated with the second image frame; wherein generating the second image frame comprises accessing the data buffer and modifying a process for generating the plurality of portions of the second image frame in dependence on the predictions of quality stored in the data buffer.

10. The image processing method of any preceding claim, wherein generating the second image frame in dependence on the prediction of quality comprises selecting, in dependence on the prediction of quality of the second image frame, one of a plurality of image generation techniques for generating the second image frame based on the one or more first image frames.

11. The image processing method of claim 10, wherein determining the prediction of quality of the second image frame comprises determining a prediction of quality of the second image frame if the second image frame was generated based on the one or more first image frames using a first image generation technique; and wherein generating the second image frame comprises selecting a different, second, image generation technique for generating the second image frame based on the one or more first image frames.

12. The image processing method of any preceding claim, wherein generating the second image frame in dependence on the prediction of quality comprises generating the second image frame based on the one or more first image frames, and performing one or more postprocessing operations on the second image frame in dependence on the prediction of quality.Attorney Docket No. 116457-1550661Client Ref. No.: SYP356387WO01 13. The image processing method of any preceding claim, wherein the data relating to one or more first image frames comprises one or more selected from the list consisting of: image data, motion data, temporal data, and context data.

14. The image processing method of any preceding claim, wherein the machine learning model is trained using training data comprising pairs of: data relating to one or more third images frames, and indicators of quality of fourth image frames generated based on the one or more third image frames.

15. The image processing method of claim 14, wherein the indicators of quality of the fourth image frames are obtained by evaluating the fourth image frames against ground truth image frames corresponding to the fourth image frames.

16. One or more non-transitory computer-readable media storing computer executable instructions that, upon execution by one or more processors of a system, cause the system to perform the method of any one of the preceding claims.

17. An image processing system comprising:an input processor configured to input data relating to one or more first image frames for content to a machine learning model;the machine learning model, the machine learning model being configured to determine a prediction of quality of at least part of a second image frame for the content if the second image frame was generated based on the one or more first image frames; andan image generation processor configured to generate, in dependence on the prediction of quality, the second image frame at least partly based on the one or more first image frames.

18. The image processing system of claim 17, wherein generating the second image frame based on the one or more first image frames comprises at least one of interpolating and extrapolating the second image frame using the one or more first image frames.

19. The image processing system of claim 17 or 18, wherein generating the second image frame in dependence on the prediction of quality comprises modifying a process for generating the second image frame in dependence on the prediction of quality.

20. The image processing system of any preceding claim, wherein generating the second image frame in dependence on the prediction of quality comprises generating at least partAttorney Docket No. 116457-1550661Client Ref. No.: SYP356387WO01 of the second image frame based on the one or more first image frames if the prediction of quality is above a predetermined threshold.

21. The image processing system of claim 20, wherein, if the prediction of quality is below the predetermined threshold, generating the second image frame comprises rendering the second image frame.

22. The image processing system of any preceding claim, wherein determining the prediction of quality of the second image frame comprises determining a prediction of quality for a plurality of portions of the second image frame if the portions of the second image frame were generated based on the one or more first image frames.

23. The image processing system of claim 22, wherein generating the second image frame in dependence on the prediction of quality comprises:for a portion of the second image frame for which the prediction of quality is above a predetermined threshold, generating the portion of the second image frame based on the one or more first image frames; andfor a portion of the second image frame for which the prediction of quality is below the predetermined threshold, rendering the portion of the second image frame.

24. The image processing system of claim 22 or 23, further comprising modifying the predetermined threshold for at least one portion of the second image frame in dependence on gaze data indicative of a location of gaze of a user of the second image frame.

25. The image processing system of claim 22 or 24, further comprising storing the prediction of quality for each of the plurality of portions of the second image frame in a data buffer associated with the second image frame; wherein generating the second image frame comprises accessing the data buffer and modifying a process for generating the plurality of portions of the second image frame in dependence on the predictions of quality stored in the data buffer.

26. The image processing system of any preceding claim, wherein generating the second image frame in dependence on the prediction of quality comprises selecting, in dependence on the prediction of quality of the second image frame, one of a plurality of image generation techniques for generating the second image frame based on the one or more first image frames.

27. The image processing system of claim 26, wherein determining the prediction of quality of the second image frame comprises determining a prediction of quality of the secondAttorney Docket No. 116457-1550661Client Ref. No.: SYP356387WO01 image frame if the second image frame was generated based on the one or more first image frames using a first image generation technique; and wherein generating the second image frame comprises selecting a different, second, image generation technique for generating the second image frame based on the one or more first image frames.

28. The image processing system of any preceding claim, wherein generating the second image frame in dependence on the prediction of quality comprises generating the second image frame based on the one or more first image frames, and performing one or more postprocessing operations on the second image frame in dependence on the prediction of quality.

29. The image processing system of any preceding claim, wherein the data relating to one or more first image frames comprises one or more selected from the list consisting of: image data, motion data, temporal data, and context data.

30. The image processing system of any preceding claim, wherein the machine learning model is trained using training data comprising pairs of: data relating to one or more third images frames, and indicators of quality of fourth image frames generated based on the one or more third image frames.

31. The image processing system of claim 30, wherein the indicators of quality of the fourth image frames are obtained by evaluating the fourth image frames against ground truth image frames corresponding to the fourth image frames.