Image processing method and system for setting a bitrate ladder

WO2026202641A1PCT designated stage Publication Date: 2026-10-01SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2026/052490
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-03-13
Publication Date
2026-10-01

Smart Images

  • Figure IB2026052490_01102026_PF_FP_ABST
    Figure IB2026052490_01102026_PF_FP_ABST
Patent Text Reader

Abstract

There is provided an image processing method and system for setting a bitrate ladder during transmission of image content. The method comprises generating one or more first image frames for image content in dependence on a first bitrate ladder; estimating motion information in a current scene of the image content, where the motion information is indicative of movement of one or more objects in the current scene; selecting a second bitrate ladder in dependence on the motion information; and generating one or more second image frames for the image content in dependence on the second bitrate ladder instead of the first bitrate ladder.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No.: 116457-1550002

[0002] Client Ref. No.: SYP356379WO01

[0003] IMAGE PROCESSING METHOD AND SYSTEM FOR SETTING A BITRATE LADDER PRIORITY CLAIM

[0004] This international application claims priority to UK Patent Application No. 2504344.9, filed on March 25, 2025, the disclosure of which is herein incorporated by reference in its entirety for all purposes.

[0005] FIELD

[0006] The present invention relates to an image processing method and system.

[0007] BACKGROUND

[0008] When transmitting image content over a network, the transmission can be improved using adaptive bitrate streaming. A bitrate ladder can be provided such that the image content is available for transmission at various bitrates, each of which may provide various resolutions and / or frame rates. The transmission can then be tailored to the prevailing network conditions or device capabilities by selecting one of the available bitrates from the bitrate ladder to provide appropriate playback of the image content.

[0009] However, in some scenarios, the network conditions or device capabilities may not allow an appropriate playback of the image content from the provided bitrates available in the bitrate ladder, thereby causing a reduction in the perceived quality of the image content, for example due to excessive lag or reduced resolution.

[0010] The present invention seeks to mitigate or alleviate these problems.

[0011] SUMMARY

[0012] Various aspects and features of the present invention are defined in the appended claims and within the text of the accompanying description.

[0013] Some embodiments include an image processing method for setting a bitrate ladder during transmission of image content, the image processing method including: generating one or more first image frames for image content in dependence on a first bitrate ladder; estimating motion information in a current scene of the image content, the motion information being indicative of movement of one or more objects in the current scene; selecting a second bitrate ladder in dependence on the motion information; and generating one or more second image frames for the image content in dependence on the second bitrate ladder instead of the first bitrate ladder.Attorney Docket No.: 116457-1550002

[0014] Client Ref. No.: SYP356379WO01 In some embodiments, selecting the second bitrate ladder in dependence on the motion information comprises comparing the movement of one or more objects in the current scene to a movement threshold.

[0015] In some embodiments, in response to determining that the movement of the one or more objects in the current scene does not exceed the movement threshold, the second bitrate ladder is selected such that the one or more second image frames are rendered at a higher resolution than the first bitrate ladder for the same available bandwidth.

[0016] In some embodiments, in response to determining that the movement of the one or more objects in the current scene exceeds the movement threshold, the second bitrate ladder is selected such that the one or more second image frame are rendered at a lower resolution than the first bitrate ladder for the same available bandwidth.

[0017] In some embodiments, selecting the second bitrate ladder comprises selecting one of a plurality of candidate bitrate ladders in dependence on the motion information.

[0018] In some embodiments, each of the plurality of candidate bitrate ladders is associated with a different degree of motion in the scene, and wherein the plurality of candidate bitrate ladders are generated using a machine learning model trained to determine, for a plurality of degrees of motion in the scene, a bitrate ladder that optimises quality of frames generated using the bitrate ladder. In some embodiments, the selecting is performed in response to generation of an update trigger while transmitting the one or more first image frames over a network.

[0019] In some embodiments, generating the update trigger at predetermined sampling points in the image content.

[0020] In some embodiments, the predetermined sampling points are between each of a plurality of scenes of the image content.

[0021] In some embodiments, predetermined sampling points are at a regular time interval during the image content.

[0022] In some embodiments, the regular time interval is based on a predetermined number of frames. In some embodiments, estimating motion information in the current scene is performed in dependence on motion vectors associated with one or more preceding frames.

[0023] In some embodiments, obtaining the motion vectors from information stored in velocity buffers.Attorney Docket No.: 116457-1550002

[0024] Client Ref. No.: SYP356379WO01 In some embodiments, estimating motion information in the current scene is performed in dependence on scene metadata associated with the current scene.

[0025] In some embodiments, the image content is a video game, and the method comprises identifying the scene metadata based on game state data indicative of a current gameplay state of the video game.

[0026] In some embodiments, identifying the scene metadata comprises comparing the game state data to recorded game state data associated with one or more previous instances of the video game being played, the recorded game state data being associated with the scene metadata.

[0027] In some embodiments, the game state data comprises one or more of a selection of settings including but not limited to player viewpoint, player location, player movement, number of objects, and environment lighting.

[0028] In some embodiments, estimating motion information in the current scene is performed using a machine learning model trained to generate a dynamicity value indicative of movement of the one or more objects in the current scene in dependence on one or more sample frames of the current scene.

[0029] Some embodiments include a system that includes: one or more processors; and one or more computer-readable media storing instructions which, when executed by the one or more processors, cause the system to perform part or all of the operations and / or methods disclosed herein.

[0030] Some embodiments include one or more non-transitory computer-readable media storing instructions which, when executed by one or more processors, cause a system to perform part or all of the operations and / or methods disclosed herein.

[0031] BRIEF DESCRIPTION OF THE DRAWINGS

[0032] A more complete appreciation of the disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings, wherein:

[0033] Figure l is a schematic diagram illustrating an example of an entertainment device;

[0034] Figure 2 is sequence of steps illustrating an example of selecting a bitrate ladder in dependence on motion information;Attorney Docket No.: 116457-1550002

[0035] Client Ref. No.: SYP356379WO01 Figures 3A and 3B are examples illustrating distributions of predetermined sampling points in image content;

[0036] Figures 4A, 4B, and 4C are examples illustrating how motion information may be estimated for a current scene;

[0037] Figures 5 A and 5B are examples illustrating use of game data to identify a current scene;

[0038] Figure 6 is a schematic diagram illustrating an image processing apparatus in accordance with one embodiment of the disclosure; and

[0039] Figure 7 is a sequence of steps illustrating an example of receiving image content.

[0040] DESCRIPTION OF THE EMBODIMENTS

[0041] An image processing method and system for setting a bitrate ladder during transmission of image content are disclosed. In the following description, a number of specific details are presented in order to provide a thorough understanding of the embodiments of the present invention. It will be apparent, however, to a person skilled in the art that these specific details need not be employed to practice the present invention. Conversely, specific details known to the person skilled in the art are omitted for the purposes of clarity where appropriate.

[0042] In an example embodiment of the present invention, a suitable system and / or platform for implementing the methods and techniques herein may be an entertainment device.

[0043] Referring now to the drawings, wherein like reference numerals designate identical or corresponding parts, Figure 1 shows an example of an entertainment device 10 which may be a computer or video game console, for example.

[0044] The entertainment device 10 comprises a central processor 20. The central processor 20 may be a single or multi core processor. The entertainment device also comprises a graphical processing unit or GPU 30. The GPU can be physically separate to the CPU, or integrated with the CPU as a system on a chip (SoC).

[0045] The GPU, optionally in conjunction with the CPU, may process data and generate video images (image data) and optionally audio for output via an AV output. Optionally, the audio may be generated in conjunction with or instead by an audio processor (not shown).

[0046] The video and optionally the audio may be presented to a television or other similar device. Where supported by the television, the video may be stereoscopic. The audio may be presented to a home cinema system in one of a number of formats such as stereo, 5.1 surround sound or 7.1 surroundAttorney Docket No.: 116457-1550002

[0047] Client Ref. No.: SYP356379WO01 sound. Video and audio may likewise be presented to a head mounted display unit 120 worn by a user 1.

[0048] The entertainment device also comprises RAM 40, and may have separate RAM for each of the CPU and GPU, and / or may have shared RAM. The or each RAM can be physically separate, or integrated as part of an SoC. Further storage is provided by a disk 50, either as an external or internal hard drive, or as an external solid state drive, or an internal solid state drive.

[0049] The entertainment device may transmit or receive data via one or more data ports 60, such as a USB port, Ethernet® port, Wi-Fi® port, Bluetooth® port or similar, as appropriate. It may also optionally receive data via an optical drive 70.

[0050] Audio / visual outputs from the entertainment device are typically provided through one or more A / V ports 90, or through one or more of the wired or wireless data ports 60.

[0051] An example of a device for displaying images output by the entertainment device is the head mounted display ‘HMD’ 120 worn by the user 1. The images output by the entertainment device may be displayed using various other devices - e.g. using a conventional television display connected to A / V ports 90.

[0052] Where components are not integrated, they may be connected as appropriate either by a dedicated data link or via a bus 100.

[0053] Interaction with the device is typically provided using one or more handheld controllers 130, 130A and / or one or more VR controllers 130A-L,R in the case of the HMD. The user typically interacts with the system, and any content displayed by, or virtual environment rendered by the system, by providing inputs via the handheld controllers 130, 130 A. For example, when playing a game, the user may navigate around the game virtual environment by providing inputs using the handheld controllers 130, 130A.

[0054] In embodiments of the present disclosure, the entertainment device 10 generates one or more images of a virtual environment for display (e.g. via a television or the HMD 120).

[0055] Figure 1 therefore provides an example of a data processing apparatus suitable for executing an application such as a video game and generating images for the video game for display. Images may be output via a display device such as a television or other similar monitor and / or an HMD (e.g. HMD 120). More generally, user inputs can be received by the data processing apparatus and an instance of a video game can be executed accordingly with images being rendered for display to the user.Attorney Docket No.: 116457-1550002

[0056] Client Ref. No.: SYP356379WO01 Embodiments of the present disclosure relate to setting a bitrate ladder for transmission of image content. The image content may be, in some examples, video game content that is streamed to or from the entertainment device of Figure 1. For example, one example of a streaming network may include one of the entertainment devices 10 as a transmitting device, and another one of the entertainment devices 10 as a reception device. A bitrate ladder is used during adaptive bitrate streaming, such that the image content may be made available at a plurality of different bitrates for selection by the reception device. The selection may be based on network conditions or device capabilities or other factors. Hence, when generating one or more image frames for image content in dependence on a bitrate ladder, the image frames may be generated at a plurality of different resolutions and / or framerates to be transmitted at each bitrate specified in the bitrate ladder. In accordance with the present disclosure, the image processing method includes estimating motion information in a current scene of the image content for setting a different bitrate ladder during the transmission of the image content. Hence, during transmission of the image content, one or more first image frames are generated in dependence on a first bitrate ladder, while one or more second image frames (e.g. next image frames after the first image frames) are generated in dependence on a second bitrate ladder. By dynamically selecting the second bitrate ladder, the second image frames may be generated in a way that is more suitable for transmission. Further, dynamically selecting the bitrate ladder based on motion estimation allows making more efficient use of the available bandwidth and providing improved image quality at the receiving device. In some examples, the current scene may be a current frame, or set of preceding frames including the current frame. In other examples, the current scene may be predefined sections of the image content itself. For example, sequences of frames may be associated with a scene identifier, such that a first scene has a first set of characteristics (e.g. lighting, movement, environment, etc), whereas a second scene has a second, different set of characteristics. For example, a first scene may be a dynamic action scene including more movement, whereas a second scene may be a static dialogue scene with less movement. According to the present disclosure, the motion information is indicative of movement of one or more objects in the current scene. The movement of objects may be movement of those objects through an (virtual) environment (e.g. a game virtual environment) or may be the relative movement of those objects relative to a moving camera or viewpoint. Based on the motion information, a second bitrate ladder is selected that is suitable for transmitting image content (or at least more suitable for transmitting the current scene than the first bitrate ladder).Attorney Docket No.: 116457-1550002

[0057] Client Ref. No.: SYP356379WO01 The image processing method may therefore be embodied as shown in Figure 2 illustrating a sequence of steps 200. At step 210, the image frames of the image content are generated in dependence on a currently selected (first) bitrate ladder. In step 220, the motion information in the current scene is estimated and in step 230, another (second) bitrate ladder is selected in dependence on the motion information. The method then loops back to step 210 to resume the generation of image frames in dependence on the newly selected (second) bitrate ladder. It will be appreciated that the sequence 200 may be repeated as necessary to re-select new bitrate ladders as various scenes of the image content are being transmitted. In some cases, generating the image frames at step 210 may comprise transmitting the image frames to a receiving device, e.g. over a network. In some examples, the movement of the one or more objects in the current scene may be compared to a movement threshold, and the second bitrate ladder is selected in dependence on such a comparison. In particular, where the movement of one or more objects is low (e.g. the movement threshold is not exceeded), then the second bitrate ladder may be selected to be more appropriate for scenes with less movement. For example, when there is less movement in the scene, encoding of the image data can be made more efficient by re-using information from previous frames. Thus, a bitrate ladder that assigns higher image resolution to a given available network bandwidth (i.e. bandwidth for transmission of image content to a receiving device) may be selected. Such a bitrate ladder may therefore cause the second image frames to be rendered at a higher resolution for a given available bandwidth. In this way, for example images may continue being rendered at a high resolution (e.g. 1920 x 1080 pixels) even as bandwidth is reduced; as compared to a first bitrate ladder which may reduce resolution (e.g. to 1280 x 720 pixels) for that same network bandwidth. Additionally or alternatively, less movement may allow for a lower framerate to still result in acceptable perceived image quality. Hence, the second bitrate ladder may also involve lowering the framerate to permit higher resolution for the image content.

[0058] Conversely, where the movement of one or more objects is high (e.g. the movement threshold is exceeded), then the second bitrate ladder may be selected to be more appropriate for scenes with more movement. Such a bitrate ladder may cause the second image frames to be rendered at a lower resolution for a given available bandwidth, which may allow for a higher or sustained framerate. By maintaining the higher framerate, the movement of the objects can be more easily conveyed to the viewer, thereby improving the viewer’s perceived quality of the current scene in the image content.

[0059] For efficient selection of the bitrate ladders, a plurality of candidate bitrate ladders may be predetermined, from which the second bitrate ladder is selected in dependence on the motionAttorney Docket No.: 116457-1550002

[0060] Client Ref. No.: SYP356379WO01 information. For example, each of the plurality of candidate bitrate ladders may be associated with a value indicative of motion information. When selecting from the candidate bitrate ladders, each associated value may be compared to the estimated motion information from the current scene until a match (or a closest-match) is determined.

[0061] Selection of one of the candidate bitrate ladders may also be performed by a machine learning model. In particular, a machine learning model may be trained to determine, for a plurality of degrees of motion in the scene, a bitrate ladder that optimises quality of frames generated using the bitrate ladder. Some examples of such a machine learning model may use reinforcement learning, which is a type of machine learning directed to training an artificial intelligence agent to take actions in an environment that maximize the notion of a cumulative reward. During reinforcement learning, the agent interacts with the environment, and learns from the results of its actions, thus allowing the agent to progressively improve its decision-making. Some examples may use this to optimise a cost function of perceived quality at different bandwidths, so that a given available bandwidth can be used to select an appropriate bitrate ladder from among the candidate bitrate ladders to improve the perceived quality. Further detail on reinforcement learning and other example configurations of machine learning models will be discussed later in this disclosure.

[0062] According to these examples, the selection of a new bitrate ladder based on a current scene of the image content therefore allows for improved perceived quality of the image content by a viewer. It will be appreciated that selecting a new bitrate ladder more frequently will allow the bitrate ladder to be more suitable for more of the image content, since the bitrate ladder will be selected according to what is currently being transmitted in the current scene. The frequency of selection may be controlled by performing the selecting step in response to generation of an update trigger while transmitting the one or more first image frames over a network.

[0063] The update triggers may be generated in a number of different ways. In some examples, the update trigger may be generated in the reception (i.e. receiving) device that is receiving the image content over the network. In particular, the reception device may detect a degradation in quality of the image content when using a current bitmap ladder, and therefore generates an update trigger to request the transmitting device to select a new bitrate ladder. Alternatively, in examples where the image content relates to a video game, the player using the reception device may provide an input that is recognised as being unsuitable for a current bitrate ladder. For example, if the player opens a pause menu during an action gameplay sequence, an update trigger may be generated to cause the selection of a bitrate ladder that is better for transmitting image frames of the pause menu (i.e.Attorney Docket No.: 116457-1550002

[0064] Client Ref. No.: SYP356379WO01 little or no movement), rather than continuing with the bitrate ladder currently selected for the action gameplay sequence. User inputs that generate the update trigger (e.g. entering or leaving pause menu) may be predetermined.

[0065] In other examples, the update trigger is generated in the transmitting device that is transmitting the image content over the network. Update triggers may be generated at predetermined sampling points in the image content. The predetermined sampling points may be distributed in various different ways. Some implementations of distributing sampling points may be used regardless of the content being transmitted, e.g. a “interval-based” implementation. In such an example, predetermined sampling points may be distributed at regular time intervals throughout the image content. For example, the regular time interval may be based on a predetermined number of frames; for instance, the sampling points may occur every N (e.g. 2, 5, 10, or 100) frames in the image content.

[0066] Figure 3 A illustrates an example of image content 310 comprising three different scenes (each of which may be associated with different environments, lighting, characters, action, etc).

[0067] The image content 310 is transmitted starting with scene l. The frames of scene l are generated in dependence on the first bitrate ladder, which was selected at the beginning of scene l in response to the predetermined sampling point 312. For example, scene l may relate to a stealth gameplay sequence, in which objects are estimated to be moving relatively slowly. Hence, the bitrate ladder selected at the predetermined sampling point 312 may be selected to provide higher resolutions and / or lower framerates at a given bandwidth.

[0068] After a time interval (tl), which may be defined by a set number of frames of the image content, the predetermined sampling point 314 is encountered, triggering the estimation of the motion information in the current scene for selection of a new bitrate ladder. It will be appreciated that since scene l is still the current scene, the newly selected bitrate ladder may be substantially the same, or similar, as the bitrate ladder selected at predetermined sampling point 312.

[0069] After another time interval (t2), the predetermined sampling point 316 is encountered during the transmission of scene_2, thereby triggering the estimation of the motion information in scene_2 for selection of a new bitrate ladder. For example, scene_2 may relate to an action gameplay sequence, in which objects are estimated to be moving relatively quickly. Hence, the bitrate ladder selected at the predetermined sampling point 316 may be selected to provide a lower resolution and higher framerate at a given bandwidth. It will be appreciated that the predetermined sampling point 316 occurs part-way through scene_2. Hence, some of scene_2 may have already beenAttorney Docket No.: 116457-1550002

[0070] Client Ref. No.: SYP356379WO01 transmitted using a less suitable bitrate ladder (i.e. the bitrate ladder selected for scene l at predetermined sampling point 314). Nonetheless, by selecting a new bitrate ladder at predetermined sampling point 316, the remainder of scene_2 can still be transmitted using a more suitable bitrate ladder.

[0071] After another time interval (t3), the predetermined sampling point 318 is encountered during transmission of scene_3, thereby triggering the estimation of the motion information in scene_3 for selection of a new bitrate ladder. For example, scene_3 may be a dialogue cinematic scene, which may feature close-ups of characters’ faces and less movement of objects. Accordingly, the bitrate ladder selected at the predetermined sampling point 318 may be selected to provide a higher resolution and / or lower framerate at a given bandwidth.

[0072] In the example of Figure 3A, the selection of the bitrate ladder based on estimated motion information at regular intervals of the image content allows for the bitrate ladder to be more suitable for the image content more often. Hence, the perceived quality of the transmitted image content is improved. It will be appreciated however, that selecting new bitrate ladders more regularly (i.e. with a reduced time interval) would reduce the risk of frames being transmitted using a less suitable bitrate ladder. For example, if predetermined sampling point 316 were encountered earlier in scene_2, then fewer frames of scene_2 would be transmitted using a less suitable bitrate ladder.

[0073] In another implementation, the predetermined sampling points may be positioned between scenes of the image content, e.g. at the point where one scene ends and the next scene begins. Some examples of a “scene-based” implementation may require some analysis of the image content prior to transmission to identify the points where one scene ends and the next scene begins. Alternatively, a significant change in characteristics between scenes of the image content (e.g. environment, lighting, etc) may be detected in real time while the frames are being transmitted, thereby marking the point where one scene ends and the next scene begins. Figure 3B illustrates an example of image content 320 comprising the same three scenes as in image content 310 above. As above, since the scenes are associated with different environments, lighting, characters, action, etc, it can be anticipated the level of movement between scenes may vary, in which case the bitrate ladder can be adjusted in accordance with the present disclosure.

[0074] The image content 320 is transmitted starting with scene l. The frames of scene l are generated in dependence on the first bitrate ladder, which was selected at the beginning of scene l in response to the predetermined sampling point 322. As above, scene l may relate to a stealth gameplay sequence, in which objects are estimated to be moving relatively slowly. Hence, theAttorney Docket No.: 116457-1550002

[0075] Client Ref. No.: SYP356379WO01 bitrate ladder selected at the predetermined sampling point 322 may be selected to provide a higher resolution and / or lower framerate at a given available bandwidth. For example, the resolution of scene_l may be set to 1920 x 1080 pixels with a framerate of 30 frames per second.

[0076] As the transmission reaches the end of scene l, the predetermined sampling point 324 is encountered which indicates the beginning of scene_2 and hence causes a new bitrate ladder to be selected as described above. As above, scene_2 may relate to an action gameplay sequence, in which objects are estimated to be moving relatively quickly. Hence, the bitrate ladder selected at the predetermined sampling point 324 may be selected to provide a lower resolution and / or higher framerate at a given available bandwidth than the bitrate ladder used for scene l. For example, the resolution of scene_2 may be set to 1280 x 720 pixels with a framerate of 60 frames per second. As the transmission reaches the end of scene_2, the predetermined sampling point 326 is encountered which indicates the beginning of scene_3 and hence causes a new bitrate ladder to be selected as described above. As above, scene_3 may be a dialogue cinematic scene, which may feature close-ups of characters’ faces and less movement of objects. Accordingly, the bitrate ladder selected at the predetermined sampling point 318 may be selected to provide a higher resolution and / or lower framerate at a given available bandwidth than the bitrate ladder used for scene_2. For example, the resolution of scene_3 may be set to 2560 x 1440 pixels at 25 frames per second. It will therefore be appreciated that the scene-based implementation allows for more timely selection of a new bitrate ladder, which means that most or all of each scene is transmitted using a suitable bitrate ladder. Accordingly, the perceived quality of the transmitted image content is improved. Meanwhile, the interval-based implementation allows for greater flexibility because it can be used with any image content without knowledge of where different scenes begin and end. It will also be appreciated that both of the scene-based and interval-based implementations may be used in combination. For example, a predetermined sampling point may be positioned at the beginning of each scene as shown in Figure 3B, with additional predetermined sampling points positioned throughout each scene at regular time intervals, similar to that shown in Figure 3A. Accordingly, these examples provide several configurations in which the image processing method may control the timing for estimating the motion information and selecting a new bitrate ladder. The way in which the motion information is estimated may also vary between examples. In some examples, the motion information may be estimated in dependence on motion vectors associated with one or more preceding frames. Figure 4A illustrates two example frames of the image content: a preceding frame 410 and a current frame 420. The frames include an object 412 which is moving,Attorney Docket No.: 116457-1550002

[0077] Client Ref. No.: SYP356379WO01 and hence appears in a different position in each frame. This movement is defined by a motion vector 414 which describes the transformation of the object 412 between one frame and the next frame. It will be appreciated that the motion vector 414 may be a summation of a plurality of motion vectors, each attributed to individual pixels of the object 412. The motion vector 414 may be generated in a number of different ways. In some examples, motion vectors may be generated as part of image encoding while transmitting preceding frames, e.g. using the MPEG-4 video compression standard. Where the image content relates to a video game, the motion vectors may be acquired directly from the game state data. Motion vectors may for example be obtained from velocity buffers. The motion vector 414 may therefore be used to estimate the motion information in the current frame 420 for selecting a suitable bitrate ladder regardless of what type of image content is being transmitted. This has the advantage of estimating the motion information in real time without any preceding analysis of the content.

[0078] In other examples, further information that may be derived from the image content itself may also be used for selection of a bitrate ladder.

[0079] In some examples, scene metadata may be stored in association with each scene, comprising a scene classification and the motion information associated with each scene classification. Figure 4B illustrates one such set of scene metadata 430 associated with the scenes described in Figures 3A and 3B.

[0080] In the scene metadata 430, scene l has been classified as “stealth gameplay”, which corresponds to “medium” motion. For example, it may be expected that there will be less movement relative to “action gameplay”, and a bitrate ladder may be selected to provide a higher resolution. Similarly, scene_2 has been classified as “action gameplay” with “high” motion, hence a bitrate ladder may be selected to provide a higher framerate by reducing the resolution. Lastly, scene_3 has been classified as “dialogue cinematic” with “low” motion, hence a bitrate ladder may be selected to provide higher resolution.

[0081] Accordingly, in some examples, when the predetermined sampling points 322, 324 and 326 are encountered, the scene metadata 430 may be looked up to estimate the motion information that has been predefined for each type of scene classification.

[0082] The scene metadata 430 may be obtained from a developer of the image content or from quality assurance (QA) testing of the image content. In some examples, where the image content relates to a video game, sections of the video game (corresponding to the scenes of image content) may be tagged with a scene classification by a game developer. For example, while developing theAttorney Docket No.: 116457-1550002

[0083] Client Ref. No.: SYP356379WO01 game, the developer knows that scene_3 is to be a dialogue cinematic scene, and so the scene classifications may be pre-set in the image content.

[0084] Alternatively or additionally, sections of the video game may be tagged with an amount of motion detected during test gameplay in QA testing or by early play-throughs of a game. During test gameplay, the image transmission quality may be monitored to identify sections of gameplay in which a higher or lower resolution would be suitable based on detected motion information (e.g. using the motion vectors described above). Those sections of the video game may then be tagged as being associated with particular amounts of movement thereby generating scene metadata. Figure 4C illustrates an example where the scene metadata 440 comprises scenes that are tagged based on the detected motion information that may be looked up in similar ways as with the scene metadata 430.

[0085] In some examples, when generating the scene metadata 430 or 440, the different scenes may be identified based on game state data. In some examples, the game state data comprises a set of data values, which are indicative of a current state of the gameplay that is being experienced by a player. For example, the game state data may comprise one or more of a player’s viewpoint or location in a game environment, a player’s movement through a game environment, the number of objects in the game environment, and the lighting in the game environment. This information may be used to distinguish between various scenes to utilise the scene metadata 430 or 440 for more accurate estimation of the motion information.

[0086] Figure 5 A illustrates an example where a player 510 may rotate their viewpoint between a region of low movement 512 and a region of high movement 514. Although a current gameplay sequence may be an action sequence with a high amount of movement, if the player 510 faces away from the moving objects in the region 514, then the current scene of the image content (as viewed in the region 512) may not actually contain any moving objects. Accordingly, the player’s viewpoint may be used to distinguish between different scenes in the scene metadata. For example, the region 512 may be associated (e.g. by a developer or by QA tests as described above) with a particular scene having less motion in scene metadata. Hence, when selecting a new bitrate ladder, the scene metadata may be looked up to identify that the scene is associated with less movement, and the new bitrate ladder may be selected for higher resolution. Similarly, the region 514 may be associated with another scene having more movement. Hence, when selecting a new bitrate ladder, the metadata 440 is looked up to identify that the scene is associated with more movement, and the new bitrate ladder may be selected for higher framerate.Attorney Docket No.: 116457-1550002

[0087] Client Ref. No.: SYP356379WO01 Figure 5B illustrates an example where the player 520 may move from the outside to the inside of an enclosed space 526 such as a building. In this example, the outside of the enclosed space 526 is associated with a region of higher movement 522 due to other objects moving through that environment, while inside of the enclosed space 526 is associated with a region of lower movement 524. Therefore, the location of the player 520 may be used in the same way as the above to distinguish between different scenes in the image content and the bitrate ladder may be selected accordingly.

[0088] It will be appreciated that other data values of the game state data mentioned above may also be used in a similar way. In particular, the movement of the player through the environment and / or the number of objects present in the environment may be used to identify the amount of movement between scenes. The lighting in an environment may also be used to identify how much movement would be visible to the viewer.

[0089] Various aspects of the present disclosure may be implemented using a machine learning model, such as a supervised machine learning model. In some examples, a machine learning model may be trained to generate a dynamicity value indicative of movement of the one or more objects in the current scene in dependence on one or more sample frames of the current scene. The sample frames may be a current frame and / or one or more frames immediately preceding the current frame which are input to the machine learning model to assess the movement of the objects and generate the dynamicity value. In yet further examples, some frame information may be available prior to the frames actually being generated, hence the sample frames may also include at least some information relating to a frame that follows the current frame. The dynamicity value may then be used as an estimation of the motion information for selection of a bitrate ladder as in previous examples. Indeed, the selection may use a further machine learning model trained to identify an ‘optimised’ bitrate ladder for a given amount of movement (e.g. the dynamicity value or the scene metadata described previously). It will be appreciated that ‘optimised’ in this context does not necessarily require that the selected bitrate ladder is the most optimal bitrate ladder that is possible. In some examples, the machine learning model may be trained to identify a bitrate ladder that is more optimised than the current bitrate ladder, for transmission of the current scene.

[0090] The supervised learning model is trained using labelled training data to learn a function that maps inputs (typically provided as feature vectors) to outputs (i.e. labels). The labelled training data comprises pairs of inputs and corresponding output labels. The output labels are typically provided by an operator to indicate the desired output for each input. The supervised learning modelAttorney Docket No.: 116457-1550002

[0091] Client Ref. No.: SYP356379WO01 processes the training data to produce an inferred function that can be used to map new (i.e. unseen) inputs to a label.

[0092] The input data (during training and / or inference) may comprise various types of data, such as numerical values, images, video, text, or audio. Raw input data may be pre-processed to obtain an appropriate feature vector used as input to the model - for example, features of an image or audio input may be extracted to obtain a corresponding feature vector. It will be appreciated that the type of input data and techniques for pre-processing of the data (if required) may be selected based on the specific task the supervised learning model is used for. In the example of the motion estimation model, the input data may one or more frames of various scenes of the image content, whereas for the bitrate ladder selection model, the input data may be the motion information.

[0093] Once prepared, the labelled training data set is used to train the supervised learning model. During training the model adjusts its internal parameters (e.g. weights) so as to optimize (e.g. minimize) an error function, aiming to minimize the discrepancy between the model’s predicted outputs and the labels provided as part of the training data. In some cases, the error function may include a regularization penalty to reduce overfitting of the model to the training data set.

[0094] The supervised learning model may use one or more machine learning algorithms in order to learn a mapping between its inputs and outputs. Example suitable learning algorithms include linear regression, logistic regression, artificial neural networks, decision trees, support vector machines (SVM), random forests, and the K-nearest neighbour algorithm.

[0095] Once trained, the supervised learning model may be used for inference - i.e. for predicting outputs for previously unseen input data. The supervised learning model may perform classification and / or regression tasks. In a classification task, the supervised learning model predicts discrete class labels for input data, and / or assigns the input data into predetermined categories, such as the scene classifications described previously (for a motion estimation model) assigns the scene to an ‘optimised’ bitrate ladder. In a regression task, the supervised learning model predicts labels that are continuous values, such as a dynamicity value indicative of an amount of movement of the objects in the current scene.

[0096] In some cases, limited amounts of labelled data may be available for training of the model (e.g. because labelling of the data is expensive or impractical). In such cases, the supervised learning model may be extended to further use unlabelled data and / or to generate labelled data.

[0097] Considering using unlabelled data, the training data may comprise both labelled and unlabelled training data, and semi-supervised learning may be used to learn a mapping between the model’sAttorney Docket No.: 116457-1550002

[0098] Client Ref. No.: SYP356379WO01 inputs and outputs. For example, a graph-based method such as Laplacian regularization may be used to extend a SVM algorithm to Laplacian SVM in order to perform semi -supervised learning on the partially labelled training data.

[0099] Considering generating labelled data, an active learning model may be used in which the model actively queries an information source (such as a user, or operator) to label data points with the desired outputs. Labels are typically requested for only a subset of the training data set thus reducing the amount of labelling required as compared to fully supervised learning. The model may choose the examples for which labels are requested - for example, the model may request labels for data points that would most change the current model, or that would most reduce the model's generalization error. Semi-supervised learning algorithms may then be used to train the model based on the partially labelled data set.

[0100] In contrast to the supervised learning model, a reinforced learning (RL) model typically comprises an action-reward feedback loop. The feedback loop comprises: an environment, state, agent, policy, action, and reward. The environment is the system with which the agent interacts and in which the agent operates - for example, the environment may be a virtual environment of a game. The state represents the current conditions in the environment. The agent receives the state as an input and takes an action which may affect the environment and change the state of the environment. The agent takes the action based on its policy which is a mapping from states of the environment to actions of the agent. The policy may be deterministic or stochastic. The reward represents feedback from the environment to the action taken by the agent. The reward provides an indication (typically in the form of a numerical value) of the desirability of the result of the agent’s action. The reward may comprise positive signals to reward desirable behaviour of the agent and / or negative signals to penalize undesirable behaviour of the agent. For example, selection of a bitrate ladder resulting in higher perceived quality of the image content would be desirable behaviour, whereas selection of a bitrate ladder resulting in lower perceived quality of the image content would be undesirable behaviour.

[0101] Through multiple iterations of action-reward feedback loop, the agent aims to maximise the total cumulative reward it receives, thus learning how to take optimal actions in the environment. The reinforcement learning process thus allows the agent to learn an optimal policy that maximizes the cumulative reward. The cumulative award may be estimated using a value function which estimates the expected return starting from a given state or from a given state and action. Using the cumulative reward in the reinforcement learning process allows the agent to consider longterm effects of its policy.Attorney Docket No.: 116457-1550002

[0102] Client Ref. No.: SYP356379WO01 A reinforcement learning algorithm may be used to refine the agent’ s policy and the value function over iterations of the action-reward feedback loop. The learning algorithm may rely on a model of the environment (e.g. based on Markov Decision Processes (MDPs)) or be model-free. Example suitable model-free reinforcement learning algorithms include Q-leaming, SARSA (State- Action-Reward-State-Action), Deep Q-Networks (DQNs), or Deep Deterministic Policy Gradient (DDPG).

[0103] It will be appreciated that the agent will typically engage in both exploration and exploitation of the environment in which it operates. In exploration, the agent takes typically random actions to gather information about the environment and identify potentially desirable actions (i.e. actions that maximise cumulative reward). In exploitation, the agent takes actions that are expected to maximise reward (e.g. by selecting the action based on the agent’s latest policy). Various techniques may be used to control the proportion of explorative and exploitative actions taken by the agent - for example, a predetermined probability of taking an explorative action in a given iteration of the feedback loop may be set (and optionally reduced over time to allow the agent to shifts more towards exploitation over time to maximise cumulative reward in view of diminishing returns for further exploration).

[0104] In some cases, the RL model may be configured to learn from feedback provided by a user. Utilising user feedback in this way may allow the agent to improve its choice of actions and better align with user preferences. For example, reinforcement learning from human feedback (RLHF) techniques may be used. RLHF includes training a reward model based on user feedback and using this model for determining the reward in the reinforcement learning process described above. The user feedback may be received in various forms depending on the specific reinforcement learning problem being solved - for example, the feedback may be received in the form of a user ranking of instances of the agent’s actions. RLHF thus allows incorporating user feedback into the reinforcement learning process. RLHF approaches may be advantageous where it is easier for a user than for an algorithm to assess the quality of the machine learning model’s output (e.g. for generative artificial intelligence RL models).

[0105] It will be appreciated that the methods described above may be carried out on conventional hardware suitably adapted as applicable by software instruction or by the inclusion or substitution of dedicated hardware. Thus, the required adaptation to existing parts of a conventional equivalent device may be implemented in the form of a computer program product comprising processor implementable instructions stored on a non-transitory machine-readable medium such as a floppy disk, optical disk, hard disk, solid state disk, PROM, RAM, flash memory or any combination ofAttorney Docket No.: 116457-1550002

[0106] Client Ref. No.: SYP356379WO01 these or other storage media, or realised in hardware as an ASIC (application specific integrated circuit) or an FPGA (field programmable gate array) or other configurable circuit suitable to use in adapting the conventional equivalent device. Separately, such a computer program may be transmitted via data signals on a network such as an Ethernet, a wireless network, the Internet, or any combination of these or other networks.

[0107] It will be appreciated that the methods may be carried out using an image processing system for setting a bitrate ladder during transmission of image content. Figure 6 illustrates an example of an image processing apparatus 600 in accordance with one or more embodiments of the present disclosure.

[0108] The image processing apparatus 600 may be provided as part of a user device (such as the entertainment device of Figure 1) and / or as part of a server device. The image processing apparatus 600 may be implemented in a distributed manner using two or more respective processing devices that communicate via a wired and / or wireless communications link. The image processing apparatus 600 may be implemented as a special purpose hardware device or a general-purpose hardware device operating under suitable software instruction. The image processing apparatus 600 may be implemented using any suitable combination of hardware and software.

[0109] The image processing apparatus 600 comprises a generation processor 610, an estimation processor 620, and a selection processor 630. The operations discussed throughout the present disclosure may be performed by the image processing apparatus 600. In particular, the generation processor 610 performs operations including generating one or more images for the image content in dependence on a selected bitrate ladder. The estimation processor 620 performs operations including estimating motion information in a current scene of the image content where the motion information is indicative of movement of one or more objects in the current scene. The selection processor 630 performs operations including selecting a new bitrate ladder in dependence on the motion information. Accordingly, the generation processor 610 then generates one or more image frames for the image content in dependence on the new bitrate ladder instead of the previous bitrate ladder.

[0110] Of course, the functionality of these processors may be realised by any suitable number of processors located at any suitable number of devices as appropriate rather than requiring a one-to-one mapping between the functionality and a device or processor.

[0111] Much of the above disclosure is relevant for transmitting image content over a network. In other embodiments of the present disclosure, there is provided a complementary image processingAttorney Docket No.: 116457-1550002

[0112] Client Ref. No.: SYP356379WO01 method for receiving image content. The method comprises receiving one or more first image frames for image content in dependence on a first bitrate ladder, and (subsequently) receiving one or more second image frames for the image content in dependence on a second bitrate ladder selected in dependence on motion information indicative of movement of one or more objects in a current scene.

[0113] The image processing method may therefore be embodied as shown in Figure 7 illustrating a sequence of steps 700. At step 710, the image frames (i.e. of the image content) are received in dependence on the first bitrate ladder. It will be appreciated that in a working system, the first bitrate ladder may be that which is selected for performance of step 210 described in Figure 2 above. At step 720, there is an optional step of generating a bitrate ladder update trigger, e.g. at the reception device. As described in previous examples, while the selection of a bitrate ladder may be predominantly performed by a transmitter device, the selection may nonetheless be triggered by the reception device, for example in response to detecting degradation image quality of the received image frames, or specific user inputs. However, in other examples, the selection is triggered by the transmitting device itself, for example in response to encountering the predetermined sampling points of Figures 3 A and 3B. Regardless of how the selection is triggered, once a new (second) bitrate ladder has been selected, step 730 includes receiving the image frames in dependence on the second bitrate ladder. At this point, the method may optionally repeat steps 720 and step 730 as appropriate, whenever another bitrate ladder is to be selected.

[0114] The techniques described above may be implemented in hardware, software or combinations of the two. In the case that a software-controlled data processing apparatus is employed to implement one or more features of the embodiments, it will be appreciated that such software, and a storage or transmission medium such as a non-transitory machine-readable storage medium by which such software is provided, are also considered as embodiments of the disclosure.

[0115] The foregoing discussion discloses and describes merely exemplary embodiments of the present invention. As will be understood by those skilled in the art, the present invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. Accordingly, the disclosure of the present invention is intended to be illustrative, but not limiting of the scope of the invention, as well as other claims. The disclosure, including any readily discernible variants of the teachings herein, defines, in part, the scope of the foregoing claim terminology such that no inventive subject matter is dedicated to the public.

Claims

Attorney Docket No.: 116457-1550002Client Ref. No.: SYP356379WO01 CLAIMSWHAT IS CLAIMED IS:

1. An image processing method for setting a bitrate ladder during transmission of image content, the image processing method comprising:generating one or more first image frames for image content in dependence on a first bitrate ladder;estimating motion information in a current scene of the image content, the motion information being indicative of movement of one or more objects in the current scene;selecting a second bitrate ladder in dependence on the motion information; and generating one or more second image frames for the image content in dependence on the second bitrate ladder instead of the first bitrate ladder.

2. The image processing method of claim 1, wherein selecting the second bitrate ladder in dependence on the motion information comprises comparing the movement of one or more objects in the current scene to a movement threshold.

3. The image processing method of claim 2, wherein in response to determining that the movement of the one or more objects in the current scene does not exceed the movement threshold, the second bitrate ladder is selected such that the one or more second image frames are rendered at a higher resolution than the first bitrate ladder for the same available bandwidth.

4. The image processing method of claim 2 or claim 3, wherein in response to determining that the movement of the one or more objects in the current scene exceeds the movement threshold, the second bitrate ladder is selected such that the one or more second image frame are rendered at a lower resolution than the first bitrate ladder for the same available bandwidth.

5. The image processing method of any preceding claim, wherein selecting the second bitrate ladder comprises selecting one of a plurality of candidate bitrate ladders in dependence on the motion information.

6. The image processing method of claim 5, wherein each of the plurality of candidate bitrate ladders is associated with a different degree of motion in the scene, and wherein the plurality of candidate bitrate ladders are generated using a machine learning model trained to determine, forAttorney Docket No.: 116457-1550002Client Ref. No.: SYP356379WO01 a plurality of degrees of motion in the scene, a bitrate ladder that optimises quality of frames generated using the bitrate ladder.

7. The image processing method of any preceding claim, wherein the selecting is performed in response to generation of an update trigger while transmitting the one or more first image frames over a network.

8. The image processing method of claim 7, comprising generating the update trigger at predetermined sampling points in the image content.

9. The image processing method of claim 8, wherein the predetermined sampling points are between each of a plurality of scenes of the image content.

10. The image processing method of claim 8, wherein the predetermined sampling points are at a regular time interval during the image content.

11. The image processing method of claim 10, wherein the regular time interval is based on a predetermined number of frames.

12. The image processing method of any preceding claim, wherein estimating motion information in the current scene is performed in dependence on motion vectors associated with one or more preceding frames.

13. The image processing method of claim 12, comprising obtaining the motion vectors from information stored in velocity buffers.

14. The image processing method of any preceding claim, wherein estimating motion information in the current scene is performed in dependence on scene metadata associated with the current scene.

15. The image processing method of claim 14, wherein the image content is a video game, and the method comprises identifying the scene metadata based on game state data indicative of a current gameplay state of the video game.

16. The image processing method of claim 15, wherein identifying the scene metadata comprises comparing the game state data to recorded game state data associated with one or more previous instances of the video game being played, the recorded game state data being associated with the scene metadata.Attorney Docket No.: 116457-1550002Client Ref. No.: SYP356379WO01 17. The image processing method of claim 15 or claim 16, wherein the game state data comprises one or more of:player viewpoint;player location;player movement;number of objects; andenvironment lighting.

18. The image processing method of any of claims 1 to 11, wherein estimating motion information in the current scene is performed using a machine learning model trained to generate a dynamicity value indicative of movement of the one or more objects in the current scene in dependence on one or more sample frames of the current scene.

19. An image processing system for setting a bitrate ladder during transmission of image content, the image processing system comprising:a generation processor configured to generate one or more first image frames for image content in dependence on a first bitrate ladder;an estimation processor configured to estimate motion information in a current scene of the image content, the motion information being indicative of movement of one or more objects in the current scene; anda selection processor configured to select a second bitrate ladder in dependence on the motion information;wherein the generation processor is configured to generate one or more second image frames for the image content in dependence on the second bitrate ladder instead of the first bitrate ladder.

20. The image processing system of claim 19, wherein selecting the second bitrate ladder in dependence on the motion information comprises comparing the movement of one or more objects in the current scene to a movement threshold.

21. The image processing system of claim 20, wherein in response to determining that the movement of the one or more objects in the current scene does not exceed the movement threshold, the second bitrate ladder is selected such that the one or more second image frames are rendered at a higher resolution than the first bitrate ladder for the same available bandwidth.Attorney Docket No.: 116457-1550002Client Ref. No.: SYP356379WO01 22. The image processing system of claim 20 or claim 21, wherein in response to determining that the movement of the one or more objects in the current scene exceeds the movement threshold, the second bitrate ladder is selected such that the one or more second image frame are rendered at a lower resolution than the first bitrate ladder for the same available bandwidth.

23. The image processing system of any preceding claim, wherein selecting the second bitrate ladder comprises selecting one of a plurality of candidate bitrate ladders in dependence on the motion information.

24. The image processing system of claim 23, wherein each of the plurality of candidate bitrate ladders is associated with a different degree of motion in the scene, and wherein the plurality of candidate bitrate ladders are generated using a machine learning model trained to determine, for a plurality of degrees of motion in the scene, a bitrate ladder that optimises quality of frames generated using the bitrate ladder.

25. The image processing system of any preceding claim, wherein the selecting is performed in response to generation of an update trigger while transmitting the one or more first image frames over a network.

26. The image processing system of claim 25, comprising generating the update trigger at predetermined sampling points in the image content.

27. The image processing system of claim 26, wherein the predetermined sampling points are between each of a plurality of scenes of the image content.

28. The image processing system of claim 26, wherein the predetermined sampling points are at a regular time interval during the image content.

29. The image processing system of claim 28, wherein the regular time interval is based on a predetermined number of frames.

30. The image processing system of any preceding claim, wherein estimating motion information in the current scene is performed in dependence on motion vectors associated with one or more preceding frames.

31. The image processing system of claim 30, comprising obtaining the motion vectors from information stored in velocity buffers.Attorney Docket No.: 116457-1550002Client Ref. No.: SYP356379WO01 32. The image processing system of any preceding claim, wherein estimating motion information in the current scene is performed in dependence on scene metadata associated with the current scene.

33. The image processing system of claim 32, wherein the image content is a video game, and the method comprises identifying the scene metadata based on game state data indicative of a current gameplay state of the video game.

34. The image processing system of claim 33, wherein identifying the scene metadata comprises comparing the game state data to recorded game state data associated with one or more previous instances of the video game being played, the recorded game state data being associated with the scene metadata.

35. The image processing system of claim 33 or claim 34, wherein the game state data comprises one or more of:player viewpoint;player location;player movement;number of objects; andenvironment lighting.

36. The image processing system of any of claims 19 to 29, wherein estimating motion information in the current scene is performed using a machine learning model trained to generate a dynamicity value indicative of movement of the one or more objects in the current scene in dependence on one or more sample frames of the current scene.

37. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:generating one or more first image frames for image content in dependence on a first bitrate ladder;estimating motion information in a current scene of the image content, the motion information being indicative of movement of one or more objects in the current scene;selecting a second bitrate ladder in dependence on the motion information; and generating one or more second image frames for the image content in dependence on the second bitrate ladder instead of the first bitrate ladder.Attorney Docket No.: 116457-1550002Client Ref. No.: SYP356379WO01 38. The non-transitory computer-readable medium of claim 37, wherein selecting the second bitrate ladder in dependence on the motion information comprises comparing the movement of one or more objects in the current scene to a movement threshold.

39. The non-transitory computer-readable medium of claim 38, wherein in response to determining that the movement of the one or more objects in the current scene does not exceed the movement threshold, the second bitrate ladder is selected such that the one or more second image frames are rendered at a higher resolution than the first bitrate ladder for the same available bandwidth.

40. The non-transitory computer-readable medium of claim 38 or claim 39, wherein in response to determining that the movement of the one or more objects in the current scene exceeds the movement threshold, the second bitrate ladder is selected such that the one or more second image frame are rendered at a lower resolution than the first bitrate ladder for the same available bandwidth.

41. The non-transitory computer-readable medium of any preceding claim, wherein selecting the second bitrate ladder comprises selecting one of a plurality of candidate bitrate ladders in dependence on the motion information.

42. The non-transitory computer-readable medium of claim 41, wherein each of the plurality of candidate bitrate ladders is associated with a different degree of motion in the scene, and wherein the plurality of candidate bitrate ladders are generated using a machine learning model trained to determine, for a plurality of degrees of motion in the scene, a bitrate ladder that optimises quality of frames generated using the bitrate ladder.

43. The non-transitory computer-readable medium of any preceding claim, wherein the selecting is performed in response to generation of an update trigger while transmitting the one or more first image frames over a network.

44. The non-transitory computer-readable medium of claim 43, comprising generating the update trigger at predetermined sampling points in the image content.

45. The non-transitory computer-readable medium of claim 44, wherein the predetermined sampling points are between each of a plurality of scenes of the image content.

46. The non-transitory computer-readable medium of claim 44, wherein the predetermined sampling points are at a regular time interval during the image content.Attorney Docket No.: 116457-1550002Client Ref. No.: SYP356379WO01 47. The non-transitory computer-readable medium of claim 46, wherein the regular time interval is based on a predetermined number of frames.

48. The non-transitory computer-readable medium of any preceding claim, wherein estimating motion information in the current scene is performed in dependence on motion vectors associated with one or more preceding frames.

49. The non-transitory computer-readable medium of claim 48, comprising obtaining the motion vectors from information stored in velocity buffers.

50. The non-transitory computer-readable medium of any preceding claim, wherein estimating motion information in the current scene is performed in dependence on scene metadata associated with the current scene.

51. The non-transitory computer-readable medium of claim 50, wherein the image content is a video game, and the method comprises identifying the scene metadata based on game state data indicative of a current gameplay state of the video game.

52. The non-transitory computer-readable medium of claim 51 , wherein identifying the scene metadata comprises comparing the game state data to recorded game state data associated with one or more previous instances of the video game being played, the recorded game state data being associated with the scene metadata.

53. The non-transitory computer-readable medium of claim 51 or claim 52, wherein the game state data comprises one or more of:player viewpoint;player location;player movement;number of objects; andenvironment lighting.

54. The non-transitory computer-readable medium of any of claims 37 to 47, wherein estimating motion information in the current scene is performed using a machine learning model trained to generate a dynamicity value indicative of movement of the one or more objects in the current scene in dependence on one or more sample frames of the current scene.