Signal processing system

GB2628763BActive Publication Date: 2026-08-20V NOVA INT LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
GB2023004870
Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2026-08-20
Estimated Expiration
2043-03-31

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000002_0000
    Figure 00000002_0000
  • Figure 00000002_0001
    Figure 00000002_0001
Patent Text Reader

Abstract

A system for encoding a multi-layer video stream, the system comprising: a content generator configured to output a first content relating to at least one frame of a video at a first level of quality
Need to check novelty before this filing date? Find Prior Art

Description

FIELD OF THE INVENTION The present disclosure relates generally to low latency communication systems, for example to low latency communication systems providing video communication. Moreover, the present disclosure also relates to methods of operating aforesaid low latency systems. BACKGROUND Systems that encode data, communicate encoded data and decode encoded data are known. In video streaming systems, it is highly desirable that latency between operations is as small as possible. Such latency can be influenced by a plurality of factors, and can result in a rendition of an input signal being a temporally inaccurate representation of the original input signal. Aforesaid latency arises on account of a combination of several temporal delays, namely: (i) latency arising in the encoder when encoding the input signal, for example depending on a complexity of the input signal; (ii) latency arising in the data communication network; and (iii) latency arising in the decoder when decoding the encoded data to generate corresponding data for use in generating the rendition. A conventional approach to address latency is to employ a data communication network with sufficient bandwidth and channels so that noticeable latency is not encountered under operating conditions. However, such an approach is not potentially technically feasible in many circumstances, or is prohibitively expensive. In many situations, latency may, as a priority, be more important to keep as low as possible than maintaining image resolution and / or colour resolution, for example at times when data communication network congestion becomes more severe. As such, it would be desirable to reduce temporal latency in systems, for example systems supporting interaction between users or remote control systems that are required to react quickly to changing events. Examples of such systems where low latency is a priority include extended reality systems in which it is generally desirable for content created to be visible by the end user as soon as possible. General hierarchical codecs are known, such as progression image format (PJPEG), VC-6 (standardised as ST-2117) and LCEVC. Typically these formats and coding approaches reuse information across layers and separate information between the layers themselves. It has previously been described, for example in patent publication GB2601720 that it may be desirable to send base encoded data to the decoder as soon as it is ready as a means to reduce latency. However, such a solution reduces latency only within the coding part of the pipeline. Certain ones of the scalable or hierarchical codecs are known to increase bitrate and also latency. While this is not always the case, the use of such codecs is not always desirable when latency and bitrate are the main objectives of system design. It is known in industry to address latency issues by requesting a resolution of an image or video according to the available bandwidth or bitrate. Such solutions are typically referred to as bitrate ladders. In these solutions, a set of predetermined versions of content are created, with one being chosen for transmission according to the available conditions. This may be thought of as circumventing latency issues as opposed to solutions designed to reduce latency effectively. It remains a longstanding objective to reduce latency and bitrate of image and video streaming systems. It remains an underlying problem to balance the sometimes conflicting objectives of performance and experience (e.g. bitrate, latency, quality of video and resolution); however, in certain scenarios, such as XR, latency is the overriding objective. SUMMARY OF INVENTION According to a first aspect there is provided a system for encoding a multi-layer video stream, the system comprising: a content generator configured to output a first content relating to at least one frame of a video at a first level of quality and a second content relating to the at least one frame at a second level of quality, the second level of quality being different from the first level of quality; and an encoder configured to encode the first and second contents and output encoded data. The system reduces latency due to the omission of a downsampling step. That is, the overall encoding time is reduced because the downsampling operation is not performed. Preferably, the system does not comprise a downsampler. Preferably, the system does not perform downsampling to generate either the first content or the second content. Preferably the encoder is configured to encode the first content at least partially simultaneously as the content generator outputs the second content, such that the content generator and the encoder function temporarily simultaneously when processing a given frame. Herein when we refer to simultaneously we refer to functionally simultaneously. That is, one operation may be performed before the other operation has been completed. In other words, two operations are being performed at least partly at the same time. Even if operations are performed in a serial processing manner, we mean that the two operations are performed concurrently. The functions of the content generator and the encoder have a timing overlap, generating the low resolution data first and starting encoding of the low resolution data before the high resolution data has been rendered. Advantageously, the temporarily simultaneous operation of the content generator and encoder processing a given frame helps to reduce latency arising due to processing delays occurring as a result of serial processing due to waiting for the entire frame to be rendered before it is passed to the encoder. The first level of quality may be a resolution and the second level of quality may be a higher resolution than the first level of quality. In alternative embodiments, the first and / or second level of quality may be bitdepth or other similar quality such as visual fidelity, resolution (number of pixels), image data bit-depth (e.g., for high- nits HDR), frame rate, stereoscopy, focal distance (e.g., for varifocal adjustments), haptics, picture quality (finer quantisation), temporal quality (high frames per second) etc. The encoded data may be encoded in a tiled manner or a hierarchical layered manner. The encoded data may be structured to conform to a LCEVC standard or a VC-6 standard. The encoding may be an enhancement coding in which a base layer generates a base stream and an enhancement coding generates an enhancement stream for combination with the base stream at the decoder. Examples of the one or more base codecs include, for example, AVC, HEVC, VP9, EVC, AV1 and may be implemented in software or hardware as is commonplace in this field. In some examples, where for example the hierarchical coding scheme is LCEVC, the encoder may be configured to encode the first content using a base encoding layer to generate base encoded data and wherein the encoder is configured to encode the second content using an enhancement encoding layer to generate enhancement data representing a difference between a decoded rendition of the base encoded data and the second content. Optionally, the outputting second content comprises the content generator being configured to upscale the first content, for example to generate the second content. Preferably, the upscale is a smart upscale. By smart upscale we mean an upscaling utilising machine learning or Al, such as a neural network upsampling using a convolutional neural network. The smart upscale may be a different upscale to a typical 'encoder' upscale performed (for example by the encoder) on the first (low quality) content (for example render). In some examples, the smart upscale may be a higher quality upscale than the encoder upscale and / or the smart upscale may be a more computationally complex upscale than the encoder upscale. In some cases, for a given frame, the content generator upscales the first content at least partially simultaneously as the content generator outputs the first content, such that the content generator temporarily simultaneously generates the first and second contents. In some examples, these two operations can be performed by using different hardware blocks of the same chip. In this case, temporarily simultaneously generating the low and high resolution layers results in better utilisation of available resources and overall latency and throughput will be improved because the content generator can start producing the next frame earlier as soon as the previous frame is processed. In preferred examples, the encoder is configured to encode the first and second contents without converting the second content from the second level of quality to the first level of quality. In other words, the encoder does not downsample the input content at a higher resolution for passing to a base coding layer at a lower resolution. The encoder is modified to receive two input frames at two different levels of quality and then pass the lower of the two to the base coding layer and the higher of the two to the enhancement coding layer. According to a further aspect there is provided an encoder for encoding a multilayer video stream, the encoder configured to: receive, from a content generator, a first content relating to at least one frame of a video at a first level of quality; receive, from a content generator, a second content relating to the at least one frame of the video at a second level of quality, the first level of quality being different from the second level of quality; begin encoding the first content before fully receiving the second content. Preferably, the encoder does not comprise a downsampler. The encoder may be further configured to generate encoded data at least partially simultaneously as the encoder receives, from the content generator, the second content. The encoder may be configured to encode the first content to generate encoded data. The first level of quality may be a resolution and the second level of quality may be a higher resolution than the first level of quality. The encoding may be in a tiled manner or a hierarchical layered manner. The encoding may produce encoded data structured to conform to a LCEVC standard or a VC-6 standard. The encoder may be further configured to encode the first content using a base encoding layer to generate base encoded data and wherein the encoder is configured to encode the second content using an enhancement encoding layer to generate enhancement data representing a difference between a decoded rendition of the base encoded data and the second content. The encoder may be configured to encode the first and second contents without converting the second content from the second level of quality to the first level of quality. According to a further aspect there may be provided a content generator for producing a video, the content generator configured to: produce a first content relating to at least one frame of a video at a first level of quality; produce a second content relating to the at least one frame of the video at a second level of quality, the second level of quality being higher than the first level of quality; send the first content and the second content to an encoder. Preferably the content generator does not comprise a downsampler. Preferably the content generator does not perform downsampling. The content generator may be further configured to send the first content to the encoder before the content generator finishes producing the second content. The content generator may be configured to send the first content to the encoder at least partially simultaneously as the content generator produces the second content. The level of quality may be a resolution and the second level of quality may be a higher resolution than the first level of quality. The producing the second content may comprise the content generator being configured to upscale the first content. The upscale may be a smart upscale. The content generator may be configured to upscale the first content at least partially simultaneously as the content generator outputs the first content. According to a further aspect there may be provided a method of encoding a multilayer video stream, the method comprising: receiving a first content relating to at least one frame of a video at a first level of quality; receiving a second content relating to the at least one frame of the video at a second level of quality, the first level of quality being different from the second level of quality; beginning encoding the first content before fully receiving the second content. Preferably the method does not comprise a downsampling step. The method may further comprise encoding the first content at least partially simultaneously as receiving the second content. The encoding may be in a tiled manner or a hierarchical layered manner. The encoding may produce encoded data structured to conform to a LCEVC standard or a VC-6 standard. The method may further comprise encoding the first content using a base encoding layer to generate base encoded data and encoding the second content using an enhancement encoding layer to generate enhancement data representing a difference between a decoded rendition of the base encoded data and the second content. The method may further comprise encoding the first and second contents without converting the second content from the second level of quality to the first level of quality. According to a further aspect there may be provided a method of generating content for a video, the method comprising: producing a first content relating to at least one frame of a video at a first level of quality; producing a second content relating to the at least one frame at a second level of quality, the second level of quality being different from the first level of quality; outputting the first and second content to an encoder. Preferably the method does not comprise a downsampling step. The method may further comprise: sending the first content to the encoder before the content generator finishes producing the second content. The producing the second content may comprise upscaling the first content. The upscaling may be a smart upscale. Upscaling the first content may be performed at least partially simultaneously as the content generator outputs the first content. In all methods, the first level of quality may be a resolution and the second level of quality may be a higher resolution than the first level of quality. The content generator may be a renderer. The content generated, or output, by the content generator may be a render. According to another aspect there is provided a non-transitory computer-readable storage medium having computer-readable instructions stored thereon, the computer-readable instructions being executable by a computerized device comprising processing hardware to execute a method as previously described. According to another aspect there is provided an encoder as previously described, the encoder being comprised in virtual reality equipment, medical imaging equipment, machine vision equipment and / or or a mobile communications device. There may be provided virtual reality equipment comprising an encoder as previously described. There may be provided medical imaging equipment comprising an encoder as previously described. There may be provided machine vision equipment comprising an encoder as previously described. There may be provided a mobile communications device comprising an encoder as previously described. According to another aspect there is provided virtual reality equipment comprising an encoder as described above and / or a content generator as described above. According to another aspect there is provided medical imaging equipment comprising an encoder as described above and / or a content generator as described above. According to another aspect there is provided machine vision equipment comprising an encoder as described above and / or a content generator as described above. According to another aspect there is provided a mobile communications device comprising an encoder as described above and / or a content generator as described above. BRIEF DESCRIPTION OF DRAWINGS Embodiments of the present invention will be described by way of example only with reference to the accompanying drawings in which: Figure 1 is an illustration of a hierarchical data structure; Figure 1a is a known approach to hierarchically encoding a frame render; Figure 2 is a schematic diagram of a system for signal processing; Figure 3 is a schematic diagram of a system for signal processing; Figure 4 is a flow diagram of operations performed by a content generator; Figure 5 is a schematic diagram of a prior art system for signal processing; Figure 6 is a schematic diagram of a system for signal processing; Figure 7 is a schematic diagram of a system for signal processing; and Figure 8 is a schematic diagram of a system for signal processing. DETAILED DESCRIPTION Embodiments of the disclosure comprise encoders and decoders that employ a tiered hierarchical approach to representing signals, for example video signals, as corresponding encoded data. The tiered hierarchical approach employs a base layer and one or more enhancement layers. In a case of LCEVC (“low complexity enhancement video coding”), an LCEVC encoder employs a base encoder, for example implemented in hardware, for generating base layer encoded data that is capable of providing a low resolution rendition of a given video signal, and one or more enhancement encoders, for example implemented using software that is executable using computing hardware, to generate enhancement layer encoded data that includes residual data that can be used to enhance the low resolution rendition to generate enhanced images of higher quality than the low resolution rendition. In order to provide backward compatibility, the base encoder is optionally implemented using known encoding hardware conforming to well-established standards, for example H.264, H.265, MPEG-2, MPEG-4, MPEG-5, VP-9, AV-1. In such a situation, the quality of rendition that is achievable at a given decoder depends upon operation of the one or more enhancement encoders whose operation can be dynamically controlled within software. The one or more enhancement layers include residual data that is generated by employing, during encoding, a combination of image downsampling and upsampling transformations followed by a subtraction operation. In some cases, the downsampling and upsampling transformations are mutually asymmetrical in nature to provide the residuals with particularly preferred entropy characteristics that provide for highly efficient data compression during encoding. Moreover, a quantisation operation can also be employed during encoding to control an amount of data needing to be encoded. Alternatively, in a case of the VC-6 standard ST-2117, a tiered hierarchy of layers is employed with a base layer and one or more enhancement layers, although no attempt is made for the base layer to be backward compatible with known encoding standards. A representation of a hierarchical structure for representing an image is depicted in Figure 1, wherein a base layer is denoted by 10, and enhancement layers are denoted by 20(1) to 20(N), wherein N is an integer. There may be a single enhancement layer wherein N=1, or there may be a plurality of enhancement layers wherein N>1. Use of such a base layer 10 and one or more enhancement layers 20 enables scalability within a codec arrangement by adjusting a degree of quantization used, or by selectively electing to employ only a certain number of enhancement layers. Optionally, when data communication network congestion is severe, certain enhancement layers are omitted in data transmitted from a given encoder to a corresponding given decoder to reduce an amount of data that needs to be communicated. It is possible to optimise signal processing and transmission of video frames by means of layered encoding methods, i.e., methods that comprise a layer of data representing the signal at a lower level of quality and one or more additional layers of data comprising information to reconstruct a rendition of the signal at a higher level of quality. By way of non-limiting example, additional layers of data allow the decoder device to enhance one or more quality aspects of the signal comprising visual fidelity, resolution, image data bit-depth (e.g., for high-nits HDR), frame rate, stereoscopy, focal distance (e.g., for varifocal adjustments), haptics, etc. One or more of the additional data layers can be treated as optional and can be safely dropped after encoding without compromising the quality of the base layer data. Latency is generally determined by two factors: i) the number of operations that need to be carried out; and ii) how quick the operations are to perform. As well as the actual number of operations being performed, there may also be a dependency between different operations in that some operations cannot start until the output of another operation has been completed. This method of serial processing introduces delays into the system. Overall system latency can be reduced by minimising the amount of wait time between operations and processing as many operations as possible in parallel. Within a system, a temporal delay may be introduced as a result of an encoding delay, De(LoQ, F), where LoQ corresponds to Level of Quality (colour resolution, spatial pixel resolution) and F corresponds to frame rate. Using an encoding scheme that allows for parallel processing in the system as a whole means that the delay De and therefore latency can be reduced. In a typical content creation pipeline, i.e. a graphics or video pipeline, a model is first generated and the scene is then rendered into a resulting image, i.e. the render, which is then passed to an encoder, such as those identified above, to reduce its size for subsequent transmission. Such content generation pipelines can be more complex, or more simple, depending on the implementation but the general blocks remain similar. This is exemplified in extended reality pipelines in which the rendering stage plays a critical role the content creation. These pipelines will be well understood by the skilled addressee. Generally, renderers generate images from a 2D or 3D data input (e.g. images or videos), the resulting image being referred to as the render. Typically renderers do not make a distinction between high resolution and low resolution renders; instead the renderer simply outputs a single high resolution image. This is because renderers are designed to produce a render at one resolution in the most efficient way. Historically, rendering is performed on the end user device. As such, video encoding and / or network delivery did not pose a significant problem for interactive rendering applications (e.g. games and visualisation apps) because the main process that was required was to render an image e.g. a scene from a video and show it on the user device screen. The ever-increasing complexity and processing demand of games, however, combined with user devices becoming more lightweight and portable, means that this rendering approach was becoming less feasible. For example, a complex interactive application may need to be displayed on a user device that is not powerful enough to render it on its own. A solution to this problem is cloud rendering, where rendering happens on a powerful server in a datacentre after which the video in question is encoded and transmitted over a network. In this situation, the end user device only needs to decode and display the video. A current approach to rendering and then encoding that render is illustrated in Figure 1a. As illustrated, the integration of the rendering application with an encoder results in the rendering and encoding being performed separately. A full resolution frame render 106 is produced by the renderer 110. This is then passed to the encoder for encoding, illustrated here in this example as an LCEVC encoder. In the encoding part of the pipeline, the full resolution input frame is first downsampled 103 before being passed to a base coder where it is encoded 108 to produce a base stream. The base encoded version of the downsampled input frame is then decoded by the base coder 114 before being upsampled 118. The upsampled decoded base rendition is compared to the original full resolution input frame and encoded by the enhancement layer encoding 122 to generate an enhancement stream. Together the enhancement stream and base stream can be combined at the decoder to generate the original full resolution frame render. The LCEVC encoder uses a downsample stage to generate a low resolution version of the input frame in order to pass it to the base layer. In accordance with the principles of the present disclosure, rendering and video encoding can be performed more efficiently if the interaction of these operations is optimised towards working with each other. The present invention relates to a system in which a renderer produces both a low resolution rendition and a high resolution rendition for a given frame, instead of just a high resolution rendition. Typically, a renderer may be configured to convert a 3D model into a 2D image. A renderer may be a component of a 3D modelling, animation, and rendering module. Renderers are generally responsible for plotting the positions, colours, and textures of objects in a 3D model and creating the image that will be seen on a screen. This process is usually done by calculating the lighting, shadows, and reflections of a 3D model. Atypical renderer may be configured to produce one or more (often high-quality, realistic) images for use in a display. The display may be a XR display (e.g. virtual reality, V.R., headsets) ora more ‘conventional’ 2D display (e.g. a TV). Renderers are typically used in gaming, architectural visualization, and other interactive applications. In other words, a renderer is generally used to create realistic images of objects and environments with lighting, shadows, and textures. For example, a renderer can be used to create a 3D image of a room with furniture, decorations, and other objects, including shadows and reflections. Additionally, a renderer can create detailed images of characters and other objects with textures and colours. Specific examples of how a 2D renderer can be used in a VR headset include creating detailed environments such as a cityscape with buildings, streets, trees, and other objects. A renderer can also be used to create detailed characters and objects with realistic skin and clothing textures. The renderer can also be used to create detailed lighting effects such as shadows and reflections, and to create realistic textures for objects such as walls, floors, and furniture. There are a variety of renderers available, from open source and free solutions to commercial renderers. Some popular renderers on the market include Blender, Maya, V-Ray, Redshift, RenderMan, Arnold, OctaneRender, Unreal Engine, Unity, and CryEngine. Other renderers, such as OSVR Render Manager, are designed specifically for VR headsets, providing optimized performance and compatibility with various headsets. In addition, some renderers, such as AMD ProRender, provide support for physically-based rendering and GPU acceleration for faster rendering times. Renderers including the aforementioned renderers may be configured to provide features such as photorealistic graphics, advanced lighting effects, and global illumination. In ‘comparative’ coding schemes, a high quality image may be rendered, this is fed into the comparative coding scheme (and is processed, e.g. downsampled, during the comparative coding scheme). In general, the described embodiments (e.g. the described coding schemes) utilise a high quality imagine and a low quality image. Embodiments of the invention utilise (i.e. instruct) a renderer (such as one of the above described renderers) to generate a low quality image (e.g. low resolution) and a high quality image. In short, comparative schemes use the renderer to generate a high image and then downsample this to produce a low quality image, in described schemes a renderer is used to generate (both) a high quality image and a low quality image. Depending on the capabilities of the renderer, the renderer may perform this in a number of ways, two examples are given below: 1) Generate the low quality image and the high quality image separately (i.e. optionally in parallel). The low quality image may be generated (i.e. ‘ rendered’) in a shorter amount of time than the high quality. 2) generate the low quality image and then generate the high quality image by utilising the low quality image (e.g. using the low quality image and further information to generate the high quality image). In both examples, the low quality image is available before the high quality image is available. Therefore the low quality image can be fed into the coding scheme at an earlier time than if a high quality image was fed into the coding scheme, thus reducing latency. Moreover, because a low quality image and a high quality image are generated by the renderer, this removes the need to downsample a high quality image in order to generate the low quality image (as is performed in comparative coding schemes), thus further reducing latency and also reducing computing power. The renderer of the present invention is particularly useful for gaming, VFX, or data visualisations in which all of the video content is created from scratch by the renderer. The renderer takes in various data and parameters as inputs including but not limited to meshes, coordinates, textures, etc which are used to produce the resulting render. Figure 4 illustrates operation of the proposed renderer. The proposed system also includes an encoder which receives both the low and high resolution renders as inputs to the encoder, instead of just the high resolution render. Typical encoders receive only a single input per frame, the size of the input being equal to the high resolution (enhancement layer) size. This is then used by the encoder to produce a low resolution encoded bitstream (base layer) by downsampling, for example as shown in Figure 5, as well as a high resolution encoded bitstream (enhancement layer). The encoder of the present invention receives two inputs for each frame: one for low resolution (base layer) and one for high resolution (enhancement layer). These are both used by the encoder in order to produce a low resolution encoded bitstream (base layer) as well as a high resolution encoded bitstream (enhancement layer). The encoder of the present invention therefore avoids the need to perform downsampling to obtain the low resolution layer from the high resolution input, because both resolution are received by the encoder. Notably in preferred implementations of this concept the encoder can start encoding the low resolution layer as soon as the low resolution render is ready and received by the encoder. In this way, the encoder does not need to wait for the high resolution render before beginning encoding. Examples of the implementation of an enhancement encoder are presented in patent publication WO2022 / 023747 and the LCEVC SDK provided by V-Nova. In examples of the present disclosure, the API may be modified to receive two input frames and pass these two input frames to each functional module, skipping the downsampling stage illustrated in Figure 1a. Several latency benefits are achieved by the present concept. In one example, the system reduces latency due to the omission of the downsampling step. That is, the overall encoding time is reduced because the downsampling operation is not performed. Of course, this may be compromised by an additional time taken in creating a second render, however it may be that the renderer is more powerful than the encoder, resulting in significant latency reductions. In addition, the conventional method of performing a high resolution render and downsampling takes more time than the present method of producing a low resolution render. In preferred implementations of the present concept, low resolution (e.g. base layer) operations may be begun before high resolution (e.g. enhancement layer) operations. In fact, not only may they be begun, but they may be mostly performed before the high resolution operations. The base layer data stream contains the most important information, whereas the enhancement layer data stream contains residual information to improve the quality of the base layer. The base layer therefore provides a basic level of perceived quality while the enhancement layers can be used to incrementally improve this quality. The base layer is encoded, decoded, and upsampled before it is ready to be compared to the original input to generate residuals. The enhancement layer encodes the difference between the original input and the upscaled reconstruction of the base layer. It is therefore possible to reduce latency by receiving the low resolution render before the high resolution render, so that operations to be performed on the low resolution render can begin as soon as this render is ready. In otherwords, the low resolution render is available earlier than the contemporary high resolution render. The base layer operations can then begin sooner on that low resolution render than in conventional designs. The output of the base layer is therefore ready to be used by the enhancement layer at a much earlier time, i.e. the operations at the enhancement layer that are dependent on the base layer processing can be performed, or are ready to be performed, much earlier. Rendering the low resolution render and then having base layer data ready for output earlier than enhancement layer data is ready allows the possibility of taking advantage of separate base and enhancement layer network streams. In this scenario, not only does the renderer pass the low resolution frame render to the base coder as soon as it is ready, but the base stream is also output as soon as it is ready, without waiting for the enhancement layer to complete processing to be combined with the base stream. Another advantage of the system is that the low resolution render that is generated first may be used to generate the high resolution render more efficiently and more quickly, compared to generating the high resolution render alone. This may have the effect that the time to generate the low and high resolution renders is less than the time taken to independently generate the low resolution render and a high resolution render. Specific details of the implementation of the invention will now be described, with reference to Figure 2. A multi-layer video signal includes multiple individual frames which comprise data that provides for a rendition of the frame in a minimum of two different levels of quality. In the following, the description will make reference to different resolution, but any other quality aspect of the signal could be used instead. For each frame of the input signal 102, a renderer 110 of the system 100 generates low resolution data 104 corresponding to a low resolution layer of the multi-layer video signal, for example at a base layer resolution, for example as shown in Figure 4. Optionally, as shown for example in Figures 2 and 3, the renderer receives an input signal 102 using which the renderer produces the render. As soon as the low resolution data 104 is generated the low resolution data 104 may be sent to a first encoder 108, for example a base layer encoder for encoding before the renderer 110 has finished generating high resolution data 106 corresponding to a high resolution layer of the multi-layer video signal, for example at an enhancement layer resolution. In this way, the first encoder receives data of the frame that is at a lower resolution quickly with low latency. An encoding process starts as soon as the data to be encoded by a particular encoder is ready. By implementing this approach within the system 100, an encoding process occurring within the first encoder 108 may be temporally overlapped with a rendering process within the renderer 110. As shown in Figure 2, the renderer 110 completes the rendering process, to generate the high resolution data 106 of the same frame, temporally in parallel with the encoding process occurring in the first encoder 108, thereby effectively making the render and encode processes operate at least partially in parallel, which reduces the overall time to carry out these processes. While the rendering process completes (for example, while the renderer produces the high resolution output), the first encoder 108 processes the low resolution data 104 to output a first encoded stream 112. The first encoded stream 112 is decoded by a first decoder 114, and the output decoded signal 116 is up-sampled by an upsampler 118. These operations may take place at least partially in parallel with the renderer producing the high resolution output. The resulting up-sampled signal 120 is processed together with the high resolution data 106, which is produced by the renderer, by a second encoder 122 to generate a second encoded stream 124. In some examples, the up-sampling operation is not performed and therefore the decoded signal 116 is passed straight to the second encoder 122. For example, a base layer and an enhancement layer may be of the same resolution but they may have the same or different bitdepth. In the latter case the enhancement layer will be the correction of the base layer reconstruction artifacts. If a residual signal is generated by the second encoder 122 this will correct for any difference between the high resolution data input signal 106 and the output of the upsampler 118 (or base decoding process or other correction or conversion stage performed after base decoding and before combination with the full resolution frame render). In a further optional implementation of the proposed concept, instead of the renderer producing both a low resolution frame render and a high resolution frame render, as set out in the above examples, the renderer may produce a low resolution frame render which is then upscaled to produce a high resolution frame render. The low resolution frame render may be passed to the base coding layer before being compared to the input upscaled, resolution frame render at the enhancement layer. This upscaling may be thought of as moving the upscaling stage from a downsampling at the encoder stage to an upsampling at the renderer stage. There are numerous benefits to such a concept. Firstly, some renderers may not have the capability of generating a full resolution render. To address this, the renderer can generate the low resolution render from the input signal and then perform an upscale to generate the high resolution render. The lower resolution render and the upscaled high resolution render are input into the second encoder as before. The upscale enables renderers which are not able to generate a 4K resolution render to instead render in 1080p, upscale the 1080p render, and then perform encoding as before. Further, the upscale may be a ‘smart upscaler’, for example as shown in Figure 3. By ‘smart upscaling’ we mean an upscaler utilising Al or machine learning technologies, such as neural networks. Such machine learning based upscalers are known in the art. The smart upscaler 126, as part of the renderer, may be different to the standard upscalers used in encoders and decoders. In particular the smart upscaler may be more powerful and more accurate. The system 100 can therefore use standard upscalers in the encoder and decoder (the standard upscalers using less power and being quicker) and then the residuals can be used to reconstruct the smart upscaled render. In practical implementations of the pipeline, the renderer may be implemented in powerful software approaches, e.g. powerful CPUs or GPUs, whereas the encoding and decoding may be performed using dedicated modules. The use of a ‘smart’ upscaling function at the renderer allows for the functionality of the renderer to be leveraged while the overall functionality is not limited by the capability of the encoder and decoder. In general the ‘smart upscaler’ can be characterised as more resource demanding, such that their operations cannot be performed during decoding on a wide range of end user devices. In particular they may be based on a large CNN that requires a lot of memory and powerful hardware for realtime processing. They can still be used to produce high resolution version of input content for applications described. Returning to our analogy of moving the scaling process from the encoder to the renderer it can be seen that the more powerful scaling of the renderer is used while only a limited version is needed at the encoder, with the residuals of the enhancement layer providing the difference between the two. Combining the concepts proposed together, it may be seen that the low resolution frame render in this embodiment can be passed to the base coding layer before the upscaling process has completed. Thus reducing overall latency because the base coding layer does not need to wait for that stage to complete before operating on the downsampled full resolution frame render. Similarly to the examples given above, optionally, that base encoded low resolution rendition may then be output before the corresponding enhancement layer operations have completed on the upscaled frame render. The present disclosure provides a system and methods which allows different encoding stages to start when the relevant renders are complete. Different parts of a frame can therefore be operated on at the same time. Thus, embodiments of the present disclosure are capable of providing low latency multi-layer signal processing, for example of a ‘layered’ video codec. In addition, embodiments of the present disclosure are capable of providing low latency multilayer signal processing via use of a video codec that encodes images in a tiled manner, for example as employed in VC-6 standard. Furthermore, there is disclosed a computer program product comprising a non-transitory computer-readable storage medium having computer-readable instructions stored thereon, the computer-readable instructions being executable by a computerised device comprising processing hardware to execute a method for operating the system 100 as described in the foregoing. Although the present disclosure has been described with reference to a renderer, it will be understood that in some cases a form of content already exists that is available to be processed (e.g. encoded), for example some content that is being produced for a streaming service. In this case, because the content already exists, there is no need for the initial content to be rendered for example from an input model. Thus, the component described above as a renderer takes the more general form of a content generator, wherein the content generator is able to generate at least two outputs, wherein the outputs have different levels of qualities, for example different resolutions. As discussed above, the content generator and the encoder are each associated with two different levels of quality, for example resolution. The content generator generates two outputs each having different levels of quality and these outputs are passed to the encoder. In other words, the encoder receives two different resolution streams; the encoder does not receive a high resolution stream and downsample to produce the low resolution stream. As a result of the content generator outputting two different resolutions, a downsampling step does not need to be performed. In some examples, the content generator starts with a single first resolution (e.g. provided by an app or another source of content) and then the content generator generates a second resolution output using the initial first resolution available to the content generator. In one example, the content generator may have low resolution content available and then generate the high resolution content using the low resolution content, for example by (smart) upscaling as illustrated in Figures 3 and 7. In another example, the content generator may have high resolution content available and then generate the low resolution content using the high resolution content. In the latter example, the generation of low resolution content from the high resolution content is done using methods that do not involve downsampling. Looking at Figure 6 in particular, an exemplary content generator is illustrated for use when specific content needs to be generated in real time for a specific user (eg for use in a game or VR application), rather than generating static content such as a TV show where the same content may be sent to many users. In this example, the content generator 200 generates a high resolution source image 202 and a low resolution source image 204. The low resolution source image 204 is sent to a base encoder 206 to produce a base encoded low resolution image 208. This is then passed to a base decoder 210 to produce a decoded base encoded low resolution image 212 which can then be upsampled by an upsampler 214 to generate an upsampled decoded base encoded image 216 which is now at a higher resolution than the original low resolution source image 204. The upsampled image 216 and the high resolution source image 202 are compared to generate residuals 220, which can then be encoded 222 as the enhancement stream before being transmitted to a user 224. The output 208 from the base encoder 206 can also be transmitted to the user 226 as a separate stream or this can be combined with the enhancement stream before being sent to the user. As can be seen from Figure 6, the usual encoding process remains unchanged. A latency saving is made due to the lack of need to downsample to produce the low resolution source content 204 (as can be clearly seen when comparing Figure 5 to Figure 6), because this is generated initially by the content generator 200. A further latency saving is made because, as discussed previously, the low resolution image 204 will often be ready before the high resolution image 202 and so the low resolution image 204 can be output from the content generator 200 and passed to the base encoder 206 without having to wait for the high resolution image 202 to be generated (which is not needed by the base encoder 206). Therefore, operations on the low resolution image can begin relatively earlier compared to conventional encoding schemes. Turning to Figures 7 and 8, two methods of generating high resolution content from low resolution content are illustrated. In the method illustrated in Figure 7, the low resolution source image 204 is upsampled using a smart upsampler 228 to generate a high resolution upsampled image 230. In the method illustrated in Figure 8, the low resolution source image 204 is enhanced using an enhancer 232 to generate a high resolution enhanced image 234. The rest of the method in Figures 7 and 8 is the same as illustrated in Figure 6. The smart upsampling method of Figure 7 can be used to adjust (e.g. increase) the resolution of original content eg using Al. For example an animated film that was originally at resolution 1080p can be upsampled to 4K resolution. The smart upsampling may take place in a cloud-based computing system which is able to generate a better quality image, (for example a higher resolution image) than would be achieved using a desktop PC. Advantageously, in this example, the decoder only needs to perform the low quality upsample (for example a low resolution upsample) and combine the residuals in order to effectively generate the smart upsampled high resolution content. This method and system can therefore be considered to comprise operations carried out by a first upsampler and a second upsampler, wherein one of the upsamplers (eg the upsampler which is able to generate the high resolution “source” content) is not part of the decoder or performed by the decoder. Although smart upscalers have been mainly been described as upsamplers utilising machine learning and / or Al, this example helps to illustrate that what is being described as a smart upscaler may be any upscaler that is better performing in some way (e.g. more powerful and I or more accurate and I or better quality, and so forth) compared to the 'other upsampler' 214. As such, the smart upsampler may not necessarily utilise machine learning and I or Al to be considered a smart upsampler. Rather, the smart upsampler 228 is considered smart due its difference from (in particular, its better performance than) the other upsampler 214. The enhanced method of Figure 8 is an alternative way of adjusting (eg increasing) the level of quality of a source image. However, in this case, the adjusting is not limited to an adjustment in resolution. For example the colour could be adjusted i.e. enhanced from SDR to HDR. This method can be useful for example when restoring an old or damaged film. All the methods and systems, and their associated components, described herein all contribute towards decreasing latency. Additionally the computing and processing burden on the encoder and I or decoder is reduced by removing at least the initial step of downsampling. As a result of omitting the downsampling, all methods and systems described herein provide the advantage of allowing the encoder to carry out fewer processing steps therefore reducing computer demands. This also allows the low resolution content to be obtained more quickly (generated by the content generator rather than using downsampling) and so latency is reduced.

Claims

1. A system for encoding a multi-layer video stream, the system comprising: a renderer configured to output a first content relating to at least one frame of a video at a first level of quality and a second content relating to the at least one frame at a second level of quality, the second level of quality being different from the first level of quality; andan encoder configured to encode the first and second contents and output encoded data;wherein the encoder is configured to encode the first content at least partially simultaneously as the renderer outputs the second content, such that the renderer and the encoder function temporarily simultaneously when processing a given frame.

2. The system of claim 1, in which the first level of quality is a resolution and in which the second level of quality is a higher resolution than the first level of quality.

3. The system of any preceding claim, wherein the encoder is configured to encode the first content using a base encoding layer to generate base encoded data and wherein the encoder is configured to encode the second content using an enhancement encoding layer to generate enhancement data representing a difference between a decoded rendition of the base encoded data and the second content.

4. The system of any preceding claim wherein the outputting the second content comprises the renderer being configured to upscale the first content.

5. The system of claim 4 wherein the upscale is a smart upscale.

6. The system of claims 4 or 5 wherein the renderer is configured to upscalethe first content at least partially simultaneously as the renderer outputs the first content.

7. The system of any preceding claim, wherein the encoder is configured to encode the first and second contents without converting the second content from the second level of quality to the first level of quality.

8. An encoder for encoding a multi-layer video stream, the encoder configured to:receive, from a renderer, a first content relating to at least one frame of a video at a first level of quality;receive, from the renderer, a second content relating to the at least one frame of the video at a second level of quality, the first level of quality being different from the second level of quality;begin encoding the first content before fully receiving the second content; andgenerate encoded data at least partially simultaneously as the encoder receives, from the renderer, the second content.

9. The encoder of claim 8, wherein the first level of quality is a resolution and in which the second level of quality is a higher resolution than the first level of quality.

10. The encoder of any of claims 8 to 9, wherein the encoder is further configured to encode the first content using a base encoding layer to generate base encoded data and wherein the encoder is configured to encode the second content using an enhancement encoding layer to generate enhancement data representing a difference between a decoded rendition of the base encoded data and the second content.

11. The encoder of any of claims 8 to 10, wherein the encoder is configured to encode the first and second contents without converting the second content from the second level of quality to the first level of quality.

12. A renderer for producing a multi-layer video stream, the renderer configured to:produce a first content relating to at least one frame of a video at a first level of quality;produce a second content relating to the at least one frame of the video at a second level of quality, the second level of quality being higher than the first level of quality;send the first content and the second content to an encoderfurther configured to send the first content to the encoder before the renderer finishes producing the second content.

13. The renderer of claim 12, in which the first level of quality is a resolution and in which the second level of quality is a higher resolution than the first level of quality.

14. The renderer of any of claims 12 to 13, wherein the producing the secondcontent comprises the renderer being configured to upscale the first content.

15. The renderer of claim 14, wherein the upscale is a smart upscale.

16. The renderer of claim 14 or 15, wherein the renderer is configured toupscale the first content at least partially simultaneously as the renderer outputs the first content.

17. A method of encoding a multi-layer video stream, the method comprising:receiving, from a renderer, a first content relating to at least one frame of a video at a first level of quality;receiving, from the renderer, a second content relating to the at least one frame of the video at a second level of quality, the first level of quality being different from the second level of quality;beginning encoding the first content before fully receiving the second content;the method further comprising encoding the first content at least partially simultaneously as receiving the second content.

18. The method of claim 17, further comprising encoding the first content using a base encoding layer to generate base encoded data and encoding the second content using an enhancement encoding layer to generate enhancementdata representing a difference between a decoded rendition of the base encoded data and the second content.

19. The method of any of claims 17 to 18, further comprising: encoding the first and second contents without converting the second content from the second level of quality to the first level of quality.

20. A method of generating content for a multi-layer video stream, the method carried out by a rendered and comprising the steps of:producing a first content relating to at least one frame of a video at a first level of quality;producing a second content relating to the at least one frame at a second level of quality, the second level of quality being different from the first level of quality;outputting the first and second content to an encoder; andsending the first content to the encoder before the renderer finishes producing the second content.

21. The method of claim 20, wherein the producing the second content comprises upscaling the first content.

22. The method of claim 21, wherein the upscaling is a smart upscale.

23. The method of claim 22, wherein upscaling the first content is performedat least partially simultaneously as the renderer outputs the first content.

24. The method of any of claims 17 to 23, in which the level of quality is a resolution and in which the second level of quality is a higher resolution than the first level of quality.

25. A non-transitory computer-readable storage medium having computer-readable instructions stored thereon, the computer-readable instructions being executable by a computerized device comprising processing hardware to execute a method according to any of claims 17 to 23.

26. An encoder according to any of claims 8 to 11, the encoder being comprised in virtual reality equipment, medical imaging equipment, machine vision equipment and / or or a mobile communications device.

27. Any one or more of virtual reality equipment, medical imaging equipment, 5 machine vision equipment, and I a mobile communications device comprising an encoder according to any of claims 8 to 11 and / or a renderer according to any of claims 12 to 16.

Citation Information

Patent Citations

  • ViewUS20070286293A1onEspacenetopensinnewtab

  • ViewWO2023/000179A1onEspacenetopensinnewtab

  • ViewJPH1118085AonEspacenetopensinnewtab

  • ViewUS20170359586A1onEspacenetopensinnewtab

  • ViewGB2611129AonEspacenetopensinnewtab