Method, device, system, and non-transitory computer-readable storage medium for rendering overlays in layered encoded video sequences
By aligning glyph positions across resolutions using a ratio-based mapping technique, the method addresses misalignment issues in hierarchical video coding, improving compression efficiency and reducing bit rate.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2026-03-10
AI Technical Summary
The challenge in hierarchical video coding is the precise alignment of overlays, such as text or graphics, across different resolutions, which can lead to misalignment and increased bitrate due to variations in rendering processes like positioning, hinting, and kerning.
A method for rendering overlays in layered video sequences by aligning glyph positions across different resolutions using a ratio-based mapping technique, ensuring accurate alignment and reducing residual error, thereby minimizing the enhancement layer's bit rate.
This method achieves consistent rendering of glyphs across resolutions, reducing residual error and bit rate, enhancing compression efficiency and conserving bandwidth and storage resources.
Smart Images

Figure 2026041665000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to layered video coding, and in particular to a method, device, and non-transitory computer-readable storage medium for rendering overlays within a layered encoded video sequence. [Background technology]
[0002] The advent of hierarchical video coding has significantly improved the efficiency and flexibility of video streaming technologies. Layered coding, such as Low Complexity Enhancement Video Coding (LCEVC), is a coding technique that encodes video data in multiple layers, enabling the delivery of video content at various resolutions from a single coding source. This approach starts with a base layer containing a low-resolution version of the video and adds one or more enhancement layers that provide the information necessary to reconstruct the video at higher resolutions. Such scalable methods are advantageous, for example, for adaptive streaming technologies that need to adjust to changing network conditions and device capabilities.
[0003] A challenge in hierarchical coding is the rendering and alignment of overlays, such as text or graphics, across different resolutions. Overlays rendered directly at the resolution at which they are displayed often achieve better visual quality than those scaled from a higher resolution. For example, rendering text at a high resolution and then downscaling it can lose clarity and sharpness. Therefore, applying different overlays directly to both the base layer and enhancement layer may be desirable to maintain high visual quality at both resolutions, especially when using the base layer as its own data stream.
[0004] However, this approach introduces complications when considering how to achieve precise alignment of the overlay between different resolutions. Slight variations in the rendering process, such as positioning, hinting, and kerning, can cause misalignment, which can lead to mismatches between the base layer and the enhancement layer. These mismatches, even if small, can accumulate and result in noticeable misalignments, adversely affecting visual quality and increasing the bitrate required to encode the enhancement layer.
[0005] So, in this context, there is a need for improvement. Summary of the Invention
[0006] In view of the above, it would be advantageous to overcome or at least reduce one or more of the above-mentioned disadvantages as set out in the accompanying independent patent claims.
[0007] According to a first aspect of the present disclosure, a method for rendering an overlay in a hierarchically coded video sequence having an enhancement layer and a base layer includes: rendering the video sequence at a first resolution and a second resolution, one of the first resolution and the second resolution being lower than the other of the first resolution and the second resolution; rendering a first overlay in the video sequence at the first resolution, the first overlay including a first pattern of one or more glyphs, the first glyph of the one or more glyphs being rendered at a first pixel location in the video sequence at the first resolution; and rendering a first overlay in the video sequence at the second resolution. a video sequence having a higher resolution and a second overlay having a higher resolution, wherein the second overlay includes a first pattern of one or more glyphs, and the rendering of the second overlay is controlled to render first glyphs of the first pattern at second pixel positions within a video sequence at a second resolution, the second pixel positions being obtained by mapping the first pixel positions to the second pixel positions according to a ratio between the first resolution and the second resolution; encoding a video sequence having a lower resolution in a base layer; and encoding a residual between the video sequence having a higher resolution and the video sequence having a lower resolution in an enhancement layer.
[0008] Advantageously, this method achieves alignment of at least one glyph of the first pattern between the two resolutions, facilitating consistent rendering of the glyph pattern at both the lower and higher resolutions. Aligning the rendering locations (first pixel location and second pixel location) of the first glyph can reduce the difference between the low-resolution and high-resolution versions of the video sequence. This reduces residual error and reduces the bit rate required to encode the enhancement layer. As a result, this method achieves a more streamlined and efficient encoding process that increases compression efficiency and conserves bandwidth and storage resources.
[0009] As used herein, the term "ratio" refers to the proportional relationship between a first resolution and a second resolution. This ratio is employed to map pixel locations of a glyph from the first resolution to the second resolution. Specifically, when a glyph is rendered at a particular pixel location in the first resolution, the ratio determines the corresponding pixel location in the second resolution. For example, if the second resolution is twice the width and height of the first resolution, the ratio (or scaling factor) between the first and second resolutions is 1:2, meaning that pixel coordinates (second pixel locations) in the second resolution are scaled twice as large as pixel coordinates (first pixel locations) in the first resolution. Note that the ratio is not limited to being the same for both width and height. For example, the ratio may be 1:2 for width and 1:1.5 for height. In this case, pixel locations of a glyph may be scaled differently horizontally compared to vertically. Advantageously, such mapping facilitates accurate rendering of the first glyph at the same relative position at both resolutions, enabling alignment of the first glyph between the first overlay and the second overlay. If scaling involves non-natural number scaling, for example, scaling according to a ratio of 1:1.5, the second pixel position obtained by mapping the first pixel position according to the ratio may be a sub-pixel position, which can then be rounded to the nearest pixel position within the video sequence at the second resolution. In this case, the alignment between the first overlay and the second overlay is still good enough to achieve the desired bitrate reduction.
[0010] As used herein, a "glyph" refers to a visual symbol or character that is rendered as part of an overlay within an animation sequence. Glyphs can represent text, icons, or other graphic elements that are overlaid on an animation sequence to provide additional information or visual effect.
[0011] As used herein, a "glyph rendered at a pixel location" and similar expressions refer to a specific placement of a glyph at specific coordinates within a video sequence at a relevant resolution. The pixel location that is aligned between two resolutions may be, for example, the top-left pixel location of a first glyph in a first overlay and a second overlay, or any other suitable pixel location that serves as a reference point (such as a visual or geometric center location, a bottom-right location, etc.). This reference point is used to ensure that glyphs maintain consistent alignment when rendered at different resolutions, facilitating accurate mapping and reducing discrepancies between overlays rendered at different resolutions.
[0012] In the context of the present disclosure, the terms "first," "second," "third," etc. do not necessarily denote a sequential order or priority. Instead, these terms are merely used to identify and distinguish between different features, elements, or steps in the description. The terms are intended to provide clarity and should not be construed as implying a particular sequence or hierarchy unless otherwise expressly stated.
[0013] In some examples, the first pattern includes a second glyph, and rendering the first overlay includes rendering the second glyph at a third pixel position within the video sequence at a first resolution, and rendering the second overlay is controlled to render the second glyph at a fourth pixel position within the video sequence at a second resolution, the fourth pixel position being obtained by mapping the third pixel position to the fourth pixel position according to a ratio.
[0014] Therefore, by aligning the rendering pixel positions of multiple glyphs, the compression efficiency of the hierarchical encoding process can be further improved.
[0015] In some examples, the rendering position of each glyph in the first pattern is aligned between the two overlays according to the ratio between the first and second resolutions. However, in other cases, only a subset of the glyphs in the first pattern are aligned with this technique. The determination of the number of glyphs to align can be made by considering several factors, including, for example, the computational overhead of the alignment process, the visual appearance of the glyphs in the second overlay (ensuring that the characteristics of the rendering pattern, such as kerning and spacing, are aesthetically pleasing), and the benefit of reducing bit rate and improving compression efficiency.
[0016] In some examples, the first pattern includes a first group and a second group of glyphs, each group of glyphs including multiple glyphs, the first glyph being part of the first group and the second glyph being part of the second group. In some examples, the groups correspond to words. In some examples, the groups are separated by white space in the first pattern, and different groups are determined using the white space as a boundary rule. In some examples, other appropriate boundary rules are used. The boundary rules may be language-specific. For example, some languages do not use spaces between words, and in such languages, other rules may be applied to detect groups, such as using invisible characters / glyphs such as "zero width space" (ZWSP) characters as boundary rules.
[0017] In some examples, each group (eg, word) is aligned as described above using at least one glyph per group.
[0018] In some examples, rendering the second overlay includes using a glyph layout algorithm to render a glyph that is different from the first glyph in the first group of glyphs, and the pixel position of each of the glyphs that is different from the first glyph in the first group of glyphs is determined using the glyph layout algorithm and the second pixel position. Similarly, rendering a glyph that is different from the second glyph in the second group of glyphs can be achieved using a glyph layout algorithm. Advantageously, by aligning each group of glyphs using a subset of the glyphs in the group (e.g., one glyph) and rendering the remaining glyphs in the group using a standard typesetting or text layout algorithm (which may also be referred to as a text shaping algorithm), each word can be aligned consistently across resolutions. The rendering positions of glyphs within a particular group (i.e., glyphs that are not specifically controlled using a mapping technique) are determined using a layout algorithm so that kerning and other typographic details are visually appealing to users. An additional advantage of this approach is that it may help maintain readability of text across different resolutions. This approach balances the need for precise alignment with processing complexity and visual appearance, ensuring efficient compression while maintaining high visual quality and readability.
[0019] In some examples, the first pattern includes a first group of glyphs and a second group of glyphs, where the first glyphs and the second glyphs are part of the first group. As a result, as described above, the ratio between the first resolution and the second resolution can be used to align multiple glyphs within a group / word. Advantageously, as described above, compression efficiency can be increased.
[0020] In some examples, the first resolution is a lower resolution. Mapping from a lower resolution to a higher resolution can advantageously reduce the likelihood of pixel misalignment due to sub-pixel positioning. When glyph positions are mapped between resolutions, the resulting coordinates at the mapped resolution may fall between pixel boundaries, creating sub-pixel positions. These sub-pixel positions may be rounded to the nearest integer, resulting in misalignment of the glyphs. When mapping from a lower resolution to a higher resolution, such misalignments are less noticeable and therefore have less impact on compression efficiency than in the case of downscaling, where the mismatches are greater and can have a significant impact on compression efficiency.
[0021] In some examples, the first overlay includes a second pattern of one or more glyphs, and the second overlay includes a third pattern of one or more glyphs, the third pattern being different from the second pattern, and the method further includes determining a first pixel area required to render the second pattern within the video sequence at the first resolution and determining a second pixel area required to render the third pattern within the video sequence at the first resolution, and the rendering of the first overlay is controlled to render the one or more glyphs of the first pattern outside of both the first pixel area and the second pixel area.
[0022] Advantageously, by predetermining the areas required for the dynamic portions of the overlay (i.e., the second and third patterns) and ensuring that the static portions (first pattern) are rendered outside of these areas, this method maintains a clear separation between the static and dynamic content. This separation facilitates rendering the static portions away from the dynamic portions, regardless of whether the dynamic portions appear larger in the first resolution or the second resolution. As a result, this technique facilitates the static content of the overlay (first pattern) not interfering with the dynamic content at either resolution, while remaining rendered in a corresponding position between the first and second resolutions, as described above. This approach preserves the visual integrity of the video sequence and increases compression efficiency.
[0023] In some examples, the method further includes transmitting a base layer in the first data stream and transmitting the base layer and the enhancement layer in the second data stream.
[0024] Transmitting the base layer in one data stream and both the base layer and enhancement layers in a second data stream enables scalable video streaming, allowing devices with lower bandwidth or processing power to receive only the base layer and ensure basic video playback. Meanwhile, devices with higher bandwidth and processing power can receive both layers and benefit from enhanced video quality. This dual-stream method also provides flexibility in network conditions, because as network bandwidth fluctuates, one of the streams takes priority to maintain continuous playback. Furthermore, because the base layer includes an overlay, information provided in the overlay remains available in both data streams.
[0025] In some examples, a method includes transmitting a data stream including a base layer and an enhancement layer over a communication channel, receiving an indication of network congestion on the communication channel, and adjusting transmission of the data stream to not include the enhancement layer.
[0026] Indication of network congestion on a communication channel can be implemented using various techniques. For example, network performance metrics such as packet loss, latency, and jitter can be monitored. If these metrics exceed predefined thresholds, an indication of congestion may be triggered.
[0027] By dynamically adjusting transmission to exclude enhancement layers during network congestion, this example can ensure that the base layer is still delivered, maintaining uninterrupted video streaming. Because the base layer includes an overlay, information provided in the overlay remains available in the data stream.
[0028] According to a second aspect of the present disclosure, the above object is achieved by a non-transitory computer-readable storage medium having stored thereon instructions for performing the method according to the first aspect when executed on a device having processing capabilities.
[0029] According to a third aspect of the present disclosure, the above object is achieved by providing a device for rendering an overlay in a hierarchically coded video sequence having an enhancement layer and a base layer, the device comprising: rendering the video sequence at a first resolution and a second resolution, one of the first resolution and the second resolution being a lower resolution compared to the other of the first resolution and the second resolution; rendering a first overlay in the video sequence at the first resolution, the first overlay including a first pattern of one or more glyphs, a first glyph of the one or more glyphs being rendered at a first pixel location in the video sequence at the first resolution; and rendering a first overlay in the video sequence at the second resolution. This is achieved by a device configured for rendering a second overlay, the second overlay including a first pattern of one or more glyphs, the rendering of the second overlay being controlled to render first glyphs of the first pattern at second pixel positions within a video sequence at a second resolution, the second pixel positions being obtained by mapping the first pixel positions to the second pixel positions according to a ratio between the first resolution and the second resolution; encoding a video sequence having a lower resolution in a base layer; and encoding a residual between the video sequence having a higher resolution and the video sequence having a lower resolution in an enhancement layer.
[0030] In some examples, the device of the third aspect is a camera and the video sequence is captured by the camera.
[0031] According to a fourth aspect of the present disclosure, the above object is achieved by a system comprising a first device of the third aspect and a second device, wherein the first device is configured for transmitting a base layer on a first data stream and transmitting the base layer and an enhancement layer on a second data stream, and the second device is configured to receive the first data stream and the second data stream and use the first data stream for a first purpose and the second data stream for a second, different purpose.
[0032] Thus, a second device can receive both data streams and utilize the first data stream for one purpose and the second data stream for a different purpose. For example, the first data stream, containing only the base layer, can be used for real-time, low-bandwidth applications such as live video monitoring, where maintaining continuous playback even under network congestion is important. Meanwhile, the second data stream, containing the enhancement layer, can be used for video recording. In recording scenarios, the emphasis is on capturing the highest possible video quality rather than real-time playback, so slight delays and higher bandwidth usage are acceptable. This dual-stream approach increases flexibility and efficiency, allowing the system to adapt to changing network conditions and application requirements, ensuring, for example, that both real-time performance and high-quality video recording are achievable.
[0033] Other objectives include recording both streams and implementing different retention policies for the recordings. For example, the first data stream, which requires less storage space, can be retained for a longer period, while the second data stream, which requires more storage due to its higher quality, can be retained for a shorter period. This approach allows for efficient use of storage resources, ensuring that essential lower-resolution recordings are available for a longer period, while higher-quality recordings are preserved for immediate but short-term needs.
[0034] According to a fifth aspect of the present disclosure, the above object is achieved by a system including a first device of the third aspect, a second device, and a third device, wherein the first device is configured to transmit a base layer on a first data stream and a base layer and an enhancement layer on a second data stream, the second device is configured to receive the first data stream, and the third device is configured to receive the second data stream. This dual-stream approach allows the system to adapt to changing network conditions and application requirements. This allows different devices to handle different data streams depending on their needs. For example, a second device that can prioritize low-bandwidth and real-time applications can utilize the first data stream. Meanwhile, a third device that can focus on applications requiring higher video quality can handle the second data stream. This setup promotes improved performance in various use cases and allows for efficient management of resources and network capacity.
[0035] The second, third, fourth and fifth aspects may generally have the same features and advantages as the first aspect.
[0036] The above, as well as additional objects, features, and advantages of the present invention will be better understood through the following illustrative and non-limiting detailed description of embodiments of the present disclosure, taken in conjunction with the accompanying drawings, in which like reference numerals are used for like elements, and in which: [Brief explanation of the drawings]
[0037] [Figure 1] FIG. 1 is a diagram of a device implementing techniques for rendering overlays in a hierarchically coded video sequence having an enhancement layer and a base layer, according to an embodiment. [Figure 2]2 is a diagram of a system including the first device of FIG. 1 and a second device configured to receive a hierarchically encoded video sequence in multiple data streams from the first device, according to an embodiment. [Figure 3] 2 is a diagram of a system including the first device of FIG. 1 and a second device and a third device each configured to receive a hierarchically encoded video sequence from the first device in a separate data stream, according to an embodiment. [Figure 4] FIG. 1 illustrates misalignment between glyphs rendered in overlay at different resolutions. [Figure 5] 5 illustrates how the misalignment of FIG. 4 is reduced using alignment techniques described herein, according to an embodiment. [Figure 6] FIG. 10 illustrates alignment of overlays containing patterns with glyphs rendered at different resolutions, according to an embodiment. [Figure 7] FIG. 10 illustrates alignment between overlays containing both dynamic and static patterns of glyphs, according to an embodiment. [Figure 8] 1 is a flowchart of a method for rendering an overlay in a hierarchically coded video sequence having an enhancement layer and a base layer, according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0038] Layered coding is a video compression technique that improves efficiency by organizing data into multiple layers. The base layer contains the essential video information for basic playback and offers a lower resolution and bitrate to ensure compatibility with a wide range of devices and network conditions. This layer ensures that the video can still be viewed with acceptable quality even when bandwidth is limited. Meanwhile, the enhancement layer contains additional data (residual data) that refines and enhances the video quality, providing higher resolution and better visual detail. When both layers are available, they work together to provide an improved viewing experience.
[0039] If the base layer can be viewed alone, it may be advantageous to include some overlay information in this layer. Overlays often include elements that provide context or supplemental information to the video content, such as subtitles, annotations, or graphics. Embedding such overlays in the base layer allows viewers to still access such information even if only the base layer is available, for example, due to bandwidth constraints or device limitations.
[0040] If the intention is to provide both a base layer alone (i.e., a layer in the first stream) and a base layer enhanced by an enhancement layer (i.e., a layer in the second stream), it is preferable to use different overlays (containing the same glyphs / information) for each resolution. For example, if the overlay contains text, it is often better to render the overlay directly at the native resolution of each layer, rather than rendering the overlay at the highest resolution and scaling it down, or vice versa.
[0041] A potential problem is that glyphs may not align perfectly due to the separate rendering of the overlay for each resolution, which increases the bit size of the enhancement layer. Such misalignments can be caused by positioning differences, hinting, kerning, and other typographic adjustments.
[0042] FIG. 4 shows an example of misalignment that can result from hinting. Hinting involves adjusting the display of a vector-based glyph (character) to more precisely align it with the screen's pixel grid. In FIG. 4, a single glyph 410 (the letter "T") is used as an example. The left side of FIG. 4 shows the glyph 410 positioned without hinting on pixel grids of two different resolutions (illustrated by one-dimensional lines 412). The top side shows the lower first resolution 106, and the bottom side shows the higher second resolution 104. On the right side of FIG. 4, hinting is applied. As can be seen, the glyph 410 is positioned slightly to the right in the lower resolution 106 compared to its position in the higher resolution 104. While only one glyph 410 is shown for simplicity, it should be noted that in an overlay containing multiple glyphs (i.e., a first pattern), small sub-pixel differences between the positions of each individual glyph will accumulate. At the end of a pattern with multiple glyphs, the glyphs may be significantly misaligned. This misalignment results in a larger difference between the overlay at the first resolution 104 and the overlay at the second resolution 106, leading to a larger bitrate due to the larger residual. In Figure 4, the example focuses on misalignment along the x-dimension, and the drawing only shows the x-coordinate for simplicity. However, it is important to note that similar misalignment can also occur along the y-dimension.
[0043] This disclosure provides techniques for achieving two overlays, one at a lower resolution and one at a higher resolution, while simultaneously reducing or minimizing the size of the enhancement layer. This is achieved by controlling the rendering of one overlay based on the rendering position of one or more glyphs in the other overlay. In other words, the rendering of an overlay at one resolution is used to guide the rendering of an overlay at the other resolution. FIG. 1 illustrates, by way of example, a device (system, component, etc.) 100 that implements such a technique.
[0044] Device 100 receives a video sequence 102 including a plurality of image frames. Device 100 is configured to render an overlay within a hierarchically encoded video sequence having an enhancement layer and a base layer. Thus, device 100 can render video sequence 102 in a first (e.g., lower) resolution 106 and a second (e.g., higher) resolution 104. To this end, device 100 includes a video scaler component 108 configured to scale the video sequence 102 in the higher resolution 104 to the lower resolution 106. In this manner, device 100 is configured to provide the video sequence 102 in the first resolution 106 and the second resolution 104 using video scaler component 108 according to a ratio (proportionality) between the first resolution and the second resolution. The proportionality is then used to guide the rendering of an overlay on the video sequence in the second resolution, as described herein. The ratio may be predetermined or may be configurable within device 100.
[0045] It should be noted that in some instances (not shown in FIG. 1), the original video sequence 102 is first scaled to provide a higher resolution 104 .
[0046] In the example shown in FIG. 1, the lower resolution 106 is used to guide the rendering of overlays at the higher resolution 104. However, in other examples, the higher resolution 104 can also be used to guide the rendering at the lower resolution 106. Additionally, while the examples of FIGS. 1-7 are limited to two resolutions (low and high), the techniques described herein are extendable to hierarchical coding using more than two layers. In such cases, the rendering of overlays in any one of the layers / resolutions can be used to guide the rendering of overlays in the remaining layers. For example, in a three-layer coding scenario, the rendering of overlays in the middle layer can be used to guide the rendering of overlays in both the base layer and the top layer.
[0047] Device 100 includes one or more overlay rendering components. For ease of explanation, FIG. 1 illustrates a first overlay rendering component 111 responsible for rendering first overlay 116 and a second overlay rendering component 110 responsible for rendering second overlay 114, along with guidance 138 from first overlay rendering component 111. However, in other examples, device 100 may include a single overlay rendering component configured to render both first overlay 116 and second overlay 114.
[0048] 1 , a first rendering component 111 renders a first overlay 116 including a first pattern 118 of one or more glyphs in an animated sequence (in this example, the pattern includes a string of characters spelling out the word "Text") at a first, lower, resolution 106. Thus, each glyph (character) in the pattern 118 is rendered at the first resolution 106 at a respective pixel location in the first overlay 116. A second rendering component 110 renders a second overlay 114 including the same first pattern 118 of one or more glyphs in the animated sequence at a second, higher resolution 104. The device 100 is configured such that the rendering of the second overlay 114 is controlled using at least one of the pixel locations of the glyphs rendered in the first overlay 116. In this way, the rendering of the second overlay 114 is controlled to render the first glyph of the first pattern 118 at pixel locations 138 guided by the pixel locations of the first glyph rendered in the first overlay 116.
[0049] The mapping between positions at different resolutions is performed according to the ratio or scale factor between the first and second resolutions. For example, if the ratio is 1:2 (i.e., the second resolution is twice the first resolution), then pixel position (x, y) at the first resolution can be mapped to pixel position (2x, 2y) at the second resolution to ensure alignment. Different ratios result in other mapping rules.
[0050] Figure 5 visualizes an example of how misalignment between overlays of different resolutions 106, 104 can be reduced by controlling the rendering, as described in conjunction with Figure 4. Like Figure 4, Figure 5 illustrates only the x coordinate for simplicity, and focuses only on the x dimension in the example describing how the rendering of the glyphs is guided. However, it is important to note that the guidance process can also be applied along the y dimension.
[0051] For example, consider the letter "T" in the first pattern 118 of FIG. 1. In the example of FIG. 5, this glyph 410 is rendered at a first pixel location 502 (x=1) in the first overlay at the lower resolution 106. Using the determined 1:2 ratio, the rendering of the first pattern 118 in the second overlay at the higher resolution 104 can be controlled so that the glyph 410 is rendered at a second pixel location 504 (x=2), which aligns the location of the glyph 410 between the two resolutions 106, 104. Compare this to the example of FIG. 4, where the use of kerning alone results in a misaligned rendering of the glyph (x=1 in both resolutions). In FIG. 5, the top-left pixel location 502 of the glyph 410 is used for alignment purposes. However, this is merely an example, and other locations for the glyph 410 can also be used for alignment. For example, the alignment point can be the center pixel location, the bottom left pixel location, or the centroid of the glyph 410. By choosing different alignment points, the method can be fine-tuned to better suit specific typographic requirements and visual consistency across different resolutions.
[0052] Returning to FIG. 1 , second overlay 114 is overlaid onto the video sequence at second resolution 104 to result in video sequence 122. For simplicity, this step is not included in FIG. 1 . First overlay 116 is overlaid onto the video sequence at first resolution 106 to result in video sequence 120. For simplicity, this step is not included in FIG. 1 . This video sequence 120 is then encoded into a base layer 136 using a base codec 124, such as AVC (H.264), HEVC (H.265), VP9, or AV1. Additionally, base layer 136 is decoded (not shown in FIG. 1 ) into a decoded base layer 123 (corresponding to video sequence 120 at first resolution 106 including first overlay 116), and decoded base layer 123 is upscaled using video scaling component 121 to create an upscaled version 126 at second resolution 104. A residual 130, representing the difference between the upscaled version 126 and the video sequence 122, is determined using a residual determination component 128. This residual 130 is then encoded into an enhancement layer 134 using an LCEVC enhancement codec 132. It should be noted that while LCEVC is used as an example, other hierarchical codecs such as SVC (Scalable Video Coding) for H.264 / AVC, SHVC (Scalable High-efficiency Video Coding) for H.265 / HEVC, or progressive JPEG may also be used depending on the desired output format.
[0053] In some examples, device 100 is a camera. In these examples, video sequence 102 may be captured by the camera. In other examples, device 100 is coupled to a camera that captures video sequence 102. In yet other examples, device 100 receives video sequence 102 from external or internal storage.
[0054] In general, a device (e.g., a camera, a server) implementing components 108, 110, 111, 121, 124, 128, or 132 of FIG. 1 may include circuits configured to implement the components, and more specifically, their functionality. The described features of device 100 may be advantageously implemented in one or more computer programs executable on a programmable system including at least one programmable processor coupled to receive data and instructions from and transmit data and instructions to a data storage system. The computer program may, for example, execute instructions for implementing the techniques described herein, and the instructions may be stored on a non-transitory computer-readable storage medium. Processors suitable for executing a program of instructions include, by way of example, both general-purpose and special-purpose microprocessors, and the sole processor or one of multiple processors or cores of any type of computer. The processor may be supplemented by or incorporated in an application-specific integrated circuit (ASIC). In some examples (not shown in FIG. 1 ), the components and functionality described herein are implemented in multiple connected devices.
[0055] 2 illustrates an example system 200 including a first device 100 (e.g., the device shown in FIG. 1) and a second device 206. The first device is configured to transmit a base layer in a first data stream 202 and both the base layer and enhancement layers in a second data stream 204. The second device 206 is configured to receive both the first data stream 202 and the second data stream 204, using each for different purposes. For example, the second device 206 can use the low-bitrate first data stream 202 for real-time display, such as for monitoring purposes, and the high-bitrate second data stream 204 for storage purposes. In other scenarios, the second device 206 can store both data streams 202 and 204 and apply different retention policies to each.
[0056] In some examples (not shown in FIG. 2 ), the first device 200 is configured to transmit only one data stream to the second device 206: either the base layer alone or both the base layer and enhancement layers. The selection between these options is based on an indication of network congestion on the communication channel used to transmit the data stream. For example, in the absence of an indication of network congestion, the first device 200 may transmit a data stream including both the base layer and enhancement layers to the second device. However, upon detecting network congestion, the first device 200 may adjust its transmission to transmit only the base layer, excluding the enhancement layers. Advantageously, such a system facilitates reliable delivery of essential video content in the base layer even under adverse network conditions, thereby maintaining uninterrupted playback. Additionally, the system may optimize bandwidth usage and enable enhancement layers to be transmitted only when the network can support them, thereby improving overall streaming efficiency and quality.
[0057] FIG. 3 illustrates an example system 300 including a first device 100 (e.g., the device shown in FIG. 1), a second device 306, and a third device 308. Similar to the setup of FIG. 2, the first device 100 is configured to transmit a base layer in a first data stream 302 and both the base layer and enhancement layers in a second data stream 304. In this example, the second device 306 is configured to receive the first data stream 302, and the third device 308 is configured to receive the second data stream 304. One possible implementation of this system 300 includes real-time monitoring and high-quality recording. The second device 306, which receives the first data stream 302, can be used for real-time monitoring applications where low latency and continuous playback are important. This setup ensures that essential video content is delivered reliably, even under varying network conditions. Meanwhile, the third device 308, which receives the second data stream 304, can be used for high-quality recording or broadcasting, taking advantage of the enhanced video quality provided by the additional enhancement layers.
[0058] It should be noted that the examples shown in Figures 2 and 3 can be combined in any suitable way to create a versatile and adaptive video streaming system. In addition, the network congestion management techniques described herein can be integrated into such a setup, or into either of the systems shown in Figure 2 or 3, etc.
[0059] 6-7 illustrate, by way of example, additional details that may be implemented for the above-described overlay alignment techniques. For example, as shown in Figure 6, the rendering of the second overlay (at the second resolution 104) may be performed using multiple pixel locations from the rendering of the glyphs in the first overlay at the first resolution 106. In Figure 6, the pattern of the glyphs 118 includes the text "CAM NW."
[0060] In some examples, the rendering of the second overlay at the second resolution 104 is guided by at least two pixel locations of the glyphs rendered at the first resolution 106. In addition to determining a first pixel location 502a for the first glyph (“C”) in the first overlay, a third pixel location 502b for the second glyph (“N”) in the first overlay is also determined. The rendering of the second overlay at the second resolution 104 is then controlled such that the rendering of the first glyph (“C”) of the first pattern 118 is guided to the second pixel location 504a at the second resolution 104, where the second pixel location 504a is obtained by mapping the first pixel location 502a to the second pixel location 504a according to a ratio between the first resolution 106 and the second resolution 106. Similarly, the rendering of the second overlay at the second resolution 104 is then controlled such that the rendering of the second glyph ("N") of the first pattern 118 is directed to the fourth pixel location 504b at the second resolution 104, which is obtained by mapping the third pixel location 502b to the fourth pixel location 504b according to a ratio. For example, if the ratio is 1:2, then the x-value of the first pixel location 502a is 2 and the x-value of the third pixel location 502b is 6, resulting in an x-value of the second pixel location 504a being 4 and an x-value of the fourth pixel location 504b being 12 (using the mapping rule 2x).
[0061] In some examples, such as that shown in FIG. 6 , first pattern 118 includes a first group of glyphs (“CAM”) and a second group of glyphs (“NW”) separated by white space. As noted above, white space is merely one way of identifying groups of glyphs for the techniques described herein. Other suitable characters or markers, such as “zero width space” (ZWSP) characters, may be employed to define boundaries between groups of glyphs. These group boundary rules help organize the glyphs in first pattern 118 into separate segments (e.g., words or other groups, such as groups of letters and groups of numbers) and may be particularly useful for rendering and alignment purposes.
[0062] In such an example, the rendering of the second overlay at the second resolution 104 may be guided such that the pixel locations of the first glyph "C" (part of the first group "CAM") at the second resolution 104 are aligned based on the pixel locations of the corresponding first glyph at the first resolution 106, and similarly, the rendering location of the second glyph "N" (part of the second group "NW") at the second resolution 104 is guided based on the pixel locations of the corresponding second glyph at the first resolution 106. This method facilitates alignment of each group of glyphs across different resolutions and reduces potential misalignments caused by scaling, kerning, and hinting. In some cases, the rendering of the second overlay involves rendering the glyphs ("A" and "M") of the first group ("CAM") using a glyph layout algorithm (text shaping). Here, pixel locations of glyphs ("A" and "M") different from the first glyph ("C") are determined using this algorithm and the second pixel location (504a). Thus, the rendering of the first glyph ("C") is precisely guided by its rendering location 502a in the first overlay at the first resolution 106 and the ratios as described above. The remaining glyphs in the group ("A" and "M") are then rendered based on the guided rendering location 504a of the glyph ("C") and a glyph layout algorithm that implements characteristics such as kerning and hinting to determine their pixel locations. A similar approach is followed for the second group of glyphs ("NW"). In this way, the group or word as a whole is aligned between resolutions 106 and 104, but the visualization of each group follows typographical rules. This approach ensures that each glyph in the group is positioned accurately according to typographic standards, maintaining visual consistency while simultaneously achieving alignment at different resolutions.
[0063] In some examples, not shown in FIG. 6, a subset of glyphs (e.g., multiple glyphs), or all glyphs in a group of glyphs, are aligned by guiding each glyph at the second resolution 104 based on the rendering position of the same glyph at the first resolution 106.
[0064] In some cases, the overlay includes a combination of static and dynamic patterns of glyphs. The static patterns remain the same in both the first and second resolutions, while the dynamic patterns vary depending on which resolution is rendered. For example, an overlay indicating the bitrate of a data stream (e.g., data streams 202 and 204 in FIG. 2 or data streams 302 and 304 in FIG. 3) may include a static portion such as “Mbit / second” and a dynamic portion such as “X,” where “X” represents the actual bitrate value and varies between resolutions. Similarly, an overlay indicating the resolution of a video sequence may include a static portion such as “MPixels” and a dynamic portion such as “Y,” where “Y” represents the resolution value and varies between resolutions. In such an example, the dynamic portion may occupy different amounts of space depending on its value, potentially complicating alignment of the static portion. FIG. 7 illustrates, by way of example, a technique for maintaining alignment of the static portion to reduce the bitrate required to encode an enhancement layer.
[0065] 7, the overlay represents the bit rate of the data stream and includes dynamic portions (second and third patterns) 706, 708 that vary depending on the resolution 106, 104, and one static portion (first pattern) 118 that does not vary between the resolutions 106 and 104. To facilitate alignment of the static portion 118 between the resolutions 106 and 104, the following technique can be used: In this example, the second pattern includes the text "12.3" and the third pattern includes the text "227.4".
[0066] To facilitate alignment of the static first pattern 118 ("MB / s"), a first pixel area 702 required to render a second pattern 706 in the video sequence at the first resolution 106 is determined. Additionally, a second pixel area 704 required to render a third pattern 708 in the video sequence at the first resolution 106 is determined. The rendering of the first pattern 118 in the first overlay is then controlled so that all glyphs of the first pattern 118 are rendered outside the combined area of the first pixel area 702 and the second pixel area 704. This combined area corresponds to the larger of the two pixel areas (in this case, the second pixel area 704), ensuring that the static pattern is positioned beyond the maximum space occupied by the dynamic pattern.
[0067] In this example, the first pixel region and the second pixel region cover a portion of the first overlay from x=1 to x=5. This means that the first glyph ("M") of the first pattern 118 is rendered at x=6 in the first resolution 106. As a result (according to the 1:2 ratio described above), the first glyph ("M") of the first pattern 118 is rendered at x=11 in the second resolution 104. The dynamic patterns (the second pattern 706 in the overlay at the first resolution 106 and the third pattern 708 in the overlay at the second resolution 104) are rendered in corresponding positions between the first and second overlays, but occupy different amounts of space in each overlay.
[0068] As a result, the distance between the second dynamic pattern 706 and the first static pattern 118 in the first overlay is greater than the distance between the third dynamic pattern 708 and the first static pattern 118 in the second overlay. However, aligning the static portions 118 of the overlay reduces the bit rate required to encode the enhancement layer compared to maintaining the same distance between the static and dynamic portions of the overlay across both resolutions 106 and 104. This approach can improve compression efficiency while accommodating the varying sizes of dynamic content.
[0069] FIG. 8 is a flowchart of an example method 800 for rendering overlays in a hierarchically coded video sequence having an enhancement layer and a base layer.
[0070] The method 800 includes a step S802 of representing the video sequence in a first resolution and a second resolution, one of the first resolution and the second resolution being lower than the other of the first resolution and the second resolution. In some examples, the first resolution is lower than the second resolution, and in some examples, the second resolution is lower than the first resolution.
[0071] The method includes rendering S804 a first overlay at a first resolution, the first overlay including a first pattern of one or more glyphs in the video sequence. The rendering of the first pattern may be performed using a glyph layout algorithm (i.e., a text shaping algorithm).
[0072] In some examples, the overlay includes a static first pattern that is the same regardless of the resolution to which the overlay belongs, and a dynamic portion, a second and third pattern, that differs depending on the resolution, such that the first overlay includes the second pattern and the second overlay (rendered within the video sequence at the second resolution) includes the third pattern. In these examples, method 800 may include controlling S806 the rendering of the static pattern in the first overlay using the determined pixel area in the first overlay for rendering the dynamic pattern. The controlling S806 includes determining a first pixel area required to render the second pattern within the video sequence at the first resolution and determining a second pixel area required to render the third pattern within the video sequence at the first resolution, wherein the rendering of the first overlay is controlled to render one or more glyphs of the first pattern outside both the first pixel area and the second pixel area.
[0073] The method 800 further includes determining S808 at least a first pixel location where a first glyph of the one or more glyphs of the first pattern is to be rendered within the video sequence at the first resolution.
[0074] The method 800 further includes rendering S810 a second overlay within the video sequence at a second resolution, the second overlay including the first pattern as described above.
[0075] The rendering of the second overlay is controlled S812 according to the determined S808 pixel locations for the glyphs rendered in the video sequence at the first resolution. For each of the determined S808 pixel locations, a corresponding pixel location of the same glyph is controlled S812 in the video sequence at the second resolution. For example, the rendering of the second overlay is controlled S812 to render first glyphs of the first pattern at second pixel locations in the video sequence at the second resolution, the second pixel locations being obtained by mapping the first pixel locations to the second pixel locations according to a ratio between the first resolution and the second resolution to render the first glyphs of the first pattern at the determined second pixel locations in the video sequence at the second resolution.
[0076] In some examples, not all glyphs of the first pattern that are rendered in the animation sequence at the second resolution are controlled using the corresponding rendering positions of the glyphs in the animation sequence at the first resolution. In these examples, rendering the second overlay may include rendering S814 the remaining glyphs (i.e., glyphs that are not controlled according to step S812) using a glyph layout algorithm (i.e., a text shaping algorithm), where the pixel positions of each of the glyphs that are not controlled according to step S812 are determined using the glyph layout algorithm and the positions of the glyphs that are controlled according to step S812.
[0077] The method further includes encoding S816 a video sequence having a lower resolution in the base layer, and encoding S818 a residual between the video sequence having a higher resolution and the video sequence having a lower resolution in the enhancement layer.
[0078] The above-described embodiments should be understood as illustrative of the present invention. Further embodiments of the present invention are contemplated. For example, explicit congestion notification (ECN) can be another indicator of network congestion. It should be understood that any feature described in connection with any one embodiment can be used alone or in combination with other features described, and can also be used in combination with any other feature or features of the embodiments, or any other combination of embodiments. Furthermore, equivalents and modifications not described above may be employed without departing from the scope of the present invention, as defined in the appended claims.
Claims
1. 1. A method for rendering an overlay in a hierarchically coded video sequence having an enhancement layer and a base layer, comprising: Rendering a video sequence at a first resolution and a second resolution, one of the first resolution and the second resolution being lower than the other of the first resolution and the second resolution, wherein the rendering includes scaling the video sequence of a higher resolution to a lower resolution. and after said scaling, said method comprises: rendering a first overlay including a first pattern of one or more glyphs within the animation sequence at the first resolution, wherein a first glyph of the one or more glyphs is rendered at a first pixel location within the animation sequence at the first resolution; rendering a second overlay within the video sequence at the second resolution, the second overlay including the first pattern of one or more glyphs; rendering a second overlay, wherein the rendering of the second overlay is controlled by guidance from the rendering of the first overlay to render the first glyphs of the first pattern at second pixel locations within the animation sequence at the second resolution, the second pixel locations being obtained by mapping the first pixel locations to the second pixel locations according to a ratio between the first resolution and the second resolution; encoding the video sequence having the lower resolution in the base layer; encoding a residual between the video sequence having the higher resolution and the video sequence having the lower resolution in an enhancement layer; The method further comprises:
2. the first pattern includes a second glyph, and rendering the first overlay includes rendering the second glyph at a third pixel location within the video sequence at the first resolution; 2. The method of claim 1, wherein the rendering of the second overlay is controlled to render the second glyph at a fourth pixel location within the animation sequence at the second resolution, the fourth pixel location being obtained by mapping the third pixel location to the fourth pixel location according to the ratio.
3. 3. The method of claim 2, wherein the first pattern includes a first group and a second group of glyphs, each group of glyphs including a plurality of glyphs, the first glyph being part of the first group and the second glyph being part of the second group.
4. 4. The method of claim 3, wherein the rendering of the second overlay includes rendering glyphs that are different from the first glyph in the first group of glyphs using a glyph layout algorithm, and pixel locations of each of the glyphs that are different from the first glyph in the first group of glyphs are determined using the glyph layout algorithm and the second pixel locations.
5. 3. The method of claim 2, wherein the first pattern includes a first group of glyphs and a second group of glyphs, the first glyphs and the second glyphs being part of the first group.
6. The method of claim 1 , wherein the first resolution is the lower resolution.
7. the first overlay includes a second pattern of one or more glyphs, and the second overlay includes a third pattern of one or more glyphs, the third pattern being different from the second pattern, and the method further comprising: determining a first pixel area required to render the second pattern within the video sequence at the first resolution, and determining a second pixel area required to render the third pattern within the video sequence at the first resolution; 2. The method of claim 1, wherein the rendering of the first overlay is controlled to render the one or more glyphs of the first pattern outside of both the first pixel region and the second pixel region.
8. transmitting the base layer in a first data stream and transmitting the base layer and the enhancement layer in a second data stream; The method of claim 1 further comprising:
9. transmitting a data stream including the base layer and the enhancement layer over a communication channel; receiving an indication of network congestion on the communication channel; adjusting the transmission of the data stream to not include the enhancement layer; The method of claim 1 further comprising:
10. A non-transitory computer-readable storage medium having stored thereon instructions for performing the method of claim 1 when executed on a device having processing capabilities.
11. 1. A device for rendering an overlay in a hierarchically coded video sequence having an enhancement layer and a base layer, the device comprising: Rendering a video sequence at a first resolution and a second resolution, one of the first resolution and the second resolution being lower than the other of the first resolution and the second resolution, wherein the rendering includes scaling the video sequence of a higher resolution to a lower resolution. and after said scaling, said device is configured for rendering a first overlay including a first pattern of one or more glyphs within the animation sequence at the first resolution, wherein a first glyph of the one or more glyphs is rendered at a first pixel location within the animation sequence at the first resolution; rendering a second overlay within the video sequence at the second resolution, the second overlay including the first pattern of one or more glyphs; rendering a second overlay, wherein the rendering of the second overlay is controlled by guidance from the rendering of the first overlay to render the first glyphs of the first pattern at second pixel locations within the animation sequence at the second resolution, the second pixel locations being obtained by mapping the first pixel locations to the second pixel locations according to a ratio between the first resolution and the second resolution; encoding the video sequence having the lower resolution in the base layer; encoding a residual between the video sequence having the higher resolution and the video sequence having the lower resolution in an enhancement layer; The device is further configured for:
12. The device of claim 11 , which is a camera and the video sequence is captured by the camera.
13. A system comprising a first device according to claim 11 and a second device, wherein the first device: transmitting the base layer in a first data stream and transmitting the base layer and the enhancement layer in a second data stream; configured for the second device is configured to receive the first data stream and the second data stream, and to use the first data stream for a first purpose and the second data stream for a second, different purpose.
14. 12. A system comprising a first device according to claim 11, a second device, and a third device, wherein the first device: transmitting the base layer in a first data stream and transmitting the base layer and the enhancement layer in a second data stream; configured for The system, wherein the second device is configured to receive the first data stream and the third device is configured to receive the second data stream.