MÉTODO DE DECODIFICAÇÃO DE UM SINAL DE VÍDEO
Patent Information
- Application Number
- BR112025020202
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-24
- Filing Date
- 2024-03-22
- Publication Date
- 2026-08-04
Smart Images

Figure 00000037_0000 
Figure 00000038_0000 
Figure 00000039_0000
Abstract
Description
1 / 31 “METHOD OF DECODING A VIDEO SIGNAL” FIELD OF THE INVENTION
[0001] The present invention relates to methods for use in video encoding technology. BACKGROUND
[0002] In an enhancement-type coding algorithm, such as Low Complexity Enhancement Video Coding (LCEVC) MPEG-5 Part 2, one or more layers of residual data can be used to improve the performance of a base coding algorithm.
[0003] In an enhancement-type encoder, residual data is calculated based on a comparison of a base decoded video signal and an original input video signal. In general, the base decoded video signal is the result of encoding and decoding the video signal according to a base codec. Each image element (such as a pixel) of the original input video signal may have a different value from the corresponding image element of the base decoded video signal. In general, the differences may comprise a mixture of higher values (positive residual data values) for some image elements, lower values (negative residual data values) for other image elements, and zero values.
[0004] The encoder produces a video signal comprising a base-encoded video signal and an enhancement-encoded video signal. The base-encoded video signal comprises the original input video signal encoded by the base codec. The enhancement-encoded video signal comprises the residual data.
[0005] In an enhancement-type decoder, residual data is used to recover the original input video signal based on the base decoded video signal. This typically involves adding the positive values of the residual data to the base decoded video signal and subtracting the negative values of the residual data from the base decoded video signal.
[0006] Encoding and decoding can be used Petition 870250106234, dated 11 / 19 / 2025, page 9 / 48 2 / 31 to compress and / or protect content communicated over a network, such as in a streaming service. Alternatively, encoding and decoding can be used in other data transport / transmission contexts, such as physical media (e.g., DVDs, portable flash memory).
[0007] The encoder can, for example, be implemented in a content creation, content distribution, or content streaming service.
[0008] The decoder can, for example, be implemented in consumer hardware for viewing decoded content, such as in a display device (e.g., a television) or in a separate device to receive encoded content and provide decoded content to the display device (e.g., a set-top box or a DVD player).
[0009] In WO 2020 / 188272, which is incorporated herein by reference, the residual data produced by the encoder and used by the decoder is reduced. This is achieved using time prediction. In time prediction, a time buffer in the decoder stores residual data associated with a previous frame of the video signal. To obtain a subsequent frame of the video signal, the decoder reads the residual data from the time buffer and retrieves the subsequent frame of the input video signal by combining the data from the subsequent frame of the base decoded video signal with the residual data stored in the buffer and any data from the subsequent frame of the enhanced encoded video signal.
[0010] When temporal prediction is used, subsequent frame data of the enhancement-encoded video signal may be residual data calculated based on a comparison of an original input video signal with the sum of a base decoded video signal and the residual data stored in the temporal buffer. In other words, subsequent frame data of the enhancement-encoded video signal may comprise adjustments to the residual data stored in the temporal buffer so that the data Petition 870250106234, dated 11 / 19 / 2025, page 10 / 48 3 / 31 residuals become applicable to the subsequent framework.
[0011] This has the advantage that, when the residual data of the subsequent frame is identical or similar to the residual data of the previous frame, a reduced amount of residual data needs to be included in the enhanced encoded video signal.
[0012] In WO 2020 / 188272, the encoder can control the use of time prediction in the decoder. This control can be achieved with parameters in the video signal, including a “temporal_enabled” parameter (specifying whether time prediction should be used when decoding a frame) and a “time update bit” parameter (specifying whether the time buffer should be updated). Updating the time buffer includes setting the residual data of the time buffer to zero values.
[0013] In general, it is desirable to reduce the size of the encoded video signal and reduce the processing resources required to encode and decode the video signal. In this application, it is specifically considered how to perform time prediction with reduced processing resources without increasing the size of the encoded video signal. SUMMARY OF THE INVENTION
[0014] According to a first aspect, a method for decoding a video signal is described comprising: obtaining an encoded frame from the video signal, the encoded frame comprising a base encoding and an enhancement encoding; determining whether a time buffer should be updated; and if it is determined that the time buffer should be updated: reconstructing a decoded frame based on the base encoding and enhancement encoding of the encoded frame, without reading the time buffer; and writing a set of new time predictions to the time buffer, based on the enhancement encoding of the encoded frame.
[0015] The base coding and the enhancement coding of the encoded frame can be stored and / or transmitted separately. For example, in a video signal, the base coding of the frame Petition 870250106234, dated 11 / 19 / 2025, page 11 / 48 4 / 31 encoded can be part of a base-encoded video signal, and the encoded frame enhancement encoding can be part of an enhancement encoded video signal. This allows for greater flexibility between the base-encoded video signal and the enhancement encoded video signal. For example, the enhancement encoded video signal can have a higher frame rate than the base-encoded video signal, and the base encoding can be reused for multiple encoded frames. Similarly, the base-encoded video signal can have a higher frame rate than the enhancement encoded video signal, and the enhancement encoding can be reused for multiple encoded frames.
[0016] Alternatively, the base coding and enhancement coding of the encoded frame can be stored and / or transmitted together. For example, the base coding and enhancement coding are part of a single combined video signal.
[0017] According to the claimed features, a temporal update operation is merged with the temporal prediction of a frame. This eliminates a previously necessary operation of writing zero values to the temporal buffer during the temporal update operation and reading those same zero values from the temporal buffer in the subsequent temporal prediction operation. This reduces the processing resources required for decoding a video signal compatible with temporal prediction.
[0018] Temporal predictions can be residual. Specifically, temporal predictions can be residuals generated or intended for a frame, which are at least partially reused in the decoding process of a subsequent frame.
[0019] The recording of the set of new temporal predictions in the temporal buffer, based on the enhancement coding of the encoded frame, may specifically include: obtaining a set of residuals from the enhancement coding of the encoded frame, where the residuals are differences between: a base decoded frame obtained by decoding the Petition 870250106234, dated 11 / 19 / 2025, page 12 / 48 5 / 31 base encoding; and an input frame, where the base encoding is generated (i.e., was generated by an encoder) by encoding the input frame; and the writing of the residue set to the temporal buffer. The residues can then be read from the temporal buffer when decoding a subsequent frame. In other words, the input frame is a frame that was previously encoded by an encoder to generate the encoded frame. In other words,
[0020] In some embodiments, each frame of the video signal has a frame area divided into a plurality of tiles and / or a plurality of blocks, and the residual set includes residual values for a subset of the plurality of tiles or a subset of the plurality of blocks, wherein the recording of the set of new temporal predictions in the temporal buffer, based on the encoded frame enhancement coding, further includes recording a zero value in the temporal buffer for a tile of the frame area that is not in the subset of the plurality of tiles and / or for a block of the frame area that is not in the subset of the plurality of blocks. In other words, a zero value may be recorded in a block or tile for which no residual (non-zero) value was included in the encoded frame enhancement coding.
[0021] In this case, a “block” is a coding unit for encoding and decoding. For example, the size of a block may depend on a directional decomposition transformation used in encoding and decoding.
[0022] Here, a “tile” is a group of blocks that cover a region of a frame. Timing prediction can be applied to an entire frame or to individual tiles within a frame. A tile size parameter can be included as overhead in the video signal. Increasing the tile size can reduce the overhead of timing prediction control per tile, while reducing the tile size can increase timing prediction flexibility.
[0023] Furthermore, a canvas can be partitioned into a plurality of “planes”, and each plane can be partitioned into tiles and / or blocks. The planes can, for example, be color channels that combine to generate a Petition 870250106234, dated 11 / 19 / 2025, page 13 / 48 6 / 31 multicolored image.
[0024] In tile- and / or block-divided implementations, the method may comprise: constructing a complete set of new timing predictions for all tiles and / or all blocks in the frame area, consisting of the set of residuals obtained from the encoded frame enhancement coding and zero values for any tiles and / or blocks for which no residual value is included in the encoded frame enhancement coding; and writing the complete set of new timing predictions to the timing buffer. More specifically, the residual values of an entire frame area may be written to the timing buffer in a single step.
[0025] In some embodiments, the method also includes: if it is determined that the time buffer should not be updated: reading the time buffer to obtain a set of previous time predictions; and reconstructing the decoded frame based on the base coding and the enhancement coding of the encoded frame and based on the set of previous time predictions. The previous time predictions may include a set of residuals for a previous frame in the video signal preceding the encoded frame. In addition, the enhancement coding of the encoded frame may include adjustments to the set of previous time predictions so that it is applicable to the base coding to reconstruct the decoded frame.
[0026] In some embodiments, the method also includes: receiving an update parameter in the video signal, where the determination of the temporal buffer update is based on the update parameter. The update parameter may be part of the encoded frame. Alternatively, the update parameter may be part of a video signal control component, separate from the encoded frame. For example, the update parameter may be included in the video signal in a manner similar to the “temporal_enabled” parameter of WO 2020 / 188272.
[0027] According to a second aspect, the following disclosure provides a decoder for decoding a video signal, the Petition 870250106234, dated 11 / 19 / 2025, page 14 / 48 7 / 31 decoder comprising: a memory comprising a time buffer configured to store a set of time predictions; and a processor configured to: obtain an encoded frame from the video signal, the encoded frame comprising a base encoding and an enhancement encoding; determine whether to update a time buffer; and if it is determined that the time buffer should be updated: reconstruct a decoded frame based on the base encoding and enhancement encoding of the encoded frame, without reading the time buffer; and write a set of new time predictions to the time buffer, based on the enhancement encoding of the encoded frame. In other words, the decoder is configured to execute a method according to the first aspect.
[0028] The processor can be configured to write the set of new time predictions to the time buffer, based on the enhancement coding of the encoded frame, by: obtaining a set of residues from the enhancement coding of the encoded frame, wherein the residues are differences between: a base decoded frame obtained by decoding the base coding; and an input frame, wherein the base coding is generated (i.e., was generated by an encoder) by encoding the input frame; and writing the set of residues to the time buffer.
[0029] In some embodiments, each frame of the video signal has a frame area divided into a plurality of tiles and / or a plurality of blocks, and the residual set includes residual values for a subset of the plurality of tiles or a subset of the plurality of blocks, where the processor is configured to write the set of new temporal predictions to the temporal buffer, based on the enhancement encoding of the encoded frame, by writing a zero value to the temporal buffer for a tile of the frame area that is not in the subset of the plurality of tiles and / or a block of the frame area that is not in the subset of the plurality of tiles.
[0030] Optionally, the processor is configured to write the set of new timing forecasts to the timing buffer, based on Petition 870250106234, dated 11 / 19 / 2025, p. 15 / 48 8 / 31 Encoded frame enhancement coding, by means of: constructing a complete set of new temporal predictions for all tiles and / or all tiles in the frame area, consisting of the set of residuals obtained from the encoded frame enhancement coding and zero values for any tiles and / or blocks for which there are no residual values included in the encoded frame enhancement coding; and recording the complete set of new temporal predictions in the temporal buffer.
[0031] The processor can also be configured to: determine that the time buffer should not be updated; read the time buffer to obtain a set of previous time predictions; and reconstruct the decoded frame based on the base encoding and enhancement encoding of the encoded frame and based on the set of previous time predictions. Optionally, the previous time predictions comprise a set of residues for a previous frame in the video signal preceding the encoded frame.
[0032] The processor can be configured to: receive a refresh parameter in the video signal; and determine whether the time buffer should be refreshed based on the refresh parameter. The refresh parameter can be part of the encoded frame. Alternatively, the refresh parameter can be part of a video signal control component, separate from the encoded frame.
[0033] According to a third aspect, a decoder is encoded to execute a method according to the first aspect or configured to execute a method according to the first aspect using a combination of one or more processors that execute instructions from computer programs to execute one part of the method and application-specific hardware configured to execute another part of the method.
[0034] According to a fourth aspect, a computer program is provided comprising instructions which, when executed by one or more processors, cause the processors to execute a method in accordance with the first aspect. Petition 870250106234, dated 11 / 19 / 2025, page 16 / 48 9 / 31
[0035] According to a fifth aspect, a computer-readable storage medium is provided that stores instructions which, when executed by one or more processors, cause the processors to execute a method in accordance with the first aspect.
[0036] According to a sixth aspect, a chipset configured to execute a method according to the first aspect is provided. For example, the chipset may include at least one ASIC adapted to execute all or part of the method according to the first aspect. The chipset may also include memory that stores instructions which, when executed by one or more processors, cause the processors to perform an additional part of a method according to the first aspect.
[0037] According to a seventh aspect, a decoder is provided that includes a decoder according to the second aspect or the third aspect.
[0038] According to an eighth aspect, a method is provided of adapting a decoder, chipset, or decoder to provide a decoder, chipset, or decoder configured to perform a method according to the first aspect. For example, after a customer purchases a device, the device may be adapted by means of the manufacturer providing an update to the device, wherein the execution of the update constitutes the manufacture of a device with the capabilities of one of the previous aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a schematic diagram illustrating an example of a background coding process using time prediction;
[0040] Figure 2 is a schematic diagram illustrating an example of a background decoding process using time prediction;
[0041] Figures 3A and 3B are schematic diagrams that each illustrate an example of the background of an additional coding process;
[0042] Figures 4A and 4B are schematic diagrams. Petition 870250106234, dated 11 / 19 / 2025, page 17 / 48 10 / 31 which each illustrate an example of the background of a further decoding process;
[0043] Figure 5 is a flowchart that schematically illustrates an example of a method for updating a time buffer in a decoder;
[0044] Figure 6 is a flowchart that schematically illustrates a method for updating a time buffer in a decoder, according to the invention. DETAILED DESCRIPTION
[0045] This document describes a backward-compatible hybrid encoding technology.
[0046] The examples described here provide a flexible, adaptable, highly efficient, and computationally inexpensive encoding format that combines a different video encoding format, a base codec (e.g., AVC, HEVC, or any other present or future codec) with at least two levels of encoded data enhancement.
[0047] The general structure of the encoding scheme uses a downsampled source signal encoded with a base codec and adds one or more levels of correction data to the decoded output of the base codec to generate a corrected image.
[0048] In some cases, the base encoding is decodable by a hardware decoder, while the enhancement encoding is suitable for software processing.
[0049] Although the decoded output of the base codec is not intended for viewing, it is a fully decoded video at a lower resolution, making the output compatible with existing decoders and, when deemed appropriate, also usable as a lower resolution output.
[0050] Figure 1 shows a first example of encoder 100. The components illustrated can also be implemented as steps in a corresponding encoding process. Petition 870250106234, dated 11 / 19 / 2025, page 18 / 48 11 / 31
[0051] In encoder 100, a full-resolution input video 102 is processed to generate several encodings. A first encoding (base 110 encoding) is produced by feeding a base 106 encoder (e.g., AVC, HEVC, or any other codec) with a downsampled version of the input video, which is produced by downsampling 104 the input video 102. A second encoding (level 1 encoding 116, an example of enhancement encoding) is produced by applying an encoding operation 114 to the residuals obtained by the difference 112 between the reconstructed base codec video and the downsampled version of the input video. The reconstructed base codec video is obtained by decoding the output of the base 106 encoder with a base 108 decoder.A third encoding (level 2 encoding 128, another example of enhancement encoding) is produced by processing 126 the residues obtained by the difference 124 between an oversampled version of a corrected version of the reconstructed base codec video and the input video 102. The corrected version of the reconstructed base codec video is obtained by combining 120 the reconstructed base codec video and the residues obtained by applying a decoding operation 118 to level 1 encoding 116.
[0052] The level 1 encoding operation 114 operates with an optional level 1 temporal buffer 130, which can be used to apply temporal processing as described later. The level 2 encoding operation 126 also operates with an optional level 2 temporal buffer 132, which can be used to apply temporal processing as described later. The level 1 temporal buffer 130 and the level 2 temporal buffer 132 can operate under the control of a temporal selection component 134. The temporal selection component 134 can receive one or more input videos 102 and the downsampling output 104 to select a temporal mode. This is explained in more detail in later examples.
[0053] Figure 2 shows a first example of a 200 decoder. The components illustrated can also be implemented Petition 870250106234, dated 11 / 19 / 2025, page 19 / 48 12 / 31 as steps in a corresponding decoding process. The decoder receives the three streams (a base 210 encoding, a level 1 encoding 216, and a level 2 encoding 228) generated by an encoder, such as encoder 100 in Figure 1, along with headers 236 containing other decoding information. The base 210 encoding is decoded by a base 208 decoder corresponding to the base decoder used in the encoder, and its output is combined 238 with the decoded residues obtained by decoding 240 the level 1 encoding 216. The combined video is enhanced 242 and combined 244 with the decoded residues obtained by applying a decoding operation 246 to the level 2 encoding 228.
[0054] Figures 3A and 3B show different variations of a second example 300, 380 encoder. The second example 300, 380 encoder may comprise an implementation of the first example 100 encoder from Figure 1. In the examples in Figures 3A and 3B, the coding steps are expanded in more detail to provide an example of how the steps can be performed. Figure 3A illustrates a first variation with timing prediction provided only at the second level of the enhancement process, i.e., with respect to level 2 coding. Figure 3B illustrates a second variation with timing prediction performed at both enhancement levels (i.e., levels 1 and 2).
[0055] Base 310 encoding is substantially created by a process as explained with reference to Figure 1. That is, an input video 302 has a 304 downsampling (i.e., a 304 downsampling operation is applied to the input video 302 to generate a downsampled input video). The downsampled video obtained by the 304 downsampling of the input video 302 is then encoded using a first base encoder 306 (i.e., an encoding operation is applied to the downsampled input video to generate base 310 encoding using a first encoder or base encoder 306). Preferably, the first encoder or base encoder 306 is a Petition 870250106234, dated 11 / 19 / 2025, page 20 / 48 13 / 31 codec suitable for hardware decoding. Base 310 encoding can be called base layer or base level.
[0056] As noted above, the enhancement-encoded video signal may comprise two or more enhancement encodings. A first level of enhancement (described here as “level 1”) provides a correction data set that can be combined with a decoded version of the base encoding to generate a corrected image. This first enhancement encoding is illustrated in Figures 1 and 3 as level 1 316 encoding. The enhancement encoding may be generated by an enhancement encoder. The enhancement encoder may be different from the base 306 encoder used to generate the base 310 encoding.
[0057] To generate the level 1 316 encoding, the base 310 encoding is decoded using a base 308 decoder (i.e., a decoding operation is applied to the base 310 encoding to generate a decoded base). The difference 312 between the decoded base and the downsampled input video obtained by downsampling 304 of the input video 302 is then created (i.e., a subtraction operation 312 is applied to the downsampled input video and the decoded base to generate a first set of residuals). Here, the term residuals is used in the same way known in the art, i.e., the error between a reference frame and a desired frame. Here, the reference frame is the decoded base and the desired frame is the downsampled input video.Thus, the residues used in the first level of enhancement can be considered as corrected video, since they “correct” the decoded base to the downsampled input video that was used in the base encoding operation.
[0058] The difference 312 is then encoded to generate the Level 1 encoding 316 (i.e., an encoding operation is applied to the first set of residues to generate a first enhancement encoding 316).
[0059] In the implementation example of Figures 3A and 3B, Petition 870250106234, dated 11 / 19 / 2025, page 21 / 48 14 / 31 The coding operation comprises several steps, each of which is optional and preferential and offers specific benefits.
[0060] In Figure 3, the steps include a transformation step 336, a quantization step 338 and an entropy encoding step 340.
[0061] Although not shown in the Figures, in some examples, the coding process identifies whether the waste sorting mode is selected. If the waste sorting mode is selected, the waste sorting step can be performed (i.e., a waste sorting operation can be performed in the first waste sorting step to generate a sorted set of waste). The sorted set of waste sorted can be filtered so that not all waste is coded in the first enhancement coding 316 (or correction flow).
[0062] The first set of residues, or the first sorted or filtered set of residues, are then transformed 336, quantized 338, and entropy-coded 340 to produce Level 1 coding 316 (i.e., a transformation operation 336 is applied to the first set of residues or the first filtered set of residues, depending on whether sorting mode is selected or not, to generate a transformed set of residues; a quantization operation 338 is applied to the transformed set of residues to generate a quantized set of residues; and an entropy coding operation 340 is applied to the quantized set of residues to generate the first level of enhancement coding 316). Preferably, the entropy coding operation 340 can be a Huffman coding operation or a run-length coding operation, or both.Optionally, a control operation (not shown in the Figures) can be applied to the quantized set of residuals to correct for the effects of the classification operation.
[0063] As noted above, the enhancement-encoded video signal may comprise a first enhancement level 316 and a second enhancement level 328. The first enhancement level Petition 870250106234, dated 11 / 19 / 2025, page 22 / 48 15 / 31 316 can be considered a corrected stream. The second enhancement level, 328, can be considered an additional enhancement level that converts the corrected stream back into the original input video.
[0064] The additional enhancement level 328 is created by encoding an additional set of residues that are the difference 324 between an oversampled version of a decoded level 1 encoding and the input video 302.
[0065] In Figure 3, the quantized (or controlled) set of residues is inversely quantized 342 and inversely transformed 344 before an unlocking filter (not shown in the Figures) is optionally applied to generate a first decoded set of residues (i.e., an inverse quantization operation 342 is applied to the first quantized set of residues to generate a first dequantized set of residues; an inverse transformation operation 344 is applied to the first dequantized set of residues to generate a first detransformed set of residues; and an unlocking filter operation is optionally applied to the first detransformed set of residues to generate a first decoded set of residues). The unlocking filter step is optional depending on the transformation 336 applied and comprises applying a weighted mask to each block of the first detransformed set of residues 344.
[0066] The decoded base is combined 320 with the first decoded set of residues (i.e., a 320 addition operation is performed on the decoded base and the first decoded set of residues). As illustrated in Figures 3A and 3B, this combination is then subjected to 322 sampling.
[0067] A difference operation 324 is applied to generate an additional set of residues by comparison with the input video 302. The additional set of residues is then encoded as Level 2 encoding 328. As with Level 1 encoding 316, the encoding applied to Level 2 residues may comprise several steps. Figure 3A illustrates the steps as Petition 870250106234, dated 11 / 19 / 2025, page 23 / 48 16 / 31 time forecast, transformation 348, quantization 350 and entropy coding 352.
[0068] Although not shown in the Figures, in some examples, the coding process identifies whether the waste sorting mode is selected. If the waste sorting mode is selected, the waste sorting step can be performed (i.e., a waste sorting operation can be performed on the additional waste set to generate an additional sorted waste set). The highest sorted waste set can be filtered so that not all waste is coded in Level 2 328 coding.
[0069] The additional waste set or the additional sorted waste set is subsequently transformed 348 (i.e., a transformation operation 348 is performed on the additional sorted waste set to generate an additional transformed waste set). As illustrated, the transformation operation 348 may use a predicted coefficient or a predicted average derived before topsampling 322.
[0070] Figure 3A shows a variation of the second encoder example 300 where time prediction is performed as part of the level 2 encoding process. Time prediction is performed using the time selection component 334 and the level 2 time buffer 332. The time selection component 334 can determine a time processing mode, as described in more detail below, and control the use of the level 2 time buffer 332 accordingly. For example, if no time processing is performed, the time selection component 334 can indicate that the contents of the level 2 time buffer 332 should be set to 0. Figure 3B shows a variation of the second encoder example 380 where time prediction is performed as part of both the level 1 and level 2 encoding processes. In Figure 3B, a level 1 time buffer 330 is provided in addition to the level 2 time buffer 332.Although not shown, other variations in which temporal processing is performed at level 1 but not at level 2 are also possible.
[0071] When the weather forecast is selected, the Petition 870250106234, dated 11 / 19 / 2025, page 24 / 48 17 / 31 The second example of encoder 300, 380 from Figures 3A or 3B can further modify the coefficients (i.e., the output of transformed residues by a transformation component) by subtracting a corresponding set of coefficients derived from an appropriate time buffer. The corresponding set of coefficients may comprise a set of coefficients for the same spatial area (e.g., the same encoding unit located in a frame) that are derived from a previous frame (e.g., coefficients for the same area in a previous frame). These coefficients may be derived or otherwise obtained from a time buffer. The coefficients obtained from a time buffer may be referred to here as time coefficients. The subtraction may be applied by a subtraction component, such as the third subtraction components 354 and 356 (for levels 2 and 1 respectively).This timing prediction step will be described in more detail in later examples. In summary, when timing prediction is applied, the encoded coefficients correspond to a difference between one frame and another in the video signal. The other frame can be a previous or subsequent frame (or block within the frame). Timing prediction can be selectively applied to groups of encoding units (here called “tiles”) based on control information, and the application of timing prediction in a decoder can be done by sending additional control information along with the encodings (for example, in the headers).
[0072] As shown in Figures 3A and 3B, when the weather forecast is active, each transformed coefficient can be: Δ = Factual - ^buffer
[0073] where the temporal buffer can store data associated with a previous frame. Temporal prediction can be performed for one color plane or for multiple color planes. In general, subtraction can be applied as an element-by-element subtraction for a video “frame” where the frame elements represent transformed coefficients, where the transformation is applied with respect to a given n by n encoding unit size. Petition 870250106234, dated 11 / 19 / 2025, page 25 / 48 18 / 31 (e.g., 2x2 or 4x4). The difference resulting from the time forecast (e.g., the delta above) can be stored in the buffer for use in a subsequent frame. Therefore, in fact, the residual resulting from the time forecast is a residual coefficient with respect to the buffer. Although Figures 3A and 3B show the time forecast being performed after the transformation operation, it can also be performed after the quantization operation. This can avoid the need to apply the level 2 inverse quantization component 358 and / or the level 1 inverse quantization component 360.Thus, as illustrated in Figures 3A and 3B and described above, the output of the second example encoders 300, 380 after the execution of an encoding process is a base encoding 310 and one or more enhancement encodings which preferably comprise a level 1 encoding 316 for a first level of enhancement and a level 2 encoding 328 for an additional or second level of enhancement.
[0074] Figures 4A and 4B illustrate the respective variations of a second 400, 480 decoder example. The variations of the second 400, 480 decoder example can be implemented respectively to correspond to the first 200 decoder example in Figure 2. As is clearly identifiable, the steps and components of the decoding are expanded in more detail to provide an example of how the decoding can be performed. As in Figures 3A and 3B, Figure 4A illustrates a variation where timing prediction is used only for the second level (i.e., level 2) and Figure 4B illustrates a variation where timing prediction is used at both levels (i.e., levels 1 and 2). As before, other variations are provided (e.g., level 1 but not level 2), where the shape of the configuration can be controlled using signaling information.
[0075] As shown in the example in Figure 4B, in the decoding process, the 480 decoder can analyze the 436 headers (e.g., containing global configuration data, image configuration data, and other data blocks) and configure the decoder based on these 436 headers. To recreate the input video, the 400, 480 decoder can decode each Petition 870250106234, dated 11 / 19 / 2025, page 26 / 48 19 / 31 one of the base encodings 410, the first enhancement encoding 416 and the additional enhancement encoding 428. The frames of the video signal can be synchronized and then combined to derive the decoded video 448.
[0076] In each decoding process, enhancement encodings can go through the entropy decoding steps 450, 452, inverse quantization 454, 456 and inverse transformation 458, 460 to recreate a set of residues.
[0077] The decoding processes of Figures 4A and 4B comprise the recovery of a matrix of entropy-decoded quantized coefficients, representing a first level of enhancement, and the output of an L-1 residual matrix. The entropy-decoded quantized coefficients in this case are obtained by applying the entropy decoding operation 450 to the L-1 encoding 416. The decoding processes of Figures 4A and 4B also include the recovery of a matrix of output samples from a base 408 decoder.The decoding processes in Figures 4A and 4B further comprise the application of a 454 dequantization process to the set of entropy-decoded quantized coefficients to derive a set of dequantized coefficients, the application of a 458 transformation process to the set of dequantized coefficients, and optionally, the application of a filter process (not shown in Figures 4A and 4B) to generate the L-1 residual set representing a first level of enhancement, which can be called the preliminary residual set. In this case, the 454 dequantization process is applied to the entropy-decoded quantized coefficients for the respective blocks of a level 1 encoding frame 416, and the 458 transformation process (which can be called the inverse transformation operation) is applied to the output of the 454 dequantization process for the respective blocks of the frame.The decoding processes for Figures 4A and 4B also include recreating an image by combining 462 the L-1 residue matrix with the output sample matrix of the base 408 decoder. The decoding processes for Figures 4A and 4B include applying a transformation process 458 to a set. Petition 870250106234, dated 11 / 19 / 2025, page 27 / 48 20 / 31 of predetermined transformation processes according to a signaled parameter. For example, transformation process 458 can be applied to a 2x2 encoding unit or a 4x4 encoding unit. An encoding unit can be referred to here as a block of elements in a matrix, in this case, the L-1 residue matrix.
[0078] The decoding processes in Figures 4A and 4B include the recovery of an entropy-decoded quantized coefficient matrix representing an additional level of enhancement and the output of a residual matrix. In the decoding processes shown in Figures 4A and 4B, the additional level of enhancement is a second level of enhancement and the output residual matrix is an L-2 residual matrix. The method in Figures 4A and 4B also includes the recovery of the L-1 residual matrix of the first level of enhancement corresponding to the entropy-decoded quantized coefficient matrix representing an additional level of enhancement. The method in Figures 4A and 4B also includes the application of a higher sampling process 464 to the residual matrix of the first level of enhancement.In Figures 4A and 4B, the 464 upsampling process is applied to the combination of the L-1 residual matrix from the first enhancement level and the corresponding matrix of output samples from the base 408 decoder.
[0079] In Figures 4A and 4B, the 464 upsampling process is a modified upsampling process in which a modifier is added to a residual. The step of adding a modifier can be performed as part of the 460 transformation process. Alternatively, since the 460 transformation process involves a linear transformation, the step of adding a modifier can be performed as part of the modified 464 upsampling process, as shown in Figures 4A and 4B. The step of adding a modifier therefore results in a modification of a residual. The modification can be performed based on a location of the residual in a frame. The modification can be a predetermined value.
[0080] In Figure 4A, weather forecasting is applied Petition 870250106234, dated 11 / 19 / 2025, page 28 / 48 21 / 31 during level 2 decoding. In the example in Figure 4A, the time forecast is controlled by a time forecast component 466. In this variation, the control information for the time forecast is extracted from the level 2 encoding 428 to the time forecast component 466, as indicated by the arrow. In other implementations, such as those shown in Figure 4B, the control information for the time forecast may be sent separately from the level 2 encoding 428, for example, in the headers 436. The time forecast component 466 controls the use of the level 2 time buffer 432, for example, it can determine a time mode and control the time update, as described in later examples. The contents of the time buffer 432 can be updated based on the data from a previous residual frame. When the time buffer 432 is applied, the contents of the buffer are added 468 to the second set of residuals.In Figure 4A, the contents of the temporal buffer 432 are added 468 to the output of a level 2 decoding component 446 (which, in Figure 4A, implements entropy decoding 452, inverse quantization 456, and inverse transformation 460). In other examples, the contents of the temporal buffer may represent any set of intermediate decoding data, and as such, the addition 468 can be moved appropriately to apply the contents of the temporal buffer at an appropriate stage (for example, if the temporal buffer is applied at the dequantized coefficient stage, the addition 468 can be located before the inverse transformation 460). The second temporally corrected set of residuals is then combined 470 with the upsampling output 464 to generate the decoded video 448. The decoded video 448 is at a level 2 spatial resolution, which may be higher than a level 1 spatial resolution.The second set of residues applies a correction to the reconstructed video with oversampling (visible), where the correction adds fine details again and improves the sharpness of lines and features.
[0081] Transformation processes 458 and 460 can be selected from a set of predetermined transformation processes according to a signaled parameter. For example, transformation process 460 Petition 870250106234, dated 11 / 19 / 2025, page 29 / 48 22 / 31 can be applied to a 2x2 block of elements in the L-2 residue matrix or to a 4x4 block of elements in the L-2 residue matrix.
[0082] Figure 4B shows a variation of the second decoder example 480. In this case, the timing prediction control data is received by a timing prediction component 466 from the headers 436. The timing prediction component 466 controls both level 1 and level 2 timing prediction, but in other examples, separate control components can be provided for both levels if desired. Figure 4B shows how the second reconstructed set of residuals that are added 468 to the output of the level 2 decoding component 446 can be fed back 469 to be stored in the level 2 timing buffer 432 for a next frame (the feedback is omitted in Figure 4A for clarity). A level 1 timing buffer 430 operating similarly to the level 2 timing buffer 432 described above is also shown, and the feedback loop for the buffer is shown in this Figure.The contents of the level 1 temporal buffer 430 are added to the level 1 residual processing pipeline via a sum 472, and the sum of the level 1 residuals is fed back 473 to the level 1 temporal buffer 430. Again, the position of this sum 472 can vary along the level 1 residual processing pipeline, depending on where the temporal forecast is applied (for example, if applied in the transformed coefficient space, it may be located before the level 1 inverse transformation component 458).
[0083] Figure 4B shows two ways in which temporal control information can be signaled to the decoder. The first way is via headers 436, as described above. A second way, which can be used as an alternative or additional signaling pathway, is via data encoded in the residuals themselves. Figure 4B shows a case where data 474 can be encoded in a transformed HH coefficient and therefore can be extracted after entropy decoding 452. This data 474 can be extracted from the level 2 residual processing pipeline and passed to the temporal forecasting component 466. Petition 870250106234, dated 11 / 19 / 2025, page 30 / 48 23 / 31
[0084] Each enhancement encoding or both enhancement encodings can be encapsulated in one or more enhancement bitstreams using a set of Network Abstraction Layer Units (NALUs). NALUs are intended to encapsulate the enhancement bitstream to apply the enhancement to the correct reconstructed base frame. The NALU may, for example, contain a reference index to the NALU containing the bitstream of the reconstructed base frame to which the enhancement should be applied. In this way, the enhancement can be synchronized with the base stream and the frames of each bitstream combined to produce the decoded output video (i.e., the residues of each frame of the enhancement level are combined with the frame of the decoded base stream). A group of frames can represent multiple NALUs.
[0085] Each frame can be composed of three different planes representing a different color component, for example, each component of a three-channel YUV video can have a different plane. Each plane can then have residual data related to a given enhancement level, for example, a Y plane can have a level 1 residual data set and a level 2 residual data set. In certain cases, for example, for monochrome signals, there may be only one plane; in this case, the terms frame and plane can be used interchangeably. Level 1 residual data and level 2 residual data can be divided as follows. The residual data is divided into blocks whose size depends on the size of the transformation used. The blocks are, for example, a block of 2x2 elements if a 2x2 directional decomposition transformation is used, or a block of 4x4 elements if a 4x4 directional decomposition transformation is used.A tile is a group of blocks that cover a region of a frame (for example, an M by N region, which can be a square region). A tile is, for example, a 32x32 tile. In this way, each frame can be divided into a plurality of tiles, and each tile within that plurality of tiles can be divided into a plurality of blocks. In the case of color video, each frame can be divided into multiple planes, where each... Petition 870250106234, dated 11 / 19 / 2025, p. 31 / 48 The 24 / 31 plan is divided into several tiles, and each tile in the tile set is further divided into several blocks.
[0086] With regard to Figures 4A and 4B, the following example refers to a time prediction process applied during level 2 decoding. However, it should be considered that the following time prediction process can be applied additionally or alternatively during level 1 decoding.
[0087] In this example, the 400, 480 decoder is configured to receive a temporal_enabled parameter that specifies whether temporal prediction should be used when decoding an image. The temporal_enabled parameter can be referred to in this document as a first parameter with a first value indicating that temporal processing is enabled. The value of the temporal_enabled parameter can have a bit length of one bit. In this example, a value of 1 specifies that temporal prediction will be used when decoding an image, and a value of 0 specifies that temporal prediction will not be used when decoding an image. The temporal_enabled parameter can be received once for an image group associated with the encoded streams discussed above, the image group being a collection of successive images in an encoded video stream.The temporal_enabled parameter can be included in the time control information signaled to the decoder, for example, via the 436 headers, as described above.
[0088] Inputs for a weather forecasting process may include:
[0089] • an nTbS parameter that specifies the size of the current transformation block. For example, nTbS equals 2 when a 2x2 directional decomposition transformation is to be used, and nTbS equals 4 when a 4x4 directional decomposition transformation process is to be used.
[0090] • a temporal_tile_intra_signalling_enabled parameter that specifies whether time forecasting should be Petition 870250106234, dated 11 / 19 / 2025, pp. 32 / 48 25 / 31 used when decoding a specific tile.
[0091] In the decoding process described in this document, the generation of the decoded video can be performed in blocks. In this way, the generation of a block of elements in a frame of the decoded video can be performed without the use of another block of elements in the same frame of the decoded video that was generated previously. For this reason, the timing prediction process can be executed in parallel for all blocks of elements in a frame, instead of sequentially executing the timing prediction process for each block of elements in the frame.
[0092] The description above Figures 1 to 4B explains the timing forecast and the contents of the timing buffer(s) according to examples. The following methods explain how the timing buffer can be “updated”, contrasting the approach in a background example, shown in Figure 5, with a new approach, as shown in Figure 6.
[0093] With regard to Figure 5, this method comprises two independent procedures: an update procedure (steps S510 and S520) and a frame decoding procedure (steps S610 to S640).
[0094] In step S510, the 400, 480 decoder obtains a temporal_refresh parameter that specifies whether the 432 temporal buffer should be updated for the frame. If a frame includes multiple planes, the update may be applied to all planes in the frame (i.e., the frame that includes the planes) or to specific indicated planes. The value of the temporal_refresh parameter can have a bit length of one bit. In this example, a value of 1 specifies that the 432 temporal buffer should be updated for the frame, and a value of 0 indicates that the 432 temporal buffer should not be updated for the frame. The temporal_refresh parameter can be received once for each image in the encoded video stream. The temporal_refresh parameter can be included in the temporal control information signaled to the decoder, for example, via the 436 headers, as described above.If the decoder has multiple 430, 432 time buffers, the `temporal_refresh` parameter can be used to update all buffers. Petition 870250106234, dated 11 / 19 / 2025, pp. 33 / 48 26 / 31 update one or more individual time buffers.
[0095] In step S520, decoder 400, 480 performs a temporal update, writing zeros to temporal buffer 432. The update can be applied to the entire frame area or to one or more specific blocks, tiles, and planes and applied to the temporal buffer(s) of one or more enhancement layers 430, 432, according to the temporal_refresh parameter.
[0096] In the S610 step, the 400, 480 decoder obtains an encoded frame from the video signal (e.g., bytestream).
[0097] In step S620, to decode the encoded frame, decoder 400, 480 reads the contents of the temporal buffer(s) 430, 432. The contents of the temporal buffer are all zero when step S620 is executed, and therefore the temporal buffer has no effect on the decoding. However, step S620 occurs because the decoding method (steps S610 to S640) is executed independently of the update (steps S510 to S520).
[0098] Next, in step S630, decoder 400, 480 decodes the encoded frame according to one of the methods described above. This method makes no assumptions about the values stored in the temporal buffer(s) and therefore steps 468, 472 of adding the contents of the temporal buffer(s) are executed regardless of the fact that these contents are currently zero values.
[0099] Finally, in step S640, decoder 400, 480 updates the time buffer(s) with the most recent residual values. This can be implemented as feedback 473 and 469 for the L-1 and L-2 time buffers 430, 432 in Figure 4B.
[0100] As mentioned above, the background example method in Figure 5 contains several inefficiencies as a result of performing the update operation independently of decoding the next frame. The time buffer is read in step S620, despite containing zero values. Furthermore, step S630 involves performing arithmetic operations to combine the enhancement layer of the decoded frame with the time buffer, despite the Petition 870250106234, dated 11 / 19 / 2025, pp. 34 / 48 27 / 31 fact that the time buffer contains zero values at the moment.
[0101] In view of these inefficiencies observed by the inventors, a new method is proposed, as illustrated in Figure 6.
[0102] With regard to Figure 6, step S710 may be the same as step S510 described earlier. The 400, 480 decoder obtains a temporal_refresh parameter that specifies whether the 432 temporal buffer should be updated for the frame. If a frame includes multiple planes, the update may be applied to all planes in the frame (i.e., the frame that includes the planes) or to specific indicated planes. The value of the temporal_refresh parameter can have a bit length of one bit. In this example, a value of 1 specifies that the 432 temporal buffer should be updated for the frame, and a value of 0 indicates that the 432 temporal buffer should not be updated for the frame. The temporal_refresh parameter may be received once for each image in the encoded video stream. The temporal_refresh parameter may be included in the temporal control information signaled to the decoder, for example, via the 436 headers, as described above.If the decoder has multiple 430, 432 time buffers, the temporal_refresh parameter can be used to refresh all buffers or to refresh one or more individual time buffers.
[0103] However, after step S710, the 400, 480 decoder does not update the time buffer until the next frame decoding. In other words, the 400, 480 decoder skips step S520 of writing something to the time buffer at that point. The 400, 480 decoder can optionally store an indication in memory that an update is pending. Alternatively, the decoder can have a live execution state and can add an update indicator to its live execution state.
[0104] In step S720, the 400, 480 decoder obtains an encoded frame from the video signal (e.g., bytestream). This is similar to step S610 from the previous example, except that the 400, 480 decoder remembers that a temporal update is pending (e.g., referring to an indication stored in memory or included in its program state). Petition 870250106234, dated 11 / 19 / 2025, pp. 35 / 48 28 / 31
[0105] For example, the 400, 480 decoder can wait to refresh the temporal buffer by storing a “dirty” flag associated with the temporal buffer in working memory. The “dirty” flag can be one bit long and, in fact, can be a copy of the temporal_refresh parameter. When the 400, 480 decoder decodes the next frame in steps S710 to S740, the 400, 480 decoder detects the “dirty” flag and responds by executing a combined decoding / refresh procedure as follows.
[0106] In step S730, decoder 400, 480 reconstructs the decoded frame based on the base encoding and enhancement encoding of the encoded frame, without reading the temporal buffer that needs to be updated. For example, referring to Figure 4B, if the temporal_refresh parameter indicates updating the L-1 temporal buffer 430, decoder 400, 480 will skip the steps of reading the L-1 temporal buffer and add 472 the contents of the level 1 temporal buffer 430 to the level 1 residual processing pipeline. Similarly, referring to Figure 4B, if the temporal_refresh parameter indicates updating the L-2 temporal buffer 432, decoder 400, 480 will skip the steps of reading the L-2 temporal buffer and add 468 the contents of the level 2 temporal buffer 432 to the level 2 residual processing pipeline.
[0107] Finally, in step S740, decoder 400, 480 writes the most recent residual values to the time buffer(s). This can be implemented as feedback 473 and 469 to the L-1 and L-2 time buffers 430, 432 in Figure 4B.
[0108] However, unlike the “update” in step S640, step S740 is a temporal buffer update. More specifically, the updated area can be the entire frame area or it can be one or more specific blocks, tiles, and planes, and the update can be applied to the temporal buffer(s) of one or more enhancement layers 430, 432, according to the temporal_refresh parameter. Therefore, in a case where the encoded frame enhancement encoding includes residual values for only one Petition 870250106234, dated 11 / 19 / 2025, pp. 36 / 48 29 / 31 subset s1 of the tiles, blocks, or planes of the frame, the 400, 480 decoder still writes data to the entire updated area of the temporal buffer. This update can be performed as a single write operation for the entire updated area. This update may include writing zero values to the temporal buffer for any tiles, blocks, or planes of the frame for which no residual value was explicitly included in the enhancement encoding. In other words, the update may include writing zero values to the temporal buffer for any tiles, blocks, or planes that are not part of subset s1.
[0109] In the background example of Figure 5, the separate update (step S520), read (step S620), and update (S640) steps involve copying data to / from the time buffer. Instead, by performing a single write step (S740), a significant amount of memory bandwidth is freed up for other decoder operations.
[0110] Furthermore, in the context of video streaming, the time between steps S710 and S720 (or between steps S510 and S610) is likely to be short, and there may be no waiting at all. As a result, the processing resources saved by avoiding a separate update procedure (steps S520 etc.) and instead performing a combined decoding / update procedure have a considerable effect on the decoder's maximum frame rate.
[0111] Step 740 can alternatively be executed before step S730, because step S730 does not need to include reading the temporal buffer.
[0112] Although one or more embodiments have been described in relation to LCEVC, aspects of the invention can also be implemented in other hierarchical coding schemes. In these embodiments, the “base encoder” may correspond to an encoder configured to encode a low layer (corresponding to a low quality) of a hierarchical coding scheme. In such embodiments, the “enhancement encoder” may correspond to an encoder configured to encode a Petition 870250106234, dated 11 / 19 / 2025, pp. 37 / 48 30 / 31 high layer (corresponding to a high quality, i.e., a higher quality than the low layer) of the hierarchical coding scheme. For example, the base encoder might correspond to an encoder configured to encode a lower layer of the hierarchical coding scheme, and the enhancement encoder might correspond to an encoder configured to encode an enhancement layer (e.g., the first) of the hierarchical coding scheme. More generally, the base encoder might correspond to an encoder configured to encode the nth layer of a hierarchical coding scheme, and the enhancement encoder might correspond to an encoder configured to encode the (n+1)th layer of the hierarchical coding scheme.
[0113] A base encoder can be a multi-layer encoder. In particular, a base encoder can include a base encoder and one or more enhancement encoders.
[0114] An enhancement encoder can include multiple encoders. In particular, an enhancement encoder can include multiple enhancement encoders.
[0115] A base encoding can be the output of one or more encoding layers. For example, a base encoding can be a first encoding layer combined with one or more additional encoding layers (e.g., enhancement).
[0116] An enhancement code may include one or more enhancement layers.
[0117] For example, a base encoding might be a base layer (i.e., output by a single-layer codec such as HEVC, VVC, etc.) combined with a first layer of LCEVC residue, while the enhancement encoding might be a second layer of LCEVC residue. In another example, a base encoding might be a lower layer encoded according to the SMPTE VC-6 standard combined with one or more VC-6 enhancement layers, while the enhancement encoding might Petition 870250106234, dated 11 / 19 / 2025, pp. 38 / 48 31 / 31 being one or more layers of enhancement of the "upper" layer of the VC-6 standard.
[0118] The above embodiments should be understood as illustrative examples. Other embodiments are provided. It should be understood that any feature described in relation to any embodiment may be used alone or in combination with other described features, and may also be used in combination with one or more features of any other embodiment, or any combination of any other embodiment. Furthermore, equivalents and modifications not described above may also be employed within the scope of the appended claims. Petition 870250106234, dated 11 / 19 / 2025, pp. 39 / 48
Claims
1 / 5 CLAIMS 1. A method for decoding a video signal, characterized in that it comprises: obtaining an encoded frame from the video signal, the encoded frame comprising a base coding and an enhancement coding; determining whether a time buffer should be updated; and if it is determined that the time buffer should be updated: reconstructing a decoded frame based on the base coding and enhancement coding of the encoded frame, without reading the time buffer; and writing a set of new time predictions to the time buffer, based on the enhancement coding of the encoded frame.
2. A method according to claim 1, characterized in that recording the set of new temporal predictions in the temporal buffer, based on the enhancement coding of the encoded frame, comprises: obtaining a set of residues from the enhancement coding of the encoded frame, wherein the residues are differences between: a decoded base frame obtained by decoding the base coding; and an input frame, wherein the base coding is generated by encoding the input frame; and recording the set of residues in the temporal buffer.
3. Method, according to claim 2, characterized in that each frame of the video signal has a frame area divided into a plurality of clippings and / or a plurality of blocks, and the set of residuals includes residual values for a subset of the plurality of clippings or a subset of the plurality of blocks, wherein recording the set of new temporal predictions in the temporal buffer, based on the enhancement encoding of the encoded frame, further comprises: recording a zero value in the temporal buffer for a clipping of the frame area that is not in the subset of the plurality of clippings and / or for a block of the frame area that is not in the subset of the plurality of blocks.
4. Method, according to claim 3, characterized in that it comprises: constructing a complete set of new temporal predictions for all clippings and / or all blocks of the frame area, consisting of the set of residuals obtained from the encoded frame enhancement coding and zero values for any clippings and / or blocks for which no residual value is included in the encoded frame enhancement coding; and recording the complete set of new temporal predictions in the temporal buffer.
5. A method, according to any of the preceding claims, characterized in that it further comprises: if it is determined that the temporal buffer should not be updated: reading the temporal buffer to obtain a set of previous temporal predictions; and reconstructing the decoded frame based on the base encoding and enhancement encoding of the encoded frame and based on the set of previous temporal predictions.
6. Method, according to claim 5, characterized in that the previous time predictions comprise a set of residuals for a previous frame in the video signal preceding the encoded frame.
7. A method, according to any of the preceding claims, characterized in that it further comprises: receiving a refresh parameter in the video signal, wherein the determination of the temporal buffer refresh is based on the refresh parameter.
8. Method according to claim 7, characterized in that the update parameter is part of the coded frame. Petition 870250085514, dated 09 / 22 / 2025, pp. 107 / 112 3 / 5 9. Method according to claim 7, characterized in that the update parameter is part of a video signal control component, separate from the encoded frame.
10. Decoder for decoding a video signal, the decoder is characterized by the fact that it comprises: a memory comprising a time buffer configured to store a set of time predictions; and a processor configured to: obtain an encoded frame from the video signal, the encoded frame comprising a base encoding and an enhancement encoding; determine whether to update a time buffer; and if it is determined that the time buffer should be updated: reconstruct a decoded frame based on the base encoding and enhancement encoding of the encoded frame, without reading the time buffer; and write a set of new time predictions to the time buffer, based on the enhancement encoding of the encoded frame.
11. Decoder, according to claim 10, characterized in that the processor is configured to write the set of new timing predictions to the time buffer, based on the enhancement coding of the encoded frame, in order to: obtain a set of residues from the enhancement coding of the encoded frame, wherein the residues are differences between a base decoded frame obtained by decoding the base coding and an input frame, wherein the base coding is generated by encoding the input frame; and write the set of residues to the time buffer.
12. Decoder, according to claim 11, characterized in that each frame of the video signal has a frame area Petition 870250085514, dated 09 / 22 / 2025, pp. 108 / 112 4 / 5 divided into a plurality of clippings and / or a plurality of blocks, and the residual set includes residual values for a subset of the plurality of clippings or a subset of the plurality of blocks, wherein the processor is configured to write the set of new temporal predictions to the temporal buffer, based on the enhancement encoding of the encoded frame, by: writing a zero value to the temporal buffer for a clipping of the frame area that is not in the subset of the plurality of clippings and / or a block of the frame area that is not in the subset of the plurality of clippings.
13. Decoder, according to claim 12, characterized in that the processor is configured to write the set of new time predictions to the time buffer, based on the encoded frame enhancement coding, to: construct a complete set of new time predictions for all clips and / or all blocks in the frame area, consisting of the set of residuals obtained from the encoded frame enhancement coding and zero values for any clips and / or blocks for which no residual value is included in the encoded frame enhancement coding; and write the complete set of new time predictions to the time buffer.
14. Decoder, according to any one of claims 10 to 13, characterized in that the processor is further configured to: determine that the temporal buffer should not be updated; read the temporal buffer to obtain a set of previous temporal predictions; and reconstruct the decoded frame based on the base encoding and enhancement encoding of the encoded frame and based on the set of previous temporal predictions.
15. Decoder, according to claim 14, characterized in that the previous time predictions comprise a set of residues for a previous frame in the video signal preceding the encoded frame.
16. Decoder, according to any one of claims 10 to 15, characterized in that the processor is configured to: receive a refresh parameter in the video signal; and determine whether the time buffer should be refreshed based on the refresh parameter.
17. Decoder, according to claim 16, characterized in that the update parameter is part of the encoded frame.
18. Decoder, according to claim 16, characterized in that the update parameter is part of a video signal control component, separate from the encoded frame. Petition 870250085514, dated 09 / 22 / 2025, pp. 110 / 112