Image and video encoding and decoding.

BR112025021033A2Pending Publication Date: 2026-08-25
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112025021033
Authority / Receiving Office
BR · BR
Patent Type
Applications
Publication Date
2026-08-25

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

1 / 73 “IMAGE AND VIDEO ENCODING AND DECODING” FIELD OF THE INVENTION

[0001] The present invention relates to the encoding and decoding of image and video partitioning data. FUNDAMENTALS OF THE INVENTION

[0002] The Joint Video Expert Team (JVET), a collaborative team formed by MPEG and VCEG from ITU-T Study Group 16, has released a new video coding standard called Versatile Video Coding (VVC). The goal of VVC is to provide significant improvements in compression performance compared to the existing HEVC standard (i.e., typically twice the previous performance). Key target applications and services include, among others, 360-degree and high dynamic range (HDR) video. Particular effectiveness has been demonstrated in ultra-high-definition (UHD) video test material. Thus, compression efficiency gains well beyond the 50% target for the final standard can be expected.

[0003] Since the end of the VVC v1 standardization, JVET has begun an exploration phase, establishing an exploration software (ECM). It brings together additional tools and enhancements to existing tools based on the VVC standard to achieve better coding efficiency. SUMMARY OF THE INVENTION

[0004] According to a first aspect of the invention, a method is provided for decoding or encoding image data from or into a bitstream, the bitstream including data indicating a partitioning of the image data into a plurality of blocks according to an encoding tree, wherein the blocks in the encoding tree can be divided according to a plurality of splits, including a quadtree split. The method comprises determining a value indicating a quadtree split depth of a current block and restricting the possible splits for the current block based on a (current) quadtree depth and at least one other value (e.g., where the other value represents a limit of some kind). Petition 870250088574, dated 09 / 30 / 2025, page 12 / 119 2 / 73

[0005] In general terms, the advantages include a reduction in encoder complexity thanks to the reduction in the number of partitions evaluated on the encoder side (constrained based on the quadtree split depth) and an improvement in encoding efficiency on the decoder thanks to the reduction in signaling required for partitioning, making the decoding process more efficient.

[0006] In a second aspect according to the present invention, a method is provided for encoding or decoding image data into or from a bitstream, the bitstream including data indicating a partitioning of the image data into a plurality of blocks according to an encoding tree, wherein the blocks in the encoding tree can be divided according to a plurality of divisions, including a quadtree division, the method comprising: determining a value indicating a quadtree division depth of a current block; and allowing only one quadtree division of the current block from among the plurality of divisions if the value indicating the quadtree division depth does not meet one or more conditions. For example, if the value indicating the quadtree division depth does not exceed a limit (value).The advantage is a reduction in encoding time, as it is not necessary to evaluate partitioning for QT depths lower than that value. The second advantage is improved encoding efficiency, as it is not necessary to transmit any relative syntax elements before QT depths lower than the value. But, of course, there is a trade-off between the bitrate saved and the quality. Despite this, surprisingly significant gains are achieved with the proposed modification.

[0007] The plurality of divisions may also include a ternary tree division, a binary tree division, and no division.

[0008] In a third aspect according to the present invention, a method is provided for encoding or decoding image data into or from a bitstream, the bitstream including data indicating a partitioning of the image data into a plurality of blocks according to a tree of Petition 870250088574, dated 09 / 30 / 2025, p. 13 / 119 3 / 73 encoding, in which the blocks in the encoding tree can be split according to a plurality of splits, including a binary tree split, a ternary tree split, a quadtree split, and no split, the method comprising: determining a value indicating the quadtree split depth of a current block; and not allowing binary tree split based on one or more criteria, the criterion including that the value indicating the quadtree split depth does not meet one or more conditions. For example, the value indicating the quadtree split depth does not exceed a limit (value). The advantage is a reduction in encoding time, as only the No Split mode needs to be evaluated and the QT split node when the QT depth is less than that value.A second advantage is an improvement in encoding efficiency, since only the No Division flags before the QT depth smaller than that value need to be transmitted.

[0009] The method may also include allowing quadtree splitting, ternary tree splitting, and no splitting when the value indicating the quadtree split depth does not exceed the limit.

[0010] Alternatively, the method may also include not allowing ternary tree splitting, but allowing quadtree splitting and not splitting when the value indicating the quadtree split depth does not exceed the limit.

[0011] Optionally, only quadtree splitting and no splitting are allowed, and not binary tree splitting or quadtree splitting based on additional criteria, while still allowing TT splitting if the condition is met.

[0012] In an additional optional feature of the third aspect, one or more criteria may include that a ternary tree split is not performed at a level above the current block level, where, if the ternary tree split has not been performed, do not allow the binary tree split, and if the ternary tree split has been performed, allow the binary tree split for the current block.

[0013] The method may include executing a first method according to any embodiment of the second aspect; and executing a second method of Petition 870250088574, dated 09 / 30 / 2025, page 14 / 119 4 / 73 according to any modality of the third aspect, wherein the limit to allow only a quadtree split is a first limit and the limit to not allow a binary tree split is a second limit, wherein the first and second limits are different.

[0014] Optionally, the second method includes the additional option feature of the third aspect and further comprises executing a third method in accordance with the third aspect, wherein the limit for the third method is a third limit different from the first and second limits.

[0015] Optionally, in any of the above aspects and modalities, the value indicating the quadtree split depth may be a quadtree split depth value that crosses said limit(s) if it is less than a predetermined quadtree split depth value.

[0016] The value indicating the quadtree split depth may, alternatively or additionally, be related to or be a block size of the current block and which crosses said limit(s) if it is greater than a predetermined block size.

[0017] The limit value(s) may be based on one or more parameters associated with an area (temporal) in one or more temporal frames, a temporal frame being a frame different from the frame containing the current block. A temporal frame may be a reference frame.

[0018] The (temporal) area may be an area that is colocated with the current block.

[0019] The (temporal) area may comprise a plurality of blocks in different positions.

[0020] The (temporal) area can be an area larger than the current block.

[0021] The (temporal) area may be larger than the current block in a coding tree unit, CTU.

[0022] A central position of the current block can be used to determine the (temporal) area in the time frame or in a time frame. Petition 870250088574, dated 09 / 30 / 2025, page 15 / 119 5 / 73

[0023] The (temporal) area can encompass the entire area of ​​the temporal frame or of a temporal frame.

[0024] A temporal frame or a temporal frame may be a frame with the same temporal ID as the current frame that includes the current block. The frame with the same temporal ID may be the nearest frame with the same temporal ID.

[0025] A timeframe or a timeframe can be a frame with the same quantization parameter as the current frame that includes the current block.

[0026] A time frame or a time frame can be a frame used for predicting the (temporal) motion vector.

[0027] The timeframe or a timeframe can be a frame that is the closest reference frame to the current frame that includes the current block.

[0028] The first area of ​​a first time frame and a second area of ​​a second time frame can be used to obtain the limit value.

[0029] The limit value(s) may be based on an average quadtree depth determined from the area (temporal).

[0030] The limit value(s) may be (based on) a minimum quadtree depth determined from the (temporal) area.

[0031] The limit value(s) can be determined based on a maximum multitree depth or an average multitree depth from the temporal area.

[0032] The threshold value can be determined using the maximum or average encoding tree depth from the time area.

[0033] The limit value(s) may be based on a value transmitted in a header associated with the encoding tree unit of the current block.

[0034] The value transmitted in the header of the current encoding tree unit is used to predict the threshold (or a threshold value) using one or more other values ​​transmitted in a header associated with another encoding tree unit.

[0035] The value transmitted in the header can be used to determine a plurality or each limit value.

[0036] Optionally, one of the criteria is not applied when a condition is Petition 870250088574, dated 09 / 30 / 2025, page 16 / 119 6 / 73 achieved. For example, when the maximum multitree depth of the timeframe is less than the current maximum multitree depth, the first criterion is not applied.

[0037] The following aspects of the invention may be used independently or in combination with any of the aspects of the invention and embodiments mentioned above.

[0038] The method may involve adding an integer offset to obtain the limit value or each limit value.

[0039] The limit value or a limit value can be adapted based on a slice type corresponding to the time area in the reference frame.

[0040] The threshold value or a threshold value can be adapted based on the maximum multitree depth and the maximum multitree depth of the temporal area frame.

[0041] The threshold value or a threshold value can be adapted based on the maximum multitree depth and the maximum multitree depth of the temporal area.

[0042] The threshold value or a threshold value can be adapted based on the minimum quadtree size of the temporal area frame or the minimum, maximum, or average quadtree size of the temporal area.

[0043] The threshold value or a threshold value can be adapted according to whether the current frame has a lower QP than the frame that includes the time area.

[0044] The limit value or a limit value can be adapted according to whether the current frame that includes the current block is a reference frame or not.

[0045] Optionally, the threshold value or a threshold value is based on block size statistics in the time area.

[0046] The parameters used to determine the threshold value or a threshold value can be obtained from blocks at the encoding tree unit level.

[0047] Optionally, one or more steps of the method are not executed if the image data is related to a low-latency configuration or if at least one flag is transmitted in at least one header. Petition 870250088574, dated 09 / 30 / 2025, page 17 / 119 7 / 73 related to the image.

[0048] Optionally, one or more steps of the method are not executed if the image data is related to the display content.

[0049] Optionally, not executing one or more steps or features does not mean enabling or disabling such features. Not executing can be for one or more blocks in the image or any other predetermined region or part of the image.

[0050] Optionally, one or more steps of the method are not executed when the image data is related to the display content and when one or more reference frames are an intra frame.

[0051] Image data can be considered related to the display content when one or more tools are enabled and / or flagged in the bitstream, where the one or more tools may include any of the following: Palette, Intra Block Copy (IBC), Transform Skip, and other tools associated with the display content.

[0052] One or more tools can be enabled according to data associated with a hierarchical level in the bitstream.

[0053] The data can be one or more of a sequence parameter set (SPS), image parameter set (PPS), image header, and slice header.

[0054] One or more tools can be enabled and / or flagged for at least one block in the image or in a reference frame (thus indicating the display content).

[0055] One or more tools can be enabled and / or flagged for at least one block in a temporal area of ​​the reference frame.

[0056] Optionally, one or more steps of the method are not executed when the number of blocks with one or more tools enabled and / or flagged is greater than a predetermined number.

[0057] Optionally, the one or more steps not performed are one or more Petition 870250088574, dated 09 / 30 / 2025, page 18 / 119 8 / 73 steps associated with at least one of the first method and the second method, as established in the aspects and modalities described above.

[0058] Not executing one or more steps of the second method comprises disabling at least one of the steps where one or more criteria include that a ternary tree split is not performed at a level above the current block level and, if the ternary tree split has not been performed, not allowing the binary tree split and, if the ternary tree split has been performed, allowing the binary tree split for the current block; and not executing one or more of the steps associated with ai) determining a value indicating a quadtree split depth of a current block;and not allow binary tree splitting based on one or more criteria, the criterion or criteria including that the value indicating the depth of the quadtree split does not exceed a limit, ii) allow quadtree splitting, ternary tree splitting and no splitting when the value indicating the depth of the quadtree split does not exceed the limit, and iii) not allow ternary tree splitting, but allow quadtree splitting and no splitting when the value indicating the depth of the quadtree split does not exceed the limit.

[0059] Optionally, the threshold value or a threshold value is based on block size statistics in the time area.

[0060] The parameters used to determine the threshold value or a threshold value can be obtained from blocks at the encoding tree unit level.

[0061] In a fourth aspect of the present invention, a method is provided for encoding or decoding image data into or from a bitstream, the bitstream including data indicating a partitioning of the image data into a plurality of blocks according to an encoding tree, wherein the blocks in the encoding tree can be divided according to a plurality of splits, including a ternary tree split and a binary tree split, the method comprising: determining whether a ternary tree split has been performed at a level above the current block level and, if the ternary tree split has been performed, allowing the binary tree split for the current block. Petition 870250088574, dated 09 / 30 / 2025, page 19 / 119 9 / 73 The advantage of this aspect is an overall improvement in coding efficiency with only a small increase in coding time.

[0062] The determination can be made by obtaining the multitree depth value at the level above the current block.

[0063] According to a fifth aspect of the present invention, a method is provided for encoding image data into a bitstream, the bitstream including data indicating a partitioning of the image data into a plurality of blocks according to an encoding tree, wherein the blocks in the encoding tree may be divided according to a plurality of divisions, including a binary tree division and a quadtree division, and no division, wherein the binary tree division may be a horizontal and vertical division, the method comprising: in a case where horizontal and vertical binary tree divisions are determined as not permitted to determine the division of a current block of the encoding tree, performing the control such that the quadtree division is tested for the current block before the binary tree division.

[0064] This provides an improvement in coding efficiency, with and without the proposed method, without increasing coding time. This is because the encoder optimizations are related to avoiding QT testing when BT is detected as better. Therefore, when TT is the only MTT split available, these optimizations are not efficient.

[0065] Optionally, quadtree splitting is tested before all splits, except for not splitting the current block.

[0066] According to a sixth aspect of the present invention, a bitstream is provided including data indicating the partitioning of the image data into a plurality of blocks according to an encoding tree, wherein the blocks in the encoding tree can be divided according to a plurality of splits, which can be a horizontal or vertical split, the method comprising: not including a split in a test list for a current block of the encoding tree when it is determined that the split is not available. Petition 870250088574, dated 09 / 30 / 2025, page 20 / 119 10 / 73 This provides an improvement in coding efficiency, with or without the proposed method of the aspects or modalities mentioned above, and without increasing coding time.

[0067] In a first aspect of high-level syntax according to the invention, a method is provided for decoding or encoding video or image data from or into a bitstream, wherein a (high-level) syntax element (flag) for the entire sequence is transmitted to enable or disable any of the aspects or modes mentioned above. For example, a flag may be transmitted within the SPS, which controls whether the method is enabled or not.

[0068] In a second high-level syntax aspect according to the invention, a method is provided for decoding or encoding video or image data from or into a bitstream, a (high-level) syntax element (flag) for a set of images is transmitted or obtained to enable or disable any of the aspects or modes mentioned above. For example, a flag may be transmitted within the PPS to control enabling.

[0069] In a third high-level syntax aspect according to the invention, a method is provided for decoding or encoding video or image data from or into a bitstream, wherein a (high-level) syntax element (flag) for each image is transmitted or obtained to enable or disable any of the aspects or modes mentioned above. For example, the flag is transmitted within the image header (PH).

[0070] In a fourth aspect of high-level syntax according to the invention, a method is provided for decoding or encoding video or image data from or into a bitstream, wherein a high-level syntax element (flag) is transmitted or obtained at a lower level, in a slice or block, to enable or disable any of the aspects or modes mentioned above. For example, the flag is transmitted within the slice header (SH).

[0071] Depending on the mode, a signal is transmitted and / or Petition 870250088574, dated 09 / 30 / 2025, page 21 / 119 11 / 73 obtained, which indicates the result of the determination in relation to the limit. In other words, a signal is transmitted or obtained to indicate each criterion that has been met.

[0072] The advantage of high-level syntax signaling aspects and modes is the coder's flexibility to determine the best compromise.

[0073] Furthermore, it is understood that the above embodiments may be combined, where practicable, to form other embodiments in accordance with the invention.

[0074] For example, in a further aspect of the invention, a method is provided for encoding or decoding image data into or from a bitstream, the bitstream including data indicating a partitioning of the image data into a plurality of blocks according to an encoding tree, wherein the blocks in the encoding tree can be divided according to a plurality of divisions, including a quadtree division, the method comprising: determining a value indicating a quadtree division depth of a current block, allowing only a quadtree division depth of the current block among the plurality of modes that is less than a first (limit) value (a first criterion) based on a minimum quadtree depth of an area of ​​a reference frame,If the quadtree split depth of the current block is less than a second value (threshold) (a second criterion) based on a value indicating an average quadtree depth in the reference frame area, i) allow quadtree splitting, ternary splitting, and no splitting, and optionally ii) if ternary splitting has been used at a higher level in the encoding tree of the current encoding unit, also allow binary splitting.

[0075] In another aspect, according to the invention, for each block, the partitioning permissions of divisions are predicted according to the minimum QT division and the average QT division obtained from a temporal area. Thus, when the current QT depth is less than the minimum temporal QT depth minus 1, only QT division is allowed, and when the QT depth Petition 870250088574, dated 09 / 30 / 2025, page 22 / 119 12 / 73 current is less than the temporal average QT depth minus 1, no splitting, QT splitting and TT splitting are allowed, and BT splitting is allowed if TT is selected on a parent node.

[0076] In at least some aspects and modalities, reference is made to binary division. Such binary division may include horizontal binary division and / or vertical binary division. In aspects and modalities, reference is made to ternary division. Such ternary division may include horizontal ternary division and / or vertical ternary division.

[0077] Furthermore, although the above embodiments refer to binary tree (horizontal and vertical), ternary tree (horizontal and vertical), quadtree and divisions of a coding unit or CTU without the possibility of division, it is understood that the invention is not thus limited and other modes may be considered. For example, other geometric divisions may be considered in different numbers of blocks and restricted according to one or more criteria mentioned in the aspects and embodiments above.

[0078] Furthermore, although the terms binary splitting, ternary splitting and quadtree splitting are used, it is generally understood that these refer to partitioning modes in which a block (e.g., a coding unit or coding tree unit) is split or divided into two, three or four subblocks, respectively.

[0079] Other aspects of the invention relate to corresponding encoding methods, an encoding device, a decoding device and an operable computer program for executing the decoding and / or encoding methods of the invention.

[0080] Other aspects of the invention are provided by the independent and dependent claims.

[0081] The program may be provided in isolation or may be executed on, by, or in a carrier medium. The carrier medium may be non-transient, for example, a storage medium, in particular a computer-readable storage medium. The carrier medium may also be transient, for example, a signal or Petition 870250088574, dated 09 / 30 / 2025, page 23 / 119 13 / 73 another means of transmission. The signal can be transmitted via any suitable network, including the Internet. Other features of the invention are characterized by independent and dependent claims.

[0082] Any feature in one aspect of the invention can be applied to other aspects of the invention in any appropriate combination. In particular, aspects of the method can be applied to aspects of the apparatus, and vice versa.

[0083] Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features herein should be interpreted accordingly.

[0084] Any device feature, as described herein, may also be provided as a method feature, and vice versa. As used herein, means features plus function may alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.

[0085] It should also be understood that specific combinations of the various features described and defined in any aspects of the invention may be implemented and / or provided and / or used independently.

[0086] Other aspects of the invention are provided by the independent and dependent claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] By way of example, reference will be made to the attached drawings, in which:

[0088] Figure 1 is a diagram for use in explaining a coding structure used in HEVC.

[0089] Figure 2 is a block diagram that schematically illustrates a data communication system in which one or more embodiments of the invention can be implemented.

[0090] Figure 3 is a block diagram illustrating the components of a processing device in which one or more embodiments of the invention may be implemented. Petition 870250088574, dated 09 / 30 / 2025, page 24 / 119 14 / 73

[0091] Figure 4 is a diagram illustrating the functional elements of an encoder according to the embodiments of the invention.

[0092] Figure 5 is a diagram illustrating the functional elements of a decoder according to the embodiments of the invention.

[0093] Figure 6 shows blocks positioned relative to an actual block, including a colocalized block.

[0094] Figure 7 illustrates a temporal random access GOP structure for 33 frames with related Temporal ID and POC.

[0095] Figure 8 illustrates the 6 possible ways of dividing VVC.

[0096] Figure 9 illustrates maxBtSize and maxMttDepht.

[0097] Figure 10 illustrates an example of the minQtSize variable.

[0098] Figure 11 illustrates some partitioning constraints.

[0099] Figure 12 illustrates incomplete CTUs at the edges of a frame.

[0100] Figure 13 illustrates encoding by maxMttDepth settings based on the temporal ID.

[0101] Figure 14 illustrates one embodiment of the invention.

[0102] Figure 15 illustrates one embodiment of the invention.

[0103] Figure 16 illustrates one of the various temporal positions.

[0104] Figure 17 is a diagram showing a system comprising an encoder or a decoder and a communication network according to the embodiments of the present invention.

[0105] Figure 18 is a schematic block diagram of a computing device for implementing one or more embodiments of the invention.

[0106] Figure 19 is a diagram illustrating a network camera system.

[0107] Figure 20 is a diagram illustrating a smartphone. DETAILED DESCRIPTION

[0108] Figure 1 refers to a coding structure used in the High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC) video standards. A video sequence 1 is composed of a succession of digital images i. Each of these digital images is represented by one or Petition 870250088574, dated 09 / 30 / 2025, page 25 / 119 15 / 73 plus matrices. The matrix coefficients represent pixels.

[0109] An image 2 of the sequence can be divided into slices 3. A slice can, in some cases, constitute an entire image. These slices are divided into non-overlapping Coding Tree Units (CTUs). A Coding Tree Unit (CTU) is the basic processing unit of the High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC) video standards and conceptually corresponds, in structure, to the macroblock units used in several previous video standards. A CTU is also sometimes called a Largest Coding Unit (LCU). A CTU has luma and chroma component parts, each of which is called a Coding Tree Block (CTB). These different color components are not shown in Figure 1.

[0110] A CTU typically has a size of 64 pixels χ 64 pixels for HEVC, but for VVC this size can be 128 pixels χ 128 pixels. Each CTU can, in turn, be iteratively divided into smaller Coding Units (CUs) of variable size 5 using a quadtree decomposition (QT).

[0111] Coding units are the elementary coding elements and consist of two types of subunits called Prediction Unit (PU) and Transform Unit (TU). The maximum size of a PU or TU is equal to the size of the CU. A Prediction Unit corresponds to the partition of the CU for predicting pixel values. Several different partitions of a CU into PUs are possible, as shown in Figure 6, including a partition into 4 square PUs and two different partitions into 2 rectangular PUs. A Transform Unit is an elementary unit subjected to spatial transformation using DCT. A CU can be partitioned into TUs based on a quadtree representation 7.

[0112] Each slice is embedded in a Network Abstraction Layer (NAL) unit. In addition, the video sequence encoding parameters are stored in dedicated NAL units, called parameter sets. In HEVC and H.264 / AVC, two types of parameter sets (NAL units) are Petition 870250088574, dated 09 / 30 / 2025, page 26 / 119 16 / 73 employees: first, a Sequence Parameter Set (SPS) NAL unit, which gathers all the parameters that remain unchanged throughout the entire video sequence. Typically, it handles the encoding profile, video frame size, and other parameters. Second, a Picture Parameter Set (PPS) NAL unit includes parameters that can change from one picture (or frame) to another in a sequence. HEVC also includes a Video Parameter Set (VPS) NAL unit that contains parameters describing the overall structure of the bitstream. The VPS is a type of parameter set defined in HEVC and applies to all layers of a bitstream. A layer can contain multiple temporal sublayers, and all version 1 bitstreams are restricted to a single layer.HEVC has certain layered extensions for scalability and multiview, which will allow multiple layers, with a base layer from version 1 that is backward compatible.

[0113] Other ways of splitting an image have been introduced in VVC, including subimages, which are independently encoded groups of one or more slices.

[0114] Figure 2 illustrates a data communication system in which one or more embodiments of the invention may be implemented. The data communication system comprises a transmitting device, in this case a server 201, which is operable to transmit data packets from a data stream to a receiving device, in this case a client terminal 202, via a data communication network 200. The data communication network 200 may be a Wide Area Network (WAN) or a Local Area Network (LAN). Such a network may be, for example, a wireless network (Wi-Fi / 802.11a or b or g), an Ethernet network, an Internet network, or a mixed network composed of several different networks. In a specific embodiment of the invention, the data communication system may be a digital television transmission system in which the server 201 sends the same data content to multiple clients.

[0115] The data stream 204 provided by server 201 may consist of Petition 870250088574, dated 09 / 30 / 2025, page 27 / 119 17 / 73 multimedia data representing video and audio data. The audio and video data streams, in some embodiments of the invention, can be captured by server 201 using a microphone and a camera, respectively. In some embodiments, the data streams can be stored on server 201 or received by server 201 from another data provider or generated on server 201. Server 201 is equipped with an encoder for encoding video and audio streams, in particular, to provide a compressed bitstream for transmission that is a more compact representation of the data presented as input to the encoder.

[0116] To obtain a better ratio between the quality of transmitted data and the quantity of transmitted data, video data compression can be, for example, in accordance with the HEVC, H.264 / AVC, VVC format or the data format generated by the ECM.

[0117] Client 202 receives the transmitted bitstream and decodes the reconstructed bitstream to reproduce video images on a display device and audio data on a loudspeaker.

[0118] Although a streaming scenario is considered in the example in Figure 2, it should be noted that, in some embodiments of the invention, data communication between an encoder and a decoder can be performed using, for example, a media storage device, such as an optical disc.

[0119] In one or more embodiments of the invention, a video image is transmitted with representative offset offset data for application to the reconstructed pixels of the image, in order to provide filtered pixels in a final image.

[0120] Figure 3 schematically illustrates a processing device 300 configured to implement at least one embodiment of the present invention. The processing device 300 may be a device such as a microcomputer, a workstation, or a lightweight portable device. The device 300 comprises a communication bus 313 connected to:

[0121] - a central processing unit 311, such as a Petition 870250088574, dated 09 / 30 / 2025, page 28 / 119 18 / 73 microprocessor, called CPU;

[0122] - a read-only memory 306, called ROM, for storing computer programs to implement the invention;

[0123] - a random access memory 312, referred to as RAM, for storing the executable code of the method of the embodiments of the invention, as well as registers adapted to record variables and parameters necessary to implement the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to the embodiments of the invention; and

[0124] - a communication interface 302 connected to a communication network 303 through which the digital data to be processed is transmitted or received.

[0125] Optionally, the 300 unit may also include the following components:

[0126] - a data storage medium 304, such as a hard disk, for storing computer programs to implement methods of one or more embodiments of the invention and data used or produced during the implementation of one or more embodiments of the invention;

[0127] - a disk drive 305 to a disk 306, the disk drive being adapted to read data from disk 306 or to write data to said disk;

[0128] - a screen 309 for displaying data and / or serving as a graphical user interface, via a keyboard 310 or any other pointing means.

[0129] The device 300 can be connected to various peripherals, such as, for example, a digital camera 320 or a microphone 308, each connected to an input / output board (not shown) to provide multimedia data to the device 300.

[0130] The communication bus provides communication and interoperability between the various elements included in or connected to the 300 device. The bus representation is not limiting and, in particular, the central processing unit is operable to communicate instructions to any element. Petition 870250088574, dated 09 / 30 / 2025, page 29 / 119 19 / 73 of apparatus 300 directly or by means of another element of apparatus 300.

[0131] The 306 disc can be replaced by any information medium, such as, for example, a rewritable or non-rewritable compact disc (CD-ROM), a ZIP disc or a memory card and, in general terms, by a means of information storage readable by a microcomputer or microprocessor, integrated or not into the device, possibly removable and adapted to store one or more programs whose execution allows the implementation of the method of encoding a sequence of digital images and / or the method of decoding a bit stream according to the invention.

[0132] The executable code may be stored in read-only memory 306, on hard disk 304, or on a removable digital medium, such as, for example, a disk 306, as described above. According to a variant, the executable code of the programs may be received via the communication network 303, via the interface 302, to be stored on one of the storage media of the device 300 before being executed, such as hard disk 304.

[0133] The central processing unit 311 is adapted to control and direct the execution of instructions or parts of the software code of the program or programs according to the invention, instructions which are stored in one of the storage media mentioned above. At initialization, the program or programs stored in non-volatile memory, for example, on the hard disk 304 or in read-only memory 306, are transferred to the random access memory 312, which then contains the executable code of the program or programs, as well as registers to store the variables and parameters necessary for the implementation of the invention.

[0134] In this embodiment, the device is a programmable device that uses software to implement the invention. However, alternatively, the present invention may be implemented in hardware (for example, in the form of an Application-Specific Integrated Circuit or ASIC).

[0135] Figure 4 illustrates a block diagram of an encoder according to Petition 870250088574, dated 09 / 30 / 2025, page 30 / 119 20 / 73 with at least one embodiment of the invention. The encoder is represented by connected modules, each module being adapted to implement, for example, in the form of programming instructions to be executed by the CPU 311 of the device 300, at least one corresponding step of a method that implements at least one embodiment of encoding an image from a sequence of images according to one or more embodiments of the invention.

[0136] An original sequence of digital images i0a in401 is received as input by encoder 400. Each digital image is represented by a set of samples, sometimes also called pixels (hereinafter referred to as pixels).

[0137] A 410 bitstream is generated by the 400 encoder after the encoding process is implemented. The 410 bitstream comprises a plurality of encoding units or slices, each slice comprising a slice header for transmitting encoding values ​​of the encoding parameters used to encode the slice and a slice body, comprising encoded video data.

[0138] The input digital images i0a in401 are divided into pixel blocks by module 402. The blocks correspond to image parts and can have variable sizes (e.g., 4 χ 4, 8 χ 8, 16 χ 16, 32 χ 32, 64 χ 64, 128 χ 128 pixels, and various sizes of rectangular blocks can also be considered). An encoding mode is selected for each input block. Two families of encoding modes are provided: encoding modes based on spatial prediction encoding (intra prediction) and encoding modes based on temporal prediction (inter encoding, Merge, SKIP). The possible encoding modes are tested.

[0139] Module 403 implements an intra-prediction process, in which the block to be encoded is predicted by a predictor computed from pixels in the neighborhood of said block to be encoded. An indication of the selected intra-predictor and the difference between the given block and its predictor is encoded to provide a residual if intra-encoding is selected.

[0140] Time prediction is implemented by the estimation module of Petition 870250088574, dated 09 / 30 / 2025, page 31 / 119 21 / 73 motion 404 and by the motion compensation module 405. First, a reference image from a set of reference images 416 is selected, and a portion of the reference image, also called the reference area or image portion, which is the area closest (closest in terms of pixel value similarity) to the given block to be encoded, is selected by the motion estimation module 404. The motion compensation module 405 then predicts the block to be encoded using the selected area. The difference between the selected reference area and the given block, also called the residual block, is computed by the motion compensation module 405. The selected reference area is indicated using a motion vector.

[0141] Thus, in both cases (spatial and temporal prediction), a residual is computed by subtracting the predictor from the original block.

[0142] In the Intra prediction implemented by module 403, a prediction direction is encoded. In the Inter prediction implemented by modules 404, 405, 416, 418 and 417, at least one motion vector or data to identify such a motion vector is encoded for the temporal prediction.

[0143] The relevant information for the motion vector and residual block is encoded if Inter prediction is selected. To further reduce the bit rate, assuming the motion is homogeneous, the motion vector is encoded by difference with respect to a motion vector predictor. The motion vector predictors from a set of motion information predictor candidates are obtained from the motion vector field 418 by a motion vector prediction and encoding module 417.

[0144] The encoder 400 further comprises a selection module 406 for selecting the encoding mode by applying an encoding cost criterion, such as a rate-distortion criterion. To further reduce redundancies, a transform (such as DCT) is applied by the transform module 407 to the residual block. The resulting transformed data is then quantized by the quantization module 408 and entropy-encoded by the entropic encoding module 409. Finally, the encoded residual block of the current block being encoded is Petition 870250088574, dated 09 / 30 / 2025, page 32 / 119 22 / 73 inserted into bit stream 410.

[0145] Encoder 400 also performs decoding of the encoded image in order to produce a reference image (e.g., those in Reference Images / Figures 416) for motion estimation of subsequent images. This allows the encoder and decoder receiving the bitstream to have the same reference frames (reconstructed images or image parts are used). Inverse quantization (“dequantization”) module 411 performs inverse quantization (“dequantization”) of the quantized data, followed by an inverse transform by the inverse transform module 412. The intra prediction module 413 uses the prediction information to determine which predictor to use for a given block, and the motion compensation module 414 adds the residual obtained by module 412 to the reference area obtained from the set of reference images 416.

[0146] Post-filtering is then applied by module 415 to filter the reconstructed frame (image or image parts) of pixels. In embodiments of the invention, an SAO mesh filter is used, in which compensation offsets are added to the pixel values ​​of the reconstructed pixels of the reconstructed image. It is understood that post-filtering does not always need to be performed. Furthermore, any other type of post-filtering can also be performed in addition to, or instead of, SAO mesh filtering.

[0147] Figure 5 illustrates a block diagram of a decoder 60 that can be used to receive data from an encoder according to an embodiment of the invention. The decoder is represented by connected modules, each module being adapted to implement, for example, in the form of programming instructions to be executed by the CPU 311 of the device 300, a corresponding step of a method implemented by the decoder 60.

[0148] The decoder 60 receives a bit stream 61 composed of encoded units (for example, data corresponding to a block or a coding unit), each consisting of a header containing information about the encoding parameters and a body containing the video data. Petition 870250088574, dated 09 / 30 / 2025, page 33 / 119 23 / 73 encoded. As explained in relation to Figure 4, the encoded video data is entropy-encoded, and the indices of the motion vector predictors are encoded, for a given block, in a predetermined number of bits. The received encoded video data is entropy-decoded by modulo 62. The residual data is then dequantized by modulo 63, and then an inverse transform is applied by modulo 64 to obtain pixel values.

[0149] The mode data that indicate the encoding mode are also decoded by entropy and, based on the mode, an INTRA or INTER type decoding is performed on the encoded blocks (units / sets / groups) of image data.

[0150] In the case of INTRA mode, an INTRA predictor is determined by the intra prediction module 65 based on the intra prediction mode specified in the bitstream.

[0151] If the mode is INTER, motion prediction information is extracted from the bitstream to find (identify) the reference area used by the encoder. The motion prediction information comprises the reference frame index and the motion vector residual. The motion vector predictor is added to the motion vector residual by the motion vector decoding module 70 to obtain the motion vector. The various motion predictor tools used in VVC are discussed in more detail below, with reference to Figures 6 to 10.

[0152] The motion vector decoding module 70 applies motion vector decoding to each current block encoded by motion prediction. Once a motion vector predictor index is obtained for the current block, the actual motion vector value associated with the current block can be decoded and used to apply motion compensation by module 66. The reference image portion indicated by the decoded motion vector is extracted from a reference image 68 to apply motion compensation 66. The motion vector field data 71 is updated with Petition 870250088574, dated 09 / 30 / 2025, page 34 / 119 24 / 73 the decoded motion vector to be used in predicting subsequent decoded motion vectors.

[0153] Finally, a decoded block is obtained. When appropriate, post-filtering is applied by post-filtering module 67. A decoded video signal 69 is finally obtained and provided by decoder 60.

[0154] Random Access Configuration

[0155] Figure 7 shows a temporal random access GOP structure for 33 consecutive frames from 0 to 32. The length of the vertical line representing each frame corresponds to its temporal ID (e.g., the longest length corresponds to temporal ID 0 and the shortest length to temporal ID 5). Frames with temporal ID 0 are the highest in the temporal hierarchy because they can be decoded independently of all other frames with a higher temporal ID value. Similarly, frames with temporal ID 1 are second in the temporal hierarchy and can be decoded independently of all other frames with a higher temporal ID, and so on for the other temporal IDs. In other words, a frame with a specific temporal ID can be decoded independently of frames with higher temporal ID values, but may be dependent on frames with lower temporal IDs. This is known as temporal scalability.

[0156] This parameter is similar to hierarchy depth, but hierarchy depth does not imply decoding independence for all other frames with greater depth.

[0157] VVC Partitioning

[0158] VVC Partitioning has a specific block partitioning. For a tree node, 6 splits are possible, as illustrated in Figure 8:

[0159] - The quad QT,801 split, which divides a block into 4 square blocks of equal size.

[0160] - The binary division BT with its two possible subdivisions 802 and 803:

[0161] - vertical binary division, 802, SPLIT_BT_VER

[0162] - horizontal binary division, 803, SPLIT_BT_HOR Petition 870250088574, dated 09 / 30 / 2025, page 35 / 119 25 / 73

[0163] - the ternary division TT with its 2 possible subdivisions 804 and 805, where the block is divided into 3 blocks with a larger strip in the middle:

[0164] - vertical ternary division, 804, SPLIT_TT_VER

[0165] - horizontal ternary division, 805, SPLIT_TT_HOR

[0166] - The No Division, 806, which terminates a tree node so that there is no division.

[0167] In this description, a block can be a CTU and / or CU or, more generally, any unit in a coding tree.

[0168] Division Control Variables (VVC)

[0169] For a given block, not all possible splits are allowed. The available splits depend on several conditions. These conditions depend on several defined split control variables. A first set of variables defines the maximum and minimum size of the block / node:

[0170] • CTU Size: Corresponds to the size of the root node of a quadtree (e.g., 256 x 256, 128 x 128, 64 x 64, 32 x 32, 16 x 16 luma samples);

[0171] · maxBtSize: This is the maximum allowed size of the root node of the binary tree, that is, the maximum size of a leaf node of the quadtree that can be partitioned by binary division. A current block can be divided using a BT division if the height and width of the current block are less than or equal to maxBtSize. Figure 9 illustrates the concept of maxBtSize, where maxBtSize is the size of the leaf nodes of quadtree 902 of a CTU 901.

[0172] · minBtSize: This is the minimum allowed size for a leaf node in a binary tree; that is, the minimum width or height of a binary leaf node. Thus, a current block can be split using a horizontal BT split if its height is greater than minBtSize. And a current block can be split using a vertical BT split if its width is greater than minBtSize.

[0173] · maxTtSize: This is the maximum allowed size for the root node of the ternary tree, that is, the maximum size of a quadtree leaf node that can be partitioned by ternary splitting. A current block can be split using a TT split if the height and width of the current block are less than or equal to Petition 870250088574, dated 09 / 30 / 2025, p. 36 / 119 26 / 73 maxTtSize.

[0174] · minTtSize: Represents the minimum allowed size for a leaf node in a ternary tree (TT); that is, the minimum width or height of a binary leaf node. But, unlike BT Splitting, a minimum TT partition size is considered to be allowed. Thus, a current block can be split using a horizontal TT split if its height is greater than twice minTtSize. And a current block can be split using a vertical TT split if its width is strictly greater than twice minTtSize.

[0175] • minQtSize: This is the minimum allowed size of the quadtree leaf node (QT); Therefore, for the current block, if the width of the current block is not greater than the minQtSize, the QT split mode is not allowed. Figure 10 illustrates an example of minQtSize. Considering a CTU 128, in the illustrated example, the minQtSize is equal to 16.

[0176] There is no definition of maxQtSize, therefore it corresponds to the size of the CTU.

[0177] The minimum allowed block size for width and height is 4.

[0178] A set of depths is also defined.

[0179] • Depth: is the depth in the tree. In the VVC specification, a leaf is a terminal node of a tree that is a root node of a tree with depth 0. This means that for each split, this value is incremented (by 1).

[0180] • mttDepth: is the depth of the multitree. The multitree includes BT and divisions. TT.

[0181] • maxMttDepth is defined in the VVC specification as the maximum allowed depth for multitree. Therefore, mttDepth is greater than or equal to maxMttDepth. Figure 14 illustrates the concept of maxMttDepth.

[0182] In VVC, these variables are defined independently for Luma and Chroma.

[0183] In VTM software and ECM software, there are several other variables corresponding to depths.

[0184] The variable currBtDepth is the current number of BT divisions used for Petition 870250088574, dated 09 / 30 / 2025, page 37 / 119 27 / 73 reach the current tree node (or current block). The variable currMttDepth is the current number of BT and TT splits used to reach the current tree node (or current block). The variable maxBtDepth corresponds to the maxMttDepth variable of the VVC specification. currQtDepth is the current number of QT splits used to reach the current tree node (or current block). MaxBtDepth: is the maximum allowed depth of the binary tree, i.e., the lowest level at which binary splitting can occur, where the leaf node of the quadtree is the root (e.g., 3).

[0185] VVC Split Control Syntax Elements

[0186] To define the values ​​of these different variables, some high-level syntax elements are passed in SPS, as illustrated in the following table of SPS syntax elements. TABLE 1 seq_parameter_set_rbsp() { Descritor sps_log2_min_luma_coding_block_size_minus2 ue(v) sps_partition_constraints_override_enabled_flag u(1) sps_log2_diff_min_qt_min_cb_intra_slice_luma ue(v) sps_max_mtt_hierarchy_depth_intra_slice_luma ue(v) if( sps_max_mtt_hierarchy_depth_intra_slice_luma != 0 ) { sps_log2_diff_max_bt_min_qt_intra_slice_luma ue(v) s ps_log 2_d iffm axttm i n_qt_i nt ra_s 1 i ce_l u ma ue(v)} if( sps_chroma_format_idc != 0 ) sps_qtbtt_dual_tree_intra_flag u(1) if( sps_qtbtt_dual_tree_intra_flag) { sps_log2_diff_min_qt_min_cb_intra_slice_chroma ue(v) sps_max_mtt_hierarchy_depth_intra_slice_chroma ue(v) if(sps_max_mtt_hierarchy_depth_intra_slice_chroma != 0){ s ps_l og 2_d iff_max_bt_m i n_qt_i nt ra_s 1 ice_c h roma ue(v) s ps_l og 2_d iffmaxttm i n_qt_i nt ra_s 1 i ce_c h roma ue(v)}} s ps_log 2_d iffm i nqtrn i n_c b_i nte r_sl i ce ue(v) sps_max_mtt_hierarchy_depth_inter_slice ue(v) if( sps_max_mtt_hierarchy_depth_inter_slice != 0) { s ps_log 2_d iffm axbtm in_qt_i nte r_s 1 ice ue(v) sps_log 2_d iff_m ax_tt_m i n_qt_i nter_slice ue(v) Petition 870250088574, dated 09 / 30 / 2025, p. 38 / 119 28 / 73 }

[0187] When sps_partition_constraints_override_enabled_flag is enabled in SPS, some image header syntax elements are passed to update the partitioning variables, as illustrated in the following table of PH syntax elements. TABLE 2 picture_header_structure() { Descritor if( sps_partition_constraints_override_enabled_flag) ph_partition_constraints_override_flag u(1) if( ph_intra_slice_allowed_flag) { if( ph_partition_constraints_override_flag) { ph_log2_diff_min_qt_min_cb_intra_slice_luma ue(v) ph_max_mtt_hierarchy_depth_intra_slice_luma ue(v) if( ph_max_mtt_hierarchy_depth_intra_slice_luma != 0 ) { ph_log2_diff_max_bt_min_qt_intra_slice_luma ue(v) ph_log2_diff_max_tt_min_qt_intra_slice_luma ue(v)} if( sps_qtbtt_dual_tree_intra_flag) { p hj og2_d iffm i nqtrn i n_c b_i nt ra_s 1 ice_c h rom a ue(v) ph_max_mtt_hierarchy_depth_intra_slice_chroma ue(v) if( ph_max_mtt_hierarchy_depth_intra_slice_chroma != 0) { ph_log2_d iff_max_bt_m i n_qt_i ntra_sl ice_ch roma ue(v) ph_log2_diff_max_tt_min_qt_intra_slice_chroma ue(v)}}}} if( ph_inter_slice_allowed_flag) { if( ph_partition_constraints_override_flag) { ph_log2_diff_min_qt_min_cb_inter_slice ue(v) ph_max_mtt_hierarchy_depth_inter_slice ue(v) if(ph_max_mtt_hierarchy_depth_inter_slice != 0 ) { p h_l og2_d iffmaxbtm i n_qt_i nte r_sl i ce ue(v) p hj og 2_d iffm a xttrn in_qt_inter_slice ue(v)}} Petition 870250088574, dated 09 / 30 / 2025, page 39 / 119 29 / 73

[0188] VVC Coding Division Mode

[0189] In VVC, the coding split mode is transmitted in the 'coding_tree' coding tree, as illustrated in the following syntax table, where the conditionally parsed flags, split_cu_flag, split_qt_flag, mtt_split_cu_vertical_flag and mtt_split_cu_binary flag, define the split of a CU. TABLE 3 co- ding_tree( xO, yO, cbWidth, cbHeight, qgOnY, qgOnC, cbSubdiv, cqtDepth, mttDepth, depthOffset, partldx, treeTypeCurr, modeTypeCurr) { Descritor if( (allowSplitBtVer | | allowSplitBtHor | | allowSplitTtVer | | allowSplitTtHor | | allowSplitQt) && (xO + cbWidth <= pps_pic_width_in_luma_samples) && ( yO + cbHeight <= pps_pic_height_in_luma_samples)) split_cu_flag ae(v) if( pps_cu_qp_delta_enabled_flag && qgOnY && cbSubdiv <= CuQpDeltaSubdiv ){ IsCuQpDeltaCoded = 0 CuQpDeltaVal = 0 CuQgTopLeftX = xO CuQgTopLeftY = yO} if( sh_cu_chroma_qp_offset_enabled_flag && qgOnC && cbSubdiv <= CuChromaQpOffsetSubdiv) { IsCuChromaQpOffsetCoded = 0 CuQpOffsetcb = 0 CuQpOffsetcr = 0 CuQpOffsetcbcr = 0} if( split_cu_flag) { if( (allowSplitBtVer || allowSplitBtHor || allowSplitTtVer || allowSplitTtHor) && allowSplitQt) split_qt_flag ae(v) if( !split_qt_flag) { if( (allowSplitBtHor || allowSplitTtHor) && (allowSplitBtVer || allowSplitTtVer)) mtt_split_cu_vertical_flag ae(v) if( (allowSplitBtVer &&allowSplitTtVer && mtt_split_cu_vertical_flag ) | | ( allowSplitBtHor && allowSplitTtHor && !mtt_split_cu_vertical_flag )) mtt_split_cu_binary_flag ae(v)} Petition 870250088574, dated 09 / 30 / 2025, page 40 / 119 30 / 73 if( ModeTypeCondition = = 1 ) modeType = MODE_TYPE_INTRA else if( ModeTypeCondition = = 2) { non_inter_flag ae(v) modeType = non_inter_flag ? MODE_TYPE_INTRA : MODE_TYPE_INTER} else modeType = modeTypeCurr

[0190] VVC Division Restrictions

[0191] VVC partitioning has several restrictions. These restrictions are mainly aimed at preventing the same partitioning after several consecutive splits. Figure 11 illustrates some of these restrictions. The idea is to avoid the same partitioning with BT and TT. As illustrated in Figure 11(a), two consecutive vertical BT splits are allowed, but a vertical TT split followed by a vertical BT split in the central block is not allowed, as illustrated in Figure 11(b).

[0192] Similarly, as illustrated in Figure 11(c), two consecutive horizontal BT splits are allowed, but a horizontal TT split followed by a horizontal BT split in the central block is not allowed, as illustrated in Figure 11(d).

[0193] In VVC, there are additional restrictions for the minimum chroma block size and for the maximum TT and BT block size for the interblock size. These restrictions have been removed by the ECM software.

[0194] Chroma Partitioning

[0195] In VVC, Chroma partitioning can be inferred based on Luma partitioning, but this can be disabled. For example, according to the dual-tree mode, the Chroma partitioning tree is independent of the Luma tree. However, there are some restrictions.

[0196] The tree may also be partially dependent on Luma partitioning for CCLM mode; otherwise, it is independent.

[0197] Image Frontier

[0198] The frame resolution is not always equal to an integer multiple of the CTU size. Consequently, there may be incomplete CTUs at the edges of the frame, as illustrated in Figure 12, where CTUs 1201-1206 are incomplete. Petition 870250088574, dated 09 / 30 / 2025, page 41 / 119 31 / 73 due to the lower and right frame limits 1207, 1208. In VVC, in contrast to previous standards, split signaling is allowed at the image boundary. The splitting process at the boundary is applied until the encoding tree node represents a CU located entirely within an image. However, some splits are inferred (not transmitted). Consequently, different variables, such as maxMttDepth, minQtDepth, and minQtSize, are increased or decreased and are possibly different from those used for splits outside the boundary.

[0199] QT BT TT Encoding Selection

[0200] In VTM and ECM software, several encoder-side optimizations are used for QT BT TT encoding selection.

[0201] One of these optimizations includes determining whether QT splitting is tested before BT splitting.

[0202] The condition is that at least one CU to the left of or above the current coding tree node has a QT depth greater than the QT depth of the current coding tree node; and if the width of the CU represented by the current coding tree node is greater than minQtSize * 2.

[0203] If this condition is true, the QT will come before the BT and the divisions will be handled in the following order:

[0204] -No Division

[0205] -QT

[0206] -BT Horizontal

[0207] -BT Vertical

[0208] -TT Horizontal

[0209] -TT Vertical

[0210] Otherwise, the order will be:

[0211] -No Division

[0212] -BT Horizontal

[0213] -BT Vertical

[0214] -TT Horizontal Petition 870250088574, dated 09 / 30 / 2025, p. 42 / 119 32 / 73

[0215] -TT Vertical

[0216] -QT

[0217] This order is important because, according to some optimizations, several divisions will not be tested depending on the results of the first modes tested. Therefore, when QT is tested last, there are many occasions when it will not be evaluated.

[0218] maxMttDepth

[0219] The maximum MTT depth has a significant impact on encoder complexity. Common test conditions for ECM have been updated to reduce encoding through different maxMttDepth settings, as illustrated in Figure 13. In this configuration, maxMttDepth is smaller for some temporal IDs for high resolutions or low QP settings.

[0220] maxBtSize Adaptive

[0221] In VTM and ECM, there is a frame-level encoding option that defines the maximum bit size according to the average block sizes of previously encoded frames with the same depth (=> same temporal ID in the case of CTC RA). The average block size is compared to the limits according to the following pseudocode: if( dBlkSize < AMAXBT_TH32 ) { newMaxBtSize = 32; } else if( dBlkSize < AMAXBT_TH64 ) { newMaxBtSize = 64; } else if( dBlkSize < AMAXBT_TH128 ) { newMaxBtSize = 128; Petition 870250088574, dated 09 / 30 / 2025, p. 43 / 119 33 / 73} else { newMaxBtSize = 256; }

[0222] Where AMAXBT_TH32 equals 15, AMAXBT_TH64 equals 30, and AMAXBT_TH128 equals 60. This method decreases the maximum BT block size when the average block size is small and increases it when it is large.

[0223] MODALITIES

[0224] Possible split modes are restricted according to the current QT depth and other values.

[0225] In one embodiment, the possible partitioning of a block is restricted to a limited number of split modes when the current QT split does not reach a value. In other words, when the current QT split is less than a value or greater than a value, that is, a value that indicates that the QT split does not exceed a limit (value). More precisely, which of the BT, TT, and No Split modes is allowed to split the block depends on the value.

[0226] In this mode, the QT split can be indicated by a block size of the current QT split or by a QT split depth.

[0227] The advantage of this mode is the reduction in encoder complexity, thanks to the reduction in the number of partitions evaluated on the encoder side, and the improvement in encoding efficiency, thanks to the reduction in partitioning signaling.

[0228] Solution 1: Division in QT Only

[0229] According to one solution, the partitioning (splitting) of a block is restricted so as to allow only QT splitting when the current QT split does not reach a value. In one embodiment, only QT splitting for a current block is allowed when the QT depth for a block is less than a value. Therefore, BT, TT, and No Split are not allowed.

[0230] Alternatively, only the QT split for a current block is Petition 870250088574, dated 09 / 30 / 2025, page 44 / 119 34 / 73 is allowed when the block size is greater than a certain value. This is effective when there is a relationship between the block size and the QT depth. In some implementations, there is a direct relationship, but for others, for example, the block size for a given QT depth value may depend on another parameter; for example, in VVC, it also depends on the CTU size.

[0231] Figure 14(a) illustrates this mode. In this figure, the block size is considered to be the maximum between the height and the width, in that order. For example, to be usable for the block at the edge of the image, where square CTU is not possible, the minimum between the width and the height must also be considered. Or, alternatively, only the width or the height. In this figure, the value to be compared with the QT Depth is defined as 3 (for the block size, this corresponds to size 32). As illustrated, for the current QT depths 0, 1, and 2, only the QT split mode is allowed.

[0232] The advantage is a reduction in encoding time, since no partitioning for QT depths lower than this value needs to be evaluated. The second advantage is an improvement in encoding efficiency, since no relative syntax elements before QT depths lower than the value need to be transmitted. But, of course, it is a compromise between the bitrate saved and the quality. Despite this, surprisingly significant gains are achieved with the proposed modification.

[0233] For example, the split_cu_flag syntax element does not need to be passed and therefore the signaling overhead is reduced.

[0234] The CTU size is adapted block by block

[0235] In an alternative to the previous method, the CTU size is determined block by block based on a value representing, for example, a QT depth or a block size. For example, for a current block, another value is obtained representing a QT depth. Additionally, a CTU size related to another value is obtained. Based on both values, the CTU for the current block is obtained. And the current QT depth starts based on that value. Petition 870250088574, dated 09 / 30 / 2025, page 45 / 119 35 / 73

[0236] Solution 2: QT Only and No Division

[0237] In another embodiment, a solution is provided in which only QT splitting and No Splitting for a current block are allowed when the QT depth for a block is less than a value. Therefore, BT and TT splits are not allowed. Alternatively, only QT splitting and No Splitting for a current block are allowed when the block size of a block is greater than a value.

[0238] Figure 14(b) illustrates this mode. In Figure 14(a), the value to be compared with the QT Depth is defined as 3. As illustrated, for the current QT depths 0, 1 and 2, only the QT split and No Split modes are allowed.

[0239] The advantage is the reduction in encoding time, since only the No Division mode needs to be evaluated and the QT division node when the QT depth is less than this value. A second advantage is the improvement in encoding efficiency, since only the No Division flags prior to the QT depth being less than this value need to be transmitted.

[0240] QT Only, No Division, TT

[0241] In one mode, only QT, TT, and No Split for a current block are allowed when the QT depth of a block is less than a value or when the block size is greater than a value. Therefore, only BT split is not allowed.

[0242] Figure 15 illustrates this mode. As in Figures 14(a) and 14(b), the value to be compared with the QT Depth is defined as 3. As illustrated, for the current QT depths 0, 1 and 2, only the BT division is not allowed.

[0243] The advantage is also a reduction in encoding time, as in previous modes, along with an improvement in encoding efficiency.

[0244] Another feature related to the BT division permission depending on the TT division at a higher level Petition 870250088574, dated 09 / 30 / 2025, page 46 / 119 36 / 73

[0245] In one mode, only QT, TT, and No Split for a current block are allowed, and if TT split has been selected at a higher level, BT split is also allowed. To determine if TT has been allowed, the mttDepth value is checked to see if it is greater than 0.

[0246] The advantage, compared to the previous method, is an improvement in coding efficiency with a small impact on coding execution time.

[0247] Combination of Solutions

[0248] In one mode, only QT splitting for a current block is allowed when the QT depth for a block is less than the value “val1” and only QT Split and No Split modes for a current block are allowed when the QT depth for a block is less than another value “val2” and, additionally or alternatively, only QT Split, TT and No Split modes for a current block are allowed when the QT depth for a block is less than another value “val3” and, additionally or alternatively, if TT splitting has been selected at a higher level, BT splitting is also allowed.

[0249] Other combinations may be considered. This modality is illustrated by the following pseudocode: If (QTDepth < val1) { Only QT Split is allowed. If (QTDepth < val2) { QT Split allowed No splits allowed. If (QTDepth < val3) { Petition 870250088574, dated 09 / 30 / 2025, page 47 / 119 37 / 73 QT Split allowed No Split allowed TT Split allowed}

[0250] The advantage is a more optimized reduction of encoder complexity and an increase in encoding efficiency. Obviously, the improvement in encoding efficiency is achieved thanks to an optimized selection of values ​​(“val1”, “val2”, “val3”).

[0251] When there are multiple solutions, the values ​​are ordered.

[0252] In one modality, the values ​​“val1”, “val2”, “val3” are ordered according to the following formula: val1 <= val2 <= val3

[0253] val1 is less than or equal to val2 and val3 is less than or equal to val2. Therefore, in this case, the smaller the value, the more restricted the possible division modes.

[0254] The advantage is an optimized reduction in encoder complexity and an increase in encoding efficiency.

[0255] This mode is dedicated to comparing the QT depth with a value. If, instead, the block size is considered, the order can be dictated according to the following conditional formula: val1 >= val2 >= val3

[0256] Modalities Related to the Temporal Area

[0257] Obtained from a temporal area

[0258] In one embodiment, for one or more of the criteria described above, the QT depth of the current block is compared to a value determined from a temporal area. In other words, from an area of ​​a reference frame that is not the current frame. The temporal area, in some embodiments, may have the same position and / or the same size as the current block (e.g., colocalized) or the same position and / or size as a current CTU (largest block) with a coding tree that includes the current block. More generally, the temporal area may correspond to or be associated with the current block of Petition 870250088574, dated 09 / 30 / 2025, page 48 / 119 38 / 73 current situation somehow.

[0259] For example, the QT depth for the temporal area (e.g., colocalized block) is considered for the first criterion, where only QT is available, according to the following pseudocode: If (QTDepth < QTDepthCol-1) { Only QT Split allowed.

[0260] Where QTDepthCol is the QT depth of the temporal area (e.g., colocalized block).

[0261] The advantage is an optimized reduction in coding execution time and an improvement in coding efficiency.

[0262] More details about the different possible time areas for obtaining value are presented below.

[0263] Colocalized Area

[0264] In one embodiment, the block used to determine the temporal value of QT depth or block size is a temporally colocalized area or block, that is, an area or block in the reference frame that is colocalized with the current block. An example of a possible definition of a colocalized area or block is shown in Figure 6. As mentioned above, colocalized in this context can mean an area or block that has the same position (e.g., has the same origin) as the block in the current frame or both the same size and the same position (as in the example in Figure 6). Alternatively, it can mean an area or block in the reference frame that is encompassed by the area of ​​the current block when projected onto the reference frame. It can also be said that the temporal area corresponds to the current block (i.e., is colocalized) even when the area completely encompasses the area of ​​the current block if it were projected directly onto the reference frame.This situation can arise, for example, when the partitioning of the reference frame is different from the current frame and a rounding operation is needed to determine which one it is. Petition 870250088574, dated 09 / 30 / 2025, p. 49 / 119 Block 39 / 73 in the reference frame should be considered to correspond to or be co-located with the current block.

[0265] Plurality of Positions

[0266] In one embodiment, a plurality of block positions in the reference frame is used to determine a temporal value of the QT depth or block size. For example, the C, TL, TR, BL, and BR positions in Figure 16 can be considered. In this figure, the C position is the center of the colocalized temporal block. The TL, TR, BL, and BR positions are, respectively, the Upper Left, Upper Right, Lower Left, and Lower Right positions around the colocalized temporal block. Other adjacent block positions are also possible.

[0267] Compared to the previous method, the current one provides a greater reduction in encoding time and increases encoding efficiency, as the determined temporal QT depth or block size is often more reliable.

[0268] An area larger than the area of ​​the current block

[0269] In one embodiment, a temporal value of QT depth or block size is determined based on a temporal area larger in size (area) than the actual block. For example, the temporal area is a colocalized CTU.

[0270] Compared to the two previous methods, more blocks can be considered, therefore the trade-off between reducing encoding time and encoding efficiency is better.

[0271] The center of the current block is used to determine the temporal positions.

[0272] In one embodiment, to determine the colocalized block, multiple temporal blocks, or a temporal area, the center of the current block is considered. This is further illustrated by the 'Center' block shown within the colocalized area in Figure 6 or by position c in Figure 16.

[0273] The advantage is a better understanding, since the center is the best Petition 870250088574, dated 09 / 30 / 2025, page 50 / 119 40 / 73 position to represent the current block.

[0274] Alternatively, when the center of the block is outside the current frame, the upper left position is considered.

[0275] Full Frame

[0276] In one modality, a temporal value of QT depth or block size is determined based on all blocks in a temporal frame.

[0277] The advantage of this method is the simplification of the process for determining the temporal QT depth value or block size, but it is less efficient because it is less content-adapted compared to previous methods.

[0278] Frame with the same temporal ID

[0279] In one modality, the colocalized block, multiple temporal blocks, or a temporal area originate from a frame with the same temporal ID.

[0280] For the Random Access configuration example, as shown in Figure 7, if the current frame has a temporal ID equal to 4, another encoded / decoded frame with the same temporal ID is used to determine the values ​​of the proposed method.

[0281] Frames with the same temporal ID often have the same encoding parameters, especially if they have the same or similar QP and the same spatial distances relative to their reference frames. Therefore, they are very interesting for predicting QT depth, as this data is correlated with QP and the spatial distance between frames.

[0282] Closest frame with the same temporal ID

[0283] In one modality, the temporal block or several co-located temporal blocks or a temporal area come from the nearest frame with the same temporal ID.

[0284] For the Random Access configuration example, as shown in Figure 7, the closest frame with the same temporal ID is (generally) more correlated than the others. Therefore, the result is better. Petition 870250088574, dated 09 / 30 / 2025, page 51 / 119 41 / 73

[0285] Frame or a reference frame with the same QP

[0286] In one modality, the colocalized temporal block or blocks or temporal area originate from a frame or a reference frame with the same QP. Ideally, a reference frame with the same QP.

[0287] As mentioned above, QP has a significant influence on block partitioning. Therefore, with a frame having the same QP, temporal QT depth or block size is a better predictor.

[0288] The reference frame is the same as that used for the prediction of the temporal motion vector.

[0289] In one embodiment, the colocalized temporal block or blocks, or a temporal area, come from the reference frame that is used for the prediction of the temporal motion vector. This may be the first reference of Reference List 0 or the first reference frame of List 1, according to a flag transmitted in the image header or in the slice header.

[0290] Surprisingly, this mode offers the best compromise between reducing encoder time and encoding efficiency, even though this reference frame has a lower QP. However, it is closer to the current frame compared to all frames with the same temporal ID.

[0291] Nearest reference frame

[0292] In one modality, the temporal block or blocks, or a temporal area, are derived from the nearest reference frame.

[0293] As explained in the previous embodiment, the distance to the current reference frame seems more interesting for the trade-off between reducing encoder time and coding efficiency, even though frames with the same PQ statistically have more correlations between their PQ depths.

[0294] More than one frame of reference

[0295] In one embodiment, two reference frames are considered and two temporal areas, or 2 sets of multiple blocks, or 2 colocalized blocks, are used to determine 2 temporal QT depths or 2 block sizes. Petition 870250088574, dated 09 / 30 / 2025, page 52 / 119 42 / 73 They are then used to determine a QT depth or a block size. For example, the minimum QT depth of the 2 temporal areas can be considered.

[0296] More than two reference frames can also be considered.

[0297] The advantage is a better compromise between reducing encoder time and encoding efficiency, as the QT depth value is computed from more data. This is particularly efficient when both reference frames have the same time distance, but increases the number of memory accesses.

[0298] Optional resources related to the value(s) (limit)

[0299] The other value is the average QT Depth value for a temporal area.

[0300] In one modality, the QT Depth of the current block is compared with an average of the QT Depth values ​​determined from a time area.

[0301] For example, the QT depth values ​​of the time positions in Figure 20 are considered to calculate the average value AverageQTDepthTempo, which can be used for the first criterion thanks to the following pseudocode: If (QTDepth < AverageQTDepthTime-2) { Only QT Split is allowed.

[0302] Compared to the previous method, this method provides a greater reduction in encoding time and increases encoding efficiency, since the average temporal QT depths are often more reliable.

[0303] In an alternative approach, a larger time area may be considered.

[0304] The other value is the minimum QT Depth value of a temporal area.

[0305] In one modality, the QT Depth of the current block is compared Petition 870250088574, dated 09 / 30 / 2025, p. 53 / 119 43 / 73 with minimum value of QT Depth values ​​determined from a temporal area.

[0306] For example, the first criterion can be written as the following pseudocode: If (QTDepth < MinQTDepthTime-1) { Only QT Split is allowed.

[0307] The advantage is that the minimum temporal QT depth provides a value for which there is a very low probability that the QT depth for the current block will be less than that value. Therefore, it offers a highly optimized value and, consequently, an excellent compromise between reducing encoding time and encoding efficiency.

[0308] The other value is the value computed thanks to the maximum MTT depth or the average MTT depth of a temporal area.

[0309] In one mode, the value compared with the QT depth of the current block for the defined criteria is computed according to the maximum MTT depth of a temporal area “MaxTempoMTTDepth”.

[0310] Alternatively, the average MTT depth can be considered.

[0311] The other value is the value computed thanks to the maximum or average depth of a temporal area.

[0312] In one modality, the value compared with the QT depth of the current block for the defined criteria is computed according to the maximum depth of the encoding tree used in a temporal area “MaxTempoDepth”. Or, alternatively, according to the average temporal depths.

[0313] An example of the two previous modalities can be the second criterion provided by the following pseudocode: If (QTDepth < (MaxTempoDepth -MaxTempoMTTDepth) ) { QT is allowed Petition 870250088574, dated 09 / 30 / 2025, page 54 / 119 44 / 73 No Split is allowed.

[0314] This alternative method offers a good compromise between reducing coding time and coding efficiency.

[0315] Modality with a particularly advantageous combination of features

[0316] In one mode, QT splitting is the only split allowed when the QT depth of the current block value is less than the minimum of the QT depth values ​​of a temporal area “MinQTDepthTempo” minus 1. And only the QT, TT, and No Split modes for a current block are allowed if the average depth is reached. If TT splitting has been selected at a higher level, BT splitting is also allowed when the QT depth of the current block value is less than the average of the QT depth values ​​of a temporal area “AverageQTDepthTempo” minus 1. If (QTDepth < MinQTDepthTime - 1) { Only QT split allowed. If (QTDepth < AverageQTDepthTime - 1) { QT Split allowed No splitting allowed. TT split allowed If (TT was previously selected in the tree) { BT split allowed}}

[0317] Note that “If (QTDepth < MinQTDepthTempo - 1)” can be written Petition 870250088574, dated 09 / 30 / 2025, p. 55 / 119 45 / 73 as “If (QTDepth+1 < MinQTDepthTempo)”, similarly, “If (QTDepth < AverageQTDepthTempo - 1)” can be written as “If (QTDepth +1 < AverageQTDepthTempo in po)”.

[0318] This specific combination provides the best compromise between encoding efficiency and encoder execution time reduction. In fact, thanks to the limit set at MinQTDepthTempo - 1, an almost systematic bitrate saving is obtained. Thanks to the difference between MinQTDepthTempo and AverageQTDepthTempo, the selection between QT division and TT division is well adapted, since these are the division modes with the highest number of subdivisions. Eventually, if TT is not selected or QT is above the average time, there is a high probability that BT or TT division can be selected.

[0319] According to a value transmitted in a header corresponding to the current CTU

[0320] In one embodiment, a list of values ​​is transmitted directly in the high-level header as the image header or slice header for each CTU. For example, an image header table ph_QTDepthLimit[] is transmitted. The size of this table is equal to the number of CTUs in the image.

[0321] The advantage of this modality is that the encoder can select the compromise between encoding efficiency and encoder complexity with an adapted encoder criterion. Furthermore, compared with temporal area-based modalities, analysis and decoding do not need to access temporal information.

[0322] The values ​​can be predicted from each other

[0323] In one modality, the transmitted values ​​are predicted from each other.

[0324] For example, the ph_QTDepthLimit [N] value associated with CTU number N is predicted by the ph_QTDepthLimit [N-1] value associated with CTU number N-1. And only the residual (ph_QTDepthLimit [N] - ph_QTDepthLimit [N-1]) is transmitted.

[0325] The advantage is the reduction in the rate of these syntax elements in the image header. Petition 870250088574, dated 09 / 30 / 2025, p. 56 / 119 46 / 73

[0326] A set of values ​​can be passed to all criteria (e.g., all conditional constraints in split modes)

[0327] When considering multiple criteria, only one set of values ​​needs to be transmitted, instead of one per criterion.

[0328] For example, when using 2 criteria, the pseudocode could be: If (QTDepth < ph_QTDepthLimit [N]) { Only QT Split is allowed. If (QTDepth < ph_QTDepthLimit [N] + 1) { QT Split allowed No splitting allowed. TT split allowed If (TT was previously selected in the tree) { BT split allowed}}

[0329] The advantage is the reduced rate compared to a solution where multiple values ​​are transmitted.

[0330] Other Variants

[0331] The value is adapted between MinQTDepthCol or QTDepthCol / +2, 1, 0, -1, -2 etc.:

[0332] In one mode, the value for the current QT depth (or block size) is adapted according to one or more parameters. According to the example of the preferred mode, an “OFFSET” is added to the value representing the temporal QT depth or the minimum QT depth, as in the following pseudocode: Petition 870250088574, dated 09 / 30 / 2025, p. 57 / 119 47 / 73 If (QTDepth < MinQTDepthTime - 1 + OFFSET) { Only QT Split is allowed. If (QTDepth < AverageQTDepthTime - 1 + OFFSET) { QT Split allowed No Split allowed TT split allowed If (TT was selected previously in the tree) { BT split allowed}}

[0333] Note that two displacements can also be considered, of which one can be selected.

[0334] The advantage is a better compromise between reducing encoder time and encoding efficiency, since the method is adapted for some parameters. Especially for those listed below.

[0335] Toggle between “No BT” only and “No BT and TT”.

[0336] In one embodiment, for the second criterion, switching between methods is enabled according to certain parameters. For example, when the “Cond” condition is false, BT and TT are not allowed; otherwise, only BT is not allowed, as in the following pseudocode: If (QTDepth < AverageQTDepthTime - 1) { QT Split allowed No Split allowed If(Cond) Petition 870250088574, dated 09 / 30 / 2025, page 58 / 119 48 / 73 { TT split allowed If (TT was selected previously in the tree) { BT split allowed}}}

[0337] For example, the “Cond” condition could be the maximum multitree depth of the timeframe is less than the current maximum multitree depth.

[0338] Avoid one or more criteria

[0339] In one modality, one of the criteria is not applied when a condition is met. For example, when the maximum multitree depth of the timeframe is less than the current maximum multitree depth, the first criterion is not applied. In fact, if the maximum multitree depth increases, QTdepth must be reduced. Therefore, there is a risk of not achieving an interesting coding efficiency.

[0340] In fact, for some configurations, both criteria are not useful.

[0341] Time frame slice type

[0342] In one modality, a parameter for adapting the method is the type of time area slice in the reference frame. For example, when the time area comes from an Intra slice, the partitioning is different, as Intra prediction works better when block sizes are small. Therefore, in this case, the “OFFSET” value is set to -1 to take into account that the QT depth of an Intra frame is generally greater than that of an Inter frame for equivalent quality. The threshold may also depend on other parameters used for encoding the Intra frame, such as the maximum MTT depth and the minimum QT size.

[0343] MaxMttDepth of the current frame and the temporal area frame Petition 870250088574, dated 09 / 30 / 2025, page 59 / 119 49 / 73

[0344] In one embodiment, the method is adapted according to the maxMttDepth of the current frame and the maxMttDepth of the temporal area frame. For example, the “OFFSET” offset is set to +1 when the current maxMttDepth is less than the maxMttDepth of the temporal area frame (and -1 if it is greater). With this setting, “OFFSET equal to +1”, the QT depth inference is increased with the first criterion. In fact, if maxMttDepth decreases, it is more likely that the ideal minimum block size cannot be achieved. Therefore, it is preferable to increase the QT depth.

[0345] MaxMttDepth of the current frame and maxMttDepth of the temporal area

[0346] In one embodiment, the method is adapted according to the maxMttDepth of the current frame and the maxMttDepth of the temporal area. For example, the maxMttDepth of the temporal area is the average of the mttDepth of the temporal area. (The minimum or maximum can also be considered). MaxMttDepth for the frame of the temporal area is the value for the entire frame and can be passed in the image header PH_MaxMttDepth. MaxMttDepth of the temporal area is the maximum mttDepth of the block or blocks contained in the temporal area (e.g., colocalized).

[0347] This method is better than the previous one because it is more tailored to the content, rather than a coding parameter that was not defined based on the target quality content.

[0348] MaxMttDepth of the current frame and maxMttDepth of the temporal area reference frame are equal

[0349] In one embodiment, when maxMttDepth of the current frame and maxMttDepth of the temporal area reference frame are equal, the method is adapted according to maxMttDepth of the current frame and maxMttDepth of the temporal area.

[0350] This mode is particularly efficient when the CTU size is equal to 256.

[0351] MinQtSize

[0352] In one embodiment, the method is adapted based on the minQtSize. In fact, the minQtSize influences the QT depth of each block, therefore, this Petition 870250088574, dated 09 / 30 / 2025, page 60 / 119 The 50 / 73 parameter impacts the efficiency of the criterion.

[0353] As with the two previous methods for maxMttDepth, the minQtSize of the current frame can be compared with the minQtSize of the time area frame or with the minimum / maximum or average of the minQtSize of the time area.

[0354] If the current frame has a QP greater than the frame of the temporal area

[0355] In one modality, the method is adapted when the current frame has a lower QP than the frame of the temporal area. For example, the OFFSET value is set to +1 or more if the QP difference is high. In fact, a block in a frame with a higher QP has, on average, a lower QT depth. Therefore, the minimum QT depth of that frame with a higher QP is lower than if that frame had a lower QP. Consequently, the criterion values ​​must be increased.

[0356] Whether the current framework is a reference framework or not

[0357] In one modality, the method is adapted when the current framework is a reference framework or not.

[0358] For example, the OFFSET value has a positive value when the current frame is not a reference frame. In fact, thanks to a larger offset, the encoding will decrease, but if the frame quality decreases, this will not impact the future frame. This is especially true for the second criterion, where BT and TT are not allowed.

[0359] Disabled for display content encoding

[0360] In one embodiment, the proposed method is adapted to display content encoding. In particular, the method is disabled for display content. It was found that when the sequence content contains display content, it appears more difficult to predict the partitioning parameters.

[0361] Alternatively, the number of IBC blocks is computed in the time area and, according to a limit, the method is disabled for the current block.

[0362] Partially disable weather prediction

[0363] In one embodiment, the proposed method is partially disabled when display content is detected. This embodiment and the following ones are Petition 870250088574, dated 09 / 30 / 2025, page 61 / 119 51 / 73 described independently of the detection or presence of display content, which will be described later.

[0364] The advantage of this approach is an improvement in encoding efficiency and a compromise in complexity. In fact, as mentioned earlier, it is difficult to predict the partitioning parameters of display content sequences. Thanks to partial prediction, the reduction in encoding time is maintained and losses for these sequences are reduced.

[0365] Disable Solution 1: QT Division Only

[0366] In one embodiment, when the current block is considered to need to be adapted for display content encoding, the restriction of only allowing QT splitting when the current QT splitting does not reach a disabled value (Solution 1) is not allowed. In this embodiment, at least one other partitioning prediction solution, as described previously, remains enabled.

[0367] The advantage is an increase in encoding efficiency for blocks containing display content, while maintaining the reduction in encoding time thanks to at least one other partitioning prediction solution / mode, as described previously.

[0368] Disable Solution 2: QT Only and No Division

[0369] In one embodiment, when the current block is considered to need to be adapted for display content encoding, the restriction of allowing QT and No Split only for a current block when the QT depth of a block is less than a value (Solution 2) is not allowed. In this embodiment, at least one other partitioning prediction solution, as described previously, remains enabled.

[0370] The advantage is an increase in encoding efficiency for blocks containing display content, while maintaining the reduction in encoding time thanks to at least one other partitioning prediction solution / implementation, as described previously.

[0371] Disable Solution 2.2: QT Only, No Division, TT Petition 870250088574, dated 09 / 30 / 2025, page 62 / 119 52 / 73

[0372] In one embodiment, when the current block is considered to need to be adapted for display content encoding, the restriction of allowing only QT, TT, and No Split for a current block when the QT depth for a block is less than a value or when the block size is greater than a value is not allowed. In this embodiment, at least one other partitioning prediction solution, as described previously, remains enabled.

[0373] The advantage is an increase in encoding efficiency for blocks containing display content, while maintaining the reduction in encoding time thanks to at least one other partitioning prediction solution / implementation, as described previously.

[0374] Disable Solution 2 + enable BT split depending on TT split

[0375] In one mode, when it is considered that the current block should be adapted for display content encoding, the restriction of permission only to splitting into QT, TT, and No Split for a current block, and if the TT split has been selected at a higher level, the BT split is also allowed, is not allowed. To determine if TT has been allowed, the mttDepth value is checked to see if it is greater than 0. In this mode, at least one other partitioning prediction solution, as described previously, remains enabled.

[0376] The advantage is an increase in encoding efficiency for blocks containing display content, while maintaining the reduction in encoding time thanks to at least one other partitioning prediction solution / mode, as described previously.

[0377] The following example illustrates an implementation of the proposed modalities based on a previous illustration: If ((QTDepth < val1) && !(IsScreenContent)) { Only QT Split allowed Petition 870250088574, dated 09 / 30 / 2025, page 63 / 119 53 / 73} If ((QTDepth < val2) && !(IsScreenContent)) { QT Split allowed No Split allowed} If (QTDepth < val3) { QT Split allowed No Split allowed TT Split allowed}

[0378] In this example, && is the AND operator, ! is the NOT operator, and IsScreenContent is a boolean that is equal to TRUE when the current block should be adapted for display content encoding and equal to FALSE otherwise.

[0379] Adaptation when display content is detected

[0380] As mentioned earlier, enabling or disabling one or more modes, or all modes of temporal partitioning prediction, depends on the detection (presence) of display content. The following modes describe some methods for detecting whether the current block potentially contains display content. Each of these methods is applicable to the modes already described, which depend on the detection or implicit presence of display content.

[0381] When Palette, IBC, Transform Skip or other SCC tools are enabled

[0382] In one mode, when Palette, IBC, or Transform Skip are enabled, the current block is considered to potentially contain display content. Other future efficient encoding tools for display content may also be considered.

[0383] In fact, the IBC and Palette modes are some very useful modes for Petition 870250088574, dated 09 / 30 / 2025, page 64 / 119 54 / 73 Display content encoding. And Transform Skip is the efficient way to encode residual texture from display content. Therefore, when these tools are enabled or allowed, there is a chance that the encoding block is in a display content area. But IBC and Transform Skip are known to offer encoding efficiency for natural content.

[0384] Therefore, in an alternative mode, the detection of the display content depends only on the Palette mode.

[0385] The advantage of these modes is an increase in encoding efficiency compared to signaling the enable or disable modes for the encoding block or for a slice, image or sequence, since enabling / disabling depends on the data present in the bitstream.

[0386] If flagged

[0387] In one embodiment, the current block is considered to potentially contain display content when Palette, IBC, or Transform Skip are signaled. And, as mentioned earlier, this may depend on signaling only the Palette mode or on a future display content encoding tool.

[0388] The advantage is the same as in the previous option.

[0389] SPS, PPS, image header, slice header

[0390] In one embodiment, enabling / disabling IBC, Palette, or Transform Skip mode, or future encoding tools for display content, is used to determine whether the current block is potentially within the display content area. Enabling / disabling can be signaled in SPS, PPS, image header, or slice header.

[0391] For example, as described earlier, the boolean IsScreenContent can be the SPS flag sps_palette_enabled_flag. Alternatively, as defined, this boolean can be changed to a PPS flag pps_palette_enabled_flag, an image header flag ph_palette_enabled_flag, or a slice header flag sh_palette_enabled_flag. Petition 870250088574, dated 09 / 30 / 2025, p. 65 / 119 55 / 73

[0392] The advantage is the same as in the previous option.

[0393] If flagged for the reference frame containing the temporal area

[0394] In one embodiment, the current block is detected as a potential display content area if Palette, IBC, or Transform Skip mode, or future display content encoding tools, are present within the reference frame containing the temporal area. Each encoding tool may be considered, or several may be considered. Specifically, only Palette mode may be considered. In this embodiment, one or more encoding tools may be enabled at the slice level, image level, or PPS level for the reference frame.

[0395] The advantage is the same as the previous method and is more precise.

[0396] If at least one of the SCC content tools is detected

[0397] In an alternative embodiment to the previous one, if at least one block has been coded with the display content coding tool or more display content coding tools, the current block may be considered as potential display content. As before, alternatively, only Palette may be considered.

[0398] Compared to previous methods, detection is better at actual selection and not just at enabling or disabling coding tools. But, compared to previous methods, this one is more complex, as it requires checking all blocks.

[0399] If at least one of the SCC content tools is detected in the reference frame

[0400] In one modality, in addition to the characteristics of the previous modality, the blocks considered are the blocks of the reference framework.

[0401] Compared to the previous method, detection is better because the reference frame is the most relevant frame for encoding the current frame.

[0402] If at least one of the SCC content tools is detected in the temporal area Petition 870250088574, dated 09 / 30 / 2025, p. 66 / 119 56 / 73

[0403] In an additional embodiment, the reference frame blocks considered for detection are the blocks within a (or the) temporal area. Alternatively, the colocalized block may be considered, or the colocalized block and its neighboring blocks.

[0404] Compared to the previous method, detection is better because the temporal area or other alternatives are more relevant as they are temporally closer to the current block. Furthermore, this method is especially more efficient for mixed content sequences (containing a display content portion and a natural content portion). However, this method is also more complex, as it is necessary to check the internal blocks of each temporal area. Unlike the reference frame, where blocks only need to be checked once.

[0405] Based on the number of occurrences

[0406] In one embodiment, the current block is detected as a potential display content area if a certain number of Palette, IBC, or Transform Skip mode tools or future display content encoding tools are present within the reference frame containing the temporal area. In this embodiment, either the reference frame or the temporal area can be considered. The occurrences of these encoding tools are computed for the considered area (reference frame or temporal area). As mentioned earlier, Palette mode is an efficient mode only for display content encoding, therefore its occurrence can be considered equal to 1 for detection. For IBC, this mode is more suitable for display content encoding than for natural content encoding. For example, if more than 10% of the blocks in the considered area are encoded with this mode, display content is detected.Similarly, for Transform Skip tools. For future display content encoding tools, appropriate statistics may be considered.

[0407] Alternatively, the number of pixels (or samples) that use these encoding tools can be considered instead of the occurrences. Petition 870250088574, dated 09 / 30 / 2025, page 67 / 119 57 / 73

[0408] The advantage of this implementation is an improvement in coding efficiency, since the occurrences should offer the possibility of using IBC and Transform Skip more efficiently for detection. Furthermore, this implementation is better suited for detecting natural content and displaying mixed content. However, the computation of occurrences and their storage is more complex than in the previous implementation.

[0409] Disable fully or partially when SCC is detected and when the reference is an intra-frame

[0410] In one embodiment, the proposed temporal partitioning prediction is disabled entirely or partially when it is detected that the current block potentially contains display content and that the reference frame containing the temporal area is an intra frame. Alternatively, this restriction may be applied to one solution and not to another solution / embodiment. In other words, the additional restriction that the reference frame is an intra frame may be selectively applied to one or more of the solutions described above. Alternatively, it may be applied whenever display content detection is used.

[0411] The advantage of this mode is an improvement in encoding efficiency and a reduction in complexity compared to previous modes. In fact, intra-frames containing display content are encoded with small blocks. For example, to match the letters of a text thanks to Palette mode or IBC mode. On the other hand, since movement is very low or localized only to the mouse pointer, inter-frames contain large blocks. Therefore, it is not suitable for this proposed temporal partitioning prediction to predict inter-partitioning thanks to intra-partitioning. However, partitioning prediction works between inter-frames for display content.

[0412] To illustrate the different modes, some examples are provided. In the following example, display content detection is based on the sps_palette_enabled_flag string parameter flag, but this can be overridden by each display content detection mode as described. Petition 870250088574, dated 09 / 30 / 2025, p. 68 / 119 58 / 73 previously.

[0413] In the following example, when Palette mode is enabled for the sequence (sps_palette_enabled_flag), the solution where only QT splitting is allowed is not applied, but the solution where all splits are allowed except BT splitting is allowed. If ((QTDepth +1 < MinQTDepthTime) && !( sps_palette_enabled_flag)) { Only QT split allowed. If (QTDepth +1 < AverageQTDepthTime) { QT Split allowed No split allowed. TT split allowed. If (TT was previously selected in the tree) { BT split allowed}}

[0414] In the following example, when Palette mode is enabled for the sequence (sps_palette_enabled_flag), the solution that allows only QT splitting is not applied only when Palette mode is enabled and the current reference frame is an intra-frame; otherwise, it is allowed. And the second criterion, which does not allow only BT splitting, is not applied if Palette mode is enabled. In this example, when Palette mode is enabled, the proposed temporal prediction of partitioning is completely disabled when the reference frame is an intra-frame. And when the reference frame is not an intra-frame, only the first part of the prediction is allowed. Petition 870250088574, dated 09 / 30 / 2025, p. 69 / 119 59 / 73 If ((QTDepth +1 < MinQTDepthTime) && (!(sps_palette_enabled_flag) || !(RefFrame->IntraFrame))) { Only QT split allowed. If ((QTDepth +1 < AverageQTDepthTime) && !(sps_palette_enabled_flag)) { QT Split allowed No split allowed. TT split allowed. If (TT was previously selected in the tree) { BT split allowed}}

[0415] In the following example, when Palette mode is enabled for the sequence (sps_palette_enabled_flag), the solution that only allows QT splitting is not applied only when Palette mode is enabled. And the second criterion, that only BT splitting is disabled, is not applied if Palette mode is enabled and if the current reference frame is an intra frame; otherwise, it is allowed. In this example, when Palette mode is enabled, the proposed temporal prediction of partitioning is completely disabled when the reference frame is an intra frame. And when the reference frame is not an intra frame, only the second part of the prediction is allowed. If ((QTDepth +1 < MinQTDepthTime) && !(sps_palette_enabled_flag)) { Only QT split allowed. Petition 870250088574, dated 09 / 30 / 2025, pp. 70 / 119 60 / 73 If ((QTDepth +1 < AverageQTDepthTime) && (!(sps_palette_enabled_flag) || !(RefFrame->IntraFrame))) { QT Split allowed No Split allowed TT split allowed If (TT was selected previously in the tree) { BT split allowed}}

[0416] Disabled for low latency configuration

[0417] In one embodiment, the proposed method is adapted for low latency configuration. In particular, the method is disabled.

[0418] Disabled using a flag

[0419] In one embodiment, the proposed method is enabled or disabled using at least one flag transmitted in at least one header.

[0420] Not based on QTDepthCol, but based on block size statistics, as in the VTM “Adaptive Maximum BT Size” encoding option.

[0421] In one embodiment, the proposed method is based on the statistics of the block sizes in the temporal area. For example, the average block sizes in the temporal area can be considered. For example, if said average multiplied by the maximum multitree depth of the current image is less than the current block size (or a value indicative of the current block size), the first criterion is applied.

[0422] Determination by CTU instead of by block.

[0423] In one embodiment, the criteria parameters are determined based on the CTU rather than based on the block, so that the same parameters Petition 870250088574, dated 09 / 30 / 2025, pp. 71 / 119 61 / 73 should be used for all CTU blocks. For example, MinQTDepthTempo and AverageQTDepthTempo are determined for each CTU and are the same for all CTU blocks. This can be easily achieved by considering the current CTU center for each CTU block.

[0424] This reduces the trade-off between reducing coding execution time and coding efficiency, but is better suited when combined with other methods.

[0425] Solution 3: When a TT division has been previously selected in the division, BT is allowed.

[0426] In a modality, independent or combinable with any previous modalities, if the TT division was selected at a higher level, the BT division is also permitted.

[0427] To determine if TT was allowed, the mttDepth value is checked to see if it is greater than 0. Alternatively, if the difference between mttDepth and btDepth is equal to 0 (this can be used for specific minTtSize, maxTtSize, minBtSize, maxBtSize settings).

[0428] The advantage of this method is an improvement in coding efficiency with a small impact on increasing coding time.

[0429] High-Level Syntax Signaling

[0430] The method is enabled for the entire sequence.

[0431] In one embodiment, a high-level syntax flag for the entire sequence is passed to enable or disable the proposed method. For example, a flag might be passed within the SPS to indicate that the method is enabled.

[0432] The method is enabled for a set of images

[0433] In one embodiment, a high-level syntax flag for an image set is passed to enable or disable the proposed method. For example, a flag might be passed within the PPS to indicate enabling at the image level.

[0434] The method is enabled for each image Petition 870250088574, dated 09 / 30 / 2025, p. 72 / 119 62 / 73

[0435] In one embodiment, a high-level syntax flag for each image is passed to enable or disable the proposed method. For example, a flag might be passed within the image header (PH).

[0436] The method is enabled for lower levels

[0437] In one embodiment, a high-level syntax flag at lower levels, for example, at the slice or block level, is passed to enable or disable the proposed method. For example, a flag can be passed within a slice header (SH).

[0438] For each criterion

[0439] In one embodiment, a flag for each criterion is transmitted. For example, the flag might indicate whether a specific restriction should be applied, and this might be transmitted by an encoder and obtained by the decoder from the bitstream. For example, a PH_Restriction_QT image header flag is transmitted. And the first criterion, as described earlier, is applied as follows: If (PH_Restriction_QT) { If (QTDepth < MinQTDepthTime - 1) { Only QT split allowed.

[0440] The advantage of all the high-level syntax signaling modes mentioned above is the coder's flexibility to determine the best compromise.

[0441] Encoding Options

[0442] If BTH and BTV are not allowed, then QT comes before BT.

[0443] In one embodiment, on the encoder side, when the test order of different divisions is determined, the QT variable before BT is set to true if BT Horizontal and BT Vertical are not allowed. Petition 870250088574, dated 09 / 30 / 2025, page 73 / 119 63 / 73

[0444] This provides an improvement in coding efficiency, with and without the proposed method, without increasing coding time. This is because the encoder optimizations are related to avoiding QT testing when BT is detected as better. Therefore, when TT is the only MTT division available, these optimizations are not efficient.

[0445] Possible divisions are not added to the list of disallowed divisions.

[0446] In one mode, unavailable splits are not added to the test list when they are unavailable. For example, when the QT variable before BT is true, the pseudocode is applied to populate the “listSplits” split list: if(No Split is allowed) add( No Split, listSplits) if(QT is allowed) add( QT, listSplits) if(BT Horizontal is allowed) add( BT Horizontal, listSplits) if(BT Vertical is allowed) add( BT Vertical, listSplits) if(TT Horizontal is allowed) add( TT Horizontal, listSplits) if(TT Vertical is allowed) add( TT Vertical, listSplits)

[0447] The similar pseudocode is used when the QT variable before BT is false.

[0448] As in the previous mode, this provides an improvement Petition 870250088574, dated 09 / 30 / 2025, page 74 / 119 64 / 73 in coding efficiency, with or without the proposed method in the modalities described above, without increasing coding time. This gain is related to some encoder optimizations that can avoid TT or QT evaluation according to certain conditions and if some divisions are previously in the list.

[0449] If this condition is true, QT will precede BT: the divisions will be handled in the following order:

[0450] -No Division

[0451] -QT

[0452] -BT Horizontal

[0453] -BT Vertical

[0454] -TT Horizontal

[0455] -TT Vertical

[0456] Otherwise, the order will be:

[0457] -No Division

[0458] -BT Horizontal

[0459] -BT Vertical

[0460] -TT Horizontal

[0461] -TT Vertical

[0462] -QT

[0463] Others

[0464] In one embodiment, other partitioning division modes may be applied and the method may be adapted accordingly.

[0465] All the modalities described can be combined, unless explicitly indicated otherwise. In fact, many combinations are synergistic and can produce efficiency gains greater than the sum of their parts.

[0466] Implementation of the Invention

[0467] Figure 17 shows a system 191, 195 comprising at least one encoder 150 or decoder 100 and a communication network 199, according to embodiments of the present invention. According to a Petition 870250088574, dated 09 / 30 / 2025, page 75 / 119 In mode 65 / 73, system 195 serves to process and provide content (e.g., video and audio content for display / output or streaming of video / audio content) to a user who has access to decoder 100, for example, through a user interface of a user terminal comprising the decoder 100 or a user terminal communicable with the decoder 100. Such a user terminal may be a computer, a mobile phone, a tablet, or any other type of device capable of providing / displaying the content (provided / transmitted) to the user. System 195 obtains / receives a bit stream 101 (in the form of a continuous stream or a signal - e.g., while previous video / audio is being displayed / emitted) through the communication network 199.According to one embodiment, the system 191, 195 is for processing content and storing the processed content, for example, video and audio content processed for display / broadcast / streaming at a later time. The system 191, 195 obtains / receives content comprising an original sequence of images 151, which is received and processed (including filtering with a deblocking filter according to the present invention) by the encoder 150, and the encoder 150 generates a bitstream 101 that must be communicated to the decoder 100 via a communication network 191.The bitstream 101 is then communicated to the decoder 100 in various ways, for example, it may be previously generated by the encoder 150 and stored as data in a storage device on the communication network 199 (e.g., on a server or cloud storage) until a user requests the content (i.e., the bitstream data) from the storage device, at which point the data is communicated / transmitted to the decoder 100 from the storage device. The system 191, 195 may also comprise a content provider device to provide / transmit to the user (e.g., communicating data to a user) an interface to be displayed on a user terminal, content information for the content stored in the storage device (e.g., the title of the content and other metadata / storage location data to identify, select and request the content) and. Petition 870250088574, dated 09 / 30 / 2025, pp. 76 / 119 66 / 73 to receive and process a user request for content, so that the requested content can be delivered / transmitted from the storage device to the user's terminal. Alternatively, the encoder 150 generates the bitstream 101 and communicates / transmits it directly to the decoder 100 as and when the user requests the content. The decoder 100 then receives the bitstream 101 (or a signal) and performs filtering with a deblocking filter according to the invention to obtain / generate a video signal 109 and / or audio signal, which is then used by a user terminal to provide the requested content to the user.

[0468] Any step of the method / process according to the invention or functions described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the steps / functions may be stored or transmitted by means of one or more instructions, code or program, or by a computer-readable medium, and executed by one or more hardware-based processing units, such as a programmable computing machine, which may be a PC (“Personal Computer”), a DSP (“Digital Signal Processor”), a circuit, a set of circuits, a processor and memory, a general-purpose microprocessor or central processing unit, a microcontroller, an ASIC (“Application-Specific Integrated Circuit”), a field-programmable logic array (FPGAs), or other equivalent set of integrated or discrete logic circuits.Consequently, the term "processor," as used herein, may refer to any of the structures above or to any other structure suitable for implementing the techniques described herein.

[0469] The embodiments of the present invention can also be implemented by a wide variety of devices or apparatus, including a wireless device, an integrated circuit (IC), or an array of ICs (e.g., a chip set). Various components, modules, or units are described herein to illustrate functional aspects of devices / apparatuses configured to perform these embodiments, but do not necessarily require implementation by different Petition 870250088574, dated 09 / 30 / 2025, p. 77 / 119 67 / 73 hardware units. Instead, multiple modules / units can be combined into a codec hardware unit or provided by a set of interoperable hardware units, including one or more processors in conjunction with appropriate software / firmware.

[0470] The embodiments of the present invention can be carried out by a computer of a system or apparatus that reads and executes computer executable instructions (e.g., one or more programs) stored in a storage medium to execute the modules / units / functions of one or more of the embodiments described above and / or that includes one or more processing units or circuits to execute the functions of one or more of the embodiments described above, and by a method performed by the computer of the system or apparatus, for example, by reading and executing the computer executable instructions from the storage medium to execute the functions of one or more of the embodiments described above and / or by controlling one or more processing units or circuits to execute the functions of one or more of the embodiments described above.The computer may include a network of separate computers or separate processing units for reading and executing computer-executable instructions. Computer-executable instructions may be supplied to the computer, for example, from a computer-readable medium, such as a network communication medium or a tangible storage medium. The communication medium may be a signal / bit stream / carrier wave. The tangible storage medium is a “non-transient computer-readable storage medium” which may include, for example, one or more of a hard disk, random access memory (RAM), read-only memory (ROM), distributed computing system storage, an optical disc (such as a compact disc (CD), a digital versatile disc (DVD) or a Blu-ray disc (BD)™), a flash memory device, a memory card and the like.At least some of the steps / functions can also be implemented in hardware by a dedicated machine or component, such as an FPGA (“Programmable Gate Array”). Petition 870250088574, dated 09 / 30 / 2025, pp. 78 / 119 68 / 73 Field”) or an ASIC (“Application-Specific Integrated Circuit”).

[0471] Figure 18 is a schematic block diagram of a 3600 computing device for implementing one or more embodiments of the invention. The 3600 computing device may be a device such as a microcomputer, a workstation, or a lightweight portable device.The computing device 3600 comprises a communication bus connected to: - a central processing unit (CPU) 3601, such as a microprocessor; - a random access memory (RAM) 3602 for storing the executable code of the embodiments of the invention, as well as registers adapted to record variables and parameters necessary to implement the method of encoding or decoding at least part of an image according to the embodiments of the invention, whose memory capacity can be expanded by an optional RAM connected to an expansion port, for example; - a read-only memory (ROM) 3603 for storing computer programs to implement the embodiments of the invention; - a network interface (NET) 3604 is typically connected to a communication network through which the digital data to be processed is transmitted or received.The network interface (NET) 3604 can be a single network interface or composed of a set of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Data packets are written to the network interface for transmission or read from the network interface for reception under the control of the software application running on the CPU 3601; - a user interface (UI) 3605 can be used to receive input from a user or display information to a user; - a hard disk (HD) 3606 can be provided as a mass storage device; - an input / output (IO) module 3607 can be used to receive / send data to / from external devices, such as a video source or monitor. The executable code can be stored in ROM 3603, on HD 3606, or on a removable digital medium, such as a disk.According to one variant, the executable code of the programs can be received through a communication network, via NET 3604, to be... Petition 870250088574, dated 09 / 30 / 2025, pp. 79 / 119 69 / 73 stored in one of the storage media of the communication device 3600, such as the HD 3606, before being executed. The CPU 3601 is adapted to control and direct the execution of instructions or parts of the software code of the program or programs according to the embodiments of the invention, these instructions being stored in one of the storage media mentioned above. After initialization, the CPU 3601 is capable of executing instructions from the main RAM memory 3602 relating to a software application, after these instructions have been loaded from the program ROM 3603 or the HD 3606, for example. Such a software application, when executed by the CPU 3601, performs the steps of the method according to the invention.

[0472] It is also understood that, according to another embodiment of the present invention, a decoder, according to the embodiment mentioned above, is provided in a user terminal, such as a computer, a mobile phone, a tablet or any other type of device (for example, a display device) capable of providing / displaying content to a user. According to another embodiment, an encoder according to an embodiment mentioned above is provided in an image capture device that also comprises a camera, a video camera or a network camera (for example, a closed-circuit television or a video surveillance camera) that captures and provides the content for the encoder to encode. Two examples are provided below with reference to Figures 19 and 20.

[0473] Figure 19 is a diagram illustrating a 3700 network camera system, including a 3702 network camera and a 202 client handset.

[0474] The 3702 network camera includes an image generation unit 3706, an encoding unit 3708, a communication unit 3710 and a control unit 3712.

[0475] Network camera 3702 and client device 202 are mutually connected so that they can communicate over network 200.

[0476] The 3706 imaging unit includes a lens and an image sensor (e.g., a charge-coupled device (CCD) or a semiconductor). Petition 870250088574, dated 09 / 30 / 2025, pp. 80 / 119 70 / 73 complementary metal oxide (CMOS)) and captures an image of an object and generates image data based on the image. This image can be a still image or a video image.

[0477] The 3708 encoding unit encodes image data using the encoding methods explained above or a combination of the encoding methods described above.

[0478] The communication unit 3710 of the network camera 3702 transmits the image data encoded by the encoding unit 3708 to the client device 202.

[0479] In addition, communication unit 3710 receives commands from client device 202. The commands include commands to set parameters for encoding unit 3708.

[0480] Control unit 3712 controls other units in network camera 3702 according to commands received by communication unit 3712.

[0481] The client device 202 includes a communication unit 3714, a decoding unit 3716 and a control unit 3718.

[0482] The communication unit 3714 of client device 202 transmits commands to network camera 3702.

[0483] In addition, communication unit 3714 of client device 202 receives the encoded image data from network camera 3712.

[0484] The 3716 decoding unit decodes the encoded image data using the decoding methods explained above, or a combination of the decoding methods explained above.

[0485] The control unit 3718 of client device 202 controls other units in client device 202 according to user operation or commands received by communication unit 3714.

[0486] The control unit 3718 of the client device 202 controls a display device 2120 to display an image decoded by the decoding unit 3716.

[0487] The control unit 3718 of client device 202 also controls a Petition 870250088574, dated 09 / 30 / 2025, pp. 81 / 119 71 / 73 display device 2120 to display the GUI (Graphical User Interface) to assign the parameter values ​​for the network camera 3702, including the parameters for encoding the encoding unit 3708.

[0488] The control unit 3718 of client device 202 also controls other units in client device 202 according to the user operation entered in the GUI displayed by display device 2120.

[0489] The control unit 3718 of client device 202 controls the communication unit 3714 of client device 202 to transmit commands to network camera 3702 that assign parameter values ​​to network camera 3702, according to the user operation entered in the GUI displayed by display device 2120.

[0490] Figure 20 is a diagram illustrating a 3800 smartphone.

[0491] The 3800 smartphone includes a communication unit 3802, a decoding unit 3804, a control unit 3806 and a display unit 3808.

[0492] Communication unit 3802 receives the encoded image data via network 200.

[0493] Decoding unit 3804 decodes the encoded image data received by communication unit 3802.

[0494] The 3804 decoding / encoding unit decodes / encodes the encoded image data using the decoding methods explained above.

[0495] The 3806 control unit controls other units in the 3800 smartphone according to a user operation or commands received by the 3806 communication unit.

[0496] For example, control unit 3806 controls a display unit 3808 to display an image decoded by decoding unit 3804. Smartphone 3800 can also comprise sensors 3812 and an image recording device 3810. In this way, smartphone 3800 can record images and encode them (using a method described above).

[0497] The 3800 smartphone can subsequently decode the images Petition 870250088574, dated 09 / 30 / 2025, page 82 / 119 72 / 73 encoded (using a method described above) and display them by means of display unit 3808 - or transmit the encoded images to another device by means of communication unit 3802 and network 200.

[0498] Alternatives and Modifications

[0499] Although the present invention has been described with reference to embodiments, it should be understood that the invention is not limited to the embodiments described. Those skilled in the art will understand that various alterations and modifications may be made without departing from the scope of the invention as defined in the appended claims. All features described in this descriptive report (including any claims, abstract and appended drawings) and / or all steps of any method or process so described may be combined in any combination, except combinations in which at least some of these features and / or steps are mutually exclusive. Each feature described in this descriptive report (including any claims, abstract and appended drawings) may be substituted for alternative features serving the same, equivalent or similar purpose, unless expressly stated otherwise.Thus, unless expressly stated otherwise, each feature described is only an example from a generic series of equivalent or similar features.

[0500] It is also understood that any result of comparison, determination, evaluation, selection, execution, realization or consideration described above, for example, a selection made during an encoding or filtering process, may be indicated or determinable / inferred from data in a bitstream, for example, a flag or data indicating the result, so that the indicated or determined / inferred result may be used in processing instead of actually performing the comparison, determination, evaluation, selection, execution, realization or consideration, for example, during a decoding process.

[0501] In claims, the word “includes” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality. The simple Petition 870250088574, dated 09 / 30 / 2025, page 83 / 119 73 / 73 The fact that different features are cited in mutually different dependent claims does not indicate that a combination of these features cannot be advantageously used.

[0502] The reference numbers appearing in the claims are merely illustrative and shall not have a limiting effect on the scope of the claims. Petition 870250088574, dated 09 / 30 / 2025, page 84 / 119

Claims

1 / 5 CLAIMS 1. A method for encoding or decoding image data into or from a bitstream, the bitstream including data indicating a partitioning of the image data into a plurality of blocks according to an encoding tree, wherein the blocks in the encoding tree can be divided according to a plurality of divisions, including a quadtree division, the method characterized in that it comprises: determining a value indicating a quadtree division depth of a current block; and allowing only one quadtree division of the current block among the plurality of divisions if the value indicating the quadtree division depth does not exceed a limit.

2. A method according to claim 1, characterized in that the plurality of splits additionally comprises a ternary tree split, a binary tree split, and no split.

3. A method for encoding or decoding image data into or from a bitstream, the bitstream including data indicating a partitioning of the image data into a plurality of blocks according to an encoding tree, wherein blocks in the encoding tree can be split according to a plurality of splits, including a binary tree split, a ternary tree split, a quadtree split and no split, the method characterized in that it comprises: determining a value indicating a quadtree split depth of a current block; and not allowing binary tree splitting based on one or more criteria, the one or more criteria including that the value indicating the quadtree split depth does not exceed a limit.

4. Method, according to claim 3, characterized in that one or more criteria include that a ternary tree split is not performed at a level above the current block level and, Petition 870250088574, dated 09 / 30 / 2025, pp. 114 / 119 2 / 5 if the ternary tree split has not been performed, not allowing the binary tree split and, if the ternary tree split has been performed, allowing the binary tree split for the current block.

5. A method for decoding a bit stream, characterized in that it comprises: executing a first method according to claim 1; and executing a second method according to claim 3, wherein the limit for allowing only quadtree splitting is a first limit and the limit for not allowing binary tree splitting is a second limit, wherein the first and second limits are different.

6. Method, according to claim 1, characterized in that the value indicating the depth of the quadtree split is a block size of the current block and crosses said limit(s) if it is greater than a predetermined block size.

7. Method according to claim 1, characterized in that the limit value is based on one or more parameters associated with a time area in one or more reference frames, a reference frame being a frame different from the frame containing the current block.

8. Method according to claim 7, characterized in that the area is an area that is co-located with the current block.

9. Method according to claim 7, characterized in that the area comprises a plurality of blocks in different positions.

10. Method according to claim 7, characterized in that the area is an area having a size larger than the current block.

11. Method according to claim 10, characterized in that the area has a size larger than the current block in a coding tree unit, CTU.

12. Method, according to claim 7, characterized in that the central position of the current block is used to determine the area in the reference frame or in a reference frame.

13. Method according to claim 7, characterized in that the area covers the entire area of ​​the reference frame or of a reference frame.

14. Method, according to claim 7, characterized in that the reference frame or a reference frame is a frame with the same temporal ID as the current frame that includes the current block.

15. Method, according to claim 14, characterized in that the frame with the same temporal ID is the closest frame with the same temporal ID.

16. Method according to claim 7, characterized in that the reference frame or a reference frame is a reference frame with the same quantization parameter as the current frame that includes the current block.

17. A method according to claim 7, characterized in that the reference frame or a reference frame is a frame that is used for predicting temporal motion vectors.

18. Method according to claim 7, characterized in that the reference frame or a reference frame is a frame that is the closest reference frame to the current frame that includes the current block.

19. Method according to claim 1, characterized in that a first area of ​​a first reference frame and a second area of ​​a second reference frame are used to obtain the limit value.

20. Method according to claim 7, characterized in that the limit value is based on an average quadtree depth determined from the temporal area.

21. Method, according to claim 7, characterized in that the limit value is based on a minimum quadtree depth determined from the temporal area.

22. Method, according to claim 7, characterized in that the limit value is determined based on a maximum multitree depth or an average multitree depth from the temporal area.

23. Method, according to claim 7, characterized in that the limit value is determined using the maximum or average depth of the encoding tree from the temporal area.

24. Method according to claim 1, characterized in that the limit value is based on a value transmitted in a header associated with the encoding tree unit of the current block.

25. A method according to claim 24, characterized in that the value transmitted in the header of the current encoding tree unit is used to predict the threshold value using one or more other values ​​transmitted in a header associated with another encoding tree unit.

26. Method according to claim 20, characterized in that the threshold value can be adapted based on the maximum multitree depth and the maximum multitree depth of the temporal area frame.

27. Method according to claim 20, characterized in that the threshold value can be adapted based on the maximum multitree depth and the maximum multitree depth of the temporal area.

28. Method according to claim 20, characterized in that the limit value can be adapted based on the minimum quadtree size of the temporal area frame or on the minimum, maximum or average quadtree size of the temporal area.

29. Method, according to claim 20, characterized in that the limit value can be adapted according to whether the current frame has a lower QP than the frame that includes the temporal area.

30. Method, according to claim 1, characterized in that the limit value is based on block size statistics in the time area.

31. Method, according to claim 1, characterized in that one or more steps of the method are not executed if image data related to a low-latency configuration is used or if at least one flag is transmitted in at least one image-related header. Petition 870250088574, dated 09 / 30 / 2025, pp. 117 / 119 5 / 5 32. Method according to claim 1, characterized in that one or more steps of the method are not performed when the image data are related to the display content and when one or more reference frames are an intra-frame.

33. A method according to claim 32, characterized in that the image data is related to the display content when one or more tools are enabled and / or signaled in the bitstream, wherein the one or more tools may include any of a Palette tool, Intra Block Copy (IBC), Transform Skip, and other tool associated with the display content.

34. Device for decoding image data from a bitstream, characterized in that it is configured to perform the method as defined in any one of claims 1 to 33.

35. Device for encoding image data into a bitstream, characterized in that it is configured to execute the method as defined in any one of claims 1 to 33.

36. Non-transient computer-readable medium, characterized in that it comprises instructions which, when executed by a computer, cause the computer to perform the method as defined in any of claims 1 to 33. Petition 870250088574, dated 09 / 30 / 2025, pp. 118 / 119