Method, apparatus, and computer program product for encoding and decoding digital media content

By deriving and applying scale, shift, and rounding parameters to approximate division operations, the method addresses inefficiencies in video encoding and decoding, reducing computational load and enhancing efficiency.

JP2025518355APending Publication Date: 2025-06-12NOKIA TECHNOLOGIES OY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024571889
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-07
Filing Date
2023-03-21
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing video encoding and decoding systems face inefficiencies in performing division operations, which are complex and require additional logic circuits or processing cycles, especially in hardware-based implementations.

Method used

The method involves deriving scale, shift, and rounding parameters to approximate the result of a division operation using simpler operations like multiplication, addition, and bit-shift, allowing for the use of an undersampled lookup table to determine an approximated output.

Benefits of technology

This approach significantly reduces the computational load and memory requirements for video encoders and decoders, while maintaining high accuracy in approximating division operations, thereby enhancing encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025518355000001_ABST
    Figure 2025518355000001_ABST
Patent Text Reader

Abstract

Embodiments relate to a method for encoding / decoding, the method comprising encoding / decoding a picture comprising a number of samples (910), wherein the encoding / decoding phase comprises a division operation, determining a numerator and a denominator (920), determining an approximated output of the division operation between the numerator and the denominator (930), which comprises deriving a scale parameter using a piecewise approximation, deriving a shift parameter and a rounding parameter, applying the scale parameter, the shift parameter, and the rounding parameter to the numerator, and using the approximated output of the division operation in the encoding / decoding phase (940). Embodiments also relate to an apparatus and a computer program product for performing this method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This solution generally relates to the encoding and decoding of digital media content such as video or still image data.

Background Art

[0002] This section is intended to explain the background or context of the present invention as described in the claims. The descriptions herein may include concepts that could be pursued, but are not necessarily concepts that have been previously devised or pursued. Thus, unless otherwise indicated herein, the content described in this section is not prior art to the description and claims in this application, and is not admitted to be prior art by virtue of its inclusion in this section.

[0003] A video encoding system may be composed of an encoder that converts an input video into a compressed representation suitable for storage / transmission, and a decoder that can restore the compressed video representation back into a viewable form. The encoder may discard some information within the original video sequence in order to represent the video in a more compact form, for example, to enable the storage / transmission of video information at a bit rate lower than that which might otherwise be required.

Summary of the Invention

[0004] The scope of protection sought for various embodiments of the present invention is defined by the independent claims. If there are embodiments and features described herein that fall outside the scope of the independent claims, they should be construed as useful examples for understanding the various embodiments of the present invention.

[0005] The various aspects include a method, an apparatus, and a computer-readable medium storing a computer program, characterized by the description of the independent claims. The various embodiments are disclosed in the dependent claims.

[0006] According to a first aspect, means for encoding / decoding a picture including a plurality of samples, the encoding / decoding phase including a division operation, means for encoding / decoding, means for determining a numerator and a denominator, and means for determining an approximated output of the division operation between the numerator and the denominator, ○ means for deriving a scale parameter using piecewise approximation; ○ means for deriving a shift parameter and a rounding parameter; ○ means for applying the scale parameter, the shift parameter, and the rounding parameter to the numerator; A device is provided, comprising means for determining and means for using the approximated output of the division operation in the encoding / decoding phase.

[0007] According to a second aspect, encoding / decoding a picture including a plurality of samples, the encoding / decoding phase including a division operation, encoding / decoding, determining a numerator and a denominator, and determining an approximated output of the division operation between the numerator and the denominator, ○ deriving a scale parameter using piecewise approximation; ○ deriving a shift parameter and a rounding parameter; ○ applying the scale parameter, the shift parameter, and the rounding parameter to the numerator; A method is provided, including determining and using the approximated output of the division operation in the encoding / decoding phase.

[0008] According to a third aspect, there is provided an apparatus comprising at least one processor and a memory including computer program code, wherein the memory and the computer program code are configured to cause the at least one processor to cause the apparatus to, at least, encode / decode a picture including several samples, the encoding / decoding phase including a division operation, determine a numerator and a denominator, determine an approximated output of the division operation between the numerator and the denominator, and the apparatus further ○ derive a scale parameter using piecewise approximation, ○ derive a shift parameter and a rounding parameter, ○ apply the scale parameter, the shift parameter, and the rounding parameter to the numerator, and use the approximated output of the division operation in the encoding / decoding phase.

[0009] According to a fourth aspect, there is provided a computer program product including computer program code, wherein the computer program code is configured to cause an apparatus or a system to, when executed on at least one processor, encode / decode a picture including several samples, the encoding / decoding phase including a division operation, determine a numerator and a denominator, determine an approximated output of the division operation between the numerator and the denominator, and the apparatus further ○ derive a scale parameter using piecewise approximation, ○ derive a shift parameter and a rounding parameter, ○ apply the scale parameter, the shift parameter, and the rounding parameter to the numerator, and use the approximated output of the division operation in the encoding / decoding phase.

[0010] According to one embodiment, the encoding / decoding phase includes predicting at least one sample of the picture.

[0011] According to one embodiment, the encoding / decoding phase includes filtering.

[0012] According to one embodiment, an additional shift parameter is applied to the molecule.

[0013] According to one embodiment, it is determined whether the bit precisions of the numerator, denominator, scale parameter, and output parameter are different. If they are different, the bit precisions of the output and scale parameter are used when determining the approximated output.

[0014] According to one embodiment, the scale parameter is determined using one of interpolation operations, polynomial processing, and linear interpolation.

[0015] According to one embodiment, the scale parameter is determined using an interpolation operation between at least two values determined based on a table look-up.

[0016] According to an embodiment, the scale parameter is adjusted by a value for reducing the range of the denominator.

[0017] According to one embodiment, a computer program product is embodied on a non-transitory computer-readable medium.

[0018] Hereinafter, various embodiments will be described in more detail with reference to the accompanying drawings.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7a

Figure 7b

Figure 8

Figure 9

Figure 10

[0020] In the following, several embodiments will be described in the context of one video coding configuration. However, it should be noted that the embodiments of the present invention are not necessarily limited to this specific configuration. The embodiments discussed herein relate to intra prediction in video or still image coding using sparse linear component regression, which can even be signal processing other than image / video compression.

[0021] The following description and drawings are illustrative and should not be construed as limiting unnecessarily. Specific details are provided for a thorough understanding of the present disclosure. However, in some cases, well-known details or conventional details are not described so as not to obscure the description. A reference to one embodiment or an embodiment in the present disclosure may refer to the same embodiment, but not necessarily so, and such a reference means at least one of the embodiments.

[0022] References to "one embodiment" or "an embodiment" in this specification mean that the particular features, structures, or characteristics described in connection with that embodiment are included in at least one embodiment of the present disclosure.

[0023] A video codec consists of an encoder and a decoder. The encoder is configured to convert an input video into a compressed representation suitable for storage / transmission. The decoder can restore the compressed video representation back to a viewable form. The encoder may discard some information within the original video sequence in order to represent the video in a more compact form, such as a lower bitrate.

[0024] In most cases, the basic unit of each of the input to the encoder and the output of the decoder is a picture. The picture given as input to the encoder may also be called a source picture, and the picture decoded by the decoder may also be called a decoded picture or a reconstructed picture.

[0025] A source picture and a decoded picture are each composed of one or more sample arrays, for example, composed of one of the following sets of sample arrays. - Luma (Y) only (monochrome) - Luma and two chromas (YCbCr or YCgCo) - Green, blue, and red (GBR, also known as RGB) - Arrays representing other undefined monochrome or trichromatic sampling (e.g., YZX, also known as XYZ)

[0026] A picture can be defined as either a frame or a field. A frame is composed of a matrix of luma samples and, optionally, corresponding chroma samples. A field is a set of sample lines taken every other line of a frame and can be used as an encoder input when the source signal is interlaced. There may be no chroma sample array (so monochrome sampling can be used), or the chroma sample array may be subsampled compared to the luma sample array.

[0027] A bitstream can be defined as a sequence of bits that forms a representation of encoded pictures and associated data that form one or more encoded video sequences, which in some encoding formats or standards can be in the form of a network abstraction layer (NAL) unit stream or a byte stream. After a first bitstream, a second bitstream can follow within the same logical channel, such as within the same file or within the same connection of a communication protocol. A base stream (in the context of video encoding) can be defined as a sequence of one or more bitstreams. In some encoding formats or standards, the end of the first bitstream can be indicated by a specific NAL unit, which can also be called an end of bitstream (EOB) NAL unit and is the last NAL unit of the bitstream.

[0028] The phrase "along a bitstream" (e.g., a phrase indicating along a bitstream) or the phrase "along an encoded unit of a bitstream" (e.g., a phrase indicating along an encoded tile) may be used in the claims and the described embodiments to refer to transmission, signaling, or storage in such a way that "out-of-band" data is associated with, but not included in, the bitstream or the encoded unit. Phrases such as decoding along a bitstream or decoding along an encoded unit of a bitstream may refer to decoding the referenced out-of-band data associated with the bitstream or the encoded unit (which may be obtained from out-of-band transmission, signaling, or storage). For example, the phrase "along a bitstream" may be used when the bitstream is included in a container file such as a file conforming to the ISO base media file format, and specific file metadata is stored in the file in a way that associates metadata with the bitstream, such as in a box within a sample entry of a track containing the bitstream, a sample group of a track containing the bitstream, or a time-specified metadata track associated with a track containing the bitstream.

[0029] Hybrid video coders, such as ITU-T H.263 and H.264, can encode video information in two phases. First, pixel values within a particular picture region (or “block”) are predicted, for example, by motion compensation means (finding and indicating an area within one of the encoded video frames that closely corresponds to the block being encoded) or spatial means (using pixel values surrounding the block to be encoded in a specified way). In the first phase, predictive encoding can be applied, for example, as so-called sample prediction and / or so-called syntax prediction. In sample prediction, pixel values or sample values within a particular picture region or “block” are predicted. These pixel values or sample values can be predicted using, for example, one or more of motion compensation or intra prediction mechanisms. Second, the prediction error, that is, the difference between the predicted pixel block and the original pixel block, is encoded. This can be performed by using a specified transform (such as a discrete cosine transform (DCT) or a variant thereof) to transform the difference in pixel values, quantizing the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the coder can control the balance between the accuracy of the pixel representation (image quality) and the size of the resulting encoded video representation (file size or transmission bitrate).

[0030] An example of the encoding process is shown in FIG. 1. FIG. 1 shows the image to be encoded (I n ), the predicted representation of the image block (P’ n ), the prediction error signal (D n ), the reconstructed prediction error signal (D’ n ), the pre-reconstructed image (I’ n ), the final reconstructed image (R’ n ), the transform (T) and inverse transform (T -1 ), the quantization (Q) and inverse quantization (Q -1 ), the entropy encoding (E), the reference frame memory (RFM:reference frame memory), the inter prediction (Pinter ) Intra prediction (P intra ) mode selection (MS), and filtering (F).

[0031] In some video codecs such as H.265 / HEVC, a video picture is divided into coding units (CUs) that cover the area of the picture. A CU includes one or more prediction units (PUs) that define the prediction process for samples within the CU and one or more transform units (TUs) that define the prediction error coding process for samples within the CU. A CU can be composed of a square sample block having a size selectable from a set of predefined possible CU sizes. A CU having the maximum allowable size can be referred to as the largest coding unit (LCU) or coding tree unit (CTU), and a video picture can be divided into non-overlapping CTUs. A CTU can be further divided into a combination of smaller CUs, for example, by recursively dividing the CTU and the resulting CUs. Each resulting CU can be associated with at least one PU and at least one TU. Each PU and TU can be further divided into smaller PUs and TUs to increase the granularity of the prediction process and the prediction error coding process, respectively. Prediction information is associated with each PU, which defines what prediction is applied to the pixels within that PU (e.g., motion vector information for an inter-predicted PU and intra-prediction direction information for an intra-predicted PU). Similarly, information (e.g., including DCT coefficient information) that describes the prediction error decoding process for samples within that TU is associated with each TU. Whether prediction error coding is applied to each CU can be signaled at the CU level. If there is no prediction error residue associated with a CU, it can be considered that there is no Tu for that CU. The division of an image into CUs and the division of a CU into PUs and TUs can be signaled in a bitstream, enabling a decoder to reproduce the intended structure of these units.

[0032] The decoder can reconstruct the output video by applying prediction means similar to those of the encoder to form a predicted representation of the pixel block (using motion information or spatial information created by the encoder and stored in the compressed representation), and applying prediction error decoding (the inverse operation of prediction error encoding that restores the quantized prediction error signal to the spatial pixel domain). After applying the prediction means and the prediction error decoding means, the decoder is configured to sum the predicted signal and the prediction error signal (pixel values) to form an output video frame. The decoder (and the encoder) can also apply additional filtering means to improve the quality of the output video and then pass the output video for display and / or store it as a prediction reference for future frames in the video sequence. An example of the decoding process is shown in FIG. 2. FIG. 2 shows the predicted representation (P’ n ) of the image block, the reconstructed prediction error signal (D’ n ), the pre-reconstructed image (I’ n ), the final reconstructed image (R’ n ), the inverse transform (T -1 ), the inverse quantization (Q -1 ), the entropy decoding (E -1 ), the reference frame memory (RFM), the prediction (inter or intra) (P), and the filtering (F).

[0033] Instead of, or in addition to, an approach that uses sample value prediction and transform coding to represent coded sample values, color palette-based coding can be used. Palette-based coding refers to a group of approaches in which a palette, i.e., a set of colors and associated indices, is defined and the value of each sample within an encoding unit is represented by indicating an index within the palette. Palette-based coding can achieve good coding efficiency in encoding units with a relatively small number of colors (e.g., an image region representing computer screen content such as text or simple graphics). To improve the coding efficiency of palette coding, various types of palette index prediction approaches can be utilized, or the palette indices can be run-length coded to enable more efficient representation of larger homogeneous regions. Also, when a CU contains sample values that do not repeatedly appear within the CU, escape coding can be used. Escape-coded samples are transmitted without reference to a palette index. Instead, the value can be individually indicated for each escape-coded sample.

[0034] When a CU is coded in palette mode, various prediction strategies are used to exploit the correlation between pixels within the CU. For example, the mode information can be signaled for each row or pixel and indicates any of the following. The mode can be a horizontal mode, which means that a single palette index is signaled and the entire pixel line shares this index. The mode can be a vertical mode, where the entire pixel line is the same as the line above and no further information is signaled. The mode can be a normal mode, where for each pixel position, a flag is signaled indicating whether it is the same as one of the pixels to the left and other pixels, and if not, the color index itself is transmitted separately.

[0035] In a video codec, motion information can be represented using motion vectors associated with each motion-compensated image block. Each of these motion vectors can represent the displacement between an image block in the picture to be coded (encoder side) or decoded (decoder side) and a prediction source block in one of the coded or decoded pictures. To efficiently represent the motion vectors, they can be coded in difference with respect to the predicted motion vectors for each block. In a video codec, the predicted motion vectors can be created in a predefined way, such as by calculating the median value of the coded or decoded motion vectors of adjacent blocks. Other ways to create a motion vector prediction are to generate a list of prediction candidates from adjacent blocks and / or co-located blocks in the temporal reference pictures, and signal the selected candidate as the motion vector predictor. In addition to predicting the value of the motion vector, the reference index of the coded / decoded picture can also be predicted. The reference index can be predicted from adjacent blocks and / or co-located blocks in the temporal reference pictures. Furthermore, high efficiency video coding can employ an additional motion information coding / decoding mechanism often called the merging / merge mode, in which all motion field information including the motion vector and the corresponding reference picture index in each available reference picture list is predicted and used without change. Similarly, the prediction of the motion field information can be performed using the motion field information of adjacent blocks and / or co-located blocks in the temporal reference pictures, and the motion field information used can be signaled from a list of lists of motion field candidates filled with the motion field information of available adjacent / co-located blocks.

[0036] The video codec may support motion compensation prediction from one source picture (unidirectional prediction) and two sources (bidirectional prediction). In the case of unidirectional prediction, a single motion vector may be applied. In the case of bidirectional prediction, two motion vectors may be signaled, and the motion compensation predictions from the two sources may be averaged to create a final sample prediction. In the case of weighted prediction, the relative weights of the two predictions may be adjusted, or a signaled offset may be added to the prediction signal.

[0037] In addition to applying motion compensation for inter-picture prediction, a similar approach can be applied to intra-picture prediction. In this case, the displacement vector indicates where a block of samples can be copied from the same picture to form the prediction of the block to be encoded or decoded. This type of intra-block copy method can significantly improve the encoding efficiency when there are repeating structures such as text or other graphics within the frame.

[0038] In a video codec, the prediction residual after motion compensation or intra prediction can first be transformed using a transform kernel (such as DCT, Discrete Cosine Transform), and then encoded. This is because there is often still some correlation in the residual, and the transform can reduce this correlation and often achieve more efficient encoding.

[0039] The video encoder can use a Lagrangian cost function to find an optimal encoding mode, such as a desired macroblock mode and related motion vectors. This type of cost function uses a weighting factor λ to combine the (exact or estimated) image distortion due to the irreversible encoding method and the (exact or estimated) amount of information required to represent the pixel values within the image region. C = D + λR (Equation 1) Here, C is the Lagrangian cost to be minimized, D is the image distortion (e.g., mean squared error) at the mode and motion vector under consideration, and R is the number of bits required to represent the data necessary for the decoder to reconstruct the image block (including the amount of data representing the motion vector candidates).

[0040] Scalable video coding refers to an encoding structure in which a single bitstream can include multiple representations of content at different bitrates, resolutions, or frame rates. In these cases, the receiver can extract the desired representation according to its characteristics (e.g., the resolution most suitable for the display device). Alternatively, a server or network element can extract a portion of the bitstream transmitted to the receiver according to, for example, network characteristics or the processing capabilities of the receiver. A scalable bitstream can be composed of a "base layer" that provides the lowest available quality video and one or more enhancement layers that improve the video quality when received and decoded together with the lower layer. To improve the encoding efficiency of the enhancement layer, the encoded representation of the enhancement layer can depend on the lower layer. For example, the motion information and mode information of the enhancement layer can be predicted from the lower layer. Similarly, the pixel data of the lower layer can be used to create predictions for the enhancement layer.

[0041] A scalable video codec for quality scalability (also known as signal-to-noise ratio or SNR) and / or spatial scalability can be implemented as follows. For the base layer, a conventional non-scalable video encoder and decoder are used. The reconstructed / decoded pictures of the base layer are included in the reference picture buffer for the enhancement layer. In H.264 / AVC, HEVC, and similar codecs that use reference picture lists for inter prediction, the decoded pictures of the base layer can be inserted into the reference picture list for encoding / decoding of the enhancement layer pictures, similar to the decoded reference pictures of the enhancement layer. As a result, the encoder can select the base layer reference picture as an inter prediction reference and indicate its use in the encoded bitstream using a reference picture index. The decoder decodes from the bitstream, for example from the reference picture index, that the base layer picture is being used as an inter prediction reference for the enhancement layer. The decoded base layer picture is called an inter-layer reference picture when it is used as a prediction reference for the enhancement layer.

[0042] In addition to quality scalability, the following scalability modes exist. - Spatial scalability: The base layer pictures are encoded at a lower resolution than the enhancement layer pictures. - Bit-depth scalability: The base layer pictures are encoded at a lower bit depth (e.g., 8 bits) than the enhancement layer pictures (e.g., 10 bits or 12 bits). - Chroma format scalability: The enhancement layer pictures provide a higher chroma fidelity than the base layer pictures (e.g., encoded in 4:4:4 chroma format compared to 4:2:0 format for the base layer).

[0043] In the cases of the above-mentioned scalability, by encoding the enhancement layer using the base layer information, the additional bitrate overhead can be minimized.

[0044] Scalability can be achieved in two ways: a) introducing a new coding mode for performing prediction of pixel values or syntax from the lower layer of the scalable representation, or b) putting the picture of the lower layer into the reference picture buffer (decoded picture buffer, DPB) of the upper layer. Approach a) is more flexible and can generally provide better coding efficiency. However, approach b), i.e., reference frame-based scalability, can be implemented very efficiently while still realizing most of the benefits of the resulting coding efficiency with minimal changes to the single-layer codec. The reference frame-based scalability codec can be implemented by using the same hardware or software implementation for all layers and only processing the DPB management by external means.

[0045] To enable the use of parallel processing, the image can be divided into independently encodable and decodable image segments (slices or tiles). A slice can refer to an image segment composed of a certain number of basic coding units processed in the default coding or decoding order, and a tile can refer to an image segment defined as a rectangular image region that is processed as at least somewhat individual frames.

[0046] Video can be encoded in the YUV or YCbCr color space, which is considered to reflect the characteristics of the human visual system. These color spaces allow for the use of lower-quality representations in the Cb and Cr channels because human perception is not very sensitive to the fidelity of the color differences represented by these channels.

[0047] Various parts of a video codec may require the use of division operations or approximations of such operations. For example, in the Versatile Video Coding (VVC / H.266) standard, a Cross-Component Linear Model (CCLM) is used as a linear model for predicting samples of chroma channels (e.g., Cb and Cr) based on reconstructed luma samples. The prediction model of the CCLM is pred C (i,j)=α·rec L ’(i,j)+β and can be, where pred C (i,j) represents the predicted chroma sample within the coding unit, and rec L ’(i,j) represents the downsampled reconstructed luma sample of the same coding unit.

[0048] The parameters α and β are derived as follows using up to four adjacent chroma samples and their corresponding downsampled luma samples.

Equation

[0049] Thus, when calculating the CCLM parameters, a division operation is usually required. The division operation can be implemented using a lookup table.

[0050] On the other hand, in a cross-component convolutional model (CCCM), a division operation is usually also required when filtering the prediction coefficients.

[0051] Division operations are relatively complex operations to perform. In a hardware-based video codec implementation, additional logic circuits are required, and in a software-based implementation operating on a general-purpose processor, additional cycles are required to execute division operations.

[0052] For cross-component linear model prediction (CCLM) used in VVC / H.266 and ECM3.0 video codecs, etc., the division operations required use a table-lookup-based approach to approximate the division operations. In this approach, a lookup table with 2^4 = 16 entries is used to derive a 4-bit significant value that approximates the reciprocal of the input with an accuracy of 1 / 2^4.

[0053] Embodiments of the present invention aim to determine scale values and shift values that can be used to approximate division operations in video encoding applications, etc. This method involves decoupling the accuracy and size of the entries in the lookup table used to generate an approximate reciprocal of the input data. As a result, when using a lookup table to derive an M-bit mantissa value, the size of the lookup table is set to be smaller than 2^M entries. The mantissa values of the intermediate entries are not included in the table and are derived by rounding to the existing entries in the table or by interpolation using a set of existing entries in the lookup table. Both rounding to existing entries and interpolation between existing entries provide a piecewise representation of the relationship between the input and output scale values.

[0054] An alternative way to derive the value of an entry is to use a polynomial (e.g., power of 2) function. Specifically, it has been found that when the form of the polynomial function and its parameters are carefully chosen, by implementing the polynomial function using integer arithmetic with low-bit precision, accurate results with minimal impact of parameter quantization errors are produced. This enables a video encoder or decoder to generate a highly accurate estimate of the result of a division operation while minimizing the computational load and the memory requirements of the required look-up tables. The performance can be further improved by creating a segmented model where different segments of the input value use different approximations and each segment has its own polynomial function.

[0055] A method operating in accordance with the present invention uses an undersampled look-up table to determine an approximated output of a division operation. That is, the aim is to calculate an output value that approximates the result of a division operation between a numerator nom and a denominator denom: output≒nom / denom using simpler operations such as multiplication, addition, and bit-shift operations. Thus, embodiments of the present invention provide an inverse function without division. Generally, the output is approximated by deriving scale, round, and shift parameters, and the output can be calculated as follows. output=(nom*scale+round)>>shift

[0056] If the inputs (nom and denom) and output of the operation use fixed-point arithmetic, there can be an additional shift of the nom value to align the fractional bits of the input and output. output=((nom<<shiftDecim)*scale+round)>>shift

[0057] In such a case, the value of shiftDecim can be set as a function of the number of bits used to represent the fractional part of a number. The value of shiftDecim can also depend on other parameters such as the precision of the scale value used in the above equation.

[0058] The rounding parameter round in the equation can be used to improve the calculation precision and can be set, for example, as follows. round = 1< <shift>>1. Or round = 1 << (shift - 1), or round = 0

[0059] Generally, when the nom, denom, scale, and output parameters have different bit precisions, the output value can be calculated using the following formula, where N_output and N_scale are the bit precision or number of bits representing the fractional part of the numerical values of the output parameter and the scale parameter, respectively. Also, the rounding parameter is calculated using (shift + N_scale). output = ((nom << (N_denom + N_output)) * scale + round) >> (shift + N_scale + N_nom) round = 1 << (shift + N_scale + N_nom - 1)

[0060] An alternative way to calculate the output is as follows. delta_shift = (shift + N_scale + N_nom) - (N_denom + N_output) output = (delta_shift > 0)? (nom * scale + round) >> (delta_shift) : (nom * scale + round) << (delta_shift) round = (delta_shift > 0)? 1 << (delta_shift - 1) : 0

[0061] Setting the scale to zero can of course be omitted from the formula, and while the resulting precision decreases somewhat, the number of operations required to calculate the output value is also somewhat reduced.

[0062] The shift parameter shift can be set in various ways. For example, it can be set equal to the base-2 logarithm of the denominator denom. The base-2 logarithm can be set to be rounded down to the nearest integer value. Alternatively, it can be rounded to the nearest integer value or rounded up to the nearest integer value. Alternatively, the shift can be set to include terms related to the precision of the scale parameter. For example, if the scale parameter is represented as an M-bit integer, the shift parameter can be set to cancel it out by adding M to the base-2 logarithm of the denominator. shift = floorLog2(denom), or shift = floorLog2(denom) + M Here, floorLog2 is a function that returns the nearest integer less than or equal to the base-2 logarithm of the input parameter denom.

[0063] The scale value can be determined in various ways. One alternative is to use a look-up table that represents values from 1 to 1 / 2 with an accuracy of M bits (or N_scale bits). That is, for a selected M, a table where the value 2^M represents 1 and 2^(M - 1) represents 1 / 2. Since values in such a range vary between 2^(M - 1) and 2^M, the values in the table can be subtracted by 2^(M - 1) to reduce the storage bit depth required for the table. In that case, when reading a value from the table, 2^(M - 1) can be added back to the read value to restore the full-range value. Such addition can also be performed by a bitwise "or" operation between the value read from the table and the value of 2^(M - 1).

[0064] The size of the table is selected to be smaller than 2^M so that it can operate with a small amount of memory compared to a full-size table with 2^M entries. For example, a table of size 2^N or (2^N)+1 can be selected, where N<M. The difference D between the table entry precision parameter M and the table size parameter N represents the subsampling ratio between the "full" 2^M table typically required for an M-bit scale value and the actual table size 2^N. Setting: D = M - N The subsampling ratio 2^(M - N) between the tables can be given as 2^D.

[0065] Entries within such a lookup table can be determined, for example, using the following pseudocode. D = M - N d = 1 << M index = 0 while d <= (2 << M) { table[index] = round(double(1 << (2 * M)) / double(d)) index += 1 d += 1 << D }

[0066] Here, << is used to represent a bitwise left shift operation, and x += y is used to represent adding the value of y to the value of x. double(x) is used to represent casting an integer to a double-precision floating-point number, and round(x) represents rounding to the nearest integer value.

[0067] As a result, the lookup table T can have values (for M = 14 and N = 8, for example) as shown in Figure 3.

[0068] In other examples of this case, the values M = 12 and N = 8 are selected, and when selected, the lookup table T will be as shown in Figure 4.

[0069] As a further example, when M = 14 and N = 6 are selected, the lookup table T can be set as shown in, for example, FIG. 5.

[0070] As a further example, when M = 14 and N = 2 are selected, the lookup table T can be set as shown in, for example, FIG. 6.

[0071] It is understood that different parameters of M and N as well as different algorithms can be used to generate such a lookup table. For example, different rounding approaches can be used, or different calculation precisions can be used to determine the values in the lookup table.

[0072] An alternative way to calculate the entries in the lookup table is to use linear regression. There are 2^D - 1 scale values missing between two entries of the lookup table. In other words, there are 2^D - 1 scale values missing before and after each entry. Therefore, a linear regression model can be calculated for each side of each entry, and the value of that entry point can be evaluated based on those two regression models to calculate two values. And the value of that entry can be set to the average of these two values. For example, when M = 14 and N = 2, there are four segments, and the regression models of the first and second segments are shown in FIGS. 7a and 7b respectively. The values of the second entry using the first and second regression models are evaluated to be 12996.72 (i.e., -0.7964 * 20480 + 29306.85) and 13039.67 (i.e., -0.5317 * 20480 + 23929.43) respectively. And the final value of that entry can be the average of these two values, i.e., (12996.72 + 12996.72) / 2 = 13018.12, which can be rounded down to 13018. This value is slightly lower than the value created by the aforementioned method (i.e., 13107). As a result of this method, the lookup table T can be set as follows, for example. T[2^2 + 1]= { 16259, 13018, 10872, 9331, 8168 };

[0073] The table constructed in this way represents scale values subsampled by a factor of 2^D, and the missing 2^D - 1 scale values between the scale values included in the table can be determined using interpolation. Alternatively, to reduce complexity at the expense of output accuracy, the denominator can be mapped to the nearest available entry in the table or truncated to map to the nearest available entry in the table. Using linear 2 - tap interpolation, the approximate output of the division operation between the numerator nom and the denominator denom can be determined, for example, as follows. / / Stage 1 shift0 = floorLog2(denom) round0 = 1 < <shift0>>1 normDiff = (((denom << M) + round0) >> shift0) & ((1 << M) - 1); / / Stage 2 D = M - N diffFull = normDiff >> D diffFrac = normDiff & ((1 << D) - 1) tapRound = 1 << <d>>1 scale0 = T[diffFull] scale1 = T[diffFull + 1] scale = scale0 + ((diffFrac * (scale1 - scale0) + tapRound) >> D) / / Stage 3 output = ((nom << (FP_BITS - M)) * scale + round0) >> shift0

[0074] In this case, the sequence of operations can be considered to have three stages as described above. In stage 1, the shift0 and normDiff parameters are determined. This represents the tabulation of the input denominator denom into the data range. The normDiff parameter is set to map the denominator denom to the precision of the table entry precision parameter M. In stage 2, normDiff is split into the full entry parameter diffFull and the fractional entry parameter diffFrac. The full entry parameter diffFull is used to identify an entry in table T, and the fractional entry parameter diffFrac is used to determine the final scale parameter based on the entry read from table T or a parameter generated based on the entry read from table T. The parameters scale0 and scale1 are determined using table T at the entry positions diffFull and diffFull + 1 respectively. And the final scale parameter scale is calculated using linear interpolation based on scale0, scale1, and the fractional entry part diffFrac. In the last stage 3, the output is calculated based on the scale, rounding, and shift parameters. When using a fixed-point representation where the number of fractional bits is given as FP_BITS, the numerator nom can be adjusted here by left-shifting by (FP_BITS - M) before multiplying by the scale. Alternatively, the final shift value can be set to shift0 + M, and the final rounding parameter can be calculated based on the new final shift. In this type of alternative, the output of stage 3 can be determined as follows. / / Stage 3, alternative 1 shift = shift0 + M round = 1< <shift>>1 output = ((nom << FP_BITS) * scale + round) >> shift

[0075] Alternatively, if fixed-point representation is not used in the calculation, it can be further simplified as follows. / / Stage 3, alternative 2 shift = shift0 + M round = 1 << <shift>>1 output=(nom*scale+round)>>shift

[0076] In all of these examples, the rounding can of course be set to be done in different ways. For example, the parameter round or round0 can be set to 0, or shift or other values derived using other parameters can be set.

[0077] The interpolation in stage 2 can also be performed in other ways. For example, two or more entries from table T can be included in the calculation using different coefficients based on the value of diffFrac. Also, for example, as follows, a weighted sum of the entries read from table T or the parameters derived from the values of table T can be calculated and performed. / / Stage 2, alternative example 1 D = M - N diffFull = normDiff >> D diffFrac = normDiff & ((1 << D) - 1) tapSum = 1 << D tapRound = 1 < <d>>1 scale0 = T[diffFull] scale1 = T[diffFull + 1] w0 = tapSum - diffFrac w1 = diffFrac scale = (w0 * scale0 + w1 * scale1 + tapRound) >> D

[0078] Instead of, or in addition to, the interpolation operation used in stage 2 to improve the scale value, alternative approaches such as a method based on the derivative of the scale function can be used.

[0079] As a further simplified approach, the interpolation performed in stage 2 can be omitted. In this case, the table size parameter N can be advantageously used during the calculation of the normDiff parameter in stage 1, and the table entry precision parameter M can only be used when determining the final shift and optional rounding parameters. This example can be expressed in pseudocode as follows. shift0 = floorLog2(denom) round0 = 1 < <shift0>>1 normDiff = (((denom << N) + round0) >> shift0) & ((1 << N) - 1) scale = T[normDiff] shift = shift0 + M round = 1< <shift>>1 output=(nom*scale+round)>>shift

[0080] Also, when operating using a fixed-point number with the fractional bits given as FP_BITS, the final shift can be made equal to floorLog2(denom), and the effect of the decimal point can be handled as part of the scaling of the numerator nom as follows. shift=floorLog2(denom) round=1< <shift>>1 normDiff = (((denom << N) + round) >> shift) & ((1 << N) - 1) scale = T[normDiff] output = ((nom << (FP_BITS - M) * scale + round) >> shift

[0081] If the same division operation using the same denominator needs to be performed on several different numerators, the scale value and the shift value can be calculated once and reused when performing the division operation. In this way, most of the necessary operations are related to the calculation of the scale value and the shift value, and the actual operations for calculating the output only require multiplication by the scale, shifting by shift, and optional additional shifting and rounding, so the number of calculations can be significantly reduced.

[0082] An alternative way to calculate the scale is to use piecewise polynomial estimation or polynomial interpolation. In this case, the range of the input data is divided into multiple regions (for example, 4 or 8 regions), and the scale value is estimated using a polynomial function. For example, in the case of a quadratic polynomial function, the scale value can be estimated using the following formula. shift0 = floorLog2(denom) round0 = 1< <shift0>>1 normDiff = (((denom << M) + round0) >> shift0) & ((1 << M) - 1); scale = a2 * (normDiff - b2) ^ 2 + a1 * (normDiff - b1) + a0

[0083] Here, a0, a1, a2, b1, and b2 are pre - defined values of the model in each region, and "^" is the square operation. The values of a0, a1, and a2 can be calculated using regression and MSE minimization. The parameter b2 can be defined to minimize or reduce the dynamic range of (normDiff - b2). For example, b2 can be set to the middle value of normDiff in that region.

[0084] To take advantage of integer operations, the parameters can be converted to integer values with a specific bit precision. In that case, the scale value can be calculated as follows, where n1 and n2 are the number of precision bits of a1 and a2 respectively. In this case, these parameters can be quantized to their precision bits. The value of n2 can be selected such that the original (fractional) parameter of a2 is quantized to an integer value with the minimum quantization error. scale = a2 * (((normDiff - b2) ^ 2) >> n2) + a1 * ((normDiff - b1) >> n1) + a0

[0085] The squared term can have a large dynamic range. To limit that dynamic range and the number of bits required for arithmetic operations, the squared term can be shifted right by a pre - defined number of shifts. In a special case, this shift number can be set to the bit depth of the normDiff parameter. In a general case, to maintain optimal accuracy for the calculation of the scale parameter, the following formula can be used and the right shift can be performed in two steps. scale = ((a2 * (((normDiff - b2) ^ 2) >> n21)) >> n22) + ((a1 * ((normDiff - b1) >> n11)) >> n12) + a0

[0086] In a special case, b1 and b2 can have the same value and be equal to b. In this case, the scale value is calculated using the following formula. The value of b can be selected to optimize the value of a0, a1, or a2. For example, if the value of b in the region is carefully selected, a1 can be a power of 2 (positive or negative power), and the resulting term a1*(normDiff - b) can be implemented by a shift (left or right shift) operation. Another advantage of a1 having a value that is a power of 2 is that the error in quantizing a1 to an integer value is reduced. scale=a2*(normDiff - b)^2+a1*(normDiff - b)+a0

[0087] As a result, when there are R = 4 regions and M = 14, the input value denom is in the range {16384, 32767}, the exact scale is in the range {16384, 8192}, and the region value is calculated using the following formula. region=(denom - 16384)>>(M - floorLog2(R))

[0088] The value of the region parameter can be calculated in various ways, where & is the bitwise AND operation. For example, region=((denom<<floorLog2(R))>>M)&((1<<floorLog2(R)-1)

[0089] Next, based on the region, appropriate parameters can be used for calculating the scale.

[0090] For example, in the case of 4 regions (R = 4), the parameters are as follows. a2[4]={182,99,60,39}; a0[4]={12348,11570,11926,13273}; b[4]={21850,23198,22434,19170}; region=r=(denom - 16384)>>(M - 2) normDiff = (((denom << M) + round0) >> shift0) & ((1 << (M + 1)) - 1); normDiff2 -= b[r]; scale = ((a2[r] * ((normDiff2 * normDiff2) >> (10))) >> (12)) - (normDiff2 >> 1) + a0[r];

[0091] As another example, when there are R = 8 regions and M = 14, the input value denom is in the range {16384, 32767}, the exact scale is in the range {16384, 8192}, and the scale value is estimated using the following formula. a2[8] = {214, 153, 113, 86, 67, 53, 43, 35}; a0[8] = {12784, 12054, 11670, 11583, 11764, 12195, 12870, 13782}; b[8] = {21206, 22336, 23008, 23176, 22792, 21808, 20176, 17850}; a0[4] = {12348, 11570, 11926, 13273}; b[4] = {21850, 23198, 22434, 19170}; region = r = (denom - 16384) >> (M - 2) normDiff = (((denom << M) + round0) >> shift0) & ((1 << (M + 1)) - 1); normDiff2 -= b[r]; scale = ((a2[r] * ((normDiff2 * normDiff2) >> (10))) >> (12)) - (normDiff2 >> 1) + a0[r];

[0092] The values of the parameters b and a2 for a given M can be calculated based on the value of denom or normDiff, for example, using a quadratic equation. normDiff3 = (normDiff - 19314) b = (-288 * normDiff3 * normDiff3) >> 22 + normDiff3 >> 1 + 22326 normDiff4 = (normDiff - 26569) a2 = (4 * normDiff3 * normDiff3) >> 22 + normDiff3 >> 7 + 55

[0093] Alternatively, the intermediate value normDiff can be calculated using a slightly different formula, and its range is {0, 16383}, and the region value is calculated using the following formula. region = normDiff >> (M - floorLog2(R))

[0094] Next, based on the region, appropriate parameters can be used for scale calculation.

[0095] For example, in the case of 4 regions (R = 4), the parameters are as follows. a2[4] = {182, 99, 60, 39}; a0[4] = {12348, 11570, 11926, 13273}; b[4] = {5466, 6814, 6050, 2786}; normDiff = (((denom << M) + round0) >> shift0) & ((1 << M) - 1); region = r = normDiff >> (M - 2) normDiff2 -= b[r]; scale = ((a2[r] * ((normDiff2 * normDiff2) >> (10))) >> (12)) - (normDiff2 >> 1) + a0[r];

[0096] As another example, when there are 8 regions (R = 8) and M = 14, the input value denom is in the range of {16384, 32767}, the exact scale is in the range of {16384, 8192}, and the scale value is estimated using the following formula. a2[8] = {214, 153, 113, 86, 67, 53, 43, 35}; a0[8] = {12784, 12054, 11670, 11583, 11764, 12195, 12870, 13782}; b[8] = {4822, 5952, 6624, 6792, 6408, 5424, 3792, 1466}; a0[4] = {12348, 11570, 11926, 13273}; b[4] = {21850, 23198, 22434, 19170}; normDiff = (((denom << M) + round0) >> shift0) & ((1 << M) - 1); region = r = normDiff >> (M - 2) normDiff2 -= b[r]; scale = ((a2[r] * ((normDiff2 * normDiff2) >> (10))) >> (12)) - (normDiff2 >> 1) + a0[r];

[0097] The values of parameters such as a2, a1, a0, and b for different M' values can be calculated using scaling based on one specific M value such as M = 14. a0' = (M' > M)? (a0 << (M' - M)) : (a0 >> (M - M')) a1' = (M' > M)? (a1 << (M' - M)) : (a1 >> (M - M')) a2' = (M' > M)? (a2 << (M' - M)) : (a2 >> (M - M')) b' = (M' > M)? (b << (M' - M)) : (b >> (M - M'))

[0098] The values of parameters b and a2 for a given M can be calculated based on the value of denom or normDiff, for example, using a quadratic equation. normDiff3 = (normDiff - 2930) b = (-288 * normDiff3 * normDiff3) >> 22 + normDiff3 >> 1 + 5942 normDiff4 = (normDiff - 10185) a2 = (4 * normDiff3 * normDiff3) >> 22 + normDiff3 >> 7 + 55

[0099] In one embodiment, the entire range of normDiff can be unevenly divided, and the a2, a1, a0, and b parameters can be optimized for each region, resulting in less truncation or quantization error and easier implementation. The ranges used to calculate each of those parameters can be overlapping regions. For example, in the case of M = 14, when normDiff is within the range of {16383, 25472}, those parameters can be set to the following values, which can be easily implemented with low quantization error. a2 = 128 (which is a power of 2 and can be implemented by a 7-bit right shift) a1 = -1 a0 = 17745 b = 14895

[0100] The fixed parameters of the above formula can be made smaller or larger by one unit or several units based on the method by which their values are quantized to integer values in order to optimize the division approximation error and / or software / hardware implementation.

[0101] The functions described in this specification can be used in various contexts. For example, using the calculated output, the exact average sample value or the estimated average sample value of a sample set can be determined. In that case, the sum of the sample values in the set can be used as the numerator nom, and the number of samples in the set can be used as the denominator denom. As another example, this method can be used to determine the filter coefficient in the least squares sense. In some algorithms used in video and image coding, for example, it is necessary to decompose a matrix using LDL decomposition or Cholesky decomposition. All of them require performing a division operation in order to perform the selected decomposition. In such cases, by replacing the complete division operation with the method proposed in this specification, the deviation from the mathematical optimal solution can be minimized while significantly reducing the complexity.

[0102] According to one embodiment, the output value approximating the result of the division operation between the numerator and the denominator is determined using a scale value and a shift value, and the scale value is determined using an interpolation operation.

[0103] According to one embodiment, the output value approximating the result of the division operation between the numerator and the denominator is determined using a scale value and a shift value, and the scale value is determined using polynomial interpolation.

[0104] According to one embodiment, the output value approximating the result of the division operation between the numerator and the denominator is determined using a scale value and a shift value, and the scale value is determined using linear interpolation.

[0105] According to one embodiment, the output value approximating the result of the division operation between the numerator and the denominator is determined using a scale value and a shift value, and the scale value is determined using an interpolation operation between at least two values determined based on table lookup.

[0106] The method according to one embodiment is shown in FIG. 8. This method generally includes receiving a numerator and a denominator 810, and determining an approximated output of a division operation between the numerator and the denominator 820, where this determination includes deriving a scale parameter using piecewise approximation, deriving a shift parameter and a rounding parameter, and applying the scale parameter, the shift parameter, and the rounding parameter to the numerator. Each step can be performed by respective modules of a computer system.

[0107] The method according to one embodiment is shown in FIG. 9. This method is encoding / decoding a picture including a plurality of samples 910, where the encoding / decoding phase includes a division operation, encoding / decoding 910, determining a numerator and a denominator 920, and determining an approximated output of a division operation between the numerator and the denominator 930, where this determination includes deriving a scale parameter using piecewise approximation, deriving a shift parameter and a rounding parameter, and applying the scale parameter, the shift parameter, and the rounding parameter to the numerator, determining 930, and using a division operation in the encoding / decoding phase 940, and generally includes. Each step can be performed by respective modules of a computer system.

[0108] An apparatus according to one embodiment comprises means for carrying out either of the methods of FIG. 8 or FIG. 9. The means comprises at least one processor and a memory including computer program code, and the processor may further comprise a processor circuit. The memory and the computer program code are configured to cause the at least one processor to execute the method according to various embodiments in the apparatus.

[0109] FIG. 10 shows an apparatus according to an embodiment. The general structure of this apparatus will be described along the functional blocks of the system. Some functions can be executed by a single physical device. For example, all calculation procedures can be executed by a single processor as needed. The data processing system of the apparatus according to an example of FIG. 10 includes a main processing unit 1000, a memory 1002, a storage device 1004, an input device 1006, an output device 1008, and a graphics subsystem 1010, all of which are interconnected via a data bus 1012. A client can be understood as a client device or a software client operating on the device.

[0110] The main processing unit 1000 is a processing unit configured to process data within the data processing system. The main processing unit 1000 may include or be implemented as one or more processors or processor circuits. The memory 1002, the storage device 1004, the input device 1006, and the output device 1008 may include other components recognized by those skilled in the art. The memory 1002 and the storage device 1004 store data in the data processing system 1000. The computer program code exists in the memory 1002 to implement, for example, a machine learning process. The input device 1006 inputs data into the system, and the output device 1008 receives data from the data processing system and transfers the data to, for example, a display. Although the data bus 1012 is shown as a single line, it can be any combination of a processor bus, a PCI bus, a graphical bus, and an ISA bus. Therefore, those skilled in the art can easily understand that this apparatus can be any data processing device, such as a computer device, a personal computer, a server computer, a mobile phone, a smartphone, or an Internet access device, such as an Internet tablet computer.

[0111] Various embodiments can be implemented with the aid of computer program code that resides in a memory and causes a related apparatus to execute the method. For example, a device can include circuitry and electronics for processing, receiving, and transmitting data, computer program code in a memory, and a processor that, when executing the computer program code, causes the device to implement the features of one embodiment. Additionally, a network device such as a server can include circuitry and electronics for processing, receiving, and transmitting data, computer program code in a memory, and a processor that, when executing the computer program code, causes the network device to implement the features of various embodiments.

[0112] If desired, the various functions discussed herein can be performed in a different order and / or simultaneously with each other. Further, if desired, one or more of the above-described functions and embodiments can be optional or can be combined.

[0113] While various aspects of the embodiments are set forth in the independent claims, other aspects include other combinations of the features of the described embodiments and / or the dependent claims and the features of the independent claims, and do not include only the combinations explicitly recited in the claims.

[0114] It should also be noted that, while exemplary embodiments have been described above in this specification, these descriptions should not be construed in a limiting sense. Rather, there are several variations and modifications that can be made without departing from the scope of the disclosure as defined in the appended claims. < / shift> < / shift> < / d> < / shift> < / shift> < / d> < / shift>

Claims

1. Means for encoding / decoding a picture containing a plurality of samples, wherein a phase of the encoding / decoding includes a division operation, the means for encoding / decoding, means for determining a numerator and a denominator, means for determining an approximated output of the division operation between the numerator and the denominator, means for deriving a scale parameter using piecewise approximation, means for deriving a shift parameter and a rounding parameter, means for determining, comprising means for applying the scale parameter, the shift parameter, and the rounding parameter to the numerator, means for using the approximated output of the division operation in the phase of the encoding / decoding, An apparatus comprising the above.

2. The apparatus according to claim 1, wherein the phase of the encoding / decoding includes predicting at least one sample of the picture.

3. The apparatus according to claim 1, wherein the phase of the encoding / decoding includes filtering.

4. The apparatus according to claim 1, 2, or 3, further comprising means for applying an additional shift parameter to the numerator.

5. The apparatus according to any one of claims 1 to 4, further comprising means for determining if bit precisions of the numerator, denominator, scale parameter, and output parameter are different, and if so, using the bit precisions of the output and the scale parameter when determining the approximated output.

6. The apparatus according to any one of claims 1 to 5, further comprising means for determining the scale parameter using one of interpolation operation, polynomial processing, and linear interpolation.

7. The apparatus according to claim 6, further comprising means for determining the scale parameter using the interpolation operation between at least two values determined based on table look-up.

8. The apparatus according to claim 7, further comprising means for adjusting the scale parameter with a value for narrowing a range of the denominator.

9. Encoding / decoding a picture containing a plurality of samples, wherein a phase of the encoding / decoding includes a division operation, the encoding / decoding, determining a numerator and a denominator, determining an approximated output of the division operation between the numerator and the denominator, Deriving a scale parameter using a piecewise approximation, deriving a shift parameter and a rounding parameter, determining, including applying the scale parameter, the shift parameter, and the rounding parameter to the numerator, a method including using the approximated output of the division operation in the encoding / decoding phase. **Claim 10** The method according to claim 9, wherein the encoding / decoding phase includes predicting at least one sample of the picture. **Claim 11** The method according to claim 9 or 10, wherein the encoding / decoding phase includes filtering. **Claim 12** The method according to any one of claims 9 to 11, further including determining if the bit precisions of the numerator, denominator, scale parameter, and output parameter are different, and if so, using the bit precisions of the output and the scale parameter when determining the approximated output. **Claim 13** The method according to any one of claims 9 to 12, further including determining the scale parameter using one of interpolation operations, polynomial processing, and linear interpolation. **Claim 14** The method according to claim 13, further including determining the scale parameter using the interpolation operation between at least two values determined based on a table lookup. **Claim 15** An apparatus comprising at least one processor and a memory including computer program code, wherein the memory and the computer program code cause the at least one processor to cause the apparatus to at least encode / decode a picture including some samples, the encoding / decoding phase including a division operation, determine a numerator and a denominator, determine an approximated output of the division operation between the numerator and the denominator, and the apparatus further derives a scale parameter using a piecewise approximation, derives a shift parameter and a rounding parameter, applies the scale parameter, the shift parameter, and the rounding parameter to the numerator, uses the approximated output of the division operation in the encoding / decoding phase, is configured as such.

Citation Information

Patent Citations

  • Method and apparatus for division-free intra-prediction

    US20220038689A1

  • Method and apparatus for low-complexity bi-directional intra prediction in video encoding and decoding

    US20220086487A1

  • Method and apparatus for performing an arithmetic coding for data symbols

    WO2015102432A1

  • Filter apparatus and methods

    WO2018149995A1

  • Piecewise modeling for linear component sample prediction

    WO2020127956A1