Adaptive clipping in model parameter derivation method for video compression

By cutting the current template and reference template during the video encoding and decoding process, the adaptive cropping technology of derives the model parameters, the problem of poor encoding efficiency and quality in the existing technology is solved, and more efficient video compression is achieved.

CN120359751APending Publication Date: 2025-07-22TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005256.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-22
Filing Date
2024-10-25
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology is difficult to effectively use model parameters for adaptive cropping during compression, resulting in poor encoding efficiency and quality.

Method used

By receiving the encoded information of the picture sequence, a prediction sample of the current block is generated using model-based prediction technology, and the current template and reference template are cropped to derive the adaptive cropping technology of model parameters and improve coding efficiency.

Benefits of technology

Improve the encoding efficiency and quality during video encoding and decoding, and optimize the video compression effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359751A_ABST
    Figure CN120359751A_ABST
Patent Text Reader

Abstract

A method includes receiving a bitstream of encoded information for a picture. The encoded information indicates prediction of a current block using a model-based prediction technique that generates a prediction sample of the current block based on a model, in which at least one reconstructed sample of a reference block is input, the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block. The method further comprises the following steps: executing a clipping operation on at least one of the current template and the reference template to obtain a clipped template sample; deriving at least one parameter value of the at least one parameter of the model from the clipped template sample; and generating at least one prediction sample of the current block by using the model.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporated herein by reference

[0001] This application claims priority to U.S. Patent Application No. 18 / 923,643, filed on October 22, 2024, entitled "Adaptive Cropping in Model Parameter Derivation Methods for Video Compression", which claims priority to U.S. Provisional Application No. 63 / 602,343, filed on November 22, 2023, entitled "Adaptive Cropping in Model Parameter Derivation Methods for Video Compression". The entire disclosure of the prior application is incorporated herein by reference. Technical Field

[0002] Embodiments of the present disclosure generally relate to video encoding and decoding. Background Art

[0003] The background description provided herein is intended to present the background of the present disclosure generally. The extent to which the work of the presently named inventors, which is described in the background art section and in various aspects of this specification, was carried out does not indicate that it was prior art at the time of filing of the present disclosure, and is never expressly or implicitly admitted to be prior art of the present disclosure.

[0004] Image / video compression can help transmit image / video data across different devices, storage, and networks while minimizing quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancies. In one example, a video codec can use a technique called intra prediction, which can compress an image based on spatial redundancy. For example, intra prediction can use reference data from the current picture being reconstructed to predict samples. In another example, a video codec can use a technique called inter prediction, which can compress an image based on temporal redundancy. For example, inter prediction can predict samples in the current picture from a previously reconstructed picture through motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention

[0005] Aspects of the present disclosure include bitstreams, methods, and apparatuses for video encoding / decoding. In some embodiments, the apparatus for video encoding / decoding includes processing circuitry

[0006] Some aspects of the present disclosure provide a method for video decoding. The method includes: receiving a bitstream of encoded information of a picture sequence, the encoded information indicating the prediction of a current block in a current picture using a model-based prediction technique, the model-based prediction technique generating a prediction sample of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block. The method further includes: performing at least one cropping operation on at least one of the current template and the reference template to obtain cropped template samples; deriving at least one parameter value of the at least one parameter of the model according to the cropped template samples; and generating at least one prediction sample of the current block by using the model with the at least one parameter set to the at least one parameter value.

[0007] Some aspects of the present disclosure provide a method for video encoding. The method includes: determining to encode a current block in a current picture using a model-based prediction technique, the model-based prediction technique generating a prediction sample of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block. The method further includes: performing at least one cropping operation on at least one of the current template and the reference template to obtain cropped template samples; deriving at least one parameter value of the at least one parameter of the model according to the cropped template samples; and encoding the current block based on the model with the at least one parameter set to the at least one parameter value.

[0008] Aspects of the present disclosure also provide an apparatus for video encoding / decoding.

[0009] Aspects of the present disclosure also provide a method for processing visual media data. In this method, a bitstream of visual media data is processed according to format rules. For example, the bitstream can be a bitstream decoded / encoded by any decoding and / or encoding method described herein. The format rules can specify at least one constraint of the bitstream and / or at least one process to be performed by a decoder and / or an encoder.

[0010] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any one of the video decoding / encoding methods described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0012] Figure 1 It is a schematic diagram of an example of a block diagram of a video processing system in some embodiments.

[0013] Figure 2 It is a schematic diagram of an example of a block diagram of a decoder.

[0014] Figure 3 It is a schematic diagram of an example of a block diagram of an encoder.

[0015] Figure 4 Shows the positions of spatial merge candidates according to an embodiment of the present disclosure.

[0016] Figure 5 Shows candidate pairs considering redundancy checks for spatial merge candidates according to an embodiment of the present disclosure.

[0017] Figure 6 Shows an example motion vector scaling of temporal merge candidates in some embodiments.

[0018] Figure 7 Shows example candidate positions of temporal merge candidates in some examples.

[0019] Figure 8 Shows templates in some embodiments.

[0020] Figure 9 Shows a flowchart outlining a decoding process according to some aspects of the present disclosure.

[0021] Figure 10 Shows a flowchart outlining a decoding process according to other aspects of the present disclosure.

[0022] Figure 11 Shows a flowchart outlining an encoding process according to some aspects of the present disclosure.

[0023] Figure 12 Shows a flowchart outlining an encoding process according to other aspects of the present disclosure.

[0024] Figure 13 Is a schematic diagram of a computer system according to one aspect. Detailed Description

[0025] Figure 1 Shows a block diagram of a video processing system (100) in some examples. The video processing system (100) is an example of an application for video encoders and video decoders in the disclosed subject matter, streaming environments. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0026] A video processing system (100) includes an acquisition subsystem (113), and the acquisition subsystem may include a video source (101) such as a digital camera. The video source creates an uncompressed video picture stream (102). In an embodiment, the video picture stream (102) includes samples taken by the digital camera. Compared with the encoded video data (104) (or the encoded video bitstream), the video picture stream (102) is depicted as a thick line to emphasize the high data volume of the video picture stream. The video picture stream (102) can be processed by an electronic device (120), and the electronic device (120) includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the video picture stream (102), the encoded video data (104) (or the encoded video bitstream) is depicted as a thin line to emphasize the lower data volume of the encoded video data (104) (or the encoded video bitstream), which can be stored on a streaming server (105) for future use. At least one streaming client subsystem, such as Figure 1 the client subsystem (106) and the client subsystem (108) in can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and generates an output video picture stream (111) that can be presented on a display (112) (such as a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), the video data (107), and the video data (109) (such as the video bitstream) may be encoded according to certain video coding / compression standards. Embodiments of such standards include ITU-T H.265. In an embodiment, a video coding standard under development is informally referred to as Versatile Video Coding (VVC), and this application can be used in the context of the VVC standard.

[0027] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).

[0028] Figure 2Shows an example of a block diagram of a video decoder (210). The video decoder (210) may be provided in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 the video decoder (110) in the embodiment.

[0029] The receiver (231) may receive at least one encoded video sequence to be decoded by the video decoder (210), for example, included in a bitstream. In one aspect, one encoded video sequence is received at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that can be forwarded to their respective using entities (not labeled). The receiver (231) may separate the encoded video sequence from other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be provided outside the video decoder (210) (not labeled). In other cases, a buffer memory (not labeled) is provided outside the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be configured inside the video decoder (210) to, for example, handle the playback timing. And when the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous synchronous network, it may not be necessary to configure the buffer memory (215), or the buffer memory may be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (215) may also be required, which may be relatively large and may have an adaptive size, and may be at least partially implemented in an operating system or a similar element (not labeled) outside the video decoder (210).

[0030] The video decoder (210) may include a parser (220) to reconstruct symbols (221) according to the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (210), and potential information for controlling a display device (212) (e.g., a display screen), etc., which is not a component of the electronic device (230), but may be coupled to the electronic device (230), such as Figure 2As shown in. The control information for the display device may be a parameter set segment (not labeled) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (220) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (220) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Unit (CU), blocks, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0031] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0032] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (221) may involve at least two different units. Which units are involved and the way they are involved may be controlled by the subgroup control information parsed by the parser (220) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (220) and at least two units below are not described.

[0033] In addition to the functional blocks already mentioned, the video decoder (210) may be conceptually divided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and may be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide them into the functional units below.

[0034] The first unit is a scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives, from a parser (220), quantized transform coefficients as symbols (221) and control information, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) may output a block including sample values, and the sample values may be input into an aggregator (255).

[0035] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed part of the current picture. Such predictive information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a surrounding block having the same size and shape as the block being reconstructed, using reconstructed information extracted from a current picture buffer (258). For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, based on each sample, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0036] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-coded and potentially motion-compensated block. In this case, a motion compensation prediction unit (253) may access a reference picture memory (257) to extract samples for prediction. After motion-compensating the extracted samples according to the symbol (221), these samples may be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (which is referred to as a residual sample or a residual signal in this case), thereby generating output sample information. The motion compensation prediction unit (253) obtaining prediction samples from an address within the reference picture memory (257) may be controlled by a motion vector, and the motion vector is in the form of the symbol (221) for use by the motion compensation prediction unit (253), and the symbol (221) includes, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (257), a motion vector prediction mechanism, etc., when using sub-sample accurate motion vectors.

[0037] The output samples of the aggregator (255) can be employed by various loop filtering techniques in the loop filter unit (256). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream), and the parameters can be available to the loop filter unit (256) as symbols (221) from the parser (220). However, in other embodiments, video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) portions of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0038] The output of the loop filter unit (256) can be a sample stream that can be output to the display device (212) and stored in the reference picture memory (257) for subsequent inter-picture prediction.

[0039] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture has been fully reconstructed and the encoded picture is identified (e.g., by the parser (220)) as a reference picture, the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.

[0040] The video decoder (210) can perform decoding operations according to a predetermined video compression technique or standard (e.g., ITU-T H.265). The encoded video sequence can conform to the syntax specified by the video compression technique or standard used in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the meaning of the profile recorded in the video compression technique or standard. Specifically, a profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under that profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.

[0041] In an embodiment, the receiver (231) may receive additional (redundant) data together with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (210) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, a temporal, spatial, or signal noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.

[0042] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is provided in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) may be used to replace Figure 1 the video encoder (103) in the embodiment.

[0043] The video encoder (303) may receive video samples from a video source (301) (which is not Figure 3 part of the electronic device (320) in the embodiment), and the video source may capture video images to be encoded by the video encoder (303). In another embodiment, the video source (301) is part of the electronic device (320).

[0044] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 Y CrCB, RGB,...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device that stores previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as at least two separate pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, and each pixel may include at least one sample depending on the sampling structure, color space, etc. used. The following focuses on the description of samples.

[0045] According to an embodiment, a video encoder (303) may encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other desired time constraint. Enforcing an appropriate encoding speed is a function of a controller (350). In some embodiments, the controller (350) controls and is functionally coupled to other functional units as described below. For the sake of brevity, couplings are not labeled in the figures. Parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be used for other suitable functions that relate to optimizing the video encoder (303) for a particular system design.

[0046] In some embodiments, the video encoder (303) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data. The reconstructed sample stream (sample data) is input into a reference picture memory (334). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is also used in some related technologies.

[0047] The operation of the "local" decoder (333) may be the same as that of the "remote" decoder that has been described in detail above in connection with Figure 2 the video decoder (310). However, briefly referring additionally to Figure 2 , when symbols are available and the entropy encoder (345) and the parser (220) can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implementable in the local decoder (333).

[0048] In an embodiment, in addition to parsing / entropy decoding that exists in the decoder, the decoder technology also exists in the corresponding encoder in the same or substantially the same functional form. Therefore, this application focuses on decoder operations. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. In some areas, a more detailed description is provided below.

[0049] During operation, in some embodiments, the source encoder (330) may perform motion-compensated predictive coding. With reference to at least one previously encoded picture designated as a "reference picture" in the video sequence, the motion-compensated predictive coding performs predictive coding on the input picture. In this way, the encoding engine (332) encodes the difference between the pixel blocks of the input picture and the pixel blocks of the reference picture, and the reference picture can be selected as the prediction reference for the input picture.

[0050] The local video decoder (333) may decode the encoded video data of the picture that can be designated as a reference picture based on the symbols created by the source encoder (330). The operation of the encoding engine (332) may be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 3 not shown), the reconstructed video sequence is generally a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that can be performed by the video decoder on the reference picture and may store the reconstructed reference picture in the reference picture memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference picture, which has the same content (without transmission errors) as the reconstructed reference picture to be obtained by the remote video decoder.

[0051] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) may search in the reference picture memory (334) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new picture. The predictor (335) may operate block by block based on sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (335), it may be determined that the input picture may have prediction references obtained from at least two reference pictures stored in the reference picture memory (334).

[0052] The controller (350) may manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0053] The outputs of all the above functional units can be entropy encoded in an entropy encoder (345). The entropy encoder (345) losslessly compresses the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.

[0054] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) to prepare for transmission over a communication channel (360), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can merge the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0055] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures can typically be assigned to any of the following picture types:

[0056] An intra picture (I picture) can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.

[0057] A predictive picture (P picture) can be encoded and decoded using intra prediction or inter prediction, which uses motion vectors and reference indices to predict the sample values of each block.

[0058] A bi - predictive picture (B picture) can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values of each block. Similarly, at least two predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0059] Source pictures can typically be spatially subdivided into at least two sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (already encoded) blocks, which are determined according to the encoding assignments of the corresponding pictures applied to the blocks. For example, blocks of an I picture can be non-prediction-encoded, or the blocks can be prediction-encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be prediction-encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be prediction-encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.

[0060] The video encoder (303) can perform encoding operations according to a predetermined video encoding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (303) can perform various compression operations, including prediction encoding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard used.

[0061] In an embodiment, the transmitter (340) can transmit additional data when transmitting the encoded video. The source encoder (330) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0062] The captured video can be at least two source pictures (video pictures) in a time sequence. Intra picture prediction (often simplified to intra prediction) utilizes the spatial correlation in a given picture, while inter picture prediction utilizes the (temporal or other) correlation between pictures. In an embodiment, the particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using at least two reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0063] In some embodiments, bidirectional prediction techniques can be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the decoded order (but may be past and future respectively in the display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block can be predicted by a combination of the first reference block and the second reference block.

[0064] In addition, the merge mode technique can be used in inter-picture prediction to improve coding efficiency.

[0065] According to some embodiments disclosed in the present application, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression. The CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be split into at least one coding unit (CU) by a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an embodiment, each CU is analyzed to determine the prediction type for the CU, such as an inter-prediction type or an intra-prediction type. In addition, depending on the temporal and / or spatial predictability, the CU is split into at least one prediction unit (PU). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed on a prediction block basis. Taking the luminance prediction block as an example of the prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0066] Note that any suitable technology may be used to implement video encoders (103) and (303) and video decoders (110) and (210). In an embodiment, at least one integrated circuit may be used to implement video encoders (103) and (303) and video decoders (110) and video decoder (210). In another embodiment, at least one processor executing software instructions may be used to implement video encoders (103) and video encoder (303) and video decoders (110) and video decoder (210).

[0067] Some aspects of the present disclosure provide techniques for adaptive cropping in model parameter derivation for video compression. In some examples, these techniques may be used for coded information derivation in inter prediction coding.

[0068] Various inter prediction modes may be used in video coding. For example, in VVC, for an inter prediction CU, the motion parameters may include at least one MV, at least one reference picture index, a reference picture list usage index, and additional information for certain coding features to be used for generating inter prediction samples. The motion parameters may be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU may be associated with a PU and may have no significant residual coefficients, no coded motion vector difference or MV difference (e.g., MVD) or reference picture index. A merge mode may be specified, where the motion parameters of the current CU are obtained from at least one neighboring CU, including spatial and / or temporal candidates, and optionally including additional information introduced, for example, in VVC. The merge mode is applicable not only to skip mode but may also be applied to inter prediction CUs. In an example, an alternative to the merge mode is the explicit transmission of motion parameters, where each CU explicitly signals at least one MV, the corresponding reference picture index for each reference picture list, and a reference picture list usage flag, and other information.

[0069] In an embodiment, for example, in VVC, the VVC Test Model (VTM) reference software includes at least one modified inter prediction coding tool, which includes: extended merge prediction, merged motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, sub-block based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luminance sample MV storage and 8×8 motion field compression), bi-directional prediction with CU level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), etc. Inter prediction and related methods are described in detail below.

[0070] In some examples, extended merge prediction can be used. In an example, such as in VTM4, a merge candidate list is constructed by sequentially including the following five types of candidates: at least one spatial motion vector predictor (MVP) from at least one spatially adjacent CU, at least one temporal MVP from at least one collocated CU, at least one history-based MVP (HMVP) from a first-in-first-out (FIFO) table, at least one pairwise average MVP, and at least one zero MV.

[0071] The size of the merge candidate list can be signaled in the slice header. In an example, in VTM4, the maximum allowed size of the merge candidate list is 6. For each CU encoded in the merge mode, the index of the best merge candidate (e.g., the merge index) can be encoded using truncated unary binarization (TU). The first binary digit (bin) of the merge index can be encoded using context (e.g., context-adaptive binary arithmetic coding (CABAC)), and the other bins can be encoded using bypass coding.

[0072] Some examples of the generation process of merge candidates for each category are provided below. In an embodiment, at least one spatial candidate is derived as follows. The derivation of spatial merge candidates in VVC can be the same as that in HEVC. In an example, up to four merge candidates are selected from the candidates located at the positions shown in Figure 4 .

[0073] Figure 4 Shows the positions of spatial merge candidates according to an embodiment of the present disclosure. Referring to Figure 4 , the derivation order is B1, A1, B0, A0, and B2. Position B2 is considered only if any of the CUs at positions A0, B0, B1, and A1 are unavailable (e.g., because the CU belongs to another slice or another tile) or are intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates to ensure that candidates with the same motion information are excluded from the candidate list, thereby improving the coding efficiency.

[0074] To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs linked by the arrows in Figure 5 are considered, and a candidate is added to the candidate list only if the corresponding candidates used for the redundancy check do not have the same motion information.

[0075] Figure 5 Shows the candidate pairs considered for the redundancy check of spatial merge candidates according to an embodiment of the present disclosure. Referring to Figure 5, the pairs linked by respective arrows include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Accordingly, candidates at positions B1, A0, and / or B2 can be compared with candidates at position A1, and candidates at positions B0 and / or B2 can be compared with candidates at position B1.

[0076] In an embodiment, at least one temporal candidate is derived as follows. In an example, only one temporal merge candidate is added to the candidate list. Figure 6 An example motion vector scaling of a temporal merge candidate is shown. To derive a temporal merge candidate for a current CU (611) in a current picture (601), a scaled MV (621) can be derived based on a collocated CU (612) belonging to a collocated reference picture (604) (e.g., Figure 6 as indicated by the dashed line in ). The reference picture list used to derive the collocated CU (612) can be signaled explicitly in the slice header. As Figure 6 indicated by the dashed line in, a scaled MV (621) of the temporal merge candidate can be obtained. The scaled MV (621) can be scaled from the MV of the collocated CU (612) using picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (602) of the current picture (601) and the current picture (601). The POC distance td can be defined as the POC difference between the collocated reference picture (604) of the collocated picture (603) and the collocated picture (603). The reference picture index of the temporal merge candidate can be set to zero. The collocated picture is a reference picture that serves as a source picture for temporal motion information derivation. The collocated picture can be identified in one of two lists (referred to as list0 or list1). In some examples, the encoder can determine the collocated picture and signal the collocated picture using an appropriate syntax technique.

[0077] Figure 7 Example candidate positions (e.g., C0 and C1) of a temporal merge candidate for a current CU are shown. A position of the temporal merge candidate can be selected from candidate positions C0 and C1. Candidate position C0 is at the lower right corner of the collocated CU (710) of the current CU. Candidate position C1 is at the center of the collocated CU (710) of the current CU. If the CU at candidate position C0 is unavailable, intra-coded, or outside the current row of the CTU, candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, inter-coded, and within the current row of the CTU, candidate position C0 is used to derive the temporal merge candidate.

[0078] According to some aspects of the present disclosure, a model-based prediction method can be used in inter-frame prediction or intra-frame prediction. For inter-frame prediction, the model-based prediction method can use a model (e.g., formula, function) to generate samples of a current block in a current picture based on reference samples of a reference block in a reference picture. For intra-frame prediction, the model-based prediction method can use a model (e.g., formula, function) to generate a first color component of a current block based on a second color component of the current block. In some examples, parameters in the model for the model-based prediction method can be derived based on a template of the current block.

[0079] In some examples, local illumination compensation (LIC) is used as an inter-frame prediction technique to model local illumination changes between a current block and a reference block of the current block by using a linear function. The reference block is located in a reference picture and can be pointed to by a motion vector (MV). The model is applied to the reference block to generate a predicted block. The parameters of the linear formula can include a scaling factor α and an offset β, and the linear formula can be represented as α×p[x,y]+β to compensate for illumination changes, where p[x,y] represents a reference sample at position [x,y] in the reference block (also referred to as the predicted block), and the reference block is pointed to by the MV of the current block. In some examples, the scaling factor α and the offset β can be derived based on a template of the current block and a corresponding reference template of the reference block by using the least squares method. Therefore, in addition to signaling an LIC flag to indicate the use of LIC, no signaling overhead is required. The scaling factor α and the offset β derived based on the template of the current block can be referred to as a template-based parameter set.

[0080] In some examples, LIC is used for uni-directional prediction between CUs. In some examples, intra-frame neighboring samples of the current block (neighboring samples predicted using intra-frame prediction) can be used for LIC parameter derivation. In some examples, LIC is disabled for blocks with luminance samples less than 32. In some examples, for non-sub-block modes (e.g., non-affine modes), LIC parameter derivation is performed based on the modulo-block samples of the current CU rather than the partial modulo-block samples of the top-left 16×16 unit. In some examples, LIC parameter derivation is performed based on partial modulo-block samples (e.g., partial modulo-block samples for the top-left 16×16 unit). In some examples, the template samples of the reference block are determined by using motion compensation (MC) with the MV of the block without rounding it to integer pixel accuracy.

[0081] In some examples, cross-component prediction can be used as an intra-frame prediction technique. Cross-component prediction can include a first technique called cross-component linear model (CCLM), a second technique called multi-model linear model (MMLM), a third technique called convolutional cross-component model (CCCM), and a fourth technique called gradient linear model (GLM).

[0082] For example, the first technique CCLM is used to reduce cross-component redundancy. In CCLM, by using a linear model (also known as a linear formula), such as using Equation (1), chrominance samples are predicted based on the reconstructed luma samples of the same CU: pred C (i,j) = a·rec L ′(i,j)+b Equation (1) where pred C (i,j) represents the predicted chrominance sample in the CU, and rec L (i,j) represents the downsampled reconstructed luma sample of the same CU. The CCLM linear model includes parameters (a and b), which in the example can be derived using up to four adjacent chrominance samples and their corresponding downsampled luma samples. In the example, up to four adjacent chrominance samples and their corresponding downsampled luma samples are referred to as the template of the CU.

[0083] In some examples, based on the positions of adjacent chrominance samples, CCLM can include different modes called LM_T (LM top mode or upper mode LM_A), LM_L (LM left mode), and LM_LT (LM top-left mode or upper-left mode LM_LA or just the LM mode). For example, if the size of the current chrominance block is W×H, then W’ and H’ can be set for various modes in CCLM. When the LM mode (also known as LM_LT or LM_LA) is applied, W’ = W and H’ = H; when the LM-A mode is applied, W’ = W + H; when the LM-L mode is applied, H’ = H + W.

[0084] Note that MMLM, CCCM, and GLM also use functions for prediction. The parameters of the functions can be derived based on the template.

[0085] Note that the following description uses inter prediction to illustrate the techniques of the model-based prediction method, and these techniques can be appropriately used for intra prediction.

[0086] According to some aspects of the present disclosure, a model-based prediction method can build a model, such as a linear model or a non-linear model, based on already available data, and then apply the model to subject data to improve its quality. The model-based prediction method works under the assumption that the data used to derive the model has a high correlation with the subject data, and thus the model is reliable in many cases. For example, an inter-frame prediction correction model is designed to minimize the distortion between a current block and its predicted block (generated from a reference block in a corresponding reference picture based on the inter-frame prediction correction model). For example, an inter-frame prediction correction model (also referred to as Method A, model-based inter-frame prediction technique, function-based inter-frame prediction technique) can apply a formula, such as a non-linear formula, a linear formula, etc., using the original reference block in the reference picture as the input of the formula to generate a predicted block for the current block in the current picture. For example, the inter-frame prediction correction technique can generate a prediction for a sample in the current block based on a formula that takes at least one predicted sample in the reference picture as the input. The formula can include linear terms or non-linear terms, and can include at least one parameter that can be derived. For example, the non-linear term or the linear term can be defined by a formula that has a set of parameters (e.g., represented by α below without loss of generality) i which are derived based on the template of the current block and the template of the reference block of the current block by minimizing the difference between the template of the current block and the template of the reference block (e.g., by using the least squares method).

[0087] Note that LIC is one of such inter-frame prediction correction techniques.

[0088] In some examples, the formula is a linear formula and can be represented as where n is a non-negative integer, p(x i ,y i ) is the predicted sample at position (x i ,y i ) in the reference picture, and this predicted sample is pointed to based on the MV associated with the current block. In addition, a set of predicted samples represented by p(x i ,y i ) (where i = 0,..., n) can be a set of predicted samples around the corresponding sample in the reference samples, and these predicted samples are pointed to based on the MV of the current sample to be predicted. In some examples, the parameters α i and β can be derived by minimizing the difference between the template of the current block (also referred to as the current block template) and the template of the predicted block of the current block (also referred to as the predicted block template) (e.g., by using the least squares method). The template of the current block consists of spatially adjacent reconstructed samples of the current block, and the template of the predicted block consists of spatially adjacent reconstructed samples of the predicted block.

[0089] Note that, although linear models are used to illustrate some embodiments in the present disclosure, non-linear models can also be used instead of linear models.

[0090] Figure 8 A schematic diagram of templates in some examples is shown. For example, template (810) is referred to as an L-shaped template T L , and includes adjacent samples above the current block (also referred to as the current coding block), in the left column, and at the upper left corner; template (820) is referred to as an above and left template T a+l , and includes adjacent samples in the above row and left column of the current block; template (830) is referred to as an above template T a , and includes adjacent samples in the above row of the current block; template (840) is referred to as a left template T l , and includes adjacent samples in the left column of the current block. It should be noted that the template may include Figure 8 adjacent samples of other suitable shapes not shown in

[0091] In some examples, multiple candidate template types can also be supported, and one candidate template type is selected to derive the parameters of a linear or non-linear model. The syntax can be signaled in the bitstream (e.g., at the block level) to indicate which candidate template type is selected.

[0092] In some examples, control flags associated with model-based inter prediction techniques can be signaled in the bitstream (e.g., at the block level) to indicate whether to apply the model-based inter prediction technique to the current block. Note that, in some examples, the value of the control flag can also be inherited from other coding blocks. More specifically, a first control flag of the model-based inter prediction technique associated with the current block is inherited from a second control flag of the model-based inter prediction technique associated with at least one other coding block. Additionally, the control flag can be derived at the coding block level to adaptively determine whether to apply the model-based inter prediction technique.

[0093] In some examples, a first control flag of the model-based inter prediction technique associated with the current block is inherited from a second control flag of the model-based inter prediction technique associated with at least one other coding block. In some examples, the encoded information of the model-based inter prediction technique can be derived from adjacent coding blocks, non-adjacent coding blocks, or coding blocks that store the encoded information in a buffer.

[0094] Note that, in the present disclosure, without loss of generality, in an example, at least one term parameter refers to parameters α i and β for determining a linear or non-linear model for deriving a prediction block. The term template type refers to different template shapes, such as but not limited to Figure 8One of the template types for parameter derivation in non - linear or linear formulas as shown.

[0095] According to some aspects of the present disclosure, an additional clipping process can be applied together with the derivation of parameters in a model - based prediction method. Note that any suitable clipping technique can be applied in the additional clipping process. In some examples, the additional clipping process uses an adaptive clipping technique.

[0096] Generally, the clipping range determines the minimum and maximum values of the samples after the clipping process. For example, the clipping range (x, y) includes the upper and lower limits of the clipping range, or the minimum value x and the maximum value y. For example, when the bit depth is set to 8 bits, each sample of the subject signal can be clipped within the clipping range of (0, 255), and the clipping process of the input z can be represented by Equation (1):

[0097] In an example of adaptive clipping, at least two different clipping ranges can be used. A first clipping operation can be performed on a first part of the video data based on a first clipping range among at least two different clipping ranges; and a second clipping operation can be performed on a second part of the video data based on a second clipping range different from the first clipping range.

[0098] Note that in some examples, mapping (also known as shaping) techniques are used in video coding. Mapping techniques are used to better utilize the distribution of sample codeword values in pictures in the video. The mapping function used in the mapping technique can change the value range of the samples.

[0099] Mapping and inverse mapping can be performed outside the decoding loop. For example, the mapping process can be applied to the input samples of the encoder before core coding; the inverse mapping process can be applied to the output samples of the decoder on the decoder side.

[0100] Mapping and inverse mapping can also be performed within the decoding loop. In some examples, a technique called luminance mapping with chroma scaling (LMCS) is used for in - loop shaping. The mapping (also known as shaping) of the luminance or chroma signal is implemented inside the encoding loop.

[0101] In some examples of LMCS, at the encoder, after applying the mapping function to the original samples (to be encoded) and the predicted samples respectively, a residual signal before quantization is generated. For example, the mapping function is applied to the original samples to generate mapped original samples, and the mapping function is also applied to the predicted samples to generate mapped predicted samples. The residual signal is calculated as the difference between the mapped original samples and the mapped predicted samples. The residual of a block can be transformed into transform coefficients, and the transform coefficients can be quantized. At the decoder, the residual can be calculated through dequantization and inverse transformation. Additionally, the mapping function is applied to the predicted samples to generate mapped predicted samples, and then the mapped predicted samples are combined with the residual signal to form mapped reconstructed samples. Furthermore, the inverse mapping function is applied to the mapped reconstructed samples to generate reconstructed samples.

[0102] According to some aspects of the present disclosure, a clipping process (also referred to as an additional clipping process) can be applied together with model-based prediction techniques, such as during model parameter derivation and / or reconstruction using a derived model. For example, the encoder / decoder can perform at least one clipping operation on at least one of the current template and the reference template to obtain clipped template samples, and derive at least one parameter value of at least one parameter of the model based on the clipped template samples.

[0103] In some aspects, the clipping process can be applied to no template / one template / or two templates before model parameter derivation. In some aspects, the clipping process can be applied after applying the derived model.

[0104] In some examples, before model parameter derivation, an additional clipping process is applied to one of the current template of the current block and the reference template of the reference block. The result of the clipping process can be used as an input for model parameter derivation to derive a model for model-based prediction techniques.

[0105] In some examples, before model parameter derivation, an additional clipping process is applied to both the current template of the current block and the reference template of the reference block. The result of the clipping process can be used as an input for model parameter derivation to derive a model for model-based prediction techniques.

[0106] In some examples, before model parameter derivation, the additional clipping process is not applied to either the current template of the current block or the reference template of the reference block, but is applied to the samples of the reference block. The result of the clipping process can be input into the model to calculate the predicted samples of the current block.

[0107] In some examples, before model parameter derivation, the additional clipping process is not applied to either the current template of the current block or the reference template of the reference block, but is applied to the predicted block of the current block, which is the result of applying the derived model to the reference block.

[0108] In some embodiments, before model parameter derivation, an additional cropping process is applied to the current template of the current block and the reference template of the reference block without any additional conditions. In some examples, the additional cropping process is always applied to these two templates regardless of whether the current template of the current block and the reference template of the reference block are in the same domain (value range).

[0109] In some embodiments, if and only if both the current template of the current block and the reference template of the reference block are in the original domain (e.g., at least one mapping function without a transformed domain is applied to any template), before model parameter derivation, an additional cropping process is applied to the current template of the current block and the reference template of the reference block. In an example, when some mapping methods (such as LMCS in VVC) are enabled in the decoder, if and only if the current template and the reference template are collected from at least one picture area not affected by LMCS, the additional cropping process is applied to the current template of the current block and the reference template of the reference block. For example, LMCS is not applied to the area including the current template and the reference template.

[0110] In some embodiments, it is determined which one of the current template of the current block and / or the reference template of the reference block has a domain different from the original picture. For at least one template having a domain different from the original picture, an additional cropping process is applied before model parameter derivation. For example, an inspection process of the current template and the reference template can be performed to determine whether the current template and / or the reference template is affected by a mapping function. When the current template of the current block has a domain different from the original picture (e.g., is affected by a mapping function), an additional cropping process is applied to the current template before model parameter derivation. When the reference template of the reference block has a domain different from the original picture (e.g., is affected by a mapping function), an additional cropping process is applied to the reference template before model parameter derivation. When both the current template and the reference template have a domain different from the original picture (e.g., are affected by a mapping function), an additional cropping process is applied to the current template and the reference template before model parameter derivation.

[0111] In some examples, one of the current template of the current block and the reference template of the reference block has a domain different from the original picture and is called a different-domain template. In an example, after applying the inverse mapping function to the different-domain template, an additional cropping process is applied to the different-domain template. In the example of LMCS, at the decoder, the residual can be calculated by dequantization and inverse transformation. In addition, a mapping function is applied to the predicted sample to generate a mapped prediction sample, and then the mapped prediction sample is combined with the residual to form a mapped reconstruction sample. In addition, the inverse mapping function is applied to the mapped reconstruction sample of the different-domain template to generate a reconstruction sample of the different-domain template, and an additional cropping process is applied to the reconstruction sample of the different-domain template. Then the cropped reconstruction sample is used as the input for model parameter derivation.

[0112] In another example, an additional cropping process is applied to different domain templates before applying the inverse mapping function to the different domain templates. In the example of LMCS, at the decoder, the residual can be calculated by dequantization and inverse transformation. Additionally, a mapping function is applied to the predicted samples to generate mapped predicted samples, and then the mapped predicted samples are combined with the residual to form mapped reconstruction samples. An additional cropping process is applied to the mapped reconstruction samples of different domain templates to generate cropped mapped reconstruction samples of different domain templates. Additionally, an inverse mapping function is applied to the cropped mapped reconstruction samples of different domain templates to generate reconstruction samples of different domain templates. Then the reconstruction samples are used as the input for model parameter derivation. In this example, there is an additional step of inverse mapping the cropping parameters using the same mapping function as the template. In the example, the sample mapping function is applied to the upper and lower limit values of the first cropping range of the reconstruction samples to obtain the second cropping range of the mapped reconstruction samples. The second cropping range is used for the additional cropping process to be applied to the mapped reconstruction samples of different domain templates.

[0113] In another example, an additional cropping process is applied to different domain templates instead of applying the inverse mapping function to the template. For example, the inverse mapping function is not applied to the template at all. In the example of LMCS, at the decoder, the residual can be calculated by dequantization and inverse transformation. Additionally, a mapping function is applied to the predicted samples to generate mapped predicted samples, and then the mapped predicted samples are combined with the residual to form mapped reconstruction samples. An additional cropping process is applied to the mapped reconstruction samples of different domain templates to generate reconstruction samples of different domain templates. Then the reconstruction samples are used as the input for model parameter derivation.

[0114] In some embodiments, it is determined which one of the current template of the current block and / or the reference template of the reference block has the same domain as the original picture, and then an additional cropping process is applied to at least one template having the same domain as the original picture before model parameter derivation. For example, an inspection process of the current template and the reference template can be performed to determine whether the current template and / or the reference template is affected by the mapping function. When the current template of the current block has the same domain as the original picture (e.g., not affected by the mapping function), the additional cropping process is applied to the current template before model parameter derivation. When the reference template of the reference block has the same domain as the original picture (e.g., not affected by the mapping function), the additional cropping process is applied to the reference template before model parameter derivation. When both the current template and the reference template have the same domain as the original picture (e.g., not affected by the mapping function), the additional cropping process is applied to both the current template and the reference template before model parameter derivation.

[0115] In some examples, one of the two templates has the same domain as the original picture and is referred to as the same-domain template. In the example, before the model parameters are generated, an additional cropping process is applied to the same-domain template. For example, the reconstructed samples of the same-domain template are not affected by any mapping process, and an additional cropping process is applied to the reconstructed samples of the same-domain template to generate cropped reconstructed samples, and the cropped reconstructed samples are used as the input for model parameter derivation.

[0116] In another example, an additional step of forward mapping is applied to the same-domain template, followed by an additional cropping process. This additional step of forward mapping uses the same mapping function as the template. Before the cropping process, the same mapping function is applied to the cropping range. For example, the mapping function is applied to the same-domain template to generate mapped template samples, and then an additional cropping is applied to the mapped template samples to generate cropped mapped samples. In the example, the cropped mapped samples are used as the input for model parameter derivation. In the example, a first cropping range is provided in the value range of the original picture, and the mapping function is applied to the upper and lower limits of the first cropping range to determine the upper and lower limits of the second cropping range, and the second cropping range is used in the additional cropping applied to the mapped template samples.

[0117] In some embodiments, the cropping range of the additional cropping is signaled at an appropriate level (such as the picture parameter set (PPS) level, picture header level, sub-picture level, slice header level, and / or block level).

[0118] In some embodiments, the cropping range is derived on the decoder side. The cropping derivation can depend on the input of samples from the current template or the reference template. For example, the cropping range derivation is based on the average sample value of one of the current template and the reference template.

[0119] In some examples, the reference template is indicated by the motion vectors between different frames (for inter-frame prediction). In some examples, the reference template is indicated by the block vectors within the same frame (for intra-frame prediction).

[0120] Figure 9 A flowchart showing an overview process (900) according to an embodiment of the present disclosure is presented. Process (900) can be used in a device such as a video decoder. In various embodiments, process (900) is executed by a processing circuit, such as a processing circuit that performs the functions of video decoder (110), a processing circuit that performs the functions of video decoder (210), etc. In some embodiments, process (900) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes process (900). The process starts at (S901) and proceeds to (S910).

[0121] At (S910), a bitstream of encoded information of a picture sequence is received, the encoded information indicating prediction of a current block in a current picture using a model-based prediction technique, the model-based prediction technique generating prediction samples of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block.

[0122] At (S920), at least one cropping operation is performed on at least one of the current template and the reference template to obtain cropped template samples.

[0123] At (S930), at least one parameter value of the at least one parameter of the model is derived based on the cropped template samples.

[0124] At (S940), at least one prediction sample of the current block is generated by employing the model with the at least one parameter set to the at least one parameter value.

[0125] In some examples, a cropping operation is performed on reference samples of the reference block to obtain cropped reference samples; and the model is applied to the cropped reference samples to generate prediction samples of the current block. In some examples, after generating the prediction samples by employing the model, a cropping operation is performed on the prediction samples of the current block.

[0126] According to one aspect of the present disclosure, a first cropping operation is applied to the current template to obtain cropped current template samples; and a second cropping operation is applied to the reference template to obtain cropped reference template samples. In some examples, it is not checked whether the reference template and the current template are in the same value range.

[0127] According to one aspect of the present disclosure, it is checked whether the current template and the reference template are affected by a mapping function (e.g., which may change the value range). In some examples, when neither the current template nor the reference template is affected by the mapping function, a first cropping operation is performed on the current template to obtain cropped current template samples, and a second cropping operation is performed on the reference template to obtain cropped reference template samples. In some examples, it is checked whether the mapping function is enabled. When the mapping function is enabled, it is checked whether the current template and the reference template are affected by the mapping function. When the current template and the reference template are not affected by the mapping function, the first cropping operation and the second cropping operation are performed.

[0128] According to another aspect of the present disclosure, check which one of the current template and / or the reference template has a value range different from that of the original picture of the current picture. When one of the current template and the reference template has a value range different from that of the original picture, perform a cropping operation on the template to obtain a cropped template sample of the template. In some examples, after applying the inverse mapping function to the template, the cropping operation is applied to the template. In some examples, before applying the inverse mapping function to the template, the cropping operation is applied to the template, and the cropping range of the cropping operation is calculated according to the mapping function associated with the inverse mapping function. In some examples, the cropping operation is applied to the template without applying the inverse mapping function to the template before or after the cropping operation.

[0129] According to another aspect of the present disclosure, check which one of the current template and / or the reference template has the same value range as the original picture of the current picture. When one of the current template and the reference template has the same value range as the original picture, before deriving the at least one parameter value, perform a cropping operation on the template to obtain a cropped template sample of the template. In some examples, after applying the mapping function to the template, the cropping operation is applied to the template. In an example, the cropping range of the cropping operation is calculated according to the mapping function.

[0130] According to another aspect of the present disclosure, decode a signal indicating a cropping range from the bitstream. The signal can be signaled at one of a picture level, a sub-picture level, a slice level, a tile level, and / or a block level.

[0131] According to another aspect of the present disclosure, derive the cropping range of the cropping operation based on the sample values (e.g., the average sample values of the current template and / or the reference template) of the current template and / or the reference template.

[0132] Note that the reference template can be located according to the motion vector for inter prediction and / or the block vector for intra prediction.

[0133] Then, the process proceeds to (S999) and ends.

[0134] The process (900) can be appropriately adjusted. At least one step in the process (900) can be modified and / or omitted. At least one other step can be added. Any suitable order of implementation can be used.

[0135] Figure 10A flowchart showing an overview process (1000) according to an embodiment of the present disclosure is presented. The process (1000) can be used in a video encoder. In various embodiments, the process (1000) is executed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. In some embodiments, the process (1000) is implemented by software instructions. Thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1000). The process starts at (S1001) and proceeds to (S1010).

[0136] At (S1010), a bitstream of encoded information of a picture sequence is received. The encoded information indicates that a model-based prediction technique is used to predict a current block in a current picture. The model-based prediction technique generates prediction samples for the current block based on a model, where at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block.

[0137] At (S1020), at least one parameter value of the at least one parameter of the model is derived based on the current template and the reference template.

[0138] At (S1030), at least one prediction sample for the current block is generated by adopting the model with the at least one parameter set to the at least one parameter value.

[0139] At (S1040), after the prediction samples for the current block are generated by adopting the model, a clipping operation is applied to the prediction samples of the current block.

[0140] Then, the process proceeds to (S1099) and ends.

[0141] The process (1000) can be adjusted appropriately. At least one step in the process (1000) can be modified and / or omitted. At least one other step can be added. Any suitable order of implementation can be used.

[0142] Figure 11 A flowchart showing an overview process (1100) according to an embodiment of the present disclosure is presented. The process (1100) can be used in a video decoder. In various embodiments, the process (1100) is executed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some embodiments, the process (1100) is implemented by software instructions. Thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1100). The process starts at (S1101) and proceeds to (S1110).

[0143] In (S1110), it is determined to encode a current block in a current picture using a model-based prediction technique that generates a prediction sample of the current block based on a model, where at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block.

[0144] In (S1120), at least one cropping operation is performed on at least one of the current template and the reference template to obtain cropped template samples.

[0145] In (S1130), at least one parameter value of the at least one parameter of the model is derived based on the cropped template samples.

[0146] In (S1140), the current block is encoded based on the model with the at least one parameter set to the at least one parameter value.

[0147] In some examples, a cropping operation is performed on reference samples of the reference block to obtain cropped reference samples; and the model is applied to the cropped reference samples to generate a prediction sample of the current block. In some examples, after generating the prediction sample by adopting the model, a cropping operation is performed on the prediction sample of the current block.

[0148] According to one aspect of the present disclosure, a first cropping operation is applied to the current template to obtain cropped current template samples; and a second cropping operation is applied to the reference template to obtain cropped reference template samples. In some examples, it is not checked whether the reference template and the current template are in the same value range.

[0149] According to one aspect of the present disclosure, it is checked whether the current template and the reference template are affected by a mapping function (e.g., which may change the value range). In some examples, when neither the current template nor the reference template is affected by the mapping function, a first cropping operation is performed on the current template to obtain cropped current template samples, and a second cropping operation is performed on the reference template to obtain cropped reference template samples. In some examples, it is checked whether the mapping function is enabled. When the mapping function is enabled, it is checked whether the current template and the reference template are affected by the mapping function. When the current template and the reference template are not affected by the mapping function, the first cropping operation and the second cropping operation are performed.

[0150] According to another aspect of the present disclosure, check which one of the current template and / or the reference template has a value range different from that of the original picture of the current picture. When one of the current template and the reference template has a value range different from that of the original picture, perform a cropping operation on the template to obtain a cropped template sample of the template. In some examples, after applying the inverse mapping function to the template, the cropping operation is applied to the template. In some examples, before applying the inverse mapping function to the template, the cropping operation is applied to the template, and the cropping range of the cropping operation is calculated according to the mapping function associated with the inverse mapping function. In some examples, the cropping operation is applied to the template without applying the inverse mapping function to the template before or after the cropping operation.

[0151] According to another aspect of the present disclosure, check which one of the current template and / or the reference template has the same value range as the original picture of the current picture. When one of the current template and the reference template has the same value range as the original picture, before deriving the at least one parameter value, perform a cropping operation on the template to obtain a cropped template sample of the template. In some examples, after applying the mapping function to the template, the cropping operation is applied to the template. In an example, the cropping range of the cropping operation is calculated according to the mapping function.

[0152] According to another aspect of the present disclosure, encode (signal) a signal indicating the cropping range into the bitstream. This signal can be signaled at one level among the picture level, sub-picture level, slice level, tile level, and / or block level.

[0153] According to another aspect of the present disclosure, derive the cropping range of the cropping operation based on the sample values of the current template and / or the reference template (e.g., the average sample values of the current template and / or the reference template).

[0154] Note that the reference template can be located according to the motion vector for inter prediction and / or the block vector for intra prediction.

[0155] Then, the process proceeds to (S1199) and ends.

[0156] The process (1100) can be appropriately adjusted. At least one step in the process (1100) can be modified and / or omitted. At least one other step can be added. Any suitable order of implementation can be used.

[0157] Figure 12A flowchart showing an overview process (1200) according to an embodiment of the present disclosure is shown. The process (1200) can be used in a video decoder. In various embodiments, the process (1200) is executed by a processing circuit, such as a processing circuit that executes the functions of a video decoder (120), a processing circuit that executes the functions of a video decoder (210), etc. In some embodiments, the process (1200) is implemented by software instructions. Thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1200). The process starts at (S1201) and proceeds to (S1210).

[0158] At (S1210), it is determined to encode a current block in a current picture using a model-based prediction technique, where the model-based prediction technique generates prediction samples of the current block based on a model, and at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block.

[0159] At (S1220), at least one parameter value of the at least one parameter of the model is derived based on the current template and the reference template.

[0160] At (S1230), at least one prediction sample of the current block is generated by employing the model, where the at least one parameter is set to the at least one parameter value.

[0161] At (S1240), after generating the prediction samples by employing the model, a clipping operation is performed on the prediction samples of the current block.

[0162] At (S1250), the current block is encoded based on the model with the at least one parameter set to the at least one parameter value.

[0163] Then, the process proceeds to (S1299) and ends.

[0164] The process (1200) can be adjusted appropriately. At least one step in the process (1200) can be modified and / or omitted. At least one other step can be added. Any suitable order of implementation can be used.

[0165] According to one aspect of the present disclosure, a method for processing visual media data is provided. In this method, a bitstream of visual media data is processed according to format rules. For example, the bitstream can be a bitstream decoded / encoded by any decoding and / or encoding method described herein. The format rules can specify at least one constraint of the bitstream and / or at least one process to be performed by the decoder and / or encoder.

[0166] For example, the bitstream includes encoded information of at least one picture, the encoded information indicating prediction of a current block in a current picture using a model-based prediction technique, the model-based prediction technique generating prediction samples of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block. The formatting rules specify: performing at least one cropping operation on at least one of the current template and the reference template to obtain cropped template samples; deriving at least one parameter value of the at least one parameter of the model based on the cropped template samples; and generating at least one prediction sample of the current block by using the model with the at least one parameter set to the at least one parameter value.

[0167] The above techniques can be implemented as computer software by computer-readable instructions and physically stored in at least one computer-readable medium. For example, Figure 13 A computer system (1300) is shown, which is adapted to implement certain embodiments of the disclosed subject matter.

[0168] The computer software can be encoded by any suitable machine code or computer language, creating code including instructions through mechanisms such as assembly, compilation, and linking, and the instructions can be directly executed by at least one computer central processing unit (CPU), graphics processing unit (GPU), etc., or executed through decoding, microcode, etc.

[0169] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0170] Figure 13 The components shown for the computer system (1300) are exemplary and are not used to impose any limitations on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Nor should the configuration of the components be construed as having any dependence on or requirement for any one component or combination thereof shown in the exemplary embodiments of the computer system (1300).

[0171] A computer system (1300) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to the input of at least one human user through tactile inputs (such as keyboard input, swiping, data glove movement), audio inputs (such as voice, applause), visual inputs (such as gestures), and olfactory inputs (not shown). The human-machine interface device may also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0172] The human-machine interface input device may include at least one of the following (only one is shown): keyboard (1301), mouse (1302), touchpad (1303), touch screen (1310), data glove (not shown), joystick (1305), microphone (1306), scanner (1307), camera (1308).

[0173] The computer system (1300) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of at least one human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (such as tactile feedback through the touch screen (1310), data glove (not shown), or joystick (1305), but there may also be tactile feedback devices that do not serve as input devices), audio output devices (such as speakers (1309), headphones (not shown)), visual output devices (such as screens (1310) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light-emitting diode screens, each of which may or may not have touch screen input functionality and each of which may or may not have tactile feedback functionality - some of which may output two-dimensional visual output or output above three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).

[0174] The computer system (1300) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical discs with CD / DVD (CD / DVD ROM / RW) (1320) or similar media (1321), thumb drives (1322), removable hard disk drives or solid state drives (1323), traditional magnetic media such as tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.

[0175] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the disclosed subject matter does not include a transmission medium, a carrier wave, or other transitory signals.

[0176] The computer system (1300) may also include an interface (1354) to at least one communication network (1355). For example, the network can be wireless, wired, optical. The network can also be a local area network, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, etc. The network also includes local area networks such as Ethernet, wireless local area network, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), wired or wireless wide area digital television networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANBus), etc. Some networks typically require an external network interface adapter for connection to certain common data ports or peripheral buses (1349) (e.g., the USB port of the computer system (1300)); other systems are typically integrated into the core of the computer system (1300) by connecting to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks, the computer system (1300) can communicate with other entities. The communication can be unidirectional, only for receiving (e.g., wireless television), unidirectional only for sending (e.g., CAN bus to certain CAN bus devices), or bidirectional, e.g., via a local or wide area digital network to other computer systems. Each of the above networks and network interfaces may use certain protocols and protocol stacks.

[0177] The above-mentioned human-machine interface device, human-accessible storage device, and network interface can be connected to the core (1340) of the computer system (1300).

[0178] The core (1340) may include at least one central processing unit (CPU) (1341), a graphics processing unit (GPU) (1342), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1343), a hardware accelerator for specific tasks (1344), a graphics adapter (1350), etc. These devices, as well as read-only memory (ROM) (1345), random access memory (1346), internal mass storage (such as an internal non-user-accessible hard disk drive, solid state drive, etc.) (1347), etc. may be connected via a system bus (1348). In some computer systems, the system bus (1348) may be accessed in the form of at least one physical plug so as to allow for expansion via additional central processing units, graphics processing units, etc. Peripheral devices may be directly attached to the system bus (1348) of the core or connected via a peripheral bus (1349). In an example, a screen (1310) may be connected to the graphics adapter (1350). The architecture of the peripheral bus includes external peripheral component interconnect PCI, universal serial bus USB, etc.

[0179] The CPU (1341), GPU (1342), FPGA (1343), and accelerator (1344) may execute certain instructions which, when combined, may constitute the above-mentioned computer code. The computer code may be stored in the ROM (1345) or RAM (1346). Transitional data may also be stored in the RAM (1346), while permanent data may be stored in, for example, the internal mass storage (1347). Fast storage and retrieval of any memory device may be achieved by using a cache memory which may be closely associated with at least one CPU (1341), GPU (1342), mass storage (1347), ROM (1345), RAM (1346), etc.

[0180] The computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure or may be of the kind well-known and available to those skilled in the field of computer software.

[0181] By way of example and not limitation, a computer system having an architecture (1300), and in particular a core (1340), can provide the functionality of a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in at least one tangible computer-readable medium. Such a computer-readable medium can be the medium associated with the user-accessible mass storage described above, as well as a specific memory of the non-volatile core (1340), such as the on-core mass storage (1347) or ROM (1345). The software implementing the various embodiments of the present disclosure can be stored in such a device and executed by the core (1340). Depending on specific needs, the computer-readable medium can include one or more storage devices or chips. The software can cause the core (1340), and in particular the processors therein (including the CPU, GPU, FPGA, etc.), to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in the RAM (1346) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality that is logically hardwired or otherwise included in circuitry (e.g., an accelerator (1344)), which can operate in place of or in conjunction with the software to execute specific processes or specific parts of specific processes described herein. In appropriate cases, references to software can include logic, and vice versa. In appropriate cases, references to a computer-readable medium can include circuitry (such as an integrated circuit (IC)) that stores the software for execution, circuitry that includes the execution logic, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0182] As used in this disclosure, "at least one" or "one of" is intended to include any one or combination of the recited elements. For example, a reference to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A through C is intended to include only A, only B, only C, or any combination thereof. A reference to one of A or B and one of A and B is intended to include A or B or (A and B). Where applicable, the use of "one of" does not exclude any combination of the recited elements, such as when the elements are not mutually exclusive.

[0183] Although the present disclosure has been described with respect to at least two exemplary embodiments, various changes, permutations, and various equivalent substitutions of the embodiments are within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

[0184] The foregoing disclosure also includes the following features. These features can be combined in various ways and are not limited to the combinations mentioned below.

[0185] (1) A method for video decoding, the method comprising: receiving a bitstream of encoded information of a picture sequence, the encoded information indicating the use of a model-based prediction technique to predict a current block in a current picture, the model-based prediction technique generating a prediction sample of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block; performing at least one cropping operation on at least one of the current template and the reference template to obtain cropped template samples; deriving at least one parameter value of the at least one parameter of the model according to the cropped template samples; and generating at least one prediction sample of the current block by using the model with the at least one parameter set to the at least one parameter value.

[0186] (2) The method according to feature (1), the method further comprising: performing a cropping operation on a reference sample of the reference block to obtain a cropped reference sample; and applying the model to the cropped reference sample to generate a prediction sample of the current block.

[0187] (3) The method according to any one of features (1) to (2), the method further comprising: after generating the prediction sample by using the model, performing a cropping operation on the prediction sample of the current block.

[0188] (4) The method according to any one of features (1) to (3), the method further comprising: performing a first cropping operation on the current template to obtain a cropped current template sample; and performing a second cropping operation on the reference template to obtain a cropped reference template sample.

[0189] (5) The method according to any one of features (1) to (4), the method further comprising: checking whether the current template and the reference template are affected by a mapping function; and when neither the current template nor the reference template is affected by the mapping function, performing a first cropping operation on the current template to obtain a cropped current template sample, and performing a second cropping operation on the reference template to obtain a cropped reference template sample.

[0190] (6) The method according to any one of features (1) to (5), the method further comprising: checking whether the mapping function is enabled; when the mapping function is enabled, checking whether the current template and the reference template are affected by the mapping function; and when the current template and the reference template are not affected by the mapping function, performing the first cropping operation and the second cropping operation.

[0191] (7) The method according to any one of features (1) to (6), the method further comprising: checking which one of the current template and / or the reference template has a value range different from that of the original picture of the current picture; and when one of the current template and the reference template has a value range different from that of the original picture, performing a cropping operation on the template to obtain a cropped template sample of the template.

[0192] (8) The method according to any one of features (1) to (7), the method further comprising: applying the cropping operation to the template after applying the inverse mapping function to the template.

[0193] (9) The method according to any one of features (1) to (8), the method further comprising: applying the cropping operation to the template before applying the inverse mapping function to the template.

[0194] (10) The method according to any one of features (1) to (9), the method further comprising: calculating a cropping range of the cropping operation according to a mapping function associated with the inverse mapping function.

[0195] (11) The method according to any one of features (1) to (10), the method further comprising: applying the cropping operation to the template without applying the inverse mapping function to the template before or after the cropping operation.

[0196] (12) The method according to any one of features (1) to (11), the method further comprising: checking which one of the current template and / or the reference template has a value range the same as that of the original picture of the current picture; and when one of the current template and the reference template has a value range the same as that of the original picture, performing a cropping operation on the template to obtain a cropped template sample of the template before deriving the at least one parameter value.

[0197] (13) The method according to any one of features (1) to (12), the method further comprising: applying the cropping operation to the template after applying the mapping function to the template.

[0198] (14) The method according to any one of features (1) to (13), the method further comprising: calculating a cropping range of the cropping operation according to the mapping function.

[0199] (15) The method according to any one of features (1) to (14), the method comprising: decoding a signal indicating a cropping range, the signal being one of a picture level, a sub-picture level, a slice level, a tile level, and / or a block level.

[0200] (16) The method according to any one of features (1) to (15), the method comprising: deriving a cropping range of the cropping operation based on an average sample value of the current template and / or the reference template.

[0201] (17) The method according to any one of features (1) to (16), the method further comprising: locating the reference template according to a motion vector for inter-frame prediction and / or a block vector for intra-frame prediction.

[0202] (18) A method for video coding, the method comprising: determining to encode a current block in a current picture using a model-based prediction technique, the model-based prediction technique generating a prediction sample of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, the model comprising at least one parameter derived based on a current template of the current block and a reference template of the reference block; performing at least one cropping operation on at least one of the current template and the reference template to obtain cropped template samples; deriving at least one parameter value of the at least one parameter of the model according to the cropped template samples; and encoding the current block based on the model with the at least one parameter set to the at least one parameter value.

[0203] (19) The method according to feature (18), the method further comprising: performing a cropping operation on a reference sample of the reference block to obtain a cropped reference sample; and applying the model to the cropped reference sample to generate a prediction sample of the current block.

[0204] (20) The method according to any one of features (18) to (19), the method further comprising: generating at least one prediction sample of the current block by adopting the model with the at least one parameter set to the at least one parameter value; and performing a cropping operation on the prediction sample of the current block after generating the prediction sample by adopting the model.

[0205] (21) The method according to any one of features (18) to (20), the method further comprising: performing a first cropping operation on the current template to obtain a cropped current template sample, and performing a second cropping operation on the reference template to obtain a cropped reference template sample.

[0206] (22) The method according to any one of features (18) to (21), the method further comprising: checking whether the current template and the reference template are affected by a mapping function; and when neither the current template nor the reference template is affected by the mapping function, performing a first cropping operation on the current template to obtain a cropped current template sample, and performing a second cropping operation on the reference template to obtain a cropped reference template sample.

[0207] (23) The method according to any one of features (18) to (22), the method further comprising: checking whether the mapping function is enabled; when the mapping function is enabled, checking whether the current template and the reference template are affected by the mapping function; and when the current template and the reference template are not affected by the mapping function, performing the first cropping operation and the second cropping operation.

[0208] (24) The method according to any one of features (18) to (23), the method further comprising: checking which of the current template and / or the reference template has a value range different from that of the original picture of the current picture; and when one of the current template and the reference template has a value range different from that of the original picture, performing a cropping operation on the template to obtain a cropped template sample of the template.

[0209] (25) The method according to any one of features (18) to (24), the method further comprising: applying the cropping operation to the template after applying an inverse mapping function to the template.

[0210] (26) The method according to any one of features (18) to (25), the method further comprising: applying the cropping operation to the template before applying an inverse mapping function to the template.

[0211] (27) The method according to any one of features (18) to (26), the method comprising: calculating a cropping range of the cropping operation according to a mapping function associated with the inverse mapping function.

[0212] (28) The method according to any one of features (18) to (27), the method comprising: applying the cropping operation to the template without applying an inverse mapping function to the template before or after the cropping operation.

[0213] (29) The method according to any one of features (18) to (28), the method further comprising: checking which one of the current template and / or the reference template has a value range identical to that of the original picture of the current picture; and when one of the current template and the reference template has a value range identical to that of the original picture, before deriving the at least one parameter value, performing a cropping operation on the template to obtain a cropped template sample of the template.

[0214] (30) The method according to any one of features (18) to (29), the method further comprising: after applying the mapping function to the template, applying the cropping operation to the template.

[0215] (31) The method according to any one of features (18) to (30), the method further comprising: calculating a cropping range of the cropping operation according to the mapping function.

[0216] (32) The method according to any one of features (18) to (31), the method further comprising: encoding a signal indicating the cropping range, the signal being signaled at one of a picture level, a sub - picture level, a slice level, a tile level, and / or a block level.

[0217] (33) The method according to any one of features (18) to (32), the method further comprising: deriving the cropping range of the cropping operation based on an average sample value of the current template and / or the reference template.

[0218] (34) The method according to any one of features (18) to (33), the method further comprising: locating the reference template according to a motion vector for inter - frame prediction and / or a block vector for intra - frame prediction.

[0219] (35) A method for processing visual media data, the method comprising: the bitstream includes encoded information of a current block in a current picture, the encoded information indicating prediction of the current block in the current picture using a model-based prediction technique, the model-based prediction technique generating a prediction sample of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, the model including at least one parameter derived based on a current template of the current block and a reference template of the reference block; and processing a bitstream of visual media data according to formatting rules; the formatting rules specifying: performing at least one cropping operation on at least one of the current template and the reference template to obtain cropped template samples; deriving at least one parameter value of the at least one parameter of the model according to the cropped template samples; and generating at least one prediction sample of the current block by employing the model with the at least one parameter set to the at least one parameter value.

[0220] (36) A method for video decoding, the method comprising: receiving a bitstream of encoded information of a picture sequence, the encoded information indicating prediction of a current block in a current picture using a model-based prediction technique, the model-based prediction technique generating a prediction sample of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, the model including at least one parameter derived based on a current template of the current block and a reference template of the reference block; deriving at least one parameter value of the at least one parameter of the model according to the current template and the reference template; generating at least one prediction sample of the current block by employing the model with the at least one parameter set to the at least one parameter value; and performing a cropping operation on the prediction sample of the current block after generating the prediction sample by employing the model.

[0221] (37) A method for video encoding, the method comprising: determining to encode a current block in a current picture using a model-based prediction technique, the model-based prediction technique generating a prediction sample of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, the model including at least one parameter derived based on a current template of the current block and a reference template of the reference block; deriving at least one parameter value of the at least one parameter of the model according to the current template and the reference template; generating at least one prediction sample of the current block by employing the model with the at least one parameter set to the at least one parameter value; performing a cropping operation on the prediction sample of the current block after generating the prediction sample by employing the model; and encoding the current block based on the model with the at least one parameter set to the at least one parameter value.

[0222] (38) A video decoding device, comprising a processing circuit configured to execute the method according to any one of features (1) to (17) and feature (36).

[0223] (39) A video encoding device, comprising a processing circuit configured to execute the method according to any one of features (18) to (34) and feature (37).

[0224] (30) A non - volatile computer - readable storage medium for storing instructions which, when executed by at least one processor, cause the at least one processor to execute the method according to any one of features (1) to (37).

Claims

1. A method for video decoding, characterized in that, Comprising: Receiving a bitstream of encoded information of a picture sequence, the encoded information indicating prediction of a current block in a current picture using a model-based prediction technique, the model-based prediction technique generating prediction samples of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block; Performing at least one cropping operation on at least one of the current template and the reference template to obtain cropped template samples; Deriving at least one parameter value of the at least one parameter of the model according to the cropped template samples; And Generating at least one prediction sample of the current block by employing the model with the at least one parameter set to the at least one parameter value.

2. The method according to claim 1, wherein Further comprising: Performing a cropping operation on reference samples of the reference block to obtain cropped reference samples; And Applying the model to the cropped reference samples to generate prediction samples of the current block.

3. The method according to claim 1, wherein Further comprising: After generating the prediction samples by employing the model, performing a cropping operation on the prediction samples of the current block.

4. The method according to claim 1, wherein The performing of the at least one cropping operation further comprises: Performing a first cropping operation on the current template to obtain cropped current template samples; and Performing a second cropping operation on the reference template to obtain cropped reference template samples.

5. The method according to claim 1, wherein Further comprising: Checking whether the current template and the reference template are affected by a mapping function; And When neither the current template nor the reference template is affected by the mapping function, Performing a first cropping operation on the current template to obtain cropped current template samples, and Performing a second cropping operation on the reference template to obtain cropped reference template samples.

6. The method according to claim 5, wherein Further comprising: Checking whether the mapping function is enabled; When the mapping function is enabled, checking whether the current template and the reference template are affected by the mapping function; And When the current template and the reference template are not affected by the mapping function, performing the first cropping operation and the second cropping operation.

7. The method according to claim 1, wherein Further comprising: Checking which one of the current template and / or the reference template has a value range different from that of an original picture of the current picture; And When one of the current template and the reference template has a value range different from that of the original picture, performing a cropping operation on the template to obtain cropped template samples of the template.

8. The method according to claim 7, characterized in that The performing of the cropping operation on the template further comprises: After applying an inverse mapping function to the template, applying the cropping operation to the template.

9. The method according to claim 7, wherein The performing of the cropping operation on the template further comprises: Before applying an inverse mapping function to the template, applying the cropping operation to the template.

10. The method according to claim 9, characterized in that Further comprising: Calculating a cropping range of the cropping operation according to a mapping function associated with the inverse mapping function.

11. The method according to claim 7, wherein The performing of the cropping operation on the template further comprises: Apply the cropping operation to the template without applying an inverse mapping function to the template before or after the cropping operation.

12. The method according to claim 1, wherein Further comprising: Checking which of the current template and / or the reference template has the same value range as the original picture of the current picture; And When one of the current template and the reference template has the same value range as the original picture, performing a cropping operation on the template before deriving the at least one parameter value to obtain a cropped template sample of the template.

13. The method according to claim 12, characterized in that, Further comprising: After applying the mapping function to the template, applying the cropping operation to the template.

14. A method for video encoding, characterized in that, Comprising: Determining to encode a current block in a current picture using a model-based prediction technique that generates a prediction sample of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block; Performing at least one cropping operation on at least one of the current template and the reference template to obtain a cropped template sample; Deriving at least one parameter value of the at least one parameter of the model according to the cropped template sample; And Encoding the current block based on the model with the at least one parameter set to the at least one parameter value.

15. A method for processing visual media data, characterized in that, The method comprises: Processing a bitstream of visual media data according to format rules; The bitstream includes encoded information of a current block in a current picture, the encoded information indicating a prediction of the current block in the current picture using a model-based prediction technique that generates a prediction sample of the current block based on a model, wherein at least one reconstructed sample of a reference block is input into the model, and the model includes at least one parameter derived based on a current template of the current block and a reference template of the reference block; and The format rules specify: Performing at least one cropping operation on at least one of the current template and the reference template to obtain a cropped template sample; Deriving at least one parameter value of the at least one parameter of the model according to the cropped template sample; and Generating at least one prediction sample of the current block by employing the model with the at least one parameter set to the at least one parameter value.