SEI information for film particle synthesis

By introducing film particle characteristic SEI messages and alpha channel information SEI messages, combined with labeled area SEI messages, film particle synthesis is dynamically controlled, and the region and intensity selection problems of film particle processing in mixed content scenarios are solved, improving coding efficiency and user experience.

CN120266479APending Publication Date: 2025-07-04TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005013.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-09-25
Filing Date
2024-09-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing video encoding technologies are difficult to achieve efficient compression while maintaining the effect of film particles, especially in mixed content scenarios, where the area and intensity of film particles cannot be effectively identified and applied.

Method used

By introducing film particle characteristics assisted enhancement information SEI messages and alpha channel information SEI messages, combined with labeled area SEI messages, dynamically control the application of film particle synthesis, and use alpha map sample values to determine whether and how to apply film particle processing.

Benefits of technology

It realizes the selective application of film particles to different regions in video encoding, improves coding efficiency and user experience, and maintains the authenticity and clarity of artistic effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005414838190000151
    Figure BDA0005414838190000151
  • Figure BDA0005414838190000152
    Figure BDA0005414838190000152
  • Figure BDA0005414838190000156
    Figure BDA0005414838190000156
Patent Text Reader

Abstract

A method and apparatus includes computer code configured to cause at least one processor to acquire video data including at least one encoded picture; reconstructing an encoded picture, the encoded picture comprising a film particle characteristic auxiliary enhancement information (SEI) message, an alpha channel information (SEI) message, and at least first and second samples and at least third and fourth alpha map samples, the first sample is spatially co-located with the third alpha map sample, and the second sample is spatially co-located with the first sample; applying film particle synthesis to the first sample based on the first value of the third sample; and determining that film particle synthesis is not applied to the second sample based on a second value of the fourth sample, wherein the first value and the second value are different from each other.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 540,559, filed on September 26, 2023; U.S. Provisional Application No. 63 / 542,074, filed on October 2, 2023; U.S. Provisional Application No. 63 / 599,457, filed on November 15, 2023; and U.S. Application No. 18 / 895,594, filed on September 25, 2024, the disclosures of which are hereby incorporated by reference in their entireties. Technical Field

[0002] The disclosed subject matter relates to video encoding and decoding, and more particularly, to the application of film grain or similar noise to a portion of a picture controlled by metadata including alpha or depth maps and metadata encoded in a VSEI annotation region SEI message. Background Art

[0003] For decades, video encoding and decoding using inter - picture prediction with motion compensation has been known. Uncompressed digital video can include a series of pictures, each picture having a spatial dimension such as 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (also informally called frame rate), such as 60 pictures per second or 60 Hz. Uncompressed video has high bit - rate requirements. For example, a 1080p60 4:2:0 video (1920x1080 luminance sample resolution at 60 Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video would require more than 600 GB of storage space.

[0004] One purpose of video encoding and decoding is to reduce the redundancy of the input video signal through compression. Compression can help reduce the requirements for the above - mentioned bandwidth or storage space, which can be reduced by two or more orders of magnitude in some cases. Lossless compression, lossy compression, and combinations of both can be employed. Lossless compression is a technique for reconstructing an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal can be used for the intended application. Lossy compression is widely used in video. The amount of allowable distortion depends on the application. For example, users of some consumer streaming applications can tolerate higher distortion compared to users of television applications. The achievable compression ratio reflects that higher allowed / tolerated distortion can result in a higher compression ratio.

[0005] Video encoders and decoders can utilize several broad categories of techniques, such as including: motion compensation, transformation, quantization, and entropy coding, some of which will be described below.

[0006] Film Grain Synthesis (also known as FGS) is a tool aimed at preserving the impression of original film grain, mainly for content shot with chemical film (as opposed to digital cameras) in a digital compression video environment. The digitized input material can be pre-filtered, which can remove film grain, but helps achieve better compression efficiency in subsequent encoding steps compared to input material including noise such as film grain. An artificial approximation of the film grain can be re-inserted after reconstruction. The amount and characteristics of the noise can be part of the video bitstream in the form of metadata. For example, ITU-T Recommendation H.274 includes Film Grain Characteristics SEI messages.

[0007] ITU Rec.H.274 also includes an Alpha Channel Information SEI message, which can be used to define the application of alpha channel information that may be present in, for example, an H.266 bitstream. Summary of the Invention

[0008] Including a method and apparatus, the method and apparatus including: a memory configured to store computer program code, and at least one processor configured to access the computer program code and operate according to the instructions of the computer program code. The computer program is configured to cause the processor to implement: an acquisition code configured to cause the at least one processor to acquire video data including at least one encoded picture; a reconstruction code configured to cause the at least one processor to reconstruct the encoded picture, the encoded picture including a Film Grain Characteristics Auxiliary Enhancement Information SEI message, an Alpha Channel Information SEI message, and at least a first sample and a second sample and at least a third alpha map sample and a fourth alpha map sample, wherein after reconstruction, the first sample is spatially co-located with the third alpha map sample, and the second sample is spatially co-located with the first sample; an application code configured to cause the at least one processor to apply Film Grain Synthesis to the first sample based on a first value of the third sample; and a determination code configured to cause the at least one processor to determine not to apply Film Grain Synthesis to the second sample based on a second value of the fourth sample, wherein the first value and the second value are different from each other.

[0009] The Film Grain Characteristics SEI message may specify the coordinates of the upper left corner of the bounding box of the i-th region of the encoded picture, as well as the width and height.

[0010] The Film Grain Characteristics SEI message may specify the coordinates, the width, and the height through any one of the fgr_region_top[i] syntax, the fgr_region_left[i] syntax, the fgr_region_width[i] syntax, and the fgr_region_height[i] syntax.

[0011] Any of the following: the value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicWidthInLumaSamples - 1, inclusive of 0 and PicWidthInLumaSamples - 1, the value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicHeightInLumaSamples - 1, inclusive of 0 and PicHeightInLumaSamples - 1, the value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive of 0 and PicWidthInLumaSamples - fgr_region_left[i], and the value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive of 0 and PicHeightInLumaSamples - fgr_region_top[i].

[0012] Each of the following: the value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicWidthInLumaSamples-1, inclusive of 0 and PicWidthInLumaSamples-1, the value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicHeightInLumaSamples-1, inclusive of 0 and PicHeightInLumaSamples-1, the value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive of 0 and PicWidthInLumaSamples - fgr_region_left[i], and the value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive of 0 and PicHeightInLumaSamples - fgr_region_top[i].

[0013] Both the first sample and the second sample are rectangular.

[0014] The first sample and the second sample at least partially overlap each other. Description of the Drawings

[0015] Other features, properties, and various advantages of the disclosed subject matter will become further apparent from the following detailed description and the accompanying drawings, in which:

[0016] Figure 1 is a schematic diagram of a computer environment according to an embodiment;

[0017] Figure 2 is a simplified block diagram of media processing according to an embodiment;

[0018] Figure 3 is a simplified schematic diagram of decoding according to an embodiment;

[0019] Figure 4 is a simplified schematic diagram of encoding according to an embodiment;

[0020] Figure 5 is a simplified schematic diagram of a NAL unit and an SEI header according to an embodiment;

[0021] Figure 6Is a schematic diagram of a video scene according to an embodiment, the video scene including movie content that requires film grain synthesis and computer-generated content that does not require film grain synthesis;

[0022] Figure 7 Is a schematic diagram of a modified film grain characteristic SEI message according to an embodiment;

[0023] Figure 8 Is a schematic diagram of a modified alpha channel information SEI message according to an embodiment;

[0024] Figure 9 Is a schematic diagram of a video sequence including an annotation area SEI message according to an embodiment;

[0025] Figure 10 Is a schematic diagram of a modified film grain characteristic SEI message according to an embodiment; and

[0026] Figure 11 Is a schematic diagram of a computer system according to an embodiment. Detailed Description

[0027] Techniques for per-sample applications of film grain synthesis using SEI messages are disclosed.

[0028] The proposed features discussed below can be used alone or in any combination. Additionally, embodiments can be implemented by a processing circuit (e.g., at least one processor or at least one integrated circuit). In one example, at least one processor executes a program stored in a non-volatile computer-readable medium.

[0029] Today's video signals typically consist of at least two sources. For example, a movie with commentary can consist of the movie content itself and the surrounding frames with commentary. The movie may have been shot using chemical film technology and may thus include film grain after digitization. To achieve reasonable compression, the film grain can be removed by pre-filtering. On the other hand, the surrounding frames are created digitally and do not include film grain. After encoding, transmission, and reconstruction, the resulting reconstructed picture does not include the noise associated with film grain. This is preferable for the surrounding frames but not for the movie, as the film grain here has been removed by pre-filtering. Therefore, a technique is needed in which film grain can be selectively inserted into parts of the reconstructed picture while leaving the unselected parts of the picture without film grain. More complex scenarios may involve including multiple movies with different film grain characteristics in a synthetic picture. In such a scenario, different film grain re-insertion parameters can be advantageously applied to different parts of the reconstructed picture.

[0030] Figure 1is a simplified block diagram of a communication system 100 according to an embodiment disclosed in the present application. The communication system 100 may include at least two terminal devices 102 and 103 interconnected by a network 105. For one-way data transmission, the first terminal device 103 may encode video data at a local location for transmission over the network 105 to another terminal device 102. The second terminal device 102 may receive the encoded video data of another terminal device from the network 105, decode the encoded video data, and display the recovered video data. One-way data transmission is more common in applications such as media services.

[0031] Figure 1 A second pair of terminal devices 101 and 104 is shown, which is provided to support two-way transmission of encoded video that may occur, for example, during a video conference. For two-way data transmission, each terminal device 101 and 104 may encode video data collected at a local location for transmission over the network 105 to another terminal device. Each terminal device 101 and 104 may also receive the encoded video data transmitted by another terminal device, may decode the encoded video data, and may display the recovered video data on a local display device.

[0032] In Figure 1 the terminal devices 101, 102, 103, and 104 may be servers, personal computers, and smart phones, but the principles disclosed in the present application are not limited thereto. The embodiments disclosed in the present application are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network 105 represents any number of networks that convey encoded video data between the terminal devices 101, 102, 103, and 104, including, for example, wired (wired) and / or wireless communication networks. The communication network 105 may exchange data in circuit-switched and / or packet-switched channels. The network may include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of the present application, unless otherwise explained hereinafter, the architecture and topology of the network 105 may be irrelevant to the operation disclosed in the present application. The network 105 may include a media-aware network element (MANE), which may be included in the transmission path, for example, between the terminal devices 101 and 104. The purpose of the MANE may be to selectively forward portions of the media data in response to network congestion, media switching, media mixing, archiving, and similar tasks typically performed by service providers rather than end users. Such a MANE is capable of parsing and reacting to a limited portion of the media transmitted over the network, such as syntax elements related to the network abstraction layer of a video coding technology or standard.

[0033] As an example of an application of the disclosed subject matter, Figure 2Shows the placement of a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0034] A streaming system may include an acquisition subsystem 203, which may include a video source 201 such as a digital camera that creates an uncompressed video sample stream 213, for example. The sample stream 213 may be emphasized as having a high data volume compared to an encoded video bitstream and may be processed by an encoder 202 coupled to the video source 201. The encoder 202 may include hardware, software, or a combination of both to implement or enforce aspects of the disclosed subject matter described in more detail below. The encoded video bitstream 204 may be emphasized as having a lower data volume compared to the sample stream and may be stored on a streaming server 205 for future use. At least one streaming client 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. The client 212 includes a video decoder 211 that decodes an incoming copy 208 of the encoded video bitstream and produces an output video sample stream 210 that may be presented on a display 209 or other display device (not depicted). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to certain video coding / compression standards. Implementations of these standards are as described above and are further described herein. Examples of implementations of these standards include ITU-T Recommendations H.265 and H.266. The disclosed subject matter may be used in the context of the VVC standard.

[0035] Figure 4 Is a functional block diagram of a video decoder 300 according to an embodiment of the present application.

[0036] The receiver 302 may receive at least one encoded video sequence to be decoded by the video decoder 300; in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from the channel 301, which may be a hardware / software link to a storage device storing the encoded video data. The receiver 302 may receive the encoded video data and other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not shown). The receiver 302 may separate the encoded video sequence from the other data. To prevent network jitter, the buffer memory 303 may be coupled between the receiver 302 and the entropy decoder / parser 304 (hereinafter referred to as "parser"). When the receiver 302 receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory 303, or the buffer memory may be made smaller. For use on a best-effort packet network such as the Internet, the buffer memory 303 may be required, which may be relatively large and may have an adaptive size.

[0037] The video decoder 300 may include a parser 304 to reconstruct symbols 313 from the entropy-coded video sequence. The categories of these symbols include information for managing the operation of the video decoder 300 and potential information for controlling a display device such as the display 312, which is not part of the decoder but may be coupled to the decoder. The control information for the display device may be in the form of Supplemental Enhancement Information (SEI messages) or Video Usability Information (VUI). The parser 304 may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technology or standard and may follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser 304 may extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and so on. The entropy decoder / parser may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0038] The parser 304 can perform entropy decoding / parsing operations on the video sequence received from the buffer memory 303 to create symbols 313. The parser 304 can receive the encoded data and selectively decode specific symbols 313. In addition, the parser 304 can determine whether to provide a specific symbol 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.

[0039] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter picture and intra picture, inter block and intra block) and other factors, the reconstruction of the symbol 313 may involve at least two different units. Which units are involved and the way of involvement can be controlled by subgroup control information parsed by the parser 304 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser 304 and at least two units below are not described.

[0040] In addition to the functional blocks already mentioned, the video decoder 300 can be conceptually divided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide into the functional units below.

[0041] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives the quantized transform coefficients as symbols 313 and control information from the parser 304, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit 305 can output a block including sample values, and the sample values can be input into the aggregator 310.

[0042] In some cases, the output samples of the scaler / inverse transform unit 305 may belong to an intra-coded block; that is, a block that does not use the predictive information from the previously reconstructed picture but may use the predictive information from the previously reconstructed part of the current picture. Such predictive information can be provided by the intra picture prediction unit 307. In some cases, the intra picture prediction unit 307 uses the reconstructed information extracted from the current (partially reconstructed) picture buffer 309 to generate surrounding blocks having the same size and shape as the block being reconstructed. In some cases, the aggregator 310 adds the prediction information generated by the intra prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305 based on each sample.

[0043] In other cases, the output samples of the scaler / inverse transform unit 305 may belong to an inter-coded and potentially motion-compensated block. In such cases, the motion compensation prediction unit 306 may access the reference picture memory 308 to extract samples for prediction. After motion-compensating the extracted samples according to the symbol 313, these samples may be added by the aggregator 310 to the output of the scaler / inverse transform unit (referred to as residual samples or a residual signal in this case), thereby generating output sample information. The motion compensation prediction unit obtaining prediction samples from an address within the reference picture memory may be controlled by a motion vector, and the motion vector is in the form of the symbol 313 for use by the motion compensation prediction unit, the symbol 313 including, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.

[0044] The output samples of the aggregator 310 may be employed by various loop filtering techniques in the loop filter unit 311. Video compression techniques may include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream, and the parameters may be available to the loop filter unit 311 as the symbol 313 from the parser 304, but may also respond to meta-information obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and in response to previously reconstructed and loop-filtered sample values.

[0045] The output of the loop filter unit 311 may be a sample stream that may be output to the display device 312 and stored in the reference picture memory 557 for subsequent inter-picture prediction.

[0046] Once fully reconstructed, some encoded pictures may be used as reference pictures for future prediction. Once an encoded picture has been fully reconstructed and the encoded picture is identified (by, for example, the parser 304) as a reference picture, the current picture buffer 309 may become part of the reference picture memory 308, and a new current picture buffer may be reallocated before starting to reconstruct subsequent encoded pictures.

[0047] The video decoder 300 may perform decoding operations according to a predetermined video compression technique, such as that documented in the ITU-T Recommendation H.266 standard. An encoded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the encoded video sequence follows the video compression technique or standard specified in a video compression technique document or standard, particularly the syntax specified by a profile therein. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further restricted by the Hypothetical Reference Decoder (HRD) specification and metadata signaled in the encoded video sequence for HRD buffer management.

[0048] In an embodiment, the receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder 300 to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0049] Figure 4 is a functional block diagram of a video encoder 400 according to an embodiment disclosed in the present application.

[0050] The video encoder 400 may receive video samples from a video source 401 (not part of the encoder), which may capture video images to be encoded by the video encoder 400.

[0051] The video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by a video encoder (303). The digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 Y CrCb, RGB, etc.) and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 401 may be a camera that captures local image information as a video sequence. The video data may be provided as at least two separate pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include at least one sample depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0052] According to an embodiment, the video encoder 400 may encode and compress pictures of the source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 402. The controller 402 controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, the couplings are not labeled in the figure. The parameters set by the controller 402 may include rate control related parameters (picture skipping, quantizer, λ value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller 402 as they may be related to the video encoder 400 optimized for a certain system design.

[0053] Some video encoders operate in a manner that is readily recognizable to those skilled in the art as an "encoding loop". As a simple description, the encoding loop can consist of an encoding portion of encoder 400 (hereinafter referred to as the "source encoder") that is responsible for creating symbols based on the input picture to be encoded and reference pictures, and an (local) decoder 406 embedded within video encoder 400. The decoder 400 reconstructs the symbols to create sample data that the (remote) decoder would also create (since in the video compression techniques contemplated in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input into reference picture memory 405. Since the decoding of the symbol stream produces a bit-exact result that is independent of the decoder location (local or remote), the content of the reference picture memory is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. The basic principles of this reference picture synchronization (and the drift that occurs in cases where synchronization cannot be maintained, for example, due to channel errors) are well-known to those skilled in the art.

[0054] The operation of the "local" decoder 406 can be the same as the operation of the "remote" decoder 300 described in detail above in conjunction with Figure 3 However, briefly referring additionally to Figure 4 , when the symbols are available and the entropy encoder 408 and parser 304 can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding portion of video decoder 300, including channel 301, receiver 302, buffer memory 303, and parser 304, may not be fully implementable in local decoder 406.

[0055] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. The description of the encoder technology can be simplified because the encoder technology is inverse to the decoder technology described comprehensively. More detailed descriptions are only needed in certain areas and are provided below.

[0056] As part of its operation, source encoder 403 can perform motion compensation predictive coding. Referencing at least one previously encoded frame in the video sequence designated as a "reference frame", the motion compensation predictive coding performs predictive coding on the input frame. In this way, encoding engine 407 encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, and the reference frame can be selected as the prediction reference for the input frame.

[0057] The local video decoder 406 can decode the encoded video data of the frame that can be specified as a reference frame based on the symbols created by the source encoder 403. The operation of the encoding engine 407 can be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 4 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that can be performed by the video decoder on the reference frame and can store the reconstructed reference frame in the reference picture memory 405, which can be a cache, for example. In this way, the video encoder 400 can locally store a copy of the reconstructed reference frame that has the same content as the reconstructed reference frame that will be obtained by the remote video decoder (without transmission errors).

[0058] The predictor 404 can perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 can search the reference picture memory 405 for sample data (as a candidate reference pixel block) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new picture. The predictor 404 can operate on a per-pixel block basis of the sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 404, it can be determined that the input picture can have prediction references taken from at least two reference pictures stored in the reference picture memory 405.

[0059] The controller 402 can manage the encoding operations of the source encoder 403 (which can be a video encoder), including, for example, setting parameters and subgroup parameters for encoding the video data.

[0060] The outputs of all the above functional units can be entropy encoded in the entropy encoder 408. The entropy encoder performs lossless compression on the symbols generated by various functional units according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.

[0061] The transmitter 409 can buffer the encoded video sequence created by the entropy encoder 408 to prepare for transmission through the communication channel 411, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter 409 can merge the encoded video data from the source encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0062] The controller 402 can manage the operation of the video encoder 400. During encoding, the controller 402 can assign a certain type of encoded picture to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures can generally be assigned to any of the following frame types:

[0063] An intracoded picture (I picture), which can be a picture that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intracoded pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and characteristics.

[0064] A predictive picture (P picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.

[0065] A bi-predictive picture (B picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, at least two predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0066] Source pictures can generally be spatially subdivided into at least two sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (encoded) blocks, which are determined according to the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-predictively encoded, or the blocks can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be predictively encoded with reference to a previously encoded reference picture either through spatial prediction or through temporal prediction. Blocks of a B picture can be predictively encoded with reference to one or two previously encoded reference pictures either through spatial prediction or through temporal prediction.

[0067] The encoder 400, which can be a video encoder for example, can perform encoding operations according to a predetermined video encoding technique or standard such as the ITU-T Recommendation H.266. In operation, the video encoder 400 can perform various compression operations, including predictive encoding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard used.

[0068] In an embodiment, the transmitter 409 may transmit additional data when transmitting the encoded video. The source encoder 403 may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.

[0069] The compressed video may be enhanced by supplementary enhancement information in the video bitstream, for example, in the form of supplementary enhancement information (SEI) messages or video usability information (VUI). The video coding standard may include a normative part for SEI and VUI. The SEI and VUI information may also be specified in a separate specification referenced by the video coding specification.

[0070] Reference Figure 5 Example 500 of shows an exemplary layout of an encoded video sequence (CVS) according to H.266. The encoded video sequence is subdivided into network abstraction layer units (NAL units). The exemplary NAL unit 501 may include a NAL unit header 502, which in turn includes the following 16 bits: The forbidden_zero_bit 503 and nuh_reserved_zero_bit 504 may not be used by H.266 and may be zero in a NAL unit compliant with H.266. The three-bit nuh_layer_id 505 may indicate the layer (spatial, SNR, or multi-view enhancement) to which the NAL unit belongs. The five-bit nuh_nal_unit_type defines the type of the NAL unit. In H.266, 22 NAL unit type values are defined for the NAL unit types defined in H.266, six NAL unit types are reserved, four NAL unit type values are unspecified and may be used by specifications other than H.266. Finally, three bits of the NAL unit header indicate the temporal layer nuh_temporal_id_plus1 506 to which the NAL unit belongs.

[0071] The encoded picture may include at least one video coding layer (VCL) NAL unit and zero or at least two non-VCL NAL units. The VCL NAL unit may contain encoded data that conceptually belongs to the video coding layer as described above. The non-VCL NAL units may contain data that conceptually does not belong to the video coding layer. Taking H.266 as an example, they may be classified as (1) parameter sets, (2) picture headers (PH_NUT), (3) NAL units, (4) prefix and suffix SEI Nal unit types (PREFIX_SEI_NUT and SUFFIX_SEI_NUT), (5) filler data NAL unit type FD_NUT, and (6) reserved and unspecified NAL unit types, as follows.

[0072] (1) Parameter sets, which include the information required for the decoding process and can be applied to more than one coded picture. Parameter sets and conceptually similar NAL units can be NAL unit types such as DCI_NUT (Decoding Capability Information (DCI)), VPS_NUT (Video Parameter Set (VPS), establishing layer relationships, etc.), SPS_NUT (Sequence Parameter Set (SPS), establishing parameters used and kept constant in the entire coded video sequence (CVS)), PPS_NUT (Picture Parameter Set (PPS), establishing parameters used and kept constant within the coded picture), and PREFIX_APS_NUT and SUFFIX_APS_NUT (prefix and suffix adaptation parameter sets). Parameter sets can include the information required by the decoder to decode VCL NAL units and are thus referred to as "standard" NAL units here.

[0073] (2) Picture Header (PH_NUT), which is also a "standard" NAL unit.

[0074] (3) NAL units that mark certain positions in the NAL unit stream. These include NAL units with NAL unit types AUD_NUT (Access Unit Delimiter), EOS_NUT (End of Sequence), and EOB_NUT (End of Bitstream). These are non-standard and also called informative in the sense that although compatible decoders need to be able to receive them in the NAL unit stream, they are not required during the decoding process.

[0075] (4) Prefix and suffix SEI Nal unit types (PREFIX_SEI_NUT and SUFFIX_SEI_NUT), which represent NAL units containing prefix and suffix supplementary enhancement information. In H.266, these NAL units are informative as they are not required for the decoding process.

[0076] (5) The filler data NAL unit type FD_NUT indicates filler data; it can be random and data that can be used to "waste" bits in the NAL unit stream or bitstream, which may be necessary for transmission in certain synchronous transmission environments.

[0077] (6) Reserved and undefined NAL unit types.

[0078] Still refer to Figure 5, shows the layout of a NAL unit stream in decoding order 510, which contains encoded pictures 511. The encoded pictures 511 contain some of the types of NAL units introduced earlier. Somewhere early in the NAL unit stream, the DCI 512, VPS 513, and SPS 514 can be combined to establish the parameters that the decoder can use to decode the encoded pictures (including the encoded pictures 511 of the NAL unit stream) of an encoded video sequence (CVS).

[0079] The encoded pictures 511 can contain, in the order described or any other order compliant with the video coding technology or standard in use (here H.266): a prefix APS 516, a picture header (PH) 517, a prefix SEI 518, at least one VCL NAL unit 519, and a suffix SEI 520.

[0080] The prefix and suffix SEI NAL units 518 and 520 were motivated during the standard development process because for some SEI messages, the content of the message is known before the encoding of a given picture begins, while other content is only known after the picture is encoded. By allowing certain SEI messages to appear earlier or later in the NAL unit stream of an encoded picture through the prefix and suffix SEIs, buffering can be avoided. As an example, in the encoder, the sampling time of the picture to be encoded is known before the picture is encoded, so the picture timing SEI message can be a prefix SEI message 516. On the other hand, the decoded picture hash SEI message (which contains the hash of the sample values of the decoded picture and can be used, for example, to debug the encoder implementation) is a suffix SEI message 518 because the encoder cannot calculate the hash of the reconstructed samples before the picture is encoded. The positions of the prefix and suffix SEI NAL units are not limited to their positions in the NAL unit stream. The phrases "prefix" and "suffix" can imply which encoded pictures or NAL units the prefix / suffix SEI messages can belong to, and details of such applicability can be specified, for example, in the semantic description of a given SEI message.

[0081] Still referring to Figure 5, shows a simplified syntax diagram of a NAL unit containing a prefix or suffix SEI message 520. This syntax is a container format for at least two SEI messages that can be carried in one NAL unit. For clarity, the details of the emulation prevention syntax specified in H.266 are omitted here. Like other NAL units, the SEI NAL unit starts with a NAL unit header 521. Following the header is at least one SEI message; two are shown as 530, 531 and are described below. Each SEI message within the SEI NAL unit includes: an 8-bit payload_type_byte 522 that specifies one of 256 different SEI types; an 8-bit payload_size_byte 523 that specifies the number of bytes of the SEI payload; and a payload of payload_size - byte number of bytes 524. This structure can be repeated until the payload_type_byte is observed to be equal to 0xff, which indicates the end of the NAL unit. The syntax of the payload 524 depends on the SEI message and can be of any length between 0 and 255 bytes.

[0082] Reference Figure 6 Example 600 of, shows a still image that is obtained from a black-and-white movie taken on a chemical film 601. The lighting conditions in this particular shot are adverse (nighttime shooting), so many film grains can be observed when viewing the original movie content. Nowadays, it is desirable to reproduce these film grains as it is directly related to the original content and may even be part of the original artistic expression. In contrast, a computer screenshot 602 relies on clear, noise-free information to represent, for example, the lowercase letters on the right sidebar. In such a usage scenario, it may be acceptable that the potentially noisy mountain view in the background is not faithfully presented.

[0083] Now consider mixing these two contents in the same screen layout 603 and possibly in the same video stream. Some parts of the content benefit from film grains, while other parts (such as the letters on the right sidebar) do not.

[0084] This does not show distracting details in screenshot 604. The movie image 605 benefits from the reapplication of film grain, while the background 606 and especially areas of high-spatial computer-generated details (such as the string 607) do not benefit from film grain. Note that the screen layout may also include a window labeled "Alarm" 608 that partially obscures the displayed movie 605. The window is depicted as rectangular, but can be of any shape in modern systems. This is shown here to emphasize that indicating the movie area 605 by a rectangle encoded as metadata (such as by its corner coordinates) may not be sufficient to identify the screen areas that need film grain processing. Those skilled in the art can easily observe other scenarios where non-rectangular (or non-uniform) areas that may need film grain processing need to be identified. For example, consider a scenario that shows a room with a TV playing a movie, where the TV is at an angle with respect to the viewer's parallax. A sample-by-sample based solution is needed to indicate film grain.

[0085] Using a mask for each sample to identify film grain may also be useful in cases such as: · For artistic intent (focusing on a face), to draw more attention of the viewer to a specific area in the picture (producers used to use shallow depth of field to blur the background) · To restore texture, such as the grass on a football field, which is particularly useful when using frame interpolation in a TV (film grain will maintain the perception of details) · To provide realism only on the face / skin in a video conferencing scenario (if grain is applied over the whole image, it will look like old-fashioned video: this may also be a problem if there is screen content in it) · Synthetic objects in special effects in movie production may have too soft a texture, in which case such a film grain mask will be used as a post-production tool, such as a plugin for Apple Final Cut Pro or DaVinci Resolve. · Video games have cinematic animations that would benefit from the local application of film grain. · To control exposure and contrast by applying higher film grain to shadows and highlights. · To limit the perception of visual artifacts in the picture.

[0086] Brief reference Figure 2, consider an application scenario where, in the transmitting system 203, the content source 201 is not an actual camera but an image synthesizer that puts together a video sequence 202 covering a scene comparable to the synthetic scene 603 shown. In other words, the signal presented to the encoder 202 is a synthetic scene, some parts of which benefit from denoising and film grain recognition and encoding in the form of metadata such as SEI messages, while other parts do not. Such an encoded stream can potentially be transmitted 204, 208 via a server 205 to the receiving system 212 and the decoder 211 located therein. The receiving system can decode the metadata identifying samples that require film grain processing and apply it to the reconstructed bitstream 210 via a post-processing step.

[0087] Summarizing the above use case discussion, what is needed is: 1. A mechanism for identifying at least one sample of a reconstructed encoded picture to which film grain processing should be applied for an optimal user experience. Such samples can be represented by rectangular regions, but can also have more complex shapes. Additionally, it may make sense to a) apply different film grain characteristics to different samples, and / or b) apply film grain at different intensities for different samples. 2. A mechanism for describing film grain characteristics. 3. A mechanism for combining the above two mechanisms, where the sample selection / intensity mechanism is applicable to film grain.

[0088] Response 1: The sample selection mechanism is readily available in video coding standards and related metadata standards. As an example, the "annotated region" SEI is available, which allows the annotation of regions of a picture identified by four corners using metadata that can be processed by a renderer. Another option could be to use an alpha map. An alpha map can be a reconstructed picture with a single plane, which in the current case can be used as a boolean or integer value for each corresponding sample of the reconstructed picture. The alpha planes can have the same spatial dimensions as the reconstructed picture, or they can have different sizes, in which case they can be enlarged / shrunk to map each sample of the reconstructed picture to a sample (or a filtered synthesis of at least two samples) of the alpha plane.

[0089] Response 2: The VSEI specifies film grain SEI messages, and similar information can be obtained in the VUI of the video coding specification.

[0090] Response 3: Currently, no such combination is available in video coding or metadata specifications.

[0091] In an embodiment, an alpha map is used to indicate whether a film grain process should be applied to spatially corresponding samples of a reconstructed picture. Assuming that the sample values of the alpha plane are 10-bit integers in the value range of 0 to 1023 (as is the case when using H.266), alpha plane values below 512 may indicate that film grain synthesis should not be applied, while values above 511 may indicate that film grain synthesis should be applied. For a scene associated with picture layout 608, such an alpha map can be well compressed - perhaps compressed to a few bytes.

[0092] In an embodiment, the alpha map may follow a similar design principle as described above, but may include integer values instead of integer-encoded boolean values. For each sample, the integer value may represent the application strength of the film grain.

[0093] Referring Figure 7 to example 700 of [], in order to recognize the alpha map as a mechanism for indicating the use or strength of film grain on a per-sample basis, the existing alpha channel information (ACI) SEI message of H.274 can be modified as follows: 1. The syntax remains unchanged 2. In the semantics of the film grain synthesis information SEI, insert a paragraph as Figure 7 shown, which modifies the FGS blending based on the alpha channel information. 3. In the semantics of the alpha channel information SEI, insert a paragraph as Figure 8 shown to introduce the previously reserved value "3" and associate it with the use of the alpha map for film grain synthesis.

[0094] Referring Figure 7 , an excerpt of the semantic definition of the modified film grain characteristic SEI message according to an embodiment is shown. The modifications are shown by the underlined text. For clarity, some parts of the semantics are removed and represented by "[...]" 702. For most parts, the semantics remain unchanged 703. However, if the current picture unit (PU) contains an alpha channel control indication 704 with the (newly defined) value "3", the blending equation changes such that the alpha channel sample value Iaux[][][] 705 is included in the addition or multiplication 706, 707.

[0095] In the same or another embodiment, referring Figure 8 to example 800 of [], an alpha channel control SEI 801 is shown, and some parts 802 are omitted for clarity. The underlined text represents the content added / modified relative to the published H.274 specification.

[0096] A new value “3” of alpha_channel_use_idc is introduced. Its semantics is that when found, the sample values in the reconstructed auxiliary picture (the standard term for alpha map in this context) are used for film grain synthesis. Specifically, when a value equal to 3 is found, this may indicate that before calculating the film grain value, the analog film grain value at the same position and color component sample should be multiplied by the interpreted sample value of the decoded auxiliary picture.

[0097] In an embodiment, instead of extending the previously existing alpha channel information to add a new mode, a new type of SEI message can be specifically generated to mask the regions in the picture where film grain is applicable. The new type of SEI message can contain a syntax element (e.g., fgs_aux_weighting_idc) that indicates whether the analog film grain value at the same position and color component sample should be weighted using a weighting factor determined based on the associated auxiliary picture. For example, when the value of fgs_aux_weighting_idc is equal to 0, the analog film grain value at the same position and color component sample is not weighted using the weighting factor determined based on the associated auxiliary picture. On the other hand, if the value of fgs_aux_weighting_idc is equal to 1, the analog film grain value at the same position and color component sample is not weighted using the weighting factor determined based on the associated auxiliary picture. Specifically, before calculating the film grain value, the analog film grain value at the same position and color component sample in the picture can be multiplied by the interpreted sample value of the decoded auxiliary picture.

[0098] In one embodiment, if the current PU contains the new type of SEI message, the analog film grain value will be weighted by the auxiliary data value as follows: - If fg_blending_mode_id is equal to 0, the blending mode is the addition mode specified as follows: - Otherwise (fg_blending_mode_id is equal to 1), the blending mode is the multiplication mode specified as follows: where Iaux[c][x][y] represents the sample value at coordinates x, y of color component c of the decoded auxiliary picture.

[0099] Film Grain Synthesis (also known as FGS) is a tool designed to preserve the original film grain while maintaining encoding efficiency due to its randomness and complexity. Although used to preserve artistic intent or mitigate visual artifacts, the film grain does not appear relevant across an entire sequence or even within an image. Various aspects of this contribution address the problem of locally applying film grain synthesis on a per-image basis by using auxiliary data as a mask.

[0100] Film grain is a well-known feature that originated from small grains of silver halide crystals on film and is easily recognizable as a cinematic art feature. The film grain effect also gives games a movie-like feel by simulating the grainy visual effects presented in some movies. If used locally and softly in places where it is valuable to the viewer, it creates a sense of realism. In this case, those artificial defects make the synthetic content more authentic. Finally, film grain can also be used to soften visual artifacts generated during the capture process or during the compression of the video signal.

[0101] In many cases, film grain provides a certain degree of realism only in certain parts of the picture (e.g., on skin textures), while keeping the background soft.

[0102] In its simplest approach, the spatial adaptation of film grain synthesis combines the calculated film grain values with a binary mask conveyed in the auxiliary data. In this case, each film grain value is applied as is in the image or not applied at all.

[0103] Another approach is to use a mask to define a weighting factor that follows the CTU structure of the main picture and apply more film grain in the lower frequency regions.

[0104] Another approach still relies on the same process and uses a mask that has a weighting factor for the film grain values. Such a mask can be a depth map transmitted as auxiliary data for 3D rendering purposes. In this scenario, film grain is applied to the foreground objects while the background remains unchanged. This configuration is commonly used to increase depth perception.

[0105] The film grain characteristic SEI message is defined as the grain of the synthetic content without encoding it. Currently, it lacks spatial adaptation to the image.

[0106] The VSEI specification also defines the alpha channel information (ACI) SEI message, which defines how to combine the auxiliary data with the output image for alpha blending.

[0107] The ACI SEI message has two defined modes in alpha_channel_use_idc: Mode 0: Multiplication mode, and Mode 1: Non-multiplication mode. The third mode (Mode 3, since Mode 2 is set to "unspecified") can be defined as the multiplicative use of the auxiliary data with the corresponding analog film grain defined by the film grain characteristic SEI message.

[0108] When interpreting the film grain characteristic SEI message, if there is also an ACI SEI message for the same picture where alpha_channel_use_idc is equal to 3, the associated auxiliary picture is used as a weighting factor for the film grain values.

[0109] The following clauses depict the affected changes in the VSEI specification.

[0110] fg_blending_mode_id identifies the blending mode used to blend the analog film grain with the input image specified in Table 6. fg_blending_mode_id shall be in the range from 0 to 1 (including the end values). The values 2 and 3 of fg_blending_mode_id are reserved for future use by ITU T ISO / IEC and shall not be present in the bitstream of this version that complies with this specification. The decoder shall ignore the film grain characteristic SEI message where fg_blending_mode_id is equal to 2 or 3.

[0111] According to the value of fg_blending_mode_id, the blending mode is specified as follows: - If fg_blending_mode_id is equal to 0, the blending mode is the addition mode specified as follows: - Otherwise (fg_blending_mode_id is equal to 1), the blending mode is the multiplication mode specified as follows: where represents the sample value at coordinates x, y of color component c of the input image G[c][x][y] is the analog film grain value at the same location and color component, and fgBitDepth[c] is the number of bits used for each sample in the fixed-length unsigned binary representation of the arrays Igrain[c][x][y], and G[c][x][y], where c = 0..2, x = 0..PicWidthInLumaSamples - 1, y = 0..PicHeightInLumaSamples - 1.

[0112] If the current PU contains an ACI SEI message with alpha_channel_use_idc equal to 3 (as defined in Clause 8.23.2), the simulated film grain value shall be weighted by the auxiliary data value as follows: - If fg_blending_mode_id is equal to 0, the blending mode is the additive mode specified as follows: - Otherwise (fg_blending_mode_id is equal to 1), the blending mode is the multiplicative mode specified as follows: where Iaux[c][x][y] represents the sample value at coordinates x, y of color component c of the decoded auxiliary picture.

[0113] The Alpha Channel Information (ACI) SEI message provides information about alpha channel sample values and information about post-processing applied to the decoded alpha plane that is encoded in an auxiliary picture of type AUX_ALPHA and at least one associated primary picture.

[0114] When the CVS does not contain an SDI SEI message (for at least one i value, sdi_aux_id[i] is equal to 1), no picture in the CVS shall be associated with an ACI SEI message.

[0115] When an AU contains an SDI SEI message (for at least one i value, sdi_aux_id[i] is equal to 1) and an ACI SEI message, the SDI SEI message shall precede the ACI SEI message in decoding order.

[0116] When an access unit contains an auxiliary picture picA in a layer with nuh_layer_id equal to nuhLayerIdA, where the layer is indicated by an SDI SEI message as an alpha auxiliary layer, the alpha channel sample values of picA persist in output order until at least one of the following conditions is true: - The next (in output order) picture with nuh_layer_id equal to nuhLayerIdA is output. - The CLVS containing the auxiliary picture picA ends. - The bitstream ends. - The CLVS of any associated primary layer of the auxiliary picture layer with nuh_layer_id equal to nuhLayerIdA ends.

[0117] The following semantics apply separately to each nuh_layer_id targetLayerId in the nuh_layer_id values to which the ACI SEI message applies.

[0118] An alpha_channel_cancel_flag equal to 1 indicates that the SEI message cancels the persistence of any previous ACI SEI message in output order applicable to the current layer. An alpha_channel_cancel_flag equal to 0 indicates that the ACI follows.

[0119] Let currPic be the picture associated with the ACI SEI message. The semantics of the ACI SEI message remain unchanged for the current layer in output order until at least one of the following conditions is true: - A new CLVS for the current layer begins. - The bitstream ends. - Output a picture in the current layer in the AU associated with the ACI SEI message that is in output order after the current picture.

[0120] An alpha_channel_use_idc equal to 0 indicates that, for the purpose of alpha blending, after output from the decoding process, during the display process, the decoded samples of the associated primary picture shall be multiplied by the interpreted sample values of the decoded auxiliary picture. An alpha_channel_use_idc equal to 1 indicates that, for the purpose of alpha blending, after output from the decoding process, during the display process, the decoded samples of the associated primary picture shall not be multiplied by the interpreted sample values of the decoded auxiliary picture. An alpha_channel_use_idc equal to 2 indicates that the use of the auxiliary picture is not specified. An alpha_channel_use_idc equal to 3 indicates that, before calculating the film grain value, the simulated film grain value at the same location and color component sample shall be multiplied by the interpreted sample values of the decoded auxiliary picture. Values of alpha_channel_use_idc greater than 3 are reserved for future use by ITU-T|ISO / IEC. When absent, the value of alpha_channel_use_idc is inferred to be equal to 2. The decoder shall ignore alpha channel information SEI messages with alpha_channel_use_idc greater than 3.

[0121] Refer briefly again to Figure 2, consider an application scenario where, in the sending system 203, the content source 201 is not an actual camera but an image synthesizer that puts together the video sequence 202, and the scene covered by the video sequence 202 is comparable to the synthetic scene 603 shown. In other words, the signal presented to the encoder 202 is a synthetic scene, some parts of which benefit from denoising and film grain recognition as well as encoding in the form of metadata such as SEI messages, while other parts do not. This encoded stream can potentially be transmitted 204, 208 via the server 205 to the receiving system 212 and the decoder 211 located therein. The receiving system can decode the metadata identifying the samples that require film grain processing and apply it to the reconstructed bitstream 210 through a post-processing step.

[0122] Summarizing the above use case discussion, what is needed and provided by the embodiments is: 1. A mechanism for identifying at least one region of a reconstructed encoded picture to which film grain processing should be applied for obtaining an optimal user experience. Such regions can be rectangular regions but can overlap. Additionally, it may make sense to apply different film grain characteristics to different regions. 2. A mechanism for describing film grain characteristics. 3. A mechanism for combining the above two mechanisms, where for regions that require film grain, appropriate film grain characteristics (including possibly no film grain characteristics) can be applied; that is, film grain can be synthesized in a post-processing step.

[0123] Response 1: The "Annotated Region" SEI (AR) is available in the VSEI specification, and for video coding techniques that do not use VSEI - or in cases where VSEI is available but insufficient for the application - a person skilled in the art can design a mechanism equivalent to AR. Note that, among other things, AR can specify at least two regions in an ordered list and associate each region with a string selected by the encoder.

[0124] Response 2: The VSEI specifies the Film Grain SEI Message (FGS), and similar information can be obtained in the VUI of the video coding specification. At least two FGS messages can be in a Picture Unit (PU), and the decoder / receiver (211) can use the order in which they appear in the bitstream to associate a given one of the at least two FGSs with one of the regions defined in the AR.

[0125] Response 3: A new mechanism is needed to establish a combination between at least one FGS and at least one AR to associate the film grain characteristics conveyed by a given FGS with the regions defined in the AR.

[0126] In an embodiment, referring to Figure 9 Example 900 of, the bitstream may include, for example, at least one FGS (one FGS 902 is shown here) and an AR 903 in decoding order 906. The term "bitstream" here may refer to at least one encoded picture or picture unit (PU); a sensible design choice may be a CVS, but even the scope of just a single picture is feasible, even if this means potentially resending the same SEI on each picture in the sequence. This can be an option because both the FGS and the AR are small compared to the encoded pictures (whose quality makes the use of FGS desirable).

[0127] Following two SEI messages are NAL units that make up the content of the encoded picture. Two pictures 904, 905 are shown here. Both of these pictures can be within the scope of the FGS and AR SEI messages 902, 903 before them.

[0128] The FGS 902 can be encoded using the H.274 film grain characteristic SEI syntax and semantics.

[0129] To support at least two FGSs in a CVS, it may be necessary to modify the scope of the FGS relative to that currently defined in the VSEI specification such that at least two FGSs can be applied to a single picture. This modification can be done in the standard text, for example, by a) defining a new FGS SEI message that has a different scope definition and / or has the option of carrying at least two film grain characteristics instead of a single one; b) signaling through a dedicated new "different scope SEI" that may need to be in the bitstream order before the FGS, c) implying through a marking mechanism within the AR, as described below. Each of these methods and other methods that may occur to those skilled in the art has advantages and disadvantages that are not further described herein.

[0130] Except for possible modifications to the persistence scope of the FGS, no other modifications may be required.

[0131] The syntax of the AR message 903 can also remain unchanged. However, in an embodiment, when a predefined string is found in the "ar_label[][]" field of the AR message, the semantics of the AR need to be modified to add a specific interpretation of the AR message.

[0132] Referring to Figure 6 and Figure 9, three AR regions are required to represent the regions in picture 604 that may require different FGS features, namely the regions labeled 605, 608, and 609. Specifically, starting from the default of not using film grain for all other content in 604, film grain must be enabled for regions 605 and 609, while film grain must be specifically disabled for region 608. Only the relevant part of the AR 703 content is depicted graphically, following the syntax of AR according to the VSEI standard, and its content can be as follows:

[0133] Control information 911: ar_cancel_flag = 0 […]

[0134] Label definitions 912, 913

[0135] Object definitions 914, 915, 916

[0136] According to an embodiment, the two label strings 912, 913 are defined by a standard rather than by an application. One string can refer to disabling FGS for the region associated with it. For example, the string can be "FGS-D". The other string can refer to enabling FGS and may also refer to certain FGS parameters, such as an intensity value. The string can be in a format such as "FGS-E%03d" (following the C programming language syntax, where %03d indicates three digits forming an integer value). Or, if only FGS on / off signaling is required, the string can be selected as "FGS-E". Those skilled in the art can design other such strings. However, to emphasize again, it may be necessary to specify the strings in a standard, especially in the updated annotation region SEI semantic part, so the choice of the encoder may be restricted.

[0137] Three AR objects (914, 915, 916) may be required, each AR object including a bounding box, where the bounding box can be set to the top corner / left corner coordinates, and the width and height of the regions (605, 608, 609) respectively.

[0138] The ar_label string of the label defined here can be associated with the ar_objects defining the regions through the mapping mechanism defined in the AR SEI message. This relationship is depicted graphically here by arrows 917, 918, 919.

[0139] Through this association, the decoder and the receiving system have sufficient information to identify whether and how to apply FGS to each sample of the reconstructed pictures 904, 905.

[0140] A process that can be performed in the decoder can involve the following steps: The process will be and is, according to an embodiment, the following process: Step 1: Parameter initialization Step 2: Scan the tags and look for "FGS-E" Step 3: Scan the object and look for at least one annotation area marked with FGS. Step 4: If at least one area (object) with FGS is found, apply the simulated film grain value only to the areas associated with the FGS-E and FGS-D tags, and keep the rest of the image unchanged. Otherwise, apply the film grain synthesis to the entire picture by default.

[0141] Reference Figure 10 , within the VSEI specification, the above mechanism can be encoded by inserting the underlined text 1002 into the FGS message semantics 1001. For clarity, the irrelevant parts of the FGS message semantics 1003 have been removed.

[0142] Summarizing the above use case discussion, what is needed and provided by the embodiments is: 1. A mechanism for identifying at least one sample of a reconstructed encoded picture to which film grain processing should be applied in order to obtain the best user experience. Such samples can be represented by rectangular regions, but can also have more complex shapes. Additionally, it may make sense to a) apply different film grain characteristics to different samples, and / or b) apply the film grain with different intensities for different samples. 2. A mechanism for describing film grain characteristics. 3. A mechanism for combining the above two mechanisms, where the sample selection / intensity mechanism is applicable to the film grain.

[0143] In addition, for the 3.38.2 film grain region characteristics SEI message semantics according to the exemplary embodiments, the SEI message provides a parameterized model for the film grain synthesis process to the decoder. Before displaying the decoded picture, the film grain synthesis process should be applied to the decoded picture. The use of this SEI message requires the definition of the following variables:

[0144] - The picture width and picture height in units of luma samples, denoted as PicWidthInLumaSamples and PicHeightInLumaSamples respectively in this document.

[0145] - When the syntax element rdf_separate_colour_description_present_flag of the film grain region characteristics SEI message is equal to 0, the following additional variables:

[0146] - The chroma format indicator, denoted as ChromaFormatIdc in this document, as described in Clause 7.3.

[0147] - The bit depth of the samples of the luma component, denoted as BitDepthY in this document, and when ChromaFormatIdc is not equal to 0, the bit depth of the samples of the two associated chroma components, denoted as BitDepthC in this document.

[0148] The film grain model specified in the film grain region characteristics SEI message is represented as being applied to a decoded picture with a 4:4:4 color format, where the luma and chroma bit depths correspond to the luma and chroma bit depths of the film grain model, and the same color representation domain as the identified film grain model is used. When the color format of the decoded video is not 4:4:4, or the decoded video uses luma or chroma bit depths different from those of the film grain model, or uses a color representation domain different from that of the identified film grain model, an unspecified conversion process is expected to be applied to convert the decoded picture into a form represented as applying the film grain model. Since there is no requirement to use a specific method to perform the film grain generation function used in the display process, the decoder can (if necessary) down-convert the model information for chroma in order to simulate film grain for other chroma formats (4:2:0 or 4:2:2), rather than up-converting the decoded video (using a method not specified in this specification) before performing film grain generation.

[0149] According to an embodiment, fgr_region_top[i], fgr_region_left[i], fgr_region_width[i], and fgr_region_height[i] respectively specify the coordinates of the upper left corner of the bounding box of the i-th region in the picture, as well as the width and height.

[0150] The value of fgr_region_left[i] should be in the range from 0 to PicWidthInLumaSamples - 1, inclusive.

[0151] The value of fgr_region_top[i] should be in the range from 0 to PicHeightInLumaSamples - 1, inclusive.

[0152] The value of fgr_region_width[i] should be in the range from 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive.

[0153] The value of fgr_region_height[i] should be in the range from 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive.

[0154] The above technologies, such as the technology for the SEI message of the above film grain synthesis mask, can be implemented as computer software by computer-readable instructions and physically stored in at least one computer-readable medium, or implemented by at least one specially configured hardware processor. For example, Figure 11 FIG. 1100 shows a computer system, which is suitable for implementing certain embodiments of the disclosed subject matter.

[0155] The computer software can be encoded in any suitable machine code or computer language, and code including instructions is created through mechanisms such as assembly, compilation, and linking. The instructions can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through decoding, microcode, etc.

[0156] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0157] Figure 11The components shown for computer system 1100 are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Nor should the configuration of the components be construed as having any dependency or requirement on any one component or combination thereof shown in the exemplary embodiments of computer system 1100.

[0158] Computer system 1100 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to the input of at least one human user through tactile input (such as keyboard input, swiping, data glove movement), audio input (such as sound, applause), visual input (such as gestures), and olfactory input (not shown). The human-machine interface device may also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0159] The human-machine interface input device may include at least one of the following (only one is shown): keyboard 1101, mouse 1102, touchpad 1103, touch screen 1110, joystick 1105, microphone 1106, scanner 1108, camera 1107.

[0160] Computer system 1100 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of at least one human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (such as tactile feedback through touch screen 1110 or joystick 1105, but there may also be tactile feedback devices that do not serve as input devices), audio output devices (such as speakers 1109, headphones (not shown)), visual output devices (such as screen 1110 including a cathode ray tube screen, liquid crystal screen, plasma screen, organic light emitting diode screen, each of which has or does not have touch screen input function, each of which has or does not have tactile feedback function - some of which may output two-dimensional visual output or output above three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).

[0161] The computer system 1100 may also include human-accessible storage devices and associated media, such as optical media including high-density read-only / rewritable optical discs (CD / DVD ROM / RW) 1120 with CD / DVD 1111 or similar media, thumb drives 1122, removable hard disk drives or solid state drives 1123, conventional magnetic media such as tapes and floppy disks (not shown), ROM / ASIC / PLD-based dedicated devices such as security dongles (not shown), and so on.

[0162] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0163] The computer system 1100 may also include an interface 1199 to at least one communication network 1198. For example, the network 1198 may be wireless, wired, optical. The network 1198 may also be a local area network, wide area network, metropolitan area network, in-vehicle network, and industrial network, real-time network, delay-tolerant network, and so on. The network 1198 also includes local area networks such as Ethernet, wireless local area network, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), television cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial television broadcasting), in-vehicle and industrial networks (including CANBus), and so on. Some networks 1198 typically require an external network interface adapter for connection to certain common data ports or peripheral buses (1150 and 1151) (e.g., the USB port of the computer system 1100); other systems are typically integrated into the core of the computer system 1100 by connecting to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks 1198, the computer system 1100 can communicate with other entities. The communication may be unidirectional, only for receiving (e.g., wireless television), unidirectional only for sending (e.g., CAN bus to certain CAN bus devices), or bidirectional, e.g., via a local or wide area digital network to other computer systems. Each of the above networks and network interfaces may use certain protocols and protocol stacks.

[0164] The above-described human-machine interface devices, human-accessible storage devices, and network interfaces may be connected to the core 1140 of the computer system 1100.

[0165] The core 1140 may include at least one central processing unit (CPU) 1141, a graphics processing unit (GPU) 1142, a graphics adapter 1117, a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) 1143, a hardware accelerator 1144 for specific tasks, etc. These devices, as well as a read-only memory (ROM) 1145, a random access memory 1146, an internal mass storage (such as an internal non-user-accessible hard disk drive, a solid state drive, etc.) 1147, etc. may be connected via a system bus 1148. In some computer systems, the system bus 1148 may be accessed in the form of at least one physical plug so as to be expandable by an additional central processing unit, a graphics processing unit, etc. Peripheral devices may be directly attached to the system bus 1148 of the core or connected via a peripheral bus 114. The architecture of the peripheral bus includes an external peripheral component interconnect PCI, a universal serial bus USB, etc.

[0166] The CPU 1141, GPU 1142, FPGA 1143, and accelerator 1144 may execute certain instructions which, when combined, may constitute the above-mentioned computer code. The computer code may be stored in the ROM 1145 or the RAM 1146. Transitional data may also be stored in the RAM 1146, while permanent data may be stored in, for example, the internal mass storage 1147. Fast storage and retrieval of any memory device may be achieved by using a cache memory which may be closely associated with at least one CPU 1141, GPU 1142, mass storage 1147, ROM 1145, RAM 1146, etc.

[0167] The computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure or may be of the kind well-known and available to those skilled in the computer software art.

[0168] By way of example and not limitation, a computer system having an architecture 1100, and in particular a core 1140, can provide the functionality of a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in at least one tangible computer-readable medium. Such a computer-readable medium can be a medium associated with the user-accessible mass storage described above, as well as a specific memory of the non-volatile core 1140, such as the core internal mass storage 1147 or the ROM 1145. The software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core 1140. Depending on specific needs, the computer-readable medium can include one or more storage devices or chips. The software can cause the core 1140, and in particular the processor therein (including a CPU, GPU, FPGA, etc.), to execute a specific process or a specific part of a specific process described herein, including defining a data structure stored in the RAM 1146 and modifying such a data structure according to a software-defined process. Additionally or alternatively, the computer system can provide logic hardwired or otherwise included in a circuit (e.g., the functionality in the accelerator 1144), which can replace the software or operate together with the software to execute a specific process or a specific part of a specific process described herein. In appropriate cases, a reference to software can include logic, and vice versa. In appropriate cases, a reference to a computer-readable medium can include a circuit (such as an integrated circuit (IC)) storing the software for execution, a circuit including the executing logic, or both. The present disclosure includes any suitable combination of hardware and software.

[0169] Although the present disclosure has been described with respect to at least two exemplary embodiments, various changes, permutations, and various equivalent replacements of the embodiments are within the scope of the present disclosure. It should thus be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

[0170] The foregoing disclosure also includes the features described below. These features can be combined in various ways and are not limited to the combinations mentioned below.

[0171] (1) A method for video decoding, the method comprising: obtaining video data including at least one encoded picture; reconstructing the encoded picture, the encoded picture including a film grain characteristic auxiliary enhancement information SEI message, an alpha channel information SEI message, and at least a first sample and a second sample and at least a third alpha map sample and a fourth alpha map sample, wherein after reconstruction, the first sample and the third alpha map sample, and the second sample and the first sample are spatially co-located; applying film grain synthesis to the first sample based on a first value of the third sample; and determining not to apply film grain synthesis to the second sample based on a second value of the fourth sample, wherein the first value and the second value are different from each other.

[0172] (2) The method according to feature (1), wherein the film grain characteristic SEI message specifies the coordinates of the upper left corner of the bounding box of the i-th region of the picture, as well as the width and height.

[0173] (3) The method according to any one of features (1) to (2), wherein the film grain characteristic SEI message specifies the coordinates by any one of the fgr_region_top[i] syntax, the fgr_region_left[i] syntax, the fgr_region_width[i] syntax, and the fgr_region_height[i] syntax.

[0174] (4) The method according to any one of features (1) to (3), wherein any one of the following: the value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - 1, including 0 and PicWidthInLumaSamples - 1, the value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - 1, including 0 and PicHeightInLumaSamples - 1, the value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - fgr_region_left[i], including 0 and PicWidthInLumaSamples - fgr_region_left[i], and the value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - fgr_region_top[i], including 0 and PicHeightInLumaSamples - fgr_region_top[i].

[0175] (5) The method according to any one of features (1) to (4), wherein each of the following: the value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - 1, including 0 and PicWidthInLumaSamples - 1, the value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - 1, including 0 and PicHeightInLumaSamples - 1, the value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - fgr_region_left[i], including 0 and PicWidthInLumaSamples - fgr_region_left[i], and the value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - fgr_region_top[i], including 0 and PicHeightInLumaSamples - fgr_region_top[i].

[0176] (6) The method according to any one of features (1) to (5), wherein both the first sample and the second sample are rectangular.

[0177] (7) The method according to any one of features (1) to (6), wherein the first sample and the second sample at least partially overlap each other.

[0178] (8) A method for video encoding, the method comprising: obtaining video data including at least one picture; and encoding the video data such that reconstructing the at least one picture includes: reconstructing the at least one picture based on a film grain characteristic assisted enhancement information SEI message, an alpha channel information SEI message, and at least a first sample and a second sample and at least a third alpha map sample and a fourth alpha map sample, wherein after reconstruction, the first sample and the third alpha map sample are spatially co-located, and the second sample and the first sample are spatially co-located, wherein reconstructing the at least one picture further includes: applying film grain synthesis to the first sample based on a first value of the third sample, and determining not to apply film grain synthesis to the second sample based on a second value of the fourth sample, wherein the first value and the second value are different from each other.

[0179] (9) The method according to feature (8), wherein the film grain characteristic SEI message specifies the coordinates of the upper left corner of the bounding box of the i-th region of the picture, as well as the width and height.

[0180] (10) The method according to any one of features (8) to (9), wherein the film grain characteristic SEI message specifies the coordinates by any one of the fgr_region_top[i] syntax, the fgr_region_left[i] syntax, the fgr_region_width[i] syntax, and the fgr_region_height[i] syntax.

[0181] (11) The method according to any one of features (8) to (10), wherein any one of the following: the value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - 1, including 0 and PicWidthInLumaSamples - 1, the value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - 1, including 0 and PicHeightInLumaSamples - 1, the value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - fgr_region_left[i], including 0 and PicWidthInLumaSamples - fgr_region_left[i], and the value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - fgr_region_top[i], including 0 and PicHeightInLumaSamples - fgr_region_top[i].

[0182] (12) The method according to any one of features (8) to (11), wherein each of the following: the value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples-1, inclusive of 0 and PicWidthInLumaSamples-1, the value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples-1, inclusive of 0 and PicHeightInLumaSamples-1, the value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive of 0 and PicWidthInLumaSamples - fgr_region_left[i], and the value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive of 0 and PicHeightInLumaSamples - fgr_region_top[i].

[0183] (13) The method according to any one of features (8) to (12), wherein both the first sample and the second sample are rectangular.

[0184] (14) The method according to any one of features (8) to (13), wherein the first sample and the second sample at least partially overlap each other.

[0185] (15) A method for processing visual media data, the method comprising: performing a conversion between a visual media file and a bitstream of visual media data according to format rules, the format rules indicating: obtaining video data including at least one encoded picture; reconstructing the encoded picture, the encoded picture including a film grain characteristic auxiliary enhancement information SEI message, an alpha channel information SEI message, and at least a first sample and a second sample and at least a third alpha map sample and a fourth alpha map sample, wherein after reconstruction, the first sample is spatially co-located with the third alpha map sample, and the second sample is spatially co-located with the first sample; applying film grain synthesis to the first sample based on a first value of the third sample; and determining not to apply film grain synthesis to the second sample based on a second value of the fourth sample, wherein the first value and the second value are different from each other.

[0186] (16) The method according to feature (15), wherein the film grain characteristic SEI message specifies the coordinates of the upper left corner of the bounding box of the i-th region of the picture, as well as the width and height.

[0187] (17) The method according to any one of features (15) to (16), wherein the film grain characteristic SEI message specifies the coordinates by any one of the fgr_region_top[i] syntax, the fgr_region_left[i] syntax, the fgr_region_width[i] syntax, and the fgr_region_height[i] syntax.

[0188] (18) The method according to any one of features (15) to (17), wherein any one of the following: the value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - 1, inclusive of 0 and PicWidthInLumaSamples - 1, the value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - 1, inclusive of 0 and PicHeightInLumaSamples - 1, the value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive of 0 and PicWidthInLumaSamples - fgr_region_left[i], and the value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive of 0 and PicHeightInLumaSamples - fgr_region_top[i].

[0189] (19)A method according to any one of features (15) to (18), wherein each of the following: the value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples-1, including 0 and PicWidthInLumaSamples-1, the value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples-1, including 0 and PicHeightInLumaSamples-1, the value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples-fgr_region_left[i], including 0 and PicWidthInLumaSamples-fgr_region_left[i], and the value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples-fgr_region_top[i], including 0 and PicHeightInLumaSamples-fgr_region_top[i].

[0190] (20)A method according to any one of features (15) to (19), wherein both the first sample and the second sample are rectangular.

[0191] (21)A video decoding apparatus, comprising a processing circuit configured to perform the method according to any one of features (1) to (7).

[0192] (22)A video encoding apparatus, comprising a processing circuit configured to perform the method according to any one of features (8) to (15).

[0193] (23)A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method according to any one of features (1) to (20).

Claims

1. A method for video decoding, characterized in that, The method includes: Obtaining video data including at least one encoded picture; Reconstructing the encoded picture, the encoded picture including a film grain characteristic assist enhancement information SEI message, an alpha channel information SEI message, and at least a first sample and a second sample and at least a third alpha map sample and a fourth alpha map sample, wherein after reconstruction, the first sample and the third alpha map sample, and the second sample and the first sample are spatially co-located; Applying film grain synthesis to the first sample based on a first value of the third sample; and Determining not to apply film grain synthesis to the second sample based on a second value of the fourth sample, wherein the first value and the second value are different from each other.

2. The method according to claim 1, wherein: The film grain characteristic SEI message specifies the coordinates of the upper left corner of the bounding box of the i-th region of the picture, as well as the width and height.

3. The method according to claim 2, wherein: The film grain characteristic SEI message specifies the coordinates by any one of the fgr_region_top[i] syntax, the fgr_region_left[i] syntax, the fgr_region_width[i] syntax, and the fgr_region_height[i] syntax.

4. The method according to claim 3, characterized in that, Any one of the following: The value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - 1, inclusive of 0 and PicWidthInLumaSamples - 1, The value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - 1, inclusive of 0 and PicHeightInLumaSamples - 1, The value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive of 0 and PicWidthInLumaSamples - fgr_region_left[i], and The value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive of 0 and PicHeightInLumaSamples - fgr_region_top[i].

5. The method according to claim 3, wherein Each of the following: The value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicWidthInLumaSamples - 1, inclusive of 0 and PicWidthInLumaSamples - 1, The value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicHeightInLumaSamples - 1, inclusive of 0 and PicHeightInLumaSamples - 1, The value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive of 0 and PicWidthInLumaSamples - fgr_region_left[i], and The value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive of 0 and PicHeightInLumaSamples - fgr_region_top[i].

6. The method according to claim 1, wherein, Both the first sample and the second sample are rectangular.

7. The method according to claim 1, wherein, The first sample and the second sample at least partially overlap each other.

8. A video encoding method, characterized in that, The method includes: Obtaining video data including at least one picture; and Encoding the video data such that reconstructing the at least one picture includes: reconstructing the at least one picture based on a film grain characteristic assisted enhancement information SEI message, an alpha channel information SEI message, and at least a first sample and a second sample and at least a third alpha map sample and a fourth alpha map sample, wherein after reconstruction, the first sample is spatially co-located with the third alpha map sample, and the second sample is spatially co-located with the first sample, wherein reconstructing the at least one picture further includes: applying film grain synthesis to the first sample based on a first value of the third sample, and determining not to apply film grain synthesis to the second sample based on a second value of the fourth sample, wherein the first value and the second value are different from each other.

9. The method according to claim 8, wherein, The film grain characteristic SEI message specifies the coordinates of the upper left corner of the bounding box of the i-th region of the picture, as well as the width and height.

10. The method according to claim 9, wherein, The film grain characteristic SEI message specifies the coordinates by any one of the fgr_region_top[i] syntax, the fgr_region_left[i] syntax, the fgr_region_width[i] syntax, and the fgr_region_height[i] syntax.

11. The method according to claim 10, wherein Any one of the following: The value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicWidthInLumaSamples - 1, inclusive of 0 and PicWidthInLumaSamples - 1, The value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicHeightInLumaSamples - 1, inclusive of 0 and PicHeightInLumaSamples - 1, The value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive of 0 and PicWidthInLumaSamples - fgr_region_left[i], and The value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive of 0 and PicHeightInLumaSamples - fgr_region_top[i].

12. The method according to claim 10, wherein Each of the following: The value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicWidthInLumaSamples - 1, inclusive of 0 and PicWidthInLumaSamples - 1, The value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicHeightInLumaSamples - 1, inclusive of 0 and PicHeightInLumaSamples - 1, The value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range of 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive of 0 and PicWidthInLumaSamples - fgr_region_left[i], and The value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive of 0 and PicHeightInLumaSamples - fgr_region_top[i].

13. The method according to claim 8, wherein, both the first sample and the second sample are rectangular.

14. The method according to claim 8, wherein, the first sample and the second sample at least partially overlap each other.

15. A method for processing visual media data, characterized in that, The method includes: performing a conversion between a visual media file and a bitstream of visual media data according to format rules, the format rules indicating: obtaining video data including at least one encoded picture; reconstructing the encoded picture, the encoded picture including a film grain characteristic auxiliary enhancement information SEI message, an alpha channel information SEI message, and at least a first sample and a second sample and at least a third alpha map sample and a fourth alpha map sample, wherein after reconstruction, the first sample is spatially co-located with the third alpha map sample, and the second sample is spatially co-located with the first sample; applying film grain synthesis to the first sample based on a first value of the third sample; and determining not to apply film grain synthesis to the second sample based on a second value of the fourth sample, wherein the first value and the second value are different from each other.

16. The method according to claim 15, wherein, the film grain characteristic SEI message specifies the coordinates of the upper left corner of the bounding box of the i-th region of the picture, as well as the width and height.

17. The method according to claim 16, wherein, the film grain characteristic SEI message specifies the coordinates by any one of the fgr_region_top[i] syntax, the fgr_region_left[i] syntax, the fgr_region_width[i] syntax, and the fgr_region_height[i] syntax.

18. The method according to claim 17, wherein Any one of the following: the value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - 1, inclusive of 0 and PicWidthInLumaSamples - 1, the value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - 1, inclusive of 0 and PicHeightInLumaSamples - 1, The value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive of 0 and PicWidthInLumaSamples - fgr_region_left[i], and The value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive of 0 and PicHeightInLumaSamples - fgr_region_top[i].

19. The method according to claim 17, wherein Each of the following: The value of the fgr_region_top[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - 1, inclusive of 0 and PicWidthInLumaSamples - 1, The value of the fgr_region_left[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - 1, inclusive of 0 and PicHeightInLumaSamples - 1, The value of the fgr_region_width[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicWidthInLumaSamples - fgr_region_left[i], inclusive of 0 and PicWidthInLumaSamples - fgr_region_left[i], and The value of the fgr_region_height[i] syntax is represented in the film grain characteristic SEI message as being in the range from 0 to PicHeightInLumaSamples - fgr_region_top[i], inclusive of 0 and PicHeightInLumaSamples - fgr_region_top[i].

20. The method according to claim 15, wherein Both the first sample and the second sample are rectangular.