Descriptors used for film grain synthesis
By introducing SEI messages to indicate film grain information in video coding and combining them with intra-frame and inter-frame prediction techniques, the problem of video quality differences in film grain synthesis is solved, and video optimization and artifact reduction are achieved in machine consumption environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2025-01-09
- Publication Date
- 2026-07-31
AI Technical Summary
Existing video coding technologies struggle to effectively indicate the type, purpose, and importance of film grain when processing film grain synthesis, leading to differences in video quality between machine and human consumption and potentially introducing artifacts.
By introducing supplemental enhancement information (SEI) messages into the encoded video stream, the type, purpose, and importance of film grain are indicated for sample processing in specific areas, and video encoding and decoding are performed in combination with intra-frame prediction and inter-frame prediction techniques.
It achieves video quality optimization in machine-consumed environments, reduces artifacts, meets the machine's analysis needs for high-contrast content, and maintains video quality acceptable to humans.
Smart Images

Figure CN122498151A_ABST
Abstract
Description
[0001] Related applications This application claims priority to U.S. Application No. 19 / 014,008, filed January 8, 2025, which in turn claims priority to U.S. Provisional Application No. 63 / 619,306, filed January 9, 2024, entitled “DESCRIPTORS FOR FILM GRAIN SYNTHESIS”. The entire disclosure of the prior applications is incorporated herein by reference. Technical Field
[0002] This disclosure describes aspects generally related to video coding, including film grain composition. Background Technology
[0003] The background description provided herein is for the purpose of presenting the overall context of this disclosure. The extent of the work of the currently named inventors described in this background section and in various aspects of this specification does not imply that it was prior art at the time of filing, nor is it expressly or implied that it was acknowledged as prior art to this disclosure.
[0004] Image / video compression helps transmit image / video data between different devices, storage devices, and networks with minimal quality degradation. In some examples, video codecs can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which compresses images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current image being reconstructed to predict samples. In another example, a video codec can use a technique called inter-frame prediction, which compresses images based on temporal redundancy. For example, inter-frame prediction can use motion compensation to predict samples in the current image based on previously reconstructed images. Motion compensation can be indicated by motion vectors (MV). Summary of the Invention
[0005] Various aspects of this disclosure include methods and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding / encoding includes processing circuitry.
[0006] A video decoding method includes: receiving an encoded video stream including encoded information of an encoded picture, and receiving a Supplemental Enhancement Information (SEI) message associated with the encoded picture; and reconstructing the encoded picture associated with the SEI message. The SEI message indicates one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain, applied to a first region in the encoded picture including one or more first samples.
[0007] A video encoding method includes: encoding an image in a video stream; and encoding a Supplemental Enhancement Information (SEI) message associated with the image. The SEI message indicates one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain, applied to a first region in the image including one or more first samples.
[0008] Some aspects of this disclosure provide a method for processing visual media data. The method includes processing a bitstream of visual media data according to format rules. The bitstream includes encoding information of an encoded image and Supplemental Enhancement Information (SEI) messages associated with the encoded image. The SEI message indicates one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain, applied to a first region in the encoded image including one or more first samples. The format rules specify that the encoded image associated with the SEI message is reconstructed.
[0009] According to another aspect of this disclosure, an apparatus is provided. The apparatus includes processing circuitry. The processing circuitry can be configured to perform any of the described methods for video decoding / encoding.
[0010] This disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the described methods for video decoding / encoding. Attached Figure Description
[0011] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which: Figure 1 Block diagrams of communication systems in some examples are shown.
[0012] Figure 2 Block diagrams of some example video processing systems are shown.
[0013] Figure 3 Block diagrams of video decoders in some examples are shown.
[0014] Figure 4 Block diagrams of video encoders are shown in some examples.
[0015] Figure 5 The layout of encoded video sequences (CVS) is shown in some examples.
[0016] Figure 6An example of an SEI message according to one aspect of this disclosure is shown. The SEI message includes information indicating the type of film grain, the importance of the film grain, and / or the purpose of the film grain.
[0017] Figure 7 An example of the syntax of an SEI message according to one aspect of this disclosure is shown.
[0018] Figure 8A Examples of possible values for a syntax element (e.g., fge_film_grain_type) according to one aspect of this disclosure are shown.
[0019] Figure 8B Examples of possible values for a syntactic element (e.g., fge_film_grain_purpose) according to one aspect of this disclosure are shown.
[0020] Figure 8C Examples of possible values for a syntax element (e.g., fge_film_grain_essentiality) according to one aspect of this disclosure are shown.
[0021] Figure 9 A flowchart outlining a process (900) according to one aspect of this disclosure is shown.
[0022] Figure 10 A flowchart outlining a process (1000) according to one aspect of this disclosure is shown.
[0023] Figure 11 It is a schematic diagram based on one aspect of a computer system. Detailed Implementation
[0024] Some aspects of this disclosure provide techniques for video encoding and decoding (e.g., encoding and decoding). In some examples, such techniques are used in SEI messages that indicate descriptors used for film grain composition.
[0025] Video coding techniques compress video data. Video coding and decoding using inter-frame picture prediction and motion compensation have been used in various examples. In some examples, uncompressed digital video may comprise a series of pictures, each with a spatial size of, for example, 1920×1080 luminance samples and associated chrominance samples. The series of pictures may have a fixed or variable picture rate (informally, also called a frame rate) of, for example, 60 pictures per second or 60Hz. Uncompressed video has specific bit rate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (at a 60Hz frame rate, with a 1920×1080 luminance sample resolution) requires close to 1.5 Gbit / s of bandwidth. One hour of such video would require more than 600GB of storage space.
[0026] Video encoding and decoding technologies (e.g., encoding and decoding techniques) can reduce redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth or storage space requirements, in some cases by two orders of magnitude or more. Lossless compression, lossy compression, and combinations thereof can be employed. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may differ from the original signal, but the distortion between the original and reconstructed signals is small enough that the reconstructed signal can be used for the intended application. In the case of video, lossy compression is widely used. The tolerable amount of distortion depends on the application; for example, users of some consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can be reflected: higher tolerable / acceptable distortion results in a higher compression ratio.
[0027] Video encoders and decoders can utilize techniques from a wide range of categories, including, for example, motion compensation, transform, quantization, and entropy coding, some of which will be described below.
[0028] This document discloses systems and methods for video encoding and decoding in at least one processor. For example, a video decoding method may include: receiving at least one picture including an SEI message, the SEI message including at least one syntax element indicating at least one of a forward truncation function and a backward truncation function; and reconstructing an encoded picture including samples, wherein the numbering range of the samples is determined by at least one of the forward truncation function and the backward truncation function indicated by the SEI message.
[0029] In certain environments, such as video encoding and decoding for machine consumption (as opposed to human consumption), it may not be necessary for the bitstream to exceed a quality threshold based on human perception. Instead, its quality may be sufficient for machine consumption, even if it is insufficient for human consumption.
[0030] In some examples, techniques for reducing the range of sample values are used, which can employ preprocessing performed before encoding and corresponding postprocessing performed after decoding. According to one aspect of this disclosure, such processing of a natural sequence can result in artifacts, such as banding artifacts, which are not pleasing to human receivers but may be acceptable for machine consumption (e.g., image analysis of high-contrast content such as barcodes). In some examples, the bit depth of a VVC image can be reduced from 10 bits to 5 bits—corresponding to 32 grayscale and color component levels. For example, in the original image, 10 bits are used to represent grayscale levels, while after the bit depth reduction, 5 bits (e.g., the 5 most important bits out of 10) are used to represent grayscale levels. In some examples, when such preprocessing is performed before encoding, the decoder and associated processor need to know how to perform the preprocessing. Some aspects of this disclosure provide techniques for providing such information for video encoding and decoding.
[0031] Figure 1 A block diagram of a communication system (100) is shown in some examples. The system (100) includes at least two terminals interconnected via a network (150), for example... Figure 1 The diagram shows a first terminal (110) and a second terminal (120). In some examples, one-way data transmission is performed in the communication system (100). In one example, for one-way data transmission, the first terminal (110) may encode video data at a local location for transmission over a network (150) to the second terminal (120). The second terminal (120) may receive encoded video data from another terminal (e.g., the first terminal (110)) from the network (150), decode the encoded data, and display the recovered video data. It should be noted that one-way data transmission is typically used in applications such as media services.
[0032] Figure 1 A second terminal pair, such as a third terminal (130) and a fourth terminal (140), is also shown, which are configured to support bidirectional transmission of encoded video, for example, that may occur during a video conference. For bidirectional data transmission, each of the third terminal (130) and the fourth terminal (140) can encode video data acquired at a local location for transmission over a network (150) to the other terminal. Each of the third terminal (130) and the fourth terminal (140) can also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the recovered video data on a local display device.
[0033] It should be noted that, although in Figure 1In this disclosure, terminals (110), (120), (130), and (140) are shown as servers, personal computers, and smartphones, but this disclosure is not limited to such terminal examples. Aspects of this disclosure may include applications in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (150) refers to any number of networks, including, for example, wired and / or wireless communication networks, that transmit encoded video data between terminals (110), (120), (130), and (140). Network (150) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this discussion, the architecture and topology of network (150) may be irrelevant to the operation of this disclosure unless otherwise stated below. In some examples, the network (150) includes a Media Aware Network Unit (MANE, 160), which may be included in a transmission path, for example, between a third terminal (130) and a fourth terminal (140). In some examples, the MANE (160) may selectively forward a portion of media data in response to network congestion, media switching, media mixing, archiving, and similar tasks typically performed by service providers (rather than end users). Such a MANE is capable of resolving and reacting to a limited portion of the media transmitted over the network (e.g., syntactic elements associated with video codec technologies or standard network abstraction layers).
[0034] Figure 2 Block diagrams of some example video processing systems (200) are shown. The video processing system (200) is an application example of the disclosed subject matter, namely a video encoder and video decoder located in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0035] The video processing system (200) includes an acquisition subsystem (213), which may include a video source (201), such as a digital camera, which creates, for example, an uncompressed video image stream (202). In one example, the video image stream (202) includes samples captured by a digital camera. The video image stream (202), depicted as a thick line to emphasize its high data volume, may be processed by an electronic device (220), which includes a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data (204) (or encoded video stream), depicted as a thin line to emphasize its lower data volume, may be stored on a streaming server (205) for future use. One or more streaming client subsystems, such as Figure 2 Client subsystems (206) and (208) can access a streaming server (205) to retrieve copies (207) and (209) of encoded video data (204). Client subsystem (206) may include, for example, a video decoder (210) in an electronic device (230). The video decoder (210) decodes the incoming copy (207) of the encoded video data and creates an output video picture stream (211) that can be displayed on a monitor (212) (e.g., a display screen) or other presentation device (not depicted). In some streaming systems, the encoded video data (204), (207), and (209) (e.g., a video stream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed topics are applicable in the context of VVC.
[0036] It should be noted that electronic devices (220) and (230) may include other components (not shown). For example, electronic device (220) may include a video decoder (not shown), and electronic device (230) may also include a video encoder (not shown).
[0037] Figure 3 An exemplary block diagram of a video decoder (310) is shown. The video decoder (310) may be included in an electronic device (330). The electronic device (330) may include a receiver (331) (e.g., receiving circuitry). The video decoder (310) may be used in place of... Figure 2 The example video decoder (210).
[0038] The receiver (331) may receive, for example, one or more coded video sequences included in the bitstream that will be decoded by the video decoder (310). In one aspect, one coded video sequence is received at a time, wherein the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences may be received from a channel (301), which may be a hardware / software link to a storage device storing the coded video data. The receiver (331) may receive coded video data and other data, such as coded audio data and / or auxiliary data streams that may be forwarded to their respective user entities (not depicted). The receiver (331) may separate the coded video sequences from other data. To prevent network jitter, a buffer (315) may be coupled between the receiver (331) and the entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). In some applications, the buffer (315) is part of the video decoder (310). In other applications, the buffer (315) may be located outside the video decoder (310) (not depicted). In other applications, a buffer memory (not depicted) may be provided externally to the video decoder (310) to, for example, prevent network jitter. Additionally, another buffer memory (315) may be provided internally to the video decoder (310) to, for example, handle playback timing. When the receiver (331) receives data from a store / forward device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (315) may not be necessary, or it may be made smaller. For use on packet-switched networks such as the Internet, the buffer memory (315) may be required. The buffer memory (315) may be relatively large, advantageously having an adaptive size, and may be implemented at least partially in the operating system or in a similar component (not depicted) external to the video decoder (310).
[0039] The video decoder (310) may include a parser (320) to reconstruct symbols (321) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (310) and potential information for controlling a presentation device such as a presentation device (312) (e.g., a display screen), which is not part of the electronic device (330) but may be coupled to it. Figure 3As shown. The control information used for the presentation device may be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (320) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may be performed according to video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (320) may extract a subgroup parameter set of at least one subgroup of pixels in the encoded video sequence for use in the video decoder based on at least one parameter corresponding to a group. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The parser (320) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0040] The parser (320) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (315) to create symbols (321).
[0041] Depending on the type of encoded video picture or a subset of encoded video pictures (e.g., inter-frame pictures and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of the symbol (321) may involve multiple different units. Which units are involved and how they are involved can be controlled by the parser (320) through subgroup control information parsed from the encoded video sequence. For brevity, the flow of such subgroup control information between the parser (320) and the various units described below is not depicted.
[0042] In addition to the functional blocks already mentioned, the video decoder (310) can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into multiple functional units as described below.
[0043] The first unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives quantization transform coefficients as symbols (321) from the parser (320) and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (351) can output a block containing sample values, which can be input into the aggregator (355).
[0044] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to an intra-coded block. An intra-coded block is a block that does not use prediction information from a previously reconstructed image, but can use prediction information from a previously reconstructed portion of the current image. Such prediction information may be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses surrounding reconstructed information extracted from the current picture buffer (358) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (358) buffers a partially reconstructed current image and / or a fully reconstructed current image. In some cases, the aggregator (355) adds the prediction information generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) based on each sample.
[0045] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to blocks of inter-frame coding and potential motion compensation. In this case, the motion compensation prediction unit (353) may access the reference image memory (357) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (321) belonging to the block, these samples may be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (in this case, referred to as residual samples or residual signals) to generate output sample information. The extraction of prediction samples by the motion compensation prediction unit (353) from the address in the reference image memory (357) may be controlled by motion vectors, which may be available to the motion compensation prediction unit (353) in the form of symbols (321), which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values extracted from the reference image memory (357) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0046] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), which can be used by the loop filter unit (356) as symbols (321) from the parser (320). Video compression may also respond to metadata obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0047] The output of the loop filter unit (356) can be a sample stream that can be output to the presentation device (312) and stored in the reference image memory (357) for future inter-frame image prediction.
[0048] Once fully reconstructed, certain encoded images can be used as reference images for future predictions. For example, once the encoded image corresponding to the current image has been fully reconstructed and the encoded image (by, for example, the parser (320)) is identified as the reference image, the current image buffer (358) can become part of the reference image memory (357), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.
[0049] The video decoder (310) can perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T Recommendation H.265. The encoded video sequence may conform to the syntax specified by the video compression technology or standard in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the configuration file documented in the video compression technology or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technology or standard as the only tools available under that configuration file. For compliance, it may also be necessary to keep the complexity of the encoded video sequence within the limits defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata managed by the HRD buffer, which is represented as a signal in the encoded video sequence.
[0050] On one hand, the receiver (331) can receive additional (redundant) data when receiving the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (310) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0051] Figure 4 An exemplary block diagram of a video encoder (403) is shown. The video encoder (403) is included in an electronic device (420). The electronic device (420) includes a transmitter (440) (e.g., transmission circuitry). The video encoder (403) can be used in place of... Figure 2 The video encoder (203) in the example.
[0052] The video encoder (403) can obtain data from the video source (401) (not...). Figure 4 In one example, an electronic device (420) receives video samples, and a video source (401) can capture video images that will be encoded by a video encoder (403). In another example, the video source (401) is part of the electronic device (420).
[0053] A video source (401) can provide a sequence of source video samples to be encoded by a video encoder (403) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (401) can be a storage device storing previously prepared video. In a video conferencing system, the video source (401) can be a camera that captures local image information as a video sequence. The video data can be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves can be organized into a spatial pixel array, where each pixel can include one or more samples, depending on the sampling structure, color space, etc. used. The following focuses on describing the samples.
[0054] According to one aspect, the video encoder (403) can encode and compress images of a source video sequence into an encoded video sequence (443) in real time or under any other required time constraints. Implementing an appropriate encoding rate is a function of the controller (450). In some aspects, the controller (450) controls and is functionally coupled to other functional units described below. For clarity, the coupling is not depicted in the figures. Parameters set by the controller (450) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (450) can be configured to have other suitable functions related to the video encoder (403) optimized for a particular system design.
[0055] In some respects, the video encoder (403) is configured to operate within an encoding loop. As an oversimplification, in one example, the encoding loop may include a source encoder (430) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (433) embedded within the video encoder (403). The decoder (433) reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder can also create sample data. The reconstructed sample stream (sample data) is input to a reference image memory (434). Since decoding of the symbol stream produces bit-accurate results independent of the decoder's location (local or remote), the contents of the reference image memory (434) are also bit-accurately corresponding between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same sample values that the decoder will "see" when using the prediction during decoding. This fundamental principle of reference image synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related techniques.
[0056] The operation of the “local” decoder (433) can be combined with, for example, the operations already described above. Figure 3 The video decoder (310) described in detail is the same as a "remote" decoder. However, a brief additional reference is provided. Figure 3 When symbols are available and the entropy encoder (445) and parser (320) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (310), including the buffer (315) and parser (320), may not be fully implemented in the local decoder (433).
[0057] On the one hand, aside from the parsing / entropy decoding present in the decoder, the decoder techniques exist in the corresponding encoder in the same or substantially the same functional form. Therefore, the subject matter disclosed focuses on decoder operation. The description of the encoder techniques can be simplified, as the encoder techniques are inverses of the fully described decoder techniques. In certain areas, more detailed descriptions are provided below.
[0058] During operation, in some examples, the source encoder (430) may perform motion-compensated predictive coding, which predictively encodes the input image by referencing one or more previously encoded images from the video sequence designated as "reference images." In this way, the encoding engine (432) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.
[0059] The local video decoder (433) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (430). Advantageously, the operation of the encoding engine (432) can be a lossy process. When the encoded video data can be decoded by the video decoder (430), Figure 4 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (433) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in the reference image memory (434). In this way, the video encoder (403) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.
[0060] The predictor (435) can perform a prediction search against the encoding engine (432). That is, for a new image to be encoded, the predictor (435) can search in the reference image memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new image. The predictor (435) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (435), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (434).
[0061] The controller (450) can manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.
[0062] The outputs of all the above functional units can be entropy encoded in the entropy encoder (445). The entropy encoder (445) converts the symbols generated by the various functional units into a encoded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, and arithmetic coding.
[0063] The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) in preparation for transmission via a communication channel (460), which may be a hardware / software link to a storage device capable of storing the encoded video data. The transmitter (440) can combine the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0064] The controller (450) manages the operation of the video encoder (403). During encoding, the controller (450) can assign a specific encoding image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types: Intra-frame pictures (I-pictures) are pictures that can be encoded and decoded without using any other pictures in the sequence as prediction sources. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0065] A predictive image (P-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses a motion vector and a reference index to predict the sample value for each block.
[0066] Bidirectional predictive images (B-images) can be images that can be encoded and decoded using intra-frame or inter-frame prediction, which uses two motion vectors and a reference index to predict sample values for each block. Similarly, multiple predictive images can be used to reconstruct a single block using more than two reference images and associated metadata.
[0067] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 samples), and each block is encoded sequentially. These blocks can be predictively coded with reference to other (already coded) blocks, which are determined by the coding assignments of the corresponding images applied to the blocks. For example, blocks of an I-image can be non-predictively coded, or blocks of an I-image can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be predictively coded with reference to a previously coded reference image via spatial or temporal prediction. Blocks of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial or temporal prediction.
[0068] The video encoder (403) can perform encoding operations according to a predetermined video coding technology or standard, such as ITU-T H.266 Recommendation. In operation, the video encoder (403) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technology or standard used.
[0069] On one hand, the transmitter (440) can transmit additional data while transmitting encoded video. The source encoder (430) can include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, video availability information (VUI) parameter set fragments, etc.
[0070] The captured video can be presented as multiple source images (video images) in a time-series format. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In one example, a specific image being encoded / decoded is segmented into blocks; this specific image being encoded / decoded is called the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector may have a third dimension that identifies the reference image.
[0071] In some respects, bidirectional prediction techniques can be used for inter-frame image prediction. According to this technique, two reference images are used, such as a first reference image and a second reference image that precede the current image in the video in decoding order (but may be past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. The block can be predicted using a combination of the first and second reference blocks.
[0072] In addition, merging mode techniques can be used for inter-frame image prediction to improve coding efficiency.
[0073] According to some aspects of this disclosure, predictions such as inter-frame picture prediction and intra-frame picture prediction are performed on a block-by-block basis. For example, according to the High-Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU consists of three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU can be recursively partitioned into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be partitioned into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type used for the CU, such as inter-frame prediction or intra-frame prediction. CUs are further partitioned into one or more prediction units (PUs) based on temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one respect, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. Using a luma prediction block as an example, a prediction block comprises a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0074] It should be noted that any suitable technology can be used to implement the video encoder (203) and video encoder (403), as well as the video decoder (210) and video decoder (310). On one hand, the video encoder (203) and video encoder (403), as well as the video decoder (210) and video decoder (310), can be implemented using one or more integrated circuits. On the other hand, the video encoder (203) and video encoder (403), as well as the video decoder (210) and video decoder (310), can be implemented using one or more processors that execute software instructions.
[0075] Video encoders and video decoders can utilize techniques from a wide range of categories, including, for example, motion compensation, transform, quantization, entropy coding, and carrying supplementary information (e.g., metadata that describes the images in the encoded bitstream).
[0076] On one hand, compressed video and / or images can be enhanced in the video stream by supplemental enhancement information, such as in the form of Supplemental Enhancement Information (SEI) messages or VUIs. In some examples, video coding standards may include specifications for SEI and VUI. In some examples, SEI and VUI information may also be specified in a separate specification that can be referenced by the video coding standard.
[0077] The techniques used in video coding standards may include one or more SEI messages that enable, for example, the carrying of supplementary information to the encoded video within the encoded bitstream. Such SEI information may or may not be directly related to the video coding process, for example, as specified by video standards (e.g., H.264|AVC, H.265|HEVC, H.266|VVC, etc.). In many cases, the information in the SEI message may be related to an application process that cooperates with or immediately follows the video decoding process. For example, such applications may include a rendering process that uses certain SEI messages to adjust the brightness or color space of decoded video frames (also called pictures) before being displayed by a display device.
[0078] On the one hand, within current standards that utilize SEI messages (e.g., H.264|AVC, H.265|HEVC, and H.266|VVC), SEI messages can be divided into two categories: the first category of SEI messages that can affect the video decoding process; and the second category of SEI messages that do not affect the video decoding process (e.g., for external applications). SEI messages that do not affect the decoding process can be specified in separate specifications, such as the specification entitled "General Supplemental Enhancement Information Messages for Encoding Video Bitstreams" (VSEI). SEI messages that can affect the decoding process can be specified in master coding specifications, such as "General Video Coding".
[0079] This disclosure includes video encoding and decoding, such as applying film grain or similar noise to a portion of an image, the application of which is controlled by metadata, including metadata encoded in SEI messages (e.g., VSEI annotation area SEI messages).
[0080] On one hand, film grain synthesis (FGS) is a tool designed to preserve the impression of the original film grain, for example, for content shot using chemical film (as opposed to digital cameras) in a digitally compressed video environment. The digitized input material can be pre-filtered, which removes film grain, but this can contribute to better compression efficiency in subsequent encoding steps compared to input material that includes noise such as film grain. Artificial approximations of film grain can be reinserted after reconstruction. The amount and characteristics of the noise can be part of the video bitstream, for example, in the form of metadata. For example, ITU-T Recommendation H.274 includes a Film Grain Characteristics (FGC) SEI message.
[0081] On one hand, the parameters specified in the FGC SEI message are generally designed to reproduce the physical properties of a particular film. Some examples of popular grain (or film grain) include: 8mm film, which has coarse grain and was used, for example, for private film recording in the 1970s and 80s; 16mm film, which has noticeable grain and was used for documentaries and television (TV) programs; 35mm film, which has finer grain that can be identified as a historical film effect; 65-70mm film, which was used for blockbuster film production, such as those requiring VFX, and so on. Of the grain of 8mm, 16mm, 35mm, and 65-70mm film, 65-70mm film has the finest grain.
[0082] In one example, using a signal to indicate the type of particle can inform the client of the type of particles to be added to the video.
[0083] On the one hand, applying noise to video can address different objectives. These objectives can be artistic intents. Within this category, there are two scenarios: (i) film grain has been removed before encoding, and the FGC SEI message aims to reproduce the original version; or (ii) grain is selected to provide an arbitrary artistic intent similar to the video effect. Another objective of applying film grain is to mask visual artifacts and restore some texture effects, for example, in areas where loop filtering is critical.
[0084] When describing film grain, it may be helpful to use signals to indicate how important post-processing of FGS is to the target service. Importance levels may include (i) high importance for maintaining the original artistic intent, (ii) moderate importance as a subjectively preferred version, and (iii) low importance limited to, for example, obscuring artifacts.
[0085] In some examples, the video signal may consist of multiple sources. For instance, a film with narration may consist of the film content and surrounding frames with narration. The film may be shot using chemical film techniques, and therefore may include film grain after digitization. To achieve reasonable compression, the film grain of the film can be removed by pre-filtering. The surrounding frames can be created digitally and do not contain film grain. After encoding, transmission, and reconstruction, the resulting reconstructed image does not include noise associated with the film grain of the film, which may be directed at the surrounding frames but not at the film itself, since the film grain of the film has been removed by pre-filtering. Therefore, a technique can be used in which film grain can be selectively inserted into a portion of the reconstructed image, while leaving the unselected portions of the reconstructed image grain-free. More complex scenarios may involve multiple films with different film grain characteristics in the composite image. In such scenarios, it is advantageous to apply different film grain re-insertion parameters separately to different portions of the reconstructed image.
[0086] Figure 5 The diagram illustrates the layout of, for example, a encoded video sequence (CVS) according to H.266. The encoded video sequence is subdivided into Network Abstraction Layer (NAL) units (NAL units), for example... Figure 5 The NAL unit (501) in the text. Figure 5 In the example, the NAL cell (501) may include a NAL cell header (502). In some examples, the NAL cell header (502) includes 16 bits. Figure 5 In the example, the NAL cell header (502) includes a first bit (e.g., forbidden_zero_bit) (503) and a second bit (e.g., nuh_reserved_zero_bit) (504). In one example, the first and second bits are not used by H.266 and can be set to zero in an H.266 compliant NAL cell.
[0087] exist Figure 5In the example, the NAL unit header (502) includes a syntax element nuh_layer_id (505) with multiple bits (e.g., 6 bits). In one example, 3 bits of nuh_layer_id (505) may indicate the (spatial, SNR, or multi-view enhancement) layer to which the NAL unit (501) belongs. The NAL unit header (502) includes nuh_nal_unit_type (506) with 5 bits that defines the type of the NAL unit (501). In some examples (e.g., H.266), of the 32 values represented by 5 bits, 22 NAL unit type values are defined for the NAL unit type, 6 NAL unit types are reserved, and 4 NAL unit type values are unspecified and can be used by specifications other than H.266. The NAL unit header (502) includes nuh_temporal_id_plus1 (507) with 3 bits to indicate the temporal layer to which the NAL unit (501) belongs.
[0088] NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units may include data representing sample values in video frames, while non-VCL NAL units may contain associated additional information, such as one or more parameter sets, SEI, etc.
[0089] In some examples, a encoded picture may include one or more VCL NAL units and zero or more non-VCL NAL units. VCL NAL units may contain encoded data that conceptually belongs to the video coding layer as previously described. Non-VCL NAL units may contain data that conceptually does not belong to the video coding layer. For example, using H.266 as an example, non-VCL NAL units can be categorized as follows: (1) Parameter set, which includes information that can be used in the decoding process and applied to more than one encoded picture. The parameter set and the conceptually similar NAL unit can be NAL unit type (NUT), such as DCI_NUT (Decoding Capability Information (DCI)), VPS_NUT (Video Parameter Set (VPS), especially for establishing layer relationships), SPS_NUT (Sequence Parameter Set (SPS), especially for establishing parameters used and kept constant throughout the encoded video sequence CVS), PPS_NUT (Picture Parameter Set (PPS), especially for establishing parameters used and kept constant within the encoded picture), PREFIX_APS_NUT and SUFFIX_APS_NUT (Prefix and Suffix Adaptation Parameter Sets). The parameter set may include the information required by the decoder to decode the VCL NAL unit, and is therefore referred to here as a "normalized" NAL unit.
[0090] (2) Image header (PH_NUT), which is also a "normalized" NAL unit.
[0091] (3) NAL units that mark certain positions in the NAL unit stream. The third category includes NAL units with NAL unit types AUD_NUT (Access Unit Separator), EOS_NUT (End of Sequence), and EOB_NUT (End of Stream). NAL units of the third category are considered non-normalized units, also known as informative units, in the sense that a compliant decoder does not need to use them in its decoding process. In one example, a compliant decoder may need to be able to receive NAL units of the third category in the NAL unit stream.
[0092] (4) Prefix and suffix SEI NAL unit types (PREFIX_SEI_NUT and SUFFIX_SEI_NUT), which indicate NAL units containing prefix and suffix supplementary enhancement information. In some examples (e.g., in H.266), the fourth category of NAL units is informative because it is not necessary to use the fourth category of NAL units in the decoding process.
[0093] (5) Padding data NAL cell type FD_NUT, which indicates padding data. Padding data can be random data and can be used to “waste” bits in the NAL cell stream or bit stream, which may be necessary for transmissions in certain isochronous transmission environments.
[0094] (6) Reserved and unspecified NAL cell types.
[0095] Figure 5 Also shown in some examples is the layout of the NAL unit stream (510) in decoding order (590). The NAL unit stream (510) includes encoded pictures (511). The NAL unit stream (510) includes DCI (512), VPS (513), and SPS (514) located somewhere earlier than the encoded pictures (511). DCI (512), VPS (513), and SPS (514) can be combined to establish parameters that the decoder can use to decode the encoded pictures of the encoded video sequence (CVS), including the encoded pictures (511) in the NAL unit stream (510).
[0096] exist Figure 5 In the examples, the order is as depicted or conforms to the video coding technology or standard used (e.g., Figure 5 In any other order of H.266 shown, the encoded picture (511) may include a prefix APS (516), a picture header (PH) (617), a prefix SEI (518), one or more VCL NAL units (519), and a suffix SEI (520).
[0097] In some examples, during standard development, the prefix SEI NAL unit (518) and suffix SEI NAL unit (520) are configured such that, for some SEI messages, the content of the message can be known before the encoding of a given image begins, while in other examples, the content can only be known once the image has been encoded. Allowing certain SEI messages to appear earlier or later in the NAL unit stream of the encoded image by using prefix and suffix SEI avoids buffering. For example, in an encoder, the sampling time of the image to be encoded is known before the image is encoded, so the image timing SEI message can be a prefix SEI message (518). On the other hand, the decoded image hash SEI message (which contains the hash of the sample values of the decoded image and can be used, for example, to debug the encoder implementation) is a suffix SEI message (520) because the encoder cannot compute the hash of the reconstructed sample before the image has been encoded. The positions of the prefix and suffix SEI NAL units are not limited to their positions in the NAL unit stream. The phrases “prefix” and “suffix” can imply what coded picture or NAL unit a prefix / suffix SEI message may belong to, for example, details of this applicability can be specified in the semantic description of a given SEI message.
[0098] Figure 5 A diagram is also shown illustrating the syntax of a NAL unit (551) containing prefixed or suffixed SEI messages. The syntax is a container format for multiple SEI messages that can be carried within a single NAL unit (also called an SEI NAL unit). For clarity, details of the spoofing prevention syntax specified in H.266 are omitted here. For other NAL units, the SEI NAL unit may begin with a NAL unit header (521). Following the NAL unit header (521) are one or more SEI messages, such as... Figure 5 The first SEI message (530) and the second SEI message (540) are in the NAL unit (551). Each SEI message within the NAL unit (551) may include an 8-bit payload_type_byte, which specifies one of 256 different SEI types, such as those specified by [unclear text - likely a typo]. Figure 5 The payload_type_byte (532) and payload_type_byte (542) are shown in the NAL unit (551). Each SEI message within the NAL unit (551) may include an 8-bit payload_size_byte, which specifies the number of bytes of the SEI payload, for example, by Figure 5 The payload_size_byte (533) and payload_size_byte (543) are shown in the example. Each SEI message within the NAL unit (551) may include an SEI payload, which has a number of bytes specified by payload_size_byte, for example... Figure 5 The payload (534) and payload (544) are specified in the structure. In some examples, this structure may be repeated until the payload_type_byte is observed to be equal to 0xff, which indicates the end of the NAL unit. The syntax of the payload may depend on the SEI message and may have any suitable length, such as between 0 and 255 bytes.
[0099] In some examples, a first mechanism may be used, which further utilizes information related to the type of film grain, the necessity (or importance) of the film grain, and the purpose of the film grain to guide the receiver of the video stream. Information related to the type of film grain, the importance of the film grain, and the purpose of the film grain may be included in SEI messages used in the FGS process (e.g., film grain characteristic SEI message, FGS extended SEI message, etc.).
[0100] In some examples, a second mechanism for describing the grain characteristics of film can be used.
[0101] In some examples, a third mechanism that combines the first and second mechanisms described above may be used.
[0102] On the one hand, regarding the first mechanism, a new SEI can be defined, or descriptors for the type of film grain, the importance of the film grain, the purpose of the film grain, etc., can be added to the existing SEI. The purpose is to explicitly provide additional descriptors that guide the video receiver based on the type, importance, and / or purpose of the film grain.
[0103] On one hand, regarding the second mechanism, film grain SEI messages can be specified, such as film grain feature SEI messages in VSEI, and similar information can be available in the VUI of the video coding specification. Multiple film grain SEI messages can reside in a picture unit (PU), and the order in which multiple film grain SEI messages appear in the bitstream can be used by the decoder / receiver (e.g., decoder (210)) to associate a given film grain SEI message among multiple film grain SEI messages with a region defined, for example, in an annotation region (AR).
[0104] In one respect, regarding the second mechanism, the mechanism can be used to create a binding between one or more film grain SEI messages described in the second mechanism and additional descriptors described in the first mechanism, so as to associate the film grain characteristics conveyed by a given film grain SEI message with a region as described in the first mechanism.
[0105] Figure 6An example of an SEI message (601) according to one aspect of this disclosure is shown. The SEI message (601) includes information indicating the type of film grain, the importance of the film grain, and / or the purpose of the film grain. Reference Figure 6 The SEI message (601) is used in the FGS process. In one example, the SEI message (601) may be called the FGS Extended SEI message because the SEI message (601) can extend the active film grain characteristic SEI message.
[0106] In one example, for example Figure 6 As shown, the FGS extended SEI message (601) can define control information and other information related to the application of film grain in the FGS currently active for the image. The image can be associated with the SEI message (601). Other information may include film grain descriptors (e.g., film grain descriptor information) (e.g., a descriptor may describe the type, importance, and / or purpose of the film grain) to address the use case.
[0107] Alternatively, in another example, not as... Figure 6 The additional SEI message of the FGS extended SEI message (601) shown can be used to provide film grain descriptor information (e.g., the film grain descriptor information can describe the type, importance, and / or purpose of the film grains) related to the FGS SEI message that is currently active for the picture. Therefore, in this case, two SEI messages can be used, including the FGS SEI message for the picture (e.g., the film grain characteristics SEI message) and the SEI message describing the type, importance, and / or purpose of the film grains.
[0108] refer to Figure 6 The FGS extended SEI message (601) carries an additional film grain descriptor (e.g., type, importance, and / or purpose of the film grain), and control information (602) limits the persistence of the current FGS extended SEI message (601). A flag (610) can be signaled to indicate the presence of area information (e.g., rectangular area information). A sample-based area flag (611) can be signaled to indicate that an active Alpha Channel Information (ACI) SEI message will be used to determine the sample-based area, in which FGS (or the FGS procedure) is applied.
[0109] refer to Figure 6 The SEI message (601) may include a syntax element (608) indicating the number of one or more regions that are active for a region-based application of FGS. Figure 6In the example shown, one or more regions include object regions 1 through 3. For one or more regions defined by rectangles, object region 1 (603) may define the bounding box of a region for a first rectangular region of the image, object region 2 (604) may define the bounding box of a region for a second rectangular region of the image, and object region 3 (605) may define the bounding box of a third rectangular region of the image. Each region (e.g., (603), (604), and (605)) may be accompanied by a flag (606) to indicate whether FGS is applied to the corresponding region. A separate flag (609) may indicate whether FGS is applied (or not applied) to regions not defined by regions (603) through (605).
[0110] To address the use of additional descriptors (e.g., type, importance, and / or purpose of film grain), Figure 6 The SEI message (601) shown also includes a flag (612). The flag (612) may indicate the presence of additional descriptor information (e.g., the type, importance, and / or purpose of film grain) in the SEI message (601). Descriptors (613) to (615) are shown to indicate information related to the type, purpose, and importance of film grain, respectively.
[0111] Figure 7 An example of the syntax (700) for a SEI message (e.g., an FGS extended SEI message) (601) according to one aspect of this disclosure is shown. In one example, the SEI message (e.g., an FGS extended SEI message) (601) is associated with an image. In another example, a video stream includes an SEI message (601) and an image. A cancellation flag (e.g., fge_cancel_flag) (701) in the syntax (700) can be used to disable the persistence of a previously processed SEI message (e.g., a previously processed FGS extended SEI message). If the cancellation flag (701) is set to a value indicating 'true', processing of the current FGS extended SEI message can be completed. Otherwise, if the cancellation flag (701) indicates that processing should continue, for example, if the cancellation flag (701) is set to a value indicating 'false', a flag (e.g., fge_spatial_adaptation_information_present_flag) (702) can receive a value indicating whether FGS region-based or sample-based application adaptation information is enabled. If flag (702) indicates that FGS is enabled for region-based or sample-based application adaptation information, then flag (e.g., fge_region_based_adaptation_flag) (703) receives a value that indicates whether FGS is enabled for region-based applications.
[0112] If flag (703) indicates that FGS is enabled for a region-based application, then a syntax element (e.g., fge_default_grain_enabled) (704) receives a value indicating whether FGS will be applied to modify regions not covered by the regions specified in the SEI message (601). If flag (703) indicates that FGS is enabled for a region-based application, then a number (e.g., fge_active_regions_number) (705) receives a value indicating the number of active regions for the FGS-based application. For each active region indicated by a number (705) (the i-th region in the image is indicated by an index [i]), the following index variables may be received: variables (e.g., fge_ar_bounding_box_top[i]) (706), variables (e.g., fge_ar_bounding_box_left[i]) (707), variables (e.g., fge_ar_bounding_box_width[i]) (708), variables (e.g., fge_ar_bounding_box_height[i]) (709), and flags (e.g., fge_file_grain_enabled_flag[i]) (710). Index variables (706) through (709) may specify the coordinates of the top-left corner, width, and height of the bounding box of the i-th region in the image, respectively. In one example, variable (706) uses a signal to represent the coordinates of the top-right corner. In one example, variable (707) uses a signal to represent the coordinates of the bottom-left corner. In one example, variable (708) uses a signal to represent the width of the bounding box of a rectangular region (e.g., the i-th region) within the image, and variable (709) uses a signal to represent the height of the bounding box of the rectangular region (e.g., the i-th region) within the image. Variable (710) can use a signal to indicate whether FGS is applied to the region (e.g., the i-th region) within the image. Variable (710) can enable or disable the application of FGS to sample values appearing in overlapping bounding box regions.
[0113] In one example, the syntax (700) includes an Alpha channel information flag (e.g., fge_alpha_channel_adaptation_flag) (711). In one example, the Alpha channel information flag (e.g., fge_alpha_channel_adaptation_flag) (711) is located after the above processing for the region-based FGS. The Alpha channel information flag (e.g., fge_alpha_channel_adaptation_flag) (711) may receive a value indicating whether the application of the ACI SEI message to the FGS will be weighted if the ACI SEI message is currently active for the image.
[0114] According to one aspect of this disclosure, the SEI message (601) associated with an image may indicate descriptor information, such as one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain in the payload of the SEI message (601). In one example, the descriptor information (e.g., one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain) is applied to a first region in the image that includes one or more first samples. The image associated with the SEI message (601) can be reconstructed.
[0115] In one example, the SEI message (e.g., the FGS Extended SEI message) (601) indicates spatial information about a first region. For example, the spatial information of the first region includes its location and size. [Reference] Figure 7 The first region is the i-th region, and its position is indicated by variables (e.g., fge_ar_bounding_box_top[i]) (706) and variables (e.g., fge_ar_bounding_box_left[i]) (707). The size of the first region is indicated by variables (e.g., fge_ar_bounding_box_width[i]) (708) and variables (e.g., fge_ar_bounding_box_height[i]) (709). The first region can have any suitable shape. In one example, the spatial information of the first region indicates that the first region is rectangular, for example, by... Figure 7 The variables (706) to (709) are shown in the table. In one example, the image includes a second region that is different from the first region, and the first and second regions are not the same in the descriptor information.
[0116] Flags (e.g., fge_description_information_present_flag) (712) may indicate whether descriptor information is available in the payload of the SEI message (601). In one example, the SEI message (601) includes a flag (712) indicating that one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain are available in the payload of the SEI message (601).
[0117] In one example, after the above processing (via ACI SEI message) for region-based FGS and sample-based FGS, a flag (e.g., fge_description_information_present_flag) (712) may receive a value indicating whether descriptor information is available in the payload of the SEI message (601). If the flag (712) indicates that descriptor information is available, one or more syntax elements in the syntax (700) may be used to indicate the descriptor information. The one or more syntax elements may include syntax elements (e.g., fge_film_grain_type) (713), syntax elements (e.g., fge_film_grain_purpose) (714), syntax elements (e.g., fge_film_grain_estentiality) (715), etc. The syntax element (e.g., fge_film_grain_type) (713) may indicate the type of film grain; for example, the syntax element (e.g., fge_film_grain_type) (713) receives a value indicating the type of film grain. Syntax elements (e.g., fge_film_grain_purpose) (714) can indicate the purpose of film grains; for example, syntax element (714) receives a value indicating the purpose of film grains. Syntax elements (e.g., fge_film_grain_estentiality) (715) can indicate the importance of film grains; for example, syntax element (715) receives a value indicating the importance of film grains.
[0118] In one example, film grain extension is handled by a flag (e.g., fge_persistence_flag) (716). The flag (716) can receive a value that signals the range in which the current film grain extension SEI message (601) persists.
[0119] Figures 8A to 8C It shows the relationship with Figure 7 The numerical values corresponding to the descriptor semantics described in the text can be used to distinguish descriptors for film grain.
[0120] The film grain type (also referred to as the type of film grain) can be any suitable film grain type. For example, the film grain type can be one of the following: soft film grain, organic film grain, vivid film grain, 8mm film grain, 16mm film grain, 35mm film grain, and 65-70mm film grain. The various types of film grain can be indicated by different values of the syntax element (713). Figure 8A Examples of possible values for a syntax element (e.g., fge_film_grain_type) (713) according to one aspect of this disclosure are shown. Table (801) shows an example of the relationship between the values (e.g., film grain type values 0 to 255) and the corresponding types of film grain. Figure 8A In the example shown, the value of the syntax element fge_film_grain_type (713) can be described by the film grain type values 0 to 255 in Table (801). Value (0) corresponds to the soft film grain type, value (1) corresponds to the organic film grain type, value (2) corresponds to the vivid film grain type, value (3) corresponds to the 8mm film grain type, value (4) corresponds to the 16mm film grain type, value (5) corresponds to the 35mm film grain type, value (6) corresponds to the 65-70mm film grain type, and values (7 to 255) are reserved, for example, for future use by the standards development organization. Table (801) shows some examples of film grain types, and, for example, the film grain type can be appropriately modified based on the application.
[0121] The purpose of film grain (also referred to as the use of film grain) can be any suitable purpose of film grain. For example, the purpose of film grain can be one of the following: a purpose for simulating original film grain, a purpose for artistic effects in film, a purpose for artistic effects in video, a purpose for artistic effects in games, and a purpose for visual artifact occlusion. The various purposes of film grain can be indicated by different values of the syntax element (e.g., fge_film_grain_purpose) (714). Figure 8B Examples of possible values for a syntax element (e.g., fge_film_grain_purpose) (714) according to one aspect of this disclosure are shown. Table (802) shows an example of the relationship between the values (e.g., film grain purpose values 0 to 15) and the corresponding purposes of film grain. Figure 8BIn the example shown, the value of the syntax element fge_film_grain_purpose (714) can be described by the film grain purpose values 0 to 15 in Table (802). Value (0) corresponds to the original film grain simulation, value (1) corresponds to the artistic effect for film, value (2) corresponds to the artistic effect for video, value (3) corresponds to the artistic effect for game, value (4) corresponds to the purpose of visual artifact occlusion, and values (5 to 15) are reserved for future use by standards development organizations, for example. Table (802) shows some examples of film grain purposes, and, for example, the film grain purpose can be appropriately modified based on the application.
[0122] Film grain importance (also referred to as film grain significance) can be any suitable film grain necessity (or importance). Film grain significance can indicate the importance of film grains. For example, film grain significance can be indicated by one of a number of values from 0 to 3. Different film grain significance can be indicated by different values of the syntax element (e.g., fge_film_grain_essentiality) (715). Figure 8C Examples of possible values for a syntax element (e.g., fge_film_grain_essentiality) (715) according to one aspect of this disclosure are shown. Table (803) shows examples of the relationship between values (e.g., film grain importance values from 0 to 3) and the corresponding importance of film grain. Figure 8C In the example shown, the value of the syntax element fge_film_grain_essentiality (715) can be described by the film grain importance values 0 to 3 in Table (803). The value (0) has the lowest importance, the value (1) has a higher importance than the value (0) but a lower importance than the value (2), the value (2) has a higher importance than the value (1) but a lower importance than the value (3), and the value (3) has the highest importance. Table (803) shows some examples of film grain importance, and, for example, the importance of film grain can be appropriately modified based on the application.
[0123] In the above description, the SEI message (601) includes, for example, descriptor information indicated by syntax elements (712) to (715) and spatial information of a first region indicated, for example, by variables (706) to (710). In another example, the first SEI message includes descriptor information indicating the type, purpose, and / or importance of film grain, but does not indicate spatial information of the first region. Conversely, the second SEI message is also associated with the image. The second message indicates spatial information of the first region and does not include descriptor information indicating the type, purpose, and / or importance of film grain.
[0124] Figure 9 A flowchart outlining a process (900) according to one aspect of this disclosure is shown. The process (900) can be used with a video decoder. In various aspects, the process (900) is executed by processing circuitry, such as processing circuitry that performs the functions of video decoder (210), processing circuitry that performs the functions of video decoder (310), etc. In some aspects, the process (900) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (900). The process begins at (S901) and proceeds to (S910).
[0125] At (S910), an encoded video stream including encoded information of an encoded image can be received, and a Supplemental Enhancement Information (SEI) message associated with the encoded image can be received. The SEI message may indicate one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain, applied to a first region in the encoded image that includes one or more first samples.
[0126] In one example, the SEI message (e.g., a film grain synthesis (FGS) extended SEI message) indicates spatial information about a first region. The spatial information about the first region may include its location and size. In one example, the spatial information about the first region indicates that the first region is rectangular.
[0127] In one example, another SEI message (e.g., an FGS SEI message) associated with the encoded image is received. The FGS SEI message indicates spatial information about the first region, and the SEI message differs from the FGS SEI message.
[0128] In one example, the SEI message indicates the type of film grain, and the film grain type is one of the following: soft film grain, organic film grain, vivid film grain, 8mm film grain, 16mm film grain, 35mm film grain, and 65-70mm film grain.
[0129] In one example, the SEI message indicates the purpose of the film grain, and the purpose of the film grain is one of the following: for simulating original film grain, for artistic effects in film, for artistic effects in video, for artistic effects in game, and for visual artifact occlusion.
[0130] In one example, the SEI message indicates the importance of film grain, and the importance of film grain is indicated by one of a plurality of values including 0 to 3.
[0131] In one example, the SEI message includes flags indicating that one or more of the following are available in the SEI message payload: (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain.
[0132] In one example, the encoded image includes a second region that is different from the first region.
[0133] In one example, the encoded video stream includes SEI messages.
[0134] At (S920), the encoded image associated with the SEI message can be reconstructed.
[0135] Then, the process proceeds to (S999) and terminates.
[0136] The process (900) can be adjusted as appropriate. One or more steps in the process (900) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.
[0137] Figure 10 A flowchart outlining a process (1000) according to one aspect of this disclosure is shown. The process (1000) can be used with a video decoder. In various aspects, the process (1000) is executed by processing circuitry, such as processing circuitry that performs the functions of a video encoder (203), processing circuitry that performs the functions of a video encoder (403), etc. In some aspects, the process (1000) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (1200). The process begins at (S1001) and proceeds to (S1010).
[0138] At (S1010), the images in the video stream are encoded.
[0139] At (S1020), a Supplemental Enhancement Information (SEI) message associated with the image can be encoded. The SEI message indicates that one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain are applied to a first region in the image that includes one or more first samples.
[0140] In one example, the SEI message is a Film Grain Synthesis (FGS) Extended SEI message, which indicates the spatial information of the first region.
[0141] In one example, another SEI message associated with the image, such as an FGS SEI message, is encoded. The FGS SEI message indicates spatial information about the first region, and the SEI message is different from the FGS SEI message.
[0142] In one example, the SEI message indicates the type of film grain, and the film grain type is one of the following: soft film grain, organic film grain, vivid film grain, 8mm film grain, 16mm film grain, 35mm film grain, and 65-70mm film grain.
[0143] In one example, the SEI message indicates the purpose of the film grain, and the purpose of the film grain is one of the following: for simulating original film grain, for artistic effects in film, for artistic effects in video, for artistic effects in game, and for visual artifact occlusion.
[0144] In one example, the SEI message indicates the importance of film grain, and the importance of film grain is indicated by one of a plurality of values including 0 to 3.
[0145] In one example, the SEI message includes flags indicating that one or more of the following are available in the SEI message payload: (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain.
[0146] In one example, the video stream includes SEI messages.
[0147] Then, the process proceeds to (S1099) and terminates.
[0148] The process (1000) can be adjusted as appropriate. One or more steps in the process (1000) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.
[0149] A method for processing visual media data is disclosed. The method may include processing a bitstream of visual media data according to format rules. The bitstream may include encoding information of an encoded image and a Supplemental Enhancement Information (SEI) message associated with the encoded image, wherein the SEI message indicates one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain, applied to a first region in the encoded image including one or more first samples. The format rules specify: reconstruct the encoded image associated with the SEI message.
[0150] The aforementioned technologies, including those for descriptors used in film grain synthesis, can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 11 A computer system (1100) suitable for implementing certain aspects of the disclosed subject matter is shown.
[0151] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly used to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or through interpretation, microcode execution, etc.
[0152] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0153] Figure 11 The components of the computer system (1100) shown are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of computer software implementing the aspects of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any component or combination of components shown in the exemplary aspects of the computer system (1100).
[0154] The computer system (1100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, images captured from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0155] Human-machine interface input devices may include one or more of the following (only one of each is depicted): keyboard (1101), mouse (1102), touchpad (1103), touch screen (1110), data glove (not shown), joystick (1105), microphone (1106), scanner (1107), camera (1108).
[0156] The computer system (1100) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (1110), a data glove (not shown), or a joystick (1105), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (1109), headphones (not depicted)), and visual output devices (e.g., screens (1110) including CRT screens, LCD screens, plasma screens, OLED screens, each screen having or not having touch screen input functionality, each screen having or not having tactile feedback functionality, some of which are capable of outputting two-dimensional visual output or more than three-dimensional output via devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).
[0157] The computer system (1100) may also include human-machine-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1120) having media such as CD / DVD (1121), finger drives (1122), removable hard disk drives or solid-state drives (1123), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0158] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0159] The computer system (1100) may also include an interface (1154) to one or more communication networks (1155). The network may be, for example, a wireless network, a wired network, or an optical network. The network may further be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a latency-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system (1100)) attached to some general-purpose data port or peripheral bus (1149); other network interfaces are typically integrated into the core of the computer system (1100) by attaching to a system bus as described below (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). The computer system (1100) can use any of these networks to communicate with other entities. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., a CANBus connected to certain CANBus devices), or bidirectional, such as connecting to other computer systems using a local area network (LAN) or wide area network (WAN) digital network. Certain protocols and protocol stacks can be used on each of those networks and network interfaces as described above.
[0160] The human-machine interface device, human-machine accessible storage device and network interface mentioned above can be attached to the kernel (1140) of the computer system (1100).
[0161] The core (1140) may include one or more central processing units (CPU) (1141), graphics processing units (GPUs) (1142), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (1143), hardware accelerators (1144) for certain tasks, graphics adapters (1150), etc. These devices, as well as read-only memory (ROM) (1145), random access memory (1146), and internal mass storage (1147) such as internal non-user-accessible hard disk drives, SSDs, etc., may be connected via a system bus (1148). In some computer systems, the system bus (1148) may be accessed in the form of one or more physical plugs to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (1148) or attached to the core's system bus (1148) via a peripheral bus (1149). In one example, a screen (1110) may be connected to a graphics adapter (1150). Peripheral bus architectures include PCI, USB, etc.
[0162] The CPU (1141), GPU (1142), FPGA (1143), and accelerator (1144) can execute certain instructions that can be combined to form the computer code mentioned above. This computer code can be stored in ROM (1145) or RAM (1146). Transient data can also be stored in RAM (1146), while permanent data can be stored, for example, in internal mass storage (1147). Fast storage and retrieval to any storage device can be achieved by using a cache, which can be closely associated with one or more CPUs (1141), GPUs (1142), mass storage (1147), ROM (1145), RAM (1146), etc.
[0163] Computer-readable media may have computer code thereon that performs various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.
[0164] As an example, and not a limitation, a computer system (1100) having an architecture, particularly a kernel (1140), can provide functionality because one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as memory of certain non-transitory kernels (1140), such as internal kernel mass storage (1147) or ROM (1145). Software implementing aspects of this disclosure can be stored in such devices and executed by the kernel (1140). Depending on specific needs, the computer-readable media may include one or more storage devices or chips. The software can cause the kernel (1140), and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes described herein or specific portions of such processes, including defining data structures stored in RAM (1146) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system provides functionality, through hard-wired or otherwise embodied logic in the circuitry (e.g., the accelerator (1144)), that the circuitry may replace or operate with the software to perform a particular process described herein or to perform a particular portion of the particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., an integrated circuit (IC)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.
[0165] As used in this disclosure, "at least one of..." or "one of..." is intended to include any one or a combination of the listed elements. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A to C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B, and to one of A and B, are intended to include either A or B or (A and B). Where applicable, the use of "one of..." does not exclude any combination of the listed elements, for example, when the elements are not mutually exclusive.
[0166] While several exemplary aspects have been described in this disclosure, there are modifications, substitutions, and various alternatives that fall within the scope of this disclosure. Therefore, it should be appreciated that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.
[0167] The aforementioned disclosure also covers the features mentioned below. These features can be combined in various ways, and are not limited to the combinations mentioned below.
[0168] (1) A video decoding method, the method comprising: receiving an encoded video stream including encoded information of an encoded picture, and receiving a supplementary enhancement information (SEI) message associated with the encoded picture, wherein the SEI message indicates one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain applied to a first region in the encoded picture including one or more first samples; and reconstructing the encoded picture associated with the SEI message.
[0169] (2) According to the method of feature (1), wherein the SEI message is a film grain synthesis (FGS) extended SEI message, and the FGS extended SEI message indicates the spatial information of the first region.
[0170] (3) According to the method of feature (2), the spatial information of the first region includes the location and size of the first region.
[0171] (4) The method according to any one of features (2) and (3), wherein the spatial information of the first region indicates that the first region is a rectangle.
[0172] (5) The method according to feature (1), wherein the method includes: receiving a film grain synthesis (FGS) SEI message associated with an encoded picture, the FGS SEI message indicating spatial information of a first region, and the SEI message being different from the FGS SEI message.
[0173] (6) The method according to any one of features (1) to (5), wherein the SEI message indicates the type of film grain, and the type of film grain is one of the following: soft film grain type, organic film grain type, vivid film grain type, 8mm film grain type, 16mm film grain type, 35mm film grain type and 65-70mm film grain type.
[0174] (7) The method according to any one of features (1) to (6), wherein the SEI message indicates the purpose of the film grains, and the purpose of the film grains is one of the following: a purpose for simulating original film grains, a purpose for artistic effects in film, a purpose for artistic effects in video, a purpose for artistic effects in game, and a purpose for visual artifact occlusion.
[0175] (8) The method according to any one of features (1) to (7), wherein the SEI message indicates the importance of film grains and the importance of film grains is indicated by one of a plurality of values including 0 to 3.
[0176] (9) The method according to any one of features (1) to (8), wherein the SEI message includes a flag indicating that one or more of (i) the type of film grain, (ii) the purpose of the film grain and (iii) the importance of the film grain are available in the payload of the SEI message.
[0177] (10) The method according to any one of features (1) to (9), wherein the encoded image includes a second region different from the first region.
[0178] (11) The method according to any one of features (1) to (10), wherein the encoded video stream includes SEI messages.
[0179] (12) A video coding method comprising: encoding an image in a video stream; and encoding a supplementary enhancement information (SEI) message associated with the image, wherein the SEI message indicates one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain, applied to a first region in the image including one or more first samples.
[0180] (13) According to the method of feature (12), wherein the SEI message is a film grain synthesis (FGS) extended SEI message, and the FGS extended SEI message indicates the spatial information of the first region.
[0181] (14) The method according to feature (12), wherein the method includes: encoding a film grain synthesis (FGS) SEI message associated with the image, wherein the FGS SEI message indicates spatial information of a first region and the SEI message is different from the FGS SEI message.
[0182] (15) The method according to any one of features (12) to (14), wherein the SEI message indicates the type of film grain, and the type of film grain is one of soft film grain type, organic film grain type, vivid film grain type, 8mm film grain type, 16mm film grain type, 35mm film grain type and 65-70mm film grain type.
[0183] (16) The method according to any one of features (12) to (15), wherein the SEI message indicates the purpose of the film grains, and the purpose of the film grains is one of the following: a purpose for simulating original film grains, a purpose for artistic effects in film, a purpose for artistic effects in video, a purpose for artistic effects in game, and a purpose for visual artifact occlusion.
[0184] (17) The method according to any one of features (12) to (16), wherein the SEI message indicates the importance of film grains and the importance of film grains is indicated by one of a plurality of values including 0 to 3.
[0185] (18) The method according to any one of features (12) to (17), wherein the SEI message includes a flag indicating that one or more of (i) the type of film grain, (ii) the purpose of the film grain and (iii) the importance of the film grain are available in the payload of the SEI message.
[0186] (19) The method according to any one of features (12) to (18), wherein the encoded video stream includes SEI messages.
[0187] (20) A method for processing visual media data, the method comprising: processing a bitstream of visual media data according to a format rule, wherein the bitstream includes encoding information of an encoded picture, a supplementary enhancement information (SEI) message associated with the encoded picture, wherein the SEI message indicates one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain applied to a first region in the encoded picture including one or more first samples, and the format rule specifies: reconstructing the encoded picture associated with the SEI message.
[0188] (21) An apparatus for video decoding, the apparatus including processing circuitry configured to perform a method according to any one of features (1) to (11).
[0189] (22) An apparatus for video encoding, the apparatus including processing circuitry configured to perform a method according to any one of features (12) to (19).
[0190] (23) A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause at least one processor to perform a method according to any one of features (1) to (20).
Claims
1. A video decoding method, comprising: The system receives an encoded video stream including encoded information of an encoded image, and receives a Supplemental Enhancement Information (SEI) message associated with the encoded image, wherein the SEI message indicates one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain, applied to a first region in the encoded image including one or more first samples; and Reconstruct the encoded image associated with the SEI message.
2. The method according to claim 1, wherein, The SEI message is a Film Grain Synthesis (FGS) Extended SEI message, which indicates the spatial information of the first region.
3. The method according to claim 2, wherein, The spatial information of the first region includes the location and size of the first region.
4. The method according to claim 2, wherein, The spatial information of the first region indicates that the first region is rectangular.
5. The method according to claim 1, wherein, The method includes: receiving a Film Grain Synthesis (FGS) SEI message associated with the coded image. The FGS SEI message indicates the spatial information of the first region, and The SEI message is different from the FGS SEI message.
6. The method according to any one of claims 1 to 5, wherein, The SEI message indicates the type of film grain, and the type of film grain is one of the following: soft film grain, organic film grain, vivid film grain, 8mm film grain, 16mm film grain, 35mm film grain, and 65-70mm film grain.
7. The method according to any one of claims 1 to 6, wherein, The SEI message indicates the purpose of the film grains, and the purpose of the film grains is one of the following: for simulating original film grains, for artistic effects in film, for artistic effects in video, for artistic effects in game, and for visual artifact occlusion.
8. The method according to any one of claims 1 to 7, wherein, The SEI message indicates the importance of the film grains, and the importance of the film grains is indicated by one of a plurality of values including 0 to 3.
9. The method according to any one of claims 1 to 8, wherein, The SEI message includes flags indicating that one or more of the following in the payload of the SEI message are available: (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain.
10. The method according to any one of claims 1 to 9, wherein, The encoded image includes a second region that is different from the first region.
11. The method according to any one of claims 1 to 10, wherein, The encoded video stream includes the SEI message.
12. A video coding method, comprising: Encode images in the video stream; as well as Encoding a supplemental enhancement information (SEI) message associated with the image, wherein the SEI message indicates one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain applied to a first region in the image that includes one or more first samples.
13. The method according to claim 12, wherein, The SEI message is a Film Grain Synthesis (FGS) Extended SEI message, which indicates the spatial information of the first region.
14. The method of claim 12, further comprising: The film grain synthesis (FGS) SEI message associated with the image is encoded, wherein the FGS SEI message indicates spatial information of the first region, and the SEI message is different from the FGS SEI message.
15. A method for processing visual media data, the method comprising: The bitstream of the visual media data is processed according to format rules, wherein... The bitstream includes encoding information of the encoded image and Supplemental Enhancement Information (SEI) messages associated with the encoded image, wherein the SEI message indicates one or more of (i) the type of film grain, (ii) the purpose of the film grain, and (iii) the importance of the film grain, applied to a first region in the encoded image including one or more first samples. The formatting rules specify that the encoded image associated with the SEI message should be reconstructed.