Enhanced information messages by essential assistance of profile
The video decoder's processing of SEI message types is controlled through configuration file indicators, which solves the problem of inflexible SEI message processing in the prior art, and realizes a more efficient and flexible video decoding process.
Patent Information
- Application Number
- CN202480004495.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-19
- Filing Date
- 2024-09-20
- Publication Date
- 2025-06-03
AI Technical Summary
When existing video encoding and decoding technologies process auxiliary enhanced information (SEI) messages in multifunctional video encoding (VVC), there is a lack of effective mechanisms to distinguish and process different types of SEI messages, resulting in limited flexibility and efficiency of the decoding process.
The value of the configuration file indicator indicates whether the decoder needs to process a specific SEI message type, thereby realizing the decoding of the video data. Specifically, the configuration file indicator may indicate whether the decoder must process or ignore certain SEI message types.
This technical method improves the flexible processing capability of video decoder for different SEI message types, ensuring that the syntax and decoding process of video encoding and decoding is enhanced without changing the video encoding specification syntax, thereby improving the efficiency and adaptability of video encoding and decoding.
Smart Images

Figure HDA0005366302530000011 
Figure HDA0005366302530000021 
Figure HDA0005366302530000031
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the priority of U.S. Provisional Application No. 63 / 539,776, filed on September 21, 2023, and U.S. Patent Application No. 18 / 889,642, filed on September 19, 2024, the entire disclosures of which are incorporated herein by reference. Technical Field
[0003] The disclosed subject matter relates to video encoding and decoding, and more particularly, to Supplementary Enhancement Information (SEI) messages that a receiver must process based on the value of a profile indicator. Background Art
[0004] For nearly a decade, video encoding and decoding using inter - frame prediction with motion compensation has been well - known. Uncompressed digital video can consist of a series of pictures, each with a spatial dimension of, for example, luminance samples of 1920×1080 and associated chrominance samples. This series of pictures can have a fixed or variable picture rate (informally also called frame rate), such as 60 pictures per second or 60 Hz. Uncompressed video has a high bitrate requirement. For example, a 1080p60 4:2:0 video with 8 bits per sample (luminance sample resolution of 1920×1080 at 60 Hz frame rate) requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 GByte of storage space.
[0005] One purpose of video encoding and decoding is to reduce redundancy in the input video signal through compression. Compression helps reduce the above - mentioned bandwidth or storage space requirements, and in some cases, can reduce them by two orders of magnitude or more. Lossless compression and lossy compression can be employed, or a combination of both. Lossless compression refers to a technique where an exact copy of the original signal can be reconstructed from the compressed signal. When using lossy compression, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal is useful for the intended application. In the video domain, lossy compression is widely used. The degree of tolerable distortion depends on the application; for example, users of some consumer streaming applications may be more tolerant of higher distortion than users of television transmission applications. The achievable compression ratio can reflect that the higher the allowed / tolerable distortion, the higher the achievable compression ratio.
[0006] Video encoders and decoders can utilize several broad categories of techniques, including, for example, motion compensation, transformation, quantization, and entropy coding, some of which will be described below.
[0007] Some video coding and decoding specifications use profile indications, such as using integer values in control structures such as sequence parameter sets to indicate the profile being used, in order to select video coding tools compliant with those allowed in the bitstream from a superset of video coding tools permitted by the video coding specification syntax and semantics.
[0008] Some video codecs support auxiliary SEI messages. Historically, SEI messages have not been necessary for the decoding of sample values, and in most cases, the decoder can ignore SEI messages without compromising the decoding process, although the user experience may be affected. Summary of the Invention
[0009] Including a method and an apparatus, the apparatus includes a memory configured to store computer program code, and one or more processors configured to access the computer program code and operate in accordance with the instructions of the computer program code. The computer program is configured to cause the processor to perform: obtaining code that is configured to cause at least one processor to obtain video data including at least one encoded picture; identifying, by a decoder, at least one first SEI message type included with the video data from among a plurality of SEI message types that need to be processed by the decoder, the first SEI message type being selected based on the value of a profile indicator; and decoding the video data by the decoder based on the first SEI message type.
[0010] The profile indicator may represent a profile for video coding for machine (VCM).
[0011] Obtaining the video data may include: obtaining the video data from a bitstream including network abstraction layer (NAL) units, the NAL units containing a plurality of SEI messages.
[0012] The bitstream may identify at least some of the SEI messages in the NAL units as VVC (Versatile Video Coding) test model (VTM) bitstream SEIs, and at least some of the other SEI messages in the NAL units as non-VTM bitstream SEIs.
[0013] The value of the profile indicator may indicate whether VTM bitstream SEIs are required for decoding the video data.
[0014] The value of the profile indicator may indicate whether at least one first SEI message type must be processed by the decoder.
[0015] The value of the profile indicator may indicate whether at least one first SEI message type is to be ignored by the decoder. Brief Description of the Drawings
[0016] Other features, aspects, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0017] Figure 1 is a schematic diagram of a computer environment according to an embodiment;
[0018] Figure 2 is a simplified block diagram of media processing according to an embodiment;
[0019] Figure 3 is a simplified illustration of decoding according to an embodiment;
[0020] Figure 4 is a simplified illustration of encoding according to an embodiment;
[0021] Figure 5 is a simplified illustration of a NAL unit and SEI header according to an embodiment;
[0022] Figure 6 is a schematic diagram of a bitstream including mandatory SEI messages and non-mandatory SEI messages according to an embodiment; and
[0023] Figure 7 is a simplified diagram of computer characteristics according to an embodiment. DETAILED DESCRIPTION
[0024] The features mentioned in the following discussion can be used alone or in any combination. Additionally, these embodiments can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0025] In the context of the ongoing machine video coding projects in JVET and MPEG, there is a need for a mechanism to maintain the basic syntax structure of video coding and decoding specifications (such as Versatile Video Coding (H.266 / VVC)), and ideally to enhance the syntax of video coding and decoding without changing the syntax of H.266 itself (or at least not majorly), and still allow changes to the decoding process.
[0026] Figure 1 FIG. 41 shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The communication system 100 may include at least two terminals 102 and 103 interconnected by a network 105. For unidirectional data transmission, a first terminal 103 may locally encode video data for transmission to another terminal 102 via the network 105. A second terminal 102 may receive the encoded video data of another terminal from the network 105, decode the encoded data, and display the restored video data. Unidirectional data transmission is common in applications such as media services.
[0027] Figure 1 Shows second terminals 101 and 104 for supporting two-way transmission of encoded video (e.g., during a video conference). For two-way data transmission, each of terminals 101 and 104 can encode video data collected locally and transmit it over network 105 to the other terminal. Each of terminals 101 and 104 can also receive the encoded video data transmitted by the other terminal, decode the encoded data, and can display the recovered video data on a local display device.
[0028] In Figure 1 , terminals 101, 102, 103, and 104 may be shown as servers, personal computers, and smart phones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. Network 105 represents any number of networks that convey encoded video data between terminals 101, terminal 102, terminal 103, and terminal 104, including, for example, wired and / or wireless communication networks. Communication network 105 can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this application, unless otherwise stated hereinafter, the architecture and topology of network 105 may be immaterial to the operation of the present disclosure. Network 105 may include Media Aware Network Elements (MANEs), which may be included in, for example, the transmission path between terminal 101 and terminal 104. The purpose of the MANEs may be to selectively forward portions of the media data in response to network congestion, media switching, media mixing, archiving, and similar tasks typically performed by service providers rather than end users. These MANEs are capable of parsing and responding to a limited portion of the media transmitted over the network, such as syntax elements related to the network abstraction layer of a video coding technology or standard.
[0029] As an example application of the disclosed subject matter, Figure 2 Shows the placement of video encoders and decoders in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0030] A streaming system may include an acquisition subsystem 203, which may include a video source 201 such as a digital camera that creates a stream of, for example, uncompressed video samples 213. Compared to an encoded video bitstream, the sample stream 213 may be emphasized as a high-data-volume sample stream and may be processed by an encoder 202 coupled to the video source 201, which may be, for example, the camera described above. The encoder 202 may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter as described in more detail below. Compared to the sample stream, the encoded video bitstream 204 may be emphasized as a low-data-volume encoded video bitstream and may be stored in a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. The client 212 may include a video decoder 211 that decodes an incoming copy of the encoded video bitstream 208 and creates an output video sample stream 210 that can be presented on a display 209 or other rendering device (not depicted). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to certain video encoding / compression standards. Examples of these standards are as described above and further described below. Examples of these standards include ITU-T H.265 and H.266. The disclosed subject matter may be used in the context of VVC.
[0031] Figure 3 It may be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.
[0032] A receiver 302 may receive one or more encoded video sequences to be decoded by the decoder 300; in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel 301, which may be a hardware / software link to a storage device storing the encoded video data. The receiver 302 may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not depicted). The receiver 302 may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory 303 may be coupled between the receiver 302 and an entropy decoder / parser 304 (hereinafter referred to as "parser"). When the receiver 302 receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer 303 may not be needed, or the buffer 303 may be made smaller. For use on a service packet network such as the Internet, the buffer 303 may be needed, which may be relatively large and advantageously have an adaptive size.
[0033] Video decoder 300 may include a parser 304 to reconstruct symbols 313 from an entropy-coded video sequence. The categories of these symbols include information for managing the operation of video decoder 300 and potential information for controlling a rendering device such as display 312, which is not part of the decoder but may be coupled to the decoder. The control information for the rendering device may be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). Parser 304 may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be according to video coding techniques or standards and may follow various principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 304 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The entropy decoder / parser may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0034] Parser 304 may perform entropy decoding / parsing operations on the video sequence received from buffer 303 to create symbols 313. Parser 304 may receive the encoded data and selectively decode specific symbols 313. Additionally, parser 304 may determine whether a specific symbol 313 is to be provided to motion compensation prediction unit 306, scaler / inverse transform unit 305, intra prediction unit 307, or loop filter 311.
[0035] Depending on the type of the encoded video picture or a portion of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol 311 may involve multiple different units. Which units are involved and the manner of involvement may be controlled by parser 304 through subgroup control information parsed from the encoded video sequence. For clarity, such subgroup control information flows between parser 304 and the multiple units below are not depicted.
[0036] In addition to the functional blocks already mentioned, decoder 300 can conceptually be subdivided into a number of functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the number of functional units below.
[0037] The first unit is scaler / inverse transform unit 305. Scaler / inverse transform unit 305 receives, from parser 304, quantized transform coefficients as symbols 313 and control information, including the transform to be used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit can output a block including sample values, which can be input into aggregator 310.
[0038] In some cases, the output samples of scaler / inverse transform unit 305 can belong to intra-coded blocks; that is, blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by intra-picture prediction unit 307. In some cases, intra-picture prediction unit 307 generates a surrounding block of the same size and shape as the block being reconstructed using reconstructed information extracted from current (partially reconstructed) picture memory 309. In some cases, aggregator 310 adds the prediction information generated by intra-picture prediction unit 307 to the output sample information provided by scaler / inverse transform unit 305, based on each sample.
[0039] In other cases, the output samples of scaler / inverse transform unit 305 can belong to inter-coded and potentially motion-compensated blocks. In this case, motion compensation prediction unit 306 can access reference picture memory 308 to extract samples for prediction. After motion-compensating the extracted samples according to symbols 313 belonging to the block, these samples can be added by aggregator 310 to the output of scaler / inverse transform unit 305 (in this case, called residual samples or residual signal), thereby generating output sample information. The extraction of prediction samples by the motion compensation unit from an address within the reference picture memory can be controlled by a motion vector, which can be in the form of symbols 313 for use by the motion compensation unit, such as X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory, a motion vector prediction mechanism, etc., when using sub-sampled accurate motion vectors.
[0040] The output samples of aggregator 310 can be adopted by various loop filtering techniques in loop filter unit 311. Video compression techniques can include in-loop filter techniques, which are controlled by parameters included in the encoded video bitstream, and the parameters can be provided to loop filter unit 311 as symbols 313 from parser 304; however, video compression can also respond to meta-information obtained during the decoding of previous (in decoding order) parts of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0041] The output of loop filter unit 311 can be a sample stream, which can be output to rendering device 312 and stored in reference picture memory 557 for future inter-picture prediction.
[0042] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture (e.g., by parser 304) is identified as a reference picture, the current reference picture 309 can become part of reference picture buffer 308, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0043] Video decoder 300 can perform decoding operations according to predetermined video compression techniques described in standards such as ITU-T H.266. The encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and specifically as specified in the profile. For compliance, it may also be required that the complexity of the encoded video sequence be within the range defined by the tier of the video compression technique or standard. In some cases, the tier limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the tier can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0044] In one embodiment, receiver 302 can receive additional (redundant) data when receiving the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by video decoder 300 to properly decode the data and / or more accurately reconstruct the original video data. The additional data can take forms such as temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0045] Figure 4 It may be a functional block diagram of a video encoder 400 according to an embodiment of the present disclosure.
[0046] The encoder 400 may receive video samples from a video source 401 (which is not part of the encoder), and the video source may capture video images to be encoded by the encoder 400.
[0047] The video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder 303. The digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 Y CrCb, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 401 may be a camera capturing local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0048] According to one embodiment, the encoder 400 may encode and compress pictures of the source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 402. The controller controls the other functional units as described below and is functionally coupled to these units. For clarity, the couplings are not depicted in the figure. The parameters set by the controller may include rate control-related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller 402, which may involve optimizing the video encoder 400 for a certain system design.
[0049] Some video encoders operate in a "coding loop" manner that is readily recognizable to those skilled in the art. As a minimalist description, the coding loop can consist of the following: the coding portion of an encoder 400 (hereinafter referred to as the "source encoder") (responsible for creating symbols based on input pictures and reference pictures to be encoded) and a (local) decoder 406 embedded in the encoder 400, which reconstructs the symbols to create sample data that the (remote) decoder will also create (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input into the reference picture memory 405. Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content of the reference picture buffer is also bit-exactly corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is known to those skilled in the art.
[0050] The operation of the "local" decoder 406 can be the same as that of the "remote" decoder 300 described in detail above in conjunction Figure 3 with. However, briefly referring Figure 4 to, since the symbols are available and the entropy encoder 408 and the parser 304 are capable of losslessly encoding / decoding the symbols into the encoded video sequence, the entropy decoding portion of the decoder 300, which includes the channel 301, the receiver 302, the buffer 303, and the parser 304, may not be fully implementable in the local decoder 406.
[0051] What can be observed at this point is that, except for the parsing / entropy decoding that needs to be present in the decoder, the decoder technology exists in the corresponding encoder in a substantially identical functional form. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. More detailed descriptions are only needed in certain areas and are provided below.
[0052] As part of the operation, the source encoder 403 can perform motion-compensated predictive coding, referring to one or more previously encoded frames in the video sequence designated as "reference frames", and this motion-compensated predictive coding performs predictive coding on the input frame. In this way, the coding engine 407 encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, and the reference frame can be selected as the prediction reference for the input frame. The local video decoder 406 can decode the encoded video data of the frame that can be designated as a reference frame based on the symbols created by the source encoder 403. The operation of the coding engine 407 can advantageously be a lossy process. When the encoded video data can be in the video decoder ( Figure 4When decoded at a location (not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder 406 duplicates the decoding process that can be performed by the video decoder on the reference frames and enables the reconstructed reference frames to be stored in the reference picture memory 405. In this way, the encoder 400 can locally store a copy of the reconstructed reference frames, which has common content (in the absence of transmission errors) with the reconstructed reference frames to be obtained by the remote video decoder.
[0053] The predictor 404 can perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 can search the reference picture memory 405 for sample data (as a candidate reference pixel block) or some metadata, such as reference picture motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. The predictor 404 can operate block by block on the sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 404, it can be determined that the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory 405.
[0054] The controller 402 can manage the encoding operations of the source encoder 403, which can be, for example, a video encoder, and the operations include, for example, setting parameters and subgroup parameters for encoding the video data.
[0055] The outputs of all the above functional units can be entropy encoded in the entropy encoder 408. The entropy encoder performs lossless compression on the symbols generated by the various functional units according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence. The transmitter 409 can buffer the encoded video sequence created by the entropy encoder 408 to prepare for transmission through the communication channel 411, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter 409 can merge the encoded video data from the source encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0056] The controller 402 can manage the operations of the encoder 400. During encoding, the controller 402 can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures can generally be assigned to one of the following frame types:
[0057] An Intra picture (I picture) is a picture that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of Intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and characteristics.
[0058] A Predictive picture (P picture) is a picture that can be encoded and decoded using Intra prediction or Inter prediction, where the Intra prediction or Inter prediction uses at most one motion vector and reference index to predict the sample values of each block.
[0059] A Bi-directional Predictive picture (B picture) is a picture that can be encoded and decoded using Intra prediction or Inter prediction, where the Intra prediction or Inter prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0060] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the coding assignments applied to the corresponding pictures of the blocks. For example, blocks of an I picture can be non-predictively encoded, or blocks of an I picture can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or Intra prediction). Pixel blocks of a P picture can be predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.
[0061] The encoder 400 (which can be, for example, a video encoder) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.266 recommendation. In operation, the encoder 400 can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0062] In an embodiment, the transmitter 409 can transmit additional data when transmitting the encoded video. The source encoder 403 can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.
[0063] In a video bitstream, the compressed video can be enhanced by auxiliary enhancement information, e.g., in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI). Video coding standards can include specification sections for SEI and VUI. SEI and VUI information can also be specified in a separate specification that can be referenced by the video coding specification.
[0064] Reference Figure 5 Example 500 in [reference] shows an example layout of a Coded Video Sequence (CVS) according to H.266. The coded video sequence is subdivided into Network Abstraction Layer (NAL) units. An exemplary NAL unit 501 can include a NAL unit header (NAL_unit_header) 502, which in turn includes the following 16 bits: The forbidden_zero_bit 503 and nuh_reserved_zero_bit 504 may not be used in H.266 and can be zero in a NAL unit compliant with H.266. The 3 bits of nuh_layer_id 505 can indicate the layer (spatial, SNR, or multi-view enhancement) to which the NAL unit belongs. The 5 bits of nuh_nal_unit_type define the type of the NAL unit. In H.266, 22 NAL unit type values are defined for the NAL unit types defined in H.266, 6 NAL unit types are reserved, 4 NAL unit type values are unspecified and can be used by specifications other than H.266. Finally, the 3 bits of the NAL unit header indicate the temporal layer nuh_temporal_id_plus1 506 to which the NAL unit belongs.
[0065] A coded picture can include one or more Video Coding Layer (VCL) NAL units and zero or more non-VCL NAL units. VCL NAL units can include coded data that conceptually belongs to the video coding layer as described above. Non-VCL NAL units can include data that conceptually does not belong to the video coding layer. Taking H.266 as an example, they can be classified into the following categories: (1) parameter sets, (2) picture headers (PH_NUT), (3) NAL units, (4) prefix SEI NAL unit type and suffix SEI NAL unit type (PREFIX_SEI_NUT and SUFFIX_SEI_NUT), (5) filler data NAL unit type FD_NUT, and (6) reserved and unspecified NAL unit types, as shown below:
[0066] (1) Parameter sets, including information that may be necessary for the decoding process, which can be applied to multiple encoded pictures. Parameter sets and conceptually similar NAL units can be the following NAL unit types, for example: DCI_NUT (Decoding Capability Information, DCI), VPS_NUT (Video Parameter Set, VPS, which is used, among other things, to establish layer relationships, etc.), SPS_NUT (Sequence Parameter Set, SPS, which is used, among other things, to establish parameters used and kept unchanged in the entire encoded video sequence CVS, etc.), PPS_NUT (Picture Parameter Set, PPS, which is used, among other things, to establish parameters used and kept unchanged in the encoded picture, etc.), and PREFIX_APS_NUT and SUFFIX_APS_NUT (prefix adaptive parameter set and suffix adaptive parameter set). Parameter sets can include information necessary for the decoder to decode VCL NAL units and are therefore referred to here as "standard" NAL units.
[0067] (2) Picture header (PH_NUT), which is also a "standard" NAL unit.
[0068] (3) NAL units that mark certain positions in the NAL unit stream. These NAL units include NAL units with NAL unit types such as AUD_NUT (Access Unit Delimiter), EOS_NUT (End of Sequence), and EOB_NUT (End of Bitstream). These are non-standard and are also called informative because compliant decoders do not require these NAL units in their decoding process, although the decoder needs to be able to receive these NAL units in the NAL unit stream.
[0069] (4) Prefix SEI NAL unit type and suffix SEI NAL unit type (PREFIX_SEI_NUT and SUFFIX_SEI_NUT), which indicate NAL units including prefix and suffix auxiliary enhancement information. In H.266, these NAL units are informative because they are not necessary for the decoding process.
[0070] (5) The filler data NAL unit type FD_NUT indicates filler data; this data can be random and can be used to "waste" bits in the NAL unit stream or bitstream, which may be necessary for transmission in certain isochronous transmission environments.
[0071] (6) Reserved and undefined NAL unit types.
[0072] Still refer toFigure 5 , which shows the layout of the NAL unit stream arranged in decoding order 510, including the encoded picture 511, which includes some of the types of NAL units introduced above. At the front position of the NAL unit stream, the DCI 512, VPS 513, and SPS 514 can jointly establish the parameters that the decoder can use to decode the encoded pictures (including the encoded picture 511 of the NAL unit stream) of the encoded video sequence (CVS).
[0073] The encoded picture 511 may (in the depicted order or any other order conforming to the video coding technology or standard used here, which is H.266) include: a prefix APS 516, a picture header (PH) 517, a prefix SEI 518, one or more VCL NAL units 519, and a suffix SEI 520.
[0074] The prefix SEI NAL unit 518 and the suffix SEI NAL unit 520 were introduced during the standard development process because for some SEI messages, the content of the message is known before encoding a given picture begins, while other content is only known after the picture encoding is complete. With the prefix and suffix SEIs, certain SEI messages are allowed to appear early or late in the NAL unit stream of the encoded picture, thus avoiding buffering. As an example, in the encoder, the sampling time of the picture to be encoded is known before the picture is encoded, so the picture timing SEI message can be the prefix SEI message 516. On the other hand, the decoded picture hash SEI message (including the hash value of the sample values of the decoded picture and which can be used, for example, to debug the encoder implementation) is the suffix SEI message 518 because the encoder cannot calculate the hash value of the reconstructed samples until after the picture encoding is complete. The positions of the prefix SEI NAL unit and the suffix SEI NAL unit are not limited to their positions in the NAL unit stream. "Prefix" and "suffix" can imply which encoded pictures or NAL units the prefix / suffix SEI messages are related to, and the details of such applicability can be specified, for example, in the semantic description of a given SEI message.
[0075] Still referring to Figure 5, shows a simplified syntax diagram of a NAL unit including a prefix or suffix SEI message 520. This syntax is a container format for carrying multiple SEI messages in a single NAL unit. For clarity, details of the emulation prevention syntax specified in H.266 are omitted here. Like other NAL units, the SEI NAL unit starts with a NAL unit header (NAL_unit_header) 521. After the header is one or more SEI messages; two SEI messages 530 and 531 are depicted and described below. Each SEI message in the SEI NAL unit includes: an 8-bit payload type byte (payload_type_byte 522) for specifying one of 256 different SEI types; an 8-bit payload size byte (payload_size_byte) 523 for specifying the number of bytes of the SEI payload; and a Payload (payload) 524 of payload_size_byte bytes. This structure can be repeated until a payload_type_byte equal to 0xff is observed, which indicates the end of the NAL unit. The syntax of Payload 524 depends on the SEI message and its length can range from 0 to 255 bytes. In the MPEG machine video coding project, one approach is to use an intra codec (referred to as the "VTM codec") different from the H.266 intra picture coding mechanism, which allows the use of, for example, AI-based techniques. Since the samples reconstructed by the VTM codec can be used for predicting other pictures encoded in the bitstream, the reconstruction process must be handled by a decoder configured for the VTM codec. Traditionally, doing so would require standardizing the VTM codec as a component of H.266 in the JVET process and creating a new profile in which this intra codec replaces the existing H.266 IRAP coding mechanism. Due to some procedural and principle reasons, the relevant projects of MPEG and JVET decided not to adopt this approach. Therefore, a mechanism is needed that only makes minimal changes to the VVC syntax but allows the syntax of the VTM codec to be included in the VVC bitstream.
[0076] In one embodiment, the bitstream of the VTM codec can be included in one or more SEI messages configured for this purpose. The details of this inclusion can be based on a format that splits the NAL units created / consumed by the VTM codec into segments that can be handled by the SEI message syntax of H.266 (whose payload has a limited size, e.g., 255 bytes), along with appropriate splitting and aggregation rules to recreate the sub-bitstream used by / from the VTM codec to / from the SEI messages.
[0077] The current problem is that, according to the current H.266 specification, SEI messages are not used for decoding sample information. This is contrary to the concept that samples reconstructed based on SEI messages carrying a VTM codec sub-stream are used to reconstruct H.266 inter-frame / B pictures.
[0078] Reference Figure 6 In Example 600 of [reference], in the same or another embodiment, the H.266 bitstream may include a parameter set 602, such as a sequence parameter set, which includes a profile indicator 603 pointing to a VCM profile. The bitstream may include a plurality of NAL units, and the NAL units include SEI messages with identifiers that label these SEI messages as VTM bitstream SEIs 604. The bitstream also includes (one or more) NAL units containing other types of SEI messages 605 that are not VTM bitstream SEIs, and H.266 applies the normal rules to these SEI messages - including that these SEI messages are not necessary for the decoding process of sample values. Other NAL units 606 may include the H.266 bitstream. The value of the profile indicator 603 may indicate that the VTM bitstream SEI is necessary for the decoding process of the VTM codec and is also necessary for the decoding of the entire bitstream including VTM-encoded intra pictures and H.266-encoded inter pictures. If the profile indicator 603 has a predefined value indicating a VTM codec SEI message, the relevant SEI messages must be processed by the decoder, particularly by the VTM decoder. If the profile indicator 603 has a different value, the normal operating rules of the SEI messages may be applied, which means that the decoder may ignore the VTP bitstream SEI messages even if they are included in the bitstream.
[0079] Using the VTM encoder bitstream as the payload for required SEI messages is only an example. Similar mechanisms can also be used for technologies unrelated to VCM. For example, similar to the VTM-encoded pictures in VTM-encoded SEI messages, enhancement layers conforming to codecs different from H.266 can be carried in required SEI messages.
[0080] The techniques described above (such as for required auxiliary enhancement information messages via profiles) can be implemented in the following ways: using computer software with computer-readable instructions and physically stored in one or more computer-readable media; or by one or more specially configured hardware processors. For example, Figure 7 FIG. shows a computer system 700 suitable for implementing certain embodiments of the disclosed subject matter.
[0081] The computer software can be encoded using any suitable machine code or computer language, which can be subject to assembly, compilation, linking, or similar mechanisms to create code including instructions that can be directly executed by a computer's central processing unit (CPU), graphics processing unit (GPU), etc., or executed through interpretation, microcode execution, etc.
[0082] These instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0083] Figure 7 The components of the computer system 700 shown are exemplary in nature and are not intended to place any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Additionally, the configuration of the components should not be construed as having any dependency or requirement related to any one or combination of the components shown in the exemplary aspects of the computer system 700.
[0084] The computer system 700 may include certain human-machine interface input devices. Such human-machine interface input devices can respond to one or more human users through inputs such as, for example: tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., speech, clapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices can also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, captured images from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0085] The human-machine interface input devices may include one or more of the following (only one of each is shown): keyboard 701, mouse 702, touchpad 703, touch screen 710, joystick 705, microphone 706, scanner 708, camera 707. The computer system 700 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include haptic output devices (e.g., haptic feedback of the touch screen 710, or the joystick 705, but may also be haptic feedback devices that are not input devices), audio output devices (e.g., speakers 709, headphones (not depicted)), visual output devices (e.g., screen 710 including CRT screens, LCD screens, plasma screens, OLED screens, each screen having or not having touch screen input function, each screen having or not having haptic feedback function, some of which are capable of outputting two-dimensional visual output or output beyond three dimensions through means such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)), and printers (not depicted).
[0086] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 720 with media 711 such as CD / DVD, thumb drives 722, removable hard disk drives or solid state drives 723, traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc. Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0087] The computer system 700 may also include an interface 799 leading to one or more communication networks 798. The network 798 may be, for example, a wireless network, a wired network, or an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of the network 798 include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Certain networks 798 typically require an external network interface adapter (e.g., the USB port of the computer system 700) attached to certain general-purpose data ports or peripheral buses (750 and 751); other network interfaces are typically integrated into the core of the computer system 700 by attaching to the system bus as described below (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). The computer system 700 may use any of these networks 798 to communicate with other entities. Such communication may be only one-way reception (e.g., broadcast television), only one-way transmission (e.g., CANBus connected to certain CANBus devices), or two-way, for example, using a local area network or a wide area digital network to connect to other computer systems. Certain protocols and protocol stacks may be used on each of the networks and network interfaces described above. The above-described human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core 740 of the computer system 700.
[0088] The core 740 may include one or more central processing units (CPUs) 741, a graphics processing unit (GPU) 742, a graphics adapter 717, a dedicated programmable processing unit in the form of a Field Programmable Gate Area (FPGA) 743, a hardware accelerator 744 for certain tasks, etc. These devices, as well as a read-only memory (ROM) 745, a random access memory 746, and an internal mass storage 747 such as an internal non-user-accessible hard disk drive, SSD, etc., may be connected via a system bus 748. In some computer systems, the system bus 748 may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the system bus 748 of the core or attached to the system bus 748 of the core via a peripheral bus 749. The architecture of the peripheral bus includes PCI, USB, etc.
[0089] The CPU 741, GPU 742, FPGA 743, and accelerator 744 can execute certain instructions that, when combined, can constitute the aforementioned computer code. The computer code can be stored in the ROM 745 or the RAM 746. Transitional data can also be stored in the RAM 746, while permanent data can be stored in the internal mass storage 747. Fast storage and retrieval to any storage device can be achieved by using a cache, which can be closely associated with one or more of the following: one or more CPUs 741, GPUs 742, mass storage 747, ROM 745, RAM 746, etc.
[0090] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code specially designed and constructed for the purposes of this disclosure, or the medium and the computer code can be of the type well-known and available to those skilled in the field of computer software.
[0091] By way of example, and not limitation, because software embodied in one or more tangible computer-readable media is executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.), a computer system 700 having an architecture, particularly having a core 740, can provide functionality. Such computer-readable media can be media associated with the user-accessible mass storage introduced above, as well as certain memories of the non-transitory core 740, such as the core internal mass storage 747 or the ROM 745. The software implementing the embodiments of this disclosure can be stored in such devices and executed by the core 740. Depending on specific requirements, the computer-readable medium can include one or more storage devices or chips. The software can cause the core 740), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes described herein or specific portions of the specific processes described herein, including defining data structures stored in the RAM 746 and modifying such data structures according to processes defined by the software. Additionally or alternatively, because of logic hardwired or otherwise embodied in a circuit (e.g., accelerator 744), a computer system can provide functionality, and this circuit can replace the software or operate together with the software to execute specific processes described herein or specific portions of the specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to the computer-readable medium can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or including both. This disclosure encompasses any suitable combination of hardware and software.
[0092] Although the present disclosure describes some exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be realized that those skilled in the art will be able to devise many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.
Claims
1. A method comprising: Acquire video data, wherein the video data includes at least one encoded picture; identifying, by a decoder, at least one first supplemental enhancement information (SEI) message type included with the video data from a plurality of SEI message types that need to be processed by the decoder, wherein the first SEI message type is selected based on a value of a profile indicator; and The video data is decoded by the decoder based on the first SEI message type.
2. The method according to claim 1, in, The profile indicator indicates a video coding (VCM) profile for the machine.
3. The method according to claim 1, in, Acquiring the video data includes: acquiring the video data from a bitstream including a NAL unit, where the NAL unit includes multiple SEI messages.
4. The method according to claim 3, in, The code stream identifies at least some of the SEI messages in the SEI messages of the NAL unit as VVC test model (VTM) code stream SEI, and identifies at least some of the other SEI messages in the SEI messages of the NAL unit as non-VTM code stream SEI.
5. The method according to claim 4, wherein: The value of the profile indicator indicates whether the VTM codestream SEI is required for decoding the video data.
6. The method according to claim 1, in, The value of the profile indicator indicates whether the at least one first SEI message type has to be processed by the decoder.
7. The method according to claim 1, in, The value of the profile indicator indicates whether the at least one first SEI message type is to be ignored by the decoder.
8. A video decoding device, the device comprising: at least one memory configured to store computer program code; At least one processor is configured to access the computer program code and operate according to the instructions of the computer program code, wherein the computer program code comprises: Acquire video data, wherein the video data includes at least one encoded picture; identifying, by a decoder, at least one first supplemental enhancement information (SEI) message type included with the video data from a plurality of SEI message types that need to be processed by the decoder, wherein the first SEI message type is selected based on a value of a profile indicator; and The video data is decoded by the decoder based on the first SEI message type.
9. The device according to claim 8, in, The profile indicator indicates a video coding (VCM) profile for the machine.
10. The device according to claim 8, in, Acquiring the video data: comprising acquiring the video data from a bitstream comprising a NAL unit, wherein the NAL unit comprises a plurality of SEI messages.
11. The device according to claim 10, in, The codestream identifies at least some of the SEI messages in the NAL unit as VVC test model (VTM) codestream SEI, and identifies at least some of the other SEI messages in the SEI messages in the NAL unit as non-VTM codestream SEI.
12. The device according to claim 11, in, The value of the profile indicator indicates whether the VTM codestream SEI is required for decoding the video data.
13. The device according to claim 8, in, The value of the profile indicator indicates whether the at least one first SEI message type has to be processed by the decoder.
14. The device according to claim 8, in, The value of the profile indicator indicates whether the at least one first SEI message type is to be ignored by the decoder.
15. A non-transitory computer-readable medium storing a program, the program causing a computer to perform the following operations: Acquire video data, wherein the video data includes at least one encoded picture; identifying, by a decoder, at least one first supplemental enhancement information (SEI) message type included with the video data from a plurality of SEI message types that need to be processed by the decoder, wherein The first SEI message type is selected based on a value of a profile indicator; and The video data is decoded by the decoder based on the first SEI message type.
16. The method according to claim 15, in, The profile indicator indicates a video coding (VCM) profile for the machine.
17. The method according to claim 15, in, Acquiring the video data includes: acquiring the video data from a bitstream including a NAL unit, where the NAL unit includes multiple SEI messages.
18. The method according to claim 17, in, The code stream identifies at least some of the SEI messages in the SEI messages of the NAL unit as VVC test model (VTM) code stream SEI, and identifies at least some of the other SEI messages in the SEI messages of the NAL unit as non-VTM code stream SEI.
19. The method according to claim 18, wherein: The value of the profile indicator indicates whether the VTM codestream SEI is required for decoding the video data.
20. The method according to claim 15, in, The value of the profile indicator indicates whether the at least one first SEI message type has to be processed by the decoder.