Large SEI messages
Through segmentation technology, the SEI message payload greater than 255 bytes is divided into small segments, and the segmentation relationship is identified by controlling information, the problem of limited SEI message payload in H.266/VVC is solved, and effective processing and decoding of large SEI payloads are realized.
Patent Information
- Application Number
- CN202480004386.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-26
- Filing Date
- 2024-09-27
- Publication Date
- 2025-05-30
AI Technical Summary
In existing video encoding standards such as H.266/VVC, the SEI message payload is limited to 255 bytes and cannot effectively process data exceeding 255 bytes, such as containers of data structures such as EXIF.
Through segmentation technology, SEI message payloads greater than 255 bytes are divided into segments of 255 bytes or less, and the relationship between segments is identified through control information to achieve effective processing of large SEI payloads.
Effective processing of SEI message payloads over 255 bytes is realized, data loss and decoding errors are avoided, and video encoding flexibility and compatibility are improved.
Smart Images

Figure CN120077663A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 541,223, filed on September 28, 2023, and U.S. Application No. 18 / 897,918, filed on September 26, 2024, the entire disclosures of which are incorporated herein by reference. Technical field
[0003] The disclosed subject matter relates to video encoding and decoding, and more particularly, to the syntax and semantic mechanisms for adding Supplementary Information Enhancement (SEI) payloads of more than 255 bytes to a VVC encoded video bitstream. Background art
[0004] For decades, video encoding and decoding using inter - picture prediction with motion compensation has been known. Uncompressed digital video may include a series of pictures, each picture having a spatial dimension of, for example, 1920×1080 luminance samples and associated chrominance samples. The series of pictures has a fixed or variable picture rate (also informally called the frame rate), such as 60 pictures per second or 60 Hz. Uncompressed video has very high bitrate requirements. For example, a 1080p60 4:2:0 video with 8 bits per sample (1920x1080 luminance sample resolution, 60 Hz frame rate) requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video would require more than 600 GB of storage space.
[0005] One purpose of video encoding and decoding is to reduce the redundant information of the input video signal through compression. Video compression can help reduce the requirements for the above - mentioned bandwidth or storage space, and in some cases, can reduce it by two or more orders of magnitude. Both lossless and lossy compression, as well as combinations of the two, can be employed. Lossless compression refers to a technique for reconstructing an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal can be used for the intended application. Lossy compression is widely used in video. The amount of allowable distortion depends on the application. For example, users of some consumer streaming applications can tolerate higher distortion compared to users of television applications. The achievable compression ratio reflects that higher allowed / tolerated distortion can result in a higher compression ratio.
[0006] Video encoders and decoders can utilize several major categories of techniques, such as including: motion compensation, transformation, quantization, and entropy coding, some of which will be introduced below. Summary of the invention
[0007] Including a method and an apparatus, the apparatus includes a memory and at least one processor. Among them, the memory is used to store computer program code, and the at least one processor is used to access the computer program code and operate according to the instructions of the computer program code. The computer program is used to cause the processor to execute code acquisition, code recognition, and code decoding. The code acquisition is used to cause the at least one processor to acquire video data including at least one encoded picture. The code recognition is used to cause the at least one processor to identify at least one first supplementary information enhancement (SEI) message through a decoder. The first SEI message indicates variables, and the variables specify the type payloadType and size payloadSize of the payload of the at least one SEI message, and are specified in bytes. And the code decoding is used to cause the at least one processor to decode the video data through the decoder based on the first SEI message.
[0008] Obtaining the video data may include: receiving a network abstraction layer (NAL) unit, where the type of the NAL unit is different from the type of the NAL unit in the at least one first SEI message.
[0009] The received NAL unit may at least include a length header field, where the length header field specifies a value greater than 256 bytes; where the maximum size of a payload of an SEI message may be less than or equal to 256 bytes, and decoding the video data may include: extracting a payload greater than 256 bytes from the NAL unit.
[0010] The at least one first SEI message includes a start segment, and the start segment includes the payload of the start segment.
[0011] Decoding the video data may include: receiving a second SEI message including an end segment, and the end segment includes the payload of the end segment.
[0012] Decoding the video data may include: extracting the start segment and the end segment.
[0013] Decoding the video data may include: assembling a large SEI payload based on the extracted start segment and the extracted end segment, where the start segment is at the beginning of the assembled large SEI payload, and the end segment is at the end of the assembled large SEI payload, and the large SEI payload is at least greater than 256 bytes. Description of the Drawings
[0014] Other features, aspects, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0015] Figure 1 is a schematic diagram of a computer environment according to one embodiment;
[0016] Figure 2 is a simplified block diagram of media processing according to one embodiment;
[0017] Figure 3 is a simplified illustration of decoding according to one embodiment;
[0018] Figure 4 is a simplified illustration of encoding according to another embodiment;
[0019] Figure 5 is a simplified illustration of a Network Abstraction Layer (NAL) unit and SEI header according to one embodiment;
[0020] Figure 6 is a schematic diagram of a NAL unit type table according to another embodiment;
[0021] Figure 7 is a schematic diagram of a large SEI message syntax using different NAL unit types according to one embodiment;
[0022] Figure 8 is a schematic diagram of fragmentation syntax according to one embodiment; and
[0023] Figure 9 is a simplified diagram of computer characteristics according to one embodiment. DETAILED DESCRIPTION
[0024] Techniques are disclosed for extending the payload of SEI messages beyond the 255 - byte limit specified by the H.266 syntax through fragmentation. Various techniques are disclosed for indicating the relationships between fragments and other SEI messages that may be part of NAL units.
[0025] The H.266 / VVC syntax allows a maximum of 255 bytes in the SEI payload because the SEI payload_size_byte is an 8 - bit fixed - length codeword and the value 0xff is reserved for certain purposes. However, the SEI message mechanism may be useful for data larger than 255 bytes, for example, as a container for data structures such as EXIF for thumbnails or the like. Thus, techniques are needed to allow SEI message payloads to exceed 255 bytes.
[0026] The proposed features discussed below can be used individually or in any combination. Additionally, these embodiments can be implemented by a processing circuit (e.g., at least one processor or at least one integrated circuit). In one example, at least one processor executes a program stored in a non-volatile computer-readable medium.
[0027] In the context of the ongoing machine video coding projects in JVET and MPEG, there is a need for a mechanism to maintain the basic syntax structure of video codec specifications, such as Versatile Video Coding (H.266 / VVC). Ideally, the syntax of the video codec can be enhanced without involving changes to the syntax of H.266 itself (or at least not in a major way), and still allow for variations in the decoding process.
[0028] Figure 1 A simplified block diagram of a communication system 100 according to an embodiment of the present application is shown. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional transmission of data, the first terminal 103 may encode video data at a local location for transmission via the network 105 to another terminal 102. The second terminal 102 may receive the encoded video data of another terminal from the network 105, decode the encoded data, and display the recovered video data. Unidirectional data transmission may be common in media service applications and the like.
[0029] Figure 1 A second pair of terminals 101 and 104 is shown, which is provided to support two-way transmission of encoded video that may occur, for example, during a video conference. For two-way transmission of data, each of the terminals 101 and 104 may encode video data captured at a local location for transmission via the network 105 to the other terminal. Each of the terminals 101 and 104 may also receive the encoded video data sent by the other terminal, decode the encoded data, and display the recovered video data on a local display device.
[0030] In Figure 1Among them, the terminals 101, 102, 103, and 104 can be shown as servers, personal computers, and smart phones, but the principles of this application are not limited thereto. Embodiments of this application can be applied to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network 105 represents any number of networks that transmit encoded video data between the terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. The communication network 105 can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise explained below, the architecture and topology of the network 105 may be unimportant for the operation of this application. The network 105 can include a Media Aware Network Element (MANE), which can be included, for example, in the transmission path between the terminals 101 and 104. The purpose of the MANE can be to selectively forward portions of the media data in response to network congestion, media switching, media mixing, archiving, and similar tasks typically performed by service providers rather than end users. Such a MANE may be able to parse and react to a limited portion of the media transmitted over the network, such as syntax elements related to video coding techniques or the network abstraction layer of a standard.
[0031] As an example, Figure 2 shows the placement of video encoders and video decoders in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, and so on.
[0032] A streaming system may include a capture subsystem 203, which may include a video source 201, such as a digital camera, to create an uncompressed video sample stream 213, for example. When compared with an encoded video bitstream, the sample stream 213 may be emphasized as having a high data volume and may be processed by an encoder 202 coupled to the video source 201, which may be a camera as discussed above. The encoder 202 may include hardware, software, or a combination thereof to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video bitstream 204 may be emphasized as having a lower data volume compared to the sampled stream and may be stored on a streaming server 205 for future use. At least one streaming client 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. The client 212 may include a video decoder 211 that decodes an incoming copy of the encoded video bitstream 208 and creates an outgoing video sample stream 210 that may be rendered on a display 209 or other rendering device (not shown). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to certain video coding / compression standards. Examples of such standards are described above and further herein. Examples of such standards include ITU-T Recommendations H.265 and H.266. The disclosed subject matter may be used in the context of VVC.
[0033] Figure 3 It may be a functional block diagram of a decoder 300 according to an embodiment of the present invention.
[0034] A receiver 302 may receive at least one codec video sequence to be decoded by the decoder 300; in the same or another embodiment, one encoded video sequence at a time, where the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence may be received from a channel 301, which may be a hardware / software link to a storage device storing the encoded video data. The receiver 302 may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver 302 may separate the encoded video sequence from the other data. To counter network jitter, a buffer 303 may be coupled between the receiver 302 and an entropy decoder / parser 304 (hereinafter referred to as "parser"). When the receiver 302 receives data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer 303 may not be needed, or the buffer 303 may be small. For use on a best-effort packet network such as the Internet, the buffer 303 may be needed, which may be relatively large and may advantageously have an adaptive size.
[0035] The decoder 300 may include a parser 304 for reconstructing symbols 313 from an entropy-coded video sequence. The categories of these symbols include information for managing the operation of the decoder 300 and information that may be used to control a rendering device such as a display 312. The display 312 is not a part of the decoder but may be coupled thereto. The control information for the rendering device may be in the form of supplementary enhancement information (SEI messages) or a video usability information (VUI) parameter set segment (not shown). The parser 304 may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be according to a video coding technology or standard and may follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract a set of subgroup parameters for at least one subgroup in a subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. Subgroups may include group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transformation unit (TU), prediction unit (PU), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0036] The parser 304 may perform entropy decoding / parsing operations on the video sequence received from the buffer 303 to create symbols 313. The parser 304 may receive the encoded data and selectively decode specific symbols 313. In addition, the parser 304 may determine whether to provide a specific symbol 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.
[0037] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols 313 may involve at least two different units. Which units are involved and the way they are involved may be controlled by subgroup control information parsed by the parser 304 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser 304 and at least two units below are not described.
[0038] In addition to the functional blocks already mentioned, the decoder 300 can conceptually be subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.
[0039] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives quantized transform coefficients and control information from the parser 304, including the transform to be used, block size, quantization factor, quantization scaling matrix, etc. as at least one symbol 313. It can output blocks including sample values, which can be input into the aggregator 310.
[0040] In some cases, the output samples of the scaler / inverse transform unit 305 can belong to intra-coded blocks, that is: blocks that do not use prediction information from a previous reconstructed picture, but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by the intra picture prediction unit 307. In some cases, the intra picture prediction unit 307 uses the already reconstructed surrounding information obtained from the current (partially reconstructed) picture 309 to generate a block having the same size and shape as the block being reconstructed. In some cases, the aggregator 310 adds the prediction information generated by the intra picture prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305 based on each sample.
[0041] In other cases, the output samples of the scaler / inverse transform unit 305 can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit 306 can access the reference picture buffer 308 to extract samples for prediction. After motion-compensating the extracted samples according to the symbol 313, these samples can be added by the aggregator 310 to the output of the scaler / inverse transform unit (referred to as residual samples or residual signals in this case), thereby generating output sample information. The motion compensation prediction unit obtaining prediction samples from an address within the reference picture memory can be controlled by a motion vector, and the motion vector is in the form of the symbol 313 for use by the motion compensation prediction unit, and the symbol 313 includes, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and so on.
[0042] The output samples of aggregator 310 can be adopted by various loop filtering techniques in loop filter unit 311. Video compression techniques can include in-loop filter techniques, which are controlled by parameters included in the encoded video bitstream, and the parameters can be used for loop filter unit 311 as symbols 313 from parser 304. However, in other embodiments, video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) portions of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0043] The output of loop filter unit 311 can be a sample stream, which can be output to display device 312 and stored in reference picture memory 557 for subsequent inter-picture prediction.
[0044] Once some encoded pictures are fully reconstructed, they can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by parser 304), the current picture 309 can become part of reference picture buffer 308, and a new current picture memory can be reallocated before starting to reconstruct the next encoded picture.
[0045] Decoder 300 can perform decoding operations according to predetermined video compression techniques that can be recorded in standards such as ITU-T Rec.H.266. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, as it conforms to the syntax of the video compression technique or standard specified in the video compression technique document or standard, particularly in the profile document thereof. Compliance also requires that the complexity of the encoded video sequence be within the range defined at the video compression technique or standard level. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (e.g., megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further restricted by assuming reference decoder (HRD) specifications and metadata for HRD buffer management signaled in the encoded video sequence.
[0046] In an embodiment, receiver 302 can receive additional (redundant) data along with the encoded video. The additional data can be part of the encoded video sequence. The additional data can be used by decoder 300 to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0047] Figure 4It may be a functional block diagram of a video encoder 400 according to an embodiment of the present application.
[0048] The encoder 400 may receive video samples from a video source 401 (not part of the encoder). The video source 401 may capture at least one video image to be encoded by the encoder 400.
[0049] The video source 401 may provide a source video sequence in the form of a stream of digital video samples to be encoded by the encoder (303). The stream of digital video samples may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 Y CrCb, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 401 may be a camera that captures local image information as a video sequence. The video data may be provided as at least two separate pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as spatial arrays of pixels, where each pixel may include at least one sample depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0050] According to one embodiment, the encoder 400 may encode and compress pictures of the source video sequence in real time or under any other time constraints required by the application into an encoded video sequence 410. Enforcing an appropriate encoding speed is a function of the controller 402. The controller controls other functional units as described below and is functionally coupled to these units. For clarity, the couplings are not depicted. The parameters set by the controller may include parameters related to rate control (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller 402 as they may be related to the video encoder 400 optimized for a specific system design.
[0051] Some video encoders operate in a manner that is readily recognizable to those skilled in the art as a "coding loop". As a gross oversimplification, the coding loop can consist of an encoding portion of encoder 400 (hereinafter referred to as the "source encoder") that is responsible for creating symbols based on an input picture to be encoded and at least one reference picture, and an (local) decoder 406 embedded within encoder 400 that reconstructs the symbols to create sample data that a (remote) decoder would also create (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). This reconstructed sample stream is input into reference picture memory 405. Since the decoding of the symbol stream results in a bit-exact result independent of the decoder location (local or remote), the reference picture buffer content is also bit-exact between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same as the sample values "seen" by the decoder when using the prediction during decoding. The basic principles of reference image synchronization (and the resulting drift, if synchronization cannot be maintained, e.g., due to channel errors) are well known to those skilled in the art.
[0052] The operation of the "local" decoder 406 can be the same as that of the "remote" decoder 300 described in detail above in conjunction Figure 3 with. However, briefly referring Figure 4 to, since the symbols are available and the entropy encoder 408 and parser 304 perform lossless encoding of the symbols of the encoded video sequence, the entropy decoding portion of decoder 300, including channel 301, receiver 302, buffer 303, and parser 304, may not be fully implementable in local decoder 406.
[0053] It can be observed at this point that any decoder technology other than parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. The description of encoder technology can be simplified since encoder technology is reciprocal to the decoder technology described in detail. More detailed description is only needed in certain areas and is provided below.
[0054] As part of its operation, source encoder 403 can perform motion-compensated predictive coding that predicts and encodes an input frame by referring to at least one previously encoded frame designated as a "reference frame" in the video sequence. In this way, coding engine 407 encodes the difference between a pixel block of the input frame and pixel blocks of at least one reference frame that can be selected as a prediction reference for the input frame.
[0055] The local video decoder 406 can decode the encoded video data of the frames that can be designated as reference frames based on the symbols created by the source encoder 403. The operation of the encoding engine 407 can advantageously be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 4 not shown), the reconstructed video sequence may generally be a copy of the source video sequence, but with some errors. The local video decoder 406 replicates the decoding process that can be performed by the video decoder on the reference frames and can cause the reconstructed reference frames to be stored in the reference picture memory 405, which can be a cache, for example. In this way, the encoder 400 can locally store a copy of the reconstructed reference frames, which has the same content as the reconstructed reference frames obtained by the remote video decoder (in the absence of transmission errors).
[0056] The predictor 404 can perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 can search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new picture. The predictor 404 can operate block by block based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 404, it can be determined that the input picture can have prediction references taken from at least two reference pictures stored in the reference picture memory 405.
[0057] The controller 402 can manage the encoding operations of the source encoder 403, which can be a video encoder, for example, including setting parameters and subgroup parameters for encoding the video data.
[0058] The outputs of all the above functional units can be entropy encoded in the entropy encoder 408. The entropy encoder performs lossless compression on the symbols generated by various functional units according to techniques well known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0059] The transmitter 409 can buffer the encoded video sequence created by the entropy encoder 408 to prepare for transmission over the communication channel 411, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter 409 can merge the encoded video data from the source encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0060] The controller 402 may manage the operation of the encoder 022. During encoding, the controller 402 may assign a certain type of encoded picture to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures may generally be assigned to any of the following frame types:
[0061] Intra pictures (I pictures), which may be pictures that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and characteristics.
[0062] Predictive pictures (P pictures), which may be pictures that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0063] Bi - predictive pictures (B pictures), which may be pictures that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, at least two predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.
[0064] Source pictures may generally be spatially subdivided into at least two sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block - by - block. These blocks may be predictively encoded with reference to other (encoded) blocks, which are determined according to the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture may be non - predictively encoded, or the blocks may be predictively encoded with reference to already - encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be non - predictively encoded with reference to one previously - encoded reference picture either through spatial prediction or through temporal prediction. Blocks of a B picture may be non - predictively encoded with reference to one or two previously - encoded reference pictures either through spatial prediction or through temporal prediction.
[0065] The encoder 400 (which may be, for example, a video encoder) may perform encoding operations according to a predetermined video encoding technique or standard such as the ITU - T H.266 recommendation. In operation, the encoder 400 may perform various compression operations, including predictive encoding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0066] In an embodiment, the transmitter 409 may transmit additional data when transmitting the encoded video. The source encoder 403 may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set segments, etc.
[0067] The compressed video may be enhanced in the video bitstream by supplementary enhancement information, for example in the form of supplementary enhancement information (SEI) messages or video usability information (VUI). The video coding standard may include a normative part for SEI and VUI. The SEI and VUI information may also be specified in a separate specification referenced by the video coding specification.
[0068] Reference Figure 5 Example 500 of shows an exemplary layout of a coded video sequence (CVS) according to H.266. The coded video sequence is subdivided into network abstraction layer units (NAL units). The exemplary NAL unit 501 may include a nal_unit_header 502, which in turn includes the following 16 bits: The forbidden_zero_bit 503 and nuh_reserved_zero_bit 504 may not be used by H.266 and may be zero in a NAL unit compliant with H.266. The three bits of nuh_layer_id 505 may indicate the layer (spatial, SNR, or multi-view enhancement) to which the NAL unit belongs. The five bits of nuh_nal_unit_type define the type of the NAL unit. In H.266, 22 NAL unit type values are defined for the NAL unit types defined in H.266, six NAL unit types are reserved, and four NAL unit type values are unspecified and may be used by specifications other than H.266. Finally, the three bits of the NAL unit header indicate the temporal layer nuh_temporal_id_plus1 506 to which the NAL unit belongs.
[0069] The encoded picture may contain at least one Video Coding Layer (VCL) NAL unit and zero or at least two non-VCL NAL units. The VCL NAL unit may contain encoded data that conceptually belongs to the video coding layer as described above. The non-VCL NAL unit may contain data that conceptually does not belong to the video coding layer. Taking H.266 as an example, they can be classified into (1) parameter sets, (2) picture headers (PH_NUT), (3) NAL units, (4) prefix and suffix SEI Nal unit types (PREFIX_SEI_NUT and SUFFIX_SEI_NUT), (5) filler data NAL unit type FD_NUT, and (6) reserved and unspecified NAL unit types, as follows.
[0070] (1) Parameter sets, including information required for the decoding process, can be applied to at least two encoded pictures. Parameter sets and conceptually similar NAL units can be NAL unit types, for example, DCI_NUT (Decoding Capability Information (DCI)), VPS_NUT (Video Parameter Set (VPS), establishing layer relationships, etc.), SPS_NUT (Sequence Parameter Set (SPS), establishing, among other things, the parameters used and remaining constant throughout the encoded video sequence CVS), PPS_NUT (Picture Parameter Set (PPS), establishing the parameters used within the encoded picture and remaining constant), and PREFIX_APS_NUT and SUFFIX_APS_NUT (prefix and suffix adaptive parameter sets). The parameter set may include information required by the decoder to decode the VCL NAL unit, so it is called a "canonical" NAL unit here.
[0071] (2) Picture headers (PH_NUT), which are also a "standard" NAL unit.
[0072] (3) NAL units that mark certain positions in the NAL unit stream. These include NAL units of NAL unit types AUD_NUT (Access Unit Delimiter), EOS_NUT (End of Sequence), and EOB_NUT (End of Bitstream). These are non-canonical and are also called informative. Because a compliant decoder does not require them during the decoding process, although it needs to be able to receive them in the NAL unit stream.
[0073] (4) PREFIX_SEI_NUT and SUFFIX_SEI_NUT for prefix and suffix SEI NAL unit types represent NAL units containing prefix and suffix supplementary enhancement information. In H.266, these NAL units are informative as they are not essential for the decoding process.
[0074] (5) The FD_NUT for filler data NAL unit type represents filler data; it can be random data and can be used to "waste" bits in the NAL unit stream or bitstream, which may be necessary for transmission in certain isochronous transport environments.
[0075] is random and can be used to "waste" bits in the NAL unit stream or bitstream, which may be necessary for transmission over certain synchronous transport environments.
[0076] (6) Reserved and undefined NAL unit types.
[0077] Still refer to Figure 5 , showing the layout of the NAL unit stream in decoding order 510. This NAL unit stream contains the encoded picture 511. This encoded picture 511 contains some of the types of NAL units introduced previously. Somewhere early in the NAL unit stream, DCI 512, VPS 513, and SPS 514 can be combined to set up the decoder. This decoder can be used to decode the parameters of at least two encoded pictures in the encoded video sequence (CVS), including the encoded picture 511 of the NAL unit stream.
[0078] The encoded picture 511 can contain, in the order shown, or in any other order conforming to the video coding technology or standard in use (here: H.266): Prefix-APS 516, Picture Header 517, Prefix-SEI 518, at least one VCL NAL unit 519, and Suffix-SEI 520.
[0079] Prefix and suffix SEI NAL units 518 and 520 are motivated during standardization because for some SEI messages, the content of the message is known before the start of encoding of a given picture, while other content is only known after the picture is encoded. With prefix and suffix SEI, certain SEI messages are allowed to appear earlier or later in the NAL unit stream of an encoded picture, which can avoid buffering. For example, in an encoder, the sampling time of the picture to be encoded is known before the picture is encoded, so the picture timing SEI message can be a prefix SEI message (Prefix-SEI) 516. On the other hand, the hash SEI message of the decoded picture is a suffix SEI message (Suffix-SEI) 518 because the encoder cannot compute the hash of the reconstructed samples before the picture is encoded, so the message contains the hash of the sample values of the decoded picture, which can be used, for example, to debug the implementation of the encoder. The positions of the prefix and suffix SEI NAL units may not be limited to their positions in the NAL unit stream. The phrases "prefix" and "suffix" may mean which encoded pictures or NAL units the prefix / suffix SEI messages may belong to, and the details of such applicability can be specified, for example, in the semantic description of a given SEI message.
[0080] Still referring to Figure 5 , a simplified syntax diagram of a NAL unit containing a prefix or suffix SEI message 520 is shown. This syntax is a container format for carrying at least two SEI messages in one NAL message. For clarity, the details of the emulation prevention syntax specified in H.266 are omitted here. Like other NAL units, the SEI NAL unit starts with a NAL unit header 521. After the header is at least one SEI message. Two are depicted as 530, 531 and are described below. Each SEI message within the SEI NAL unit includes an 8-bit payload_type_byte 522 that specifies one of 256 different SEI types; an 8-bit payload_size_byte 523 that specifies the number of bytes of the SEI payload, and a payload of payload_size-byte bytes 524. This structure can be repeated until the payload_type_byte is observed to be equal to 0xff, which indicates the end of the NAL unit. The syntax of the payload 524 depends on the SEI message and can be any length between 0 and 255 bytes.
[0081] The following description uses H.266 / VVC and its SEI message syntax as an example; however, similar techniques can be applied to other video compression technologies, especially H.264 and H.265 and future video compression technologies.
[0082] To overcome the 255 - byte size limit of the SEI payload, two basic design alternatives are considered according to an embodiment.
[0083] The first design alternative according to an embodiment is to use the currently available NAL unit types for the SEI message syntax, which can use at least two bytes as the syntax element currently called payload_size_byte. In this design alternative, some drawbacks of the SEI message syntax can also be fixed, as described below. The original SEI message definition and the new SEI message syntax can co - exist.
[0084] The second design alternative according to an embodiment is to split the payload larger than 255 bytes into segments of 255 bytes or less as needed or desired. Control information may be required to identify the first, any intermediate shards, and the last fragment of the segmented payload for various purposes, including error recovery.
[0085] Reference Figure 6 and Figure 7 Examples 600 and 700 of show the syntax and semantics of an exemplary implementation of the first design alternative.
[0086] In one embodiment, in the NAL unit type definition table 601, which is reproduced only in part 602 omitting code points not relevant to the disclosed subject matter, a previously reserved NAL unit type (here as example 26) can be assigned to represent a large supplementary enhancement information RBSP, also known as lsei_rbsp. Doing so means that there is only one unassigned reserved NAL unit type 27 604.
[0087] An example syntax of the RBSP of the new NAL unit type is as Figure 7 shown. The lsei_rbsp 701 can have a syntax similar to the sei_rbsp of H.266, except that it references at least one lsei_message() 703 in a do loop 702.
[0088] Similarly in Figure 7 the syntax of lsei_message() 710 is shown. This syntax may be completely different from the H.266 sei_message() syntax introduced below. For clarity, some syntax mechanisms implemented in H.266 for purposes such as preventing emulation are omitted.
[0089] The lsei_message() syntax 710 may include, for example, the syntax element lsei_position 711. This syntax element may be used for the same purposes as the currently distinguished NAL unit types PREFIX_SEI_NUT and POSTFIX_SEI_NUT. In other words, for long SEI messages, only one such NAL unit type may be used instead of two NAL units, and the first two bits that are easy to find and parse in the payload of the RBSP are used to distinguish between prefix and suffix SEI messages. For future scalability, two bits are recommended. A value of 0 may indicate a prefix large NAL unit, while a value of 1 may indicate a suffix large NAL unit.
[0090] The syntax element lsei_relevance 712 may indicate the relevance of the SEI message selected by the encoder, with abstract values available, for example, between 0 and 3, where 0 may be the least relevant and 3 may be the most relevant. Recent projects in MPEG have started using SEI messages. For some applications and possibly some profiles, it is necessary not only to process SEI messages for a good user experience (like current SEI messages), but also to maintain the integrity of the decoding process outside the H.266 decoder, but still within the non-H.266 environment of the specification. Specifications outside H.266 may state that certain SEIs must be available and processed at the H.266 decoder for forwarding to entities downstream of the H.266 decoder. At least two network middleboxes cannot discard such SEIs. lsei_relevance can be used for all these purposes. The relevance may follow the semantics defined in the SEI list SEI of H.266.
[0091] In the exemplary syntax, four bits 713 are reserved for byte alignment and future extension.
[0092] The 8 bits reserved for lsei_payload_type_byte 714 may have semantics similar to the payload_type_byte of H.266. The 8 bits allow for 256 different messages.
[0093] payload_size_16bits 715 allows for a payload size of up to 64k bytes.
[0094] The lsei_payload() syntax 720 may include an if-then-elseif chain 721 of various defined payload types, followed by byte alignment of the RBSP 722.
[0095] From the perspective of coding efficiency, the second design solution described above may be less efficient and may not be as concise as the first solution, but it has the advantage of not requiring one of the two unallocated NAL unit types.
[0096] According to an embodiment, each large SEI message consists of variables specifying the type payloadType and size payloadSize of the large SEI message payload. The payload of the large SEI message is specified in Appendix D. The derived large SEI message payload size payloadSize is specified in bytes and shall be equal to the number of RBSP bytes in the large SEI message payload.
[0097] Note - The NAL unit byte sequence containing the large SEI message may include at least one emulation prevention byte (represented by the emulation_prevention_three_byte syntax element). Since the payload size of the large SEI message is specified in terms of RBSP bytes, the number of emulation prevention bytes is not included in the large SEI payload size payloadSize.
[0098] According to an embodiment, lsei_position indicates whether the SEI message corresponds to PREFIX_SEI_NUT and SUFFIX_SEI_NUT. An lsei_position equal to 0 indicates that the SEI message is considered a PREFIX_SEI_NUT. An lsei_position equal to 1 indicates that the SEI message is considered a SUFFIX_SEI_NUT. The values 3 and 4 of lsei_position are reserved for future use and shall be ignored.
[0099] According to an embodiment, lsei_relevance indicates the relevance of the SEI message to the target application. The range of lsei_relevance is from 0 to 3, with 0 being the lowest relevance and 3 being the highest relevance.
[0100] Note - The relevance of the SEI message is arbitrarily determined and its use is specified by the target application.
[0101] According to an embodiment, lsei_reserved is reserved for future use and shall be ignored.
[0102] According to an embodiment, lsei_payload_type_byte is the byte of the payload type of a large SEI message. payloadType = lsei_payload_type_byte. And payload_size_16bits is the payload size (in bits) of a large SEI message. payloadSize = payload_size_16bits.
[0103] It is proposed to define a large SEI message header indicating: the payload size in a fixed-length 16-bit syntax element, allowing signaling of any payload up to 64 kB. The signaling of the large SEI position can be used for the same purposes as the currently differentiated prefix and suffix SEI NAL unit types, meaning only one NAL unit type is needed (1 bit is sufficient, but 2 bits allow for defining future scenarios). SEI message relevance, indicating the importance of the SEI message for a given application. Using 2 bits will allow signaling of 4 levels of relevance information (0 being the least relevant and 3 being the most relevant). For byte alignment, 4 reserved bits can be used in the future. For example, these bits can be used to signal payload sizes greater than 64 kB.
[0104] Therefore, for such SEI messages, at least the problems in the HEVC and VVC specifications can be improved.
[0105] Reference Figure 8 , in one embodiment, an SEI payload is shown, such as 900 bytes 801, which is above the 255-byte limit allowed in the H.266 syntax. The payload is divided into four segments: a first segment 802, two intermediate segments 803, 804, and a last segment 805. The size of each of these segments (as well as the associated header information / control information, see below) can be chosen so as not to exceed the 255-byte limit. Thus, each segment can be contained in its own SEI message. None, some, or all of the segments can have a header 806 containing control information.
[0106] For the purpose of fragmentation, a syntax may be needed to know where the fragmentation chain starts and ends. Without such a syntax, it may not be clear what is in the fragmented payload and where the payload boundaries are when at least two fragments appear successively. There are many different options for encoding such start and end information, some of which will be briefly introduced.
[0107] In a first option according to an embodiment, each segment may be preceded by a header 810 that includes at least a start bit 811 and an end bit 812. The header may be padded 813 until byte alignment is achieved, such that the padded segment becomes a simple byte-oriented memcpy() operation.
[0108] In a second option, the start segment may include a header 820 that may include at least two segments 821 that together form a payload. To improve fault tolerance, this field may count down to zero for intermediate segments and the last segment, and count down to 0 for the last segment. Using all 8 bits of the minimum byte-aligned header size, there may be 256 segments, which would allow for an SEI payload of nearly 64k bytes.
[0109] In a third option, no header is required. Instead, the first, intermediate, and last segments may be marked by selecting three different values for the SEI payload_type_byte 522. For example, the first segment 802 may be encoded as an SEI message with a payload_type_byte equal to 60, the second and third intermediate segments may be encoded with a payload_type_byte equal to 61, and the third intermediate segment may be encoded with a payload_type_byte equal to 62.
[0110] A combination of these mechanisms is possible. For example, it may make sense to use a header with start / end bits to mark the start and end segments, and different payload_type_bytes to mark the intermediate packets.
[0111] In the receiver, using this segmentation can be relatively straightforward. After receiving a segment marked as the start segment, for example, by any one or a combination of the three mechanisms described above, the decoder may copy the payload of the segment to the start of an assembly buffer. The decoder may further set itself to a state in which it expects an end segment or an intermediate segment. After receiving an end or intermediate segment, the decoder may copy the payload of the segment to the assembly buffer after the payload of the previous segment. If the segment is an intermediate segment, the internal state of the decoder may remain unchanged, while if the decoder receives an end segment, after copying, the decoder may forward the contents of the assembly buffer to the using entity and reset the state to expect a start segment. Error conditions can be easily detected by a state machine.
[0112] The techniques described above for large SEI messages may be implemented as computer software using computer-readable instructions and physically stored in at least one computer-readable medium. For example, Figure 9A computer system 900 is shown, which is suitable for implementing certain embodiments of the disclosed subject matter.
[0113] The computer software can be encoded in any suitable machine code or computer language, and code including instructions is created through mechanisms such as assembly, compilation, and linking. The instructions can be directly executed by at least one computer central processing unit (CPU), graphics processing unit (GPU), etc., or executed through decoding, microcode, etc.
[0114] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0115] Figure 9 The components shown for computer system 900 are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependence on or requirement for any one component or combination thereof shown in the exemplary embodiments of computer system 900.
[0116] Computer system 900 may include certain human-machine interface input devices. Such human-machine interface input devices can respond to the input of at least one human user through tactile input (such as keyboard input, swiping, data glove movement), audio input (such as sound, applause), visual input (such as gestures), and olfactory input (not shown). The human-machine interface device can also be used to capture certain media, which does not have to be directly related to the conscious input of humans, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0117] The human-machine interface input device may include at least one of the following (only one of which is drawn): keyboard 901, mouse 902, touchpad 903, touch screen 910, joystick 905, microphone 906, scanner 908, camera 907.
[0118] The computer system 900 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of at least one human user through, for example, haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include haptic output devices (e.g., haptic feedback through the touch screen 910 or the joystick 905, but there may also be haptic feedback devices that do not serve as input devices), audio output devices (e.g., speakers 909, headphones (not shown)), visual output devices (e.g., a screen 910 including a cathode ray tube screen, a liquid crystal screen, a plasma screen, an organic light emitting diode screen, each of which has or does not have touch screen input functionality, each of which has or does not have haptic feedback functionality - some of which may output two-dimensional visual output or output above three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).
[0119] The computer system 900 may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical discs (CD / DVD ROM / RW 920) with CD / DVD 911 or similar media, thumb drives 922, removable hard disk drives or solid state drives 923, traditional magnetic media such as tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.
[0120] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0121] The computer system 900 may also include an interface 999 to at least one communication network 998. For example, the network 998 may be wireless, wired, or optical. The network 998 may also be a local area network, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, and so on. The network 998 also includes local area networks such as Ethernet, wireless local area network, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), wired or wireless wide area digital television networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANBus), and so on. Some networks 998 typically require an external network interface adapter for connection to certain common data ports or peripheral buses (950 and 951) (e.g., the USB port of the computer system 900); other systems are typically integrated into the core of the computer system 900 by connecting to the system bus described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks 998, the computer system 900 can communicate with other entities. The communication may be unidirectional, only for receiving (e.g., wireless television), unidirectional only for sending (e.g., CAN bus to certain CAN bus devices), or bidirectional, e.g., to other computer systems via a local or wide area digital network. Each of the above networks and network interfaces may use certain protocols and protocol stacks.
[0122] The above-described human-machine interface device, human-accessible storage device, and network interface may be connected to the core 940 of the computer system 900.
[0123] The core 940 may include at least one central processing unit (CPU) 941, a graphics processing unit (GPU) 942, a graphics adapter 917, a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) 943, a hardware accelerator 944 for specific tasks, and so on. These devices, as well as a read-only memory (ROM) 945, a random access memory 946, an internal mass storage (e.g., an internal non-user-accessible hard disk drive, a solid-state drive, etc.) 947, and so on, may be connected via a system bus 948. In some computer systems, the system bus 948 may be accessed in the form of at least one physical plug for expansion with additional central processing units, graphics processing units, and so on. Peripherals may be directly attached to the system bus 948 of the core or connected via a peripheral bus 949. The architecture of the peripheral bus includes an external controller interface PCI, a universal serial bus USB, and so on.
[0124] The CPU 941, GPU 942, FPGA 943, and accelerator 944 can execute certain instructions, which, when combined, can constitute the aforementioned computer code. This computer code can be stored in the ROM 945 or RAM 946. Transitional data can also be stored in the RAM 946, while permanent data can be stored in, for example, the internal mass storage 947. Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with at least one of the CPU 941, GPU 942, mass storage 947, ROM 945, RAM 946, etc.
[0125] The computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this application or can be of the kind well-known and available to those skilled in the art of computer software.
[0126] By way of example and not limitation, a computer system having the architecture 900, and in particular the core 940, can provide the functionality of a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in at least one tangible computer-readable medium. Such a computer-readable medium can be a medium associated with the aforementioned user-accessible mass storage, as well as a specific memory of the non-volatile core 940, such as the core internal mass storage 947 or ROM 945. The software implementing the various embodiments of this application can be stored in such a device and executed by the core 940. Depending on specific requirements, the computer-readable medium can include one or more storage devices or chips. The software can cause the core 940, and in particular the processors therein (including the CPU, GPU, FPGA, etc.), to execute the specific processes or specific portions of specific processes described herein, including defining data structures stored in the RAM 946 and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality that is logically hardwired or otherwise included in a circuit (e.g., the accelerator 944), which can operate in place of or in conjunction with the software to execute the specific processes or specific portions of specific processes described herein. In appropriate cases, references to software can include logic and vice versa. In appropriate cases, references to the computer-readable medium can include a circuit (such as an integrated circuit (IC)) that stores the software for execution, a circuit that contains the execution logic, or both. This application encompasses any suitable combination of hardware and software.
[0127] Although the present application has described at least two exemplary embodiments, various changes, permutations, and various equivalent substitutions of the embodiments are within the scope of the present application. Therefore, it should be understood that those skilled in the art can design various systems and methods, which, although not explicitly shown or described herein, embody the principles of the present application and are thus within the spirit and scope of the present application.
[0128] The above disclosure also includes the features described below. These features can be combined in various ways and are not limited to the combinations mentioned below.
[0129] (1) A method for video decoding, the method comprising: obtaining video data including at least one encoded picture; identifying, by a decoder, at least one first supplementary enhancement information (SEI) message, the at least one first SEI message indicating variables that specify a type (payloadType) and a size (payloadSize) of a payload of the at least one SEI message, and specifying in bytes; and decoding, by the decoder, the video data based on the first SEI message.
[0130] (2) The method of feature (1), wherein obtaining the video data includes: receiving a network abstraction layer (NAL) unit, wherein a type of the NAL unit is different from a type of the NAL unit in the at least one first SEI message.
[0131] (3) The method according to any one of features (1) to (2), wherein the received NAL unit includes at least one length header field, wherein the length header field specifies a value greater than 256 bytes; a maximum size of a payload of an SEI message is less than or equal to 256 bytes, and decoding the video data includes: extracting a payload greater than 256 bytes from the NAL unit.
[0132] (4) The method according to any one of features (1) to (3), wherein the at least one first SEI message includes a start segment, and the start segment includes a payload of the start segment.
[0133] (5) The method according to any one of features (1) to (4), wherein decoding the video data further includes: receiving a second SEI message including an end segment, and the end segment includes a payload of the end segment.
[0134] (6) The method according to any one of features (1) to (5), wherein decoding the video data further includes: extracting the start segment and the end segment.
[0135] (7) The method according to any one of features (1) to (6), wherein decoding the video data further includes assembling a large SEI payload based on the extracted start segment and the extracted end segment, wherein the start segment is at the start of the assembled large SEI payload, and the end segment is at the end of the assembled large SEI payload, and the large SEI payload is at least greater than 256 bytes.
[0136] (8) A method for video encoding, the method comprising: obtaining video data including at least one picture;
[0137] encoding the video data and the at least one picture such that the encoded video data indicates an encoded version of the at least one picture and at least one first supplementary enhancement information SEI message, wherein the at least one first SEI message indicates variables that specify the type payloadType and size payloadSize of the payload of the at least one SEI message, and is specified in bytes.
[0138] (9) The method of feature (8), wherein the encoded video data indicates a network abstraction layer NAL unit, and wherein the type of the NAL unit is different from the type of the NAL unit in the at least one first SEI message.
[0139] (10) The method according to any one of features (8) to (9), wherein the received NAL unit includes at least one length header field, wherein the length header field specifies a value greater than 256 bytes; wherein the maximum size of a SEI message payload is less than or equal to 256 bytes, and the encoded video data is configured to be decoded based on a payload greater than 256 bytes extracted from the NAL unit.
[0140] (11) The method according to any one of features (8) to (10), wherein the at least one first SEI message includes a start segment, and the start segment includes the payload of the start segment.
[0141] (12) The method according to any one of features (8) to (11), wherein the encoded video data indicates a second SEI message including an end segment, and the end segment includes the payload of the end segment.
[0142] (13) The method according to any one of features (8) to (12), wherein the encoded video data is configured to be decoded based on extracting the start segment and the end segment.
[0143] (14) The method according to any one of features (8) to (13), wherein the encoded video data is configured to be decoded based on assembling a large SEI payload from the extracted start segment and the extracted end segment, wherein the start segment is at the start of the assembled large SEI payload, and the end segment is at the end of the assembled large SEI payload, and the large SEI payload is at least greater than 256 bytes.
[0144] (15) A method of processing visual media data, the method comprising: performing a conversion between a visual media file and a bitstream of visual media data according to format rules, wherein the format rules indicate processing of at least one first supplementary enhancement information (SEI) message, wherein the at least one first SEI message indicates variables that specify a type payloadType and a size payloadSize of the payload of the at least one SEI message, and are specified in bytes.
[0145] (16) The method of feature (15), wherein the format rules further indicate processing of a network abstraction layer (NAL) unit, wherein the type of the NAL unit is different from the type of the NAL unit in the at least one first SEI message.
[0146] (17) The method according to any one of features (15) to (16), wherein the NAL unit includes at least one length header field, wherein the length header field specifies a value greater than 256 bytes; wherein the maximum size of a payload of an SEI message is less than or equal to 256 bytes, and the format rules further indicate: extracting a payload greater than 256 bytes from the NAL unit.
[0147] (18) The method according to any one of features (15) to (17), wherein the at least one first SEI message includes a start segment, and the start segment includes a payload of the start segment.
[0148] (19) The method according to any one of features (15) to (18), wherein the format rules further indicate processing of a second SEI message including an end segment, and the end segment includes a payload of the end segment.
[0149] (20) The method according to any one of features (15) to (19), wherein the format rules further indicate processing the start segment and the end segment, and assembling a large SEI payload based on the extracted start segment and the extracted end segment, wherein the start segment is at the start of the assembled large SEI payload, and the end segment is at the end of the assembled large SEI payload, and the large SEI payload is at least greater than 256 bytes.
[0150] (21) A video decoding device, comprising a processing circuit configured to perform the method according to any one of features (1) to (7).
[0151] (22) A video encoding device, comprising a processing circuit configured to perform the method according to any one of features (8) to (15).
[0152] (23) A non - volatile computer - readable storage medium storing instructions, wherein when the instructions are executed by at least one processor, the at least one processor is caused to perform the method according to any one of features (1) to (20).
Claims
1. A video decoding method, characterized in that: The method comprises: Obtaining video data including at least one encoded picture; identifying, by a decoder, at least one first supplemental enhancement information (SEI) message, the at least one first SEI message indicating variables that specify a type payloadType and a size payloadSize of a payload of the at least one SEI message, and that are specified in bytes; and, The video data is decoded by the decoder based on the first SEI message.
2. The method according to claim 1, characterized in that Obtaining the video data includes receiving a network abstraction layer NAL unit, wherein the type of the NAL unit is different from the NAL unit type in the at least one first SEI message.
3. The method according to claim 2, characterized in that The received NAL unit includes at least one length header field, wherein the length header field specifies a value greater than 256 bytes; The maximum size of an SEI message payload is less than or equal to 256 bytes, and Decoding the video data includes extracting a payload greater than 256 bytes from the NAL unit.
4. The method according to claim 1, characterized in that: in, The at least one first SEI message includes a start segment including a payload of the start segment.
5. The method according to claim 4, characterized in that Decoding the video data further includes receiving a second SEI message including an end segment, the end segment including a payload of the end segment.
6. The method according to claim 5, characterized in that Decoding the video data further includes extracting the start segment and the end segment.
7. The method according to claim 6, characterized in that Decoding the video data further assembles a large SEI payload based on the extracted start fragment and the extracted end fragment, wherein the start fragment is at the beginning of the assembled large SEI payload and the end fragment is at the end of the assembled large SEI payload, and the large SEI payload is at least larger than 256 bytes.
8. A video encoding method, characterized in that: The method comprises: Obtaining video data including at least one picture; encoding the video data and the at least one picture so that the encoded video data indicates an encoded version of the at least one picture and at least one first supplemental enhancement information (SEI) message, The at least one first SEI message indicates a variable, and the variable specifies the type payloadType and size payloadSize of the payload of the at least one SEI message, and is specified in bytes.
9. The method according to claim 8, characterized in that The coded video data indicates a network abstraction layer NAL unit, wherein a type of the NAL unit is different from a NAL unit type in the at least one first SEI message.
10. The method according to claim 9, characterized in that The received NAL unit includes at least one length header field, wherein the length header field specifies a value greater than 256 bytes; The maximum size of an SEI message payload is less than or equal to 256 bytes, and The encoded video data is configured to be decoded based on a payload greater than 256 bytes extracted from the NAL unit.
11. The method according to claim 8, characterized in that in, The at least one first SEI message includes a start segment including a payload of the start segment.
12. The method according to claim 11, characterized in that The coded video data indication includes a second SEI message of an end segment, wherein the end segment includes a payload of the end segment.
13. The method according to claim 12, characterized in that The encoded video data is configured to be decoded based on extracting the start segment and the end segment.
14. The method according to claim 13, characterized in that The encoded video data is configured to be decoded by assembling a large SEI payload based on the extracted start fragment and the extracted end fragment, wherein the start fragment is at the beginning of the assembled large SEI payload and the end fragment is at the end of the assembled large SEI payload, and the large SEI payload is at least larger than 256 bytes.
15. A method for processing visual media data, characterized in that: The method comprises: performing conversion between a visual media file and a code stream of visual media data according to a format rule, wherein the format rule indicates processing of at least one first supplemental enhancement information (SEI) message, The at least one first SEI message indicates a variable, and the variable specifies the type payloadType and size payloadSize of the payload of the at least one SEI message, and is specified in bytes.
16. The method according to claim 15, characterized in that The format rule further indicates processing of a network abstraction layer NAL unit, wherein the type of the NAL unit is different from the type of the NAL unit in the at least one first SEI message.
17. The method according to claim 16, characterized in that The NAL unit includes at least one length header field, wherein the length header field specifies a value greater than 256 bytes; The maximum size of an SEI message payload is less than or equal to 256 bytes, and The format rule further indicates that a payload larger than 256 bytes is extracted from the NAL unit.
18. The method according to claim 15, characterized in that in, The at least one first SEI message includes a start segment including a payload of the start segment.
19. The method according to claim 18, characterized in that The format rule further instructs processing of a second SEI message including an end segment, wherein the end segment includes a payload of the end segment.
20. The method according to claim 19, characterized in that The format rule further instructs processing the start segment and the end segment, and assembling a large SEI payload based on the extracted start segment and the extracted end segment, wherein the start segment is at the beginning of the assembled large SEI payload and the end segment is at the end of the assembled large SEI payload, and the large SEI payload is at least larger than 256 bytes.