Media file generation / reception method and apparatus supporting sample-unit random access, and method of transmitting media file
By generating and receiving media files that support random access on a sample-by-sample basis, and constraining the current sample to be an IRAP sub-frame and a TemporalId value of 0, the storage and transmission cost issues of high-resolution images are solved, achieving efficient media file management and seamless reproduction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2021-12-14
- Publication Date
- 2026-05-15
AI Technical Summary
With the increasing demand for high-resolution and high-quality images, existing technologies face challenges in terms of storage and transmission costs, especially on mobile devices with limited hardware and network resources, requiring more efficient image compression and file processing technologies to support random access.
By generating and receiving media files that support random access on a sample-by-sample basis, constraining the current sample to include IRAP sub-pictures, and ensuring that the current sample's TemporalId value is 0, efficient storage and transmission of media files are achieved.
It achieves efficient storage and transmission of media files during random access, addressing current technical challenges, and supports seamless media content reproduction of high-resolution and high-quality images.
Smart Images

Figure CN116569557B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to methods and apparatus for generating / receiving media files, and more specifically, to methods and apparatus for generating / receiving media files that support random access on a sample-by-sample basis, and to methods for transmitting media files generated by the methods / apparatus for generating the media files disclosed herein. Background Technology
[0002] Recently, there has been an increasing demand for high-resolution and high-quality images, such as 360-degree images. As image resolution or quality increases, file size or frame rate also increases, inevitably leading to higher storage and transmission costs. Furthermore, with the growing popularity of mobile devices such as smartphones and tablet PCs, the demand for network-based multimedia services is rapidly increasing. However, there are limitations in the hardware and network resources available for multimedia services.
[0003] Therefore, there is a need for efficient image compression and file processing technologies to store and transmit image data more effectively. Summary of the Invention
[0004] Technical issues
[0005] According to this disclosure, the purpose of this disclosure is to provide methods and apparatus for generating / receiving media files that support random access on a sample-by-sample basis.
[0006] According to this disclosure, the purpose of this disclosure is to provide a method and apparatus for generating / receiving media files, wherein, when random access occurs, the current sample is constrained to include IRAP sub-screens.
[0007] According to this disclosure, the purpose of this disclosure is to provide a method and apparatus for generating / receiving media files, wherein, when random access occurs, the TemporalId value of the current sample is constrained to 0.
[0008] One object of this disclosure is to provide a method for transmitting media files generated by a media file generation method or device according to this disclosure.
[0009] One object of this disclosure is to provide a recording medium for storing media files generated by a media file generation method or apparatus according to this disclosure.
[0010] One object of this disclosure is to provide a recording medium for storing media files received by a media file receiving device according to this disclosure and used for reconstructing images.
[0011] The technical problems solved by this disclosure are not limited to those described above. Other technical problems not described herein will become clear to those skilled in the art through the following description.
[0012] Technical solution
[0013] A media file receiving method according to one aspect of this disclosure may include the steps of: obtaining one or more tracks and sample groups from a media file; and processing video data in the media file by reconstructing samples included in the tracks based on the sample groups. Based on the presence of samples in the tracks mapped to a predetermined type of stream access point sample group or random access recovery point sample group, the current sample in the mapped samples may be constrained to include at least one intra-frame random access point (IRAP) subframe, and based on the current sample including non-IRAP subframes, samples belonging to the same coding layer video sequence (CLVS) as the current sample and following the current sample in decoding order may be constrained to include at least one IRAP subframe having the same subframe index value as the non-IRAP subframe.
[0014] According to another aspect of this disclosure, a media file receiving device may include a memory and at least one processor. The at least one processor may obtain one or more tracks and sample groups from the media file and process video data in the media file by reconstructing samples included in the tracks based on the sample groups. Based on the presence of samples in the tracks mapped to a predetermined type of stream access point sample group or random access recovery point sample group, the current sample in the mapped samples may be constrained to include at least one intra-frame random access point (IRAP) subframe; and based on the current sample including non-IRAP subframes, samples belonging to the same coding layer video sequence (CLVS) as the current sample and following the current sample in decoding order may be constrained to include at least one IRAP subframe having the same subframe index value as the non-IRAP subframe.
[0015] A media file generation method according to another aspect of this disclosure may include the following steps: encoding the video data; generating one or more tracks and sample groups for the encoded video data; and generating the media file based on the generated tracks and sample groups. Based on the presence of samples in the tracks mapped to a predetermined type of stream access point sample group or random access recovery point sample group, the current sample in the mapped samples may be constrained to include at least one intra-frame random access point (IRAP) subframe, and based on the current sample including non-IRAP subframes, samples belonging to the same coding layer video sequence (CLVS) as the current sample and following the current sample in decoding order may be constrained to include at least one IRAP subframe having the same subframe index value as the non-IRAP subframe.
[0016] According to another aspect of this disclosure, a media file generation apparatus may include a memory and at least one processor. The at least one processor may encode video data; generate one or more tracks and sample groups for the encoded video data; and generate a media file based on the generated tracks and sample groups. Based on the presence of samples in the tracks mapped to a predetermined type of stream access point sample group or random access recovery point sample group, the current sample in the mapped samples may be constrained to include at least one intra-frame random access point (IRAP) subframe; and based on the current sample including non-IRAP subframes, samples belonging to the same coding layer video sequence (CLVS) as the current sample and following the current sample in decoding order may be constrained to include at least one IRAP subframe having the same subframe index value as the non-IRAP subframe.
[0017] In another aspect of the media file transmission method according to this disclosure, a media file generated by the media file generation method or device of this disclosure can be transmitted.
[0018] According to another aspect of this disclosure, a computer-readable recording medium may store media files generated by the media file generation method or apparatus of this disclosure.
[0019] The features described above in this brief overview are merely exemplary aspects of the following detailed description of this disclosure and do not limit the scope of this disclosure.
[0020] Beneficial effects
[0021] According to this disclosure, methods and apparatus for generating / receiving media files that support random access on a sample-by-sample basis can be provided.
[0022] According to this disclosure, a method and apparatus for generating / receiving media files can be provided, wherein, when random access occurs, the current sample is constrained to include IRAP sub-screens.
[0023] According to this disclosure, a method and apparatus for generating / receiving media files can be provided, wherein, when random access occurs, the TemporalId value of the current sample is constrained to 0.
[0024] According to this disclosure, a method for sending media files generated by a media file generation method or device according to this disclosure can be provided.
[0025] According to this disclosure, a recording medium may be provided for storing media files generated by a media file generation method or apparatus according to this disclosure.
[0026] According to this disclosure, a recording medium can be provided for storing media files received by a media file receiving device according to this disclosure and used for reconstructing images.
[0027] Those skilled in the art will understand that the effects achievable through this disclosure are not limited to those specifically described above, and that other advantages of this disclosure will become clearer from the detailed description. Attached Figure Description
[0028] Figure 1 This is a diagram that schematically illustrates a media file sending / receiving system according to an embodiment of the present disclosure.
[0029] Figure 2 This is a flowchart illustrating a method for sending media files.
[0030] Figure 3 This is a flowchart illustrating a method for receiving media files.
[0031] Figure 4 This is a diagram illustrating an image encoding device according to one embodiment of the present disclosure.
[0032] Figure 5 This is a diagram illustrating an image decoding device according to one embodiment of the present disclosure.
[0033] Figure 6 This is a diagram illustrating an example of a layered structure for an encoded image / video.
[0034] Figure 7 This is a diagram illustrating an example of a media file structure.
[0035] Figure 8 This is an example Figure 7 A diagram illustrating an example of a trak box structure.
[0036] Figure 9 This is a diagram illustrating an example of image signal structure.
[0037] Figure 10 This is a diagram illustrating an example of a sub-screen with mixed NAL units in a sample based on an existing VVC file format.
[0038] Figure 11 This is a diagram illustrating an example of a sub-screen having a hybrid NAL unit in a sample according to one embodiment of the present disclosure.
[0039] Figure 12 This is a flowchart illustrating a method for determining sample characteristics according to one embodiment of the present disclosure.
[0040] Figure 13 This is a flowchart illustrating a method for receiving media files according to one embodiment of the present disclosure.
[0041] Figure 14This is a flowchart illustrating a method for generating a media file according to one embodiment of the present disclosure.
[0042] Figure 15 This is a diagram illustrating an embodiment of the present disclosure that can be applied to a content streaming system. Detailed Implementation
[0043] In the following, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0044] In describing this disclosure, detailed descriptions of relevant known functions or constructions will be omitted if they unnecessarily obscure the scope of this disclosure. In the accompanying drawings, portions irrelevant to the description of this disclosure are omitted, and similar reference numerals are assigned to similar portions.
[0045] In this disclosure, when a component is “connected,” “linked,” or “coupled” to another component, it may include not only a direct connection but also an indirect connection with intermediate components. Furthermore, when a component “comprises” or “has” other components, unless otherwise stated, it is intended to include, but not exclude, other components.
[0046] In this disclosure, the terms first, second, etc., are used only for the purpose of distinguishing one component from other components, and do not limit the order or importance of the components unless otherwise stated. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0047] In this disclosure, the distinguishing components are intended to clearly describe each feature and do not imply that the components must be separate. That is, multiple components may be integrated and implemented in a single hardware or software unit, or a single component may be distributed and implemented in multiple hardware or software units. Therefore, unless otherwise specified, implementations of these component integrations or distributions are included within the scope of this disclosure.
[0048] In this disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included within the scope of this disclosure. In addition, embodiments that include other components besides those described in the various embodiments are also included within the scope of this disclosure.
[0049] This disclosure relates to the encoding and decoding of images. Unless redefined in this disclosure, the terms used herein may have the general meaning commonly used in the art to which this disclosure pertains.
[0050] In this disclosure, a "picture" generally refers to a unit representing an image within a specific time period, and a slice / tile is a coding unit that constitutes part of a picture; a picture can consist of one or more slices / tiles. Additionally, a slice / tile may include one or more coding tree units (CTUs).
[0051] In this disclosure, "pixel" or "pixel" can refer to the smallest unit that constitutes a frame (or image). Furthermore, "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, or it can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0052] In this disclosure, "unit" can refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. In some cases, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "region." Generally, an M×N block may include a set (or array) of samples (or transform coefficients) with M columns and N rows.
[0053] In this disclosure, "current block" can mean one of "current coding block," "current coding unit," "coding target block," "decoding target block," or "processing target block." When performing prediction, "current block" can mean "current prediction block" or "prediction target block." When performing transform (inverse transform) / quantization (dequantization), "current block" can mean "current transform block" or "transform target block." When performing filtering, "current block" can mean "filter target block."
[0054] Furthermore, in this disclosure, unless explicitly stated as a chroma block, "current block" may mean a block that includes both luma component blocks and chroma component blocks, or "the luma block of the current block." The luma component block of the current block can be represented by an explicit description including terms such as "luma block" or "current luma block." Similarly, "the chroma component block of the current block" can be represented by an explicit description including terms such as "chroma block" or "current chroma block."
[0055] In this disclosure, the terms “ / ” or “,” can be interpreted as indicating “and / or”. For example, “A / B” and “A, B” can mean “A and / or B”. Furthermore, “A / B / C” and “A / B / C” can mean “at least one of A, B and / or C”.
[0056] In this disclosure, the term "or" should be interpreted to indicate "and / or". For example, the expression "A or B" can include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in this disclosure, "or" should be interpreted to indicate "additionally or alternatively".
[0057] Overview of Media File Sending / Receiving Systems
[0058] Figure 1 This is a schematic illustration of a media file sending / receiving system according to one embodiment of the present disclosure.
[0059] Reference Figure 1 The media file sending / receiving system 1 may include a sending device A and a receiving device B. In some embodiments, the media file sending / receiving system 1 may support adaptive streaming based on MPEG-DASH (HTTP Dynamic Adaptive Streaming), thereby supporting seamless media content reproduction.
[0060] The transmitting device A may include a video source 10, an encoder 20, an encapsulation unit 30, a transmitting processor 40, and a transmitter 45.
[0061] Video source 10 can generate or acquire media data such as video or images. For this purpose, video source 10 may include a video / image capturing device and / or a video / image generating device, or it may be connected to an external device to receive media data.
[0062] Encoder 20 can encode media data received from video source 10. Encoder 20 can perform a series of processes such as prediction, transformation, and quantization according to video codec standards for compression and coding efficiency (e.g., the Universal Video Coding (VVC) standard). Encoder 20 can output the encoded media data as a bitstream.
[0063] Encapsulation unit 30 can encapsulate encoded media data and / or media data-related metadata. For example, encapsulation unit 30 can encapsulate data in file formats (e.g., ISO BMFF or Common Media Application Format (CMAF)) or process data in segmented form. In some embodiments, media data encapsulated in the form of a file (hereinafter referred to as a "media file") can be stored in a storage unit (not shown). The media file stored in the storage unit can be read by the transmitting processor 40 and transmitted to the receiving device B according to an on-demand, non-real-time (NRT) or broadband method.
[0064] The transmitting processor 40 can generate an image signal by processing the media file according to any transmitting method. The media file transmitting method can include broadcasting and broadband methods.
[0065] Depending on the broadcast method, media files can be sent using either the MPEG Media Transfer (MMT) protocol or the One-Way Real-Time Object Transfer (ROUTE) protocol. The MMT protocol can be a transport protocol that supports media streaming regardless of file format or codec in an IP-based network environment. When using the MMT protocol, media files can be processed in a Media Processing Unit (MPU) based on MMT, and then sent according to the MMT protocol. The ROUTE protocol is an extension of One-Way File Transfer (FLUTE) and can be a transport protocol that supports real-time transmission of media files. When using the ROUTE protocol, media files can be processed into one or more segments based on MPEG-DASH, and then sent according to the ROUTE protocol.
[0066] According to the broadband approach, media files can be sent over a network using HTTP (Hypertext Transfer Protocol). Information sent via HTTP can include signaling metadata, segmentation information, and / or non-real-time (NRT) service information.
[0067] In some implementations, the sending processor 40 may include an MPD generator 41 and a segment generator 42 to support adaptive media streaming.
[0068] MPD generator 41 can generate a Media Presentation Description (MPD) based on a media file. An MPD is a file that includes detailed information about the media presentation and can be expressed in XML format. The MPD can provide signaling metadata such as identifiers for each segment. In this case, receiving device B can dynamically obtain segments based on the MPD.
[0069] Segment generator 42 can generate one or more segments based on a media file. Segments may include actual media data and may have a file format such as ISO BMFF. Segments may be included in a representation of the image signal, and as described above, segments can be identified based on MPD.
[0070] In addition, the transmitting processor 40 can generate image signals according to the MPEG-DASH standard based on the generated MPD and segments.
[0071] Transmitter 45 can send the generated image signal to receiving device B. In some embodiments, transmitter 45 can send the image signal to receiving device B via an IP network according to the MMT standard or the MPEG-DASH standard. According to the MMT standard, the image signal sent to receiving device B may include a Presentation Information Document (PI) that includes reproduction information of media data. According to the MPEG-DASH standard, the image signal sent to receiving device B may include the aforementioned MPD as reproduction information of media data. However, in some embodiments, the MPD and segments may be sent to receiving device B separately. For example, a first image signal including the MPD may be generated by transmitting device A or an external server and sent to receiving device B, while a second image signal including segments may be generated by transmitting device A and sent to receiving device B.
[0072] Furthermore, despite Figure 1 The transmitting processor 40 and the transmitter 45 are illustrated as separate elements, but in some embodiments, they can be implemented as a single, integrated element. Furthermore, the transmitting processor 40 can be implemented as an external device (e.g., a DASH server) separate from the transmitting device A. In this case, the transmitting device A can operate as a source device that generates media files by encoding media data, and the external device can operate as a server device that generates image signals by processing media data according to any transmission protocol.
[0073] Next, receiving device B may include receiver 55, receiving processor 60, decapsulation unit 70, decoder 80, and renderer 90. In some embodiments, receiving device B may be an MPEG-DASH based client.
[0074] Receiver 55 can receive image signals from transmitting device A. Image signals according to the MMT standard may include PI documents and media files. Additionally, image signals according to the MPEG-DASH standard may include MPDs and segments. In some implementations, MPDs and segments may be transmitted separately using different image signals.
[0075] The receiver processor 60 can extract / parse media files by processing the received image signals according to the transmission protocol.
[0076] In some implementations, the receiving processor 60 may include an MPD parsing unit 61 and a segmentation parsing unit 62 to support adaptive media stream transmission.
[0077] MPD parsing unit 61 can obtain MPD from the received image signal and parse the obtained MPD to generate commands required for segmentation. Furthermore, MPD parsing unit 61 can obtain media data reproduction information (e.g., color conversion information) based on the parsed MPD.
[0078] The segmentation parsing unit 62 can obtain segments based on the parsed MPD and parse the obtained segments to extract the media file. In some embodiments, the media file may have a file format such as ISO BMFF or CMAF.
[0079] The decapsulation unit 70 can decapsulate the extracted media file to obtain media data and associated metadata. The obtained metadata may be in the form of frames or tracks in a file format. In some embodiments, the decapsulation unit 70 may receive the metadata required for decapsulation from the MPD parsing unit 61.
[0080] Decoder 80 can decode the acquired media data according to a video codec standard (e.g., the VVC standard). To this end, decoder 80 can perform a series of processes such as prediction, inverse quantization, and inverse transform corresponding to the operations of encoder 20.
[0081] The renderer 90 can render media data such as decoded video or images. The rendered media data can be reproduced through a display unit (not shown).
[0082] The methods for sending / receiving media files will be described in detail below.
[0083] Figure 2 This is a flowchart illustrating a method for sending media files.
[0084] In one example Figure 2 Each step can be made by Figure 1 The transmitting device A performs this operation. Specifically, step S210 can be performed by... Figure 1 The encoder 20 performs the operation. Furthermore, steps S220 and S230 can be performed by the transmitting processor 40. Additionally, step S240 can be performed by the transmitter 45.
[0085] Reference Figure 2 The transmitting device can encode media data such as video or images (S210). The media data can be captured / generated by the transmitting device or obtained from an external device (e.g., a camera, video archive, etc.). The media data can be encoded in bitstream form according to a video codec standard (e.g., the VVC standard).
[0086] The transmitting device can generate an MPD and one or more segments based on the encoded media data (S220). As described above, the MPD may include detailed information about media presentation. The segments may include the actual media data. In some embodiments, the media data may be encapsulated in a file format such as ISO BMFF or CMAF and included in the segments.
[0087] The transmitting device can generate an image signal including the generated MPD and segments (S230). In some embodiments, the image signal can be generated separately for each of the MPD and segments. For example, the transmitting device can generate a first image signal including the MPD and a second image signal including the segments.
[0088] The transmitting device can send the generated image signal to the receiving device (S240). In some embodiments, the transmitting device can transmit the image signal using a broadcast method. In this case, an MMT protocol or a ROUTE protocol can be used. Alternatively, the transmitting device can transmit the image signal using a broadband method.
[0089] In addition, although Figure 2 In this embodiment, the MPD and the image signal including the MPD are described as being generated and transmitted by the transmitting device (steps S220 to S240). However, in some embodiments, the MPD and the image including the MPD may be generated and transmitted by an external server different from the transmitting device.
[0090] Figure 3 This is a flowchart illustrating a method for receiving media files.
[0091] In the example, Figure 3 Each step can be made by Figure 1 The receiving device B performs the operation. Specifically, step S310 can be performed by receiver 55. Furthermore, step S320 can be performed by receiving processor 60. Additionally, step S330 can be performed by decoder 80.
[0092] Reference Figure 3 The receiving device can receive image signals from the transmitting device (S310). Image signals according to the MPEG-DASH standard may include MPDs and segments. In some implementations, MPDs and segments can be received separately from different image signals. For example, they can be received from... Figure 1 The transmitting device or external server receives the first image signal, including the MPD, and can receive it from... Figure 1 The transmitting device receives a second image signal that includes segments.
[0093] The receiving device can extract the MPD and segments from the received image signal, and parse the extracted MPD and segments (S320). Specifically, the receiving device can parse the MPD to generate commands needed to obtain the segments. Then, the receiving device can obtain the segments based on the parsed MPD, and parse the obtained segments to obtain media data. In some embodiments, the receiving device can perform decapsulation on the media data in the file format to obtain media data from the segments.
[0094] The receiving device can decode media data such as acquired video or images (S330). The receiving device can perform a series of processes such as inverse quantization, inverse transform, and prediction to decode the media data. Then, the receiving device can render the decoded media data and reproduce the media data on a display.
[0095] The image encoding / decoding device will be described in detail below.
[0096] Overview of Image Encoding Devices
[0097] Figure 4 This is a diagram that schematically illustrates an image encoding device according to an embodiment of the present disclosure. Figure 4 Image encoding device 400 can be used with reference Figure 1 The encoder 20 of the described transmitting device A corresponds to this.
[0098] Reference Figure 4 The image encoding device 400 may include an image segmenter 410, a subtractor 415, a transformer 420, a quantizer 430, a dequantizer 440, an inverse transformer 450, an adder 455, a filter 460, a memory 470, an inter-frame prediction unit 480, an intra-frame prediction unit 485, and an entropy encoder 490. The inter-frame prediction unit 480 and the intra-frame prediction unit 485 may be collectively referred to as "predictors". The transformer 420, quantizer 430, dequantizer 440, and inverse transformer 450 may be included in a residual processor. The residual processor may also include a subtractor 415.
[0099] In some implementations, all or at least some of the components of the image encoding device 400 may be configured by a single hardware component (e.g., an encoder or a processor). Furthermore, the memory 470 may include a decoded screen buffer (DPB) and may be configured by a digital storage medium.
[0100] Image segmenter 410 can segment an input image (or picture or frame) input to image encoding device 400 into one or more processing units. For example, a processing unit may be called an encoding unit (CU). Encoding units can be obtained by recursively segmenting encoding tree units (CTUs) or maximum encoding units (LCUs) according to a quadtree / binary tree / tritree (QT / BT / TT) structure. For example, an encoding unit can be segmented into multiple encoding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the segmentation of encoding units, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary tree structure. The encoding process according to this disclosure can be performed based on the final encoding unit that is no longer segmented. The maximum encoding unit can be used as the final encoding unit, or a deeper encoding unit obtained by segmenting the maximum encoding unit can be used as the final encoding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). Prediction units and transform units can be partitioned or segmented from the final coding unit. Prediction units can be sample prediction units, and transform units can be units used to derive transform coefficients and / or units used to derive residual signals from transform coefficients.
[0101] The prediction unit (inter-frame prediction unit 480 or intra-frame prediction unit 485) can perform prediction on the block to be processed (the current block) and generate a prediction block that includes prediction samples of the current block. The prediction unit can determine whether to apply intra-frame prediction or inter-frame prediction to the current block or CU unit. The prediction unit can generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 490. The information about the prediction can be encoded in the entropy encoder 490 and output as a bitstream.
[0102] Intra-prediction unit 485 can predict the current block by referencing samples in the current frame. Depending on the intra-prediction mode and / or intra-prediction technique, the reference samples may be located among the neighbors of the current block or may be placed separately. Intra-prediction modes may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, directional modes may include, for example, 33 or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-prediction unit 485 can determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.
[0103] Inter-frame prediction unit 480 can deduce the prediction block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, dual prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc. The reference frame including the temporally neighboring block may be referred to as a juxtaposed frame (colPic). For example, inter-frame prediction unit 480 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to deduce the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame prediction unit 480 can use the motion information of neighboring blocks as the motion information of the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, the motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be signaled by encoding the motion vector difference and an indicator of the motion vector predictor. The motion vector difference can refer to the difference between the motion vector of the current block and the motion vector predictor.
[0104] The prediction unit can generate a prediction signal based on various prediction methods and techniques described below. For example, the prediction unit can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame prediction and inter-frame prediction simultaneously to predict the current block. A prediction method that simultaneously applies both intra-frame prediction and inter-frame prediction to predict the current block can be called Combined Intra-Frame and Inter-Frame Prediction (CIIP). Furthermore, the prediction unit can perform Intra-Frame Block Copy (IBC) to predict the current block. Intra-Frame Block Copy can be used for content image / video coding in games, such as Screen Content Coding (SCC). IBC is a method of predicting the current frame using a previously reconstructed reference block in the current frame at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current frame can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction because the reference block is derived within the current frame. That is, IBC can use at least one inter-frame prediction technique described in this disclosure.
[0105] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. Subtractor 415 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the prediction unit from the input image signal (original block or original sample array). The generated residual signal can be transmitted to converter 420.
[0106] Transformer 420 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented graphically. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size or to blocks of variable size instead of square.
[0107] Quantizer 430 quantizes the transform coefficients and transmits them to entropy encoder 490. Entropy encoder 490 encodes the quantized signal (information about the quantized transform coefficients) and outputs a bitstream. This information about the quantized transform coefficients can be referred to as residual information. Quantizer 430 can rearrange the block-type quantized transform coefficients into a one-dimensional vector based on the coefficient scan order and generate information about the quantized transform coefficients based on this one-dimensional vector form.
[0108] The entropy encoder 490 can perform various encoding methods (e.g., Exponential Columbus, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc.). The entropy encoder 490 can encode information required for video / image reconstruction (e.g., values of syntax elements, etc.) other than the quantization transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at Network Abstraction Layer (NAL) units. The video / image information may also include information about various parameter sets (e.g., Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS)). Furthermore, the video / image information may also include general constraint information. The signaled information, transmitted information, and / or syntax elements described in this disclosure can be encoded and included in the bitstream through the above encoding process.
[0109] The bitstream can be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting the signal output from the entropy encoder 490 and / or a storage unit (not shown) for storing the signal may be included as internal / external components of the image encoding device 400. Alternatively, a transmitter may be provided as a component of the entropy encoder 490.
[0110] The quantized transform coefficients output from quantizer 430 can be used to generate residual signals. For example, the residual signals (residual blocks or residual samples) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients using dequantizer 440 and inverse transformer 450.
[0111] Adder 455 adds the reconstructed residual signal to the prediction signal output from inter-frame prediction unit 480 or intra-frame prediction unit 485 to generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). If the block to be processed has no residual (e.g., in the case of applying skip mode), the predicted block can be used as a reconstructed block. Adder 455 may be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.
[0112] Furthermore, luminance mapping with chroma scaling (LMCS) is applicable during image encoding and / or reconstruction.
[0113] Filter 460 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 460 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 470, specifically in the DPB of memory 470. Various filtering methods can include, for example, deblocking filtering, sample adaptive offsetting, adaptive loop filtering, bilateral filtering, etc. Filter 460 can generate various filtering-related information and transmit the generated information to entropy encoder 490, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 490 and output as a bitstream.
[0114] The modified reconstructed frame transmitted to memory 470 can be used as a reference frame in inter-frame prediction unit 480. When inter-frame prediction is applied by image encoding device 400, prediction mismatch between image encoding device 400 and image decoding device can be avoided and coding efficiency can be improved.
[0115] The DPB of memory 470 can store modified reconstructed frames for use as reference frames in inter-frame prediction unit 480. Memory 470 can store motion information of blocks from which motion information in the current frame is derived (or encoded) and / or motion information of already reconstructed blocks in the frame. The stored motion information can be transmitted to inter-frame prediction unit 480 and used as motion information for spatially or temporally adjacent blocks. Memory 470 can store reconstructed samples of reconstructed blocks in the current frame and can transmit the reconstructed samples to intra-frame prediction unit 485.
[0116] Overview of image decoding devices
[0117] Figure 5 This is a diagram that schematically illustrates an image decoding device according to an embodiment of the present disclosure. Figure 5 Image decoding device 500 can be used with reference Figure 1 The decoder 80 of the described receiving device B corresponds to this.
[0118] Reference Figure 5 The image decoding device 500 may include an entropy decoder 510, a dequantizer 520, an inverse transformer 530, an adder 535, a filter 540, a memory 550, an inter-frame prediction unit 560, and an intra-frame prediction unit 565. The inter-frame prediction unit 560 and the intra-frame prediction unit 565 may be collectively referred to as "predictors". The dequantizer 520 and the inverse transformer 530 may be included in a residual processor.
[0119] According to an implementation, all or at least some of the components of the image decoding device 500 can be configured by hardware components (e.g., a decoder or a processor). Furthermore, the memory 550 may include a decoded screen buffer (DPB) or may be configured by a digital storage medium.
[0120] The image decoding device 500, having received a bitstream including video / image information, can perform operations related to... Figure 4 The image is reconstructed by processing corresponding to the processing performed by the image encoding device 100. For example, the image decoding device 500 can use a processing unit applied in the image encoding device to perform decoding. Therefore, the decoding processing unit can be, for example, an encoding unit. The encoding unit can be obtained by segmenting a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 500 can be reproduced by a reproduction device (not shown).
[0121] Image decoding device 500 can receive from Figure 4The image encoding device generates signals in the form of a bitstream. The received signals can be decoded by an entropy decoder 510. For example, the entropy decoder 510 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets (e.g., adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS)). Furthermore, the video / image information may also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements notified / received by signals described in this disclosure can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 510 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine the context model using information about the target syntax element, decoding information of neighboring blocks and the target block, or information about symbols / bins decoded in the previous stage, perform arithmetic decoding on the bins based on the determined context model by predicting the occurrence probability of the bins, and generate symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The prediction-related information in the information decoded by the entropy decoder 510 can be provided to the prediction units (inter-frame prediction unit 560 and intra-frame prediction unit 565), and the residual value of entropy decoding performed in the entropy decoder 510, i.e., the quantization transform coefficients and related parameter information, can be input to the dequantizer 520. In addition, the filtering information in the information decoded by the entropy decoder 510 can be provided to the filter 540. Furthermore, the receiver (not shown) for receiving signals output from the image encoding device can be further configured as an internal / external element of the image decoding device 500, or the receiver can be a component of the entropy decoder 510.
[0122] Furthermore, the image decoding apparatus according to this disclosure can be referred to as a video / image / screen decoding apparatus. The image decoding apparatus can be divided into an information decoder (video / image / screen information decoder) and a sample decoder (video / image / screen sample decoder). The information decoder may include an entropy decoder 510. The sample decoder may include at least one of a dequantizer 520, an inverse transformer 530, an adder 535, a filter 540, a memory 550, an inter-frame prediction unit 560, or an intra-frame prediction unit 565.
[0123] The dequantizer 520 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 520 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the image encoding device. The dequantizer 520 can obtain the transform coefficients by performing dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information).
[0124] The inverse transformer 530 can perform inverse transformation on the transform coefficients to obtain the residual signal (residual block, residual sample array).
[0125] The prediction unit can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The prediction unit can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from the entropy decoder 510, and can determine a specific intra-frame / inter-frame prediction mode (prediction technique).
[0126] Similar to that described in the prediction unit of the image coding device 100, the prediction unit can generate a prediction signal based on various prediction methods (techniques) described later.
[0127] Intra-prediction unit 565 can predict the current block by referring to samples in the current frame. The description of intra-prediction unit 485 also applies to intra-prediction unit 565.
[0128] The inter-frame prediction unit 560 can deduce the prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, dual prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, the inter-frame prediction unit 560 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.
[0129] Adder 535 generates a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including inter-frame prediction unit 560 and / or intra-frame prediction unit 565). If the block to be processed has no residual (e.g., in the case of applying skip mode), the prediction block can be used as a reconstruction block. The description of adder 155 also applies to adder 535. Adder 535 may be referred to as a reconstructor or reconstruction block generator. The generated reconstruction signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.
[0130] Furthermore, Luminance Mapping with Chroma Scaling (LMCS) is applicable during the image decoding process.
[0131] Filter 540 can improve the quality of subjective / objective images by applying filtering to the reconstructed signal. For example, filter 540 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 550, specifically in the DPB of memory 550. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0132] The (modified) reconstructed frame stored in the DPB of memory 550 can be used as a reference frame in inter-frame prediction unit 560. Memory 550 can store motion information of blocks from which motion information in the current frame is derived (or decoded) and / or motion information of already reconstructed blocks in the frame. The stored motion information can be transmitted to inter-frame prediction unit 560 to be used as motion information for spatially or temporally neighboring blocks. Memory 550 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to intra-frame prediction unit 565.
[0133] In this disclosure, the embodiments described in the filter 460, inter-frame prediction unit 480 and intra-frame prediction unit 485 of the image encoding device 400 can be equally or correspondingly applied to the filter 540, inter-frame prediction unit 560 and intra-frame prediction unit 565 of the image decoding device 500.
[0134] The quantizer of an encoding device derives the quantized transform coefficients by applying quantization to the transform coefficients, and the dequantizer of either the encoding or decoding device derives the transform coefficients by applying dequantization to the quantized transform coefficients. In video coding, the quantization rate can be changed, and the compression ratio can be adjusted using the changed quantization rate. From an implementation perspective, considering complexity, a quantization parameter (QP) can be used instead of the quantization rate directly. For example, a quantization parameter with integer values from 0 to 63 can be used, and each quantization parameter value can correspond to the actual quantization rate. Furthermore, the quantization parameter QP for the luma component (luma sample) can be set differently. Y Quantization parameter QP of chromaticity components (chromaticity samples) C .
[0135] During quantization, the transform coefficient C can be received as input and divided by the quantization rate Q. step Furthermore, the quantization transform coefficients C' can be derived from this. In this case, considering computational complexity, the quantization rate is multiplied by scaling to form an integer, and shift operations can be performed according to the values corresponding to the scaling values. Quantization scaling can be derived based on the product of the quantization rate and the scaling value. That is, quantization scaling can be derived from QP. In this case, the quantization transform coefficients C' can be derived from this by applying quantization scaling to the transform coefficients C.
[0136] Dequantization is the inverse of quantization, and the quantization transform coefficients C' can be multiplied by the quantization rate Q. step Therefore, the reconstructed transform coefficients C' are derived based on this. In this case, level scaling can be derived from the quantization parameters, and the level scaling can be applied to the quantized transform coefficients C', thereby deriving the reconstructed transform coefficients C'. Due to losses during the transform and / or quantization process, the reconstructed transform coefficients C' can be slightly different from the original transform coefficients C. Therefore, even the encoding device can perform dequantization in the same way as the decoding device.
[0137] Furthermore, adaptive frequency-weighted quantization (IFQ) can be applied, where the quantization intensity is adjusted according to the frequency. IFQ corresponds to methods that apply different quantization intensities based on frequency. In IFQ, a predefined quantization scaling matrix can be used to apply different quantization intensities according to the frequency. That is, the quantization / dequantization process described above can be further performed based on the quantization scaling matrix.
[0138] For example, different quantization scaling matrices can be used depending on the size of the current block and / or whether the prediction mode applied to generate the residual signal for the current block is inter-frame prediction or intra-frame prediction. The quantization scaling matrix can also be called the quantization matrix or the scaling matrix. The quantization scaling matrix can be predefined. Additionally, the frequency quantization scaling information for the quantization scaling matrix used for frequency adaptive scaling can be constructed / encoded by the encoding device and signaled to the decoding device. This frequency quantization scaling information can be called quantization scaling information. The frequency quantization scaling information can include scaling list data (scaling_list_data).
[0139] Based on the scaling list data, the quantization scaling matrix can be derived. Additionally, the frequency quantization scaling information may include presence flags indicating whether scaling list data exists. Alternatively, when the scaling list data is signaled at a higher level (e.g., SPS), it may also include information indicating whether the scaling list data has been modified at a lower level (e.g., PPS or tile group header).
[0140] Figure 6 This is a diagram illustrating an example of the layered structure for encoding images / videos.
[0141] Encoded images / videos are classified into a Video Coding Layer (VCL) for image / video decoding and processing, a lower-layer system for sending and storing encoded information, and a Network Abstraction Layer (NAL) that exists between the VCL and the lower-layer system and is responsible for network adaptation functions.
[0142] In VCL, VCL data that includes compressed image data (slice data) can be generated, or additional enhancement information (SEI) messages required for image decoding processing or parameter sets that include information such as picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS) can be generated.
[0143] In NAL, header information (NAL unit header) can be added to the raw byte sequence payload (RBSP) generated in VCL to generate NAL units. In this case, RBSP refers to the slice data, parameter set, and SEI message generated in VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.
[0144] like Figure 6As shown, NAL units can be classified into VCL NAL units and non-VCL NAL units based on the type of RBSP generated in the VCL. A VCL NAL unit can refer to a NAL unit that includes information about the image (slice data), while a non-VCL NAL unit can refer to a NAL unit that includes information required for decoding the image (parameter set or SEI message).
[0145] VCL NAL units and non-VCL NAL units can be appended with header information and transmitted over the network according to the data standard of the underlying system. For example, NAL units can be modified to have a data format with a predetermined standard (e.g., H.266 / VVC file format, RTP (Real-Time Transport Protocol), or TS (Transport Stream)) and transmitted over various networks.
[0146] As described above, within a NAL unit, the NAL unit type can be specified based on the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and signaled. For example, this can be broadly categorized into VCL NAL unit types and non-VCL NAL unit types based on whether the NAL unit includes image information (slice data). VCL NAL unit types can be further subdivided based on the characteristics / type of the image included in the VCL NAL unit, and non-VCL NAL unit types can be subdivided based on the type of parameter set.
[0147] The following is an example of the VCL NAL unit type based on the screen type.
[0148] - "IDR_W_RADL" and "IDR_N_LP": VCL NAL unit types for Instantaneous Decoding Refresh (IDR) frames, which are the types of IRAP (Intra-Frame Random Access Point) frames;
[0149] An IDR frame can be the first frame in the bitstream in decoding order or a frame following the first frame. A frame with a NAL unit type such as "IDR_W_RADL" can have one or more Random Access Decodeable Leading (RADL) frames associated with it. In contrast, a frame with a NAL unit type such as "IDR_N_LP" does not have any leading frames associated with it.
[0150] - "CRA_NUT": VCL NAL unit type for Pure Random Access (CRA) screens, which is the type for IRAP screens;
[0151] A CRA (Cross-Action Rendering) frame can be the first frame in the bitstream in decoding order or a frame following the first frame. A CRA frame can be associated with a RADL (Random Access Skip Precursor) frame or a RASL (Random Access Skip Precursor) frame.
[0152] - "GDR_NUT": VCL NAL unit type for Random Access Gradual Decoding Refresh (GDR) screen;
[0153] - "STSA_NUT": VCL NAL unit type for the Random Access Step-by-Step Time Sublayer Access (STSA) screen;
[0154] - "RADL_NUT": VCL NAL unit type of the RADL frame used as a lead frame;
[0155] - "RASL_NUT": VCL NAL unit type of the RASL screen used as a lead screen;
[0156] - "TRAIL_NUT": VCL NAL unit type for the rear view;
[0157] The rear view is a non-IRAP view, which can be in the output order after the IRAP view or GDR view associated with the rear view, and can also be in the decoding order after the IRAP view associated with the rear view.
[0158] Next, examples of non-VCL NAL unit types based on parameter set types are as follows.
[0159] - "DCI_NUT": Non-VCL NAL unit type including decoding capability information (DCI).
[0160] - "VPS_NUT": Includes non-VCL NAL unit types such as Video Parameter Set (VPS).
[0161] - "SPS_NUT": Non-VCL NAL unit type including Sequence Parameter Set (SPS).
[0162] - "PPS_NUT": Non-VCL NAL unit type including Picture Parameter Set (PPS).
[0163] - "PREFIX_APS_NUT", "SUFFIX_APS_NUT": Non-VCL NAL unit types including Adaptive Parameter Set (APS).
[0164] - "PH_NUT": Non-VCL NAL unit type including the screen header.
[0165] The NAL unit type described above can be identified by predefined syntax information (e.g., nal_unit_type) included in the NAL unit header.
[0166] Furthermore, in this disclosure, the image / video information encoded in bitstream form may include not only frame segmentation information, intra / inter-frame prediction information, residual information, and / or in-loop filtering information, but also slice header information, frame header information, APS information, PPS information, SPS information, VPS information, and / or DCI. Additionally, the encoded image / video information may also include general constraint information (GCI) and / or NAL unit header information. According to embodiments of this disclosure, the encoded image / video information can be encapsulated into a media file of a predetermined format (e.g., ISO BMFF) and transmitted to a receiving device.
[0167] Media files
[0168] Encoded image information can be configured (or formatted) based on a predetermined media file format to generate a media file. For example, encoded image information can be used to form a media file (segmentation) based on one or more NAL units / sample entries for the encoded image information.
[0169] Media files may include sample entries and tracks. In one example, a media file may include various records, and each record may include information related to the media file format or information related to the image. In one example, one or more NAL units may be stored in a configuration record (or decoder configuration record) field in the media file. Additionally, the media file may include operation point records and / or operation point group boxes. In this disclosure, a decoder configuration record supporting Multifunction Video Coding (VVC) may be referred to as a VVC decoder configuration record. Similarly, an operation point record supporting VVC may be referred to as a VVC operation point record.
[0170] In media file formats, the term "sample" can refer to all the data associated with a single time or a single element of any of the three sample arrays (Y, Cb, Cr) representing a picture. When used in the context of a track (media file format), "sample" can refer to all the data associated with a single time of the track. Here, time can correspond to decoding time or composition time. Furthermore, when used in the context of a picture (e.g., a luminance sample), "sample" can refer to a single element of any of the three sample arrays representing the picture.
[0171] Figure 7 This is a diagram illustrating an example of a media file structure.
[0172] As mentioned above, standardized media file formats can be defined for storing and transmitting media data such as audio, video, or images. In some implementations, media files may have a file format based on the ISO Basic Media File Format (ISOBMFF).
[0173] A media file may include one or more boxes. Here, a box can be a data block or object containing media data or metadata related to the media data. Within a media file, boxes can form a hierarchical structure. Therefore, a media file can have a format suitable for storing and / or transmitting large volumes of media data. Furthermore, a media file can have a structure that facilitates access to specific media data.
[0174] Reference Figure 7 The media file 700 may include the ftyp box 710, the moov box 720, the moof box 730, and the mdat box 740.
[0175] The ftyp box 710 may include information about the file type, file version, and / or compatibility of the media file 700. In some implementations, the ftyp box 710 may be located at the beginning of the media file 700.
[0176] The moov box 720 may include metadata describing the media data in the media file 700. In some implementations, the moov box 720 may reside at the top level of the metadata-related boxes. Furthermore, the moov box 720 may include header information of the media file 700. For example, the moov box 720 may include decoder configuration records as decoder configuration information.
[0177] The moov box 720 is a sub-box and may include the mvhd box 721, the trak box 722, and the mvex box 723.
[0178] The mvhd box 721 may include presentation-related information of the media data in the media file 700 (e.g., media creation time, modification time, cycle, etc.).
[0179] The trak box 722 can include metadata about the tracks of the media data. For example, the trak box 722 can include streaming-related information, rendering-related information, and / or access-related information for audio or video tracks. Multiple trak boxes 722 may exist depending on the number of tracks present in the media file 700. See below for further details. Figure 8 An example describing the structure of the trak box 722.
[0180] The mvex box 723 may include information about the presence of one or more movie clips in the media file 700. A movie clip may be a portion of media data obtained by partitioning media data in the media file 700. A movie clip may include one or more encoded frames. For example, a movie clip may include one or more groups of frames (GOPs), and each group of frames may include multiple encoded frames or frames. Movie clips may be stored in each of the mdat boxes 740-1 to 740-N (where N is an integer greater than or equal to 1).
[0181] Moof frames 730-1 to 730-N (where N is an integer greater than or equal to 1) may include metadata of the movie clip, i.e., mdat frames 740-1 to 740-N. In some implementations, moof frames 730-1 to 730-N may exist at the top level of the metadata-related frames of the movie clip.
[0182] mdat frames 740-1 to 740-N may include actual media data. Depending on the number of movie clips present in the media file 700, multiple mdat frames 740-1 to 740-N may exist. Each of mdat frames 740-1 to 740-N may include one or more audio or video samples. In one example, a sample may refer to an Access Unit (AU). When a decoder configuration record is stored in a sample entry, the decoder configuration record may include the size of a length field indicating the length of the Network Abstraction Layer (NAL) unit to which each sample belongs, as well as a set of parameters.
[0183] In some implementations, the media file 700 may be processed, stored, and / or transmitted in segments. Segments may include an initialization segment I_seg and a media segment M_seg.
[0184] The initialization segment I_seg can be an object-type data unit that includes initialization information for accessing the representation. The initialization segment I_seg can include the aforementioned ftyp box 710 and / or moov box 720.
[0185] Media segment M_seg can be an object-type data unit comprising time-divided media data of a streaming service. Media segment M_seg can include the aforementioned moof frames 730-1 to 730-N and mdat frames 740-1 to 740-N. Although Figure 7 Not shown in the figure, but the media segment M_seg may also include: a styp box containing information related to the segment type and an sidx box (optional) containing identification information of the sub-segments included in the media file 700.
[0186] Figure 8 This is an example Figure 7 A diagram illustrating an example of a trak box structure.
[0187] Reference Figure 8 The trak box 800 can include the tkhd box 810, the tref box 820, and the mdia box 830.
[0188] The tkhd box 810 is a track header box and may include header information (e.g., the creation / modification time of the corresponding track, track identifier, etc.) of the track indicated by the trak box 800 (hereinafter referred to as the "corresponding track").
[0189] The tref box 820 is a track reference box and may include reference information for the corresponding track (e.g., the track identifier of another track referenced by the corresponding track).
[0190] The mdia box 830 may include information and objects describing the media data in the corresponding track. In some implementations, the mdia box 830 may include a minf box 840 providing information about the media data. Furthermore, the minf box 840 may include a stbl box 850 including metadata for a sample of the media data.
[0191] stbl frame 850 is a sample table frame and can include information such as the position and time of samples in the track. The reader can determine the sample type, sample size and offset within the container based on the information provided by stbl frame 850, and locate the samples in the correct time order.
[0192] stbl frame 850 may include one or more sample entry frames 851 and 852. Sample entry frames 851 and 852 may provide various parameters for a specific sample. For example, a sample entry frame for a video sample may include the width, height, resolution, and / or frame count of the video sample. Additionally, a sample entry frame for an audio sample may include the channel count, channel layout, and / or sampling rate of the audio sample. In some embodiments, sample entry frames 851 and 852 may be included within a sample description frame (not shown) in stbl frame 850. The sample description frame may provide detailed information about the encoding type applied to the sample and any initialization information required for that encoding type.
[0193] Additionally, stbl box 850 may include one or more sample-to-group boxes 853 and 854 and one or more sample-to-group description boxes 855 and 856.
[0194] Sample-to-group boxes 853 and 854 can indicate the sample group to which a sample belongs. For example, sample-to-group boxes 853 and 854 can include a grouping type syntax element (e.g., `grouping_type`) indicating the type of sample group. Furthermore, sample-to-group boxes 853 and 854 can include one or more sample group entries. Sample group entries can include a sample count syntax element (e.g., `sample_count`) and a group description index syntax element (e.g., `group_description_index`). Here, the sample count syntax element can indicate the number of consecutive samples to which the corresponding group description index is applied. Sample groups can include Stream Access Point (SAP) sample groups, Random Access Recovery Point sample groups, etc., and their details will be described later.
[0195] Sample group description boxes 855 and 856 can provide a description of the sample group. For example, sample group description boxes 855 and 856 can include a grouping type syntax element (e.g., grouping_type). Sample group description boxes 855 and 856 can correspond to sample-to-group boxes 853 and 854 with the same grouping type syntax element value. Furthermore, sample group description boxes 855 and 856 can include one or more sample group description entries. Sample group description entries can include "spor" sample group description entries, "minp" sample group description entries, "roll" sample group description entries, etc.
[0196] As mentioned above Figure 7 and Figure 8 As described, media data can be encapsulated into media files according to file formats such as ISO BMFF. Additionally, the media files can be transmitted to the receiving device as image signals according to MMT or MPEG-DASH standards.
[0197] Figure 9 This is a diagram illustrating an example of image signal structure.
[0198] Reference Figure 9 The image signal conforms to the MPEG-DASH standard and may include MPD 910 and multiple representations 920-1 to 920-N.
[0199] MPD 910 is a file that includes detailed information about media presentation and can be expressed in XML format. MPD 910 may include information about multiple representations of 920-1 to 920-N (e.g., bitrate of streaming content, image resolution, frame rate, etc.) and information about the URLs of HTTP resources (e.g., initialization segments and media segments).
[0200] This means that each of the numbers from 920-1 to 920-N (where N is an integer greater than 1) can be divided into multiple segments S-1 to SK (where K is an integer greater than 1). Here, the multiple segments S-1 to SK can correspond to the above reference... Figure 7 The initialization segments and media segments are described. The Kth segment SK can represent the last movie segment in each of the representations 920-1 to 920-N. In some implementations, the number of segments S-1 to SK included in each of the representations 920-1 to 920-N (that is, the value of K) can be different from each other.
[0201] Each of segments S-1 to SK may include actual media data such as one or more video or image samples. The characteristics of the video or image samples included in each of segments S-1 to SK may be described by MPD 910.
[0202] Each of the segments S-1 to SK has a unique URL (Uniform Resource Locator) and can therefore be accessed and reconstructed independently.
[0203] In addition, three types of basic streams can be defined for storing VVC content. First, a video basic stream that does not include any parameter sets can be defined. In this case, all parameter sets can be stored in one or more sample entries. Second, a basic stream that can include parameter sets can be defined, and a video and parameter set basic stream that can include parameter sets stored in one or more sample entries can be defined. Third, a non-VCL basic stream that includes non-VCL NAL units synchronized with the basic stream carried in the video track can be defined. In this case, the non-VCL track may not include parameter sets from sample entries.
[0204] Carriage of sub-images within the track.
[0205] The VVC file format defines several types of tracks.
[0206] -VVC Track: A VVC track can be represented by including NAL units in samples and sample entries (possibly by referencing VVC tracks of other sub-layers that include the VVC bitstream, and possibly by referencing VVC sub-picture tracks). When a VVC track references a VVC sub-picture track, the VVC track can be called a VVC base track.
[0207] -VVC Non-VCL Track: Adaptive Parameter Set (APS), LMCS (Luminance Map with Chroma Scaling), or scaling list parameters carrying the ALF (Adaptive Loop Filter), as well as other non-VCL NAL units, can be stored in a separate track from the track containing VCL NAL units and transmitted through that track. VVC Non-VCL track can refer to this type of track.
[0208] - VVC Subpicture Track: A VVC subpicture track can contain a sequence of one or more VVC subpictures forming a rectangular region, or a sequence of one or more complete slices. Additionally, a sample of a VVC subpicture track can contain one or more complete subpictures consecutively in decoding order, or one or more complete slices forming a rectangular region consecutively in decoding order. VVC subpictures or slices included in any sample of a VVC subpicture track can be consecutively in decoding order.
[0209] Furthermore, VVC non-VCL tracks and VVC sub-picture tracks enable optimized delivery of VVC video in streaming applications. Each track can be carried in its own DASH representation. Additionally, for decoding and rendering subsets of tracks, the DASH representations containing subsets of VVC sub-picture tracks and the DASH representations containing non-VCL tracks can be requested segment by segment by the client. In this way, redundant transmission of APS and other non-VCL NAL units can be avoided.
[0210] Reconstruct PU from samples in the VVC track of the reference VVC sub-picture track.
[0211] VVC track samples can be resolved into picture units (PUs) that include the following NAL units.
[0212] -AUD NAL cell (if present in the sample); Access cell delimiter (AUD) NAL cell can be the first NAL cell in the sample.
[0213] - When the sample is the first sample in a sample sequence associated with the same sample entry: the parameter set and SENAL unit contained in the sample entry.
[0214] - When at least one NAL unit of type nal_unit_type equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31 (NAL units of this type cannot precede the first VCL NAL unit in the PU) exists in the sample: NAL units in the sample are up to and exclude the first NAL unit among these NAL units; otherwise, all NAL units in the sample are excluded.
[0215] - The contents of the time-aligned (by decoding time) sample resolved from each referenced VVC subpicture track; VVC subpicture tracks are resolved in the order of the VVC subpicture tracks referenced in the 'subp' track reference (when num_subpic_ref_idx is equal to 0 in the same group entry of the 'spor' sample group entry mapped to the sample) or in the order specified in the 'spor' sample group description entry mapped to the sample (when num_subpic_ref_idx is greater than 0 in the same group entry of the 'spor' sample group entry mapped to the sample); all DCI, OPI, VPS, SPS, PPS, AUD, PH, EOS, EOB and other access unit (AU) level or picture level non-VCL NAL units are excluded; track references can be resolved as described below.
[0216] Furthermore, when the referenced VVC sub-picture track is associated with a VVC non-VCL track, the resolved sample of the VVC sub-picture track may contain non-VCL NAL units (if any) of the time-aligned samples in the VVC non-VCL track.
[0217] - All NAL units in the sample whose nal_unit_type is equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31.
[0218] If num_subpic_ref_idx in the 'spor' sample group description entry mapped to the sample is equal to 0, then each track reference in the 'subp' box can be resolved as follows. Otherwise, each instance of the track reference subp_track_ref_idx in the 'spor' sample group description entry mapped to the sample can be resolved as follows.
[0219] Each sample from the VVC basic orbit resolved from the 'subp' orbit reference can form a rectangular region without holes (i.e., the rectangular region is completely covered by the sample) and without overlap (i.e., the samples in the rectangular region cover different regions without overlap).
[0220] In addition, if the track reference points to the track ID of the VVC sub-screen track, the track reference can be resolved into the VVC sub-screen track.
[0221] Otherwise (i.e., when the orbital reference points to the 'alte' orbital group), the orbital reference can be resolved to any orbital in the 'alte' orbital group, and when a particular orbital reference index value is resolved to a particular orbital in a previous sample, it will be resolved to any of the following in the current sample.
[0222] - Same specific track, or
[0223] - Any other orbitals in the same 'alte' orbital group that contain synchronized samples aligned with the current sample time.
[0224] VVC sub-picture tracks within the same 'alte' track group need to be independent of any other VVC sub-picture tracks referenced by the same VVC base track to avoid decoding mismatch, and therefore the following constraints can be applied.
[0225] - All VVC sub-screen tracks contain VVC sub-screens.
[0226] - Sub-screen boundaries are similar to screen boundaries.
[0227] If the reader selects a VVC sub-picture track that contains a set of sub-picture ID values that are either the initial selection or different from the previous selection, the following steps can be performed:
[0228] - The 'spor' sample group description entries can be examined to infer whether the PPS or SPS NAL unit needs to be changed; SPS changes are only possible at the start of CLVS.
[0229] - When the 'spor' sample group description entry indicates that the start code emulation prevention byte exists before or within the sub-screen ID in the NAL unit, derive the original byte sequence payload (RBSP) from the NAL unit (i.e., the start code emulation prevention byte can be removed); after the overriding in the next step, start code emulation prevention is performed again.
[0230] The reader can use the bit position and sub-screen ID length information in the 'spor' sample group entry to infer which bits are rewritten to update the sub-screen ID to the selected sub-screen ID.
[0231] - When initially selecting the sub-screen ID value of PPS or SPS, the reader needs to rewrite PPS or SPS using the selected sub-screen ID value in the reconstructed access unit (AU).
[0232] - When the sub-screen ID value of a PPS or SPS changes compared to a previous PPS or SPS (respectively) with the same PPS ID value or SPS ID value, the reader needs to include copies of the previous PPS and SPS (if no PPS or SPS with the same PPS ID value or SPS ID value exists in the access unit AU, respectively); in addition, the reader needs to rewrite the PPS or SPS (respectively) using the updated sub-screen ID value in the reconstructed access unit (AU).
[0233] When there is a 'minp' sample group description entry for a sample that is mapped to the VVC basic orbit, the following operations can be applied.
[0234] - You can examine the 'minp' sample group description entry to infer the value of pps_mixed_nalu_types_in_pic_flag.
[0235] - If the derived value is different from the value of the previous PPS NAL cell with the same PPS ID in the reconstructed bitstream, then the following can be applied.
[0236] - When the PPS is not included in the picture unit through the above steps, the reader needs to include a copy of the PPS with the updated pps_mixed_nalu_types_in_pic_flag value in the reconstructed picture unit (PU).
[0237] - The reader can use the bit positions in the 'minp' sample group entry to infer which bit was overwritten to update pps_mixed_nalu_types_in_pic_flag.
[0238] Stream access point sample group
[0239] Flow Access Point (SAP) sample groups can be used to provide information about all SAPs. In the following disclosure, Flow Access Point Sample Groups will be abbreviated as 'sap' sample groups. 'sap' sample groups can be defined in standard documents such as ISO / IEC 14496-12. Specific examples of the syntax for specifying the grouping type of 'sap' sample groups, `grouping_type_parameter`, are shown in Table 1 below.
[0240] [Table 1]
[0241]
[0242] Referring to Table 1, the syntax grouping_type_parameter can include the syntax elements target_layers and layer_id_method_idc.
[0243] The syntax element `target_layers` can specify the target layers for a specific SAP system. The semantics of `target_layers` can be determined based on the value of the syntax element `layer_id_method_idc`. For example, when `layer_id_method_idc` is 0, `target_layers` can be retained.
[0244] The syntax element `layer_id_method_idc` can specify the semantics of `target_layers`. A `layer_id_method_idc` value of 0 specifies that the target layer consists of all layers represented by tracks. In contrast, the semantics of a non-zero `layer_id_method_idc` can be specified by the obtained media file specification.
[0245] When layer_id_method_idc equals 0, SAP can interpret it as follows:
[0246] - If the sample entry type is 'vvc1' or 'vvil' and the track does not contain any sub-layers with TemporalId of 0, SAP can specify access to all sub-layers present in the track.
[0247] - Alternatively, SAP can specify access to all layers present in the track.
[0248] For example, when the sample entry type is 'vvc1' or 'vvil' and the track does not contain any sub-layers with TemporalId of 0, the STSA screen with the lowest TemporalId present in the track can act as an SAP.
[0249] The semantics of layer_id_method_idc equaling 1 can be defined in standard documents such as ISO / IEC 14496-15.
[0250] Gradual Decoding Refresh (GDR) frames in a VVC bitstream are typically indicated by SAP type 4 in the 'sap' sample group. The VVC standard can support subframes with different VCL NAL unit types within the same encoded frame. GDR can be obtained by updating the subframes indexed by each subframe to IRAP subframes within the frame range. However, the VVC standard does not specify the decoding process to begin with a frame that has mixed VCL NAL unit types.
[0251] The characteristics of samples within a media file can be defined as follows.
[0252] -Condition 1: The set of picture parameters (PPS) in the VVC track where the sample reference pps_mixed_nalu_types_in_pic_flag is equal to 1 (i.e., each picture of the reference PPS has a mixed NAL unit type).
[0253] -Condition 2: For each subpick index i in the range from 0 to sps_num_subpics_minus1, all of the following sub-conditions are satisfied.
[0254] 2-1) sps_subpic_treated_as_pic_flag[i] equals 1 (that is, the i-th subpic is treated as a picture).
[0255] 2-2) At least one intra-frame random access point (IRAP) subframe with the same subframe index i exists in or after the current sample within the same coding layer video sequence (CLVS).
[0256] When all of the above conditions are met, the following sample characteristics will be applied to the sample.
[0257] - Sample characteristic 1: The sample can be indicated as a Type 4 SAP sample; here, a Type 4 SAP sample can contain a GDR image with ph_recovery_poc_cnt greater than 0.
[0258] - Sample Feature 2: Samples can be mapped to 'roll' sample group description entries with the following roll_distance value, which is correct for decoding processes that omit decoding of sub-screens with specific sub-screen indices before the existence of IRAP sub-screens.
[0259] When using the above 'sap' sample group, the 'sap' sample group will be used on all tracks carrying the same VVC bitstream.
[0260] Random access recovery point sample group
[0261] Random access recovery point sample sets can be used to provide information about recovery points for Gradual Decode Refresh (GDR). In the following, and throughout this disclosure, random access recovery point sample sets will be abbreviated as 'roll' sample sets.
[0262] When 'roll' sample groups are used with VVC tracks, the syntax and semantics of the grouping_type_parameter can be defined in standard documents such as ISO / IEC 14496-12. Specific examples are shown in Table 1 above.
[0263] When the target layer of a sample mapped to the 'roll' sample group is a GDR image, layer_id_method_idc can be set to 0 or 1.
[0264] When `layer_id_method_idc` equals 0, the 'roll' sample group can specify the behavior of all layers present in the track. Furthermore, the semantics of `layer_id_method_idc` equaling 1 can be defined, for example, in standard documents such as ISO / IEC 14496-15. For instance, when `layer_id_method_idc` equals 1, each bit in the `target_layers` field can specify the layers carried in the track. Since the field is only 28 bits long, the indication of SAPs within a track can be constrained to a maximum of 28 layers. Each bit of the field, starting from the least significant bit (LSB), is mapped in ascending order of the `layer_id` values in the list of `layer_id` values associated with the sample.
[0265] Alternatively, when all frames in the target layer of a sample mapped to the 'roll' sample group are not GDR frames, a layer_id_method_idc of 2 or 3 can be used. In this case, the following frame characteristics can be applied to frames in the target layer other than GDR frames.
[0266] - Picture characteristic 1: The referenced PPS has pps_mixed_nalu_types_in_pic_flag equal to 1 (that is, each picture of the referenced PPS has a mixed NAL unit type).
[0267] -Frame Feature 2: For each sub-frame index i in the range from 0 to sps_num_subpics_minus1,
[0268] 2-1) sps_subpic_treated_as_pic_flag[i] equals 1 (that is, the i-th subpic is treated as a picture).
[0269] 2-2) There is at least one IRAP subframe with the same subframe index i in the current sample within the same coding layer video sequence (CLVS) or after the current sample.
[0270] When layer_id_method_idc equals 2, the 'roll' sample group can specify the behavior of all layers present in the track. Furthermore, the semantics of layer_id_method_idc equaling 3 can be specified in standard documents such as ISO / IEC 14496-15.
[0271] When the reader starts decoding using samples marked with layer_id_method_idc equal to 2 or 3, the reader needs to modify the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), and Picture Header (PH) NAL units of the reconstructed bitstream as follows.
[0272] - Any SPS referenced by the sample has a sps_gdr_enabled_flag equal to 1 (i.e., GDR screens can be enabled and exist in CLVS).
[0273] - Any PPS referenced by the sample has pps_mixed_nalu_types_in_pic_flag equal to 0 (i.e., each frame of the referenced PPS does not have a mixed NAL unit type).
[0274] - All VCL NAL units of the access unit reconstructed from the sample have a nal_unit_type equal to GDR_NUT.
[0275] - Any frame header from an access unit reconstructed from a sample has a ph_gdr_pic_flag equal to 1 (i.e., the current frame is a GDR frame) and a ph_recovery_poc_cnt value corresponding to the roll_distance of the 'roll' sample group description entry to which the sample is mapped. Here, ph_recovery_poc_cnt specifies the recovery point of the decoded frames in the output order.
[0276] Based on the above modifications, a bitstream starting with a sample that is marked as belonging to a sample group with a layer_id_method_idc of equal to 2 or 3 can satisfy bitstream consistency.
[0277] When the 'roll' sample group is associated with a dependency layer but not with a reference layer, the sample group can indicate the features applied when all reference layers of the dependency layer are available and being decoded. The sample group can be used to initiate the decoding of the prediction layer.
[0278] Problems with existing technology
[0279] The VVC file format allows samples containing frames with mixed NAL units to be mapped to 'roll' sample groups. Details are defined below.
[0280] -Condition 1: The sample reference in the VVC track is a PPS with pps_mixed_nalu_types_in_pic_flag equal to 1 (i.e., each frame of the reference PPS has a mixed NAL unit type).
[0281] -Condition 2: For each subpick index i in the range from 0 to sps_num_subpics_minus1, all of the following sub-conditions are satisfied.
[0282] 2-1) sps_subpic_treated_as_pic_flag[i] equals 1 (that is, the i-th subpic is treated as a picture).
[0283] 2-2) At least one IRAP subframe with the same subframe index i exists in the current sample within the same coding layer video sequence (CLVS) or after the current sample.
[0284] When all the conditions are met, the following sample characteristics can be applied to the sample.
[0285] - Sample characteristic 1: The sample can be indicated as a Type 4 Flow Access Point (SAP) sample.
[0286] - Sample Feature 2: Samples can be mapped to 'roll' sample group description entries with the following roll_distance value, which is correct for decoding processes that omit decoding of sub-screens with specific sub-screen indices before the existence of IRAP sub-screens.
[0287] Furthermore, when all frames in the target layer of a sample mapped to the 'roll' sample group are not GDR frames, a layer_id_method_idc of 2 or 3 can be used. In this case, the following frame characteristics can be applied to frames in the target layer other than GDR frames.
[0288] - Picture characteristic 1: The referenced PPS has pps_mixed_nalu_types_in_pic_flag equal to 1 (that is, each picture of the referenced PPS has a mixed NAL unit type).
[0289] -Frame Feature 2: For each sub-frame index i in the range from 0 to sps_num_subpics_minus1,
[0290] 2-1) sps_subpic_treated_as_pic_flag[i] equals 1 (that is, the i-th subpic is treated as a picture).
[0291] 2-2) There is at least one IRAP subframe with the same subframe index i in the current sample within the same coding layer video sequence (CLVS) or after the current sample.
[0292] Figure 10 A specific example of a sub-picture based on picture characteristics is shown.
[0293] Figure 10 This is a diagram illustrating an example of a sub-screen with mixed NAL units in a sample based on an existing VVC file format.
[0294] Reference Figure 10 The mdat box 1000 can include the first sample Sample_0 through the fourth sample Sample_3. However, this is merely an example for ease of description, and therefore... Figure 10 As shown, the first sample (Sample_0) through the fourth sample (Sample_3) can be included in two or more mdat boxes.
[0295] The first samples (Sample_0) through the fourth samples (Sample_3) can constitute a coded layer video sequence (CLVS). Additionally, each of the first samples (Sample_0) through the fourth samples (Sample_3) can indicate an access unit (AU). In the following text, it is assumed that the first samples (Sample_0) through the fourth samples (Sample_3) are mapped to a 'roll' sample group, and that the layer_id_method_idc for the first samples (Sample_0) through the fourth samples (Sample_3) is 2 or 3 (i.e., all frames in the target layer included in the first samples (Sample_0) through the fourth samples (Sample_3) are not GDR frames).
[0296] The first sample (Sample_0) to the fourth sample (Sample_3) can each include a first sub-pic (Subpic_0) to a fourth sub-pic (Subpic_3). Based on the above assumption, since each picture including the first sub-pic (Subpic_0) to the fourth sub-pic (Subpic_3) is not a GDR picture, the above-mentioned picture characteristic 2 can be applied to each picture.
[0297] Specifically, based on the above-mentioned screen characteristic 2-1, each sub-screen in each screen from the first sub-screen Subpic_0 to the fourth sub-screen Subpic_3 can be regarded as a screen.
[0298] Furthermore, according to the screen characteristic 2-2 described above, for each sub-screen (Subpic_0 to Subpic_3) in CLVS (i.e., for each sub-screen with the same sub-screen index), at least one sub-screen can be an IRAP sub-screen. For example, at least one of the first sub-screen (Subpic_0), at least one of the second sub-screen (Subpic_1), at least one of the third sub-screen (Subpic_2), and at least one of the fourth sub-screen (Subpic_3) included in the first sample (Sample_0) to the fourth sample (Sample_3) can be an IRAP sub-screen. Therefore, as Figure 10 As shown, it is possible for only the fourth sample, Sample_3, to include an IRAP sub-screen. In this case, since the first sample, Sample_0, to the third sample, Sample_2, only include non-IRAP sub-screens, random access to the first sample, Sample_0, to the third sample, Sample_2 cannot be performed correctly.
[0299] Therefore, according to the existing VVC file format, a picture with mixed NAL unit types can be mapped to a 'roll' sample group regardless of the type of the NAL units in the picture. For example, a sample with a mixed NAL unit type of RASL_NUT and RADL_NUT (i.e., only including non-IRAP subpicks) can also be mapped to a 'roll' sample group. However, to support random access, the file format should be designed to only allow pictures with mixed NAL unit types, and at least one of the NAL unit types should be an IRAP type mapped to the 'roll' sample group (e.g., CRA_NUT, IDR_W_RADL, or IDR_N_LP). Therefore, random access to samples without IRAP types cannot be performed as correctly as samples with mixed NAL unit types RASL_NUT and RADL_NUT in a 'roll' sample group. Additionally, this problem may also occur for type 4 'sap' sample groups.
[0300] To address the aforementioned issues, according to embodiments of this disclosure, the current sample may include at least one IRAP sub-screen or may have a TemporalId equal to 0.
[0301] Embodiments of this disclosure may include at least one of the following aspects. According to the embodiments, each aspect may be implemented individually or in combination of both or more.
[0302] (Aspect 1): Samples mapped to the 'roll' sample group / frames that are not GDR frames will have at least one NAL unit of type IRAP. Here, the IRAP type may include CRA_NUT, IDR_W_RADL, and IDR_N_LP.
[0303] In other words, a sample mapped to the 'roll' sample group / a non-GDR image should meet at least all of the following conditions.
[0304] - Condition 1: The referenced PPS has pps_mixed_nalu_types_in_pic_flag equal to 1 (i.e., each frame of the referenced PPS has a mixed NAL unit type).
[0305] -Condition 2: For each subpick index i in the range from 0 to sps_num_subpics_minus1, sps_subpic_treated_as_pic_flag[i] is equal to 1 (i.e., the i-th subpick is treated as a picture).
[0306] -Condition 3: There is at least one IRAP sub-screen.
[0307] (Aspect 2) When the i-th sub-picture in the current picture is not an IRAP sub-picture, if the current picture is mapped to the 'roll' sample group and is not a GDR picture, the i-th sub-picture in the CLVS in the picture after the current picture according to the decoding order can be an IRAP sub-picture.
[0308] (Aspect 3): A GDR frame that is not mapped to the 'roll' sample group should be a frame whose temporal identifier (i.e., TemporalId) is equal to 0. In other words, the nuh_temporal_id_plus1 value of all NAL cells in the frame should be equal to 1.
[0309] The embodiments of this disclosure based on the above aspects will be described in detail below.
[0310] Implementation Method 1
[0311] Embodiment 1 of this disclosure can be provided based on aspects 1 and 2 above.
[0312] According to Implementation 1, samples mapped to the stream access point ('sap') sample group or the random access point ('roll') sample group can have predetermined sample characteristics. Details are as follows.
[0313] (1) Sample group of flow access points
[0314] The sample group of Stream Access Points ('sap') according to Implementation Method 1 can be used to provide information about all SAPs. The basic content of the 'sap' sample group is as described above, and in the following text, the differences from the existing VVC file format will be emphasized.
[0315] According to Implementation Method 1, the characteristics of samples within a media file can be defined as follows.
[0316] -Condition 1: The set of picture parameters (PPS) in the VVC track whose sample reference pps_mixed_nalu_types_in_pic_flag is equal to 1 (i.e., each picture of the reference PPS has a mixed NAL unit type).
[0317] -Condition 2: For each subpick index i in the range from 0 to sps_num_subpics_minus1, all of the following sub-conditions are satisfied.
[0318] 2-1) sps_subpic_treated_as_pic_flag[i] equals 1 (that is, the i-th subpic is treated as a picture).
[0319] 2-2) There is at least one IRAP sub-screen in the current sample.
[0320] 2-3) For each of the sub-pictures that are not IRAP sub-pictures in the current sample, at least one IRAP sub-picture exists in a sample that belongs to the same CLVS as the current sample and is in the decoding order after the current sample.
[0321] When all the above conditions are met, the following sample characteristics can be applied to the sample.
[0322] - Sample characteristic 1: The sample can be indicated as a Type 4 Flow Access Point (SAP) sample; here, a Type 4 SAP sample can include a GDR screen with ph_recovery_poc_cnt greater than 0.
[0323] - Sample Feature 2: Samples can be mapped to 'roll' sample group description entries with the following roll_distance value, which is correct for decoding processes that omit decoding of sub-screens with specific sub-screen indices before the existence of IRAP sub-screens.
[0324] Here, applying sample characteristics to a sample can mean that a sample can be mapped to a 'sap' sample group.
[0325] Furthermore, according to one implementation, applying sample characteristics to a sample can mean that all of the above conditions can be applied as image characteristics to samples mapped to the 'sap' sample group.
[0326] According to Implementation 1, in samples mapped to the 'sap' sample group of type 4, the current sample may include at least one IRAP sub-frame. Additionally, when the current sample includes a non-IRAP sub-frame, samples within the same CLVS that follow the current sample in decoding order may include IRAP sub-frames with the same sub-frame index as the non-IRAP sub-frames. Therefore, unlike existing VVC file formats, samples of the hybrid NAL unit type with only IRAP type can be mapped to the 'sap' sample group of type 4.
[0327] (2) Random access recovery point sample group
[0328] The random access recovery point ('roll') sample set according to Implementation Method 1 can be used to provide information about recovery points in GDR. The basic content of the 'roll' sample set is as described above, and in the following text, the differences from the existing VVC file format will be emphasized.
[0329] When all frames in the target layer of a sample included in the 'roll' sample group are not GDR frames, a layer_id_method_idc of 2 or 3 can be used. In this case, according to Implementation 1, the following frame characteristics can be applied to frames in the target layer other than GDR frames.
[0330] - Picture characteristic 1: The referenced PPS has pps_mixed_nalu_types_in_pic_flag equal to 1 (that is, each picture of the referenced PPS has a mixed NAL unit type).
[0331] -Frame Feature 2: For each sub-frame index i in the range from 0 to sps_num_subpics_minus1,
[0332] 2-1) sps_subpic_treated_as_pic_flag[i] equals 1 (that is, the i-th subpic is treated as a picture).
[0333] 2-2) There is at least one IRAP sub-screen in the current sample.
[0334] 2-3) For each of the sub-pictures that are not IRAP sub-pictures in the current sample, at least one IRAP sub-picture exists in a sample that belongs to the same CLVS as the current sample and is in the decoding order after the current sample.
[0335] According to one implementation, a frame characteristic can refer to a condition used to map a specific sample (e.g., a sample in the target layer that does not include GDR frames) to a 'roll' sample group. For example, a specific sample can be mapped to a 'roll' sample group when every frame in the specific sample satisfies all the frame characteristics. In contrast, if every frame in the specific sample does not satisfy at least one of the frame characteristics, the specific sample may not be mapped to a 'roll' sample group.
[0336] Figure 11 A specific example of a sub-picture based on picture characteristics is shown.
[0337] Figure 11 This is a diagram illustrating an example of a sub-screen having a hybrid NAL unit in a sample according to one embodiment of the present disclosure.
[0338] Reference Figure 11 The mdat box 1000 can include the first sample Sample_0 through the fourth sample Sample_3. In the following text, it is assumed that the first sample Sample_0 through the fourth sample Sample_3 are mapped to the 'roll' sample group, and the layer_id_method_idc for the first sample Sample_0 through the fourth sample Sample_3 is equal to 2 or 3 (i.e., all frames included in the target layer of the first sample Sample_0 through the fourth sample Sample_3 are not GDR frames).
[0339] The first sample (Sample_0) to the fourth sample (Sample_3) can each include a first sub-pic (Subpic_0) to a fourth sub-pic (Subpic_3). Based on the above assumption, since each picture including the first sub-pic (Subpic_0) to the fourth sub-pic (Subpic_3) is not a GDR picture, the above-mentioned picture characteristic 2 can be applied to each picture.
[0340] Specifically, based on the above-mentioned screen characteristics 2-1, each sub-screen in each screen, from the first sub-screen Subpic_0 to the fourth sub-screen Subpic_3, can be considered as a screen.
[0341] In addition, according to the above-mentioned image characteristic 2-2, the first sub-image Subpic_0 in the first sample Sample_0, which is the current sample, can be an IRAP sub-image.
[0342] Furthermore, since the second sub-pic_1 to the fourth sub-pic_3 in the first sample Sample_0 are all non-IRAP sub-pics, the aforementioned picture characteristics 2-3 can be applied to the second sample Sample_1 to the fourth sample Sample_3. Therefore, the second sub-pic_1 in the second sample Sample_1 can be an IRAP sub-pic. Additionally, the third sub-pic_2 in the third sample Sample_2 can be an IRAP sub-pic. Furthermore, the fourth sub-pic_3 in the fourth sample Sample_3 can be an IRAP sub-pic.
[0343] Therefore, as Figure 11 As shown, each of the first sample (Sample_0) to the fourth sample (Sample_3) can include an IRAP sub-screen. Therefore, with Figure 10 Unlike other cases, random access to any sample Sample_0 through Sample_3 can be performed correctly.
[0344] As described above, according to Implementation 1, among samples in the target layer that do not include GDR frames (i.e., layer_id_method_idc = 2 or 3), the current sample may include at least one IRAP sub-frame. Additionally, when the current sample includes a non-IRAP sub-frame, samples within the same CLVS that follow the current sample in decoding order may include IRAP sub-frames with the same sub-frame index as the non-IRAP sub-frames. Therefore, unlike existing VVC file formats, samples of mixed NAL unit types with only IRAP type can be mapped to 'roll' sample groups.
[0345] The method for determining whether to apply the 'sap' / 'roll' sample feature according to Implementation 1 is as follows: Figure 12 As shown.
[0346] Figure 12 This is a flowchart illustrating a method for determining sample characteristics according to one embodiment of the present disclosure. Figure 12 Each step can be performed by a media file generating device and / or a media file receiving device. In the following description, the media file receiving device will be used as the basis for the description. Figure 12 Each step.
[0347] Reference Figure 12 The media file receiving device can determine whether the frame in the target sample has a mixed NAL unit type (S1210). In one example, this determination can be performed based on a predetermined flag (e.g., pps_mixed_nalu_types_in_pic_flag) in the PPS referenced by the frame. For example, when pps_mixed_nalu_types_in_pic_flag equals 1, the media file receiving device can determine that the frame has a mixed NAL unit type. In contrast, when pps_mixed_nalu_types_in_pic_flag equals 0, the media file receiving device can determine that the frame does not have a mixed NAL unit type. Furthermore, the above step S1210 can correspond to condition 1 related to the above-mentioned sample characteristics or frame characteristics 1.
[0348] If the image does not have a mixed NAL unit type (S1210 'No'), the media file receiving device may not apply the above 'sap' / 'roll' sample characteristics to the target sample (S1260).
[0349] In contrast, when the picture has a mixed NAL unit type ("Yes" in S1210), the media file receiving device can determine whether each subpicture included in the picture is considered a picture (S1220). In one example, this determination can be performed based on a predetermined flag in the SPS referenced by the picture (e.g., sps_subpic_treated_as_pic_flag). For example, when sps_subpic_treated_as_pic_flag is 1, the media file receiving device can determine that each subpicture i is considered a picture. In contrast, when sps_subpic_treated_as_pic_flag is 0, the media file receiving device can determine that each subpicture is not considered a picture.
[0350] If each sub-picture is not considered a picture ('No' in S1220), the media file receiving device may not apply the aforementioned 'sap' / 'roll' sample characteristics to the target sample (S1260).
[0351] In contrast, when each sub-picture is considered a picture ('yes' in S1220), the media file receiving device can determine whether the current sample includes at least one IRAP sub-picture (S1230).
[0352] If the current sample does not include at least one IRAP sub-screen (S1230 'No'), the media file receiving device may not apply the aforementioned 'sap' / 'roll' sample characteristics to the target sample (S1260).
[0353] In contrast, if the current sample includes at least one IRAP sub-picture ("Yes" in S1230), the media file receiving device can determine the sub-picture index i for the non-IRAP sub-picture present in the current sample, and whether the IRAP sub-picture with sub-picture index i exists in the target sample after the current sample in the same CLVS according to the decoding order (S1240).
[0354] If an IRAP sub-screen exists (is 'yes' in S1240), the media file receiving device can apply the aforementioned 'sap' / 'roll' sample characteristics to the target sample (S1250). In other words, the target sample can be mapped to the 'sap' / 'roll' sample group.
[0355] In contrast, if the IRAP sub-screen does not exist (No in S1240), the media file receiving device may not apply the aforementioned 'sap' / 'roll' sample characteristics to the target sample (S1260).
[0356] Furthermore, steps S1220 to S1240 described above can correspond to condition 2 related to the aforementioned sample characteristic or image characteristic 2. Specifically, step S1220 can correspond to condition 2-1 or image characteristic 2-1. Furthermore, step S1230 can correspond to condition 2-2 or image characteristic 2-2. Furthermore, step S1240 can correspond to condition 2-3 or image characteristic 2-3.
[0357] According to Embodiment 1 of this disclosure, samples of mixed NAL unit types with only IRAP type can be mapped to 'sap' or 'roll' sample groups. Therefore, random access on a sample-by-sample basis can be correctly performed.
[0358] Implementation Method 2
[0359] Implementation 2 of this disclosure can be provided based on aspect 3 above.
[0360] According to Implementation 2, samples mapped to the stream access point ('sap') sample group or the random access point ('roll') sample group can have predetermined sample characteristics. Details are as follows.
[0361] (1) Sample group of flow access points
[0362] The sample group of Stream Access Points ('sap') according to Implementation Method 2 can be used to provide information about all SAPs. The basic content of the 'sap' sample group is as described above, and in the following text, the differences from the existing VVC file format will be emphasized.
[0363] According to Implementation Method 2, the characteristics of the samples within the media file can be defined as follows.
[0364] -Condition 1: The sample reference pps_mixed_nalu_types_in_pic_flag in the VVC track is equal to 1 (i.e., each frame of the reference PPS has a mixed NAL unit type).
[0365] -Condition 2: For each subpick index i in the range from 0 to sps_num_subpics_minus1, all of the following sub-conditions are satisfied.
[0366] 2-1) sps_subpic_treated_as_pic_flag[i] is 1 (that is, the i-th subpic is treated as a picture).
[0367] 2-2) The current sample or belongs to the same CLVS as the current sample and there is at least one IRAP sub-screen with the same sub-screen index i in the samples after the current sample (i.e., for each sub-screen index i, there is at least one IRAP sub-screen in the samples belonging to the same CLVS).
[0368] 2-3) The TemporalID value of the current sample is equal to 0; here, TemporalID can refer to the time identifier of the NAL cell in the current sample; when the TemporalID value of the current sample is equal to 0, the current sample can include NAL cells of IRAP type such as CRA_NUT, IDR_W_RADL or IDR_N_LP.
[0369] When all the above conditions are met, the following sample characteristics can be applied to the sample.
[0370] - Sample characteristic 1: The sample can be indicated as a type 4 SAP sample; here, a type 4 SAP sample can include GDR screens with ph_recovery_poc_cnt greater than 0.
[0371] - Sample Feature 2: Samples can be mapped to 'roll' sample group description entries with the following roll_distance value, which is correct for decoding processes that omit decoding of sub-screens with specific sub-screen indices before the existence of IRAP sub-screens.
[0372] Here, applying sample characteristics to a sample means that a sample can be mapped to a 'sap' sample group.
[0373] Furthermore, according to one implementation, applying sample characteristics to a sample can mean that all of the above conditions can be applied as image characteristics to samples mapped to the 'sap' sample group.
[0374] According to implementation method 2, in samples mapped to the 'sap' sample group of type 4, the TemporalID value of the current sample can be equal to 0. In this case, the current sample includes an IRAP type NAL unit and therefore can include at least one IRAP sub-screen. Additionally, when the current sample includes a non-IRAP sub-screen, samples within the same CLVS that follow the current sample in decoding order can include IRAP sub-screens with the same sub-screen index as the non-IRAP sub-screens. Therefore, unlike existing VVC file formats, samples of the mixed NAL unit type with only IRAP type can be mapped to the 'sap' sample group of type 4. Thus, random access on a sample-by-sample basis can be correctly performed.
[0375] (2) Random access recovery point sample group
[0376] The random access recovery point ('roll') sample set according to Implementation Method 2 can be used to provide information about recovery points in GDR. The basic content of the 'roll' sample set is as described above, and in the following text, the differences from the existing VVC file format will be emphasized.
[0377] When all frames in the target layer of a sample mapped to the 'roll' sample group are not GDR frames, a layer_id_method_idc of 2 or 3 can be used. In this case, the following frame characteristics can be applied to frames in the target layer other than GDR frames.
[0378] - Picture characteristic 1: The referenced PPS has pps_mixed_nalu_types_in_pic_flag equal to 1 (that is, each picture of the referenced PPS has a mixed NAL unit type).
[0379] -Frame Feature 2: For each sub-frame index i in the range from 0 to sps_num_subpics_minus1,
[0380] 2-1) sps_subpic_treated_as_pic_flag[i] equals 1 (that is, the i-th subpic is treated as a picture).
[0381] 2-2) The current sample or belongs to the same CLVS as the current sample and there is at least one IRAP sub-screen with the same sub-screen index i in the samples after the current sample (i.e., for each sub-screen index i, there is at least one IRAP sub-screen in the samples belonging to the same CLVS).
[0382] 2-3) The TemporalID value of the current sample is equal to 0; here, TemporalID can refer to the time identifier of the NAL cell in the current sample; when the TemporalID value of the current sample is equal to 0, the current sample can include NAL cells of IRAP type such as CRA_NUT, IDR_W_RADL or IDR_N_LP.
[0383] According to one implementation, a frame characteristic can refer to a condition used to map a specific sample (e.g., a sample in the target layer that does not include GDR frames) to a 'roll' sample group. For example, a specific sample can be mapped to a 'roll' sample group when every frame in the specific sample satisfies all the frame characteristics. In contrast, if every frame in the specific sample does not satisfy at least one of the frame characteristics, the specific sample may not be mapped to a 'roll' sample group.
[0384] As described above, according to Implementation 2, in samples that do not include GDR frames in the target layer (i.e., layer_id_method_idc = 2 or 3), the TemporalID value of the current sample can be 0. Therefore, the current sample can include at least one IRAP sub-frame. Additionally, when the current sample includes a non-IRAP sub-frame, samples within the same CLVS that follow the current sample in decoding order can include IRAP sub-frames with the same sub-frame index as the non-IRAP sub-frames. Therefore, unlike existing VVC file formats, samples of mixed NAL unit types with only IRAP type can be mapped to 'roll' sample groups. Thus, random access on a sample-by-sample basis can be correctly performed.
[0385] The media file generation / receiving method according to one embodiment of the present disclosure will be described in detail below.
[0386] Figure 13 This is a flowchart illustrating a method for receiving media files according to one embodiment of the present disclosure. Figure 13 Each step can be performed by the media file receiving device. In one example, the media file receiving device can correspond to... Figure 1 Receiving device B.
[0387] Reference Figure 13 The media file receiving device can obtain one or more tracks and sample groups from the media file received by the self-media file generating / sending device (S1310). In one example, the media file may have a file format such as ISO Basic Media File Format (ISO BMFF) or Common Media Application Format (CMAF).
[0388] The media file receiving device can process video data in a media file by reconstructing samples included in the track based on a sample set (S1320). Here, video data processing may include a process of decapsulating the media file, a process of obtaining video data from the decapsulated media file, and a process of decoding the obtained video data according to a video codec standard (e.g., the VVC standard).
[0389] In one implementation, based on the existence of samples in the track that are mapped to a predetermined type (e.g., type 4) of a stream access point ('sap') sample group or a random access recovery point ('roll') sample group, the current sample in the mapped sample can be constrained to include at least one intra-frame random access point (IRAP) sub-picture.
[0390] Additionally, in one implementation, based on the presence of samples in the track that are mapped to a predetermined type (e.g., type 4) of a stream access point ('sap') sample group or a random access recovery point ('roll') sample group, when the current sample includes a non-IRAP subframe, samples belonging to the same coding layer video sequence (CLVS) as the current sample and following the current sample in decoding order can be constrained to include at least one IRAP subframe with the same subframe index value as the non-IRAP subframe.
[0391] In one implementation, each frame included in the mapping sample can be constrained to have a hybrid NAL unit type.
[0392] In one implementation, each sub-picture included in the mapping sample can be constrained to be considered a picture.
[0393] In the above implementation, the mapped sample may have only a hybrid NAL unit type of IRAP type.
[0394] In one implementation, all frames in the target layer of a sample mapped to a random access recovery point sample group may not be Gradually Decoded Refresh (GDR) frames (i.e., layer_id_method_idc = 2 or 3).
[0395] In one implementation, the TemporalID value of the current sample can be constrained to 0.
[0396] Figure 14 This is a flowchart illustrating a media file generation method according to one embodiment of the present disclosure. Figure 14 Each step can be performed by a media file generation device. In one example, the media file generation device can correspond to... Figure 1 Transmitting device A.
[0397] Reference Figure 14 The media file generation device can encode video data (S1410). In one example, the video data can be encoded according to a video codec standard (e.g., the VVC standard) through a prediction, transformation, and quantization process.
[0398] The media file generation device can generate one or more tracks and sample groups for the encoded video data (S1420).
[0399] The media file generation device can generate media files based on the generated tracks and sample sets (S1430). In one example, the media file can have a file format such as ISO Basic Media File Format (ISO BMFF) or Common Media Application Format (CMAF).
[0400] In one implementation, based on the existence of samples in the track that are mapped to a predetermined type (e.g., type 4) of a stream access point ('sap') sample group or a random access recovery point ('roll') sample group, the current sample in the mapped sample can be constrained to include at least one intra-frame random access point (IRAP) sub-picture.
[0401] Additionally, in one implementation, based on the presence of samples in the track that are mapped to a predetermined type (e.g., type 4) of a stream access point ('sap') sample group or a random access recovery point ('roll') sample group, when the current sample includes a non-IRAP subframe, samples belonging to the same coding layer video sequence (CLVS) as the current sample and following the current sample in decoding order can be constrained to include at least one IRAP subframe with the same subframe index value as the non-IRAP subframe.
[0402] In one implementation, each frame included in the mapping sample can be constrained to have a hybrid NAL unit type.
[0403] In one implementation, each sub-picture included in the mapping sample can be constrained to be considered a picture.
[0404] In the above embodiments, the mapping sample may only have a hybrid NAL unit type including the IRAP type.
[0405] In one implementation, all frames in the target layer of a sample that is mapped to a random access recovery point sample group may not be Gradually Decoded Refresh (GDR) frames (i.e., layer_id_method_idc = 2 or 3).
[0406] In one implementation, the TemporalID value of the current sample can be constrained to 0.
[0407] The generated media files can be sent to a media file receiving device via recording media or a network.
[0408] As described above, according to embodiments of this disclosure, only samples of mixed NAL cell types with IRAP type can be mapped to 'sap' or 'roll' sample groups. Therefore, random access on a sample-by-sample basis can be correctly performed.
[0409] Figure 15 This is a diagram illustrating how embodiments of this disclosure can be applied to content streaming systems.
[0410] like Figure 15 As shown, the content streaming system applying the embodiments of this disclosure may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0411] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream and then sends the bitstream to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.
[0412] The bitstream can be generated by an image encoding method or image encoding device applying the embodiments of this disclosure, and the stream server can temporarily store the bitstream during the sending or receiving of the bitstream.
[0413] A streaming server sends multimedia data to a user's device based on a request from a web server, and the web server acts as a medium for informing the user of the service. When a user requests a service from the web server, the web server can deliver it to the streaming server, and the streaming server can send the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between devices in the content streaming system.
[0414] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0415] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, board PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital televisions, desktop computers, digital signage, etc.
[0416] In a content streaming system, each server can operate as a distributed server, in which case the data received from each server can be distributed.
[0417] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on a device or computer.
[0418] Industrial applicability
[0419] The embodiments disclosed herein can be used to generate and send / receive media files.
Claims
1. A media file receiving method performed by a media file receiving device for receiving media files of a predetermined format, the media file including video data, the media file receiving method comprising the following steps: Obtain one or more tracks and sample groups from the media file; as well as The video data in the media file is processed by reconstructing samples included in the track based on the sample set. Specifically, this includes samples within the track that are mapped to a predetermined type of stream access point sample group or random access recovery point sample group. The current sample in the mapping sample is constrained to include at least one intra-frame random access point (IRAP) sub-frame, and Based on the fact that the current sample includes non-IRAP sub-frames, samples belonging to the same coding layer video sequence (CLVS) as the current sample and following the current sample in decoding order are constrained to include at least one IRAP sub-frame with the same sub-frame index value as the non-IRAP sub-frame.
2. The media file receiving method according to claim 1, wherein, Each frame included in the mapping sample is constrained to have a hybrid network abstraction layer (NAL) unit type, which allows the frame to include sub-frames with different NAL unit types from each other.
3. The media file receiving method according to claim 1, wherein, Each sub-picture included in the mapping sample is constrained to be considered a picture.
4. The media file receiving method according to claim 1, wherein, All frames in the target layer of the sample mapped to the random access recovery point sample group are not progressively decoded and refreshed GDR frames.
5. The media file receiving method according to claim 1, wherein, The mapping sample has only a hybrid network abstraction layer (NAL) unit type of IRAP type, which allows the picture included in the mapping sample to include sub-pictures with different NAL unit types from each other.
6. The media file receiving method according to claim 1, wherein, The TemporalID value of the current sample is constrained to 0.
7. A media file receiving device, the media file receiving device comprising a memory and at least one processor, in, The at least one processor is configured to: Obtain one or more tracks and sample groups from the media file; and The video data in the media file is processed by reconstructing samples included in the track based on the sample group. Specifically, this includes samples within the track that are mapped to a predetermined type of stream access point sample group or random access recovery point sample group. The current sample in the mapping sample is constrained to include at least one intra-frame random access point (IRAP) sub-frame, and Based on the fact that the current sample includes non-IRAP sub-frames, samples belonging to the same coding layer video sequence (CLVS) as the current sample and following the current sample in decoding order are constrained to include at least one IRAP sub-frame with the same sub-frame index value as the non-IRAP sub-frame.
8. A media file generation method performed by a media file generation device for generating media files of a predetermined format, the media file comprising video data, the media file generation method comprising the following steps: The video data is encoded; Generate one or more tracks and sample groups for encoded video data; as well as The media file is generated based on the generated tracks and sample groups. Specifically, this includes samples within the track that are mapped to a predetermined type of stream access point sample group or random access recovery point sample group. The current sample in the mapping sample is constrained to include at least one intra-frame random access point (IRAP) sub-frame, and Based on the fact that the current sample includes non-IRAP sub-frames, samples belonging to the same coding layer video sequence (CLVS) as the current sample and following the current sample in decoding order are constrained to include at least one IRAP sub-frame with the same sub-frame index value as the non-IRAP sub-frame.
9. The media file generation method according to claim 8, wherein, Each frame included in the mapping sample is constrained to have a hybrid network abstraction layer (NAL) unit type, which allows the frame to include sub-frames with different NAL unit types from each other.
10. The media file generation method according to claim 8, wherein, Each sub-picture included in the mapping sample is constrained to be considered a picture.
11. The media file generation method according to claim 8, wherein, All frames in the target layer of the sample mapped to the random access recovery point sample group are not progressively decoded and refreshed GDR frames.
12. The media file generation method according to claim 8, wherein, The mapping sample has only a hybrid network abstraction layer (NAL) unit type of IRAP type, which allows the picture included in the mapping sample to include sub-pictures with different NAL unit types from each other.
13. The media file generation method according to claim 8, wherein, The TemporalID value of the current sample is constrained to 0.
14. A method for sending a media file generated by the media file generation method of claim 8.
15. A media file generation device, the media file generation device comprising a memory and at least one processor, in, The at least one processor is configured to: Encode the video data; Generate one or more tracks and sample groups for the encoded video data; and Media files are generated based on the generated tracks and sample groups. Specifically, this includes samples within the track that are mapped to a predetermined type of stream access point sample group or random access recovery point sample group. The current sample in the mapping sample is constrained to include at least one intra-frame random access point (IRAP) sub-frame, and Based on the fact that the current sample includes non-IRAP sub-frames, samples belonging to the same coding layer video sequence (CLVS) as the current sample and following the current sample in decoding order are constrained to include at least one IRAP sub-frame with the same sub-frame index value as the non-IRAP sub-frame.