Method and device for processing a file including video data, and method and device for generating a file including video data
Patent Information
- Application Number
- BR112019024653
- Authority / Receiving Office
- BR · BR
- Patent Type
- Patents
- Current Assignee / Owner
- Publication Date
- 2026-08-25
Smart Images

Figure 00000089_0000 
Figure 00000089_0001 
Figure 00000090_0000
Abstract
Description
1 / 82 METHOD AND DEVICE FOR PROCESSING A FILE INCLUDING VIDEO DATA, AND METHOD AND DEVICE FOR GENERATING A FILE INCLUDING VIDEO DATA
[001] This application claims the benefit of US Provisional Patent Application Serial Number 62 / 511,189, filed May 25, 2017, the full content of which is incorporated herein by reference. TECHNICAL FIELD
[002] This disclosure relates to the storage and transport of encoded video data. BACKGROUND
[003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, direct digital broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptops or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video games, radio or satellite cellular phones, video teleconferencing devices, and the like. Digital video devices implement video compression techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 (also called High Efficiency Video Coding (HEVC)), and extensions of such standards to transmit and receive digital video information more efficiently.
[004] Video compression techniques perform spatial and / or temporal prediction for Petition 870250044067, dated 05 / 28 / 2025, page 6 / 185 2 / 82 Reduce or eliminate the inherent redundancy in video sequences. For block-based video coding, a video frame or slice can be divided into macroblocks. Each macroblock can be further divided. Macroblocks in an intra-coded frame or slice (I) are coded using spatial prediction with respect to neighboring macroblocks. Macroblocks in an inter-coded frame or slice (P or B) can use spatial prediction with respect to neighboring macroblocks in the same frame or slice, or temporal prediction with respect to other reference samples.
[005] After video data encoding, the video data can be packaged for transmission or storage. The video data can be assembled into a video file conforming to any of several standards, such as the International Organization for Standardization (ISO) base media file format and its extensions, such as AVO. SUMMARY
[006] In one example, a method includes processing a file containing fisheye video data, the file containing a syntax structure containing a plurality of syntax elements that specify attributes of the fisheye video data, wherein the plurality of syntax elements includes: a first syntax element that explicitly indicates whether the fisheye video data is monoscopic or stereoscopic and one or more syntax elements that implicitly indicate whether the fisheye video data is monoscopic or stereoscopic; determine, based on the first element Petition 870250044067, dated 05 / 28 / 2025, page 7 / 185 3 / 82 of syntax, whether the fisheye video data is monoscopic or stereoscopic; and render, based on the determination, the fisheye video data as monoscopic or stereoscopic.
[007] In another example, a device includes a memory configured to store at least a portion of a file containing fisheye video data, the file including a syntax structure including a plurality of syntax elements that specify attributes of the fisheye video data, wherein the plurality of syntax elements includes: a first syntax element that explicitly indicates whether the fisheye video data is monoscopic or stereoscopic and one or more syntax elements that implicitly indicate whether the fisheye video data is monoscopic or stereoscopic; and one or more processors configured to: determine, based on the first syntax element, whether the fisheye video data is monoscopic or stereoscopic; and render, based on the determination, the fisheye video data as monoscopic or stereoscopic.
[008] In another example, a method includes obtaining fisheye video data and extrinsic parameters from cameras used to capture the fisheye video data; determining, based on the extrinsic parameters, whether the fisheye video data is monoscopic or stereoscopic; and encoding, in a file, the fisheye video data and a syntax structure, including a plurality of syntax elements that specify attributes of the fisheye video data, wherein the plurality of syntax elements includes: a first syntax element that indicates Petition 870250044067, dated 05 / 28 / 2025, page 8 / 185 4 / 82 explicitly state whether the fisheye video data is monoscopic or stereoscopic and one or more syntax elements that explicitly indicate the extrinsic parameters of the cameras used to capture the fisheye video data.
[009] In another example, a device includes a memory configured to store fisheye video data; and one or more processors configured to: obtain extrinsic parameters from cameras used to capture the fisheye video data; determine, based on the extrinsic parameters, whether the fisheye video data is monoscopic or stereoscopic; and encode, in a file, the fisheye video data and a syntax structure, including a plurality of syntax elements that specify attributes of the fisheye video data, wherein the plurality of syntax elements includes: a first syntax element that explicitly indicates whether the fisheye video data is monoscopic or stereoscopic and one or more syntax elements that explicitly indicate the extrinsic parameters of the cameras used to capture the fisheye video data.
[0010] Details of one or more examples are shown in the accompanying drawings and description below. Other features, objects and advantages will be evident from the description, drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] FIGS. 1A and 1B are block diagrams illustrating examples of devices for capturing omnidirectional image content, according to one or more exemplary techniques described in this disclosure. Petition 870250044067, dated 05 / 28 / 2025, p. 9 / 185 5 / 82
[0012] FIG. 2 is an example of an image that includes multiple fisheye images.
[0013] FIGS. 3-8 are conceptual diagrams illustrating various extrinsic parameters and omnidirectional imaging fields of view, according to one or more examples of techniques described in this disclosure.
[0014] FIG. 9 is a block diagram illustrating an exemplary system that implements techniques for continuous media data flow over a network.
[0015] FIG. 10 is a conceptual diagram illustrating elements of exemplary multimedia content.
[0016] FIG. 11 is a block diagram illustrating elements of an example video file, which may correspond to a segment of a representation.
[0017] FIG. 12 is a flowchart that illustrates an example technique for processing a file that includes fisheye video data, according to one or more techniques of this disclosure.
[0018] FIG. 13 is a flowchart that illustrates an example technique for generating a file that includes fisheye video data, according to one or more techniques of this disclosure. DETAILED DESCRIPTION
[0019] The example techniques described in this disclosure relate to the processing of files representing omnidirectional video or image data. When omnidirectional media content is consumed with certain devices (e.g., a head-mounted display and headphones), only the portions of the media that correspond to the user's viewing orientation are processed. Petition 870250044067, dated 05 / 28 / 2025, page 10 / 185 6 / 82 rendered, as if the user were in the location where and when the media was captured (e.g., where the cameras were). One of the most popular forms of omnidirectional media applications is omnidirectional video, also known as 360-degree video. Omnidirectional video is typically captured by multiple cameras that cover up to 360 degrees of the scene.
[0020] In general, omnidirectional video is formed from a sequence of omnidirectional images. Therefore, the example techniques described in this disclosure are described with respect to the generation of omnidirectional image content. Then, for omnidirectional video content, these omnidirectional images can be displayed sequentially. In some examples, a user may wish to take only one omnidirectional image (for example, as a snapshot of the user's entire 360-degree environment), and the techniques described in this disclosure are also applicable to these example cases.
[0021] Omnidirectional video can be stereoscopic or monoscopic. When the video is stereoscopic, a different image is displayed to each eye, so that the viewer perceives depth. As such, stereoscopic video is typically captured using two cameras facing each direction. When the video is monoscopic, the same image is shown to both eyes.
[0022] Video data can be considered fisheye video data, where it is captured using one or more fisheye lenses (or generated to appear as if it was captured using one or more fisheye lenses). Petition 870250044067, dated 05 / 28 / 2025, page 11 / 185 7 / 82 A fisheye lens can be a wide-angle lens that produces strong visual distortion intended to create a wide panoramic or hemispherical image.
[0023] The techniques can be applied to captured video content, virtual reality, and generally to video and image display. The techniques can be used on mobile devices, but the techniques should not be considered limited to mobile applications. In general, the techniques can be for virtual reality applications, video game applications, or other applications where a 360-degree spherical image / video environment is desired.
[0024] In some examples, omnidirectional image content can be captured with a camera device that includes two fisheye lenses. Where the two fisheye lenses are positioned on opposite sides of the camera device to capture opposite parts of the image content sphere, the image content can be monoscopic and cover the entire 360-degree video sphere. Similarly, where the two fisheye lenses are positioned on the same side of the camera device to capture the same part of the image content sphere, the image content can be stereoscopic and cover half of the 360-degree video sphere. The images generated by the cameras are circular images (e.g., one image frame includes two circular images).
[0025] FIGS. 1A and 1B are block diagrams illustrating examples of devices for capturing omnidirectional image content, according to one or more exemplary techniques described in this disclosure. Petition 870250044067, dated 05 / 28 / 2025, page 12 / 185 8 / 82 As illustrated in FIG. 1A, computing device 10A is a video capture device that includes fisheye lens 12A and fisheye lens 12B located on opposite sides of computing device 10A to capture monoscopic image content covering the entire sphere (e.g., full 360-degree video content). As illustrated in FIG. 1B, computing device 10B is a video capture device that includes fisheye lens 12C and fisheye lens 12D located on the same side of computing device 10B for stereoscopic image content covering approximately half of the sphere.
[0026] As described above, a camera device includes a plurality of fisheye lenses. Some examples of camera devices include two fisheye lenses, but the example techniques are not limited to two fisheye lenses. One example of a camera device might include 16 lenses (e.g., a set of 16 cameras for filming 3D VR content). Another example of a camera device might include eight lenses, each with a 195-degree angle of view (e.g., each lens captures 195 degrees of the 360 degrees of image content). Another example of a camera device might include three or four lenses. Some examples might include a 360-degree lens that captures 360 degrees of image content.
[0027] The example techniques described in this disclosure are generally described with respect to two fisheye lenses capturing omnidirectional image / video. However, the example techniques are not limited to this. The example techniques may be applicable to the example camera devices that include a Petition 870250044067, dated 05 / 28 / 2025, page 13 / 185 9 / 82 plurality of lenses (e.g., two or more), even if the lenses are not fisheye lenses and a plurality of fisheye lenses. For example, the techniques in the example describe ways to stitch captured images together, and the techniques may be applicable to examples where there is a plurality of images captured from a plurality of lenses (which may be fisheye lenses, as an example). Although the techniques in the example are described in relation to two fisheye lenses, the techniques in the example are not limited to this and are applicable to the various types of cameras used to capture omnidirectional images / videos.
[0028] The techniques of this disclosure can be applied to video files conforming to video data encapsulated according to any ISO base media file format (e.g., ISOBMFF, ISO / IEC 14496-12) and other ISOBMFF-derived file formats, including MPEG-4 file format (ISO / IEC 14496-15), Third Generation Partnership Project (3GPP) file format (3GPP TS 26.244), and file formats for AVO and HEVC families of video codecs (ISO / IEC 14496-15) or other similar video file formats.
[0029] ISOBMFF is used as the basis for many codec encapsulation formats, such as the AVC file format, as well as for many multimedia container formats, such as the MPEG-4 file format, the 3GPP (3GP) file format, and the DVB file format. In addition to continuous media such as audio and video, static media such as images, as well as metadata, can be stored in an ISOBMFF-compliant file. Files Petition 870250044067, dated 05 / 28 / 2025, page 14 / 185 10 / 82 structured according to ISOBMFF can be used for various purposes, including local media file playback, progressive download of a remote file, segments for Dynamic Adaptive Streaming over HTTP (DASH), containers for content to be streamed and their packaging instructions, and recording of received media streams in real time.
[0030] A box is the elementary syntax structure in ISOBMFF, including a four-character encoded box type, the box's byte count, and the payload. An ISOBMFF file consists of a sequence of boxes, and boxes can contain other boxes. A movie box (moov) contains the metadata for the continuous media streams present in the file, each represented in the file as a track. The metadata for a track is placed in a Track box (trak), while the media content of a track is either placed in a Media Data box (mdat) or directly in a separate file. The media content of the tracks consists of a sequence of samples, such as audio or video access units.
[0031] When rendering fisheye video data, it may be desirable for a video decoder to determine whether the fisheye video data is monoscopic or stereoscopic. For example, the way in which the video decoder displays and / or joins the circular images included in a fisheye video data image depends directly on whether the fisheye video data is monoscopic or stereoscopic.
[0032] In some examples, a video decoder can determine if the fisheye video data is Petition 870250044067, dated 05 / 28 / 2025, page 15 / 185 11 / 82 Monoscopic or stereoscopic based on one or more syntax elements included in a file that implicitly indicate whether the fisheye video data described by the file is monoscopic or stereoscopic. For example, a video encoder might include syntax elements in the file that describe extrinsic parameters (e.g., physical / local attributes such as yaw angle, pitch angle, roll angle, and one or more spatial offsets) of each of the cameras used to capture the fisheye video data, and the video decoder might process the extrinsic parameters to calculate whether the video data described by the file is monoscopic or stereoscopic. As an example, if processing the extrinsic parameters leads to an indication that the cameras are facing the same direction and have a spatial offset, the video decoder might determine that the video data is stereoscopic.As another example, if processing the extrinsic parameters indicates that the cameras are facing opposite directions, the video decoder may determine that the video data is monoscopic. However, in some examples, it may be desirable for the video decoder to be able to determine whether the fisheye video data is monoscopic or stereoscopic without having to perform these additional, resource-intensive calculations. Additionally, in some examples, processing the extrinsic parameters may not produce the correct (i.e., as intended) classification of the video data as monoscopic or stereoscopic. For example, the values of the extrinsic parameters... Petition 870250044067, dated 05 / 28 / 2025, page 16 / 185 12 / 82 can be corrupted during encoding / transit / decoding. As another example, extrinsic parameters may not have been accurately supplied to the encoder.
[0033] According to one or more techniques of this disclosure, a video encoder may include a first syntax element in a file that explicitly indicates whether the fisheye video data described by the file is monoscopic or stereoscopic. The file may include the first syntax element in addition to the syntax elements that are the extrinsic parameters of the fisheye video data. In this way, the file may include an explicit indication (i.e., the first syntax element) and an implicit indication (i.e., the syntax elements that are the extrinsic parameters) of whether the fisheye video data described by the file is monoscopic or stereoscopic. In this way, a video decoder can accurately determine whether the video data is monoscopic or stereoscopic (e.g., without the need to perform additional calculations based on the extrinsic parameters).
[0034] ISOBMFF specifies the following track types: a media track, which contains an elementary media stream, a hint track, which includes media transmission instructions or represents a received packet stream, and a timed metadata track, which includes time-synchronized metadata.
[0035] Although originally designed for storage, ISOBMFF has proven to be very valuable for streaming, for example, for progressive download or DASH. For streaming purposes, the film fragments defined in Petition 870250044067, dated 05 / 28 / 2025, page 17 / 185 13 / 82 ISOBMFF can be used.
[0036] The metadata for each track includes a list of sample description entries, each providing the encoding or encapsulation format used in the track and the initialization data needed to process that format. Each sample is associated with one of the track's sample description entries.
[0037] ISOBMFF allows specifying sample-specific metadata with various mechanisms. Specific boxes within the Sample Table (stbl) box have been standardized to address common needs. For example, a Sample Sync (stss) box is used to list the random access samples of the range. The sample grouping mechanism allows mapping samples according to a four-character grouping type into sample groups that share the same property specified as a sample group description entry in the file. Several grouping types have been specified in ISOBMFF.
[0038] Virtual reality (VR) is the ability to be virtually present in a non-physical world created by rendering natural and / or synthetic images and sounds correlated by the movements of the immersed user, allowing interaction with that world. With the recent progress achieved in rendering devices, such as head-mounted displays (FDVID) and VR video creation (generally also known as 360-degree video), a significant quality of experience can be offered. VR applications include gaming, training, education, sports video, online shopping, adult entertainment and Petition 870250044067, dated 05 / 28 / 2025, page 18 / 185 14 / 82 and so on.
[0039] A typical VR system may include one or more of the following components, which may perform one or more of the following steps: 1) A camera array, which typically consists of several individual cameras pointing in different directions and ideally collectively covering all viewpoints around the camera array. 2) Image stitching, in which video images taken by multiple individual cameras are synchronized in the time domain and stitched together in the space domain to create a spherical video, but mapped to a rectangular format, such as an equi-rectangular (like a world map) or cube map. 3) The video in mapped rectangular format is encoded / compressed using a video codec, for example, H.265 / HEVC or H.264 / AVC. 4) The compressed video bitstream can be stored and / or encapsulated in a media format and transmitted (possibly only the subset that covers the area being viewed by a user) over a network to a receiver. 5) The receiver receives the video bitstream(s) or part thereof, possibly encapsulated in a format, and sends the decoded video signal or part thereof to a Petition 870250044067, dated 05 / 28 / 2025, page 19 / 185 15 / 82 rendering device. 6) The rendering device could be, for example, an FDVID, which can track head movement and even the timing of eye movement and render the corresponding part of the video, so that an immersive experience is delivered to the user.
[0040] The Omnidirectional Media Application Format (OMAF) is being developed by MPEG to define a media application format that enables omnidirectional media applications, focusing on VR applications with 360° video and associated audio. OMAF specifies the projection and region-by-region packaging that can be used for converting a spherical or 360° video into a two-dimensional rectangular video, followed by how to store omnidirectional media and associated metadata using the ISO base media file format (ISOBMFF), and how to encapsulate, signal, and transmit omnidirectional media using Dynamic Adaptive Streaming over HTTP (DASH), and finally, which video and audio codecs, as well as media encoding settings, can be used for compression and playback of the omnidirectional media signal.
[0041] Projection and region packing are the processes used on the content production side to generate 2D video images from the sphere signal for projected omnidirectional video. Projection usually follows stitching, which can generate the sphere signal from multiple images captured by the camera. Petition 870250044067, dated 05 / 28 / 2025, page 20 / 185 16 / 82 for each video image. Projection can be a crucial step in VR video processing. Typical projection types include equi-rectangular and cubemap. The OMAF International Draft Standard (DIS) only supports the equi-rectangular projection type. Region packing is an optional step after projection (in the display window on the content production side). Region packing allows manipulation (resizing, repositioning, rotation, and mirroring) of any rectangular region of the compressed image before encoding.
[0042] FIG. 2 is an example of an image that includes multiple fisheye images. OMAF DIS supports a VR / 360 fisheye video format, in which, instead of applying projection and optionally region packing to generate the 2D video before encoding, for each access unit, the circular images from the capture cameras are embedded directly into a 2D image. For example, as shown in FIG. 2, the first fisheye image 202 and the second fisheye image 204 are embedded in the 2D image 200.
[0043] This fisheye video can then be encoded and the bitstream can be encapsulated in an ISOBMFF file and can be further encapsulated as a DASH representation. Additionally, the fisheye video property, including parameters indicating the fisheye video characteristics, can be signaled and used to correctly render the 360 video on the client side. One advantage of the VR / 360 fisheye video approach is that it supports low-cost VR content generated by users on mobile devices. Petition 870250044067, dated 05 / 28 / 2025, page 21 / 185 17 / 82
[0044] In OMAF DIS, the use of the omnidirectional fisheye video scheme for the restricted video sample input type 'resv' indicates that the decoded images are fisheye video images. The use of the omnidirectional fisheye video scheme is indicated by the scheme type being 'fodv' (omnidirectional fisheye video) in the SchemeTypeBox. The fisheye video format is indicated by the FisheyeOmnidirectionalVideoBox contained in the SchemelnformationBox, which is included in the RestrictedSchemelnfoBox that is included in the sample input. In the current draft of OMAF DIS (Information Technology - Immersive Media Encoded Representation (MPEG-I) - Part 2: Omnidirectional Media Format, ISO / IEC EDIS 14496-15: 2014 (E), ISO / IEC JTC 1 / SC 29 / WG 11, W16824, 2014-01-13, hereinafter the current draft of OMAF DIS), one and only one FisheyeOmnidirectionalVideoBox must be present in the SchemelnformationBox when the schema type is 'fodv'.When FisheyeOmnidirectionalVideoBox is present in the SchemelInformationBox, StereoVideoBox and RegionWisePackingBox will not be present in the same SchemelInformationBox. The FisheyeOmnidirectionalVideoBox, as specified in clause 6 of the OMAF DIS, contains the FisheyeOmnidirectionalVideoInfo() syntax structure which contains the fisheye video property parameters.
[0045] The syntax of the FisheyeOmnidirectionalVideoInfo() syntax structure in the current OMAF DIS draft is as follows. Petition 870250044067, dated 05 / 28 / 2025, p. 22 / 185 18 / 82 aligned(8) class FisheyeOmnidirectionalVídeoInfo( ) { bit(24) reserved = 0; unsigned int(8) in a circular images; for(i=0, i< num_circular_images; i++) { unsigned int(32) image_center_x; unsigned int(32) image_center_y, unsigned int(32) full_radius; unsigned int(32) picture_radius; unsigned int(32) scene_radius; unsigned int(32) image_rotation; bit(30) reserved = 0; unsigned int(2) image_flip, unsigned int(32) image_scale_axis_angle; unsigned int(32) image_scale_x; unsigned int(32) image_scale_y; unsigned int(32) field of view; bit(l 6) reserved = 0; unsigned int (16) num_angle_for_displaying_fov; for(j=0; j< num angle for displaying fov; j++) { Petição 870250044067, de 28 / 05 / 2025, pág. 23 / 185 19 / 82 unsigned int(32) displayed fov; unsigned int(32) overlapped_fov; signed int(32) camera_center_yaw; signed int(32) camera_center pitch, signed int(32) camera_center_roll; unsigned int(32) camera_center_offset_x; unsigned int(32) camera_center_offset_y; unsigned int(32) camera_center_offset_z; bit(l 6) reserved = 0; unsigned int(l 6) num_polynomial coefficeients, for(j=0; j< num_polynomial_coefficients, j++) { unsigned int(32) polynomial_coefficient_K; bit( 16) reserved = 0; unsigned int (16) num local fov region; for(j=0;j<num local fov region; j++) { unsigned int(32) start radius; unsigned int(32) end_radius; signed int(32) start_angle; signed int(32) end_angle; unsigned int(32) radius delta; signed int(32) angle_delta; for(rad=start_radius; rad<= end_radius; rad+=radius_delta) { for(ang=start_angle; ang<= ang_radius; ang+=angle_delta) { unsigned int(32) local_fov_weight; bit(16) reserved = 0; unsigned int(l 6) num_polynomial_coefficients_lsc; for(j=0; j< num_polynomial coefficients 1 sc; j++) { unsigned int (32) polynomial coefficient K lsc R; unsigned int (32) polynomial coefficient K lsc G; Petition 870250044067, dated 05 / 28 / 2025, p. 24 / 185 20 / 82 unsigned int (32) polynomial_coefFicient_K_lsc_B; bit(24) reserved = 0; unsigned int(8) num_deadzones; for(i=0, i< numdeadzones, i++) { unsigned int(16) deadzone J eftjaorizontal_offset; unsigned int(l 6) deadzone_top_vertical_offset; unsigned int(l 6) deadzone_width; unsigned int(16) deadzone height;
[0046] The semantics of the FisheyeOmnidirectionalVideoInfo() syntax structure in the current draft of OMAF DIS is as follows.
[0047] num_circular_images specifies the number of circular images in the encoded image of each sample to which this box applies. Normally, the value is 2, but other non-zero values are also possible.
[0048] image_center_x is a fixed-point value 16.16 that specifies the horizontal coordinate, in luma samples, of the center of the circular image in the encoded image of each sample to which this box applies.
[0049] image_center_y is a fixed-point value 16.16 that specifies the vertical coordinate, in luma samples, of the center of the circular image in the encoded image of each sample to which this box applies.
[0050] full_radius is a fixed-point value Petition 870250044067, dated 05 / 28 / 2025, p. 25 / 185 21 / 82 16.16 which specifies the radius, in luma samples, from the center of the circular image to the edge of the complete round image.
[0051] picture_radius is a fixed-point value of 16.16 that specifies the radius, in luma samples, from the center of the circular image to the nearest edge of the image. The circular fisheye image may be cropped by the camera image. Therefore, this value indicates the radius of a circle in which the pixels are usable.
[0052] scene_radius is a fixed-point value 16.16 that specifies the radius, in luma samples, from the center of the circular image to the nearest edge of the area in the image where it is guaranteed that there are no obstructions in the camera body itself and that within the enclosed area there is no lens distortion too great for stitching.
[0053] image_rotation is a fixed-point value 16.16 that specifies the amount of rotation, in degrees, of the circular image. The image can be rotated by + / - 90 degrees, or + / - 180 degrees, or any other value.
[0054] image_flip specifies whether and how the image has been flipped and therefore a reverse flip operation needs to be applied. The value 0 indicates that the image has not been flipped. The value 1 indicates that the image has been flipped vertically. The value 2 indicates that the image has been flipped horizontally. The value 3 indicates that the image has been flipped both vertically and horizontally.
[0055] Image_scale_axis_angle, image_scale_x and image_scale_y are three fixed-point values that specify whether and how the image has been scaled along an axis. The axis is defined by a single angle, as indicated by the image scale axis angle value. Petition 870250044067, dated 05 / 28 / 2025, page 26 / 185 22 / 82 in degrees. An angle of 0 degrees means that a horizontal vector is perfectly horizontal and a vertical vector is perfectly vertical. The values of image_scale_x and image_scale_y indicate the scale ratios in the directions parallel and orthogonal, respectively, to the axis.
[0056] field_of_view is a fixed-point value 16.16 that specifies the field of view of the fisheye lens, in degrees. A typical value for a hemispherical fisheye lens is 180.0 degrees.
[0057] num_angle_for_displaying_fov specifies the number of angles. Depending on the value of num_angle_for_displaying_fov, various displayed fov and overlaid fov values are set at equal intervals, starting at 12 o'clock and going clockwise.
[0058] displayed_fov specifies the displayed field of view and the corresponding image area of each fisheye camera image, overlapped_fov specifies the region that includes overlapping regions, which are generally used for merging, in terms of the field of view between multiple circular images. The values of displayed_fov and overlapped_fov are less than or equal to the field of view value.
[0059] NOTE: The field of view value is determined by the physical properties of each fisheye lens, while the displayed_fov and overlapped_fov values are determined by the configuration of multiple fisheye lenses. For example, when the num_circular_images value is equal to 2 and two lenses are located symmetrically, the displayed_fov and overlapped_fov values can be set to 180 and 190. Petition 870250044067, dated 05 / 28 / 2025, p. 27 / 185 23 / 82 respectively, by default. However, the value can be changed depending on the lens configuration and content characteristics. For example, if the stitch quality with the displayed_fov values (left camera = 170 and right camera = 190) and the overlapped_fov values (left camera = 185 and right camera = 190) is better than the quality with the default values (180 and 190), or if the physical configuration of the cameras is asymmetrical, unequal values of displayed_fov and overlapped_fov may be obtained. Furthermore, when dealing with multiple (N>2) fisheye images, a single displayed_fov value cannot specify the exact area of each fisheye image. As shown in FIG. 6, the displayed_fov (602) varies according to the direction. To handle multiple (N>2) fisheye images, the num_angle_for_displaying_fov is introduced. For example, if this value is equal to 12, the fisheye image will be divided into 12 sectors where each sector angle is 30 degrees.
[0060] camera_center_yaw specifies the yaw angle, in units of 2-16 degrees, of the point where the center pixel of the circular image in the encoded image of each sample is projected onto a spherical surface. This is the first of 3 angles that specify the extrinsic camera parameters relative to the global coordinate axes. camera_center_yaw must be in the range -180 * 216 to 180 * 216 - 1, inclusive.
[0061] camera_center_pitch specifies the yaw angle, in units of 2-16 degrees, of the point where the center pixel of the circular image in the encoded image of each sample is projected onto a spherical surface. Petition 870250044067, dated 05 / 28 / 2025, page 28 / 185 24 / 82 Camera_center_pitch should be in the range of -90 * 216 to 90 * 216, inclusive.
[0062] camera_center_roll specifies the roll angle, in units of 2-16 degrees, of the point where the center pixel of the circular image in the encoded image of each sample is projected onto a spherical surface. camera_center_roll must be in the range of -180 * 216 to 180 * 216, inclusive.
[0063] camera_center_offset_x, camera_center_offset_y and camera_center_offset_z are fixed-point values 8,24 that indicate the XYZ offset values from the origin of the unit sphere onto which the pixels in the circular image in the encoded image are projected. camera_center_offset_x, camera_center_offset_y and camera_center_offset_z must be in the range of -1.0 to 1.0, inclusive.
[0064] Num_polynomial_coefficients is an integer that specifies the number of polynomial coefficients present. The list of polynomial coefficients K are 8.24 fixed-point values that represent the coefficients in the polynomial that specify the transformation of fisheye space into a non-stored planar image.
[0065] num_local_fov_region specifies the number of local adjustment regions with different fields of view.
[0066] start_radius, end_radius, start_angle and end_angle specify the region for local adjustment / distortion to change the actual field of view for local display, start_radius and end_radius are fixed-point values 16.16 that specify the minimum and maximum radius values, Petition 870250044067, dated 05 / 28 / 2025, page 29 / 185 25 / 82 start_angle and end_angle specify the minimum and maximum values of angles that begin at 12 o'clock and increase clockwise, in units of 2-16 degrees, start_angle and end_angle must be in the range of -180 * 216 to 180 * 2161, inclusive.
[0067] Radius_delta is a fixed-point value 16.16 which specifies the delta radius value for representing a different field of view for each ray.
[0068] angle_delta specifies the value of the delta angle, in units of 2-16 degrees, for representing a different field of view for each angle.
[0069] Local_fov_weight is an 8.24 fixed-point format that specifies the weighting value for the field of view at the position specified by start_radius, end_radius, start_angle, end_angle, angle index 1, and radius index j. A positive value for local fov weight specifies expanding the field of view, while a negative value specifies contracting the field of view.
[0070] num_polynomial_coefficients_lsc must be the order of the polynomial approximation of the lens shading curve. [ 0071] polynomial_coefficient_K_lsc_R, polynomial_coefficient_K_lsc_G, and polynomial_coefficient_K_lsc_B are 8,24 fixed-point formats that specify the LSC parameters to compensate for the shading artifact that reduces color along the radial direction. The compensation weight (w) to be multiplied to the original color is approximated as a curve function of the image center radius using a Petition 870250044067, dated 05 / 28 / 2025, page 30 / 185 26 / 82 polynomial expression. It is formulated as w = Σ =1 Pi-' r1-1, where p indicates the value of the coefficient equal to polynomial_coefficient_K_lsc_R, polynomial_coefficient_K_lsc_G or polynomial_coefficient_K_lsc_B, er indicates the value of the radius after normalization by full_radius. N is equal to the value of num_polynomial_coefficients_lsc.
[0072] num_deadzones is an integer that specifies the number of dead zones in the encoded image of each sample to which this box applies. [ 0 07 3] deadzone_left_horizontal_offset, deadzone_top_vertical_offset, deadzone_width, and deadzone_height are integers that specify the position and size of the rectangular dead zone area where pixels are not used. deadzone_left_horizontal_offset and deadzone_top_vertical_offset specify the horizontal and vertical coordinates, respectively, in luma samples, of the upper-left corner of the dead zone in the encoded image. deadzone_width and deadzone_height specify the width and height, respectively, in luma samples, of the dead zone. To save bits for video representation, all pixels in a dead zone must be set to the same pixel value, for example, all black.
[0074] FIGS. 3-8 are conceptual diagrams illustrating various aspects of the syntax and semantics above. FIG. 3 illustrates the syntax of image_center_x and image_center_y as center 302, the syntax full_radius as full radius 304, picture_radius as frame radius 306 of frame 300, scene_radius as scene radius 308 (e.g., the unobstructed area of the camera body 510). FIG. 4 Petition 870250044067, dated 05 / 28 / 2025, p. 31 / 185 Figure 27 / 82 illustrates the displayed_fov (i.e., displayed field of view (FOV)) for two fisheye images. FOV 402 represents a 170-degree field of view, and FOV 404 represents a 190-degree field of view. Figure 5 illustrates the displayed_fov (i.e., displayed field of view (FOV)) and overlapped_fov for multiple fisheye images (e.g., N > 2). FOV 502 represents a first field of view, and FOV 504 represents a second field of view. Figure 6 is an illustration of the syntax camera_center_offset_x (ox*), camera_center_offset_y (oy), and camera_center_offset_z (oz). Figure 7 is a conceptual diagram illustrating parameters with respect to the local field of view. Figure 8 is a conceptual diagram illustrating an example of a local field of view.
[0075] The fisheye video signaling in the current OMAF DIS draft may present one or more disadvantages.
[0076] As an example of a disadvantage of fisheye video signaling in the current OMAF DIS draft, when there are two circular images in each fisheye video image (for example, where num_circular_images in the FisheyeOmnidirectionalVideoInfo() syntax structure is equal to 2), depending on the values of the camera's extrinsic parameters (for example, camera_center_yaw, camera_center_pitch, camera_center_center_ro11, camera_center_offset_x, camera_center_offset_y, and camera_center_offset_z), the fisheye video can be monoscopic or stereoscopic. In particular, when the two cameras capturing the two circular images are on the same side, the fisheye video is stereoscopic covering Petition 870250044067, dated 05 / 28 / 2025, page 32 / 185 28 / 82 covers approximately half (i.e., 180 degrees horizontally) of the sphere, and when the two cameras are on opposite sides, the fisheye video is monoscopic but covers approximately the entire sphere (i.e., 360 degrees horizontally). For example, when two sets of extrinsic camera parameter values are as follows, the fisheye video is monoscopic: Ioconjunto: camera_center_yaw = 0 degrees (+ / - 5 degrees) camera_center_pitch = 0 degrees (+ / - 5 degrees) camera_center_roll = 0 degrees (+ / - 5 degrees) camera_center_offset_x = 0 mm (+ / - 3 mm) camera_center_offset_y = 0 mm (+ / - 3 mm) camera_center_offset_z = 0 mm (+ / - 3 mm) 2nd set: camera_center_yaw = 180 degrees (+ / - 5 degrees) camera_center_pitch = 0 degrees (+ / - 5 degrees) camera_center_roll = 0 degrees (+ / - 5 degrees) camera_center_offset_ X = 0 mm ( + / - 3 mm) camera_ _center _of f set_ y = 0 mm ( + / - 3 mm) camera_ _center _of f set_ z = 0 mm ( + / - 3 mm)
[0077] The values of the parameter above can correspond to the extrinsic parameter values of the computing device's camera 10B in FIG. 1B. As another example, when two sets of extrinsic camera parameter values are as follows, the fisheye video is stereoscopic: Ioconjunto: camera_center_yaw = 0 degrees (+ / - 5 degrees) camera_center_pitch = 0 degrees (+ / - 5 degrees) Petition 870250044067, dated 05 / 28 / 2025, page 33 / 185 29 / 82 camera_center_roll = 0 degrees (+ / - 5 degrees) camera_center_offset_x = 0 mm (+ / - 3mm) camera_center_offset_y = 0 mm (+ / - 3mm) camera_center_offset_z = 0 mm (+ / - 3mm) 2nd set: camera_center_yaw = 0 degrees (+ / - 5 degrees) camera_center_pitch = 0 degrees (+ / - 5 degrees) camera_center_roll = 0 degrees (+ / - 5 degrees) camera_center_offset_x = 64 mm (+ / - 3 mm) camera_center_offset_y = 0 mm (+ / - 3 mm) camera_center_offset_z = 0 mm (+ / - 3 mm)
[0078] The above parameter values may correspond to the extrinsic parameter values of the computing device 10A camera in FIG. 1A. Note that the X distance of the stereo offset of 64 mm is equivalent to 2.5 inches, which is the average distance between human eyes.
[0079] In other words, information about whether the fisheye video is monoscopic or stereoscopic is hidden (i.e., implicitly encoded, but not explicitly). However, for high-level system purposes, such as content selection, it may be desirable for this information to be readily accessible at the file format level and in DASH (for example, so that the entity performing the content selection function does not need to parse much information in the FisheyeOmnidirectionalVideoInfo() syntax structure to determine whether the corresponding video data is monoscopic or stereoscopic).
[0080] As another example of a disadvantage of Petition 870250044067, dated 05 / 28 / 2025, page 34 / 185 30 / 82 Fisheye video signaling in the current OMAF DIS draft, when there are two circular images in each fisheye video frame (for example, where circular images in the FisheyeOmnidirectionalVideoInfo() syntax structure equals 2) and the fisheye video is stereoscopic, a client device can determine which of the two circular images is the left eye view and which is the right eye view from the camera's extrinsic parameters. However, it may not be desirable for the client device to need to determine which image is the left eye view and which is the right eye view from the camera's extrinsic parameters.
[0081] As another example of a disadvantage of fisheye video signaling in the current OMAF DIS draft, when there are four fisheye cameras, two on each side, to capture the fisheye video, the fisheye video would be stereoscopic, covering the entire sphere. In this case, there are four circular images in each fisheye video image (for example, the number of circular images in the FisheyeOmnidirectionalVideoInfo() syntax structure is equal to 4). In these cases (when there are more than two circular images in each fisheye video image), a client device can determine the pairing of certain two circular images belonging to the same view from the camera's extrinsic parameters. However, it may not be desirable for the client device to have to determine the pairing of certain two circular images belonging to the same view from the camera's extrinsic parameters.
[0082] As another example of a disadvantage of fisheye video signaling in the current OMAF DIS draft, for the projected omnidirectional video, the information of Petition 870250044067, dated 05 / 28 / 2025, page 35 / 185 31 / 82 Coverage (e.g., what area of the sphere is covered by the video) is explicitly signaled using the CoverageInformationBox or, when not present, inferred by the client device as being the entire sphere. Although the client device can determine coverage from extrinsic camera parameters, again, explicit signaling of this information may be desirable for easy access.
[0083] As another example of a disadvantage of fisheye video signaling in the current OMAF DIS draft, the use of region-based packing is allowed for projected omnidirectional video, but not allowed for fisheye video. However, some benefits of region-based packing applicable to projected omnidirectional video may also be applicable to fisheye video.
[0084] As another example of a disadvantage of fisheye video signaling in the current OMAF DIS draft, omnidirectional video transport designed in sub-image tracks is specified in OMAF DIS using composition track grouping. However, omnidirectional fisheye video transport in sub-image tracks is not supported.
[0085] This revelation presents solutions to the problems above. Some of these techniques can be applied independently and some of them can be applied in combination.
[0086] According to one or more techniques of this disclosure, a file that includes fisheye video data may include an explicit indication of whether the fisheye video data is monoscopic or stereoscopic. In Petition 870250044067, dated 05 / 28 / 2025, p. 36 / 185 32 / 82 In other words, an indication of whether the fisheye video is monoscopic or stereoscopic can be explicitly signaled. In this way, a client device can avoid having to infer whether the fisheye video data is monoscopic or stereoscopic from other parameters.
[0087] As an example, one or more of the initial 24 bits in the syntax structure FisheyeOmnidirectionalVideoInfo() can be used to form a field (e.g., a one-bit flag) to indicate monoscopic or stereoscopic video. The field being equal to a specific first value (e.g., 0) can indicate that the fisheye video is monoscopic, and the field being equal to a specific second value (e.g., 1) can indicate that the fisheye video is stereoscopic. For example, when generating a file including fisheye video data, a content preparation device (e.g., content preparation device 20 from FIG. 20) can be used.9 and, in one specific example, the encapsulation unit 30 of the content preparation device 20) can encode a box (e.g., a FisheyeOmnidirectionalVideoBox) within the file, the box including a syntax structure (e.g., a FisheyeOmnidirectionalVideoInfo()) that contains parameters for the fisheye video data and the syntax structure, including an explicit indication of whether the fisheye video data is monoscopic or stereoscopic. A client device (e.g., client device 40 in FIG. 9, and in one specific example, the retrieval unit 52 of client device 40) can process the data in the same way. Petition 870250044067, dated 05 / 28 / 2025, page 37 / 185 File 33 / 82 to obtain explicit indication of monoscopic or stereoscopic.
[0088] As another example, a new box containing a field, for example, a one-bit flag, some reserved bits, and possibly some other information can be added to the FisheyeOmnidirectionalVideoBox to signal whether the video is monoscopic or stereoscopic. The field being equal to a specific first value (e.g., 0) can indicate that the fisheye video is monoscopic, and the field being equal to a specific second value (e.g., 1) can indicate that the fisheye video is stereoscopic. For example, when generating a file including fisheye video data, a content preparation device (e.g., content preparation device 20 in FIG. 9) can encode a first box (e.g., a FisheyeOmnidirectionalVideoBox) within the file, the first box including a second box that includes a field that explicitly indicates whether the fisheye video data is monoscopic or stereoscopic. A client device (e.g., client device 40 in FIG. 9) can then encode a second box containing a field that explicitly indicates whether the fisheye video data is monoscopic or stereoscopic.9) can also process the file to obtain an explicit indication of whether it is monoscopic or stereoscopic.
[0089] As another example, a field (for example, a one-bit flag) can be added directly to the FisheyeOmnidirectionalVideoBox to signal this indication. For example, when generating a file including fisheye video data, a content preparation device (for example, content preparation device 20 of FIG. 9) can encode a box (for Petition 870250044067, dated 05 / 28 / 2025, page 38 / 185 34 / 82 example, a Fisheye Omnidirectional Video Box) within the file, the box including a field that explicitly indicates whether the fisheye video data is monoscopic or stereoscopic. A client device (e.g., client device 40 in FIG. 9) can also process the file to obtain the explicit indication of monoscopic or stereoscopic.
[0090] According to one or more techniques of this disclosure, a file that includes fisheye video data as a plurality of circular images may include, for each respective circular image of the plurality of circular images, an explicit indication of a respective view identification (view ID). In other words, for each circular image of fisheye video data, a field may be added indicating the view ID of each circular image. When there are only two view ID values 0 and 1, then the view with view ID equal to 0 may be the left view and the view with view ID equal to 1 may be the right view.For example, where the plurality of circular images includes only two circular images, one circular image from the plurality of circular images with a predetermined first video ID is a left view, and one circular image from the plurality of circular images with a predetermined second video ID is a right view. In this way, a client device can avoid having to infer which circular image is the left view and which is the right view from other parameters.
[0091] The flagged view ID can also be used to indicate the pairing relationship of any Petition 870250044067, dated 05 / 28 / 2025, p. 39 / 185 35 / 82 two circular images belonging to the same view (for example, as both may have the same ID value of the flagged view).
[0092] To allow view IDs to be easily accessed without needing to parse all fisheye video parameters for circular images, a loop can be added after the num_circular_images field and before the existing loop. For example, file processing could include parsing, in a first loop, the explicit indications of the view IDs; and parsing, in a second loop after the first loop, other parameters of the circular images. For example, the view IDs and other parameters could be parsed as follows: aligned(8) class FisheyeOmnidirectionalVideoInfo() { bit(24) reserved = 0; unsigned int(8) num_circular_images; for(i=0; i< num_circular_images; i++) { unsigned int(8) viewid; bit(24) reserved = 0; } for(i=0; i< num circular images; i++) { unsigned int(32) image center x; unsigned int(32) image_center_y; Petition 870250044067, dated 05 / 28 / 2025, p. 40 / 185 36 / 82
[0093] view_id indicates the identifier of the view to which the circular image belongs. When there are only two view_id values, 0 and 1, for all circular images, circular images with view_id equal to 0 may belong to the left view, and circular images with view_id equal to 1 may belong to the right view. In some examples, the syntax element(s) that indicates the view identifier (i.e., the syntax element that indicates which circular image is the left view and which is the right view) may be the same as the syntax element that explicitly indicates whether the fisheye video data is stereoscopic or monoscopic.
[0094] According to one or more techniques of this disclosure, a file that includes fisheye video data may include a first box that contains a syntax structure containing parameters for the fisheye video data and, optionally, may include a second box that indicates coverage information for the fisheye video data. For example, the coverage information signaling for omnidirectional fisheye video may be added to the FisheyeOmnidirectionalVideoBox. For example, the CoverageInformationBox, as defined in OMAF DIS, may optionally be contained within the FisheyeOmnidirectionalVideoBox, and when not present in the FisheyeOmnidirectionalVideoBox, full sphere coverage for the fisheye video may be inferred. When CoverageInformationBox is present in the FisheyeOmnidirectionalVideoBox, the spherical region represented by the fisheye video images may be a region specified by two yaw circles and two Petition 870250044067, dated 05 / 28 / 2025, page 41 / 185 37 / 82 tilt circles. In this way, a client device can avoid having to infer coverage information from other parameters.
[0095] According to one or more techniques of this disclosure, a file that includes fisheye video data may include a first box that includes schema information, the first box may include a second box that includes a syntax structure containing parameters for the fisheye video data, and the first box may optionally include a third box that indicates whether the fisheye video data images are region-packed. For example, the RegionWisePackingBox may optionally be included in the SchemelnformationBox for omnidirectional fisheye video (e.g., when FisheyeOmnidirectionalVideoBox is present in the SchemelnformationBox, RegionWisePackingBox may be present in the same SchemelnformationBox). In this case, the presence of the RegionWisePackingBox indicates that the fisheye video images are region-compressed and require decompression before rendering.In this way, one or more benefits of region-specific packaging can be obtained in fisheye video.
[0096] According to one or more techniques of this disclosure, subpicture composition can be used to group omnidirectional fisheye video, and a specification regarding the application of subpicture composition grouping can be added to omnidirectional fisheye video. This specification can be applied when any of the tracks mapped to the subpicture composition track group has a type of Petition 870250044067, dated 05 / 28 / 2025, page 42 / 185 38 / 82 sample entry equal to 'resv' and a schema type equal to 'fodv' in the SchemeTypeBox included in the sample entry. In this case, each composite image is a packed image that has the fisheye format indicated by any FisheyeOmnidirectionalVideoBox and the region-wise packing format indicated by any RegionWisePackingBox included in the sample entries of the tracks mapped to the subimage composition track group. Additionally, the following may apply:
[0097] 1) Each track mapped to this grouping must have a sample entry type equal to 'resv'. The schema type must be equal to 'fodv' in the SchemeTypeBox included in the sample entry.
[0098] 2) The content of all instances of FisheyeOmnidirectionalVideoBox included in the sample entries of the tracks mapped to the same subimage composition track group must be identical.
[0099] 3) The content of all RegionWisePackingBox instances included in the sample entries of tracks mapped to the same subimage composition track group must be identical.
[00100] In HTTP streaming, the most frequently used operations include HEAD, GET, and partial GET. The HEAD operation retrieves a header of a file associated with a given Uniform Resource Locator (URL) or Uniform Resource Name (URN), without retrieving a payload associated with the URL or URN. The GET operation retrieves an entire file associated with a given URL or URN. The partial GET operation takes a range of bytes as an input parameter and retrieves a continuous number of bytes. Petition 870250044067, dated 05 / 28 / 2025, p. 43 / 185 39 / 82 of a file, where the number of bytes corresponds to the received byte range. Thus, movie fragments can be provided for HTTP streaming, because a partial GET operation can obtain one or more individual movie fragments. Within a movie fragment, there can be multiple track fragments of different ranges. In HTTP streaming, a media presentation can be a structured collection of data that is accessible to the client. The client can request and download media data information to present a streaming service to a user.
[00101] In the 3GPP streaming data example that uses HTTP streaming, there may be multiple representations for video and / or audio data of the multimedia content. As explained below, different representations may correspond to different encoding characteristics (e.g., different profiles or levels of a video encoding standard), different encoding standards or extensions of encoding standards (such as multiview and / or scalable extensions), or different bitrates. The manifest of such representations can be defined in a Media Presentation Description (MPD) data structure. A media presentation may correspond to a structured collection of data that is accessible to an HTTP streaming client device. The HTTP streaming client device may request and download media data information to present a streaming service to a user of the client device.A media presentation can be described in the MPD data structure, which may include MPD updates. Petition 870250044067, dated 05 / 28 / 2025, page 44 / 185 40 / 82
[00102] A media presentation may contain a sequence of one or more periods. Each period may extend to the beginning of the next period, or to the end of the media presentation in the case of the last period. Each period may contain one or more representations for the same media content. A representation may be one of a series of alternative encoded versions of audio, video, timed text, or other such data. Representations may differ by encoding types, for example, by bitrate, resolution, and / or codec for video data and bitrate, language, and / or codec for audio data. The term representation may be used to refer to a section of encoded audio or video data corresponding to a particular period of the multimedia content encoded in a specific manner.
[00103] Representations of a specific period can be assigned to a group indicated by an attribute in the MPD indicating an adaptation set to which the representations belong. Representations in the same adaptation set are generally considered as substitutes for each other, where a client device can dynamically and easily switch between these representations, for example, to perform bandwidth adaptation. For example, each representation of the video data for a given period can be assigned to the same adaptation set, so that any of the representations can be selected for decoding to present multimedia data, such as video data or audio data, of the multimedia content for the period. Petition 870250044067, dated 05 / 28 / 2025, page 45 / 185 41 / 82 corresponding. 0 media content within a period can be represented by any representation of group 0, if present, or the combination of, at most, one representation of each non-zero group, in some examples. Timing data for each representation of a period can be expressed relative to the period's start time.
[00104] A representation may include one or more segments. Each representation may include an initialization segment, or each segment of a representation may be self-initializing. When present, the initialization segment may contain initialization information for accessing the representation. In general, the initialization segment does not contain media data. A segment may be uniquely referenced by an identifier, such as a Uniform Resource Locator (URL), Uniform Resource Name (URN), or Uniform Resource Identifier (URI). MPD may provide the identifiers for each segment. In some instances, MPD may also provide byte ranges in the form of a range attribute, which may correspond to the data for a segment within a file accessible by URL, URN, or URI.
[00105] Different representations can be selected for the substantially simultaneous retrieval of different types of media data. For example, a client device might select an audio representation, a video representation, and a timed text representation from which to retrieve segments. In some examples, the client device might select specific sets of adaptations to perform the Petition 870250044067, dated 05 / 28 / 2025, page 46 / 185 42 / 82 bandwidth adaptation. That is, the client device can select an adaptation set including video representations, an adaptation set including audio representations, and / or an adaptation set including timed text. Alternatively, the client device can select adaptation sets for certain media types (e.g., video), and directly select representations for other media types (e.g., audio and / or timed text).
[00106] FIG. 9 is a block diagram illustrating an exemplary system 10 that implements techniques for continuous media data flow over a network. In this example, the system 10 includes the content preparation device 20, the server device 60, and the client device 40. The client device 40 and server device 60 are communicatively coupled by the network 74, which may comprise the Internet. In some examples, the content preparation device 20 and server device 60 may also be coupled by the network 74 or another network, or they may be directly communicatively coupled. In some examples, the content preparation device 20 and server device 60 may comprise the same device.
[00107] The content preparation device 20, in the example of FIG. 9, comprises an audio source 22 and a video source 24. The audio source 22 may comprise, for example, a microphone that produces electrical signals representative of captured audio data to be encoded by the audio encoder 26. Alternatively, the audio source 22 may comprise a media of Petition 870250044067, dated 05 / 28 / 2025, page 47 / 185 43 / 82 storage that stores previously recorded audio data, an audio data generator, such as a computerized synthesizer, or any other audio data source. The video source 24 may comprise a video camera that produces video data to be encoded by a video encoder 28, a storage medium encoded with previously recorded video data, a video data generation unit, such as a computer graphics source, or any other video data source. The content preparation device 20 is not necessarily communicatively coupled to the server device 60 in all examples, but may store the multimedia content to a separate medium that is read by the server device 60.
[00108] Audio and video data may comprise analog or digital data. Analog data may be digitized before being encoded by audio encoder 26 and / or video encoder 28. Audio source 22 may obtain audio data from a speaking participant while the speaking participant is speaking, and video source 24 may simultaneously obtain video data from the speaking participant. In other examples, audio source 22 may comprise a computer-readable storage medium comprising stored audio data, and video source 24 may comprise a computer-readable storage medium comprising stored video data. Thus, the techniques described in this disclosure may be applied to live, streaming, and video data. Petition 870250044067, dated 05 / 28 / 2025, page 48 / 185 44 / 82 real-time or archived, pre-recorded audio and video data.
[00109] Audio frames that correspond to video frames are generally audio frames containing audio data that was captured (or generated) by audio source 22 simultaneously with video data captured (or generated) by video source 24 that is contained within the video frames. For example, while a speaking participant generally produces audio data by speaking, audio source 22 captures the audio data and video source 24 captures the video data of the speaking participant at the same time, that is, while audio source 22 is capturing the audio data. Thus, an audio frame can temporally correspond to one or more specific video frames.Thus, an audio frame corresponding to a video frame generally corresponds to a situation where audio and video data were captured simultaneously, and for which an audio frame and a video frame comprise, respectively, the audio and video data that were captured at the same time.
[00110] In some examples, audio encoder 26 may encode a timestamp in each encoded audio frame that represents a moment when the audio data for the encoded audio frame was recorded, and similarly, video encoder 28 may encode a timestamp in each encoded video frame that represents a moment when the video data for the encoded video frame was recorded. In these examples, an audio frame corresponding to a frame of Petition 870250044067, dated 05 / 28 / 2025, page 49 / 185 45 / 82 video may comprise an audio frame comprising a timestamp and a video frame comprising the same timestamp. The content preparation device 20 may include an internal clock from which the audio encoder 26 and / or video encoder 28 may generate the timestamps, or which the audio source 22 and video source 24 may use to associate audio and video data, respectively, with a timestamp.
[00111] In some examples, audio source 22 may send data to audio encoder 26 corresponding to a time when the audio data was recorded, and video source 24 may send data to video encoder 28 corresponding to a time when the video data was recorded. In some examples, audio encoder 26 may encode a sequence identifier in the encoded audio data to indicate a relative temporal ordering of the encoded audio data, but without necessarily indicating an absolute time at which the audio data was recorded, and similarly, a video encoder 28 may also use sequence identifiers to indicate a relative temporal ordering of the encoded video data. Likewise, in some examples, a sequence identifier may be mapped or otherwise correlated with a timestamp.
[00112] Audio encoder 26 typically produces a stream of encoded audio data, while video encoder 28 produces a stream of encoded video data. Each individual data stream (whether audio or video) can be referred to as an elementary stream. One Petition 870250044067, dated 05 / 28 / 2025, page 50 / 185 46 / 82 An elementary stream is a single digitally encoded (possibly compressed) component of a representation. For example, the encoded video or audio portion of the representation can be an elementary stream. An elementary stream can be converted into a packetized elementary stream (PES) before being encapsulated within a video file. Within the same representation, the stream ID can be used to distinguish the PES packets belonging to one elementary stream from another. The basic data unit of an elementary stream is a packetized elementary stream (PES) packet. Thus, encoded video data generally corresponds to elementary video streams. Similarly, audio data corresponds to one or more respective elementary streams.
[00113] Many video coding standards, such as ITU-T H.264 / AVC and the High Efficiency Video Coding (HEVC) standard, define the syntax, semantics, and decoding process of error-free bitstreams, any of which corresponds to a given profile or level. Video coding standards generally do not specify the encoder, but the encoder has the task of ensuring that the generated bitstreams are compatible with a decoder's standard. In the context of video coding standards, a profile corresponds to a subset of algorithms, features or tools, and the constraints that apply to them. As defined by the H.264 standard, for example, a profile is a subset of the entire bitstream syntax that is specified by the H.264 standard. A level corresponds to the limitations on the decoder's resource consumption, such as, Petition 870250044067, dated 05 / 28 / 2025, page 51 / 185 47 / 82 for example, decoder memory and computation, which are related to image resolution, bit rate, and block processing rate. A profile can be signaled with a profile ide (profile indicator) value, while a level can be signaled with a level ide (level indicator) value.
[00114] The H.264 standard, for example, recognizes that, within the limits imposed by the syntax of a given profile, it is still possible to require a large variation in the performance of encoders and decoders depending on the values assumed by the syntax elements in the bitstream, such as the specified size of the decoded images. The H.264 standard also recognizes that, in many applications, it is neither practical nor economical to implement a decoder capable of handling all hypothetical uses of the syntax within a particular profile. Thus, the H.264 standard defines a level as a specific set of constraints imposed on the values of the syntax elements in the bitstream. These constraints can be simple limits on values. Alternatively, these constraints can take the form of constraints on arithmetic combinations of values (e.g., image width multiplied by image height multiplied by the number of images decoded per second).The H.264 standard also stipulates that individual implementations may support a different level for each supported profile.
[00115] A decoder according to a profile normally supports all the features defined in the profile. For example, as an encoding feature, the encoding of image B is not supported. Petition 870250044067, dated 05 / 28 / 2025, page 52 / 185 48 / 82 in the H.264 / AVC baseline profile, but it is supported in other H.264 / AVC profiles. A decoder according to a level must be able to decode any bitstream that does not require resources beyond the limitations defined in the level. The definitions of profiles and levels can be useful for ease of interpretation. For example, during video transmission, a pair of profile and level definitions can be negotiated and agreed upon for an entire transmission session. More specifically, in H.264 / AVC, a level can define limitations on the number of macroblocks that need to be processed, the size of the decoded image buffer (DPB), the encoded image buffer (CEC), the vertical motion vector range, the maximum number of motion vectors per two consecutive MBs, and whether a B-block can have sub-macroblock partitions smaller than 8x8 pixels. In this way, a decoder can determine if the decoder is capable of correctly decoding the bitstream.
[00116] In the example of FIG. 9, the encapsulation unit 30 of the content preparation device 20 receives elementary streams comprising encoded video data from the video encoder 28 and elementary streams comprising encoded audio data from the audio encoder 26. In some examples, the video encoder 28 and the audio encoder 26 may each include packers to form PES packets from encoded data. In other examples, the video encoder 28 and the audio encoder 26 may each interface with their respective packers to form PES packets from encoded data. In still other examples, the Petition 870250044067, dated 05 / 28 / 2025, page 53 / 185 49 / 82 encapsulation unit 30 may include packers to form PES packets of encoded audio and video data.
[00117] Video encoder 28 can encode video data from multimedia content in various ways, to produce different representations of the multimedia content at various bitrates and with various characteristics, such as pixel resolutions, frame rates, conformance to various encoding standards, conformance to various profiles and / or profile levels for various encoding standards, representations having one or more display modes (e.g., for two-dimensional or three-dimensional playback), or other such characteristics. A representation, as used in the present disclosure, may comprise audio data, video data, text data (e.g., for closed captions), or other similar data. The representation may include an elementary stream, such as an elementary audio stream or an elementary video stream. Each PES packet may include a stream ID that identifies the elementary stream to which the PES packet belongs.The encapsulation unit 30 is responsible for assembling elementary streams into video files (e.g., segments) of different representations.
[00118] The encapsulation unit 30 receives PES packets for elementary streams of a representation of the audio encoder 26 and video encoder 28 and forms corresponding network abstraction layer (NAL) units from the PES packets. Encoded video segments can be organized into NAL units, which Petition 870250044067, dated 05 / 28 / 2025, page 54 / 185 50 / 82 NALs provide a network-friendly video representation addressing applications such as video telephony, storage, broadcast, or streaming. NAL units can be categorized into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL units may contain the core compression engine and may include block, macroblock, and / or slice-level data. Other NAL units may be non-VCL NAL units. In some examples, an image encoded at a given point in time, typically presented as a primary encoded image, may be contained within an access unit, which may include one or more NAL units.
[00119] Non-VCL NAL units may include parameter set NAL units and SEI NAL units, among others. Parameter sets may contain sequence-level header information (in sequence parameter sets (SPS)) and rarely changing image-level header information (in image parameter sets (PPS)). With parameter sets (e.g., PPS and SPS), information that rarely changes does not need to be repeated for each sequence or image, thus coding efficiency can be improved. Furthermore, the use of parameter sets can allow out-of-band transmission of important header information, avoiding the need for redundant transmissions for error resilience. In out-of-band transmission examples, parameter set NAL units may be transmitted on a different channel than other NAL units, such as SEI NAL units. Petition 870250044067, dated 05 / 28 / 2025, p. 55 / 185 51 / 82
[00120] Supplemental Enhancement Information (SEI) may contain information that is not necessary for decoding the encoded image samples from NAL VCL units, but may assist in processes related to decoding, display, error resilience, and other purposes. SEI messages may be contained in non-VCL NAL units. SEI messages are the normative part of some standard specifications and are therefore not always mandatory for applying the standard-compliant decoder. SEI messages may be sequence-level SEI messages or image-level SEI messages. Some sequence-level information may be contained in SEI messages, such as scalability information SEI messages in the SVC example and view scalability information messages in MVC. These exemplary SEI messages may convey information about, for example, the extraction of operating points and operating point characteristics.Furthermore, encapsulation unit 30 can form a manifest file, such as a Media Presentation Descriptor (MPD), which describes the characteristics of the representations. Encapsulation unit 30 can format the MPD according to the Extensible Markup Language (XML).
[00121] The encapsulation unit 30 can provide data for one or more multimedia content representations, along with the manifest file (e.g., the MPD) for output interface 32. Output interface 32 can comprise a network interface or an interface for writing to a storage medium, such as Petition 870250044067, dated 05 / 28 / 2025, page 56 / 185 52 / 82 a Universal Serial Bus (USB) interface, a CD or DVD recorder, an interface for magnetic or flash storage media, or other interfaces for storing or transmitting media data. The encapsulation unit 30 can provide data from each of the multimedia content representations to the output interface 32, which can send the data to the server device 60 via network transmission or storage media. In the example of FIG. 9, the server device 60 includes storage media 62 that stores various multimedia contents 64, each including a respective manifest file 66 and one or more representations 68A68N (representations 68). In some examples, the output interface 32 can also send data directly to the network 74.
[00122] In some examples, 68 representations may be separated into adaptation sets. That is, various subsets of 68 representations may include respective common sets of features, such as codec, profile and level, resolution, number of views, file format for segments, text type information that may identify a language or other text features to be presented with the representation and / or audio data to be decoded and presented, for example, by loudspeakers, camera angle information that may describe a camera angle or real-world camera perspective of a scene for representations in the adaptation set, classification information that describes the suitability of the content for specific audiences, or similar. Petition 870250044067, dated 05 / 28 / 2025, page 57 / 185 53 / 82
[00123] Manifest file 66 may include data indicating the subsets of representations 68 corresponding to certain adaptation sets, as well as common characteristics for the adaptation sets. Manifest file 66 may also include data representing individual characteristics, such as bitrates, for the individual representations of adaptation sets. In this way, an adaptation set may provide a simplified network bandwidth adaptation. Representations in an adaptation set may be indicated through child elements of an adaptation set element in manifest file 66. In some examples, manifest file 66 may include some or all of the FisheyeOmnidirectionalVideoInfo() data discussed here, or similar data.Alternatively, the representation segments 68 may include some or all of the FisheyeOmnidirectionalVideoInfo data discussed here, or similar data.
[00124] Server device 60 includes request processing unit 70 and network interface 72. In some examples, server device 60 may include a plurality of network interfaces. Furthermore, any and all features of server device 60 may be implemented in other devices of a content delivery network, such as routers, bridges, proxy devices, switches, or other devices. In some examples, the intermediate devices of a content delivery network may cache multimedia content data 64, and include components that Petition 870250044067, dated 05 / 28 / 2025, page 58 / 185 54 / 82 substantially conform to those of the server device 60. In general, the network interface 72 is configured to send and receive data over the network 74.
[00125] The request processing unit 70 is configured to receive network requests from client devices, such as client device 40, for data from storage media 62. For example, the request processing unit 70 may implement the Hypertext Transfer Protocol (HTTP) version 1.1, as described in RFC 2616, Hypertext Transfer Protocol HTTP / 1.1, by R. Fielding et al, Network Working Group, IETF, June 1999. That is, the request processing unit 70 may be configured to receive HTTP GET requests or partial GET requests and provide multimedia content data 64 in response to the requests. Requests may specify a segment of one of the representations 68, for example, using a segment URL. In some examples, requests may also specify one or more byte ranges of the segment, thus comprising partial GET requests.The request processing unit 70 can also be configured for HTTP service HEAD requests to provide header data from a segment of one of the representations 68. In any case, the request processing unit 70 can be configured to process requests to provide requested data to a requesting device, such as client device 40.
[00126] Alternatively, order processing unit 70 can be configured to deliver media data via a protocol of Petition 870250044067, dated 05 / 28 / 2025, page 59 / 185 55 / 82 broadcast or multicast, such as eMBMS. Content preparation device 20 can create DASH segments and / or subsegments substantially in the same way as described, but server device 60 can provide these segments or subsegments using eMBMS or another broadcast or multicast network transport protocol. For example, request processing unit 70 can be configured to receive a multicast group participation request from client device 40. That is, server device 60 can advertise an Internet Protocol (IP) address associated with a multicast group to client devices, including client device 40, associated with specific media content (e.g., a live event broadcast). Client device 40, in turn, can submit a request to join the multicast group.This request can be propagated throughout the entire network 74, for example, the routers that make up the network 74, so that the routers direct the traffic destined for the IP address associated with the multicast group to the subscribing client devices, such as client device 40.
[00127] As illustrated in the example in FIG. 9, multimedia content 64 includes a manifest file 66, which may correspond to a media presentation description (MPD). The manifest file 66 may contain descriptions of different alternative representations 68 (e.g., video services with different qualities) and the description may include, for example, codec information, a profile value, a level value, a bitrate, and other descriptive characteristics of the representations 68. The Petition 870250044067, dated 05 / 28 / 2025, page 60 / 185 56 / 82 client device 40 can retrieve the MPD from a media presentation to determine how to access segments of representations 68.
[00128] In particular, retrieval unit 52 may retrieve configuration data (not shown) from client device 40 to determine the decoding capabilities of video decoder 48 and video output rendering capabilities 44. The configuration data may also include any or all of a language preference selected by a user of client device 40, one or more camera perspectives corresponding to depth preferences set by a user of client device 40, and / or a sorting preference selected by a user of client device 40. Retrieval unit 52 may comprise, for example, a web browser or a media client configured to send HTTP GET and partial GET requests. Retrieval unit 52 may correspond to software instructions executed by one or more processors or processing units (not shown) of client device 40.In some examples, all or part of the functionality described with respect to recovery unit 52 may be implemented in hardware, or a combination of hardware, software and / or firmware, where the required hardware may be provided to execute instructions for the software or firmware.
[00129] Recovery unit 52 can compare the decoding and rendering capabilities of the client device 40 to the representation characteristics 68 indicated by the file information. Petition 870250044067, dated 05 / 28 / 2025, page 61 / 185 57 / 82 manifest 66. For example, retrieval unit 52 can determine whether client device 40 (such as video output) is capable of rendering stereoscopic data or only monoscopic data. Retrieval unit 52 can initially retrieve at least a portion of manifest file 66 to determine representation characteristics 68. For example, retrieval unit 52 can request a portion of manifest file 66 that describes characteristics of one or more adaptation sets. Retrieval unit 52 can select a subset of representations 68 (e.g., an adaptation set) that has characteristics that can be satisfied by the encoding and rendering capabilities of client device 40.The recovery unit 52 can then determine bitrates for representations in the adaptation set, determine a currently available amount of network bandwidth, and retrieve segments from one of the representations that have a bitrate that can be satisfied by the network bandwidth. According to the techniques of this disclosure, the recovery unit 52 can retrieve data indicating whether the corresponding fisheye video data is monoscopic or stereoscopic. For example, the recovery unit 52 can retrieve an initial portion of a file (e.g., a segment) from one of the representations 68 to determine whether the file includes monoscopic or stereoscopic fisheye video data and determine whether it wants to retrieve video data from the file, or a different file from a different representation 68, depending on whether the client device 40 is capable of rendering the video data. Petition 870250044067, dated 05 / 28 / 2025, page 62 / 185 58 / 82 fisheye from the file.
[00130] In general, higher bitrate representations can produce superior quality video playback, while lower bitrate representations can offer sufficient quality video playback when available network bandwidth decreases. Thus, when available network bandwidth is relatively high, the retrieval unit 52 can retrieve data from relatively high bitrate representations, while when available network bandwidth is low, the retrieval unit 52 can retrieve data from relatively low bitrate representations. In this way, the client device 40 can transmit multimedia data over the network 74, while also adapting to changes in network bandwidth availability 74.
[00131] Alternatively, retrieval unit 52 can be configured to receive data according to a broadcast or multicast network protocol, such as eMBMS or IP multicast. In these examples, retrieval unit 52 can submit a request to join a multicast network group associated with specific media content. After joining the multicast group, retrieval unit 52 can receive data from the multicast group without additional requests being issued to the server device 60 or content preparation device 20. Retrieval unit 52 can submit a request to leave the multicast group when the multicast group data is no longer needed, for example, to stop playback or to change channels. Petition 870250044067, dated 05 / 28 / 2025, page 63 / 185 59 / 82 different multicast group.
[00132] Network interface 54 can receive and provide segment data of a selected representation to recovery unit 52, which in turn can provide the segments to decapsulation unit 50. Decapsulation unit 50 can decapsulate elements of a video file into constituent PES streams, unpack the PES streams to recover encoded data, and send the encoded data either to audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio or video sequence, for example, as indicated by headers of the stream's PES packets. Audio decoder 46 decodes the encoded audio data and sends the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and sends the decoded video data, which may include a plurality of views of a stream, to video output 44.
[00133] The video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, recovery unit 52 and decapsulation unit 50 can each be implemented as any one of a variety of suitable processing circuits, as the case may be, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware or any combination thereof. Each Petition 870250044067, dated 05 / 28 / 2025, page 64 / 185 60 / 82 video encoder 28 and video decoder 48 may be included in one or more encoders or decoders, each of which may be integrated as part of a combined video encoder / decoder (codec). Similarly, each of the video encoder 26 and video decoder 46 may be included in one or more encoders or decoders, each of which may be integrated as part of a combined CODEC. Equipment including a video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, recovery unit 52 and / or decapsulation unit 50 may comprise an integrated circuit, a microprocessor and / or a wireless communication device, such as a mobile phone.
[00134] Client device 40, server device 60, and / or content preparation device 20 can be configured to operate in accordance with the techniques of this disclosure. For example purposes, this disclosure describes these techniques with respect to client device 40 and server device 60. However, it should be understood that content preparation device 20 can be configured to perform these techniques instead of (or in addition to) server device 60.
[00135] The encapsulation unit 30 can form NAL units comprising a header that identifies a program to which the NAL unit belongs, as well as a payload, for example, audio data, video data, or data describing the transport or flow. Petition 870250044067, dated 05 / 28 / 2025, page 65 / 185 61 / 82 program to which the NAL unit corresponds. For example, in H.264 / AVC, a NAL unit includes a 1-byte header and a variable-size payload. A NAL unit including video data in its payload can comprise various levels of video data granularity. For example, a NAL unit can comprise a video data block, a plurality of blocks, a video data slice, or an entire video data image. The encapsulation unit 30 can receive encoded video data from the video encoder 28, in the form of PES packets of elementary streams. The encapsulation unit 30 can associate each elementary stream with a corresponding program.
[00136] The encapsulation unit 30 can also assemble access units from a plurality of NAL units. In general, an access unit may comprise one or more NAL units to represent a video data frame, as well as audio data corresponding to the frame when such audio data is available. An access unit generally includes all NAL units for an output time instant, for example, all audio and video data for a time instant. For example, if each view has a frame rate of 20 frames per second (fps), then each time instant may correspond to a time interval of 0.05 seconds. During this time interval, the specific frames for all views of the same access unit (the same time instant) can be processed simultaneously.In one example, an access unit might comprise an image encoded at a given point in time, which can be presented as an encoded image. Petition 870250044067, dated 05 / 28 / 2025, page 66 / 185 62 / 82 primary.
[00137] Consequently, an access unit can comprise all audio and video frames from a common time instant, for example, all views corresponding to time X. This revelation also refers to an encoded image of a specific view as a display component. That is, a display component can include an encoded image (or frame) for a specific view at a specific time. Thus, an access unit can be defined as comprising all display components from a common time instant. The decoding order of the access units does not necessarily have to be the same as the output or display order.
[00138] A media presentation may include a media presentation description (MPD), which may contain descriptions of different alternative representations (e.g., video services with different qualities) and the description may include, for example, codec information, a profile value, and a level value. An MPD is an example of a manifest file, such as manifest file 66. The client device 40 may retrieve the MPD of a media presentation to determine how to access movie fragments from the various presentations. Movie fragments may be located in the movie fragment boxes (moof boxes) of the video files.
[00139] The manifest file 66 (which may include, for example, an MPD) may announce the availability of the segments of the representations 68. That is, the MPD may include information indicating the time of Petition 870250044067, dated 05 / 28 / 2025, page 67 / 185 63 / 82 wall clock in which a first segment of one of the representations 68 becomes available, as well as information indicating the durations of the segments within the representations 68. In this way, the retrieval unit 52 of the client device 40 can determine when each segment is available, based on the start time as well as the durations of the segments preceding a specific segment.
[00140] After encapsulation unit 30 has assembled NAL units and / or access units on a video file based on the received data, encapsulation unit 30 passes the video file to output interface 32 for output. In some examples, encapsulation unit 30 may store the video file locally or send the video file to a remote server via output interface 32, instead of sending the video file directly to the client device 40. Output interface 32 may comprise, for example, a transmitter, a transceiver, a device for recording data on computer-readable media such as, for example, an optical drive, a magnetic media drive (e.g., floppy disk drive), a Universal Serial Bus (USB), a network interface, or another output interface.The output interface 32 outputs the video file to a computer-readable medium, such as a broadcast signal, magnetic media, optical media, memory, a flash drive, or other computer-readable media.
[00141] Network interface 54 can receive a NAL unit or access unit via network 74 and Petition 870250044067, dated 05 / 28 / 2025, page 68 / 185 64 / 82 provides the NAL unit or access unit to the decapsulation unit 50, via the recovery unit 52. The decapsulation unit 50 can decapsulate elements of a video file into constituent PES streams, unpack the PES streams to recover encoded data, and send the encoded data either to the audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio or video sequence, for example, as indicated by the headers of the stream's PES packets. The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, while the video decoder 48 decodes the encoded video data and sends the decoded video data, which may include a plurality of views of a stream, to the video output 44.
[00142] In this way, client device 40 represents an example of a device for retrieving media data, the device including a device for retrieving fisheye video data files, as described above and / or as claimed below. For example, client device 40 can retrieve fisheye video data files and / or render fisheye video based on a determination of whether the fisheye video is monoscopic or stereoscopic, and this determination can be based on a syntax element that explicitly specifies whether the fisheye video is monoscopic or stereoscopic.
[00143] Similarly, content preparation device 20 represents a device for generating fisheye video data files, as described Petition 870250044067, dated 05 / 28 / 2025, page 69 / 185 65 / 82 above and / or as claimed below. For example, content preparation device 20 may include, in fisheye video data, a syntax element that explicitly specifies whether the fisheye video is monoscopic or stereoscopic.
[00144] FIG. 10 is a conceptual diagram illustrating elements of exemplary multimedia content 120. Multimedia content 120 may correspond to multimedia content 64 (FIG. 9), or other multimedia content stored on storage media 62. In the example in FIG. 10, multimedia content 120 includes media presentation description (MPD) 122 and a plurality of representations 124A-124N (representations 124). Representation 124A includes optional header data 126 and segments 128A-128N (segments 128), while representation 124N includes optional header data 130 and segments 132A-132N (segments 132). The letter N is used to designate the last film fragment in each of the representations 124, for convenience. In some examples, there may be a different number of film fragments among the 124 representations.
[00145] MPD 122 may comprise a data structure separate from representations 124. MPD 122 may correspond to manifest file 66 of FIG. 9. In general, MPD 122 may include data that typically describes characteristics of representations 124, such as encoding and rendering characteristics, adaptation sets, a profile to which MPD 122 corresponds, text type information, camera angle information, Petition 870250044067, dated 05 / 28 / 2025, page 70 / 185 66 / 82 classification information, trick mode information (e.g., information indicating representations that include temporal subsequences), and / or information to retrieve remote periods (e.g., for the insertion of targeted advertising into media content during playback).
[00146] Header 126 data, when present, may describe the characteristics of segments 128, for example, temporal locations of random access points (RAPs, also known as flow access points (SAPs)), which of the segments 128 includes random access points, byte offsets for the random access points within segments 128, uniform resource locators (URLs) of segments 128, or other aspects of segments 128. Header 130 data, when present, may describe similar characteristics for segments 132. Additionally or alternatively, these characteristics may be fully included within MPD 122.
[00147] Segments 128 and 132 include one or more encoded video samples, each of which may include frames or slices of video data. Each of the encoded video samples of segment 128 may have similar characteristics, for example, height, width, and bandwidth requirements. These characteristics may be described by the MPD 122 data, although such data is not illustrated in the example in FIG. 10. MPD 122 may include the characteristics as described by the 3GPP specification, with the addition of any or all of the signaled information described herein. Petition 870250044067, dated 05 / 28 / 2025, page 71 / 185 67 / 82 revelation.
[00148] Each of segments 128, 132 can be associated with a unique Uniform Resource Locator (URL). Thus, each of segments 128, 132 can be independently retrieved using a streaming network protocol such as DASH. In this way, a target device, such as client device 40, can use an HTTP GET request to retrieve segments 128 or 132. In some examples, client device 40 can use partial HTTP GET requests to retrieve specific byte ranges from segments 128 or 132.
[00149] FIG. 11 is a block diagram illustrating elements of an example video file 150, which may correspond to a segment of a representation, such as one of the segments 114, 124 of FIG. 10. Each of the segments 128, 132 may include data that substantially conforms to the arrangement of the data illustrated in the example of FIG. 11. It can be said that video file 150 encapsulates a segment. As described above, video files according to the ISO base media file format and their extensions store data in a series of objects, called boxes. In the example of FIG. 11, the video file 150 includes file type box (FTYP) 152, movie box (MOOV) 154, segment index boxes (sidx) 162, movie fragment boxes (MOOF) 164 and movie fragment random access box (MFRA) 166. Although FIG.11 represents an example of a video file; it should be understood that other media files may include other types of media data (e.g., audio data, programmed text data, or...). Petition 870250044067, dated 05 / 28 / 2025, page 72 / 185 68 / 82 similar) that are structured similarly to the video file data 150, according to the ISO base media file format and its extensions.
[00150] The file type box (FTYP) 152 generally describes a file type for video file 150. The file type box 152 may include data that identifies a specification describing a best use for video file 150. The file type box 152 may alternatively be placed before the MOOV box 154, the film fragment boxes 164 and / or the MFRA box 166.
[00151] In some examples, a segment, such as video file 150, may include an MPD update box (not shown) before the FTYP box 152. The MPD update box may include information indicating that an MPD corresponding to a representation including video file 150 should be updated, along with information to update the MPD. For example, the MPD update box may provide a URI or URL to a resource to be used to update the MPD. As another example, the MPD update box may include data to update the MPD. In some examples, the MPD update box may immediately follow a segment type (STYP) box (not shown) of video file 150, where the STYP box may define a segment type for video file 150.
[00152] The MOOV 154 box, in example 2 of FIG. 11, includes the movie header box (MVHD) 156, the track box (TRAK) 158 and one or more movie extension boxes (MVEX) 160. In general, the MVHD 156 box can Petition 870250044067, dated 05 / 28 / 2025, page 73 / 185 69 / 82 describe general characteristics of the video file 150. For example, the MVHD 156 box may include data describing when the video file 150 was originally created, when the video file 150 was last modified, a timescale for the video file 150, a playback duration of the video file 150, or other data that generally describes the video file 150.
[00153] The TRAK 158 box may include data for a track of the video file 150. The TRAK 158 box may include a track header box (TKHD) that describes characteristics of the track corresponding to the TRAK 158 box. In some examples, the TRAK 158 box may include encoded video images, while in other examples, the encoded video images of the track may be included in the film fragments 164, which may be referenced by data from the TRAK 158 box and / or sidx boxes 162.
[00154] In some examples, video file 150 may include more than one track. Therefore, MOOV box 154 may include a number of TRAK boxes equal to the number of tracks in video file 150. TRAK box 158 may describe characteristics of a track from the corresponding video file 150. For example, TRAK box 158 may describe temporal and / or spatial information for the corresponding track. A TRAK box similar to TRAK box 158 of MOOV box 154 may describe characteristics of a parameter set track, when encapsulation unit 30 (FIG. 9) includes a parameter set track in a video file, such as video file 150. Petition 870250044067, dated 05 / 28 / 2025, page 74 / 185 70 / 82 150. The encapsulation unit 30 can signal the presence of sequence-level SEI messages in the parameter set range within the TRAK box that describes the parameter set range.
[00155] MVEX boxes 160 can describe characteristics of the corresponding film fragments 164, for example, to signal that the video file 150 includes film fragments 164, in addition to the video data included in the MOOV box 154, if any. In the context of streaming video data, encoded video images can be included in the film fragments 164 instead of in the MOOV box 154. Consequently, all encoded video samples can be included in the film fragments 164 instead of in the MOOV box 154.
[00156] The MOOV 154 box may contain a number of MVEX 160 boxes equal to the number of film fragments 164 in the video file 150. Each of the MVEX 160 boxes may describe characteristics of one of the corresponding video fragments 164. For example, each MVEX box may include a Movie Extension Header (MEHD) box that describes a temporal duration for a corresponding film fragment 164.
[00157] Furthermore, according to the techniques of this disclosure, video file 150 may include a Fisheye Omnidirectional Video Box in a Scheme Information Box, which may be included in the MOOV box 154. In some examples, the Fisheye Omnidirectional Video Box may be included in the TRAK box 158, if different tracks of video file 150 may include monoscopic or stereoscopic fisheye video data. In some examples, the Petition 870250044067, dated 05 / 28 / 2025, page 75 / 185 71 / 82 Fisheye Omnidirectional VideoBox can be included in the FOV 157 box.
[00158] As noted above, encapsulation unit 30 can store a sequence data set in a video sample that does not include actual encoded video data. A video sample can generally correspond to an access unit, which is a representation of an encoded image at a specific time instance. In the context of AVC, the encoded image may include one or more VCL NAL units containing the information to construct all the pixels of the access unit and other associated non-VCL NAL units, such as SEI messages. Therefore, encapsulation unit 30 can include a sequence data set, which may include sequence-level SEI messages, in one of the film fragments 164.The encapsulation unit 30 can also signal the presence of a sequence data set and / or SEI messages at the sequence level as being present in one of the film fragments 164 within one of the MVEX boxes 160 corresponding to the film fragments 164.
[00159] SIDX boxes 162 are optional elements of the video file 150. That is, video files conforming to the 3GPP file format or other file formats do not necessarily include SIDX boxes 162. Following the example of the 3GPP file format, a SIDX box can be used to identify a subsegment of a segment (for example, a segment contained in video file 150). The 3GPP file format defines a subsegment as a set Petition 870250044067, dated 05 / 28 / 2025, page 76 / 185 72 / 82 independent of one or more consecutive movie fragment boxes with corresponding media data boxes, and a media data box containing data referenced by a movie fragment box must follow the movie fragment and precede the next movie fragment box containing information about the same track. The 3GPP file format also indicates that a SIDX box contains a sequence of references to subsegments of the (sub)segment documented by the box. The referenced subsegments are contiguous at presentation time. Similarly, the bytes referenced by a segment index box are always contiguous within the segment. The referenced size provides the count of the number of bytes in the referenced material.
[00160] SIDX boxes 162 generally provide representative information for one or more subsegments of a segment included in the video file 150. For example, this information may include playback times at which the subsegments begin and / or end, byte offsets for the subsegments, whether the subsegments include (e.g., begin with) a stream access point (SAP), a type for the SAP (e.g., whether the SAP is an instant decoder update (IDR) image, a clean random access (CRA) image, a broken link access (BLA) image, or similar), a position of the SAP (in terms of playback time and / or byte offset) in the subsegment, and the like.
[00161] Film fragments 164 may include one or more encoded video images. In some examples, film fragments 164 may include one or Petition 870250044067, dated 05 / 28 / 2025, page 77 / 185 73 / 82 more picture groups (GOPs), each of which may include a number of encoded video images, for example, frames or images. Additionally, as described above, the 164 film fragments may include sequence data sets in some instances. Each of the 164 film fragments may include a Film Fragment Header Box (MFHD, not shown in FIG. 11). The MFHD box may describe characteristics of the corresponding film fragment, such as a sequence number for the film fragment. The 164 film fragments may be included in sequence number order in the 150 video file.
[00162] The MFRA 166 box can describe random access points within the 164 movie fragments of the 150 video file. This can aid in the execution of trick modes, such as performing searches at specific temporal locations (i.e., playback times) within a segment encapsulated by the 150 video file. The MFRA 166 box is generally optional and does not need to be included in video files, in some instances. Similarly, a client device, such as client device 40, does not necessarily need to reference the MFRA 166 box to correctly decode and display the video data from the 150 video file. The MFRA 166 box can include a number of track fragment random access (TFRA) boxes (not shown) equal to the number of tracks in the 150 video file or, in some instances, equal to the number of media tracks (e.g., tracks without cue) in the 150 video file.
[00163] In some examples, the fragments of Petition 870250044067, dated 05 / 28 / 2025, page 78 / 185 74 / 82 film 164 may include one or more stream access points (SAPs), such as IDR images. Similarly, the MFRA box 166 may provide indications of locations in video file 150 of the SAPs. Consequently, a temporal subsequence of video file 150 may be formed from SAPs of video file 150. The temporal subsequence may also include other figures, such as P-frames and / or B-frames that depend on the SAPs. The frames and / or slices of the temporal subsequence may be arranged within segments, so that frames / slices of the temporal subsequence that depend on other frames / slices of the subsequence can be decoded appropriately. For example, in hierarchical data arrangement, data used for prediction for other data may also be included in the temporal subsequence.
[00164] FIG. 12 is a flowchart illustrating an example technique for processing a file that includes fisheye video data, according to one or more techniques of this disclosure. The techniques in FIG. 12 are described as being performed by client device 40 of FIG. 9, although devices with configurations different from client device 40 may perform the technique in FIG. 12.
[00165] Client device 40 can receive a file including fisheye video data and a syntax structure that includes a plurality of syntax elements that specify attributes of the fisheye video data (1202). As discussed above, the plurality of syntax elements can include a first syntax element that explicitly indicates whether the video data Petition 870250044067, dated 05 / 28 / 2025, p. 79 / 185 75 / 82 fisheye are monoscopic or stereoscopic and one or more syntax elements implicitly indicate whether the fisheye video data is monoscopic or stereoscopic. The one or more syntax elements implicitly indicating whether the fisheye video data is monoscopic or stereoscopic may be syntax elements that explicitly indicate extrinsic parameters of the cameras used to capture the fisheye video data. In some examples, the one or more syntax elements may be, or may be similar to, syntax elements included in the FisheyeOmnidirectionalVideoInfo() syntax structure in the current OMAF DIS draft. In some examples, the first syntax element may be included in an initial bit set of the syntax structure (e.g., initial 24 bits of the FisheyeOmnidirectionalVideoInfo() syntax structure).
[00166] Client device 40 can determine, based on the first syntax element, whether the fisheye video data is monoscopic or stereoscopic (1204). A value of the first syntax element can explicitly indicate whether the fisheye video data is monoscopic or stereoscopic. As an example, where the first syntax element has a first value, client device 40 can determine that the fisheye video data is monoscopic (i.e., that the circular images included in the fisheye video data images have aligned optical axes and are facing opposite directions). As another example, where the second syntax element has a first value, client device 40 can determine that the fisheye video data is Petition 870250044067, dated 05 / 28 / 2025, p. 80 / 185 76 / 82 stereoscopic (i.e., the circular images included in the fisheye video data have parallel optical axes and are facing the same direction)
[00167] As discussed above, although it may be possible for client device 40 to determine whether fisheye video data is monoscopic or stereoscopic based on syntax elements that explicitly indicate extrinsic parameters of the cameras used to capture the fisheye video data, such a calculation may increase the computational load on client device 40. Therefore, to reduce the calculations performed (and thus the computational resources used), client device 40 can determine whether fisheye video data is monoscopic or stereoscopic based on the first syntax element.
[00168] Client device 40 can render fisheye video data based on the determination. For example, in response to the determination that the fisheye video data is monoscopic, client device 40 can render the fisheye data as monoscopic (1206). For example, client device 40 can determine a viewport (i.e., a portion of the sphere in which a user is looking), identify a portion of the circular images of the fisheye video data that corresponds to the viewport, and display the same portion of the circular images to each of the viewer's eyes. Similarly, in response to the determination that the fisheye video data is stereoscopic, client device 40 can render the fisheye data as stereoscopic (1208). For example, the device Petition 870250044067, dated 05 / 28 / 2025, page 81 / 185 77 / 82 client 40 can determine a viewport (that is, a portion of the sphere in which a user is looking), identify a corresponding portion of each of the circular images in the fisheye video data that correspond to the viewport, and display the corresponding portions of the circular images to the viewer's eyes.
[00169] FIG. 13 is a flowchart illustrating an example technique for generating a file that includes fisheye video data, according to one or more techniques of this disclosure. The techniques in FIG. 13 are described as being performed by content preparation device 20 of FIG. 9, although devices with configurations different from content preparation device 20 may perform the technique in FIG. 13.
[00170] Content preparation device 20 can obtain fisheye video data and extrinsic parameters from cameras used to capture the fisheye video data (1302). For example, content preparation device 20 can obtain images from the fisheye video data that are encoded using a video codec, such as HEVC. Each of the images from the fisheye video data can include a plurality of circular images that correspond to an image captured by a different camera with a fisheye lens (for example, an image can include a first circular image captured through the fisheye lens 12A of FIG. IA and a second circular image captured through the fisheye lens 12B of FIG. IA). The extrinsic parameters can specify various attributes of the cameras. For example, the extrinsic parameters can specify a yaw angle, a Petition 870250044067, dated 05 / 28 / 2025, page 82 / 185 78 / 82 tilt angle, a rotation angle, and one or more spatial displacements of each of the cameras used to capture the fisheye video data.
[00171] Content preparation device 20 can determine, based on extrinsic parameters, whether fisheye video data is monoscopic or stereoscopic (1304). As an example, when two sets of camera extrinsic parameter values are as follows, content preparation device 20 can determine that the fisheye video is stereoscopic: Ioconjunto: camera_center_yaw = 0 degrees (+ / - 5 degrees) camera_center _pitch = 0 degrees (+ / - 5 degrees) camera_center roll = 0 degrees (+ / - 5 degrees) camera_center_offset_x = 0 mm (+ / - 3 mm) camera_center_offset_y = 0 mm (+ / - 3 mm) camera_center_offset_z = 0 mm (+ / - 3 mm) 2nd set: camera_center_yaw = 0 degrees (+ / - 5 degrees) camera_center _pitch = 0 degrees (+ / - 5 degrees) camera_center roll = 0 degrees (+ / - 5 degrees) camera_center_offset_x = 64 mm (+ / - 3 mm) camera_center_offset_y = 0 mm (+ / - 3 mm) camera_center_offset_z = 0 mm (+ / - 3 mm)
[00172] As another example, when two sets of extrinsic camera parameter values are as follows, content preparation device 20 can determine that the fisheye video is monoscopic: Ioconjunto: camera_center_yaw = 0 degrees (+ / - 5 degrees) Petition 870250044067, dated 05 / 28 / 2025, page 83 / 185 79 / 82 camera_center_pitch = 0 degrees (+ / - 5 degrees) camera_center_roll = 0 degrees (+ / - 5 degrees) camera_ center_offset_x = 0 mm ( + / - 3 mm camera_ center_offset_y = 0 mm ( + / - 3 mm camera center offset z = 0 mm ( + / - 3 mm 2nd set: camera_center_yaw = 180 degrees (+ / - 5 degrees) camera_center _pitch = 0 degrees (+ / - 5 degrees) camera_center roll = 0 degrees (+ / - 5 degrees) camera_center_offset_x = 0 mm (+ / - 3 mm) camera_center_offset_y = 0 mm (+ / - 3 mm) camera_center_offset_z = 0 mm (+ / - 3 mm)
[00173] 0 The content preparation device may encode in a file, the fisheye video data and a syntax structure including a plurality of syntax elements that specify attributes of the fisheye video data (1306). The plurality of syntax elements includes: a first syntax element that explicitly indicates whether the fisheye video data is monoscopic or stereoscopic, and one or more syntax elements that explicitly indicate the extrinsic parameters of the cameras used to capture the fisheye video data.
[00174] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored in or transmitted via one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. The medium Petition 870250044067, dated 05 / 28 / 2025, page 84 / 185 Computer-readable media may include computer-readable storage media, which corresponds to tangible media, such as data storage media or communication media, including any means that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, computer-readable media may generally correspond to (1) tangible computer-readable storage media that is non-transient or (2) a communication medium, such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for the implementation of the techniques described in this disclosure. A computer program product may include computer-readable media.
[00175] By way of example, and not as a limitation, such computer-readable storage media may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory or any other means that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly called a computer-readable medium. For example, if instructions are transmitted from a site, server, or other remote source via a coaxial cable, fiber optic cable, twisted pair, line of Petition 870250044067, dated 05 / 28 / 2025, page 85 / 185 81 / 82 Digital subscriber (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of transmission media. It should be understood, however, that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but are instead directed at tangible, non-transient storage media. Disk and diskette, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc where floppy disks generally reproduce data magnetically, while disks reproduce data optically with lasers. Combinations of the foregoing should also be included within the scope of computer-readable media.
[00176] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent discrete or logic circuits. Consequently, the term processor as used herein may refer to any of the preceding structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some respects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a Petition 870250044067, dated 05 / 28 / 2025, page 86 / 185 82 / 82 combined codec. Also, the techniques can be fully implemented in one or more circuits or logic elements.
[00177] The techniques of the present disclosure can be implemented in a wide variety of devices or apparatus, including a wireless monotone, an integrated circuit (TC), or an array of ICs (e.g., a chip set). Various components, modules, or units are described in the present disclosure to emphasize the functional aspects of devices configured to perform the disclosed techniques, but do not necessarily need to be implemented by different hardware units. Instead, as described above, various units can be combined into a codec hardware unit or provided by a collection of interoperable hardware units, including one or more processors, as described above, in conjunction with the appropriate software and / or firmware.
[00178] Several examples have been described. These and other examples are within the scope of the following claims. Petition 870250044067, dated 05 / 28 / 2025, page 87 / 185
Claims
1 / 6 CLAIMS 1. A method for processing a file including video data, the method characterized in that it comprises: processing a file including fisheye video data, the file including a FisheyeOmnidirectionalVideoInfo syntax structure of the OMAF standard including a plurality of syntax elements specifying attributes of the fisheye video data, wherein the plurality of syntax elements includes: a first syntax element explicitly indicating whether the fisheye video data is monoscopic or stereoscopic (1202), and one or more syntax elements implicitly indicating whether the fisheye video data is monoscopic or stereoscopic, wherein the first syntax element is enclosed in a set of initial bits of syntax structure, wherein the set of initial bits is 24 bits long; determining, based on the first syntax element, whether the fisheye video data is monoscopic or stereoscopic (1204);and output, based on the determination, the fisheye video data to render as monoscopic (1206) or stereoscopic (1208).
2. Method, according to claim 1, characterized in that the file includes a box that contains the syntax structure.
3. Method according to claim 2, characterized in that the box is a first box that is included in a second box that includes schema information, the method further comprising: Petition 870260061807, dated 06 / 24 / 2026, page 6 / 31 2 / 6 determining whether the first box includes a third box that indicates whether images from the fisheye video data are packaged from region to region.
4. A method according to claim 3, characterized in that it further comprises: in response to the determination that the first box includes the third box, unpacking the images from the fisheye video data before rendering the images from the fisheye video data; or in response to the determination that the first box does not include the third box, rendering the images from the fisheye video data without unpacking the images from the fisheye video data.
5. Method according to claim 4, characterized in that the first box is a SchemelInformationBox, the second box is a FisheyeOmnidirectionalVideoBox and the third box is a RegionWisePackingBox.
6. Method according to claim 1, characterized in that the syntax structure additionally includes a second syntax element that specifies a number of circular images included in each image of the fisheye video data.
7. Method, according to claim 6, characterized in that the syntax structure comprises, for each respective circular image, a respective third syntax element indicating a view identifier of the respective circular image.
8. Method, according to claim 1, characterized in that the syntax structure is external to the video encoding layer data, VCL, encapsulated by the file.
9. A method according to claim 1, characterized in that determining whether fisheye video data is monoscopic or stereoscopic comprises: determining, based on the first syntax element and independently of syntax elements that implicitly indicate whether fisheye video data is monoscopic or stereoscopic, whether the fisheye video data is monoscopic or stereoscopic.
10. Method for generating a file including video data, the method characterized in that it comprises: obtaining fisheye video data and extrinsic parameters from cameras used to capture the fisheye video data (1302); determining, based on the extrinsic parameters, whether the fisheye video data is monoscopic or stereoscopic (1304); and encoding, in a file, the fisheye video data and a FisheyeOmnidirectionalVideoInfo syntax structure of the OMAF standard including a plurality of syntax elements that specify attributes of the fisheye video data (1306), wherein the plurality of syntax elements includes: a first syntax element that explicitly indicates whether the fisheye video data is monoscopic or stereoscopic and one or more syntax elements that implicitly indicate whether the fisheye video data is monoscopic or stereoscopic. Petition 870260061807, dated 06 / 24 / 2026, p.8 / 31 4 / 6 stereoscopic, where encoding the first syntax element comprises encoding the first syntax element in a set of initial bits of the syntax structure, where the set of initial bits is 24 bits long.
11. Method according to claim 10, characterized in that the file includes a box that contains the syntax structure.
12. Device for processing video data, the device characterized in that it comprises: a memory configured to store at least a portion of a file including fisheye video data, the file including a FisheyeOmnidirectionalVideoInfo syntax structure of the OMAF standard including a plurality of syntax elements that specify attributes of the fisheye video data, wherein the plurality of syntax elements includes: a first syntax element that explicitly indicates whether the fisheye video data is monoscopic or stereoscopic and one or more syntax elements that implicitly indicate whether the fisheye video data is monoscopic or stereoscopic; wherein the first syntax element is enclosed in a set of initial bits of syntax structure, wherein the set of initial bits is 24 bits long;and one or more processors configured to: determine, based on the first syntax element, whether the fisheye video data is monoscopic or stereoscopic; and output, based on the determination, the fisheye video data to render as monoscopic or stereoscopic.
13. Device for generating a file including video data, the device characterized in that it comprises: a memory configured to store fisheye video data; and one or more processors configured to: obtain extrinsic parameters from cameras used to capture the fisheye video data (1302); determine, based on the extrinsic parameters, whether the fisheye video data is monoscopic or stereoscopic;and to encode, in a file, the fisheye video data and a FisheyeOmnidirectionalVideoInfo syntax structure of the OMAF standard including a plurality of syntax elements that specify attributes of the fisheye video data, wherein the plurality of syntax elements includes: a first syntax element that explicitly indicates whether the fisheye video data is monoscopic or stereoscopic and one or more syntax elements that implicitly indicate whether the fisheye video data is monoscopic or stereoscopic, wherein, to encode the first syntax element, the one or more processors are configured to encode the first syntax element in a set of initial bits of the syntax structure, wherein the set of initial bits is 24 bits long.
14. Device according to claim 12 or 13, characterized in that it further comprises means for carrying out the method as defined in any one of claims 2 to 11.