Media file encapsulation methods, media file decapsulation methods, and related equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-14
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]相关技术无法在文件封装中对不同的应用场景进行区分,这给用户侧的处理带来了不必要的麻烦
[0016]在本公开的一些实施例所提供的技术方案中,通过在生成对应应用场景下的媒体码流的封装文件时,在封装文件中扩展第一应用场景类型字段,通过该第一应用场景类型字段来指示该媒体码流对应的应用场景,由此使得在媒体文件的封装中即可对不同媒体码流的不同应用场景进行区分,一方面,将该封装文件发送至第一设备时,该第一设备可以根据该封装文件中的第一应用场景类型字段即可区分该媒体码流的应用场景,从而可以根据该媒体码流对应的应用场景确定对该媒体码流采取何种解码或者渲染方式,可以节约第一设备的运算能力和资源;另一方面,由于在封装阶段即可确定媒体码流的应用场景,因此即使第一设备不具备媒体码流的解码能力,也可以确定该媒体码流对应的应用场景,而不需要等到解码该媒体码流之后才能区分。
Smart Images

Figure CN116248642B_ABST
Abstract
Description
[0001] This application is a divisional application filed on October 14, 2020, with application number 202011098190.7 and invention title "Media File Packaging Method, Media File Depackaging Method and Related Equipment". Technical Field
[0002] This disclosure relates to the field of media file encapsulation and decapsulation technology, and more specifically, to a media file encapsulation method, a media file decapsulation method, a media file encapsulation apparatus, a media file decapsulation apparatus, an electronic device, and a computer-readable storage medium. Background Technology
[0003] Immersive media refers to media content that provides users with an immersive experience; it can also be called immersive media. Broadly speaking, immersive media uses audio and video technology to create a feeling of being physically present in a scene. For example, when users wear VR (Virtual Reality) headsets, they experience a strong sense of immersion in the scene.
[0004] Immersive media has a variety of applications, and the operational steps and processing capabilities required for users to decapsulate, decode, and render immersive media in different application scenarios are obviously different.
[0005] The relevant technologies cannot distinguish between different application scenarios in file encapsulation, which brings unnecessary trouble to the user-side processing.
[0006] Therefore, there is a need for a new method for encapsulating media files, a method for decapsulating media files, a device for encapsulating media files, a device for decapsulating media files, an electronic device, and a computer-readable storage medium.
[0007] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure. Summary of the Invention
[0008] This disclosure provides a media file encapsulation method, a media file decapsulation method, a media file encapsulation apparatus, a media file decapsulation apparatus, an electronic device, and a computer-readable storage medium, which can distinguish different application scenarios in the media file encapsulation process.
[0009] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0010] This disclosure provides a method for encapsulating media files. The method includes: determining target media content; obtaining the media stream of the target media content in a corresponding application scenario; encapsulating the media stream in the corresponding application scenario to generate an encapsulation file for the media stream in the corresponding application scenario, the encapsulation file including a first application scenario type field, the first application scenario type field indicating the application scenario corresponding to the media stream; and sending the encapsulation file to a first device so that the first device can obtain the application scenario corresponding to the media stream based on the first application scenario type field and determine the decoding or rendering method of the media stream.
[0011] This disclosure provides a method for decapsulating media files. The method includes: receiving a packaged file of a media stream of target media content in a corresponding application scenario, the packaged file including a first application scenario type field, the first application scenario type field indicating the application scenario corresponding to the media stream; decapsulating the packaged file to obtain the first application scenario type field; obtaining the application scenario corresponding to the media stream based on the first application scenario type field; and determining the decoding or rendering method of the media stream based on the application scenario corresponding to the media stream.
[0012] This disclosure provides a media file encapsulation apparatus, comprising: a media content determination unit for determining target media content; a media stream acquisition unit for acquiring the media stream of the target media content in a corresponding application scenario; a media stream encapsulation unit for encapsulating the media stream in the corresponding application scenario and generating an encapsulation file of the media stream in the corresponding application scenario, the encapsulation file including a first application scenario type field, the first application scenario type field indicating the application scenario corresponding to the media stream; and an encapsulation file sending unit for sending the encapsulation file to a first device, so that the first device obtains the application scenario corresponding to the media stream based on the first application scenario type field and determines the decoding or rendering method of the media stream.
[0013] This disclosure provides a media file decapsulation device, comprising: a capsulation file receiving unit, configured to receive a capsulation file of a media stream of target media content in a corresponding application scenario, the capsulation file including a first application scenario type field, the first application scenario type field indicating the application scenario corresponding to the media stream; a file decapsulation unit, configured to decapsulate the capsulation file to obtain the first application scenario type field; an application scenario obtaining unit, configured to obtain the application scenario corresponding to the media stream based on the first application scenario type field; and a decoding and rendering determining unit, configured to determine the decoding or rendering method of the media stream based on the application scenario corresponding to the media stream.
[0014] This disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the media file encapsulation method or video encoding method as described in the above embodiments.
[0015] This disclosure provides an electronic device, including: at least one processor; and a storage device configured to store at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the media file encapsulation method or video encoding method as described in the above embodiments.
[0016] In some embodiments of this disclosure, the technical solutions provided include extending a first application scenario type field into the encapsulation file when generating the media stream encapsulation file for a corresponding application scenario. This first application scenario type field indicates the application scenario corresponding to the media stream, thereby enabling the different application scenarios of different media streams to be distinguished during the encapsulation of the media file. On the one hand, when the encapsulation file is sent to the first device, the first device can distinguish the application scenario of the media stream based on the first application scenario type field in the encapsulation file, and thus determine the appropriate decoding or rendering method for the media stream based on the application scenario, saving the computing power and resources of the first device. On the other hand, since the application scenario of the media stream can be determined at the encapsulation stage, even if the first device does not have the ability to decode the media stream, it can still determine the application scenario corresponding to the media stream without having to wait until the media stream is decoded to distinguish it.
[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0019] Figure 1 A schematic diagram of three degrees of freedom is shown.
[0020] Figure 2 A schematic diagram of a three-degree-of-freedom+ system is shown.
[0021] Figure 3 A schematic diagram of a six-degree-of-freedom system is shown.
[0022] Figure 4 A flowchart illustrating a method for encapsulating media files according to an embodiment of the present disclosure is shown.
[0023] Figure 5 A flowchart illustrating a method for encapsulating media files according to an embodiment of the present disclosure is shown.
[0024] Figure 6 A flowchart illustrating a method for encapsulating media files according to an embodiment of the present disclosure is shown.
[0025] Figure 7 The illustration shows a schematic diagram of a six-degree-of-freedom media splicing method according to an embodiment of the present disclosure.
[0026] Figure 8 The illustration shows a schematic diagram of a six-degree-of-freedom media left-right splicing method according to an embodiment of the present disclosure.
[0027] Figure 9 A schematic illustration of a six-DOF media depth according to an embodiment of the present disclosure is shown. Figure 1 A diagram illustrating the 4x4 resolution splicing method.
[0028] Figure 10 A flowchart illustrating a method for encapsulating media files according to an embodiment of the present disclosure is shown.
[0029] Figure 11 A flowchart illustrating a method for encapsulating media files according to an embodiment of the present disclosure is shown.
[0030] Figure 12 A flowchart illustrating a method for encapsulating media files according to an embodiment of the present disclosure is shown.
[0031] Figure 13 The illustration shows a schematic diagram of a first multi-view video stitching method according to an embodiment of the present disclosure.
[0032] Figure 14 The illustration shows a schematic diagram of a second multi-view video stitching method according to an embodiment of the present disclosure.
[0033] Figure 15 A flowchart illustrating a method for decapsulating media files according to an embodiment of the present disclosure is shown.
[0034] Figure 16 A block diagram of a media file encapsulation apparatus according to an embodiment of the present disclosure is shown schematically.
[0035] Figure 17 A block diagram of a media file decapsulation apparatus according to an embodiment of the present disclosure is shown schematically.
[0036] Figure 18 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0037] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.
[0038] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0039] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in at least one hardware module or integrated circuit, or in different network and / or processor devices and / or microcontroller devices.
[0040] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0041] First, some of the terms used in the embodiments of this disclosure will be explained.
[0042] Point cloud: A point cloud is a set of randomly distributed discrete points in space that represent the spatial structure and surface properties of a three-dimensional object or scene. A point cloud refers to the geometry of a massive number of three-dimensional points. Each point in a point cloud has at least three-dimensional positional information, and depending on the application scenario, it may also have additional attributes such as color, material, or other information like reflectivity. Typically, each point in a point cloud has the same number of additional attributes. For example, a point cloud obtained based on laser measurement principles includes three-dimensional coordinates (XYZ) and laser reflection intensity; a point cloud obtained based on photogrammetry principles includes three-dimensional coordinates (XYZ) and color information (RGB, red, green, blue); a point cloud obtained by combining laser measurement and photogrammetry principles includes three-dimensional coordinates (XYZ), laser reflection intensity, and color information (RGB).
[0043] Point clouds can be divided into two main categories based on their applications: machine-perceived point clouds, which can be used in scenarios such as autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, and disaster relief robots; and human-perceived point clouds, which can be used in scenarios such as digital cultural heritage, free-viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.
[0044] Point clouds can be classified according to the acquisition method: the first type is static point cloud, where the object is stationary and the device for acquiring the point cloud is also stationary; the second type is dynamic point cloud, where the object is moving but the device for acquiring the point cloud is stationary; and the third type is dynamically acquired point cloud, where the device for acquiring the point cloud is moving.
[0045] PCC: Point Cloud Compression. Point clouds are collections of massive amounts of points. Storing this point cloud data not only consumes a lot of memory but is also inconvenient for transmission. Related technologies do not have enough bandwidth to support the direct transmission of point clouds at the network layer without compression. Therefore, compressing point clouds is essential.
[0046] G-PCC: Geometry-based Point Cloud Compression. G-PCC compresses point clouds based on geometric features. It is used for compressing static point clouds (type 1) and dynamically acquired point clouds (type 3). The resulting point cloud media can be called point cloud media compressed based on geometric features, or simply G-PCC point cloud media.
[0047] V-PCC: Video-based Point Cloud Compression, is a point cloud compression method based on traditional video coding. V-PCC compresses type II dynamic point clouds, and the resulting point cloud media can be referred to as point cloud media compressed using traditional video coding methods, or simply V-PCC point cloud media.
[0048] Sample: A unit of encapsulation in the media file encapsulation process. A media file is composed of many samples. Taking video media as an example, a sample of video media is usually a video frame.
[0049] DoF: Degree of Freedom. In a mechanical system, it refers to the number of independent coordinates, including translational degrees of freedom, rotational degrees of freedom, and vibrational degrees of freedom. In this embodiment of the disclosure, it refers to the degrees of freedom for motion and content interaction supported by the user when watching immersive media.
[0050] 3DoF: refers to three degrees of freedom, which means the user's head rotates around the XYZ axes. Figure 1 A schematic diagram of three degrees of freedom is shown. (For example...) Figure 1 As shown, at a certain location or point, one can rotate on all three axes, turning their head, tilting it up and down, or swaying it. Through this three-degrees-of-freedom experience, users can be fully immersed in a scene from 360 degrees. If it's static, it can be understood as a panoramic image. If the panoramic image is dynamic, it's a panoramic video, or VR video. However, VR videos have certain limitations; users cannot move or choose any location to view the content.
[0051] 3DoF+: In addition to the three degrees of freedom, users also have a limited number of degrees of freedom to move along the XYZ axes. It can also be called restricted six degrees of freedom, and the corresponding media stream can be called restricted six degrees of freedom media stream. Figure 2 A schematic diagram of a three-degree-of-freedom+ system is shown.
[0052] 6DoF: In addition to the three degrees of freedom, users also have the freedom to move freely along the XYZ axes. The corresponding media stream can be called a six-degree-of-freedom media stream. Figure 3The diagram illustrates a six-degrees-of-freedom (6DoF) scenario. 6DoF media refers to six-DOF video, meaning video that allows users to freely move their viewpoint along the XYZ axes in three-dimensional space and freely rotate it around the XYX axes, providing a high degree of freedom for viewing. 6DoF media is a combination of video footage from different spatial perspectives captured by a camera array. To facilitate the expression, storage, compression, and processing of 6DoF media, the data is represented as a combination of the following information: texture maps captured by multiple cameras, depth maps corresponding to the texture maps from the multiple cameras, and corresponding 6DoF media content description metadata. The metadata includes parameters of the multiple cameras, as well as descriptive information such as the 6DoF media's stitching layout and edge protection. At the encoding end, the texture map information and corresponding depth map information from the multiple cameras are stitched together, and the description data of the stitching method is written into the metadata according to the defined syntax and semantics. The stitched multi-camera depth map and texture map information are encoded using planar video compression and transmitted to the terminal for decoding. The resulting 6DoF virtual viewpoint is then synthesized to provide the user with the 6DoF media viewing experience.
[0053] Volumetric media is a type of immersive media, which may include volumetric video. Volumetric video is a three-dimensional data representation. Since current mainstream encoding is based on two-dimensional video data, the raw volumetric video data needs to be converted from three-dimensional to two-dimensional before encoding, encapsulation, and transmission at the system level. During the presentation of volumetric video content, the two-dimensional data needs to be converted back to three-dimensional data to represent the final volumetric video. How volumetric video is represented in a two-dimensional plane directly affects system-level encapsulation, transmission, and the final content presentation of the volumetric video.
[0054] An atlas indicates region information on a 2D (2-dimensional) planar frame, region information in a 3D (3-dimensional) rendering space, and the mapping relationship between the two, as well as the necessary parameter information required for the mapping. An atlas includes tiles and a collection of associated information about each tile's corresponding region in the 3D space of the volumetric data. A patch is a rectangular region in the atlas, associated with volumetric information in 3D space. Tiles are generated by processing the component data of the 2D representation of the volumetric video. Based on the position of the volumetric video represented in the geometric component data, the 2D planar region containing the 2D representation of the volumetric video is divided into multiple rectangular regions of different sizes. Each rectangular region is a tile, and each tile contains the necessary information for back-projecting that rectangular region into 3D space. Tiles are packaged to generate an atlas, placing them in a 2D grid and ensuring that the effective parts of each tile do not overlap. Tiles generated from a single volumetric video can be packaged into one or more atlases. Based on the atlas data, corresponding geometric data, attribute data, and placeholder data are generated. These data are then combined to form the final representation of the volumetric video in a two-dimensional plane. The geometric component is mandatory, the placeholder component is conditionally mandatory, and the attribute component is optional.
[0055] AVS: Audio Video Coding Standard.
[0056] ISOBMFF: ISO Based Media File Format, a media file format based on the ISO (International Standards Organization) standard. ISOBMFF is a media file encapsulation standard, and the most typical ISOBMFF file is the MP4 (Moving Picture Experts Group 4) file.
[0057] Depth map: As a way of representing three-dimensional scene information, the gray value of each pixel in the depth map can be used to characterize the distance of a point in the scene from the camera.
[0058] The media file encapsulation method provided in this disclosure can be executed by any electronic device. In the following example description, an example is given for use on the server side of an immersive system, but this disclosure is not limited thereto.
[0059] Figure 4 A flowchart illustrating a media file encapsulation method according to an embodiment of the present disclosure is shown. Figure 4 As shown, the method provided in this disclosure embodiment may include the following steps.
[0060] In step S410, the target media content is determined.
[0061] In this embodiment of the disclosure, the target media content can be any one or a combination of video, audio, images, etc. In the following examples, video is used as an example for illustration, but this disclosure is not limited to this.
[0062] In step S420, the media bitstream of the target media content in the corresponding application scenario is obtained.
[0063] In this embodiment of the disclosure, the media stream may include any media stream rendered in 3D space, such as a six-degree-of-freedom (6DoF) media stream or a restricted six-degree-of-freedom (3DoF+) media stream. The following examples use 6DoF media as an example. The method provided in this embodiment of the disclosure can be applied to applications such as 6DoF media content recording, on-demand playback, live streaming, communication, program editing, and production.
[0064] Immersive media can be categorized into 3DoF media, 3DoF+ media, and 6DoF media based on the degree of freedom users have when consuming target media content. 6DoF media can include multi-view video and point cloud media.
[0065] Point cloud media can be further divided into point cloud media compressed based on traditional video coding methods (i.e., V-PCC) and point cloud media compressed based on geometric features (G-PCC) based on coding methods.
[0066] Multi-view video is typically captured by an array of cameras from multiple angles (also known as viewpoints) of the same scene, forming a texture map that includes the scene's texture information (color information, etc.) and a depth map that includes depth information (spatial distance information, etc.). Combined with the mapping information from 2D planar frames to 3D rendering space, this constitutes 6DoF media that can be consumed by the user.
[0067] As can be seen from the relevant technologies, 6DoF media has a variety of applications. The operation steps and processing capabilities required for users to decapsulate, decode and render 6DoF media in different application scenarios are obviously different.
[0068] For example, multi-view video and V-PCC use one set of encoding rules, while G-PCC uses another set. Since their encoding standards are different, their decoding processes will definitely differ.
[0069] For example, although multi-view video and V-PCC use the same encoding standard, one renders images into 3D space, while the other renders a set of points into 3D space, so there will be some differences. Additionally, multi-view video requires texture maps and depth maps, while V-PCC, in addition to these, may also require placeholder maps, which is another difference.
[0070] In step S430, the media stream corresponding to the application scenario is encapsulated to generate an encapsulation file for the media stream corresponding to the application scenario. The encapsulation file includes a first application scenario type field, which indicates the application scenario corresponding to the media stream.
[0071] For example, this disclosure can differentiate between different 6DoF media application scenarios for applications of 6DoF media.
[0072] Since the industry currently defines 6DoF media uniformly as volumetric media, failing to differentiate between different application scenarios within the file encapsulation will cause unnecessary trouble for user-side processing. For example, if these different application scenarios cannot be distinguished within the media file encapsulation, it is necessary to decode the media stream first, which wastes computing resources. Furthermore, some intermediate nodes, such as CDN (Content Delivery Network) nodes, do not have decoding capabilities.
[0073] As mentioned above, these different applications have different processing methods and need to be distinguished. The advantage of distinguishing them in file encapsulation is that this information can be obtained at a high level of the media file, thereby saving computing resources. At the same time, it also allows some intermediate nodes that do not have decoding capabilities, such as CDN nodes, to obtain this information.
[0074] In step S440, the encapsulated file is sent to the first device so that the first device can obtain the application scenario corresponding to the media stream based on the first application scenario type field and determine the decoding or rendering method of the media stream.
[0075] In this embodiment of the disclosure, the first device can be any intermediate node or any user terminal consuming the media stream, and this disclosure does not limit it.
[0076] The media file encapsulation method provided in this disclosure extends a first application scenario type field into the encapsulation file when generating the encapsulation file for the media stream in a corresponding application scenario. This first application scenario type field indicates the application scenario corresponding to the media stream, thereby enabling the different application scenarios of different media streams to be distinguished during the media file encapsulation process. On the one hand, when the encapsulation file is sent to a first device, the first device can distinguish the application scenario of the media stream based on the first application scenario type field in the encapsulation file, and thus determine the appropriate decoding or rendering method for the media stream based on the application scenario, saving the computing power and resources of the first device. On the other hand, since the application scenario of the media stream can be determined at the encapsulation stage, even if the first device does not have the ability to decode the media stream, it can still determine the application scenario corresponding to the media stream without having to wait until the media stream is decoded to make the distinction.
[0077] Figure 5 A flowchart illustrating a media file encapsulation method according to an embodiment of the present disclosure is shown. Figure 5 As shown, the method provided in this disclosure embodiment may include the following steps.
[0078] Figure 5 Steps S410-S420 in the embodiments can refer to the above embodiments.
[0079] exist Figure 5 In the embodiments, the above Figure 4 Step S430 in the embodiment may further include the following steps.
[0080] In step S431, the first application scenario type field is added to the volumetric visual media header data box (e.g., the VolumetricVisualMediaHeaderBox exemplified below) of the target media file format data box.
[0081] In this embodiment of the disclosure, in order to enable corresponding identification of media files based on application scenarios such as 6DoF media, several descriptive fields can be added at the system level, including field extensions at the file encapsulation level. For example, in the following illustration, an extended ISOBMFF data box (as the target media file format data box) is used as an example, but this disclosure is not limited to this.
[0082] In step S432, the value of the first application scenario type field is determined according to the application scenario corresponding to the media stream.
[0083] In an exemplary embodiment, the value of the first application scenario type field may include any one of the following: a first value (e.g., "0") indicating that the media stream is a multi-view video with non-large-scale atlas information; a second value (e.g., "1") indicating that the media stream is a multi-view video with large-scale atlas information; a third value (e.g., "2") indicating that the media stream is point cloud media compressed based on traditional video encoding methods; and a fourth value (e.g., "3") indicating that the media stream is point cloud media compressed based on geometric features.
[0084] It should be understood that the value of the first application scenario type field is not limited to indicating the above application scenarios. It can indicate more or fewer application scenarios and can be set according to actual needs.
[0085] Figure 5 Step S440 in the embodiment can refer to the above embodiment.
[0086] The media file encapsulation method provided in this disclosure can, by distinguishing different application scenarios of 6DoF media, enable the first device consuming 6DoF media to make targeted strategy selections in the decapsulation, decoding, and rendering stages of 6DoF media.
[0087] Figure 6 A flowchart illustrating a media file encapsulation method according to an embodiment of the present disclosure is shown. Figure 6 As shown, the method provided in this disclosure embodiment may include the following steps.
[0088] Figure 6 Steps S410-S420 in the embodiments can refer to the above embodiments.
[0089] exist Figure 6 In the embodiment, step S430 in the above embodiment may further include the following step S4321, that is, during encapsulation, the media stream is determined to be a multi-view video with large-scale image atlas information by using the first application scenario type field.
[0090] In step S4321, the media stream under the corresponding application scenario is encapsulated to generate an encapsulation file for the media stream under the corresponding application scenario. The encapsulation file includes a first application scenario type field (e.g., application_type as exemplified below). The value of the first application scenario type field is a second value indicating that the media stream is a multi-view video with large-scale atlas information.
[0091] For multi-view videos, the mapping information from 2D planar frames to 3D rendering space determines the 6DoF experience. There are two methods for indicating this mapping relationship. One method defines an atlas to finely divide the 2D planar regions, thus indicating the mapping relationship of these small 2D regions to 3D space. This is called non-large-scale atlas information, and the corresponding multi-view videos are non-large-scale atlas information multi-view videos. The other method is coarser, directly identifying the depth map and texture map generated by each camera from the perspective of the acquisition device (using cameras as examples). It then reconstructs the mapping relationship of the corresponding 2D depth map and texture map in 3D space based on the parameters of each camera. This is called large-scale atlas information, and the corresponding multi-view videos are large-scale atlas information multi-view videos. It's important to understand that large-scale and non-large-scale atlas information are relative terms and do not directly limit specific dimensions.
[0092] Camera parameters are typically divided into extrinsic and intrinsic parameters. Extrinsic parameters usually include information such as the camera's shooting position and angle, while intrinsic parameters usually include information such as the camera's optical center position and focal length.
[0093] Therefore, it can be seen that multi-view videos in 6DoF media can further include multi-view videos with large-scale atlas information and multi-view videos with non-large-scale atlas information. That is, 6DoF media has a variety of application forms, and the operation steps and processing capabilities required by users to decapsulate, decode and render 6DoF media in different application scenarios are obviously different.
[0094] For example, the difference between large-scale atlas information and non-large-scale atlas information lies in the granularity of the mapping and rendering from 2D regions to 3D space. For instance, large-scale atlas information might map 6 2D puzzle pieces to 3D space, while non-large-scale atlas information might map 60 puzzle pieces to 3D space. Therefore, the complexity of the mapping algorithm will certainly differ; the algorithm for large-scale atlas information will be simpler than that for non-large-scale atlas information.
[0095] In particular, for multi-view videos, if the mapping relationship between their 2D regions and 3D space is obtained from camera parameters, that is, if it is a multi-view video with large-scale atlas information, then it is not necessary to define a smaller mapping relationship between 2D regions and 3D space in the encapsulation file.
[0096] When the media stream is a multi-view video containing large-scale image atlas information, the method may further include the following steps.
[0097] In step S601, if the media stream is encapsulated in a single track, a large-scale atlas flag (e.g., large_scale_atlas_flag) is added to the bitstream sample entry (e.g., V3CbitstreamSampleEntry as illustrated below, but this disclosure is not limited thereto) of the target media file format data box.
[0098] In step S602, if the large-scale atlas identifier indicates that the media stream is a multi-view video with large-scale atlas information, then add an identifier of the number of cameras that acquired the media stream (e.g., camera_count as exemplified below) and an identifier of the number of viewpoints corresponding to the cameras contained in the current file of the media stream (e.g., camera_count_contained as exemplified below) to the bitstream sample entry.
[0099] In step S603, the resolution of the texture map and depth map captured by the camera corresponding to the current file (e.g., camera_resolution_x and camera_resolution_y as exemplified below) is added to the bitstream sample entry.
[0100] Continue to refer to Figure 6 Furthermore, the method further includes at least one of the following steps S604-S607.
[0101] In step S604, a downsampling factor (e.g., depth_downsample_factor as exemplified below) of the depth map acquired from the viewpoint corresponding to the camera contained in the current file is added to the bitstream sample entry.
[0102] In step S605, the offset of the top left vertex of the texture map captured by the camera corresponding to the current file relative to the origin of the planar frame in the large-scale atlas information is added to the bitstream sample entry (e.g., texture_vetex_x and texture_vetex_y as exemplified below).
[0103] In step S606, the offset of the top left vertex of the depth map captured by the camera corresponding to the current file relative to the origin of the planar frame in the large-scale atlas information is added to the bitstream sample entry (e.g., depth_vetex_x and depth_vetex_y as exemplified below).
[0104] In step S607, the guard band width (e.g., padding_size_texture and padding_size_depth as exemplified below) of the texture map and depth map captured by the camera corresponding to the current file is added to the bit stream sample entry.
[0105] In this embodiment of the disclosure, `padding_size_texture` and `padding_size_depth` define the size of the edge protection region for each texture map and depth map, in order to protect against abrupt edge changes during compression of the stitched image. The values of `padding_size_texture` and `padding_size_depth` represent the width of the edge protection region for the texture map and depth map, respectively. A value of 0 for `padding_size_texture` and `padding_size_depth` indicates that there is no edge protection.
[0106] Continue to refer to Figure 6 When the media stream is a multi-view video with large-scale image atlas information, the method may further include the following steps.
[0107] In step S608, if the media stream is encapsulated in a multi-track manner, a large-scale atlas identifier is added to the sample entry of the target media file format data box.
[0108] In step S609, if the large-scale atlas identifier indicates that the media stream is a multi-view video with large-scale atlas information, then add an identifier of the number of cameras that acquired the media stream and an identifier of the number of viewpoints corresponding to the cameras contained in the current file of the media stream to the sample entry.
[0109] In step S610, the resolution of the texture map and depth map captured by the camera corresponding to the current file is added to the sample entry.
[0110] Continue to refer to Figure 6 Furthermore, the method further includes at least one of the following steps S611-S614.
[0111] In step S611, a downsampling factor is added to the sample entry for the depth map captured by the camera viewpoint contained in the current file.
[0112] In step S612, the offset of the top left vertex of the texture map captured by the camera corresponding to the current file relative to the origin of the planar frame in the large-scale atlas information is added to the sample entry.
[0113] In step S613, the offset of the top left vertex of the depth map captured by the camera corresponding to the current file relative to the origin of the planar frame in the large-scale atlas information is added to the sample entry.
[0114] In step S614, a guard band width is added to the sample entry for the texture map and depth map captured by the camera corresponding to the viewpoint contained in the current file.
[0115] Figure 6 Step S440 in the embodiment can refer to the above embodiment.
[0116] In this embodiment of the disclosure, the six_dof_stitching_layout field can be used to indicate the stitching method of the depth map and texture map captured by each camera in the 6DoF media, which is used to identify the stitching layout of the texture map and depth map of the 6DoF media. The specific values can be found in Table 1 below.
[0117] Table 1 6DoF Media Splicing Layout
[0118]
[0119]
[0120] Figure 7 The illustration shows a schematic diagram of a six-degree-of-freedom media splicing method according to an embodiment of the present disclosure.
[0121] When the value of six_dof_stitching_layout is 0, the 6DoF media stitching mode is top-to-bottom stitching, such as... Figure 7 As shown, in the top-bottom stitching mode, texture maps captured by multiple cameras (e.g., Figure 7 The viewpoint 1 texture map, viewpoint 2 texture map, viewpoint 3 texture map, and viewpoint 4 texture map are arranged sequentially above the image, while the corresponding depth maps (e.g., Figure 7 The depth maps (view 1, view 2, view 3, and view 4) are arranged in order below the image.
[0122] The resolution of the stitched 6DoF media is set to nWidth×nHeight. The reconstruction module can then use the values of camera_resolution_x and camera_resolution_y to calculate the layout positions of the texture maps and depth maps of each camera, thereby further utilizing the texture map and depth map information of multiple cameras to reconstruct the 6DoF media.
[0123] Figure 8The illustration shows a schematic diagram of a six-degree-of-freedom media left-right splicing method according to an embodiment of the present disclosure.
[0124] When the value of six_dof_stitching_layout is 1, the 6DoF media stitching mode is left-right stitching, such as... Figure 8 As shown, in the left-right stitching mode, texture maps captured by multiple cameras (e.g., Figure 8 The viewpoint 1 texture map, viewpoint 2 texture map, viewpoint 3 texture map, and viewpoint 4 texture map are arranged sequentially on the left side of the image, while the corresponding depth maps (e.g., Figure 8 The depth maps (view 1, view 2, view 3, and view 4) are arranged in order on the right side of the image.
[0125] Figure 9 A schematic illustration of a six-DOF media depth according to an embodiment of the present disclosure is shown. Figure 1 A diagram illustrating the 4x4 resolution splicing method.
[0126] When the value of six_dof_stitching_layout is 2, the 6DoF media stitching mode is depth-based. Figure 1 / 4 downsampling splicing, such as Figure 9 As shown, depth Figure 1 In the / 4 downsampling stitching method, the depth map (e.g.) Figure 9 The depth maps (view 1, view 2, view 3, and view 4) are downsampled at 1 / 4 resolution and then stitched onto the texture map (e.g., ...). Figure 9 The lower right of the texture maps (view 1, view 2, view 3, and view 4) in the image. If the depth map stitching cannot fill the rectangular area of the final stitched image, the remaining part is filled with a blank image.
[0127] The media file encapsulation method provided in this disclosure can not only distinguish different application scenarios of 6DoF media, but also enable the first device consuming 6DoF media to make targeted strategy selections in the decapsulation, decoding, and rendering stages of 6DoF media. Furthermore, for multi-view video applications in 6DoF media, a method is proposed to indicate information related to multi-view video depth maps and texture maps in the file encapsulation, making the encapsulation combination of depth maps and texture maps from different perspectives of multi-view video more flexible.
[0128] In an exemplary embodiment, the method may further include: generating a target description file for the target media content, the target description file including a second application scenario type field, the second application scenario type field indicating the application scenario corresponding to the media stream; and sending the target description file to the first device, so that the first device determines the target encapsulation file of the target media stream from the encapsulation file of the media stream according to the second application scenario type field.
[0129] Sending the encapsulated file to the first device so that the first device can determine the application scenario corresponding to the media stream based on the first application scenario type field may include: sending the target encapsulated file to the first device so that the first device can determine the target application scenario of the target media stream based on the first application scenario type field in the target encapsulated file.
[0130] Figure 10 A flowchart illustrating a media file encapsulation method according to an embodiment of the present disclosure is shown. Figure 10 As shown, the method provided in this disclosure embodiment may include the following steps.
[0131] Figure 10 Steps S410-S430 in the embodiments can refer to the above embodiments, and may further include the following steps.
[0132] In step S1010, a target description file for the target media content is generated. The target description file includes a second application scenario type field (e.g., v3cAppType as exemplified below). The second application scenario type field indicates the application scenario corresponding to the media stream.
[0133] In this embodiment of the disclosure, several descriptive fields are added at the system layer. In addition to the field extensions at the file encapsulation layer mentioned above, field extensions at the signaling transmission layer can also be performed. In the following embodiment, the example of supporting DASH (Dynamic adaptive streaming over HTTP (HyperText Transfer Protocol)) MPD (Media Presentation Description) signaling (as the target description file) is used to illustrate the definition of 6DoF media application scenario type indicators and large-scale atlas indicators.
[0134] In step S1020, the target description file is sent to the first device so that the first device can determine the target encapsulation file of the target media stream from the encapsulation file of the media stream according to the second application scenario type field.
[0135] In step S1030, the target encapsulation file is sent to the first device so that the first device can determine the target application scenario of the target media stream based on the first application scenario type field in the target encapsulation file.
[0136] The media file encapsulation method provided in this disclosure can not only determine the application scenario corresponding to the media bitstream through a first application scenario type field in the encapsulation file, but also determine the application scenario corresponding to the media bitstream through a second application scenario type field in the target description file. In this way, the first device can first determine what kind of media bitstream it needs to obtain based on the second application scenario type field in the target description file, and then request the corresponding target media bitstream from the server. This can reduce the amount of data transmission, and at the same time, the requested target media bitstream can match the actual capabilities of the first device. After the first device receives the requested target media bitstream, it can further determine the target application scenario of the target media bitstream based on the first application scenario type field in the encapsulation file, so as to know which decoding and rendering method should be adopted, thereby reducing computing resources.
[0137] Figure 11 A flowchart illustrating a media file encapsulation method according to an embodiment of the present disclosure is shown. Figure 11 As shown, the method provided in this disclosure embodiment may include the following steps.
[0138] Figure 11 Steps S410-S430 in the embodiments can refer to the above embodiments.
[0139] Figure 11 In the embodiments, the above Figure 10 Step S1010 in the embodiment may further include the following steps.
[0140] In step S1011, the second application scenario type field is added to the target description file of the target media content for dynamic adaptive streaming media transmission based on the Hypertext Transfer Protocol.
[0141] In step S1012, the value of the second application scenario type field is determined according to the application scenario corresponding to the media stream.
[0142] Figure 11 Steps S1020 and S1030 in the embodiments can refer to the above embodiments.
[0143] The following is an example illustrating the media file encapsulation method proposed in this disclosure. Taking 6DoF media as an example, the method proposed in this disclosure can be used for 6DoF media application scenario indication and may include the following steps:
[0144] 1. Identify media files according to their application scenarios in 6DoF media.
[0145] 2. Specifically, for multi-view videos, it is determined whether the mapping from 2D planar frames to 3D space is performed on a unit basis, i.e., the mapping from 2D planar frames to 3D space is performed on a unit basis, i.e., the mapping is performed on a unit basis, i.e., the texture map and depth map captured by each camera are captured by each camera. This is called large-scale atlas information. If it is necessary to further divide the texture map and depth map captured by each camera into a more detailed division, indicating the mapping of the divided 2D small region set to 3D space, this is called non-large-scale atlas information.
[0146] 3. If the mapping of multi-view video from 2D planar frames to 3D space is performed on a per-camera basis, then the relevant information of the different camera outputs should be indicated in the encapsulation file.
[0147] This embodiment can add several descriptive fields at the system layer, including field extensions at the file encapsulation level and field extensions at the signaling transmission level, to support the above steps of this disclosure embodiment. The following example uses extended ISOBMFF data boxes and DASH MPD signaling to define application type indicators and large-scale atlas indicators for 6DoF media, as detailed below (where the extended parts are indicated in italics).
[0148] I. ISOBMFF Data Box Expansion
[0149] The mathematical operators and their precedence used in this section are based on those in the C programming language. Unless otherwise specified, numbering and counting are conventionally based on starting from 0.
[0150]
[0151]
[0152]
[0153] In this embodiment of the disclosure, the first application scenario type field application_type indicates the application scenario type of the 6DoF media, and the specific values include, but are not limited to, those shown in Table 2 below:
[0154] Table 2
[0155] 0 Multi-view video (not large-scale image atlas information) 1 Multi-view video (large-scale image collection information) 2 Point cloud media compression based on traditional video encoding methods 3 Point cloud media based on geometric features
[0156] The large-scale atlas flag indicates whether the atlas information is large-scale atlas information, that is, whether the atlas information can be obtained solely through camera parameters and other related information. Here, it is assumed that when large-scale_atlas_flag equals 1, it indicates multi-view video (large-scale atlas information), and when it equals 0, it indicates multi-view video (non-large-scale atlas information).
[0157] It should be noted that, as shown in Table 2 above, the first application scenario type field `application_type` is already sufficient to indicate whether it is a multi-view video with large-scale atlas information. Considering that the indication of `application_type` is relatively high-level, `large_scale_atlas_flag` is added for easier parsing. Only one is needed in practice, but because it's uncertain which field will be used, this information is redundant.
[0158] Here, `camera_count` indicates the total number of cameras capturing 6DoF media, referred to as the camera count identifier for capturing the media stream. The value of `camera_number` ranges from 1 to 255. `camera_count_contained` represents the number of camera viewpoints contained in the current file of the 6DoF media, referred to as the camera viewpoint count identifier contained in the current file.
[0159] Here, `padding_size_depth` represents the guard band width of the depth map. `padding_size_texture` is the guard band width of the texture map. During video encoding, guard bands are typically added to improve the error tolerance of video decoding; these are extra pixels padded at the edges of image frames.
[0160] `camera_id` represents the camera identifier corresponding to each viewpoint. `camera_resolution_x` and `camera_resolution_y` represent the resolution width and height of the texture map and depth map captured by the camera, respectively, indicating the resolution in the X and Y directions captured by the corresponding camera. `depth_downsample_factor` represents the downsampling factor of the corresponding depth map; the actual resolution width and height of the depth map are half of the resolution width and height captured by the camera. depth _downsample_factor.
[0161] depth_vetex_x and depth_vetex_y represent the X and Y component values of the offset of the top-left vertex of the corresponding depth map relative to the origin (top-left vertex of the plane frame) of the plane frame, respectively.
[0162] texture_vetex_x and texture_vetex_y represent the X and Y component values of the offset of the top-left vertex of the corresponding texture map relative to the origin (top-left vertex of the plane frame), respectively.
[0163] II. DASH MPD Signaling Extension
[0164] The second application scenario type field v3cAppType can be extended in the table shown in Table 3 of the DASH MPD signaling.
[0165] Table 3—Semantics of Representation Element
[0166]
[0167]
[0168] Corresponding to the above Figure 7 In this example, it is assumed that there is a multi-view video A on the server side, and the atlas information of the multi-view video A is large-scale atlas information.
[0169] At this time: application_type = 1;
[0170] large_scale_atlas_flag=1: camera_count=4; camera_count_contained=4;
[0171] padding_size_depth=0; padding_size_texture=0;
[0172] {camera_id=1; camera_resolution_x=100; camera_resolution_y=100;
[0173] depth_downsample_factor = 0; texture_vetex = (0, 0); depth_vetex = (0, 200)} / / View 1 texture map and view 1 depth map
[0174] {camera_id=2; camera_resolution_x=100; camera_resolution_y=100;
[0175] depth_downsample_factor = 0; texture_vetex = (100, 0); depth_vetex = (100, 200)} / / View 2 texture map and view 2 depth map
[0176] {camera_id=3; camera_resolution_x=100; camera_resolution_y=100;
[0177] depth_downsample_factor = 0; texture_vetex = (0, 100); depth_vetex = (0, 300) / / View 3 texture map and view 3 depth map
[0178] {camera_id=4; camera_resolution_x=100; camera_resolution_y=100;
[0179] depth_downsample_factor = 0; texture_vetex = (100, 100); depth_vetex = (100, 300)} / / View 4 texture map and view 4 depth map
[0180] The above system description corresponds to Figure 7 The data composition of each region of the planar frame.
[0181] Corresponding to the above Figure 8 In this example, it is assumed that there is a multi-view video A on the server side, and the atlas information of the multi-view video A is large-scale atlas information.
[0182] At this time: application_type = 1;
[0183] large_scale_atlas_flag=1: camera_count=4; camera_count_contained=4;
[0184] padding_size_depth=0; padding_size_texture=0;
[0185] {camera_id=1; camera_resolution_x=100; camera_resolution_y=100;
[0186] depth_downsample_factor=0; texture_vetex=(0,0); depth_vetex=(200,0)}
[0187] {camera_id=2; camera_resolution_x=100; camera_resolution_y=100;
[0188] depth_downsample_factor=0; texture_vetex=(100,0); depth_vetex=(300,0)}
[0189] {camera_id=3; camera_resolution_x=100; camera_resolution_y=100;
[0190] depth_downsample_factor=0; texture_vetex=(0,100); depth_vetex=(200,100)}
[0191] {camera_id=4; camera_resolution_x=100; camera_resolution_y=100;
[0192] depth_downsample_factor=0; texture_vetex=(100,100); depth_vetex=(300,100)}
[0193] The above system description corresponds to Figure 8 The data composition of each region of the planar frame.
[0194] Corresponding to the above Figure 9 In this example, it is assumed that there is a multi-view video A on the server side, and the atlas information of the multi-view video A is large-scale atlas information.
[0195] At this time: application_type = 1;
[0196] large_scale_atlas_flag=1: camera_count=4; camera_count_contained=4;
[0197] padding_size_depth=0; padding_size_texture=0;
[0198] {camera_id=1; camera_resolution_x=100; camera_resolution_y=100;
[0199] depth_downsample_factor=1; texture_vetex=(0,0); depth_vetex=(0,200)}
[0200] {camera_id=2; camera_resolution_x=100; camera_resolution_y=100;
[0201] depth_downsample_factor=1; texture_vetex=(100,0); depth_vetex=(50,200)}
[0202] {camera_id=3; camera_resolution_x=100; camera_resolution_y=100;
[0203] depth_downsample_factor=1; texture_vetex=(0,100); depth_vetex=(100,200)}
[0204] {camera_id=4; camera_resolution_x=100; camera_resolution_y=100;
[0205] depth_downsample_factor=1; texture_vetex=(100,100); depth_vetex=(150,200)}
[0206] The above system description corresponds to Figure 9 The data composition of each region of the planar frame.
[0207] It should be noted that `padding_size_depth` and `padding_size_texture` do not have an absolute range of values, and different values do not affect the method provided in this embodiment. This scheme only indicates the size of `padding_size_depth` and `padding_size_texture`. The reason why `padding_size_depth` and `padding_size_texture` are the same is determined by the encoding algorithm and is unrelated to the method provided in this embodiment.
[0208] Among them, camera_resolution_x and camera_resolution_y are used to calculate the actual resolution width and height of the depth map, which is the resolution of each camera. Multi-view video is shot by multiple cameras, and different cameras can have different resolutions. Here, the resolution width and height of all views are set to 100 pixels for the example only for the sake of illustration, and are not actually limited to this.
[0209] It is understood that the method provided in this disclosure is not limited to the above combinations and can provide corresponding instructions for any combination.
[0210] Once the client installed on the first device receives the encapsulated file of the multi-view video sent by the server, it can map each region of the multi-view video's planar frames to the texture maps and depth maps of different cameras by parsing the corresponding fields in the encapsulated file. Then, by decoding the camera parameter information in the media stream of the multi-view video, it can restore each region of the planar frames to the 3D rendering area, thereby consuming the multi-view video.
[0211] Corresponding to the above Figure 10 The following example illustrates this concept. Assume that for the same target media content, the server possesses three different forms of 6DoF media: multi-view video A (large-scale atlas information), V-PCC point cloud media B, and G-PCC point cloud media C. When encapsulating these three media streams, the server assigns corresponding values to the `application_type` field in the `VolumetricVisualMediaHeaderBox` data box. Specifically, for multi-view video A: `application_type = 1`; for V-PCC point cloud media B: `application_type = 2`; and for G-PCC point cloud media C: `application_type = 3`.
[0212] Meanwhile, the MPD file describes the application scenario types of the three representations: multi-view video A (large-scale atlas information), V-PCC point cloud media B, and G-PCC point cloud media C. Specifically, the values of the v3cAppType field are as follows: multi-view video A: v3cAppType = 1; V-PCC point cloud media B: v3cAppType = 2; G-PCC point cloud media C: v3cAppType = 3.
[0213] Then, the server sends the target description file corresponding to the MPD signaling to the client installed on the first device.
[0214] After the client receives the target description file corresponding to the MPD signaling sent by the server, it requests the target encapsulation file of the target media stream corresponding to the application scenario type, based on the client device's capabilities and presentation requirements. Assuming the first device has low client processing power, the client requests the target encapsulation file of multi-view video A.
[0215] The server will then send the target encapsulated file of the multi-view video A to the client of the first device.
[0216] After receiving the target encapsulated file of the multi-view video A sent by the server, the client on the first device determines the application scenario type of the current 6DoF media file based on the application_type field in the VolumetricVisualMediaHeaderBox data box, and then processes it accordingly. Different application scenario types will have different decoding and rendering algorithms.
[0217] Taking multi-view video as an example, if application_type=1, it means that the atlas information of the multi-view video is based on the depth map and texture map captured by the camera. Therefore, the client can process the multi-view video using a relatively simple processing algorithm.
[0218] It should be noted that in other embodiments, in addition to DASH MPD, similar signaling files can be extended to indicate the application scenario type of different media files in the signaling files.
[0219] In an exemplary embodiment, obtaining the media bitstream of the target media content in a corresponding application scenario may include: receiving a first encapsulation file of a first multi-view video sent by a second device and a second encapsulation file of a second multi-view video sent by a third device; decapsulating the first encapsulation file and the second encapsulation file respectively to obtain the first multi-view video and the second multi-view video; decoding the first multi-view video and the second multi-view video respectively to obtain a first depth map and a first texture map in the first multi-view video and a second depth map and a second texture map in the second multi-view video; and obtaining a merged multi-view video based on the first depth map, the second depth map, the first texture map, and the second texture map.
[0220] The second device may be equipped with a first number of cameras, and the third device may be equipped with a second number of cameras. The second device and the third device respectively use their respective cameras to capture and shoot multi-view videos of the same scene to obtain the first multi-view video and the second multi-view video.
[0221] The first encapsulation file and the second encapsulation file may each include the first application scenario type field, and the values of the first application scenario type field in the first encapsulation file and the second encapsulation file are respectively used to represent the second values of the first multi-view video and the second multi-view video as multi-view videos of large-scale atlas information.
[0222] Figure 12 A flowchart illustrating a media file encapsulation method according to an embodiment of the present disclosure is shown. Figure 12 As shown, the method provided in this disclosure embodiment may include the following steps.
[0223] In step S1210, the first encapsulation file of the first multi-view video sent by the second device and the second encapsulation file of the second multi-view video sent by the third device are received.
[0224] In step S1220, the first encapsulated file and the second encapsulated file are decapsulated to obtain the first multi-view video and the second multi-view video.
[0225] In step S1230, the first multi-view video and the second multi-view video are decoded respectively to obtain the first depth map and the first texture map in the first multi-view video, and the second depth map and the second texture map in the second multi-view video.
[0226] In step S1240, a merged multi-view video is obtained based on the first depth map, the second depth map, the first texture map, and the second texture map.
[0227] In step S1250, the multi-view videos are encapsulated and merged to generate an encapsulated file of the merged multi-view videos. The encapsulated file includes a first application scenario type field, which indicates the second value of the multi-view videos whose application scenario is large-scale atlas information.
[0228] In step S1260, the encapsulated file is sent to the first device so that the first device can obtain the application scenario corresponding to the merged multi-view video according to the first application scenario type field and determine the decoding or rendering method of the merged multi-view video.
[0229] The following is combined with Figure 13 and 14 right Figure 12The method provided in the embodiment is illustrated by example. Assume the second device and the third device are drone A and drone B respectively (but this disclosure is not limited to this), and assume that drone A and drone B are each equipped with two cameras (i.e., both the first and second numbers are equal to 2, but this disclosure is not limited to this and can be set according to the actual scenario). Using drone A and drone B to capture and film the same scene from multiple perspectives, the first encapsulated file corresponding to the first multi-view video captured by drone A during the process of capturing and producing the first multi-view video is as follows:
[0230] application_type = 1;
[0231] large_scale_atlas_flag=1: camera_count=4; camera_count_contained=2;
[0232] padding_size_depth=0; padding_size_texture=0;
[0233] {camera_id=1; camera_resolution_x=100; camera_resolution_y=100;
[0234] depth_downsample_factor = 1; texture_vetex = (0,0); depth_vetex = (0,100) / / View 1 texture map and view 1 depth map
[0235] {camera_id=2; camera_resolution_x=100; camera_resolution_y=100;
[0236] depth_downsample_factor = 1; texture_vetex = (100, 0); depth_vetex = (100, 100)} / / View 2 texture map and view 2 depth map
[0237] The above system description corresponds to Figure 13 The data composition of each region of the planar frame is illustrated here using an example of top-to-bottom splicing.
[0238] During the process of drone B acquiring and producing the second multi-view video, the corresponding second encapsulated file for the second multi-view video is as follows:
[0239] application_type = 1;
[0240] large_scale_atlas_flag=1: camera_count=4; camera_count_contained=2;
[0241] padding_size_depth=0; padding_size_texture=0;
[0242] {camera_id=3; camera_resolution_x=100; camera_resolution_y=100;
[0243] depth_downsample_factor = 1; texture_vetex = (0,0); depth_vetex = (0,100) / / View 3 texture map and view 3 depth map
[0244] {camera_id=4; camera_resolution_x=100; camera_resolution_y=100;
[0245] depth_downsample_factor = 1; texture_vetex = (100, 0); depth_vetex = (100, 100)} / / View 4 texture map and view 4 depth map
[0246] The above system description corresponds to, for example: Figure 14 The data composition of each region in the planar frame shown.
[0247] On the server side, after receiving the first and second encapsulated files captured by different drones, the server decapsulates and decodes the first and second encapsulated files, merges all the depth maps and texture maps, and assumes that the depth maps are downsampled to obtain the merged multi-view video.
[0248] Depth maps are less important than texture maps. Downsampling can reduce the amount of data. This disclosure describes such a scenario, but limits the scope of the scenario.
[0249] After encapsulating the merged multi-view videos, you can obtain the encapsulated file shown below:
[0250] application_type = 1;
[0251] large_scale_atlas_flag=1: camera_count=4; camera_count_contained=4;
[0252] padding_size_depth=0; padding_size_texture=0;
[0253] {camera_id=1; camera_resolution_x=100; camera_resolution_y=100;
[0254] depth_downsample_factor=1; texture_vetex=(0,0); depth_vetex=(0,200)}
[0255] {camera_id=2; camera_resolution_x=100; camera_resolution_y=100;
[0256] depth_downsample_factor=1; texture_vetex=(100,0); depth_vetex=(50,200)}
[0257] {camera_id=3; camera_resolution_x=100; camera_resolution_y=100;
[0258] depth_downsample_factor=1; texture_vetex=(0,100); depth_vetex=(100,200)}
[0259] {camera_id=4; camera_resolution_x=100; camera_resolution_y=100;
[0260] depth_downsample_factor=1; texture_vetex=(100,100); depth_vetex=(150,200)}
[0261] The above system description corresponds to the above. Figure 9 The data composition of each region in the planar frame shown.
[0262] Once the client on the first device receives the encapsulated file of the merged multi-view video from the server, it can parse the corresponding fields in the encapsulated file to map each region of the planar frame of the merged multi-view video to the texture maps and depth maps of different cameras. Then, by decoding the camera parameter information in the media stream of the merged multi-view video, it can restore each region of the planar frame to the 3D rendering area, thereby consuming the merged multi-view video.
[0263] The media file encapsulation method provided in this disclosure, for multi-view video applications in 6DoF media, proposes a method to indicate information related to the depth map and texture map of the multi-view video in the file encapsulation, making the encapsulation combination of depth maps and texture maps from different perspectives of the multi-view video more flexible. It can support different application scenarios. As described in the above embodiments, some scenarios involve different devices shooting, resulting in encapsulation into two files. The method provided in this disclosure can associate these two files and consume them together. Otherwise, in the above embodiments, only the two files can be presented separately, and they cannot be presented together.
[0264] The media file decapsulation method provided in this disclosure can be executed by any electronic device. In the following examples, an example is given for use in an intermediate node or a first device (e.g., a player) in an immersive system, but this disclosure is not limited thereto.
[0265] Figure 15 A flowchart illustrating a method for decapsulating media files according to an embodiment of this disclosure is shown schematically. Figure 15 As shown, the method provided in this disclosure embodiment may include the following steps.
[0266] In step S1510, the encapsulation file of the media stream of the target media content in the corresponding application scenario is received. The encapsulation file includes a first application scenario type field, which indicates the application scenario corresponding to the media stream.
[0267] In an exemplary embodiment, the method may further include: receiving a target description file of the target media content, the target description file including a second application scenario type field, the second application scenario type field indicating the application scenario corresponding to the media stream; and determining a target encapsulation file of the target media stream from the encapsulation file of the media stream according to the second application scenario type field.
[0268] The process of receiving the encapsulation file of the media stream of the target media content in the corresponding application scenario may include: receiving the target encapsulation file to determine the target application scenario of the target media stream according to the first application scenario type field in the target encapsulation file.
[0269] In step S1520, the encapsulated file is decapsulated to obtain the first application scenario type field.
[0270] In step S1530, the application scenario corresponding to the media stream is obtained according to the first application scenario type field.
[0271] In step S1540, the decoding or rendering method of the media stream is determined according to the application scenario corresponding to the media stream.
[0272] In an exemplary embodiment, if the value of the first application scenario type field is a second value indicating that the media stream is a multi-view video with large-scale atlas information, the method may further include: parsing the encapsulated file to obtain the mapping relationship between the texture map and depth map captured by the camera corresponding to the viewpoint contained in the media stream and the planar frames in the large-scale atlas information; decoding the media stream to obtain the camera parameters in the media stream; and presenting the multi-view video in three-dimensional space according to the mapping relationship and the camera parameters.
[0273] Other aspects of the media file decapsulation method provided in this disclosure can be found in the media file encapsulation methods described in the other embodiments above.
[0274] The media file encapsulation device provided in this disclosure can be installed on any electronic device. In the following example description, it is illustrated by setting it on the server side of an immersive system, but this disclosure is not limited thereto.
[0275] Figure 16 A block diagram schematically illustrates a media file encapsulation apparatus according to an embodiment of the present disclosure. Figure 16 As shown, the media file encapsulation device 1600 provided in this embodiment may include a media content determination unit 1610, a media stream acquisition unit 1620, a media stream encapsulation unit 1630, and an encapsulated file sending unit 1640.
[0276] In this embodiment, the media content determination unit 1610 can be used to determine target media content. The media stream acquisition unit 1620 can be used to acquire the media stream of the target media content in a corresponding application scenario. The media stream encapsulation unit 1630 can be used to encapsulate the media stream in the corresponding application scenario and generate an encapsulation file of the media stream in the corresponding application scenario. The encapsulation file includes a first application scenario type field, which indicates the application scenario corresponding to the media stream. The encapsulation file sending unit 1640 can be used to send the encapsulation file to a first device, so that the first device can obtain the application scenario corresponding to the media stream based on the first application scenario type field and determine the decoding or rendering method of the media stream.
[0277] The media file encapsulation apparatus provided in this disclosure extends a first application scenario type field into the encapsulation file when generating the encapsulation file for the media stream in a corresponding application scenario. This first application scenario type field indicates the application scenario corresponding to the media stream, thereby enabling the different application scenarios of different media streams to be distinguished during the media file encapsulation process. On the one hand, when the encapsulation file is sent to a first device, the first device can distinguish the application scenario of the media stream based on the first application scenario type field in the encapsulation file, and thus determine the appropriate decoding or rendering method for the media stream based on the application scenario, saving the computing power and resources of the first device. On the other hand, since the application scenario of the media stream can be determined at the encapsulation stage, even if the first device does not have the ability to decode the media stream, it can still determine the application scenario corresponding to the media stream without having to wait until the media stream is decoded to make the distinction.
[0278] In an exemplary embodiment, the media stream encapsulation unit 1630 may include: a first application scenario type field addition unit, which can be used to add the first application scenario type field to the volumetric visible media header data box of the target media file format data box; and a first application scenario type field value determination unit, which can be used to determine the value of the first application scenario type field according to the application scenario corresponding to the media stream.
[0279] In an exemplary embodiment, the value of the first application scenario type field may include any one of the following: a first value indicating that the media stream is a multi-view video with non-large-scale atlas information; a second value indicating that the media stream is a multi-view video with large-scale atlas information; a third value indicating that the media stream is point cloud media compressed based on traditional video encoding methods; and a fourth value indicating that the media stream is point cloud media compressed based on geometric features.
[0280] In an exemplary embodiment, if the value of the first application scenario type field is equal to the second value, the media file encapsulation device 1600 may further include: a single-track large-scale atlas identifier adding unit, which can be used to add a large-scale atlas identifier to the bitstream sample entry of the target media file format data box if the media stream is encapsulated in a single track; a single-track camera viewpoint identifier adding unit, which can be used to add an identifier of the number of cameras acquiring the media stream and an identifier of the number of viewpoints corresponding to the cameras contained in the current file of the media stream to the bitstream sample entry if the large-scale atlas identifier indicates that the media stream is a multi-view video with large-scale atlas information; and a single-track texture depth map resolution adding unit, which can be used to add the resolution of the texture map and depth map acquired by the viewpoints corresponding to the cameras contained in the current file to the bitstream sample entry.
[0281] In an exemplary embodiment, the media file encapsulation device 1600 may further include at least one of the following: a single-track downsampling factor adding unit, which can be used to add a downsampling factor of the depth map captured by the camera corresponding to the viewpoint contained in the current file to the bitstream sample entry; a single-track texture map offset adding unit, which can be used to add an offset of the upper left vertex of the texture map captured by the camera corresponding to the viewpoint contained in the current file to the origin of the planar frame in the large-scale atlas information to the bitstream sample entry; a single-track depth map offset adding unit, which can be used to add an offset of the upper left vertex of the depth map captured by the camera corresponding to the viewpoint contained in the current file to the origin of the planar frame in the large-scale atlas information to the bitstream sample entry; and a single-track guard band width adding unit, which can be used to add a guard band width of the texture map and depth map captured by the camera corresponding to the viewpoint contained in the current file to the bitstream sample entry.
[0282] In an exemplary embodiment, if the value of the first application scenario type field is equal to the second value, the media file encapsulation device 1600 may further include: a multi-track large-scale atlas identifier adding unit, which can be used to add a large-scale atlas identifier to the sample entry of the target media file format data box if the media stream is encapsulated according to multi-track; a multi-track camera viewpoint identifier adding unit, which can be used to add an identifier of the number of cameras capturing the media stream and an identifier of the number of viewpoints corresponding to the cameras contained in the current file of the media stream to the sample entry if the large-scale atlas identifier indicates that the media stream is a multi-view video with large-scale atlas information; and a multi-track texture depth map resolution adding unit, which can be used to add the resolution of the texture map and depth map captured by the viewpoints corresponding to the cameras contained in the current file to the sample entry.
[0283] In an exemplary embodiment, the media file encapsulation device 1600 may further include at least one of the following: a multi-track downsampling factor adding unit, which can be used to add a downsampling factor of the depth map captured by the camera corresponding to the viewpoint contained in the current file to the sample entry; a multi-track texture map offset adding unit, which can be used to add an offset of the upper left vertex of the texture map captured by the camera corresponding to the viewpoint contained in the current file to the origin of the planar frame in the large-scale atlas information to the sample entry; a multi-track depth map offset adding unit, which can be used to add an offset of the upper left vertex of the depth map captured by the camera corresponding to the viewpoint contained in the current file to the origin of the planar frame in the large-scale atlas information to the sample entry; and a multi-track guard band width adding unit, which can be used to add a guard band width of the texture map and depth map captured by the camera corresponding to the viewpoint contained in the current file to the sample entry.
[0284] In an exemplary embodiment, the media file encapsulation apparatus 1600 may further include: a target description file generation unit, configured to generate a target description file for the target media content, the target description file including a second application scenario type field, the second application scenario type field indicating the application scenario corresponding to the media stream; and a target description file sending unit, configured to send the target description file to the first device, so that the first device determines the target encapsulation file of the target media stream from the encapsulation file of the media stream according to the second application scenario type field. The encapsulation file sending unit 1640 may include: a target encapsulation file sending unit, configured to send the target encapsulation file to the first device, so that the first device determines the target application scenario of the target media stream according to the first application scenario type field in the target encapsulation file.
[0285] In an exemplary embodiment, the target description file generation unit may include: a second application scenario type field addition unit, which can be used to add the second application scenario type field to the target description file of the target media content based on the Hypertext Transfer Protocol for dynamic adaptive streaming media transmission; and a second application scenario type field value determination unit, which can be used to determine the value of the second application scenario type field according to the application scenario corresponding to the media stream.
[0286] In an exemplary embodiment, the media stream acquisition unit 1620 may include: a package file receiving unit, configured to receive a first package file of a first multi-view video sent by a second device and a second package file of a second multi-view video sent by a third device; a package file decapsulation unit, configured to decapsulate the first package file and the second package file respectively to obtain the first multi-view video and the second multi-view video; a multi-view video decoding unit, configured to decode the first multi-view video and the second multi-view video respectively to obtain a first depth map and a first texture map in the first multi-view video and a second depth map and a second texture map in the second multi-view video; and a multi-view video merging unit, configured to obtain a merged multi-view video based on the first depth map, the second depth map, the first texture map, and the second texture map.
[0287] In an exemplary embodiment, the second device may be equipped with a first number of cameras, and the third device may be equipped with a second number of cameras. The second device and the third device can respectively use their respective cameras to capture and shoot multi-view videos of the same scene to obtain the first multi-view video and the second multi-view video. The first encapsulation file and the second encapsulation file may each include a first application scenario type field, and the values of the first application scenario type field in the first encapsulation file and the second encapsulation file can be used to represent second values indicating that the first multi-view video and the second multi-view video are multi-view videos of large-scale atlas information.
[0288] In an exemplary embodiment, the media stream may include a six-degree-of-freedom media stream and a restricted six-degree-of-freedom media stream.
[0289] The specific implementation of each unit in the media file encapsulation device provided in this embodiment can be referred to the content of the media file encapsulation method described above, and will not be repeated here.
[0290] The media file decapsulation device provided in this disclosure can be installed in any electronic device. In the following example description, it is illustrated by setting it in the intermediate node of an immersive system or the first device (e.g., the player end), but this disclosure is not limited thereto.
[0291] Figure 17 A block diagram schematically illustrates a media file decapsulation apparatus according to an embodiment of the present disclosure. Figure 17 As shown, the media file decapsulation device 1700 provided in this embodiment may include a file receiving unit 1710, a file decapsulation unit 1720, an application scenario acquisition unit 1730, and a decoding and rendering determination unit 1740.
[0292] In this embodiment of the disclosure, the encapsulation file receiving unit 1710 can be used to receive the encapsulation file of the media stream of the target media content in a corresponding application scenario. The encapsulation file includes a first application scenario type field, which indicates the application scenario corresponding to the media stream. The file decapsulation unit 1720 can be used to decapsulate the encapsulation file to obtain the first application scenario type field. The application scenario obtaining unit 1730 can be used to obtain the application scenario corresponding to the media stream based on the first application scenario type field. The decoding and rendering determining unit 1740 can be used to determine the decoding or rendering method of the media stream based on the application scenario corresponding to the media stream.
[0293] In an exemplary embodiment, if the value of the first application scenario type field is a second value indicating that the media stream is a multi-view video with large-scale atlas information, the media file decapsulation device 1700 may further include: a captcha file parsing unit, which can be used to parse the captcha file to obtain the mapping relationship between the texture map and depth map captured by the camera corresponding to the viewpoint contained in the media stream and the planar frames in the large-scale atlas information; a media stream decoding unit, which can be used to decode the media stream to obtain the camera parameters in the media stream; and a multi-view video presentation unit, which can be used to present the multi-view video in three-dimensional space according to the mapping relationship and the camera parameters.
[0294] In an exemplary embodiment, the media file decapsulation apparatus 1700 may further include: a target description file receiving unit, configured to receive a target description file of the target media content, the target description file including a second application scenario type field, the second application scenario type field indicating the application scenario corresponding to the media stream; and a target encapsulation file determining unit, configured to determine the target encapsulation file of the target media stream from the encapsulation file of the media stream according to the second application scenario type field. The encapsulation file receiving unit 1710 may include: a target application scenario determining unit, configured to receive the target encapsulation file to determine the target application scenario of the target media stream according to a first application scenario type field in the target encapsulation file.
[0295] The specific implementation of each unit in the media file decapsulation device provided in this embodiment can be referred to the content of the media file decapsulation method described above, and will not be repeated here.
[0296] It should be noted that although several units of the device for performing actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0297] This disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the media file encapsulation method described in the above embodiments.
[0298] This disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the media file decapsulation method as described in the above embodiments.
[0299] This disclosure provides an electronic device, including: at least one processor; and a storage device configured to store at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the media file encapsulation method as described in the above embodiments.
[0300] This disclosure provides an electronic device, including: at least one processor; and a storage device configured to store at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the media file decapsulation method as described in the above embodiments.
[0301] Figure 18 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown.
[0302] It should be noted that, Figure 18 The illustrated electronic device 1800 is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0303] like Figure 18As shown, the electronic device 1800 includes a central processing unit (CPU) 1801, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1802 or programs loaded from storage portion 1808 into random access memory (RAM) 1803. The RAM 1803 also stores various programs and data required for system operation. The CPU 1801, ROM 1802, and RAM 1803 are interconnected via a bus 1804. An input / output (I / O) interface 1805 is also connected to the bus 1804.
[0304] The following components are connected to I / O interface 1805: an input section 1806 including a keyboard, mouse, etc.; an output section 1807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1808 including a hard disk, etc.; and a communication section 1809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to I / O interface 1805 as needed. Removable media 1811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1810 as needed so that computer programs read from them can be installed into storage section 1808 as needed.
[0305] In particular, according to embodiments of this disclosure, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1809, and / or installed from removable medium 1811. When the computer program is executed by central processing unit (CPU) 1801, it performs various functions defined in the methods and / or apparatus of this application.
[0306] It should be noted that the computer-readable storage medium disclosed herein may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having at least one wire, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF (Radio Frequency), etc., or any suitable combination thereof.
[0307] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods, apparatus, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing at least one executable instruction for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0308] The units described in the embodiments of this disclosure can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the unit itself.
[0309] On the other hand, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable storage medium carries one or more programs that, when executed by the electronic device, cause the electronic device to perform the methods described in the following embodiments. For example, the electronic device may perform... Figure 4 or Figure 5 or Figure 6 or Figure 10 or Figure 11 or Figure 12 or Figure 15 The steps shown.
[0310] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this disclosure.
[0311] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0312] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for encapsulating media files, characterized in that, include: Identify the target media content; Obtain the media bitstream of the target media content in the corresponding application scenario; Encapsulate the media stream for the corresponding application scenario and generate an encapsulation file for the media stream for the corresponding application scenario. The encapsulation file includes a first application scenario type field, which indicates the application scenario corresponding to the media stream. The encapsulation file is sent to the first device so that the first device can obtain the application scenario corresponding to the media stream based on the first application scenario type field in the encapsulation file and determine the decoding or rendering method of the media stream. Add the first application scenario type field to the target media file format data box.
2. The media file encapsulation method according to claim 1, characterized in that, Also includes: If the media stream is encapsulated in a single track, then: Add the number of cameras acquiring the media stream and the number of camera viewpoints corresponding to the cameras contained in the current file of the media stream to the bitstream sample entry of the target media file format data box. Add the location information of the texture map and depth map captured by the camera corresponding to the current file to the bitstream sample entry.
3. The media file encapsulation method according to claim 2, characterized in that, It also includes at least one of the following: Add a downsampling factor to the depth map captured by the camera viewpoint contained in the current file at the bitstream sample inlet; Add the offset of the top left vertex of the texture map captured by the camera corresponding to the current file relative to the origin of the planar frame in the media bitstream to the bitstream sample entry. Add the offset of the top left vertex of the depth map captured by the camera corresponding to the current file relative to the origin of the plane frame in the media bitstream to the bitstream sample entry. Add the guard band width of the texture map and depth map captured by the camera corresponding to the current file in the bitstream sample inlet.
4. The media file encapsulation method according to claim 1, characterized in that, The method further includes: If the media stream is packaged in a multi-track manner, then add the number of cameras that acquired the media stream and the number of camera viewpoints corresponding to the cameras contained in the current file of the media stream to the sample entry of the target media file format data box. Add the location information of the texture map and depth map captured by the camera corresponding to the current file in the sample entry.
5. The media file encapsulation method according to claim 1, characterized in that, The method further includes: Generate a target description file for the target media content, the target description file including an application scenario type field, the application scenario type field indicating the application scenario corresponding to the media stream; The target description file is sent to the first device so that the first device can determine the target encapsulation file of the target media stream from the encapsulation file of the media stream according to the application scenario type field; Sending the encapsulated file to a first device, so that the first device can determine the application scenario corresponding to the media stream based on the encapsulated file, including: The target encapsulation file is sent to the first device so that the first device can determine the target application scenario of the target media stream based on the target encapsulation file.
6. The media file encapsulation method according to claim 5, characterized in that, Generating the target description file for the target media content includes: Add the application scenario type field to the target description file of the target media content for dynamic adaptive streaming media transmission based on the Hypertext Transfer Protocol; The value of the application scenario type field is determined based on the application scenario corresponding to the media stream.
7. The media file encapsulation method according to claim 1, characterized in that, Obtaining the media bitstream of the target media content in the corresponding application scenario includes: Receive the first encapsulated file of the first multi-view video sent by the second device and the second encapsulated file of the second multi-view video sent by the third device; Decapsulate the first encapsulated file and the second encapsulated file respectively to obtain the first multi-view video and the second multi-view video; Decode the first multi-view video and the second multi-view video respectively to obtain the first depth map and the first texture map in the first multi-view video, and the second depth map and the second texture map in the second multi-view video; A merged multi-view video is obtained based on the first depth map, the second depth map, the first texture map, and the second texture map.
8. The media file encapsulation method according to claim 7, characterized in that, The second device is equipped with a first number of cameras, and the third device is equipped with a second number of cameras. The second device and the third device respectively use their respective cameras to capture and shoot multi-view videos of the same scene to obtain the first multi-view video and the second multi-view video.
9. A method for decapsulating media files, characterized in that, include: The encapsulation file of the media stream of the target media content in the corresponding application scenario is received. The encapsulation file includes a first application scenario type field, which indicates the application scenario corresponding to the media stream. De-encapsulate the encapsulation file and obtain the application scenario corresponding to the media stream based on the first application scenario type field in the encapsulation file; Based on the application scenario corresponding to the media stream, determine the decoding or rendering method of the media stream; The target media file format data box includes the first application scenario type field.
10. The media file decapsulation method according to claim 9, characterized in that, in, If the media stream is encapsulated in a single track, the bitstream sample entry of the target media file format data box includes an identifier of the number of cameras that acquired the media stream and an identifier of the number of camera viewpoints corresponding to the current file in the media stream. The bitstream sample entry includes the position information of the texture map and depth map captured by the camera's viewpoint contained in the current file.
11. The media file decapsulation method according to claim 9, characterized in that, The method further includes: Parse the encapsulated file to obtain the mapping relationship between the texture map and depth map captured by the camera's corresponding viewpoint in the media stream and the planar frames in the media stream; Decode the media stream to obtain the camera parameters in the media stream; Based on the mapping relationship and the camera parameters, the media stream is presented in three-dimensional space.
12. The method for decapsulating media files according to any one of claims 9 to 11, characterized in that, The method further includes: Receive the target description file of the target media content, the target description file includes an application scenario type field, the application scenario type field indicates the application scenario corresponding to the media stream; The target encapsulation file of the target media stream is determined from the encapsulation file of the media stream based on the application scenario type field. The encapsulation file for the media stream of the target media content in the corresponding application scenario includes: The target encapsulation file is received to determine the target application scenario of the target media stream based on the target encapsulation file.
13. A media file encapsulation device, characterized in that, include: The media content determination unit is used to determine the target media content; The media stream acquisition unit is used to acquire the media stream of the target media content in the corresponding application scenario; A media stream encapsulation unit is used to encapsulate media streams under a corresponding application scenario and generate an encapsulation file for the media stream under the corresponding application scenario. The encapsulation file includes a first application scenario type field, which indicates the application scenario corresponding to the media stream. A package file sending unit is used to send the package file to a first device, so that the first device can obtain the application scenario corresponding to the media stream based on the first application scenario type field in the package file, and determine the decoding or rendering method of the media stream. The device is also used to add the first application scenario type field to the target media file format data box.
14. The media file encapsulation apparatus according to claim 13, characterized in that, Also includes: A single-track camera viewpoint identifier adding unit is used to add, if the media stream is packaged in a single track, the number of cameras acquiring the media stream and the number of viewpoints corresponding to the cameras contained in the current file of the media stream to the bitstream sample entry of the target media file format data box. The single-track texture depth map resolution adding unit is used to add the position information of the texture map and depth map captured by the camera corresponding to the viewpoint contained in the current file to the bitstream sample inlet.
15. A media file decapsulation device, characterized in that, include: A package file receiving unit is used to receive a package file of the media stream of the target media content in a corresponding application scenario. The package file includes a first application scenario type field, which indicates the application scenario corresponding to the media stream. The file decapsulation unit is used to decapsulate the encapsulated file and obtain the application scenario corresponding to the media stream according to the first application scenario type field in the encapsulated file. The decoding and rendering determination unit is used to determine the decoding or rendering method of the media stream based on the application scenario corresponding to the media stream. The target media file format data box includes the first application scenario type field.
16. The media file decapsulation apparatus according to claim 15, characterized in that, in, If the media stream is encapsulated in a single track, the bitstream sample entry of the target media file format data box includes an identifier of the number of cameras that acquired the media stream and an identifier of the number of camera viewpoints corresponding to the current file in the media stream. The bitstream sample entry includes the position information of the texture map and depth map captured by the camera's viewpoint contained in the current file.
17. An electronic device, characterized in that, include: At least one processor; A storage device configured to store at least one program, which, when executed by the at least one processor, causes the at least one processor to implement a media file encapsulation method as described in any one of claims 1 to 8 or a media file decapsulation method as described in any one of claims 9 to 12.
18. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the media file encapsulation method as described in any one of claims 1 to 8 or the media file decapsulation method as described in any one of claims 9 to 12.
19. A computer program product comprising a computer program that, when executed by a processor, implements a media file encapsulation method as described in any one of claims 1 to 8 or a media file decapsulation method as described in any one of claims 9 to 12.