Media file encapsulation and decapsulation method, device, equipment and storage medium

By determining the recommended window in the file encapsulation device and associating its characteristic information to generate media files, the problems of waste of decoding resources and low efficiency in the prior art are solved, and a more efficient decoding process is achieved.

CN115941995BActive Publication Date: 2025-08-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110970077.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-23
Publication Date
2025-08-08
Estimated Expiration
2041-08-23

AI Technical Summary

Technical Problem

The existing video stream encapsulation method cannot effectively recommend the media resources corresponding to the window, resulting in waste of decoding resources and low decoding efficiency.

Method used

By obtaining the content of the immersive media, determining the recommended window, and associating its characteristic information with the media file, generating a media file, sending instructions information to the file decapsulation device to indicate the metadata of the window, allowing the decapsulation device to request the relevant media file.

Benefits of technology

Save broadband and decoding resources, improve decoding efficiency, and improve the decoding performance of immersive media.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115941995B_ABST
    Figure CN115941995B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device, and storage medium for media file encapsulation and decapsulation, the method comprising: obtaining the content of immersive media, and determining a recommended window of the immersive media based on the content of the immersive media; determining characteristic information of the immersive media corresponding to the recommended window; associating the recommended window with the characteristic information of the immersive media corresponding to the recommended window to generate a media file of the immersive media; and sending a first indication message to a file decapsulation device, the first indication message being used to indicate metadata of the recommended window, the metadata of the recommended window including the characteristic information of the immersive media corresponding to the recommended window. That is, by associating the recommended window with the characteristic information of the immersive media corresponding to the recommended window, the file decapsulation device can request the media file corresponding to the recommended window for consumption based on the characteristic information of the immersive media corresponding to the recommended window, thereby saving broadband and decoding resources and improving decoding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of video processing technology, and in particular to a method, apparatus, device, and storage medium for media file encapsulation and decapsulation. Background Art

[0002] Immersive media refers to media content that can provide consumers with an immersive experience. Based on the degree of freedom users have when consuming media content, immersive media can be divided into 3-degree-of-freedom (DoF) media, 3DoF+ media, and 6DoF media.

[0003] After encapsulating immersive media, the file encapsulation device sends recommended viewport information to the user, allowing them to consume the media resources corresponding to the recommended viewport. However, with current video stream encapsulation methods, while the file encapsulation device can recommend viewports to the file decapsulation device, it cannot recommend the corresponding media resources to the file decapsulation device. This results in wasted decoding resources and low decoding efficiency. Summary of the Invention

[0004] The present application provides a media file encapsulation and decapsulation method, apparatus, device and storage medium. The file decapsulation device can request media files associated with the recommended window, thereby saving bandwidth and decoding resources and improving decoding efficiency.

[0005] In a first aspect, the present application provides a media file encapsulation method, applied to a file encapsulation device, the method comprising:

[0006] Acquiring content of immersive media, and determining a recommended viewport for the immersive media based on the content of the immersive media;

[0007] Determining characteristic information of the immersive media corresponding to the recommended window;

[0008] Associating the recommendation window with characteristic information of the immersive media corresponding to the recommendation window to generate a media file of the immersive media;

[0009] First indication information is sent to a file decapsulation device, where the first indication information is used to indicate metadata of the recommended viewport, where the metadata of the recommended viewport includes feature information of the immersive media corresponding to the recommended viewport.

[0010] In a second aspect, the present application provides a media file encapsulation method, applied to a file decapsulation device, the method comprising:

[0011] receiving first indication information sent by a file encapsulation device, the first indication information being used to indicate metadata of a recommended viewport, the metadata of the recommended viewport including characteristic information of an immersive media corresponding to the recommended viewport, the recommended viewport being determined based on content of the immersive media;

[0012] In response to the first indication information, it is determined whether to request metadata of the recommended viewport.

[0013] In a third aspect, the present application provides a media file encapsulation device, which is applied to a file encapsulation device, and includes:

[0014] an acquisition unit, configured to acquire content of the immersive media and determine a recommended viewport for the immersive media based on the content of the immersive media;

[0015] a processing unit, configured to determine characteristic information of the immersive media corresponding to the recommended viewport;

[0016] an encapsulation unit, configured to associate the recommendation window with characteristic information of the immersive media corresponding to the recommendation window, and generate a media file of the immersive media;

[0017] The transceiver unit is configured to send first indication information to a file decapsulation device, where the first indication information is used to indicate metadata of the recommended viewport, where the metadata of the recommended viewport includes feature information of the immersive media corresponding to the recommended viewport.

[0018] In a fourth aspect, the present application provides a media file decapsulation device, which is applied to a file decapsulation device, and the device includes:

[0019] a transceiver unit, configured to receive first indication information sent by a file encapsulation device, the first indication information being used to indicate metadata of a recommended view, the metadata of the recommended view including characteristic information of an immersive media corresponding to the recommended view, the recommended view being determined based on content of the immersive media;

[0020] The processing unit is configured to determine, in response to the first indication information, whether to request metadata of the recommended window.

[0021] In a fifth aspect, the present application provides a file encapsulation device, comprising: a processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method of the first aspect.

[0022] In a sixth aspect, the present application provides a file decapsulation device, comprising: a processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method of the second aspect.

[0023] In a seventh aspect, a computing device is provided, comprising: a processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method of the first aspect and / or the second aspect.

[0024] In an eighth aspect, a computer-readable storage medium is provided for storing a computer program, wherein the computer program enables a computer to execute the method of the first aspect and / or the second aspect.

[0025] In summary, in the present application, the file encapsulation device obtains the content of the immersive media and determines the recommended window of the immersive media based on the content of the immersive media; determines the characteristic information of the immersive media corresponding to the recommended window; associates the recommended window with the characteristic information of the immersive media corresponding to the recommended window to generate a media file of the immersive media; and sends a first indication message to the file decapsulation device, where the first indication message is used to indicate the metadata of the recommended window, and the metadata of the recommended window includes the characteristic information of the immersive media corresponding to the recommended window. That is, the present application associates the recommended window with the characteristic information of the immersive media corresponding to the recommended window, so that after the file decapsulation device obtains the metadata of the recommended window, it can request the media file of the immersive media corresponding to the recommended window for consumption based on the characteristic information of the immersive media corresponding to the recommended window, without having to apply for the entire media file of the immersive media for consumption, thereby saving broadband and decoding resources and improving decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0027] Figure 1 A schematic diagram of three degrees of freedom is shown schematically;

[0028] Figure 2 A schematic diagram of three degrees of freedom + is shown schematically;

[0029] Figure 3 A schematic diagram of six degrees of freedom is shown schematically;

[0030] Figure 4A An architectural diagram of an immersive media system provided in one embodiment of the present application;

[0031] Figure 4B A schematic diagram of the content flow of V3C media provided in one embodiment of the present application;

[0032] Figure 5 An interactive flow chart of a media file encapsulation and decapsulation method provided in an embodiment of the present application;

[0033] Figure 6 An interactive flow chart of a media file encapsulation and decapsulation method provided in an embodiment of the present application;

[0034] Figure 7 A schematic diagram of a multi-track container provided in accordance with an embodiment of the present application;

[0035] Figure 8 An interactive flow chart of a media file encapsulation and decapsulation method provided in an embodiment of the present application;

[0036] Figure 9 A schematic structural diagram of a media file encapsulation device provided in one embodiment of the present application;

[0037] Figure 10 A schematic diagram of the structure of a media file decapsulation device provided in one embodiment of the present application;

[0038] Figure 11 It is a schematic block diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0040] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0041] The embodiments of the present application relate to data processing technology for immersive media.

[0042] Before introducing the technical solution of this application, the following is an introduction to the relevant knowledge of this application:

[0043] Multi-view / multi-viewpoint video: This refers to video captured from multiple angles using a multi-camera array and containing depth information. Multi-view / multi-viewpoint video, also known as free-viewpoint / free-viewpoint video, is an immersive media that provides a six-degree-of-freedom experience.

[0044] Point cloud: A point cloud is a set of randomly distributed discrete points in space that represent the spatial structure and surface properties of a three-dimensional object or scene. Each point in a point cloud has at least 3D position information and, depending on the application scenario, may also have color, material, or other information. Typically, each point in a point cloud has the same number of additional attributes.

[0045] V3C volumetric media: Visual volumetric video-based coding media refers to immersive media that captures three-dimensional visual content and provides a 3DoF+ or 6DoF viewing experience. It uses traditional video encoding and contains volumetric video tracks in the file encapsulation. This includes multi-view video and video-coded point clouds.

[0046] PCC: Point Cloud Compression, point cloud compression.

[0047] G-PCC: Geometry-based Point Cloud Compression, point cloud compression based on geometric model.

[0048] V-PCC: Video-based Point Cloud Compression, point cloud compression based on traditional video coding.

[0049] Atlas: Indicates the area information on the 2D plane frame, the area information of the 3D presentation space, the mapping relationship between the two, and the necessary parameter information required for the mapping.

[0050] Track: A media file is a collection of media data in the process of media file encapsulation. A media file can be composed of multiple tracks. For example, a media file can contain a video track, an audio track, and a subtitle track.

[0051] Component track refers to the point cloud geometry data track or point cloud attribute data track.

[0052] Sample: A sample is a unit of media file encapsulation. A media track consists of many samples. For example, a sample in a video track is usually a video frame.

[0053] DoF: Degree of Freedom. In a mechanical system, this refers to the number of independent coordinates. In addition to translational degrees of freedom, this also includes rotational and vibrational degrees of freedom. In this embodiment, it refers to the degrees of freedom that allow users to move and interact with content while viewing immersive media.

[0054] 3DoF: Three degrees of freedom, refers to the three degrees of freedom of the user's head rotating around the XYZ axes. Figure 1 The schematic diagram of the three degrees of freedom is shown schematically. Figure 1 As shown, at a certain location or point, you can rotate on three axes: you can turn your head, tilt it up and down, and even tilt it. This three-degree-of-freedom experience allows users to immerse themselves in a scene 360 degrees. If it's static, it can be considered a panoramic image. If the panoramic image is dynamic, it's a panoramic video, or VR video. However, VR videos have certain limitations: users can't move around or choose a specific location to view.

[0055] 3DoF+: In addition to the three degrees of freedom (DOF), users also have limited freedom of movement along the X, Y, and Z axes. This can also be called restricted six degrees of freedom, and the corresponding media stream can be called a restricted six-DOF media stream. Figure 2 A schematic diagram of three degrees of freedom + is shown schematically.

[0056] 6DoF: In addition to the three degrees of freedom, users also have the freedom to move freely along the X, Y, and Z axes. The corresponding media stream can be called a six-degree-of-freedom media stream. Figure 3 A schematic diagram of six degrees of freedom is shown. 6DoF media refers to six-degree-of-freedom video, meaning it provides users with a high-degree-of-freedom viewing experience, allowing them to freely move their viewpoints along the X, Y, and Z axes of three-dimensional space, as well as freely rotate their viewpoints around the X, Y, and X axes. 6DoF media is a combination of videos captured from different spatial perspectives by a camera array. To facilitate the expression, storage, compression, and processing of 6DoF media, 6DoF media data is represented as a combination of the following information: texture maps captured by multiple cameras, depth maps corresponding to the multi-camera texture maps, and corresponding 6DoF media content description metadata. This metadata includes multi-camera parameters and information describing the 6DoF media's stitching layout and edge protection. At the encoder, the multi-camera texture map information and corresponding depth map information are spliced together, and the description of the splicing method is written into the metadata according to defined syntax and semantics. The spliced multi-camera depth map and texture map information are encoded using a planar video compression method. After being transmitted to the terminal for decoding, the user's requested 6DoF virtual viewpoint is synthesized, providing the user with a 6DoF media viewing experience.

[0057] AVS: Audio Video Coding Standard, audio and video coding standard.

[0058] ISOBMFF: ISO Based Media File Format, a media file format based on the ISO (International Standard Organization) standard. ISOBMFF is a media file encapsulation standard, and the most typical ISOBMFF file is the MP4 (Moving Picture Experts Group 4) file.

[0059] DASH: dynamic adaptive streaming over HTTP, dynamic adaptive streaming over HTTP is an adaptive bitrate streaming technology that enables high-quality streaming media to be delivered over the Internet through traditional HTTP network servers.

[0060] MPD: media presentation description, media presentation description signaling in DASH, used to describe media segment information.

[0061] HEVC: High Efficiency Video Coding, international video coding standard HEVC / H.265.

[0062] VVC: versatile video coding, international video coding standard VVC / H.266.

[0063] Intra(picture)Prediction: Intra-frame prediction.

[0064] Inter(picture)Prediction: Inter-frame prediction.

[0065] SCC: screen content coding, screen content coding.

[0066] Immersive media refers to media content that provides consumers with an immersive experience. Based on the degree of freedom users have when consuming media content, immersive media can be categorized into 3DoF media, 3DoF+ media, and 6DoF media. Common 6DoF media include multi-view video and point cloud media.

[0067] Multi-view videos are usually captured by a camera array from multiple angles, forming the scene's texture information (color information, etc.) and depth information (spatial distance information, etc.). Combined with the mapping information from 2D plane frames to 3D presentation space, this constitutes 6DoF media that can be consumed on the user side.

[0068] A point cloud is a collection of randomly distributed discrete points in space that represent the spatial structure and surface properties of a three-dimensional object or scene. Each point in a point cloud contains at least 3D position information and, depending on the application scenario, may also contain color, material, or other information. Typically, each point in a point cloud has the same number of additional attributes.

[0069] Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes, and therefore have a wide range of applications, including virtual reality (VR) games, computer-aided design (CAD), geographic information systems (GIS), autonomous navigation systems (ANS), digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive telepresence, and three-dimensional reconstruction of biological tissues and organs.

[0070] Point clouds are primarily acquired through computer generation, 3D laser scanning, and 3D photogrammetry. Computers can generate point clouds of virtual three-dimensional objects and scenes. 3D scanning can obtain point clouds of static, real-world three-dimensional objects or scenes, generating millions of point clouds per second. 3D imaging can obtain point clouds of dynamic, real-world three-dimensional objects or scenes, generating tens of millions of point clouds per second. Furthermore, in the medical field, MRI, CT, and electromagnetic positioning information can be used to obtain point clouds of biological tissues and organs. These technologies reduce the cost and time required to acquire point cloud data while improving data accuracy. Changes in point cloud data acquisition methods have made it possible to acquire large amounts of point cloud data. With the continuous accumulation of large-scale point cloud data, efficient storage, transmission, publication, sharing, and standardization of point cloud data have become key to point cloud applications.

[0071] After encoding point cloud media, the encoded data stream needs to be encapsulated and transmitted to the user. Correspondingly, the point cloud media player must first decapsulate the point cloud file, then decode it, and finally present the decoded data stream. Therefore, obtaining specific information during the decapsulation process can improve the efficiency of the decoding process to a certain extent, thereby providing a better experience for the presentation of point cloud media.

[0072] Figure 4AThis is an architectural diagram of an immersive media system provided in one embodiment of the present application. Figure 4A As shown, the immersive media system includes an encoding device and a decoding device. The encoding device may refer to a computer device used by the provider of the immersive media, and the computer device may be a terminal (such as a PC (Personal Computer), a smart mobile device (such as a smart phone), etc.) or a server. The decoding device may refer to a computer device used by the user of the immersive media, and the computer device may be a terminal (such as a PC (Personal Computer), a smart mobile device (such as a smart phone), a VR device (such as a VR helmet, VR glasses, etc.)). The data processing process of the immersive media includes the data processing process on the encoding device side and the data processing process on the decoding device side.

[0073] The data processing process on the encoding device side mainly includes:

[0074] (1) The acquisition and production process of immersive media content;

[0075] (2) The process of encoding and file packaging of immersive media. The data processing process on the decoding device side mainly includes:

[0076] (3) The process of decapsulating and decoding immersive media files;

[0077] (4) Rendering process of immersive media.

[0078] In addition, the transmission process of immersive media between the encoding device and the decoding device can be carried out based on various transmission protocols. The transmission protocols here may include but are not limited to: DASH (Dynamic Adaptive Streaming over HTTP, dynamic adaptive streaming media transmission) protocol, HLS (HTTP Live Streaming, dynamic bit rate adaptive transmission) protocol, SMTP (Smart Media Transport Protocol, smart media transmission protocol), TCP (Transmission Control Protocol, transmission control protocol), etc.

[0079] The following will be combined Figure 4A , each process involved in the data processing of immersive media is introduced in detail.

[0080] 1. Data processing process on the encoding device side:

[0081] (1) The process of acquiring and producing media content for immersive media.

[0082] 1) The process of acquiring media content of immersive media.

[0083] A real-world audiovisual scene (A) is captured by audio sensors and a camera array or camera rig with multiple lenses and sensors. This capture produces a set of digital image / video (Bi) and audio (Ba) signals. The cameras / lenses typically cover all directions around a central point of the camera array or rig, hence the term 360-degree video.

[0084] In one implementation, the capture device may refer to a hardware component provided in the encoding device, for example, the capture device may refer to a microphone, camera, sensor, etc. of a terminal. In another implementation, the capture device may also be a hardware device connected to the encoding device, for example, a camera connected to a server.

[0085] The capture device may include, but is not limited to, an audio device, a camera device, and a sensor device. The audio device may include an audio sensor, a microphone, etc. The camera device may include a standard camera, a stereo camera, a light field camera, etc. The sensor device may include a laser device, a radar device, etc.

[0086] Multiple capture devices can be deployed at specific locations in real space to simultaneously capture audio and video content from different angles within that space. The captured audio and video content is synchronized in both time and space. The media content collected by the capture devices is called immersive media raw data.

[0087] 2) The production process of media content for immersive media.

[0088] The captured audio content itself is suitable for audio encoding for immersive media. The captured video content undergoes a series of production processes before it becomes suitable for video encoding for immersive media. The production process includes:

[0089] ① Stitching. Since the captured video content is shot by the capture device at different angles, stitching refers to stitching the video content shot at various angles into a complete video that can reflect the 360-degree visual panorama of the real space. In other words, the stitched video is a panoramic video (or spherical video) represented in three-dimensional space.

[0090] ② Projection. Projection refers to the process of mapping a spliced 3D video onto a 2D image. The resulting 2D image is called a projected image. Projection methods include, but are not limited to, latitude and longitude projection and regular hexahedron projection.

[0091] ③Region encapsulation. The projected image can be encoded directly, or the projected image can be region encapsulated before encoding. In practice, it is found that in the data processing process of immersive media, encoding the two-dimensional projection image after region encapsulation can greatly improve the video coding efficiency of the immersive media. Therefore, the region encapsulation technology is widely used in the video processing process of immersive media. The so-called region encapsulation refers to the process of performing conversion processing on the projection image by region. The region encapsulation process converts the projection image into an encapsulated image. The region encapsulation process specifically includes: dividing the projection image into multiple mapping regions, and then performing conversion processing on the multiple mapping regions to obtain multiple encapsulated regions, and mapping the multiple encapsulated regions to a 2D image to obtain an encapsulated image. Among them, the mapping region refers to the region obtained by division in the projection image before performing region encapsulation; the encapsulation region refers to the region located in the encapsulated image after performing region encapsulation.

[0092] The conversion process may include, but is not limited to, mirroring, rotating, rearranging, upsampling, downsampling, changing the resolution of the region, and moving the region.

[0093] It should be noted that since capture devices can only capture panoramic video, after such video is processed by the encoding device and transmitted to the decoding device for appropriate data processing, users on the decoding device can only view 360-degree video information by performing certain specific actions (such as head rotation). Non-specific actions (such as moving the head) do not produce corresponding video changes, resulting in a poor VR experience. Therefore, it is necessary to provide additional depth information that matches the panoramic video to achieve a better immersion and a more optimal VR experience. This involves 6DoF (six degrees of freedom) production technology. When users can move relatively freely in a simulated scene, it is called 6DoF. When using 6DoF production technology to produce immersive media video content, the capture device generally uses a light field camera, laser equipment, radar equipment, etc. to capture point cloud data or light field data in space. During the above production processes ①-③, some specific processing is also required, such as cutting and mapping the point cloud data and calculating depth information.

[0094] Images (Bi) of the same time instance are concatenated, possibly rotated, projected and mapped onto the packed picture (D).

[0095] (2) The process of encoding and file packaging of immersive media.

[0096] The captured audio content can be directly audio-encoded to form an audio code stream for immersive media. After the above-mentioned production process ①-② or ①-③, the projected image or encapsulated image is video-encoded to obtain a video code stream for immersive media. For example, the packaged picture (D) is encoded into an encoded image (Ei) or an encoded video bit stream (Ev). The captured audio (Ba) is encoded into an audio bit stream (Ea). Then, according to a specific media container file format, the encoded images, videos and / or audio are combined into a media file (F) for file playback or a sequence of initialization segments and media segments for streaming (Fs). The encoding device also includes metadata, such as projection and region information, into the file or fragment to facilitate the presentation of the decoded packaged picture.

[0097] It should be noted here that if 6DoF production technology is used, a specific encoding method (such as point cloud encoding) needs to be used for encoding during the video encoding process. The audio stream and the video stream are encapsulated in a file container according to the file format of immersive media (such as ISOBMFF (ISO Base Media File Format, ISO base media file format)) to form a media file resource of immersive media. The media file resource can be a media file or a media fragment to form a media file of immersive media; and according to the file format requirements of immersive media, the media presentation description information (MPD) is used to record the metadata of the media file resource of the immersive media. The metadata here is a general term for information related to the presentation of immersive media. The metadata may include description information of the media content, description information of the window, and signaling information related to the presentation of media content, etc. As Figure 4A As shown, the encoding device stores the media presentation description information and media file resources formed after the data processing process.

[0098] The immersive media system supports data boxes, which are data blocks or objects containing metadata. Specifically, a data box contains metadata about the corresponding media content. Immersive media can include multiple data boxes, such as a Sphere Region Zooming Box, which contains metadata describing sphere region zooming information; a 2D Region Zooming Box, which contains metadata describing 2D region zooming information; a Region Wise Packing Box, which contains metadata describing the corresponding information during the region packing process, and so on.

[0099] The fragment Fs is delivered to the player using a delivery mechanism.

[0100] 2. Data processing on the decoding device side:

[0101] (3) The process of decapsulating and decoding immersive media files;

[0102] The decoding device can obtain the media file resources and corresponding media presentation description information of the immersive media from the encoding device through the recommendation of the encoding device or adaptively and dynamically according to the user needs of the decoding device. For example, the decoding device can determine the user's orientation and position based on the user's head / eye / body tracking information, and then dynamically request the encoding device to obtain the corresponding media file resources based on the determined orientation and position. The media file resources and media presentation description information are transmitted from the encoding device to the decoding device through a transmission mechanism (such as DASH, SMT). The file decapsulation process on the decoding device side is the opposite of the file encapsulation process on the encoding device side. The decoding device decapsulates the media file resources according to the file format requirements of the immersive media to obtain audio streams and video streams. The decoding process on the decoding device side is the opposite of the encoding process on the encoding device side. The decoding device performs audio decoding on the audio stream to restore the audio content.

[0103] In addition, the decoding process of the video stream by the decoding device includes the following:

[0104] ① Decoding the video stream to obtain a planar image; based on the metadata provided by the media presentation description information, if the metadata indicates that the immersive media has performed a region encapsulation process, the planar image is a packaged image; if the metadata indicates that the immersive media has not performed a region encapsulation process, the planar image is a projected image;

[0105] ② If the metadata indicates that the immersive media has performed a regional encapsulation process, the decoding device will perform regional decapsulation on the encapsulated image to obtain a projected image. Here, regional decapsulation is the opposite of regional encapsulation. Regional decapsulation refers to the process of performing inverse conversion processing on the encapsulated image according to the region. Regional decapsulation converts the encapsulated image into a projected image. The process of regional decapsulation specifically includes: performing inverse conversion processing on multiple encapsulated regions in the encapsulated image according to the instructions of the metadata to obtain multiple mapping regions, and mapping the multiple mapping regions to a 2D image to obtain a projected image. Inverse conversion processing refers to processing that is opposite to the conversion processing. For example, if the conversion processing refers to a counterclockwise rotation of 90 degrees, then the inverse conversion processing refers to a clockwise rotation of 90 degrees.

[0106] ③ Reconstruct the projected image according to the media presentation description information to convert it into a 3D image. The reconstruction process here refers to the process of re-projecting the two-dimensional projected image into a 3D space.

[0107] (4) Rendering process of immersive media.

[0108] The decoding device renders the audio content obtained by audio decoding and the 3D image obtained by video decoding according to the metadata related to rendering and viewport in the media presentation description information. Once the rendering is completed, the playback output of the 3D image is realized. In particular, if 3DoF and 3DoF+ production technologies are adopted, the decoding device mainly renders the 3D image based on the current viewpoint, parallax, depth information, etc. If 6DoF production technology is adopted, the decoding device mainly renders the 3D image in the viewport based on the current viewpoint. Among them, the viewpoint refers to the user's viewing position, the parallax refers to the difference in line of sight between the user's two eyes or the difference in line of sight caused by movement, and the viewport refers to the viewing area.

[0109] The immersive media system supports data boxes, which are data blocks or objects containing metadata. Specifically, a data box contains metadata about the corresponding media content. Immersive media can include multiple data boxes, such as a Sphere Region Zooming Box, which contains metadata describing sphere region zooming information; a 2D Region Zooming Box, which contains metadata describing 2D region zooming information; and a Region Wise Packing Box, which contains metadata describing the corresponding information during the region packing process.

[0110] For example Figure 4A As shown, the file (F) output by the encoding device is the same as the file (F') input by the decoding device. The decoding device processes the file (F') or the received fragments (F's) to extract the encoded bitstream (E'a, E'v and / or E'i) and parse the metadata. Viewport-related video data can be carried in multiple tracks, which can be rewritten in the bitstream before decoding and merged into a single video bitstream E'v. The audio, video and / or image are then decoded into decoded signals (B'a is an audio signal, and D' is an image / video signal). Based on the current viewing direction or viewport, as well as information such as projection, spherical coverage, rotation and area in the metadata, the decoded image / video (D') is displayed on the screen of a head-mounted display or any other display device. The current viewing direction is determined by head tracking information and / or eye tracking information. At the same time, the decoded audio signal (B'a) is rendered and listened to by the user through headphones, for example. In addition to the rendering of video and audio signals, the current viewing direction can also be used to optimize decoding. In viewport-dependent delivery, the current viewing direction is also passed to the policy module, which determines which video track to receive based on the viewing direction.

[0111] Figure 4B A schematic diagram of the content flow of V3C media provided in an embodiment of the present application is shown as follows: Figure 4BAs shown, the immersive media system includes a file encapsulator and a file decapsulator. In some embodiments, the file encapsulator can be understood as the encoding device mentioned above, and the file decapsulator can be understood as the decoding device mentioned above.

[0112] A real-world or synthetic visual scene (A) is captured by a set of cameras, a camera device with multiple lenses and sensors, or a virtual camera. The result is source volume data (B). One or more volume frames are encoded as a V3C bitstream, consisting of an atlas bitstream, at most one occupancy bitstream, a geometry bitstream, and zero or more attribute bitstreams (Ev).

[0113] One or more encoded bitstreams are then packaged into a media file (F) for local playback or a sequence of initialization segments and media segments (Fs) for streaming, based on a specific media container file format. The media container file format is the ISO base media file format specified in ISO / IEC 14496-12. The file encapsulator may also include metadata within the file or segment. The segments Fs are delivered to the player using a delivery mechanism.

[0114] The file (F) output by the file encapsulator is the same file (F') that was taken as input by the file decapsulator. The file decapsulator processes the file (F') or the received fragments (F's) to extract the encoded bitstream (E'v) and parse the metadata. The V3C bitstream is then decoded into a decoded signal (D'). Based on the current viewing direction or viewport, the decoded signal (D') is reconstructed, rendered, and displayed on the screen of a head-mounted display or any other display device. The current viewing direction is determined by head tracking information and / or eye tracking information. In viewport-related delivery, the current viewing direction is also passed to the strategy module, which determines the track to receive based on the viewing direction.

[0115] The above process works for both real-time and on-demand use cases.

[0116] The following is an introduction to the grammatical structures involved in the embodiments of this application:

[0117] 1.1.1 External camera information

[0118] 1.1.1.1 Syntax

[0119]

[0120] 1.1.1.2 Semantics

[0121] cam_pos_x, cam_pos_y, and cam_pos_z: represent the x, y, and z coordinates of the camera position in meters in the global reference coordinate system, respectively. These values should be represented in 32-bit binary floating point format, where the 4 bytes are in big-endian order and parsed according to the parsing process specified in IEEE 754.

[0122] cam_quat_x, cam_quat_y, and cam_quat_z: These represent the x, y, and z components of the camera rotation expressed as a quaternion. These values should be between –2 30 to 2 30 The range includes 2 30 and 2 30 When no rotation component is present, its value shall be inferred to be equal to 0. The value of the rotation component can be calculated as follows:

[0123] qX=cam_quat_x÷2 30 ,

[0124] qY=cam_quat_y÷2 30 ,

[0125] qZ=cam_quat_z÷2 30 .

[0126] The fourth component qW of the current camera model rotation expressed using a quaternion is calculated as follows:

[0127] qW=Sqrt(1–(qX 2 +qY 2 +qZ 2 ))

[0128] The point (w,x,y,z) represents a rotation around the axis pointed by the vector (x,y,z) by an angle of 2*cos^{-1}(w)=2*sin^{-1}(sqrt(x^{2}+y^{2}+z^{2})).

[0129] Note that, consistent with ISO / IEC FDIS 23090-5, qW is always positive. If a negative qW is required, all three syntax elements, cam_quat_x, cam_quat_y, and cam_quat_z, can be expressed with opposite signs, which is equivalent.

[0130] 1.1.2 Camera intrinsic information

[0131] 1.1.2.1 Syntax

[0132]

[0133]

[0134] 1.1.2.2 Semantics

[0135] camera_id: is an identifier number used to identify the camera parameters of a given viewport.

[0136] camera_type: Indicates the projection mode of the viewport camera. A value of 0 specifies ERP projection. A value of 1 specifies perspective projection. A value of 2 specifies orthographic projection. Values in the range of 3 to 255 are reserved for future use by ISO / IEC.

[0137] erp_horizontal_fov: Specifies the longitude extent of the ERP projection corresponding to the horizontal size of the viewport area, in radians. This value should be in the range of 0 to 2π.

[0138] erp_vertical_fov: Specifies the latitude extent of the ERP projection corresponding to the vertical size of the viewport area, in radians. This value should be in the range of 0 to π.

[0139] perspective_horizontal_fov: Specifies the horizontal field of view of the perspective projection in radians. Values should be in the range of 0 and π. Perspective aspect ratio specifies the relative aspect ratio of the perspective projection (horizontal / vertical) viewport. The value should be represented in 32-bit binary floating point format, where the 4 bytes are in big-endian order and parsed according to the parsing process specified in IEEE 754.

[0140] ortho_aspect_ratio: Specifies the relative aspect ratio of the orthographic projection (horizontal / vertical) viewport. This value should be represented in 32-bit binary floating point format, where the 4 bytes are in big-endian order and parsed according to the parsing process specified in IEEE754.

[0141] ortho_horizontal_size: Specifies the horizontal size of the orthogonal in meters. This value should be represented in 32-bit binary floating point format, where the 4 bytes are in big-endian order and parsed according to the parsing process specified in IEEE 754.

[0142] clipping_near_plane and clipping_far_plane: Indicates the near and far depths (or distances) of the near and far clipping planes (in meters) based on the viewport. These values should be represented in 32-bit binary floating point format, where the 4 bytes are in big-endian order and parsed according to the parsing process specified in IEEE 754.

[0143] 1.1.3 Viewport Information

[0144] 1.1.3.1 Syntax

[0145]

[0146]

[0147] 1.1.3.2 Semantics

[0148] center_view_flag: is a flag that indicates whether the signaled viewport position corresponds to the center of the viewport or to one of the two stereo positions of the viewport. A value of 1 indicates that the signaled viewport position corresponds to the center of the viewport. A value of 0 indicates that the signaled viewport position corresponds to one of the two stereo positions of the viewport.

[0149] left_view_flag: A flag indicating whether the viewport information being sent corresponds to the left stereo position of the right stereo position of the viewport. A value of 1 indicates that the signaled viewport information corresponds to the left stereo position of the viewport. A value of 0 indicates that the signaled viewport information corresponds to the right stereo position of the viewport.

[0150] extCamInfo: is an instance of the external camera information structure, used to define the external camera parameters of the viewport.

[0151] intCamInfo: is an instance of the internal camera information structure, defining the intrinsic camera parameters of the viewport.

[0152] 1.2 Viewport Information Timed Metadata Track

[0153] 1.2.1 General

[0154] This clause describes the use of a timed metadata track to send viewport information in the V3C transport format, consisting of intrinsic and extrinsic camera parameters, including viewport position and rotation information and viewport camera parameters. To represent the viewport information in a V3C bitstream, the viewport information timed metadata track only references the relevant V3C atlas track, and not the V3C video component track directly.

[0155] Viewport information timed metadata tracks containing a "cdtg" track reference collectively describe the referenced track and track group. When a timed metadata track is linked to one or more V3C atlas tracks with a "cdsc" track reference, it describes each V3C atlas track individually.

[0156] Any sample in the viewport information timed metadata track can be marked as a sync sample. For a specific sample in the timed metadata track, if at least one media sample with the same decoding time in the referenced V3C atlas track is a sync sample, then the specific sample should be marked as a sync sample. Otherwise, the sample may or may not be marked as a sync sample.

[0157] 1.2.2 Viewport Information Example Entry

[0158] 1.2.2.1 Definition

[0159] Data box type: '6vpt'

[0160] Contained in:Sample Description Box ('stsd')

[0161] Is it mandatory: No

[0162] Quantity: 0 or 1

[0163] A sample entry for viewport information associated with the V3C transport format is defined by ViewportInfoSampleEntry.

[0164] A viewport info sample entry should contain a ViewportInfoConfigurationBox describing the viewport type and (if applicable to all samples for the track) the intrinsic and / or extrinsic camera parameters.

[0165] The codec parameter value for this track as defined in RFC 6381 SHOULD be set to "6vpt".

[0166] 1.2.2.2 Syntax

[0167]

[0168]

[0169] 1.2.2.3 Semantics

[0170] viewport_type: indicates the viewport type of all samples corresponding to the current sample entry. Its value meaning is shown in Table 1 below.

[0171] Table 1

[0172]

[0173] viewport_description: A null-terminated string providing a textual description of the recommended viewport.

[0174] dynamic_int_camera_flag: A value of 0 indicates that the camera intrinsic parameters of all samples corresponding to the current sample entry are fixed. If dynamic_ext_camera_flag is 0, dynamic_int_camera_flag must also be 0.

[0175] dynamic_ext_camera_flag: A value of 0 indicates that the camera extrinsics of all samples corresponding to the current sample entry remain unchanged.

[0176] For viewport_type equal to 3, the timed metadata indicates the recommended initial viewport information when playing the associated V3C media track, consisting of the initial viewport position and rotation. When intending to start playback of a media track using another viewport, the initial viewport position (cam_pos_x, cam_pos_y, cam_pos_z) is equal to (0,0,0) relative to the world coordinate axes and the initial view rotation (cam_quat_x, cam_quat_y, cam_quat_z) is equal to (0,0,0) relative to the world coordinate axes. This metadata track should be present and associated with the media track. In the absence of this type of metadata, for the initial viewport, cam_pos_x, cam_pos_y, cam_pos_z, cam_quat_x, cam_quat_y, and cam_quat_z should all be inferred to be equal to 0.

[0177] 1.2.3 Viewport Information Example Format

[0178] Each viewport sample comes with a set of viewports of the type defined in the associated sample entry. The parameters for each viewport include the extrinsic and intrinsic camera information parameters described by IntCameraInfoStruct and ExtCameraInfoStruct. While the extrinsic camera information parameters described by ExtCameraInfoStruct are expected to be present in every sample, the intrinsic camera parameters described by IntCameraInfoStruct are only present in a sample if the intrinsic camera parameters signaled in an earlier sample are no longer applicable.

[0179] If not modified, extrinsic or intrinsic camera parameters previously defined for a viewport from an earlier sample will remain unchanged.

[0180] 1.2.3.1 Syntax

[0181]

[0182] 1.2.3.2 Semantics

[0183] If the viewport info timed metadata track is present, the extrinsic camera parameters represented by ExtCameraInfoStruct() shall be present at the sample entry or sample level. The following two conditions shall not occur at the same time: dynamic_ext_camera_flag[i] is equal to 0 and camera_extrinsic_flag[i] is equal to 0 for all samples.

[0184] num_viewports: Indicates the number of viewports signaled in the sample.

[0185] viewport_id[i]: is the identifier used to identify the i-th viewport.

[0186] viewport_cancel_flag[i]: equal to 1, indicating that the viewport with id viewport_id[i] is canceled. The viewport information indicating the i-th viewport is as follows.

[0187] camera_intrinsic_flag[i]: equal to 1 indicates that the intrinsic camera parameters are present in the i-th viewport of the current sample. If dynamic_int_camera_flag[i] is equal to 0, this should be equal to 0. In addition, when camera_extrinsic_flag[i] is equal to 0, it should be set to 0.

[0188] camera_extrinsic_flag[i]: equal to 1 indicates that the extrinsic camera parameters are present in the i-th viewport of the current sample. If dynamic_ext_camera_flag[i] is equal to 0, it should be equal to 0.

[0189] As can be seen above, current technology defines the viewport structure of immersive media and its associated temporal metadata. However, it fails to integrate the viewport with viewpoint selection and point cloud segment selection at different quality levels. This prevents file decapsulation devices from requesting only media resources related to the recommended viewport. This results in wasted decoding resources and low decoding efficiency.

[0190] In order to solve the above technical problems, the present application associates the recommendation window with the characteristic information of the immersive media corresponding to the recommendation window, that is, the characteristic information of the immersive media corresponding to the recommendation window is included in the metadata of the recommendation window. In this way, after the file decapsulation device obtains the metadata of the recommendation window, it can request the media file of the immersive media corresponding to the recommendation window for consumption based on the characteristic information of the immersive media corresponding to the recommendation window, without requesting the entire media file of the immersive media for consumption, thereby saving broadband and decoding resources and improving decoding efficiency.

[0191] The following describes the technical solutions of the embodiments of the present application in detail through some embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0192] Figure 5 An interactive flow chart of a media file encapsulation and decapsulation method provided in an embodiment of the present application, such as Figure 5As shown, the method includes the following steps:

[0193] S501: The file packaging device obtains the content of the immersive media and determines a recommended window of the immersive media according to the content of the immersive media.

[0194] In some embodiments, the file encapsulation device is also referred to as a video encapsulation device or a video encoding device.

[0195] The embodiment of the present application does not limit the specific type of immersive media, and it can be any existing type of immersive media.

[0196] In an exemplary embodiment, the immersive media may be multi-perspective video media.

[0197] In another example, the immersive media may be point cloud media.

[0198] In another example, the immersive media includes both multi-perspective video media and point cloud media.

[0199] When the immersive media is multi-view video media, the content of the immersive media is also referred to as multi-view video data. When the immersive media is point cloud media, the content of the immersive media is also referred to as point cloud data.

[0200] In the embodiment of the present application, the file encapsulation device obtains the immersive media content in the following ways, but is not limited to:

[0201] Method 1: The file encapsulation device obtains immersive media content from a capture device. For example, the file encapsulation device obtains multi-view video data captured by cameras with multiple viewpoints, or obtains point cloud data from a point cloud capture device.

[0202] Method 2: The file encapsulation device obtains the immersive media content from the storage device. For example, after a camera with multiple perspectives captures multi-perspective video data, the multi-perspective video data is stored in the storage device. Alternatively, after a point cloud acquisition device captures point cloud data, the point cloud data is stored in the storage device.

[0203] The embodiment of the present application does not limit the method of determining the recommended window of immersive media according to the content of the immersive media. For details, reference may be made to existing technologies and will not be repeated here.

[0204] S502: The file packaging device determines characteristic information of the immersive media corresponding to the recommended window.

[0205] The characteristic information of the immersive media described in the embodiments of the present application can be understood as information that can uniquely indicate the immersive media. For example, if the immersive media is multi-view video media, the characteristic information of the immersive media can include perspective information or camera information corresponding to the immersive media.

[0206] In one embodiment, if the immersive media is multi-view video media, and the video tracks of the immersive media are divided according to viewpoints or viewpoint groups, then S502 includes:

[0207] S502-A: The file packaging device determines the perspective information of the multi-perspective video media corresponding to the recommended window as the characteristic information of the immersive media corresponding to the recommended window.

[0208] The perspective information of the multi-perspective video media corresponding to the recommended window includes at least one of viewpoint group information, viewpoint information, and camera information of the multi-perspective video media corresponding to the recommended window.

[0209] In one example, if the perspective information of the multi-perspective video media corresponding to the recommended window is viewpoint group information, the viewpoint group information includes: the number of viewpoint groups associated with the recommended window and the identifier of the viewpoint group associated with the recommended window.

[0210] In an example, if the perspective information of the multi-perspective video media corresponding to the recommended window is viewpoint information, the viewpoint information includes: the number of viewpoints associated with the recommended window and identifiers of the viewpoints associated with the recommended window.

[0211] In an example, if the perspective information of the multi-perspective video media corresponding to the recommended window is camera information, the camera information includes: the number of cameras associated with the recommended window and the identifiers of the cameras associated with the recommended window.

[0212] In one embodiment, if the immersive media is point cloud media, and the point cloud media is packaged according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, then S502 includes:

[0213] S502-B: The file packaging device determines the replaceable group selection information of the point cloud blocks corresponding to the recommended window as feature information of the immersive media corresponding to the recommended window.

[0214] The replaceable group selection information of the point cloud block includes at least one of identification information of the component track corresponding to the point cloud block and a quality level corresponding to the point cloud block.

[0215] It should be noted that the component track can be understood as the track that encapsulates the data code stream of the point cloud block, where the component track can include: Occ. track, Geo track, Att track, etc.

[0216] In a possible implementation, the replaceable group selection information of the point cloud block further includes at least one of the number of replaceable groups corresponding to the point cloud block, the identifier of the replaceable group, and the number of component tracks selected from the replaceable group.

[0217] As can be seen from the above, according to the above method, for different types of immersive media, characteristic information of the immersive media corresponding to the recommended window is determined, and then the following S503 is executed.

[0218] S503: The file packaging device associates the recommended window with characteristic information of the immersive media corresponding to the recommended window to generate a media file of the immersive media.

[0219] In one embodiment, if the characteristic information of the immersive media corresponding to the recommended window is perspective information of the multi-perspective video media corresponding to the recommended window, the above S503 includes:

[0220] S503-A: The file packaging device associates the recommended view with the view information of the multi-view video media corresponding to the recommended view.

[0221] In this step, the file encapsulation device associates the recommended window with the perspective information of the multi-view video media corresponding to the recommended window. This can be understood as adding the perspective information of the multi-view video media corresponding to the recommended window to the metadata of the recommended window. This allows the file decapsulation device to determine the perspective information associated with the recommended window based on the metadata of the recommended window and subsequently request decoding of the media file associated with the perspective information of the recommended window, thereby saving bandwidth and decoding resources and improving decoding efficiency.

[0222] In some embodiments, if the encapsulation standard of the above media file is ISOBMFF, the perspective information data structure of the multi-perspective video media corresponding to the recommended window is as follows:

[0223]

[0224] Among them, num_view_groups: indicates the number of viewpoint groups associated with the recommended window.

[0225] view_group_id: an identifier indicating a view group.

[0226] num_views: Indicates the number of viewpoints associated with the recommended viewport. The value of num_view_groups and the value of num_views cannot be 0 at the same time.

[0227] view_id: an identifier indicating the viewpoint to which the recommended viewport is associated.

[0228] Alternatively, a camera identifier can be used instead of a viewpoint identifier, as follows:

[0229] num_cameras: indicates the number of shooting cameras.

[0230] camera_id: Indicates the identifier of each capturing camera.

[0231] Optionally, the perspective information data structure of the multi-perspective video media corresponding to the recommended viewport may be added to the metadata sample of the recommended viewport.

[0232] In one embodiment, if the characteristic information of the immersive media corresponding to the recommended window is the replaceable group selection information of the point cloud segment corresponding to the recommended window, the above S503 includes:

[0233] S503-B: The file packaging device associates the recommended window with the replaceable group selection information of the point cloud block corresponding to the recommended window.

[0234] In this step, the file encapsulation device associates the recommended window with the alternative group selection information for the point cloud segment corresponding to the recommended window, which can be understood as adding the alternative group selection information for the point cloud segment corresponding to the recommended window to the metadata of the recommended window. This allows the file decapsulation device to determine the alternative group selection information associated with the recommended window based on the metadata of the recommended window, such as the identification information of the component track associated with the recommended window and / or the quality level of the point cloud segment corresponding to the recommended window, and then request decoding of the component track associated with the recommended window, or request decoding of the component track corresponding to the quality level of the point cloud segment corresponding to the recommended window, thereby saving bandwidth and decoding resources and improving decoding efficiency.

[0235] If the packaging standard of the media file of the above point cloud media is ISOBMFF, the current standard defines the replaceable group selection information data structure as follows:

[0236]

[0237] Among them, alternative_type: the difference attribute type of the alternative track. Depending on the value of the difference type, the track can have one or more difference attributes.

[0238] quality_ranking: quality ranking information. The smaller the value of this field, the higher the quality of the corresponding track.

[0239] lossless_flag: If the value of this field is 0, it indicates that the corresponding track uses lossy coding; if the value of this field is 1, it indicates that the corresponding track uses lossless coding.

[0240] Bitrate: Bitrate information, indicating the bitrate of the corresponding track.

[0241] Framerate: Frame rate information, indicating the frame rate of the corresponding track.

[0242] codec_type: Coding type, indicating the coding type of the corresponding track.

[0243] Optionally, the present application can add the number of alternative groups corresponding to the point cloud block corresponding to the recommended window, the identifier of the alternative group, the number of component tracks selected from the alternative group, the identifier information of the component tracks corresponding to the point cloud block, the quality level corresponding to the point cloud block, etc. to the above-mentioned existing alternative group selection information data structure. The specific data structure is as follows:

[0244]

[0245] Among them, num_alternative_groups: the number of alternative groups corresponding to the point cloud block.

[0246] alternate_group_id: Indicates the identifier of each alternate group.

[0247] num_selections: Indicates the number of component tracks to select from the replaceable group.

[0248] track_id: indicates the identification information of the component track corresponding to the point cloud block.

[0249] AlternativeInfoStruct: indicates the quality level corresponding to the point cloud block.

[0250] Optionally, the data structure of the replaceable group selection information of the point cloud blocks corresponding to the recommended viewport may be added to the metadata sample of the recommended viewport.

[0251] S504: The file encapsulation device sends first indication information to the file decapsulation device. The first indication information is used to indicate metadata of the recommended window. The metadata of the recommended window includes feature information of the immersive media corresponding to the recommended window.

[0252] S505: The file decapsulation device determines whether to request metadata of the recommended window in response to the first indication information.

[0253] Specifically, the file encapsulation device, according to the methods of S501 to S503 above, adds characteristic information of the immersive media corresponding to the recommended view to the metadata of the recommended view, thereby generating a metadata track for the recommended view. Next, the file encapsulation device sends first indication information to the file encapsulation device. This first indication information may be DASH signaling, indicating the metadata of the recommended view. For example, the first indication information includes a track identifier for the metadata track of the recommended view. Upon receiving this first indication information, the file decapsulation device determines whether to request metadata for the recommended view based on the current application scenario.

[0254] The media file encapsulation and decapsulation method of the embodiment of the present application is as follows: the file encapsulation device obtains the content of the immersive media and determines the recommended window of the immersive media based on the content of the immersive media; determines the characteristic information of the immersive media corresponding to the recommended window; associates the recommended window with the characteristic information of the immersive media corresponding to the recommended window to generate a media file of the immersive media; and sends a first indication message to the file decapsulation device, the first indication message being used to indicate the metadata of the recommended window, the metadata of the recommended window including the characteristic information of the immersive media corresponding to the recommended window. That is, the present application associates the recommended window with the characteristic information of the immersive media corresponding to the recommended window, so that after the file decapsulation device obtains the metadata of the recommended window, it can request the media file of the immersive media corresponding to the recommended window for consumption based on the characteristic information of the immersive media corresponding to the recommended window, without having to request the entire media file of the immersive media for consumption, thereby saving broadband and decoding resources and improving decoding efficiency.

[0255] Figure 6 An interactive flow chart of a media file encapsulation and decapsulation method provided in an embodiment of the present application, such as Figure 6 As shown, the method includes the following steps:

[0256] S601: The file packaging device obtains the content of the immersive media and determines a recommended window of the immersive media according to the content of the immersive media.

[0257] S602: The file packaging device determines characteristic information of the immersive media corresponding to the recommended window.

[0258] S603: The file packaging device associates the recommended window with characteristic information of the immersive media corresponding to the recommended window to generate a media file of the immersive media.

[0259] S604: The file encapsulation device sends first indication information to the file decapsulation device. The first indication information is used to indicate metadata of the recommended window. The metadata of the recommended window includes feature information of the immersive media corresponding to the recommended window.

[0260] The above steps S601 to S604 are the same as the above steps S501 to S504. Please refer to the description of the above steps S501 to S504 and will not be repeated here.

[0261] S605: The file decapsulation device sends first request information to the file encapsulation device in response to the first instruction information, where the first request information is used to request metadata of the recommended window.

[0262] Specifically, after receiving the first indication information, the file decapsulation device determines whether to request metadata for the recommended window based on the current application scenario. For example, if the file decapsulation device determines that the current network is poor, the device is malfunctioning, or the recommended window is not being consumed, then the metadata for the recommended window is not requested. If the file decapsulation device determines that the current network is good, or the recommended window is being consumed, then the metadata for the recommended window is requested.

[0263] S606: The file encapsulation device sends the metadata track of the recommended window to the file decapsulation device according to the first request information.

[0264] S607: The file decapsulation device decapsulates and decodes the metadata track of the recommended window to obtain metadata of the recommended window.

[0265] Specifically, the file decapsulation device decapsulates the metadata track of the recommended window to obtain a code stream of the metadata of the recommended window, and then decodes the code stream of the metadata of the recommended window to obtain the metadata of the recommended window.

[0266] S608: The file decapsulation device sends second request information to the file encapsulation device according to the characteristic information of the immersive media corresponding to the recommended window in the metadata of the recommended window.

[0267] The second request information is used to request a media file of the immersive media corresponding to the recommended window.

[0268] In some embodiments, if the immersive media is multi-view video media, and the video tracks of the immersive media are divided according to viewpoints or viewpoint groups, and the recommended window is associated with the viewpoint information of the multi-view video media corresponding to the recommended window, that is, the metadata of the recommended window includes the viewpoint information of the multi-view video media corresponding to the recommended window, then the above S608 includes S608-A:

[0269] S608-A. The file decapsulation device sends a second request message to the file encapsulation device according to the perspective information of the multi-perspective video media corresponding to the recommended window. The second request message includes the perspective information of the multi-perspective video media corresponding to the recommended window.

[0270] For example, if the viewing angle information is viewpoint information, the second request information includes viewpoint identification information, so that the file encapsulation device can send the media file corresponding to the viewpoint to the file decapsulation device based on the viewpoint identification information. If the viewing angle information is viewpoint group information, the second request information includes viewpoint group identification information, so that the file encapsulation device can send the media file corresponding to the viewpoint group to the file decapsulation device based on the viewpoint group identification information. If the viewing angle information is camera information, the second request information includes camera identification information, so that the file encapsulation device can send the media file corresponding to the camera to the file decapsulation device based on the camera identification information.

[0271] In some embodiments, if the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window, that is, the metadata of the recommended window includes the replaceable group selection information of the point cloud block corresponding to the recommended window, then the above S608 includes S608-B:

[0272] S608-B. The file decapsulation device sends second request information to the file encapsulation device according to the replaceable group selection information of the point cloud block corresponding to the recommended window. The second request information includes the replaceable group selection information of the point cloud block.

[0273] It can be seen from the above embodiments that the replaceable group selection information of the point cloud block includes at least one of the identification information of the component track corresponding to the point cloud block and the quality level corresponding to the point cloud block.

[0274] In one example, if the replaceable group selection information of the point cloud block includes identification information of the component track corresponding to the point cloud block, the corresponding second request information includes identification information of the component track corresponding to the point cloud block, so that the file encapsulation device can send the component track to the file decapsulation device based on the identification information of the component track.

[0275] In one example, if the replaceable group selection information of the point cloud block includes the quality level corresponding to the point cloud block, the corresponding second request information includes the quality level corresponding to the point cloud block, so that the file encapsulation device can send the component track corresponding to the quality level to the file decapsulation device according to the quality level.

[0276] S609: The file encapsulation device sends the media file of the immersive media corresponding to the recommended window to the file decapsulation device according to the second request information.

[0277] In some embodiments, if the second request information includes perspective information of the multi-perspective video media corresponding to the recommended window, then the above S609 includes S609-A:

[0278] S609-A: Send the media file corresponding to the viewing angle information to the file decapsulation device.

[0279] For example, if the viewing angle information is viewpoint information and the second request information includes viewpoint identification information, the file encapsulation device can send the media file corresponding to the viewpoint to the file decapsulation device based on the viewpoint identification information. If the viewing angle information is viewpoint group information and the second request information includes viewpoint group identification information, the file encapsulation device can send the media file corresponding to the viewpoint group to the file decapsulation device based on the viewpoint group identification information. If the viewing angle information is camera information and the second request information includes camera identification information, the file encapsulation device can send the media file corresponding to the camera to the file decapsulation device based on the camera identification information.

[0280] In some embodiments, if the second request information includes alternative group selection information of the point cloud segment, the above S609 includes:

[0281] S609-B. If the replaceable group selection information of the point cloud block includes the identification information of the component track corresponding to the point cloud block, the component track corresponding to the point cloud block is sent to the file decapsulation device; or, if the replaceable group selection information of the point cloud block includes the quality level corresponding to the point cloud block, the component track corresponding to the quality level is sent to the file decapsulation device.

[0282] In one example, if the second request information includes identification information of a component track corresponding to the point cloud block, the file encapsulation device may send the component track to the file decapsulation device according to the identification information of the component track.

[0283] In one example, if the second request information includes a quality level corresponding to the point cloud block, the file encapsulation device may send the component track corresponding to the quality level to the file decapsulation device according to the quality level.

[0284] S610: The file decapsulation device decapsulates and decodes the media file of the immersive media corresponding to the recommended window to obtain the content of the immersive media corresponding to the recommended window.

[0285] Specifically, after obtaining the immersive media file corresponding to the recommended window according to the above steps, the file decapsulation device decapsulates the media file corresponding to the recommended window to obtain the immersive media stream corresponding to the recommended window. The immersive media stream corresponding to the recommended window is then decoded to obtain the immersive media content corresponding to the recommended window. The specific decapsulation and decoding methods can be referenced in existing technologies and will not be further described here.

[0286] Furthermore, the media file packaging method provided in the embodiment of the present application is described below through specific examples.

[0287] In Example 1, if the immersive media is a multi-view video, the encapsulation process specifically includes the following steps:

[0288] Step 11: The file packaging device determines a recommended viewing window for the multi-view video based on the content of the multi-view video;

[0289] Step 12: If the atlas information track of the multi-view video is divided according to viewpoint groups, the file encapsulation device associates the recommended viewport of the multi-view video with the corresponding viewpoint group information to generate a media file F1.

[0290] Step 13: The file encapsulation device generates a recommended view window metadata track, wherein the recommended view window metadata includes viewpoint group information corresponding to the recommended window;

[0291] Step 14: The file encapsulation device sends first indication information to the file decapsulation device. The first indication information may be DASH signaling, and the first indication information is used to indicate metadata of the recommended view window.

[0292] Step 15: The file decapsulation device sends a first request message to the file encapsulation device according to the first instruction message, where the first request message is used to request metadata of the recommended window.

[0293] Step 16: The file encapsulation device sends the metadata track of the recommended window to the file decapsulation device;

[0294] Step 17: The file decapsulation device decodes the metadata track of the recommended view window to obtain the viewpoint group information corresponding to the recommended view window included in the metadata of the recommended view window;

[0295] Step 18: The file decapsulation device requests and consumes the media resources corresponding to the recommended window based on its own network conditions and decoding capabilities and the viewpoint group information corresponding to the recommended window.

[0296] For example, assuming the client presents according to the recommended viewport 1, and the viewpoint group associated with viewport 1 is view_group 1, the file decapsulation device sends a second request message to the file encapsulation device, which includes the identification information of view_group 1. The file encapsulation device finds the corresponding atlas track of tile 0 through view_group 1 and sends the atlas track of tile 0 to the file decapsulation device. The file decapsulation device directly decodes the component track associated with atlas track tile 0 for consumption.

[0297] As can be seen from the above, the embodiment of the present application associates the recommended window with the perspective information of the multi-perspective video media corresponding to the recommended window, so that the file decapsulation device directly requests the corresponding media resources, saving bandwidth and decoding resources.

[0298] Example 2: If the immersive media is point cloud media, the encapsulation process specifically includes the following steps:

[0299] Step 21: The file packaging device determines a recommended viewport for the point cloud media based on the content of the point cloud media.

[0300] Step 22: If the compression method of the point cloud media is VPCC, and the point cloud media is organized according to point cloud tiles, and the point cloud tiles have different quality levels, then the recommended window is associated with the replaceable group selection information of the corresponding point cloud tile to generate a media file F2.

[0301] The replaceable group selection information of the point cloud block includes at least one of identification information of the component track corresponding to the point cloud block and a quality level corresponding to the point cloud block.

[0302] Step 23: The file encapsulation device generates a recommended view window metadata track, wherein the recommended view window metadata includes replaceable group selection information of the point cloud blocks corresponding to the recommended view window;

[0303] Step 24: The file encapsulation device sends first indication information to the file decapsulation device. The first indication information may be DASH signaling, and the first indication information is used to indicate metadata of the recommended view window.

[0304] Step 25: The file decapsulation device sends a first request message to the file encapsulation device according to the first instruction message, where the first request message is used to request metadata of the recommended window.

[0305] Step 26: The file encapsulation device sends the metadata track of the recommended window to the file decapsulation device;

[0306] Step 27: The file decapsulation device decodes the metadata track of the recommended viewport to obtain the replaceable group selection information of the point cloud blocks corresponding to the recommended viewport included in the metadata of the recommended viewport;

[0307] Step 28: The file decapsulation device requests and consumes the media resources corresponding to the recommended window based on its own network conditions and decoding capabilities and the replaceable group selection information of the point cloud blocks corresponding to the recommended window.

[0308] For example, assuming that the client presents according to the recommended viewport1, and viewport1 is associated with the alternative group selection information of the point cloud block, the file decapsulation device sends a second request message to the file encapsulation device, and the second request message includes the alternative group selection information (AlternativesSelectInfoStruct) of the point cloud block. The file encapsulation device can find all the alternative groups through the alternate_group_id in the AlternativesSelectInfoStruct, and then select the corresponding component track from each alternative group according to the AlternativeInfoStruct or track_id, and send the selected component track to the file decapsulation device for decoding and consumption.

[0309] For example Figure 7 As shown, tile0 corresponds to three replaceable groups, and each replacement group includes two component tracks. For example, component track 1 and component track 1' form a replacement group. Optionally, component track 1 is Occ.Track, and component track 1' is Occ.Track'. Component track 2 and component track 2' form a replacement group. Optionally, component track 2 is Geo.Track, and component track 2' is Geo.Track'. Component track 3 and component track 3' form a replacement group. Optionally, component track 3 is Att.Track, and component track 3' is Att.Track'. Similarly, tile 1 has three replacement groups, each containing two component tracks. For example, component track 11 and component track 11' form a replacement group. Optionally, component track 11 is Occ.Track and component track 11' is Occ.Track'. Component track 12 and component track 12' form a replacement group. Optionally, component track 12 is Geo.Track and component track 12' is Geo.Track'. Component track 13 and component track 13' form a replacement group. Optionally, component track 13 is Att.Track and component track 13' is Att.Track'. The tracks in one replacement group have different quality levels.

[0310] If the recommended view corresponds to point cloud block 0 and point cloud block 1, where point cloud block 0 corresponds to tile 0 and point cloud block 1 corresponds to tile 1, and the quality level corresponding to point cloud block 0 is 0 and the quality level corresponding to point cloud block 1 is 1, then the file decapsulation device carries quality level 0 and quality level 1 in the second request information. The file packaging device queries tile 0 and tile 1 based on quality level 0 and quality level 1, and sends the component tracks of the three replaceable groups corresponding to tile 0 and the component tracks of the three replaceable groups corresponding to tile 1 to the file decapsulation device. Based on the position information of the recommended view, the file decapsulation device may select component tracks with better quality from the three replaceable groups corresponding to tile 0, such as Occ.Track, Geo.Track, and Att.Track, for decoding, but select component tracks with worse quality from the three replaceable groups corresponding to tile 1, such as Occ.Track', Geo.Track', and Att.Track', for decoding.

[0311] As can be seen from the above, the embodiment of the present application associates the recommended window with the replaceable group selection information of the point cloud block corresponding to the recommended window, so that the file decapsulation device directly requests the corresponding media resources, saving bandwidth and decoding resources.

[0312] In some embodiments, if the recommended window is associated with the perspective information of the multi-perspective video media corresponding to the recommended window, the file encapsulation device also adds a first flag to the metadata of the recommended window, and the first flag is used to indicate that the recommended window is associated with the perspective information of the multi-perspective video media corresponding to the recommended window.

[0313] At this time, the file decapsulation device further includes: determining whether the metadata of the recommended view includes the first flag before sending the second request information to the file encapsulation device based on the view information of the multi-view video media corresponding to the recommended view;

[0314] Correspondingly, the above S608-A includes: when the file decapsulation device determines that the metadata of the recommended window includes the first flag, the file decapsulation device sends second request information to the file encapsulation device according to the perspective information of the multi-perspective video media corresponding to the recommended window.

[0315] That is, in this embodiment, if it is determined that the metadata of the recommended window includes the first flag, it indicates that the metadata of the recommended window includes the view information of the multi-view video media corresponding to the recommended window, and the view information of the multi-view video media corresponding to the recommended window is obtained, and then S608-A is executed. If it is determined that the metadata of the recommended window does not include the first flag, it indicates that the metadata of the recommended window does not include the view information of the multi-view video media corresponding to the recommended window, and S608-A is not executed, thereby avoiding unnecessary data processing and saving decoding resources.

[0316] In some embodiments, if the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window, the file encapsulation device also adds a second flag in the metadata of the recommended window, and the second flag is used to indicate that the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window.

[0317] At this time, the above-mentioned file decapsulation device also includes, before sending the second request information to the file encapsulation device based on the replaceable group selection information of the point cloud block corresponding to the recommended window: determining whether the metadata of the recommended window includes a second flag, and the second flag is used to indicate that the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window.

[0318] Correspondingly, the above S608-B includes: when the file decapsulation device determines that the metadata of the recommended window includes the second flag, it sends a second request information to the file encapsulation device according to the replaceable group selection information of the point cloud block corresponding to the recommended window.

[0319] That is, in this embodiment, if it is determined that the metadata of the recommended window includes the second flag, it indicates that the metadata of the recommended window includes the alternative group selection information for the point cloud segment corresponding to the recommended window, and then the alternative group selection information for the point cloud segment corresponding to the recommended window is obtained, and then S608-B is executed. If it is determined that the metadata of the recommended window does not include the second flag, it indicates that the metadata of the recommended window does not include the alternative group selection information for the point cloud segment corresponding to the recommended window, and S608-A is not executed, thereby saving decoding resources.

[0320] In a possible implementation, when the first flag or the second flag is added to the metadata of the recommended window, a sample format of the metadata of the recommended window is as follows:

[0321]

[0322]

[0323] If the viewport info metadata track exists, the camera extrinsic information ExtCameraInfoStruct() should be present in the sample entry or in the sample. The following conditions must not occur: dynamic_ext_camera_flag is 0 and camera_extrinsic_flag[i] is 0 in all samples.

[0324] num_viewports: Indicates the number of viewports indicated in the sample.

[0325] viewport_id[i]: indicates the identifier of the corresponding viewport.

[0326] viewport_cancel_flag[i]: A value of 1 indicates that the window with the viewport identifier value viewport_id[i] is canceled.

[0327] camera_intrinsic_flag[i]: A value of 1 indicates that camera intrinsics exist for the i-th window in the current sample. If dynamic_int_camera_flag is 0, this field must be 0. Also, when camera_extrinsic_flag[i] is 0, this field must be 0.

[0328] camera_extrinsic_flag[i]: A value of 1 indicates that camera extrinsic parameters exist for the i-th window in the current sample. If dynamic_ext_camera_flag is 0, this field must be 0.

[0329] view_id_flag[i]: a value of 1 indicates that the i-th view in the current sample is associated with the corresponding view information, for example, the recommended view is associated with the view information of the multi-view video media corresponding to the recommended view.

[0330] alter_info_flag[i]: A value of 1 indicates that the i-th window in the current sample is associated with the corresponding alternative group selection information, for example, the recommended window is associated with the alternative group selection information of the point cloud block corresponding to the recommended window.

[0331] In this embodiment, the i-th window may be a recommended window.

[0332] Figure 8 An interactive flow chart of a media file encapsulation and decapsulation method provided in an embodiment of the present application, such as Figure 8 As shown, the method includes the following steps:

[0333] S701: The file packaging device obtains content of the immersive media and determines a recommended window of the immersive media according to the content of the immersive media.

[0334] S702: The file packaging device determines characteristic information of the immersive media corresponding to the recommended window.

[0335] S703: The file packaging device associates the recommended window with characteristic information of the immersive media corresponding to the recommended window to generate a media file of the immersive media.

[0336] S704: The file encapsulation device sends first indication information to the file decapsulation device. The first indication information is used to indicate metadata of the recommended window. The metadata of the recommended window includes feature information of the immersive media corresponding to the recommended window.

[0337] The above steps S701 to S704 are the same as the above steps S501 to S504. Please refer to the description of the above steps S501 to S504 and will not be repeated here.

[0338] S705: The file decapsulation device sends first request information to the file encapsulation device in response to the first instruction information, where the first request information is used to request metadata of the recommended window.

[0339] S706: The file encapsulation device sends the metadata track of the recommended window to the file decapsulation device according to the first request information.

[0340] S707: The file decapsulation device decapsulates and decodes the metadata track of the recommended window to obtain metadata of the recommended window.

[0341] The above steps S705 to S707 are the same as the above steps S605 to S607. Please refer to the description of the above steps S605 to S607 and will not be repeated here.

[0342] S708: The file decapsulation device sends third request information to the file encapsulation device.

[0343] The third request information is used to request the media file of the entire immersive media.

[0344] S709: The file encapsulation device sends the media file of the immersive media to the file decapsulation device according to the third request information.

[0345] In this embodiment, the file decapsulation device requests the entire immersive media media file, and then decodes part of the media file according to actual needs.

[0346] S710: The file decapsulation device decapsulates and decodes the media file of the immersive media corresponding to the recommended window according to the characteristic information of the immersive media corresponding to the recommended window, to obtain the content of the immersive media corresponding to the recommended window.

[0347] In some embodiments, if the immersive media is multi-view video media, and the video tracks of the immersive media are divided according to viewpoints or viewpoint groups, and the recommended window is associated with the viewpoint information of the multi-view video media corresponding to the recommended window, that is, the metadata of the recommended window includes the viewpoint information of the multi-view video media corresponding to the recommended window, then the above S710 includes S710-A1 and S710-A2:

[0348] S710-A1. The file decapsulation device searches for a media file corresponding to the perspective information in the received media file of the immersive media according to the perspective information of the multi-perspective video media corresponding to the recommended window.

[0349] S710-A2: The file decapsulation device decapsulates and decodes the media file corresponding to the retrieved viewing angle information to obtain immersive media content corresponding to the recommended viewport.

[0350] For example, when the perspective information is viewpoint information, the file decapsulation device queries the media file corresponding to the viewpoint information from the received media file of the immersive media, decapsulates and decodes the media file corresponding to the viewpoint information, and obtains the immersive media content corresponding to the recommended window.

[0351] For example, when the perspective information is viewpoint group information, the file decapsulation device queries the media file corresponding to the viewpoint group information from the media file of the received immersive media, decapsulates and decodes the media file corresponding to the viewpoint group information, and obtains the immersive media content corresponding to the recommended window.

[0352] For example, when the viewing angle information is camera information, the file decapsulation device queries the media file corresponding to the camera information from the received media file of the immersive media, decapsulates and decodes the media file corresponding to the camera information, and obtains the content of the immersive media corresponding to the recommended window.

[0353] In some embodiments, before the above-mentioned S710-A1, that is, before searching for the media file corresponding to the perspective information in the received media files of the immersive media based on the perspective information of the multi-perspective video media corresponding to the recommended window, the method of this embodiment also includes: determining whether the metadata of the recommended window includes a first flag, and the first flag is used to indicate that the recommended window is associated with the perspective information of the multi-perspective video media corresponding to the recommended window.

[0354] When it is determined that the metadata of the recommended viewport includes the first flag, the above S710-A1 is executed, that is, according to the viewport information of the multi-viewport video media corresponding to the recommended viewport, the media file corresponding to the viewport information is searched in the received media files of the immersive media.

[0355] That is, in this embodiment, if it is determined that the metadata of the recommended window includes the first flag, it indicates that the metadata of the recommended window includes the view information of the multi-view video media corresponding to the recommended window, and then the view information of the multi-view video media corresponding to the recommended window is obtained, and then S710-A1 is executed. If it is determined that the metadata of the recommended window does not include the first flag, it indicates that the metadata of the recommended window does not include the view information of the multi-view video media corresponding to the recommended window, and S710-A1 is not executed, thereby avoiding unnecessary data processing and saving decoding resources.

[0356] In some embodiments, if the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window, that is, the metadata of the recommended window includes the replaceable group selection information of the point cloud block corresponding to the recommended window, then the above S710 includes S710-B1 and S710-B2:

[0357] S710-B1. The file decapsulation device searches for a media file corresponding to the replaceable group selection information in the received media files of the immersive media according to the replaceable group selection information of the point cloud block corresponding to the recommended viewport.

[0358] S710-B2. The file decapsulation device decapsulates and decodes the media file corresponding to the queried replaceable group selection information to obtain immersive media content corresponding to the recommended window.

[0359] It can be seen from the above embodiments that the replaceable group selection information of the point cloud block includes at least one of the identification information of the component track corresponding to the point cloud block and the quality level corresponding to the point cloud block.

[0360] In one example, if the replaceable group selection information of the point cloud block includes identification information of the component track corresponding to the point cloud block, the file decapsulation device queries the media file corresponding to the component track from the media file of the received immersive media, and decapsulates and decodes the media file corresponding to the component track to obtain the content of the immersive media corresponding to the recommended window.

[0361] In one example, if the replaceable group selection information of the point cloud block includes the quality level corresponding to the point cloud block, the file decapsulation device queries the media file corresponding to the quality level from the received media file of the immersive media, decapsulates the media file corresponding to the quality level and then decodes it to obtain the content of the immersive media corresponding to the recommended window.

[0362] In some embodiments, based on the replaceable group selection information of the point cloud block corresponding to the recommended window, before searching for the media file corresponding to the replaceable group selection information in the received media file of the immersive media, the method also includes: determining whether the metadata of the recommended window includes a second flag, wherein the second flag is used to indicate that the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window.

[0363] When it is determined that the metadata of the recommended window includes the second flag, the above S710-B1 is executed, that is, according to the replaceable group selection information of the point cloud block corresponding to the recommended window, the media file corresponding to the replaceable group selection information is searched in the received media files of the immersive media.

[0364] That is, in this embodiment, if it is determined that the metadata of the recommended window includes the second flag, it indicates that the metadata of the recommended window includes the alternative group selection information of the point cloud segment corresponding to the recommended window, and then the alternative group selection information of the point cloud segment corresponding to the recommended window is obtained, and then S710-B1 is executed. If it is determined that the metadata of the recommended window does not include the second flag, it indicates that the metadata of the recommended window does not include the alternative group selection information of the point cloud segment corresponding to the recommended window, and S710-B1 is not executed, thereby saving decoding resources.

[0365] Furthermore, the media file packaging method provided in the embodiment of the present application is described below through specific examples.

[0366] In Example 1, if the immersive media is a multi-view video, the encapsulation process specifically includes the following steps:

[0367] Step 31: The file packaging device determines a recommended viewing window for the multi-view video based on the content of the multi-view video;

[0368] Step 32: If the atlas information track of the multi-view video is divided according to viewpoint groups, the file encapsulation device associates the recommended viewport of the multi-view video with the corresponding viewpoint group information to generate a media file F1.

[0369] Step 33: The file encapsulation device generates a recommended view window metadata track, wherein the recommended view window metadata includes viewpoint group information corresponding to the recommended window;

[0370] Step 34: The file encapsulation device sends first indication information to the file decapsulation device. The first indication information may be DASH signaling, and the first indication information is used to indicate metadata of the recommended view window.

[0371] Step 35: The file decapsulation device sends a first request message to the file encapsulation device according to the first instruction message, where the first request message is used to request metadata of the recommended window.

[0372] Step 36: The file encapsulation device sends the metadata track of the recommended window to the file decapsulation device;

[0373] Step 37: The file decapsulation device decodes the metadata track of the recommended view window to obtain the viewpoint group information corresponding to the recommended view window included in the metadata of the recommended view window;

[0374] Step 38: The file decapsulation device decodes and consumes the media resources corresponding to the recommended viewport based on its own network conditions and decoding capabilities and in combination with the viewpoint group information corresponding to the recommended viewport.

[0375] For example, assuming that the recommended viewport is viewport1, and the viewpoint group associated with viewport1 is view_group1, the file decapsulation device searches for the atlas track corresponding to view_group1 as tile0 in the media file of the entire immersive media requested, and directly decodes the component track associated with atlas track tile0 for consumption.

[0376] As can be seen from the above, the embodiment of the present application associates the recommended window with the perspective information of the multi-perspective video media corresponding to the recommended window, so that the file decapsulation device directly decodes the corresponding media resources for consumption, saving bandwidth and decoding resources.

[0377] Example 2: If the immersive media is point cloud media, the encapsulation process specifically includes the following steps:

[0378] Step 41: The file packaging device determines a recommended viewport for the point cloud media based on the content of the point cloud media.

[0379] Step 42: If the compression method of the point cloud media is VPCC, and the point cloud media is organized according to point cloud tiles, and the point cloud tiles have different quality levels, then the recommended window is associated with the replaceable group selection information of the corresponding point cloud tile to generate a media file F2.

[0380] The replaceable group selection information of the point cloud block includes at least one of identification information of the component track corresponding to the point cloud block and a quality level corresponding to the point cloud block.

[0381] Step 43: The file encapsulation device generates a recommended view window metadata track, wherein the recommended view window metadata includes replaceable group selection information of the point cloud blocks corresponding to the recommended view window;

[0382] Step 44: The file encapsulation device sends first indication information to the file decapsulation device. The first indication information may be DASH signaling, and the first indication information is used to indicate metadata of the recommended view window.

[0383] Step 45: The file decapsulation device sends a first request message to the file encapsulation device according to the first instruction message, where the first request message is used to request metadata of the recommended window.

[0384] Step 46: The file encapsulation device sends the metadata track of the recommended window to the file decapsulation device;

[0385] Step 47: The file decapsulation device decodes the metadata track of the recommended view to obtain the replaceable group selection information of the point cloud blocks corresponding to the recommended view included in the metadata of the recommended view.

[0386] Step 48: The file decapsulation device decodes and consumes the media resources corresponding to the recommended window based on its own network conditions and decoding capabilities and in combination with the replaceable group selection information of the point cloud block corresponding to the recommended window.

[0387] For example, the recommended viewport is viewport1, and viewport1 is associated with the alternative group selection information (AlternativesSelectInfoStruct) of the point cloud block. In this way, the file decapsulation device finds all alternative groups through the alternate_group_id in the AlternativesSelectInfoStruct in the media file of the entire immersive media requested, and then selects the corresponding component track from each alternative group according to the AlternativeInfoStruct or track_id for decoding and consumption.

[0388] For example Figure 7 As shown, if the recommended window corresponds to point cloud block 0 and point cloud block 1, where point cloud block 0 corresponds to tile0 and point cloud block 1 corresponds to tile1, and the quality level corresponding to point cloud block 0 is 0 and the quality level corresponding to point cloud block 1 is 1, then the file decapsulation device may select a component track with better quality from the three replaceable groups corresponding to tile0, such as Occ.Track, Geo.Track and Att.Track, for decoding based on the position information of the recommended window, but select a component track with worse quality from the three replaceable groups corresponding to tile1, such as Occ.Track', Geo.Track' and Att.Track' for decoding.

[0389] As can be seen from the above, the embodiment of the present application associates the recommendation window with the replaceable group selection information of the point cloud block corresponding to the recommendation window, so that the file decapsulation device directly decodes the corresponding media resources for consumption, saving bandwidth and decoding resources.

[0390] It should be understood that Figures 5 to 8 This is only an example of the present application and should not be considered as limiting the present application.

[0391] The preferred embodiments of the present application are described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the technical concept of the present application, a variety of simple modifications can be made to the technical solution of the present application, and these simple modifications all fall within the scope of protection of the present application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. In order to avoid unnecessary repetition, the present application will not further explain various possible combinations. For another example, the various different embodiments of the present application can also be arbitrarily combined, and as long as they do not violate the ideas of the present application, they should also be regarded as the contents disclosed in the present application.

[0392] Combined with the above Figure 5 and Figure 8 , describes the method embodiment of the present application in detail, and the following is combined with Figures 9 to 11 , describe in detail the device embodiments of the present application.

[0393] Figure 9 This is a schematic structural diagram of a media file encapsulation device provided in one embodiment of the present application. The device 10 is applied to a file encapsulation device and includes:

[0394] an acquisition unit 11, configured to acquire content of the immersive media and determine a recommended window of the immersive media according to the content of the immersive media;

[0395] A processing unit 12 is configured to determine characteristic information of the immersive media corresponding to the recommended viewport;

[0396] an encapsulation unit 13, configured to associate the recommendation window with characteristic information of the immersive media corresponding to the recommendation window, and generate a media file of the immersive media;

[0397] The transceiver unit 14 is configured to send first indication information to the file decapsulation device, where the first indication information is used to indicate metadata of the recommended viewport, where the metadata of the recommended viewport includes feature information of the immersive media corresponding to the recommended viewport.

[0398] In some embodiments, the immersive media includes at least one of multi-perspective video media and point cloud media.

[0399] In some embodiments, the transceiver unit 14 is further used to receive a first request message sent by the file decapsulation device, the first request message being used to request metadata of the recommended window; receive a second request message sent by the file decapsulation device, the second request message being used to request a media file of the immersive media corresponding to the recommended window; and send the media file of the immersive media corresponding to the recommended window to the file decapsulation device according to the second request message.

[0400] In some embodiments, the transceiver unit 21 is used to receive a first request message sent by the file decapsulation device, the first request message being used to request metadata of the recommended window; based on the first request message, the metadata track of the recommended window is sent to the file decapsulation device; receive a third request message sent by the file decapsulation device, the third request message being used to request media files of the immersive media; based on the third request message, the media files of the immersive media are sent to the file decapsulation device.

[0401] In some embodiments, if the immersive media is multi-view video media, and the video tracks of the immersive media are divided according to viewpoints or viewpoint groups, the processing unit 12 is specifically configured to determine the viewpoint information of the multi-view video media corresponding to the recommended viewport as the characteristic information of the immersive media corresponding to the recommended viewport;

[0402] The encapsulation unit 13 is specifically configured to associate the recommended viewport with the viewport information of the multi-viewport video media corresponding to the recommended viewport.

[0403] In some embodiments, the perspective information of the multi-perspective video media corresponding to the recommended viewport includes at least one of viewpoint group information, viewpoint information, and camera information of the multi-perspective video media corresponding to the recommended viewport.

[0404] In some embodiments, if the perspective information of the multi-perspective video media corresponding to the recommended window is viewpoint group information, the viewpoint group information includes: the number of viewpoint groups associated with the recommended window and the identifier of the viewpoint group associated with the recommended window.

[0405] In some embodiments, if the perspective information of the multi-perspective video media corresponding to the recommended window is viewpoint information, the viewpoint information includes: the number of viewpoints associated with the recommended window and identifiers of the viewpoints associated with the recommended window.

[0406] In some embodiments, if the perspective information of the multi-perspective video media corresponding to the recommended window is camera information, the camera information includes: the number of cameras associated with the recommended window and the identifier of the camera associated with the recommended window.

[0407] In some embodiments, the transceiver unit 14 is configured to send the media file corresponding to the perspective information to the file decapsulation device if the second request information includes perspective information of the multi-perspective video media corresponding to the recommended window.

[0408] In some embodiments, if the recommended window is associated with the perspective information of the multi-perspective video media corresponding to the recommended window, the encapsulation unit 13 is also used to add a first flag in the metadata of the recommended window, and the first flag is used to indicate that the recommended window is associated with the perspective information of the multi-perspective video media corresponding to the recommended window.

[0409] In some embodiments, if the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, the processing unit 12 is specifically configured to determine the replaceable group selection information of the point cloud block corresponding to the recommended window as the characteristic information of the immersive media corresponding to the recommended window, where the replaceable group selection information of the point cloud block includes at least one of identification information of the component track corresponding to the point cloud block and the quality level corresponding to the point cloud block;

[0410] The encapsulation unit 13 is specifically configured to associate the recommended view window with the replaceable group selection information of the point cloud block corresponding to the recommended view window.

[0411] In some embodiments, the replaceable group selection information of the point cloud block corresponding to the recommended window also includes: at least one of the number of replaceable groups corresponding to the point cloud block, the identifier of the replaceable group, and the number of component tracks selected from the replaceable group.

[0412] In some embodiments, if the second request information includes the replaceable group selection information of the point cloud block, the above-mentioned transceiver unit 14 is specifically used to send the component track corresponding to the point cloud block to the file decapsulation device if the replaceable group selection information of the point cloud block includes the identification information of the component track corresponding to the point cloud block; or

[0413] If the replaceable group selection information of the point cloud block includes a quality level corresponding to the point cloud block, the component track corresponding to the quality level is sent to the file decapsulation device.

[0414] In some embodiments, if the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window, the above-mentioned encapsulation unit 13 is specifically used to add a second flag in the metadata of the recommended window, and the second flag is used to indicate that the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window.

[0415] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Figure 9The device 10 shown can execute the method embodiment corresponding to the file encapsulation device, and the aforementioned and other operations and / or functions of each module in the device 10 are respectively for implementing the method embodiment corresponding to the file encapsulation device. For the sake of brevity, they are not repeated here.

[0416] Figure 10 This is a schematic diagram of the structure of a media file decapsulation device provided in an embodiment of the present application. The device 20 is applied to a file decapsulation device, and the device 20 includes:

[0417] a transceiver unit 21 configured to receive first indication information sent by a file encapsulation device, wherein the first indication information is used to indicate metadata of a recommended viewport, wherein the metadata of the recommended viewport includes characteristic information of an immersive media corresponding to the recommended viewport, wherein the recommended viewport is determined based on content of the immersive media;

[0418] The processing unit 22 is configured to determine whether to request metadata of the recommended window in response to the first indication information.

[0419] In some embodiments, the immersive media includes at least one of multi-perspective video media and point cloud media.

[0420] In some embodiments, the transceiver unit 21 is used to send a first request message to the file encapsulation device if it is determined to request the metadata of the recommended window, the first request message being used to request the metadata of the recommended window; receive the metadata track of the recommended window sent by the file encapsulation device, decapsulate and then decode the metadata track of the recommended window to obtain the metadata of the recommended window; send a second request message to the file encapsulation device based on the characteristic information of the immersive media corresponding to the recommended window in the metadata of the recommended window, the second request message being used to request the media file of the immersive media corresponding to the recommended window; receive the media file of the immersive media corresponding to the recommended window sent by the file encapsulation device, decapsulate and then decode the media file of the immersive media corresponding to the recommended window to obtain the content of the immersive media corresponding to the recommended window.

[0421] In some embodiments, if the immersive media is multi-perspective video media, and the video track of the immersive media is divided according to viewpoints or viewpoint groups, the recommended window is associated with the perspective information of the multi-perspective video media corresponding to the recommended window, then the transceiver unit 21 is also used to send a second request information to the file encapsulation device according to the perspective information of the multi-perspective video media corresponding to the recommended window, and the second request information includes the perspective information of the multi-perspective video media corresponding to the recommended window.

[0422] In some embodiments, the processing unit 22 is further configured to determine whether the metadata of the recommended viewport includes a first flag, the first flag being configured to indicate that the recommended viewport is associated with perspective information of the multi-perspective video media corresponding to the recommended viewport;

[0423] The transceiver unit 21 is configured to send second request information to the file encapsulation device according to the perspective information of the multi-perspective video media corresponding to the recommended viewport when the processing unit 22 determines that the metadata of the recommended viewport includes the first flag.

[0424] In some embodiments, if the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window, and the replaceable group selection information of the point cloud block includes the identification information of the component track corresponding to the point cloud block and at least one of the quality levels corresponding to the point cloud block, then the transceiver unit 21 is specifically used to send a second request information to the file encapsulation device according to the replaceable group selection information of the point cloud block corresponding to the recommended window, and the second request information includes the replaceable group selection information of the point cloud block.

[0425] In some embodiments, the processing unit 22 is further configured to determine whether the metadata of the recommended viewport includes a second flag, wherein the second flag is configured to indicate that the recommended viewport is associated with the replaceable group selection information of the point cloud segment corresponding to the recommended viewport;

[0426] The transceiver unit 21 is specifically configured to send a second request message to the file encapsulation device according to the replaceable group selection information of the point cloud block corresponding to the recommended window when the processing unit 22 determines that the metadata of the recommended window includes a second flag.

[0427] In some embodiments, if it is determined to request metadata of the recommended view, the transceiver unit 21 is configured to send a first request message to the file encapsulation device, the first request message being used to request metadata of the recommended view; receive the metadata track of the recommended view sent by the file encapsulation device, decapsulate and decode the metadata track of the recommended view to obtain metadata of the recommended view; send a third request message to the file encapsulation device, the third request message being used to request a media file of the immersive media; and receive the media file of the immersive media sent by the file encapsulation device.

[0428] The processing unit 22 is further configured to decapsulate and decode the media file of the immersive media corresponding to the recommended window according to the characteristic information of the immersive media corresponding to the recommended window, so as to obtain the content of the immersive media corresponding to the recommended window.

[0429] In some embodiments, if the immersive media is multi-perspective video media, and the video track of the immersive media is divided according to viewpoints or viewpoint groups, and the recommended window is associated with the perspective information of the multi-perspective video media corresponding to the recommended window, then the transceiver unit 21 is used to query the media files of the received immersive media for the media files corresponding to the perspective information of the multi-perspective video media corresponding to the recommended window; the processing unit 22 is specifically used to decapsulate and decode the media files corresponding to the queried perspective information to obtain the content of the immersive media corresponding to the recommended window.

[0430] In some embodiments, the processing unit 22 is further used to determine whether the metadata of the recommended window includes a first flag, where the first flag is used to indicate that the recommended window is associated with the perspective information of the multi-perspective video media corresponding to the recommended window; when it is determined that the metadata of the recommended window includes the first flag, the media file corresponding to the perspective information is queried in the received media files of the immersive media according to the perspective information of the multi-perspective video media corresponding to the recommended window.

[0431] In some embodiments, if the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window, and the replaceable group selection information of the point cloud block includes the identification information of the component track corresponding to the point cloud block and at least one of the quality levels corresponding to the point cloud block, then the processing unit 22 is used to query the received media files of the immersive media for the media file corresponding to the replaceable group selection information according to the replaceable group selection information of the point cloud block corresponding to the recommended window; decapsulate and decode the media file corresponding to the queried replaceable group selection information to obtain the content of the immersive media corresponding to the recommended window.

[0432] In some embodiments, the processing unit 22 is used to determine whether the metadata of the recommended window includes a second flag, and the second flag is used to indicate that the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window; when it is determined that the metadata of the recommended window includes the second flag, according to the replaceable group selection information of the point cloud block corresponding to the recommended window, the media file corresponding to the replaceable group selection information is queried in the received media files of the immersive media.

[0433] In some embodiments, the perspective information of the multi-perspective video media corresponding to the recommended viewport includes at least one of viewpoint group information, viewpoint information, and camera information of the multi-perspective video media corresponding to the recommended viewport.

[0434] In some embodiments, if the perspective information of the multi-perspective video media corresponding to the recommended window is viewpoint group information, the viewpoint group information includes: the number of viewpoint groups associated with the recommended window and the identifier of the viewpoint group associated with the recommended window.

[0435] In some embodiments, if the perspective information of the multi-perspective video media corresponding to the recommended window is viewpoint information, the viewpoint information includes: the number of viewpoints associated with the recommended window and identifiers of the viewpoints associated with the recommended window.

[0436] In some embodiments, if the perspective information of the multi-perspective video media corresponding to the recommended window is camera information, the camera information includes: the number of cameras associated with the recommended window and the identifier of the camera associated with the recommended window.

[0437] In some embodiments, the replaceable group selection information of the point cloud block corresponding to the recommended window also includes: at least one of the number of replaceable groups corresponding to the point cloud block, the identifier of the replaceable group, and the number of component tracks selected from the replaceable group.

[0438] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Figure 10 The device 20 shown can execute the method embodiment corresponding to the server, and the aforementioned and other operations and / or functions of each module in the device 20 are respectively for implementing the method embodiment corresponding to the file decapsulation device. For the sake of brevity, they are not repeated here.

[0439] The apparatus of the embodiment of the present application is described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.

[0440] Figure 11 It is a schematic block diagram of a computing device provided in an embodiment of the present application. The computing device can be the above-mentioned file encapsulation device or file decapsulation device, or the computing device has the functions of a file encapsulation device and a file decapsulation device.

[0441] like Figure 11 As shown, the computing device 40 may include:

[0442] The memory 41 and the memory 42 are configured to store computer programs and transfer the program code to the memory 42. In other words, the memory 42 can call and run the computer program from the memory 41 to implement the method in the embodiment of the present application.

[0443] For example, the memory 42 may be used to execute the above method embodiments according to the instructions in the computer program.

[0444] In some embodiments of the present application, the memory 42 may include but is not limited to:

[0445] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0446] In some embodiments of the present application, the memory 41 includes but is not limited to:

[0447] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0448] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 41 and executed by the memory 42 to implement the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the video production device.

[0449] like Figure 11 As shown, the computing device 40 may further include:

[0450] The transceiver 40 and the transceiver 43 may be connected to the memory 42 or the memory 41 .

[0451] The memory 42 can control the transceiver 43 to communicate with other devices. Specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 43 may include a transmitter and a receiver. The transceiver 43 may further include an antenna, and the number of antennas may be one or more.

[0452] It should be understood that the various components in the video production device are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0453] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment. In other words, the present application also provides a computer program product containing instructions, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment.

[0454] When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0455] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0456] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0457] Modules described as separate components may or may not be physically separate, and components displayed as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected based on actual needs to achieve the purpose of the present embodiment. For example, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module.

[0458] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A media file encapsulation method, characterized in that: Applied to a file encapsulation device, the method includes: Acquiring content of immersive media, and determining a recommended viewport for the immersive media based on the content of the immersive media, the immersive media comprising at least one of multi-view video media and point cloud media; Determining characteristic information of the immersive media corresponding to the recommended window; Associating the recommendation window with characteristic information of the immersive media corresponding to the recommendation window to generate a media file of the immersive media; Sending first indication information to a file decapsulation device, where the first indication information is used to indicate metadata of the recommended viewport, where the metadata of the recommended viewport includes feature information of the immersive media corresponding to the recommended viewport; If the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, then determining the characteristic information of the immersive media corresponding to the recommended viewport includes: The replaceable group selection information of the point cloud block corresponding to the recommended window is determined as the characteristic information of the immersive media corresponding to the recommended window, and the replaceable group selection information of the point cloud block includes at least one of the identification information of the component track corresponding to the point cloud block and the quality level corresponding to the point cloud block.

2. The method according to claim 1, characterized in that If the immersive media is multi-view video media, and the video tracks of the immersive media are divided according to viewpoints or viewpoint groups, then determining the characteristic information of the immersive media corresponding to the recommended viewport includes: determining the perspective information of the multi-perspective video media corresponding to the recommended window as characteristic information of the immersive media corresponding to the recommended window; The associating the recommended window with characteristic information of the immersive media corresponding to the recommended window includes: The recommended viewport is associated with viewpoint information of the multi-viewpoint video media corresponding to the recommended viewport.

3. The method according to claim 2, characterized in that If the recommended viewport is associated with the viewport information of the multi-viewport video media corresponding to the recommended viewport, the method further includes: A first flag is added to metadata of the recommended viewport, where the first flag is used to indicate that the recommended viewport is associated with view information of the multi-view video media corresponding to the recommended viewport.

4. The method according to claim 2, characterized in that The perspective information of the multi-perspective video media corresponding to the recommended viewport includes at least one of viewpoint group information, viewpoint information, and camera information of the multi-perspective video media corresponding to the recommended viewport.

5. The method according to claim 4, characterized in that If the perspective information of the multi-perspective video media corresponding to the recommended window is viewpoint group information, the viewpoint group information includes: the number of viewpoint groups associated with the recommended window and the identifier of the viewpoint group associated with the recommended window.

6. The method according to claim 4, characterized in that If the perspective information of the multi-perspective video media corresponding to the recommended window is viewpoint information, the viewpoint information includes: the number of viewpoints associated with the recommended window and identifiers of the viewpoints associated with the recommended window.

7. The method according to claim 4, characterized in that If the perspective information of the multi-perspective video media corresponding to the recommended window is camera information, the camera information includes: the number of cameras associated with the recommended window and the identifiers of the cameras associated with the recommended window.

8. The method according to claim 1, characterized in that If the recommended viewport is associated with the alternative group selection information of the point cloud segment corresponding to the recommended viewport, the method further includes: A second flag is added to the metadata of the recommended view window, where the second flag is used to indicate that the recommended view window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended view window.

9. The method according to claim 8, characterized in that The alternative group selection information of the point cloud block corresponding to the recommended window further includes: at least one of the number of alternative groups corresponding to the point cloud block, the identifier of the alternative group, and the number of component tracks selected from the alternative group.

10. The method according to any one of claims 1 to 9, characterized in that The method further comprises: receiving first request information sent by the file decapsulation device, where the first request information is used to request metadata of the recommended window; sending the metadata track of the recommended window to the file decapsulation device according to the first request information; receiving second request information sent by the file decapsulation device, where the second request information is used to request a media file of the immersive media corresponding to the recommended window; According to the second request information, the media file of the immersive media corresponding to the recommended window is sent to the file decapsulation device.

11. The method according to claim 10, characterized in that If the second request information includes perspective information of the multi-perspective video media corresponding to the recommended viewport, sending the media file of the immersive media corresponding to the recommended viewport to the file decapsulation device according to the second request information includes: The media file corresponding to the viewing angle information is sent to the file decapsulation device.

12. The method according to claim 10, characterized in that If the second request information includes the replaceable group selection information of the point cloud segment, sending the media file of the immersive media corresponding to the recommended viewport to the file decapsulation device according to the second request information includes: If the replaceable group selection information of the point cloud block includes identification information of the component track corresponding to the point cloud block, the component track corresponding to the point cloud block is sent to the file decapsulation device; or If the replaceable group selection information of the point cloud block includes a quality level corresponding to the point cloud block, the component track corresponding to the quality level is sent to the file decapsulation device.

13. The method according to any one of claims 1 to 9, characterized in that The method further comprises: receiving first request information sent by the file decapsulation device, where the first request information is used to request metadata of the recommended window; sending the metadata track of the recommended window to the file decapsulation device according to the first request information; receiving third request information sent by the file decapsulation device, where the third request information is used to request the media file of the immersive media; The media file of the immersive media is sent to the file decapsulation device according to the third request information.

14. A method for decapsulating a media file, characterized in that: Applied to a file decapsulation device, the method includes: Receiving first indication information sent by a file encapsulation device, the first indication information is used to indicate metadata of a recommended viewport, the metadata of the recommended viewport including characteristic information of immersive media corresponding to the recommended viewport, the recommended viewport being determined based on content of the immersive media, the immersive media including at least one of multi-view video media and point cloud media, and if the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, then the characteristic information includes replaceable group selection information of the point cloud block corresponding to the recommended viewport, the replaceable group selection information including at least one of identification information of a component track corresponding to the point cloud block and a quality level corresponding to the point cloud block; In response to the first indication information, it is determined whether to request metadata of the recommended viewport.

15. The method according to claim 14, characterized in that The method further comprises: If it is determined to request the metadata of the recommended window, sending first request information to the file encapsulation device, where the first request information is used to request the metadata of the recommended window; receiving the metadata track of the recommended viewport sent by the file encapsulation device, decapsulating and decoding the metadata track of the recommended viewport to obtain metadata of the recommended viewport; sending, to the file encapsulation device, second request information for requesting a media file of the immersive media corresponding to the recommended window according to feature information of the immersive media corresponding to the recommended window in the metadata of the recommended window; The media file of the immersive media corresponding to the recommended window is received from the file encapsulation device, and the media file of the immersive media corresponding to the recommended window is decapsulated and then decoded to obtain the content of the immersive media corresponding to the recommended window.

16. The method according to claim 15, characterized in that If the immersive media is multi-view video media, and the video tracks of the immersive media are divided according to viewpoints or viewpoint groups, and the recommended viewport is associated with viewpoint information of the multi-view video media corresponding to the recommended viewport, then sending the second request information to the file encapsulation device based on feature information of the immersive media corresponding to the recommended viewport in metadata of the recommended viewport includes: According to the perspective information of the multi-perspective video media corresponding to the recommended window, second request information is sent to the file encapsulation device, where the second request information includes the perspective information of the multi-perspective video media corresponding to the recommended window.

17. The method according to claim 16, characterized in that Before sending the second request information to the file encapsulation device based on the perspective information of the multi-perspective video media corresponding to the recommended viewport, the method further includes: determining whether metadata of the recommended viewport includes a first flag, wherein the first flag is used to indicate that the recommended viewport is associated with perspective information of a multi-view video medium corresponding to the recommended viewport; The sending second request information to the file encapsulation device according to the perspective information of the multi-perspective video media corresponding to the recommended window includes: When it is determined that the metadata of the recommended viewport includes the first flag, second request information is sent to the file encapsulation device according to the view information of the multi-view video media corresponding to the recommended viewport.

18. The method according to claim 15, characterized in that If the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, the recommended view is associated with the replaceable group selection information of the point cloud block corresponding to the recommended view, and the replaceable group selection information of the point cloud block includes at least one of identification information of the component track corresponding to the point cloud block and the quality level corresponding to the point cloud block, then sending the second request information to the file encapsulation device based on the characteristic information of the immersive media corresponding to the recommended view in the metadata of the recommended view includes: According to the replaceable group selection information of the point cloud block corresponding to the recommended window, second request information is sent to the file encapsulation device, where the second request information includes the replaceable group selection information of the point cloud block.

19. The method according to claim 18, characterized in that Before sending the second request information to the file encapsulation device based on the replaceable group selection information of the point cloud blocks corresponding to the recommended viewport, the method further includes: determining whether metadata of the recommended viewport includes a second flag, wherein the second flag is used to indicate that the recommended viewport is associated with alternative group selection information of a point cloud segment corresponding to the recommended viewport; The sending second request information to the file encapsulation device according to the replaceable group selection information of the point cloud block corresponding to the recommended window includes: When it is determined that the metadata of the recommended viewport includes a second flag, second request information is sent to the file encapsulation device according to the replaceable group selection information of the point cloud block corresponding to the recommended viewport.

20. The method according to claim 14, wherein The method further comprises: If it is determined to request the metadata of the recommended window, sending first request information to the file encapsulation device, where the first request information is used to request the metadata of the recommended window; receiving the metadata track of the recommended viewport sent by the file encapsulation device, decapsulating and decoding the metadata track of the recommended viewport to obtain metadata of the recommended viewport; Sending third request information to the file encapsulation device, where the third request information is used to request the media file of the immersive media; receiving a media file of the immersive media sent by the file encapsulation device; According to the characteristic information of the immersive media corresponding to the recommended window, the media file of the immersive media corresponding to the recommended window is decapsulated and then decoded to obtain the content of the immersive media corresponding to the recommended window.

21. The method according to claim 20, characterized in that If the immersive media is multi-view video media, and the video tracks of the immersive media are divided according to viewpoints or viewpoint groups, and the recommended window is associated with the viewpoint information of the multi-view video media corresponding to the recommended window, then decapsulating and decoding the media file of the immersive media corresponding to the recommended window based on the characteristic information of the immersive media corresponding to the recommended window to obtain the content of the immersive media corresponding to the recommended window includes: According to the perspective information of the multi-perspective video media corresponding to the recommended window, searching the received media files of the immersive media for a media file corresponding to the perspective information; The media file corresponding to the retrieved viewing angle information is decapsulated and then decoded to obtain immersive media content corresponding to the recommended viewport.

22. The method according to claim 21, characterized in that Before searching, based on the perspective information of the multi-perspective video media corresponding to the recommended viewport, the received media files of the immersive media for a media file corresponding to the perspective information, the method further includes: determining whether metadata of the recommended viewport includes a first flag, wherein the first flag is used to indicate that the recommended viewport is associated with perspective information of a multi-view video medium corresponding to the recommended viewport; The step of searching, based on the perspective information of the multi-perspective video media corresponding to the recommended viewport, the received media files of the immersive media for a media file corresponding to the perspective information includes: When it is determined that the metadata of the recommended viewport includes the first flag, the received media files of the immersive media are searched for a media file corresponding to the perspective information according to the perspective information of the multi-perspective video media corresponding to the recommended viewport.

23. The method according to claim 20, characterized in that If the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, the recommended window is associated with the replaceable group selection information of the point cloud block corresponding to the recommended window, and the replaceable group selection information of the point cloud block includes at least one of identification information of the component track corresponding to the point cloud block and the quality level corresponding to the point cloud block, then, based on the characteristic information of the immersive media corresponding to the recommended window, the media file of the immersive media corresponding to the recommended window is decapsulated and then decoded to obtain the content of the immersive media corresponding to the recommended window, including: searching, according to the replaceable group selection information of the point cloud segment corresponding to the recommended viewport, for a media file corresponding to the replaceable group selection information in the received media files of the immersive media; The media file corresponding to the queried alternative group selection information is decapsulated and then decoded to obtain immersive media content corresponding to the recommended window.

24. The method according to claim 23, wherein Before searching the received media files of the immersive media for the media files corresponding to the alternative group selection information of the point cloud segment corresponding to the recommended viewport, the method further includes: determining whether metadata of the recommended viewport includes a second flag, wherein the second flag is used to indicate that the recommended viewport is associated with alternative group selection information of a point cloud segment corresponding to the recommended viewport; The step of searching, according to the replaceable group selection information of the point cloud segment corresponding to the recommended viewport, the received media files of the immersive media for a media file corresponding to the replaceable group selection information includes: When it is determined that the metadata of the recommended viewport includes the second flag, the received media files of the immersive media are searched for a media file corresponding to the replaceable group selection information according to the replaceable group selection information of the point cloud segment corresponding to the recommended viewport.

25. The method according to claim 16 or 21, characterized in that The perspective information of the multi-perspective video media corresponding to the recommended viewport includes at least one of viewpoint group information, viewpoint information, and camera information of the multi-perspective video media corresponding to the recommended viewport.

26. The method according to claim 25, characterized in that If the perspective information of the multi-perspective video media corresponding to the recommended window is viewpoint group information, the viewpoint group information includes: the number of viewpoint groups associated with the recommended window and the identifier of the viewpoint group associated with the recommended window.

27. The method according to claim 25, characterized in that If the perspective information of the multi-perspective video media corresponding to the recommended window is viewpoint information, the viewpoint information includes: the number of viewpoints associated with the recommended window and identifiers of the viewpoints associated with the recommended window.

28. The method according to claim 25, characterized in that If the perspective information of the multi-perspective video media corresponding to the recommended window is camera information, the camera information includes: the number of cameras associated with the recommended window and the identifiers of the cameras associated with the recommended window.

29. The method according to claim 18 or 23, characterized in that The alternative group selection information of the point cloud block corresponding to the recommended window further includes: at least one of the number of alternative groups corresponding to the point cloud block, the identifier of the alternative group, and the number of component tracks selected from the alternative group.

30. A media file packaging device, characterized in that: Applied to a file encapsulation device, the device comprises: an acquisition unit, configured to acquire content of immersive media and determine a recommended viewport for the immersive media based on the content of the immersive media, wherein the immersive media includes at least one of multi-view video media and point cloud media; a processing unit, configured to determine characteristic information of the immersive media corresponding to the recommended viewport; an encapsulation unit, configured to associate the recommendation window with characteristic information of the immersive media corresponding to the recommendation window, and generate a media file of the immersive media; a transceiver unit, configured to send first indication information to a file decapsulation device, wherein the first indication information is used to indicate metadata of the recommended viewport, wherein the metadata of the recommended viewport includes characteristic information of the immersive media corresponding to the recommended viewport; If the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, then the processing unit is specifically used to determine the replaceable group selection information of the point cloud block corresponding to the recommended window as the characteristic information of the immersive media corresponding to the recommended window, and the replaceable group selection information of the point cloud block includes the identification information of the component track corresponding to the point cloud block and at least one of the quality levels corresponding to the point cloud block.

31. A media file decapsulation device, characterized in that: Applied to a file decapsulation device, the device comprises: a transceiver unit, configured to receive first indication information sent by a file encapsulation device, the first indication information being used to indicate metadata of a recommended viewport, the metadata of the recommended viewport including characteristic information of immersive media corresponding to the recommended viewport, the recommended viewport being determined based on content of the immersive media, the immersive media including at least one of multi-view video media and point cloud media, and if the immersive media is point cloud media, and the point cloud media is encapsulated according to point cloud blocks, and each point cloud block of the point cloud media has a different quality level, then the characteristic information includes replaceable group selection information of the point cloud block corresponding to the recommended viewport, the replaceable group selection information including at least one of identification information of a component track corresponding to the point cloud block and a quality level corresponding to the point cloud block; The processing unit is configured to determine, in response to the first indication information, whether to request metadata of the recommended window.

32. A file packaging device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 13.

33. A file decapsulation device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 14 to 29.

34. A computing device, characterized in that include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 13 or 14 to 29.

35. A computer-readable storage medium, characterized in that Used to store a computer program, the computer program causing a computer to execute the method according to any one of claims 1 to 13 or 14 to 29.

Citation Information

Patent Citations

  • Video transmission method, client side and server

    CN108111899A