Point Cloud Media File Encapsulation Method, Device, Equipment and Storage Medium

By adding quality level indication information to point cloud media files, the problem that file encapsulation devices cannot selectively consume some point cloud media is solved, achieving more efficient consumption selectivity and user experience, and saving resources.

CN116137664BActive Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111362971.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-17
Publication Date
2025-07-25
Estimated Expiration
2041-11-17

AI Technical Summary

Technical Problem

In the prior art, file packaging equipment cannot selectively consume part of point cloud media based on the quality level of point cloud tracks, resulting in poor selective consumption of point cloud media.

Method used

Adding quality level indication information in the point cloud media file is used to indicate the quality level of different tracks in the media file, the mass of different tracks in any combination, and the mass of sample groups within the track, allowing the file decapsulation device to selectively consume part of the track and/or part of the samples in the track.

Benefits of technology

It improves the consumption selectivity and flexibility of point cloud media, saves bandwidth and decoding resources, improves user experience, and improves decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116137664B_ABST
    Figure CN116137664B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device and storage medium for encapsulating a point cloud media file. The method includes: a file encapsulation device obtains a bitstream after encoding the point cloud content; encapsulates the bitstream of the point cloud content to obtain a media file of the point cloud content; wherein, the media content includes at least one track and quality level indication information, and the quality level indication information is used to indicate at least one of the quality levels of different tracks within a replaceable group in the media file, the quality levels of any combination of different tracks, and the quality levels of sample groups within a track. That is, the quality level indication information in the present application indicates the quality of different tracks and / or the quality of samples in the tracks, so that the file de-encapsulation device selectively consumes some tracks and / or some samples in the tracks in the point cloud media according to the quality level indication information, thereby improving the consumption flexibility of the point cloud media.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of video processing, and in particular, to a method, apparatus, device, and storage medium for encapsulating point cloud media files. Background Art

[0002] Immersive media refers to media content that can bring an immersive experience to consumers. According to the degree of freedom of users when consuming media content, immersive media can be divided into 3-degree-of-freedom (DoF) media, 3DoF+ media, and 6DoF media.

[0003] Immersive media includes point cloud media. Due to differences in acquisition devices, encoding devices, encoding methods, etc., point cloud media includes point cloud tracks with different quality levels. However, currently, file encapsulation devices indicate point cloud tracks with different quality levels in point cloud media, so that when file decapsulation devices consume point cloud media, they cannot selectively consume part of the point cloud media according to the quality levels of the point cloud tracks, resulting in poor selectable consumption of point cloud media. Summary of the Invention

[0004] The present application provides a method, apparatus, device, and storage medium for encapsulating point cloud media files to improve the selectable consumption of point cloud media, enhance the user experience, and reduce decoding and transmission resources.

[0005] In a first aspect, the present application provides a method for encapsulating a point cloud media file, which is applied to a file encapsulation device. The method includes:

[0006] Obtain the coded stream after encoding the point cloud content;

[0007] Encapsulate the coded stream of the point cloud content to obtain a media file of the point cloud content;

[0008] Wherein, the media content includes at least one track and quality level indication information, and the quality level indication information is used to indicate at least one of the quality levels of different tracks within a replaceable group, the quality levels of different tracks in any combination, and the quality levels of sample groups within a track in the media file.

[0009] In a second aspect, the present application provides a method for decapsulating a point cloud media file, which is applied to a file decapsulation device. The method includes:

[0010] Obtain quality level indication information, where the quality level indication information is used to indicate at least one of the quality of different tracks within a replaceable group, the quality of different tracks in any combination, and the quality of sample groups within a track in a media file, and the media file is obtained by encapsulating the coded stream of the point cloud content;

[0011] Obtain a target file to be decoded according to the quality level indication information;

[0012] Unpackage the target file to obtain a target bitstream to be decoded;

[0013] Decode the target bitstream to obtain target point cloud content.

[0014] In a third aspect, the present application provides a point cloud media file packaging device, which is applied to a file packaging device. The device includes:

[0015] An acquisition unit, configured to acquire a bitstream after encoding point cloud content;

[0016] A packaging unit, configured to package the bitstream of the point cloud content to obtain a media file of the point cloud content;

[0017] Wherein, the media content includes at least one track and quality level indication information, and the quality level indication information is used to indicate at least one of the quality levels of different tracks within a replaceable group in the media file, the quality levels of any combination of different tracks, and the quality levels of sample groups within a track.

[0018] In a fourth aspect, the present application provides a point cloud media file unpackaging device, which is applied to a file unpackaging device. The device includes:

[0019] An acquisition unit, configured to acquire quality level indication information, where the quality level indication information is used to indicate at least one of the quality of different tracks within a replaceable group in a media file, the quality of any combination of different tracks, and the quality of sample groups within a track, and the media file is obtained by packaging a bitstream of point cloud content; and obtain a target file to be decoded according to the quality level indication information;

[0020] An unpackaging unit, configured to unpack the target file to obtain a target bitstream to be decoded;

[0021] A decoding unit, configured to decode the target bitstream to obtain target point cloud content.

[0022] In a fifth aspect, the present application provides a file packaging device, including: a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method of the first aspect.

[0023] In a sixth aspect, the present application provides a file unpackaging device, including: a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method of the second aspect.

[0024] In a seventh aspect, a computing device is provided, including: a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method of the first aspect and / or the second aspect.

[0025] In an eighth aspect, a computer-readable storage medium is provided, which is used to store a computer program, and the computer program causes a computer to execute the method of the first aspect and / or the second aspect.

[0026] In a ninth aspect, a computer program product is provided, including computer program instructions, and the computer program instructions cause a computer to execute the method in any one of the above first aspect and / or the second aspect or its various implementation manners.

[0027] In a tenth aspect, a computer program is provided, which when running on a computer, causes the computer to execute the method in any one of the above first aspect and / or the second aspect or its various implementation manners.

[0028] In summary, in this application, the file encapsulation device obtains the bitstream after encoding the point cloud content; encapsulates the bitstream of the point cloud content to obtain the media file of the point cloud content; wherein, the media content includes at least one track and quality level indication information, and the quality level indication information is used to indicate at least one of the quality levels of different tracks within a replaceable group in the media file, the quality levels of any combination of different tracks, and the quality levels of sample groups within a track. That is, the quality level indication information in this application indicates the quality of different tracks and / or the quality of samples in the tracks, so that the file de-encapsulation device can selectively consume some tracks and / or some samples in the track of the point cloud media according to the quality level indication information, thereby improving the consumption selectivity and flexibility of the point cloud media, enhancing the user experience, and avoiding decoding unnecessary media files, thus saving bandwidth and decoding resources and improving the decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 A schematic diagram of three degrees of freedom is schematically shown;

[0031] Figure 2 A schematic diagram of three degrees of freedom + is schematically shown;

[0032] Figure 3 A schematic diagram of six degrees of freedom is schematically shown;

[0033] Figure 4A It is an architecture diagram of an immersive media system provided by an embodiment of the present application;

[0034] Figure 4B It is a schematic diagram of the content flow of GPCC media provided by an embodiment of the present application;

[0035] Figure 5A It is a schematic diagram of a replaceable group;

[0036] Figure 5B It is another schematic diagram of a replaceable group;

[0037] Figure 6 It is a flowchart of a method for encapsulating a point cloud media file provided by an embodiment of the present application;

[0038] Figure 7 It is a flowchart of a method for decompressing a point cloud media file provided by an embodiment of the present application;

[0039] Figure 8 It is an interaction diagram of a method for encapsulating and decompressing a point cloud media file provided by an embodiment of the present application;

[0040] Figure 9 It is a schematic structural diagram of a device for encapsulating a point cloud media file provided by an embodiment of the present application;

[0041] Figure 10 It is a schematic structural diagram of a device for decompressing a point cloud media file provided by an embodiment of the present application;

[0042] Figure 11 It is a schematic block diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0044] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0045] The embodiments of this application relate to data processing technologies for immersive media.

[0046] Before introducing the technical solution of this application, the relevant knowledge of this application will be introduced below:

[0047] Multi-viewpoint / multi-view video: refers to a video with depth information captured by multiple camera arrays from multiple angles. Multi-viewpoint / multi-view video is also called free-viewpoint / free-view video, which is an immersive media that provides a six-degree-of-freedom experience.

[0048] Point cloud: A point cloud is a set of unordered discrete points in space that represents the spatial structure and surface properties of a three-dimensional object or scene. Each point in the point cloud has at least three-dimensional position information, and may also have color, material or other information depending on the application scenario. Usually, each point in the point cloud has the same number of additional attributes.

[0049] V3C volumetric media: visual volumetric video-based coding media, refers to immersive media that captures visual content from three-dimensional space and provides a 3DoF+ or 6DoF viewing experience, is encoded using traditional video coding, and includes a volumetric video type track in the file encapsulation, including multi-view video, video-coded point cloud, etc.

[0050] PCC: Point Cloud Compression, point cloud compression.

[0051] G-PCC: Geometry-based Point Cloud Compression, geometry-based point cloud compression.

[0052] V-PCC: Video-based Point Cloud Compression, point cloud compression based on traditional video coding.

[0053] Atlas: Indicates the region information on the 2D plane frame, the region information of the 3D rendering space, as well as the mapping relationship between the two and the necessary parameter information required for the mapping.

[0054] Track: Orbit, a collection of media data during the media file encapsulation process. A media file can consist of multiple tracks. For example, a media file can contain a video track, an audio track, and a subtitle track.

[0055] Component track refers to the point cloud geometry data track or the point cloud attribute data track.

[0056] Sample: Encapsulation unit during the media file encapsulation process. A media track consists of many samples. For example, a sample of a video track is usually a video frame.

[0057] DoF: Degree of Freedom. In a mechanical system, it refers to the number of independent coordinates. In addition to the degrees of freedom of translation, there are also degrees of freedom of rotation and vibration. In the embodiments of the present application, it refers to the degrees of freedom that support the movement of the user when watching immersive media and generate content interaction.

[0058] 3DoF: That is, three degrees of freedom, referring to three degrees of freedom for the user's head to rotate around the XYZ axes. Figure 1 Schematically shows a schematic diagram of three degrees of freedom. As Figure 1 shown, at a certain place, a certain point can rotate on the three axes. One can turn the head, lower the head up and down, or shake the head. Through the experience of three degrees of freedom, the user can be immersed in a scene 360 degrees. If it is static, it can be understood as a panoramic picture. If the panoramic picture is dynamic, it is a panoramic video, that is, a VR video. However, VR videos have certain limitations. The user cannot move and cannot choose any place to view.

[0059] 3DoF+: That is, on the basis of three degrees of freedom, the user also has the degrees of freedom to make limited movements along the XYZ axes. It can also be called restricted six degrees of freedom, and the corresponding media bitstream can be called a restricted six degrees of freedom media bitstream. Figure 2 Schematically shows a schematic diagram of three degrees of freedom +.

[0060] 6DoF: That is, on the basis of three degrees of freedom, the user also has the degrees of freedom to move freely along the XYZ axes, and the corresponding media bitstream can be called a six degrees of freedom media bitstream. Figure 3A schematic diagram showing six degrees of freedom is presented. Among them, 6DoF media refers to 6-degree-of-freedom video, which means that the video can provide users with a high-degree-of-freedom viewing experience of freely moving the viewing point in the XYZ axis directions of the three-dimensional space and freely rotating the viewing point around the XYX axes. 6DoF media is a combination of videos from different perspectives in space collected by a camera array. To facilitate the expression, storage, compression, and processing of 6DoF media, the 6DoF media data is expressed as a combination of the following information: texture maps collected by multiple cameras, depth maps corresponding to the texture maps of multiple cameras, and corresponding 6DoF media content description metadata, which contains parameters of multiple cameras and description information such as the stitching layout and edge protection of 6DoF media. At the encoding end, the texture map information and the corresponding depth map information of multiple cameras are stitched, and the description data of the stitching method is written into the metadata according to the defined syntax and semantics. The stitched depth maps and texture map information of multiple cameras are encoded through planar video compression and transmitted to the terminal for decoding, and then the 6DoF virtual viewpoints requested by the user are synthesized to provide the user with a viewing experience of 6DoF media.

[0061] AVS: Audio Video Coding Standard, an audio and video coding standard.

[0062] ISOBMFF: ISO Based Media File Format, a media file format based on the ISO (International Standard Organization) standard. ISOBMFF is a packaging standard for media files, and the most typical ISOBMFF file is the MP4 (Moving Picture Experts Group 4) file.

[0063] DASH: dynamic adaptive streaming over HTTP, a dynamic adaptive streaming technology based on HTTP, which enables high-quality streaming media to be delivered over the Internet through a traditional HTTP web server.

[0064] MPD: media presentation description, a media presentation description signaling in DASH, used to describe media segment information.

[0065] HEVC: High Efficiency Video Coding, the international video coding standard HEVC / H.265.

[0066] VVC: Versatile Video Coding, the international video coding standard VVC / H.266.

[0067] Intra(picture) Prediction: Intra-frame prediction.

[0068] Inter(picture) Prediction: Inter-frame prediction.

[0069] SCC: Screen Content Coding, screen content coding.

[0070] Immersive media refers to media content that can bring an immersive experience to consumers. Immersive media can be classified into 3DoF media, 3DoF+ media, and 6DoF media according to the degree of freedom of users when consuming media content. Common 6DoF media include multi-view video and point cloud media.

[0071] Multi-view video is usually captured by an array of cameras from multiple angles to form the texture information (such as color information) and depth information (such as spatial distance information) of the scene. Together with the mapping information from 2D planar frames to 3D presentation space, it constitutes 6DoF media that can be consumed on the user side.

[0072] A point cloud is a set of discrete points in space that are irregularly distributed and represent the spatial structure and surface properties of a three-dimensional object or scene. Each point in the point cloud has at least three-dimensional position information, and may also have color, material, or other information depending on the application scenario. Usually, each point in the point cloud has the same number of additional attributes.

[0073] Point clouds can flexibly and conveniently represent the spatial structure and surface properties of three-dimensional objects or scenes, and thus are widely used, including virtual reality (VR) games, computer-aided design (CAD), geographic information system (GIS), autonomous navigation system (ANS), digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive telepresence, three-dimensional reconstruction of biological tissues and organs, etc.

[0074] The acquisition of point clouds mainly has the following ways: computer generation, 3D laser scanning, 3D photogrammetry, etc. Computers can generate point clouds of virtual three-dimensional objects and scenes. 3D scanning can obtain point clouds of static real-world three-dimensional objects or scenes, and can acquire millions of point clouds per second. 3D photography can obtain point clouds of dynamic real-world three-dimensional objects or scenes, and can acquire tens of millions of point clouds per second. In addition, in the medical field, point clouds of biological tissue organs can be obtained from MRI, CT, and electromagnetic positioning information. These technologies have reduced the cost and time cycle of point cloud data acquisition and improved the accuracy of the data. The revolution in the way of point cloud data acquisition has made it possible to acquire a large amount of point cloud data. Along with the continuous accumulation of large-scale point cloud data, the efficient storage, transmission, publishing, sharing, and standardization of point cloud data have become the key to point cloud applications.

[0075] After encoding the point cloud media, it is necessary to encapsulate the encoded data stream and transmit it to the user. Correspondingly, on the point cloud media player side, it is necessary to first decompose the encapsulation of the point cloud file, then decode it, and finally present the decoded data stream. Therefore, in the decompression encapsulation link, after obtaining specific information, the efficiency of the decoding link can be improved to a certain extent, thus bringing a better experience for the presentation of point cloud media.

[0076] Figure 4A This is an architecture diagram of an immersive media system provided by an embodiment of the present application. As Figure 4A shown, the immersive media system includes an encoding device and a decoding device. The encoding device may refer to the computer device used by the provider of the immersive media, and this computer device may be a terminal (such as a PC (Personal Computer), a smart mobile device (such as a smart phone), etc.) or a server. The decoding device may refer to the computer device used by the user of the immersive media, and this computer device may be a terminal (such as a PC (Personal Computer), a smart mobile device (such as a smart phone), a VR device (such as a VR helmet, VR glasses, etc.)). The data processing process of the immersive media includes the data processing process on the encoding device side and the data processing process on the decoding device side.

[0077] The data processing process on the encoding device side mainly includes:

[0078] (1) The process of acquiring and producing the media content of the immersive media;

[0079] (2) The process of encoding and file encapsulation of the immersive media. The data processing process on the decoding device side mainly includes:

[0080] (3) The process of file decompression encapsulation and decoding of the immersive media;

[0081] (4) Rendering process of immersive media.

[0082] In addition, there is a transmission process of immersive media between the encoding device and the decoding device. This transmission process can be carried out based on various transmission protocols, and the transmission protocols here may include but are not limited to: DASH (Dynamic Adaptive Streaming over HTTP) protocol, HLS (HTTP Live Streaming) protocol, SMTP (Smart Media Transport Protocaol), TCP (Transmission Control Protocol), etc.

[0083] The following will be combined with Figure 4A to introduce each process involved in the data processing process of immersive media in detail.

[0084] I. Data processing process at the encoding device end:

[0085] (1) Acquisition and production process of the media content of immersive media.

[0086] 1) Acquisition process of the media content of immersive media.

[0087] In one implementation, the capture device may refer to a hardware component provided in the encoding device. For example, the capture device refers to the microphone, camera, sensor, etc. of the terminal. In another implementation, the capture device may also be a hardware device connected to the encoding device, such as a camera connected to the server.

[0088] The capture device may include but is not limited to: audio devices, camera devices, and sensing devices. Among them, the audio devices may include audio sensors, microphones, etc. The camera devices may include ordinary cameras, stereo cameras, light field cameras, etc. The sensing devices may include laser devices, radar devices, etc.

[0089] The number of capture devices may be multiple. These capture devices are deployed at some specific positions in the real space to simultaneously capture the audio content and video content at different angles in this space. The captured audio content and video content are synchronized both in time and space. The media content collected by the capture device is called the original data of immersive media.

[0090] 2) Production process of the media content of immersive media.

[0091] The captured audio content itself is content suitable for audio encoding of immersive media. The captured video content can become content suitable for video encoding of immersive media only after a series of production processes, and the production processes include:

[0092] ① Stitching. Since the captured video content is captured by the capturing device from different angles, stitching refers to stitching the video content captured from these various angles into a complete video that can reflect the 360-degree visual panorama of the real space, that is, the stitched video is a panoramic video (or spherical video) represented in three-dimensional space.

[0093] ② Projection. Projection refers to the process of mapping a three-dimensional video formed by stitching onto a two-dimensional (3-Dimension, 2D) image, and the 2D image formed by projection is called a projection image; the projection methods can include but are not limited to: longitude and latitude map projection, regular hexahedron projection.

[0094] ③ Region encapsulation. The projection image can be directly encoded, or it can be encoded after region encapsulation. In practice, it is found that in the data processing of immersive media, encoding after region encapsulation of the two-dimensional projection image can greatly improve the video encoding efficiency of immersive media. Therefore, the region encapsulation technology is widely applied to the video processing of immersive media. Region encapsulation refers to the process of performing conversion processing on the projection image by region, and the region encapsulation process converts the projection image into an encapsulated image. The process of region encapsulation specifically includes: dividing the projection image into multiple mapping regions, then performing conversion processing on the multiple mapping regions respectively to obtain multiple encapsulated regions, and mapping the multiple encapsulated regions into a 2D image to obtain the encapsulated image. Among them, the mapping region refers to the region obtained by division in the projection image before region encapsulation; the encapsulated region refers to the region located in the encapsulated image after region encapsulation.

[0095] The conversion processing can include but are not limited to: mirroring, rotation, rearrangement, upsampling, downsampling, changing the resolution of the region, and moving, etc.

[0096] It should be noted that since only panoramic videos can be captured by the capture device, after such videos are processed by the encoding device and transmitted to the decoding device for corresponding data processing, users on the decoding device side can only view 360-degree video information by performing some specific actions (such as head rotation), while performing non-specific actions (such as moving the head) cannot obtain corresponding video changes, resulting in a poor VR experience. Therefore, it is necessary to additionally provide depth information that matches the panoramic video to enable users to obtain a better immersion and a better VR experience, which involves 6DoF (Six Degrees of Freedom) production technology. When users can move relatively freely in the simulated scene, it is called 6DoF. When using 6DoF production technology to produce video content of immersive media, the capture device generally selects a light field camera, a laser device, a radar device, etc. to capture point cloud data or light field data in space, and some specific processes are also required during the execution of the above production processes ①-③, such as the processes of cutting and mapping the point cloud data, the calculation process of depth information, etc.

[0097] (2) The process of encoding and file encapsulation of immersive media.

[0098] The captured audio content can be directly audio-encoded to form the audio bitstream of immersive media. After the above production processes ①-② or ①-③, the projection image or the encapsulated image is video-encoded to obtain the video bitstream of immersive media. For example, the packed picture (D) is encoded as the encoded image (Ei) or the encoded video bitstream (Ev). The captured audio (Ba) is encoded as the audio bitstream (Ea). Then, according to a specific media container file format, the encoded image, video, and / or audio are combined into a media file (F) for file playback or a sequence of initialization segments and media segments (Fs) for streaming. The encoding device side also includes metadata, such as projection and region information, into the file or segment, which helps to present the decoded packed picture.

[0099] It should be noted here that if 6DoF production technology is adopted, a specific encoding method (such as point cloud encoding) needs to be used during video encoding. The audio bitstream and the video bitstream are encapsulated in a file container according to the file format of immersive media (such as ISOBMFF (ISO Base Media File Format)) to form a media file resource of immersive media. This media file resource can be a media file or a media segment to form a media file of immersive media; and the metadata of this media file resource of immersive media is recorded according to the requirements of the file format of immersive media using Media Presentation Description (MPD). The metadata here is a general term for information related to the presentation of immersive media. This metadata can include description information of media content, description information of the viewport, and signaling information related to the presentation of media content, etc. As Figure 4A shown, the encoding device will store the media presentation description information and the media file resource formed after the data processing process.

[0100] The immersive media system supports data boxes (Box). A data box refers to a data block or object that includes metadata, that is, the data box contains the metadata of the corresponding media content. Immersive media can include multiple data boxes. For example, it includes a Sphere Region Zooming Box, which contains metadata for describing sphere region zooming information; a 2D Region Zooming Box, which contains metadata for describing 2D region zooming information; a Region Wise Packing Box, which contains metadata for describing the corresponding information during the region packing process, and so on.

[0101] II. Data processing process at the decoding device end:

[0102] (3) The process of unpacking and decoding the file of immersive media;

[0103] The decoding device can adaptively and dynamically obtain the media file resources of immersive media and the corresponding media presentation description information from the encoding device according to the recommendation of the encoding device or the user requirements on the decoding device side. For example, the decoding device can determine the orientation and position of the user based on the detection information of the user's head / eyes / body, and then dynamically request the corresponding media file resources from the encoding device based on the determined orientation and position. The media file resources and the media presentation description information are transmitted from the encoding device to the decoding device through a transmission mechanism (such as DASH, SMT). The process of file de-encapsulation on the decoding device side is the reverse of the file encapsulation process on the encoding device side. The decoding device de-encapsulates the media file resources according to the file format requirements of immersive media to obtain the audio bitstream and the video bitstream. The decoding process on the decoding device side is the reverse of the encoding process on the encoding device side. The decoding device decodes the audio bitstream to restore the audio content.

[0104] In addition, the decoding process of the decoding device for the video bitstream includes the following:

[0105] ① Decode the video bitstream to obtain a planar image; according to the metadata provided by the media presentation description information, if the metadata indicates that the immersive media has performed the region encapsulation process, the planar image refers to the encapsulated image; if the metadata indicates that the immersive media has not performed the region encapsulation process, the planar image refers to the projection image;

[0106] ② If the metadata indicates that the immersive media has performed the region encapsulation process, the decoding device performs region de-encapsulation on the encapsulated image to obtain the projection image. Here, the region de-encapsulation is the reverse of the region encapsulation. The region de-encapsulation refers to the process of performing inverse transformation processing on the encapsulated image according to regions. The region de-encapsulation converts the encapsulated image into a projection image. The process of region de-encapsulation specifically includes: performing inverse transformation processing on multiple encapsulated regions in the encapsulated image respectively according to the indication of the metadata to obtain multiple mapped regions, and mapping the multiple mapped regions to a 2D image to obtain the projection image. The inverse transformation processing refers to the processing that is the reverse of the transformation processing. For example, if the transformation processing refers to a 90-degree counterclockwise rotation, then the inverse transformation processing refers to a 90-degree clockwise rotation.

[0107] ③ Reconstruct the projection image according to the media presentation description information to convert it into a 3D image. Here, the reconstruction processing refers to the process of re-projecting the two-dimensional projection image into the 3D space.

[0108] (4) The rendering process of immersive media.

[0109] The decoding device renders the audio content obtained by audio decoding and the 3D image obtained by video decoding according to the metadata related to rendering and window in the media presentation description information. Once the rendering is completed, the playback output of the 3D image is realized. In particular, if the production technologies of 3DoF and 3DoF+ are adopted, the decoding device mainly renders the 3D image based on the current viewpoint, parallax, depth information, etc. If the production technology of 6DoF is adopted, the decoding device mainly renders the 3D image within the window based on the current viewpoint. Among them, the viewpoint refers to the user's viewing position point, the parallax refers to the line-of-sight difference generated by the user's binocular vision or the line-of-sight difference generated due to movement, and the window refers to the viewing area.

[0110] The immersive media system supports data boxes (Box). A data box refers to a data block or object including metadata, that is, the data box contains the metadata of the corresponding media content. The immersive media can include multiple data boxes. For example, it includes a Sphere Region Zooming Box, which contains metadata for describing sphere region zooming information; a 2D Region Zooming Box, which contains metadata for describing 2D region zooming information; a Region Wise Packing Box, which contains metadata for describing the corresponding information in the region packing process, etc.

[0111] Figure 4B The following is a schematic diagram of the content process of the GPCC media provided by an embodiment of the present application. As Figure 4B shown, the immersive media system includes a file encapsulation device and a file decapsulation device. In some embodiments, the file encapsulation device can be understood as the above-mentioned encoding device, and the file decapsulation device can be understood as the above-mentioned decoding device.

[0112] The visual scene (A) of the real world is captured by a set of cameras or a camera device with multiple lenses and sensors. The acquisition result in the source point cloud data (B). One or more point cloud frames are encoded into an encoded G-PCC bitstream, including an encoded geometry bitstream and an attribute bitstream (E). Then, according to a specific media container file format, one or more encoded bitstreams are combined into a media file (F) for file playback or a sequence of an initialization segment and media segments for streaming (Fs). In the present application, the media container file format is the ISO base media file format specified in ISO / IEC 14496-12. The file encapsulation device can also include metadata into the file or segment. The segment Fs is delivered to the player using a delivery mechanism.

[0113] The file (F) output by the file encapsulation device is the same as the file (F') input to the file decapsulation device. The file decapsulation device processes the file (F') or the received segments (F's), extracts the encoded bitstream (E'), and parses the metadata. Then, the G-PCC bitstream is decoded into a decoded signal (D'), and the point cloud data is generated from the decoded signal (D'). Optionally, depending on the current viewing position, viewing direction, or viewport determined by various types of sensors (such as the head), the point cloud data is rendered and displayed on the screen of a head-mounted display or any other display device, possibly with position or eye movement detection sensors. In addition to being used by the player to access the appropriate part of the decoded point cloud data, the current viewing position or viewing direction can also be used for decoding optimization. In viewport-related delivery, the current viewing position and viewing direction are also passed to the policy module, which determines the tracks to be received.

[0114] The above process applies to both real-time and on-demand use cases.

[0115] Figure 4B The parameters in it are defined as follows:

[0116] E / E': The encoded G-PCC bitstream;

[0117] F / F': The media file including the track format specification, which may contain constraints on the elementary streams included in the track samples.

[0118] A replaceable group, that is, multiple tracks with the same content can be replaced with each other during presentation due to reasons such as encoding method and bit rate. These tracks can form a replaceable group, and the tracks within a replaceable group can be identified by a replaceable group ID. When consuming, the file decapsulation device (such as a client) consumes one track in the replaceable group each time.

[0119] Figure 5A It is a schematic diagram of a replaceable group, including two tracks with replaceable group ID = 10 forming a replaceable group and being replaceable with each other, including two tracks with replaceable group ID = 11 forming a replaceable group and being replaceable with each other, and including three tracks with replaceable group ID = 12 forming a replaceable group and being replaceable with each other.

[0120] Figure 5BFor another schematic diagram of the alternative group, both media track 1 and media track 2 include the identification of the alternative group, that is, alternative group = 1. In this way, it can be determined that media track 1 and media track 2 belong to an alternative group and are media tracks that can be replaced identically in this alternative group. Among them, media track 1 includes a geometric component sub-track and an attribute component sub-track, and media track 2 includes a geometric component sub-track and an attribute component sub-track. The encoding of the attribute component depends on the geometric component. The encoding methods of media track 1 and media track 2 are different. For example, media track 1 is encoded using a lossless encoding method based on GPCC, and media track 2 is encoded using a lossy encoding method based on GPCC.

[0121] In some embodiments, there is a method for indicating the quality level of each track in the alternative group, for example, indicated by the following GPCC alternative quality data box.

[0122] GPCC alternative quality data box

[0123] Definition

[0124] Data box type: 'gpaq'

[0125] Contained in: GPCC sample entry ('gpe1', 'gpeg', 'gpc1', 'gpcg', 'gpeb', 'gpcb') or GPCC item sample entry

[0126] Is it mandatory: No

[0127] Quantity: 0 or 1

[0128] The GPCC alternative quality data box is used to indicate the quality difference information between different tracks within the same alternative group. When this box exists in the track, the quality level information of the relevant track is determined by the content provider. (is used to indicate the quality difference between tracks within the same alternative group. When this box is presented, the quality level is ranked by the content provider).

[0129] Syntax

[0130] aligned(8) class GPCCAlternativesQualityBox extends FullBox('gpaq', 0, 0) {

[0131] unsigned int(8) quality_ranking;

[0132] }

[0133] Semantics

[0134] quality_ranking indicates the quality ranking information of the corresponding track. The lower the value of this field, the higher the quality ranking. (indicates quality ranking information of the corresponding track. The lower value of this field suggests a higher quality).

[0135] The above embodiments indicate the quality rankings of different tracks within the same entity group. However, for tracks that are not within a replaceable group or for media files without a replaceable group, the quality rankings of the tracks are not indicated. As a result, it is not possible to selectively consume part of the point cloud media, leading to poor selectable consumption of the point cloud media.

[0136] To solve the above technical problems, the present application adds quality ranking indication information to the point cloud media file. This quality ranking indication information is used to indicate at least one of the quality of different tracks within a replaceable group in the media file, the quality of any combination of different tracks, and the quality of sample groups within a track. That is, the quality ranking indication information of the present application indicates the quality of different tracks and / or the quality of samples in the tracks, enabling the file unpacking device to selectively consume part of the tracks and / or part of the samples in the tracks in the point cloud media according to this quality ranking indication information. As a result, the selectable consumption and flexibility of the point cloud media are improved, the user experience is enhanced, and the decoding of unnecessary media files is avoided, thereby saving bandwidth and decoding resources and improving the decoding efficiency.

[0137] The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These several embodiments below can be combined with each other, and for the same or similar concepts or processes, they may not be repeated in some embodiments.

[0138] Figure 6 It is a flowchart of a method for encapsulating a point cloud media file provided by an embodiment of the present application, as Figure 6 shown. The method includes the following steps:

[0139] S601. The file encapsulation device obtains the coded bitstream of the point cloud content.

[0140] In some embodiments, the file encapsulation device is also referred to as a point cloud encapsulation device or a point cloud encoding device.

[0141] In some embodiments, the point cloud content is also referred to as point cloud or point cloud data.

[0142] In the embodiments of the present application, the ways for the file encapsulation device to obtain the coded bitstream of the point cloud content include but are not limited to the following:

[0143] Way 1: The file encapsulation device obtains the point cloud content from the acquisition device. For example, the file encapsulation device obtains the point cloud content from the point cloud acquisition device, encodes the point cloud content, and obtains the bitstream of the point cloud content.

[0144] Way 2: The file encapsulation device obtains the coded bitstream of the point cloud content from the storage device. For example, after the encoding device encodes the point cloud content, it stores it in the storage device, and the file encapsulation device reads the bitstream of the point cloud content from the storage device.

[0145] S602: The file encapsulation device encapsulates the bitstream of the point cloud content to obtain the media file of the point cloud content.

[0146] Among them, the media content includes at least one track and quality level indication information, and the quality level indication information is used to indicate at least one of the quality levels of different tracks within the replaceable group, the quality levels of any combination of different tracks, and the quality levels of sample groups within the track in the media file.

[0147] The file encapsulation device groups the obtained bitstream of the point cloud content to obtain the media file of the point cloud content. The media file includes at least one track. Optionally, the media file further includes a replaceable group.

[0148] Due to the change in the stability of the acquisition device, or the change in the acquisition environment, such as the change in light, or the difference in the performance of the encoding device, or the difference in the encoding method, etc., the quality of the generated tracks is different.

[0149] For example, the media file includes Track 1 and Track 2. The contents corresponding to Track 1 and Track 2 are both colors, but the color contents corresponding to Track 1 and Track 2 are different. Therefore, Track 1 and Track 2 cannot form a replaceable group, and the encoding methods used for the color contents corresponding to Track 1 and Track 2 during encoding are also different, resulting in different quality levels of Track 1 and Track 2. For this, in the embodiments of the present application, the quality level indication information can be used to indicate the quality levels of Track 1 and Track 2, so that the file decapsulation device can selectively decode Track 1 and / or Track 2 according to the quality levels of Track 1 and Track 2, improving the decoding selectivity of the media file and enhancing the user experience.

[0150] For another example, the media file includes Track 3 and Track 4, and the corresponding point cloud content of Track 3 and Track 4 is the same. However, due to the non-fixed capabilities of the acquisition device, there are differences in the quality of the point cloud content corresponding to Track 3 and Track 4. Specifically, the quality of the point cloud content varies in units of samples. For this, in the embodiments of the present application, the quality level indication information can be used to indicate the sample groups with different qualities in Track 3 and Track 4. For example, the quality of the samples in Sample Group 1 in Track 3 is different from that of the samples in Sample Group 3 in Track 4. The quality level indication information can indicate the quality levels of Sample Group 1 in Track 3 and Sample Group 3 in Track 4, so that the file demultiplexing device can selectively decode Sample Group 1 in Track 3 or decode Sample Group 3 in Track 4, that is, the present application can selectively decode some samples in the media file, further improving the selectivity of media file decoding.

[0151] In some embodiments, the above quality level indication information is used to indicate the quality levels of different tracks within a replaceable group in the media file, or is used to indicate the quality levels of different tracks in any combination in the media file, or is used to indicate one of the quality levels of sample groups within a track in the media file.

[0152] In some embodiments, if the media file includes multiple tracks, for example, 10 tracks, and these multiple tracks include a replaceable group, then the quality level indication information of the present application can indicate the quality levels of different tracks in the replaceable group in the media file. Optionally, if there are also different tracks in non-replaceable groups in the media file, and these different tracks have different quality levels within the same grade range, then the above quality level indication information can also indicate the quality levels of different tracks in any combination that are not in the replaceable group in the media file. Optionally, if the sample groups in at least two tracks in the media file have different quality levels within the same grade range, then the above quality level indication information can also indicate the quality levels of the sample groups within the tracks in the media file.

[0153] It should be noted that the specific content indicated by the quality level indication information in the embodiments of the present application is determined according to the specific quality levels of the tracks and the specific quality levels of the samples in the tracks. The present application does not limit this.

[0154] In some embodiments, the quality level indication information includes a quality sorting object flag and a quality sorting field.

[0155] Among them, the quality sorting object flag is used to indicate the object for which the quality level sorting is performed.

[0156] In the embodiments of the present application, the above object can be tracks within the same replaceable group, or can be tracks that are not within the same replaceable group. Optionally, the above object further includes sample groups or samples within a track.

[0157] In this application, different numerical values can be assigned to the quality ranking object flag so that the quality ranking object flag indicates different objects for which the quality level ranking is targeted.

[0158] Exemplary one, if the value of the quality ranking object flag is the first numerical value, it indicates that the quality level ranking is for the tracks within the same replaceable group.

[0159] Example two, if the value of the quality ranking object flag is the second numerical value, it indicates that the quality level ranking is for different tracks of any combination carrying the same ranking group identifier.

[0160] Optionally, the ranking group identifier is used to indicate the track group identifier applicable to the quality level ranking.

[0161] In the above example two, when the value of the quality ranking object flag is the second numerical value, the file decapsulation device can determine that the quality level ranking is for different tracks of any combination carrying the same ranking group identifier, and then selectively decode different tracks from the different tracks carrying the same ranking group identifier, thereby achieving selective decoding of different tracks within the non-replaceable group.

[0162] This application does not limit the specific values of the above first numerical value and second numerical value.

[0163] Optionally, the first numerical value is 1.

[0164] Optionally, the second numerical value is 0.

[0165] Optionally, the quality ranking object flag can be represented by the field default_alternatives_ranking_flag.

[0166] Optionally, the ranking group identifier can be represented by the field ranking_group_id.

[0167] Optionally, the above ranking_group_id can also represent any identifier of different tracks corresponding to the quality level ranking.

[0168] In some embodiments, if the above quality level ranking is for different tracks within the same track group, the track group identifier can also be used to replace the above ranking_group_id.

[0169] In some embodiments, if the above quality level ranking is for different entities within the same entity group, the entity group identifier can also be used to replace the above ranking_group_id.

[0170] The above quality ranking field is used to indicate the quality level information of the object. Optionally, the smaller the value of this field, the higher the quality level.

[0171] In some embodiments, the quality ranking field may be represented by the field quality_ranking.

[0172] The media file of the embodiment of the present application includes quality level indication information, and the quality level indication information includes a quality ranking object flag and a quality ranking field. In this way, the file de-encapsulation device can determine the quality level information of the object targeted by the quality level ranking according to the quality ranking object flag and the quality ranking field, and then decode some objects in the media file according to the quality level information of the object, so as to achieve selective consumption of the media file.

[0173] In some embodiments, in addition to the above quality ranking object flag and quality ranking field, the quality level indication information of the embodiment of the present application further includes a ranking unit flag, and the ranking unit flag is used to indicate the unit of the quality level ranking.

[0174] In this embodiment, the content included in the quality level indication information is as follows:

[0175] {quality ranking object flag, quality ranking field, ranking unit flag}.

[0176] In the present application, the units involved in the quality level ranking include tracks and samples.

[0177] In some embodiments, the unit of the quality level ranking can be indicated by assigning different numerical values to the quality level ranking.

[0178] In one example, if the value of the ranking unit flag is the third numerical value, it indicates that the quality level ranking is performed in units of tracks.

[0179] In one example, if the value of the ranking unit flag is the fourth numerical value, it indicates that the quality level ranking is performed in units of samples within the track.

[0180] Optionally, when the ranking unit flag indicates that the quality level ranking is performed in units of samples within the track, the quality level of the sample is indicated by the quality ranking (quality_ranking) field within the sample group.

[0181] The embodiment of the present application does not limit the specific values of the above third numerical value and fourth numerical value.

[0182] Optionally, the above third numerical value is 0.

[0183] Optionally, the above fourth numerical value is 1.

[0184] In some embodiments, the sorting unit flag may be represented by the field sample_ranking_flag.

[0185] In some embodiments, when the value of the sorting unit flag is the fourth value (e.g., 1), the quality level indication information further includes a quality level information sample group, and the quality level information sample group includes a quality ranking field for indicating the quality level information of the samples within the sample group.

[0186] In this embodiment, the content included in the quality level indication information is as follows:

[0187] {quality ranking object flag, sorting unit flag = fourth value, {quality level information sample group: quality level information}}.

[0188] In some embodiments, when the quality level indication information includes the quality level information sample group, the sample group division method for the tracks within the same quality level ranking range is the same. That is, the sample groups of different tracks should have the same presentation time.

[0189] For example, track 1 includes two sample groups, namely sample groups 11 and 12, and track 2 includes sample groups 21 and 22. Among them, the samples in sample group 11 and sample group 21 belong to the same grade range but are different. For example, the quality level of the samples in sample group 11 is 1, and the quality level of the samples in sample group 21 is 2. The sample group division methods of sample group 11 in track 1 and sample group 21 in track 2 are the same, thereby ensuring mutual replaceability and ensuring normal presentation when the user selects sample group 11 or sample group 21 for consumption.

[0190] In some embodiments, the quality level information sample group may be represented by the field GPCCQualityInfoSampleGroupEntry.

[0191] In one example, the code of the quality level information sample group is as follows:

[0192] Syntax

[0193]

[0194] Semantics

[0195] quality_ranking is used to indicate the quality level information of the samples within the sample group. The smaller the value of this field, the higher the quality level. When GPCCQualityInfoSampleGroupEntry exists, for the tracks within the same quality level ranking range, the division of their sample groups should be the same, that is, the sample groups of different tracks should have the same presentation time.

[0196] In some embodiments, when the value of the above-mentioned quality ranking object flag is the second value (for example, default_alternatives_ranking_flag = 0), the quality level indication information further includes a ranking group identifier (for example, ranking_group_id), and this ranking group identifier is used to indicate the track group identifier or entity group identifier applicable to the quality level ranking.

[0197] The embodiments of the present application do not limit the location where the above-mentioned quality level indication information exists in the media file. For example, it can be at the track header of the media file.

[0198] In some embodiments, the quality level indication information is carried in the quality information data box in the media file.

[0199] In some embodiments, if the encapsulation standard of the above-mentioned media file is ISOBMFF, the quality information data box is extended as follows:

[0200] Add relevant fields to the point cloud quality information data box

[0201] Data box type: 'gpaq'

[0202] Contained in: GPCCSampleEntry('gpe1', 'gpeg', 'gpc1', 'gpcg', 'gpeb', 'gpcb') or GPCCTileSampleEntry

[0203] Mandatory: No

[0204] Quantity: 0 or 1

[0205] The GPCCQualityBox is used to indicate the quality information of the point cloud track. When this data box exists in the track, the quality level information of the relevant track is determined by the content provider.

[0206] Syntax

[0207]

[0208]

[0209] Semantics

[0210] quality_ranking indicates the quality level information of the corresponding track and / or sample group. The smaller the value of this field, the higher the quality level.

[0211] When the value of default_alternatives_ranking_flag is 1, it indicates that the quality level ranking is only for the tracks within the same replaceable group. When the value is 0, it indicates that the quality level ranking is for all tracks carrying the same ranking_group_id.

[0212] When the value of sample_ranking_flag is 0, it indicates that the quality level ranking is carried out on a per-track basis. When the value is 1, it indicates that the quality level ranking is carried out on a per-sample basis within the track. When the value of this field is 1, the quality level of the sample is indicated by the quality_ranking field within the sample group.

[0213] ranking_group_id indicates the track group ID for the scope where the quality level ranking takes effect.

[0214] In some embodiments, if the quality level indication information is located in the quality information data box in the media file, the quality ranking field can reuse the quality ranking field in the quality data box.

[0215] This quality information data box is used to indicate the quality information of the point cloud track.

[0216] In some embodiments, when the value of the quality ranking object flag is the second value (for example, default_alternatives_ranking_flag = 1), indicating that the quality level ranking is for all tracks carrying the same ranking_group_id, in addition to separately adding ranking_group_id to indicate the scope where the quality level ranking takes effect, the present application can also indicate the scope where the quality level ranking takes effect through the track group and entity group. This is because the samples within the track group all include the track group identifier, and thus the track group identifier can be used to replace ranking_group_id. Or, because the entities within the entity group all include the entity group identifier, the entity group identifier can be used to replace ranking_group_id.

[0217] In a possible implementation manner, the quality level indication information further includes a quality information track group data box, and the quality information track group data box is used to indicate that the quality level ranking is applicable to the tracks within the track group.

[0218] In this implementation, during file encapsulation, multiple tracks within the same quality level range are encapsulated in one track group. For example, the quality levels of Track 1, Track 2, and Track 3 are 1, 3, and 2 respectively. In this way, Track 1, Track 2, and Track 3 are encapsulated in the same Track Group 01. In this way, Track 1, Track 2, and Track 3 all include the identifier of Track Group 01. That is to say, these 3 tracks belonging to the same quality level range have the same sorting group identifier, and this sorting group identifier is the identifier of Track Group 01. Furthermore, there is no need to separately set a sorting group identifier for these 3 tracks, thereby reducing the codewords of the quality level indication information and improving the encapsulation efficiency.

[0219] Optionally, the track identifier in the quality information track group data box refers to the track group identifier in the track group type data box included in the media file, thereby avoiding the quality information track group data box from carrying unnecessary data and preventing redundancy.

[0220] In some embodiments, the quality level field (quality_ranking) and the sorting unit flag (sample_ranking_flag) in the above quality level indication information can be placed in the above quality information track group data box. That is, if the quality level indication information includes a quality information track group data box, then the quality sorting field and the sorting unit flag are included in the quality information track group data box.

[0221] In some embodiments, the quality information track group can be implemented through the following code:

[0222]

[0223] Among them, TrackGroupTypeBox:

[0224]

[0225]

[0226] If the quality level indication information contains GPCCQualityInfoTrackGroupBox and there are tracks with the same track_group_id in GPCCQualityInfoTrackGroupBox, its quality level information takes effect within the scope of this track group.

[0227] quality_ranking indicates the quality level information of the corresponding track. The smaller the value of this field, the higher the quality level.

[0228] When the value of sample_ranking_flag is 0, it indicates that the quality level ranking is performed on a track-by-track basis. When the value is 1, it indicates that the quality level ranking is performed on a sample-by-sample basis within a track. When the value of this field is 1, the quality level of the sample is indicated by the quality_ranking field within the sample group.

[0229] In another possible implementation, the quality level indication information further includes a quality information entity group data box, which is used to indicate that the quality level ranking applies to the entities within the entity group.

[0230] In this implementation, during file encapsulation, multiple entities within the same quality level range are encapsulated within an entity group. For example, if the quality levels of Entity 1, Entity 2, and Entity 3 are 1, 3, and 2 respectively, then Entity 1, Entity 2, and Entity 3 are encapsulated within the same entity group 01. In this way, Entity 1, Entity 2, and Entity 3 all include the identifier of entity group 01. That is to say, these 3 tracks belonging to the same quality level range have the same sorting group identifier, which is the identifier of entity group 01. Thus, there is no need to separately set a sorting group identifier for these 3 entities, thereby reducing the codewords of the quality level indication information and improving the encapsulation efficiency.

[0231] Optionally, the entity group identifier in the quality information entity group data box refers to the entity group identifier in the entity group type data box included in the media file, thereby avoiding the quality information entity group data box carrying unnecessary data and preventing redundancy.

[0232] In some embodiments, the quality level field (quality_ranking) and the sorting unit flag (sample_ranking_flag) in the above quality level indication information can be placed in the above quality information entity group data box. That is, if the quality level indication information includes a quality information entity group data box, then the quality sorting field and the sorting unit flag are included in the quality information entity group data box.

[0233] In some embodiments, the quality information entity group can be implemented by the following code:

[0234]

[0235]

[0236] Among them, EntityToGroupBox:

[0237]

[0238] For the entities in GPCCQualityInfoEntityToGroupBox, their quality level information takes effect within the scope of the entity group.

[0239] quality_ranking indicates the quality level information of the corresponding entity. The smaller the value of this field, the higher the quality level.

[0240] When sample_ranking_flag has a value of 0, it indicates that the quality level ranking is carried out on an entity basis. When the value is 1, it indicates that the quality level ranking is carried out on the samples within the entity. When the value of this field is 1, the quality level of the sample is indicated by the quality_ranking field within the sample group.

[0241] In the embodiments of the present application, the grouping indicating the scope of quality information taking effect is achieved through the track group and the entity group. The method is simple and enriches the form of quality level indication information.

[0242] In some embodiments, the present application can also be extended with a track selection data box to indicate the difference information between point cloud tracks.

[0243] Exemplarily, the extended track selection data box is as follows:

[0244] Syntax

[0245]

[0246]

[0247] Semantics

[0248] switch_group is an integer used to specify a group or a set of tracks. If this field is 0 (default value) or the track selection data box does not exist, there is no information about whether the track can be used for playback or streaming switch. If the integer is not 0, then the tracks that can be switched with each other should be the same. Tracks belonging to the same switch group should belong to the same replaceable group. A replaceable group may have only one member.

[0249] attribute_list is a list that, until the end of the data box, is attribute information. The attributes in this list should be used as the description of the track or the differentiation criterion for tracks in the same alternate or replaceable group. Each differentiating attribute is associated with a pointer to the field or information that differentiates the tracks.

[0250] Attributes

[0251] The attribute descriptions are shown in Table 1:

[0252] Table 1

[0253]

[0254] The attribute distinctions are as shown in Table 2 below:

[0255] Table 2

[0256]

[0257]

[0258] Among them, the attribute description is used to characterize the modified track, and the attribute distinction is used to distinguish the tracks belonging to the same alternative or switching group. The pointer of the attribute distinction is used to indicate the position information for distinguishing this track from other tracks with the same attributes.

[0259] For the method of encapsulating a point cloud media file according to an embodiment of the present application, the file encapsulation device obtains the bitstream after encoding the point cloud content; encapsulates the bitstream of the point cloud content to obtain the media file of the point cloud content; wherein, the media content includes at least one track and quality level indication information, and the quality level indication information is used to indicate at least one of the quality levels of different tracks within a replaceable group in the media file, the quality levels of any combination of different tracks, and the quality levels of sample groups within a track. That is, the quality level indication information of the present application can indicate the quality of different tracks and / or the quality of samples in the tracks, so that the file de-encapsulation device selectively consumes some tracks and / or some samples in the track of the point cloud media according to the quality level indication information, thereby improving the consumption selectivity and flexibility of the point cloud media, enhancing the user experience, and avoiding decoding unnecessary media files, thus saving bandwidth and decoding resources and improving the decoding efficiency.

[0260] Figure 7 It is a flowchart of the method for de-encapsulating a point cloud media file provided by an embodiment of the present application, as Figure 7 shown, and the method includes the following steps:

[0261] S701. The file de-encapsulation device obtains the quality level indication information.

[0262] Among them, the quality level indication information is used to indicate at least one of the quality of different tracks within a replaceable group in the media file, the quality of any combination of different tracks, and the quality of sample groups within a track.

[0263] Among them, the media file is obtained by encapsulating the bitstream of the point cloud content, and includes at least one track and quality level indication information.

[0264] The implementation manners of the above S701 include but are not limited to the following several types:

[0265] In Method 1, the file encapsulation device sends a first signaling to the file decapsulation device, and the first signaling includes quality level indication information. After receiving the first signaling, the file decapsulation device parses the first signaling to obtain the quality level indication information carried in the first signaling.

[0266] Optionally, the above first signaling is a DASH signaling.

[0267] In Method 2, the file decapsulation device obtains a media file, and the media file includes quality level indication information. Among them, there are various ways for the file decapsulation device to obtain the media file. For example, after the file encapsulation device generates the media file, it sends the media file to the file decapsulation device. Or, the file encapsulation device stores the generated media file in a storage device, such as in a cloud server, and the file decapsulation device reads the media file from the storage device.

[0268] S702: The file decapsulation device obtains a target file to be decoded according to the quality level indication information.

[0269] In some embodiments, when the file decapsulation device reads the quality level indication information from the first signaling, the file decapsulation device sends a first request information to the file encapsulation device according to the quality level indication information, and the first request information includes an identifier of the target file; the file encapsulation device sends the target file in the media file to the file decapsulation device according to the identifier of the target file carried in the first request information.

[0270] In some embodiments, if the file decapsulation device has obtained the media file, the file decapsulation device can obtain the target file to be decoded from the media file according to the quality level indication information.

[0271] In some embodiments, the quality level indication information includes a quality sorting object flag and a quality sorting field. Among them, the quality sorting object flag is used to indicate the object for which the quality level sorting is performed, and the quality sorting field is used to indicate the quality level information of the object. In this way, the file decapsulation device can determine the identifier of the target file according to the quality sorting object flag and the quality sorting field, and then obtain the target file to be decoded from the media file according to the identifier of the target file.

[0272] In some embodiments, when the value of the quality sorting object flag is a first numerical value, it indicates that the quality level sorting is for the tracks within the same replaceable group;

[0273] When the value of the quality sorting object flag is a second numerical value, it indicates that the quality level sorting is for any combination of different tracks carrying the same ranking group identifier.

[0274] In some embodiments, when the value of the quality sorting object flag is the second value, the quality level indication information further includes a ranking group identifier (ranking_group_id), and the ranking group identifier is used to indicate the track group identifier or entity group identifier applicable to the quality level sorting.

[0275] Optionally, the first value is 1.

[0276] Optionally, the second value is 0.

[0277] Exemplarily, when the value of the quality sorting object flag (default_alternatives_ranking_flag) is the first value, the file decompression device can determine the identifier of the target file to be decoded from the alternative group according to the quality sorting field.

[0278] For example, taking track C2 and track C3 as an example for illustration, the quality level indication information corresponding to track C2 and track C3 is as follows:

[0279] C2: {track_id = 2; default_alternatives_ranking_flag = 1;

[0280] sample_ranking_flag = 0; quality_ranking = 0; alternative_group_id = 100};

[0281] C3: {track_id = 3; default_alternatives_ranking_flag = 1;

[0282] sample_ranking_flag = 1; quality_ranking = 1; alternative_group_id = 100}.

[0283] As can be seen from the above, track C2 and track C3 are a group of alternative groups, and the quality of track C2 is better than that of track C3. In this way, when the network resources of the file decompression device are relatively good, the file decompression device can select to request decoding of track C2 for decoding.

[0284] When the value of the quality sorting object flag is the second value, the file decompression device can determine the identifier of the target file to be decoded from different tracks of any combination carrying the same ranking group identifier according to the quality sorting field.

[0285] For example, taking track C2 and track C3 as an example for illustration, the quality level indication information corresponding to track C2 and track C3 is as follows:

[0286] C2: {track_id = 2; default_alternatives_ranking_flag = 0;

[0287] sample_ranking_flag = 0; quality_ranking = 0; ranking_group_id = 100};

[0288] C3: {track_id = 3; default_alternatives_ranking_flag = 0;

[0289] sample_ranking_flag = 0; quality_ranking = 1; ranking_group_id = 100}.

[0290] As can be seen from the above, track C2 and track C3 are not alternative groups, and the ranking_group_id corresponding to both track C2 and track C3 is 100. In this way, one track can be selected from track C2 and track C3 as the track to be decoded according to the value of the quality level field. For example, when the network resources of the file unpacking device are poor, track C3 can be selected to request decoding.

[0291] In some embodiments, the quality level indication information further includes a ranking unit flag, and the ranking unit flag is used to indicate the unit of quality level ranking.

[0292] For example, if the value of the ranking unit flag is the third value, it indicates that the quality level ranking is performed in units of tracks.

[0293] For another example, if the value of the ranking unit flag is the fourth value, it indicates that the quality level ranking is performed in units of samples within a track.

[0294] Based on this, if the value of the ranking unit flag is the third value, the file unpacking device can determine the target track identifier to be decoded according to the quality ranking field. If the value of the ranking unit flag is the fourth value, the file unpacking device can determine the target sample identifier to be decoded according to the quality ranking field.

[0295] In some embodiments, if the value of the ranking unit flag is the fourth value, the quality level indication information further includes a quality level information sample group, and the quality level information sample group includes a quality ranking field, and the quality ranking field is used to indicate the quality level information of the samples within the sample group.

[0296] Optionally, the third value is 0.

[0297] Optionally, the fourth value is 1.

[0298] For example, the quality level indication information corresponding to track C2 and track C3 is as follows:

[0299] C2: {track_id = 2; alternative_group_id = 100; default_alternatives_ranking_flag = 1; sample_ranking_flag = 1}{GPCCQualityInfoSampleGroup21: quality_ranking = 0; GPCCQualityInfoSampleGroup22: quality_ranking = 1};

[0300] C3: {track_id = 3; alternative_group_id = 100; default_alternatives_ranking_flag = 1; sample_ranking_flag = 1}{GPCCQualityInfoSampleGroup31: quality_ranking = 1; GPCCQualityInfoSampleGroup32: quality_ranking = 0}.

[0301] Track C2 includes two groups of quality level information samples, namely GPCCQualityInfoSampleGroup21 and GPCCQualityInfoSampleGroup22, and track C3 includes two groups of quality level information samples, namely GPCCQualityInfoSampleGroup31 and GPCCQualityInfoSampleGroup32. Among them, the quality levels of GPCCQualityInfoSampleGroup21 and GPCCQualityInfoSampleGroup31 belong to the same quality level range. Among them, the quality_ranking of GPCCQualityInfoSampleGroup21 = 0, indicating that the quality level of each sample in GPCCQualityInfoSampleGroup21 is 0, and the quality_ranking of GPCCQualityInfoSampleGroup31 = 1, indicating that the quality level of each sample in GPCCQualityInfoSampleGroup31 is 1. Among them, the quality levels of GPCCQualityInfoSampleGroup22 and GPCCQualityInfoSampleGroup32 belong to the same quality level range. Among them, the quality_ranking of GPCCQualityInfoSampleGroup22 = 1, indicating that the quality level of each sample in GPCCQualityInfoSampleGroup22 is 1, and the quality_ranking of GPCCQualityInfoSampleGroup32 = 0, indicating that the quality level of each sample in GPCCQualityInfoSampleGroup32 is 0.

[0302] In this way, the file de-encapsulation device can determine the identifier of the target file to be decoded according to the above quality level indication information. For example, if the network resources of the file de-encapsulation device are sufficient, the file de-encapsulation device can select GPCCQualityInfoSampleGroup21 in track C2 and GPCCQualityInfoSampleGroup32 in C3 to be decoded as the target file to be decoded.

[0303] In some embodiments, if the quality level indication information includes a quality level information sample group, the sample group division method for the tracks within the same quality level sorting range is the same. For example, the sample division methods of GPCCQualityInfoSampleGroup21 in track C2 and GPCCQualityInfoSampleGroup31 in track C3 are the same, and the sample division methods of GPCCQualityInfoSampleGroup22 in track C2 and GPCCQualityInfoSampleGroup32 in track C3 are the same, ensuring that the sample groups of different tracks should have the same presentation time.

[0304] In some embodiments, when the quality level indication information is located in the quality information data box in the media file, the quality sorting field multiplexes the quality sorting field in the quality data box, and the quality information data box is used to indicate the quality information of the track. In this way, the file decompression device can obtain the quality level indication information from this quality information data box.

[0305] In some embodiments, if the value of the quality sorting object flag is the second value, the quality level indication information further includes a quality information track group data box, where the quality information track group data box is used to indicate that the quality level sorting applies to the tracks within the track group. In this way, the file decompression device can determine that the quality level sorting applies to the tracks within the track group according to the quality information track group data box.

[0306] In some embodiments, if the value of the quality sorting object flag is the second value, the quality level indication information further includes a quality information entity group data box, where the quality information track group data box is used to indicate that the quality level sorting applies to the entities within the entity group. In this way, the file decompression device can determine that the quality level sorting applies to the entities within the entity group according to the quality information entity group data box.

[0307] In some embodiments, if the quality level indication information includes a quality information track group data box, the quality sorting field and the sorting unit flag are included in the quality information track group data box.

[0308] In some embodiments, if the quality level indication information includes a quality information entity group data box, the quality sorting field and the sorting unit flag are included in the quality information entity group data box.

[0309] According to the above method, after obtaining the target file, the file decompression device performs the following steps to implement the decoding of the target file.

[0310] S703. The file decompression device decompresses the target file to obtain the target bitstream to be decoded.

[0311] S704. The file decapsulation device decodes the target bitstream to obtain the target point cloud content.

[0312] Specifically, after receiving the target file, the file decapsulation device first decapsulates the target file to obtain the decapsulated target bitstream, and then decodes the bitstream to obtain the decoded target point cloud content.

[0313] In some embodiments, the decoding of the attribute component is based on the geometric component. At this time, before decoding the attribute component track, the geometric component track is decoded first.

[0314] In the point cloud media file decapsulation method provided by the embodiments of the present application, the file decapsulation device obtains quality level indication information, which is used to indicate at least one of the quality of different tracks within a replaceable group in the media file, the quality of any combination of different tracks, and the quality of sample groups within a track. The media file is obtained by encapsulating the bitstream of the point cloud content; according to the quality level indication information, the target file to be decoded is obtained; the target file is decapsulated to obtain the target bitstream to be decoded; the target bitstream is decoded to obtain the target point cloud content. That is, the file decapsulation device of the present application can selectively consume some tracks and / or some samples in the track of the point cloud media according to the quality level indication information, thereby improving the consumption selectivity and flexibility of the point cloud media, enhancing the user experience, and avoiding decoding unnecessary media files, thus saving bandwidth and decoding resources and improving the decoding efficiency.

[0315] Figure 8 It is an interaction schematic diagram of the point cloud media file encapsulation and decapsulation method provided by an embodiment of the present application. As Figure 8 shown, it includes:

[0316] S801. The file encapsulation device obtains the bitstream after encoding the point cloud content.

[0317] The above S801 refers to the description of S601 above and will not be elaborated here.

[0318] S802. The file encapsulation device encapsulates the bitstream of the point cloud content to obtain the media file of the point cloud content.

[0319] Among them, the media content includes at least one track and quality level indication information, which is used to indicate at least one of the quality levels of different tracks within a replaceable group in the media file, the quality levels of any combination of different tracks, and the quality levels of sample groups within a track.

[0320] In some embodiments, the quality level indication information includes a quality sorting object flag and a quality sorting field. The quality sorting object flag is used to indicate the object for which the quality level sorting is performed, and the quality sorting field is used to indicate the quality level information of the object.

[0321] In some embodiments, the quality level indication information further includes a sorting unit flag, and the sorting unit flag is used to indicate the unit of the quality level sorting.

[0322] In some embodiments, when the value of the sorting unit flag is the fourth value, the quality level indication information further includes a quality level information sample group, and the quality level information sample group includes a quality sorting field, and the quality sorting field is used to indicate the quality level information of the samples within the sample group.

[0323] The above S802 refers to the description of the above S602 and will not be elaborated here.

[0324] S803. The file encapsulation device generates a first signaling, and the first signaling includes quality level indication information.

[0325] S804. The file encapsulation device sends the first signaling to the file decapsulation device.

[0326] S805. The file decapsulation device receives the first signaling sent by the file encapsulation device and parses the first signaling to obtain the quality level indication information.

[0327] S806. The file decapsulation device sends a first request message to the file encapsulation device according to the quality level indication information, and the first request message includes the identifier of the target file.

[0328] Optionally, the above target file may be a part of the media file or the entire media file.

[0329] S807. The file encapsulation device sends the target file in the media file to the file decapsulation device according to the identifier of the target file.

[0330] S808. The file decapsulation device decapsulates the target file to obtain the target bitstream to be decoded.

[0331] S809. The file decapsulation device decodes the target bitstream to obtain the target point cloud content.

[0332] To further illustrate the technical solutions of the embodiments of the present application, specific examples are given below for illustration.

[0333] Example 1, Multi-track Quality Sorting

[0334] Suppose the file encapsulation device includes the point cloud content F. After the point cloud content F is encoded and encapsulated, the following tracks are obtained, namely the geometry track Component1, the attribute track Component2, and the attribute track Component3.

[0335] Step 11. Assume that both attribute tracks C2 and C3 are color attributes, but their color contents are different. Due to different encoding parameter settings, the quality levels of C2 and C3 are 0 and 1 respectively. Since C2 and C3 have different contents and do not belong to the same replaceable group, in file encapsulation, the values of each field are as follows:

[0336] C1: {track_id = 1};

[0337] C2: {track_id = 2; default_alternatives_ranking_flag = 0;

[0338] sample_ranking_flag = 0; quality_ranking = 0; ranking_group_id = 100};

[0339] C3: {track_id = 3; default_alternatives_ranking_flag = 0;

[0340] sample_ranking_flag = 0; quality_ranking = 1; ranking_group_id = 100}.

[0341] Step 12. The file encapsulation device generates a first signaling according to the quality level indication information in Step 11, and the first signaling includes the quality level indication information.

[0342] Optionally, the first signaling can be a signaling description file of a point cloud resource set.

[0343] Step 13. The file encapsulation device sends the first signaling to the file decapsulation device.

[0344] Step 14. The file decapsulation device determines the target file to be decoded according to the quality level indication information in the first signaling.

[0345] For example, the client A corresponding to the file decapsulation device is an on-demand user. Since the network bandwidth is sufficient, to ensure the best viewing experience, it requests to consume C1 + C2.

[0346] For another example, the client B corresponding to the file decapsulation device is an on-demand user. Due to limited network bandwidth and device capabilities, it requests to consume C1 + C3.

[0347] For another example, the client C corresponding to the file decapsulation device plays locally after downloading the complete file. When decoding, it decodes C1 + C2 for consumption according to the quality level information.

[0348] Example 2. Quality ranking of multi-track sample groups

[0349] Suppose the file encapsulation device includes point cloud content F. After the point cloud content F is encoded and encapsulated, the following tracks are obtained, namely geometric track Component1, attribute track Component2, and attribute track Component3:

[0350] Step 21, suppose that attribute tracks C2 and C3 are mutually replaceable tracks. However, due to the fact that the acquisition capabilities of the device during the acquisition process are not fixed, the quality information of C2 and C3 exists differences in units of samples. Therefore, in file encapsulation, the values of each field are as follows:

[0351] C1: {track_id = 1}

[0352] C2: {track_id = 2; alternative_group_id = 100; default_alternatives_ranking_flag = 1; sample_ranking_flag = 1}{GPCCQualityInfoSampleGroup21: quality_ranking = 0; GPCCQualityInfoSampleGroup22: quality_ranking = 1};

[0353] C3: {track_id = 3; alternative_group_id = 100; default_alternatives_ranking_flag = 1; sample_ranking_flag = 1}{GPCCQualityInfoSampleGroup31: quality_ranking = 1; GPCCQualityInfoSampleGroup32: quality_ranking = 0}.

[0354] Step 22, the file encapsulation device generates a first signaling according to the quality level information in Step 21 and sends it to the file decapsulation device.

[0355] Optionally, the first signaling is a signaling description file of the point cloud resource set.

[0356] Step 23, the file encapsulation device sends the first signaling to the file decapsulation device.

[0357] Step 24, the file decapsulation device determines the target file to be decoded according to the quality level indication information in the first signaling.

[0358] For example, after the client A corresponding to the file decapsulation device downloads the complete media file and plays it locally, during decoding, according to the quality level information, when consuming the samples corresponding to consumption sample groups 21 and 31, the samples in 21 (higher quality) are selected for consumption; when consuming the samples corresponding to consumption sample groups 22 and 32, the samples in 32 (higher quality) are selected for consumption.

[0359] It should be understood that Figures 6 to 8 this is only an example of the present application and should not be construed as a limitation to the present application.

[0360] The preferred embodiments of the present application have been described in detail above with reference to the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the technical concept scope of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all fall within the protection scope of the present application. For example, in the various specific technical features described in the above specific embodiments, they can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present application does not separately explain various possible combination methods. For another example, any combination can be made between various different embodiments of the present application, as long as it does not violate the idea of the present application, it should also be regarded as the content disclosed by the present application.

[0361] As described above in conjunction with Figure 6 and Figure 8 , the method embodiments of the present application have been described in detail. Below, in conjunction with Figures 9 to 11 , the device embodiments of the present application will be described in detail.

[0362] Figure 9 FIG. is a schematic structural diagram of a point cloud media file encapsulation device provided by an embodiment of the present application. The device 10 is applied to a file encapsulation device. The device 10 includes:

[0363] An acquisition unit 11, configured to acquire a bitstream after encoding the point cloud content;

[0364] An encapsulation unit 12, configured to encapsulate the bitstream of the point cloud content to obtain a media file of the point cloud content;

[0365] Wherein, the media content includes at least one track and quality level indication information, and the quality level indication information is used to indicate at least one of the quality levels of different tracks within a replaceable group, the quality levels of any combination of different tracks, and the quality levels of sample groups within a track in the media file.

[0366] In some embodiments, the quality level indication information includes any one of a quality ranking object flag, a quality ranking field, and a ranking unit flag. The quality ranking object flag is used to indicate the object for which the quality level ranking is performed. The quality ranking field is used to indicate the quality level information of the object. The ranking unit flag is used to indicate the unit of the quality level ranking.

[0367] In some embodiments, if the value of the quality ranking object flag is a first numerical value, it indicates that the quality level ranking is for the tracks within the same replaceable group;

[0368] If the value of the quality ranking object flag is a second numerical value, it indicates that the quality level ranking is for different tracks of any combination carrying the same ranking group identifier.

[0369] In some embodiments, if the value of the ranking unit flag is a third numerical value, it indicates that the quality level ranking is performed in units of tracks;

[0370] If the value of the ranking unit flag is a fourth numerical value, it indicates that the quality level ranking is performed in units of samples within the track.

[0371] In some embodiments, if the value of the ranking unit flag is a fourth numerical value, the quality level indication information further includes a quality level information sample group. The quality level information sample group includes the quality ranking field. The quality ranking field is used to indicate the quality level information of the samples within the sample group, and the sample group division method for the tracks within the same quality level ranking range is the same.

[0372] In some embodiments, if the value of the quality ranking object flag is the second numerical value, the quality level indication information further includes a ranking group identifier. The ranking group identifier is used to indicate the track group identifier or entity group identifier applicable to the quality level ranking.

[0373] In some embodiments, if the quality level indication information is located in the quality information data box in the media file, the quality ranking field reuses the quality ranking field in the quality data box. The quality information data box is used to indicate the quality information of the track.

[0374] In some embodiments, if the value of the quality ranking object flag is the second numerical value, the quality level indication information further includes a quality information track group data box, or a quality information entity group data box;

[0375] Among them, the quality information track group data box is used to indicate that the quality level sorting is applicable to the tracks within the track group, and the quality information track group data box is used to indicate that the quality level sorting is applicable to the entities within the entity group.

[0376] In some embodiments, if the quality level indication information includes the quality information track group data box, then the quality sorting field and the sorting unit flag are included in the quality information track group data box; or,

[0377] if the quality level indication information includes the quality information entity group data box, then the quality sorting field and the sorting unit flag are included in the quality information entity group data box.

[0378] In some embodiments, the encapsulation unit 12 is further configured to generate a first signaling, where the first signaling includes the quality level indication information; send the first signaling to the file de-encapsulation device; receive a first request information sent by the file de-encapsulation device, where the first request information includes an identifier of a target file; and send the target file in the media file to the file de-encapsulation device according to the identifier of the target file.

[0379] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 9 the illustrated device 10 can execute the method embodiment corresponding to the file encapsulation device, and the foregoing and other operations and / or functions of each module in the device 10 respectively implement the method embodiment corresponding to the file encapsulation device. For the sake of brevity, they will not be elaborated here.

[0380] Figure 10 FIG. 18 is a schematic structural diagram of a point cloud media file de-encapsulation device provided in an embodiment of the present application. The device 20 is applied to a file de-encapsulation device. The device 20 includes:

[0381] An obtaining unit 21, configured to obtain quality level indication information, where the quality level indication information is used to indicate at least one of the quality of different tracks within a replaceable group in a media file, the quality of any combination of different tracks, and the quality of a sample group within a track. The media file is obtained by encapsulating a bitstream of point cloud content; and obtain a target file to be decoded according to the quality level indication information;

[0382] A de-encapsulation unit 22, configured to de-encapsulate the target file to obtain a target bitstream to be decoded;

[0383] A decoding unit 23, configured to decode the target bitstream to obtain target point cloud content.

[0384] In some embodiments, the obtaining unit 21 is specifically configured to receive a first signaling sent by a file encapsulation device, parse the first signaling to obtain the quality level indication information, where the first signaling includes the quality level indication information; according to the quality level indication information, send a first request message to the file encapsulation device, the first request message including an identifier of a target file; and receive the target file sent by the file encapsulation device.

[0385] In some embodiments, the obtaining unit 21 is specifically configured to obtain the media file from a file encapsulation device, and obtain the quality level indication information in the media file; according to the quality level indication information, obtain the target file to be decoded from the media file.

[0386] In some embodiments, the quality level indication information includes any one of a quality sorting object flag, a quality sorting field, and a sorting unit flag. The quality sorting object flag is used to indicate an object for which the quality level sorting is performed. The quality sorting field is used to indicate the quality level information of the object. The sorting unit flag is used to indicate a unit for the quality level sorting.

[0387] In some embodiments, if the value of the quality sorting object flag is a first numerical value, it indicates that the quality level sorting is for tracks within the same replaceable group;

[0388] If the value of the quality sorting object flag is a second numerical value, it indicates that the quality level sorting is for different tracks of any combination carrying the same ranking group identifier.

[0389] In some embodiments, if the value of the sorting unit flag is a third numerical value, it indicates that the quality level sorting is performed in units of tracks;

[0390] If the value of the sorting unit flag is a fourth numerical value, it indicates that the quality level sorting is performed in units of samples within a track.

[0391] In some embodiments, if the value of the sorting unit flag is a fourth numerical value, the quality level indication information further includes a quality level information sample group. The quality level information sample group includes the quality sorting field, and the quality sorting field is used to indicate the quality level information of samples within the sample group. The sample group division method for tracks within the same quality level sorting range is the same.

[0392] In some embodiments, if the value of the quality sorting object flag is the second numerical value, the quality level indication information further includes a sorting group identifier, and the sorting group identifier is used to indicate a track group identifier or an entity group identifier applicable to the quality level sorting.

[0393] In some embodiments, if the quality level indication information is located in the quality information data box of the media file, the quality sorting field multiplexes the quality sorting field in the quality data box, and the quality information data box is used to indicate the quality information of the track.

[0394] In some embodiments, if the value of the quality sorting object flag is the second value, the quality level indication information further includes a quality information track group data box or a quality information entity group data box;

[0395] Wherein, the quality information track group data box is used to indicate that the quality level sorting is applicable to the tracks within the track group, and the quality information track group data box is used to indicate that the quality level sorting is applicable to the entities within the entity group.

[0396] In some embodiments, if the quality level indication information includes the quality information track group data box, the quality sorting field and the sorting unit flag are included in the quality information track group data box; or,

[0397] if the quality level indication information includes the quality information entity group data box, the quality sorting field and the sorting unit flag are included in the quality information entity group data box.

[0398] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 10 The illustrated device 20 can execute the corresponding method embodiment of the server, and the foregoing and other operations and / or functions of each module in the device 20 are respectively for implementing the corresponding method embodiment of the file unpacking device. For the sake of brevity, they will not be elaborated here.

[0399] The device of the embodiments of the present application has been described above from the perspective of functional modules in conjunction with the drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in software form. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0400] Figure 11 It is a schematic block diagram of a computing device provided by an embodiment of the present application. The computing device can be the above-mentioned file encapsulation device, or file decapsulation device, or the computing device has the functions of a file encapsulation device and a file decapsulation device.

[0401] As Figure 11 shown, the computing device 40 may include:

[0402] Memories 41 and 42. The memory 41 is used to store a computer program and transmit the program code to the memory 42. In other words, the memory 42 can call and run the computer program from the memory 41 to implement the method in the embodiment of the present application.

[0403] For example, the memory 42 can be used to execute the above method embodiment according to the instructions in the computer program.

[0404] In some embodiments of the present application, the memory 42 may include, but is not limited to:

[0405] General-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.

[0406] In some embodiments of the present application, the memory 41 includes, but is not limited to:

[0407] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synch Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0408] In some embodiments of the present application, the computer program can be divided into one or more modules, and the one or more modules are stored in the memory 41 and executed by the memory 42 to complete the method provided by the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the video production device.

[0409] As Figure 11 shown, the computing device 40 may further include:

[0410] A transceiver 40, and the transceiver 43 can be connected to the memory 42 or the memory 41.

[0411] Among them, the memory 42 can control the transceiver 43 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 43 can include a transmitter and a receiver. The transceiver 43 can further include an antenna, and the number of antennas can be one or more.

[0412] It should be understood that the components in the video production device are connected through a bus system. Among them, the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.

[0413] This application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the methods in the above method embodiments. Or rather, the embodiments of this application also provide a computer program product containing instructions. When the instructions are executed by the computer, the computer executes the methods in the above method embodiments.

[0414] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that contains one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0415] Those of ordinary skill in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0416] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0417] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of the present application, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0418] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for encapsulating a point cloud media file, characterized in that Applied to a file encapsulation device, the method includes: Obtaining the bitstream after encoding the point cloud content; Encapsulating the bitstream of the point cloud content to obtain the media file of the point cloud content; Wherein, the media file includes at least one track and quality level indication information, and the quality level indication information is used to indicate at least one of the quality levels of different tracks within a replaceable group in the media file, the quality levels of any combination of different tracks, and the quality levels of sample groups within a track. The quality level indication information includes a sorting unit flag, and the sorting unit flag is used to indicate the unit of quality level sorting. If the sorting unit flag indicates that the quality level sorting is performed in units of samples within a track, then the quality level indication information further includes a quality level information sample group, and the quality level information sample group includes a quality sorting field, and the quality sorting field is used to indicate the quality level information of the samples within the sample group, and the sample group division methods of tracks within the same quality level sorting range are the same.

2. The method according to claim 1, wherein The quality level indication information includes any one of a quality sorting object flag and a quality sorting field. The quality sorting object flag is used to indicate the object for which the quality level sorting is performed, and the quality sorting field is used to indicate the quality level information of the object.

3. The method according to claim 2, wherein If the value of the quality sorting object flag is a first numerical value, it indicates that the quality level sorting is for tracks within the same replaceable group; if the value of the quality sorting object flag is a second numerical value, it indicates that the quality level sorting is for any combination of different tracks carrying the same ranking group identifier; Or, If the value of the sorting unit flag is a third numerical value, it indicates that the quality level sorting is performed in units of tracks; if the value of the sorting unit flag is a fourth numerical value, it indicates that the quality level sorting is performed in units of samples within a track.

4. The method according to claim 3, characterized in that If the value of the quality sorting object flag is the second numerical value, then the quality level indication information further includes a sorting group identifier, and the sorting group identifier is used to indicate the track group identifier or entity group identifier applicable to the quality level sorting.

5. The method according to claim 4, characterized in that, If the quality level indication information is located in the quality information data box in the media file, then the quality sorting field reuses the quality sorting field in the quality information data box, and the quality information data box is used to indicate the quality information of the track.

6. The method according to claim 3, characterized in that, If the value of the quality sorting object flag is the second numerical value, then the quality level indication information further includes a quality information track group data box or a quality information entity group data box; Wherein, the quality information track group data box is used to indicate that the quality level sorting is applicable to the tracks within the track group, and the quality information track group data box is used to indicate that the quality level sorting is applicable to the entities within the entity group.

7. The method according to claim 6, characterized in that, If the quality level indication information includes the quality information track group data box, then the quality sorting field and the sorting unit flag are included in the quality information track group data box; or, When the quality level indication information includes the quality information entity group data box, the quality sorting field and the sorting unit flag are included in the quality information entity group data box.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: generating a first signaling that includes the quality level indication information; sending the first signaling to a file de-encapsulation device; receiving first request information sent by the file de-encapsulation device, where the first request information includes an identifier of a target file; sending the target file in the media file to the file de-encapsulation device according to the identifier of the target file.

9. A method for unpacking a point cloud media file, characterized in that, When applied to a file de-encapsulation device, the method includes: obtaining quality level indication information, where the quality level indication information is used to indicate at least one of the quality of different tracks within a replaceable group in a media file, the quality of any combination of different tracks, and the quality of sample groups within a track. The quality level indication information includes a sorting unit flag, and the sorting unit flag is used to indicate the unit of the quality level sorting. When the sorting unit flag indicates that the quality level sorting is performed in units of samples within a track, the quality level indication information further includes a quality level information sample group, and the quality level information sample group includes a quality sorting field. The quality sorting field is used to indicate the quality level information of the samples within the sample group, and the sample group division methods of tracks within the same quality level sorting range are the same. The media file is obtained by encapsulating a bitstream of point cloud content; obtaining a target file to be decoded according to the quality level indication information; de-encapsulating the target file to obtain a target bitstream to be decoded; decoding the target bitstream to obtain target point cloud content.

10. The method according to claim 9, wherein The obtaining of the quality level indication information includes: receiving a first signaling sent by a file encapsulation device and parsing the first signaling to obtain the quality level indication information, where the first signaling includes the quality level indication information; The obtaining of the target file to be decoded according to the quality level indication information includes: sending first request information to the file encapsulation device according to the quality level indication information, where the first request information includes an identifier of a target file; receiving the target file sent by the file encapsulation device.

11. The method according to claim 9, wherein The obtaining of the quality level indication information includes: obtaining the media file from a file encapsulation device and obtaining the quality level indication information in the media file; The obtaining of the target file to be decoded according to the quality level indication information includes: obtaining the target file to be decoded from the media file according to the quality level indication information.

12. The method according to claim 9, characterized in that The quality level indication information includes either a quality sorting object flag or a quality sorting field. The quality sorting object flag is used to indicate the object for which the quality level sorting is performed, and the quality sorting field is used to indicate the quality level information of the object.

13. The method according to claim 12, wherein If the value of the quality sorting object flag is the first numerical value, it indicates that the quality level sorting is for the tracks within the same replaceable group; if the value of the quality sorting object flag is the second numerical value, it indicates that the quality level sorting is for different tracks of any combination carrying the same ranking group identifier. Or, If the value of the sorting unit flag is the third numerical value, it indicates that the quality level sorting is performed in units of tracks; if the value of the sorting unit flag is the fourth numerical value, it indicates that the quality level sorting is performed in units of samples within the tracks.

14. The method according to claim 13, wherein If the value of the quality sorting object flag is the second numerical value, the quality level indication information further includes a sorting group identifier, and the sorting group identifier is used to indicate the track group identifier or entity group identifier applicable to the quality level sorting.

15. The method according to claim 14, wherein If the quality level indication information is located in the quality information data box in the media file, the quality sorting field multiplexes the quality sorting field in the quality information data box, and the quality information data box is used to indicate the quality information of the tracks.

16. The method according to claim 13, characterized in that If the value of the quality sorting object flag is the second numerical value, the quality level indication information further includes a quality information track group data box or a quality information entity group data box; wherein, the quality information track group data box is used to indicate that the quality level sorting is applicable to the tracks within the track group, and the quality information track group data box is used to indicate that the quality level sorting is applicable to the entities within the entity group.

17. The method according to claim 16, wherein If the quality level indication information includes the quality information track group data box, the quality sorting field and the sorting unit flag are included in the quality information track group data box; or, If the quality level indication information includes the quality information entity group data box, the quality sorting field and the sorting unit flag are included in the quality information entity group data box.

18. A point cloud media file encapsulation device, characterized in that Applied to a file encapsulation device, the device includes: An acquisition unit, configured to acquire the coded bitstream of the point cloud content; An encapsulation unit, configured to encapsulate the coded bitstream of the point cloud content to obtain a media file of the point cloud content; wherein, the media file includes at least one track and quality level indication information, and the quality level indication information is used to indicate at least one of the quality levels of different tracks within the replaceable group, the quality levels of different tracks of any combination, and the quality levels of sample groups within the tracks in the media file. The quality level indication information includes a sorting unit flag, and the sorting unit flag is used to indicate the unit of the quality level sorting. If the sorting unit flag indicates that the quality level sorting is performed in units of samples within the tracks, the quality level indication information further includes a quality level information sample group, and the quality level information sample group includes a quality sorting field, and the quality sorting field is used to indicate the quality level information of the samples within the sample group, and the sample group division methods of the tracks within the same quality level sorting range are the same.

19. A point cloud media file de-encapsulation device, characterized in that, Applied to a file de-encapsulation device, the device includes: An acquisition unit, configured to acquire quality level indication information, where the quality level indication information is used to indicate at least one of the quality of different tracks within a replaceable group in a media file, the quality of any combination of different tracks, and the quality of a sample group within a track. The quality level indication information includes a sorting unit flag, and the sorting unit flag is used to indicate the unit of the quality level sorting. If the sorting unit flag indicates that the quality level sorting is performed based on samples within a track, the quality level indication information further includes a quality level information sample group. The quality level information sample group includes a quality sorting field, and the quality sorting field is used to indicate the quality level information of the samples within the sample group. And the sample group division methods of the tracks within the same quality level sorting range are the same. The media file is obtained by encapsulating a bitstream of point cloud content; and according to the quality level indication information, obtain a target file to be decoded; A de-encapsulation unit, configured to de-encapsulate the target file to obtain a target bitstream to be decoded; A decoding unit, configured to decode the target bitstream to obtain target point cloud content.

20. A computing device, characterized in that, Comprising: A processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 8 or 9 to 17.

21. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program causes a computer to execute the method according to any one of claims 1 to 8 or 9 to 17.

Citation Information

Patent Citations

  • Point cloud coding and decoding method and device, computer readable medium and electronic equipment

    CN115150384A

  • Methods for timed metadata priority rank signaling for point clouds

    US20210019936A1

  • Method for processing 360 image data on basis of sub picture, and device therefor

    WO2019199024A1