A data processing method, apparatus, equipment, and medium for point cloud media.

By using a point cloud media file encapsulation method centered on time domain information, the problem of the single encapsulation mode in the existing technology is solved, enabling flexible transmission and decoding of point cloud media and improving the flexibility and efficiency of data processing.

CN116781675BActive Publication Date: 2026-05-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-03-08
Publication Date
2026-05-26

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

This application provides a data processing method, apparatus, device, and computer-readable storage medium for point cloud media. The method includes: acquiring point cloud frames of the point cloud media; encapsulating the point cloud frames into a time-domain hierarchical track based on the time-domain information of the point cloud frames to obtain a media file of the point cloud media. The time-domain information includes time-domain hierarchical information and frame rate information; any sample in the time-domain hierarchical track contains complete data of at least one point cloud frame of the point cloud media. Therefore, encapsulating point cloud media files based on time-domain information (such as time-domain hierarchical information or frame rate information) enriches the encapsulation methods for point cloud media, allowing content consumption devices to flexibly select the required point cloud media files for transmission and decoding consumption according to their needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a data processing method for point cloud media, a data processing device for point cloud media, a data processing equipment for point cloud media, and a computer-readable storage medium. Background Technology

[0002] With the continuous development of science and technology, it is now possible to obtain large amounts of highly accurate point cloud data at a relatively low cost and in a short period of time. As large-scale point cloud data continues to accumulate, the efficient storage, transmission, distribution, sharing, and standardization of point cloud data have become hot topics in point cloud application research.

[0003] Currently, point cloud media encapsulation modes include a multi-track encapsulation mode. In this mode, point cloud media is encapsulated within a geometric component track and one or more attribute component tracks. However, in practice, it has been found that the multi-track encapsulation mode only supports track division based on component type, resulting in a relatively limited encapsulation method. Summary of the Invention

[0004] This application provides a data processing method, apparatus, device, and computer-readable storage medium for point cloud media, which can enrich the encapsulation methods of point cloud media.

[0005] On one hand, embodiments of this application provide a data processing method for point cloud media, including:

[0006] Obtain the media file of the point cloud media. The media file includes a temporal hierarchical track. The temporal hierarchical track is obtained by encapsulating the point cloud frames of the point cloud media based on temporal information. The temporal information includes temporal hierarchical information and frame rate information. Any sample in the temporal hierarchical track contains the complete data of at least one point cloud frame of the point cloud media.

[0007] Decode media files to present point cloud media.

[0008] In this embodiment, a media file of point cloud media is obtained. The media file includes a temporal hierarchical track, which is obtained by encapsulating the point cloud frames of the point cloud media based on temporal information. The temporal information includes temporal hierarchical information and frame rate information. Any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media. The media file is decoded to present the point cloud media. It is evident that encapsulating point cloud media files based on temporal information (such as temporal hierarchical information or frame rate information) enriches the encapsulation methods of point cloud media. Content consumption devices can flexibly select the required point cloud media files for transmission and decoding consumption according to their needs.

[0009] On one hand, embodiments of this application provide a data processing method for point cloud media, including:

[0010] Acquire point cloud frames from point cloud media;

[0011] Based on the temporal information of the point cloud frame, the point cloud frame is encapsulated into a temporal hierarchical track to obtain the media file of the point cloud media; the temporal information includes temporal hierarchical information and frame rate information; any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media.

[0012] In this embodiment, point cloud frames of point cloud media are obtained; based on the temporal information of the point cloud frames, they are encapsulated into temporal hierarchical tracks to obtain the media file of the point cloud media. The temporal information includes temporal hierarchical information and frame rate information; any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media. It is evident that encapsulating point cloud media files using temporal information (such as temporal hierarchical information and frame rate information) as the core enriches the encapsulation methods of point cloud media, allowing content consumption devices to flexibly select the required point cloud media files for transmission and decoding consumption according to their needs.

[0013] On one hand, embodiments of this application provide a data processing apparatus for point cloud media, including:

[0014] The acquisition unit is used to acquire the media file of the point cloud media. The media file includes a temporal hierarchical track, which is obtained by encapsulating the point cloud frames of the point cloud media based on temporal information. The temporal information includes temporal hierarchical information and frame rate information. Any sample in the temporal hierarchical track contains the complete data of at least one point cloud frame of the point cloud media.

[0015] The processing unit is used to decode media files to present point cloud media.

[0016] In one implementation, the sample ingress of the temporal hierarchical orbit uses a point cloud compressed sample ingress based on a geometric model; and the sample ingress type of the temporal hierarchical orbit is a target type.

[0017] Among them, the point cloud compression sample entry based on the geometric model is obtained by expanding the volumetric visual sample entry.

[0018] In one implementation, the time-domain hierarchical orbit includes N samples, where N is a positive integer;

[0019] Each of the N samples is a subset of point cloud frames of the point cloud media at a temporal level.

[0020] In one implementation, the complete data of a point cloud frame consists of one or more coded content units, which are used to store the geometric coded data or attribute coded data of the point cloud frame.

[0021] In one implementation, the time-domain hierarchical orbit includes a first orbit and a second orbit;

[0022] The samples in the first track correspond to the first time domain level, and the samples in the second track correspond to the second time domain level.

[0023] In one implementation, the decoding time of samples in the first track is different from that of samples in the second track.

[0024] In one implementation, the first time domain level is higher than the second time domain level, and the first track is indexed to the second track.

[0025] In one implementation, the media file further includes a sample description data box, which includes one or more type description fields for a sample entry type, the type description fields being used to indicate the sample entry type.

[0026] In one implementation, the acquisition unit is further configured to:

[0027] Obtain the transmission signaling file of the point cloud media. The transmission signaling file includes at least one adaptive set, and the at least one adaptive set includes the representation of the time-domain level track.

[0028] In one implementation, the acquisition unit is used to acquire the media file of the point cloud media, specifically for:

[0029] Based on the description information of at least one adaptive set, determine the media files required to render point cloud media;

[0030] The media files of the determined point cloud media are retrieved using streaming transmission.

[0031] In one implementation, point cloud media includes multiple media segments;

[0032] If a point cloud frame encapsulates at least two media segments in a time-domain level track, then the at least two media segments include an initialization media segment.

[0033] The initial media segment includes a configuration record for a point cloud compressor decoder based on a geometric model.

[0034] In one implementation, the time-domain level track includes a frame rate indication field for indicating the frame rate of the time-domain level track.

[0035] On one hand, embodiments of this application provide a data processing apparatus for point cloud media, including:

[0036] The acquisition unit is used to acquire point cloud frames of point cloud media;

[0037] The processing unit is used to encapsulate the point cloud frame into a time-domain hierarchical track based on the time-domain information of the point cloud frame to obtain the media file of the point cloud media; the time-domain information includes time-domain hierarchical information and frame rate information; any sample in the time-domain hierarchical track contains complete data of at least one point cloud frame of the point cloud media.

[0038] In one implementation, the number of time-domain level tracks is M, the M time-domain level tracks belong to P time-domain levels, M is a positive integer, and P is a positive integer less than or equal to M; the processing unit is further configured to:

[0039] Slicing a media file to obtain multiple media segments; and,

[0040] The transmission signaling file that generates the media file includes at least one adaptive set, and M time-domain level tracks are represented as P representations under at least one adaptive set according to the time-domain level.

[0041] In one implementation, the media file is encapsulated in a single-track encapsulation mode;

[0042] The transmission signaling file includes at least one adaptive set, which includes a representation of the track at the time-domain level and a representation of the track in the single-track encapsulation mode.

[0043] In one implementation, the representation corresponding to the time-domain hierarchical orbit carries a time-domain hierarchical descriptor;

[0044] In single-track encapsulation mode, the track representation does not carry time-domain hierarchical descriptors.

[0045] In one implementation, both the representation of the time-domain level track and the representation of the track in the single-track encapsulation mode carry an encapsulation mode descriptor.

[0046] The value of the encapsulation mode descriptor carried by the representation corresponding to the time-domain level track is different from the value of the encapsulation mode descriptor carried by the representation corresponding to the track in the single-track encapsulation mode.

[0047] In one implementation, the transmission signaling file further includes a preselection set, which includes preselected elements, and the preselected elements include preselected component attributes and codec attributes;

[0048] Preselected component attributes are used to indicate the identifier of one or more representations in at least one adaptive set;

[0049] The codec attribute is used to indicate the sample entry type of the corresponding track as indicated by the preselected component attribute.

[0050] In one implementation, if the sample entry type of the corresponding track indicated by the preselected component attribute is a target type, the processing unit is used to generate a transmission signaling file for the media file, specifically for:

[0051] Configure the value of the codec attribute to the identifier corresponding to the target type.

[0052] In one implementation, the media file is encapsulated in a multitrack encapsulation mode; the transmission signaling file includes at least one adaptive set and a preselected set, the preselected set including preselected elements, the preselected elements including preselected component attributes and codec attributes;

[0053] The preselected component attribute is used to indicate the identifier of at least one adaptive set;

[0054] The codec attribute is used to indicate the sample entry type of the track in the adaptive set indicated by the preselected component attribute.

[0055] In one implementation, if the sample entry type of the corresponding track indicated by the preselected component attribute is a target type, the processing unit is used to generate a transmission signaling file for the media file, specifically for:

[0056] Configure the value of the codec attribute to the identifier corresponding to the target type.

[0057] Accordingly, this application provides a computer device, the device comprising:

[0058] A processor is used to load and execute computer programs;

[0059] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned data processing method for point cloud media.

[0060] Accordingly, this application provides a computer-readable storage medium storing a computer program adapted to be loaded by a processor and executed by the above-described point cloud media data processing method.

[0061] Accordingly, this application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned point cloud media data processing method.

[0062] In this embodiment, point cloud frames of point cloud media are obtained; based on the temporal information of the point cloud frames, they are encapsulated into temporal hierarchical tracks to obtain the media file of the point cloud media. The temporal information includes temporal hierarchical information and frame rate information; any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media. It is evident that encapsulating point cloud media files using temporal information (such as temporal hierarchical information or frame rate information) as the core can enrich the encapsulation methods of point cloud media, enabling content consumption devices to flexibly select the required point cloud media files for transmission and decoding consumption according to their needs. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0064] Figure 1a A schematic diagram of a 6DoF circuit provided in an embodiment of this application;

[0065] Figure 1b A schematic diagram of a 3DoF method provided in an embodiment of this application;

[0066] Figure 1c A schematic diagram of 3DoF+ provided for an embodiment of this application;

[0067] Figure 1d A data processing architecture diagram for point cloud media provided in this application embodiment;

[0068] Figure 1e This is a schematic diagram of the structure of a sample provided in an embodiment of this application;

[0069] Figure 1f A schematic diagram of a multi-track container provided in an embodiment of this application;

[0070] Figure 2 A flowchart illustrating a data processing method for point cloud media provided in this application embodiment;

[0071] Figure 3 A flowchart illustrating another point cloud media data processing method provided in this application embodiment;

[0072] Figure 4 A schematic diagram of the structure of a point cloud media data processing device provided in an embodiment of this application;

[0073] Figure 5A schematic diagram of the structure of another point cloud media data processing device provided in an embodiment of this application;

[0074] Figure 6 This is a schematic diagram of the structure of a content consumption device provided in an embodiment of this application;

[0075] Figure 7 This is a schematic diagram of the structure of a content creation device provided in an embodiment of this application. Detailed Implementation

[0076] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0077] The following describes some technical terms used in the embodiments of this application:

[0078] I. Immersive Media:

[0079] Immersive media refers to media files that provide immersive content, allowing viewers to experience visual, auditory, and other sensory sensations reminiscent of the real world. Based on the degree of freedom viewers have when consuming the media content, immersive media can be categorized as: 6DoF (Degree of Freedom) immersive media, 3DoF immersive media, and 3DoF+ immersive media.

[0080] II. Point Clouds:

[0081] A point cloud is a set of randomly distributed discrete points in space that represent the spatial structure and surface properties of a three-dimensional object or scene. Each point in a point cloud has at least three-dimensional positional information and, depending on the application, may also have color, material, or other information. Typically, each point in a point cloud has the same number of additional attributes.

[0082] III. Point Cloud Media:

[0083] Point cloud media is a typical 6DoF immersive media. It can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes, and is therefore widely used in projects such as Virtual Reality (VR) games, Computer-Aided Design (CAD), Geographic Information Systems (GIS), Autonomous Navigation Systems (ANS), digital cultural heritage, free-viewpoint broadcasting, 3D immersive telepresence, and 3D reconstruction of biological tissues and organs.

[0084] IV. Track:

[0085] A track is a collection of media data during the media file encapsulation process. A media file can consist of one or more tracks. For example, a media file can typically contain a video track, an audio track, and a subtitle track.

[0086] V. Sample:

[0087] A sample is a unit of encapsulation in the media file encapsulation process. A track consists of many samples. For example, a video track can consist of many samples. A sample is usually a video frame.

[0088] 6. ISOBMFF (ISO Based Media File Format): This is a media file encapsulation standard. A typical ISOBMFF file is an MP4 file.

[0089] 7. DASH (Dynamic Adaptive Streaming over HTTP): is an adaptive bitrate technology that enables high-quality streaming media to be delivered over the Internet through traditional HTTP web servers.

[0090] 8. MPD (Media Presentation Description, in DASH) is used to describe media segment information in a media file.

[0091] 9. Representation: refers to a combination of one or more media components in DASH. For example, a video file of a certain resolution can be regarded as a representation; in this application, a video file of a certain temporal level can be regarded as a representation.

[0092] 10. Adaptation Sets: These are collections of one or more video streams in DASH. An Adaptation Set can contain multiple representations.

[0093] This application relates to immersive media data processing technology. Some concepts in the immersive media data processing process will be introduced below. In particular, the following embodiments of this application will use free-viewpoint video as an example for immersive media.

[0094] Figure 1a This diagram illustrates a 6DoF implementation as provided in this application. 6DoF is categorized into window 6DoF, omnidirectional 6DoF, and 6DoF. Window 6DoF restricts the viewer's rotational movement along the X and Y axes, and translation along the Z axis; for example, the viewer cannot see outside the window frame or pass through the window. Omnidirectional 6DoF restricts the viewer's rotational movement along the X, Y, and Z axes; for example, the viewer cannot freely move through the 3D 360° VR content within the restricted movement area. 6DoF allows the viewer to translate freely along the X, Y, and Z axes; for example, the viewer can move freely within the 3D 360° VR content. Similar to 6DoF are 3DoF and 3DoF+ production techniques. Figure 1b This is a schematic diagram of a 3DoF implementation provided in an embodiment of this application; as shown... Figure 1b As shown, 3DoF refers to the viewer of immersive media being fixed at the center point in a three-dimensional space, while the viewer's head rotates along the X, Y, and Z axes to view the images provided by the media content. Figure 1c This is a schematic diagram of a 3DoF+ embodiment provided in this application, as shown below. Figure 1c As shown, 3DoF+ refers to the ability of immersive media viewers to move their heads within a limited space based on 3DoF to view the images provided by the media content when the virtual scene provided by the immersive media has a certain depth information.

[0095] With the continuous development of science and technology, it is now possible to obtain large amounts of highly accurate point cloud data at a relatively low cost and in a short period of time. Point cloud data acquisition methods include computer generation, 3D laser scanning, and 3D photogrammetry. Specifically, point cloud data can be obtained by acquiring visual scenes of the real world through acquisition devices (a set of cameras or a camera device with multiple lenses and sensors). 3D laser scanning can obtain point clouds of static real-world 3D objects or scenes, acquiring millions of point cloud data per second; 3D photogrammetry can obtain point clouds of dynamic real-world 3D objects or scenes, acquiring tens of millions of point cloud data per second. Furthermore, in the medical field, point cloud data of biological tissues and organs can be obtained through magnetic resonance imaging (MRI), computed tomography (CT), and electromagnetic positioning information. Additionally, point cloud data can also be directly generated by computers based on virtual 3D objects and scenes; for example, computers can generate point cloud data for virtual 3D objects and scenes. With the continuous accumulation of large-scale point cloud data, the efficient storage, transmission, publication, sharing, and standardization of point cloud data have become crucial for point cloud applications.

[0096] Figure 1d This is a data processing architecture diagram for point cloud media provided in an embodiment of this application. For example... Figure 1d As shown, the data processing process on the content production device side mainly includes: (1) the acquisition process of media content from point cloud data; (2) the encoding and file encapsulation process of point cloud data. The data processing process on the content consumption device side mainly includes: (3) the file decapsulation and decoding process of point cloud data; (4) the rendering process of point cloud data. In addition, the transmission process of point cloud media between the content production device and the content consumption device is involved. This transmission process can be based on various transmission protocols, including but not limited to: DASH (Dynamic Adaptive Streaming over HTTP) protocol, HLS (HTTP Live Streaming) protocol, SMTP (Smart Media Transport Protocol), TCP (Transmission Control Protocol), etc.

[0097] The data processing procedure for point cloud media is described in detail below:

[0098] (1) Obtain the media content of point cloud media.

[0099] From the perspective of how point cloud media content is acquired, it can be divided into two methods: acquiring sound and visual scenes from the real world through capture devices, and generating them through computers. In one implementation, the capture device can refer to a hardware component installed in the content production equipment, such as a microphone, camera, or sensor on the terminal. In another implementation, the capture device can also be a hardware device connected to the content production equipment, such as a camera connected to a server; used to provide the content production equipment with point cloud data media content acquisition services. The capture device can include, but is not limited to, audio devices, camera devices, and sensing devices. Audio devices can include audio sensors, microphones, etc. Camera devices can include ordinary cameras, stereo cameras, light field cameras, etc. Sensing devices can include laser devices, radar devices, etc. Multiple capture devices can be used, deployed at specific locations in the real space to simultaneously capture audio and video content from different angles within that space, with the captured audio and video content remaining synchronized in both time and space. Due to the different acquisition methods, the compression encoding methods corresponding to the media content of different point cloud data may also differ.

[0100] (2) The process of encoding and encapsulating media content in point cloud media.

[0101] Currently, geometry-based point cloud compression (GPCC) is commonly used to encode acquired point cloud data, resulting in a geometry-based compressed bitstream (including encoded geometric bitstream and attribute bitstream). The encapsulation modes of geometry-based compressed bitstream include single-track encapsulation and multi-track encapsulation.

[0102] Single-track encapsulation mode refers to encapsulating point cloud bitstreams in the form of a single track. In single-track encapsulation mode, a sample will contain one or more coded content units (such as a geometric coded content unit and multiple attribute coded content units). The advantage of single-track encapsulation mode is that a single-track encapsulated point cloud file can be obtained based on the point cloud bitstream without much processing. Figure 1e This is a schematic diagram of the structure of a sample provided in an embodiment of this application, such as... Figure 1e As shown, in the single-track encapsulation mode, the parameter set, geometric information, and attribute information are encapsulated in a single sample.

[0103] Multi-track encapsulation mode refers to encapsulating point cloud bitstreams in the form of multiple tracks. In multi-track encapsulation mode, each track contains one component in the point cloud bitstream, namely a geometric component track and one or more attribute component tracks. The advantage of multi-track encapsulation is that encapsulating different components separately allows the client to select the required components for transmission and decoding consumption according to its own needs. Figure 1f A schematic diagram of a multi-track container provided in this application embodiment is shown below. Figure 1f As shown, track 1 contains an encoded geometric information bitstream but not an encoded attribute information bitstream; track 2 contains an encoded attribute information bitstream but not an encoded geometric information bitstream.

[0104] (3) The process of decapsulating and decoding point cloud media files;

[0105] Content consumption devices can obtain media file resources and corresponding media presentation description information from point cloud data through content production devices. The media file resources and media presentation description information of the point cloud data are transmitted from the content production device to the content consumption device via a transmission mechanism (such as DASH or SMT). The file decapsulation process on the content consumption device side is the reverse of the file encapsulation process on the content production device side. The content consumption device decapsulates the media file resources according to the file format requirements of the point cloud media to obtain the encoded bitstream (GPCC bitstream or VPCC bitstream). The decoding process on the content consumption device side is the reverse of the encoding process on the content production device side. The content consumption device decodes the encoded bitstream to reconstruct the point cloud data.

[0106] (4) The rendering process of point cloud media.

[0107] The content consumption device renders the point cloud data obtained by decoding the GPCC bitstream based on the metadata related to rendering and windowing in the media presentation description information, obtains the point cloud frames of the point cloud media, and presents the point cloud media according to the presentation time of the point cloud frames.

[0108] In one embodiment, the content production device first samples a real-world visual scene using an acquisition device to obtain point cloud data corresponding to the real-world visual scene. Then, it encodes the acquired point cloud data using geometry-based point cloud compression (GPCC) to obtain a GPCC bitstream (including encoded geometric bitstreams and attribute bitstreams). Next, it encapsulates the GPCC bitstream to obtain a media file (i.e., point cloud media) corresponding to the point cloud data. Specifically, the content production device combines one or more encoded bitstreams into a media file for file playback, or a sequence of initialization segments and media segments for streaming, according to a specific media container file format. The media container file format refers to the ISO Basic Media File Format as specified in International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) 14496-12. In one embodiment, the content production device also encapsulates metadata into the media file or the sequence of initialization / media segments and transmits the sequence of initialization / media segments to the content consumption device via a transmission mechanism (such as a dynamic adaptive streaming media transmission interface).

[0109] On the content consumption device side: First, it receives point cloud media files sent by the content production device, including media files for playback or initialization segments and media segment sequences for streaming. Then, it decapsulates the point cloud media files to obtain an encoded GPCC bitstream. Next, it parses the encoded GPCC bitstream (i.e., decodes the encoded GPCC bitstream to obtain point cloud data). In specific implementations, the content consumption device determines the media files or media segment sequences required for presenting the point cloud media based on the current viewing position / direction. It then decodes these media files or media segment sequences to obtain the required point cloud data. Finally, based on the current viewing (window) direction, it renders the decoded point cloud data to obtain point cloud frames of the point cloud media, and presents the point cloud media on the screen of the head-mounted display or any other display device carried by the content consumption device according to the presentation time of the point cloud frames. It should be noted that the current viewing position / direction is determined by head tracking and possibly visual tracking functions. In addition to using a renderer to render point cloud data of the current object's viewing position / viewing direction, an audio decoder can also be used to decode and optimize the audio in the current object's viewing (viewport) direction.

[0110] Content creation equipment and content consumption equipment can together form a point cloud media system. Content creation equipment refers to the computer equipment used by the provider of point cloud media (e.g., the content creator of point cloud media). This computer equipment can be a terminal (such as a PC, a smart mobile device, or a smartphone) or a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Content consumption equipment refers to the computer equipment used by the user of point cloud media (e.g., the viewer of point cloud media). This computer equipment can be a terminal (such as a PC, a smart mobile device, a smartphone, VR devices such as VR headsets or VR glasses), smart home appliances, in-vehicle terminals, or aircraft.

[0111] It is understood that the data processing technology of point cloud media involved in this application can be implemented based on cloud technology; for example, using a cloud server as a content production device. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to realize the computation, storage, processing, and sharing of data.

[0112] As can be seen from the data processing process of point cloud media described above, the multi-track encapsulation mode only supports track division based on component type. In many applications, such as scenarios with a geometric component track and an attribute (color) component track, the client must obtain both the geometric component track and the attribute track simultaneously when consuming data. In such cases, the advantages of multi-track encapsulation are difficult to realize.

[0113] Based on this, this application proposes a data processing method for point cloud media. It encapsulates point cloud media files using temporal information (such as temporal hierarchy information and frame rate information) to obtain temporal hierarchy tracks, enriching the encapsulation methods for point cloud media. This allows clients to flexibly select the most suitable point cloud file for transmission and decoding consumption according to their needs. Each sample in the temporal hierarchy track contains complete data of at least one point cloud frame of the point cloud media. The sample entry point of the temporal hierarchy track uses a geometric model-based point cloud compression (GPCCSampleEntry) sample entry point. The syntax of the geometric model-based point cloud compression (GPCCSampleEntry) sample entry point is shown in Table 1.

[0114] Table 1

[0115]

[0116] As shown in Table 1, the sample entry type for the temporal-level orbit is a target type, which can be represented as 'agpl'; that is, the orbit corresponding to the 'agpl' type sample entry is the temporal-level orbit. The point cloud compression sample entry based on the geometric model is obtained by extending the volumetric visual sample entry. The GPCC decoder configuration information corresponding to the point cloud compression sample entry based on the geometric model is indicated by the config in the point cloud compression configuration data box based on the geometric model (GPCCConfigurationBox).

[0117] In one implementation, the sample entry type of a time-domain level track can be indicated by a SampleDescriptionBox, which includes one or more type description fields for the sample entry type. The type description fields are used to indicate the sample entry type, which may include 'gpll', 'gplg', and 'agpl'. The mandatory nature of the SampleDescriptionBox is conditional.

[0118] The track corresponding to a sample entry of type 'agpl' (i.e., the temporal-level track) can include N samples, where N is a positive integer. In the temporal-level track, each sample contains at least one complete point cloud frame (i.e., each sample contains all component data of at least one point cloud frame); the component data of the point cloud frame is stored in one or more coded content units of the sample, and the component data of the point cloud frame includes at least one of the following: geometric coded data, attribute coded data; for example, a sample contains one geometric coded content unit and multiple attribute coded content units.

[0119] In the track corresponding to the 'agpl' type sample entry (i.e., the temporal hierarchy track), each sample is a subset of the point cloud frames of the point cloud media at a temporal hierarchy. Specifically, the point cloud media can be divided into multiple point cloud frames based on the temporal hierarchy. These point cloud frames are encoded into a GPCC bitstream, and the samples in the temporal hierarchy track are a subset of all point cloud frames in the GPCC bitstream in the temporal domain.

[0120] In one implementation, the GPCC bitstream corresponding to the point cloud media can be encapsulated in multiple tracks (i.e., time domain level tracks) corresponding to sample entries of type 'agpl'. The samples in different time domain level tracks do not overlap in decoding time (or rendering time). For example, if the decoding time of the sample in time domain level track 1 is t1 and the rendering time is t2, and the decoding time of the sample in time domain level track 2 is t3 and the rendering time is t4, then t1≠t3 and t2≠t4.

[0121] Furthermore, within the temporal hierarchy track, samples can be assigned different temporal hierarchy identifiers (i.e., different samples are divided into different temporal hierarchy levels); for example, sample 1 has a temporal hierarchy identifier of 0, and sample 2 has a temporal hierarchy identifier of 1. Within the temporal hierarchy track, samples can correspond to a specific frame rate (e.g., 30fps). For details on the specific encapsulation methods of samples within the temporal hierarchy track, please refer to [reference needed]. Figure 1e The sample encapsulation method will not be elaborated here. Optionally, if the time domain level of time domain level track 1 is higher than the time domain level of time domain level track 2, then time domain level track 1 is indexed to time domain level track 2.

[0122] Correspondingly, in the transmission signaling file corresponding to the point cloud media, the track (i.e., the domain-level track) corresponding to the sample entry of type 'agpl' is represented in DASH as one or more representations under the same adaptive set.

[0123] In one implementation, the media file of point cloud media can be divided into multiple media segments. The point cloud bitstreams of these media segments are encapsulated in one or more time-domain level tracks according to the time-domain level. If at least two media segments are encapsulated in a time-domain level track, then these at least two media segments include an initial media segment, which contains a point cloud compression decoder configuration record based on a geometric model.

[0124] In one embodiment, the representation of the track (i.e., the domain-level track) corresponding to a sample entry of type 'agpl' can reside in the same adaptive set as the representation of the track in the single-track encapsulation mode; for example, the adaptive set may include representation 1 and representation 2, where representation 1 is the representation of track 1 corresponding to a sample entry of type 'agpl', and representation 2 is the representation of the track in the single-track encapsulation mode. In the single-track encapsulation mode, the representation of the track (i.e., the domain-level track) corresponding to a sample entry of type 'agpl' needs to be distinguished from the representation of the track in the single-track encapsulation mode. Optionally, there are three schemes:

[0125] 1. Using time-domain-dependent descriptors, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the time-domain-dependent track) needs to be distinguished from the representation of the track in single-track encapsulation mode. The representation of the track corresponding to the 'agpl' type sample entry (i.e., the time-domain-dependent track) carries a time-domain-dependent descriptor; the representation of the track in single-track encapsulation mode does not carry a time-domain-dependent descriptor. The syntax of the time-domain-dependent descriptor is shown in Table 2:

[0126] Table 2

[0127]

[0128] As shown in Table 2, the representations of tracks corresponding to 'agpl' type sample entries (i.e., i.e., i-domain level tracks) and tracks in single-track encapsulation mode can be distinguished by the time-domain level-related descriptor gpcc:@temporal_level_Ids. Specifically, the representations of tracks corresponding to 'agpl' type sample entries (i.e., i-domain level tracks) carry the time-domain level-related descriptor gpcc:@temporal_level_Ids; the representations of tracks in single-track encapsulation mode do not carry the time-domain level-related descriptor gpcc:@temporal_level_Ids.

[0129] 2. By defining a new encapsulation mode descriptor, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) needs to be distinguished from the representation of the track in the single-track encapsulation mode. Both the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) and the representation of the track in the single-track encapsulation mode carry an encapsulation mode descriptor, with different values ​​for the mode field. The syntax of the encapsulation mode descriptor is shown in Table 3:

[0130] Table 3

[0131]

[0132] As shown in Table 3, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) can be distinguished from the representation of the track in the single-track encapsulation mode by the encapsulation mode descriptor gpcc:@mode. Specifically, the value of the encapsulation mode descriptor gpcc:@mode carried by the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) is different from the value of the encapsulation mode descriptor gpcc:@mode carried by the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) has a value of 'agpl', while the single-track encapsulation mode has a value of 'gpel'.

[0133] 3. By defining a preselection set, the representation corresponding to the track (i.e., domain-level track) for sample entries of type 'agpl' needs to be distinguished from the representation corresponding to the track in single-track encapsulation mode. The preselection set can be used to indicate the representation corresponding to the track (i.e., domain-level track) for sample entries of type 'agpl'. Specifically, the representation corresponding to the track (i.e., domain-level track) for sample entries of type 'agpl' can be defined by the Preselection element in the MPD file. The Preselection Components attribute in the Preselection element is used to indicate the identifier of one or more representations in the adaptive set. The Codecs attribute in the Preselection element is used to indicate the sample entry type of the track corresponding to one or more representations indicated by the Preselection Components attribute. When the Codecs attribute in the Preselection element is 'agpl', it means that one or more representations indicated by the Preselection Components attribute in this preselection set are all representations corresponding to the track (i.e., domain-level track) for sample entries of type 'agpl'.

[0134] In another embodiment, in multi-track encapsulation mode, tracks with different sample entry types can be encapsulated in different adaptive sets, and the track corresponding to the 'agpl' type sample entry (i.e., domain-level track) can exist in a separate adaptive set. In this case, by defining a preselection set, the representation corresponding to the track corresponding to the 'agpl' type sample entry (i.e., domain-level track) needs to be distinguished from the representation corresponding to the track in single-track encapsulation mode. The preselection set can be used to indicate the representation corresponding to the track corresponding to the 'agpl' type sample entry (i.e., domain-level track). Specifically, the representation corresponding to the track corresponding to the 'agpl' type sample entry (i.e., domain-level track) can be defined by the Preselection element in the MPD file. The Preselection element includes a preselection component (@preselectionComponents) attribute and a codec (@codecs) attribute. The preselection component (@preselectionComponents) attribute in the Preselection element can indicate the identifier of one or more adaptive sets. The codec (@codecs) attribute in the Preselection element is used to indicate the sample entry type of the track in one or more adaptive sets indicated by the preselection component attribute. When the codec (@codecs) attribute in the Preselection element is 'agpl', it means that the adaptive set indicated by the preselection component (@preselectionComponents) attribute in the preselection set is the representation corresponding to the track (i.e., domain-level track) of the sample entry of type 'agpl'. In other words, the adaptive set indicated by the preselection component (@preselectionComponents) attribute in the preselection set is the adaptive set corresponding to the track (i.e., domain-level track) of the sample entry of type 'agpl'.

[0135] In this embodiment, point cloud frames of point cloud media are obtained; based on the temporal information of the point cloud frames, they are encapsulated into temporal hierarchical tracks to obtain the media file of the point cloud media. The temporal information includes temporal hierarchical information and frame rate information; any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media. It is evident that encapsulating point cloud media files using temporal information (such as temporal hierarchical information or frame rate information) as the core can enrich the encapsulation methods of point cloud media, enabling content consumption devices to flexibly select the required point cloud media files for transmission and decoding consumption according to their needs.

[0136] Figure 2A flowchart of a data processing method for point cloud media provided in this application embodiment; the method can be executed by a content consumption device in a point cloud media system, and the method includes the following steps S201 and S202:

[0137] S201. Obtain the media files of the point cloud media.

[0138] The media file includes a temporal hierarchical track, which is obtained by encapsulating point cloud frames of the point cloud media based on temporal information, including temporal hierarchical information and frame rate information. The sample entry type of the temporal hierarchical track is 'agpl', and any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media, that is, any sample in the temporal hierarchical track contains all component data of at least one point cloud frame (the component data of the point cloud frame includes at least one of the following: geometric encoded data, attribute encoded data; for example, a sample contains one geometric encoded content unit and multiple attribute encoded content units); the complete data of a point cloud frame consists of one or more encoded content units, which are used to store the geometric encoded data or attribute encoded data of the point cloud frame.

[0139] In one implementation, the temporal hierarchical track is obtained by encapsulating the point cloud bitstream corresponding to the point cloud frame of the point cloud media based on temporal hierarchical information; in another implementation, the temporal hierarchical track is obtained by encapsulating the point cloud bitstream corresponding to the point cloud frame of the point cloud media based on frame rate information. The media file may also include a SampleDescriptionBox, which includes one or more type description fields for the sample entry type. The type description fields indicate the sample entry type, which may include 'gpll', 'gplg', and 'agpl'. The sample description box is conditionally mandatory.

[0140] Specifically, the sample entry point of the temporal hierarchical orbit can use a point cloud compressed sample entry point based on a geometric model; and the sample entry point type of the temporal hierarchical orbit is a target type (the target type can be specifically represented as 'agpl' type); among which, the point cloud compressed sample entry point based on the geometric model is obtained by expanding the volumetric visual sample entry point.

[0141] The track corresponding to the 'agpl' type sample entry (i.e., the temporal-level track) can include N samples, where N is a positive integer; each sample is a subset of the point cloud frames of the point cloud media at a temporal level. Specifically, the point cloud media can be divided into multiple point cloud frames based on the temporal level. These point cloud frames are encoded into a GPCC stream, and the samples in the temporal-level track are a subset of all point cloud frames in the GPCC stream in the temporal domain.

[0142] In one implementation, the temporal hierarchy tracks include a first track and a second track; samples in the first track correspond to the first temporal hierarchy, and samples in the second track correspond to the second temporal hierarchy. The first and second temporal hierarchies are two different temporal hierarchies, and the decoding time (or rendering time) of samples in the first track is different from that of samples in the second track; that is, samples in different temporal hierarchy tracks do not overlap in decoding time or rendering time; for example, if the decoding time of a sample in temporal hierarchy track 1 is t1 and the rendering time is t2, and the decoding time of a sample in temporal hierarchy track 2 is t3 and the rendering time is t4, then t1 ≠ t3, and t2 ≠ t4. Optionally, if the first temporal hierarchy is higher than the second temporal hierarchy, then the first track is indexed to the second track.

[0143] In one embodiment, the media file of the point cloud media is obtained by the content consuming device from the complete file of the point cloud media stored locally, based on the needs of the object (such as the rendering frame rate of the point cloud media set by the object) or its own decoding capabilities.

[0144] In another embodiment, the content consumption device obtains the transmission signaling file of the point cloud media, and determines the media file required to present the point cloud media based on the description information in the transmission signaling file, the object's requirements (such as the presentation frame rate of the point cloud media set by the object), its own decoding capability, or the current network conditions (such as network transmission speed), and pulls the determined media file of the point cloud media through streaming transmission.

[0145] In one implementation, the media file is encapsulated in a single-track encapsulation mode. The transmission signaling file of the point cloud media includes at least one adaptive set. This adaptive set may include the representation corresponding to the track (i.e., the domain-level track) of the 'agpl' type sample entry, as well as the representation corresponding to the track in the single-track encapsulation mode. That is, the representation corresponding to the track (i.e., the domain-level track) of the 'agpl' type sample entry can reside in the same adaptive set as the representation corresponding to the track in the single-track encapsulation mode.

[0146] In single-track encapsulation mode, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) needs to be distinguished from the representation of the track in single-track encapsulation mode. Optionally, there are three solutions:

[0147] (1) Using time-domain hierarchical descriptors, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the time-domain hierarchical track) needs to be distinguished from the representation of the track in the single-track encapsulation mode. In this case, the representation of the time-domain hierarchical track carries the time-domain hierarchical descriptor; the representation of the track in the single-track encapsulation mode does not carry the time-domain hierarchical descriptor. The syntax of the time-domain hierarchical descriptor can be found in Table 2 above, and will not be repeated here.

[0148] (2) By defining a new encapsulation mode descriptor, the representation corresponding to the track (i.e., the time-domain level track) of the 'agpl' type sample entry needs to be distinguished from the representation corresponding to the track in the single-track encapsulation mode. In this case, both the representation corresponding to the time-domain level track and the representation corresponding to the track in the single-track encapsulation mode carry encapsulation mode descriptors; the value of the encapsulation mode descriptor carried by the representation corresponding to the time-domain level track is different from the value of the encapsulation mode descriptor carried by the representation corresponding to the track in the single-track encapsulation mode; for example, the value of the encapsulation mode descriptor gpcc:@mod carried by the representation corresponding to the track (i.e., the time-domain level track) of the 'agpl' type sample entry is 'agpl', while the value of the encapsulation mode descriptor gpcc:@mod carried by the representation corresponding to the track in the single-track encapsulation mode is 'gpel'. The syntax of the encapsulation mode descriptor can be found in Table 3 above, and will not be repeated here.

[0149] (3) By defining a preselection set, the representation corresponding to the track (i.e., time domain level track) of the 'agpl' type sample entry needs to be distinguished from the representation corresponding to the track in the single-track encapsulation mode. In this case, the transmission signaling file also includes a preselection set, which includes a Preselection element. The Preselection element includes a Preselection Components attribute and a Codecs attribute. The Preselection Components attribute is used to indicate the identifier of the representation corresponding to one or more tracks in the adaptive set, and the Codecs attribute is used to indicate the sample entry type of the track corresponding to the representation of one or more tracks indicated by the Preselection Components attribute. When the Codecs attribute in the Preselection element is 'agpl', it means that the sample entry type of the track corresponding to one or more representations indicated by the Preselection Components attribute in the preselection set is all 'agpl' type, that is, the tracks corresponding to one or more representations indicated by the Preselection Components attribute in the preselection set are all time domain level tracks.

[0150] In another implementation, the media file uses a multi-track encapsulation mode. Tracks with different sample entry types can be encapsulated in different adaptive sets. In this case, the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) can exist in a separate adaptive set. In this scenario, by defining a preselection set, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) needs to be distinguished from the representation of the track in the single-track encapsulation mode. The preselection set can be used to indicate the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track). Specifically, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) can be defined by the Preselection element in the MPD file. The Preselection element includes the preselection component (@preselectionComponents) attribute and the codec (@codecs) attribute. The preselection component (@preselectionComponents) attribute in the Preselection element indicates the identifier of one or more adaptive sets. The codec (@codecs) attribute in the Preselection element indicates the type of one or more adaptive sets indicated by the preselection component attribute. When the codec (@codecs) attribute in the Preselection element is 'agpl', it means that the adaptive set indicated by the preselection component (@preselectionComponents) attribute in the preselection set is the representation corresponding to the track (i.e., domain-level track) of the sample entry of type 'agpl'. In other words, the adaptive set indicated by the preselection component (@preselectionComponents) attribute in the preselection set is the adaptive set corresponding to the track (i.e., domain-level track) of the sample entry of type 'agpl'.

[0151] In another implementation, the media file of point cloud media can be divided into multiple media segments (each media segment contains at least one point cloud frame). The point cloud bitstreams of these media segments are encapsulated in one or more time-domain level tracks according to the time-domain level. If at least two media segments are encapsulated in a time-domain level track, then these at least two media segments contain an initial media segment, which contains a point cloud compression decoder configuration record (GPCCDecoderConfigurationRecord) based on the geometric model.

[0152] Optionally, the time-domain level track includes a frame rate indicator field (frame_rate) which indicates the frame rate of the time-domain level track; for example, when the frame rate of the time-domain level track is 30fps, frame_rate = 30fps.

[0153] S202. Decode the media file to present point cloud media.

[0154] The decoding process on the content consumption device is the reverse of the encoding process on the content production device. The content consumption device decodes the encoded bitstream to reconstruct point cloud data. It then renders the obtained point cloud data based on metadata related to rendering and windowing in the media presentation description information, obtaining point cloud frames of the point cloud media. The point cloud media is then presented according to the presentation time of the point cloud frames. For a detailed implementation of how the content consumption device decodes media files to present point cloud media, please refer to [link to relevant documentation]. Figure 1d The implementation methods for decoding and presenting Zhongdian Cloud Media will not be elaborated here.

[0155] In this embodiment, a media file of point cloud media is obtained. The media file includes a temporal hierarchical track, which is obtained by encapsulating the point cloud frames of the point cloud media based on temporal information. The temporal information includes temporal hierarchical information and frame rate information. Any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media. The media file is decoded to present the point cloud media. It is evident that encapsulating point cloud media files based on temporal information (such as temporal hierarchical information or frame rate information) enriches the encapsulation methods of point cloud media. Content consumption devices can flexibly select the required point cloud media files for transmission and decoding consumption according to their needs.

[0156] Figure 3 A flowchart of another data processing method for point cloud media provided in an embodiment of this application; the method can be executed by a content creation device in a point cloud media system, and the method includes the following steps S301 and S302:

[0157] S301. Obtain point cloud frames from point cloud media.

[0158] In one implementation, the content production device acquires the media content of the point cloud media; the specific acquisition method can be found in [reference needed]. Figure 1d The implementation method for obtaining media content from point cloud media will not be described in detail here. After obtaining the media content from point cloud media, the content production device can divide the media content from point cloud media based on the time domain level (or frame rate information) to obtain point cloud frames of point cloud media.

[0159] S302. Based on the temporal information of the point cloud frame, encapsulate the point cloud frame into a temporal hierarchical track to obtain the media file of the point cloud media.

[0160] Temporal information includes temporal hierarchy information and frame rate information. Any sample in a temporal hierarchy track contains complete data for at least one point cloud frame of the point cloud media, meaning any sample in a temporal hierarchy track contains all component data for at least one point cloud frame (the component data of a point cloud frame includes at least one of the following: geometrically encoded data, attribute-encoded data; for example, a sample contains one geometrically encoded content unit and multiple attribute-encoded content units). The complete data of a point cloud frame consists of one or more encoded content units, which are used to store the geometrically encoded data or attribute-encoded data of the point cloud frame.

[0161] In one embodiment, the content creation device encapsulating a point cloud frame into a time-domain hierarchical track based on its temporal information means that the content creation device encapsulates the point cloud bitstream corresponding to the point cloud frame into a time-domain hierarchical track based on the temporal hierarchical information of the point cloud frame; for example, encapsulating point cloud bitstreams corresponding to point cloud frames at different time-domain hierarchical levels into different time-domain hierarchical tracks. Specific encapsulation methods can be found in [reference needed]. Figure 1d The file encapsulation process will not be described in detail here. In another embodiment, the content production device encapsulates the point cloud frame into a time-domain hierarchical track based on the time-domain information of the point cloud frame. This means that the content production device encapsulates the point cloud bitstream corresponding to the point cloud frame into a time-domain hierarchical track based on the frame rate information of the point cloud frame.

[0162] In one implementation, the sample entry point of the temporal hierarchical orbit can use a point cloud compressed sample entry point based on a geometric model; and the sample entry point type of the temporal hierarchical orbit is a target type (the target type can specifically be represented as the 'agpl' type); wherein, the point cloud compressed sample entry point based on the geometric model is obtained by expanding the volumetric visual sample entry point.

[0163] The track corresponding to the 'agpl' type sample entry (i.e., the temporal-level track) can include N samples, where N is a positive integer; each sample is a subset of the point cloud frames of the point cloud media at a temporal level. Specifically, the point cloud media can be divided into multiple point cloud frames based on the temporal level. These point cloud frames are encoded into a GPCC stream, and the samples in the temporal-level track are a subset of all point cloud frames in the GPCC stream in the temporal domain.

[0164] In one implementation, the temporal hierarchy tracks include a first track and a second track; samples in the first track correspond to the first temporal hierarchy, and samples in the second track correspond to the second temporal hierarchy. The first and second temporal hierarchies are two different temporal hierarchies, and the decoding time (or rendering time) of samples in the first track is different from that of samples in the second track; that is, samples in different temporal hierarchy tracks do not overlap in decoding time or rendering time; for example, if the decoding time of a sample in temporal hierarchy track 1 is t1 and the rendering time is t2, and the decoding time of a sample in temporal hierarchy track 2 is t3 and the rendering time is t4, then t1 ≠ t3, and t2 ≠ t4. Optionally, if the first temporal hierarchy is higher than the second temporal hierarchy, then the first track is indexed to the second track.

[0165] Optionally, the time-domain level track includes a frame rate indicator field (frame_rate) which indicates the frame rate of the time-domain level track; for example, when the frame rate of the time-domain level track is 30fps, frame_rate = 30fps.

[0166] In another implementation, the number of time-domain level tracks is M, and the M time-domain level tracks belong to P time-domain levels, where M is a positive integer and P is a positive integer less than or equal to M. The content production device supports streaming transmission. After obtaining the media file of the point cloud media, the content production device can slice the media file of the point cloud media to obtain multiple media segments; and generate a transmission signaling file for the media file. The transmission signaling file includes at least one adaptive set, and the M time-domain level tracks are represented as P representations under at least one adaptive set according to the time-domain level. For example, if 3 time-domain level tracks belong to 3 different time-domain levels, then these 3 time-domain level tracks can be represented as 3 representations under an adaptive set.

[0167] In one implementation, the media file is encapsulated in a single-track encapsulation mode. The transmission signaling file of the point cloud media includes at least one adaptive set. This adaptive set may include the representation corresponding to the track (i.e., the domain-level track) of the 'agpl' type sample entry, as well as the representation corresponding to the track in the single-track encapsulation mode. That is, the representation corresponding to the track (i.e., the domain-level track) of the 'agpl' type sample entry can reside in the same adaptive set as the representation corresponding to the track in the single-track encapsulation mode.

[0168] In single-track encapsulation mode, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) needs to be distinguished from the representation of the track in single-track encapsulation mode. Optionally, there are three solutions:

[0169] (1) Using time-domain hierarchical descriptors, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the time-domain hierarchical track) needs to be distinguished from the representation of the track in the single-track encapsulation mode. In this case, the representation of the time-domain hierarchical track carries the time-domain hierarchical descriptor; the representation of the track in the single-track encapsulation mode does not carry the time-domain hierarchical descriptor. The syntax of the time-domain hierarchical descriptor can be found in Table 2 above, and will not be repeated here.

[0170] (2) By defining a new encapsulation mode descriptor, the representation corresponding to the track (i.e., the time-domain level track) of the 'agpl' type sample entry needs to be distinguished from the representation corresponding to the track in the single-track encapsulation mode. In this case, both the representation corresponding to the time-domain level track and the representation corresponding to the track in the single-track encapsulation mode carry encapsulation mode descriptors; the value of the encapsulation mode descriptor carried by the representation corresponding to the time-domain level track is different from the value of the encapsulation mode descriptor carried by the representation corresponding to the track in the single-track encapsulation mode; for example, the value of the encapsulation mode descriptor gpcc:@mod carried by the representation corresponding to the track (i.e., the time-domain level track) of the 'agpl' type sample entry is 'agpl', while the value of the encapsulation mode descriptor gpcc:@mod carried by the representation corresponding to the track in the single-track encapsulation mode is 'gpel'. The syntax of the encapsulation mode descriptor can be found in Table 3 above, and will not be repeated here.

[0171] (3) By defining a preselection set, the representation corresponding to the track (i.e., time domain level track) of the 'agpl' type sample entry needs to be distinguished from the representation corresponding to the track in the single-track encapsulation mode. In this case, the transmission signaling file also includes a preselection set, which includes a Preselection element. The Preselection element includes a Preselection Components attribute and a Codecs attribute. The Preselection Components attribute is used to indicate the identifier of the representation corresponding to one or more tracks in the adaptive set, and the Codecs attribute is used to indicate the sample entry type of the track corresponding to the representation of one or more tracks indicated by the Preselection Components attribute. When the Codecs attribute in the Preselection element is 'agpl', it means that the sample entry type of the track corresponding to one or more representations indicated by the Preselection Components attribute in the preselection set is all 'agpl' type, that is, the tracks corresponding to one or more representations indicated by the Preselection Components attribute in the preselection set are all time domain level tracks.

[0172] In one embodiment, if the sample entry type of the corresponding track indicated by the preselected component attribute is a target type (i.e., 'agpl' type), then the transmission signaling file for generating media files by the content consumption device includes: configuring the value of the codec attribute to the identifier corresponding to the target type (i.e., 'agpl').

[0173] In another implementation, the media file uses a multi-track encapsulation mode. Tracks with different sample entry types can be encapsulated in different adaptive sets. In this case, the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) can exist in a separate adaptive set. In this scenario, by defining a preselection set, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) needs to be distinguished from the representation of the track in the single-track encapsulation mode. The preselection set can be used to indicate the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track). Specifically, the representation of the track corresponding to the 'agpl' type sample entry (i.e., the domain-level track) can be defined by the Preselection element in the MPD file. The Preselection element includes the preselection component (@preselectionComponents) attribute and the codec (@codecs) attribute. The preselection component (@preselectionComponents) attribute in the Preselection element indicates the identifier of one or more adaptive sets. The codec (@codecs) attribute in the Preselection element indicates the type of one or more adaptive sets indicated by the preselection component attribute. When the codec (@codecs) attribute in the Preselection element is 'agpl', it means that the adaptive set indicated by the preselection component (@preselectionComponents) attribute in the preselection set is the representation corresponding to the track (i.e., domain-level track) of the sample entry of type 'agpl'. In other words, the adaptive set indicated by the preselection component (@preselectionComponents) attribute in the preselection set is the adaptive set corresponding to the track (i.e., domain-level track) of the sample entry of type 'agpl'.

[0174] In one embodiment, if the sample entry type of the track corresponding to the representation in the adaptive set indicated by the preselected component attribute is the target type (i.e., 'agpl' type), then the transmission signaling file for the media file generated by the content consumption device includes: configuring the value of the codec attribute to the identifier (i.e., 'agpl') corresponding to the target type.

[0175] In another implementation, the media file of point cloud media can be divided into multiple media segments (each media segment contains at least one point cloud frame). The point cloud bitstreams of these media segments are encapsulated in one or more time-domain level tracks according to the time-domain level. If at least two media segments are encapsulated in a time-domain level track, then these at least two media segments contain an initial media segment, which contains a point cloud compression decoder configuration record (GPCCDecoderConfigurationRecord) based on the geometric model.

[0176] The data processing method for point cloud media provided in this application is illustrated below through two complete examples:

[0177] In streaming media transmission scenarios, the content production device encapsulates the point cloud bitstream into multiple different time-domain level tracks (i.e., tracks with a sample entry type of 'agpl') based on the temporal information of the point cloud bitstream corresponding to the point cloud frame, such as temporal level information and frame rate information. This yields media files F1 and F2 of the point cloud media. Each temporal level track satisfies the following constraints:

[0178] a) Each sample in the time-domain hierarchical orbit contains at least one complete point cloud frame, that is, each sample contains all the component data (all geometric data and attribute data) of at least one point cloud frame.

[0179] b) The decoding time (or presentation time) of samples in different time-domain level tracks is different.

[0180] Media files F1, including:

[0181] {track1: 'agpl', frame_rate = 30fps, temporal_level_id = {0,1}}

[0182] {track2: 'agpl', frame_rate=30fps, temporal_level_id={2,3}}

[0183] {track3: 'agpl', frame_rate=30fps, temporal_level_id={4,5}}

[0184] Media file F2 includes:

[0185] { track4: 'gpel', frame_rate = 90fps}

[0186] Among them, 'agpl' and 'gpel' are used to indicate the type of track, frame_rate is used to indicate the frame rate of the track, and temporal_level_id is used to indicate the temporal level of the time-domain track.

[0187] When the content production device needs to stream media file F1, it organizes multiple time-domain level tracks (track1-track3) in media file F1 as different representations within the same adaptive set, and organizes track4 in media file F2 as a representation within another adaptive set, generating the corresponding MPD signaling file. Specifically:

[0188] <adaptationset id="1" codecs="agpl">

[0189] <representation1 frame_rate="30fps,temporal_level_id" = {0,1}>

[0190] <representation2 frame_rate="30fps,temporal_level_id" = {2,3}>

[0191] <representation3 frame_rate="30fps,temporal_level_id" = {4,5}>

[0192] < / representation3> < / representation2> < / representation1> < / adaptationset>

[0193] <adaptationset id="2" codecs="gpe1">

[0194] <representation frame_rate="90fps">

[0195]

[0196] < / representation>

[0197] < / adaptationset>

[0198] <preselection id="1" preselectioncomponents="1" codecs="agpl">

[0199] < / preselection>

[0200] in, <adaptationset id="1" codecs="agpl">Indicates: The identifier of the adaptive set is 1, and the codec type is 'agpl'; <representation1 frame_rate="30fps,temporal_level_id" ={0,1}>This indicates that track 1 corresponds to a frame rate of 30fps and a temporal hierarchy of {0,1}; similarly, <representation2 frame_rate="30fps,temporal_level_id" = {2,3}>This indicates that the frame rate of track 2 is 30fps, and the temporal level is {2,3}. <representation3 frame_rate="30fps,temporal_level_id" = {4,5}>This indicates that the frame rate of track 3 is 30fps, and the temporal level is {4,5}. <adaptationset id="2" codecs="gpe1">This indicates that the identifier for the adaptive set is 2, and the codec type is 'gpel'. <preselection id="1" preselectioncomponents="1" codecs="agpl">It indicates that the identifier of the preselected set is 1, the identifier of the adaptive set indicated by the preselected component attribute is 1 (i.e., adaptive set 1), and the type of the encoding / decoding attribute is 'agpl' (i.e., indicating that the representations in adaptive set 1 are all representations of time-domain level tracks).

[0201] Next, the content creation device sends the MPD signaling file to the content consumption device.

[0202] Content consumption devices can request and decode appropriate media resources based on one or more of the viewer's needs, network conditions, and decoding capabilities, combined with information such as the frame rate or temporal level of different tracks. For example:

[0203] Assuming that the content consumption device C1 has sufficient bandwidth and unrestricted decoding capabilities, the content consumption device C1 can directly request media segments from Adaptation Set 2 for consumption.

[0204] Assuming that content consuming device C1 has limited bandwidth, it can first request Representation1 from Adaptation Set1 for consumption. Once bandwidth becomes more abundant, it can simultaneously request both Representation1 and Representation2 from Adaptation Set1 to achieve a higher frame rate.

[0205] In local playback scenarios, the content production device encapsulates the point cloud bitstream into multiple different time-domain level tracks (i.e., tracks with sample entry type 'agpl') based on the temporal information of the point cloud bitstream corresponding to the point cloud frames, such as temporal level information and frame rate information. This yields the media file F1 of the point cloud media. Each temporal level track satisfies the following constraints:

[0206] a) Each sample in the time-domain hierarchical orbit contains at least one complete point cloud frame, that is, each sample contains all the component data (all geometric data and attribute data) of at least one point cloud frame.

[0207] b) The decoding time (or presentation time) of samples in different time-domain level tracks is different.

[0208] Media files F1, including:

[0209] {track1: 'agpl', frame_rate = 30fps, temporal_level_id = {0,1}}

[0210] {track2: 'agpl', frame_rate=30fps, temporal_level_id={2,3}}

[0211] {track3: 'agpl', frame_rate=30fps, temporal_level_id={4,5}}

[0212] Among them, 'agpl' and 'gpel' are used to indicate the type of track, frame_rate is used to indicate the frame rate of the track, and temporal_level_id is used to indicate the temporal level of the time-domain track.

[0213] The content creation device then sends the complete file F1 to the content consumption device.

[0214] Content consumption devices can decode and present corresponding tracks based on the viewer's needs, decoding capabilities, and other factors, combined with information such as frame rate or temporal hierarchy of different tracks. For example:

[0215] Assuming that the content consumption device C1 has unlimited decoding capabilities and requires a higher frame rate than the target frame rate, then the content consumption device C1 can decode all tracks (track1-track3) for consumption.

[0216] Assuming that the content consumption device C1 has limited decoding capabilities and requires a frame rate lower than the target frame rate, the content consumption device C1 can decode track1 for consumption.

[0217] In this embodiment, point cloud frames of point cloud media are obtained; based on the temporal information of the point cloud frames, they are encapsulated into temporal hierarchical tracks to obtain the media file of the point cloud media. The temporal information includes temporal hierarchical information and frame rate information; any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media. It is evident that encapsulating point cloud media files using temporal information (such as temporal hierarchical information or frame rate information) as the core can enrich the encapsulation methods of point cloud media, enabling content consumption devices to flexibly select the required point cloud media files for transmission and decoding consumption according to their needs.

[0218] The methods of the embodiments of this application have been described in detail above. In order to facilitate better implementation of the above solutions of the embodiments of this application, the apparatus of the embodiments of this application is provided below.

[0219] Please see Figure 4 , Figure 4 This is a schematic diagram of a point cloud media data processing device provided in an embodiment of this application; the point cloud media data processing device can be a computer program (including program code) running on a content consumption device, for example, the point cloud media data processing device can be application software in the content consumption device. Figure 4 As shown, the data processing device for point cloud media includes an acquisition unit 401 and a processing unit 402.

[0220] Please see Figure 4 In one exemplary embodiment, the various units are described in detail below:

[0221] The acquisition unit 401 is used to acquire the media file of the point cloud media. The media file includes a temporal hierarchical track. The temporal hierarchical track is obtained by encapsulating the point cloud frames of the point cloud media based on temporal information. The temporal information includes temporal hierarchical information and frame rate information. Any sample in the temporal hierarchical track contains the complete data of at least one point cloud frame of the point cloud media.

[0222] Processing unit 402 is used to decode media files to present point cloud media.

[0223] In one implementation, the sample ingress of the temporal hierarchical orbit uses a point cloud compressed sample ingress based on a geometric model; and the sample ingress type of the temporal hierarchical orbit is a target type.

[0224] Among them, the point cloud compression sample entry based on the geometric model is obtained by expanding the volumetric visual sample entry.

[0225] In one implementation, the time-domain hierarchical orbit includes N samples, where N is a positive integer;

[0226] Each of the N samples is a subset of point cloud frames of the point cloud media at a temporal level.

[0227] In one implementation, the complete data of a point cloud frame consists of one or more coded content units, which are used to store the geometric coded data or attribute coded data of the point cloud frame.

[0228] In one implementation, the time-domain hierarchical orbit includes a first orbit and a second orbit;

[0229] The samples in the first track correspond to the first time domain level, and the samples in the second track correspond to the second time domain level.

[0230] In one implementation, the decoding time of samples in the first track is different from that of samples in the second track.

[0231] In one implementation, the first time domain level is higher than the second time domain level, and the first track is indexed to the second track.

[0232] In one implementation, the media file further includes a sample description data box, which includes one or more type description fields for a sample entry type, the type description fields being used to indicate the sample entry type.

[0233] In one embodiment, the acquisition unit 401 is further configured to:

[0234] Obtain the transmission signaling file of the point cloud media. The transmission signaling file includes at least one adaptive set, and the at least one adaptive set includes the representation of the time-domain level track.

[0235] In one embodiment, the acquisition unit 401 is used to acquire the media file of the point cloud media, specifically for:

[0236] Based on the description information of at least one adaptive set, determine the media files required to render point cloud media;

[0237] The media files of the determined point cloud media are retrieved using streaming transmission.

[0238] In one implementation, point cloud media includes multiple media segments;

[0239] If a point cloud frame encapsulates at least two media segments in a time-domain level track, then the at least two media segments include an initialization media segment.

[0240] The initial media segment includes a configuration record for a point cloud compressor decoder based on a geometric model.

[0241] In one implementation, the time-domain level track includes a frame rate indication field for indicating the frame rate of the time-domain level track.

[0242] According to one embodiment of this application, Figure 2 The data processing method for point cloud media shown can be implemented by [the relevant authority / organization]. Figure 4 The data processing is performed by individual units within the point cloud media data processing device shown. For example, Figure 2 Step S201 shown can be performed by Figure 4 The acquisition unit 401 shown is executed, and step S202 can be performed by... Figure 4 The processing unit 402 shown executes. Figure 4 The data processing device for point cloud media shown can be composed of individual or combined units into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the data processing device for point cloud media may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0243] According to another embodiment of this application, the following can be executed by running on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 2 The computer program (including program code) involved in each step of the corresponding method shown, to construct such... Figure 4 The data processing apparatus for point cloud media shown herein, and the data processing method for point cloud media for implementing embodiments of this application, are described. A computer program may be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the computer-readable recording medium, and executed therein.

[0244] In this embodiment, a media file of point cloud media is obtained. The media file includes a temporal hierarchical track, which is obtained by encapsulating the point cloud frames of the point cloud media based on temporal information. The temporal information includes temporal hierarchical information and frame rate information. Any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media. The media file is decoded to present the point cloud media. It is evident that encapsulating point cloud media files based on temporal information (such as temporal hierarchical information or frame rate information) enriches the encapsulation methods of point cloud media. Content consumption devices can flexibly select the required point cloud media files for transmission and decoding consumption according to their needs.

[0245] Please see Figure 5 , Figure 5 This is a schematic diagram of another point cloud media data processing device provided in an embodiment of this application; the point cloud media data processing device can be a computer program (including program code) running in a content production device, for example, the point cloud media data processing device can be application software in the content production device. Figure 5 As shown, the data processing device for point cloud media includes an acquisition unit 501 and a processing unit 502. Please refer to... Figure 5 The detailed descriptions of each unit are as follows:

[0246] Acquisition unit 501 is used to acquire point cloud frames of point cloud media;

[0247] The processing unit 502 is used to encapsulate the point cloud frame into a time-domain hierarchical track according to the time-domain information of the point cloud frame to obtain the media file of the point cloud media; the time-domain information includes time-domain hierarchical information and frame rate information; any sample in the time-domain hierarchical track contains the complete data of at least one point cloud frame of the point cloud media.

[0248] In one implementation, the number of time-domain level tracks is M, the M time-domain level tracks belong to P time-domain levels, M is a positive integer, and P is a positive integer less than or equal to M; the processing unit 502 is further configured to:

[0249] Slicing a media file to obtain multiple media segments; and,

[0250] The transmission signaling file that generates the media file includes at least one adaptive set, and M time-domain level tracks are represented as P representations under at least one adaptive set according to the time-domain level.

[0251] In one implementation, the media file is encapsulated in a single-track encapsulation mode;

[0252] The transmission signaling file includes at least one adaptive set, which includes a representation of the track at the time-domain level and a representation of the track in the single-track encapsulation mode.

[0253] In one implementation, the representation corresponding to the time-domain hierarchical orbit carries a time-domain hierarchical descriptor;

[0254] In single-track encapsulation mode, the track representation does not carry time-domain hierarchical descriptors.

[0255] In one implementation, both the representation of the time-domain level track and the representation of the track in the single-track encapsulation mode carry an encapsulation mode descriptor.

[0256] The value of the encapsulation mode descriptor carried by the representation corresponding to the time-domain level track is different from the value of the encapsulation mode descriptor carried by the representation corresponding to the track in the single-track encapsulation mode.

[0257] In one implementation, the transmission signaling file further includes a preselection set, which includes preselected elements, and the preselected elements include preselected component attributes and codec attributes;

[0258] Preselected component attributes are used to indicate the identifier of one or more representations in at least one adaptive set;

[0259] The codec attribute is used to indicate the sample entry type of the corresponding track as indicated by the preselected component attribute.

[0260] In one implementation, if the sample entry type of the corresponding track indicated by the preselected component attribute is a target type, then the processing unit 502 is used to generate a transmission signaling file for the media file, specifically for:

[0261] Configure the value of the codec attribute to the identifier corresponding to the target type.

[0262] In one implementation, the media file is encapsulated in a multitrack encapsulation mode; the transmission signaling file includes at least one adaptive set and a preselected set, the preselected set including preselected elements, the preselected elements including preselected component attributes and codec attributes;

[0263] The preselected component attribute is used to indicate the identifier of at least one adaptive set;

[0264] The codec attribute is used to indicate the sample entry type of the track in the adaptive set indicated by the preselected component attribute.

[0265] In one implementation, if the sample entry type of the corresponding track indicated by the preselected component attribute is a target type, then the processing unit 502 is used to generate a transmission signaling file for the media file, specifically for:

[0266] Configure the value of the codec attribute to the identifier corresponding to the target type.

[0267] According to one embodiment of this application, Figure 3 The data processing method for point cloud media shown can be implemented by [the relevant authority / organization]. Figure 5 The data processing is performed by individual units within the point cloud media data processing device shown. For example, Figure 3 Step S301 shown can be performed by Figure 5 The acquisition unit 501 shown is executed, and step S302 can be performed by... Figure 5 The processing unit 502 shown is executed. Figure 5 The data processing device for point cloud media shown can be composed of individual or combined units into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the data processing device for point cloud media may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0268] According to another embodiment of this application, the following can be executed by running on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 3 The computer program (including program code) involved in each step of the corresponding method shown, to construct such... Figure 5 The data processing apparatus for point cloud media shown herein, and the data processing method for point cloud media for implementing embodiments of this application, are described. A computer program may be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the computer-readable recording medium, and executed therein.

[0269] In this embodiment, point cloud frames of point cloud media are obtained; based on the temporal information of the point cloud frames, they are encapsulated into temporal hierarchical tracks to obtain the media file of the point cloud media. The temporal information includes temporal hierarchical information and frame rate information; any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media. It is evident that encapsulating point cloud media files using temporal information (such as temporal hierarchical information and frame rate information) as the core enriches the encapsulation methods of point cloud media, allowing content consumption devices to flexibly select the required point cloud media files for transmission and decoding consumption according to their needs.

[0270] Figure 6 This is a schematic diagram of the structure of a content consumption device provided in an embodiment of this application; the content consumption device can be a computer device used by a user of Point Cloud Media, and the computer device can be a terminal (such as a PC, a smart mobile device (such as a smartphone), a VR device (such as a VR headset, VR glasses, etc.)). Figure 6 As shown, the content consumption device includes a receiver 601, a processor 602, a memory 603, and a display / playback device 604. Wherein:

[0271] Receiver 601 is used to enable decoding and transmission interaction with other devices, specifically for the transmission of point cloud media between the content production device and the content consumption device. That is, the content consumption device receives the relevant media resources of the point cloud media transmitted by the content production device through receiver 601.

[0272] Processor 602 (or CPU (Central Processing Unit)) is the processing core of the content production device. Processor 602 is adapted to implement one or more program instructions, specifically to load and execute one or more program instructions to achieve... Figure 2 The flowchart illustrates the data processing method for point cloud media.

[0273] Memory 603 is a memory device in the content consumption device used to store programs and media resources. It is understood that memory 603 here can include the built-in storage medium of the content consumption device, or it can include extended storage media supported by the content consumption device. It should be noted that memory 603 can be high-speed RAM or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one memory located remotely from the aforementioned processor. Memory 603 provides storage space for storing the operating system of the content consumption device. Furthermore, this storage space is also used to store computer programs, which include program instructions adapted to be called and executed by the processor to perform the various steps of the point cloud media data processing method. In addition, memory 603 can also be used to store a 3D image of the point cloud media formed after processor processing, the audio content corresponding to the 3D image, and information required for rendering the 3D image and audio content.

[0274] Display / playback device 604 is used to output rendered sound and 3D images.

[0275] Please see again Figure 6 The processor 602 may include a parser 621, a decoder 622, a converter 623, and a renderer 624; wherein:

[0276] The parser 621 is used to depackage and encapsulate the encapsulated files of the rendering media from the content production device. Specifically, it depackages the media file resources according to the file format requirements of point cloud media to obtain audio and video streams, and provides the audio and video streams to the decoder 622.

[0277] Decoder 622 decodes the audio stream to obtain audio content, which is then provided to the renderer for audio rendering. Additionally, decoder 622 decodes the video stream to obtain a 2D image. Based on the metadata provided by the media presentation description information, if the metadata indicates that the point cloud media has undergone a region encapsulation process, the 2D image refers to an encapsulated image; if the metadata indicates that the point cloud media has not undergone a region encapsulation process, the planar image refers to a projected image.

[0278] Converter 623 is used to convert 2D images into 3D images. If the point cloud media has undergone a region encapsulation process, converter 623 will first decapsulate the encapsulated image to obtain a projected image. Then, the projected image will be reconstructed to obtain a 3D image. If the rendering media has not undergone a region encapsulation process, converter 623 will directly reconstruct the projected image to obtain a 3D image.

[0279] Renderer 624 is used to render the audio content and 3D images of point cloud media. Specifically, it renders the audio content and 3D images based on the metadata related to rendering and viewport in the media presentation description information, and then outputs the rendered content to the display / playback device.

[0280] In one exemplary embodiment, the processor 602 (specifically, the devices included in the processor) executes instructions by calling one or more instructions stored in memory. Figure 2 The steps of the data processing method for point cloud media are shown. Specifically, the memory stores one or more first instructions, which are adapted to be loaded by the processor 602 and executed in the following steps:

[0281] Obtain the media file of the point cloud media. The media file includes a temporal hierarchical track. The temporal hierarchical track is obtained by encapsulating the point cloud frames of the point cloud media based on temporal information. The temporal information includes temporal hierarchical information and frame rate information. Any sample in the temporal hierarchical track contains the complete data of at least one point cloud frame of the point cloud media.

[0282] Decode media files to present point cloud media.

[0283] In one implementation, the sample entry point of the temporal hierarchical orbit uses a point cloud compressed sample entry point based on a geometric model; and the sample entry point type of the temporal hierarchical orbit is a target type.

[0284] Among them, the point cloud compression sample entry based on the geometric model is obtained by expanding the volumetric visual sample entry.

[0285] In one implementation, the time-domain hierarchical orbit includes N samples, where N is a positive integer;

[0286] Each of the N samples is a subset of point cloud frames of the point cloud media at a temporal level.

[0287] In one implementation, the complete data of a point cloud frame consists of one or more coded content units, which are used to store the geometric coded data or attribute coded data of the point cloud frame.

[0288] In one implementation, the time-domain hierarchical orbit includes a first orbit and a second orbit;

[0289] The samples in the first track correspond to the first time domain level, and the samples in the second track correspond to the second time domain level.

[0290] In one implementation, the decoding time of samples in the first track is different from that of samples in the second track.

[0291] In one implementation, the first time domain level is higher than the second time domain level, and the first track is indexed to the second track.

[0292] In one embodiment, the media file further includes a sample description data box, which includes one or more type description fields for a sample entry type, the type description fields being used to indicate the sample entry type.

[0293] In one embodiment, the computer program in memory 603 is loaded by processor 602 and further performs the following steps:

[0294] Obtain the transmission signaling file of the point cloud media. The transmission signaling file includes at least one adaptive set, and the at least one adaptive set includes the representation of the time-domain level track.

[0295] In one embodiment, the processor 602 acquires the media file of the point cloud media as follows:

[0296] Based on the description information of at least one adaptive set, determine the media files required to render point cloud media;

[0297] The media files of the determined point cloud media are retrieved using streaming transmission.

[0298] In one implementation, the point cloud media includes multiple media segments;

[0299] If a point cloud frame encapsulates at least two media segments in a time-domain level track, then the at least two media segments include an initialization media segment.

[0300] The initial media segment includes a configuration record for a point cloud compressor decoder based on a geometric model.

[0301] In one embodiment, the time-domain level track includes a frame rate indication field, which is used to indicate the frame rate of the time-domain level track.

[0302] In this embodiment, a media file of point cloud media is obtained. The media file includes a temporal hierarchical track, which is obtained by encapsulating the point cloud frames of the point cloud media based on temporal information. The temporal information includes temporal hierarchical information and frame rate information. Any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media. The media file is decoded to present the point cloud media. It is evident that encapsulating point cloud media files based on temporal information (such as temporal hierarchical information or frame rate information) enriches the encapsulation methods of point cloud media. Content consumption devices can flexibly select the required point cloud media files for transmission and decoding consumption according to their needs.

[0303] Figure 7 This is a schematic diagram of a content creation device provided in an embodiment of this application; the content creation device may be a computer device used by a provider of point cloud media, which may be a terminal (such as a PC, a smart mobile device (such as a smartphone) or a server. Figure 7 As shown, the content creation device includes a capture device 701, a processor 702, a memory 703, and a transmitter 704. Wherein:

[0304] The capture device 701 is used to acquire raw data (including audio and video content synchronized in time and space) of point cloud media from real-world sound-visual scenes. The capture device 701 may include, but is not limited to, audio devices, camera devices, and sensing devices. Audio devices may include audio sensors, microphones, etc. Camera devices may include ordinary cameras, stereo cameras, light field cameras, etc. Sensing devices may include laser devices, radar devices, etc.

[0305] Processor 702 (or CPU (Central Processing Unit)) is the processing core of the content production device. Processor 702 is adapted to implement one or more program instructions, specifically to load and execute one or more program instructions to achieve... Figure 3 The flowchart illustrates the data processing method for point cloud media.

[0306] Memory 703 is a memory device in the content creation apparatus used to store programs and media resources. It is understood that memory 703 here can include both the built-in storage medium of the content creation apparatus and extended storage media supported by the content creation apparatus. It should be noted that the memory can be high-speed RAM or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one memory located remotely from the aforementioned processor. The memory provides storage space for storing the operating system of the content creation apparatus. Furthermore, this storage space is also used to store computer programs, which include program instructions adapted to be called and executed by the processor to perform the various steps of the point cloud media data processing method. In addition, memory 703 can also be used to store point cloud media files formed after processing by the processor, which include media file resources and media presentation description information.

[0307] The transmitter 704 is used to enable transmission and interaction between the content creation device and other devices, specifically to facilitate the transmission of point cloud media between the content creation device and the content playback device. That is, the content creation device uses the transmitter 704 to transmit relevant media resources of the point cloud media to the content playback device.

[0308] Please see again Figure 7 The processor 702 may include a converter 721, an encoder 722, and a packager 723; wherein:

[0309] Converter 721 performs a series of conversion processes on captured video content to make it suitable for video encoding of point cloud media. The conversion processes may include stitching and projection; optionally, they may also include region encapsulation. Converter 721 can convert captured 3D video content into 2D images and provide them to the encoder for video encoding.

[0310] Encoder 722 is used to encode the captured audio content to form an audio bitstream of point cloud media. It is also used to encode the 2D image obtained by converter 721 to obtain a video bitstream.

[0311] The encapsulator 723 encapsulates audio and video streams into a file container according to the point cloud media file format (such as ISOBMFF) to form a point cloud media file resource. This media file resource can be a media file or a media segment forming a point cloud media file. It also records the metadata of the point cloud media file resource using media presentation description information according to the point cloud media file format requirements. The encapsulated point cloud media file obtained by the encapsulator is stored in memory and provided to the content playback device as needed for point cloud media presentation.

[0312] The processor 702 (specifically, the various components within the processor) executes instructions by calling one or more instructions from memory. Figure 4 The steps of the data processing method for point cloud media are shown. Specifically, the memory 703 stores one or more first instructions, which are adapted to be loaded by the processor 702 and executed in the following steps:

[0313] Acquire point cloud frames from point cloud media;

[0314] Based on the temporal information of the point cloud frame, the point cloud frame is encapsulated into a temporal hierarchical track to obtain the media file of the point cloud media; the temporal information includes temporal hierarchical information and frame rate information; any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media.

[0315] In one embodiment, the number of time-domain level tracks is M, the M time-domain level tracks belong to P time-domain levels, M is a positive integer, and P is a positive integer less than or equal to M; the computer program in memory 703 is loaded by processor 702 and also performs the following steps:

[0316] Slicing a media file to obtain multiple media segments; and,

[0317] The transmission signaling file that generates the media file includes at least one adaptive set, and M time-domain level tracks are represented as P representations under at least one adaptive set according to the time-domain level.

[0318] In one implementation, the media file is encapsulated in a single-track encapsulation mode;

[0319] The transmission signaling file includes at least one adaptive set, which includes a representation of the track at the time-domain level and a representation of the track in the single-track encapsulation mode.

[0320] In one implementation, the representation corresponding to the time-domain level orbit carries a time-domain level-related descriptor;

[0321] In single-track encapsulation mode, the track representation does not carry time-domain hierarchical descriptors.

[0322] In one implementation, both the representation of the time-domain level track and the representation of the track in the single-track encapsulation mode carry an encapsulation mode descriptor.

[0323] The value of the encapsulation mode descriptor carried by the representation corresponding to the time-domain level track is different from the value of the encapsulation mode descriptor carried by the representation corresponding to the track in the single-track encapsulation mode.

[0324] In one embodiment, the transmission signaling file further includes a preselection set, which includes preselected elements, and the preselected elements include preselected component attributes and codec attributes;

[0325] Preselected component attributes are used to indicate the identifier of one or more representations in at least one adaptive set;

[0326] The codec attribute is used to indicate the sample entry type of the corresponding track as indicated by the preselected component attribute.

[0327] In one implementation, if the sample entry type indicated by the preselected component attribute representing the corresponding track is a target type, then the specific implementation method for processing the transmission signaling file for generating the media file in step 702 is as follows:

[0328] Configure the value of the codec attribute to the identifier corresponding to the target type.

[0329] In one embodiment, the media file is encapsulated in a multitrack encapsulation mode; the transmission signaling file includes at least one adaptive set and a preselected set, the preselected set including preselected elements, and the preselected elements including preselected component attributes and codec attributes;

[0330] The preselected component attribute is used to indicate the identifier of at least one adaptive set;

[0331] The codec attribute is used to indicate the sample entry type of the track in the adaptive set indicated by the preselected component attribute.

[0332] In one implementation, if the sample entry type of the corresponding track in the adaptive set indicated by the preselected component attribute is the target type, then the specific implementation of the processor 702 generating the transmission signaling file for the media file is as follows:

[0333] Configure the value of the codec attribute to the identifier corresponding to the target type.

[0334] In this embodiment, point cloud frames of point cloud media are obtained; based on the temporal information of the point cloud frames, they are encapsulated into temporal hierarchical tracks to obtain the media file of the point cloud media. The temporal information includes temporal hierarchical information and frame rate information; any sample in the temporal hierarchical track contains complete data of at least one point cloud frame of the point cloud media. It is evident that encapsulating point cloud media files using temporal information (such as temporal hierarchical information and frame rate information) as the core enriches the encapsulation methods of point cloud media, allowing content consumption devices to flexibly select the required point cloud media files for transmission and decoding consumption according to their needs.

[0335] This application also provides a computer-readable storage medium storing one or more instructions, which are adapted to be loaded by a processor and executed by the data processing method for point cloud media described in the above method embodiments.

[0336] This application also provides a computer program product containing instructions that, when run on a computer, causes the computer to execute the point cloud media data processing method described in the above method embodiments.

[0337] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned point cloud media data processing method.

[0338] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.

[0339] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0340] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0341] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art will understand that all or part of the processes for implementing the above embodiments and equivalent variations made in accordance with the claims of this application are still within the scope of this application.< / preselection> < / adaptationset> < / adaptationset>

Claims

1. A data processing method for point cloud media, characterized in that, The method includes: The media file of the point cloud media is obtained. The media file is encapsulated in a multi-track encapsulation mode. The media file includes a time-domain level track. The time-domain level track is obtained by encapsulating the point cloud frames of the point cloud media based on time-domain information. The time-domain information includes time-domain level information and frame rate information. Any sample in the time-domain level track contains the complete data of at least one point cloud frame of the point cloud media. The media file is decoded to present the point cloud media.

2. The method as described in claim 1, characterized in that, The sample entry point of the time-domain hierarchical track uses a point cloud compressed sample entry point based on a geometric model; and the sample entry point type of the time-domain hierarchical track is a target type; The point cloud compression sample entry based on the geometric model is obtained by expanding the volumetric visual sample entry.

3. The method as described in claim 1, characterized in that, The time-domain hierarchical orbit includes N samples, where N is a positive integer; Each of the N samples is a subset of the point cloud frames of the point cloud media at a temporal level.

4. The method as described in claim 3, characterized in that, A complete point cloud frame consists of one or more coded content units, which are used to store the geometric coded data or attribute coded data of the point cloud frame.

5. The method as described in claim 1, characterized in that, The time-domain hierarchical orbit includes a first orbit and a second orbit; The samples in the first orbit correspond to the first time domain level, and the samples in the second orbit correspond to the second time domain level.

6. The method as described in claim 5, characterized in that, The decoding time of the samples in the first track is different from that of the samples in the second track.

7. The method as described in claim 5, characterized in that, The first time domain level is higher than the second time domain level, and the first track is indexed to the second track.

8. The method as described in claim 1, characterized in that, The media file also includes a sample description data box, which includes one or more type description fields for a sample entry type, the type description fields being used to indicate the sample entry type.

9. The method as described in claim 1, characterized in that, The method further includes: Obtain the transmission signaling file of the point cloud media, the transmission signaling file including at least one adaptive set, the at least one adaptive set including the representation corresponding to the time-domain level track.

10. The method as described in claim 9, characterized in that, The media files obtained from the point cloud media include: Based on the description information of the at least one adaptive set, determine the media file required to render the point cloud media; The media files of the determined point cloud media are retrieved using streaming transmission.

11. The method as described in claim 1, characterized in that, The point cloud media includes multiple media segments; If the time-domain level track encapsulates point cloud frames of at least two media segments, then the at least two media segments include an initialization media segment; The initial media segment includes a point cloud compression decoder configuration record based on a geometric model.

12. The method as described in claim 1, characterized in that, The time-domain level track includes a frame rate indicator field, which is used to indicate the frame rate of the time-domain level track.

13. A data processing method for point cloud media, characterized in that, The method includes: Acquire point cloud frames from point cloud media; Based on the temporal information of the point cloud frame, the point cloud frame is encapsulated into a temporal hierarchical track to obtain the media file of the point cloud media; the encapsulation mode of the media file is a multi-track encapsulation mode, and the temporal information includes temporal hierarchical information and frame rate information; any sample in the temporal hierarchical track contains the complete data of at least one point cloud frame of the point cloud media.

14. The method as described in claim 13, characterized in that, The number of time-domain level tracks is M, and the M time-domain level tracks belong to P time-domain levels, where M is a positive integer and P is a positive integer less than or equal to M; the method further includes: The media file is sliced ​​to obtain multiple media segments; and, A transmission signaling file for the media file is generated, the transmission signaling file including at least one adaptive set, wherein the M time-domain level tracks are represented as P representations under the at least one adaptive set according to the time-domain level.

15. The method as described in claim 14, characterized in that, The point cloud media also includes other media files with a single-track encapsulation mode; The transmission signaling file includes at least one adaptive set, which includes the representation corresponding to the time-domain level track and the representation corresponding to the track in the single-track encapsulation mode.

16. The method as described in claim 15, characterized in that, The representation corresponding to the time-domain hierarchical orbit carries a time-domain hierarchical related descriptor; In the single-track encapsulation mode, the track representation does not carry time-domain hierarchical descriptors.

17. The method as described in claim 15, characterized in that, Both the representation of the time-domain hierarchical track and the representation of the track in the single-track encapsulation mode carry an encapsulation mode descriptor; The value of the encapsulation mode descriptor corresponding to the time-domain hierarchical track is different from the value of the encapsulation mode descriptor corresponding to the track in the single-track encapsulation mode.

18. The method as described in claim 15, characterized in that, The transmission signaling file also includes a preselection set, which includes preselected elements, and the preselected elements include preselected component attributes and codec attributes; The preselected component attributes are used to indicate the identifier of one or more representations in the at least one adaptive set; The codec attribute is used to indicate the sample entry type of the corresponding track indicated by the preselected component attribute.

19. The method as described in claim 18, characterized in that, If the sample entry type indicated by the preselected component attribute for the corresponding track is a target type, then the transmission signaling file for generating the media file includes: Configure the value of the codec attribute to the identifier corresponding to the target type.

20. The method as described in claim 14, characterized in that, The transmission signaling file includes at least one adaptive set and a preselected set, the preselected set including preselected elements, the preselected elements including preselected component attributes and codec attributes; The preselected component attribute is used to indicate the identifier of the at least one adaptive set; The codec attribute is used to indicate the sample entry type of the track in the adaptive set indicated by the preselected component attribute.

21. The method as described in claim 20, characterized in that, If the sample entry type of the track corresponding to the representation in the adaptive set indicated by the preselected component attribute is the target type, then the transmission signaling file for generating the media file includes: Configure the value of the codec attribute to the identifier corresponding to the target type.

22. A data processing device for point cloud media, characterized in that, The data processing device for the point cloud media includes: The acquisition unit is used to acquire the media file of the point cloud media. The media file is encapsulated in a multi-track encapsulation mode. The media file includes a temporal hierarchical track. The temporal hierarchical track is obtained by encapsulating the point cloud frames of the point cloud media based on temporal information. The temporal information includes temporal hierarchical information and frame rate information. Any sample in the temporal hierarchical track contains the complete data of at least one point cloud frame of the point cloud media. The processing unit is used to decode the media file to present the point cloud media.

23. A data processing device for point cloud media, characterized in that, The data processing device for the point cloud media includes: The acquisition unit is used to acquire point cloud frames of point cloud media; The processing unit is configured to encapsulate the point cloud frame into a time-domain hierarchical track based on the time-domain information of the point cloud frame to obtain a media file of the point cloud media; the encapsulation mode of the media file is a multi-track encapsulation mode, and the time-domain information includes time-domain hierarchical information and frame rate information; any sample in the time-domain hierarchical track contains complete data of at least one point cloud frame of the point cloud media.

24. A computer device, characterized in that, include: Storage devices and processors; A memory, wherein a computer program is stored; A processor is configured to load the computer program to implement the data processing method for point cloud media as described in any one of claims 1-12; or, to load the computer program to implement the data processing method for point cloud media as described in claims 13-21.

25. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and executed as described in any one of claims 1-12; or, loaded and executed as described in claims 13-21.

26. A computer program product, characterized in that, The computing program product includes a computer program adapted to be loaded by a processor and execute the data processing method for point cloud media as described in any one of claims 1-12; or, to load and execute the data processing method for point cloud media as described in claims 13-21.