Point cloud data playing method and device, point cloud data packaging method and device and storage medium

By adding the decoding time of the fused frame during the compression process of the point cloud data, the problem of point cloud data playback jamming is solved, and timely decoding and smooth playback experience are achieved.

CN120640036APending Publication Date: 2025-09-12ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410275726.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When playing point cloud data, the decoding time of the fused frame cannot be obtained in advance, resulting in playback lag.

Method used

During the compression process of point cloud data, the decoding time of the fused frame is added to the media file or description information, so that the decoding time of the fused frame can be obtained before playback and the fused frame can be decoded in advance.

Benefits of technology

Ensure that the point cloud frames in the fused frames are decoded in a timely manner to avoid playback freezes, improve file parsing efficiency, and enhance user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640036A_ABST
    Figure CN120640036A_ABST
Patent Text Reader

Abstract

The invention provides a playing and packaging method and device of point cloud data and a storage medium, relates to the technical field of video processing, is used for solving the problem that the point cloud data is stuck during playing, and comprises the steps that a fusion frame of the point cloud data and time information of the fusion frame are acquired, and the time information of the fusion frame comprises decoding time of the fusion frame. And decoding the fusion frame based on the decoding time of the fusion frame to play the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a method, device, and storage medium for playing and packaging point cloud data. Background Art

[0002] Video compression technology aims to reduce the amount of video data, thereby lowering the storage space required by electronic devices (such as servers and terminals) and network transmission bandwidth. For example, a server compresses point cloud data composed of multiple point cloud frames and then transmits or stores the compressed point cloud data.

[0003] Currently, when compressing point cloud data, the server can fuse multiple consecutive and similar point cloud frames into a single fused frame, then encapsulate the fused frame with the remaining point cloud frames to create a media file of the point cloud data. Furthermore, each point cloud frame carries a decoding time. When playing point cloud data, the server can decode the point cloud frames based on the decoding time, ensuring that the point cloud frames are decoded before display, thus ensuring smooth playback of the point cloud data.

[0004] However, the decoding time of each point cloud frame is encapsulated in the fused frame. Therefore, when playing the fused frame of point cloud data, the point cloud frames in the fused frame cannot be decoded in advance, resulting in lag in the playback of point cloud data. Summary of the Invention

[0005] The embodiments of the present disclosure provide a method, device, and storage medium for playing and packaging point cloud data, which are used to solve the problem of lag during the playback of point cloud data.

[0006] In one aspect, a method for playing point cloud data is provided. The method includes obtaining a fused frame of the point cloud data and time information of the fused frame, wherein the time information of the fused frame includes a decoding time of the fused frame, and decoding the fused frame based on the decoding time of the fused frame to play the point cloud data.

[0007] In another aspect, a method for packaging point cloud data is provided. The method includes obtaining fused frame information of the point cloud data, the fused frame information including time information of the fused frame in the point cloud data, the fused frame time information including the decoding time of the fused frame. Based on the fused frame information and the point cloud data, a media file of the point cloud data is obtained, the media file including the fused frame and the time information of the fused frame.

[0008] On the other hand, a device for playing point cloud data is provided, which includes: an acquisition module and a processing module.

[0009] The acquisition module is used to obtain the fused frame of the point cloud data and the time information of the fused frame, wherein the time information of the fused frame includes the decoding time of the fused frame. The processing module is used to decode the fused frame based on the decoding time of the fused frame to play the point cloud data.

[0010] On the other hand, a point cloud data packaging device is provided, which includes: an acquisition module and a processing module.

[0011] The acquisition module is used to obtain fused frame information of the point cloud data. The fused frame information includes the time information of the fused frame in the point cloud data, and the time information of the fused frame includes the decoding time of the fused frame. The processing module is used to obtain a media file of the point cloud data based on the fused frame information and the point cloud data. The media file includes the fused frame and the time information of the fused frame.

[0012] In yet another aspect, a network device is provided, comprising: a memory and a processor. The memory and the processor are coupled. The memory is configured to store a computer program. When the processor executes the computer program, the method for playing and packaging point cloud data according to any of the above embodiments is implemented.

[0013] On the other hand, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the point cloud data playback and packaging method of any of the above embodiments is implemented.

[0014] On the other hand, a computer program product is provided, which includes computer program instructions, and when the computer program instructions are executed by a processor, the method for playing and packaging point cloud data of any of the above embodiments is implemented.

[0015] In an embodiment of the present disclosure, before compressing the point cloud data, the decoding time of the fused frame in the point cloud data is obtained so that the decoding time of the fused frame can be added to the media file of the point cloud data during the compression process of the point cloud data. When the point cloud data is played, the decoding time of the fused frame can be obtained by decapsulating the media file of the point cloud data, and then the fused frame is decoded based on the decoding time of the fused frame to obtain the point cloud frame in the fused frame, thereby ensuring that the point cloud frame in the fused frame can be decoded in time to avoid freezes when playing the point cloud data. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present disclosure, the following briefly introduces the drawings required for use in some embodiments of the present disclosure. Obviously, the drawings described below are only drawings of some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0017] Figure 1A flowchart of end-to-end processing of a point cloud provided in some embodiments of the present disclosure;

[0018] Figure 2 A schematic diagram of a communication system provided for some embodiments of the present disclosure;

[0019] Figure 3 A schematic flow chart of a method for packaging point cloud data provided in some embodiments of the present disclosure;

[0020] Figure 4 A schematic diagram of a point cloud single-track storage structure provided in some embodiments of the present disclosure;

[0021] Figure 5 A schematic diagram of a point cloud multi-track storage structure provided in some embodiments of the present disclosure;

[0022] Figure 6 A schematic diagram of a block-based track storage structure of a point cloud provided in some embodiments of the present disclosure;

[0023] Figure 7 A schematic flow chart of a method for playing point cloud data provided in some embodiments of the present disclosure;

[0024] Figure 8 A schematic diagram of a point cloud description information structure provided in some embodiments of the present disclosure;

[0025] Figure 9 A schematic structural diagram of a point cloud data playback device provided in some embodiments of the present disclosure;

[0026] Figure 10 A schematic structural diagram of a point cloud data packaging device provided in some embodiments of the present disclosure;

[0027] Figure 11 A schematic structural diagram of a point cloud data playback device provided in some embodiments of the present disclosure. DETAILED DESCRIPTION

[0028] The following will clearly and completely describe the technical solutions of this disclosure in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this disclosure, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of this disclosure without making any creative efforts shall fall within the scope of protection of this disclosure.

[0029] It should be noted that in this disclosure, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this disclosure as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0030] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the quantity of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features.

[0031] In the description of this disclosure, unless otherwise specified, " / " means "or." For example, A / B can mean A or B. "And / or" in this document simply describes an association relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exists simultaneously, and B exists alone. Furthermore, "at least one" means one or more, and "a plurality" means two or more.

[0032] Before introducing in detail the method for playing point cloud data provided by the embodiment of the present disclosure, the implementation environment and application scenarios of the embodiment of the present disclosure are first introduced.

[0033] First, the application scenarios of the embodiments of the present disclosure are introduced.

[0034] A 3D point cloud is a dataset consisting of a large number of 3D points, typically used to represent the surface of an object or scene in a 3D space in the real world. The application of 3D point clouds is very broad and covers many different fields. Typical application scenarios include:

[0035] 1. Computer Vision and Graphics: 3D point clouds are widely used in computer vision and graphics for tasks such as target detection, object recognition, and scene segmentation. They can be acquired through methods such as laser scanning and photogrammetry and used to model real-world objects and environments.

[0036] 2. Autonomous driving and robotics: In the field of autonomous vehicles and robotics, three-dimensional point clouds are used for environmental perception and obstacle detection.

[0037] 3. Architecture and urban planning: 3D point clouds can be used in architecture and urban planning to obtain detailed structures of buildings and urban environments through laser scanning or drones to support planning and design work.

[0038] 4. Cultural heritage protection: In the field of cultural heritage, 3D point clouds can be used to digitize and protect ancient buildings, sculptures, and other cultural heritage, and promote the protection and restoration of cultural relics by providing high-precision spatial information.

[0039] 5. Medical image processing: Three-dimensional point clouds are widely used in the field of medical imaging, such as three-dimensional modeling of patient organs to support medical diagnosis and surgical planning.

[0040] 6. Virtual Reality and Augmented Reality: 3D point clouds are used to create realistic virtual reality (VR) and augmented reality (AR) experiences. By capturing 3D information about the real world, virtual and real environments can be more naturally integrated.

[0041] 7. Industrial Manufacturing: In the industrial field, 3D point clouds can be used for quality control, product design, and process optimization. By scanning objects in the manufacturing process, real-time monitoring and analysis can be performed.

[0042] Currently, 3D point clouds are widely used in application scenarios such as autonomous driving, real-time inspection, cultural heritage, and six-degrees-of-freedom (6DoF) immersive real-time communication. Figure 1 A flowchart for end-to-end processing of point clouds in scenarios perceived by the human eye (such as cultural heritage and 6DoF immersive real-time communication) is presented, i.e., an architecture diagram based on the "Information Technology - Efficient Graphics Data Coding Part 1: System"). For machine perception scenarios (such as autonomous driving and real-time inspections), the difference is that the terminal does not need to render and display the decoded point cloud. The system flow before this is consistent with the human eye perception scenario.

[0043] like Figure 1 As shown, a real-world visual scene A is captured by a set of cameras or a camera device with multiple lenses and sensors, resulting in point cloud source data B, which is a sequence of frames consisting of a large number of point cloud frames. The point cloud encoder encodes one or more point cloud frames into a point cloud codestream E, and the file encapsulator encapsulates the file / segment according to a specific media container file format (such as ISOBMFF). If transmission is required, the file encapsulator encapsulates the point cloud codestream into media segments and provides an index file (i.e., media presentation description information). The server uses a transmission mechanism to transmit the segments Fs in the media file to the player. On the player side, the file decapsulator decapsulates the received file, extracts the encoded bitstream E', and parses the metadata. The point cloud decoder then decodes the bitstream to generate the point cloud D'. The renderer renders and displays the point cloud on the display device based on the current viewing position, viewing direction, or a viewport determined by various types of sensors (such as head, position, or eye tracking sensors).

[0044] Optionally, if local playback is required, the file encapsulator transmits the encapsulated point cloud stream E to the local player, and the point cloud is rendered and displayed on the display device through the above-mentioned file decapsulator, point cloud decoder and renderer.

[0045] However, due to the large volume of point clouds, processing and storing this data requires a lot of computing and storage resources. Therefore, in order to effectively transmit and store point clouds, it is usually necessary to apply multiple compression algorithms to reduce the data volume during the point cloud compression process. Depending on the characteristics of point clouds and application scenarios, the requirements for computing and storage networks can be divided into three categories:

[0046] 1. Static objects and scenes represented by static dense point clouds. Application scenarios include digital cultural heritage and digital twins.

[0047] 2. Objects represented by dynamic point clouds (dynamic objects), application scenarios include free viewpoint broadcasting, 3D immersive interaction, etc.

[0048] 3. Dynamic acquisition of objects represented by point clouds (dynamic acquisition). Application scenarios include high-precision maps, navigation, and autonomous driving.

[0049] The first and third types of point clouds are usually compressed using geometry-based point cloud compression (G-PCC), which is also one of the hot technologies explored by various standards organizations in the field of point cloud compression. The main idea of ​​geometry-based point cloud compression technology is to achieve efficient compression by encoding the geometric information in the point cloud. This includes encoding geometric attributes such as point coordinates, normal vectors, and colors. Generally, these technologies take into account the local characteristics and geometric structure of the point cloud to more efficiently represent the data.

[0050] Specifically, for dynamically acquired point clouds where scene changes are less noticeable, where the majority of the entire point cloud frame remains unchanged over time, while a small portion of the frame is updated over time, a fused frame compression method is often used. This method involves fusing multiple similar point cloud frames into a single frame before applying G-PCC compression technology, and then compressing and encoding the fused frame. The fused frame includes additional attribute data corresponding to the temporal information of each point in the combined frame.

[0051] The advantage of fused frames is that they can effectively reduce the amount of point cloud information in these scenarios. However, when playing or rendering on a terminal, the time information of the fused frames must be fully decoded before it can be calculated frame by frame. During playback, the terminal may involve various caching mechanisms and user operations, such as downloading and playing simultaneously, random scrolling, fast forwarding, and rewinding. These operations rely on the terminal player to obtain the decoding and playback time of each frame of the point cloud before decoding. Typically, traditional Moving Picture Experts Group (MPEG)-2 transport stream (TS) or International Organization for Standardization Base Media File Format (ISOBMFF) encapsulation tools can efficiently parse the playback and decoding time of each frame during transmission and playback.

[0052] However, when there are fused frames in the point cloud compression data, its internal time information cannot be obtained in advance in these encapsulation structures, resulting in low terminal playback flexibility and lag in the point cloud data during playback.

[0053] In order to solve the above problems, the embodiments of the present disclosure provide a method for playing and encapsulating point cloud data, which is applied to scenarios of media transmission or local playback of point cloud data including fused frames. In the case where there are fused frames in the point cloud data, the decoding time of the fused frames in the point cloud data is added to the media file of the point cloud data, or the decoding time of the fused frames in the point cloud data is added to the description information of the point cloud data, so as to obtain the decoding time of the fused frames before the playback of the point cloud data, and then the fused frames are decoded in advance during the playback of the point cloud data to obtain the point cloud frames in the fused frames. In this way, in both media transmission and local playback scenarios, the terminal can be provided with relevant parameter information of each frame in the fused frame, so that the terminal can respond to the user's playback operation in a timely manner according to information such as the time span, decoding time, and display time (i.e., playback time) of the fused frame, ensuring that the point cloud frames in the fused frame can be decoded in a timely manner, improving the efficiency of file parsing, avoiding the point cloud data from being stuck during playback, and effectively improving the user experience.

[0054] The implementation environment of the embodiments of the present disclosure is introduced below.

[0055] like Figure 2 2 is a schematic diagram of a communication system provided by an embodiment of the present disclosure. The communication system may include: a first device (such as a terminal 201) and a second device (such as a server 202).

[0056] The terminal 201 can perform wired / wireless communication with the server 202 .

[0057] Terminal 201 may request server 202 to play the point cloud data stored in server 202. Subsequently, server 202 may add the decoding time of the fused frame in the point cloud data to the description of the point cloud data when generating the description of the point cloud data, and send the description carrying the decoding time of the fused frame to terminal 201, so that terminal 201 can obtain the corresponding fused frame from server 202 in advance during the playback of the point cloud data based on the decoding time of the fused frame in the description.

[0058] Optionally, the terminal 201 can obtain the decoding time of the fused frame in the point cloud data during the process of compressing the point cloud data, and add the decoding time of the fused frame to the media file of the point cloud data, so that the terminal 201 can decode the fused frame in the point cloud data in advance based on the decoding time of the fused frame during the process of playing the media file of the point cloud data.

[0059] That is to say, the point cloud data playback method provided by the embodiment of the present disclosure can be applied to the local playback scenario of the device, and can also be applied to the remote playback scenario of streaming media between devices.

[0060] The terminal (such as terminal 201) can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook computer, or other device with transceiver functions. The embodiments of the present disclosure do not impose any particular restrictions on the specific form of the terminal. The terminal can interact with the user through one or more methods such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device.

[0061] The server (such as server 202) can be a single physical server, or a server cluster consisting of multiple servers. Alternatively, the server cluster can be a distributed cluster. Alternatively, the server can be a cloud server. The embodiments of this disclosure do not limit the specific implementation of the server.

[0062] After introducing the application scenarios and implementation environment of the embodiments of the present disclosure, the point cloud playback and packaging method provided by the embodiments of the present disclosure is described in detail below in combination with the above implementation environment.

[0063] The present disclosure provides a method for packaging point cloud data. Figure 3 As shown, the point cloud data packaging method may include: S301-S302.

[0064] S301: Obtain fusion frame information of point cloud data.

[0065] Among them, the fused frame information may include the time information of the fused frame in the point cloud data, and the time information of the fused frame may include the decoding time of the fused frame and the playback time of the fused frame. The decoding time of the fused frame is used to indicate the time for decoding the fused frame during the playback of the point cloud data, and the playback time of the fused frame is used to indicate the time for playing the fused frame during the playback of the point cloud data, and the decoding time of the fused frame is earlier than the playback time of the fused frame.

[0066] For example, if the playback time of the point cloud data is 30 minutes and the playback time of the fused frame is 20 minutes and 30 seconds, the decoding time of the fused frame in the point cloud data may be 20 minutes and 15 seconds.

[0067] That is, the time difference between the decoding time and the playback time of the video frame is greater than or equal to the time required to decode the video frame.

[0068] It should be noted that in the disclosed embodiments, a fused frame is a single frame in point cloud data that is formed by fusing multiple consecutive and similar point cloud frames. A fused frame can include multiple point cloud frames, and the point cloud frames included in the fused frame can be encoded point cloud frames. The point cloud frames included in the fused frame have a fixed playback order during the playback of the point cloud data. In other words, each point cloud frame in the fused frame has its own decoding time and playback time.

[0069] As a possible implementation method, the time information of the fused frame can also include the decoding time offset and playback time offset of each point cloud frame in the fused frame. The decoding time offset of each point cloud frame is used to characterize the offset of the decoding time of the point cloud frame relative to the decoding time of the fused frame, and the playback time offset of each point cloud frame is used to characterize the offset of the playback time of the point cloud frame relative to the playback time of the fused frame.

[0070] Exemplarily, the fused frame A includes point cloud frame B and point cloud frame C, and in the time information of the fused frame A, the playback time of the fused frame A is 5 minutes and 3 seconds, the decoding time of the fused frame A is 5 minutes and 1.2 seconds, the playback time offset of the point cloud frame B is 0 seconds, the decoding time offset of the point cloud frame B is 0.9 seconds, the playback time offset of the point cloud frame C is 1 second, and the decoding time offset of the point cloud frame C is 1.9 seconds. Then, the decoding time of the point cloud frame B is 5 minutes and 2.1 seconds, the playback time of the point cloud frame B is 5 minutes and 3 seconds, the decoding time of the point cloud frame C is 5 minutes and 3.1 seconds, and the playback time of the point cloud frame B is 5 minutes and 4 seconds.

[0071] Optionally, the fused frame information may further include the number of fused frames of the point cloud data, the number of point cloud frames contained in each fused frame in the point cloud data, and the time information of the fused frame may further include the time length of the fused frame.

[0072] Exemplarily, in combination with the above example, the time length of the fused frame in the time information of the fused frame A is 2 seconds.

[0073] It should be noted that the disclosed embodiments do not limit the method for obtaining fused frame information. For example, the point cloud data packaging device may obtain fused frame information by identifying the point cloud data. In another example, the point cloud data packaging device may receive fused frame information input by a user for point cloud data. In another example, the point cloud data packaging device may obtain fused frame information by counting the generation records of fused frames in the point cloud data.

[0074] S302: Obtain a media file of the point cloud data according to the fused frame information and the point cloud data.

[0075] The media file may be a container file of the point cloud data based on a geometrically encoded point cloud bitstream, and the media file of the point cloud data may include fusion frame information.

[0076] As a possible implementation, the point cloud data packaging device can obtain a media file of the point cloud data based on the fused frame information and the point cloud data using any of the packaging methods: single-track, multi-track, or block-based. The point cloud data media file may also include type information indicating the packaging method of the point cloud data.

[0077] It should be noted that the packaging device of the point cloud data can store the parameter information (i.e., fused frame information, including spatial position information and block division information) of the fused point cloud frame (i.e., fused frame) in a media file based on the International Organization for Standardization (ISO) basic media file format. The basic media file format can be operated with reference to the MPEG-4 Part 12 ISO Base Media File Format developed by the ISO / IEC JTC1 / SC29 / WG11 Moving Picture Experts Group (MPEG). The point cloud compression data format can be operated with reference to the MPEG-IPart 9:G-PCC point cloud compression technology based on geometric coding developed by the ISO / IEC JTC1 / SC29 / WG11 Moving Picture Experts Group (MPEG).

[0078] The following describes the three packaging methods of single track, multi-track, and block-based track with the help of the accompanying figures:

[0079] (1) Single track:

[0080] Figure 4 Schematic diagram of the point cloud single track storage structure according to an embodiment of the present disclosure. Figure 4As shown in the figure, a single geometry-coded point cloud compression track (G-PCC track) is used to store multiple elements of geometry-coded point cloud compression data. Configuration information, sequence parameter set (SPS), geometry parameter set (GPS), attribute parameter set (APS), etc. are represented by the geometry-coded point cloud compression configuration information data box (i.e., GPCC config box) in the sample entry (Sample entry) of gpe1 (used to indicate a single track). Each sample (such as sample1) contains the geometry data (i.e., geometry data unit) of a frame of point cloud and one or more attribute data (i.e., attribute data unit). The distinction between geometry data and attribute data is described by the sub-sample (i.e., sub-sample) information data box (Subsample Information box).

[0081] (2) Multi-track:

[0082] Figure 5 Schematic diagram of the point cloud multi-track storage structure according to an embodiment of the present disclosure. Figure 5 As shown, multiple elements of the point cloud compression data based on geometry coding can be stored through multiple tracks of point cloud compression based on geometry coding (such as the point cloud compression geometry data track based on geometry coding (G-PCC geometrydata track) and the point cloud compression attribute data track based on geometry coding (G-PCC attribute data track)). Taking the track where the geometry data is located (i.e., the G-PCC geometry data track) as an example, the configuration information, sequence parameter set, and geometry parameter set are described in the sample entry of gpc1 (used to indicate multiple tracks) of the track containing the geometry data. Each sample contains the geometry data of a frame of point cloud. Similarly, each type of attribute data is also stored separately through an independent track.

[0083] (3) Block-based tracks:

[0084] Figure 6 FIG. 1 is a schematic diagram of a block-based track storage structure of a point cloud according to an embodiment of the present disclosure. Figure 6As shown, point clouds of different regions (such as tile data units and tile configuration information) can be encapsulated in sample entries and samples of gpt1 (used to indicate tile tracks) of respective geometry-coded point cloud compression block tracks (such as G-PCC tile track 1 and G-PCC tile track 2) according to the spatial region division (i.e., tile division) of the 3D point cloud in the point cloud frame. The geometry-coded point cloud compression base track (i.e., G-PCC base track) is used as the access entry, wherein the configuration information, sequence parameter set, and geometry parameter set are described in the sample entry of gpeb (used to indicate the base track) containing the base track. The base track points to the geometry data track and one or more attribute data tracks through track references.

[0085] It can be understood that before compressing the point cloud data, the decoding time of the fused frame in the point cloud data is obtained so that the decoding time of the fused frame can be added to the media file of the point cloud data during the compression process of the point cloud data. When playing the point cloud data, the decoding time of the fused frame can be obtained by decapsulating the media file of the point cloud data, and then the fused frame is decoded based on the decoding time of the fused frame, ensuring that the point cloud frame in the fused frame can be decoded and played in time, improving the file parsing efficiency, avoiding the occurrence of jamming of the point cloud data during playback, and effectively improving the user experience.

[0086] The following describes the method for packaging point cloud data provided by the embodiments of the present disclosure in conjunction with specific embodiments.

[0087] When there are fused frames in the point cloud compression data (i.e., point cloud data), they can be identified in the file based on the overall time information of the fused frames. The overall time information of the fused frames is represented by the fused frame information data box.

[0088] The syntax of the fusion frame information data box is:

[0089]

[0090] Semantics:

[0091] num_combine_frames, which indicates the number of fused frames in the current track;

[0092] num_frames, which indicates the number of point cloud frames contained in the fused point cloud frame;

[0093] duration, which indicates the total duration of the fused point cloud frame.

[0094] The following introduces the packaging method of fused point cloud frames in combination with the above three packaging formats.

[0095] Implementation method 1:

[0096] When the point cloud compressed data in the file is stored in a single track structure, the fusion frame information is defined in the sample entry of the single track in the following format:

[0097] Sample entry:

[0098] Sample Entry Type:'gpe1'or'gpeg';

[0099] Container:SampleDescriptionBox;

[0100] Mandatory:A'gpe1'or'gpeg'sample entry is mandatory;

[0101] Quantity:One or more sample entries may be present;

[0102] grammar:

[0103]

[0104] Implementation 2:

[0105] When the point cloud compressed data in the file is stored in a multi-track structure, the fusion frame information is defined in the sample entry of the geometry data track in the following format:

[0106] Sample entry:

[0107] Sample Entry Type:'gpc1'or'gpcg';

[0108] Container:SampleDescriptionBox;

[0109] Mandatory:A'gpc1'or'gpcg'sample entry is mandatory;

[0110] Quantity:One or more sample entries may be present;

[0111] grammar:

[0112]

[0113] Implementation 3:

[0114] When the point cloud compressed data in the file is stored in a block track structure, the fusion frame information is defined in the sample entry of the basic track in the following format:

[0115] Sample entry:

[0116] Sample Entry Type:'gpeb'or'gpcb';

[0117] Container:SampleDescriptionBox;

[0118] Mandatory: No;

[0119] Quantity:Zero or more sample entries may be present;

[0120] grammar:

[0121]

[0122] The above describes the point cloud data packaging method. After the point cloud data is packaged, a point cloud data playback device can obtain the packaged point cloud data (i.e., the point cloud data media file) and play the point cloud data based on the packaged point cloud data. The following describes in detail the process of playing point cloud data by the point cloud data playback device.

[0123] The present disclosure provides a method for playing point cloud data. Figure 7 As shown, the method for playing the point cloud data may include: S701-S702.

[0124] S701: Acquire a fused frame of point cloud data and time information of the fused frame.

[0125] It should be noted that the embodiments of the present disclosure do not limit the scenarios for obtaining the fused frames of point cloud data and the time information of the fused frames. For example, the embodiments of the present disclosure can obtain the fused frames of point cloud data and the time information of the fused frames in the scenario of local playback. For another example, the embodiments of the present disclosure can obtain the fused frames of point cloud data and the time information of the fused frames in the scenario of remote playback of streaming media. In the scenario of local playback, the resources played by the playback device of point cloud data are locally stored resources. In the scenario of remote playback of streaming media, the resources played by the playback device of point cloud data are resources stored in other devices.

[0126] As a possible implementation, in a local playback scenario, a point cloud data playback device can retrieve the point cloud data media file from a storage space. The point cloud data media file includes fused frames and their time information. The point cloud data playback device can then decapsulate the media file to obtain the fused frames and their time information.

[0127] It is understandable that in the scenario of local playback of point cloud data, by adding the decoding time of the fused frame in the point cloud data to the media file of the point cloud data, the fused frame can be decoded in time when the point cloud data is played, thereby ensuring that the point cloud frame in the fused frame can be decoded in time, avoiding lag when playing the point cloud data.

[0128] As another possible implementation, in a remote streaming media playback scenario, a point cloud data playback device can send a point cloud data playback request message to a point cloud data server and receive a description of the point cloud data from the server. The description includes the time information of the fused frame. The point cloud data playback device can then retrieve the fused frame from the server based on the time information.

[0129] Optionally, the point cloud data playback device can determine the acquisition time of the fused frame based on the transmission delay between the point cloud data and the server and the decoding time of the fused frame, and obtain the fused frame from the server according to the acquisition time of the fused frame.

[0130] For example, if the transmission delay between the terminal (i.e., the playback device of the point cloud data) and the server (i.e., the server side) is 2 seconds, and the decoding time of the fused frame A is 12 minutes and 11 seconds of the point cloud data, the terminal can determine that the acquisition time of the fused frame A is 12 minutes and 9 seconds of the point cloud data, and when the point cloud data is played to 12 minutes and 9 seconds, the terminal obtains the fused frame A by sending a request message to the server to indicate the acquisition of the fused frame A in the point cloud data.

[0131] In an embodiment of the present disclosure, the playback request message may be a streaming data request message based on the Moving Picture Experts Group dynamic adaptive streaming over HTTP (MPEG-DASH) protocol, and the description information may be information based on the Media Presentation Description (MPD) format.

[0132] It is understandable that in the scenario where point cloud data is played remotely, by adding the decoding time of the fused frame in the point cloud data to the description information of the point cloud data, the fused frame can be obtained in time when the point cloud data is played on the other end, thereby ensuring the smoothness of the playback of the point cloud data.

[0133] S702 : Decode the fused frame based on the decoding time of the fused frame to play the point cloud data.

[0134] As a possible implementation method, the fused frame is decoded based on the decoding time of the fused frame to obtain all original point cloud frames in the point cloud data, and the original point cloud frames are played in sequence to play the point cloud data.

[0135] In the disclosed embodiments, the decoding time of a fused frame specifically indicates the time it takes to restore the fused frame into multiple point cloud frames during the playback of point cloud data. The fused frame is decoded based on the decoding time to obtain all point cloud frames included in the fused frame, and thus all original point cloud frames in the point cloud data.

[0136] For example, the fused frame of the point cloud data is formed by the fusion of point cloud frame A at 3 minutes and 11 seconds, point cloud frame B at 3 minutes and 12 seconds, and point cloud frame C at 3 minutes and 13 seconds in the point cloud data, and the decoding time of the fused frame is 2 minutes and 40 seconds. Then, the playback end can restore the fused frame to point cloud frame A, point cloud frame B, and point cloud frame C by decoding the fused frame when the point cloud data is played to 2 minutes and 40 seconds.

[0137] In the disclosed embodiment, a device for playing point cloud data can decode a fused frame based on the decoding time of the fused frame to obtain individual point cloud frames within the fused frame, and decode the point cloud frame based on the decoding time offset of the point cloud frame to obtain a decoded point cloud frame. The device for playing point cloud data can then determine the playback time of each point cloud frame based on the playback time of the fused frame and the playback time offset of each point cloud frame, and play the decoded point cloud frame based on the playback time of each point cloud frame, thereby achieving playback of a dynamic point cloud.

[0138] The following describes the playback process of the above point cloud data with reference to specific examples.

[0139] For example, point cloud data A may include: point cloud frame A, point cloud frame B, fused frame C, point cloud frame D, and fused frame E. Fused frame C includes point cloud frame F and point cloud frame G, and fused frame E includes point cloud frame H, point cloud frame I, and point cloud frame J.

[0140] Among them, the playback time of point cloud frame A in the time information of point cloud frame A is 1 second, and the decoding time of point cloud frame A is 0.1 second; the playback time of point cloud frame B in the time information of point cloud frame B is 2 seconds, and the decoding time of point cloud frame B is 1.1 seconds; the playback time of fused frame C in the time information of fused frame C is 3 seconds, and the decoding time of fused frame C is 1.2 seconds, the playback time offset of point cloud frame F is 0 seconds, and the decoding time offset of point cloud frame F is 0.9 seconds, the playback time offset of point cloud frame G is 1 second, and the decoding time offset of point cloud frame G is 1.9 seconds; in the time information of point cloud frame D, the playback time of point cloud frame D is 5 seconds, and the decoding time of point cloud frame D is 4.1 seconds; in the time information of fused frame E, the playback time of fused frame E is 6 seconds, and the decoding time of fused frame E is 4.2 seconds. The playback time offset of point cloud frame H is 0 seconds, and the decoding time offset of point cloud frame H is 0.9 seconds. The playback time offset of point cloud frame I is 1 second, and the decoding time offset of point cloud frame I is 1.9 seconds. The playback time offset of point cloud frame J is 2 seconds, and the decoding time offset of point cloud frame J is 2.9 seconds.

[0141] The point cloud data playback device can then decode point cloud frame A at 0.1 seconds and play the decoded point cloud frame A at 1 second. Next, the point cloud data playback device can decode point cloud frame B at 1.1 seconds and play the decoded point cloud frame B at 2 seconds. Simultaneously, the point cloud data playback device can decode fused frame C at 1.2 seconds to obtain point cloud frames F and G, and then sequentially decode point cloud frame F at 2.1 seconds, play the decoded point cloud frame F at 3 seconds, decode point cloud frame G at 3.1 seconds, and play the decoded point cloud frame G at 4 seconds. Thereafter, the point cloud data playback device can decode point cloud frame D at 4.1 seconds and play the decoded point cloud frame D at 5 seconds. At the same time, the point cloud data playback device can decode fusion frame C at 4.2 seconds to obtain point cloud frame H, point cloud frame I, and point cloud frame J, and then decode point cloud frame H at 5.1 seconds, play the decoded point cloud frame H at 6 seconds, decode point cloud frame I at 6.1 seconds, play the decoded point cloud frame I at 7 seconds, decode point cloud frame J at 7.1 seconds, and play the decoded point cloud frame J at 8 seconds. Based on the above decoding and playback operations, the point cloud data playback device can play point cloud data A.

[0142] It can be understood that when there are fused frames in the point cloud data, by obtaining the decoding time of the fused frames, the fused frames can be decoded in advance during the playback of the point cloud data to obtain the point cloud frames in the fused frames, thereby ensuring that the point cloud frames in the fused frames can be decoded in time to avoid freezes when playing the point cloud data.

[0143] The following describes the packaging and playback methods of point cloud data in combination with different usage scenarios (such as local playback and streaming remote playback).

[0144] In some embodiments, for local playback scenarios, the playback device of the point cloud data and the packaging device of the point cloud data are the same device. In addition, when the user's playback operation changes frequently, in order to parse the file and decode the compressed data more efficiently, the terminal (i.e., the playback device of the point cloud data) needs to be able to read the time information of each frame of the point cloud before decoding, including the decoding time and the display time, etc. This information can be identified in the file in the form of sub-samples. In this embodiment, the complete fused point cloud frame data is stored in a single sample, and the time inside the fused point cloud frame is described by the sub-sample information data box (SubSampleInformationBox) of the sample. Each sub sample represents a point cloud frame inside the fused frame, and its corresponding sub sample information represents the time information of the point cloud frame.

[0145] When SubsampleInformationBox indicates G-PCC fusion frame time information, flags is 1 and the codec_specific_parameters extension syntax is as follows:

[0146]

[0147] Semantics:

[0148] pts_offset indicates the offset of the presentation time of the frame relative to the current sample start presentation time (CTS).

[0149] dts_offset indicates the offset of the decoding time of the frame relative to the start decoding time (DTS) of the current sample.

[0150] This embodiment also includes three implementation methods. When the point cloud compressed data in the file is stored in a single track structure, the subsample information data box is defined in the single track. When the point cloud compressed data in the file is stored in a multi-track structure, the subsample information data box is defined in the geometry data track. When the point cloud compressed data in the file is stored in a block-based track structure, the subsample information data box is defined in the base track.

[0151] For locally played and rendered point cloud compressed data, when it contains fused point cloud frames, the terminal parsing process is as follows:

[0152] (1) The terminal file parser reads and determines the point cloud track type (single track, multi-track, block track) according to the sample entry type;

[0153] (2) The file parser reads the point cloud configuration information and parameter information in the sample entry of the point cloud track, point cloud geometry data track, or point cloud basic track, including SPS, APS, GPS, Tile Inventory, etc., reads the fusion frame information data box in the sample entry, and identifies the number of fusion point cloud frames contained in the current track and the time information;

[0154] (3) The terminal performs calculations based on the user's playback operation and determines the fused frame point cloud to be decoded in advance;

[0155] (4) The file parser obtains the fused frame point cloud contained in the point cloud track, inputs it into the decoder to complete decoding, and separates the individual point cloud frames in the fused frame;

[0156] (5) The terminal renders the point cloud frame by frame.

[0157] In other embodiments, for the scenario of remote playback of streaming media, the playback device of point cloud data and the packaging device of point cloud data are different devices, and point cloud fusion frame transmission signaling can be implemented between the playback device of point cloud data and the packaging device of point cloud data. The point cloud transmission based on geometric coding between the playback device of point cloud data and the packaging device of point cloud data can be described by the MPEG-DASH transmission protocol. The basic point cloud code stream can be represented by multiple adaptation sets in the DASH MPD description file, such as Figure 8 As shown, it shows a schematic diagram of a point cloud description information structure.

[0158] When a point cloud (e.g., a temporal segment (i.e., Period 1) of a volumetric video) contains fused frames, the relevant fused frame information needs to be indicated in the DASH MPD description file (i.e., Point Cloud Compression Descriptor Based on Geometry Coding (GPCCdescriptor)) of the preselection (e.g., PreSelection 1). Each point cloud track is represented in the DASH MPD file by a separate adaptation set (e.g., AdaptationSet 1 (Main Adaptation Set), AdaptationSet 2, AdaptationSet 3, and AdaptationSet 4). The @codec attribute values ​​of the adaptation set (e.g., GPCC Component descriptor, Representation 1, Representation 2) should be consistent with the sample entry attribute values ​​in the track (e.g., Geometry Data or Main Geometry (i.e., Geometry (Main)), Attribute Data (e.g., Attribute 1, Attribute 2, and Attribute 3)). The point cloud fusion frame information is represented by the point cloud fusion frame information descriptor. The descriptor is an EssentialProperty. The @schemeIdUri value is "urn:mpeg:mpegI:gpcc:2020:CombineFrameInfo". The descriptor is defined as follows:

[0159] Table 1 Descriptor definitions

[0160]

[0161]

[0162] The point cloud transmission process including the fused frame is as follows:

[0163] (1) The terminal (i.e., the point cloud data playback device) initiates a point cloud media service request based on the MPEG-DASH protocol to the server (the point cloud data packaging device) to obtain a media description file (MPD);

[0164] (2) The terminal parses the media description file and determines the main adaptation set (MainAdaptationSet) based on the @codecs type in the adaptation set;

[0165] (3) The terminal parses the main adaptation set point cloud fusion frame information descriptor and identifies the fusion frame contained in the current code stream;

[0166] (4) The terminal downloads the media segments containing the point cloud fusion frames in advance according to user needs;

[0167] (5) The terminal file parser obtains the point cloud contained in the slice file described in step 4 and inputs it into the decoder to complete decoding;

[0168] (6) Terminal rendering of part of the point cloud.

[0169] It is understandable that in order to achieve the above functions, the playback device of point cloud data and the packaging device of point cloud data include hardware structures and / or software modules corresponding to the execution of each function. It should be easy for those skilled in the art to realize that, in combination with the algorithm steps of each example described in the embodiments of the present disclosure, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present disclosure.

[0170] The embodiment of the present disclosure can divide the point cloud data playback device and the point cloud data encapsulation device into functional modules according to the above-mentioned method embodiment. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one functional module. The above-mentioned integrated module can be implemented in the form of hardware or software. It should be noted that the division of modules in the embodiment of the present disclosure is schematic and is only a logical function division. There may be other division methods in actual implementation. The following is an example of dividing each functional module corresponding to each function.

[0171] Figure 9 This is a schematic diagram of a point cloud data playback device provided by an embodiment of the present disclosure. The point cloud data playback device 900 can execute the above method embodiment. Figure 7 The method for playing point cloud data is shown in FIG. Figure 9 As shown, the point cloud data playback device 900 includes: an acquisition module 901 and a processing module 902.

[0172] The acquisition module 901 is used to acquire the fused frame of the point cloud data and the time information of the fused frame, wherein the time information of the fused frame includes the decoding time of the fused frame. The processing module 902 is used to decode the fused frame based on the decoding time of the fused frame to play the point cloud data.

[0173] Optionally, the acquisition module 901 is specifically configured to acquire a media file of the point cloud data, the media file including the fused frame and the time information of the fused frame. The processing module 902 is further configured to decapsulate the media file to obtain the fused frame and the time information of the fused frame.

[0174] Optionally, the point cloud data playback device 900 may further include a sending module 903. The sending module 903 is configured to send a point cloud data playback request message to a point cloud data server. The acquiring module 901 is further configured to receive description information of the point cloud data sent by the server, the description information including time information of the fused frame. The processing module 902 is further configured to acquire the fused frame from the server based on the time information of the fused frame.

[0175] Optionally, the time information of the fused frame also includes the decoding time offset of each point cloud frame in the fused frame, and the decoding time offset of each point cloud frame is used to represent the offset of the decoding time of the point cloud frame relative to the decoding time of the fused frame.

[0176] Optionally, processing module 902 is specifically configured to decode the fused frame based on the decoding time of the fused frame to obtain individual point cloud frames in the fused frame. Processing module 902 is also configured to decode the point cloud frame based on the decoding time offset of the point cloud frame to obtain a decoded point cloud frame. Processing module 902 is also configured to play the decoded point cloud frame.

[0177] Optionally, the time information of the fused frame also includes the playback time of the fused frame and the playback time offset of each point cloud frame in the fused frame. The playback time offset of each point cloud frame is used to represent the offset of the playback time of the point cloud frame relative to the playback time of the fused frame.

[0178] Optionally, the processing module 902 is further configured to determine the playback time of the point cloud frame based on the playback time of the fused frame and the playback time offset of the point cloud frame. The processing module 902 is further configured to play the decoded point cloud frame based on the playback time of the point cloud frame.

[0179] Optionally, the time information of the fused frame also includes the time length of the fused frame, and the media file also includes at least one of the following: the number of fused frames of the point cloud data, and the number of point cloud frames contained in each fused frame in the point cloud data.

[0180] Optionally, the media file is a container file for the point cloud data based on a geometrically encoded point cloud bitstream, and the media file further includes type information indicating a packaging method for the point cloud data. The packaging method for the point cloud data can be any of single-track, multi-track, and block-based tracks.

[0181] Optionally, the play request message is a stream data request message based on the Moving Picture Experts Group's Dynamic Adaptive Streaming (MPEG) DASH protocol, and the description information is information based on the Media Presentation Description (MPD) format.

[0182] Figure 10This is a schematic diagram of a point cloud data packaging device provided by an embodiment of the present disclosure. The point cloud data packaging device 1000 can execute the above method embodiment. Figure 3 The packaging method of point cloud data is shown in . Figure 10 As shown, the point cloud data packaging device 1000 includes: an acquisition module 1001 and a processing module 1002.

[0183] Acquisition module 1001 is configured to obtain fused frame information of point cloud data. The fused frame information includes time information of the fused frame in the point cloud data, including the decoding time of the fused frame. Processing module 1002 is configured to obtain a media file of the point cloud data based on the fused frame information and the point cloud data. The media file includes the fused frame and its time information.

[0184] Optionally, the time information of the fused frame also includes the decoding time offset of each point cloud frame in the fused frame, and the decoding time offset of each point cloud frame is used to represent the offset of the decoding time of the point cloud frame relative to the decoding time of the fused frame.

[0185] Optionally, the time information of the fused frame also includes the playback time of the fused frame and the playback time offset of each point cloud frame in the fused frame. The playback time offset of each point cloud frame is used to represent the offset of the playback time of the point cloud frame relative to the playback time of the fused frame.

[0186] Optionally, the fused frame information also includes the number of fused frames of the point cloud data and the number of point cloud frames contained in each fused frame of the point cloud data. The time information of the fused frame also includes the time length of the fused frame. The media file also includes: the number of fused frames of the point cloud data and the number of point cloud frames contained in each fused frame of the point cloud data.

[0187] Optionally, the media file is a container file of the point cloud data based on the geometrically encoded point cloud bitstream. The processing module 1002 is specifically configured to obtain the media file by encapsulating the fused frame information and the point cloud data in any of single-track, multi-track, and block-based track formats.

[0188] It should be noted that in the local playback scenario, the point cloud data playback device 900 and the point cloud data packaging device 1000 can be the same device. In the streaming media remote playback scenario, the point cloud data playback device 900 and the point cloud data packaging device 1000 can be different devices.

[0189] In the case of implementing the functions of the above-mentioned integrated modules in the form of hardware, the embodiment of the present disclosure provides another possible network device structure of the point cloud data playback device involved in the above-mentioned embodiment. Figure 11As shown, the point cloud data playback device 1100 includes: a processor 1102 and a bus 1104. Optionally, the point cloud data playback device 1100 may further include a memory 1101; and optionally, the point cloud data playback device 1100 may further include a communication interface 1103.

[0190] Processor 1102 may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of this disclosure. Processor 1102 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of this disclosure. Processor 1102 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.

[0191] The communication interface 1103 is used to connect to other devices via a communication network, such as Ethernet, wireless access network, or wireless local area network (WLAN).

[0192] The memory 1101 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0193] As a possible implementation, memory 1101 can exist independently of processor 1102. Memory 1101 can be connected to processor 1102 via bus 1104 to store instructions or program code. When processor 1102 calls and executes the instructions or program code stored in memory 1101, the point cloud data playback method provided in the embodiments of the present disclosure can be implemented.

[0194] In another possible implementation, the memory 1101 may also be integrated with the processor 1102 .

[0195] The bus 1104 may be an extended industry standard architecture (EISA) bus, etc. The bus 1104 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0196] Similarly, when the functions of the above integrated modules are realized in the form of hardware, Figure 11 The device structure shown in the figure, the embodiment of the present disclosure provides another possible network device structure of the point cloud data packaging device involved in the above embodiments, which can implement the point cloud data packaging method provided by the embodiment of the present disclosure.

[0197] Some embodiments of the present disclosure provide a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium), which stores computer program instructions. When the computer program instructions are executed on a computer, the computer executes the point cloud data playback and packaging method described in any of the above embodiments.

[0198] Exemplarily, the computer-readable storage media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes), optical disks (e.g., compact disks (CDs), digital versatile disks (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memory (EPROM), cards, sticks, or key drives, etc.). The various computer-readable storage media described in the present disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.

[0199] An embodiment of the present disclosure provides a computer program product comprising instructions. When the computer program product is run on a computer, the computer is enabled to execute the method for playing and packaging point cloud data described in any one of the above embodiments.

[0200] The above is only a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or replacements within the technical scope disclosed in the present disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A method for playing point cloud data, characterized in that: The method comprises: Acquire a fused frame of point cloud data and time information of the fused frame, where the time information of the fused frame includes a decoding time of the fused frame; The fused frame is decoded based on a decoding time of the fused frame to play the point cloud data.

2. The method according to claim 1, characterized in that The acquiring of a fused frame of point cloud data and time information of the fused frame includes: Acquire a media file of the point cloud data, where the media file includes the fused frame and time information of the fused frame; The media file is decapsulated to obtain the fused frame and time information of the fused frame.

3. The method according to claim 1, characterized in that The acquiring of a fused frame of point cloud data and time information of the fused frame includes: Sending a playback request message of the point cloud data to a server end of the point cloud data; receiving description information of the point cloud data sent by the server, where the description information of the point cloud data includes time information of the fused frame; The fused frame is obtained from the server based on time information of the fused frame.

4. The method according to claim 1, wherein The time information of the fused frame also includes the decoding time offset of each point cloud frame in the fused frame, and the decoding time offset of each point cloud frame is used to represent the offset of the decoding time of the point cloud frame relative to the decoding time of the fused frame.

5. The method according to claim 4, characterized in that The decoding of the fused frame based on the decoding time of the fused frame to play the point cloud data includes: Decoding the fused frame based on the decoding time of the fused frame to obtain each point cloud frame in the fused frame; Decoding the point cloud frame based on the decoding time offset of the point cloud frame to obtain a decoded point cloud frame; Play the decoded point cloud frame.

6. The method according to claim 5, characterized in that The time information of the fused frame also includes the playback time of the fused frame and the playback time offset of each point cloud frame in the fused frame. The playback time offset of each point cloud frame is used to represent the offset of the playback time of the point cloud frame relative to the playback time of the fused frame.

7. The method according to claim 6, characterized in that Playing the decoded point cloud frame includes: Determining the playback time of the point cloud frame based on the playback time of the fused frame and the playback time offset of the point cloud frame; The decoded point cloud frame is played based on the playback time of the point cloud frame.

8. The method according to claim 2, characterized in that The time information of the fused frame also includes the time length of the fused frame, and the media file also includes at least one of the following: the number of fused frames of the point cloud data, and the number of point cloud frames contained in each fused frame in the point cloud data.

9. The method according to claim 2, characterized in that The media file is a container file of the point cloud data based on a geometrically encoded point cloud bitstream, and the media file also includes type information for indicating the packaging method of the point cloud data; wherein the packaging method of the point cloud data is any one of a single track, a multi-track, and a block-based track.

10. The method according to claim 3, characterized in that The play request message is a stream data request message based on the Moving Picture Experts Group (MPEG) Dynamic Adaptive Streaming (DASH) protocol, and the description information is information based on the Media Presentation Description (MPD) format.

11. A method for packaging point cloud data, characterized in that: The method comprises: Acquire fusion frame information of the point cloud data, wherein the fusion frame information includes time information of the fusion frame in the point cloud data, and the time information of the fusion frame includes a decoding time of the fusion frame; A media file of the point cloud data is obtained according to the fused frame information and the point cloud data, where the media file includes the fused frame and time information of the fused frame.

12. The method according to claim 11, characterized in that The time information of the fused frame also includes the decoding time offset of each point cloud frame in the fused frame, and the decoding time offset of each point cloud frame is used to represent the offset of the decoding time of the point cloud frame relative to the decoding time of the fused frame.

13. The method according to claim 11, characterized in that The time information of the fused frame also includes the playback time of the fused frame and the playback time offset of each point cloud frame in the fused frame. The playback time offset of each point cloud frame is used to represent the offset of the playback time of the point cloud frame relative to the playback time of the fused frame.

14. The method according to claim 11, characterized in that The fused frame information further includes the number of fused frames of the point cloud data, the number of point cloud frames contained in each fused frame in the point cloud data, and the time information of the fused frame further includes the time length of the fused frame; The media file further includes: the number of fused frames of the point cloud data, and the number of point cloud frames contained in each fused frame in the point cloud data.

15. The method according to claim 11, characterized in that The media file is a container file of the point cloud data based on a geometrically encoded point cloud bitstream, and the compressed media file of the point cloud data is obtained according to the fused frame information and the point cloud data, including: The media file is obtained according to the fused frame information and the point cloud data through any packaging method of a single track, a multi-track, and a block-based track.

16. An electronic device, characterized in that: include: memory and processor; Memory and processor coupling; The memory is used to store instructions executable by the processor; When the processor executes the instructions, the method according to any one of claims 1 to 15 is performed.

17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 15.