Point cloud data playing method and packaging method, device and storage medium

By adding the decoding time information of the fused frame to the point cloud data, the problem of low terminal playback flexibility is solved, smooth playback of point cloud data is achieved, and the user experience is improved.

WO2025189921A1PCT designated stage Publication Date: 2025-09-18ZTE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/144246
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-11
Filing Date
2024-12-31
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

When processing point cloud data of fused frames, the existing technology has low terminal playback flexibility, resulting in lag and inability to respond to user playback operations in a timely manner.

Method used

Add the decoding time information of the fused frame to the point cloud data so that the fused frame can be decoded in advance before playback and the relevant parameter information of each frame can be obtained to ensure timely response to user operations and avoid lag.

Benefits of technology

The playback smoothness and file parsing efficiency of point cloud data are improved, which enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024144246_18092025_PF_FP_ABST
    Figure CN2024144246_18092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a point cloud data playing method and packaging method, a device, and a storage medium. The point cloud data playing method comprises: acquiring a fusion frame of point cloud data and time information of the fusion frame, the time information of the fusion frame comprising the decoding time of the fusion frame; and decoding the fusion frame on the basis of the decoding time of the fusion frame, so as to play the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Point cloud data playback method and packaging method, device and storage medium

[0001] This disclosure claims priority to Chinese patent application No. 202410275726.X, filed on March 11, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present disclosure relates to the field of video processing technology, and in particular to a method for playing and packaging point cloud data, a device, and a storage medium. Background Art

[0003] Video compression technology aims to reduce the amount of video data, thereby lowering the storage space required by electronic devices (such as servers and terminals) and network transmission bandwidth. For example, a server compresses point cloud data composed of multiple point cloud frames and then transmits or stores the compressed point cloud data. Summary of the Invention

[0004] In one aspect, embodiments of the present disclosure provide a method for playing point cloud data. The method comprises: obtaining a fused frame of the point cloud data and time information of the fused frame, wherein the time information of the fused frame includes a decoding time of the fused frame; and decoding the fused frame based on the decoding time of the fused frame to play the point cloud data.

[0005] In another aspect, embodiments of the present disclosure provide a method for packaging point cloud data. The method includes: obtaining fused frame information of the point cloud data, the fused frame information including time information of the fused frame in the point cloud data, the fused frame time information including the decoding time of the fused frame; and obtaining a media file of the point cloud data based on the fused frame information and the point cloud data, the media file including the fused frame and the time information of the fused frame.

[0006] In another aspect, embodiments of the present disclosure provide a device for playing point cloud data. The device includes an acquisition module and a processing module. The acquisition module is configured to acquire a fused frame of the point cloud data and time information of the fused frame, the time information of the fused frame including the decoding time of the fused frame. The processing module is configured to decode the fused frame based on the decoding time of the fused frame to play the point cloud data.

[0007] In another aspect, embodiments of the present disclosure provide a point cloud data packaging device. The point cloud data packaging device includes: an acquisition module and a processing module; the acquisition module is configured to acquire fused frame information of the point cloud data, the fused frame information including time information of the fused frame in the point cloud data, the fused frame time information including the decoding time of the fused frame; and the processing module is configured to generate a media file of the point cloud data based on the fused frame information and the point cloud data, the media file including the fused frame and the fused frame time information.

[0008] On the other hand, an embodiment of the present disclosure provides a network device, including: a memory and a processor, the memory and the processor are coupled, the memory is used to store a computer program, and when the processor executes the computer program, it implements the above-mentioned point cloud data playback method or point cloud data encapsulation method.

[0009] On the other hand, an embodiment of the present disclosure provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above-mentioned point cloud data playback method or point cloud data packaging method is implemented.

[0010] On the other hand, an embodiment of the present disclosure provides a computer program product, which includes computer program instructions, and when the computer program instructions are executed by a processor, implements the above-mentioned point cloud data playback method or point cloud data packaging method. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the present disclosure, the following briefly introduces the drawings required for use in some embodiments of the present disclosure. Obviously, the drawings described below are only drawings of some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0012] FIG1 is a flowchart of end-to-end processing of a point cloud according to some embodiments.

[0013] FIG2 is a schematic diagram of a communication system according to some embodiments.

[0014] FIG3 is a schematic flow chart of a method for packaging point cloud data according to some embodiments.

[0015] FIG4 is a schematic diagram of a point cloud single-track storage structure according to some embodiments.

[0016] FIG5 is a schematic diagram of a point cloud multi-track storage structure according to some embodiments.

[0017] FIG6 is a schematic diagram of a block-based track storage structure of a point cloud according to some embodiments.

[0018] FIG7 is a schematic flow chart of a method for playing point cloud data according to some embodiments.

[0019] FIG8 is a schematic diagram of a point cloud description information structure according to some embodiments.

[0020] FIG9 is a schematic structural diagram of a device for playing point cloud data according to some embodiments.

[0021] FIG10 is a schematic structural diagram of a point cloud data packaging device according to some embodiments.

[0022] FIG11 is a schematic structural diagram of a point cloud data playback device according to some embodiments. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions of this disclosure in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of this disclosure, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0024] It should be noted that in this disclosure, expressions such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described in this disclosure as "exemplarily" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of expressions such as "exemplarily" or "for example" is intended to present the relevant concepts in a detailed manner.

[0025] In the following, the terms "first," "second," etc. are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the quantity of the technical features indicated. Therefore, a feature specified as "first," "second," etc. may explicitly or implicitly include one or more of the features.

[0026] In the description of this disclosure, unless otherwise specified, " / " means "or." For example, A / B can mean A or B. "And / or" herein is simply a way to describe an association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can mean: only A, A and B, and only B. Furthermore, "at least one" means one or more, and "a plurality" means two or more.

[0027] Before introducing in detail the method for playing point cloud data provided by the embodiment of the present disclosure, the implementation environment and application scenarios of the embodiment of the present disclosure are first introduced.

[0028] First, the application scenarios of the embodiments of the present disclosure are introduced.

[0029] A 3D point cloud is a dataset consisting of a large number of 3D points, typically used to represent surfaces or scenes in real-world 3D space. 3D point clouds have a wide range of applications, encompassing many different fields. Typical application scenarios include the following.

[0030] 1. Computer Vision and Graphics: 3D point clouds are widely used in computer vision and graphics for tasks such as target detection, object recognition, and scene segmentation. 3D point clouds can be acquired through methods such as laser scanning and photogrammetry and used to model real-world objects and environments.

[0031] 2. Autonomous driving and robotics: In the field of autonomous vehicles and robotics, three-dimensional point clouds are used for environmental perception and obstacle detection.

[0032] 3. Architecture and urban planning: 3D point clouds can be used in architecture and urban planning to obtain detailed structures of buildings and urban environments through laser scanning or drones to support planning and design work.

[0033] 4. Cultural heritage protection: In the field of cultural heritage, 3D point clouds can be used to digitize and protect ancient buildings, sculptures, and other cultural heritage, and promote the protection and restoration of cultural relics by providing high-precision spatial information.

[0034] 5. Medical image processing: Three-dimensional point clouds are widely used in the field of medical imaging, such as three-dimensional modeling of patient organs to support medical diagnosis and surgical planning.

[0035] 6. Virtual Reality and Augmented Reality: 3D point clouds are used to create realistic virtual reality (VR) and augmented reality (AR) experiences. By capturing 3D information about the real world, virtual and real environments can be more naturally integrated.

[0036] 7. Industrial Manufacturing: In the industrial field, 3D point clouds can be used for quality control, product design, and process optimization. By scanning objects in the manufacturing process, real-time monitoring and analysis can be performed.

[0037] Currently, 3D point clouds are widely used in scenarios such as autonomous driving, real-time inspections, cultural heritage, and immersive real-time communication with six degrees of freedom (6DoF). Figure 1 shows the end-to-end processing flow for point clouds in human-perception scenarios (such as cultural heritage and 6DoF immersive real-time communication), based on the architecture diagram in "Information Technology: Efficient Graphics Data Coding Part 1: Systems." For machine-perception scenarios (such as autonomous driving and real-time inspections), the terminal does not need to render and display the decoded point cloud; the system flow prior to this is consistent with human-perception scenarios.

[0038] As shown in Figure 1, a real-world visual scene A is captured by a set of cameras or a camera device with multiple lenses and sensors, resulting in point cloud source data B, a sequence of numerous point cloud frames. The point cloud encoder encodes one or more point cloud frames into a point cloud stream E. The file encapsulator encapsulates the file / segment according to a specific media container file format (such as ISOBMFF). For transmission, the file encapsulator encapsulates the point cloud stream into media segments and provides an index file (i.e., media presentation description information). The server then uses a transport mechanism to transmit the segments Fs in the media file to the player. On the player side, the file decapsulator decapsulates the received file, extracts the encoded bitstream E', and parses the metadata. The point cloud decoder then decodes the data to generate the point cloud D'. The renderer renders the point cloud and displays it on a display device based on the current viewing position, viewing direction, or a viewport determined by various types of sensors (e.g., head, position, or eye tracking sensors).

[0039] In some embodiments, if local playback is required, the file encapsulator transmits the encapsulated point cloud stream E to the local player, and the point cloud is rendered and displayed on the display device through the above-mentioned file decapsulator, point cloud decoder and renderer.

[0040] However, due to the sheer volume of point clouds, processing and storing this data requires significant computing and storage resources. Therefore, to efficiently transmit and store point clouds, it's often necessary to apply multiple compression algorithms to reduce the data volume. Depending on the characteristics of point clouds and their application scenarios, the requirements for computing and storage networks can be categorized into the following three categories.

[0041] 1. Static objects and scenes represented by static dense point clouds. Application scenarios include digital cultural heritage and digital twins.

[0042] 2. Objects represented by dynamic point clouds (dynamic objects), application scenarios include free viewpoint broadcasting, 3D immersive interaction, etc.

[0043] 3. Dynamic acquisition of objects represented by point clouds (dynamic acquisition). Application scenarios include high-precision maps, navigation, and autonomous driving.

[0044] The first and third types of point clouds are typically compressed using geometry-based point cloud compression (G-PCC), a hot topic explored by various standards organizations. The key idea behind G-PCC is to achieve efficient compression by encoding the geometric information in the point cloud. This involves encoding geometric attributes such as the point cloud's coordinates, normals, and colors. These techniques typically take into account the local characteristics and geometric structure of the point cloud to more efficiently represent the data.

[0045] For example, for dynamically acquired point clouds where scene changes are less noticeable, where the majority of the entire point cloud frame remains unchanged over time, while a smaller portion is updated over time, fused frame compression is often used. This involves fusing multiple similar point cloud frames into a single frame before applying G-PCC compression, and then compressing and encoding the fused frame. The fused frame includes additional attribute data corresponding to the temporal information of each point in the combined frame.

[0046] The advantage of fused frames is that they can effectively reduce the amount of point cloud information in these scenarios. However, when playing or rendering on a terminal, the time information of the fused frames must be fully decoded before it can be calculated frame by frame. During playback, the terminal may involve various caching mechanisms and user operations, such as downloading and playing simultaneously, random scrolling, fast forwarding, and rewinding. These operations rely on the terminal player to obtain the decoding and playback time of each frame of the point cloud before decoding. Typically, traditional Moving Picture Experts Group (MPEG)-2 transport stream (TS) or International Organization for Standardization Base Media File Format (ISOBMFF) encapsulation tools can efficiently parse the playback and decoding time of each frame during transmission and playback.

[0047] However, when there are fused frames in the point cloud compression data, the time information inside the fused frames cannot be obtained in advance in these encapsulation structures, resulting in low terminal playback flexibility and lag in the point cloud data during playback.

[0048] In order to solve the above problems, the embodiments of the present disclosure provide a method for playing and encapsulating point cloud data, which is applied to scenarios of media transmission or local playback of point cloud data including fused frames. In the case where there are fused frames in the point cloud data, the decoding time of the fused frames in the point cloud data is added to the media file of the point cloud data, or the decoding time of the fused frames in the point cloud data is added to the description information of the point cloud data, so as to obtain the decoding time of the fused frames before the playback of the point cloud data, and then the fused frames are decoded in advance during the playback of the point cloud data to obtain the point cloud frames in the fused frames. In this way, in both media transmission and local playback scenarios, the terminal can be provided with relevant parameter information of each frame in the fused frame, so that the terminal can respond to the user's playback operation in a timely manner according to information such as the time span, decoding time, and display time (i.e., playback time) of the fused frame, ensuring that the point cloud frames in the fused frame can be decoded in a timely manner, improving the efficiency of file parsing, avoiding the point cloud data from being stuck during playback, and effectively improving the user experience.

[0049] The implementation environment of the embodiments of the present disclosure is introduced below.

[0050] As shown in Figure 2, which is a schematic diagram of a communication system according to some embodiments, the communication system may include: a first device (such as a terminal 201) and a second device (such as a server 202).

[0051] The terminal 201 can perform wired / wireless communication with the server 202 .

[0052] Terminal 201 may request server 202 to play the point cloud data stored in server 202. Subsequently, server 202 may add the decoding time of the fused frame in the point cloud data to the description of the point cloud data when generating the description of the point cloud data, and send the description carrying the decoding time of the fused frame to terminal 201, so that terminal 201 can obtain the corresponding fused frame from server 202 in advance during the playback of the point cloud data based on the decoding time of the fused frame in the description.

[0053] In some embodiments, the terminal 201 can obtain the decoding time of the fused frame in the point cloud data during the process of compressing the point cloud data, and add the decoding time of the fused frame to the media file of the point cloud data, so that the terminal 201 can decode the fused frame in the point cloud data in advance based on the decoding time of the fused frame during the process of playing the media file of the point cloud data.

[0054] That is to say, the point cloud data playback method provided by the embodiment of the present disclosure can be applied to the local playback scenario of the device, and can also be applied to the remote playback scenario of streaming media between devices.

[0055] A terminal (such as terminal 201) can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook computer, or other device with transceiver functions. The embodiments of the present disclosure do not impose any particular limitation on the form of the terminal. The terminal can interact with the user through one or more methods such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device.

[0056] The server (such as server 202) can be a single physical server, or a server cluster consisting of multiple servers. Alternatively, the server cluster can be a distributed cluster. Alternatively, the server can be a cloud server. The embodiments of this disclosure do not limit the implementation of the server.

[0057] After introducing the application scenarios and implementation environment of the embodiments of the present disclosure, the point cloud data playback method and packaging method provided by the embodiments of the present disclosure are introduced in detail below in combination with the above implementation environment.

[0058] An embodiment of the present disclosure provides a method for packaging point cloud data. As shown in FIG3 , the method for packaging point cloud data may include: S301 - S302 .

[0059] S301: Obtain fusion frame information of point cloud data.

[0060] The fused frame information may include the time information of the fused frame in the point cloud data. The time information of the fused frame may include the decoding time of the fused frame and the playback time of the fused frame. The decoding time of the fused frame is used to indicate the time when the fused frame is decoded during the playback of the point cloud data. The playback time of the fused frame is used to indicate the time when the fused frame is played during the playback of the point cloud data, and the decoding time of the fused frame is earlier than the playback time of the fused frame.

[0061] For example, if the playback time of the point cloud data is 30 minutes and the playback time of the fused frame is 20 minutes and 30 seconds, the decoding time of the fused frame in the point cloud data may be 20 minutes and 15 seconds.

[0062] That is, the time difference between the decoding time and the playback time of the video frame is greater than or equal to the time required to decode the video frame.

[0063] It should be noted that in the disclosed embodiments, a fused frame is a single frame in point cloud data that is formed by fusing multiple consecutive and similar point cloud frames. A fused frame can include multiple point cloud frames, and the point cloud frames included in the fused frame can be encoded point cloud frames. The point cloud frames included in the fused frame have a fixed playback order during the playback of the point cloud data. In other words, each point cloud frame in the fused frame has its own decoding time and playback time.

[0064] As an implementation method, the time information of the fused frame can also include the decoding time offset and playback time offset of each point cloud frame in the fused frame. The decoding time offset of each point cloud frame is used to characterize the offset of the decoding time of the point cloud frame relative to the decoding time of the fused frame, and the playback time offset of each point cloud frame is used to characterize the offset of the playback time of the point cloud frame relative to the playback time of the fused frame.

[0065] Exemplarily, the fused frame A includes point cloud frame B and point cloud frame C, and in the time information of the fused frame A, the playback time of the fused frame A is 5 minutes and 3 seconds, the decoding time of the fused frame A is 5 minutes and 1.2 seconds, the playback time offset of the point cloud frame B is 0 seconds, the decoding time offset of the point cloud frame B is 0.9 seconds, the playback time offset of the point cloud frame C is 1 second, and the decoding time offset of the point cloud frame C is 1.9 seconds. Then, the decoding time of the point cloud frame B is 5 minutes and 2.1 seconds, the playback time of the point cloud frame B is 5 minutes and 3 seconds, the decoding time of the point cloud frame C is 5 minutes and 3.1 seconds, and the playback time of the point cloud frame C is 5 minutes and 4 seconds.

[0066] In some embodiments, the fused frame information may further include the number of fused frames of the point cloud data, the number of point cloud frames contained in each fused frame in the point cloud data, and the time information of the fused frame may further include the time length of the fused frame.

[0067] Exemplarily, in combination with the above example, the time length of the fused frame in the time information of the fused frame A is 2 seconds.

[0068] It should be noted that the embodiments of the present disclosure do not limit the method for obtaining fused frame information. For example, the point cloud data packaging device may obtain fused frame information by identifying the point cloud data. In another example, the point cloud data packaging device may receive fused frame information input by a user for point cloud data. In another example, the point cloud data packaging device may obtain fused frame information by counting the generation records of fused frames in the point cloud data.

[0069] S302: Obtain a media file of the point cloud data according to the fused frame information and the point cloud data.

[0070] The media file may be a container file of a point cloud bitstream based on geometric coding of the point cloud data, and the media file of the point cloud data may include fusion frame information.

[0071] As an implementation, the point cloud data packaging device can obtain a media file of the point cloud data using any of the packaging methods, including single-track, multi-track, and block-based, based on the fused frame information and point cloud data. The media file of the point cloud data can also include type information indicating the packaging method of the point cloud data.

[0072] It should be noted that the packaging device of point cloud data can store the parameter information (i.e., fused frame information, including spatial position information and block division information) of the fused point cloud frame (i.e., fused frame) in a media file based on the International Organization for Standardization (ISO) basic media file format. The basic media file format can be operated with reference to MPEG-4 Part 12 ISO Base Media File Format developed by ISO / IEC JTC1 / SC29 / WG11 Moving Picture Experts Group (MPEG). The point cloud compression data format can be operated with reference to MPEG-I Part 9: G-PCC geometric coding-based point cloud compression technology developed by ISO / IEC JTC1 / SC29 / WG11 Moving Picture Experts Group (MPEG).

[0073] The following describes the three packaging methods of single track, multi-track and block-based track with reference to the accompanying drawings.

[0074] (1) Single track:

[0075] FIG4 is a schematic diagram of a single track storage structure of a point cloud according to some embodiments. As shown in FIG4 , multiple elements of the point cloud compression data based on geometry coding are stored through a single point cloud compression track based on geometry coding (G-PCC track). Configuration information, sequence parameter set (SPS), geometry parameter set (GPS), attribute parameter set (APS), etc. are represented by the configuration information data box (i.e., GPCC config box) of the point cloud compression based on geometry coding in the sample entry (Sample entry) of gpe1 (used to indicate a single track). Each sample (such as sample1) contains a geometry data unit (i.e., geometry data unit) of a frame of point cloud and one or more attribute data units (i.e., attribute data unit). The distinction between geometry data and attribute data is described by the sub-sample (i.e., sub-sample) information data box (Subsample Information box).

[0076] (2) Multi-track:

[0077] FIG5 is a schematic diagram of a multi-track storage structure of a point cloud according to some embodiments. As shown in FIG5 , multiple elements of point cloud compression data based on geometry coding can be stored through multi-tracks of point cloud compression based on geometry coding (such as a G-PCC geometry data track and a G-PCC attribute data track). Taking the track where the geometry data is located (i.e., the G-PCC geometry data track) as an example, the configuration information, sequence parameter set, and geometry parameter set are described in the sample entry of gpc1 (used to indicate multiple tracks) of the track containing the geometry data. Each sample contains the geometry data of a frame of point cloud. Similarly, each type of attribute data is also stored separately through an independent track.

[0078] (3) Block-based tracks:

[0079] FIG6 is a schematic diagram of a block-based track storage structure of a point cloud according to some embodiments. As shown in FIG6 , point clouds of different regions (such as tile data units and tile configuration information) can be encapsulated in sample entries and samples of gpt1 (used to indicate block tracks) of respective geometry-coded point cloud compression block tracks (such as G-PCC tile track 1 and G-PCC tile track 2) according to the spatial region division (i.e., tile division) of the 3D point cloud in the point cloud frame. With the geometry-coded point cloud compression basic track (i.e., G-PCC base track) as the access entry, configuration information, sequence parameter sets, and geometry parameter sets are described in the sample entry of gpeb (used to indicate the basic track) containing the basic track. The basic track points to the geometry data track and one or more attribute data tracks through track references.

[0080] It can be understood that before compressing the point cloud data, the decoding time of the fused frame in the point cloud data is obtained so that the decoding time of the fused frame can be added to the media file of the point cloud data during the compression process of the point cloud data. When playing the point cloud data, the decoding time of the fused frame can be obtained by decapsulating the media file of the point cloud data, and then the fused frame is decoded based on the decoding time of the fused frame, ensuring that the point cloud frame in the fused frame can be decoded and played in time, improving the file parsing efficiency, avoiding the occurrence of jamming of the point cloud data during playback, and effectively improving the user experience.

[0081] The following describes the method for encapsulating point cloud data provided by the embodiments of the present disclosure in conjunction with the embodiments.

[0082] When there are fused frames in the point cloud compression data (i.e., point cloud data), they can be identified in the file based on the overall time information of the fused frames. The overall time information of the fused frames is represented by the fused frame information data box.

[0083] The syntax of the fusion frame information data box is:

[0084] Semantics:

[0085] num_combine_frames, which indicates the number of fused frames in the current track;

[0086] num_frames, which indicates the number of point cloud frames contained in the fused point cloud frame;

[0087] duration, which indicates the total duration of the fused point cloud frame.

[0088] The following introduces the packaging method of fused point cloud frames in combination with the above three packaging formats.

[0089] Implementation method 1:

[0090] When the point cloud compressed data in the file is stored in a single track structure, the fusion frame information is defined in the sample entry of the single track in the following format:

[0091] Implementation 2:

[0092] When the point cloud compressed data in the file is stored in a multi-track structure, the fusion frame information is defined in the sample entry of the geometry data track in the following format:

[0093] Implementation 3:

[0094] When the point cloud compressed data in the file is stored in a block track structure, the fusion frame information is defined in the sample entry of the basic track in the following format:

[0095] The above describes the point cloud data packaging method. After the point cloud data is packaged, a point cloud data playback device can obtain the packaged point cloud data (i.e., the point cloud data media file) and play the point cloud data based on the packaged point cloud data. The following describes in detail the process of playing point cloud data by the point cloud data playback device.

[0096] An embodiment of the present disclosure provides a method for playing point cloud data. As shown in FIG7 , the method for playing point cloud data may include: S701 - S702 .

[0097] S701: Acquire a fused frame of point cloud data and time information of the fused frame.

[0098] It should be noted that the embodiments of the present disclosure do not limit the scenarios for obtaining the fused frames of point cloud data and the time information of the fused frames. For example, the embodiments of the present disclosure can obtain the fused frames of point cloud data and the time information of the fused frames in the scenario of local playback. For another example, the embodiments of the present disclosure can obtain the fused frames of point cloud data and the time information of the fused frames in the scenario of remote playback of streaming media. In the scenario of local playback, the resources played by the playback device of point cloud data are locally stored resources. In the scenario of remote playback of streaming media, the resources played by the playback device of point cloud data are resources stored in other devices.

[0099] As an implementation method, in a local playback scenario, a point cloud data playback device can retrieve the point cloud data media file from a storage space. The point cloud data media file includes fused frames and their time information. The point cloud data playback device can then decapsulate the media file to obtain the fused frames and their time information.

[0100] It is understandable that in the scenario of local playback of point cloud data, by adding the decoding time of the fused frame in the point cloud data to the media file of the point cloud data, the fused frame can be decoded in time when the point cloud data is played, thereby ensuring that the point cloud frame in the fused frame can be decoded in time, avoiding lag when playing the point cloud data.

[0101] As another implementation, in a remote streaming media playback scenario, a point cloud data playback device can send a point cloud data playback request message to a point cloud data server and receive a description of the point cloud data from the server. The description includes the time information of the fused frame. The point cloud data playback device can then retrieve the fused frame from the server based on the time information.

[0102] In some embodiments, the playback device of point cloud data can determine the acquisition time of the fused frame based on the transmission delay between the device and the server and the decoding time of the fused frame, and obtain the fused frame from the server according to the acquisition time of the fused frame.

[0103] For example, if the transmission delay between the terminal (i.e., the playback device of the point cloud data) and the server (i.e., the server side) is 2 seconds, and the decoding time of the fused frame A is 12 minutes and 11 seconds of the point cloud data, the terminal can determine that the acquisition time of the fused frame A is 12 minutes and 9 seconds of the point cloud data, and when the point cloud data is played to 12 minutes and 9 seconds, the terminal obtains the fused frame A by sending a request message to the server to indicate the acquisition of the fused frame A in the point cloud data.

[0104] In an embodiment of the present disclosure, the playback request message may be a streaming data request message based on the Moving Picture Experts Group dynamic adaptive streaming over HTTP (MPEG DASH) protocol, and the description information may be information based on the Media Presentation Description (MPD) format.

[0105] It is understandable that in the scenario where point cloud data is played remotely, by adding the decoding time of the fused frame in the point cloud data to the description information of the point cloud data, the fused frame can be obtained in time when the point cloud data is played on the other end, thereby ensuring the smoothness of the playback of the point cloud data.

[0106] S702 : Decode the fused frame based on the decoding time of the fused frame to play the point cloud data.

[0107] As an implementation method, the fused frame is decoded based on the decoding time of the fused frame to obtain all original point cloud frames in the point cloud data, and the original point cloud frames are played in sequence to play the point cloud data.

[0108] In the disclosed embodiment, the decoding time of a fused frame indicates the time it takes to restore the fused frame into multiple point cloud frames during the playback of point cloud data. The fused frame is decoded based on the decoding time to obtain all point cloud frames included in the fused frame, and thus all original point cloud frames in the point cloud data.

[0109] For example, the fused frame of the point cloud data is formed by the fusion of point cloud frame A at 3 minutes and 11 seconds, point cloud frame B at 3 minutes and 12 seconds, and point cloud frame C at 3 minutes and 13 seconds in the point cloud data, and the decoding time of the fused frame is 2 minutes and 40 seconds. The playback end can restore the fused frame to point cloud frame A, point cloud frame B, and point cloud frame C by decoding the fused frame when the point cloud data is played to 2 minutes and 40 seconds.

[0110] In the disclosed embodiment, a device for playing point cloud data can decode a fused frame based on the decoding time of the fused frame to obtain individual point cloud frames within the fused frame, and decode the point cloud frame based on the decoding time offset of the point cloud frame to obtain a decoded point cloud frame. The device for playing point cloud data can then determine the playback time of each point cloud frame based on the playback time of the fused frame and the playback time offset of each point cloud frame, and play the decoded point cloud frame based on the playback time of each point cloud frame, thereby achieving playback of a dynamic point cloud.

[0111] The following describes the playback process of the above point cloud data with examples.

[0112] For example, point cloud data A may include: point cloud frame A, point cloud frame B, fused frame C, point cloud frame D, and fused frame E. Fused frame C includes point cloud frame F and point cloud frame G, and fused frame E includes point cloud frame H, point cloud frame I, and point cloud frame J.

[0113] In the time information of point cloud frame A, the playback time of point cloud frame A is 1 second and the decoding time of point cloud frame A is 0.1 second; in the time information of point cloud frame B, the playback time of point cloud frame B is 2 seconds and the decoding time of point cloud frame B is 1.1 seconds; in the time information of fused frame C, the playback time of fused frame C is 3 seconds and the decoding time of fused frame C is 1.2 seconds, the playback time offset of point cloud frame F is 0 seconds and the decoding time offset of point cloud frame F is 0.9 seconds, the playback time offset of point cloud frame G is 1 second and the decoding time offset of point cloud frame G is 1. 9 seconds; in the time information of point cloud frame D, the playback time of point cloud frame D is 5 seconds, and the decoding time of point cloud frame D is 4.1 seconds; in the time information of fused frame E, the playback time of fused frame E is 6 seconds, and the decoding time of fused frame E is 4.2 seconds. The playback time offset of point cloud frame H is 0 seconds, and the decoding time offset of point cloud frame H is 0.9 seconds. The playback time offset of point cloud frame I is 1 second, and the decoding time offset of point cloud frame I is 1.9 seconds. The playback time offset of point cloud frame J is 2 seconds, and the decoding time offset of point cloud frame J is 2.9 seconds.

[0114] The point cloud data playback device can then decode point cloud frame A at 0.1 seconds and play the decoded point cloud frame A at 1 second. Next, the point cloud data playback device can decode point cloud frame B at 1.1 seconds and play the decoded point cloud frame B at 2 seconds. Simultaneously, the point cloud data playback device can decode fused frame C at 1.2 seconds to obtain point cloud frames F and G, and then sequentially decode point cloud frame F at 2.1 seconds, play the decoded point cloud frame F at 3 seconds, decode point cloud frame G at 3.1 seconds, and play the decoded point cloud frame G at 4 seconds. Thereafter, the point cloud data playback device can decode point cloud frame D at 4.1 seconds and play the decoded point cloud frame D at 5 seconds. At the same time, the point cloud data playback device can decode fused frame E at 4.2 seconds to obtain point cloud frame H, point cloud frame I, and point cloud frame J, and then decode point cloud frame H at 5.1 seconds, play the decoded point cloud frame H at 6 seconds, decode point cloud frame I at 6.1 seconds, play the decoded point cloud frame I at 7 seconds, decode point cloud frame J at 7.1 seconds, and play the decoded point cloud frame J at 8 seconds. Based on the above decoding and playback operations, the point cloud data playback device can play point cloud data A.

[0115] It can be understood that when there are fused frames in the point cloud data, by obtaining the decoding time of the fused frames, the fused frames can be decoded in advance during the playback of the point cloud data to obtain the point cloud frames in the fused frames, thereby ensuring that the point cloud frames in the fused frames can be decoded in time to avoid freezes when playing the point cloud data.

[0116] The following describes the packaging and playback methods of point cloud data in combination with different usage scenarios (such as local playback and streaming remote playback).

[0117] In some embodiments, for local playback scenarios, the playback device of the point cloud data and the packaging device of the point cloud data are the same device. In addition, when the user's playback operation changes frequently, in order to parse the file and decode the compressed data more efficiently, the terminal (i.e., the playback device of the point cloud data) needs to be able to read the time information of each frame of the point cloud before decoding, including the decoding time and the display time, etc. This information can be identified in the file in the form of sub-samples. In this embodiment, the complete fused point cloud frame data is stored in a single sample, and the time inside the fused point cloud frame is described by the sub-sample information data box (SubSampleInformationBox) of the sample. Each sub sample represents a point cloud frame inside the fused frame, and the sub sample information corresponding to the point cloud frame represents the time information of the point cloud frame.

[0118] When SubsampleInformationBox indicates G-PCC fusion frame time information, flags is 1 and the codec_specific_parameters extension syntax is as follows:

[0119] Semantics:

[0120] pts_offset indicates the offset of the presentation time of the frame relative to the current sample start presentation time (CTS).

[0121] dts_offset indicates the offset of the decoding time of the frame relative to the start decoding time (DTS) of the current sample.

[0122] This embodiment also includes three implementation methods. When the point cloud compressed data in the file is stored in a single track structure, the subsample information data box is defined in the single track. When the point cloud compressed data in the file is stored in a multi-track structure, the subsample information data box is defined in the geometry data track. When the point cloud compressed data in the file is stored in a block-based track structure, the subsample information data box is defined in the base track.

[0123] For locally played and rendered point cloud compressed data, when the point cloud compressed data contains fused point cloud frames, the terminal parsing process is as follows:

[0124] (1) The terminal file parser reads and determines the point cloud track type (single track, multi-track, block track) according to the sample entry type;

[0125] (2) The file parser reads the point cloud configuration information and parameter information in the sample entry of the point cloud track, point cloud geometry data track, or point cloud basic track, including SPS, APS, GPS, Tile Inventory, etc., reads the fusion frame information data box in the sample entry, and identifies the number of fusion point cloud frames contained in the current track and the time information;

[0126] (3) The terminal performs calculations based on the user's playback operation and determines the fused frame point cloud to be decoded in advance;

[0127] (4) The file parser obtains the fused frame point cloud contained in the point cloud track, inputs it into the decoder to complete decoding, and separates the individual point cloud frames in the fused frame;

[0128] (5) The terminal renders the point cloud frame by frame.

[0129] In other embodiments, for the scenario of remote playback of streaming media, the playback device of point cloud data and the encapsulation device of point cloud data are different devices, and point cloud fusion frame transmission signaling can be realized between the playback device of point cloud data and the encapsulation device of point cloud data, and the point cloud transmission based on geometric coding between the playback device of point cloud data and the encapsulation device of point cloud data can be described by the MPEG-DASH transmission protocol. The basic point cloud code stream can be represented by multiple adaptation sets in the DASH MPD description file, as shown in Figure 8, which shows a schematic diagram of a point cloud description information structure.

[0130] When a point cloud (e.g., a temporal segment (i.e., Period 1) of a volumetric video) contains fused frames, the relevant fused frame information needs to be indicated in the DASH MPD description file (i.e., the GPCC descriptor) of the preselection (e.g., PreSelection 1). Each point cloud track is represented in the DASH MPD file by a separate adaptation set (e.g., AdaptationSet 1 (Main), AdaptationSet 2, AdaptationSet 3, and AdaptationSet 4). The @codec attribute value of the adaptation set (e.g., the GPCC Component descriptor representation (e.g., Representation 1, Representation 2)) should be consistent with the sample entry attribute value in the track (e.g., geometry data or main geometry (i.e., Geometry (Main)) and attribute data (e.g., Attribute 1, Attribute 2, and Attribute 3)). The point cloud fusion frame information is represented by the point cloud fusion frame information descriptor. The descriptor is an EssentialProperty. The @schemeIdUri value is "urn:mpeg:mpegI:gpcc:2020:CombineFrameInfo". The descriptor is defined as follows:

[0131] Table 1 Descriptor definitions

[0132] The point cloud transmission process including the fused frame is as follows:

[0133] (1) The terminal (i.e., the point cloud data playback device) initiates a point cloud media service request based on the MPEG DASH protocol to the server (the point cloud data packaging device) to obtain a media description file (MPD);

[0134] (2) The terminal parses the media description file and determines the main adaptation set (Main Adaptation Set) based on the @codecs type in the adaptation set;

[0135] (3) The terminal parses the main adaptation set point cloud fusion frame information descriptor and identifies the fusion frame contained in the current code stream;

[0136] (4) The terminal downloads the media segments containing the point cloud fusion frames in advance according to user needs;

[0137] (5) The terminal file parser obtains the point cloud contained in the slice file described in step 4 and inputs it into the decoder to complete decoding;

[0138] (6) Terminal rendering of part of the point cloud.

[0139] It is understandable that in order to achieve the above functions, the playback device of point cloud data and the packaging device of point cloud data include hardware structures and / or software modules corresponding to the execution of each function. It should be easy for those skilled in the art to realize that, in combination with the algorithm steps of each example described in the embodiments of the present disclosure, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present disclosure.

[0140] The embodiment of the present disclosure can divide the point cloud data playback device and the point cloud data encapsulation device into functional modules according to the above-mentioned method embodiment. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one functional module. The above-mentioned integrated module can be implemented in the form of hardware or software. It should be noted that the division of modules in the embodiment of the present disclosure is schematic and is only a logical function division. There may be other division methods in actual implementation. The following is an example of dividing each functional module corresponding to each function.

[0141] FIG9 is a schematic diagram of a point cloud data playback device according to some embodiments. Point cloud data playback device 900 can execute the point cloud data playback method shown in FIG7 of the aforementioned method embodiment. As shown in FIG9 , point cloud data playback device 900 includes an acquisition module 901 and a processing module 902.

[0142] The acquisition module 901 is used to acquire the fused frame of the point cloud data and the time information of the fused frame, wherein the time information of the fused frame includes the decoding time of the fused frame. The processing module 902 is used to decode the fused frame based on the decoding time of the fused frame to play the point cloud data.

[0143] In some embodiments, the acquisition module 901 is used to acquire a media file of point cloud data, the media file including a fused frame and time information of the fused frame. The processing module 902 is further used to decapsulate the media file to obtain the fused frame and time information of the fused frame.

[0144] In some embodiments, the point cloud data playback device 900 may further include a sending module 903. The sending module 903 is configured to send a point cloud data playback request message to a point cloud data server. The acquiring module 901 is further configured to receive point cloud data description information sent by the server, the point cloud data description information including time information of the fused frame. The processing module 902 is further configured to acquire the fused frame from the server based on the time information of the fused frame.

[0145] In some embodiments, the time information of the fused frame also includes the decoding time offset of each point cloud frame in the fused frame, and the decoding time offset of each point cloud frame is used to represent the offset of the decoding time of the point cloud frame relative to the decoding time of the fused frame.

[0146] In some embodiments, processing module 902 is configured to decode the fused frame based on the decoding time of the fused frame to obtain individual point cloud frames in the fused frame. Processing module 902 is also configured to decode the point cloud frame based on the decoding time offset of the point cloud frame to obtain a decoded point cloud frame. Processing module 902 is also configured to play the decoded point cloud frame.

[0147] In some embodiments, the time information of the fused frame also includes the playback time of the fused frame and the playback time offset of each point cloud frame in the fused frame. The playback time offset of each point cloud frame is used to represent the offset of the playback time of the point cloud frame relative to the playback time of the fused frame.

[0148] In some embodiments, the processing module 902 is further configured to determine the playback time of the point cloud frame based on the playback time of the fused frame and the playback time offset of the point cloud frame. The processing module 902 is further configured to play the decoded point cloud frame based on the playback time of the point cloud frame.

[0149] In some embodiments, the time information of the fused frame further includes the time length of the fused frame, and the media file further includes at least one of the following: the number of fused frames of the point cloud data or the number of point cloud frames contained in each fused frame in the point cloud data.

[0150] In some embodiments, the media file is a container file for the point cloud data based on a geometrically encoded point cloud bitstream, and the media file further includes type information indicating a packaging method for the point cloud data. The packaging method for the point cloud data is any one of a single track, a multi-track, and a block-based track.

[0151] In some embodiments, the play request message is a streaming data request message based on the Moving Picture Experts Group (MPEG) DASH protocol, and the description information is information based on the Media Presentation Description (MPD) format.

[0152] FIG10 is a schematic diagram of a point cloud data packaging device according to some embodiments. The point cloud data packaging device 1000 can execute the point cloud data packaging method shown in FIG3 of the above method embodiment. As shown in FIG10 , the point cloud data packaging device 1000 includes an acquisition module 1001 and a processing module 1002.

[0153] Acquisition module 1001 is configured to obtain fused frame information of point cloud data. The fused frame information includes time information of the fused frame in the point cloud data, including the decoding time of the fused frame. Processing module 1002 is configured to obtain a media file of the point cloud data based on the fused frame information and the point cloud data. The media file includes the fused frame and its time information.

[0154] In some embodiments, the time information of the fused frame also includes the decoding time offset of each point cloud frame in the fused frame, and the decoding time offset of each point cloud frame is used to represent the offset of the decoding time of the point cloud frame relative to the decoding time of the fused frame.

[0155] In some embodiments, the time information of the fused frame also includes the playback time of the fused frame and the playback time offset of each point cloud frame in the fused frame. The playback time offset of each point cloud frame is used to represent the offset of the playback time of the point cloud frame relative to the playback time of the fused frame.

[0156] In some embodiments, the fused frame information also includes the number of fused frames of point cloud data, the number of point cloud frames contained in each fused frame of point cloud data, and the time information of the fused frame also includes the time length of the fused frame. The media file also includes: the number of fused frames of point cloud data, and the number of point cloud frames contained in each fused frame of point cloud data.

[0157] In some embodiments, the media file is a container file of the point cloud data based on the geometrically encoded point cloud bitstream. The processing module 1002 is configured to obtain the media file based on the fused frame information and the point cloud data by encapsulating the media file in any of single-track, multi-track, and block-based tracks.

[0158] It should be noted that in the local playback scenario, the point cloud data playback device 900 and the point cloud data packaging device 1000 can be the same device. In the streaming media remote playback scenario, the point cloud data playback device 900 and the point cloud data packaging device 1000 can be different devices.

[0159] In the case of implementing the functions of the above-mentioned integrated modules in hardware, the embodiments of the present disclosure provide an alternative network device structure for the point cloud data playback device involved in the above-mentioned embodiments. As shown in Figure 11, the point cloud data playback device 1100 includes: a processor 1102 and a bus 1104. In some embodiments, the point cloud data playback device 1100 may also include a memory 1101; in some embodiments, the point cloud data playback device 1100 may also include a communication interface 1103.

[0160] Processor 1102 may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of the present disclosure. Processor 1102 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. Processor 1102 may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of the present disclosure. Processor 1102 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0161] The communication interface 1103 is used to connect to other devices via a communication network, such as Ethernet, wireless access network, or wireless local area network (WLAN).

[0162] The memory 1101 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0163] As an implementation, the memory 1101 can exist independently of the processor 1102. The memory 1101 can be connected to the processor 1102 via a bus 1104 to store instructions or program codes. When the processor 1102 calls and executes the instructions or program codes stored in the memory 1101, the point cloud data playback method provided in the embodiments of the present disclosure can be implemented.

[0164] In another implementation, the memory 1101 may also be integrated with the processor 1102 .

[0165] Bus 1104 can be an Extended Industry Standard Architecture (EISA) bus, etc. Bus 1104 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, FIG11 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.

[0166] Similarly, when the functions of the above-mentioned integrated modules are implemented in the form of hardware, combined with the device structure shown in Figure 11, the embodiment of the present disclosure provides another network device structure of the point cloud data encapsulation device involved in the above-mentioned embodiment, which can implement the point cloud data encapsulation method provided by the embodiment of the present disclosure.

[0167] Some embodiments of the present disclosure provide a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium), which stores computer program instructions. When the computer program instructions are executed on a computer, the computer executes the point cloud data playback and packaging method described in any of the above embodiments.

[0168] For example, the computer-readable storage media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes), optical disks (e.g., compact disks (CDs), digital versatile disks (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memories (EPROMs), cards, sticks, or key drives, etc.). The various computer-readable storage media described in the present disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.

[0169] An embodiment of the present disclosure provides a computer program product comprising instructions. When the computer program product is run on a computer, the computer is enabled to execute the method for playing and packaging point cloud data described in any one of the above embodiments.

[0170] The above is only a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or replacements within the technical scope disclosed in the present disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A method for playing point cloud data, comprising: Acquire a fused frame of point cloud data and time information of the fused frame, where the time information of the fused frame includes a decoding time of the fused frame; The fused frame is decoded based on a decoding time of the fused frame to play the point cloud data.

2. The method according to claim 1, wherein The acquiring of a fused frame of point cloud data and time information of the fused frame includes: Acquire a media file of the point cloud data, where the media file includes the fused frame and time information of the fused frame; The media file is decapsulated to obtain the fused frame and time information of the fused frame.

3. The method according to claim 1, wherein The acquiring of a fused frame of point cloud data and time information of the fused frame includes: Sending a playback request message of the point cloud data to a server end of the point cloud data; receiving description information of the point cloud data sent by the server, where the description information of the point cloud data includes time information of the fused frame; The fused frame is obtained from the server based on time information of the fused frame.

4. The method according to claim 1, wherein The time information of the fused frame also includes a decoding time offset of each point cloud frame in the fused frame, and the decoding time offset of each point cloud frame is used to represent the offset of the decoding time of the point cloud frame relative to the decoding time of the fused frame.

5. The method according to claim 4, wherein The decoding of the fused frame based on the decoding time of the fused frame to play the point cloud data includes: Decoding the fused frame based on the decoding time of the fused frame to obtain each point cloud frame in the fused frame; Decoding the point cloud frame based on the decoding time offset of the point cloud frame to obtain a decoded point cloud frame; Play the decoded point cloud frame.

6. The method according to claim 5, wherein: The time information of the fused frame also includes the playback time of the fused frame and the playback time offset of each point cloud frame in the fused frame. The playback time offset of each point cloud frame is used to represent the offset of the playback time of the point cloud frame relative to the playback time of the fused frame.

7. The method according to claim 6, wherein: Playing the decoded point cloud frame includes: Determining the playback time of the point cloud frame based on the playback time of the fused frame and the playback time offset of the point cloud frame; The decoded point cloud frame is played based on the playback time of the point cloud frame.

8. The method according to claim 2, wherein: The time information of the fused frame also includes the time length of the fused frame, and the media file also includes at least one of the following: the number of fused frames of the point cloud data or the number of point cloud frames included in each fused frame in the point cloud data.

9. The method according to claim 2, wherein: The media file is a container file of the point cloud data based on a geometrically encoded point cloud bitstream, and the media file also includes type information for indicating the packaging method of the point cloud data; wherein the packaging method of the point cloud data is at least one of the following: single track, multi-track and block-based track.

10. The method according to claim 3, wherein: The play request message is a stream data request message based on the Moving Picture Experts Group (MPEG) DASH protocol, and the description information is information based on the Media Presentation Description (MPD) format.

11. A method for packaging point cloud data, comprising: Acquire fusion frame information of the point cloud data, wherein the fusion frame information includes time information of the fusion frame in the point cloud data, and the time information of the fusion frame includes a decoding time of the fusion frame; A media file of the point cloud data is obtained according to the fused frame information and the point cloud data, where the media file includes the fused frame and time information of the fused frame.

12. The method according to claim 11, wherein The time information of the fused frame also includes a decoding time offset of each point cloud frame in the fused frame, and the decoding time offset of each point cloud frame is used to represent the offset of the decoding time of the point cloud frame relative to the decoding time of the fused frame.

13. The method according to claim 11, wherein The time information of the fused frame also includes the playback time of the fused frame and the playback time offset of each point cloud frame in the fused frame. The playback time offset of each point cloud frame is used to represent the offset of the playback time of the point cloud frame relative to the playback time of the fused frame.

14. The method according to claim 11, wherein The fused frame information further includes the number of fused frames of the point cloud data, the number of point cloud frames included in each fused frame in the point cloud data, and the time information of the fused frame further includes the time length of the fused frame; The media file further includes: the number of fused frames of the point cloud data, and the number of point cloud frames included in each fused frame in the point cloud data.

15. The method according to claim 11, wherein The media file is a container file of the point cloud data based on a geometrically encoded point cloud bitstream, and the compressed media file of the point cloud data is obtained according to the fused frame information and the point cloud data, including: The media file is obtained according to the fused frame information and the point cloud data through any packaging method of a single track, a multi-track, and a block-based track.

16. An electronic device comprising: memory and processor; The memory is coupled to the processor; The memory is used to store instructions executable by the processor; When the processor executes the instructions, the method according to any one of claims 1 to 15 is implemented.

17. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer is caused to perform the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Coding and decoding method of point cloud media and related products

    CN114697668A

  • Point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device

    US20240062428A1

  • Method, device, and computer program for enhancing encoding and encapsulation of point cloud data

    WO2023111214A1