Non-time-sequence point cloud media processing method and device, equipment and storage medium

By carrying the identification of static objects in non-temporal point cloud media and using GPCC encoding processing, the problem of users being unable to determine whether the point cloud media is the same object is solved, improving processing efficiency and user experience.

CN120807834APending Publication Date: 2025-10-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510942897.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2020-11-26
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

When users request to play point cloud media, they cannot determine whether different point cloud media are of the same object, resulting in low processing efficiency.

Method used

By carrying the identification of static objects in non-sequential point cloud media and processing point cloud data using GPCC encoding, GPCC bit streams and entries encapsulating GPCC areas are generated, non-sequential point cloud media is generated, and transmitted and requested through MDP signaling.

Benefits of technology

This enables users to request point cloud media of the same static object multiple times and in a targeted manner, improving processing efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807834A_ABST
    Figure CN120807834A_ABST
Patent Text Reader

Abstract

The invention provides a non-time-sequence point cloud media processing method and device, equipment and a storage medium, and the method comprises the steps: carrying out the processing of non-time-sequence point cloud data of a static object through a GPCC coding mode, and obtaining a GPCC bit stream; packaging the GPCC bit stream to generate an entry of at least one GPCC region; entries of the at least one GPCC area are packaged, and at least one non-time-sequence point cloud media of the static object is generated; sending at least one MDP signaling of the non-time-sequence point cloud media; receiving a first request message sent by the video playing device; sending the first non-time-sequence point cloud media; wherein the items of the GPCC area are used for representing GPCC components of a three-dimensional (3D) space area corresponding to the GPCC area; the non-time-sequence point cloud media comprise the identification of the static object, so that a user can request the non-point cloud media of the same static object in a targeted manner for multiple times, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention application with the application date of November 26, 2020, the Chinese application number of 202011347626.1, and the invention name of "processing method, device and equipment of non-timed point cloud media and storage medium". TECHNICAL FIELD

[0002] Embodiments of the present application relate to the technical field of computer technology, and in particular to a processing method, device and equipment of non-timed point cloud media and storage medium. BACKGROUND

[0003] At present, the point cloud data of an object can be obtained in many ways, and the video production equipment can transmit the point cloud data to the video playing equipment in the form of point cloud media, i.e., point cloud media files, so as to play the point cloud media by the video playing equipment.

[0004] It is worth mentioning that the point cloud data of the same object can be encapsulated into different point cloud media, for example, some point cloud media is the entire point cloud media of the object, and some point cloud media is only part of the point cloud media of the object. Based on this, the user can request to play different point cloud media, however, the user does not know whether the different point cloud media is the point cloud media of the same object when requesting, thereby causing the problem of low processing efficiency. SUMMARY

[0005] The present application provides a processing method, device and equipment of timed point cloud media and storage medium, so that the user can request the non-point cloud media of the same static object in multiple times and with purpose, so as to improve the processing efficiency and user experience.

[0006] In a first aspect, the present application provides a processing method of non-timed point cloud media, comprising: obtaining non-timed point cloud data of a static object; processing the non-timed point cloud data by GPCC encoding mode to obtain a GPCC bit stream; encapsulating the GPCC bit stream to generate an entry of at least one GPCC region; encapsulating the entry of at least one GPCC region to generate at least one non-timed point cloud media of the static object; sending MDP signaling of the at least one non-timed point cloud media to a video playing equipment; receiving a first request message sent by the video playing equipment; and sending a first non-timed point cloud media to the video playing equipment according to the first request message; wherein for any one of the entries of at least one GPCC region, the entry of the GPCC region is used to represent the GPCC component of the three-dimensional (3D) space region corresponding to the GPCC region; and for any one of the at least one non-timed point cloud media, the non-timed point cloud media comprises an identifier of the static object.

[0007] In a second aspect, the present application provides a non-timed point cloud media processing method, comprising: receiving MDP signaling of at least one non-timed point cloud media; sending a first request message to a video production device; receiving a first non-timed point cloud media; and playing the first non-timed point cloud media; wherein the at least one non-timed point cloud media is obtained by processing non-timed point cloud data of a static object by using a GPCC encoding mode, obtaining a GPCC bitstream, encapsulating the GPCC bitstream, generating entries of at least one GPCC region, encapsulating the entries of the at least one GPCC region, and obtaining the at least one non-timed point cloud media; for any one of the entries of the at least one GPCC region, the entry of the GPCC region is used to represent a GPCC component of a 3D space region corresponding to the GPCC region; and for any one of the at least one non-timed point cloud media, the non-timed point cloud media comprises an identifier of the static object.

[0008] In a third aspect, the present application provides a non-timed point cloud media processing apparatus, comprising: a processing unit and a communication unit; the processing unit is configured to: obtain non-timed point cloud data of a static object; process the non-timed point cloud data by using a GPCC encoding mode to obtain a GPCC bitstream; encapsulate the GPCC bitstream to generate entries of at least one GPCC region; encapsulate the entries of the at least one GPCC region to generate at least one non-timed point cloud media of the static object; and send MDP signaling of the at least one non-timed point cloud media to a video playing device; and the communication unit is configured to: receive a first request message sent by the video playing device; and send a first non-timed point cloud media to the video playing device according to the first request message; wherein for any one of the entries of the at least one GPCC region, the entry of the GPCC region is used to represent a GPCC component of a 3D space region corresponding to the GPCC region; and for any one of the at least one non-timed point cloud media, the non-timed point cloud media comprises an identifier of the static object.

[0009] In a fourth aspect, the present application provides a non-timed point cloud media processing device, comprising: a processing unit and a communication unit; the communication unit is configured to: receive MDP signaling of at least one non-timed point cloud media; send a first request message to a video production device; receive a first non-timed point cloud media; the processing unit is configured to play the first non-timed point cloud media; wherein the at least one non-timed point cloud media is obtained by processing non-timed point cloud data of a static object by using a GPCC encoding mode, obtaining a GPCC bitstream, encapsulating the GPCC bitstream, generating entries of at least one GPCC region, encapsulating the entries of the at least one GPCC region, and obtaining the at least one non-timed point cloud media; for any one of the entries of the at least one GPCC region, the entry of the GPCC region is used to represent a GPCC component of a 3D space region corresponding to the GPCC region; and for any one of the at least one non-timed point cloud media, the non-timed point cloud media comprises an identifier of the static object.

[0010] In a fifth aspect, a video production device is provided, comprising: a processor and a memory configured to store a computer program, wherein the processor is configured to invoke and run the computer program stored in the memory to execute the method of the first aspect.

[0011] In a sixth aspect, a video playing device is provided, comprising: a processor and a memory configured to store a computer program, wherein the processor is configured to invoke and run the computer program stored in the memory to execute the method of the second aspect.

[0012] In a seventh aspect, a computer readable storage medium is provided, configured to store a computer program, wherein the computer program causes a computer to execute the method of the first aspect.

[0013] In an eighth aspect, a computer readable storage medium is provided, configured to store a computer program, wherein the computer program causes a computer to execute the method of the second aspect.

[0014] In summary, in the present application, when encapsulating the non-timed point cloud media, the video production device can carry the identifier of the static object in the non-timed point cloud media, so that the user can request the non-timed point cloud media of the same static object in multiple times and with purpose, thereby improving the user experience.

[0015] Further, in the present application, the 3D space region corresponding to the entry of the GPCC region can be divided into a plurality of sub-space regions, and in combination with the independent coding and decoding characteristics of the GPCC tile, the efficiency of the user in decoding and presenting the non-timed point cloud media is higher and the time delay is lower.

[0016] Further, the video production device can flexibly combine entries of the plurality of GPCC regions to form different non-timed point cloud media, where the non-timed point cloud media can constitute a complete GPCC frame or a partial GPCC frame. Thus, the flexibility of video production can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort.

[0018] Figure 1 An architecture schematic diagram of a non-timed point cloud media processing system provided by an example embodiment of the present application is shown;

[0019] Figure 2A An architecture schematic diagram of a non-timed point cloud media processing architecture provided by an example embodiment of the present application is shown;

[0020] Figure 2B A structure schematic diagram of a sample provided by an example embodiment of the present application is shown;

[0021] Figure 2C A structure schematic diagram of a container containing a plurality of file tracks provided by an example embodiment of the present application is shown;

[0022] Figure 2D A structure schematic diagram of a sample provided by another example embodiment of the present application is shown;

[0023] Figure 3 An interaction flowchart of a non-timed point cloud media processing method provided by an example embodiment of the present application is shown;

[0024] Figure 4A An encapsulation schematic diagram of a point cloud media provided by an example embodiment of the present application is shown;

[0025] Figure 4B An encapsulation schematic diagram of another point cloud media provided by an example embodiment of the present application is shown;

[0026] Figure 5 A schematic diagram of a non-timed point cloud media processing apparatus 500 provided by an example embodiment of the present application is shown;

[0027] Figure 6 A schematic diagram of a non-timed point cloud media processing apparatus 600 provided by an example embodiment of the present application is shown;

[0028] Figure 7is a schematic block diagram of a video production device 700 provided by an embodiment of the present application.

[0029] Figure 8 is a schematic block diagram of a video playing device 800 provided by an embodiment of the present application. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be clearly and completely described in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.

[0031] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product, or device.

[0032] Before introducing the technical solutions of the present application, the related knowledge of the present application will be introduced as follows:

[0033] The so-called point cloud refers to a set of discrete points that are irregularly distributed in space and express the spatial structure and surface attributes of a three-dimensional object or a three-dimensional scene. Point cloud data is a specific recording form of the point cloud. The point cloud data of each point in the point cloud can include geometric information (i.e., three-dimensional position information) and attribute information. The geometric information of each point in the point cloud refers to the Cartesian three-dimensional coordinate data of the point. The attribute information of each point in the point cloud can include, but is not limited to, at least one of the following: color information, material information, and laser reflection intensity information. Generally, each point in the point cloud has the same number of attribute information. For example, each point in the point cloud has two attribute information of color information and laser reflection intensity. Or each point in the point cloud has three attribute information of color information, material information, and laser reflection intensity.

[0034] With the progress and development of science and technology, a large amount of high-precision point cloud data can be obtained at a lower cost and in a shorter time period. The acquisition approach of point cloud data can include but is not limited to at least one of the following: (1) computer device generation. The computer device can generate point cloud data according to a virtual three-dimensional object and a virtual three-dimensional scene. (2) 3D (3-Dimension, three-dimensional) laser scanning acquisition. Through 3D laser scanning, point cloud data of a static real-world three-dimensional object or three-dimensional scene can be obtained, and million-level point cloud data can be obtained per second. (3) 3D photogrammetry acquisition. Through 3D photographic equipment (i.e., a group of cameras or a camera device with multiple lenses and sensors), a visual scene of the real world is collected to obtain point cloud data of the visual scene of the real world. Through 3D photography, point cloud data of a dynamic real-world three-dimensional object or three-dimensional scene can be obtained. (4) Point cloud data of biological tissue organs is obtained through medical equipment. In the medical field, point cloud data of biological tissue organs can be obtained through medical equipment such as magnetic resonance imaging (Magnetic Resonance Imaging, MRI), computed tomography (Computed Tomography, CT), and electromagnetic positioning information.

[0035] The so-called point cloud media refers to a point cloud media file formed by point cloud data. The point cloud media includes a plurality of media frames, and each media frame in the point cloud media is composed of point cloud data. The point cloud media can flexibly and conveniently express the spatial structure and surface properties of a three-dimensional object or a three-dimensional scene, and is therefore widely used in virtual reality (Virtual Reality, VR) games, computer-aided design (Computer Aided Design, CAD), geography information systems (Geography Information System, GIS), autonomous navigation systems (Autonomous Navigation System, ANS), digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive remote presentation, three-dimensional reconstruction of biological tissue organs, and the like.

[0036] The so-called non-sequential point cloud media is for the same static object, that is, for the same static object, the corresponding point cloud media is non-sequential.

[0037] Based on the above description, please refer to Figure 1 , Figure 1An architecture diagram of a non-timed point cloud media processing system provided by an example embodiment of the present application is shown, which includes a video playing device 101 and a video producing device 102. The video producing device refers to a computer device used by a non-timed point cloud media provider (for example, a non-timed point cloud media content producer), which can be a terminal (for example, a personal computer (PC), a smart mobile device (for example, a smart phone), etc.), a server, etc.; the video playing device refers to a computer device used by a non-timed point cloud media user (for example, a user), which can be a terminal (for example, a PC), a smart mobile device (for example, a smart phone), a VR device (for example, a VR helmet, a VR glasses), etc. The video producing device and the video playing device can be directly or indirectly connected through wired communication or wireless communication, which is not limited in the embodiments of the present application.

[0038] Figure 2A An architecture diagram of a non-timed point cloud media processing architecture provided by an example embodiment of the present application is shown, and the non-timed point cloud media processing scheme provided by the embodiments of the present application will be introduced below in combination with Figure 1 the non-timed point cloud media processing system shown in Figure 2A the non-timed point cloud media processing architecture shown, the non-timed point cloud media processing process includes a processing process on the video producing device side and a processing process on the video playing device side, and the specific processing process is as follows:

[0039] I. Processing process on the video producing device side:

[0040] (1) Point cloud data acquisition process.

[0041] In an implementation, from the perspective of the acquisition manner of the point cloud data, the acquisition manner of the point cloud data can be divided into two manners: one is to acquire the point cloud data by capturing a visual scene in a real world by a capturing device, and the other is to generate the point cloud data by a computer device. In an implementation, the capturing device can be a hardware component arranged in a video production device, for example, the capturing device is a camera, a sensor, etc. of a terminal. The capturing device can also be a hardware device connected to the content production device, for example, a camera connected to a server. The capturing device is used to provide the point cloud data acquisition service for the video production device, and the capturing device can include but is not limited to any one of the following: a camera device, a sensor device, a scanning device; wherein the camera device can include an ordinary camera, a stereo camera, a light field camera, etc.; the sensor device can include a laser device, a radar device, etc.; the scanning device can include a 3D laser scanning device, etc. The number of capturing devices can be multiple, and these capturing devices are deployed at some specific positions in a real space to simultaneously capture point cloud data at different angles in the space, and the captured point cloud data is kept synchronous in time and space. In another implementation, the computer device can generate point cloud data according to a virtual three-dimensional object and a virtual three-dimensional scene. Because the acquisition manners of the point cloud data are different, the compression encoding manners corresponding to the point cloud data acquired by different manners can also be different.

[0042] (2) Point cloud data encoding and encapsulation process.

[0043] In an implementation, the video production device encodes the acquired point cloud data by using a geometry-based point cloud compression (GPCC) encoding manner or a video-based point cloud compression (VPCC) encoding manner, to obtain a GPCC bitstream or a VPCC bitstream of the point cloud data.

[0044] In an implementation, taking the GPCC encoding manner as an example, the video production device encapsulates the GPCC bitstream of the encoded point cloud data by using a file track. The file track refers to an encapsulation container of the GPCC bitstream of the encoded point cloud data. The GPCC bitstream can be encapsulated in a single file track, and the GPCC bitstream can also be encapsulated in multiple file tracks. The specific conditions of the GPCC bitstream encapsulated in the single file track and the GPCC bitstream encapsulated in the multiple file tracks are as follows:

[0045] 1. The GPCC bitstream is encapsulated in a single file track. When the GPCC bitstream is transmitted in a single file track, the GPCC bitstream is required to be declared and represented according to the transmission rules of the single file track. The GPCC bitstream encapsulated in a single file track does not require further processing and can be encapsulated using the International Organization for Standardization Base Media File Format (ISOBMFF). Specifically, each sample (Sample) encapsulated in a single file track contains one or more GPCC components, which are also called GPCC components. The GPCC components can be GPCC geometric components or GPCC attribute components. The so-called sample refers to a set of encapsulation structures of one or more point clouds, that is, each sample consists of one or more Type-Length-Value ByteStream Format (TLV) encapsulation structures. Figure 2B A schematic diagram of a sample structure provided by an exemplary embodiment of the present application is shown in FIG. Figure 2B As shown, when a single file track is transmitted, the sample in the file track consists of a GPCC parameter set TLV, a geometry bitstream TLV, and an attribute bitstream TLV, and the sample is encapsulated into a single file track.

[0046] 2. GPCC bitstreams are encapsulated in multiple file tracks. When the coded GPCC geometry bitstream and the coded GPCC attribute bitstream are transmitted in different file tracks, each sample in the file track contains at least one TLV encapsulation structure that carries a single GPCC component data, and the TLV encapsulation structure does not contain the coded GPCC geometry bitstream and the coded GPCC attribute bitstream at the same time. Figure 2C FIG. 1 shows a schematic diagram of a container structure including multiple file tracks provided by an exemplary embodiment of the present application. Figure 2C As shown, encapsulated package 1 transmitted in file track 1 contains the encoded GPCC geometry bitstream but does not contain the encoded GPCC attribute bitstream; encapsulated package 2 transmitted in file track 2 contains the encoded GPCC attribute bitstream but does not contain the encoded GPCC geometry bitstream. Because the video playback device should first decode the encoded GPCC geometry bitstream during decoding, and the decoding of the encoded GPCC attribute bitstream depends on the decoded geometry information, different GPCC component bitstreams are encapsulated in separate file tracks, allowing the video playback device to access the file track carrying the encoded GPCC geometry bitstream before the encoded GPCC attribute bitstream. Figure 2DA structure diagram of a sample provided by another example embodiment of the application is shown as follows: Figure 2D As shown, when multiple file track transmissions are performed, the encoded GPCC geometry bitstream and the encoded GPCC attribute bitstream are transmitted in different file tracks, the sample in the file track is composed of a GPCC parameter set TLV and a geometry bitstream TLV, and the sample does not contain an attribute bitstream TLV, and the sample is encapsulated in any one of the multiple file tracks.

[0047] In an implementation, the obtained point cloud data is encoded and encapsulated by the video production device to form a non-timed point cloud media, which can be an entire media file of an object or a media segment of the object; and the video production device records metadata of an encapsulated file of the non-timed point cloud media by using a media presentation description information (i.e., a description signaling file) (Media presentation description, MPD) according to a file format requirement of the non-timed point cloud media. The metadata is a general term of information related to presentation of the non-timed point cloud media, and can include description information of the non-timed point cloud media, description information of a view window, and signaling information related to presentation of the non-timed point cloud media, and the like. The video production device delivers the MPD to the video playing device, so that the video playing device requests to obtain the point cloud media according to related description information in the MPD. Specifically, the point cloud media and the MPD are delivered by the video production device to the video playing device through a transmission mechanism (for example, Dynamic Adaptive Streaming over HTTP (DASH), Smart Media Transport (SMT)).

[0048] II. Data processing process of the video playing device side:

[0049] (1) Point cloud data unencapsulation and decoding process.

[0050] In an implementation, the video playing device can obtain the non-timed point cloud media by using the MPD signaling delivered by the video production device. The file unencapsulation process of the video playing device is inverse to the file encapsulation process of the video production device. The video playing device performs unencapsulation on the encapsulated file of the non-timed point cloud media according to the file format requirement of the non-timed point cloud media, to obtain an encoded bitstream (i.e., a GPCC bitstream or a VPCC bitstream). The decoding process of the video playing device is inverse to the encoding process of the video production device. The video playing device decodes the encoded bitstream to restore the point cloud data.

[0051] (2) Point cloud data rendering process.

[0052] In an implementation manner, the video playing device renders the point cloud data decoded from the GPCC bitstream according to the metadata in the MDP related to rendering and a window, and the rendering is completed, that is, the presentation of the visual scene corresponding to the point cloud data is realized.

[0053] It can be understood that the processing system of the non-time sequence point cloud media described in the embodiments of the present application is for more clearly illustrating the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. It can be known by those skilled in the art that, with the evolution of system architecture and the appearance of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0054] As described above, for the point cloud data of the same object, different point cloud media can be encapsulated, for example, some point cloud media is the entire point cloud media of the object, and some point cloud media is the partial point cloud media of the object. Based on this, the user can request to play different point cloud media, however, the user does not know whether the different point cloud media is the point cloud media of the same object when requesting, thereby causing the problem of blind request. This problem also exists for the non-point cloud media of the static object.

[0055] In order to solve the above technical problems, the present application carries the identifier of the static object in the non-point cloud media, so that the user can request the non-point cloud media of the same static object in multiple times and with purpose.

[0056] The technical solutions of the present application will be described in detail as follows:

[0057] Embodiment 1

[0058] Figure 3 An interactive flowchart of a non-time sequence point cloud media processing method provided by the embodiments of the present application is shown in FIG. 1, the execution subject of the method is a video production device and a video playing device, as shown in FIG. 1, the method comprises the following steps: Figure 3 As shown in FIG. 1, the method comprises the following steps:

[0059] S301: The video production device acquires the non-time sequence point cloud data of the static object.

[0060] S302: The video production device processes the non-time sequence point cloud data by the GPCC encoding mode, to obtain a GPCC bitstream.

[0061] S303: The video production device encapsulates the GPCC bitstream, to generate an entry of at least one GPCC region.

[0062] S304: The video production device encapsulates the entry of the at least one GPCC region, to generate at least one non-time sequence point cloud media of the static object, and each non-time sequence point cloud media comprises an identifier of the static object.

[0063] S305: The video production device sends MDP signaling of the at least one non-timed point cloud media to the video playback device.

[0064] S306: The video playback device sends a first request message.

[0065] S307: The video production device sends the first non-timed point cloud media to the video playback device according to the first request message.

[0066] S308: The video playback device plays the first non-timed point cloud media.

[0067] It should be understood that how to obtain the non-timed point cloud data of the static object and obtain the GPCC bitstream can refer to the above-mentioned related knowledge, and the present application will not repeat it here.

[0068] It should be understood that the item of any one of the at least one GPCC region in the item of the GPCC region is used to represent the GPCC component of the 3D space region corresponding to the GPCC region.

[0069] It should be understood that each GPCC region corresponds to a 3D space region of the static object, which can be the entire or partial 3D space region of the static object.

[0070] It should be understood that, as described above, the GPCC component is also referred to as a GPCC component, and the GPCC component can be a GPCC geometry component or an attribute component.

[0071] It should be understood that on the video production device side, the identifier of the static object can be defined by the following code:

[0072] aligned(8) class ObjectInfoProperty extends ItemProperty('obif') {

[0073] unsigned int(32) object_ID;}

[0074] Wherein, ObjectInfoProperty indicates the attribute of the content corresponding to the item, and both the GPCC geometry component and the attribute component can contain the attribute. If only the GPCC geometry component contains the attribute, the ObjectInfoProperty of all attribute components associated with the GPCC geometry component is the same as that of the GPCC geometry component.

[0075] object_ID indicates the identifier of the static object, and the object_ID of different items of the same static object is the same.

[0076] Optionally, the identification of the above-mentioned static object can be carried in the entry related to the GPCC geometric component in the point cloud media, or carried in the entry related to the GPCC attribute component in the point cloud media, or carried in the entry related to the GPCC geometric component in the point cloud media, and the entry related to the GPCC attribute component. This application does not impose any restrictions on this.

[0077] For example, Figure 4A A schematic diagram of a point cloud media package provided in an embodiment of the present application is shown as follows: Figure 4A As shown, the point cloud media includes: entries related to GPCC geometric components and entries related to GPCC attribute components. These entries can be associated through the GPCC entry group box in the point cloud media. Figure 4A As shown, entries related to GPCC geometry components are associated with entries related to GPCC attribute components. Entries related to GPCC geometry components may include the following entry attributes: GPCC Configuration, 3D spatial region attributes (3D spatial region or ItemSpatialInfoProperty), and static object identifiers. Entries related to GPCC attribute components may include the following entry attributes: GPCC Configuration, static object identifiers, etc.

[0078] Optionally, the GPCC configuration indicates configuration information of a decoder required to decode a corresponding entry and information related to each GPCC component, but is not limited thereto.

[0079] It is worth mentioning that the entries related to the GPCC attribute components may also include: 3D spatial area attributes, which is not limited in this application.

[0080] For example, Figure 4B This is a schematic diagram of another packaging of point cloud media provided in an embodiment of the present application. Figure 4B and Figure 4A The difference is: in Figure 4B In the point cloud media, the point cloud media includes: a GPCC geometric component related entry, and the entry is associated with two GPCC attribute component related entries. The rest of the attributes included in the GPCC geometric component related entries and the attributes included in the GPCC attribute component related entries can be referred to Figure 4A , this application will not go into details about this.

[0081] It should be understood that the identification of the static object is not limited to being carried in the attributes corresponding to the entry of each GPCC area.

[0082] It should be understood that the above-mentioned MDP signaling can refer to the relevant knowledge mentioned above in this application, and this application will not elaborate on it again.

[0083] Optionally, for any one of the at least one non-time-sequential point cloud media, the non-time-sequential point cloud media may be the entire or partial point cloud media of the static object.

[0084] It should be understood that the video playback device may send a first request message according to the above-mentioned MDP signaling to request the first non-time-sequential point cloud media.

[0085] In summary, in this application, when encapsulating non-sequential point cloud media, the video production device can carry the identification of the static object in the non-sequential point cloud media, so that the user can request the non-point cloud media of the same static object multiple times and in a purposeful manner to improve the user experience.

[0086] Example 2

[0087] In the prior art, each GPCC region entry corresponds to only one 3D spatial region. However, in this application, the 3D spatial region can be further divided. Based on this, this application updates the entry attributes and MDP signaling in the non-sequential point cloud media accordingly, as follows:

[0088] Optionally, the entry for the first GPCC region includes: a 3D spatial region entry attribute, wherein the 3D spatial region entry attribute includes: a first identifier and a second identifier. The first GPCC region is one of at least one GPCC region. The first identifier (sub_region_contained) is used to identify whether the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-space regions. The second identifier (tile_id_present) is used to identify whether the first GPCC region uses the GPCC tile encoding method.

[0089] Exemplarily, when Sub_region_contained=0, it indicates that the first 3D spatial region corresponding to the first GPCC region is not divided into multiple sub-space regions; when Sub_region_contained=1, it indicates that the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-space regions.

[0090] For example, when tile_id_present=0, it indicates that the first GPCC region does not adopt the GPCC tile coding mode. When tile_id_present=1, it indicates that the first GPCC region adopts the GPCC tile coding mode.

[0091] It should be understood that when Sub_region_contained = 1, tile_id_present = 1, i.e. when the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-space regions, the video production end must use GPCC tile encoding mode.

[0092] Optionally, if the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-space regions, the 3D spatial region entry attribute further includes, but is not limited to, the following: information of each of the multiple sub-space regions and information of the first 3D spatial region.

[0093] Optionally, for any one of the multiple sub-space regions, the information of the sub-space region includes at least one of the following, but is not limited to: an identifier of the sub-space region, position information of the sub-space region, and tile (block) identifier in the sub-space region when the first GPCC region uses GPCC tile encoding.

[0094] Optionally, the position information of the sub-space region includes, but is not limited to, the following: position information of one anchor point of the sub-space region, and lengths of the sub-space region along the X-axis, the Y-axis, and the Z-axis respectively. Alternatively, the position information of the sub-space region includes, but is not limited to, the following: position information of two anchor points of the sub-space region.

[0095] Optionally, the information of the first 3D spatial region includes at least one of the following, but is not limited to: an identifier of the first 3D spatial region, position information of the first 3D spatial region, and a number of sub-space regions included in the first 3D spatial region.

[0096] Optionally, the position information of the first 3D spatial region includes, but is not limited to, the following: position information of one anchor point of the first 3D spatial region, and lengths of the first 3D spatial region along the X-axis, the Y-axis, and the Z-axis respectively. Alternatively, the position information of the first 3D spatial region includes, but is not limited to, the following: position information of two anchor points of the first 3D spatial region.

[0097] Optionally, if the first 3D space region corresponding to the first GPCC region is divided into a plurality of sub-space regions, the 3D space region item attribute further comprises a third identifier (initial_region_id). When the third identifier takes a first value or is empty, it indicates that the item corresponding to the first GPCC is the item initially presented by the video playback device. For the first 3D space region and the sub-space regions of the first 3D space region, the video playback device initially presents the first 3D space region. When the third identifier takes a second value, it indicates that the item corresponding to the first GPCC is the item initially presented by the video playback device. For the first 3D space region and the sub-space regions of the first 3D space region, the video playback device initially presents the sub-space region corresponding to the second value in the first 3D space region.

[0098] Optionally, the first value is 0, and the second value is the identifier of the sub-space region in the first 3D space region that needs to be initially presented.

[0099] Optionally, if the first 3D space region corresponding to the first GPCC region is not divided into a plurality of sub-space regions, the 3D space region item attribute further comprises information of the first 3D space region. Optionally, the information of the first 3D space region comprises at least one of the following, but is not limited thereto: an identifier of the first 3D space region, position information of the first 3D space region, and a tile identifier in the first 3D space region when the first GPCC region is encoded using GPCC tile.

[0100] It should be understood that the position information of the first 3D space region can be referred to the above explanation of the position information of the first 3D space region, which will not be repeated here.

[0101] The update of the item attribute in the non-time sequence point cloud media will be illustrated in the form of code as follows:

[0102]

[0103]

[0104] The semantics of each field is as follows:

[0105] ItemSpatialInfoProperty represents the 3D space region attribute of the item of the GPCC region. If the item is the item corresponding to the geometry component, this attribute must be included; if the item is the item corresponding to the attribute component, the 3D space region attribute can not be included.

[0106] sub_region_contained is 1, indicating that the 3D spatial region can be further divided into multiple sub-space regions. When sub_region_contained is 1, tile_id_present must be 1. sub_region_contained is 0, indicating that the 3D spatial region is not further divided into sub-space regions.

[0107] tile_id_present is 1, indicating that the non-timed point cloud data is encoded using GPCC tile, and the tile id corresponding to the non-timed point cloud is given in this attribute.

[0108] inital_region_id indicates the ID of the initial presented space region inside the whole space region of the current entry when the current entry is the initial consumption or playback entry. If the field value is 0 or the field does not exist, the initial presented region of the entry is the whole 3D space region. If the field value is the identifier of a sub-space region, the initial presented region of the entry is the sub-space region corresponding to the identifier.

[0109] 3DSpatialRegionStruct indicates a 3D spatial region. The first 3DSpatialRegionStruct in ItemSpatialInfoProperty indicates the 3D spatial region corresponding to the entry corresponding to ItemSpatialInfoProperty, and the remaining 3DSpatialRegionStruct indicates each sub-space region in the 3D spatial region corresponding to the entry.

[0110] num_sub_regions indicates the number of sub-space regions divided in the 3D spatial region corresponding to the entry.

[0111] num_tiles indicates the number of tiles in the 3D spatial region corresponding to the entry, or the number of tiles corresponding to the sub-space region thereof.

[0112] tile_id indicates the identifier of the GPCC tile.

[0113] anchor_x, anchor_y, and anchor_z respectively indicate the x, y, and z coordinates of the anchor point of the 3D spatial region or the sub-space region of the region.

[0114] region_dx, region_dy, and region_dz respectively indicate the lengths of the 3D spatial region or the sub-space region of the region along the X-axis, Y-axis, and Z-axis.

[0115] In summary, in the present application, the 3D space region can be divided into multiple sub-space regions, and in combination with the characteristics of GPCC tile independent coding and decoding, the user can decode and present the non-timed point cloud media more efficiently and with lower latency.

[0116] Embodiment 3

[0117] As described above, the video production end can encapsulate the entries of at least one GPCC region to generate at least one non-timed point cloud media of a static object. If the entries of the at least one GPCC region are 1, the entries of the 1 GPCC region are encapsulated into 1 non-timed point cloud media. If the entries of the at least one GPCC region are N, the entries of the N GPCC regions are encapsulated into M non-timed point cloud media. N is an integer greater than 1, the value range of M is [1, N], and M is an integer. For example, if the entries of the at least one GPCC region are N, the N GPCC region entries can be encapsulated into 1 non-timed point cloud media, or encapsulated into N non-timed point cloud media, each of which includes one entry.

[0118] The fields in the second non-timed point cloud media will be described below. The second non-timed point cloud media is any one of the non-timed point cloud media including multiple entries of GPCC regions in the at least one non-timed point cloud media.

[0119] Optionally, the second non-timed point cloud media includes:

[0120] GPCC entry group box (GPCCItemGroupBox). The GPCC entry group box is used to associate multiple entries of GPCC regions, as shown in Figure 4A and 4B .

[0121] Optionally, the GPCC entry group box includes: the identification of the multiple entries of GPCC regions.

[0122] Optionally, the GPCC entry group box includes: a fourth identification (initial_item_ID). The fourth identification is the identification of the entry that is initially presented by the video playback device among the multiple entries of GPCC regions.

[0123] Optionally, the GPCC entry group box includes: a fifth identification (partial_item_flag). If the fifth identification takes the third value, it indicates that the multiple entries of GPCC regions constitute a complete GPCC frame of a static object. If the fifth identification takes the fourth value, it indicates that the multiple entries of GPCC regions constitute a partial GPCC frame of a static object.

[0124] Optionally, the third value can be 0, and the fourth value can be 1, but not limited thereto.

[0125] Optionally, the GPCC item group box includes location information of the GPCC region composed of the plurality of GPCC regions.

[0126] Exemplarily, if the plurality of GPCC regions are two regions R1 and R2, the GPCC item group box includes location information of the R1+R2 region.

[0127] The fields in the above GPCC item group box will be described by codes as follows:

[0128]

[0129] The entries contained in the GPCCItemGroupBox are entries belonging to the same static object, and the entries have a correlation relationship in rendering consumption. All the entries contained in the GPCCItemGroupBox can constitute a complete GPCC frame, or can be a part of a GPCC frame.

[0130] The initial_item_ID indicates the identification of an entry initially consumed in an entry group.

[0131] It should be noted that the initial_item_ID is only valid when the current entry group is an entry group initially requested by a user, for example, two point cloud media F1 and F2 correspond to the same static object, when a user first requests F1, the initial_item_ID in the entry group in F1 is valid, and the initial_item_ID in the entry group in F2 requested for the second time is invalid.

[0132] When the partial_item_flag is 0, it indicates that all the entries contained in the GPCCItemGroupBox and the associated entries constitute a complete GPCC frame, and when the partial_item_flag is 1, it indicates that all the entries contained in the GPCCItemGroupBox and the associated entries only constitute part of a GPCC frame.

[0133] To support the technology proposed in the present application, the corresponding signaling message also needs to be extended, for example, the MDP signaling is extended as follows:

[0134] The GPCC entry descriptor is used to describe the elements and attributes related to the GPCC entry, and the descriptor is a SupplementalProperty element.

[0135] The schemeIdUri attribute is equal to "urn:mpeg:mpegI:gpcc:2020:gpsr". The descriptor can be located at the Adaptation Set level or the Representation level.

[0136] Representation: In DASH, a combination of one or more media components, such as a video file of a certain resolution can be regarded as a Representation.

[0137] Adaptation Sets: In DASH, a set of one or more video streams, and an Adaptation Sets can contain multiple Representations.

[0138] Table 1: GPCC entry descriptor element and attribute

[0139]

[0140]

[0141]

[0142]

[0143] In summary, in the present application, the video production device can flexibly combine the entries of multiple GPCC regions to form different non-sequential point cloud media, wherein the non-sequential point cloud media can constitute a complete GPCC frame or a partial GPCC frame. Thus, the flexibility of video production can be improved. Further, when a non-sequential point cloud media includes entries of GPCC regions, the video production device can also improve the initial presentation of the entries.

[0144] Embodiments 1 to 3 will be illustrated by Embodiments 4 and 5 as follows:

[0145] Embodiment 4

[0146] Suppose that the video production device obtains non-sequential point cloud data of a certain static object, and there are four versions of point cloud media of the non-sequential point cloud data at the video production device end: a point cloud media F0 corresponding to all the non-sequential point cloud data, and point cloud media F1 to F3 corresponding to partial non-sequential point cloud data, wherein F1 to F3 correspond to 3D space regions R1 to R3 respectively. Based on this, the point cloud media F0 to F3 encapsulate the following contents:

[0147] F0: ObjectlnfoProperty: object_ID = 10;

[0148] ItemSpatialInfoProperty: sub_region_contained = 1 ; tile_id_present = 1

[0149] inital_region_id = 1001 ;

[0150] R0: 3d_region_id = 100, anchor = (0,0,0), region = (200,200,200)

[0151] num_sub_regions = 3;

[0152] SR1 : 3d_region_id = 1001, anchor = (0,0,0), region = (100,100,200);

[0153] num_tiles = 1, tile_id[] = (1);

[0154] SR2: 3d_region_id = 1002, anchor = (100,0,0), region = (100,100,200);

[0155] num_tiles = 1, tile_id[] = (2);

[0156] SR3: 3d_region_id = 1003, anchor = (0,100,0), region = (200,100,200);

[0157] num_tiles = 2, tile_id[] = (3,4);

[0158] F1 : ObjectInfoProperty: object_ID = 10;

[0159] ItemSpatialInfoProperty: sub_region_contained = 0; tile_id_present = 1;

[0160] inital_region_id = 0;

[0161] R1 : 3d_region_id = 101, anchor = (0,0,0), region = (100,100,200);

[0162] num_tiles = 1, tile_id[] = (1);

[0163] F2: ObjectlnfoProperty: object_ID = 10;

[0164] ItemSpatiallnfoProperty: sub_region_contained = 0; tile_id_present = 1

[0165] inital_region_id = 0;

[0166] R2: 3d_region_id = 102, anchor = (100, 0, 0), region = (100, 100, 200);

[0167] num_tiles = 1, tile_id[] = (2);

[0168] F3: ObjectlnfoProperty: object_ID = 10;

[0169] ItemSpatiallnfoProperty: sub_region_contained = 0; tile_id_present = 1;

[0170] inital_region_id = 0;

[0171] R3: 3d_region_id = 103, anchor = (0, 100, 0), region = (200, 100, 200);

[0172] num_tiles = 2, tile_id[] = (3, 4);

[0173] Further, the video production device sends the MPD signaling of F0-F3 to the user, wherein the Object ID, spatial region, sub-spatial region, and tile identification information are the same as in the file packaging, and will not be described here.

[0174] The network condition of user U1 is good, and the data transmission delay is low, so F0 can be requested; the network condition of user U2 is poor, and the data transmission delay is high, so F1 can be requested.

[0175] The video production device transmits F0 to the video playing device corresponding to user U1, and transmits F1 to the video playing device corresponding to user U2.

[0176] The video playback device corresponding to the user U1 receives F0, and the initial viewing area is SR1 area, and the corresponding tile ID is 1. When U1 decodes and consumes, the tile'1' can be directly decoded and consumed from the overall code stream, without the need to decode the overall file for presentation, thereby improving decoding efficiency and reducing the time required for rendering and presentation. When U1 continues to consume and views the SR2 area, the corresponding tile ID is 2, and the tile'2' corresponding to the part in the overall code stream is directly decoded and presented for consumption.

[0177] After the video playback device corresponding to the user U2 receives F1, F1 is decoded for consumption, and F2 or F3 is requested in advance for caching according to the area that the user is likely to consume next and in combination with the information in the MPD file, that is, the Object ID and the spatial area information.

[0178] Embodiment 5

[0179] It is assumed that a video production device obtains non-time sequence point cloud data of a static object, and the non-time sequence point cloud data has two versions of point cloud media F1 and F2 on the video production device. F1 contains item1 to item2, and F2 contains item3 to item4.

[0180] The point cloud media encapsulation contents of F1 and F2 are as follows:

[0181] F1:

[0182] item1: ObjectInfoProperty: object_ID=10; item_ID=101

[0183] ItemSpatialInfoProperty: sub_region_contained=0; tile_id_present=1

[0184] inital_region_id=0;

[0185] R1: 3d_region_id=1001, anchor=(0,0,0), region=(100,100,200);

[0186] num_tiles=1, tile_id[]=(1);

[0187] item2: ObjectInfoProperty: object_ID=10; item_ID=102

[0188] ItemSpatiallnfoProperty: sub_region_contained = 0; tile_id_present = 1

[0189] initial_region_id = 0;

[0190] R2: 3d_region_id = 1002, anchor = (100,0,0), region = (100,100,200);

[0191] num_tiles = 1, tile_id[] = (2);

[0192] GPCCItemGroupBox:

[0193] initial_item_ID = 101; partial_item_flag = 1;

[0194] R1+R2: 3d_region_id = 0001, anchor = (0,0,0), region = (200,100,200);

[0195] F2:

[0196] item3: ObjectlnfoProperty: object_ID = 10; item_ID = 103

[0197] ItemSpatiallnfoProperty: sub_region_contained = 0; tile_id_present = 1

[0198] initial_region_id = 0;

[0199] R3: 3d_region_id = 1003, anchor = (0,100,0), region = (100,100,200);

[0200] num_tiles = 1, tile_id[] = (3);

[0201] item4: ObjectlnfoProperty: object_ID = 10; item_ID = 104

[0202] ItemSpatiallnfoProperty: sub_region_contained = 0; tile_id_present = 1

[0203] inital_region_id = 0;

[0204] R4: 3d_region_id = 1004, anchor = (100, 100, 0), region = (100, 100, 200);

[0205] num_tiles = 1, tile_id[] = (4);

[0206] GPCCItemGroupBox:

[0207] initial_item_ID = 103; partial_item_flag = 1;

[0208] R3+R4: 3d_region_id = 0002, anchor = (0, 100, 0), region = (200, 100, 200);

[0209] The video production device sends the MPD signaling of F1-F2 to the user, wherein the Object ID, spatial region, and tile ID information are the same as in the point cloud media encapsulation, and will not be described again here.

[0210] The user U1 requests F1 consumption; and the user U2 requests F2 consumption.

[0211] The video production device respectively transmits F1 to the video playback device corresponding to the user U1, and transmits F2 to the video playback device corresponding to the user U2.

[0212] After the video playback device corresponding to U1 receives F1, the initial viewing item is item1, and the initial viewing region of item1 is the whole viewing space of item1, so U1 consumes the whole item1. Since item1 and item2 are included in F1, and correspond to tile1 and tile2 respectively, U1 can directly decode the part of the code stream corresponding to tile1 to present when consuming item1. If U1 continues to consume and watches the region of item2, the corresponding tile ID is 2, then the part corresponding to tile'2' in the whole code stream is directly decoded to present and consume. If U1 continues to consume and needs to watch the region corresponding to item3, then F2 is requested according to the MPD file. After receiving F2, the presentation and consumption are directly performed according to the region watched by the user, and the initial consumption item information and the initial viewing region information in F2 are no longer judged.

[0213] The video playback device corresponding to U2 receives F2, and initially watches item 3. The initial watching area of item 3 is the whole watching space of item 3, and therefore U2 consumes the whole item 3. Since item 3 and item 4 are included in F2, corresponding to tile 3 and tile 4 respectively, U2 can directly decode the part of the code stream corresponding to tile 3 to present when consuming item 3.

[0214] Embodiment 6

[0215] Figure 5 A schematic diagram of a non-timed point cloud media processing apparatus 500 provided by the embodiments of the present application is provided. The apparatus 500 includes a processing unit 510 and a communication unit 520. The processing unit 510 is configured to: obtain non-timed point cloud data of a static object. Process the non-timed point cloud data by using a GPCC encoding mode to obtain a GPCC bitstream. Package the GPCC bitstream to generate an entry of at least one GPCC region. Package the entry of the at least one GPCC region to generate at least one non-timed point cloud media of the static object. Send MDP signaling of the at least one non-timed point cloud media to a video playback device. The communication unit 520 is configured to: receive a first request message sent by the video playback device. According to the first request message, send a first non-timed point cloud media to the video playback device. For any one of the entries of the at least one GPCC region, the entry of the GPCC region is used to represent a GPCC component of a 3D space region corresponding to the GPCC region. For any one of the at least one non-timed point cloud media, the non-timed point cloud media includes an identifier of the static object.

[0216] Optionally, the entry of the first GPCC region includes a 3D space region entry attribute, and the 3D space region entry attribute includes a first identifier and a second identifier. The first GPCC region is one of the at least one GPCC region. The first identifier is used to identify whether the first 3D space region corresponding to the first GPCC region is divided into a plurality of sub-space regions. The second identifier is used to identify whether the first GPCC region adopts a GPCC tile encoding mode.

[0217] Optionally, if the first 3D space region corresponding to the first GPCC region is divided into a plurality of sub-space regions, the 3D space region entry attribute further includes information of the plurality of sub-space regions and information of the first 3D space region.

[0218] Optionally, for any one of the plurality of sub-space regions, the information of the sub-space region comprises at least one of: an identifier of the sub-space region, position information of the sub-space region, and a tile identifier in the sub-space region when the first GPCC region is encoded by using GPCC tile coding. The information of the first 3D space region comprises at least one of: an identifier of the first 3D space region, position information of the first 3D space region, and a number of sub-space regions included in the first 3D space region.

[0219] Optionally, if the first 3D space region corresponding to the first GPCC region is divided into a plurality of sub-space regions, the 3D space region entry attribute further comprises a third identifier. When the third identifier takes a first value or is empty, it indicates that the entry corresponding to the first GPCC is an entry initially presented by the video playback device, and when the video playback device initially presents the first 3D space region and the sub-space regions of the first 3D space region. When the third identifier takes a second value, it indicates that the entry corresponding to the first GPCC is an entry initially presented by the video playback device, and when the video playback device initially presents the first 3D space region and the sub-space regions of the first 3D space region, the second value corresponds to the sub-space region in the first 3D space region.

[0220] Optionally, if the first 3D space region corresponding to the first GPCC region is not divided into a plurality of sub-space regions, the 3D space region entry attribute further comprises information of the first 3D space region.

[0221] Optionally, the information of the first 3D space region comprises at least one of: an identifier of the first 3D space region, position information of the first 3D space region, and a tile identifier in the first 3D space region when the first GPCC region is encoded by using GPCC tile coding.

[0222] Optionally, the processing unit 510 is specifically configured to: if the entry of the at least one GPCC region is 1, encapsulate the entry of the 1 GPCC region into 1 non-timed point cloud media; and if the entry of the at least one GPCC region is N, encapsulate the entries of the N GPCC regions into M non-timed point cloud media. Wherein, N is an integer greater than 1, the value range of M is 【1, N】, and M is an integer.

[0223] Optionally, the second non-timed point cloud media comprises a GPCC entry group box. Wherein, the second non-timed point cloud media is any one of the at least one non-timed point cloud media comprising the entries of the plurality of GPCC regions, and the GPCC entry group box is used to associate the entries of the plurality of GPCC regions.

[0224] Optionally, the GPCC entry group box comprises a fourth identifier. The fourth identifier is an identifier of an entry of the plurality of GPCC entries that is initially presented by the video playback device.

[0225] Optionally, the GPCC entry group box comprises a fifth identifier. If the fifth identifier takes a third value, it indicates that the plurality of GPCC entries constitute a complete GPCC frame of a static object. If the fifth identifier takes a fourth value, it indicates that the plurality of GPCC entries constitute a partial GPCC frame of a static object.

[0226] Optionally, the GPCC entry group box comprises position information of a GPCC region formed by the plurality of GPCC regions.

[0227] Optionally, the communication unit 520 is further configured to receive a second request message sent by the video playback device, and send third non-timed point cloud media to the video playback device according to the second request message.

[0228] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, the foregoing and other operations and / or functions of each module in the device 500 are not described here again. Specifically, Figure 5 The device 500 shown can perform the method embodiments corresponding to the video production device, and the foregoing and other operations and / or functions of each module in the device 500 are respectively for realizing the method embodiments corresponding to the video production device. To be brief, they are not described here again.

[0229] The device 500 of the embodiments of the present application is described above from the perspective of functional modules in combination with the drawings. It should be understood that the functional modules can be realized by hardware, or by instructions in the form of software, or by a combination of hardware and software modules. Specifically, each step of the method embodiments in the embodiments of the present application can be completed by integrated logic circuits of hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware code processing performed by the processor, or executed by a combination of hardware and software modules in the code processing processor. Optionally, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps in the above method embodiments.

[0230] Embodiment 7

[0231] Figure 6A schematic diagram of a non-time sequence point cloud media processing apparatus 600 is provided in the embodiments of the present application. The apparatus 600 comprises a processing unit 610 and a communication unit 620. The communication unit 620 is configured to receive MDP signaling of at least one non-time sequence point cloud media, send a first request message to a video production device, and receive a first non-time sequence point cloud media. The processing unit 610 is configured to play the first non-time sequence point cloud media. The at least one non-time sequence point cloud media is obtained by processing non-time sequence point cloud data of a static object by using a GPCC encoding mode, obtaining a GPCC bitstream, encapsulating the GPCC bitstream, generating an entry of at least one GPCC region, encapsulating the entry of the at least one GPCC region, and generating the at least one non-time sequence point cloud media. For any one of the entries of the at least one GPCC region, the entry of the GPCC region is used to represent a GPCC component of a 3D space region corresponding to the GPCC region. For any one of the at least one non-time sequence point cloud media, the non-time sequence point cloud media comprises an identifier of the static object.

[0232] Optionally, the entry of the first GPCC region comprises a 3D space region entry attribute, and the 3D space region entry attribute comprises a first identifier and a second identifier. The first GPCC region is one of the at least one GPCC region. The first identifier is used to identify whether a first 3D space region corresponding to the first GPCC region is divided into a plurality of sub-space regions. The second identifier is used to identify whether the first GPCC region adopts a GPCC tile encoding mode.

[0233] Optionally, if the first 3D space region corresponding to the first GPCC region is divided into the plurality of sub-space regions, the 3D space region entry attribute further comprises information of the plurality of sub-space regions and information of the first 3D space region.

[0234] Optionally, for any one of the plurality of sub-space regions, the information of the sub-space region comprises at least one of the following: an identifier of the sub-space region, position information of the sub-space region, and a tile identifier in the sub-space region when the first GPCC region adopts the GPCC tile encoding mode. The information of the first 3D space region comprises at least one of the following: an identifier of the first 3D space region, position information of the first 3D space region, and a number of sub-space regions included in the first 3D space region.

[0235] Optionally, if the first 3D space region corresponding to the first GPCC region is divided into a plurality of sub-space regions, the 3D space region entry attribute further comprises a third identifier. When the third identifier takes a first value or is empty, it means that the entry corresponding to the first GPCC is the entry initially presented by the video playback device. For the first 3D space region and the sub-space regions of the first 3D space region, the video playback device initially presents the first 3D space region. When the third identifier takes a second value, it means that the entry corresponding to the first GPCC is the entry initially presented by the video playback device. For the first 3D space region and the sub-space regions of the first 3D space region, the video playback device initially presents the sub-space region corresponding to the second value in the first 3D space region.

[0236] Optionally, if the first 3D space region corresponding to the first GPCC region is not divided into a plurality of sub-space regions, the 3D space region entry attribute further comprises information of the first 3D space region.

[0237] Optionally, the information of the first 3D space region comprises at least one of the following: an identifier of the first 3D space region, position information of the first 3D space region, and a tile identifier in the first 3D space region when the first GPCC region is encoded using GPCC tile.

[0238] Optionally, if the entry of the at least one GPCC region is 1, the entry of the 1 GPCC region is encapsulated into 1 non-timed point cloud media. If the entry of the at least one GPCC region is N, the entries of the N GPCC regions are encapsulated into M non-timed point cloud media. Wherein, N is an integer greater than 1, the value range of M is 【1, N】, and M is an integer.

[0239] Optionally, the second non-timed point cloud media comprises a GPCC entry group box. The second non-timed point cloud media is any one of the at least one non-timed point cloud media that includes entries of multiple GPCC regions. The GPCC entry group box is used to associate the entries of the multiple GPCC regions.

[0240] Optionally, the GPCC entry group box comprises a fourth identifier. The fourth identifier is an identifier of an entry initially presented by the video playback device among the entries of the multiple GPCC regions.

[0241] Optionally, the GPCC entry group box comprises a fifth identifier. If the fifth identifier takes a third value, it means that the entries of the multiple GPCC regions constitute a complete GPCC frame of a static object. If the fifth identifier takes a fourth value, it means that the entries of the multiple GPCC regions constitute a partial GPCC frame of a static object.

[0242] Optionally, the GPCC entry group box comprises: position information of a plurality of GPCC regions.

[0243] Optionally, the communication unit 620 is further configured to send a second request message to the video production device according to the MDP signaling, and receive the second non-timed point cloud media.

[0244] Optionally, the processing unit 610 is further configured to play the second non-timed point cloud media.

[0245] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, details are not described here. Specifically, Figure 6 The device 600 shown can perform the method embodiments corresponding to the video playing device, and the foregoing and other operations and / or functions of each module in the device 600 are respectively to realize the method embodiments corresponding to the video playing device, and for the sake of brevity, details are not described here.

[0246] The device 600 of the embodiments of the present application is described above in the perspective of functional modules. It should be understood that the functional modules can be realized by hardware, or by instructions in the form of software, or by a combination of hardware and software modules. Specifically, each step of the method embodiments in the embodiments of the present application can be completed by integrated logic circuits and / or software instructions in the processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware code processing performed by the processor, or executed by a combination of hardware and software modules in the code processing processor. Optionally, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps in the above method embodiments.

[0247] Embodiment 8

[0248] Figure 7 is a schematic block diagram of a video production device 700 provided by the embodiments of the present application.

[0249] As Figure 7 shown, the video production device 700 can include:

[0250] The memory 710 is used to store computer programs and transmit the program codes to the processor 720. In other words, the processor 720 can call and run the computer programs from the memory 710 to realize the method in the embodiments of the present application.

[0251] For example, the processor 720 can be configured to perform the above-described method embodiments according to instructions in the computer program.

[0252] In some embodiments of the present application, the processor 720 can include but is not limited to:

[0253] A general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, and the like.

[0254] In some embodiments of the present application, the memory 710 includes but is not limited to:

[0255] A volatile memory and / or a non-volatile memory. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate SDRAM (DDR SDRAM), an enhanced SDRAM (ESDRAM), a synch link DRAM (SLDRAM), and a direct Rambus RAM (DR RAM).

[0256] In some embodiments of the present application, the computer program can be divided into one or more modules, which are stored in the memory 710 and executed by the processor 720 to complete the method provided in the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the video production device.

[0257] As shown in Figure 7 , the video production device can further include:

[0258] The transceiver 730 can be connected to the processor 720 or the memory 710.

[0259] The processor 720 can control the transceiver 730 to communicate with other devices, specifically, can send information or data to other devices, or receive information or data sent by other devices. The transceiver 730 can include a transmitter and a receiver. The transceiver 730 can further include an antenna, and the number of antennas can be one or more.

[0260] It should be understood that various components in the video production device are connected through a bus system, wherein the bus system includes a data bus, a power supply bus, a control bus and a state signal bus in addition to the data bus.

[0261] Embodiment 9

[0262] Figure 8 is a schematic block diagram of the video playback device 800 provided in an embodiment of the present application.

[0263] As shown in Figure 8 , the video playback device 800 can include:

[0264] The memory 810 is used to store computer programs and transmit the program codes to the processor 820. In other words, the processor 820 can call and run the computer program from the memory 810 to implement the method in the embodiments of the present application.

[0265] For example, the processor 820 can be used to execute the above-mentioned method embodiments according to the instructions in the computer program.

[0266] In some embodiments of the present application, the processor 820 can include but is not limited to:

[0267] General processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc.

[0268] In some embodiments of the present application, the memory 810 includes, but is not limited to:

[0269] volatile memory and / or non-volatile memory. The non-volatile memory can be ROM, PROM, EPROM, EEPROM, or flash memory. The volatile memory can be RAM, which is used as external cache memory. By way of example, and not limitation, many forms of RAM are available, such as SRAM, DRAM, SDRAM, DDR SDRAM, ESDRAM, SLDRAM, and DRDRAM.

[0270] In some embodiments of the present application, the computer program can be divided into one or more modules, which are stored in the memory 810 and executed by the processor 820 to complete the method provided by the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the video playback device.

[0271] As shown in Figure 8 The video playback device can further include:

[0272] a transceiver 830, which can be connected to the processor 820 or the memory 810.

[0273] The processor 820 can control the transceiver 830 to communicate with other devices, specifically, to send information or data to other devices, or to receive information or data sent by other devices. The transceiver 830 can include a transmitter and a receiver. The transceiver 830 can further include an antenna, and the number of antennas can be one or more.

[0274] It should be understood that various components in the video playback device are connected through a bus system, which includes a data bus, a power supply bus, a control bus, and a state signal bus in addition to the data bus.

[0275] The present application also provides a computer storage medium, which stores a computer program, and the computer program makes the computer execute the method of the above-mentioned method embodiments when executed by the computer. Alternatively, the present application embodiments also provide a computer program product containing instructions, which makes the computer execute the method of the above-mentioned method embodiments when executed by the computer.

[0276] When implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A storage media can be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired computer program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, or a twisted pair, as examples, then the coaxial cable, fiber optic cable, or twisted pair are included in the definition of medium. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), and Blu-Ray® disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0277] In one embodiment, the techniques described herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the software can be executed in a computer system, which can include one or more computers. The software can be stored on one or more computer readable media, such as a magnetic disk, optical disk, or solid state memory. The computer readable media can be distributed among one or more computer systems.

[0278] In several embodiments provided in the present application, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the above-described device embodiments are merely illustrative, and the division of the modules is merely a logical function division. In actual implementation, another division manner can be used, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be omitted or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed modules can be indirect coupling or communication connection through some interfaces, devices, or modules, and can be electrical, mechanical, or other forms.

[0279] The modules illustrated as separate components may or may not be physically separate, and the components illustrated as modules may or may not be physical modules, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules in various embodiments of the present application can be integrated in one processing module, or each module can exist physically separately, or two or more modules can be integrated in one module.

[0280] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for processing non-sequential point cloud media, characterized in that: include: Based on the non-temporal point cloud data of the static object, an entry of at least one GPCC area is obtained; Encapsulating entries of the at least one GPCC region to generate at least one non-temporal point cloud media of the static object; Sending a media presentation description (MPD) signaling of the at least one non-sequential point cloud media to a video playback device; receiving a first request message sent by the video playback device; Sending a first non-time-sequential point cloud media to the video playback device according to the first request message; Among them, if the first 3D spatial area corresponding to the first GPCC area in the at least one GPCC area is divided into multiple sub-space areas, the 3D spatial area entry attributes included in the entry of the first GPCC area include: a third identifier, and the third identifier is used to indicate that when the entry corresponding to the first GPCC area is the entry initially presented by the video playback device, for the first 3D spatial area and the sub-space area of ​​the first 3D spatial area, the spatial area initially presented on the video playback device.

2. The method according to claim 1, characterized in that When the value of the third identifier is the first value or empty, it indicates that when the entry corresponding to the first GPCC area is the entry initially presented by the video playback device, for the first 3D space area and the subspace area of ​​the first 3D space area, the first 3D space area is initially presented on the video playback device.

3. The method according to claim 1, characterized in that When the value of the third identifier is a second numerical value, indicating that the entry corresponding to the first GPCC area is the entry initially presented by the video playback device, for the first 3D spatial area and the subspace area of ​​the first 3D spatial area, the subspace area corresponding to the second numerical value in the first 3D spatial area is initially presented by the video playback device.

4. The method according to any one of claims 1 to 3, characterized in that The 3D space region entry attribute further includes: a first identifier and a second identifier; The first identifier is used to identify whether the first 3D space area corresponding to the first GPCC area is divided into multiple subspace areas; the second identifier is used to identify whether the first GPCC area adopts the GPCC tile encoding method.

5. The method according to claim 4, characterized in that If the first 3D space region corresponding to the first GPCC region is divided into multiple subspace regions, the 3D space region entry attribute further includes: information about each of the multiple subspace regions and information about the first 3D space region.

6. The method according to claim 5, characterized in that For any subspace area among the multiple subspace areas, the information of the subspace area includes at least one of the following: an identifier of the subspace area, location information of the subspace area, and a tile identifier in the subspace area when the first GPCC area adopts GPCC tile encoding; The information of the first 3D spatial region includes at least one of the following: an identifier of the first 3D spatial region, location information of the first 3D spatial region, and the number of sub-space regions included in the first 3D spatial region.

7. The method according to claim 4, characterized in that If the first 3D spatial region corresponding to the first GPCC region is not divided into a plurality of sub-space regions, the 3D spatial region entry attribute further includes: information about the first 3D spatial region.

8. The method according to claim 7, characterized in that The information of the first 3D spatial region includes at least one of the following: an identifier of the first 3D spatial region, location information of the first 3D spatial region, and a tile identifier in the first 3D spatial region when the first GPCC region adopts GPCC tile encoding.

9. The method according to any one of claims 1 to 3, characterized in that The step of obtaining at least one GPCC area entry based on the non-temporal point cloud data of the static object includes: Acquiring the non-time-sequential point cloud data; Processing the non-time-sequential point cloud data by using a point cloud compression GPCC encoding method based on a geometric model to obtain a GPCC bit stream; The GPCC bit stream is encapsulated to generate an entry of the at least one GPCC area.

10. The method according to any one of claims 1 to 3, characterized in that The encapsulating the entries of the at least one GPCC area to generate at least one non-sequential point cloud media of the static object includes: If the number of entries of the at least one GPCC region is one, encapsulating the entry of one GPCC region into one non-temporal point cloud media; If the number of entries of the at least one GPCC region is N, encapsulating the N GPCC region entries into M non-sequential point cloud media; Wherein, N is an integer greater than 1, the value range of M is [1, N], and M is an integer.

11. The method according to any one of claims 1 to 3, characterized in that The second non-temporal point cloud media includes: a GPCC entry group box; The second non-sequential point cloud media is any non-sequential point cloud media including entries of multiple GPCC areas in the at least one non-sequential point cloud media, and the GPCC entry group box is used to associate the entries of the multiple GPCC areas.

12. The method according to claim 11, characterized in that The GPCC entry group box includes: a fourth identifier; The fourth identifier is an identifier of an entry initially presented by the video playback device among the entries in the multiple GPCC areas.

13. The method according to claim 11, characterized in that The GPCC entry group box includes: a fifth identifier; If the value of the fifth flag is the third value, it indicates that the entries of the multiple GPCC areas constitute a complete GPCC frame of the static object; If the value of the fifth flag is the fourth value, it indicates that the entries of the multiple GPCC areas constitute a partial GPCC frame of the static object.

14. The method according to claim 11, characterized in that The GPCC entry group box includes: location information of the GPCC area composed of the multiple GPCC areas.

15. The method according to any one of claims 1 to 3, characterized in that Also includes: receiving a second request message sent by the video playback device; According to the second request message, a second non-time-sequential point cloud media is sent to the video playback device.

16. The method according to any one of claims 1 to 3, characterized in that For any one of the at least one GPCC area entries, the GPCC area entry is used to represent a GPCC component of a three-dimensional 3D space area corresponding to the GPCC area; For any one of the at least one non-time-sequential point cloud media, the non-time-sequential point cloud media includes: an identifier of the static object.

17. A device for processing non-sequential point cloud media, characterized in that: include: processing unit and communication unit; The processing unit is used for: Based on the non-temporal point cloud data of the static object, an entry of at least one GPCC area is obtained; Encapsulating entries of the at least one GPCC region to generate at least one non-temporal point cloud media of the static object; The communication unit is used for: Sending MPD signaling of the at least one non-time-sequential point cloud media to a video playback device; receiving a first request message sent by the video playback device; Sending a first non-time-sequential point cloud media to the video playback device according to the first request message; Among them, if the first 3D spatial area corresponding to the first GPCC area in the at least one GPCC area is divided into multiple sub-space areas, the 3D spatial area entry attributes included in the entry of the first GPCC area include: a third identifier, and the third identifier is used to indicate that when the entry corresponding to the first GPCC area is the entry initially presented by the video playback device, for the first 3D spatial area and the sub-space area of ​​the first 3D spatial area, the spatial area initially presented on the video playback device.

18. A video production device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 16.

19. A computer-readable storage medium, characterized in that Used to store a computer program, wherein the computer program causes a computer to execute the method according to any one of claims 1 to 16.