Method, apparatus, device and storage medium for processing non-sequential point cloud media
By carrying static object identifiers in non-temporal point cloud media and using GPCC encoding processing, the problem that users cannot determine whether the point cloud media is the same object is solved, which improves processing efficiency and user experience, and enhances the flexibility and decoding efficiency of video production.
Patent Information
- Application Number
- CN202011347626.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-11-26
AI Technical Summary
When a user requests to play different point cloud media, he cannot determine whether it is a point cloud media of the same object, resulting in low processing efficiency.
The non-temporal point cloud media carries the identification of static objects, and process point cloud data through GPCC encoding, generates GPCC bitstream and area entries, encapsulates it into a non-temporal point cloud media, and uses MDP signaling for transmission and request.
It improves the purpose and processing efficiency of point cloud media where users request the same static object, improves the user experience, and improves the flexibility and decoding efficiency of video production through flexible combination of GPCC areas.
Smart Images

Figure CN114549778B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to a method, device, equipment and storage medium for processing non-temporal point cloud media. Background Art
[0002] Currently, point cloud data of an object can be obtained in many ways, and a video production device can transmit the point cloud data to a video playback device in the form of point cloud media, that is, a point cloud media file, for the video playback device to play the point cloud media.
[0003] It is worth mentioning that for the point cloud data of the same object, different point cloud media can be encapsulated. For example, some point cloud media are the entire point cloud media of the object, while some point cloud media are only partial point cloud media of the object. Based on this, a user can request to play different point cloud media. However, when the user makes a request, the user does not know whether different point cloud media are point cloud media of the same object, resulting in a problem of low processing efficiency. Summary of the Invention
[0004] The present application provides a method, device, equipment and storage medium for processing non-temporal point cloud media, so that a user can request non-point cloud media of the same static object in multiple times and purposefully, thereby improving the processing efficiency and the user experience.
[0005] In a first aspect, the present application provides a method for processing non-temporal point cloud media, including: obtaining non-temporal point cloud data of a static object; processing the non-temporal point cloud data by a GPCC encoding method to obtain a GPCC bitstream; encapsulating the GPCC bitstream to generate entries of at least one GPCC region; encapsulating the entries of at least one GPCC region to generate at least one non-temporal point cloud media of the static object; sending an MDP signaling of at least one non-temporal point cloud media to a video playback device; receiving a first request message sent by the video playback device; and sending a first non-temporal point cloud media to the video playback device according to the first request message; wherein, for any one of the entries of at least one GPCC region, the entry of the GPCC region is used to represent the GPCC components of the three-dimensional (3D) space region corresponding to the GPCC region; and for any one of the at least one non-temporal point cloud media, the non-temporal point cloud media includes: an identifier of the static object.
[0006] Second aspect, the present application provides a method for processing non-temporal point cloud media, including: receiving MDP signaling of at least one non-temporal point cloud media; sending a first request message to a video production device; receiving a first non-temporal point cloud media; playing the first non-temporal point cloud media; wherein, the at least one non-temporal point cloud media is obtained by processing non-temporal point cloud data of a static object through GPCC encoding to obtain a GPCC bitstream, encapsulating the GPCC bitstream to generate entries of at least one GPCC region, and encapsulating the entries of the at least one GPCC region to generate the obtained at least one non-temporal point cloud media; for any one of the entries of the at least one GPCC region, the entry of the GPCC region is used to represent the GPCC components of the 3D spatial region corresponding to the GPCC region; for any one of the at least one non-temporal point cloud media, the non-temporal point cloud media includes: an identifier of a static object.
[0007] Third aspect, the present application provides a device for processing non-temporal point cloud media, including: a processing unit and a communication unit; the processing unit is configured to: obtain non-temporal point cloud data of a static object; process the non-temporal point cloud data through GPCC encoding to obtain a GPCC bitstream; encapsulate the GPCC bitstream to generate entries of at least one GPCC region; encapsulate the entries of the at least one GPCC region to generate at least one non-temporal point cloud media of the static object; send MDP signaling of the at least one non-temporal point cloud media to a video playback device; the communication unit is configured to: receive a first request message sent by the video playback device; according to the first request message, send a first non-temporal point cloud media to the video playback device; wherein, for any one of the entries of the at least one GPCC region, the entry of the GPCC region is used to represent the GPCC components of the 3D spatial region corresponding to the GPCC region; for any one of the at least one non-temporal point cloud media, the non-temporal point cloud media includes: an identifier of a static object.
[0008] Fourth aspect, the present application provides a processing device for non-temporal point cloud media, including: a processing unit and a communication unit; the communication unit is configured to: receive MDP signaling of at least one non-temporal point cloud media; send a first request message to a video production device; receive a first non-temporal point cloud media; the processing unit is configured to play the first non-temporal point cloud media; wherein, the at least one non-temporal point cloud media is obtained by processing non-temporal point cloud data of a static object through a GPCC encoding method to obtain a GPCC bitstream, encapsulating the GPCC bitstream to generate entries of at least one GPCC region, and encapsulating the entries of at least one GPCC region to generate the obtained at least one non-temporal point cloud media; for any one of the entries of at least one GPCC region, the entry of the GPCC region is used to represent the GPCC components of the 3D spatial region corresponding to the GPCC region; for any one of the at least one non-temporal point cloud media, the non-temporal point cloud media includes: an identifier of the static object.
[0009] Fifth aspect, there is provided a video production device, including: a processor and a memory, the memory is used for storing a computer program, and the processor is used for calling and running the computer program stored in the memory to execute the method of the first aspect.
[0010] Sixth aspect, there is provided a video playback device, including: a processor and a memory, the memory is used for storing a computer program, and the processor is used for calling and running the computer program stored in the memory to execute the method of the second aspect.
[0011] Seventh aspect, there is provided a computer-readable storage medium for storing a computer program, and the computer program causes a computer to execute the method of the first aspect.
[0012] Eighth aspect, there is provided a computer-readable storage medium for storing a computer program, and the computer program causes a computer to execute the method of the second aspect.
[0013] In summary, in the present application, when the video production device encapsulates non-temporal point cloud media, it can carry the identifier of the static object in the non-temporal point cloud media, so that the user can request the non-point cloud media of the same static object in multiple times and purposefully, thereby improving the user experience.
[0014] Furthermore, in the present application, the 3D spatial region corresponding to the entry of the GPCC region can be divided into multiple sub-spatial regions. Combining the characteristics of independent encoding and decoding of GPCC tiles, it can enable the user to decode and present non-temporal point cloud media with higher efficiency and lower latency.
[0015] Furthermore, the video production device can flexibly combine the entries of multiple GPCC regions to form different non-sequential point cloud media, where the non-sequential point cloud media can constitute a complete GPCC frame or a partial GPCC frame. Thereby, the flexibility of video production can be improved. Brief Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 Shows a schematic structural diagram of a processing system for a non-sequential point cloud media provided by an exemplary embodiment of the present application;
[0018] Figure 2A Shows a schematic structural diagram of a processing architecture for a non-sequential point cloud media provided by an exemplary embodiment of the present application;
[0019] Figure 2B Shows a schematic structural diagram of a sample provided by an exemplary embodiment of the present application;
[0020] Figure 2C Shows a schematic structural diagram of a container containing multiple file tracks provided by an exemplary embodiment of the present application;
[0021] Figure 2D Shows a schematic structural diagram of a sample provided by another exemplary embodiment of the present application;
[0022] Figure 3 Is an interaction flowchart of a processing method for a non-sequential point cloud media provided by an embodiment of the present application;
[0023] Figure 4A Is a schematic diagram of the encapsulation of a point cloud media provided by an embodiment of the present application;
[0024] Figure 4B Is another schematic diagram of the encapsulation of a point cloud media provided by an embodiment of the present application;
[0025] Figure 5 Is a schematic diagram of a processing device 500 for a non-sequential point cloud media provided by an embodiment of the present application;
[0026] Figure 6 Is a schematic diagram of a processing device 600 for a non-sequential point cloud media provided by an embodiment of the present application;
[0027] Figure 7It is a schematic block diagram of a video production device 700 provided by an embodiment of the present application;
[0028] Figure 8 It is a schematic block diagram of a video playback device 800 provided by an embodiment of the present application. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] Before introducing the technical solutions of the present application, the relevant knowledge of the present application will be introduced first:
[0032] The so-called point cloud refers to a set of irregularly distributed discrete points in space that express the spatial structure and surface attributes of a three-dimensional object or a three-dimensional scene. Point cloud data is the specific recording form of the point cloud. The point cloud data of each point in the point cloud can include geometric information (i.e., three-dimensional position information) and attribute information. Among them, the geometric information of each point in the point cloud refers to the Cartesian three-dimensional coordinate data of the point, and the attribute information of each point in the point cloud can include but is not limited to at least one of the following: color information, material information, and laser reflection intensity information. Usually, each point in the point cloud has the same number of attribute information; for example, each point in the point cloud has two attribute information, namely color information and laser reflection intensity information; or each point in the point cloud has three attribute information, namely color information, material information, and laser reflection intensity information.
[0033] With the progress and development of science and technology, it is currently possible to obtain a large amount of high-precision point cloud data at a relatively low cost and within a short time cycle. The acquisition methods of point cloud data can include, but are not limited to, at least one of the following: (1) Generated by computer devices. Computer devices can generate point cloud data based on virtual three-dimensional objects and virtual three-dimensional scenes. (2) Obtained by 3D (3-Dimension) laser scanning. Through 3D laser scanning, point cloud data of static real-world three-dimensional objects or three-dimensional scenes can be obtained, and millions of point cloud data can be obtained per second; (3) Obtained by 3D photogrammetry. Through 3D photographic equipment (i.e., a set of cameras or a camera device with multiple lenses and sensors), the visual scene of the real world is collected to obtain the point cloud data of the visual scene of the real world. Through 3D photography, point cloud data of dynamic real-world three-dimensional objects or three-dimensional scenes can be obtained. (4) Obtain point cloud data of biological tissue organs through medical devices. In the medical field, point cloud data of biological tissue organs can be obtained through medical devices such as magnetic resonance imaging (MRI), computed tomography (CT), and electromagnetic positioning information.
[0034] The so-called point cloud media refers to the point cloud media files formed by point cloud data. The point cloud media includes multiple media frames, and each media frame in the point cloud media is composed of point cloud data. The point cloud media can flexibly and conveniently express the spatial structure and surface attributes of three-dimensional objects or three-dimensional scenes, so it is widely used in projects such as virtual reality (VR) games, computer-aided design (CAD), geographic information system (GIS), autonomous navigation system (ANS), digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive telepresence, and three-dimensional reconstruction of biological tissue organs.
[0035] The so-called non-sequential point cloud media is for the same static object, that is, for the same static object, the corresponding point cloud media is non-sequential.
[0036] Based on the above description, please refer to Figure 1 , Figure 1The figure shows a schematic architecture diagram of a processing system for non-sequential point cloud media provided by an exemplary embodiment of the present application. The processing system 10 for non-sequential point cloud media includes a video playback device 101 and a video production device 102. Among them, the video production device refers to the computer device used by the provider of non-sequential point cloud media (such as the content producer of non-sequential point cloud media). This computer device can be a terminal (such as a personal computer (PC), a smart mobile device (such as a smart phone), etc.), a server, etc.; the video playback device refers to the computer device used by the user of non-sequential point cloud media (such as a user). This computer device can be a terminal (such as a PC), a smart mobile device (such as a smart phone), a VR device (such as a VR helmet, VR glasses), etc.). The video production device and the video playback device can be directly or indirectly connected through wired communication or wireless communication. The embodiments of the present application do not limit this here.
[0037] Figure 2A The figure shows a schematic architecture diagram of a processing architecture for non-sequential point cloud media provided by an exemplary embodiment of the present application. Next, in combination with Figure 1 the shown processing system for non-sequential point cloud media and Figure 2A the shown processing architecture for non-sequential point cloud media, the processing solution for non-sequential point cloud media provided by the embodiments of the present application will be introduced. The processing process of non-sequential point cloud media includes the processing process on the video production device side and the processing process on the video playback device side. The specific processing process is as follows:
[0038] I. Processing process on the video production device side:
[0039] (1) Process of obtaining point cloud data.
[0040] In one implementation, from the perspective of the acquisition method of point cloud data, the acquisition methods of point cloud data can be divided into two types: obtaining point cloud data by capturing the visual scene of the real world through a capture device, and generating point cloud data through a computer device. In one implementation, the capture device can be a hardware component set in a video production device. For example, the capture device is a camera, a sensor, etc. of a terminal. The capture device can also be a hardware device connected to the content production device, such as a camera connected to a server, etc. The capture device is used to provide the acquisition service of point cloud data for the video production device. The capture device can include, but is not limited to, any one of the following: a camera device, a sensing device, a scanning device; among them, the camera device can include an ordinary camera, a stereo camera, a light field camera, etc.; the sensing device can include a laser device, a radar device, etc.; the scanning device can include a 3D laser scanning device, etc. The number of capture devices can be multiple, and these capture devices are deployed at some specific positions in the real space to simultaneously capture point cloud data at different angles in this space, and the captured point cloud data is synchronized both in time and space. In another implementation, the computer device can generate point cloud data according to virtual three-dimensional objects and virtual three-dimensional scenes. Due to the different acquisition methods of point cloud data, the corresponding compression and encoding methods of the point cloud data obtained by different methods may also be different.
[0041] (2) Encoding and encapsulation process of point cloud data.
[0042] In one implementation, the video production device encodes the acquired point cloud data using a Geometry-Based Point Cloud Compression (GPCC) encoding method or a Video-Based Point Cloud Compression (VPCC) encoding method based on traditional video encoding to obtain a GPCC bitstream or a VPCC bitstream of the point cloud data.
[0043] In one implementation, taking the GPCC encoding method as an example, the video production device encapsulates the GPCC bitstream of the encoded point cloud data using a file track; the so-called file track refers to the encapsulation container of the GPCC bitstream of the encoded point cloud data; the GPCC bitstream can be encapsulated in a single file track, and the GPCC bitstream can also be encapsulated into multiple file tracks. The specific situations of the GPCC bitstream encapsulated in a single file track and the GPCC bitstream encapsulated in multiple file tracks are as follows:
[0044] 1. The GPCC bitstream is encapsulated in a single file track. When the GPCC bitstream is transmitted in a single file track, it is required that the GPCC bitstream be declared and represented according to the transmission rules of the single file track. The GPCC bitstream encapsulated in a single file track does not require further processing and can be encapsulated through the International Organization for Standardization Base Media File Format (ISOBMFF). Specifically, each sample encapsulated in a single file track contains one or more GPCC components, which are also referred to as GPCC constituents, and the GPCC constituent can be a GPCC geometric component or a GPCC attribute component. A sample refers to a set of encapsulated structures of one or more point clouds, that is, each sample consists of one or more Type-Length-Value ByteStream Format (TLV) encapsulated structures. Figure 2B shows a schematic structural diagram of a sample provided by an exemplary embodiment of the present application, as Figure 2B shown, when performing single file track transmission, the samples in the file track consist of a GPCC parameter set TLV, a geometric bitstream TLV, and an attribute bitstream TLV, and the sample is encapsulated into a single file track.
[0045] 2. The GPCC bitstream is encapsulated in multiple file tracks. When the encoded GPCC geometric bitstream and the encoded GPCC attribute bitstream are transmitted in different file tracks, each sample in the file track contains at least one TLV encapsulated structure, and the TLV encapsulated structure carries single GPCC constituent data, and the encoded GPCC geometric bitstream and the encoded GPCC attribute bitstream are not simultaneously included in the TLV encapsulated structure. Figure 2C shows a schematic structural diagram of a container containing multiple file tracks provided by an exemplary embodiment of the present application, as Figure 2C shown, the encapsulated packet 1 transmitted in file track 1 contains the encoded GPCC geometric bitstream and does not contain the encoded GPCC attribute bitstream; the encapsulated packet 2 transmitted in file track 2 contains the encoded GPCC attribute bitstream and does not contain the encoded GPCC geometric bitstream. Since the video playback device should first decode the encoded GPCC geometric bitstream during decoding, and the decoding of the encoded GPCC attribute bitstream depends on the decoded geometric information, encapsulating different GPCC component bitstreams in separate file tracks enables the video playback device to access the file track carrying the encoded GPCC geometric bitstream before the encoded GPCC attribute bitstream. Figure 2DA schematic structural diagram of a sample provided by another exemplary embodiment of the present application is shown. As Figure 2D shown, when multiple file tracks are transmitted, the encoded GPCC geometry bitstream and the encoded GPCC attribute bitstream are transmitted in different file tracks. The sample in this file track consists of the GPCC parameter set TLV and the geometry bitstream TLV, and the sample does not contain the attribute bitstream TLV. The sample is encapsulated in any one of the multiple file tracks.
[0046] In one implementation, the acquired point cloud data is encoded and encapsulated by a video production device to form a non-sequential point cloud media. The non-sequential point cloud media can be the entire media file of an object or a media segment of the object. And the video production device records the metadata of the encapsulated file of the non-sequential point cloud media according to the file format requirements of the non-sequential point cloud media by using media presentation description information (i.e., description signaling file) (Media presentation description, MPD). Here, the metadata is a general term for information related to the presentation of the non-sequential point cloud media. The metadata may include description information of the non-sequential point cloud media, description information of the viewport, and signaling information related to the presentation of the non-sequential point cloud media, etc. The video production device sends the MPD to the video playback device so that the video playback device requests to obtain the point cloud media according to the relevant description information in the MDP. Specifically, the point cloud media and the MDP are sent from the video production device to the video playback device through a transmission mechanism (such as Dynamic Adaptive Streaming over HTTP (DASH), Smart MediaTransport (SMT)).
[0047] II. Data processing process on the video playback device side:
[0048] (1) The process of unpacking and decoding point cloud data.
[0049] In one implementation, the video playback device can obtain the non-sequential point cloud media through the MDP signaling sent by the video production device. The process of unpacking the file on the video playback device side is the reverse of the process of encapsulating the file on the video production device side. The video playback device unpacks the encapsulated file of the non-sequential point cloud media according to the file format requirements of the non-sequential point cloud media to obtain the encoded bitstream (i.e., the GPCC bitstream or the VPCC bitstream). The decoding process on the video playback device side is the reverse of the encoding process on the video production device side. The video playback device decodes the encoded bitstream to restore the point cloud data.
[0050] (2) The process of rendering point cloud data.
[0051] In one implementation, the video playback device renders the point cloud data obtained by decoding the GPCC bitstream according to the metadata related to rendering and window in the MDP, and the rendering is completed to present the visual scene corresponding to the point cloud data.
[0052] It can be understood that the processing system of the non-sequential point cloud media described in the embodiments of the present application is to more clearly illustrate the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0053] As described above, for the point cloud data of the same object, it can be encapsulated into different point cloud media. For example, some point cloud media are the entire point cloud media of the object, and some point cloud media are partial point cloud media of the object. Based on this, the user can request to play different point cloud media. However, when the user requests, they do not know whether the different point cloud media are the point cloud media of the same object, resulting in the problem of blind requests. This problem also exists for non-point cloud media of static objects.
[0054] To solve the above technical problems, the present application carries the identifier of the static object in the non-point cloud media, so that the user can request the non-point cloud media of the same static object in multiple times and purposefully.
[0055] The technical solutions of the present application will be elaborated in detail below:
[0056] Embodiment 1
[0057] Figure 3 It is an interaction flowchart of a method for processing non-sequential point cloud media provided by an embodiment of the present application. The execution subject of this method is a video production device and a video playback device. As Figure 3 shown, the method includes the following steps:
[0058] S301: The video production device acquires non-sequential point cloud data of a static object.
[0059] S302: The video production device processes the non-sequential point cloud data through the GPCC encoding method to obtain a GPCC bitstream.
[0060] S303: The video production device encapsulates the GPCC bitstream to generate entries for at least one GPCC region.
[0061] S304: The video production device encapsulates the entries for at least one GPCC region to generate at least one non-sequential point cloud media of the static object, and each non-sequential point cloud media includes an identifier of the static object.
[0062] S305: The video production device sends MDP signaling of at least one non-sequential point cloud media to the video playback device.
[0063] S306: The video playback device sends a first request message.
[0064] S307: The video production device sends a first non-sequential point cloud media to the video playback device according to the first request message.
[0065] S308: The video playback device plays the first non-sequential point cloud media.
[0066] It should be understood that regarding how to obtain the non-sequential point cloud data of static objects and obtain the GPCC bitstream, reference can be made to the above relevant knowledge, and this application will not elaborate further.
[0067] It should be understood that for any entry of a GPCC region among the entries of at least one GPCC region, the entry of the GPCC region is used to represent the GPCC components of the 3D spatial region corresponding to the GPCC region.
[0068] It should be understood that each GPCC region corresponds to a 3D spatial region of the above-mentioned static object, and this 3D spatial region can be the entire or part of the 3D spatial region of the static object.
[0069] It should be understood that as described above, the GPCC component is also called a GPCC module, and this GPCC component can be a GPCC geometric component or an attribute component.
[0070] It should be understood that on the video production device side, the identifier of the static object can be defined by the following code:
[0071] aligned(8) class ObjectInfoProperty extends ItemProperty('obif') {
[0072] unsigned int(32) object_ID;}
[0073] Among them, ObjectInfoProperty indicates the property of the content corresponding to the entry, and both the GPCC geometric component and the attribute component can include this property. If only the GPCC geometric component includes this property, then the ObjectInfoProperty of all the attribute components associated with this GPCC geometric component is the same as it.
[0074] object_ID indicates the identifier of the static object, and for different entries of the same static object, its object_ID is the same.
[0075] Optionally, the identifier of the above-mentioned static object may be carried in an entry related to the GPCC geometric component in the point cloud media, or carried in an entry related to the GPCC attribute component in the point cloud media, or carried in an entry related to the GPCC geometric component in the point cloud media and an entry related to the GPCC attribute component. This application places no restrictions on this.
[0076] Exemplarily, Figure 4A is a schematic diagram of the encapsulation of a point cloud media provided by an embodiment of this application. As Figure 4A shown, the point cloud media includes: an entry related to the GPCC geometric component and an entry related to the GPCC attribute component. Among them, these entries can be associated through the GPCC entry group box in the point cloud media. As Figure 4A shown, the entry related to the GPCC geometric component is associated with the entry related to the GPCC attribute component. Among them, the entry related to the GPCC geometric component may include the following entry attributes: such as GPCC Configuration, 3D spatial region attribute (3D spatial region or ItemSpatialInfoProperty), identifier of the static object. The entry related to the GPCC attribute component may include the following entry attributes: such as GPCC Configuration, identifier of the static object, etc.
[0077] Optionally, the GPCC configuration indicates the configuration information of the decoder required to decode the corresponding entry and the information related to each GPCC component, but is not limited thereto.
[0078] It is worth mentioning that the entry related to the GPCC attribute component may also include: 3D spatial region attribute. This application places no restrictions on this.
[0079] Exemplarily, Figure 4B is another schematic diagram of the encapsulation of a point cloud media provided by an embodiment of this application. Figure 4B Differing from Figure 4A is that in Figure 4B , the point cloud media includes: an entry related to the GPCC geometric component, and this entry is associated with two entries related to the GPCC attribute component. For the respective attributes included in the remaining entries related to the GPCC geometric component and the respective attributes included in the entries related to the GPCC attribute component, reference can be made to Figure 4A , and this application will not elaborate further on this.
[0080] It should be understood that the identifier of the above-mentioned static object is not limited to being carried in the attributes corresponding to the entries in each GPCC region.
[0081] It should be understood that the above MDP signaling can refer to the relevant knowledge described above in this application, and this application will not elaborate on it further.
[0082] Optionally, for any one of the at least one non-temporal point cloud media, the non-temporal point cloud media can be the entire or partial point cloud media of the above static object.
[0083] It should be understood that the video playback device can send a first request message according to the above MDP signaling to request the first non-temporal point cloud media.
[0084] In summary, in this application, when the video production device encapsulates the non-temporal point cloud media, it can carry the identifier of the static object in the non-temporal point cloud media, so that the user can request the non-point cloud media of the same static object in multiple times and purposefully, thereby improving the user experience.
[0085] Embodiment 2
[0086] In the prior art, each entry in a GPCC region only corresponds to a 3D spatial region, while in this application, the 3D spatial region can be further divided. Based on this, this application has correspondingly updated the entry attributes in the non-temporal point cloud media and the MDP signaling, which are specifically as follows:
[0087] Optionally, the entry of the first GPCC region includes: 3D spatial region entry attributes, and the 3D spatial region entry attributes include: a first identifier and a second identifier. Among them, the first GPCC region is one of the at least one GPCC region. The first identifier (Sub_region_contained) is used to identify whether the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-spatial regions. The second identifier (tile_id_present) is used to identify whether the first GPCC region adopts the GPCC tile coding method.
[0088] Exemplarily, when Sub_region_contained = 0, it means that the first 3D spatial region corresponding to the first GPCC region is not divided into multiple sub-spatial regions. When Sub_region_contained = 1, it means that the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-spatial regions.
[0089] Exemplarily, when tile_id_present = 0, it means that the first GPCC region does not adopt the GPCC tile coding method. When tile_id_present = 1, it means that the first GPCC region adopts the GPCC tile coding method.
[0090] It should be understood that when Sub_region_contained = 1, tile_id_present = 1, that is, when the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-spatial regions, the video production end must adopt the GPCC tile coding method.
[0091] Optionally, if the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-spatial regions, the 3D spatial region entry attributes further include, but are not limited to: information of each of the multiple sub-spatial regions and information of the first 3D spatial region.
[0092] Optionally, for any one of the multiple sub-spatial regions, the information of the sub-spatial region includes at least one of the following, but is not limited to: the identifier of the sub-spatial region, the location information of the sub-spatial region, and when the first GPCC region adopts GPCC tile coding, the tile identifier in the sub-spatial region.
[0093] Optionally, the location information of the sub-spatial region includes, but is not limited to: the location information of an anchor point of the sub-spatial region, and the lengths of the sub-spatial region along the X-axis, Y-axis, and Z-axis respectively. Or, the location information of the sub-spatial region includes, but is not limited to: the location information of two anchor points of the sub-spatial region.
[0094] Optionally, the information of the first 3D spatial region includes at least one of the following, but is not limited to: the identifier of the first 3D spatial region, the location information of the first 3D spatial region, and the number of sub-spatial regions included in the first 3D spatial region.
[0095] Optionally, the location information of the first 3D spatial region includes, but is not limited to: the location information of an anchor point of the first 3D spatial region, and the lengths of the first 3D spatial region along the X-axis, Y-axis, and Z-axis respectively. Or, the location information of the first 3D spatial region includes, but is not limited to: the location information of two anchor points of the first 3D spatial region.
[0096] Optionally, if the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub - spatial regions, the 3D spatial region entry attribute further includes: a third identifier (initial_region_id). When the value of the third identifier is the first value or empty, it indicates that the entry corresponding to the first GPCC is the entry initially presented by the video playback device. For the first 3D spatial region and its sub - spatial regions, what is initially presented by the video playback device is the first 3D spatial region. When the value of the third identifier is the second value, it indicates that the entry corresponding to the first GPCC is the entry initially presented by the video playback device. For the first 3D spatial region and its sub - spatial regions, what is initially presented by the video playback device is the sub - spatial region corresponding to the second value in the first 3D spatial region.
[0097] Optionally, the above - mentioned first value is 0, and the second value is the identifier of the sub - spatial region that needs to be initially presented in the first 3D spatial region.
[0098] Optionally, if the first 3D spatial region corresponding to the first GPCC region is not divided into multiple sub - spatial regions, the 3D spatial region entry attribute further includes: information about the first 3D spatial region. Optionally, the information about the first 3D spatial region includes at least one of the following, but is not limited to: the identifier of the first 3D spatial region, the position information of the first 3D spatial region, and when the first GPCC region uses GPCC tile encoding, the tile identifier in the first 3D spatial region.
[0099] It should be understood that for the situation of the position information of the first 3D spatial region, reference can be made to the explanation of the position information of the first 3D spatial region in the above content, and this application will not elaborate further.
[0100] The following will illustrate the update of the entry attributes in the non - temporal point cloud media in the form of code in this application:
[0101]
[0102]
[0103] The semantics of each field are as follows:
[0104] ItemSpatialInfoProperty represents the 3D spatial region attribute of the entry of the GPCC region. If the entry is the entry corresponding to the geometric component, this attribute must be included; if the entry is the entry corresponding to the attribute component, this 3D spatial region attribute may not be included.
[0105] When the value of sub_region_contained is 1, it means that the 3D spatial region can be further divided into multiple sub - spatial regions. When the value of this field is 1, the value of tile_id_present must be 1. When the value of sub_region_contained is 0, it means that there is no further division of sub - spatial regions in the 3D space.
[0106] When the value of tile_id_present is 1, it means that the non - temporal point cloud data uses GPCC tile encoding, and the tile id corresponding to the non - temporal point cloud is given in this attribute.
[0107] inital_region_id represents the ID of the spatial region initially presented within the overall space of the current entry when the current entry is for initial consumption or playback. If the value of this field is 0 or this field does not exist, the initially presented region of the entry is the overall 3D spatial region. If the value of this field is the identifier of a sub - spatial region, the initially presented region of the entry is the sub - spatial region corresponding to the identifier.
[0108] 3DSpatialRegionStruct represents a 3D spatial region. The first 3DSpatialRegionStruct in ItemSpatialInfoProperty indicates the 3D spatial region corresponding to the entry of ItemSpatialInfoProperty, and the remaining 3DSpatialRegionStructs indicate each sub - spatial region within the 3D spatial region corresponding to the entry.
[0109] num_sub_regions indicates the number of sub - spatial regions divided within the 3D spatial region corresponding to the entry.
[0110] num_tiles indicates the number of tiles in the 3D spatial region corresponding to the entry, or the number of tiles in the sub - spatial region corresponding to it.
[0111] tile_id indicates the identifier of the GPCC tile.
[0112] anchor_x, anchor_y, and anchor_z respectively represent the x, y, and z coordinates of the anchor point of the 3D spatial region or its sub - spatial region.
[0113] region_dx, region_dy, and region_dz respectively represent the lengths of the 3D spatial region or its sub - spatial region along the X - axis, Y - axis, and Z - axis.
[0114] In summary, in the present application, the 3D space region can be divided into multiple sub-space regions. Combining the characteristics of independent encoding and decoding of GPCC tiles, it can enable users to decode and present non-sequential point cloud media with higher efficiency and lower latency.
[0115] Embodiment 3
[0116] As described above, the video production side can encapsulate the entries of at least one GPCC region to generate at least one non-sequential point cloud media of a static object. Among them, if the number of entries in at least one GPCC region is 1, then the entry of 1 GPCC region is encapsulated into 1 non-sequential point cloud media. If the number of entries in at least one GPCC region is N, then the entries of N GPCC regions are encapsulated into M non-sequential point cloud media. Wherein, N is an integer greater than 1, and the value range of M is [1, N], and M is an integer. For example: if the number of entries in at least one GPCC region is N, then the entries of N GPCC regions can be encapsulated into 1 non-sequential point cloud media, or encapsulated into N non-sequential point cloud media, and each non-sequential point cloud media includes one entry.
[0117] Next, the fields in the second non-sequential point cloud media will be described. Among them, the second non-sequential point cloud media is any non-sequential point cloud media among the at least one non-sequential point cloud media that includes entries of multiple GPCC regions.
[0118] Optionally, the second non-sequential point cloud media includes:
[0119] GPCC Item Group Box (GPCCItemGroupBox). Among them, the GPCC Item Group Box is used to associate the entries of multiple GPCC regions, as Figure 4A and 4B shown.
[0120] Optionally, the GPCC Item Group Box includes: the identifiers of the entries of multiple GPCC regions.
[0121] Optionally, the GPCC Item Group Box includes: a fourth identifier (initial_item_ID). Among them, the fourth identifier is the identifier of the entry that is initially presented in the video playback device among the entries of multiple GPCC regions.
[0122] Optionally, the GPCC Item Group Box includes: a fifth identifier (partial_item_flag). If the value of the fifth identifier is the third value, it means that the entries of multiple GPCC regions constitute a complete GPCC frame of a static object. If the value of the fifth identifier is the fourth value, it means that the entries of multiple GPCC regions constitute a partial GPCC frame of a static object.
[0123] Optionally, the third value may be 0 and the fourth value may be 1, but this is not limited thereto.
[0124] Optionally, the GPCC item group box includes: location information of GPCC regions composed of multiple GPCC regions.
[0125] Exemplarily, if the multiple GPCC regions are two regions R1 and R2, the GPCC item group box includes location information of the R1 + R2 region.
[0126] The following will illustrate each field in the above GPCC item group box through code:
[0127]
[0128] The items included in the GPCCItemGroupBox are items belonging to the same static object and items having an association relationship when presenting consumption. All the items included in this GPCCItemGroupBox may constitute a complete GPCC frame or may be a part of a GPCC frame.
[0129] initial_item_ID indicates the identifier of the item initially consumed within an item group.
[0130] It should be noted that this initial_item_ID is only valid when the current item group is the item group initially requested by the user. For example: the same static object corresponds to two point cloud media, namely F1 and F2. When the user requests F1 for the first time, the initial_item_ID in the item group within F1 is valid. For the second request of F2, the initial_item_ID inside it is invalid.
[0131] When the partial_item_flag takes the value of 0, it means that all the items included in the GPCCItemGroupBox and their associated items constitute a complete GPCC frame. When the value is 1, it means that all the items included in the GPCCItemGroupBox and their associated items only constitute a partial GPCC frame.
[0132] To support the technology proposed in this application, the corresponding signaling messages also need to be extended. Taking the MDP signaling as an example, the extension is as follows:
[0133] The GPCC item descriptor is used to describe the elements and attributes related to the GPCC item, and this descriptor is a SupplementalProperty element.
[0134] Its @schemeIdUri attribute is equal to "urn:mpeg:mpegI:gpcc:2020:gpsr". This descriptor can be located at the Adaptation Set level or the Representation level.
[0135] Among them, Representation: In DASH, a combination of one or more media components. For example, a video file of a certain resolution can be regarded as a Representation.
[0136] Adaptation Sets: In DASH, a set of one or more video streams. An Adaptation Sets can contain multiple Representations.
[0137] Table 1: GPCC entry descriptor elements and attributes
[0138]
[0139]
[0140]
[0141] In summary, in the present application, the video production device can flexibly combine the entries of multiple GPCC regions to form different non-temporal point cloud media, where the non-temporal point cloud media can form a complete GPCC frame or a partial GPCC frame. Thereby, the flexibility of video production can be improved. Further, when a non-temporal point cloud media includes entries of a GPCC region, the video production device can also improve the initial presented entries.
[0142] The following will illustrate Examples 1 to 3 through Example 4 and Example 5:
[0143] Example 4
[0144] Suppose the video production device obtains the non-temporal point cloud data of a certain static object. There are 4 versions of point cloud media for this non-temporal point cloud data at the video production device end: the point cloud media F0 corresponding to all the non-temporal point cloud data, and the point cloud media F1 to F3 corresponding to partial non-temporal point cloud data, where F1 to F3 correspond to the 3D space regions R1 to R3 respectively. Based on this, the encapsulation content of the point cloud media F0 to F3 is as follows:
[0145] F0: ObjectInfoProperty: object_ID = 10;
[0146] ItemSpatialInfoProperty: sub_region_contained = 1; tile_id_present = 1
[0147] inital_region_id = 1001;
[0148] R0: 3d_region_id = 100, anchor = (0,0,0), region = (200,200,200)
[0149] num_sub_regions = 3;
[0150] SR1: 3d_region_id = 1001, anchor = (0,0,0), region = (100,100,200);
[0151] num_tiles = 1, tile_id[] = (1);
[0152] SR2: 3d_region_id = 1002, anchor = (100,0,0), region = (100,100,200);
[0153] num_tiles = 1, tile_id[] = (2);
[0154] SR3: 3d_region_id = 1003, anchor = (0,100,0), region = (200,100,200);
[0155] num_tiles = 2, tile_id[] = (3,4);
[0156] F1: ObjectInfoProperty: object_ID = 10;
[0157] ItemSpatialInfoProperty: sub_region_contained = 0; tile_id_present = 1;
[0158] inital_region_id = 0;
[0159] R1: 3d_region_id = 101, anchor = (0,0,0), region = (100,100,200); num_tiles = 1, tile_id[] = (1);
[0160] F2: ObjectInfoProperty: object_ID = 10;
[0161] ItemSpatialInfoProperty: sub_region_contained = 0; tile_id_present = 1
[0162] inital_region_id = 0;
[0163] R2: 3d_region_id = 102, anchor = (100,0,0), region = (100,100,200); num_tiles = 1, tile_id[] = (2);
[0164] F3: ObjectInfoProperty: object_ID = 10;
[0165] ItemSpatialInfoProperty: sub_region_contained = 0; tile_id_present = 1; inital_region_id = 0;
[0166] R3: 3d_region_id = 103, anchor = (0,100,0), region = (200,100,200);
[0167] num_tiles = 2, tile_id[] = (3,4);
[0168] Furthermore, the video production device sends the MPD signaling of F0 to F3 to the user. The Object ID, spatial region, sub - spatial region, and tile identification information are the same as those in the file encapsulation and will not be elaborated here.
[0169] Since user U1 has good network conditions and low data transmission latency, it can request F0. User U2 has relatively poor network conditions and higher data transmission latency, so it can request F1.
[0170] The video production device transmits F0 to the video playback device corresponding to user U1 and transmits F1 to the video playback device corresponding to user U2.
[0171] After the video playback device corresponding to user U1 receives F0, the initial viewing area is the SR1 area, and the corresponding tile ID is 1. When U1 decodes and consumes, it can directly decode and present tile '1' from the overall bitstream without decoding the entire file, which improves the decoding efficiency and reduces the time required for rendering and presentation. When U1 continues to consume and views the SR2 area, the corresponding tile ID is 2, and it directly decodes and presents the corresponding part of tile '2' in the overall bitstream.
[0172] After the video playback device corresponding to user U2 receives F1, it decodes and consumes F1, and based on the area that the user may consume next, combined with the information in the MPD file, namely the Object ID and the spatial area information, it requests F2 or F3 in advance for caching.
[0173] Embodiment 5
[0174] Suppose the video production device obtains the non-temporal point cloud data of a static object. There are two versions of point cloud media for this non-temporal point cloud data at the video production device side: F1 and F2. F1 contains item1 to item2, and F2 contains item3 to item4.
[0175] The encapsulation content of the point cloud media of F1 and F2 is as follows:
[0176] F1:
[0177] item1: ObjectInfoProperty: object_ID = 10; item_ID = 101
[0178] ItemSpatialInfoProperty: sub_region_contained = 0; tile_id_present = 1
[0179] inital_region_id = 0;
[0180] R1: 3d_region_id = 1001, anchor = (0,0,0), region = (100,100,200);
[0181] num_tiles = 1, tile_id[] = (1);
[0182] item2: ObjectInfoProperty: object_ID = 10; item_ID = 102
[0183] ItemSpatialInfoProperty: sub_region_contained = 0; tile_id_present = 1
[0184] inital_region_id = 0;
[0185] R2: 3d_region_id = 1002, anchor = (100,0,0), region = (100,100,200);
[0186] num_tiles = 1, tile_id[] = (2);
[0187] GPCCItemGroupBox:
[0188] initial_item_ID = 101; partial_item_flag = 1;
[0189] R1+R2: 3d_region_id = 0001, anchor = (0,0,0), region = (200,100,200);
[0190] F2:
[0191] item3: ObjectInfoProperty: object_ID = 10; item_ID = 103
[0192] ItemSpatialInfoProperty: sub_region_contained = 0; tile_id_present = 1
[0193] inital_region_id = 0;
[0194] R3: 3d_region_id = 1003, anchor = (0,100,0), region = (100,100,200);
[0195] num_tiles = 1, tile_id[] = (3);
[0196] item4: ObjectInfoProperty: object_ID = 10; item_ID = 104
[0197] ItemSpatialInfoProperty: sub_region_contained = 0; tile_id_present = 1
[0198] initial_region_id = 0;
[0199] R4: 3d_region_id = 1004, anchor = (100, 100, 0), region = (100, 100, 200);
[0200] num_tiles = 1, tile_id[] = (4);
[0201] GPCCItemGroupBox:
[0202] initial_item_ID = 103; partial_item_flag = 1;
[0203] R3+R4: 3d_region_id = 0002, anchor = (0, 100, 0), region = (200, 100, 200);
[0204] The video production device sends the MPD signaling from F1 to F2 to the user. The Object ID, spatial region, and tileID information are the same as those in the point cloud media encapsulation and will not be elaborated here.
[0205] User U1 requests to consume F1; User U2 requests to consume F2.
[0206] The video production device transmits F1 to the video playback device corresponding to user U1 respectively, and transmits F2 to the video playback device corresponding to user U2.
[0207] After the video playback device corresponding to U1 receives F1, it initially views item1. The initial viewing region of item1 is the overall viewing space of item1. Therefore, U1 consumes the whole of item1. Since F1 contains item1 and item2, corresponding to tile1 and tile2 respectively, when U1 consumes item1, it can directly decode the partial bitstream corresponding to tile1 for presentation. If U1 continues to consume and views the region of item2, the corresponding tile ID is 2, then it directly decodes the part corresponding to '2' in the whole bitstream for presentation and consumption. If U1 continues to consume and needs to view the region corresponding to item3, it requests F2 according to the MPD file. After receiving F2, it directly presents and consumes according to the region viewed by the user, without judging the initial consumption item information and initial viewing region information in F2.
[0208] After the video playback device corresponding to U2 receives F2, it initially views item3. The initial viewing area of item3 is the overall viewing space of item3. Therefore, U2 consumes the entire item3. Since F2 contains item3 and item4, which correspond to tile3 and tile4 respectively, when U2 consumes item3, it can directly decode the partial bitstream corresponding to tile3 for presentation.
[0209] Embodiment 6
[0210] Figure 5 It is a schematic diagram of a processing device 500 for non-temporal point cloud media provided by an embodiment of the present application. The device 500 includes: a processing unit 510 and a communication unit 520. The processing unit 510 is configured to: obtain non-temporal point cloud data of a static object. Process the non-temporal point cloud data through the GPCC encoding method to obtain a GPCC bitstream. Package the GPCC bitstream to generate entries for at least one GPCC region. Package the entries for at least one GPCC region to generate at least one non-temporal point cloud media of the static object. Send an MDP signaling of at least one non-temporal point cloud media to a video playback device. The communication unit 520 is configured to: receive a first request message sent by the video playback device. According to the first request message, send a first non-temporal point cloud media to the video playback device. Wherein, for any entry of at least one GPCC region, the entry of the GPCC region is used to represent the GPCC components of the 3D space region corresponding to the GPCC region. For any non-temporal point cloud media of at least one non-temporal point cloud media, the non-temporal point cloud media includes: an identifier of a static object.
[0211] Optionally, the entry of the first GPCC region includes: 3D space region entry attributes, and the 3D space region entry attributes include: a first identifier and a second identifier. Wherein, the first GPCC region is one of at least one GPCC region. The first identifier is used to identify whether the first 3D space region corresponding to the first GPCC region is divided into multiple sub-space regions. The second identifier is used to identify whether the first GPCC region adopts the GPCC tile encoding method.
[0212] Optionally, if the first 3D space region corresponding to the first GPCC region is divided into multiple sub-space regions, the 3D space region entry attributes further include: information of each of the multiple sub-space regions and information of the first 3D space region.
[0213] Optionally, for any one of a plurality of subspace regions, the information of the subspace region includes at least one of the following: the identifier of the subspace region, the location information of the subspace region, and the tile identifier in the subspace region when the first GPCC region uses GPCC tile encoding. The information of the first 3D space region includes at least one of the following: the identifier of the first 3D space region, the location information of the first 3D space region, and the number of subspace regions included in the first 3D space region.
[0214] Optionally, if the first 3D space region corresponding to the first GPCC region is divided into a plurality of subspace regions, the 3D space region entry attribute further includes: a third identifier. When the value of the third identifier is the first value or empty, it indicates that the entry corresponding to the first GPCC is the entry initially presented by the video playback device. For the first 3D space region and the subspace regions of the first 3D space region, what is initially presented by the video playback device is the first 3D space region. When the value of the third identifier is the second value, it indicates that the entry corresponding to the first GPCC is the entry initially presented by the video playback device. For the first 3D space region and the subspace regions of the first 3D space region, what is initially presented by the video playback device is the subspace region corresponding to the second value in the first 3D space region.
[0215] Optionally, if the first 3D space region corresponding to the first GPCC region is not divided into a plurality of subspace regions, the 3D space region entry attribute further includes: the information of the first 3D space region.
[0216] Optionally, the information of the first 3D space region includes at least one of the following: the identifier of the first 3D space region, the location information of the first 3D space region, and the tile identifier in the first 3D space region when the first GPCC region uses GPCC tile encoding.
[0217] Optionally, the processing unit 510 is specifically configured to: if the number of entries of at least one GPCC region is 1, encapsulate the entries of 1 GPCC region into 1 non-sequential point cloud media. If the number of entries of at least one GPCC region is N, encapsulate the entries of N GPCC regions into M non-sequential point cloud media. Wherein, N is an integer greater than 1, and the value range of M is [1, N], and M is an integer.
[0218] Optionally, the second non-sequential point cloud media includes: a GPCC entry group box. Wherein, the second non-sequential point cloud media is any one of the non-sequential point cloud media that includes entries of multiple GPCC regions among at least one non-sequential point cloud media, and the GPCC entry group box is used to associate the entries of multiple GPCC regions.
[0219] Optionally, the GPCC entry group box includes: a fourth identifier. The fourth identifier is the identifier of the entry that is initially presented by the video playback device among the entries of multiple GPCC regions.
[0220] Optionally, the GPCC entry group box includes: a fifth identifier. If the value of the fifth identifier is the third value, it means that the entries of multiple GPCC regions constitute a complete GPCC frame of a static object. If the value of the fifth identifier is the fourth value, it means that the entries of multiple GPCC regions constitute a partial GPCC frame of a static object.
[0221] Optionally, the GPCC entry group box includes: the location information of the GPCC region composed of multiple GPCC regions.
[0222] Optionally, the communication unit 520 is further configured to: receive a second request message sent by the video playback device. According to the second request message, send a third non-temporal point cloud media to the video playback device.
[0223] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 5 The illustrated device 500 can execute the method embodiments corresponding to the video production device, and the foregoing and other operations and / or functions of each module in the device 500 are respectively for implementing the method embodiments corresponding to the video production device. For the sake of brevity, they will not be elaborated here.
[0224] In the foregoing, the device 500 of the embodiments of the present application has been described from the perspective of functional modules. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the embodiments of the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in software form. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the foregoing method embodiments.
[0225] Embodiment 7
[0226] Figure 6Schematic diagram of a processing device 600 for non-sequential point cloud media provided by an embodiment of the present application. The device 600 includes: a processing unit 610 and a communication unit 620. The communication unit 620 is configured to: receive MDP signaling of at least one non-sequential point cloud media. Send a first request message to a video production device. Receive a first non-sequential point cloud media. The processing unit 610 is configured to play the first non-sequential point cloud media. Wherein, at least one non-sequential point cloud media is obtained by processing non-sequential point cloud data of a static object through a GPCC encoding method to obtain a GPCC bitstream, encapsulating the GPCC bitstream to generate entries of at least one GPCC region, and encapsulating the entries of at least one GPCC region to generate the obtained at least one non-sequential point cloud media. For any entry of at least one GPCC region among the entries of at least one GPCC region, the entry of the GPCC region is used to represent the GPCC components of the 3D spatial region corresponding to the GPCC region. For any non-sequential point cloud media among at least one non-sequential point cloud media, the non-sequential point cloud media includes: an identifier of a static object.
[0227] Optionally, the entry of the first GPCC region includes: 3D spatial region entry attributes, and the 3D spatial region entry attributes include: a first identifier and a second identifier. Wherein, the first GPCC region is one of at least one GPCC region. The first identifier is used to identify whether the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-spatial regions. The second identifier is used to identify whether the first GPCC region adopts a GPCC tile encoding method.
[0228] Optionally, if the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-spatial regions, the 3D spatial region entry attributes further include: information of each of the multiple sub-spatial regions and information of the first 3D spatial region.
[0229] Optionally, for any sub-spatial region among the multiple sub-spatial regions, the information of the sub-spatial region includes at least one of the following: an identifier of the sub-spatial region, position information of the sub-spatial region, and a tile identifier in the sub-spatial region when the first GPCC region adopts GPCC tile encoding. The information of the first 3D spatial region includes at least one of the following: an identifier of the first 3D spatial region, position information of the first 3D spatial region, and the number of sub-spatial regions included in the first 3D spatial region.
[0230] Optionally, if the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-spatial regions, the 3D spatial region entry attribute further includes: a third identifier. When the value of the third identifier is the first numerical value or empty, it indicates that the entry corresponding to the first GPCC is the entry initially presented by the video playback device. For the first 3D spatial region and the sub-spatial regions of the first 3D spatial region, what is initially presented by the video playback device is the first 3D spatial region. When the value of the third identifier is the second numerical value, it indicates that the entry corresponding to the first GPCC is the entry initially presented by the video playback device. For the first 3D spatial region and the sub-spatial regions of the first 3D spatial region, what is initially presented by the video playback device is the sub-spatial region corresponding to the second numerical value in the first 3D spatial region.
[0231] Optionally, if the first 3D spatial region corresponding to the first GPCC region is not divided into multiple sub-spatial regions, the 3D spatial region entry attribute further includes: information about the first 3D spatial region.
[0232] Optionally, the information about the first 3D spatial region includes at least one of the following: the identifier of the first 3D spatial region, the position information of the first 3D spatial region, and the tile identifier in the first 3D spatial region when the first GPCC region uses GPCC tile encoding.
[0233] Optionally, if the number of entries in at least one GPCC region is 1, the entry of 1 GPCC region is encapsulated into 1 non-sequential point cloud media. If the number of entries in at least one GPCC region is N, the entries of N GPCC regions are encapsulated into M non-sequential point cloud media. Wherein, N is an integer greater than 1, and the value range of M is [1, N], and M is an integer.
[0234] Optionally, the second non-sequential point cloud media includes: a GPCC entry group box. Wherein, the second non-sequential point cloud media is any non-sequential point cloud media that includes entries of multiple GPCC regions among at least one non-sequential point cloud media. The GPCC entry group box is used to associate the entries of multiple GPCC regions.
[0235] Optionally, the GPCC entry group box includes: a fourth identifier. Wherein, the fourth identifier is the identifier of the entry that is initially presented by the video playback device among the entries of multiple GPCC regions.
[0236] Optionally, the GPCC entry group box includes: a fifth identifier. If the value of the fifth identifier is the third numerical value, it indicates that the entries of multiple GPCC regions constitute a complete GPCC frame of a static object. If the value of the fifth identifier is the fourth numerical value, it indicates that the entries of multiple GPCC regions constitute a partial GPCC frame of a static object.
[0237] Optionally, the GPCC entry group box includes: location information of GPCC regions composed of multiple GPCC regions.
[0238] Optionally, the communication unit 620 is further configured to send a second request message to the video production device according to the MDP signaling. Receive the second non-sequential point cloud media.
[0239] Optionally, the processing unit 610 is further configured to play the second non-sequential point cloud media.
[0240] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 6 The illustrated device 600 can execute the method embodiments corresponding to the video playback device, and the foregoing and other operations and / or functions of each module in the device 600 respectively implement the method embodiments corresponding to the video playback device. For the sake of brevity, they will not be elaborated here.
[0241] In the foregoing, the device 600 of the embodiments of the present application has been described from the perspective of functional modules. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the foregoing method embodiments.
[0242] Embodiment 8
[0243] Figure 7 is a schematic block diagram of a video production device 700 provided by an embodiment of the present application.
[0244] As Figure 7 shown, the video production device 700 may include:
[0245] A memory 710 and a processor 720. The memory 710 is used to store a computer program and transmit the program code to the processor 720. In other words, the processor 720 can call and run the computer program from the memory 710 to implement the method in the embodiments of the present application.
[0246] For example, the processor 720 can be used to execute the above method embodiments according to the instructions in the computer program.
[0247] In some embodiments of the present application, the processor 720 may include, but is not limited to:
[0248] a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.
[0249] In some embodiments of the present application, the memory 710 includes, but is not limited to:
[0250] a volatile memory and / or a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0251] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 710 and executed by the processor 720 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the video production device.
[0252] As Figure 7 shown, the video production device may further include:
[0253] A transceiver 730, which may be connected to the processor 720 or the memory 710.
[0254] Among them, the processor 720 can control the transceiver 730 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 730 may include a transmitter and a receiver. The transceiver 730 may further include an antenna, and the number of antennas may be one or more.
[0255] It should be understood that the various components in the video production device are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.
[0256] Embodiment 9
[0257] Figure 8 is a schematic block diagram of a video playback device 800 provided by an embodiment of the present application.
[0258] As Figure 8 shown, the video playback device 800 may include:
[0259] A memory 810 and a processor 820. The memory 810 is used to store a computer program and transmit the program code to the processor 820. In other words, the processor 820 can call and run the computer program from the memory 810 to implement the method in the embodiments of the present application.
[0260] For example, the processor 820 may be used to execute the above method embodiments according to the instructions in the computer program.
[0261] In some embodiments of the present application, the processor 820 may include, but is not limited to:
[0262] A general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0263] In some embodiments of the present application, the memory 810 includes, but is not limited to:
[0264] A volatile memory and / or a non-volatile memory. Among them, the non-volatile memory may be a ROM, PROM, EPROM, EEPROM or flash memory. The volatile memory may be a RAM, which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as SRAM, DRAM, SDRAM, DDR SDRAM, ESDRAM, SLDRAM and DRRAM.
[0265] In some embodiments of the present application, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory 810 and executed by the processor 820 to complete the method provided in the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the video playback device.
[0266] As Figure 8 shown, the video playback device may further include:
[0267] A transceiver 830, and the transceiver 830 may be connected to the processor 820 or the memory 810.
[0268] Among them, the processor 820 may control the transceiver 830 to communicate with other devices. Specifically, it may send information or data to other devices, or receive information or data sent by other devices. The transceiver 830 may include a transmitter and a receiver. The transceiver 830 may further include an antenna, and the number of antennas may be one or more.
[0269] It should be understood that the various components in the video playback device are connected through a bus system. Among them, the bus system includes, in addition to a data bus, a power bus, a control bus and a status signal bus.
[0270] The present application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the method of the above method embodiment. Or rather, the embodiments of the present application also provide a computer program product containing instructions. When the instructions are executed by the computer, the computer executes the method of the above method embodiment.
[0271] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0272] Those of ordinary skill in the art will realize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0273] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.
[0274] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of the present application, each functional module can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0275] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A processing method for non-temporal point cloud media, characterized in that, Including: Obtaining non-temporal point cloud data of a static object; Processing the non-temporal point cloud data by a point cloud compression GPCC encoding method based on a geometric model to obtain a GPCC bitstream; Encapsulating the GPCC bitstream to generate entries for at least one GPCC region; Encapsulating the entries for the at least one GPCC region to generate at least one non-temporal point cloud media of the static object; Sending a media presentation description MPD signaling of the at least one non-temporal point cloud media to a video playback device; Receiving a first request message sent by the video playback device; Sending a first non-temporal point cloud media to the video playback device according to the first request message; Wherein, for any entry of a GPCC region among the entries for the at least one GPCC region, the entry of the GPCC region is used to represent the GPCC components of the three-dimensional 3D spatial region corresponding to the GPCC region; for any non-temporal point cloud media among the at least one non-temporal point cloud media, the non-temporal point cloud media includes: an identifier of the static object; Wherein, if the first 3D spatial region corresponding to a first GPCC region is divided into multiple subspace regions, the 3D spatial region entry attribute included in the entry of the first GPCC region includes: a third identifier; wherein, the first GPCC region is one of the at least one GPCC region; When the value of the third identifier is a first value or empty, indicating that the entry corresponding to the first GPCC is the entry initially presented by the video playback device, for the first 3D spatial region and the subspace regions of the first 3D spatial region, the first 3D spatial region is initially presented on the video playback device; When the value of the third identifier is a second value, indicating that the entry corresponding to the first GPCC is the entry initially presented by the video playback device, for the first 3D spatial region and the subspace regions of the first 3D spatial region, the subspace region corresponding to the second value in the first 3D spatial region is initially presented on the video playback device.
2. The method according to claim 1, characterized in that The 3D spatial region entry attribute further includes: a first identifier and a second identifier; The first identifier is used to identify whether the first 3D spatial region corresponding to the first GPCC region is divided into multiple subspace regions; the second identifier is used to identify whether the first GPCC region adopts a GPCC tile encoding method.
3. The method according to claim 2, wherein If the first 3D spatial region corresponding to the first GPCC region is divided into multiple subspace regions, the 3D spatial region entry attribute further includes: information of each of the multiple subspace regions and information of the first 3D spatial region.
4. The method according to claim 3, wherein For any one of the multiple subspace regions, the information of the subspace region includes at least one of the following: an identifier of the subspace region, position information of the subspace region, and a tile identifier in the subspace region when the first GPCC region adopts GPCC tile encoding; The information of the first 3D spatial region includes at least one of the following: the identifier of the first 3D spatial region, the location information of the first 3D spatial region, the number of subspace regions included in the first 3D spatial region.
5. The method according to claim 2, wherein If the first 3D spatial region corresponding to the first GPCC region is not divided into multiple subspace regions, the 3D spatial region entry attribute further includes: the information of the first 3D spatial region.
6. The method according to claim 5, wherein The information of the first 3D spatial region includes at least one of the following: the identifier of the first 3D spatial region, the location information of the first 3D spatial region, the tile identifier in the first 3D spatial region when the first GPCC region uses GPCC tile encoding.
7. The method according to any one of claims 1-6, characterized in that, The encapsulating the entries of the at least one GPCC region to generate at least one non-temporal point cloud media of the static object includes: If the number of entries of the at least one GPCC region is 1, encapsulate the entry of 1 GPCC region into 1 non-temporal point cloud media; If the number of entries of the at least one GPCC region is N, encapsulate the entries of N GPCC regions into M non-temporal point cloud media; Wherein, N is an integer greater than 1, and the value range of M is [1, N], and M is an integer.
8. The method according to any one of claims 1 to 6, characterized in that The second non-temporal point cloud media includes: a GPCC entry group box; Wherein, the second non-temporal point cloud media is any one of the at least one non-temporal point cloud media that includes entries of multiple GPCC regions, and the GPCC entry group box is used to associate the entries of the multiple GPCC regions.
9. The method according to claim 8, wherein The GPCC entry group box includes: a fourth identifier; Wherein, the fourth identifier is the identifier of the entry that is initially presented in the video playback device among the entries of the multiple GPCC regions.
10. The method according to claim 8, wherein The GPCC entry group box includes: a fifth identifier; If the value of the fifth identifier is the third value, it indicates that the entries of the multiple GPCC regions constitute a complete GPCC frame of the static object; If the value of the fifth identifier is the fourth value, it indicates that the entries of the multiple GPCC regions constitute a partial GPCC frame of the static object.
11. The method according to claim 8, wherein The GPCC entry group box includes: the location information of the GPCC region formed by the multiple GPCC regions.
12. The method according to any one of claims 1-6, characterized in that, Further includes: Receiving a second request message sent by the video playback device; Sending the second non-temporal point cloud media to the video playback device according to the second request message.
13. A processing method for non-temporal point cloud media, characterized in that, Includes: Receiving the MPD signaling of at least one non-temporal point cloud media; Sending a first request message to the video production device; Receiving the first non-temporal point cloud media; Playing the first non-temporal point cloud media; Wherein, the at least one non-temporal point cloud media is obtained by processing the non-temporal point cloud data of the static object through the GPCC encoding method to obtain a GPCC bitstream, encapsulating the GPCC bitstream to generate entries of at least one GPCC region, and encapsulating the entries of the at least one GPCC region to generate the obtained at least one non-temporal point cloud media; For any entry of the at least one GPCC region, the entry of the GPCC region is used to represent the GPCC component of the 3D spatial region corresponding to the GPCC region; for any non-temporal point cloud media among the at least one non-temporal point cloud media, the non-temporal point cloud media includes: the identifier of the static object; Wherein, if the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-spatial regions, the 3D spatial region entry attributes included in the entry of the first GPCC region include: a third identifier; wherein, the first GPCC region is one of the at least one GPCC region; When the value of the third identifier is the first value or empty, indicating that the entry corresponding to the first GPCC is the entry initially presented by the video playback device, for the first 3D spatial region and the sub-spatial regions of the first 3D spatial region, what is initially presented by the video playback device is the first 3D spatial region; When the value of the third identifier is the second value, indicating that the entry corresponding to the first GPCC is the entry initially presented by the video playback device, for the first 3D spatial region and the sub-spatial regions of the first 3D spatial region, what is initially presented by the video playback device is the sub-spatial region corresponding to the second value in the first 3D spatial region.
14. The method according to claim 13, wherein The 3D spatial region entry attributes further include: a first identifier and a second identifier; Wherein, the first GPCC region is one of the at least one GPCC region; the first identifier is used to identify whether the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-spatial regions; the second identifier is used to identify whether the first GPCC region adopts the GPCC tile coding method.
15. The method according to claim 14, characterized in that, If the first 3D spatial region corresponding to the first GPCC region is divided into multiple sub-spatial regions, the 3D spatial region entry attributes further include: the information of each of the multiple sub-spatial regions and the information of the first 3D spatial region.
16. The method according to claim 15, wherein For any one of the multiple sub-spatial regions, the information of the sub-spatial region includes at least one of the following: the identifier of the sub-spatial region, the position information of the sub-spatial region, when the first GPCC region adopts GPCC tile coding, the tile identifier in the sub-spatial region; The information of the first 3D spatial region includes at least one of the following: the identifier of the first 3D spatial region, the position information of the first 3D spatial region, the number of sub-spatial regions included in the first 3D spatial region.
17. The method according to claim 14, characterized in that, If the first 3D spatial region corresponding to the first GPCC region is not divided into multiple sub-spatial regions, the 3D spatial region entry attributes further include: the information of the first 3D spatial region.
18. The method according to claim 17, wherein The information of the first 3D spatial region includes at least one of the following: the identifier of the first 3D spatial region, the location information of the first 3D spatial region, and when the first GPCC region uses GPCC tile encoding, the tile identifier in the first 3D spatial region.
19. The method according to any one of claims 13-18, characterized in that if the number of entries of the at least one GPCC region is 1, the entry of 1 GPCC region is encapsulated into 1 non-temporal point cloud medium; if the number of entries of the at least one GPCC region is N, the entries of N GPCC regions are encapsulated into M non-temporal point cloud media; wherein, N is an integer greater than 1, and the value range of M is [1, N], and M is an integer.
20. The method according to any one of claims 13-18, characterized in that, The second non-temporal point cloud medium includes: a GPCC entry group box; wherein, the second non-temporal point cloud medium is any one of the at least one non-temporal point cloud media that includes entries of multiple GPCC regions; the GPCC entry group box is used to associate the entries of the multiple GPCC regions.
21. The method according to claim 20, wherein The GPCC entry group box includes: a fourth identifier; wherein, the fourth identifier is the identifier of the entry that is initially presented in the entries of the multiple GPCC regions in the video playback device.
22. The method according to claim 20, characterized in that, The GPCC entry group box includes: a fifth identifier; if the value of the fifth identifier is the third value, it means that the entries of the multiple GPCC regions constitute a complete GPCC frame of the static object; if the value of the fifth identifier is the fourth value, it means that the entries of the multiple GPCC regions constitute a partial GPCC frame of the static object.
23. The method according to claim 20, wherein The GPCC entry group box includes: the location information of the GPCC region formed by the multiple GPCC regions.
24. The method according to any one of claims 13 - 18, characterized in that, It further includes: sending a second request message to the video production device according to the MPD signaling; receiving the second non-temporal point cloud medium; playing the second non-temporal point cloud medium.
25. A processing device for non-temporal point cloud media, characterized in that, It includes: a processing unit and a communication unit; The processing unit is used for: acquiring non-temporal point cloud data of a static object; processing the non-temporal point cloud data through a GPCC encoding method to obtain a GPCC bitstream; encapsulating the GPCC bitstream to generate entries of at least one GPCC region; encapsulating the entries of the at least one GPCC region to generate at least one non-temporal point cloud medium of the static object; sending the MPD signaling of the at least one non-temporal point cloud medium to the video playback device; The communication unit is used for: receiving the first request message sent by the video playback device; sending a first non-temporal point cloud medium to the video playback device according to the first request message; wherein, for any entry of the at least one GPCC region, the entry of the GPCC region is used to represent the GPCC component of the 3D spatial region corresponding to the GPCC region; for any non-temporal point cloud medium of the at least one non-temporal point cloud medium, the non-temporal point cloud medium includes: the identifier of the static object. Wherein, if the first 3D space region corresponding to the first GPCC region is divided into multiple sub - space regions, the 3D space region entry attributes included in the entry of the first GPCC region include: a third identifier; wherein, the first GPCC region is one of the at least one GPCC region; When the value of the third identifier is the first value or empty, indicating that the entry corresponding to the first GPCC is the entry initially presented by the video playback device, for the first 3D space region and the sub - space regions of the first 3D space region, what is initially presented by the video playback device is the first 3D space region; When the value of the third identifier is the second value, indicating that the entry corresponding to the first GPCC is the entry initially presented by the video playback device, for the first 3D space region and the sub - space regions of the first 3D space region, what is initially presented by the video playback device is the sub - space region corresponding to the second value in the first 3D space region.
26. A processing device for non-temporal point cloud media, characterized in that Including: A processing unit and a communication unit; The communication unit is configured to: Receive the MPD signaling of at least one non - temporal point cloud media; Send a first request message to the video production device; Receive the first non - temporal point cloud media; The processing unit is configured to play the first non - temporal point cloud media; Wherein, the at least one non - temporal point cloud media is obtained by processing the non - temporal point cloud data of a static object through the GPCC encoding method to obtain a GPCC bitstream, encapsulating the GPCC bitstream to generate entries of at least one GPCC region, and encapsulating the entries of the at least one GPCC region to generate the obtained at least one non - temporal point cloud media; For any entry of the at least one GPCC region, the entry of the GPCC region is used to represent the GPCC components of the 3D space region corresponding to the GPCC region; for any non - temporal point cloud media of the at least one non - temporal point cloud media, the non - temporal point cloud media includes: an identifier of the static object; Wherein, if the first 3D space region corresponding to the first GPCC region is divided into multiple sub - space regions, the 3D space region entry attributes included in the entry of the first GPCC region include: a third identifier; wherein, the first GPCC region is one of the at least one GPCC region; When the value of the third identifier is the first value or empty, indicating that the entry corresponding to the first GPCC is the entry initially presented by the video playback device, for the first 3D space region and the sub - space regions of the first 3D space region, what is initially presented by the video playback device is the first 3D space region; When the value of the third identifier is the second numerical value, indicating that the entry corresponding to the first GPCC is the entry initially presented by the video playback device, for the first 3D space region and the subspace regions of the first 3D space region, the subspace region corresponding to the second numerical value in the first 3D space region is initially presented by the video playback device.
27. A video production device, characterized in that, Comprising: A processor and a memory, the memory is used for storing a computer program, and the processor is used for calling and running the computer program stored in the memory to execute the method according to any one of claims 1 to 12.
28. A video playback device, characterized in that, Comprising: A processor and a memory, the memory is used for storing a computer program, and the processor is used for calling and running the computer program stored in the memory to execute the method according to any one of claims 13 to 24.
29. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program causes a computer to execute the method according to any one of claims 1 to 12.
30. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program causes a computer to execute the method according to any one of claims 13 to 24.
Citation Information
Patent Citations
Methods and apparatus for signaling spatial relationships for point cloud multimedia data tracks
TW202041020A
Information processing device and information processing method
WO2020137642A1