Data processing method, device and equipment for non-timed point cloud media

By acquiring and configuring the viewing area attribute information of non-time-series point cloud media, the problem of low transmission and consumption efficiency of non-time-series point cloud media is solved, achieving efficient transmission and flexible presentation.

CN114969394BActive Publication Date: 2025-10-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110197827.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-22
Publication Date
2025-10-28
Estimated Expiration
2041-02-22

AI Technical Summary

Technical Problem

In existing technologies, the parsing and processing efficiency of non-temporal point cloud media is relatively low, making it difficult to achieve efficient transmission and consumption.

Method used

By acquiring viewing area attribute information of non-time-series point cloud media, including indication information of recommended viewing areas, and configuring dynamic adaptive streaming media transmission signaling and attribute information data boxes, flexible presentation of non-time-series point cloud media can be supported.

Benefits of technology

It improves the transmission and consumption efficiency of non-time-series point cloud media, supports more flexible presentation formats, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114969394B_ABST
    Figure CN114969394B_ABST
Patent Text Reader

Abstract

This application provides a data processing method, apparatus, and device for non-temporal point cloud media, relating to the field of non-temporal point cloud media technology within the field of computer vision (image) technology. The data processing method for non-temporal point cloud media includes: acquiring attribute information of a viewing area corresponding to the non-temporal point cloud media; and presenting the non-temporal point cloud media based on the attribute information of the viewing area. By introducing first indication information into the attribute information of the viewing area corresponding to the non-temporal point cloud media, when indicating the existence of a recommended viewing area, it is possible to support requesting and consuming non-temporal point cloud media based on the recommended viewing area, building upon the non-temporal point cloud media encapsulation structure. This makes the transmission and consumption of non-temporal point cloud media more efficient and supports more flexible presentation formats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision (image) technology in artificial intelligence, particularly to the field of non-temporal point cloud media technology, and more specifically to data processing methods, apparatus, and devices for non-temporal point cloud media. Background Technology

[0002] With the continuous development of science and technology, it is now possible to obtain a large amount of high-precision point cloud data at a lower cost and in a shorter time period. Point cloud data is often transmitted between content production equipment and content consumption equipment in the form of point cloud media.

[0003] The transmission process of point cloud media is as follows: After encoding the point cloud media, the content production device encapsulates the encoded point cloud media to obtain an encapsulated file. The content production device then transmits the encapsulated file to the content consumption device. The content consumption device decapsulates the encapsulated file transmitted by the content production device, then decodes it, and finally presents the media file. Because point cloud media contains a large amount of point cloud data, improving the parsing and processing efficiency of point cloud media to provide a better consumption experience is a problem that the industry has been continuously working to solve. Summary of the Invention

[0004] This application provides a data processing method, apparatus, and device for non-time-series point cloud media, which enables more efficient transmission and consumption of non-time-series point cloud media and supports more flexible presentation formats for non-time-series point cloud media.

[0005] On the one hand, this application provides a data processing method for non-temporal point cloud media, including:

[0006] Obtain attribute information of the viewing area corresponding to the non-time-series point cloud media. The attribute information of the viewing area corresponding to the non-time-series point cloud media includes first indication information for indicating whether there is a recommended viewing area for the non-time-series point cloud media.

[0007] Based on the attribute information of the viewing area corresponding to the non-time-series point cloud media, the non-time-series point cloud media is presented.

[0008] On the other hand, this application provides a data processing method for non-temporal point cloud media, including:

[0009] Generate attribute information of the viewing area corresponding to the non-time-series point cloud media. The attribute information of the viewing area corresponding to the non-time-series point cloud media includes first indication information for indicating whether there is a recommended viewing area for the non-time-series point cloud media.

[0010] Based on the attribute information of the viewing area corresponding to the non-time-series point cloud media, configure the dynamic adaptive streaming media transmission DASH signaling message and the attribute information data box of the non-time-series point cloud media.

[0011] On the other hand, this application provides a data processing apparatus for point cloud media, comprising:

[0012] The acquisition unit is used to acquire attribute information of the viewing area corresponding to the non-time-series point cloud media. The attribute information of the viewing area corresponding to the non-time-series point cloud media includes first indication information for indicating whether there is a recommended viewing area for the non-time-series point cloud media.

[0013] The presentation unit is used to present the non-time-series point cloud media based on the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0014] On the other hand, this application provides a data processing apparatus for point cloud media, the method comprising:

[0015] The acquisition unit is used to generate attribute information of the viewing area corresponding to the non-time-series point cloud media. The attribute information of the viewing area corresponding to the non-time-series point cloud media includes first indication information for indicating whether there is a recommended viewing area for the non-time-series point cloud media.

[0016] The configuration unit is used to configure the dynamic adaptive streaming media transmission DASH signaling message and the attribute information data box of the non-time-series point cloud media based on the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0017] On the other hand, embodiments of this application provide a data processing device for point cloud media, the data processing device for point cloud media comprising:

[0018] Processor, adapted to implement computer instructions; and,

[0019] A computer-readable storage medium storing computer instructions adapted for loading by a processor and executing the data processing method for the point cloud media described above.

[0020] On the other hand, embodiments of this application provide a computer-readable storage medium storing computer instructions. When these computer instructions are read and executed by a processor of a computer device, the computer device performs the aforementioned data processing method for point cloud media.

[0021] The data processing method for non-time-series point cloud media provided in this application introduces attribute information of the viewing area corresponding to the non-time-series point cloud media, and introduces first indication information into the attribute information of the viewing area corresponding to the non-time-series point cloud media. When the indication is that there is a recommended viewing area for the non-time-series point cloud media, it can support requesting and consuming non-time-series point cloud media based on the recommended viewing area of ​​the non-time-series point cloud media on the basis of the non-time-series point cloud media encapsulation structure, making the transmission and consumption of non-time-series point cloud media more efficient, and supporting more flexible non-time-series point cloud media presentation formats. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic block diagram of the point cloud media data processing system provided in the embodiments of this application.

[0024] Figure 2a This is a schematic diagram of the data processing architecture for point cloud media provided in the embodiments of this application.

[0025] Figure 2b and Figure 2c This is a schematic structural diagram of the sample provided in the embodiments of this application.

[0026] Figures 3 to 7 This is a schematic flowchart of a data processing method for non-time-series point cloud media provided in an embodiment of this application.

[0027] Figure 8 and Figure 9 This is a schematic block diagram of a data processing device for non-time-series point cloud media provided in the embodiments of this application.

[0028] Figure 10 This is a schematic block diagram of a data processing device for non-time-series point cloud media provided in the embodiments of this application. Detailed Implementation

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0030] The solutions provided in this application may involve artificial intelligence technology.

[0031] Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0032] It should be understood that artificial intelligence (AI) technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0033] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0034] This application's embodiments may relate to Computer Vision (CV) technology within artificial intelligence. Computer vision is a science that studies how to enable machines to "see." More specifically, it refers to machine vision that uses cameras and computers to replace human eyes in recognizing, tracking, and measuring targets, and further performs image processing to transform the computer-processed images into those more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, attempting to establish artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0035] This application provides a technical field related to data processing of point cloud media in computer vision technology.

[0036] The following explains the concepts related to point clouds.

[0037] A point cloud is a set of discrete points in space that are randomly distributed and represent the spatial structure and surface properties of a three-dimensional object or scene.

[0038] Point cloud data is a specific record of point clouds. The point cloud data for each point can include geometric and attribute information. The geometric information of each point refers to its Cartesian three-dimensional coordinates. The attribute information for each point can include, but is not limited to, at least one of the following: color information, material information, and laser reflection intensity information. Color information can be from any color space. For example, color information can be Red, Green, Blue (RGB) information. Alternatively, color information can be luminance / chrominance (YcbCr, YUV) information. Here, Y represents luminance (Luma), Cb(U) represents blue color difference, Cr(V) represents red, and U and V represent chrominance (Chroma), which describes color difference information.

[0039] Each point in a point cloud possesses the same number of attribute information. For example, each point in a point cloud may have two attribute information: color information and laser reflection intensity. Alternatively, each point may have three attribute information: color information, material information, and laser reflection intensity. In the encapsulation process of point cloud media, the geometric information of a point can be referred to as its geometric component, and its attribute information as its attribute component. A point cloud media may include one geometric component and one or more attribute components.

[0040] Based on application scenarios, point clouds can be divided into two main categories: machine-perceived point clouds and human-perceived point clouds. Application scenarios for machine-perceived point clouds include, but are not limited to: autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, and disaster relief robots. Application scenarios for human-perceived point clouds include, but are not limited to: digital cultural heritage, free-viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction. Point clouds can be acquired through methods including, but are not limited to: computer generation, 3D laser scanning, and 3D photogrammetry. Computers can generate point clouds of virtual 3D objects and scenes. 3D scanning can obtain point clouds of static real-world 3D objects or scenes, acquiring millions of point clouds per second. 3D photography can obtain point clouds of dynamic real-world 3D objects or scenes, acquiring tens of millions of point clouds per second. Specifically, point clouds of object surfaces can be acquired using acquisition devices such as photoelectric radar, lidar, laser scanners, and multi-view cameras. Point clouds obtained based on laser measurement principles can include the 3D coordinate information of points and the laser reflection intensity of points. Point clouds obtained based on photogrammetry principles can include the three-dimensional coordinates and color information of points. Point clouds obtained by combining laser measurement and photogrammetry principles can include the three-dimensional coordinates, laser reflection intensity, and color information of points. Correspondingly, point clouds can also be classified into three types based on their acquisition methods: static point clouds, dynamic point clouds, and dynamically acquired point clouds. For static point clouds, the object is stationary, and the device acquiring the point cloud is also stationary; for dynamic point clouds, the object is moving, but the device acquiring the point cloud is stationary; for dynamically acquired point clouds, the device acquiring the point cloud is moving.

[0041] For example, in the medical field, point clouds of biological tissues and organs can be obtained using magnetic resonance imaging (MRI), computed tomography (CT), and electromagnetic positioning information. These technologies have reduced the cost and time required to acquire point clouds and improved data accuracy. This revolution in point cloud acquisition methods has made it possible to acquire large amounts of point clouds. With the continuous accumulation of large-scale point clouds, efficient storage, transmission, publishing, sharing, and standardization have become crucial for point cloud applications.

[0042] Point cloud data can be used to create point cloud media, which can be a media file. Point cloud media can include multiple media frames, each composed of point cloud data. Point cloud media can flexibly and conveniently express the spatial structure and surface properties of 3D objects or 3D scenes, and is therefore widely used. After encoding the point cloud media, the encoded stream is then encapsulated to form an encapsulated file, which can be transmitted to the user. Correspondingly, on the point cloud media player side, the encapsulated file needs to be decapsulated first, then decoded, and finally the decoded data stream is presented. The encapsulated file can also be called a point cloud file.

[0043] To date, point clouds can be encoded using a point cloud encoding framework.

[0044] Point cloud coding frameworks can be the Geometry Point Cloud Compression (G-PCC) codec framework provided by the Moving Picture Experts Group (MPEG) or the Video Point Cloud Compression (V-PCC) codec framework, or the AVS-PCC codec framework provided by the Audio Video Standard (AVS). The G-PCC codec framework can be used for compression of first-type static point clouds and third-type dynamically acquired point clouds, while the V-PCC codec framework can be used for compression of second-type dynamic point clouds. The G-PCC codec framework is also known as the point cloud codec TMC13, and the V-PCC codec framework is also known as the point cloud codec TMC2.

[0045] The data processing solution for point cloud media provided in this application embodiment.

[0046] Figure 1 This is a schematic diagram of the architecture of the point cloud media data processing system 100 provided in the embodiments of this application.

[0047] like Figure 1As shown, the point cloud media data processing system 100 includes a content consumption device 101 and a content production device 102. The content production device 102 refers to the computer equipment used by the point cloud media provider (e.g., the point cloud media content creator). This computer equipment can be a terminal (e.g., a PC, a smart mobile device, a smartphone), a server, a mobile platform (e.g., an unmanned aerial vehicle, a UAV), a robot, or any other device with point cloud media encoding and encapsulation capabilities. The content consumption device 101 refers to the computer equipment used by the point cloud media user (e.g., a user). This computer equipment can be a terminal (e.g., a PC, a smart mobile device, a VR device, a VR headset, VR glasses, etc.) with point cloud media decapsulation and decoding capabilities.

[0048] The content production device 102 and the content consumption device 101 can be directly or indirectly connected through wired or wireless communication, and this application embodiment does not impose any limitations.

[0049] Figure 2a This is a schematic diagram of the data processing architecture for point cloud media provided in the embodiments of this application. The following will combine... Figure 1 The data processing system for point cloud media shown and Figure 2a The data processing architecture of point cloud media shown herein introduces the data processing scheme for point cloud media provided in the embodiments of this application.

[0050] like Figure 2a As shown, the data processing process of point cloud media includes data processing on the content production device side and data processing on the content consumption device side. The specific processing process is as follows:

[0051] I. Data processing on the content production equipment side:

[0052] (1) The process of acquiring point cloud data.

[0053] In one implementation, from the perspective of point cloud data acquisition, point cloud data acquisition can be divided into two methods: acquiring point cloud data by capturing real-world visual scenes through a capture device and generating point cloud data through computer devices. In one implementation, the capture device can be a hardware component installed in the content creation device, such as a camera or sensor on a terminal. The capture device can also be a hardware device connected to the content creation device, such as a camera connected to a server. The capture device provides point cloud data acquisition services for the content creation device and can include, but is not limited to, any of the following: camera devices, sensing devices, and scanning devices; where camera devices can include ordinary cameras, stereo cameras, light field cameras, etc.; sensing devices can include laser devices, radar devices, etc.; and scanning devices can include 3D laser scanning devices, etc. Multiple capture devices can be deployed at specific locations in the real space to simultaneously capture point cloud data from different angles within that space, and the captured point cloud data remains synchronized both temporally and spatially. In another implementation, computer devices can generate point cloud data based on virtual 3D objects and virtual 3D scenes. Because point cloud data is acquired in different ways, the compression encoding methods corresponding to point cloud data acquired in different ways may also differ.

[0054] (2) Encoding and encapsulation process of point cloud data.

[0055] In one implementation, the content production device can use either Geometry-Based Point Cloud Compression (GPCC) or Video-Based Point Cloud Compression (VPCC) to encode the acquired point cloud data, resulting in a GPCC bitstream or a VPCC bitstream. Taking GPCC encoding as an example, the content production device encapsulates the encoded point cloud data's GPCC bitstream using file tracks. A file track refers to the encapsulation container for the encoded point cloud data's GPCC bitstream; the encapsulation container is a standard for mixing and encapsulating multimedia content (video, audio, subtitles, chapter information, etc.) generated by the encoder. Encapsulation containers simplify the synchronous playback of different multimedia content. A GPCC bitstream can be encapsulated in a single file track or multiple file tracks to form an encapsulated file. The specific cases of encapsulating a GPCC bitstream in a single file track and multiple file tracks are as follows:

[0056] ① The GPCC bitstream is encapsulated in a single file track.

[0057] When a GPCC bitstream is transmitted within a single file track, the GPCC bitstream must be declared and represented according to the transmission rules of that single file track. GPCC bitstreams encapsulated within a single file track require no further processing and can be encapsulated using the International Organization for Standardization-Based Media File Format (ISOBMFF). Specifically, each sample encapsulated within a single file track contains one or more GPCC components. A sample refers to a collection of encapsulation structures for one or more point clouds. For example, the Type-Length-Value Byte Stream Format (TLV) encapsulation structure. A sample is the unit of encapsulation in the point cloud media encapsulation process; point cloud media contains multiple samples, and a sample is typically a media frame of the point cloud media. For example, in video media, a sample is a video frame.

[0058] Figure 2b This is a schematic structural diagram of a sample provided in an embodiment of this application.

[0059] like Figure 2b As shown, when transmitting a single file track, the sample in that file track consists of the GPCC parameter set TLV, the geometric bitstream TLV, and the attribute bitstream TLV, and the sample is encapsulated into a single file track.

[0060] ② The GPCC bitstream is encapsulated in multiple file tracks.

[0061] When the encoded GPCC geometry bitstream and the encoded GPCC attribute bitstream are transmitted in different file tracks, each sample in the file track contains at least one TLV encapsulation structure that carries data for a single GPCC component, and the TLV encapsulation structure does not simultaneously contain the encoded GPCC geometry bitstream and the encoded GPCC attribute bitstream.

[0062] Assuming there are file tracks 1 and 2, sample 1 transmitted on file track 1 may contain an encoded GPCC geometry bitstream but not an encoded GPCC attribute bitstream; sample 2 transmitted on file track 2 may contain an encoded GPCC attribute bitstream but not an encoded GPCC geometry bitstream. Since the content consuming device must first decode the encoded GPCC geometry bitstream during decoding, and the decoding of the encoded GPCC attribute bitstream depends on the decoded geometry information, encapsulating the different GPCC component bitstreams in separate file tracks allows the content consuming device to access the file track carrying the encoded GPCC geometry bitstream before the encoded GPCC attribute bitstream.

[0063] Figure 2c This is a schematic structural diagram of another sample provided in the embodiments of this application.

[0064] like Figure 2c As shown, when transmitting multiple file tracks, the encoded GPCC geometric bitstream and the encoded GPCC attribute bitstream are transmitted in different file tracks. The sample in the file track consists of the GPCC parameter set TLV and the geometric bitstream TLV, but does not contain the attribute bitstream TLV. The sample is encapsulated in any of the multiple file tracks.

[0065] In one implementation, the acquired point cloud data is encoded and encapsulated by a content production device to form a point cloud media encapsulation file. This encapsulation file can be the entire media file or a media segment within a media file. The content production device records the metadata of the encapsulation file using media presentation description information according to the point cloud media file format requirements. For example, it might use a Media Presentation Description (MPD) file to record the metadata. Here, metadata refers to all information related to the presentation of the point cloud media, including descriptions of the media content, descriptions of the viewport, and signaling information related to the presentation of the media content. The content production device sends the MPD file to a content consumption device, enabling the content consumption device to request the point cloud media encapsulation file based on the relevant description information in the MDP file. Specifically, the point cloud media encapsulation file can be sent from the content production device to the content consumption device via a transmission mechanism. For example, the transmission mechanism could be Dynamic Adaptive Streaming over HTTP (DASH) or Smart Media Transport (SMT).

[0066] The content production device encapsulates the compressed point cloud data into a series of small media segments based on the Hypertext Transfer Protocol (HTTP). Each media segment has a configurable duration, typically short, but each has multiple bitrate versions, allowing for more precise network-adaptive downloading. The content consumption device adaptively selects to download and play the highest bitrate version the network can support, based on current network conditions. This ensures media quality while avoiding playback stuttering or rebuffering due to excessively high bitrates. Based on this, it can dynamically and seamlessly adapt to real-time network conditions, providing high-quality playback content with fewer stutters and significantly improving the user experience. In other words, bitrate switching is done on a segment-by-segment basis. When network bandwidth is good, the content consumption device can request the media segment with the corresponding higher bitrate; conversely, when bandwidth is poor, it downloads the media segment with the corresponding lower bitrate. Because media segments of different qualities are time-aligned, the transitions between them are smooth and natural.

[0067] A Media Presentation Description (MPD) file precisely describes the container file. An MPD file can be an Extensible Markup Language (XML) file and comprehensively describes all information about the container file, including various audio and video parameters, the duration of media segments, the bitrate and resolution of different media segments, and the corresponding Uniform Resource Locator (URL), etc. Content consuming devices first download and parse the MPD file to obtain the media segment best matched to their performance and bandwidth. An MPD file can contain one or more Adaptation Sets. For example, one Adaptation Set contains multiple video segments of the same video content at different bitrates, and another Adaptation Set contains multiple video segments of the same audio content at different bitrates. An Adaptation Set can contain multiple Representations. A representation can include a combination of one or more media contents; for example, a video file at a certain resolution can be considered a representation.

[0068] The content consumption device sends a request to the server to obtain the MPD file based on its URL. The device first parses the MPD file to obtain its content information, including video resolution, video content type, segmentation, frame rate, bitrate, and the URLs of each media segment. By analyzing this information, the device selects appropriate media segments based on the current network conditions and the client's buffer size. Then, it sends a request to the content production device to download the corresponding media segment based on the media URL and streams it. Upon receiving the received MPD file, the device decapsulates it to obtain the raw bitstream, which is then sent to the decoder for playback.

[0069] II. Data processing on the content consumption device side:

[0070] (1) The process of decapsulation and decoding of point cloud data.

[0071] In one implementation, the content consuming device can obtain the encapsulated file of point cloud media from the MDP file distributed by the content production device. The file decapsulation process on the content consuming device side is the reverse of the file encapsulation process on the content production device side. The content consuming device decapsulates the encapsulated file of the point cloud media according to the file format requirements of the point cloud media to obtain the encoded bitstream, i.e., GPCC bitstream or VPCC bitstream. The decoding process on the content consuming device side is the reverse of the encoding process on the content production device side. The content consuming device decodes the encoded bitstream to reconstruct the point cloud data. The point cloud data rendering process: In one implementation, the content consuming device renders the point cloud data obtained by decoding the GPCC bitstream based on the metadata related to rendering and viewports in the MDP file. Once rendering is complete, the visual scene corresponding to the point cloud data is presented.

[0072] In this embodiment, for the content production device, firstly, the visual scene of the real world is sampled by the acquisition device to obtain point cloud data corresponding to the visual scene of the real world; then, the acquired point cloud data is encoded using GPCC encoding or VPCC encoding to obtain a GPCC bitstream or VPCC bitstream, which may include encoded geometric bitstream and encoded attribute bitstream; next, the GPCC bitstream or VPCC bitstream is encapsulated to obtain an encapsulated file of the point cloud media, i.e., a media file or media segment. The content production device can also encapsulate metadata into the media file or media segment and send the encapsulated file of the point cloud media to the content consumption device through a transmission mechanism, such as a dynamic adaptive streaming media transmission mechanism.

[0073] For the content consumption device, the first step is to receive the encapsulated file of the point cloud media sent by the content production device. Then, the encapsulated file is decapsulated to obtain the encoded GPCC bitstream (or VPCC bitstream) and metadata. Next, the metadata in the encoded GPCC or VPCC bitstream is parsed, i.e., the encoded GPCC or VPCC bitstream is decoded to obtain point cloud data. Finally, based on the current user's viewing (window) orientation, the decoded point cloud data is rendered and displayed on the content consumption device. It should be noted that the current user's viewing (window) orientation is determined by head tracking and visual tracking functions. In addition to using a renderer to render the point cloud data in the current user's viewing (window) orientation, an audio decoder can also be used to decode and optimize the audio in the current user's viewing (window) orientation. The content production equipment encodes and encapsulates the collected point cloud data, enabling the storage and transmission of the point cloud data. The content production equipment then distributes the encapsulated point cloud media files to the content consumption equipment, enabling the publication and sharing of point cloud data. The content consumption equipment decapsulates and decodes the encapsulated point cloud media files for consumption, allowing real-world visual scenes to be presented on the content consumption equipment.

[0074] It is understood that the point cloud media data processing system described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems or scenarios.

[0075] As can be seen from the above data processing process of point cloud media, the content production device needs to encode and encapsulate the point cloud media into a point cloud media encapsulation file before it can be sent to the content consumption device. Correspondingly, the content consumption device needs to decapsulate and decode the point cloud media encapsulation file before it can render and present the point cloud media. The point cloud media data processing system provided in this application supports data boxes, such as ISOBMFF data boxes. A data box refers to a data block or object that includes metadata, that is, the data box contains the metadata of the point cloud media; a point cloud media can be associated with multiple data boxes. For example, an attribute information data box can be used to describe the attribute information of the viewing area corresponding to the point cloud media. This attribute information data box can be used to decode the encoded GPCC bitstream or VPCC bitstream.

[0076] The point cloud media involved in this application includes dynamic point cloud media and static point cloud media, with static point cloud media also referred to as non-temporal point cloud media. Currently, only a basic encapsulation method for non-temporal point cloud media is provided for static point cloud media; a scheme for determining recommended viewing areas for non-temporal point cloud media is not supported. Therefore, this application, for non-temporal point cloud media, introduces attribute information of the corresponding viewing area based on the non-temporal point cloud media encapsulation structure. This enables the summation and consumption of non-temporal point cloud media based on the attribute information of the corresponding viewing area, making the transmission and consumption of non-temporal point cloud media more efficient and supporting more flexible presentation formats.

[0077] Figure 3 This is a schematic flowchart of a point cloud media data processing method 200 provided in an embodiment of this application. This method 200 can be executed by a content consumption device in a point cloud media system, such as a content consumption client.

[0078] like Figure 3 As shown, the method 200 may include:

[0079] S210, obtain the attribute information of the viewing area corresponding to the non-time-series point cloud media. The attribute information of the viewing area corresponding to the non-time-series point cloud media includes first indication information for indicating whether the non-time-series point cloud media has a recommended viewing area. Of course, the first indication information can also be understood as indicating whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the recommended viewing area of ​​the non-time-series point cloud media.

[0080] S220, based on the attribute information of the viewing area corresponding to the non-time-series point cloud media, present the non-time-series point cloud media.

[0081] After acquiring the attribute information of the viewing area corresponding to the non-time-series point cloud media, the content consumption device can present the non-time-series point cloud media based on the specific information included in the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0082] For example, a content preparation device determines the viewing area and recommended viewing time of a point cloud file based on the content of the non-time-series point cloud media. This viewing area includes an initial viewing area and a recommended viewing area, with the recommended viewing area encompassing the initial viewing area. Based on the recommended viewing area of ​​the non-time-series point cloud media, the content preparation device generates attribute information data boxes and corresponding signaling messages during the encapsulation process. The content preparation device sends the signaling messages to a content consumption device. The content consumption device requests the corresponding encapsulated file based on the signaling messages. The content consumption device receives the encapsulated file sent by the content preparation device. Based on the signaling messages and the corresponding attribute information data boxes in the encapsulated file, the content consumption device presents the content of the non-time-series point cloud media to the user according to the initial viewing area, recommended viewing area, and recommended viewing time of the non-time-series point cloud media.

[0083] In some embodiments, if the non-temporal point cloud media does not have M recommended viewing areas, the value of the first indication information is a first value; if the non-temporal point cloud media has attribute information for the M recommended viewing areas, the value of the first indication information is a second value; M≥1. In one implementation, the attribute information of the viewing area corresponding to the non-temporal point cloud media includes the attribute information of the M recommended viewing areas; the attribute information of the M recommended viewing areas includes at least one of the following: three-dimensional spatial structure data corresponding to the M recommended viewing areas, area identifiers corresponding to the M recommended viewing areas, and title identifiers corresponding to the M recommended viewing areas. In one implementation, the attribute information of the viewing area corresponding to the non-temporal point cloud media includes the attribute information of the M recommended viewing areas; the attribute information of the viewing area corresponding to the non-temporal point cloud media also includes quantity indication information, the value of which is used to indicate the number of the M recommended viewing areas, and the number of the M recommended viewing areas is greater than 0.

[0084] When a recommended viewing area is indicated for the non-time-series point cloud media, by indicating the recommended viewing area of ​​the non-time-series point cloud media, the client can request and consume the non-time-series point cloud media according to the recommended viewing area, making the transmission and consumption of non-time-series point cloud media more efficient and supporting more flexible non-time-series point cloud media presentation formats.

[0085] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media further includes second indication information for indicating whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes an initial viewing area; if the attribute information of the viewing area corresponding to the non-time-series point cloud media does not include the initial viewing area, then the value of the second indication information is a third value; if the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the initial viewing area, then the value of the second indication information is a fourth value. In one implementation, if the non-time-series point cloud media has an initial viewing area, then the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the initial viewing area; if the non-time-series point cloud media does not have an initial viewing area, then the attribute information of the viewing area corresponding to the non-time-series point cloud media does not include the initial viewing area. Of course, the embodiments of this application are not limited thereto.

[0086] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media further includes third indication information for indicating whether the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area; if the recommended viewing area of ​​the non-time-series point cloud media does not include the initial viewing area, the value of the third indication information is a fifth value; if the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area, the value of the third indication information is a sixth value. In one implementation, if the non-time-series point cloud media has an initial viewing area and the recommended viewing area of ​​the non-time-series point cloud media does not include the initial viewing area, the value of the second indication information is a fourth value; if the non-time-series point cloud media has an initial viewing area and the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area, the value of the second indication information is either the third value or the fourth value. Of course, the embodiments of this application are not limited thereto.

[0087] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of the initial viewing area; the attribute information of the initial viewing area includes at least one of the following: the three-dimensional spatial structure data of the initial viewing area, the area identifier corresponding to the three-dimensional spatial structure data of the initial viewing area, and the title identifier corresponding to the three-dimensional spatial structure data of the initial viewing area.

[0088] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of M recommended viewing areas; the attribute information of the viewing area corresponding to the non-time-series point cloud media also includes presentation duration indication information, which is used to indicate whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the M recommended viewing areas; if the presentation duration indication information is used to indicate that the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the M recommended viewing areas, the attribute information of the viewing area corresponding to the non-time-series point cloud media also includes presentation duration information, the value of which is used to indicate the presentation duration of each of the M recommended viewing areas; M≥1.

[0089] In its implementation, this application adds several descriptive fields at the system layer, including field extensions at the file encapsulation level and the system signaling level, to support the implementation steps of this application. The following examples, using extended ISOBMFF data boxes (i.e., attribute information data boxes) and DASH signaling, define attribute information for the viewing area of ​​non-time-series point cloud files and indication signaling for the viewing area of ​​non-time-series point cloud files.

[0090] An example implementation of the syntax for the attribute information data box can be found in Table 1 below:

[0091] Table 1

[0092]

[0093]

[0094] The semantics of the syntax involved in Table 1 above are as follows:

[0095] 1. Initial viewing region indication information (initial_region_indicated):

[0096] This is used to indicate whether the attribute information data box includes attribute information for the initial viewing area of ​​non-time-series point cloud media. For example, a value of 1 indicates that the attribute information data box contains attribute information for the initial viewing area of ​​non-time-series point cloud media. A value of 0 indicates that the attribute information data box does not contain attribute information for the initial viewing area of ​​non-time-series point cloud media. For ease of description, this application refers to this initial viewing area indication information as second indication information.

[0097] 2. Recommended viewing area indication information (recommended_region_indicated):

[0098] This is used to indicate whether the attribute information data box includes attribute information for a recommended viewing area of ​​non-time-series point cloud media. For example, a value of 1 indicates that the attribute information data box contains attribute information for a recommended viewing area of ​​non-time-series point cloud media. A value of 0 indicates that the attribute information data box does not contain attribute information for a recommended viewing area of ​​non-time-series point cloud media. For ease of description, this application refers to this indication information for the recommended viewing area as the first indication information. It should be noted that when the first indication information is 1, if the recommended viewing area includes the initial viewing area, the value of the second indication information can be 0. However, this application does not impose a mandatory restriction on this; both can be set to 1 simultaneously, in which case the attribute information of the recommended viewing area can also include the attribute information of the initial viewing area.

[0099] 3. Presentation duration indication information (presentation_duration_indicated):

[0100] This indicates whether the attribute information data box includes presentation duration information corresponding to the recommended viewing area. For example, a value of 1 indicates that the attribute information data box contains information on the presentation duration of the recommended viewing area for non-time-series point cloud media. A value of 0 indicates that the attribute information data box does not contain information on the presentation duration of the recommended viewing area for non-time-series point cloud media.

[0101] 4. Three-dimensional spatial structure data (3DS spatial Region Structure):

[0102] Three-dimensional spatial structure data used to indicate the viewing area of ​​non-temporally ordered point cloud media, such as three-dimensional spatial structure data used to indicate the initial viewing area or the recommended viewing area.

[0103] 5. Quantity indication information (num_recommended_regions):

[0104] Used to indicate the number of recommended viewing areas.

[0105] 6. Presentation duration information (presentation_duration):

[0106] Used to indicate the duration of the recommended viewing area.

[0107] It should be noted that for the initial viewing area and recommended viewing area, in addition to directly indicating the 3D spatial structure data, the corresponding spatial information can also be indexed using title identifiers (tile IDs) or region identifiers (region IDs). Each title identifier corresponds to a viewing area, and each region identifier corresponds to a viewing area. The 3D spatial structure data (3DSpatialRegionStruct) can include the corresponding region identifiers.

[0108] An example implementation of the syntax for the attribute information data box can be found in Table 2 below:

[0109] Table 2

[0110]

[0111]

[0112] The semantics of the syntax involved in Table 2 above are as follows:

[0113] 7. Number of title tags (num_tiles):

[0114] Used to indicate the number of title icons corresponding to the initial viewing area or recommended viewing area.

[0115] 8. Title identifier (tile_id);

[0116] Title identifiers used to indicate the initial viewing area or recommended viewing area.

[0117] It should be understood that the meanings of the other elements in Table 2 can be found in the corresponding elements in Table 1. To avoid repetition, they will not be repeated here.

[0118] An example implementation of the syntax for the attribute information data box can be found in Table 3 below:

[0119] Table 3

[0120]

[0121]

[0122] The semantics of the syntax involved in Table 1 above are as follows:

[0123] 9. Number of region identifiers (num_regions):

[0124] The number of area identifiers used to indicate the initial viewing area or recommended viewing area.

[0125] 10. Title identifier (region_id);

[0126] This is used to indicate the area corresponding to the initial viewing area or the recommended viewing area.

[0127] It should be understood that the meanings of the other elements in Table 2 can be found in the corresponding elements in Table 1. To avoid repetition, they will not be repeated here.

[0128] It should be noted that Tables 1 to 3 are merely examples of this application and should not be construed as limiting this application. For example, in other alternative embodiments of this application, the attribute information data box can be expanded to a full box, that is, information such as a version field can be added to the attribute information data box. Furthermore, the attribute information data box in Table 1 is a data box applied to GPCC encapsulation technology, but in other alternative embodiments, the solution of this application can also be applied to VPCC encapsulation technology. The attribute information data box of non-time-series point cloud media can refer to the ISOBaseMedia File Format (ISOBMFF) data box. After acquiring the component attribute data box of the non-time-series point cloud media, the content consumption device parses the attribute information corresponding to the point cloud media according to the attribute information data box, and presents the non-time-series point cloud media based on the parsed attribute information.

[0129] For details regarding DASH signaling, please refer to Table 4 below:

[0130] Table 4

[0131]

[0132]

[0133]

[0134] The semantics of the elements involved in Table 4 above are as follows:

[0135] A descriptor is a method of representing data features, defining the syntax and semantics of those features. The Recommendation Spatial Information (RcmdSpatialInfo) descriptor is used to describe elements and attributes related to a GPCC item; this descriptor is a Supplemental Property element. For MPD files, it can contain one or more AdaptationSets. One AdaptationSet contains multiple video clips of the same video content at different bitrates, and another AdaptationSet contains multiple video clips of the same audio content at different bitrates. An AdaptationSet can contain multiple Representations. A Representation can include a combination of one or more media contents; for example, a video file at a certain resolution can be considered a Representation. This descriptor can reside at the AdaptationSet level or the Representation level. `grsi@"xxx"` indicates the elements and attributes "xxx" included in the container element of this descriptor.

[0136] It should be noted that for the initial viewing area and recommended viewing area, in addition to directly indicating the 3D spatial structure data, the corresponding spatial information can also be indexed using title identifiers (tile IDs) or region identifiers (region IDs). Each title identifier corresponds to a viewing area, and each region identifier corresponds to a viewing area. The 3D spatial structure data (3DSpatialRegionStruct) can include the corresponding region identifiers. Based on this, the corresponding DASH signaling is shown in Table 5:

[0137] Table 5

[0138]

[0139]

[0140]

[0141] Of course, the header identifier in the DASH signaling in Table 4 can also be replaced with the area identifier. To avoid repetition, this will not be elaborated here.

[0142] For specific application scenarios, the following will be combined with Figures 4 to 6 The following is a description of an example of a data processing scheme for non-time-series point cloud media provided in the embodiments of this application.

[0143] In some embodiments, S210 may include:

[0144] The system receives a Dynamic Adaptive Streaming Media Transmission (DASH) signaling message sent by a content production device. This DASH signaling message includes attribute information of the viewing area corresponding to the non-time-series point cloud media, including attribute information of the initial viewing area. Based on the attribute information of the initial viewing area of ​​the non-time-series point cloud media, the system sends an acquisition request to the content production device. This acquisition request carries target description information, which describes a target encapsulation file including the initial viewing area. The system receives the target encapsulation file returned by the content production device according to the acquisition request. The target encapsulation file includes an attribute information data box for the non-time-series point cloud media, which defines the attribute information of the viewing area corresponding to the non-time-series point cloud media. Therefore, in step S220, the target encapsulation file can be presented based on the attribute information of the viewing area corresponding to the non-time-series point cloud media in the DASH signaling message and the attribute information data box.

[0145] Figure 4 This is a schematic flowchart of a data processing method 310 for non-temporal point cloud media provided in an embodiment of this application. This method 310 can be... Figure 1 In the illustrated embodiment, the content creation device 102 and the content consumption device 101 interact and execute. For example... Figure 5 As shown, the data processing method 310 for non-time-series point cloud media may include some or all of the following:

[0146] S311, The content production device acquires a non-temporal point cloud content A. The non-temporal point cloud content A has an initial viewing area and a recommended viewing area, and each recommended viewing area has a recommended viewing time.

[0147] S312, when the content creation device encapsulates the point cloud content A, it configures the DASH signaling message for encapsulating the point cloud content A and the attribute information data box for the point cloud content A. As an example, the corresponding attribute information data box information and DASH signaling message are as follows:

[0148] F1:item1: RecommendedSpatialInfoProperty:

[0149] initial_region_indicated=1; recommended_region_indicated=1;

[0150] presentation_duration_indicated=1;

[0151] initial_region: {3d_region_id=1001, anchor=(0,0,0), region=(100,100,100)};

[0152] recommended_region:

[0153] {3d_region_id=1001, anchor=(0,0,0),

[0154] region=(100,100,100),presentation_duration=5000};{3d_region_id=1002,

[0155] anchor=(0,100,0), region=(100,100,100), presentation_duration=5000};

[0156] {3d_region_id=1003, anchor=(0,200,0),

[0157] region=(100,100,100),presentation_duration=5000}.

[0158] S313, the content creation device sends the DASH signaling message to the content consumption device. It should be noted that the information in the relevant fields of the DASH signaling message corresponds to the information in the attribute information data box; to avoid duplication, this will not be repeated here.

[0159] S314, the content consumption device requests the encapsulation file F1, which includes the initial viewing area, from the content production device according to the DASH signaling.

[0160] S315, the content creation device transmits the packaged file F1 to the content consumption device.

[0161] S316, the content consumption device presents point cloud content A to the user according to the DASH signaling and the corresponding attribute information data box information in the encapsulation file F1, based on the initial viewing area, recommended viewing area, and recommended viewing time of point cloud content A. Specifically, it first presents area 1001 (presentation time 5000ms), then area 1002 (presentation time 5000ms), and finally area 1003 (presentation time 5000ms). It should be noted that, regarding the specific presentation format, the content consumption device can directly switch the screen for the user after the presentation time has elapsed, or it can switch the screen for the user after confirmation through a prompt on the application interface; this application does not impose any restrictions on this.

[0162] In some embodiments, S210 may include:

[0163] The system receives a target encapsulation file sent by a content production device, which includes the initial viewing area of ​​the non-time-series point cloud media. The target encapsulation file includes an attribute information data box of the non-time-series point cloud media, which is used to define the attribute information of the viewing area corresponding to the non-time-series point cloud media. Based on this, in S220, the target encapsulation file can be presented based on the attribute information of the viewing area corresponding to the non-time-series point cloud media in the attribute information data box.

[0164] Figure 5 This is a schematic flowchart of a data processing method 320 for non-temporal point cloud media provided in an embodiment of this application. This method 320 can be... Figure 1 In the illustrated embodiment, the content creation device 102 and the content consumption device 101 interact and execute. For example... Figure 5 As shown, the data processing method 320 for non-time-series point cloud media may include some or all of the following:

[0165] S321, The content production device acquires a non-temporal point cloud content A. The non-temporal point cloud content A has an initial viewing area and a recommended viewing area, and each recommended viewing area has a recommended viewing time.

[0166] S322, when encapsulating the point cloud content A, the content creation device configures the attribute information data box for the point cloud content A. As an example, the corresponding attribute information data box information includes the following:

[0167] F1:item1: RecommendedSpatialInfoProperty:

[0168] initial_region_indicated=1; recommended_region_indicated=1;

[0169] presentation_duration_indicated=1;

[0170] initial_region: {3d_region_id=1001, anchor=(0,0,0), region=(100,100,100)};

[0171] recommended_region:

[0172] {3d_region_id=1001, anchor=(0,0,0),

[0173] region=(100,100,100),presentation_duration=5000};{3d_region_id=1002,

[0174] anchor=(0,100,0), region=(100,100,100), presentation_duration=5000};

[0175] {3d_region_id=1003, anchor=(0,200,0),

[0176] region=(100,100,100),presentation_duration=5000}.

[0177] S323, The content creation device transmits the packaged file F1 to the content consumption device;

[0178] S324, the content consuming device presents point cloud content A to the user according to the corresponding attribute information data box information in the encapsulation file F1, based on information such as the initial viewing area, recommended viewing area, and recommended viewing time of point cloud content A. That is, it first presents area 1001 (presentation time is 5000ms), then area 1002 (presentation time is 5000ms), and finally area 1003 (presentation time is 5000ms). It should be noted that, in terms of the specific presentation format, the content consuming device can directly switch the screen for the user after the presentation time has elapsed, or it can switch the screen for the user after the user confirms the switch through the application interface prompts. This application does not impose any restrictions on this.

[0179] In some embodiments, S210 may include:

[0180] The system receives a Dynamic Adaptive Streaming Media Transmission (DASH) signaling message sent by a content production device. This DASH signaling message includes attribute information about the viewing area corresponding to the non-time-series point cloud media, including attribute information about the initial viewing area of ​​the non-time-series point cloud media. Based on this, in step S220, an acquisition request is sent to the content production device based on the attribute information of the initial viewing area of ​​the non-time-series point cloud media. This acquisition request carries target description information, which describes a target encapsulation file including the initial viewing area. The system receives the target encapsulation file returned by the content production device according to the acquisition request. Based on the attribute information of the viewing area corresponding to the non-time-series point cloud media in the DASH signaling message, the system presents the target encapsulation file.

[0181] Figure 6This is a schematic flowchart of a data processing method 330 for non-temporal point cloud media provided in an embodiment of this application. This method 330 can be... Figure 1 In the illustrated embodiment, the content creation device 102 and the content consumption device 101 interact and execute. For example... Figure 6 As shown, the data processing method 330 for non-time-series point cloud media may include some or all of the following:

[0182] S331, The content production device acquires a non-time-series point cloud content A. The non-time-series point cloud content A has an initial viewing area and a recommended viewing area, and each recommended viewing area has a recommended viewing time.

[0183] S332, when the content creation device encapsulates the point cloud content A, it configures the DASH signaling for encapsulating the point cloud content A. As an example, the DASH signaling is as follows:

[0184] initialRegionIndicated=1; rcmdRegionIndicated=1; preDurationIndicated=1;

[0185] initial3DSpatialRegion: {3d_region_id=1001, anchor=(0,0,0), region=(100,100,100)};

[0186] rcmd3DSpatialRegion:

[0187] {3d_region_id=1001, anchor=(0,0,0),

[0188] region=(100,100,100),presentation_duration=5000};{3d_region_id=1002,

[0189] anchor=(0,100,0), region=(100,100,100), presentation_duration=5000};

[0190] {3d_region_id=1003, anchor=(0,200,0),

[0191] region=(100,100,100),presentation_duration=5000}.

[0192] S333, the content creation device sends a DASH signaling message to the content consumption device. It should be noted that the information in the relevant fields of the DASH signaling message corresponds to the information in the attribute information data box; to avoid duplication, this will not be repeated here.

[0193] S334, the content consumption device requests the encapsulation file F1, which includes the initial viewing area, from the content production device according to the DASH signaling.

[0194] S335, the content creation device transmits the packaged file F1 to the content consumption device.

[0195] S336, the content consumption device presents point cloud content A to the user according to the DASH signaling, based on information such as the initial viewing area, recommended viewing area, and recommended viewing time. Specifically, it first presents area 1001 (presentation time 5000ms), then area 1002 (presentation time 5000ms), and finally area 1003 (presentation time 5000ms). It should be noted that, regarding the specific presentation format, the content consumption device can directly switch the screen for the user after the presentation time has elapsed, or it can switch the screen for the user after confirmation through a prompt on the application interface; this application does not impose any restrictions on this.

[0196] It should be understood that the method of indicating the viewing area through three-dimensional spatial structure data is only an example of this application and should not be construed as a limitation of this application. In other embodiments of this application, the initial viewing area or recommended viewing area may also be indicated by the indication method of title label or area label.

[0197] For example, in other alternative embodiments, the information included in the attribute information data box information and / or DASH signaling involved in methods 310 to 320 can be replaced with the following information:

[0198] initialRegionIndicated=1; rcmdRegionIndicated=1; preDurationIndicated=1;

[0199] initial3DSpatialRegion: {initialTileIds:tile1,tile2};

[0200] rcmd3DSpatialRegion:

[0201] {rcmdTileIds:tile1,tile2,presentation_duration=5000};

[0202] {rcmdTileIds:tile3,tile4,presentation_duration=5000};

[0203] {rcmdTileIds:tile5,tile6,presentation_duration=5000}.

[0204] At this point, the content consumption device presents point cloud media content to the user based on the information in the attribute information data box and / or the DASH signaling, according to the initial viewing area, recommended viewing area, recommended viewing time, and other information. That is, the area corresponding to tile1+tile2 is presented first (presentation time is 5000ms), then the area corresponding to tile3+tile4 is presented (presentation time is 5000ms), and finally the area corresponding to tile5+tile6 is presented (presentation time is 5000ms).

[0205] Figure 7 This is a schematic flowchart of a point cloud media data processing method 400 provided in an embodiment of this application. This method 400 can be executed by content creation equipment in a point cloud media system. For example, devices with point cloud media encoding capabilities, such as servers, drones, and mobile terminals.

[0206] like Figure 7 As shown, the method 200 may include:

[0207] S410, Generate attribute information of the viewing area corresponding to the non-time-series point cloud media. The attribute information of the viewing area corresponding to the non-time-series point cloud media includes first indication information for indicating whether there is a recommended viewing area for the non-time-series point cloud media.

[0208] S420, based on the attribute information of the viewing area corresponding to the non-time-series point cloud media, configure the dynamic adaptive streaming media transmission DASH signaling message of the non-time-series point cloud media and the attribute information data box of the non-time-series point cloud media.

[0209] In some embodiments, if the non-temporal point cloud media does not have M recommended viewing areas, the value of the first indication information is a first value; if the non-temporal point cloud media has attribute information for the M recommended viewing areas, the value of the first indication information is a second value; M≥1. In one implementation, the attribute information of the viewing area corresponding to the non-temporal point cloud media includes the attribute information of the M recommended viewing areas; the attribute information of the M recommended viewing areas includes at least one of the following: three-dimensional spatial structure data corresponding to the M recommended viewing areas, area identifiers corresponding to the M recommended viewing areas, and title identifiers corresponding to the M recommended viewing areas. In one implementation, the attribute information of the viewing area corresponding to the non-temporal point cloud media includes the attribute information of the M recommended viewing areas; the attribute information of the viewing area corresponding to the non-temporal point cloud media also includes quantity indication information, the value of which is used to indicate the number of the M recommended viewing areas, and the number of the M recommended viewing areas is greater than 0.

[0210] When a recommended viewing area is indicated for the non-time-series point cloud media, by indicating the recommended viewing area of ​​the non-time-series point cloud media, the client can request and consume the non-time-series point cloud media according to the recommended viewing area, making the transmission and consumption of non-time-series point cloud media more efficient and supporting more flexible non-time-series point cloud media presentation formats.

[0211] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media further includes second indication information for indicating whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the initial viewing area; if the attribute information of the viewing area corresponding to the non-time-series point cloud media does not include the initial viewing area, the value of the second indication information is a third value; if the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the initial viewing area, the value of the second indication information is a fourth value.

[0212] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media further includes third indication information for indicating whether the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area; if the recommended viewing area of ​​the non-time-series point cloud media does not include the initial viewing area, the value of the third indication information is a fifth value; if the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area, the value of the third indication information is a sixth value.

[0213] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of the initial viewing area; the attribute information of the initial viewing area includes at least one of the following: the three-dimensional spatial structure data of the initial viewing area, the area identifier corresponding to the three-dimensional spatial structure data of the initial viewing area, and the title identifier corresponding to the three-dimensional spatial structure data of the initial viewing area.

[0214] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of M recommended viewing areas; the attribute information of the viewing area corresponding to the non-time-series point cloud media also includes presentation duration indication information, which is used to indicate whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the M recommended viewing areas; if the presentation duration indication information is used to indicate that the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the M recommended viewing areas, the attribute information of the viewing area corresponding to the non-time-series point cloud media also includes presentation duration information, the value of which is used to indicate the presentation duration of each of the M recommended viewing areas; M≥1.

[0215] In some embodiments, the method 400 may further include:

[0216] The system sends a Dynamic Adaptive Streaming Media Transmission (DASH) signaling message to the content consumption device. This DASH signaling message includes attribute information of the viewing area corresponding to the non-time-series point cloud media, including the attribute information of the initial viewing area of ​​the non-time-series point cloud media. The system also receives an acquisition request sent by the content consumption device to the content production device based on the attribute information of the initial viewing area of ​​the non-time-series point cloud media. This acquisition request carries target description information, which describes a target encapsulation file including the initial viewing area. The system returns the target encapsulation file to the content consumption device according to the acquisition request. The target encapsulation file includes an attribute information data box of the non-time-series point cloud media, which defines the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0217] In some embodiments, the method 400 may further include:

[0218] Send a target encapsulation file containing the initial viewing area of ​​the non-time-series point cloud media to the content consumption device. The target encapsulation file includes an attribute information data box of the non-time-series point cloud media, which is used to define the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0219] In some embodiments, the method 400 may further include:

[0220] The system sends a Dynamic Adaptive Streaming Media Transmission (DASH) signaling message to a content consumption device. This DASH signaling message includes attribute information about the viewing area corresponding to the non-time-series point cloud media, including attribute information about the initial viewing area of ​​the non-time-series point cloud media. The system also receives an acquisition request from the content consumption device based on the attribute information of the initial viewing area of ​​the non-time-series point cloud media. This acquisition request carries target description information, which describes a target encapsulation file including the initial viewing area. Finally, the system receives the target encapsulation file returned by the content production device according to the acquisition request.

[0221] Figure 8 This is a schematic diagram of the structure of a data processing device 500 for non-time-series point cloud media provided in an embodiment of this application. The data processing device 500 for non-time-series point cloud media can be used to perform… Figures 3 to 6 The corresponding steps in the data processing method for point cloud media shown.

[0222] like Figure 8 As shown, the data processing device 500 for non-time-series point cloud media may include:

[0223] The acquisition unit 510 is used to acquire attribute information of the viewing area corresponding to the non-time-series point cloud media. The attribute information of the viewing area corresponding to the non-time-series point cloud media includes first indication information for indicating whether there is a recommended viewing area for the non-time-series point cloud media.

[0224] The presentation unit 520 is used to present the non-time-series point cloud media based on the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0225] In some embodiments, if the non-temporal point cloud media does not have M recommended viewing areas, the value of the first indication information is a first value; if the non-temporal point cloud media has attribute information for the M recommended viewing areas, the value of the first indication information is a second value; M≥1. In one implementation, the attribute information of the viewing area corresponding to the non-temporal point cloud media includes the attribute information of the M recommended viewing areas; the attribute information of the M recommended viewing areas includes at least one of the following: three-dimensional spatial structure data corresponding to the M recommended viewing areas, area identifiers corresponding to the M recommended viewing areas, and title identifiers corresponding to the M recommended viewing areas. In one implementation, the attribute information of the viewing area corresponding to the non-temporal point cloud media includes the attribute information of the M recommended viewing areas; the attribute information of the viewing area corresponding to the non-temporal point cloud media also includes quantity indication information, the value of which is used to indicate the number of the M recommended viewing areas, and the number of the M recommended viewing areas is greater than 0.

[0226] When a recommended viewing area is indicated for the non-time-series point cloud media, by indicating the recommended viewing area of ​​the non-time-series point cloud media, the client can request and consume the non-time-series point cloud media according to the recommended viewing area, making the transmission and consumption of non-time-series point cloud media more efficient and supporting more flexible non-time-series point cloud media presentation formats.

[0227] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media further includes second indication information for indicating whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the initial viewing area; if the attribute information of the viewing area corresponding to the non-time-series point cloud media does not include the initial viewing area, the value of the second indication information is a third value; if the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the initial viewing area, the value of the second indication information is a fourth value.

[0228] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media further includes third indication information for indicating whether the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area; if the recommended viewing area of ​​the non-time-series point cloud media does not include the initial viewing area, the value of the third indication information is a fifth value; if the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area, the value of the third indication information is a sixth value.

[0229] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of the initial viewing area; the attribute information of the initial viewing area includes at least one of the following: the three-dimensional spatial structure data of the initial viewing area, the area identifier corresponding to the three-dimensional spatial structure data of the initial viewing area, and the title identifier corresponding to the three-dimensional spatial structure data of the initial viewing area.

[0230] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of M recommended viewing areas; the attribute information of the viewing area corresponding to the non-time-series point cloud media also includes presentation duration indication information, which is used to indicate whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the M recommended viewing areas; if the presentation duration indication information is used to indicate that the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the M recommended viewing areas, the attribute information of the viewing area corresponding to the non-time-series point cloud media also includes presentation duration information, the value of which is used to indicate the presentation duration of each of the M recommended viewing areas; M≥1.

[0231] In some embodiments, the acquisition unit 510 is specifically used for:

[0232] The system receives a Dynamic Adaptive Streaming Media Transmission (DASH) signaling message sent by a content production device. This DASH signaling message includes attribute information of the viewing area corresponding to the non-time-series point cloud media, including the attribute information of the initial viewing area of ​​the non-time-series point cloud media. Based on the attribute information of the initial viewing area of ​​the non-time-series point cloud media, the system sends an acquisition request to the content production device. This acquisition request carries target description information, which describes a target encapsulation file including the initial viewing area. The system receives the target encapsulation file returned by the content production device according to the acquisition request. This target encapsulation file includes an attribute information data box of the non-time-series point cloud media, which defines the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0233] Specifically, the presentation unit 520 is used for:

[0234] Based on the attribute information of the viewing area corresponding to the non-time-series point cloud media in the DASH signaling message and the attribute information data box of the viewing area corresponding to the non-time-series point cloud media, the target encapsulated file is presented.

[0235] In some embodiments, the acquisition unit 510 is specifically used for:

[0236] The content production device sends a target encapsulation file containing the initial viewing area of ​​the non-time-series point cloud media. The target encapsulation file includes an attribute information data box of the non-time-series point cloud media, which is used to define the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0237] Specifically, the presentation unit 520 is used for:

[0238] Based on the attribute information of the viewing area corresponding to the non-time-series point cloud media in the attribute information data box, the target encapsulated file is presented.

[0239] In some embodiments, the acquisition unit 510 is specifically used for:

[0240] Receive dynamic adaptive streaming media transmission DASH signaling message sent by the content production device; the DASH signaling message includes attribute information of the viewing area corresponding to the non-time-series point cloud media, and the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of the initial viewing area of ​​the non-time-series point cloud media.

[0241] Specifically, the presentation unit 520 is used for:

[0242] A request is sent to the content production device based on the attribute information of the initial viewing area of ​​the non-time-series point cloud media; the request carries target description information, which describes the target encapsulated file including the initial viewing area; the content production device returns the target encapsulated file according to the request; and the target encapsulated file is presented based on the attribute information of the viewing area corresponding to the non-time-series point cloud media in the DASH signaling message.

[0243] Figure 9 This is a schematic diagram of the structure of a data processing device 600 for non-time-series point cloud media provided in an embodiment of this application. The data processing device 600 for non-time-series point cloud media can be used to perform… Figures 4 to 7 The corresponding steps in the data processing method for point cloud media shown.

[0244] The generation unit 610 is used to generate attribute information of the viewing area corresponding to the non-time-series point cloud media. The attribute information of the viewing area corresponding to the non-time-series point cloud media includes first indication information for indicating whether there is a recommended viewing area for the non-time-series point cloud media.

[0245] Configuration unit 620 is used to configure the dynamic adaptive streaming media transmission DASH signaling message and the attribute information data box of the non-time-series point cloud media based on the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0246] In some embodiments, if the non-temporal point cloud media does not have M recommended viewing areas, the value of the first indication information is a first value; if the non-temporal point cloud media has attribute information for the M recommended viewing areas, the value of the first indication information is a second value; M≥1. In one implementation, the attribute information of the viewing area corresponding to the non-temporal point cloud media includes the attribute information of the M recommended viewing areas; the attribute information of the M recommended viewing areas includes at least one of the following: three-dimensional spatial structure data corresponding to the M recommended viewing areas, area identifiers corresponding to the M recommended viewing areas, and title identifiers corresponding to the M recommended viewing areas. In one implementation, the attribute information of the viewing area corresponding to the non-temporal point cloud media includes the attribute information of the M recommended viewing areas; the attribute information of the viewing area corresponding to the non-temporal point cloud media also includes quantity indication information, the value of which is used to indicate the number of the M recommended viewing areas, and the number of the M recommended viewing areas is greater than 0.

[0247] When a recommended viewing area is indicated for the non-time-series point cloud media, by indicating the recommended viewing area of ​​the non-time-series point cloud media, the client can request and consume the non-time-series point cloud media according to the recommended viewing area, making the transmission and consumption of non-time-series point cloud media more efficient and supporting more flexible non-time-series point cloud media presentation formats.

[0248] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media further includes second indication information for indicating whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the initial viewing area; if the attribute information of the viewing area corresponding to the non-time-series point cloud media does not include the initial viewing area, the value of the second indication information is a third value; if the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the initial viewing area, the value of the second indication information is a fourth value.

[0249] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media further includes third indication information for indicating whether the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area; if the recommended viewing area of ​​the non-time-series point cloud media does not include the initial viewing area, the value of the third indication information is a fifth value; if the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area, the value of the third indication information is a sixth value.

[0250] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of the initial viewing area; the attribute information of the initial viewing area includes at least one of the following: the three-dimensional spatial structure data of the initial viewing area, the area identifier corresponding to the three-dimensional spatial structure data of the initial viewing area, and the title identifier corresponding to the three-dimensional spatial structure data of the initial viewing area.

[0251] In some embodiments, the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of M recommended viewing areas; the attribute information of the viewing area corresponding to the non-time-series point cloud media also includes presentation duration indication information, which is used to indicate whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the M recommended viewing areas; if the presentation duration indication information is used to indicate that the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the M recommended viewing areas, the attribute information of the viewing area corresponding to the non-time-series point cloud media also includes presentation duration information, the value of which is used to indicate the presentation duration of each of the M recommended viewing areas; M≥1.

[0252] In some embodiments, the device 600 further includes a communication unit for:

[0253] The system sends a Dynamic Adaptive Streaming Media Transmission (DASH) signaling message to the content consumption device. This DASH signaling message includes attribute information of the viewing area corresponding to the non-time-series point cloud media, including the attribute information of the initial viewing area of ​​the non-time-series point cloud media. The system also receives an acquisition request sent by the content consumption device to the content production device based on the attribute information of the initial viewing area of ​​the non-time-series point cloud media. This acquisition request carries target description information, which describes a target encapsulation file including the initial viewing area. The system returns the target encapsulation file to the content consumption device according to the acquisition request. The target encapsulation file includes an attribute information data box of the non-time-series point cloud media, which defines the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0254] In some embodiments, the device 600 further includes a communication unit for:

[0255] Send a target encapsulation file containing the initial viewing area of ​​the non-time-series point cloud media to the content consumption device. The target encapsulation file includes an attribute information data box of the non-time-series point cloud media, which is used to define the attribute information of the viewing area corresponding to the non-time-series point cloud media.

[0256] In some embodiments, the device 600 further includes a communication unit for:

[0257] The system sends a Dynamic Adaptive Streaming Media Transmission (DASH) signaling message to a content consumption device. This DASH signaling message includes attribute information about the viewing area corresponding to the non-time-series point cloud media, including attribute information about the initial viewing area of ​​the non-time-series point cloud media. The system also receives an acquisition request from the content consumption device based on the attribute information of the initial viewing area of ​​the non-time-series point cloud media. This acquisition request carries target description information, which describes a target encapsulation file including the initial viewing area. Finally, the system receives the target encapsulation file returned by the content production device according to the acquisition request.

[0258] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details are omitted here. Specifically, the data processing device 500 for non-time-series point cloud media can correspond to the corresponding subject executing methods 200, 310, 320, or 330 of the embodiments of this application, and each unit in the data processing device 500 for point cloud media is responsible for implementing the corresponding process in the corresponding method. Similarly, the data processing device 600 for point cloud media can correspond to the corresponding subject executing methods 310, 320, 330, or 400 of the embodiments of this application, and each unit in the data processing device 600 for point cloud media is responsible for implementing the corresponding process in the corresponding method. For the sake of brevity, further details are omitted here.

[0259] It should also be understood that the various units in the point cloud media data processing apparatus involved in the embodiments of this application can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This can achieve the same operation without affecting the technical effect of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the point cloud media data processing apparatus may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by multiple units working together. According to another embodiment of this application, the point cloud media data processing apparatus involved in the embodiments of this application, and the point cloud media data processing method of the embodiments of this application, can be constructed by running a computer program (including program code) capable of performing the steps involved in the corresponding method on a general-purpose computing device including processing elements and storage elements such as a central processing unit (CPU), random access storage medium (RAM), and read-only storage medium (ROM). Computer programs can be recorded on, for example, a computer-readable storage medium, and loaded onto, through such a computer-readable storage medium. Figure 1 The corresponding method of this application embodiment is implemented in the content consumption device 101 or content production device 102 of the data processing system for point cloud media shown.

[0260] In other words, the units mentioned above can be implemented in hardware, in software instructions, or in a combination of hardware and software. Specifically, the steps of the method embodiments in this application can be completed by the integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or executed by a combination of hardware and software in the decoding processor. Optionally, the software can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads the information in the memory and completes the steps in the above method embodiments in conjunction with its hardware.

[0261] Figure 10 This is a schematic structural diagram of the point cloud media data processing device 700 provided in the embodiments of this application.

[0262] like Figure 10 As shown, the point cloud media data processing device 700 includes at least a processor 710 and a computer-readable storage medium 720. The processor 710 and the computer-readable storage medium 720 can be connected via a bus or other means. The computer-readable storage medium 720 stores a computer program 721, which includes computer instructions. The processor 710 executes the computer instructions stored in the computer-readable storage medium 720. The processor 710 is the computing and control core of the point cloud media data processing device 700, and is suitable for implementing one or more computer instructions, specifically for loading and executing one or more computer instructions to achieve corresponding method flows or corresponding functions.

[0263] As an example, processor 710 may also be referred to as a central processing unit (CPU). Processor 710 may include, but is not limited to: general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0264] As an example, the computer-readable storage medium 720 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage device; optionally, it may also be at least one computer-readable storage medium located remotely from the aforementioned processor 710. Specifically, the computer-readable storage medium 720 includes, but is not limited to, volatile memory and / or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0265] In one implementation, the data processing device 700 for the point cloud media can be... Figure 1 The content consumption device 101 in the data processing system of the point cloud media shown; the computer-readable storage medium 720 stores first computer instructions; the processor 710 loads and executes the first computer instructions stored in the computer-readable storage medium 720 to achieve... Figure 3 or Figure 5 The corresponding steps in the method embodiment shown; in specific implementation, the first computer instruction in the computer-readable storage medium 720 is loaded by the processor 710 and the corresponding steps are executed. To avoid repetition, they will not be described again here.

[0266] In one implementation, the data processing device 700 for the point cloud media can be... Figure 1The content creation device 102 in the data processing system of the point cloud media shown; the computer-readable storage medium 720 stores second computer instructions; the processor 710 loads and executes the second computer instructions stored in the computer-readable storage medium 720 to achieve... Figure 4 or Figure 5 The corresponding steps in the method embodiment shown are as follows; in specific implementation, the second computer instructions in the computer-readable storage medium 720 are loaded by the processor 710 and the corresponding steps are executed. To avoid repetition, they will not be described again here.

[0267] According to another aspect of this application, embodiments of this application also provide a computer-readable storage medium (Memory), which is a memory device in the point cloud media data processing device 700 for storing programs and data. For example, a computer-readable storage medium 720. It is understood that the computer-readable storage medium 720 here may include the built-in storage medium in the point cloud media data processing device 700, or it may include an extended storage medium supported by the point cloud media data processing device 700. The computer-readable storage medium provides storage space that stores the operating system of the point cloud media data processing device 700. Furthermore, the storage space also stores one or more computer instructions suitable for loading and execution by the processor 710, which may be one or more computer programs 721 (including program code).

[0268] According to another aspect of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. For example, computer program 721. In this case, the data processing device 700 may be a computer, and the processor 710 reads the computer instructions from the computer-readable storage medium 720, and executes the computer instructions, causing the computer to perform the data processing method for point cloud media provided in the various alternative embodiments described above.

[0269] In other words, when implemented using software, it can be implemented entirely or partially in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes of the embodiments of this application are run or the functions of the embodiments of this application are implemented. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0270] Those skilled in the art will recognize that the units and process steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0271] Finally, it should be noted that the above content is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A data processing method for non-temporal point cloud media, characterized in that, include: Obtain attribute information of the viewing area corresponding to non-time-series point cloud media. The attribute information of the viewing area corresponding to the non-time-series point cloud media includes first indication information and presentation duration indication information. The first indication information is used to indicate whether there is a recommended viewing area for the non-time-series point cloud media. The presentation duration indication information is used to indicate whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the recommended viewing area. When the first indication information indicates that the non-time-series point cloud media has a recommended viewing area, and the presentation duration indication information indicates that the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the recommended viewing area, the non-time-series point cloud media located in the recommended viewing area is presented according to the presentation duration indicated by the presentation duration indication information based on the attribute information of the viewing area corresponding to the non-time-series point cloud media; and after the presentation duration is reached, the presented content is switched directly or the presented content is switched based on user confirmation.

2. The method according to claim 1, characterized in that, If the non-time-series point cloud media does not have M recommended viewing areas, then the value of the first indication information is the first value; if the non-time-series point cloud media has the M recommended viewing areas, then the value of the first indication information is the second value; M≥1.

3. The method according to claim 2, characterized in that, The attribute information of the viewing area corresponding to the non-temporal point cloud media includes the attribute information of the M recommended viewing areas; the attribute information of the M recommended viewing areas includes at least one of the following: the three-dimensional spatial structure data corresponding to the M recommended viewing areas, the area identifier corresponding to the M recommended viewing areas, and the title identifier corresponding to the M recommended viewing areas.

4. The method according to claim 2, characterized in that, The attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of the M recommended viewing areas; the attribute information of the viewing area corresponding to the non-time-series point cloud media also includes quantity indication information, the value of which is used to indicate the number of the M recommended viewing areas, and the number of the M recommended viewing areas is greater than 0.

5. The method according to claim 1, characterized in that, The attribute information of the viewing area corresponding to the non-time-series point cloud media also includes second indication information for indicating whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the initial viewing area; If the attribute information of the viewing area corresponding to the non-time-series point cloud media does not include the initial viewing area, then the value of the second indication information is the third value; if the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the initial viewing area, then the value of the second indication information is the fourth value.

6. The method according to claim 1, characterized in that, The attribute information of the viewing area corresponding to the non-time-series point cloud media also includes third indication information for indicating whether the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area of ​​the non-time-series point cloud media; if the recommended viewing area of ​​the non-time-series point cloud media does not include the initial viewing area, the value of the third indication information is the fifth value; if the recommended viewing area of ​​the non-time-series point cloud media includes the initial viewing area, the value of the third indication information is the sixth value.

7. The method according to claim 1, characterized in that, The attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of the initial viewing area of ​​the non-time-series point cloud media; the attribute information of the initial viewing area includes at least one of the following: the three-dimensional spatial structure data of the initial viewing area, the area identifier corresponding to the three-dimensional spatial structure data of the initial viewing area, and the title identifier corresponding to the three-dimensional spatial structure data of the initial viewing area.

8. The method according to claim 1, characterized in that, If the presentation duration indication information is used to indicate that the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of M recommended viewing areas, then the attribute information of the viewing area corresponding to the non-time-series point cloud media also includes presentation duration information, and the value of the presentation duration information is used to indicate the presentation duration of each of the M recommended viewing areas; M≥1.

9. The method according to any one of claims 1 to 8, characterized in that, The acquisition of attribute information of the viewing area corresponding to non-time-series point cloud media includes: Receive dynamic adaptive streaming media transmission DASH signaling messages sent by the content production device; the DASH signaling messages include attribute information of the viewing area corresponding to the non-time-series point cloud media, and the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of the initial viewing area of ​​the non-time-series point cloud media. A request to be acquired is sent to the content production device based on the attribute information of the initial viewing area of ​​the non-time-series point cloud media; the acquisition request carries target description information, which is used to describe the target encapsulation file including the initial viewing area; The content production device returns the target encapsulation file according to the acquisition request; the target encapsulation file includes the attribute information data box of the non-time-series point cloud media, and the attribute information data box is used to define the attribute information of the viewing area corresponding to the non-time-series point cloud media; The step of presenting the non-temporal point cloud media based on the attribute information of the viewing area corresponding to the non-temporal point cloud media includes: The target encapsulated file is presented based on the attribute information of the viewing area corresponding to the non-time-series point cloud media in the DASH signaling message and the attribute information data box of the viewing area corresponding to the non-time-series point cloud media.

10. The method according to any one of claims 1 to 8, characterized in that, The step of obtaining attribute information of the viewing area corresponding to non-time-series point cloud media includes: The content production device sends a target encapsulation file containing the initial viewing area of ​​the non-time-series point cloud media. The target encapsulation file includes an attribute information data box of the non-time-series point cloud media, which is used to define the attribute information of the viewing area corresponding to the non-time-series point cloud media. The step of presenting the non-temporal point cloud media based on the attribute information of the viewing area corresponding to the non-temporal point cloud media includes: Based on the attribute information of the viewing area corresponding to the non-time-series point cloud media in the attribute information data box, the target encapsulated file is presented.

11. The method according to any one of claims 1 to 8, characterized in that, The acquisition of attribute information of the viewing area corresponding to non-time-series point cloud media includes: Receive dynamic adaptive streaming media transmission DASH signaling messages sent by the content production device; the DASH signaling messages include attribute information of the viewing area corresponding to the non-time-series point cloud media, and the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the attribute information of the initial viewing area of ​​the non-time-series point cloud media. The step of presenting the non-temporal point cloud media based on the attribute information of the viewing area corresponding to the non-temporal point cloud media includes: A request to be acquired is sent to the content production device based on the attribute information of the initial viewing area of ​​the non-time-series point cloud media; the acquisition request carries target description information, which is used to describe the target encapsulation file including the initial viewing area; The content creation device returns the target packaged file based on the acquisition request. Based on the attribute information of the viewing area corresponding to the non-time-series point cloud media in the DASH signaling message, the target encapsulated file is presented.

12. A data processing device for point cloud media, characterized in that, include: The acquisition unit is used to acquire attribute information of the viewing area corresponding to the non-time-series point cloud media. The attribute information of the viewing area corresponding to the non-time-series point cloud media includes first indication information and presentation duration indication information. The first indication information is used to indicate whether the non-time-series point cloud media has a first indication information for a recommended viewing area. The presentation duration indication information is used to indicate whether the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the recommended viewing area. The presentation unit is configured to, when the first indication information indicates that there is a recommended viewing area for the non-time-series point cloud media, and the presentation duration indication information indicates that the attribute information of the viewing area corresponding to the non-time-series point cloud media includes the presentation duration of the recommended viewing area, present the non-time-series point cloud media located within the recommended viewing area based on the attribute information of the viewing area corresponding to the non-time-series point cloud media and according to the presentation duration indicated by the presentation duration indication information; and after the presentation duration is reached, directly switch the presented content or switch the presented content based on user confirmation.

13. A data processing device for point cloud media, characterized in that, include: A processor, adapted to execute computer programs; A computer-readable storage medium storing a computer program, which, when executed by the processor, implements the data processing method for point cloud media as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Media data processing method and device

    CN109218755A

  • Dash-based streaming of point cloud content based on recommended viewports

    US20210006614A1