Media data processing method, device, equipment and readable storage medium

By encoding and encapsulating the auxiliary information and unassigned indication information of point cloud media, the assigned indication information is generated, which solves the problem of point cloud code stream redundancy and realizes efficient utilization of network transmission resources.

CN115037943BActive Publication Date: 2025-08-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210586962.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-08-08
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

In point cloud media processing, the point cloud code stream is redundant due to different types of auxiliary information, resulting in waste of network transmission resources.

Method used

By obtaining the auxiliary information of point cloud media and unassigned general indication information, encode and encapsulate it, and a media file with assigned general indication information is generated, which is used to indicate auxiliary information and reduce the redundancy of point cloud code streams.

Benefits of technology

It reduces the redundancy of point cloud code streams, reduces the waste of network transmission resources, and improves transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115037943B_ABST
    Figure CN115037943B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a media data processing method, apparatus, device and readable storage medium, the method comprising: obtaining at least one auxiliary information and unassigned general indication information of point cloud media; encoding the point cloud media to obtain a point cloud code stream; encapsulating the point cloud code stream, at least one auxiliary information and unassigned general indication information to obtain a media file; the media file includes assigned general indication information, which is obtained by assigning a value to unassigned general indication information based on at least one auxiliary information; the assigned general indication information is used to indicate at least one auxiliary information. By adopting the present application, the redundancy of point cloud code streams and the waste of network transmission resources can be reduced. The embodiments of the present invention can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a media data processing method, apparatus, device, and readable storage medium. Background Art

[0002] Immersive media refers to media content that can bring an immersive experience to business objects. Point cloud media is a typical immersive media.

[0003] In the prior art, a content production device first encodes the point cloud media to obtain a point cloud code stream, and then encapsulates the point cloud media. During encapsulation, if first auxiliary information is present, the content production device encapsulates the point cloud code stream with the first auxiliary information to obtain a first media file corresponding to the point cloud media. If second auxiliary information of a different type than the first auxiliary information is present, the point cloud code stream is encapsulated with the second auxiliary information to obtain a second media file corresponding to the point cloud media. Obviously, for different types of auxiliary information, the prior art will produce multiple different media files, but these multiple different media files all include the same point cloud code stream, resulting in redundancy of the point cloud code stream, and thus, when transmitting multiple different media files, it will lead to a waste of network transmission resources. Summary of the Invention

[0004] The embodiments of the present application provide a media data processing method, apparatus, device, and readable storage medium, which can reduce the redundancy of point cloud code streams and reduce the waste of network transmission resources.

[0005] An embodiment of the present application provides a method for processing media data, including:

[0006] Obtaining at least one auxiliary information of the point cloud media and unassigned general indication information;

[0007] Encode the point cloud media to obtain a point cloud code stream;

[0008] The point cloud code stream, at least one auxiliary information and unassigned general indication information are encapsulated to obtain a media file; the media file includes the assigned general indication information, which is obtained by assigning a value to the unassigned general indication information based on the at least one auxiliary information; the assigned general indication information is used to indicate the at least one auxiliary information.

[0009] An embodiment of the present application provides a method for processing media data, including:

[0010] Acquire a media file; the media file includes assigned general indication information obtained based on at least one auxiliary information of the point cloud media; the assigned general indication information is used to indicate the at least one auxiliary information;

[0011] The media file is decapsulated to obtain a point cloud code stream and at least one auxiliary information, and the point cloud code stream is decoded to obtain point cloud media.

[0012] An embodiment of the present application provides a media data processing device, including:

[0013] An information acquisition module, configured to acquire at least one auxiliary information of the point cloud media and unassigned general indication information;

[0014] The information encapsulation module is used to encode the point cloud media to obtain the point cloud code stream;

[0015] The information encapsulation module is also used to encapsulate the point cloud code stream, at least one auxiliary information and unassigned general indication information to obtain a media file; the media file includes the assigned general indication information, and the assigned general indication information is obtained by assigning a value to the unassigned general indication information based on at least one auxiliary information; the assigned general indication information is used to indicate at least one auxiliary information.

[0016] The assigned general indication information includes at least one of the assigned data box information and the assigned metadata track information.

[0017] Wherein, at least one auxiliary information includes auxiliary information B c ; c is a positive integer, and c is less than or equal to the total number of at least one type of auxiliary information;

[0018] Auxiliary Information B c Includes target range D for indicating point cloud media c Assistance level parameter; target range D c Including spatial range E c or time range F c ;

[0019] When the media file includes auxiliary information B c When the first auxiliary information data box is included in the sample entry of the point cloud track corresponding to the media file, the target range D c Including spatial range E c ; The assigned data box information includes the first auxiliary information data box;

[0020] Among them, the auxiliary information B c An auxiliary level parameter in is used to indicate a spatial range; a spatial range includes a spatial area respectively contained in G point cloud frames; the point cloud media includes G point cloud frames; G is a positive integer; or

[0021] When the media file includes auxiliary information B cWhen the auxiliary information metadata track is set, the target range D c Including time range F c ; The assigned metadata track information includes the auxiliary information metadata track;

[0022] Among them, the auxiliary information metadata track includes auxiliary information B c M associated sample numbers; wherein one sample number corresponds to one point cloud frame; the point cloud media includes M point cloud frames; M is a positive integer.

[0023] The first auxiliary information data box includes an auxiliary information type field whose field value is a first type value, a static data structure quantity field, and an auxiliary information data structure with default attributes;

[0024] The first type value represents auxiliary information B c The auxiliary information type;

[0025] The Static Data Structure Number field is used to indicate the total number of auxiliary information data structures with static attributes.

[0026] The value of the static data structure number field is H, indicating H auxiliary information data structures with static attributes; the H auxiliary information data structures with static attributes include auxiliary information data structures I with static attributes. j , where H and j are both positive integers, and j is less than or equal to H;

[0027] Auxiliary information data structure with static attributes I j Include field value for auxiliary level parameter L j The auxiliary level field and the target range indication field whose field value is the first indication value; the auxiliary level parameter L j Different from the default assistance level parameters; the default assistance level parameters and the assistance level parameters L j All belong to auxiliary information B c The auxiliary level parameter in ;

[0028] The first indicator value represents the assistance level parameter L j Used to indicate the auxiliary level of a spatial extent.

[0029] The auxiliary information data structure with default attributes includes an auxiliary level field whose field value is a default auxiliary level parameter and a target range indication field whose field value is a second indication value; the default auxiliary level parameter belongs to the auxiliary information B c The auxiliary level parameter in ;

[0030] The second indication value indicates that the default assistance level parameter is used to indicate an assistance level of a spatial range; the spatial range indicated by the default assistance level parameter includes a spatial range in the point cloud media except for the spatial range indicated by the auxiliary information data structure with static attributes.

[0031] Wherein, the first auxiliary information data box includes an auxiliary algorithm indication field;

[0032] When the value of the auxiliary algorithm indication field is the third indication value, it indicates that the auxiliary information B c There are auxiliary algorithms;

[0033] When the value of the auxiliary algorithm indication field is the fourth indication value, it indicates that the auxiliary information B c There is no auxiliary algorithm; the fourth indicator value is different from the third indicator value.

[0034] Wherein, when the field value of the auxiliary algorithm indication field is the third indication value, the first auxiliary information data box further includes an auxiliary algorithm type field;

[0035] When the field value of the auxiliary algorithm type field is the second type value, it indicates that the auxiliary information B c Determined by the auxiliary information detection algorithm;

[0036] When the field value of the auxiliary algorithm type field is the third type value, it indicates that the auxiliary information B c Determined by data statistics; the third type value is different from the second type value.

[0037] The media data processing device further includes:

[0038] The signaling transmission module is used to transmit the transmission signaling for the media file to the client when the first auxiliary information data box exists; the transmission instruction carries the auxiliary information description data K c ; Auxiliary information description data K c Used to instruct the client to determine the order of obtaining different media sub-files in the media file when obtaining the media file through streaming transmission; auxiliary information description data K c Based on auxiliary information B c Generated.

[0039] Among them, the auxiliary information description data K c Includes a default level indication field;

[0040] When the field value of the default level indication field is the fifth indication value, it indicates that the auxiliary level parameter corresponding to the default information indication field is the default auxiliary level parameter;

[0041] When the field value of the default level indication field is the sixth indication value, it indicates that the auxiliary level parameter corresponding to the default level indication field is not the default auxiliary level parameter; the sixth indication value is different from the fifth indication value; the auxiliary level parameter corresponding to the default level indication field belongs to the auxiliary information B c The auxiliary level parameter in .

[0042] The sample entry of the auxiliary information metadata track includes an effective range indication field and a second auxiliary information data box; the effective range indication field is used to indicate the effective range of the auxiliary level parameters corresponding to the M sample numbers; the second auxiliary information data box is used to indicate the auxiliary information B c The information with static properties in the M sample numbers belongs to the auxiliary level parameters corresponding to the auxiliary information B c The auxiliary level parameter in .

[0043] Wherein, the M sample numbers include a sample number τ, τ is a positive integer and τ is less than or equal to M;

[0044] When the field value of the effective range indication field is the seventh indication value, it indicates that the effective range of the auxiliary level parameter corresponding to the sample number τ is the auxiliary information metadata track;

[0045] When the field value of the effective range indication field is the eighth indication value, it means that the effective range of the auxiliary level parameter corresponding to the sample number τ is the point cloud frame corresponding to the sample number τ; the eighth indication value is different from the seventh indication value.

[0046] Among them, the auxiliary information metadata track includes sample number O n ; n is a positive integer and n is less than or equal to M; time range F c Including sample serial number O n The corresponding point cloud frame;

[0047] The auxiliary information metadata track includes the information for sample number O n The auxiliary level indication field of

[0048] When the field value of the auxiliary level indication field is the ninth indication value, it indicates that the auxiliary level indication field is the ninth indication value, which is the same as the sample number 0. n The associated assistance level is determined by the default assistance level parameter; the default assistance level parameter belongs to the assistance information B c The auxiliary level parameter in ;

[0049] When the field value of the auxiliary level indication field is the tenth indication value, it indicates that the auxiliary level indication field is the same as the sample number 0. n The associated assistance level is determined by the assistance information data structure; the tenth indicator value is different from the ninth indicator value.

[0050] When the field value of the auxiliary level indication field is the ninth indication value, it indicates that the default auxiliary level parameter is used to indicate the sample sequence number 0. n The assistance level of the corresponding point cloud frame.

[0051] When the field value of the auxiliary level indication field is the tenth indication value, the auxiliary information metadata track also includes the value for the sample number 0. n The data structure quantity field is used to indicate the total number of auxiliary information data structures.

[0052] The value of the data structure quantity field is P, indicating P auxiliary information data structures; the P auxiliary information data structures include an auxiliary information data structure Qr, where r and P are both positive integers and r is less than or equal to P;

[0053] Auxiliary information data structure Q r Include field value for auxiliary level parameter S r The auxiliary level field and the target range indication field; the auxiliary level parameter S r Belongs to auxiliary information B c The auxiliary level parameter in ;

[0054] When the field value of the target range indication field is the first indication value, it indicates that the auxiliary level parameter S r For indication, sample number O n The assistance level of a spatial region in the corresponding point cloud frame;

[0055] When the field value of the target range indication field is the second indication value, it indicates that the auxiliary level parameter S r For indication, sample number O n The auxiliary level of the corresponding point cloud frame; the second indication value is different from the first indication value.

[0056] Wherein, the at least one assistance information includes target assistance information; the target assistance information includes an assistance level parameter for indicating a target range;

[0057] Information encapsulation module, including:

[0058] The first encoding unit is used to optimize the encoding of the target range of the point cloud media according to the auxiliary level parameter in the target auxiliary information to obtain a point cloud code stream.

[0059] The total number of target ranges is at least two, and the at least two target ranges include a first target range and a second target range; the assistance level parameter in the target assistance information includes a first assistance level parameter corresponding to the first target range, and a second assistance level parameter corresponding to the second target range;

[0060] The first encoding unit includes:

[0061] a first determining subunit, configured to determine a first coding level for a first target range according to a first assistance level parameter, and determine a second coding level for a second target range according to a second assistance level parameter, wherein when the first assistance level parameter is greater than the second assistance level parameter, the first coding level is superior to the second coding level;

[0062] The first encoding subunit is configured to optimize and encode the first target range using a first encoding level to obtain a first sub-point cloud code stream, and optimize and encode the second target range using a second encoding level to obtain a second sub-point cloud code stream;

[0063] The first generating sub-unit is configured to generate a point cloud code stream according to the first sub-point cloud code stream and the second sub-point cloud code stream.

[0064] The at least one auxiliary information includes first auxiliary information and second auxiliary information; the first auxiliary information includes an auxiliary level parameter σ for indicating a target range of the point cloud media; the second auxiliary information includes an auxiliary level parameter ρ for indicating a target range of the point cloud media; the auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information;

[0065] Information encapsulation module, including:

[0066] The second encoding unit is used to optimize the encoding of the target range according to the auxiliary level parameter σ and the auxiliary level parameter ρ to obtain a point cloud code stream.

[0067] The total number of target ranges is at least two, and the at least two target ranges include a third target range and a fourth target range; the assistance level parameter σ includes a third assistance level parameter corresponding to the third target range and a fourth assistance level parameter corresponding to the fourth target range; the assistance level parameter ρ includes a fifth assistance level parameter corresponding to the third target range and a sixth assistance level parameter corresponding to the fourth target range;

[0068] The second encoding unit includes:

[0069] a first summing subunit, configured to perform a weighted summation on the third assistance level parameter and the fifth assistance level parameter to obtain a first total assistance level parameter corresponding to the third target range;

[0070] a second summing subunit, configured to perform a weighted summation on the fourth assistance level parameter and the sixth assistance level parameter to obtain a second total assistance level parameter corresponding to the fourth target range;

[0071] a second determining subunit, configured to determine a third coding level of a third target range based on the first total assistance level parameter, and determine a fourth coding level of a fourth target range based on the second total assistance level parameter; when the first total assistance level parameter is greater than the second total assistance level parameter, the third coding level is superior to the fourth coding level;

[0072] The second encoding subunit is configured to optimize and encode the third target range using a third encoding level to obtain a third sub-point cloud code stream, and optimize and encode the fourth target range using a fourth encoding level to obtain a fourth sub-point cloud code stream;

[0073] The second generating sub-unit is configured to generate a point cloud code stream according to the third point cloud sub-code stream and the fourth point cloud sub-code stream.

[0074] An embodiment of the present application provides a media data processing device, including:

[0075] A file acquisition module is used to acquire a media file; the media file includes assigned general indication information obtained based on at least one auxiliary information of the point cloud media; the assigned general indication information is used to indicate the at least one auxiliary information;

[0076] The code stream decoding module is used to decapsulate the media file to obtain a point cloud code stream and at least one auxiliary information, and decode the point cloud code stream to obtain point cloud media.

[0077] Wherein, the at least one auxiliary information includes target auxiliary information;

[0078] The target assistance information includes an assistance level parameter for indicating a target range of the point cloud media; the target range includes a spatial range or a temporal range;

[0079] The assigned general indication information includes at least one of assigned data box information or assigned metadata track information.

[0080] The media file includes an auxiliary information data box for indicating target auxiliary information; the assigned data box information includes the auxiliary information data box;

[0081] The media data processing device further includes:

[0082] a first determining module, configured to determine that the target range includes a spatial range when the auxiliary information data box is included at a sample entry of a point cloud track corresponding to the media file;

[0083] Among them, an auxiliary level parameter in the target auxiliary information is used to indicate a spatial range; a spatial range includes a spatial area respectively contained in T point cloud frames; the point cloud media includes T point cloud frames; and T is a positive integer.

[0084] The total number of spatial ranges is at least two, and the at least two spatial ranges include a first spatial range and a second spatial range;

[0085] The media data processing device further includes:

[0086] The first acquisition module is used to obtain the assistance level parameter U for indicating the first spatial range from the target auxiliary information. v , obtain the auxiliary level parameter U used to indicate the second spatial range v+1 ; v is a positive integer, and v is less than the total number of assistance level parameters in the target assistance information;

[0087] The second determination module is used to determine if the assistance level parameter U v Greater than the auxiliary level parameter U v+1 , it is determined that the rendering level corresponding to the first spatial range is better than the rendering level corresponding to the second spatial range;

[0088] The second determination module is also used to determine if the assistance level parameter U v Less than the auxiliary level parameter U v+1 , it is determined that the rendering level corresponding to the second spatial range is better than the rendering level corresponding to the first spatial range.

[0089] The media data processing device further includes:

[0090] a third determining module, configured to, when the media file includes an auxiliary information metadata track for indicating target auxiliary information, determine that the target range includes a time range; and the assigned metadata track information includes the auxiliary information metadata track;

[0091] The auxiliary information metadata track includes W sample numbers associated with the target auxiliary information; wherein one sample number corresponds to one point cloud frame; the point cloud media includes W point cloud frames; and W is a positive integer.

[0092] The time range includes the first point cloud frame and the second point cloud frame in the W point cloud frames;

[0093] The media data processing device further includes:

[0094] The second acquisition module is used to obtain the auxiliary level parameter X indicating the first point cloud frame in the target auxiliary information. y , get the auxiliary level parameter X used to indicate the second point cloud frame y+1 ; y is a positive integer, and y is less than the total number of auxiliary level parameters in the auxiliary information;

[0095] The fourth determining module is used to determine if the assistance level parameter X y Greater than the auxiliary level parameter X y+1, it is determined that the rendering level corresponding to the first point cloud frame is better than the rendering level corresponding to the second point cloud frame;

[0096] The fourth determining module is further configured to determine if the auxiliary level parameter X y Less than the auxiliary level parameter X y+1 , it is determined that the rendering level corresponding to the second point cloud frame is better than the rendering level corresponding to the first point cloud frame.

[0097] The media data processing device further includes:

[0098] The second acquisition module is further configured to, if the first point cloud frame includes at least two spatial regions and the at least two spatial regions correspond to different assistance levels, obtain, in the target auxiliary information, an assistance level parameter Zα corresponding to the first spatial region and an assistance level parameter Zα+1 corresponding to the second spatial region; the first spatial region and the second spatial region both belong to the at least two spatial regions; α is a positive integer and is less than the total number of assistance level parameters in the auxiliary information;

[0099] a fifth determining module, configured to determine that the rendering level corresponding to the first spatial region is superior to the rendering level corresponding to the second spatial region if the auxiliary level parameter Zα is greater than the auxiliary level parameter Zα+1;

[0100] The fifth determining module is further configured to determine that the rendering level corresponding to the second spatial region is better than the rendering level corresponding to the first spatial region if the auxiliary level parameter Zα is less than the auxiliary level parameter Zα+1.

[0101] The at least one auxiliary information includes first auxiliary information and second auxiliary information; the first auxiliary information includes an auxiliary level parameter ζ for indicating a target range of the point cloud media; the second auxiliary information includes an auxiliary level parameter ψ for indicating a target range of the point cloud media; the total number of target ranges is at least two; the at least two target ranges include a target range β and a target range δ; the auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information;

[0102] The media data processing device further includes:

[0103] A third acquisition module is configured to acquire, from the assistance level parameter ζ, an assistance level parameter ε corresponding to the target range β and an assistance level parameter φ corresponding to the target range δ;

[0104] The third acquisition module is further configured to acquire, from the assistance level parameter ψ, an assistance level parameter γ corresponding to the target range β and an assistance level parameter μ corresponding to the target range δ;

[0105] A weighted summation module is used to perform a weighted summation of the assistance level parameter ε and the assistance level parameter γ to obtain a total assistance level parameter η corresponding to the target range β;

[0106] The weighted summation module is further configured to perform a weighted summation of the assistance level parameter φ and the assistance level parameter μ to obtain a total assistance level parameter λ corresponding to the target range δ;

[0107] a sixth determining module, configured to determine, if the total assistance level parameter η is greater than the total assistance level parameter λ, that the rendering level corresponding to the target range β is superior to the rendering level corresponding to the target range δ;

[0108] The sixth determining module is further configured to determine that the rendering level corresponding to the target range δ is better than the rendering level corresponding to the target range β if the total assistance level parameter η is less than the total assistance level parameter λ.

[0109] On one hand, the present application provides a computer device, including: a processor, a memory, and a network interface;

[0110] The above-mentioned processor is connected to the above-mentioned memory and the above-mentioned network interface, wherein the above-mentioned network interface is used to provide data communication function, the above-mentioned memory is used to store computer programs, and the above-mentioned processor is used to call the above-mentioned computer program so that the computer device executes the method in the embodiment of the present application.

[0111] On one hand, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is suitable for being loaded by a processor and executing the method in the embodiment of the present application.

[0112] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium; a processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method in the embodiment of the present application.

[0113] The embodiment of the present application proposes a kind of unassigned general indication information. When at least one kind of auxiliary information of the point cloud media is obtained, the point cloud code stream corresponding to the point cloud media, at least one kind of auxiliary information and the unassigned general indication information can be encapsulated. During the encapsulation process, the unassigned general indication information can be assigned a value based on at least one kind of auxiliary information to obtain the assigned general indication information, and then a media file including the assigned general indication information can be obtained, wherein the assigned general indication information is used to indicate the at least one kind of auxiliary information mentioned above. As can be seen from the above, the embodiment of the present application can indicate any type of auxiliary information of the point cloud media by assigning a value to the unassigned general indication information, so different types of auxiliary information can be encapsulated into one media file, so it can avoid one auxiliary information corresponding to one media file, and then the redundancy of the point cloud code stream can be reduced. By transmitting a media file containing different types of auxiliary information, the waste of network transmission resources can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0115] Figure 1a is a schematic diagram of 3DoF provided in an embodiment of the present application;

[0116] Figure 1b is a schematic diagram of 3DoF+ provided in an embodiment of the present application;

[0117] Figure 1c is a schematic diagram of 6DoF provided in an embodiment of the present application;

[0118] Figure 2 This is a flow chart of an immersive media process from acquisition to consumption provided by an embodiment of the present application;

[0119] Figure 3 This is a schematic diagram of the architecture of an immersive media system provided in an embodiment of the present application;

[0120] Figure 4 1 is a flow chart of a media data processing method provided in an embodiment of the present application;

[0121] Figure 5 This is a schematic diagram of an embodiment of the present application showing that a significance level parameter is only related to time and has nothing to do with spatial regions;

[0122] Figure 6This is a schematic diagram showing that a saliency level parameter is related to a spatial region and that the saliency level parameter of the spatial region changes over time, as provided in an embodiment of the present application;

[0123] Figure 7 This is a flow diagram of a media data processing method provided by an embodiment of the present application. Figure 2 ;

[0124] Figure 8 This is a schematic diagram showing that a saliency level parameter is related to a spatial region and that the saliency level parameter related to the spatial region does not change over time, as provided in an embodiment of the present application;

[0125] Figure 9 This is a flow diagram of a media data processing method provided by an embodiment of the present application. Figure 3 ;

[0126] Figure 10 1 is a structural diagram of a media data processing device provided in an embodiment of the present application;

[0127] Figure 11 This is a schematic diagram of the structure of a media data processing device provided in an embodiment of the present application. Figure 2 ;

[0128] Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application;

[0129] Figure 13 It is a structural diagram of a data processing system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0130] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0131] The following is an introduction to some technical terms involved in the embodiments of this application:

[0132] 1. Immersive Media

[0133] Immersive media refers to media content that can provide an immersive experience, enabling business objects immersed in the media content to obtain visual, auditory and other sensory experiences in the real world. Immersive media can be divided into 3DoF media, 3DoF+ media and 6DoF media according to the degree of freedom (DoF) of business objects when consuming media content. Among them, point cloud media is a typical 6DoF media. In the embodiments of the present application, users (i.e., viewers) who consume immersive media (such as point cloud media) are collectively referred to as business objects.

[0134] 2. Point Cloud

[0135] A point cloud is a collection of randomly distributed discrete points in space that represent the spatial structure and surface properties of a three-dimensional object or scene. Each point in a point cloud contains at least 3D position information and, depending on the application scenario, may also contain color, material, or other information. Typically, each point in a point cloud has the same number of additional attributes.

[0136] Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes, and therefore have a wide range of applications, including virtual reality (VR) games, computer-aided design (CAD), geographic information systems (GIS), autonomous navigation systems (ANS), digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive telepresence, and three-dimensional reconstruction of biological tissues and organs.

[0137] Point clouds are primarily acquired through computer generation, 3D (3D) laser scanning, and 3D photogrammetry. Computers can generate point clouds of virtual 3D objects and scenes. 3D scanning can obtain point clouds of static real-world 3D objects or scenes, generating millions of point clouds per second. 3D photography can obtain point clouds of dynamic real-world 3D objects or scenes, generating tens of millions of point clouds per second. Furthermore, in the medical field, MRI (Magnetic Resonance Imaging), CT (Computed Tomography), and electromagnetic positioning information can be used to obtain point clouds of biological tissues and organs. These technologies reduce the cost and time required to acquire point cloud data while improving data accuracy. Changes in point cloud data acquisition methods have made it possible to acquire large amounts of point cloud data. With the continuous accumulation of large-scale point cloud data, efficient storage, transmission, publication, sharing, and standardization of point cloud data have become key to point cloud applications.

[0138] 3. Track:

[0139] A track is a collection of media data during the media file encapsulation process, consisting of multiple time-sequential samples. A media file can be composed of one or more tracks. For example, a media file can commonly contain a video media track, an audio media track, and a subtitle media track. Metadata information can also be included in a file as a media type in the form of a metadata media track, referred to as a metadata track in this document.

[0140] 4. Sample:

[0141] A sample is a unit of media file encapsulation. A track consists of many samples, each of which corresponds to a specific timestamp. For example, a video media track can consist of many samples, with a sample typically representing a video frame. In this embodiment of the present application, a sample in a point cloud media track can represent a point cloud frame.

[0142] 5. Sample Entry:

[0143] The sample entry is used to indicate metadata information related to all samples in the track. For example, the sample entry of a video track usually contains metadata information related to decoder initialization.

[0144] 6. Point Cloud Slice:

[0145] A point cloud slice (point cloud strip) represents a collection of syntax elements (such as geometry slices and attribute slices) of part or all of the encoded point cloud frame data.

[0146] 7. Spatial Block Area (Tile):

[0147] The hexahedral spatial block area within the boundary space area of the point cloud frame is referred to as the spatial block in this application. A spatial block consists of one or more point cloud slices, and there is no encoding and decoding dependency between spatial blocks.

[0148] 8. Auxiliary Info

[0149] Auxiliary information refers to information related to the presentation of point cloud media, including information that can assist or assist point cloud media in rendering. The embodiments of this application do not limit the type of auxiliary information, including but not limited to saliency information, credibility information, quality level and priority.

[0150] Among them, Auxiliary:

[0151] Because the human visual system naturally identifies the most prominent and salient areas in a scene during a quick scan, people are naturally drawn to the most prominent aspects of an image. The portion of an image that most captures the viewer's attention is called its salient area. Different areas of an image have varying degrees of appeal to the viewer, and the concept that characterizes the degree to which these areas attract the viewer is called saliency.

[0152] Reliability:

[0153] Due to the limitations of point cloud acquisition methods and environments, in certain environments, the points collected by the point cloud may have different qualities. This quality is also defined as credibility, that is, the quality of these points is measured by the effectiveness or trustworthiness of the collected points for point cloud applications.

[0154] Quality Grade:

[0155] The quality level is a measure of the objective quality (such as Peak Signal to Noise Ratio, abbreviated as PSNR) or subjective quality (such as human eye scoring) of the point cloud media.

[0156] Priority:

[0157] Priority refers to the priority of point cloud media processing, such as the decoding process and the transmission process.

[0158] 9. DoF (degrees of freedom):

[0159] In this application, DoF refers to the degrees of freedom of movement and content interaction supported by business objects when watching immersive media (such as point cloud media), which can include 3DoF (three degrees of freedom), 3DoF+ and 6DoF (six degrees of freedom). Among them, 3DoF refers to the three degrees of freedom of rotation of the business object's head around the x-axis, y-axis, and z-axis. 3DoF+ is based on the three degrees of freedom, and the business object also has the freedom of limited movement along the x-axis, y-axis, and z-axis. 6DoF is based on the three degrees of freedom, and the business object also has the freedom of free movement along the x-axis, y-axis, and z-axis.

[0160] 10. ISOBMFF (ISO Based Media File Format):

[0161] The media file format based on the ISO (International Standard Organization) standard is a media file encapsulation standard. A typical ISOBMFF file is an MP4 (Moving Picture Experts Group 4) file.

[0162] 11. DASH (Dynamic Adaptive Streaming over HTTP): is an adaptive bitrate technology that enables high-quality streaming media to be delivered over the Internet through traditional HTTP network servers.

[0163] 12. MPD (Media Presentation Description, media presentation description signaling in DASH) is used to describe the media segment information in the media file.

[0164] 13. Representation: refers to the combination of one or more media components in DASH. For example, a video file of a certain resolution can be regarded as a Representation.

[0165] 14. Adaptation Sets: refers to a collection of one or more video streams in DASH. An Adaptation Set can contain multiple Representations.

[0166] 15. Media Segment: A playable segment that conforms to a specific media format. Playback may require the cooperation of zero or more preceding segments and an initialization segment.

[0167] The embodiments of the present application relate to data processing technology for immersive media. Some concepts in the data processing process of immersive media will be introduced below. It should be noted that the immersive media in the subsequent embodiments of the present application are described as point cloud media as an example.

[0168] See Figure 1a , Figure 1a 3DoF is a schematic diagram of the embodiment of the present application. Figure 1a As shown, 3DoF means that the business object consuming immersive media is fixed at the center point of a three-dimensional space, and the head of the business object rotates along the X-axis, Y-axis, and Z-axis to view the picture provided by the media content.

[0169] See Figure 1b , Figure 1b This is a schematic diagram of 3DoF+ provided by the embodiment of this application. Figure 1b As shown, 3DoF+ means that when the virtual scene provided by immersive media has certain depth information, the business object head can move within a limited space based on 3DoF to view the picture provided by the media content.

[0170] See Figure 1c , Figure 1c 6DoF is a schematic diagram of the embodiment of the present application. Figure 1c As shown, 6DoF is divided into window 6DoF, omnidirectional 6DoF and 6DoF, among which window 6DoF means that the rotation movement of the business object in the X-axis and Y-axis is restricted, and the translation in the Z-axis is restricted; for example, the business object cannot see the scene outside the window frame, and the business object cannot pass through the window. Omnidirectional 6DoF means that the rotation movement of the business object in the X-axis, Y-axis and Z-axis is restricted. For example, the business object cannot freely pass through the three-dimensional 360-degree VR content in the restricted movement area. 6DoF means that the business object can freely translate along the X-axis, Y-axis and Z-axis on the basis of 3DoF. For example, the business object can move freely in the three-dimensional 360-degree VR content.

[0171] See Figure 2 , Figure 2 This is a flow chart of an immersive media from collection to consumption provided by an embodiment of the present application. Figure 2 As shown, the complete processing process for immersive media may specifically include: video acquisition, video encoding, video file encapsulation, video file transmission, video file decapsulation, video decoding and final video presentation.

[0172] Video capture is used to convert analog video into digital video and save it in a digital video file format. Specifically, video capture converts video signals (e.g., point cloud data) captured by multiple cameras from different angles into binary digital information. The binary digital information converted from the video signals is a binary data stream, which can also be referred to as the bitstream or code stream of the video signal. Video encoding, on the other hand, uses compression technology to convert files in an original video format into files in another video format. Video signals can be divided into two types: those captured by a camera and those generated by a computer. Due to different statistical characteristics, the corresponding compression encoding methods may also differ. Commonly used compression encoding methods include HEVC (High Efficiency Video Coding, an international video coding standard HEVC / H.265), VVC (Versatile Video Coding, an international video coding standard VVC / H.266), AVS (Audio Video Coding Standard, a Chinese national video coding standard), and AVS3 (the third-generation video coding standard developed by the AVS standards group).

[0173] After the video is encoded, the encoded data stream (for example, point cloud stream) needs to be encapsulated and transmitted to the business object. Video file encapsulation refers to storing the encoded and compressed video stream and audio stream in a file in a certain format according to the encapsulation format (or container, or file container). Common encapsulation formats include AVI format (Audio Video Interleaved) or ISOBMFF format. In one embodiment, the audio stream and video stream are encapsulated in a file container according to a file format such as ISOBMFF to form a media file (also called an encapsulated file, video file). The media file can be composed of multiple tracks, such as a video track, an audio track, and a subtitle track.

[0174] After the content production device performs the encoding and file encapsulation processes described above, it can transmit the media file to the client on the content consumption device. The client can then present the final video content on the client after performing reverse operations such as decapsulation and decoding. The media file can be sent to the client based on various transmission protocols, including but not limited to DASH, HLS (HTTP Live Streaming, dynamic bitrate adaptive transmission), SMTP (Smart Media Transport Protocol), TCP (Transmission Control Protocol), etc.

[0175] It's understood that the client's file decapsulation process is the inverse of the aforementioned file encapsulation process. The client can decapsulate the media file according to the file format requirements during encapsulation to obtain the audio and video streams. The client's decoding process is also the inverse of the encoding process. For example, the client can decode the video stream to restore the video content, and can also decode the audio stream to restore the audio content.

[0176] For easier understanding, please refer to Figure 3 , Figure 3 This is a schematic diagram of the architecture of an immersive media system provided by an embodiment of the present application. Figure 3 As shown, the immersive media system may include a content production device (e.g., content production device 200A) and a content consumption device (e.g., content consumption device 200B). The content production device may refer to a computer device used by a provider of point cloud media (e.g., a content producer of point cloud media). The computer device may be a terminal (e.g., a PC (Personal Computer), a smart mobile device (e.g., a smartphone), etc.) or a server. The server may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0177] A content consumption device may refer to a computer device used by a user of point cloud media (e.g., a viewer of point cloud media, i.e., a business object). The computer device may be a terminal (e.g., a PC (Personal Computer), a smart mobile device (e.g., a smartphone), a VR device (e.g., a VR helmet, VR glasses, etc.), a smart home appliance, a vehicle-mounted terminal, an aircraft, etc.), and the computer device may be integrated with a client. Content production equipment and content consumption equipment may be directly or indirectly connected via wired or wireless communication, and this application does not impose any restrictions thereon.

[0178] The client may be a client capable of displaying data information such as text, images, audio, and video, including but not limited to a multimedia client (e.g., a video client), a social client (e.g., an instant messaging client), an information application (e.g., a news client), an entertainment client (e.g., a game client), a shopping client, an in-car client, a browser, etc. The client may be a standalone client or an embedded sub-client integrated into a client (e.g., a social client), without limitation.

[0179] It is understood that the data processing technology for immersive media in this application can be implemented using cloud technology; for example, using cloud servers as content production devices. Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing.

[0180] The data processing process of point cloud media includes the data processing process on the content production device side and the data processing process on the content consumption device side.

[0181] The data processing process on the content production device side mainly includes: (1) the acquisition and production process of the media content of the point cloud media; (2) the encoding and file packaging process of the point cloud media. The data processing process on the content consumption device side mainly includes: (1) the file decapsulation and decoding process of the point cloud media; (2) the rendering process of the point cloud media. In addition, the transmission process of the point cloud media between the content production device and the content consumption device can be carried out based on various transmission protocols, which may include but are not limited to: DASH protocol, HLS protocol, SMT protocol, TCP protocol, etc.

[0182] The following will be combined Figure 3 , each process involved in the data processing of point cloud media is briefly introduced.

[0183] 1. Data processing on the content production device side:

[0184] (1) The acquisition and production process of media content of point cloud media.

[0185] 1) The process of acquiring media content of point cloud media.

[0186] The media content of point cloud media is obtained by capturing the sound and visual scenes of the real world using a capture device. In one implementation, the capture device can refer to a hardware component located in the content production device, such as a microphone, camera, or sensor on a terminal. In another implementation, the capture device can also be a hardware device connected to the content production device, such as a camera connected to a server, which provides point cloud media content acquisition services for the content production device. The capture device may include, but is not limited to, audio equipment, video equipment, and sensor equipment. Audio equipment may include audio sensors, microphones, etc. Video equipment may include ordinary cameras, stereo cameras, light field cameras, etc. Sensor equipment may include laser equipment, radar equipment, etc. There may be multiple capture devices, which are deployed at specific locations in real space to simultaneously capture audio and video content from different angles within the space, with the captured audio and video content synchronized in both time and space. In embodiments of the present application, the media content in a three-dimensional space captured by capture devices deployed at specific locations to provide a multi-degree-of-freedom (e.g., 6DoF) viewing experience may be referred to as point cloud media.

[0187] For example, take the video content of point cloud media as an example. Figure 3 As shown, a visual scene 20A (e.g., a real-world visual scene) can be captured by a camera array connected to a content production device 200A, or by a camera device having multiple cameras and sensors connected to the content production device 200A. The captured result can be source point cloud data 20B (i.e., video content of the point cloud media).

[0188] 2) The production process of media content of point cloud media.

[0189] It should be understood that the process of producing the media content of the point cloud media involved in the embodiments of the present application can be understood as the process of content production of the point cloud media, and the content production of the point cloud media here is mainly composed of content in the form of point cloud data obtained by cameras or camera arrays deployed at multiple locations. For example, the content production equipment can convert the point cloud media from a three-dimensional representation to a two-dimensional representation.

[0190] In addition, it should be noted that since panoramic videos can be captured by capture devices, such videos are processed by content production devices and transmitted to content consumption devices for corresponding data processing. Business objects on the content consumption device side need to perform some specific actions (such as head rotation) to watch 360-degree video information, while performing non-specific actions (such as moving the head) cannot obtain corresponding video changes, and the VR experience is not good. Therefore, it is necessary to provide additional depth information that matches the panoramic video to enable business objects to obtain better immersion and a better VR experience, which involves 6DoF production technology. When business objects can move more freely in a simulated scene, it is called 6DoF. When using 6DoF production technology to produce video content of point cloud media, capture devices generally use laser equipment, radar equipment, etc. to capture point cloud data in space.

[0191] (2) The process of encoding and file packaging of point cloud media.

[0192] The captured audio content can be directly audio-encoded to form an audio stream of the point cloud media. The captured video content can be video-encoded to obtain a video stream of the point cloud media. It should be noted here that if 6DoF production technology is used, a specific encoding method (such as a point cloud compression method based on traditional video encoding) needs to be used for encoding during the video encoding process. The content production device encapsulates the audio stream and the video stream in a file container according to the file format of the point cloud media (such as ISOBMFF) to form a media file resource of the point cloud media. The media file resource can be a media file of the point cloud media formed by a media file or a media fragment; and according to the file format requirements of the point cloud media, the media presentation description information (i.e., MPD) is used to record the metadata of the media file resource of the point cloud media. The metadata here is a general term for information related to the presentation of point cloud media. The metadata may include description information of the media content, description information of the window, and signaling information related to the presentation of the media content, etc. It can be understood that the content production device will store the media presentation description information (including auxiliary information in this application) and media file resources formed after the data processing process.

[0193] like Figure 3As shown, the content production device 200A performs point cloud media encoding on one or more data frames in the source point cloud data 20B, for example, using geometry-based point cloud compression (G-PCC, where PCC is point cloud compression), thereby obtaining an encoded point cloud code stream 20E (i.e., a video code stream, such as a G-PCC code stream). Subsequently, the content production device 200A can encapsulate the one or more encoded code streams into a media file 20F for local playback according to a specific media file format (such as ISOBMFF), or encapsulate them into a segment sequence 20F for streaming. s In addition, the file encapsulator in the content production device 200A may also add relevant metadata to the media file 20F or the segment sequence 20F. s Furthermore, the content production device 200A may use a certain transmission mechanism (such as DASH, SMT) to transmit the segment sequence 20F s The content consumption device 200B may be transmitted to the content consumption device 200B, or the media file 20F may be transmitted to the content consumption device 200B. In some embodiments, the content consumption device 200B may be a player.

[0194] 2. Data processing on the content consumption device side:

[0195] (3) The process of decapsulating and decoding point cloud media files.

[0196] Content consumption devices can adaptively and dynamically obtain point cloud media media file resources and corresponding media presentation description information from content production devices based on recommendations from content production devices or in response to the needs of business objects on the content consumption device side. For example, a content consumption device can determine the viewing direction and position of a business object based on the head / eye position information of the business object. Based on this determination, the content consumption device can then dynamically request the corresponding media file resources from the content production device. Media file resources and media presentation description information are transmitted from the content production device to the content consumption device via transport mechanisms (such as DASH and SMT). The file decapsulation process on the content consumption device is the inverse of the file encapsulation process on the content production device side. The content consumption device decapsulates the media file resources according to the point cloud media file format requirements (e.g., ISOBMFF) to obtain audio and video streams. The decoding process on the content consumption device side is the inverse of the encoding process on the content production device side. The content consumption device performs audio decoding on the audio stream to restore the audio content, and performs video decoding on the video stream to restore the video content.

[0197] For example, Figure 3As shown, the media file 20F output by the file encapsulator in the content production device 200A is the same as the media file 20F' input to the file decapsulator in the content consumption device 200B. The file decapsulator processes the media file 20F' or the received segment sequence 20F'. s The file is decapsulated and the encoded point cloud stream 20E' is extracted. The corresponding metadata is parsed, and then point cloud media decoding can be performed on the point cloud stream 20E' to obtain a decoded video signal 20D'. Point cloud data (i.e., restored video content) can be generated from the video signal 20D'. The media files 20F and 20F' may include track format definitions, which may include constraints on the elementary streams contained in the samples in the track.

[0198] (4) Rendering process of point cloud media.

[0199] The content consumption device renders the audio content obtained by audio decoding and the video content obtained by video decoding according to the rendering-related metadata (such as the auxiliary information of this application) in the media presentation description information corresponding to the media file resource. Once the rendering is completed, the playback output of the content is realized.

[0200] The immersive media system supports data boxes, which are data blocks or objects that include metadata. That is, the data box contains metadata for the corresponding media content. In practical applications, content production devices can use data boxes to guide content consumption devices to consume media files of point cloud media. Point cloud media can include multiple data boxes, such as ISOBMFF data boxes (ISO Base Media File Format Box, referred to as ISOBMFF Box for short), which contain metadata for describing the corresponding information when the file is encapsulated. In the embodiment of the present application, the ISOBMFF data box includes metadata for indicating auxiliary information of the point cloud media.

[0201] From the above, it can be seen that the content consumption device can dynamically obtain the media file resources corresponding to the point cloud media from the content production device. Since the media file resources are obtained by the content production device after encoding and encapsulating the captured audio and video content, after the content consumption device receives the media file resources returned by the content production device, it needs to first decapsulate the media file resources to obtain the corresponding audio and video code stream, and then decode the audio and video code stream before finally presenting the decoded audio and video content to the business object. The point cloud media here can include but is not limited to VPCC (Video-based Point Cloud Compression, point cloud compression based on traditional video coding) point cloud media and GPCC (Geometry-based Point Cloud Compression, point cloud compression based on geometric models) point cloud media.

[0202] It is understandable that auxiliary information is of great significance to the processing and analysis of images, and can greatly improve the efficiency of image processing and analysis. The embodiment of the present application does not limit the type of auxiliary information, and can be any information used to assist in rendering point cloud media, such as saliency information (including saliency level parameters and an algorithm for obtaining saliency level parameters), credibility information (including credibility level parameters and an algorithm for obtaining credibility level parameters), priority information (including priority level parameters, or priority level parameters and an algorithm for obtaining priority level parameters) and quality level information (including quality level parameters and an algorithm for obtaining the quality level parameters).

[0203] It is understandable that different types of auxiliary information are meaningful, so in actual application, the content production device may obtain auxiliary information of multiple different auxiliary information types (referred to as types). The embodiment of the present application proposes a general indication information without a value, so multiple different types of auxiliary information can be encapsulated with the point cloud code stream to obtain a media file corresponding to the point cloud media, and the media file is saved for subsequent playback, or the media file is transmitted to the content consumption device through streaming. The content consumption device decapsulates the received media file to obtain the point cloud code stream and the multiple different types of auxiliary information indicated by the assigned general indication information. Further, the point cloud code stream is decoded to obtain the point cloud media. According to the multiple different types of auxiliary information, the content consumption device renders the point cloud media and presents the media content.

[0204] Based on the above, according to different types of auxiliary information, the embodiment of the present application proposes a general indication information that can be used to indicate different types of auxiliary information. Therefore, different types of auxiliary information are encapsulated into one media file, so it is possible to avoid one auxiliary information corresponding to one media file, thereby reducing the redundancy of the point cloud code stream. By transmitting a media file containing different types of auxiliary information, it is also possible to reduce the waste of network transmission resources. In addition, through the unassigned general indication information, the embodiment of the present application can improve the efficiency of encapsulating different types of auxiliary information and the efficiency of decapsulating different auxiliary information, thereby improving the rendering effect of point cloud media according to consumer needs and presenting media content that meets consumer needs.

[0205] It should be understood that the method provided in the embodiment of the present application can be applied to the server side (i.e., the content production device side), the player side (i.e., the content consumption device side), and the intermediate nodes (e.g., SMT (Smart Media Transport) receiving entity, SMT sending entity) and other links of the immersive media system. Among them, the content production device obtains at least one auxiliary information of the point cloud media, encodes the point cloud media to obtain a point cloud code stream, and encapsulates the point cloud code stream, at least one auxiliary information, and unassigned general indication information to obtain the specific process of the media file, and the specific process of the content consumption device rendering the point cloud media based on at least one auxiliary information in the media file can be referred to below. Figure 4-Figure 9 Description of the corresponding embodiment.

[0206] Further, see Figure 4 , Figure 4 This is a flow chart of a media data processing method provided by an embodiment of the present application. The method can be performed by a content production device in an immersive media system (for example, the above Figure 3 The content production device 200A in the corresponding embodiment is executed, for example, the content production device can be a server, and the embodiment of the present application is described by taking the server execution as an example. The method can at least include the following steps S101-S102.

[0207] Step S101: Acquire at least one type of auxiliary information of point cloud media and unassigned general indication information.

[0208] Specifically, at least one auxiliary information includes auxiliary information B c ; c is a positive integer, and c is less than or equal to the total number of at least one auxiliary information; auxiliary information B c Includes target range D for indicating point cloud media c Assistance level parameter; target range D c Including spatial range E c or time range Fc .

[0209] For the acquisition and production process of point cloud media, please refer to the above Figure 3 The auxiliary information in the embodiment of the present application may include two types of information (referring to the composition of a type of auxiliary information). One type of information is auxiliary algorithm information, i.e., information for obtaining auxiliary level parameters. The embodiment of the present application does not limit the algorithm for obtaining the auxiliary level parameters, and it can be set according to the actual application scenario. The other type of information is the auxiliary level parameter used to indicate the target range.

[0210] The embodiment of the present application does not limit the total number of at least one type of auxiliary information. The total number can be 1, indicating one type of auxiliary information; the total number can be a positive integer greater than 1, for example, 2, indicating two types of auxiliary information. The two types of auxiliary information may include first auxiliary information and second auxiliary information, wherein the auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information, for example, the first auxiliary information is credibility information, and the second auxiliary information is saliency information. The saliency information may include a saliency level parameter for indicating the target range of the point cloud media (for distinction, the target range is referred to as the target range corresponding to the saliency), and an acquisition algorithm corresponding to the saliency level parameter. The credibility information may include a credibility level parameter for indicating the target range of the point cloud media (for distinction, the target range is referred to as the target range corresponding to the credibility), and an acquisition algorithm corresponding to the credibility level parameter. It will be understood that the total number of at least one type of auxiliary information can be determined based on actual application consumption needs.

[0211] In scenarios involving at least two types of auxiliary information, the server acquires multiple different types of auxiliary information, where the server acquires these different types of auxiliary information independently. For example, in the examples of credibility information and saliency information, in the embodiments of the present application, the server acquires credibility information independently of the saliency information. In this case, the target range corresponding to saliency is independent of the target range corresponding to credibility. That is, the target range corresponding to saliency can be the same as or different from the target range corresponding to credibility. Furthermore, the algorithm for acquiring saliency information is independent of the algorithm for acquiring credibility information.

[0212] It is understandable that different point cloud media require different scopes for their corresponding auxiliary level parameters. For ease of description and understanding, the auxiliary level parameters are used here as an example of saliency level parameters. For some point cloud media, the saliency level parameters corresponding to them vary only over time. That is, there are no different saliency level parameters within the point cloud frames of the point cloud media. In this case, the saliency level parameters do not need to be associated with spatial regions. For some point cloud media, the saliency level parameters corresponding to them vary only over space. That is, there are different saliency level parameters within the point cloud frames of the point cloud media. For example, the saliency level parameter corresponding to a first spatial region in a first point cloud frame is different from the saliency level parameter corresponding to a second spatial region in the first point cloud frame, but the change in the saliency level parameter within the remaining point cloud frames is the same as the change in the saliency level parameter within the first point cloud frame. In this case, the saliency level parameter does not need to be associated with the point cloud frames. For some point cloud media, the saliency level parameters corresponding to them vary over both time and space. For example, the saliency level parameter of a first point cloud frame is different from the saliency level parameter of a second point cloud frame, and different saliency level parameters exist within the first point cloud frame. Other types of assistance level parameters may be understood with reference to the above description of the saliency level parameter of saliency information.

[0213] Therefore, the embodiment of the present application proposes a general auxiliary information indication method for immersive media, especially point cloud media, which can cover the indication of different types of auxiliary information such as credibility, saliency, quality level, priority, etc. The method is implemented at the file encapsulation level and signaling transmission level.

[0214] 1. Define common indication information for indicating different scopes (including spatial scope and temporal scope);

[0215] 2. Define different types of auxiliary information acquisition algorithms;

[0216] 3. Associate auxiliary information with spatial information at the signaling transmission level;

[0217] This allows for more flexible determination of auxiliary information for different spatial regions and frames within point cloud media, thus satisfying a wider range of point cloud application scenarios. This allows the server to optimize encoding and the client to optimize transmission based on this auxiliary information. Furthermore, universal auxiliary indication information can be personalized and expanded to support a wider range of point cloud applications.

[0218] The embodiment of the present application adds several descriptive fields at the system level, including field extensions at the file encapsulation level and the transmission signaling level, to support the embodiment of the present application. The following example uses the form of an extended ISOBMFF data box to define the auxiliary information indication method for point cloud media. For field extensions at the transmission signaling level, please refer to the following Figure 7Description of DASH signaling and SMT signaling.

[0219] In the embodiment of the present application, auxiliary information of the point cloud media can be provided through an auxiliary information data structure. Please refer to Table 1, which is used to indicate the syntax of an auxiliary information data structure provided in the embodiment of the present application:

[0220] Table 1

[0221]

[0222]

[0223] The semantics of the syntax shown in Table 1 are as follows: auxiliary_info_level is the auxiliary level field, which takes a 16-bit unsigned integer value and indicates the auxiliary level parameter in the auxiliary information. For credibility information, saliency information, quality level information, and priority information, a larger value indicates a higher level of the corresponding information. associate_space_flag is the target range indicator field, which takes a 1-bit unsigned integer value. When this field takes a first indicator value (such as 1 in Table 1), it indicates that the auxiliary level parameter is the auxiliary level of a spatial region in a point cloud frame (hereinafter referred to as a specific spatial region). When this field takes a second indicator value (such as 0), it indicates that the auxiliary level parameter is not the auxiliary level of a specific spatial region. It should be noted that when associate_space_flag takes a second indicator value, it means that the corresponding auxiliary level field takes a default auxiliary level parameter and the indicated region is a non-specific spatial region. A non-specific spatial region is a spatial region other than the spatial region indicated by associate_space_flag taking the first indicator value. In addition, it should be noted that the embodiment of the present application does not limit the specific values of the first indicator value and the second indicator value, and the two indicator values only need to be different.

[0224] The region_id_ref_flag field in Table 1 is a spatial range indicator field, which takes a 1-bit unsigned integer value. When this field takes the eleventh indicator value (e.g., 1 in Table 1), it indicates that the spatial region associated with the assistance level parameter is indexed by the identifier of the spatial region. When this field takes the twelfth indicator value (e.g., 0), it indicates that the spatial region associated with the assistance level is indicated by a spatial information-related data structure. It should be noted that the embodiments of the present application do not limit the specific values of the eleventh and twelfth indicator values; the two indicator values can be different.

[0225] The spatial_region_id field in Table 1 is a 16-bit unsigned integer that indicates the spatial region identifier corresponding to the spatial region associated with the assistance level parameter. The anchor_point field indicates the anchor coordinates of the spatial region, and the bounding_info field indicates the length, width, and height information of the spatial region.

[0226] The slice_info_flag field in Table 1 is a point cloud slice information field, which takes a 1-bit unsigned integer value. When the first information value (such as 1 in Table 1) is used for this field, it indicates that the spatial region associated with the auxiliary level parameter is associated with one or more point cloud slices (referred to as associated point cloud slices). When the second information value (such as 0) is used for this field, it indicates that the spatial region associated with the auxiliary level has no associated point cloud slices. It should be noted that the embodiments of the present application do not limit the specific values of the first information value and the second information value; the two information values can be different.

[0227] In Table 1, num_slices is the point cloud slice number field, which takes a 16-bit unsigned integer value and indicates the number of point cloud slices associated with the spatial region, that is, the total number of associated point cloud slices. slice_id is the point cloud slice identifier field, which takes a 16-bit unsigned integer value and indicates the point cloud slice identifier of the associated point cloud slice.

[0228] The tile_info_flag field in Table 1 is a 1-bit unsigned integer representing the spatial tile information field. A third information value (e.g., 1 in Table 1) indicates that the spatial region associated with the assistive level is associated with one or more point cloud spatial tiles (referred to as associated spatial tiles). A fourth information value (e.g., 0) indicates that the spatial region associated with the assistive level has no associated point cloud spatial tiles. This embodiment of the present application does not limit the specific values of the third and fourth information values; they can be different.

[0229] In Table 1, num_tiles is the spatial tile information field, which takes a 16-bit unsigned integer value and indicates the number of spatial tiles associated with the spatial region, that is, the total number of associated spatial tiles. tile_id is the spatial tile identifier field, which takes a 16-bit unsigned integer value and indicates the spatial tile identifier of the associated spatial tile. It will be understood that the spatial region indicated by the assistance level parameter can be associated with either tiles or slices.

[0230] In particular, when the spatial region indicated by the assistance level parameter is indexed by a spatial region identifier, dynamic changes in the spatial range of the spatial region corresponding to the spatial region identifier do not affect the static indication of the assistance level parameter.

[0231] Unassigned general indication information may include unassigned data box information and unassigned metadata track information. Among them, unassigned data box information refers to an unassigned auxiliary information data box. In this article, the assigned auxiliary information data box is referred to as the auxiliary information data box. It can be understood that the fields included in the unassigned auxiliary information data box are the same as the fields included in the auxiliary information data box. The difference between the two is that the fields of the former are not assigned, and the fields of the latter are assigned according to the auxiliary information. Unassigned metadata track information refers to an unassigned auxiliary information metadata track. In this article, the assigned auxiliary information metadata track is referred to as the auxiliary information metadata track. It can be understood that the fields included in the unassigned auxiliary information metadata track are the same as the fields included in the auxiliary information metadata track. The difference between the two is that the fields of the former are not assigned, and the fields of the latter are assigned according to the auxiliary information.

[0232] It is known that auxiliary information may vary with space, may vary with time, and may vary with time and space. In order to satisfy the auxiliary information indicating different target ranges, the embodiment of the present application provides the above-mentioned unassigned general indication information, which is as follows. When the auxiliary level parameter of the point cloud media is only related to space and has nothing to do with time, the server only assigns a value to the unassigned auxiliary information data box to obtain an auxiliary information data box (for example, a first auxiliary information data box), which is included in the sample entrance of the point cloud track corresponding to the point cloud media. Please also refer to Table 2, which is used to indicate the syntax of an auxiliary information data box structure provided by the embodiment of the present application:

[0233] Table 2

[0234]

[0235] An auxiliary information data box can be included at the sample entry of a point cloud track, indicating time-invariant auxiliary information for that point cloud track. The number of auxiliary information boxes can be zero or more. If the point cloud media has multiple point cloud tracks, the auxiliary information data box can be placed at the sample entry of any point cloud track. When the auxiliary information data box is placed at the sample entry of the point cloud track corresponding to the point cloud media, the value of the num_static_auxiliary_struct field must be greater than 0.

[0236] The semantics of the syntax shown in Table 2 above are as follows: auxiliary_algorithm_flag is an auxiliary algorithm indication field, which takes a 1-bit unsigned integer value. When this field takes a third indication value (such as 1 in Table 2), it indicates that the auxiliary information exists and an auxiliary algorithm (i.e., an algorithm for obtaining the auxiliary information) exists. When this field takes a fourth indication value (such as 0), it indicates that the auxiliary information does not have an auxiliary algorithm. This embodiment of the application does not limit the specific values of the third indication value and the fourth indication value; the two indication values can be different.

[0237] The auxiliary_algorithm_type field in Table 2 is an 8-bit unsigned integer indicating the type of algorithm used to obtain auxiliary information. A value of the second type (e.g., 0) indicates that the auxiliary information is obtained by the auxiliary information detection algorithm. A value of the third type (e.g., 1) indicates that the auxiliary information is obtained by subjective evaluation of the object, i.e., by statistical data. Other values can be extended by the application.

[0238] The auxiliary_info_type in Table 2 is the auxiliary information type field, which takes an 8-bit unsigned integer as the value, and is used to indicate the auxiliary information type described in the auxiliary information. The meaning of the value of this field is shown in Table 3, which is a schematic table of the field value and meaning of an auxiliary information type field provided in an embodiment of the application.

[0239] Table 3

[0240] Value meaning 0 Indicates that the auxiliary information is credibility information 1 Indicates that the auxiliary information is saliency information 2 Indicates that the auxiliary information is quality level information 3 Indicates that the auxiliary information is priority information 4~127 Custom extensions 128~255 reserve

[0241] In Table 2, num_static_auxiliary_struct is the static data structure number field, which is a 16-bit unsigned integer and indicates the number of auxiliary information data structures statically associated with a specific spatial region. static_auxiliary_info is an auxiliary information data structure with static attributes, used to indicate auxiliary information statically associated with a specific spatial region. In the auxiliary information data structure with static attributes, the associate_space_flag value must be the first indication value.

[0242] The default_auxiliary_info in Table 2 is an auxiliary information data structure with default attributes, which is used to indicate default auxiliary information. In the auxiliary information data structure with default attributes, the value of associate_space_flag must be the second indication value.

[0243] It should be noted that when the point cloud track corresponds to multiple types of auxiliary information (such as saliency and credibility), each type of static auxiliary information can correspond to one auxiliary information data box.

[0244] The above describes the auxiliary information data box, which is used to describe information with static properties in the auxiliary information, such as the type of auxiliary information, the algorithm for obtaining the auxiliary information, and the auxiliary level parameters in the auxiliary information that do not change over time. When the auxiliary level parameters of the point cloud media are related to time (including only related to time, and related to both time and space), the server assists the unassigned auxiliary information data box and the unassigned auxiliary information metadata track to obtain the auxiliary information metadata track, and the sample entry of the auxiliary information metadata track contains an auxiliary information data box (for example, a second auxiliary information data box). Please also refer to Table 4, which is used to indicate the syntax of an auxiliary information metadata track structure provided in an embodiment of the present application:

[0245] Table 4

[0246]

[0247] It can be understood that the auxiliary information metadata track (also called dynamic auxiliary information metadata track, abbreviated as daui) includes the sample sequence number corresponding to the point cloud frame (also called sample) in the point cloud media. The auxiliary information metadata track is used to indicate the auxiliary information that changes with time in the point cloud track. The AuxiliaryInfoBox data box in the sample entry of the auxiliary information metadata track indicates the static auxiliary information that does not change with time in the point cloud track. In addition, when the point cloud track corresponds to multiple types of auxiliary information that changes with time (such as saliency and credibility at the same time), each type of auxiliary information can correspond to an auxiliary information metadata track.

[0248] The semantics of the syntax shown in Table 4 above are as follows: frame_range_auxiliary_info is a valid range indication field, which takes a 1-bit unsigned integer value. When this field takes the seventh indicator value (for example, 0), it indicates that the auxiliary level parameter indicated in the auxiliary information metadata track is valid within the entire auxiliary information metadata track range, that is, the auxiliary level parameters between samples can be compared with each other. When this field takes the eighth indicator value (for example, 1), it indicates that the auxiliary level parameter indicated in the auxiliary information metadata track sample is only valid within the range of the corresponding point cloud frame, that is, the auxiliary level parameters between samples cannot be compared with each other.

[0249] The default_auxiliary_flag field in Table 4 is an auxiliary level indicator field, which takes a 1-bit unsigned integer value. When the value of this field is the ninth indicator value (for example, 1), it indicates that the corresponding point cloud frame has the default auxiliary level parameter. When the value of this field is the tenth indicator value (such as 0 in Table 2), it indicates that the auxiliary level of the corresponding point cloud frame is determined by the AuxiliaryInfoStruct in the sample. This embodiment of the application does not limit the values of the ninth and tenth indicator values, and the two indicator values can be different.

[0250] The num_saliceny_struct in Table 4 is the data structure number field, which takes a value of a 16-bit unsigned integer. This field is used to indicate the total number of auxiliary information data structures.

[0251] In summary, the embodiments of the present application provide a universal indication information, which can indicate different types of auxiliary information and can indicate auxiliary information of different target ranges.

[0252] Step S102: Encode the point cloud media to obtain a point cloud code stream.

[0253] Specifically, at least one auxiliary information includes first auxiliary information and second auxiliary information; the first auxiliary information includes an auxiliary level parameter σ for indicating a target range of the point cloud media; the second auxiliary information includes an auxiliary level parameter ρ for indicating a target range of the point cloud media; the auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information; according to the auxiliary level parameter σ and the auxiliary level parameter ρ, the target range is optimized and encoded to obtain a point cloud code stream.

[0254] The total number of target ranges is at least two, and the at least two target ranges include a third target range and a fourth target range; the auxiliary level parameter σ includes a third auxiliary level parameter corresponding to the third target range, and a fourth auxiliary level parameter corresponding to the fourth target range; the auxiliary level parameter ρ includes a fifth auxiliary level parameter corresponding to the third target range, and a sixth auxiliary level parameter corresponding to the fourth target range; according to the auxiliary level parameter σ and the auxiliary level parameter ρ, the target range is optimized and encoded to obtain a point cloud code stream. The specific process may include: performing weighted summation on the third auxiliary level parameter and the fifth auxiliary level parameter to obtain a first total auxiliary level parameter corresponding to the third target range; The fourth auxiliary level parameter and the sixth auxiliary level parameter are weightedly summed to obtain a second total auxiliary level parameter corresponding to the fourth target range; the third coding level of the third target range is determined based on the first total auxiliary level parameter, and the fourth coding level of the fourth target range is determined based on the second total auxiliary level parameter; when the first total coding level parameter is greater than the second total auxiliary level parameter, the third coding level is superior to the fourth coding level; the third target range is optimized and encoded using the third coding level to obtain a third sub-point cloud codestream, and the fourth target range is optimized and encoded using the fourth coding level to obtain a fourth sub-point cloud codestream; and a point cloud codestream is generated based on the third point cloud sub-codestream and the fourth point cloud sub-codestream.

[0255] In the scenario of at least two types of auxiliary information, the process of the server processing multiple different types of auxiliary information can be independent of each other, including but not limited to the encoding process, the encapsulation process, and the transmission process. For example, the server obtains the credibility information and saliency information of the above example. In the embodiment of the present application, when encoding the point cloud media, the server may not optimize the encoding according to the auxiliary information, that is, directly encode the point cloud media, or optimize the encoding according to the credibility information, or optimize the encoding according to the saliency information, or optimize the encoding according to the saliency information and the credibility information. When encapsulating the point cloud media, the process of using general indication information to indicate the credibility information is independent of the process of using general indication information to indicate the saliency information. The same is true for the remaining processing processes.

[0256] This step uses multiple different types of auxiliary information to optimize the encoding of point cloud media. This optimization requires that the different types of auxiliary information indicate the same target range. For ease of understanding and description, this example uses two different types of auxiliary information (e.g., first auxiliary information and second auxiliary information). For other quantities of auxiliary information, refer to the description of the two different types of auxiliary information.

[0257] The server first determines the auxiliary level parameter σ (for example, a credibility level parameter) in the first auxiliary information for indicating the target range, and determines the auxiliary level parameter ρ (for example, a significance level parameter) in the second auxiliary information for indicating the target range. The total number of target ranges is at least two, and the at least two target ranges include a third target range and a fourth target range.

[0258] The server obtains, from the assistance level parameter σ, a third assistance level parameter corresponding to the third target range and a fourth assistance level parameter corresponding to the fourth target range. From the assistance level parameter ρ, the server obtains, from the assistance level parameter ρ, a fifth assistance level parameter corresponding to the third target range and a sixth assistance level parameter corresponding to the fourth target range. Furthermore, the server performs a weighted sum of the third and fifth assistance level parameters to obtain a first total assistance level parameter corresponding to the third target range. The weights corresponding to the third and fifth assistance level parameters can be set according to actual application scenarios and are not limited in this application. The server performs a weighted sum of the fourth and sixth assistance level parameters to obtain a second total assistance level parameter corresponding to the fourth target range. The weight corresponding to the fourth assistance level parameter is equal to the weight corresponding to the third assistance level parameter, and the weight corresponding to the fifth assistance level parameter is equal to the weight corresponding to the sixth assistance level parameter. Furthermore, the server determines a third coding level for the third target range based on the first total assistance level parameter and a fourth coding level for the fourth target range based on the second total assistance level parameter. When the first total assistance level parameter is greater than the second total assistance level parameter, the third coding level is superior to the fourth coding level. The server optimizes encoding of the third target range using the third encoding level to obtain a third sub-point cloud stream. Furthermore, the server optimizes encoding of the fourth target range using the fourth encoding level to obtain a fourth sub-point cloud stream. Furthermore, the server generates a point cloud stream based on the third and fourth sub-point cloud streams.

[0259] The above-mentioned optimized encoding process can be performed by the content production device, or after the content production device generates the media file and transmits it to the intermediate node, the intermediate node first decapsulates and decodes the media file to obtain the point cloud media, and then optimizes the encoding of the point cloud media based on the saliency information of the point cloud media.

[0260] Optionally, at an intermediate stage in the point cloud transmission system, if the use of the auxiliary information has been completed, the corresponding auxiliary information metadata can be removed from the media file, for example, by removing the auxiliary information data box or auxiliary information metadata track used to indicate the auxiliary information. For example, if the intermediate node re-encodes and re-packages the point cloud file based on the first auxiliary information, the metadata associated with the first auxiliary information can be removed during the re-packaging process, and the new media file only contains the second auxiliary information.

[0261] The server optimizes the encoding of point cloud media based on an auxiliary level parameter. See below. Figure 7 The description in will not be described here.

[0262] Step S103, encapsulate the point cloud code stream, at least one auxiliary information and unassigned general indication information to obtain a media file; the media file includes the assigned general indication information, which is obtained by assigning a value to the unassigned general indication information based on at least one auxiliary information; the assigned general indication information is used to indicate at least one auxiliary information.

[0263] Specifically, the assigned general indication information includes at least one of the assigned data box information and the assigned metadata track information.

[0264] Wherein, when the media file includes auxiliary information B c When the auxiliary information metadata track is set, the target range D c Including time range F c ; The assigned metadata track information includes the auxiliary information metadata track.

[0265] The sample entry of the auxiliary information metadata track includes an effective range indication field and a second auxiliary information data box; the effective range indication field is used to indicate the effective range of the auxiliary level parameters corresponding to the M sample numbers; the second auxiliary information data box is used to indicate the auxiliary information B c The information with static properties in the M sample numbers belongs to the auxiliary level parameters corresponding to the auxiliary information B c The auxiliary level parameter in .

[0266] Among them, the M sample numbers include the sample number τ, τ is a positive integer and τ is less than or equal to M; when the field value of the effective range indication field is the seventh indication value, it means that the effective range of the auxiliary level parameter corresponding to the sample number τ is the auxiliary information metadata track; when the field value of the effective range indication field is the eighth indication value, it means that the effective range of the auxiliary level parameter corresponding to the sample number τ is the point cloud frame corresponding to the sample number τ; the eighth indication value is different from the seventh indication value.

[0267] Among them, the auxiliary information metadata track includes sample number O n ; n is a positive integer and n is less than or equal to M; time range F c Including sample serial number O n The corresponding point cloud frame; the auxiliary information metadata track includes the sample number O n When the auxiliary level indication field value is the ninth indication value, it indicates that the sample number is On The associated assistance level is determined by the default assistance level parameter; the default assistance level parameter belongs to the assistance information B c When the auxiliary level indicator field value is the tenth indicator value, it indicates that the sample number is 0. n The associated assistance level is determined by the assistance information data structure; the tenth indicator value is different from the ninth indicator value.

[0268] When the field value of the auxiliary level indication field is the ninth indication value, it indicates that the default auxiliary level parameter is used to indicate the sample sequence number 0. n The assistance level of the corresponding point cloud frame.

[0269] When the field value of the auxiliary level indication field is the tenth indication value, the auxiliary information metadata track also includes the value for the sample number 0. n The data structure quantity field is used to indicate the total number of auxiliary information data structures.

[0270] The value of the data structure number field is P, indicating P auxiliary information data structures; the P auxiliary information data structures include auxiliary information data structure Q r , r, P are all positive integers and r is less than or equal to P; auxiliary information data structure Q r Include field value for auxiliary level parameter S r The auxiliary level field and the target range indication field; the auxiliary level parameter S r Belongs to auxiliary information B c When the target range indicator field value is the first indicator value, it indicates that the auxiliary level parameter S r For indication, sample number O n The auxiliary level of a spatial area in the corresponding point cloud frame; when the field value of the target range indication field is the second indication value, it indicates that the auxiliary level parameter S r For indication, sample number O n The auxiliary level of the corresponding point cloud frame; the second indication value is different from the first indication value.

[0271] In the point cloud track corresponding to the point cloud media, a point cloud frame can be called a sample. Since the server processes different auxiliary information independently, for ease of understanding and description, the following description will use one type of auxiliary information.

[0272] As can be seen from step S101, different point cloud media have different ranges of variation of auxiliary information. The point cloud media may include a first point cloud media, which includes M point cloud frames, where M is a positive integer. The auxiliary information is exemplified by the credibility information. Figure 5 , Figure 5 This is a schematic diagram of a credibility level parameter provided by an embodiment of the present application that is only related to time and has nothing to do with spatial area. Figure 5 As shown, sample 1 refers to the sample (i.e., point cloud frame) with sequence number 1 in the first point cloud media, sample 2 refers to the sample with sequence number 2 in the first point cloud media, sample 3 refers to the sample with sequence number 3 in the first point cloud media, and so on. Sample M refers to the sample with sequence number M in the first point cloud media. Figure 5 The credibility information is used as an example of auxiliary information, so Figure 5 The auxiliary level parameter in refers to the confidence level parameter.

[0273] Based on the content of the first point cloud media, the server can define the credibility information of each region of the first point cloud media. The overall credibility level of Sample 1 is a default credibility level parameter. It is understood that the immersive media system can pre-set the default credibility level parameter. This embodiment of the application does not limit the value of the default credibility level parameter, and it can be set according to the actual application scenario. The overall credibility level parameter of Sample 2 is 120, the overall credibility level parameter of Sample 3 is 110, ..., and the overall credibility level parameter of Sample M is the default credibility level parameter.

[0274] The larger the credibility level parameter is, the higher the credibility is. Therefore, in the first point cloud media, sample 2 is the most credible sample (i.e., point cloud frame), sample 3 is the second most credible sample, and the remaining samples only have the default credibility (assuming the default credibility level parameter is 100). Obviously, in the first point cloud media, the credibility level parameter is only related to time and has nothing to do with the spatial region. Figure 5 If the credibility level parameter is a point cloud frame level, the server may generate an auxiliary information metadata track as shown in Table 5 to indicate the credibility information of the first point cloud media. Table 5 is a table of an auxiliary information metadata track structure for indicating credibility information provided by an embodiment of the present application.

[0275] Table 5

[0276]

[0277]

[0278] For the meaning of each field in Table 5, please refer to the description in Table 1, Table 2, and Table 4 above, which will not be repeated here.

[0279] Furthermore, the point cloud media may include a second point cloud media, and the second point cloud media includes one or more point cloud frames (e.g., M point cloud frames). Here, the saliency information is used as an example of auxiliary information. Please refer to Figure 6 , Figure 6 : This is a schematic diagram showing that a significance level parameter is related to a spatial region and the significance level parameter of the spatial region changes over time, as provided in an embodiment of the present application. Figure 6 For the meaning of sample 1-sample M, see Figure 5 The meaning of is not explained here. Figure 6 The auxiliary information is an example of saliency information, so Figure 6 The auxiliary level parameter in refers to the significance level parameter.

[0280] According to the content of the second point cloud media, the server can define the saliency information of each area of the second point cloud media, where the saliency levels corresponding to sample 1, sample 2, sample 3, sample 4, ..., sample M are not exactly the same, that is, the saliency level parameters of the second point cloud media change over time, and there are some point cloud frames whose internal spatial areas also correspond to different saliency level parameters, that is, the saliency level parameters also change in the spatial area.

[0281] like Figure 6 As shown, the overall significance level parameter of sample 1 is the default significance level parameter, where the meaning of the default significance level parameter is as follows: Figure 5 According to the description in , the spatial area corresponding to sample 1 can be divided into two spatial areas, among which the saliency level parameter corresponding to the first spatial area in sample 1 (referred to as spatial area 1) is 2, which consists of two point cloud slices, and the point cloud slice identifiers corresponding to the two point cloud slices include point cloud slice 0 and point cloud slice 1 respectively; the saliency level parameter corresponding to the second spatial area in sample 1 (referred to as spatial area 2) is 1, which consists of two point cloud slices, and the point cloud slice identifiers corresponding to the two point cloud slices include point cloud slice 2 and point cloud slice 3 respectively; in sample 1, the saliency of the first spatial area is higher than that of the second spatial area.

[0282] like Figure 6 As shown in the figure, the overall saliency level parameter of sample 2 is 2, and the internal space of sample 2 is not divided. The overall saliency level parameter of sample 3 is the default saliency level parameter. The spatial area corresponding to sample 3 can be divided into two spatial areas, wherein the saliency level parameter corresponding to the first spatial area in sample 3 (referred to as spatial area 1) is 0, which consists of two point cloud slices, and the point cloud slice identifiers corresponding to the two point cloud slices include point cloud slice 0 and point cloud slice 1 respectively; the saliency level parameter corresponding to the second spatial area in sample 3 (referred to as spatial area 2) is 1, which consists of two point cloud slices, and the point cloud slice identifiers corresponding to the two point cloud slices include point cloud slice 2 and point cloud slice 3 respectively; in sample 3, the saliency of the first spatial area is lower than that of the second spatial area.

[0283] like Figure 6As shown, the saliency level parameters corresponding to samples 4, ..., and sample M are all default saliency level parameters. Therefore, in the second point cloud media, sample 2 is the sample with the highest saliency (i.e., point cloud frame), sample 1 has the spatial region with the highest saliency and the spatial region with the second highest saliency, sample 3 has the spatial region with the second highest saliency, and the remaining samples only have the default saliency level parameters (assuming that the default saliency level parameter is 0). Obviously, in the second point cloud media, the saliency level parameters are related to both time and spatial regions, so the server can generate an auxiliary information metadata track as shown in Table 6 to indicate the saliency information of the second point cloud media. Table 6 is a table of the structure of an auxiliary information metadata track for indicating saliency information provided in an embodiment of the present application.

[0284] Table 6

[0285]

[0286]

[0287] For the meaning of each field in Table 6, please refer to the description in Table 1, Table 2, and Table 4 above, which will not be repeated here.

[0288] The point cloud media may include a third point cloud media. The auxiliary level parameter corresponding to the third point cloud media may be only related to the spatial area and has nothing to do with time. The third point cloud media is not described in detail in this embodiment of the application. Please refer to the following Figure 7 Description in the corresponding embodiment.

[0289] The server encapsulates the point cloud code stream into a media file and indicates the auxiliary information in the form of metadata (i.e., the auxiliary information metadata track) in the media file. Since the point cloud auxiliary information changes over time, the transmission signaling does not contain the auxiliary information description data, but the auxiliary information metadata track will exist in the transmission signaling as a media resource in the form of Representation. Figure 3 It can be seen that there are two ways for the server to transmit point cloud files to the client:

[0290] a) Client C1 downloads the complete point cloud file (i.e., media file) and plays it locally.

[0291] b) The client C2 establishes streaming transmission with the server and performs presentation and consumption while receiving the point cloud file fragment Fs.

[0292] From the above, it can be seen that the embodiment of the present application can determine the auxiliary information corresponding to the point cloud media for indicating the time range, or the auxiliary information corresponding to the point cloud media for indicating the spatial range, so the embodiment of the present application can encapsulate the auxiliary information of the point cloud media together with the point cloud code stream to obtain a media file, and the auxiliary information indication method provided by the embodiment of the present application can improve the accuracy of the auxiliary information of the point cloud media, and then when rendering the point cloud media, the rendering effect of the target range can be determined by accurate auxiliary information, so the presentation effect of the point cloud media can be optimized. In addition, the present application can use a general indication information to be assigned for indication for different types of auxiliary information, so the efficiency of encapsulation of auxiliary information can be improved. In addition, by encapsulating different types of auxiliary information into one media file, the embodiment of the present application can avoid one auxiliary information corresponding to one media file, so the redundancy of the point cloud code stream can be reduced; by transmitting a media file containing different types of auxiliary information, the waste of network transmission resources can also be reduced.

[0293] Further, see Figure 7 , Figure 7 This is a flow diagram of a media data processing method provided by an embodiment of the present application. Figure 2 The method can be implemented by a content production device in an immersive media system (e.g., Figure 3 The content production device 200A in the corresponding embodiment is executed, for example, the content production device can be a server, and the embodiment of the present application is described by taking the server execution as an example. The method can at least include the following steps S201-S204.

[0294] Step S201: obtaining at least one auxiliary information of the point cloud media and unassigned general indication information; the at least one auxiliary information includes auxiliary information B c ; c is a positive integer, and c is less than or equal to the total number of at least one type of auxiliary information.

[0295] Specifically, auxiliary information B c Includes target range D for indicating point cloud media c Assistance level parameter; target range D c Including spatial range E c or time range F c .

[0296] For the specific implementation process of step S201, please refer to the above Figure 4 Step S101 in the corresponding embodiment is not described in detail here.

[0297] Step S202: Encode the point cloud media to obtain a point cloud code stream.

[0298] Specifically, at least one type of auxiliary information includes target auxiliary information; the target auxiliary information includes an auxiliary level parameter for indicating a target range; and according to the auxiliary level parameter in the target auxiliary information, the target range of the point cloud media is optimized and encoded to obtain a point cloud code stream.

[0299] The total number of target ranges is at least two, and the at least two target ranges include a first target range and a second target range; the auxiliary level parameter in the target auxiliary information includes a first auxiliary level parameter corresponding to the first target range, and a second auxiliary level parameter corresponding to the second target range; according to the auxiliary level parameter in the target auxiliary information, the target range of the point cloud media is optimized and encoded, and a specific process of obtaining a point cloud code stream may include: determining a first encoding level for the first target range according to the first auxiliary level parameter, and determining a second encoding level for the second target range according to the second auxiliary level parameter, when the first auxiliary level parameter is greater than the second auxiliary level parameter, the first encoding level is superior to the second encoding level; optimizing the encoding of the first target range through the first encoding level to obtain a first sub-point cloud code stream, and optimizing the encoding of the second target range through the second encoding level to obtain a second sub-point cloud code stream; generating a point cloud code stream according to the first sub-point cloud code stream and the second sub-point cloud code stream.

[0300] Above Figure 4 Step S102 in the previous section describes optimizing the encoding of point cloud media using multiple types of auxiliary information. This step describes optimizing the encoding of point cloud media using a single type of auxiliary information (e.g., target auxiliary information). It should be understood that the total number of target ranges is at least two, and the at least two target ranges include a first target range and a second target range.

[0301] The server obtains a first assistance level parameter corresponding to the first target range and a second assistance level parameter corresponding to the second target range from the target assistance information. Furthermore, the server determines a first encoding level for the first target range based on the first assistance level parameter, and a second encoding level for the second target range based on the second assistance level parameter. When the first assistance level parameter is greater than the second assistance level parameter, the first encoding level is superior to the second encoding level. Furthermore, the server optimizes encoding of the first target range using the first encoding level to obtain a first sub-point cloud stream, and optimizes encoding of the second target range using the second encoding level to obtain a second sub-point cloud stream. Furthermore, based on the first and second sub-point cloud streams, the server can generate a point cloud stream.

[0302] The above-mentioned optimized encoding process can be performed by the content production device, or after the content production device generates the media file and transmits it to the intermediate node, the intermediate node first decapsulates and decodes the media file to obtain the point cloud media, and then optimizes the encoding of the point cloud media based on the saliency information of the point cloud media.

[0303] In step S203, the point cloud code stream, at least one auxiliary information, and unassigned general indication information are encapsulated to obtain a media file; the media file includes the assigned general indication information, which is obtained by assigning a value to the unassigned general indication information based on the at least one auxiliary information; the assigned general indication information is used to indicate the at least one auxiliary information.

[0304] Specifically, the assigned general indication information includes at least one of the assigned data box information and the assigned metadata track information.

[0305] Wherein, when the media file includes auxiliary information B c When the auxiliary information metadata track is set, the target range D c Including time range F c ; The assigned metadata track information includes the auxiliary information metadata track.

[0306] The media file includes auxiliary information B c When the auxiliary information data box is included in the sample entry of the point cloud track corresponding to the media file, the target range D c Including spatial range E c ; Among them, auxiliary information B c An auxiliary level parameter in is used to indicate a spatial range; a spatial range includes a spatial area respectively contained in G point cloud frames; the point cloud media includes G point cloud frames; G is a positive integer.

[0307] The first auxiliary information data box includes an auxiliary information type field whose field value is a first type value, a static data structure quantity field, and an auxiliary information data structure with default attributes; the first type value represents auxiliary information B c The auxiliary information type to which it belongs; the static data structure quantity field is used to indicate the total number of auxiliary information data structures with static attributes.

[0308] The value of the static data structure number field is H, indicating H auxiliary information data structures with static attributes; the H auxiliary information data structures with static attributes include auxiliary information data structures I with static attributes. j , where H and j are both positive integers, and j is less than or equal to H; auxiliary information data structure I with static attributes jInclude field value for auxiliary level parameter L j The auxiliary level field and the target range indication field whose field value is the first indication value; the auxiliary level parameter L j Different from the default assistance level parameters; the default assistance level parameters and the assistance level parameters L j All belong to auxiliary information B c The first indicator value represents the assistance level parameter L j Used to indicate the auxiliary level of a spatial extent.

[0309] The auxiliary information data structure with default attributes includes an auxiliary level field whose field value is a default auxiliary level parameter and a target range indication field whose field value is a second indication value; the default auxiliary level parameter belongs to the auxiliary information B c The auxiliary level parameter in the point cloud media; the second indication value indicates that the default auxiliary level parameter is used to indicate the auxiliary level of a spatial range; the spatial range indicated by the default auxiliary level parameter includes the spatial range in the point cloud media except the spatial range indicated by the auxiliary information data structure with static attributes.

[0310] The first auxiliary information data box includes an auxiliary algorithm indication field; when the field value of the auxiliary algorithm indication field is the third indication value, it indicates that the auxiliary information B c There is an auxiliary algorithm; when the field value of the auxiliary algorithm indication field is the fourth indication value, it indicates that the auxiliary information B c There is no auxiliary algorithm; the fourth indicator value is different from the third indicator value.

[0311] Among them, when the field value of the auxiliary algorithm indication field is the third indication value, the first auxiliary information data box also includes an auxiliary algorithm type field; when the field value of the auxiliary algorithm type field is the second type value, it indicates that the auxiliary information B c Determined by the auxiliary information detection algorithm; when the field value of the auxiliary algorithm type field is the third type value, it means that the auxiliary information B c Determined by data statistics; the third type value is different from the second type value.

[0312] The total number of assigned general indication information is the same as the total number of at least one type of auxiliary information. The total number of assigned data box information may indicate the total number of at least one type of auxiliary information that does not change over time. The total number of assigned metadata track information may indicate the total number of at least one type of auxiliary information that changes over time.

[0313] Above Figure 4Taking the first point cloud media example, it is described that the auxiliary information of the point cloud media only changes with time; taking the second point cloud media example, it is described that the auxiliary information of the point cloud media not only changes with time, but also that for some point cloud frames, the auxiliary level parameters of the internal spatial area also change. This step uses the third point cloud media example to describe that the auxiliary information of the point cloud media is only related to the spatial area and does not change with time. Here, the auxiliary information is exemplified by saliency information. Please refer to Figure 8 , Figure 8 This is a schematic diagram of an embodiment of the present application showing that a saliency level parameter is related to a spatial region and that the saliency level parameter related to the spatial region does not change over time. Assume that the third point cloud media includes G point cloud frames, G is a positive integer, and the internal structure of each of the G point cloud frames is as follows: Figure 8 As shown. Figure 8 The auxiliary information is an example of saliency information, so Figure 5 The auxiliary level parameter in refers to the significance level parameter.

[0314] According to the content of the third point cloud media, the server can define the saliency information of each area of the third point cloud media, wherein the saliency level parameter of the first spatial area (referred to as spatial area 1) of each point cloud frame is 2, which is composed of two point cloud slices, and the point cloud slice identifiers corresponding to the two point cloud slices include point cloud slice 0 and point cloud slice 1. The saliency level parameter of the second spatial area (referred to as spatial area 2) of each point cloud frame is 1, which is composed of two point cloud slices, and the point cloud slice identifiers corresponding to the two point cloud slices include point cloud slice 2 and point cloud slice 3. Therefore, in the third point cloud media, spatial area 1 has a higher saliency, and spatial area 2 has a lower saliency. At this time, the server can provide an auxiliary information data box to indicate the saliency information in the third point cloud media. The saliency information data box can be as shown in Table 7. Table 7 is an auxiliary information data box structure table for indicating saliency information provided in an embodiment of the present application.

[0315] Table 7

[0316]

[0317]

[0318] For the meaning of each field in Table 7, please refer to the description in Table 1 and Table 2 above and will not be repeated here.

[0319] Furthermore, the server encapsulates the point cloud code stream into a point cloud file, and indicates the above-mentioned saliency information in the file in the form of metadata (ie, an auxiliary information data box).

[0320] Step S204: When there is auxiliary information B cWhen the first auxiliary information data box is generated, the transmission signaling for the media file is transmitted to the client; the transmission instruction carries the auxiliary information description data K c ; Auxiliary information description data K c Used to instruct the client to determine the order of obtaining different media sub-files in the media file when obtaining the media file through streaming transmission; auxiliary information description data K c Based on auxiliary information B c Generated.

[0321] Specifically, the auxiliary information description data K c Including a default level indication field; when the field value of the default level indication field is the fifth indication value, it indicates that the auxiliary level parameter corresponding to the default information indication field is the default auxiliary level parameter; when the field value of the default level indication field is the sixth indication value, it indicates that the auxiliary level parameter corresponding to the default level indication field is not the default auxiliary level parameter; the sixth indication value is different from the fifth indication value; the auxiliary level parameter corresponding to the default level indication field belongs to the auxiliary information B c The auxiliary level parameter in .

[0322] It is understandable that different point cloud media have different ranges of associated auxiliary level parameters. Therefore, the embodiment of the present application proposes an auxiliary information indication method for immersive media, especially point cloud media. The method adds several descriptive fields at the system layer, including field extensions at the file encapsulation level and the transmission signaling level. The field extension at the file encapsulation level has been introduced above, and the following will be given as an example in the form of DASH signaling and SMT signaling. Among them, the auxiliary information description data includes the auxiliary information descriptor defined in the DASH signaling and the auxiliary information descriptor defined in the SMT signaling, as described below.

[0323] The embodiment of the present application is expanded in the DASH signaling, and an auxiliary information descriptor is proposed. The auxiliary information descriptor (AuxiliaryInfo descriptor) is a supplemental (SupplementalProperty) element, and its @schemeIdUri attribute is "urn:avs:ims:2022:apcc". This descriptor can exist at the adaptation set level or the representation level. When it exists at the adaptation set level, the auxiliary information descriptor describes all representations in the adaptation set; when it exists at the representation level, the auxiliary information descriptor describes the corresponding representation. The AuxiliaryInfoDescriptor descriptor indicates the relevant attributes of the auxiliary information of the point cloud media. For specific attributes, please refer to Table 8. Table 8 is used to indicate the elements and attributes of an auxiliary information descriptor provided by the embodiment of the present application.

[0324] Table 8

[0325]

[0326]

[0327] Among them, N in Table 8 represents the total number of auxiliary level parameters in the auxiliary information, such as Figure 8 In the table, N=2, indicating that there are two auxiliary level parameters. M indicates that the corresponding field (such as AuxiliaryInfo@AuxiliaryLevel in Table 8) is mandatory; CM indicates that the corresponding field (such as AuxiliaryInfo@spatialRegionId in Table 8) is conditional; O indicates that the corresponding field (such as AuxiliaryInfo@tileId in Table 8) is optional.

[0328] In Table 8, unsigned represents unsigned, Short represents short integer, bool represents Boolean variable, Int represents integer, vector represents vector, and float represents floating-point type.

[0329] Another feasible transmission signaling extension is extended in the embodiment of the present application in SMT signaling and proposes an auxiliary information descriptor, which exists at the representation level and is used to describe the corresponding media resource and indicate the auxiliary information of the media resource. Please refer to Table 9, which is used to indicate an auxiliary information descriptor syntax provided by the embodiment of the present application:

[0330] Table 9

[0331]

[0332]

[0333] The semantics of the syntax shown in Table 9 are as follows: When the default_auxiliary_info_flag takes the fifth indicator value, it indicates that the information indicated in the descriptor is the default auxiliary information; when the field takes the sixth indicator value, it indicates that the information indicated in the descriptor is auxiliary information for a specific spatial region. auxiliary_info_type indicates the auxiliary information type, and its value has the same meaning as the auxiliary_info_type field in the auxiliary information data box. auxiliary_info_level indicates the level of the auxiliary information. For credibility information, saliency information, quality level information, and priority information, the larger the value of this field, the higher the level of the corresponding information. When the region_id_ref_flag takes the eleventh indicator value (for example, 1), the spatial region corresponding to the auxiliary level is indexed by the spatial region identifier; when the field takes the twelfth indicator value (for example, 0), the spatial region corresponding to the auxiliary level is directly indicated by the spatial region position information. Spatial_region_id is the spatial region identifier field, indicating the spatial region identifier. Anchor_point_x, y, z indicates the x, y, z coordinates of the spatial region anchor point. Bounding_box_x, y, z indicates the length of the spatial area along the x, y, z axis. When the related_tile_info_flag takes the third information value (for example, 1), it indicates that the spatial area associated with the auxiliary level parameter is associated with one or more spatial blocks; when it takes the fourth information value (for example, 0), it indicates that the spatial area associated with the auxiliary level has no spatial blocks associated with it. When the related_slice_info_flag takes the first information value (for example, 1), it indicates that the spatial area associated with the auxiliary level is associated with one or more point cloud slices; when it takes the second information value (for example, 0), it indicates that the spatial area associated with the auxiliary level has no point cloud slices associated with it. Num_tiles indicates the number of spatial blocks associated with the spatial area. Tile_id indicates the identifier of the associated spatial block. Num_slices indicates the number of point cloud slices associated with the spatial area. Slice_id indicates the identifier of the associated point cloud slice. The meaning of the above fields can also be found in the description in Table 1 above.

[0334] When the auxiliary level parameter in the auxiliary information changes over time, the auxiliary information of the point cloud media exists in the media file in the form of a metadata track. At this time, the transmission signaling does not include an auxiliary information descriptor or auxiliary information descriptor, but the auxiliary information metadata track or auxiliary information sample group will exist in the transmission signaling as a media resource in the form of Representation.

[0335] When the auxiliary information of point cloud media is associated with the spatial region and does not change with time, e.g. Figure 8 In the third point cloud media example, the server associates the auxiliary information with the spatial information in the transmission signaling, generates a signaling and sends it to the client. Figure 8 For the third point cloud media in the example, the transmission signaling generated by the server contains three auxiliary information descriptors:

[0336] AuxiliaryInfo descriptor1:

[0337] @defaultInfoFlag=0; @auxiliaryInfoType=1; @auxiliaryInfoLevel=2;

[0338] @regionIdRefFlag=1; @spatialRegionId=1; @sliceId=0,1;

[0339] AuxiliaryInfo descriptor2:

[0340] @defaultInfoFlag=0; @auxiliaryInfoType=1; @auxiliaryInfoLevel=1;

[0341] @regionIdRefFlag=1; @spatialRegionId=2; @sliceId=2,3

[0342] AuxiliaryInfo descriptor3:

[0343] @defaultInfoFlag=1; @auxiliaryInfoType=1; @auxiliaryInfoLevel=0;

[0344] against Figure 8For the third point cloud media in the example, when establishing streaming transmission with the server, for the client, if spatial area 1 and spatial area 2 correspond to different media resources Representation 1 and Representation 2, since spatial area 1 has a higher assistance level parameter, the client can prioritize the transmission of Representation 1 during transmission.

[0345] From the above, it can be seen that the embodiment of the present application can determine the auxiliary information corresponding to the point cloud media for indicating the time range, or the auxiliary information corresponding to the point cloud media for indicating the spatial range, so the embodiment of the present application can encapsulate the auxiliary information of the point cloud media together with the point cloud code stream to obtain a media file, and the auxiliary information indication method provided by the embodiment of the present application can improve the accuracy of the auxiliary information of the point cloud media, and then when rendering the point cloud media, the rendering effect of the target range can be determined by accurate auxiliary information, so the presentation effect of the point cloud media can be optimized. In addition, the present application can use a general indication information for indication of different types of auxiliary information, so the efficiency of encapsulation of auxiliary information can be improved. In addition, by encapsulating different types of auxiliary information into one media file, the embodiment of the present application can avoid one auxiliary information corresponding to one media file, so the redundancy of the point cloud code stream can be reduced; by transmitting a media file containing different types of auxiliary information, the waste of network transmission resources can also be reduced.

[0346] Further, see Figure 9 , Figure 9 This is a flow diagram of a media data processing method provided by an embodiment of the present application. Figure 3 The method can be implemented by a content consumption device (e.g., the aforementioned Figure 3 The method may be performed by the content consumption device 200B in the corresponding embodiment. For example, the content consumption device may be a terminal integrated with a client (such as a video client). The method may include at least the following steps S301-S302:

[0347] Step S301: Acquire a media file; the media file includes assigned general indication information obtained based on at least one auxiliary information of the point cloud media; the assigned general indication information is used to indicate at least one auxiliary information.

[0348] Specifically, the at least one auxiliary information includes target auxiliary information; the target auxiliary information includes an auxiliary level parameter for indicating a target range of the point cloud media; the target range includes a spatial range or a temporal range;

[0349] The assigned general indication information includes at least one of assigned data box information or assigned metadata track information.

[0350] The client can obtain the media file of the immersive media sent by the server and decapsulate the media file to obtain the point cloud code stream and auxiliary information of the point cloud media in the media file. It can be understood that the decapsulation process is the opposite of the encapsulation process. The client can decapsulate the media file according to the file format requirements used when encapsulating to obtain the point cloud code stream. The specific process of the server generating and sending media files can be found in the above Figure 4 The corresponding embodiments will not be described in detail here.

[0351] Step S302: decapsulate the media file to obtain a point cloud code stream and at least one auxiliary information, and decode the point cloud code stream to obtain point cloud media.

[0352] Specifically, the media file includes an auxiliary information data box for indicating target auxiliary information; the assigned data box information includes the auxiliary information data box; when the auxiliary information data box is included in the sample entry of the point cloud track corresponding to the media file, it is determined that the target range includes a spatial range; wherein, an auxiliary level parameter in the target auxiliary information is used to indicate a spatial range; a spatial range includes a spatial area respectively contained in T point cloud frames; the point cloud media includes T point cloud frames; T is a positive integer.

[0353] The total number of spatial ranges is at least two, and the at least two spatial ranges include a first spatial range and a second spatial range;

[0354] The specific process of rendering point cloud media may include: obtaining the auxiliary level parameter U for indicating the first spatial range in the target auxiliary information; v , obtain the auxiliary level parameter U used to indicate the second spatial range v+1 ; v is a positive integer, and v is less than the total number of auxiliary level parameters in the target auxiliary information; if the auxiliary level parameter U v Greater than the auxiliary level parameter U v+1 , then determine the rendering level corresponding to the first spatial range, which is better than the rendering level corresponding to the second spatial range; if the auxiliary level parameter U v Less than the auxiliary level parameter U v+1 , it is determined that the rendering level corresponding to the second spatial range is better than the rendering level corresponding to the first spatial range.

[0355] Specifically, when a media file includes an auxiliary information metadata track for indicating target auxiliary information, determining the target range includes a time range; the assigned metadata track information includes the auxiliary information metadata track; wherein the auxiliary information metadata track includes W sample numbers associated with the target auxiliary information; wherein one sample number corresponds to one point cloud frame; the point cloud media includes W point cloud frames; and W is a positive integer.

[0356] The time range includes the first point cloud frame and the second point cloud frame in the W point cloud frames; the specific process of rendering the point cloud media may include: obtaining the auxiliary level parameter X for indicating the first point cloud frame in the target auxiliary information y , get the auxiliary level parameter X used to indicate the second point cloud frame y+1 ; y is a positive integer, and y is less than the total number of auxiliary level parameters in the auxiliary information; if the auxiliary level parameter X y Greater than the auxiliary level parameter X y+1 , then the rendering level corresponding to the first point cloud frame is determined to be better than the rendering level corresponding to the second point cloud frame; if the auxiliary level parameter X y Less than the auxiliary level parameter X y+1 , it is determined that the rendering level corresponding to the second point cloud frame is better than the rendering level corresponding to the first point cloud frame.

[0357] Among them, the specific process of rendering point cloud media may also include: if the first point cloud frame includes at least two spatial areas, and the auxiliary levels corresponding to at least two spatial areas are different, then in the target auxiliary information, the auxiliary level parameter Zα corresponding to the first spatial area is obtained, and the auxiliary level parameter Zα+1 corresponding to the second spatial area is obtained; the first spatial area and the second spatial area both belong to at least two spatial areas; α is a positive integer, and α is less than the total number of auxiliary level parameters in the auxiliary information; if the auxiliary level parameter Zα is greater than the auxiliary level parameter Zα+1, then it is determined that the rendering level corresponding to the first spatial area is better than the rendering level corresponding to the second spatial area; if the auxiliary level parameter Zα is less than the auxiliary level parameter Zα+1, then it is determined that the rendering level corresponding to the second spatial area is better than the rendering level corresponding to the first spatial area.

[0358] Optionally, at least one auxiliary information includes first auxiliary information and second auxiliary information; the first auxiliary information includes an auxiliary level parameter ζ for indicating a target range of the point cloud media; the second auxiliary information includes an auxiliary level parameter ψ for indicating a target range of the point cloud media; the total number of target ranges is at least two; the at least two target ranges include a target range β and a target range δ; the auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information; in the auxiliary level parameter ζ, obtain the auxiliary level parameter ε corresponding to the target range β, obtain the auxiliary level parameter φ corresponding to the target range δ; in the auxiliary level parameter ψ, obtain the target The assistance level parameter γ corresponding to the range β is used to obtain the assistance level parameter μ corresponding to the target range δ; the assistance level parameter ε and the assistance level parameter γ are weightedly summed to obtain the total assistance level parameter η corresponding to the target range β; the assistance level parameter φ and the assistance level parameter μ are weighted summed to obtain the total assistance level parameter λ corresponding to the target range δ; if the total assistance level parameter η is greater than the total assistance level parameter λ, then it is determined that the rendering level corresponding to the target range β is better than the rendering level corresponding to the target range δ; if the total assistance level parameter η is less than the total assistance level parameter λ, then it is determined that the rendering level corresponding to the target range δ is better than the rendering level corresponding to the target range β.

[0359] It can be understood that the decoding process is the opposite of the encoding process. The client can decode the point cloud code stream according to the file format requirements used during encoding to obtain point cloud media.

[0360] After the client decapsulates and decodes the point cloud file / file fragment, it can flexibly allocate computing resources in the process of presenting and rendering the point cloud media according to the auxiliary information of the point cloud media, and optimize the presentation effect of the target range. Figure 5 In the first point cloud media shown, sample2 is the point cloud frame with the highest credibility, sample3 is the point cloud frame with the second highest credibility, and the remaining samples only have the default credibility level parameter (which can be set to 100). Therefore, the client can render sample2 and sample3 more finely, that is, the rendering level is positively correlated with the credibility level.

[0361] Please see again Figure 6In the second point cloud media of the example, sample2 is the frame with the highest significance. Sample1 contains the spatial region with the highest significance and the spatial region with the second highest significance. Sample3 contains the spatial region with the second highest significance. The remaining samples only have the default significance, that is, the reference significance level parameter. Therefore, the client can render the spatial region with the highest significance and the spatial region with the second highest significance in sample2 and sample1, and the spatial region with the second highest significance in sample3 more finely. For example, the rendering level corresponding to sample2 is better than the rendering level corresponding to sample1. In sample1, the rendering level corresponding to the spatial region with the highest significance is better than the rendering level corresponding to the spatial region with the second highest significance; and the rendering level corresponding to sample1 is better than the rendering level corresponding to sample3.

[0362] Please see again Figure 8 In the third point cloud media illustrated as an example, spatial region 1 is a region with higher saliency, so the client can render spatial region 1 more finely, that is, the rendering level corresponding to spatial region 1 is better than the rendering level corresponding to spatial region 2.

[0363] In addition, if the target ranges indicated by multiple different types of auxiliary information are the same, the embodiment of the present application can optimize the presentation of point cloud media based on the multiple different types of auxiliary information.

[0364] As can be seen from the above, the embodiments of the present application can determine the auxiliary information corresponding to the point cloud media for indicating the time range, or the auxiliary information corresponding to the point cloud media for indicating the spatial range. Therefore, the embodiments of the present application can encapsulate the auxiliary information of the point cloud media together with the point cloud code stream to obtain a media file. In addition, the auxiliary information indication method provided by the embodiments of the present application can improve the accuracy of the auxiliary information of the point cloud media. When rendering the point cloud media, the rendering effect of the target range can be determined through accurate auxiliary information, thereby optimizing the presentation effect of the point cloud media. In addition, the present application can use a universal indication information for indicating different types of auxiliary information, thereby improving the efficiency of encapsulating the auxiliary information.

[0365] See Figure 10 , Figure 10 This is a structural diagram of a media data processing device provided in an embodiment of the present application. The media data processing device can be a computer program (including program code) running on a content production device. For example, the media data processing device is an application software in the content production device; the device can be used to execute the corresponding steps in the media data processing method provided in an embodiment of the present application. Figure 10 As shown, the media data processing device 1 may include: an information acquisition module 11 and an information packaging module 12 .

[0366] An information acquisition module 11 is configured to acquire at least one auxiliary information of the point cloud media and unassigned general indication information;

[0367] The information encapsulation module 12 is used to encode the point cloud media to obtain a point cloud code stream;

[0368] The information encapsulation module is also used to encapsulate the point cloud code stream, at least one auxiliary information and unassigned general indication information to obtain a media file; the media file includes the assigned general indication information, and the assigned general indication information is obtained by assigning a value to the unassigned general indication information based on at least one auxiliary information; the assigned general indication information is used to indicate at least one auxiliary information.

[0369] The specific implementation of the information acquisition module 11 and the information packaging module 12 can be found in the above Figure 4 Steps S101 and S102 in the corresponding embodiment will not be described in detail here.

[0370] In one embodiment, the assigned general indication information includes at least one of assigned data box information or assigned metadata track information.

[0371] In one embodiment, at least one auxiliary information includes auxiliary information B c ; c is a positive integer, and c is less than or equal to the total number of at least one type of auxiliary information;

[0372] Auxiliary Information B c Includes target range D for indicating point cloud media c Assistance level parameter; target range D c Including spatial range E c or time range F c .

[0373] In one embodiment, at least one auxiliary information includes auxiliary information B c ; c is a positive integer, and c is less than or equal to the total number of at least one type of auxiliary information;

[0374] Auxiliary Information B c Includes target range D for indicating point cloud media c Assistance level parameter; target range D c Including spatial range E c or time range F c ;

[0375] When the media file includes auxiliary information B cWhen the first auxiliary information data box is included in the sample entry of the point cloud track corresponding to the media file, the target range D c Including spatial range E c ; The assigned data box information includes the first auxiliary information data box;

[0376] Among them, the auxiliary information B c An auxiliary level parameter in is used to indicate a spatial range; a spatial range includes a spatial area respectively contained in G point cloud frames; the point cloud media includes G point cloud frames; G is a positive integer; or

[0377] When the media file includes auxiliary information B c When the auxiliary information metadata track is set, the target range D c Including time range F c ; The assigned metadata track information includes the auxiliary information metadata track;

[0378] Among them, the auxiliary information metadata track includes auxiliary information B c M associated sample numbers; wherein one sample number corresponds to one point cloud frame; the point cloud media includes M point cloud frames; M is a positive integer.

[0379] In one embodiment, the first auxiliary information data box includes an auxiliary information type field whose field value is a first type value, a static data structure quantity field, and an auxiliary information data structure with default attributes;

[0380] The first type value represents auxiliary information B c The auxiliary information type;

[0381] The Static Data Structure Number field is used to indicate the total number of auxiliary information data structures with static attributes.

[0382] In one embodiment, the value of the static data structure number field is H, indicating H auxiliary information data structures with static attributes; the H auxiliary information data structures with static attributes include auxiliary information data structures I with static attributes. j , where H and j are both positive integers, and j is less than or equal to H;

[0383] Auxiliary information data structure with static attributes I j Include field value for auxiliary level parameter L j The auxiliary level field and the target range indication field whose field value is the first indication value; the auxiliary level parameter L j Different from the default assistance level parameters; the default assistance level parameters and the assistance level parameters L j All belong to auxiliary information B cThe auxiliary level parameter in ;

[0384] The first indicator value represents the assistance level parameter L j Used to indicate the auxiliary level of a spatial extent.

[0385] In one embodiment, the auxiliary information data structure with default attributes includes an assistance level field whose field value is a default assistance level parameter and a target range indication field whose field value is a second indication value; the default assistance level parameter belongs to the auxiliary information B c The auxiliary level parameter in ;

[0386] The second indication value indicates that the default assistance level parameter is used to indicate an assistance level of a spatial range; the spatial range indicated by the default assistance level parameter includes a spatial range in the point cloud media except for the spatial range indicated by the auxiliary information data structure with static attributes.

[0387] In one embodiment, the first assistance information data box includes an assistance algorithm indication field;

[0388] When the value of the auxiliary algorithm indication field is the third indication value, it indicates that the auxiliary information B c There are auxiliary algorithms;

[0389] When the value of the auxiliary algorithm indication field is the fourth indication value, it indicates that the auxiliary information B c There is no auxiliary algorithm; the fourth indicator value is different from the third indicator value.

[0390] In one embodiment, when the field value of the auxiliary algorithm indication field is the third indication value, the first auxiliary information data box further includes an auxiliary algorithm type field;

[0391] When the field value of the auxiliary algorithm type field is the second type value, it indicates that the auxiliary information B c Determined by the auxiliary information detection algorithm;

[0392] When the field value of the auxiliary algorithm type field is the third type value, it indicates that the auxiliary information B c Determined by data statistics; the third type value is different from the second type value.

[0393] Please see again Figure 10 The media data processing device 1 may further include: a signaling transmission module 13.

[0394] The signaling transmission module 13 is used to transmit the transmission signaling for the media file to the client when the first auxiliary information data box exists; the transmission instruction carries the auxiliary information description data K c ; Auxiliary information description data K cUsed to instruct the client to determine the order of obtaining different media sub-files in the media file when obtaining the media file through streaming transmission; auxiliary information description data K c Based on auxiliary information B c Generated.

[0395] The specific implementation of the signaling transmission module 13 can be found in the above Figure 7 Step S204 in the corresponding embodiment will not be described in detail here.

[0396] In one embodiment, the auxiliary information description data K c Including a default level indication field; when the field value of the default level indication field is the fifth indication value, it indicates that the auxiliary level parameter corresponding to the default information indication field is the default auxiliary level parameter; when the field value of the default level indication field is the sixth indication value, it indicates that the auxiliary level parameter corresponding to the default level indication field is not the default auxiliary level parameter; the sixth indication value is different from the fifth indication value; the auxiliary level parameter corresponding to the default level indication field belongs to the auxiliary information B c The auxiliary level parameter in .

[0397] In one embodiment, the sample entry of the auxiliary information metadata track includes an effective range indication field and a second auxiliary information data box; the effective range indication field is used to indicate the effective range of the auxiliary level parameters corresponding to the M sample numbers; the second auxiliary information data box is used to indicate the auxiliary information B c The information with static properties in the M sample numbers belongs to the auxiliary level parameters corresponding to the auxiliary information B c The auxiliary level parameter in .

[0398] In one embodiment, the M sample numbers include a sample number τ, where τ is a positive integer and τ is less than or equal to M;

[0399] When the field value of the effective range indication field is the seventh indication value, it indicates that the effective range of the auxiliary level parameter corresponding to the sample number τ is the auxiliary information metadata track;

[0400] When the field value of the effective range indication field is the eighth indication value, it means that the effective range of the auxiliary level parameter corresponding to the sample number τ is the point cloud frame corresponding to the sample number τ; the eighth indication value is different from the seventh indication value.

[0401] In one embodiment, the auxiliary information metadata track includes the sample number 0 n ; n is a positive integer and n is less than or equal to M; time range F c Including sample serial number O n The corresponding point cloud frame;

[0402] The auxiliary information metadata track includes the information for sample number O n The auxiliary level indication field of

[0403] When the field value of the auxiliary level indication field is the ninth indication value, it indicates that the auxiliary level indication field is the ninth indication value, which is the same as the sample number 0. n The associated assistance level is determined by the default assistance level parameter; the default assistance level parameter belongs to the assistance information B c The auxiliary level parameter in ;

[0404] When the field value of the auxiliary level indication field is the tenth indication value, it indicates that the auxiliary level indication field is the same as the sample number 0. n The associated assistance level is determined by the assistance information data structure; the tenth indicator value is different from the ninth indicator value.

[0405] In one embodiment, when the field value of the assistance level indication field is the ninth indication value, it indicates that the default assistance level parameter is used to indicate the sample sequence number 0. n The assistance level of the corresponding point cloud frame.

[0406] In one embodiment, when the field value of the auxiliary level indication field is the tenth indication value, the auxiliary information metadata track further includes a value for sample sequence number 0. n The data structure quantity field is used to indicate the total number of auxiliary information data structures.

[0407] In one embodiment, the value of the data structure number field is P, indicating P auxiliary information data structures; the P auxiliary information data structures include auxiliary information data structure Q r , r, P are all positive integers and r is less than or equal to P;

[0408] Auxiliary information data structure Q r Include field value for auxiliary level parameter S r The auxiliary level field and the target range indication field; the auxiliary level parameter S r Belongs to auxiliary information B c The auxiliary level parameter in ;

[0409] When the field value of the target range indication field is the first indication value, it indicates that the auxiliary level parameter S r For indication, sample number O n The assistance level of a spatial region in the corresponding point cloud frame;

[0410] When the field value of the target range indication field is the second indication value, it indicates that the auxiliary level parameter S r For indication, sample number O n The auxiliary level of the corresponding point cloud frame; the second indication value is different from the first indication value.

[0411] Please see again Figure 10 , the at least one assistance information includes target assistance information; the target assistance information includes an assistance level parameter for indicating a target range;

[0412] The information encapsulation module 12 may include: a first encoding unit 121 .

[0413] The first encoding unit 121 is configured to optimize and encode a target range of the point cloud media according to the assistance level parameter in the target auxiliary information to obtain a point cloud code stream.

[0414] The specific implementation of the first encoding unit 121 can refer to the above Figure 7 Step S202 in the corresponding embodiment will not be described in detail here.

[0415] Please see again Figure 10 , the total number of target ranges is at least two, the at least two target ranges include a first target range and a second target range; the assistance level parameter in the target assistance information includes a first assistance level parameter corresponding to the first target range, and a second assistance level parameter corresponding to the second target range;

[0416] The first encoding unit 121 may include: a first determining subunit 1211 , a first encoding subunit 1212 , and a first generating subunit 1213 .

[0417] A first determining subunit 1211 is configured to determine a first coding level within a first target range according to a first assistance level parameter, and determine a second coding level within a second target range according to a second assistance level parameter, wherein when the first assistance level parameter is greater than the second assistance level parameter, the first coding level is superior to the second coding level;

[0418] The first encoding subunit 1212 is configured to optimize and encode the first target range using a first encoding level to obtain a first sub-point cloud code stream, and optimize and encode the second target range using a second encoding level to obtain a second sub-point cloud code stream;

[0419] The first generating sub-unit 1213 is configured to generate a point cloud code stream according to the first sub-point cloud code stream and the second sub-point cloud code stream.

[0420] The specific implementation of the first determining subunit 1211, the first encoding subunit 1212 and the first generating subunit 1213 can be found in the above Figure 7 Step S202 in the corresponding embodiment will not be described in detail here.

[0421] Please see again Figure 10, the at least one auxiliary information includes first auxiliary information and second auxiliary information; the first auxiliary information includes an auxiliary level parameter σ for indicating a target range of the point cloud media; the second auxiliary information includes an auxiliary level parameter ρ for indicating a target range of the point cloud media; the auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information;

[0422] The information encapsulation module 12 may include: a second encoding unit 122 .

[0423] The second encoding unit 122 is configured to optimize and encode the target range according to the auxiliary level parameter σ and the auxiliary level parameter ρ to obtain a point cloud code stream.

[0424] The specific implementation of the second encoding unit 122 can refer to the above Figure 4 Step S102 in the corresponding embodiment will not be described in detail here.

[0425] Please see again Figure 10 , the total number of target ranges is at least two, the at least two target ranges including a third target range and a fourth target range; the assistance level parameter σ includes a third assistance level parameter corresponding to the third target range, and a fourth assistance level parameter corresponding to the fourth target range; the assistance level parameter ρ includes a fifth assistance level parameter corresponding to the third target range, and a sixth assistance level parameter corresponding to the fourth target range;

[0426] The second encoding unit 122 may include: a first summing subunit 1221 , a second summing subunit 1222 , a second determining subunit 1223 , a second encoding subunit 1224 , and a second generating subunit 1225 .

[0427] A first summing subunit 1221 is configured to perform a weighted summation on the third assistance level parameter and the fifth assistance level parameter to obtain a first total assistance level parameter corresponding to the third target range;

[0428] a second summing subunit 1222 for performing a weighted summation on the fourth assistance level parameter and the sixth assistance level parameter to obtain a second total assistance level parameter corresponding to the fourth target range;

[0429] The second determining subunit 1223 is configured to determine a third coding level within a third target range based on the first total assistance level parameter, and determine a fourth coding level within a fourth target range based on the second total assistance level parameter; when the first total assistance level parameter is greater than the second total assistance level parameter, the third coding level is superior to the fourth coding level;

[0430] The second encoding subunit 1224 is configured to perform optimized encoding on the third target range using a third encoding level to obtain a third sub-point cloud code stream, and to perform optimized encoding on the fourth target range using a fourth encoding level to obtain a fourth sub-point cloud code stream;

[0431] The second generating sub-unit 1225 is configured to generate a point cloud code stream according to the third point cloud sub-code stream and the fourth point cloud sub-code stream.

[0432] The specific implementation of the first summing subunit 1221, the second summing subunit 1222, the second determining subunit 1223, the second encoding subunit 1224 and the second generating subunit 1225 can be referred to above. Figure 4 Step S102 in the corresponding embodiment will not be described in detail here.

[0433] From the above, it can be seen that the embodiment of the present application can determine the auxiliary information corresponding to the point cloud media for indicating the time range, or the auxiliary information corresponding to the point cloud media for indicating the spatial range, so the embodiment of the present application can encapsulate the auxiliary information of the point cloud media together with the point cloud code stream to obtain a media file, and the auxiliary information indication method provided by the embodiment of the present application can improve the accuracy of the auxiliary information of the point cloud media, and then when rendering the point cloud media, the rendering effect of the target range can be determined by accurate auxiliary information, so the presentation effect of the point cloud media can be optimized. In addition, the present application can use a general indication information for indication of different types of auxiliary information, so the efficiency of encapsulation of auxiliary information can be improved. In addition, by encapsulating different types of auxiliary information into one media file, the embodiment of the present application can avoid one auxiliary information corresponding to one media file, so the redundancy of the point cloud code stream can be reduced; by transmitting a media file containing different types of auxiliary information, the waste of network transmission resources can also be reduced.

[0434] See Figure 11 , Figure 11 This is a schematic diagram of the structure of a media data processing device provided in an embodiment of the present application. Figure 2 The media data processing device may be a computer program (including program code) running on a content consumption device, for example, the media data processing device is an application software (for example, a video client) in the content consumption device; the device may be used to execute the corresponding steps in the media data processing method provided in the embodiment of the present application. Figure 11 As shown, the media data processing device 2 may include: a file acquisition module 21 and a code stream decoding module 22.

[0435] The file acquisition module 21 is used to acquire a media file; the media file includes assigned general indication information obtained based on at least one auxiliary information of the point cloud media; the assigned general indication information is used to indicate the at least one auxiliary information;

[0436] The code stream decoding module 22 is used to decapsulate the media file to obtain a point cloud code stream and at least one auxiliary information, and decode the point cloud code stream to obtain point cloud media.

[0437] The specific implementation of the file acquisition module 21 and the code stream decoding module 22 can be found in the above Figure 9 Steps S301 and S302 in the corresponding embodiment will not be described in detail here.

[0438] In one embodiment, the at least one assistance information includes target assistance information;

[0439] The target assistance information includes an assistance level parameter for indicating a target range of the point cloud media; the target range includes a spatial range or a temporal range;

[0440] The assigned general indication information includes at least one of assigned data box information or assigned metadata track information.

[0441] Please see again Figure 11 , the media file includes an auxiliary information data box for indicating target auxiliary information; the assigned data box information includes the auxiliary information data box;

[0442] The media data processing device 2 may further include:

[0443] A first determining module 23 is configured to determine that the target range includes a spatial range when the auxiliary information data box is included at a sample entry of a point cloud track corresponding to the media file;

[0444] Among them, an auxiliary level parameter in the target auxiliary information is used to indicate a spatial range; a spatial range includes a spatial area respectively contained in T point cloud frames; the point cloud media includes T point cloud frames; and T is a positive integer.

[0445] The specific implementation of the first determination module 23 can be found in the above Figure 9 Step S302 in the corresponding embodiment will not be described in detail here.

[0446] Please see again Figure 11 , the total number of spatial ranges is at least two, and the at least two spatial ranges include a first spatial range and a second spatial range;

[0447] The media data processing device 2 may further include: a first obtaining module 24 and a second determining module 25 .

[0448] The first acquisition module 24 is configured to acquire an assistance level parameter U indicating a first spatial range from the target assistance information. v, obtain the auxiliary level parameter U used to indicate the second spatial range v+1 ; v is a positive integer, and v is less than the total number of assistance level parameters in the target assistance information;

[0449] The second determining module 25 is used to determine if the assistance level parameter U v Greater than the auxiliary level parameter U v+1 , it is determined that the rendering level corresponding to the first spatial range is better than the rendering level corresponding to the second spatial range;

[0450] The second determining module 25 is further configured to determine if the assistance level parameter U v Less than the auxiliary level parameter U v+1 , it is determined that the rendering level corresponding to the second spatial range is better than the rendering level corresponding to the first spatial range.

[0451] The specific implementation of the first acquisition module 24 and the second determination module 25 can be found in the above Figure 9 Step S302 in the corresponding embodiment will not be described in detail here.

[0452] Please see again Figure 11 , the media data processing device 2 may further include: a third determining module 26.

[0453] A third determining module 26 is configured to determine, when the media file includes an auxiliary information metadata track for indicating target auxiliary information, that the target range includes a time range; and the assigned metadata track information includes the auxiliary information metadata track;

[0454] The specific implementation of the third determination module 26 can be found in the above Figure 9 Step S302 in the corresponding embodiment will not be described in detail here.

[0455] In one embodiment, the auxiliary information metadata track includes W sample numbers associated with the target auxiliary information; wherein one sample number corresponds to one point cloud frame; the point cloud media includes W point cloud frames; and W is a positive integer.

[0456] Please see again Figure 11 , the time range includes the first point cloud frame and the second point cloud frame in the W point cloud frames;

[0457] The media data processing device 2 may further include: a second obtaining module 27 and a fourth determining module 28 .

[0458] The second acquisition module 27 is used to obtain the auxiliary level parameter X indicating the first point cloud frame from the target auxiliary information. y , get the auxiliary level parameter X used to indicate the second point cloud frame y+1; y is a positive integer, and y is less than the total number of auxiliary level parameters in the auxiliary information;

[0459] The fourth determining module 28 is configured to determine if the assistance level parameter X y Greater than the auxiliary level parameter X y+1 , it is determined that the rendering level corresponding to the first point cloud frame is better than the rendering level corresponding to the second point cloud frame;

[0460] The fourth determining module 28 is further configured to determine if the assistance level parameter X y Less than the auxiliary level parameter X y+1 , it is determined that the rendering level corresponding to the second point cloud frame is better than the rendering level corresponding to the first point cloud frame.

[0461] The specific implementation of the second acquisition module 27 and the fourth determination module 28 can be found in the above Figure 9 Step S302 in the corresponding embodiment will not be described in detail here.

[0462] Please see again Figure 11 , the media data processing device 2 may further include: a fifth determining module 29.

[0463] The second acquisition module 27 is further configured to, if the first point cloud frame includes at least two spatial regions and the at least two spatial regions correspond to different assistance levels, obtain, from the target auxiliary information, an assistance level parameter Zα corresponding to the first spatial region and an assistance level parameter Zα+1 corresponding to the second spatial region; the first spatial region and the second spatial region both belong to the at least two spatial regions; α is a positive integer and is less than the total number of assistance level parameters in the auxiliary information;

[0464] a fifth determining module 29 for determining that the rendering level corresponding to the first spatial region is superior to the rendering level corresponding to the second spatial region if the assistive level parameter Zα is greater than the assistive level parameter Zα+1;

[0465] The fifth determining module 29 is further configured to determine that the rendering level corresponding to the second spatial region is superior to the rendering level corresponding to the first spatial region if the auxiliary level parameter Zα is less than the auxiliary level parameter Zα+1.

[0466] The specific implementation of the second acquisition module 27 and the fifth determination module 29 can be found in the above Figure 9 Step S302 in the corresponding embodiment will not be described in detail here.

[0467] Please see again Figure 11, the at least one auxiliary information includes first auxiliary information and second auxiliary information; the first auxiliary information includes an auxiliary level parameter ζ for indicating a target range of the point cloud media; the second auxiliary information includes an auxiliary level parameter ψ for indicating a target range of the point cloud media; the total number of target ranges is at least two; the at least two target ranges include a target range β and a target range δ; the auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information;

[0468] The media data processing device 2 may further include: a third acquisition module 30 , a weighted summation module 31 and a sixth determination module 32 .

[0469] The third acquisition module 30 is configured to acquire, from the assistance level parameter ζ, an assistance level parameter ε corresponding to the target range β and an assistance level parameter φ corresponding to the target range δ;

[0470] The third acquisition module 30 is further configured to acquire, from the assistance level parameter ψ, an assistance level parameter γ corresponding to the target range β and an assistance level parameter μ corresponding to the target range δ;

[0471] A weighted summation module 31 is configured to perform a weighted summation of the assistance level parameter ε and the assistance level parameter γ to obtain a total assistance level parameter η corresponding to the target range β;

[0472] The weighted summation module 31 is further configured to perform a weighted summation on the assistance level parameter φ and the assistance level parameter μ to obtain a total assistance level parameter λ corresponding to the target range δ;

[0473] a sixth determining module 32 for determining, if the total assistance level parameter η is greater than the total assistance level parameter λ, that the rendering level corresponding to the target range β is superior to the rendering level corresponding to the target range δ;

[0474] The sixth determining module 32 is further configured to determine that the rendering level corresponding to the target range δ is better than the rendering level corresponding to the target range β if the total assistance level parameter η is less than the total assistance level parameter λ.

[0475] The specific implementation of the third acquisition module 30, the weighted summation module 31 and the sixth determination module 32 can be found in the above Figure 9 Step S302 in the corresponding embodiment will not be described in detail here.

[0476] From the above, it can be seen that the embodiment of the present application can determine the auxiliary information corresponding to the point cloud media for indicating the time range, or the auxiliary information corresponding to the point cloud media for indicating the spatial range, so the embodiment of the present application can encapsulate the auxiliary information of the point cloud media together with the point cloud code stream to obtain a media file, and the auxiliary information indication method provided by the embodiment of the present application can improve the accuracy of the auxiliary information of the point cloud media, and then when rendering the point cloud media, the rendering effect of the target range can be determined by accurate auxiliary information, so the presentation effect of the point cloud media can be optimized. In addition, the present application can use a general indication information for indication of different types of auxiliary information, so the efficiency of encapsulation of auxiliary information can be improved. In addition, by encapsulating different types of auxiliary information into one media file, the embodiment of the present application can avoid one auxiliary information corresponding to one media file, so the redundancy of the point cloud code stream can be reduced; by transmitting a media file containing different types of auxiliary information, the waste of network transmission resources can also be reduced.

[0477] See Figure 12 , is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 12 As shown, the computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the above-mentioned computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 1005 may optionally also be at least one storage device located away from the aforementioned processor 1001. As Figure 12 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application.

[0478] In such Figure 12In the illustrated computer device 1000, the network interface 1004 provides network communication functionality; the user interface 1003 primarily provides an interface for user input; and the processor 1001 is configured to invoke a device control application stored in the memory 1005. It should be understood that the computer device 1000 described in the embodiments of this application can execute the data processing methods or apparatuses described in the preceding embodiments, and these descriptions are omitted. Furthermore, the beneficial effects of employing the same methods are also omitted.

[0479] The present application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the data processing methods or apparatuses described in the preceding embodiments, which are not described in detail here. Furthermore, the description of the beneficial effects of the same methods is not described in detail here.

[0480] The computer-readable storage medium may be the data processing device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may also include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0481] The present application also provides a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, enabling the computer device to perform the data processing methods or apparatuses described in the preceding embodiments, which are not further detailed here. Furthermore, the beneficial effects of the same methods are not further detailed here.

[0482] For further information, see Figure 13 , Figure 13 Schematic diagram of a data processing system provided in an embodiment of the present application. The data processing system 3 may include a data processing device 1a and a data processing device 2a. The data processing device 1a may be the above-mentioned Figure 10 In the corresponding embodiment of the media data processing device 1, it can be understood that the data processing device 1a can be integrated into the above Figure 3 The content production device 200A in the corresponding embodiment will not be described in detail here. Figure 11 In the corresponding embodiment, the media data processing device 2 can be understood that the data processing device 2a can be integrated into the above Figure 3 The content consumption device 200B in the corresponding embodiment will not be described in detail here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here. For technical details not disclosed in the data processing system embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0483] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.

[0484] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0485] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A media data processing method, characterized in that: include: Obtaining at least one auxiliary information of the point cloud media and unassigned general indication information; The at least one auxiliary information includes first auxiliary information and second auxiliary information; The auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information; Encoding the point cloud media to obtain a point cloud code stream; The point cloud code stream, the at least one auxiliary information and the unassigned general indication information are encapsulated to obtain a media file; the media file includes the assigned general indication information, and the assigned general indication information is obtained by assigning a value to the unassigned general indication information based on the at least one auxiliary information; the assigned general indication information is used to indicate the at least one auxiliary information.

2. The method according to claim 1, characterized in that The assigned general indication information includes at least one of assigned data box information or assigned metadata track information.

3. The method according to claim 2, characterized in that The at least one auxiliary information includes auxiliary information B c ; c is a positive integer, and c is less than or equal to the total number of the at least one auxiliary information; The auxiliary information B c including a target range D for indicating the point cloud media c Assistance level parameter; the target range D c Including spatial range E c or time range F c ; When the media file includes the auxiliary information B c When the first auxiliary information data box is included at the sample entry of the point cloud track corresponding to the media file, the target range D c Including the spatial range E c ; The assigned data box information includes the first auxiliary information data box; Among them, the auxiliary information B c An auxiliary level parameter in is used to indicate a spatial range; the spatial range includes a spatial area respectively contained in G point cloud frames; the point cloud media includes the G point cloud frames; G is a positive integer; or When the media file includes the auxiliary information B c When the auxiliary information metadata track is c Including the time range F c ; The assigned metadata track information includes the auxiliary information metadata track; The auxiliary information metadata track includes the auxiliary information B c M associated sample numbers; wherein one sample number corresponds to one point cloud frame; the point cloud media includes M point cloud frames; M is a positive integer.

4. The method according to claim 3, characterized in that The first auxiliary information data box includes an auxiliary information type field whose field value is a first type value, a static data structure quantity field, and an auxiliary information data structure with default attributes; The first type value represents the auxiliary information B c The auxiliary information type; The static data structure quantity field is used to indicate the total quantity of auxiliary information data structures with static attributes.

5. The method according to claim 4, characterized in that The value of the static data structure number field is H, indicating H auxiliary information data structures with static attributes; the H auxiliary information data structures with static attributes include auxiliary information data structures I with static attributes. j , where H and j are both positive integers, and j is less than or equal to H; The auxiliary information data structure I with static attributes j Include field value for auxiliary level parameter L j The auxiliary level field, and the target range indication field whose field value is the first indication value; the auxiliary level parameter L j Different from the default assistance level parameter; the default assistance level parameter and the assistance level parameter L j All belong to the auxiliary information B c The auxiliary level parameter in ; The first indicator value represents the assistance level parameter L j Used to indicate the auxiliary level of a spatial extent.

6. The method according to claim 4, characterized in that The auxiliary information data structure with default attributes includes an assistance level field whose field value is a default assistance level parameter, and a target range indication field whose field value is a second indication value; The default assistance level parameter belongs to the assistance information B c The auxiliary level parameter in ; The second indication value indicates that the default assistance level parameter is used to indicate an assistance level of a spatial range; The spatial range indicated by the default assistance level parameter includes a spatial range in the point cloud media except for a spatial range indicated by the auxiliary information data structure with static attributes.

7. The method according to claim 3, characterized in that The first auxiliary information data box includes an auxiliary algorithm indication field; When the value of the auxiliary algorithm indication field is the third indication value, it indicates that the auxiliary information B c There are auxiliary algorithms; When the value of the auxiliary algorithm indication field is the fourth indication value, it indicates that the auxiliary information B c There is no auxiliary algorithm; the fourth indicator value is different from the third indicator value.

8. The method according to claim 7, characterized in that When the field value of the auxiliary algorithm indication field is the third indication value, the first auxiliary information data box further includes an auxiliary algorithm type field; When the field value of the auxiliary algorithm type field is the second type value, it indicates that the auxiliary information B c Determined by the auxiliary information detection algorithm; When the field value of the auxiliary algorithm type field is the third type value, it indicates that the auxiliary information B c Determined by data statistics; the third type value is different from the second type value.

9. The method according to claim 3, characterized in that The method further comprises: When the first auxiliary information data box exists, the transmission signaling for the media file is transmitted to the client; the transmission signaling carries the auxiliary information description data K c ; The auxiliary information description data K c is used to instruct the client to determine the order of obtaining different media sub-files in the media file when obtaining the media file through streaming transmission; the auxiliary information description data K c is based on the auxiliary information B c Generated.

10. The method according to claim 9, characterized in that The auxiliary information description data K c Includes a default level indication field; When the field value of the default level indication field is the fifth indication value, it indicates that the auxiliary level parameter corresponding to the default level indication field is the default auxiliary level parameter; When the field value of the default level indication field is the sixth indication value, it indicates that the auxiliary level parameter corresponding to the default level indication field is not the default auxiliary level parameter; the sixth indication value is different from the fifth indication value; the auxiliary level parameter corresponding to the default level indication field belongs to the auxiliary information B c The auxiliary level parameter in .

11. The method according to claim 3, characterized in that The sample entry of the auxiliary information metadata track includes an effective range indication field and a second auxiliary information data box; the effective range indication field is used to indicate the effective range of the auxiliary level parameters corresponding to the M sample numbers; the second auxiliary information data box is used to indicate the auxiliary information B c The auxiliary level parameters corresponding to the M sample numbers belong to the auxiliary information B c The auxiliary level parameter in .

12. The method according to claim 11, characterized in that The M sample numbers include sample number , is a positive integer and Less than or equal to M; When the field value of the effective range indication field is the seventh indication value, it means that the sample number The effective scope of the corresponding auxiliary level parameter is the auxiliary information metadata track; When the field value of the effective range indication field is the eighth indication value, it means that the sample number The effective range of the corresponding auxiliary level parameter is the sample number corresponding point cloud frame; the eighth indication value is different from the seventh indication value.

13. The method according to claim 11, characterized in that The auxiliary information metadata track includes sample sequence number 0 n ; n is a positive integer and n is less than or equal to M; the time range F c Including sample serial number O n The corresponding point cloud frame; The auxiliary information metadata track includes the sample number O n The auxiliary level indication field of When the field value of the auxiliary level indication field is the ninth indication value, it indicates that the auxiliary level indication field is the ninth indication value. n The associated assistance level is determined by a default assistance level parameter; the default assistance level parameter belongs to the assistance information B c The auxiliary level parameter in ; When the field value of the auxiliary level indication field is the tenth indication value, it indicates that the auxiliary level indication field is the same as the sample number 0. n The associated assistance level is determined by an assistance information data structure; said tenth indicator value being different from said ninth indicator value.

14. The method according to claim 13, wherein: When the field value of the assistance level indication field is the ninth indication value, it indicates that the default assistance level parameter is used to indicate the sample sequence number 0. n The assistance level of the corresponding point cloud frame.

15. The method according to claim 13, characterized in that When the field value of the auxiliary level indication field is the tenth indication value, the auxiliary information metadata track further includes a value corresponding to the sample sequence number 0. n The data structure quantity field is used to indicate the total number of auxiliary information data structures.

16. The method according to claim 15, characterized in that The value of the data structure number field is P, indicating P auxiliary information data structures; the P auxiliary information data structures include auxiliary information data structure Q r , r, P are all positive integers and r is less than or equal to P; The auxiliary information data structure Q r Include field value for auxiliary level parameter S r The assistance level field, and the target range indication field; the assistance level parameter S r Belong to the auxiliary information B c The auxiliary level parameter in ; When the field value of the target range indication field is the first indication value, it indicates that the auxiliary level parameter S r For indication, the sample number O n The assistance level of a spatial region in the corresponding point cloud frame; When the field value of the target range indication field is the second indication value, it indicates that the auxiliary level parameter S r For indication, the sample number O n The auxiliary level of the corresponding point cloud frame; the second indication value is different from the first indication value.

17. The method according to claim 1, wherein The at least one auxiliary information includes target auxiliary information; The target assistance information includes an assistance level parameter for indicating a target range; The encoding of the point cloud media to obtain a point cloud code stream includes: The target range of the point cloud media is optimized and encoded according to the assistance level parameter in the target assistance information to obtain a point cloud code stream.

18. The method according to claim 17, characterized in that The total number of the target ranges is at least two, and the at least two target ranges include a first target range and a second target range; the assistance level parameter in the target assistance information includes a first assistance level parameter corresponding to the first target range, and a second assistance level parameter corresponding to the second target range; The step of optimizing and encoding the target range of the point cloud media according to the assistance level parameter in the target assistance information to obtain a point cloud code stream includes: determining a first coding level of the first target range according to the first assistance level parameter, and determining a second coding level of the second target range according to the second assistance level parameter, wherein when the first assistance level parameter is greater than the second assistance level parameter, the first coding level is superior to the second coding level; The first target range is optimized and encoded using the first encoding level to obtain a first sub-point cloud code stream, and the second target range is optimized and encoded using the second encoding level to obtain a second sub-point cloud code stream; The point cloud code stream is generated according to the first sub-point cloud code stream and the second sub-point cloud code stream.

19. The method according to claim 1, wherein The first auxiliary information includes an auxiliary level parameter σ for indicating a target range of the point cloud media; the second auxiliary information includes an auxiliary level parameter ρ for indicating a target range of the point cloud media; The encoding of the point cloud media to obtain a point cloud code stream includes: The target range is optimized and encoded according to the auxiliary level parameter σ and the auxiliary level parameter ρ to obtain a point cloud code stream.

20. The method according to claim 19, characterized in that The total number of the target ranges is at least two, the at least two target ranges including a third target range and a fourth target range; the assistance level parameter σ includes a third assistance level parameter corresponding to the third target range and a fourth assistance level parameter corresponding to the fourth target range; the assistance level parameter ρ includes a fifth assistance level parameter corresponding to the third target range and a sixth assistance level parameter corresponding to the fourth target range; The step of optimizing the encoding of the target range according to the auxiliary level parameter σ and the auxiliary level parameter ρ to obtain a point cloud code stream includes: performing a weighted summation on the third assistance level parameter and the fifth assistance level parameter to obtain a first total assistance level parameter corresponding to the third target range; performing a weighted summation on the fourth assistance level parameter and the sixth assistance level parameter to obtain a second total assistance level parameter corresponding to the fourth target range; determining a third coding level of the third target range based on the first total assistance level parameter, and determining a fourth coding level of the fourth target range based on the second total assistance level parameter; when the first total assistance level parameter is greater than the second total assistance level parameter, the third coding level is superior to the fourth coding level; The third target range is optimized and encoded using the third encoding level to obtain a third sub-point cloud code stream, and the fourth target range is optimized and encoded using the fourth encoding level to obtain a fourth sub-point cloud code stream. The point cloud code stream is generated according to the third sub-point cloud code stream and the fourth sub-point cloud code stream.

21. A media data processing method, characterized in that: include: Acquire a media file; the media file includes assigned general indication information obtained based on at least one auxiliary information of the point cloud media; The assigned general indication information is used to indicate the at least one auxiliary information; the at least one auxiliary information includes first auxiliary information and second auxiliary information; the auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information; The media file is decapsulated to obtain a point cloud code stream and the at least one auxiliary information, and the point cloud code stream is decoded to obtain the point cloud media.

22. The method according to claim 21, characterized in that The at least one auxiliary information includes target auxiliary information; The target assistance information includes an assistance level parameter for indicating a target range of the point cloud media; the target range includes a spatial range or a temporal range; The assigned general indication information includes at least one of assigned data box information or assigned metadata track information.

23. The method according to claim 22, characterized in that The media file includes an auxiliary information data box for indicating the target auxiliary information; the assigned data box information includes the auxiliary information data box; The method further comprises: When the auxiliary information data box is included at the sample entry of the point cloud track corresponding to the media file, determining that the target range includes the spatial range; Among them, an assistance level parameter in the target auxiliary information is used to indicate a spatial range; the spatial range includes a spatial area respectively contained in T point cloud frames; the point cloud media includes the T point cloud frames; T is a positive integer.

24. The method according to claim 23, wherein The total number of the spatial ranges is at least two, and the at least two spatial ranges include a first spatial range and a second spatial range; The method further comprises: In the target auxiliary information, obtain the auxiliary level parameter U for indicating the first spatial range v , obtain the auxiliary level parameter U for indicating the second spatial range v+1 ; v is a positive integer, and v is less than the total number of assistance level parameters in the target assistance information; If the assistance level parameter U v Greater than the assistance level parameter U v+1 , it is determined that the rendering level corresponding to the first spatial range is better than the rendering level corresponding to the second spatial range; If the assistance level parameter U v Less than the assistance level parameter U v+1 , it is determined that the rendering level corresponding to the second spatial range is better than the rendering level corresponding to the first spatial range.

25. The method according to claim 22, wherein The method further comprises: When the media file includes an auxiliary information metadata track for indicating the target auxiliary information, determining that the target range includes the time range; and the assigned metadata track information includes the auxiliary information metadata track; The auxiliary information metadata track includes W sample numbers associated with the target auxiliary information; wherein one sample number corresponds to one point cloud frame; the point cloud media includes W point cloud frames; and W is a positive integer.

26. The method according to claim 25, characterized in that The time range includes the first point cloud frame and the second point cloud frame in the W point cloud frames; The method further comprises: In the target auxiliary information, obtain the auxiliary level parameter X for indicating the first point cloud frame y , obtain the auxiliary level parameter X for indicating the second point cloud frame y+1 ; y is a positive integer, and y is less than the total number of assistance level parameters in the assistance information; If the assistance level parameter X y Greater than the assistance level parameter X y+1 , it is determined that the rendering level corresponding to the first point cloud frame is better than the rendering level corresponding to the second point cloud frame; If the assistance level parameter X y Less than the assistance level parameter X y+1 , it is determined that the rendering level corresponding to the second point cloud frame is better than the rendering level corresponding to the first point cloud frame.

27. The method according to claim 26, characterized in that The method further comprises: If the first point cloud frame includes at least two spatial regions, and the at least two spatial regions correspond to different assistance levels, then the assistance level parameter Z corresponding to the first spatial region is obtained in the target auxiliary information. α , get the auxiliary level parameter Z corresponding to the second spatial area α+1 The first spatial region and the second spatial region both belong to the at least two spatial regions; α is a positive integer, and α is less than the total number of assistance level parameters in the assistance information; If the assistance level parameter Z α Greater than the assistance level parameter Z α+1 , it is determined that the rendering level corresponding to the first spatial area is better than the rendering level corresponding to the second spatial area; If the assistance level parameter Z α Less than the assistance level parameter Z α+1 , it is determined that the rendering level corresponding to the second spatial area is better than the rendering level corresponding to the first spatial area.

28. The method according to claim 21, characterized in that The first assistance information includes an assistance level parameter for indicating a target range of the point cloud media. The second auxiliary information includes an auxiliary level parameter for indicating the target range of the point cloud media ; The total number of target ranges is at least two; at least two target ranges include the target range and target range ; The method further comprises: Parameters at the assistance level , obtain the target range Corresponding auxiliary level parameters , get the target range Corresponding auxiliary level parameters ; Parameters at the assistance level , obtain the target range Corresponding auxiliary level parameters , get the target range Corresponding auxiliary level parameters ; Parameters for the assistance level and the auxiliary level parameters Perform weighted summation to obtain the target range Corresponding total assistance level parameters ; Parameters for the assistance level and the auxiliary level parameters Perform weighted summation to obtain the target range Corresponding total assistance level parameters ; If the total assistance level parameter Greater than the total assistance level parameter , then determine the target range The corresponding rendering level is better than the target range The corresponding rendering level; If the total assistance level parameter Less than the total assistance level parameter , then determine the target range The corresponding rendering level is better than the target range The corresponding rendering level.

29. A media data processing device, characterized in that: include: An information acquisition module, configured to acquire at least one auxiliary information of the point cloud media and unassigned general indication information; The at least one auxiliary information includes first auxiliary information and second auxiliary information; the auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information; An information encapsulation module, configured to encode the point cloud media to obtain a point cloud code stream; The information encapsulation module is further used to encapsulate the point cloud code stream, the at least one auxiliary information and the unassigned general indication information to obtain a media file; the media file includes the assigned general indication information, and the assigned general indication information is obtained by assigning a value to the unassigned general indication information based on the at least one auxiliary information; the assigned general indication information is used to indicate the at least one auxiliary information.

30. A media data processing device, characterized in that: include: A file acquisition module is configured to acquire a media file; the media file includes general indication information that has been assigned a value based on at least one type of auxiliary information of the point cloud media; the general indication information that has been assigned a value is used to indicate the at least one type of auxiliary information; the at least one type of auxiliary information includes first auxiliary information and second auxiliary information; the auxiliary information type corresponding to the first auxiliary information is different from the auxiliary information type corresponding to the second auxiliary information; The code stream decoding module is used to decapsulate the media file to obtain a point cloud code stream and the at least one auxiliary information, and decode the point cloud code stream to obtain the point cloud media.

31. A computer device, characterized in that: include: processor, memory, and network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 28.

32. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 28.

Citation Information

Patent Citations

  • Data processing method, device and equipment of point cloud media and readable storage medium

    CN114116617A