A method, apparatus, device, and readable storage medium for processing media data

By determining the saliency information of point cloud media, optimizing the rendering effect in the target range, the problems of low rendering efficiency and poor results are solved, and the presentation quality of point cloud media is improved.

CN115002470BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210586954.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-07-18
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

When rendering point cloud media, the prior art focuses only on the point cloud media itself, resulting in low rendering efficiency and poor effect, affecting the presentation effect.

Method used

By determining the saliency information of point cloud media, including saliency level parameters, optimizing the rendering effect of the target range, encoding the point cloud code stream and encapsulating it into a media file.

Benefits of technology

The rendering efficiency and presentation effect of point cloud media are improved, and the rendering effect in the target range is optimized through remarkable information, which improves the experience quality of immersive media.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115002470B_ABST
    Figure CN115002470B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method, apparatus, device, and readable storage medium for media data processing. The method includes: determining saliency information of point cloud media; the saliency information includes a saliency level parameter for indicating a target range of the point cloud media; the target range includes a spatial range or a temporal range; encoding the point cloud media to obtain a point cloud bitstream, and encapsulating the point cloud bitstream and the saliency information into a media file. By using the present application, the rendering effect of the target range can be determined through the saliency information, and thus the presentation effect of the point cloud media can be optimized. The embodiments of the present invention can be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, device, and readable storage medium for processing media data. Background Art

[0002] Immersive media refers to media content that can bring an immersive experience to a business object, and point cloud media is a typical immersive media.

[0003] In the prior art, a content consumption device first unpacks a point cloud file transmitted by a content production device, then decodes it to obtain point cloud media, and finally renders and presents the point cloud media. However, when rendering the point cloud media, only the point cloud media itself is concerned. Therefore, not only is the rendering efficiency low, but the rendering effect may also be poor, which may further reduce the presentation effect of the point cloud media. Summary of the Invention

[0004] Embodiments of this application provide a method, apparatus, device, and readable storage medium for processing media data, which can determine the rendering effect of a target range through saliency information, and thus can optimize the presentation effect of point cloud media.

[0005] On the one hand, an embodiment of this application provides a method for processing media data, including:

[0006] Determine the saliency information of the point cloud media; the saliency information includes a saliency level parameter for indicating the target range of the point cloud media; the target range includes a spatial range or a temporal range;

[0007] Encode the point cloud media to obtain a point cloud bitstream, and encapsulate the point cloud bitstream and the saliency information into a media file.

[0008] On the one hand, an embodiment of this application provides a method for processing media data, including:

[0009] Obtain a media file, unpack the media file to obtain a point cloud bitstream and the saliency information of the point cloud media; the saliency information includes a saliency level parameter for indicating the target range of the point cloud media; the target range includes a spatial range or a temporal range;

[0010] Decode the point cloud bitstream to obtain the point cloud media.

[0011] On the one hand, an embodiment of this application provides a device for processing media data, including:

[0012] An information determination module, configured to determine the saliency information of the point cloud media; the saliency information includes a saliency level parameter for indicating the target range of the point cloud media; the target range includes a spatial range or a temporal range;

[0013] An information encapsulation module, which is used to encode point cloud media to obtain a point cloud bitstream, and encapsulate the point cloud bitstream and the saliency information into a media file.

[0014] Among them, the media file includes a saliency information data box for indicating saliency information. When the saliency information data box is included at the sample entry of the point cloud track corresponding to the media file, the target range includes a spatial range;

[0015] Among them, a saliency level parameter in the saliency information is used to indicate a spatial range; a spatial range includes a spatial area respectively included in A point cloud frames; the point cloud media includes A point cloud frames; A is a positive integer.

[0016] Among them, the saliency information data box includes a data structure quantity field; the data structure quantity field is used to indicate the total quantity of the saliency information data structures.

[0017] Among them, the value of the data structure quantity field is S, indicating S saliency information data structures; the S saliency information data structures include saliency information data structure B c , where S and c are both positive integers, and c is less than or equal to S;

[0018] Saliency information data structure B c includes a saliency level field with a field value of saliency level parameter D c , and a target range indication field with a field value of a first indication value; the saliency level parameter D c belongs to the saliency level parameter in the saliency information;

[0019] The first indication value indicates the saliency level of the spatial range indicated by the saliency level parameter D c used to indicate a spatial range.

[0020] Among them, the saliency information data structure B c also includes a spatial range indication field;

[0021] When the field value of the spatial range indication field is a second indication value, it means that the spatial range indicated by the saliency level parameter D c is determined by the spatial area identifier;

[0022] When the field value of the spatial range indication field is a third indication value, it means that the spatial range indicated by the saliency level parameter D c is determined by the spatial area position information; the third indication value is different from the second indication value.

[0023] Among them, when the field value of the spatial range indication field is a second indication value, the saliency information data structure B cIt further includes a spatial region identifier field; the spatial region identifier field is used to indicate the saliency level parameter D c The spatial region identifier of the indicated spatial range.

[0024] Wherein, when the field value of the spatial range indication field is the third indication value, the saliency information data structure B c It further includes a spatial region position information field; the spatial region position information field is used to indicate the saliency level parameter D c The spatial region position information of the indicated spatial range.

[0025] Wherein, the saliency information data structure B c It further includes a point cloud patch information field;

[0026] When the field value of the point cloud patch information field is the first information value, it indicates that the spatial range indicated by the saliency level parameter D c Has associated point cloud patches;

[0027] When the field value of the point cloud patch information field is the second information value, it indicates that the spatial range indicated by the saliency level parameter D c Does not have associated point cloud patches; the second information value is different from the first information value;

[0028] Wherein, when the field value of the point cloud patch information field is the first information value, the saliency information data structure B c It further includes a point cloud patch quantity field and a point cloud patch identifier field; the point cloud patch quantity field is used to indicate the total quantity of the associated point cloud patches; the point cloud patch identifier field is used to indicate the point cloud patch identifier corresponding to the associated point cloud patches.

[0029] Wherein, the saliency information data structure B c It further includes a spatial partitioning information field;

[0030] When the field value of the spatial partitioning information field is the third information value, it indicates that the spatial range indicated by the saliency level parameter D c Has associated spatial partitions;

[0031] When the field value of the point cloud patch information field is the fourth information value, it indicates that the spatial range indicated by the saliency level parameter D c Does not have associated spatial partitions; the fourth information value is different from the third information value;

[0032] Wherein, when the field value of the spatial partitioning information field is the third information value, the saliency information data structure B c It further includes a spatial partition quantity field and a spatial partition identifier field; the spatial partition quantity field is used to indicate the total quantity of the associated spatial partitions; the spatial partition identifier field is used to indicate the spatial partition identifier corresponding to the associated spatial partitions.

[0033] Among them, the salience information data box includes a salience algorithm type field;

[0034] When the field value of the salience algorithm type field is the first type value, it indicates that the salience information is determined by the salience detection algorithm;

[0035] When the field value of the salience algorithm type field is the second type value, it indicates that the salience information is determined by data statistics; the second type value is different from the first type value.

[0036] Among them, the media data processing device further includes:

[0037] A file transfer module, configured to transfer a transfer signaling for a media file to a client; the transfer instruction carries salience information description data; the salience information description data is used to instruct the client to determine the acquisition order between different media sub-files in the media file when obtaining the media file through the streaming transfer method; the salience information description data is generated based on the salience information.

[0038] Among them, when the media file includes a salience information metadata track for indicating salience information, the target range includes a time range;

[0039] Among them, the salience information metadata track includes E sample serial numbers associated with the salience information; among them, one sample serial number corresponds to one point cloud frame; the point cloud media includes E point cloud frames; E is a positive integer.

[0040] Among them, the salience information metadata track includes sample serial number F g ; g is a positive integer and g is less than or equal to E; the time range includes the point cloud frame corresponding to sample serial number F g corresponding to;

[0041] The salience information metadata track includes a salience level indication field for sample serial number F g ;

[0042] When the field value of the salience level indication field is the fourth indication value, it indicates that the salience level associated with sample serial number F g is determined by a reference salience level parameter; the reference salience level parameter belongs to the salience level parameter in the salience information;

[0043] When the field value of the salience level indication field is the fifth indication value, it indicates that the salience level associated with sample serial number F g is determined by the salience information data structure; the fifth indication value is different from the fourth indication value.

[0044] Among them, when the field value of the salience level indication field is the fourth indication value, it indicates that the reference salience level parameter is used to indicate the salience level of a time range; a time range is the point cloud frame corresponding to the sample number F g corresponding to;

[0045] The salience information metadata track also includes an effective range indication field for the sample number F g whose field value is the sixth indication value; the sixth indication value indicates that the reference salience level parameter is effective within the salience information metadata track.

[0046] Among them, when the field value of the salience level indication field is the fifth indication value, the salience information metadata track also includes an effective range indication field for the sample number F g ;

[0047] When the field value of the effective range indication field is the sixth indication value, it indicates that the salience level associated with the sample number F g is effective within the salience information metadata track;

[0048] When the field value of the effective range indication field is the seventh indication value, it indicates that the salience level associated with the sample number F g is effective within the point cloud frame corresponding to the sample number F g ; the seventh indication value is different from the sixth indication value.

[0049] Among them, when the field value of the effective range indication field is the seventh indication value, the salience information metadata track also includes a sample salience level field for the sample number F g ; the sample salience level field is used to indicate the salience level parameter of the point cloud frame corresponding to the sample number F g .

[0050] Among them, when the field value of the salience level indication field is the fifth indication value, the salience information metadata track also includes a data structure quantity field with a field value of T for the sample number F g ; the data structure quantity field is used to indicate the total quantity of the salience information data structures; T is a positive integer.

[0051] Among them, the T salience information data structures include the salience information data structure U v , where v is a positive integer and v is less than or equal to T;

[0052] The salience information data structure U v includes a salience level field with a field value of the salience level parameter W v , and a target range indication field; the salience level parameter W v belongs to the salience level parameters in the salience information;

[0053] When the field value of the target range indication field is the first indication value, it represents the saliency level parameter W v Used to indicate the sample serial number F g The saliency level of a spatial region in the corresponding point cloud frame;

[0054] When the field value of the target range indication field is the eighth indication value, it represents the saliency level parameter W v Used to indicate the sample serial number F g The saliency level of the corresponding point cloud frame; the eighth indication value is different from the first indication value.

[0055] Wherein, when T is a positive integer greater than 1, the saliency information data structure includes a saliency level field and a target range indication field with a field value of the first indication value; the field value of the saliency level field belongs to the saliency level parameter in the saliency information; the first indication value indicates that the field value of the saliency level field is used to indicate the sample serial number F g The saliency level of a spatial region in the corresponding point cloud frame.

[0056] Wherein, the sample entry of the saliency information metadata track includes a saliency information data box;

[0057] The saliency information data box includes a saliency algorithm type field; the saliency algorithm type field is used to indicate the determination algorithm type of the saliency information.

[0058] Wherein, when the media file includes Z saliency information sample groups for indicating saliency information, the target range includes a time range;

[0059] Wherein, the total number of different sample serial numbers included in each of the Z saliency information sample groups is less than or equal to H; one sample serial number is used to indicate one point cloud frame; the media file includes H point cloud frames; H is a positive integer, and Z is a positive integer and Z is less than H.

[0060] Wherein, the Z saliency information sample groups include the saliency information sample group K m ; m is a positive integer and m is less than or equal to Z; the time range includes the point cloud frame corresponding to the saliency information sample group K m The point cloud frame corresponding to the saliency information sample group K m The point cloud frame corresponding to the saliency information sample group K belongs to the H point cloud frames;

[0061] The saliency information sample group K m Includes an effective range indication field and a data structure quantity field with a field value of I; the data structure quantity field is used to indicate the total number of saliency information data structures; I is a positive integer; the I saliency information data structures are used to indicate the saliency information sample group K mThe associated significance level;

[0062] When the field value of the effective range indication field is the sixth indication value, it means that the significance level associated with the significance information sample group K m The associated significance level takes effect within the point cloud track corresponding to the media file;

[0063] When the field value of the effective range indication field is the seventh indication value, it means that the significance level associated with the significance information sample group K m The associated significance level takes effect within the significance information sample group K m ; The seventh indication value is different from the sixth indication value.

[0064] Among them, when the field value of the effective range indication field is the seventh indication value, the significance information sample group K m includes a sample significance level field; the sample significance level field is used to indicate the significance level parameter of the point cloud frame corresponding to the significance information sample group K m corresponding to.

[0065] Among them, the I significance information data structures include the significance information data structure J n , n is a positive integer, and n is less than or equal to I;

[0066] The significance information data structure J n includes a significance level field with a field value of the significance level parameter L n , and a target range indication field; the significance level parameter L n belongs to the significance level parameter in the significance information;

[0067] When the field value of the target range indication field is the first indication value, it means that the significance level parameter L n is used to indicate the significance level of a spatial region in the point cloud frame corresponding to the significance information sample group K m corresponding to;

[0068] When the field value of the target range indication field is the eighth indication value, it means that the significance level parameter L n is used to indicate the significance level of the point cloud frame corresponding to the significance information sample group K m ; The eighth indication value is different from the first indication value.

[0069] Among them, the information encapsulation module is specifically used to optimize and encode the target range of the point cloud media according to the significance level parameter in the significance information to obtain a point cloud bitstream.

[0070] Among them, the total number of target ranges is at least two, and the at least two target ranges include a first target range and a second target range; the significance level parameters in the significance information include a first significance level parameter corresponding to the first target range and a second significance level parameter corresponding to the second target range;

[0071] An information encapsulation module, including:

[0072] A level determination unit, configured to determine a first coding level of the first target range according to the first significance level parameter, and determine a second coding level of the second target range according to the second significance level parameter. When the first significance level parameter is greater than the second significance level parameter, the first coding level is superior to the second coding level;

[0073] An optimized coding unit, configured to perform optimized coding on the first target range through the first coding level to obtain a first sub-point cloud bitstream, and perform optimized coding on the second target range through the second coding level to obtain a second sub-point cloud bitstream;

[0074] A bitstream generation unit, configured to generate a point cloud bitstream according to the first sub-point cloud bitstream and the second sub-point cloud bitstream.

[0075] An embodiment of the present application provides a media data processing device on the one hand, including:

[0076] A file acquisition module, configured to acquire a media file, perform de-encapsulation on the media file to obtain a point cloud bitstream and the significance information of the point cloud media; the significance information includes a significance level parameter for indicating the target range of the point cloud media; the target range includes a spatial range or a time range;

[0077] A bitstream decoding module, configured to decode the point cloud bitstream to obtain the point cloud media.

[0078] Among them, the media data processing device further includes:

[0079] A first determination module, configured to determine that the target range includes a spatial range when the media file includes a significance information data box for indicating the significance information; one significance level parameter in the significance information is used to indicate one spatial range; one spatial range includes a spatial region respectively included in A point cloud frames; the point cloud media includes A point cloud frames; A is a positive integer.

[0080] Among them, the total number of spatial ranges is at least two, and the at least two spatial ranges include a first spatial range and a second spatial range;

[0081] The media data processing device further includes:

[0082] The first acquisition module is used to acquire, from the saliency information, the saliency level parameter O indicating the first spatial range p and acquire the saliency level parameter O indicating the second spatial range p+1 ; p is a positive integer and p is less than the total number of saliency level parameters in the saliency information;

[0083] The second determination module is used to, if the saliency level parameter O p is greater than the saliency level parameter O p+1 , determine that the rendering level corresponding to the first spatial range is better than the rendering level corresponding to the second spatial range;

[0084] The second determination module is further used to, if the saliency level parameter O p is less than the saliency level parameter O p+1 , determine that the rendering level corresponding to the second spatial range is better than the rendering level corresponding to the first spatial range.

[0085] Wherein, the media data processing device further includes:

[0086] The third determination module is used to determine that the target range includes a time range when the media file includes a saliency information metadata track indicating the saliency information;

[0087] Wherein, the saliency information metadata track includes E sample serial numbers associated with the saliency information; wherein, one sample serial number corresponds to one point cloud frame; the point cloud media includes E point cloud frames; E is a positive integer.

[0088] Wherein, the time range includes the first point cloud frame and the second point cloud frame among the E point cloud frames;

[0089] The media data processing device further includes:

[0090] The second acquisition module is used to acquire, from the saliency information, the saliency level parameter Q indicating the first point cloud frame r and acquire the saliency level parameter Q indicating the second point cloud frame r+1 ; r is a positive integer and r is less than the total number of saliency level parameters in the saliency information;

[0091] The fourth determination module is used to, if the saliency level parameter Q r is greater than the saliency level parameter Q r+1 , determine that the rendering level corresponding to the first point cloud frame is better than the rendering level corresponding to the second point cloud frame;

[0092] The fourth determination module is further used to, if the saliency level parameter Q r is less than the saliency level parameter Q r+1, the rendering level corresponding to the second point cloud frame is determined to be better than the rendering level corresponding to the first point cloud frame.

[0093] The media data processing device further includes:

[0094] A third acquisition module, configured to, if the first point cloud frame includes at least two spatial regions and the saliency levels corresponding to the at least two spatial regions are different, acquire, in the saliency information, the saliency level parameter X corresponding to the first spatial region y , and acquire the saliency level parameter X corresponding to the second spatial region y+1 ; both the first spatial region and the second spatial region belong to the at least two spatial regions; x is a positive integer and x is less than the total number of saliency level parameters in the saliency information;

[0095] A fifth determination module, configured to, if the saliency level parameter X y is greater than the saliency level parameter X y+1 , determine that the rendering level corresponding to the first spatial region is better than the rendering level corresponding to the second spatial region;

[0096] The fifth determination module is further configured to, if the saliency level parameter X y is less than the saliency level parameter X y+1 , determine that the rendering level corresponding to the second spatial region is better than the rendering level corresponding to the first spatial region.

[0097] On the one hand, the present application provides a computer device, including: a processor, a memory, and a network interface;

[0098] The above-mentioned processor is connected to the above-mentioned memory and the above-mentioned network interface. Among them, the above-mentioned network interface is used to provide a data communication function, the above-mentioned memory is used to store a computer program, and the above-mentioned processor is used to call the above-mentioned computer program so that the computer device executes the method in the embodiment of the present application.

[0099] On the one hand, the embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and the computer program is suitable for being loaded and executed by a processor to execute the method in the embodiment of the present application.

[0100] On the one hand, the embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium; a processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method in the embodiment of the present application.

[0101] In the embodiments of the present application, the saliency information of the point cloud media is first determined, where the saliency information includes a saliency level parameter for indicating the target range of the point cloud media; the target range includes a spatial range or a time range; further, by encoding the point cloud media, a point cloud bitstream can be obtained; further, the point cloud bitstream and the saliency information are encapsulated into a media file; where the saliency information in the media file can be used to determine the rendering effect of the target range when rendering the point cloud media. As can be seen from the above, the embodiments of the present application can determine the saliency information corresponding to the point cloud media for indicating the time range, or the saliency information corresponding to the point cloud media for indicating the spatial range. Therefore, the embodiments of the present application can encapsulate the saliency information of the point cloud media into the media file, and further, when rendering the point cloud media, the rendering effect of the target range can be determined through the saliency information, so the presentation effect of the point cloud media can be optimized. Description of the Drawings

[0102] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0103] Figure 1a It is a schematic diagram of 3DoF provided by the embodiments of the present application;

[0104] Figure 1b It is a schematic diagram of 3DoF+ provided by the embodiments of the present application;

[0105] Figure 1c It is a schematic diagram of 6DoF provided by the embodiments of the present application;

[0106] Figure 2 It is a schematic diagram of the process of an immersive media from collection to consumption provided by the embodiments of the present application;

[0107] Figure 3 It is a schematic diagram of the architecture of an immersive media system provided by the embodiments of the present application;

[0108] Figure 4 It is a schematic diagram of the first process of a media data processing method provided by the embodiments of the present application;

[0109] Figure 5 It is a schematic diagram of a saliency level parameter that is only related to time and not related to the spatial region provided by the embodiments of the present application;

[0110] Figure 6 It is a schematic diagram of a saliency level parameter that is related to the spatial region and the saliency level parameter of the spatial region changes with time provided by the embodiments of the present application;

[0111] Figure 7 is a flowchart of a media data processing method provided by an embodiment of the present application Figure Two ;

[0112] Figure 8 is a schematic diagram in which a saliency level parameter is related to a spatial region and the saliency level parameter related to the spatial region does not change with time provided by an embodiment of the present application;

[0113] Figure 9 is a flowchart of a media data processing method provided by an embodiment of the present application Figure Three ;

[0114] Figure 10 is a first schematic structural diagram of a media data processing device provided by an embodiment of the present application;

[0115] Figure 11 is a schematic structure of a media data processing device provided by an embodiment of the present application Figure Two ;

[0116] Figure 12 is a schematic structural diagram of a computer device provided by an embodiment of the present application;

[0117] Figure 13 is a schematic structural diagram of a data processing system provided by an embodiment of the present application. Detailed implementation manners

[0118] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0119] Next, some technical terms related to the embodiments of the present application will be introduced:

[0120] I. Immersive media:

[0121] Immersive media refers to media content that can provide an immersive experience, enabling business objects immersed in such media content to obtain sensory experiences such as vision and hearing in the real world. Immersive media can be classified into 3DoF media, 3DoF+ media, and 6DoF media according to the degree of freedom (DoF) of business objects when consuming media content. Among them, point cloud media is a typical 6DoF media. In the embodiments of this application, users (i.e., viewers) who consume immersive media (such as point cloud media) are collectively referred to as business objects.

[0122] II. Point Cloud:

[0123] A point cloud is a set of discrete points that are irregularly distributed in space and represent the spatial structure and surface attributes of a three-dimensional object or scene. Each point in the point cloud has at least three-dimensional position information, and may also have color, material, or other information depending on the application scenario. Usually, each point in the point cloud has the same number of additional attributes.

[0124] Point clouds can flexibly and conveniently represent the spatial structure and surface attributes of three-dimensional objects or scenes, and thus are widely used, including virtual reality (VR) games, computer-aided design (CAD), geographic information system (GIS), autonomous navigation system (ANS), digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive telepresence, three-dimensional reconstruction of biological tissues and organs, etc.

[0125] The main ways to obtain point clouds are as follows: computer generation, 3D (three-dimension) laser scanning, 3D photogrammetry, etc. Computers can generate point clouds of virtual three-dimensional objects and scenes. 3D scanning can obtain point clouds of static real-world three-dimensional objects or scenes, and can acquire millions of point clouds per second. 3D photography can obtain point clouds of dynamic real-world three-dimensional objects or scenes, and can acquire tens of millions of point clouds per second. In addition, in the medical field, point clouds of biological tissues and organs can be obtained from MRI (magnetic resonance imaging), CT (computed tomography), and electromagnetic positioning information. These technologies have reduced the cost and time cycle of point cloud data acquisition and improved the accuracy of the data. The transformation of point cloud data acquisition methods has made it possible to obtain a large amount of point cloud data. With the continuous accumulation of large-scale point cloud data, the efficient storage, transmission, publishing, sharing, and standardization of point cloud data have become the key to point cloud applications.

[0126] 3. Track:

[0127] Track is a collection of media data in the process of media file encapsulation, which is composed of multiple time-sequence samples. A media file can be composed of one or more tracks. For example, a common media file can contain a video media track, an audio media track and a subtitle media track. In particular, metadata information can also be included in the file as a media type in the form of metadata media track, which is referred to as metadata track in the present invention.

[0128] 4. Sample:

[0129] A sample is a unit of media file encapsulation. A track consists of many samples, each of which corresponds to a specific timestamp. For example, a video media track can consist of many samples, and a sample is usually a video frame. In the embodiment of the present application, a sample in a point cloud media track can be a point cloud frame.

[0130] 5. Sample Entry:

[0131] The sample entry is used to indicate metadata information related to all samples in the track. For example, the sample entry of a video track usually contains metadata information related to decoder initialization.

[0132] 6. Slice:

[0133] A point cloud slice (point cloud strip) represents a collection of syntax elements (such as geometry slices and attribute slices) of part or all of the encoded point cloud frame data.

[0134] 7. Space Block Area (Tile):

[0135] The hexahedral spatial block area within the boundary spatial area of the point cloud frame is referred to as the spatial block in this application. A spatial block consists of one or more point cloud slices, and there is no encoding and decoding dependency between the spatial blocks.

[0136] 8. Saliency and Saliency Detection

[0137] Since the human visual system can naturally determine the most obvious and prominent area in the current scene during a quick scan, people will naturally be attracted to the prominent part of the picture content when viewing a picture. For a picture, the part that attracts the observer's attention the most is the salient area of the picture. For a picture, different areas have different appeal to the observer, and the concept that characterizes the degree of attraction of these areas to the observer is called saliency or significance.

[0138] Salience detection, which characterizes the degree of visual attraction, is of great significance for both image processing and analysis. Human visual attention is an important mechanism in the process of humans obtaining and processing external information. It enables people to quickly screen important information and endows perception with selectivity. When combined with image processing, this mechanism can greatly improve the efficiency of existing image processing and analysis methods. Salience detection was proposed and developed based on this. Salience detection is a technical means for preprocessing images in the field of computer vision, used to find or identify salient objects in images, usually the image areas that are likely to catch people's eyes.

[0139] With the development of salience detection models, salience detection has also begun to be applied to other fields of image processing, such as image segmentation, image compression, and image recognition. The practicality of salience detection has gradually been recognized by people in continuous applications. With the rise of immersive media, the development of salience detection models has also faced new challenges. For example, for point cloud media, its point cloud frames have gone beyond the scope of images. How to perform salience detection and salience grading on points in three-dimensional space has become the latest research direction in the current immersive media field.

[0140] IX. DoF (Degree of Freedom):

[0141] In this application, DoF refers to the degrees of freedom that a service object supports for movement and content interaction when viewing immersive media (such as point cloud media), which can include 3DoF (three degrees of freedom), 3DoF+, and 6DoF (six degrees of freedom). Among them, 3DoF refers to the three degrees of freedom for the service object's head to rotate around the x-axis, y-axis, and z-axis. 3DoF+ means that on the basis of three degrees of freedom, the service object also has the degrees of freedom for limited movement along the x-axis, y-axis, and z-axis. 6DoF means that on the basis of three degrees of freedom, the service object also has the degrees of freedom for free movement along the x-axis, y-axis, and z-axis.

[0142] X. ISOBMFF (ISO Based Media File Format):

[0143] A media file format based on the ISO (International Standard Organization) standard, which is a packaging standard for media files. A relatively typical ISOBMFF file is the MP4 (Moving Picture Experts Group 4) file.

[0144] XI. DASH (Dynamic Adaptive Streaming over HTTP): An adaptive bitrate technology that enables high-quality streaming media to be delivered over the Internet via a traditional HTTP web server.

[0145] XII. MPD (Media Presentation Description, the media presentation description signaling in DASH), which is used to describe the media segment information in a media file.

[0146] XIII. Representation: It refers to the combination of one or more media components in DASH. For example, a video file with a certain resolution can be regarded as a Representation.

[0147] XIV. Adaptation Sets: It refers to the set of one or more video streams in DASH. An Adaptation Sets can contain multiple Representations.

[0148] XV. Media Segment: A playable segment that conforms to a certain media format. When playing, it may need to cooperate with zero or more previous segments and the Initialization Segment.

[0149] The embodiments of the present application relate to data processing technologies for immersive media. Some concepts in the data processing process of immersive media will be introduced below. In particular, in the subsequent embodiments of the present application, point cloud media is taken as an example of immersive media for illustration.

[0150] Please refer to Figure 1a , Figure 1a which is a schematic diagram of 3DoF provided by the embodiments of the present application. As Figure 1a shown, 3DoF means that the business object consuming immersive media has a fixed center point in a three-dimensional space, and the head of the business object rotates along the X-axis, Y-axis, and Z-axis to view the picture provided by the media content.

[0151] Please refer to Figure 1b , Figure 1b which is a schematic diagram of 3DoF+ provided by the embodiments of the present application. As Figure 1b shown, 3DoF+ means that when the virtual scene provided by the immersive media has certain depth information, the head of the business object can move within a limited space based on 3DoF to view the picture provided by the media content.

[0152] Please refer to Figure 1c ,Figure 1c is a schematic diagram of 6DoF provided by an embodiment of the present application. As Figure 1c shown, 6DoF is divided into window 6DoF, omnidirectional 6DoF, and 6DoF. Among them, window 6DoF means that the rotation and movement of the service object on the X-axis and Y-axis are restricted, and the translation on the Z-axis is restricted; for example, the service object cannot see the scene outside the window frame, and the service object cannot pass through the window. Omnidirectional 6DoF means that the rotation and movement of the service object on the X-axis, Y-axis, and Z-axis are restricted. For example, the service object cannot freely pass through the three-dimensional 360-degree VR content in the restricted movement area. 6DoF means that based on 3DoF, the service object can freely translate along the X-axis, Y-axis, and Z-axis. For example, the service object can freely walk in the three-dimensional 360-degree VR content.

[0153] Please refer to Figure 2 , Figure 2 is a schematic diagram of the process of an immersive media from collection to consumption provided by an embodiment of the present application. As Figure 2 shown, the complete processing process for immersive media may specifically include: video capture, video encoding, video file encapsulation, video file transmission, video file decapsulation, video decoding, and finally video presentation.

[0154] Among them, video capture is used to convert analog video into digital video and save it in the format of a digital video file. That is to say, video capture can convert the video signals (such as point cloud data) collected by multiple cameras from different angles into binary digital information. Among them, the binary digital information converted from the video signal is a binary data stream, and this binary digital information can also be called the bitstream or bit stream of the video signal. Video encoding refers to converting a file in the original video format into another video format file through compression technology. From the perspective of the acquisition method of video signals, video signals can be divided into two methods: captured by a camera and generated by a computer. Due to different statistical characteristics, the corresponding compression encoding methods may also be different. Commonly used compression encoding methods may specifically include HEVC (High Efficiency Video Coding, the international video coding standard HEVC / H.265), VVC (Versatile Video Coding, the international video coding standard VVC / H.266), AVS (Audio Video Coding Standard, the national video coding standard of China), AVS3 (the third-generation video coding standard launched by the AVS standard group), etc.

[0155] After video encoding, it is necessary to encapsulate the encoded data stream (e.g., point cloud bitstream) and transmit it to the service object. Video file encapsulation means storing the already encoded and compressed video bitstream and audio bitstream in a file according to a certain format according to the encapsulation format (or container, or file container). Common encapsulation formats include the AVI format (AudioVideo Interleaved) or the ISOBMFF format. In one embodiment, the audio bitstream and the video bitstream are encapsulated in a file container according to a file format such as ISOBMFF to form a media file (which can also be called an encapsulated file, a video file). The media file can be composed of multiple tracks, for example, it can include a video track, an audio track, and a subtitle track.

[0156] After the content production device executes the above encoding process and file encapsulation process, it can transmit the media file to the client on the content consumption device. After the client performs reverse operations such as de-encapsulation and decoding, the final video content can be presented on the client. Among them, the media file can be sent to the client based on various transport protocols, and the transport protocols here can include but are not limited to: DASH protocol, HLS (HTTP Live Streaming) protocol, SMTP (SmartMedia Transport Protocol), TCP (Transmission Control Protocol), etc.

[0157] It can be understood that the process of file de-encapsulation on the client is the reverse of the above file encapsulation process. The client can de-encapsulate the media file according to the file format requirements during encapsulation to obtain the audio bitstream and the video bitstream. The decoding process on the client is also the reverse of the encoding process. For example, the client can decode the video bitstream to restore the video content, and can also decode the audio bitstream to restore the audio content.

[0158] For ease of understanding, please also refer to Figure 3 , Figure 3 which is a schematic diagram of the architecture of an immersive media system provided by an embodiment of the present application. As Figure 3As shown, the immersive media system may include a content production device (e.g., content production device 200A) and a content consumption device (e.g., content consumption device 200B). The content production device may refer to the computer device used by the provider of point cloud media (e.g., the content producer of point cloud media). This computer device may be a terminal (such as a PC (Personal Computer), a smart mobile device (such as a smart phone), etc.) or a server. Among them, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0159] The content consumption device may refer to the computer device used by the user of point cloud media (e.g., the viewer of point cloud media, that is, the business object). This computer device may be a terminal (such as a PC (Personal Computer), a smart mobile device (such as a smart phone), a VR device (such as a VR helmet, VR glasses, etc.), smart home appliances, in-vehicle terminals, aircraft, etc.). This computer device is integrated with a client. The content production device and the content consumption device may be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0160] Among them, the above client may be a client with functions of displaying data information such as text, images, audio, and video, including but not limited to multimedia clients (such as video clients), social clients (such as instant messaging clients), information applications (such as news clients), entertainment clients (such as game clients), shopping clients, in-vehicle clients, browsers, etc. Among them, this client may be an independent client or an embedded sub-client integrated in a certain client (such as a social client), and no limitation is made here.

[0161] It can be understood that the data processing technology related to immersive media in this application can be realized relying on cloud technology; for example, using a cloud server as the content production device. Cloud technology is a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing.

[0162] The data processing process of point cloud media includes the data processing process on the content production device side and the data processing process on the content consumption device side.

[0163] The data processing process on the content production device side mainly includes: (1) the process of acquiring and producing the media content of the point cloud media; (2) the process of encoding and file encapsulation of the point cloud media. The data processing process on the content consumption device side mainly includes: (1) the process of file decapsulation and decoding of the point cloud media; (2) the rendering process of the point cloud media. In addition, there is a transmission process of the point cloud media between the content production device and the content consumption device, and this transmission process can be carried out based on various transmission protocols. The transmission protocols here may include but are not limited to: DASH protocol, HLS protocol, SMT protocol, TCP protocol, etc.

[0164] The following will be combined with Figure 3 , and each process involved in the data processing process of the point cloud media will be briefly introduced separately.

[0165] I. The data processing process on the content production device side:

[0166] (1) The process of acquiring and producing the media content of the point cloud media.

[0167] 1) The process of acquiring the media content of the point cloud media.

[0168] The media content of the point cloud media is obtained by a capture device collecting the real-world audio-visual scene. In one implementation, the capture device may refer to a hardware component provided in the content production device. For example, the capture device refers to the microphone, camera, sensor, etc. of the terminal. In another implementation, the capture device may also be a hardware device connected to the content production device. For example, a camera connected to a server is used to provide the acquisition service of the media content of the point cloud media for the content production device. The capture device may include but is not limited to: audio devices, imaging devices, and sensing devices. Among them, the audio devices may include audio sensors, microphones, etc. The imaging devices may include ordinary cameras, stereo cameras, light field cameras, etc. The sensing devices may include laser devices, radar devices, etc. The number of capture devices may be multiple, and these capture devices are deployed at some specific positions in the real space to simultaneously capture the audio content and video content at different angles within the space. The captured audio content and video content are synchronized both in time and space. In the embodiments of the present application, the media content of the three-dimensional space for providing a multi-degree-of-freedom (such as 6DoF) viewing experience collected by the capture devices deployed at specific positions may be referred to as point cloud media.

[0169] For example, taking the acquisition of the video content of the point cloud media as an example for illustration, such as Figure 3As shown, the visual scene 20A (such as a real-world visual scene) can be captured by a set of camera arrays connected to the content production device 200A, or can be captured by a camera device with multiple cameras and sensors connected to the content production device 200A. The acquisition result can be the source point cloud data 20B (i.e., the video content of the point cloud media).

[0170] 2) The production process of the media content of the point cloud media.

[0171] It should be understood that the production process of the media content of the point cloud media involved in the embodiments of the present application can be understood as the content production process of the point cloud media, and the content production of the point cloud media here is mainly made from content in the form of point cloud data captured by cameras or camera arrays deployed at multiple locations. For example, the content production device can convert the point cloud media from a three-dimensional representation to a two-dimensional representation.

[0172] In addition, it should be noted that since the capture device can capture panoramic videos, after such videos are processed by the content production device and transmitted to the content consumption device for corresponding data processing, the business object on the content consumption device side needs to perform some specific actions (such as head rotation) to view the 360-degree video information, while performing non-specific actions (such as moving the head) cannot obtain corresponding video changes, and the VR experience is poor. Therefore, it is necessary to additionally provide depth information matching the panoramic video to enable the business object to obtain a better immersion and a better VR experience, which involves 6DoF production technology. When the business object can move relatively freely in the simulated scene, it is called 6DoF. When using 6DoF production technology to produce the video content of the point cloud media, the capture device generally selects laser devices, radar devices, etc. to capture the point cloud data in space.

[0173] (2) The encoding and file encapsulation process of the point cloud media.

[0174] The captured audio content can be directly subjected to audio encoding to form an audio bitstream of the point cloud media. The captured video content can be subjected to video encoding to obtain a video bitstream of the point cloud media. It should be noted here that if 6DoF production technology is adopted, a specific encoding method (such as point cloud compression based on traditional video encoding) needs to be used during video encoding. The content production device encapsulates the audio bitstream and the video bitstream in a file container according to the file format of the point cloud media (such as ISOBMFF) to form a media file resource of the point cloud media. This media file resource can be a media file or a media file of the point cloud media formed by media segments; and according to the requirements of the file format of the point cloud media, media presentation description information (i.e., MPD) is used to record the metadata of this media file resource of the point cloud media. The metadata here is a general term for information related to the presentation of the point cloud media. This metadata can include description information of the media content, description information of the viewport, and signaling information related to the presentation of the media content, etc. It can be understood that the content production device will store the media presentation description information and the media file resource formed after the data processing process.

[0175] As Figure 3 shown, the content production device 200A performs point cloud media encoding on one or more data frames in the source point cloud data 20B. For example, geometry-based point cloud compression (G-PCC, where PCC is point cloud compression) is adopted to obtain the encoded point cloud bitstream 20E (i.e., the video bitstream, such as the G-PCC bitstream). Subsequently, the content production device 200A can encapsulate one or more encoded bitstreams into a media file 20F for local playback or a sequence of segments 20F for streaming transmission according to a specific media file format (such as ISOBMFF). s . In addition, the file encapsulator in the content production device 200A can also add relevant metadata to the media file 20F or the sequence of segments 20F s . Further, the content production device 200A can use a certain transmission mechanism (such as DASH, SMT) to transmit the sequence of segments 20F s to the content consumption device 200B, or transmit the media file 20F to the content consumption device 200B. In some embodiments, the content consumption device 200B can be a player.

[0176] II. Data processing process on the content consumption device side:

[0177] (3) Process of file de-encapsulation and decoding of point cloud media.

[0178] The content consumption device can adaptively and dynamically obtain the media file resources of the point cloud media and the corresponding media presentation description information from the content production device through the recommendation of the content production device or according to the requirements of the business object on the content consumption device side. For example, the content consumption device can determine the viewing direction and viewing position of the business object based on the position information of the head / eyes of the business object, and then request the corresponding media file resources from the content production device dynamically based on the determined viewing direction and viewing position. Through a transmission mechanism (such as DASH, SMT), the media file resources and the media presentation description information are transmitted from the content production device to the content consumption device. The process of file demultiplexing on the content consumption device side is the reverse of the file multiplexing process on the content production device side. The content consumption device demultiplexes the media file resources according to the file format requirements of the point cloud media (for example, ISOBMFF) to obtain the audio bitstream and the video bitstream. The decoding process on the content consumption device side is the reverse of the encoding process on the content production device side. The content consumption device decodes the audio bitstream to restore the audio content; the content consumption device decodes the video bitstream to restore the video content.

[0179] For example, as Figure 3 shown, the media file 20F output by the file multiplexer in the content production device 200A is the same as the media file 20F' input to the file demultiplexer in the content consumption device 200B. The file demultiplexer performs file demultiplexing processing on the media file 20F' or the received segment sequence 20F' s and extracts the encoded point cloud bitstream 20E', and at the same time parses the corresponding metadata. Subsequently, the point cloud bitstream 20E' can be decoded for point cloud media to obtain the decoded video signal 20D', and the point cloud data (i.e., the restored video content) can be generated from the video signal 20D'. Among them, the media file 20F and the media file 20F' may include track format definitions, which may contain constraints on the elementary streams contained in the samples in the track.

[0180] (4) The rendering process of the point cloud media.

[0181] The content consumption device renders the audio content obtained by audio decoding and the video content obtained by video decoding according to the rendering-related metadata in the media presentation description information corresponding to the media file resources. After the rendering is completed, the playback output of the content is realized.

[0182] The immersive media system supports data boxes, which refer to data chunks or objects including metadata, that is, the data boxes contain the metadata of the corresponding media content. In practical applications, content production devices can use data boxes to guide content consumption devices to consume media files of point cloud media. Point cloud media can include multiple data boxes, such as an ISOBMFF data box (ISO Base Media File Format Box, simply referred to as ISOBMFF Box), which contains metadata for describing the corresponding information during file encapsulation. In the embodiments of this application, the ISOBMFF data box includes metadata for indicating the saliency information of point cloud media.

[0183] As can be seen from the above, content consumption devices can dynamically obtain the media file resources corresponding to point cloud media from the content production device side. Since the media file resources are obtained after the content production device encodes and encapsulates the captured audio and video content, after the content consumption device receives the media file resources returned by the content production device, it needs to first unpack the media file resources to obtain the corresponding audio and video bitstreams, and then decode the audio and video bitstreams. Finally, the decoded audio and video content can be presented to the service object. The point cloud media here can include, but is not limited to, VPCC (Video-based Point Cloud Compression) point cloud media and GPCC (Geometry-based Point Cloud Compression) point cloud media.

[0184] It can be understood that saliency is of great significance to both the processing and analysis of images, and can greatly improve the efficiency of image processing and analysis. Saliency information may include a saliency level for indicating a region in an image. In some point cloud media, although the saliency level corresponding to each point cloud frame changes, the saliency level of the entire point cloud frame does not change. At this time, the saliency information does not need to be associated with the spatial region of the point cloud frame. In some point cloud media, there are spatial regions with different saliency levels within a point cloud frame, but the saliency level does not change over time, that is, the spatial regions corresponding to each frame have the same saliency level. For example, point cloud media 100 includes 100 point cloud frames, and each point cloud frame can be divided into 2 spatial regions. Among them, the saliency level parameters corresponding to the first spatial region in the first point cloud frame, the first spatial region in the second point cloud frame,..., and the first spatial region in the 100th point cloud frame are all 2, and the saliency level parameters corresponding to the second spatial region in the first point cloud frame, the second spatial region in the second point cloud frame,..., and the second spatial region in the 100th point cloud frame are all 1. At this time, the saliency level does not need to be associated with time, but only needs to be associated with the spatial range. In some point cloud media, as time and space change, the saliency level also changes. At this time, the saliency level needs to be associated with a certain spatial region of a certain point cloud frame.

[0185] Based on the above, according to whether the saliency information changes with space and time, an embodiment of the present application proposes a saliency information indicating different ranges (including spatial range and time range), so as to improve the accuracy of the saliency information. Furthermore, when applying point cloud media, more scenarios can be satisfied through the highly accurate saliency information, such as the encoding scenario, transmission scenario, and rendering scenario of point cloud media.

[0186] Specifically, after obtaining the point cloud media, the content production device can determine the saliency information of the point cloud media, and the saliency information includes a saliency level parameter for indicating a target range (time range or spatial range); encode the point cloud media to obtain a point cloud bitstream, and encapsulate the point cloud bitstream and the saliency information into a media file. In the embodiment of the present application, the saliency information may include one or more saliency level parameters, and the total number of saliency level parameters is not limited here. The total number of saliency level parameters is determined according to the actual point cloud media applied.

[0187] It can be understood that in the embodiment of the present application, the saliency information in the media file can indicate the saliency level parameter corresponding to the target range. Therefore, subsequently, the content consumption device can determine the rendering effect of the target range in the scenarios of rendering and presenting the point cloud media, so as to optimize the presentation effect of the point cloud media.

[0188] It should be understood that the method provided in the embodiments of the present application can be applied to the server side (i.e., the content production device side), the player side (i.e., the content consumption device side), and intermediate nodes (such as SMT (Smart Media Transport) receiving entities, SMT sending entities), etc. in an immersive media system. Among them, for the specific process in which the content production device determines the saliency information of the point cloud media, encodes the point cloud media to obtain a point cloud bitstream, and encapsulates the point cloud bitstream and the saliency information into a media file, and for the specific process in which the content consumption device determines the rendering effect of the target range when rendering the point cloud media based on the saliency information in the media file, reference can be made to the following Figures 4 - 9 description of the corresponding embodiments.

[0189] Furthermore, please refer to Figure 4 , Figure 4 FIG. 1 is a first schematic flowchart of a media data processing method provided in an embodiment of the present application. This method can be executed by a content production device in an immersive media system (such as the content production device 200A in the corresponding embodiment above Figure 3 ). For example, this content production device can be a server, and the embodiments of the present application are described by taking the server execution as an example. This method can at least include the following steps S101 to S102.

[0190] Step S101, determining the saliency information of the point cloud media; the saliency information includes a saliency level parameter for indicating the target range of the point cloud media; the target range includes a spatial range or a time range.

[0191] Specifically, for the acquisition process and production process of the point cloud media, reference can be made to the description above Figure 3 . Details are not described herein. The saliency information in the embodiments of the present application can include two types of information. One type of information is saliency algorithm information, that is, information for obtaining the saliency level parameter. The embodiments of the present application do not limit the algorithm for obtaining the saliency level parameter, which can be set according to the actual application scenario; the other type of information is the saliency level parameter for indicating the target range.

[0192] It can be understood that for different point cloud media, the ranges to which the corresponding saliency level parameters need to be associated are different. For example, for some point cloud media, the saliency level parameters only change with time, and at this time, the saliency level parameters do not need to be associated with a spatial region; for some point cloud media, the saliency level parameters only change with space, and at this time, the saliency level parameters do not need to be associated with a point cloud frame; for some point cloud media, the saliency level parameters change with both time and space. Therefore, for immersive media, especially point cloud media, the embodiments of the present application propose a saliency information indication method, which, at the file encapsulation level and the signaling transmission level, through

[0193] 1. Define saliency information for indicating different ranges (including spatial and temporal ranges).

[0194] 2. Define algorithms for obtaining different types of saliency information.

[0195] 3. Associate saliency information with spatial information at the signaling transmission level.

[0196] It is possible to more flexibly determine the saliency information of different spatial regions and different point cloud frames in the point cloud media, thereby meeting more point cloud application scenarios, enabling the server to perform encoding optimization based on the saliency information, and enabling the client to perform transmission optimization based on the saliency information.

[0197] The point cloud media may include a first point cloud media, and the first point cloud media includes E point cloud frames, where E is a positive integer. Among them, in the point cloud track corresponding to the point cloud media, one point cloud frame can be called a sample. Please refer to Figure 5 , Figure 5 which is a schematic diagram provided by an embodiment of the present application where the saliency level parameter is only related to time and not related to the spatial region. As Figure 5 shown, sample 1 refers to the sample (i.e., point cloud frame) with the serial number 1 in the first point cloud media sorted by time, sample 2 refers to the sample with the serial number 2 in the first point cloud media sorted by time, sample 3 refers to the sample with the serial number 3 in the first point cloud media sorted by time,..., and sample E refers to the sample with the serial number E in the first point cloud media sorted by time. According to the content of the first point cloud media, the server can define the saliency information of each region of the first point cloud media. Among them, the overall saliency level of sample 1 is the reference saliency level parameter. It can be understood that the immersive media system can preset the reference saliency level parameter, and the embodiment of the present application does not limit the value of the reference saliency level parameter, which can be set according to the actual application scenario. The overall saliency level parameter of sample 2 is 2, the overall saliency level parameter of sample 3 is 1,..., and the overall saliency level parameter of sample E is the reference saliency level parameter.

[0198] Among them, the larger the saliency level parameter, the higher the saliency. Therefore, in the first point cloud media, sample 2 is the sample (i.e., point cloud frame) with the highest saliency, sample 3 is the sample with the second highest saliency, and the remaining samples only have the reference saliency (assuming the reference saliency level parameter is 0). Obviously, in the first point cloud media, the saliency level parameter is only related to time and not related to the spatial region. Therefore, Figure 5If it is a saliency level parameter at the point cloud frame level, the server may generate a saliency information metadata track to indicate the saliency information, or may generate a saliency information sample group to indicate the above-mentioned saliency information. For the specific process of indicating the saliency information, please refer to the description of the saliency information data box and the saliency information metadata track in step S102 below.

[0199] The point cloud media may include a second point cloud media, and the second point cloud media includes one or more point cloud frames (for example, E point cloud frames). Please refer to Figure 6 , Figure 6 is a schematic diagram provided by an embodiment of the present application, in which the saliency level parameter is related to a spatial region, and the saliency level parameter of the spatial region changes over time. Among them, Figure 6 For the meaning of samples 1 - E, please refer to Figure 5 the interpretation in. Here, it will not be elaborated. According to the content of the second point cloud media, the server may define the saliency information of each region of the second point cloud media. Among them, the saliency levels corresponding to samples 1, 2, 3, 4,..., E are not completely the same, that is, the saliency level parameter of the second point cloud media changes over time, and there are some point cloud frames in which the internal spatial regions also correspond to different saliency level parameters, that is, the saliency level parameter also changes in the spatial region.

[0200] Such as Figure 6 shown, the overall saliency level parameter of sample 1 is the reference saliency level parameter. Among them, the meaning of the reference saliency level parameter is as described in Figure 5 . The spatial region corresponding to sample 1 can be divided into two spatial regions. Among them, the saliency level parameter corresponding to the first spatial region (abbreviated as spatial region 1) in sample 1 is 2, which is composed of two point cloud patches, and the point cloud patch identifiers corresponding to the two point cloud patches include point cloud patch 0 and point cloud patch 1; the saliency level parameter corresponding to the second spatial region (abbreviated as spatial region 2) in sample 1 is 1, which is composed of two point cloud patches, and the point cloud patch identifiers corresponding to the two point cloud patches include point cloud patch 2 and point cloud patch 3; in sample 1, the saliency of the first spatial region is higher than that of the second spatial region.

[0201] Such as Figure 6As shown, the significance level parameter of the overall Sample 2 is 2, and the internal space of Sample 2 is not divided. The significance level parameter of the overall Sample 3 is the reference significance level parameter, and the corresponding spatial region of Sample 3 can be divided into two spatial regions. Among them, the significance level parameter corresponding to the first spatial region (referred to as Spatial Region 1) in Sample 3 is 0, which consists of two point cloud patches, and the point cloud patch identifiers corresponding to the two point cloud patches include Point Cloud Patch 0 and Point Cloud Patch 1; the significance level parameter corresponding to the second spatial region (referred to as Spatial Region 2) in Sample 3 is 1, which consists of two point cloud patches, and the point cloud patch identifiers corresponding to the two point cloud patches include Point Cloud Patch 2 and Point Cloud Patch 3; in Sample 3, the significance of the first spatial region is lower than that of the second spatial region.

[0202] As Figure 6 shown, the significance level parameters corresponding to Sample 4, …, Sample E are all the reference significance level parameters. Therefore, in the second point cloud media, Sample 2 is the sample (i.e., the point cloud frame) with the highest significance, there are a spatial region with the highest significance and a spatial region with the second highest significance in Sample 1, there is a spatial region with the second highest significance in Sample 3, and the remaining samples only have the reference significance level parameter (assuming the reference significance level parameter is 0). Obviously, in the second point cloud media, the significance level parameter is related to both time and spatial region. Therefore, the server can generate a significance information metadata track to indicate the significance information, or can generate a significance information sample group to indicate the above significance information. For the specific process of indicating the significance information, please refer to the description of the significance information data box and the significance information metadata track in step S102 below.

[0203] Among them, the point cloud media may include a third point cloud media, and the significance level parameter corresponding to the third point cloud media may be only related to the spatial region and not related to time. The embodiments of the present application will not describe the third point cloud media in detail for the time being. Please refer to the description in the corresponding embodiments below Figure 7 corresponding.

[0204] Step S102: Encode the point cloud media to obtain a point cloud bitstream, and encapsulate the point cloud bitstream and the significance information into a media file.

[0205] Specifically, when the media file includes a significance information metadata track for indicating the significance information, the target range includes a time range; among them, the significance information metadata track includes E sample numbers associated with the significance information; one sample number corresponds to one point cloud frame; the point cloud media includes E point cloud frames; E is a positive integer.

[0206] Among them, the significance information metadata track includes sample number F g ; g is a positive integer and g is less than or equal to E; the time range includes sample number F gThe corresponding point cloud frame; the saliency information metadata track includes a saliency level indication field for sample serial number F g When the field value of the saliency level indication field is the fourth indication value, it indicates that the saliency level associated with sample serial number F g is determined by the reference saliency level parameter; the reference saliency level parameter belongs to the saliency level parameters in the saliency information; when the field value of the saliency level indication field is the fifth indication value, it indicates that the saliency level associated with sample serial number F g is determined by the saliency information data structure; the fifth indication value is different from the fourth indication value.

[0207] Among them, when the field value of the saliency level indication field is the fourth indication value, it indicates that the reference saliency level parameter is used to indicate the saliency level of a time range; a time range is the point cloud frame corresponding to sample serial number F g The saliency information metadata track also includes an effective range indication field with a field value of the sixth indication value for sample serial number F g The sixth indication value indicates that the reference saliency level parameter is effective within the saliency information metadata track.

[0208] Among them, when the field value of the saliency level indication field is the fifth indication value, the saliency information metadata track also includes an effective range indication field for sample serial number F g When the field value of the effective range indication field is the sixth indication value, it indicates that the saliency level associated with sample serial number F g is effective within the saliency information metadata track; when the field value of the effective range indication field is the seventh indication value, it indicates that the saliency level associated with sample serial number F g is effective within the point cloud frame corresponding to sample serial number F g The seventh indication value is different from the sixth indication value.

[0209] Among them, when the field value of the effective range indication field is the seventh indication value, the saliency information metadata track also includes a sample saliency level field for sample serial number F g The sample saliency level field is used to indicate the saliency level parameter of the point cloud frame corresponding to sample serial number F g Among them, when the field value of the saliency level indication field is the fifth indication value, the saliency information metadata track also includes a data structure quantity field with a field value of T for sample serial number F

[0210] The data structure quantity field is used to indicate the total quantity of the saliency information data structures; T is a positive integer. g Among them, the T saliency information data structures include the saliency information data structure U

[0211] ​v where v is a positive integer and v is less than or equal to T; the saliency information data structure U v includes a saliency level field with a field value of the saliency level parameter W v and a target range indication field; the saliency level parameter W v belongs to the saliency level parameter in the saliency information; when the field value of the target range indication field is the first indication value, it means the saliency level parameter W v is used to indicate the saliency level of a spatial region in the point cloud frame corresponding to the sample sequence number F g ; when the field value of the target range indication field is the eighth indication value, it means the saliency level parameter W v is used to indicate the saliency level of the point cloud frame corresponding to the sample sequence number F g ; the eighth indication value is different from the first indication value.

[0212] Among them, when T is a positive integer greater than 1, the saliency information data structure includes a saliency level field and a target range indication field with a field value of the first indication value; the field value of the saliency level field belongs to the saliency level parameter in the saliency information; the first indication value means that the field value of the saliency level field is used to indicate the saliency level of a spatial region in the point cloud frame corresponding to the sample sequence number F g corresponding to.

[0213] Among them, the sample entry of the saliency information metadata track includes a saliency information data box; the saliency information data box includes a saliency algorithm type field; the saliency algorithm type field is used to indicate the determination algorithm type of the saliency information.

[0214] Specifically, when the media file includes Z saliency information sample groups for indicating saliency information, the target range includes a time range; where the total number of mutually different sample sequence numbers included in the Z saliency information sample groups is less than or equal to H; one sample sequence number is used to indicate one point cloud frame; the media file includes H point cloud frames; H is a positive integer, and Z is a positive integer and Z is less than H.

[0215] Among them, the Z saliency information sample groups include the saliency information sample group K m ; m is a positive integer and m is less than or equal to Z; the time range includes the point cloud frame corresponding to the saliency information sample group K m ; the point cloud frame corresponding to the saliency information sample group K m belongs to the H point cloud frames; the saliency information sample group K m includes an effective range indication field and a data structure quantity field with a field value of I; the data structure quantity field is used to indicate the total number of saliency information data structures; I is a positive integer; the I saliency information data structures are used to indicate the saliency information sample group Km The associated significance level; when the field value of the effective range indication field is the sixth indication value, it represents the significance information sample group K m The associated significance level takes effect within the point cloud track corresponding to the media file; when the field value of the effective range indication field is the seventh indication value, it represents the significance information sample group K m The associated significance level takes effect within the significance information sample group K m and takes effect; the seventh indication value is different from the sixth indication value.

[0216] Among them, when the field value of the effective range indication field is the seventh indication value, the significance information sample group K m includes a sample significance level field; the sample significance level field is used to indicate the significance level parameter of the point cloud frame corresponding to the significance information sample group K m corresponding to it.

[0217] Among them, the I significance information data structures include the significance information data structure J n , where n is a positive integer and n is less than or equal to I; the significance information data structure J n includes a significance level field with a field value of the significance level parameter L n , and a target range indication field; the significance level parameter L n belongs to the significance level parameters in the significance information; when the field value of the target range indication field is the first indication value, it means that the significance level parameter L n is used to indicate the significance level of a spatial area in the point cloud frame corresponding to the significance information sample group K m ; when the field value of the target range indication field is the eighth indication value, it means that the significance level parameter L n is used to indicate the significance level of the point cloud frame corresponding to the significance information sample group K m ; the eighth indication value is different from the first indication value.

[0218] Among them, for the process of encoding the point cloud media to obtain the point cloud bitstream, please refer to the description in the embodiment corresponding to the above Figure 3 here and will not be elaborated. Optionally, the server side can perform encoding optimization on the point cloud media according to the significance information in the point cloud media to improve the encoding efficiency or presentation effect. This operation can be carried out during the point cloud media production stage, or after the point cloud media production is completed, during the re-encoding and encapsulation stage of the point cloud media. For the specific optimization encoding process, please refer to the description in the embodiment corresponding to the following Figure 7 here and will not be further described for the time being.

[0219] As can be seen from step S101, the server can define the saliency information of each area of the immersive media (taking point cloud media as an example in this application). Specifically, it can include: a) the algorithm type for defining the saliency level according to the way of obtaining the saliency level; b) indicating the saliency information of different ranges (including spatial range and time range) according to whether the saliency level changes with space and time. Therefore, in the system layer of this application embodiment, several descriptive fields are added, including field extensions at the file encapsulation level and the transport signaling level, to support this implementation step. This application embodiment takes the form of extending the ISOBMFF data box as an example to define the method for indicating the saliency information of point cloud media. For the field extension at the transport signaling level, please refer to the description of the DASH signaling and the SMT signaling below. Figure 7 See the descriptions of the DASH signaling and the SMT signaling in

[0220] The embodiment of this application can provide the saliency information of the point cloud media through the saliency information data structure. Please refer to Table 1 together. Table 1 is used to indicate the syntax of a saliency information data structure provided by the embodiment of this application:

[0221] Table 1

[0222]

[0223]

[0224] The semantics of the syntax shown in Table 1 above are as follows: saliency_level is the saliency level field, and its value is an 8-bit unsigned integer, indicating the saliency level parameter. The larger the value of this field, the higher the saliency. spatial_info_flag is the target range indication field, and its value is a 1-bit unsigned integer. When the value of this field is the first indication value (1 in Table 1), it means that the saliency level parameter is the saliency level of a spatial area in a point cloud frame; when the value of this field is the eighth indication value (for example, 0), it means that the saliency level parameter is the saliency level of the entire point cloud frame. It should be noted that the embodiment of this application does not limit the specific values of the first indication value and the eighth indication value, as long as the two indication values are different.

[0225] region_id_ref_flag in Table 1 is the spatial range indication field, and its value is a 1-bit unsigned integer. When the value of this field is the second indication value (1 in Table 1), it means that the spatial area associated with the saliency level parameter is indexed by the identifier of this spatial area; when the value of this field is the third indication value (for example, 0), it means that the spatial area associated with the saliency level is indicated by the spatial information related data structure (i.e., the spatial area position information). It should be noted that the embodiment of this application does not limit the specific values of the second indication value and the third indication value, as long as the two indication values are different.

[0226] The spatial_region_id in Table 1 is the spatial region identifier field, and its value is a 16-bit unsigned integer, indicating the spatial region identifier corresponding to the spatial region associated with the significance level parameter. The anchor_point indicates the anchor coordinates of the spatial region, and the bounding_info indicates the length, width, and height information of the spatial region.

[0227] The slice_info_flag in Table 1 is the point cloud slice information field, and its value is a 1-bit unsigned integer. When the value of this field is the first information value (such as 1 in Table 1), it indicates that the spatial region associated with the significance level parameter is associated with one or more point cloud slices (referred to as associated point cloud slices); when the value of this field is the second information value (such as 0), it indicates that the spatial region associated with the significance level has no related point cloud slices. It should be noted that the embodiments of the present application do not limit the specific values of the first information value and the second information value, as long as the two information values are different.

[0228] The num_slices in Table 1 is the point cloud slice quantity field, and its value is a 16-bit unsigned integer, indicating the number of point cloud slices associated with the spatial region, that is, the total number of associated point cloud slices. The slice_id is the point cloud slice identifier field, and its value is a 16-bit unsigned integer, indicating the point cloud slice identifier of the associated point cloud slice.

[0229] The tile_info_flag in Table 1 is the spatial block information field, and its value is a 1-bit unsigned integer. When the value of this field is the third information value (such as 1 in Table 1), it indicates that the spatial region associated with the significance level is associated with one or more point cloud spatial blocks (referred to as associated spatial blocks); when the value of this field is the fourth information value (such as 0), it indicates that the spatial region associated with the significance level has no related point cloud spatial blocks. The embodiments of the present application do not limit the specific values of the third information value and the fourth information value, as long as the two information values are different.

[0230] The num_tiles in Table 1 is the spatial block information field, and its value is a 16-bit unsigned integer, indicating the number of spatial blocks associated with the spatial region, that is, the total number of associated spatial blocks. The tile_id is the spatial block identifier field, and its value is a 16-bit unsigned integer, indicating the spatial block identifier of the associated spatial block. It can be understood that for the spatial region indicated by the significance level parameter, the association with the tile or slice can be selected optionally.

[0231] In particular, when indexing the spatial region indicated by the significance level parameter through the spatial region identifier, the dynamic change of the spatial region corresponding to the spatial region identifier in the spatial range does not affect the static indication of the significance level parameter.

[0232] In the embodiments of the present application, the description of the saliency level parameter of the point cloud media being only related to space and not related to time will not be given for the time being. When the saliency level of the point cloud media is related to time (including only related to time, and related to both time and space), one implementable way is to indicate the saliency information that changes with time in the saliency information metadata track. Please refer to Table 2 together. Table 2 is used to indicate the syntax of a saliency information metadata track structure provided by the embodiments of the present application:

[0233] Table 2

[0234]

[0235]

[0236] It can be understood that the saliency information metadata track (also referred to as the dynamic saliency information metadata track, abbreviated as dsai) includes the sample numbers corresponding to the point cloud frames (also called samples) in the point cloud media. The semantics of the syntax shown in Table 2 above are as follows: SaliencyInfoBox is the saliency information data box. The embodiments of the present application will not expand the description of SaliencyInfoBox for the time being. Please refer to the description in the corresponding embodiments below Figure 7 which is included at the sample entry (MetaDataSampleEntry) of the saliency information metadata track.

[0237] The default_saliency_flag in Table 2 is the saliency level indication field, and its value is an unsigned integer of 1 bit. When the value of this field is the fourth indication value (for example, 1), it means that the point cloud frame has the default saliency level (usually the default saliency level is 0), that is, the saliency level of the point cloud frame is determined by the reference saliency level parameter. When the value of this field is the fifth indication value (such as 0 in Table 2), it means that the saliency level of the point cloud frame is determined by SaliencyInfoStruct within the sample. The embodiments of the present application do not limit the values of the fourth indication value and the fifth indication value, as long as the two indication values are different.

[0238] The unified_saliency_level in Table 2 is an effective range indication field, and its value is an unsigned integer of 1 bit. When the value of this field is the sixth indication value (for example, 1), it means that the saliency level parameter indicated in the current sample is effective within the saliency information metadata track, that is, all saliency level parameters within the saliency information metadata track are indicated by the same standard. When the value of this field is the seventh indication value (0 in Table 2), it means that the saliency level parameter indicated in the current sample is only effective within the current sample (point cloud frame). The embodiments of the present application do not limit the values of the sixth indication value and the seventh indication value, as long as the two indication values are different.

[0239] The sample_saliency_level in Table 2 is a sample saliency level parameter, and its value is an unsigned integer of 8 bits. This field is used to indicate the saliency level of the sample.

[0240] The num_saliency_struct in Table 2 is a data structure quantity field, and its value is an unsigned integer of 16 bits. This field is used to indicate the total number of saliency information data structures. It can be understood that when the value of this field is greater than 1, it means that the point cloud frame includes at least two spatial regions, so the value of spatial_info_flag in SaliencyInfoStruct must be the first indication value.

[0241] Through the saliency information metadata track, the server can determine the saliency level parameters corresponding to each point cloud frame (i.e., sample). In the point cloud media, if there are some point cloud frames with the same saliency level parameters, or there are some spatial regions in the point cloud frames with the same saliency level parameters, the embodiments of the present application provide another implementable manner, which is to use a media file encapsulation sample group tool (referred to as the saliency information sample group in the embodiments of the present application) to indicate the saliency information. It can be understood that both the saliency information sample group and the saliency information metadata track include the dynamic information in the saliency information, such as the saliency level parameters that change over time, etc., while the static information in the saliency information, such as the algorithm type for obtaining the saliency level parameters, etc., is included in the saliency information data box. If the media file includes a saliency information sample group for indicating the saliency information, the saliency information data box can be included at the sample entry of the point cloud track corresponding to the point cloud media. Among them, the dynamic information and the static information in the saliency information can be set according to the actual application scenario.

[0242] For ease of understanding, please refer to Table 3 together. Table 3 is used to indicate the syntax of a saliency information sample group structure provided by the embodiments of the present application:

[0243] Table 3

[0244]

[0245]

[0246] In the saliency information sample group, only the point cloud frames without the reference saliency level parameter are organized in the form of a sample group, and the saliency level parameter of one or more point cloud frames is given. Therefore, the point cloud frames that do not belong to the saliency information sample group are the point cloud frames with the reference saliency level parameter.

[0247] The semantics of the syntax shown in Table 3 above are as follows: unified_saliency_level is an effective range indication field. When its value is the sixth indication value (1 in Table 3), it means that the saliency level parameter indicated in the saliency information sample group is effective within the point cloud track corresponding to the point cloud media, that is, all saliency level parameters within the point cloud track are indicated by the same standard. When the value of this field is the seventh indication value (for example, 0), it means that the saliency level parameter indicated in the saliency information sample group is only effective within the saliency information sample group.

[0248] sample_saliency_level in Table 3 indicates the saliency level of all samples included in the saliency information sample group.

[0249] num_saliency_struct in Table 3 indicates the total number of saliency information data structures. When the value of this field is greater than 1, it means that the point cloud frames in the saliency information sample group include at least two spatial regions. Therefore, the value of spatial_info_flag in SaliencyInfoStruct must be the first indication value.

[0250] In summary, when the saliency information of the point cloud media changes with time, the embodiments of the present application can provide two implementation manners to indicate the saliency information. One implementation manner is the saliency information metadata track, and the other implementation manner is the saliency information sample group.

[0251] Please refer to Figure 5 again. If the saliency information in the first point cloud media is only related to time and has nothing to do with the spatial region, then in the media file corresponding to the first point cloud media, the server can generate the first saliency information metadata track, and the first saliency information metadata track can be as shown in Table 4. Table 4 is a saliency information metadata track structure table provided by the embodiments of the present application.

[0252] Table 4

[0253]

[0254]

[0255] The saliency_algorithm_type in Table 4 is the saliency algorithm type field. When the value is the first type value (1 in Table 4), it indicates that the saliency level parameter is obtained through data statistics. num_saliency_struct = 0 indicates that there is no saliency spatial region that does not change with time in the first saliency information metadata track. For the meanings of other fields in Table 4, please also refer to the descriptions in Tables 1 - 2 above, which will not be elaborated here.

[0256] Please refer to Figure 6 , if the saliency information in the second point cloud media is related to the spatial region and the saliency level of the spatial region changes with time, then in the media file corresponding to the second point cloud media, the server can generate a second saliency information metadata track, which can be as shown in Table 5. Table 5 is another saliency information metadata track structure table provided by the embodiments of the present application.

[0257] Table 5

[0258]

[0259]

[0260] Compared with Table 4, there are some sample serial numbers in Table 5, such as sample 1 (sample1) and sample 3 (sample3). The point cloud frames corresponding to them are divided into multiple spatial regions. For example, the point cloud frame 1 corresponding to sample 1 has two spatial regions. The saliency level of the first spatial region is 2, and the saliency level of the second spatial region is 1. For the meanings of each field in Table 5, please also refer to the descriptions in Tables 1, 2, and 4 above, which will not be elaborated here.

[0261] Among them, the server encapsulates the point cloud bitstream into a media file and indicates the above information in the form of metadata (i.e., the dynamic saliency information metadata track) in the media file. Since the point cloud saliency information changes with time, the transmission signaling does not contain saliency information description data, but the saliency information metadata track exists as a media resource in the form of Representation in the transmission signaling. As described above Figure 3 it can be seen that there are two ways for the server to transmit the point cloud file to the client, which are respectively:

[0262] a) After the client C1 downloads the complete point cloud file (i.e., the media file), it plays locally.

[0263] b) The client C2 establishes a streaming transmission with the server and presents and consumes while receiving the point cloud file segment Fs.

[0264] From the above, it can be seen that the embodiments of the present application can determine the saliency information corresponding to the point cloud media for indicating the time range, or the saliency information corresponding to the point cloud media for indicating the spatial range. Therefore, the embodiments of the present application can encapsulate the saliency information of the point cloud media together with the point cloud code stream to obtain a media file, and the saliency information indication method provided by the embodiments of the present application can improve the accuracy of the saliency information of the point cloud media, and then when rendering the point cloud media, the rendering effect of the target range can be determined through accurate saliency information, so the presentation effect of the point cloud media can be optimized.

[0265] For further information, see Figure 7 , Figure 7 This is a flow diagram of a media data processing method provided in an embodiment of the present application. Figure Two The method may be implemented by a content production device (e.g., the above-mentioned Figure 3 The content production device 200A in the corresponding embodiment is executed, for example, the content production device can be a server, and the embodiment of the present application is described by taking the server execution as an example. The method can at least include the following steps S201-S204.

[0266] Step S201, determining saliency information of the point cloud media; the saliency information includes a saliency level parameter for indicating a target range of the point cloud media; the target range includes a spatial range or a temporal range.

[0267] For details on the acquisition and production process of point cloud media, please refer to the above Figure 3 The description in is not repeated here. The saliency information in the embodiment of the present application may include two types of information, one type of information is saliency algorithm information, that is, information for obtaining saliency level parameters; the other type of information is saliency level parameters for indicating the target range.

[0268] It is understandable that different point cloud media have different ranges of associated saliency level parameters. Therefore, the present application embodiment proposes a saliency information indication method for immersive media, especially point cloud media. The method adds several descriptive fields at the system layer, including field extensions at the file encapsulation layer and the transmission signaling layer. In the following, examples are given in the form of extended ISOBMFF data boxes, DASH signaling, and SMT signaling.

[0269] The point cloud media may include a third point cloud media, and the saliency level parameter corresponding to the third point cloud media may be only related to the spatial region and not to time. Figure 8 , Figure 8It is a schematic diagram provided by an embodiment of the present application, where the saliency level parameter is related to the spatial region, and the saliency level parameter related to the spatial region does not change with time. Assume that the third point cloud media includes A point cloud frames, where A is a positive integer, and the internal structure of each point cloud frame in the A point cloud frames is as shown in Figure 8 shown. Then, according to the content of the third point cloud media, the server can define the saliency information of each region of the third point cloud media. Among them, the saliency level parameter of the first spatial region (abbreviated as spatial region 1) of each point cloud frame is 2, which is composed of two point cloud patches, and the point cloud patch identifiers corresponding to the two point cloud patches include point cloud patch 0 and point cloud patch 1. The saliency level parameter of the second spatial region (abbreviated as spatial region 2) of each point cloud frame is 1, which is composed of two point cloud patches, and the point cloud patch identifiers corresponding to the two point cloud patches include point cloud patch 2 and point cloud patch 3.

[0270] Therefore, in the third point cloud media, spatial region 1 has a higher saliency, and spatial region 2 has a lower saliency. At this time, the server can generate a saliency information data box for indicating the saliency information that does not change with time in the point cloud track corresponding to the point cloud media.

[0271] For the scenario where the saliency information changes with time, please refer to the description of step S101 in the corresponding embodiment above, which will not be elaborated here. Figure 4 The description of step S101 in the corresponding embodiment above will not be elaborated here.

[0272] Step S202: Encode the point cloud media to obtain a point cloud bitstream.

[0273] Specifically, according to the saliency level parameter in the saliency information, optimize the encoding of the target range of the point cloud media to obtain a point cloud bitstream.

[0274] Among them, the total number of target ranges is at least two, and the at least two target ranges include a first target range and a second target range; the saliency level parameter in the saliency information includes a first saliency level parameter corresponding to the first target range and a second saliency level parameter corresponding to the second target range; the specific process of optimizing the encoding of the target range of the point cloud media according to the saliency level parameter in the saliency information to obtain a point cloud bitstream may include: determining a first encoding level of the first target range according to the first saliency level parameter, determining a second encoding level of the second target range according to the second saliency level parameter, and when the first saliency level parameter is greater than the second saliency level parameter, the first encoding level is better than the second encoding level; optimizing the encoding of the first target range through the first encoding level to obtain a first sub-point cloud bitstream, and optimizing the encoding of the second target range through the second encoding level to obtain a second sub-point cloud bitstream; generating a point cloud bitstream according to the first sub-point cloud bitstream and the second sub-point cloud bitstream.

[0275] To improve the encoding efficiency and presentation effect, in the embodiments of the present application, the point cloud media can be optimized encoded based on the saliency information of the point cloud media. Specifically as follows: The server determines the saliency level parameters corresponding to at least two target ranges, sorts the at least two saliency level parameters, sets the encoding level of the target range corresponding to the maximum saliency level parameter as the maximum encoding level, and sets the encoding level of the target range corresponding to the second highest saliency level parameter as the second highest encoding level, that is, sorts the encoding levels corresponding to the at least two target ranges in a positive order according to the sorting of the at least two saliency level parameters, and then optimally encodes the target ranges according to the encoding levels corresponding to the target ranges to obtain a point cloud bitstream.

[0276] The above optimization encoding process can be executed by the content production device, or after the content production device generates a media file and transmits it to the intermediate node, the intermediate node first unpacks and decodes the media file to obtain the point cloud media, and then optimally encodes the point cloud media based on the saliency information of the point cloud media.

[0277] Step S203, encapsulate the point cloud bitstream and the saliency information into a media file.

[0278] Specifically, the media file includes a saliency information data box for indicating the saliency information. When the saliency information data box is included at the sample entry of the point cloud track corresponding to the media file, the target range includes a spatial range; wherein, one saliency level parameter in the saliency information is used to indicate one spatial range; one spatial range includes a spatial region respectively included in A point cloud frames; the point cloud media includes A point cloud frames; A is a positive integer.

[0279] Among them, the saliency information data box includes a data structure quantity field; the data structure quantity field is used to indicate the total quantity of the saliency information data structures.

[0280] Among them, the value of the data structure quantity field is S, indicating S saliency information data structures; the S saliency information data structures include saliency information data structure B c , where S and c are both positive integers, and c is less than or equal to S; saliency information data structure B c includes a saliency level field with a field value of saliency level parameter D c and a target range indication field with a field value of a first indication value; saliency level parameter D c belongs to the saliency level parameters in the saliency information; the first indication value indicates that the saliency level parameter D c is used to indicate the saliency level of a spatial range.

[0281] Among them, saliency information data structure Bc further includes a spatial range indication field; when the field value of the spatial range indication field is a second indication value, it indicates the saliency level parameter D c the indicated spatial range is determined by a spatial region identifier; when the field value of the spatial range indication field is a third indication value, it indicates the saliency level parameter D c the indicated spatial range is determined by spatial region position information; the third indication value is different from the second indication value.

[0282] wherein, when the field value of the spatial range indication field is a second indication value, the saliency information data structure B c further includes a spatial region identifier field; the spatial region identifier field is used to indicate the spatial region identifier of the spatial range indicated by the saliency level parameter D c

[0283] wherein, when the field value of the spatial range indication field is a third indication value, the saliency information data structure B c further includes a spatial region position information field; the spatial region position information field is used to indicate the spatial region position information of the spatial range indicated by the saliency level parameter D c

[0284] wherein, the saliency information data structure B c further includes a point cloud patch information field; when the field value of the point cloud patch information field is a first information value, it indicates the saliency level parameter D c the indicated spatial range has associated point cloud patches; when the field value of the point cloud patch information field is a second information value, it indicates the saliency level parameter D c the indicated spatial range does not have associated point cloud patches; the second information value is different from the first information value; wherein, when the field value of the point cloud patch information field is a first information value, the saliency information data structure B c further includes a point cloud patch quantity field and a point cloud patch identifier field; the point cloud patch quantity field is used to indicate the total quantity of the associated point cloud patches; the point cloud patch identifier field is used to indicate the point cloud patch identifier corresponding to the associated point cloud patches.

[0285] wherein, the saliency information data structure B c further includes a spatial partitioning information field; when the field value of the spatial partitioning information field is a third information value, it indicates the saliency level parameter D c the indicated spatial range has associated spatial partitions; when the field value of the point cloud patch information field is a fourth information value, it indicates the saliency level parameter D c the indicated spatial range does not have associated spatial partitions; the fourth information value is different from the third information value; wherein, when the field value of the spatial partitioning information field is a third information value, the saliency information data structure B c ​​It also includes a spatial block quantity field and a spatial block identifier field; the spatial block quantity field is used to indicate the total quantity of associated spatial blocks; the spatial block identifier field is used to indicate the spatial block identifier corresponding to the associated spatial blocks.

[0286] Among them, the saliency information data box includes a saliency algorithm type field; when the field value of the saliency algorithm type field is the first type value, it indicates that the saliency information is determined by the saliency detection algorithm; when the field value of the saliency algorithm type field is the second type value, it indicates that the saliency information is determined by data statistics; the second type value is different from the first type value.

[0287] Embodiments of the present application can provide the saliency information of the point cloud media through the saliency information data structure. When the saliency information of the point cloud media is only related to the spatial region, the server can generate a saliency information data box in the media file to indicate the saliency information that does not change with time. Please refer to Table 6 together. Table 6 is used to indicate the syntax of a saliency information data box structure provided by embodiments of the present application:

[0288] Table 6

[0289]

[0290] The saliency information data box can be included in the sample entry of the point cloud track, indicating the saliency information that does not change with time in this point cloud track. The quantity thereof is 0 or 1. In the scenario where the point cloud media has multiple point cloud tracks, the saliency information data box can be at the sample entry of any point cloud track.

[0291] The semantics of the syntax shown in Table 6 above are as follows: saliency_algorithm_type is the saliency algorithm type field, which is used to indicate the algorithm type for obtaining the saliency information. When the value of this field is the first type value (for example, 0), it indicates that the saliency information is obtained by the algorithm; when the value of this field is the second type value (for example, 1), it indicates that the saliency information is obtained by the subjective evaluation statistics of the object, that is, obtained by data statistics; other values can be extended by the application itself.

[0292] num_saliency_struct in Table 6 is the data structure quantity field, indicating the quantity of the saliency information data structure. When the value of this field is greater than 0, the value of spatial_info_flag in SaliencyInfoStruct must be 1. When the value of this field is 0, it indicates that there is no saliency spatial region that does not change with time in the point cloud track.

[0293] To sum up, when the saliency information of the point cloud media is only related to the spatial region and does not change with time, the server can provide a saliency information data box to indicate the saliency information. Please refer to againFigure 8 In the media file corresponding to the third point cloud media, the server may generate a saliency information data box, and the saliency information data box may be as shown in Table 7. Table 7 is a structure table of a saliency information data box provided by an embodiment of the present application.

[0294] Table 7

[0295]

[0296] For the meanings of the fields in Table 7, please refer to the descriptions in Table 1 and Table 6 above, and details will not be elaborated here.

[0297] Furthermore, the server encapsulates the point cloud bitstream into a point cloud file, and indicates the above saliency information in the form of metadata (i.e., the SaliencyInfoBox data box) in the file.

[0298] Step S204: Transmit the transmission signaling for the media file to the client; the transmission instruction carries saliency information description data; the saliency information description data is used to instruct the client to determine the acquisition order between different media sub-files in the media file when obtaining the media file through the streaming transmission method; the saliency information description data is generated based on the saliency information.

[0299] The present application not only adds several descriptive fields at the file encapsulation level, but also adds several descriptive fields at the transmission signaling level. The following uses the DASH signaling and the SMT signaling as examples. Among them, the saliency information description data includes a saliency information descriptor defined in the DASH signaling and a saliency information descriptor defined in the SMT signaling, which are specifically described as follows.

[0300] The embodiment of the present application extends in the DASH signaling and proposes a saliency information descriptor. The saliency information descriptor (SaliencyInfo descriptor) is a SupplementalProperty element, and its @schemeIdUri attribute is "urn:avs:ims:2022:apcc". This descriptor may exist at the adaptation set level or the representation level. When it exists at the adaptation set level, the saliency information descriptor describes all the representations within the adaptation set; when it exists at the representation level, the saliency information descriptor describes the corresponding representation. The SaliencyInfoDescriptor descriptor indicates the relevant attributes of the saliency information of the point cloud media. For the specific attributes, please refer to Table 8 together. Table 8 is used to indicate the elements and attributes of a saliency information descriptor provided by an embodiment of the present application.

[0301] Table 8

[0302]

[0303]

[0304] Among them, N in Table 8 represents the total number of saliency level parameters in the saliency information. For example, Figure 8 in [example], N = 2, indicating that there are two saliency level parameters. M represents that its corresponding field (such as SaliencyInfo@saliencyLevel in Table 8) is a Mandatory field; CM represents its corresponding field (such as SaliencyInfo@spatialRegionId in Table 8), which is a Conditional Mandatory field; O represents that its corresponding field (such as SaliencyInfo@tileId in Table 8) is an Optional field.

[0305] unsigned in Table 8 means unsigned, Short means short integer, bool means boolean variable, Int means integer, vector means vector, and float means floating point type.

[0306] Another feasible transmission signaling extension. In the embodiments of the present application, an extension is made in the SMT signaling, and a saliency information descriptor is proposed. It exists at the representation level and is used to describe the corresponding media resource and indicate the saliency information of the media resource. Please refer to Table 9 together. Table 9 is used to indicate a syntax of the saliency information descriptor provided by the embodiments of the present application:

[0307] Table 9

[0308]

[0309]

[0310] The semantics of the syntax shown in Table 9 are as follows: Saliency_info_level indicates the saliency level, and the larger the value of this field, the higher the saliency. Region_id_ref_flag is a spatial range indication field. When the value of this field is the first indication value (e.g., 1), the spatial region corresponding to the saliency level is indexed by the spatial region identifier; when the value of this field is the eighth indication value (e.g., 0), the spatial region corresponding to the saliency level is directly indicated by the spatial region position information. Spatial_region_id is a spatial region identifier field, indicating the spatial region identifier. Anchor_point_x, y, z indicate the x, y, z coordinates of the spatial region anchor point. Bounding_box_x, y, z indicate the lengths of the spatial region along the x, y, z axes. When Related_tile_info_flag takes the value of the third information value (e.g., 1), it means that the spatial region associated with the saliency level parameter is associated with one or more spatial tiles; when it takes the value of the fourth information value (e.g., 0), it means that the spatial region associated with the saliency level has no associated spatial tiles. When Related_slice_info_flag takes the value of the first information value (e.g., 1), it means that the spatial region associated with the saliency level is associated with one or more point cloud slices; when it takes the value of the second information value (e.g., 0), it means that the spatial region associated with the saliency level has no associated point cloud slices. Num_tiles indicates the number of spatial tiles associated with the spatial region. Tile_id indicates the identifier of the associated spatial tile. Num_slices indicates the number of point cloud slices associated with the spatial region. Slice_id indicates the identifier of the associated point cloud slice. For the meanings of the above fields, refer to the description in Table 1 above.

[0311] When the saliency level parameter in the saliency information changes over time, the saliency information of the point cloud media exists in the media file in the form of a metadata track or a sample group. At this time, the saliency information descriptor or saliency information descriptor is not included in the transmission signaling, but the saliency information metadata track or saliency information sample group will exist as a media resource in the form of a Representation in the transmission signaling.

[0312] When the saliency information of the point cloud media is associated with the spatial region and does not change over time, such as Figure 8 shown in the third example of the point cloud media. At this time, the server associates the saliency information and the spatial information in the transmission signaling, generates the signaling and sends it to the client. For Figure 8 the third example of the point cloud media shown, the transmission signaling generated by the server contains 2 saliency information descriptors:

[0313] SaliencyInfo descriptor1:

[0314] SaliencyInfo@saliencyLevel = 2; SaliencyInfo@regionIdRefFlag = 1;

[0315] SaliencyInfo@spatialRegionId = 1; SaliencyInfo@sliceId = 0, 1;

[0316] SaliencyInfo descriptor2:

[0317] SaliencyInfo@saliencyLevel = 1; SaliencyInfo@regionIdRefFlag = 1;

[0318] SaliencyInfo@spatialRegionId = 2; SaliencyInfo@sliceId = 2, 3.

[0319] For Figure 8 When establishing a streaming transmission with the server for the third point cloud media shown in the example, for the client, if spatial region 1 and spatial region 2 correspond to different media resources Representation1 and Representation2, since spatial region 1 has a higher saliency, the client can give priority to ensuring the transmission of Representation1 during transmission.

[0320] The embodiment of the present application proposes a saliency information indication method for immersive media, especially point cloud media. At the file encapsulation level and the signaling transmission level, the embodiment of the present application defines saliency information in different ranges and defines different types of saliency information acquisition algorithms, and associates the saliency information and spatial information at the signaling level; therefore, it can more flexibly indicate the saliency information of different spatial regions and different point cloud frames in the point cloud media, so as to meet more point cloud application scenarios, enable the server to perform coding optimization according to the saliency information, and enable the client to perform transmission optimization according to the saliency information.

[0321] Further, please refer to Figure 9 , Figure 9 is a flowchart of a media data processing method provided by the embodiment of the present application Figure Three . This method can be executed by a content consumption device in an immersive media system (for example, the content consumption device 200B in the corresponding embodiment above Figure 3 ). For example, the content consumption device can be a terminal integrated with a client (such as a video client). This method can at least include the following steps S301 - step S302:

[0322] Step S301: Obtain a media file, perform demultiplexing on the media file to obtain a point cloud bitstream and the saliency information of the point cloud media; the saliency information includes a saliency level parameter for indicating the target range of the point cloud media; the target range includes a spatial range or a temporal range.

[0323] Specifically, the client can obtain the media file of the immersive media sent by the server and perform demultiplexing on the media file, so as to obtain the point cloud bitstream and the saliency information of the point cloud media in the media file. It can be understood that the demultiplexing process is the reverse of the multiplexing process, and the client can demultiplex the media file according to the file format requirements adopted during multiplexing to obtain the point cloud bitstream. For the specific process of the server generating and sending the media file, reference can be made to the corresponding embodiments described above Figure 4 and will not be elaborated here.

[0324] Step S302: Decode the point cloud bitstream to obtain the point cloud media.

[0325] Specifically, when the media file includes a saliency information data box for indicating the saliency information, it is determined that the target range includes a spatial range; one saliency level parameter in the saliency information is used to indicate a spatial range; one spatial range includes a spatial region respectively included in A point cloud frames; the point cloud media includes A point cloud frames; A is a positive integer.

[0326] Among them, the total number of spatial ranges is at least two, and the at least two spatial ranges include a first spatial range and a second spatial range; the specific rendering process may further include: in the saliency information, obtain the saliency level parameter O for indicating the first spatial range p , obtain the saliency level parameter O for indicating the second spatial range p+1 ; p is a positive integer, and p is less than the total number of saliency level parameters in the saliency information; if the saliency level parameter O p is greater than the saliency level parameter O p+1 , it is determined that the rendering level corresponding to the first spatial range is superior to the rendering level corresponding to the second spatial range; if the saliency level parameter O p is less than the saliency level parameter O p+1 , it is determined that the rendering level corresponding to the second spatial range is superior to the rendering level corresponding to the first spatial range.

[0327] Specifically, when the media file includes a saliency information metadata track for indicating the saliency information, it is determined that the target range includes a temporal range; among them, the saliency information metadata track includes E sample numbers associated with the saliency information; among them, one sample number corresponds to one point cloud frame; the point cloud media includes E point cloud frames; E is a positive integer.

[0328] Among them, the time range includes the first point cloud frame and the second point cloud frame in E point cloud frames; the specific rendering process may further include: in the saliency information, obtaining the saliency level parameter Q for indicating the first point cloud frame r , obtaining the saliency level parameter Q for indicating the second point cloud frame r+1 ; r is a positive integer and r is less than the total number of saliency level parameters in the saliency information; if the saliency level parameter Q r is greater than the saliency level parameter Q r+1 , then it is determined that the rendering level corresponding to the first point cloud frame is better than the rendering level corresponding to the second point cloud frame; if the saliency level parameter Q r is less than the saliency level parameter Q r+1 , then it is determined that the rendering level corresponding to the second point cloud frame is better than the rendering level corresponding to the first point cloud frame.

[0329] Among them, the specific rendering process may further include: if the first point cloud frame includes at least two spatial regions, and the saliency levels corresponding to the at least two spatial regions are different, then in the saliency information, obtaining the saliency level parameter X corresponding to the first spatial region y , obtaining the saliency level parameter X corresponding to the second spatial region y+1 ; both the first spatial region and the second spatial region belong to the at least two spatial regions; x is a positive integer, and x is less than the total number of saliency level parameters in the saliency information; if the saliency level parameter X y is greater than the saliency level parameter X y+1 , then it is determined that the rendering level corresponding to the first spatial region is better than the rendering level corresponding to the second spatial region; if the saliency level parameter X y is less than the saliency level parameter X y+1 , then it is determined that the rendering level corresponding to the second spatial region is better than the rendering level corresponding to the first spatial region.

[0330] It can be understood that the decoding process is the reverse of the encoding process. The client can decode the point cloud bitstream according to the file format requirements adopted during encoding to obtain the point cloud media.

[0331] After the client unpacks and decodes the point cloud file / file segment, it can flexibly allocate computing resources during the presentation and rendering of the point cloud media according to the saliency information of the point cloud media to optimize the presentation effect of the target range. Such as Figure 5For the first point cloud media shown, sample2 is the point cloud frame with the highest saliency, sample3 is the point cloud frame with the second highest saliency, and the remaining samples only have the initial saliency level parameter (which can be set to 0). Therefore, the client can render sample2 and sample3 more finely, that is, the rendering level has a positive relationship with the saliency level.

[0332] Please refer to Figure 6 For the second point cloud media shown as an example, sample2 is the frame with the highest saliency. In sample1, there are spatial regions with the highest saliency and spatial regions with the second highest saliency. In sample3, there are spatial regions with the second highest saliency. The remaining samples only have the default saliency, that is, the reference saliency level parameter. Therefore, the client can render sample2, the spatial regions with the highest saliency and the spatial regions with the second highest saliency in sample1, and the spatial regions with the second highest saliency in sample3 more finely. For example, the rendering level corresponding to sample2 is better than the rendering level corresponding to sample1. In sample1, the rendering level corresponding to the spatial region with the highest saliency is better than the rendering level corresponding to the spatial region with the second highest saliency; and the rendering level corresponding to sample1 is better than the rendering level corresponding to sample3.

[0333] Please refer to Figure 8 For the third point cloud media shown as an example, spatial region 1 is the region with relatively high saliency. Therefore, the client can render spatial region 1 more finely, that is, the rendering level corresponding to spatial region 1 is better than the rendering level corresponding to spatial region 2.

[0334] In summary, the embodiments of the present application can more flexibly indicate the saliency information of different spatial regions and different point cloud frames in the point cloud media, so as to meet more point cloud application scenarios, enable the server to perform encoding optimization according to the saliency information, and enable the client to perform transmission optimization according to the saliency information; during the presentation and rendering process, the client can flexibly allocate computing resources according to the saliency information to optimize the presentation effect of specific regions.

[0335] Please refer to Figure 10 , Figure 10 is the first schematic structural diagram of a media data processing device provided by an embodiment of the present application. The media data processing device can be a computer program (including program code) running on a content production device. For example, the media data processing device is an application software in the content production device; the device can be used to execute the corresponding steps in the media data processing method provided by the embodiments of the present application. As Figure 10 shown, the media data processing device 1 may include: an information determination module 11 and an information encapsulation module 12.

[0336] An information determination module 11 is used to determine saliency information of the point cloud media; the saliency information includes a saliency level parameter for indicating a target range of the point cloud media; the target range includes a spatial range or a temporal range;

[0337] The information encapsulation module 12 is used to encode the point cloud media to obtain a point cloud code stream, and encapsulate the point cloud code stream and saliency information into a media file.

[0338] The specific implementation of the information determination module 11 and the information packaging module 12 can be found in the above Figure 4 Steps S101 and S102 in the corresponding embodiment will not be described in detail here.

[0339] In one embodiment, the media file includes a saliency information data box for indicating saliency information, and when the saliency information data box is included at a sample entry of a point cloud track corresponding to the media file, the target range includes a spatial range;

[0340] Among them, a saliency level parameter in the saliency information is used to indicate a spatial range; a spatial range includes a spatial area respectively contained in A point cloud frames; the point cloud media includes A point cloud frames; A is a positive integer.

[0341] In one embodiment, the saliency information data box includes a data structure quantity field; the data structure quantity field is used to indicate the total number of saliency information data structures.

[0342] In one embodiment, the value of the data structure number field is S, indicating S saliency information data structures; the S saliency information data structures include saliency information data structure B c , where S and c are both positive integers, and c is less than or equal to S;

[0343] Prominence information data structure B c Include field value as significance level parameter D c The significance level field, and the target range indication field whose field value is the first indication value; the significance level parameter D c The significance level parameter belongs to the significance information;

[0344] The first indicator value represents the significance level parameter D c Used to indicate the significance level of a spatial extent.

[0345] In one embodiment, the saliency information data structure B c Also included is a field indicating the spatial extent;

[0346] When the field value of the spatial range indication field is the second indication value, it represents the saliency level parameter D c The indicated spatial range is determined by the spatial region identifier;

[0347] When the field value of the spatial range indication field is the third indication value, it represents the saliency level parameter D c The indicated spatial range is determined by the spatial region location information; the third indication value is different from the second indication value.

[0348] In one embodiment, when the field value of the spatial range indication field is the second indication value, the saliency information data structure B c further includes a spatial region identifier field; the spatial region identifier field is used to indicate the spatial region identifier of the indicated spatial range of the saliency level parameter D c

[0349] In one embodiment, when the field value of the spatial range indication field is the third indication value, the saliency information data structure B c further includes a spatial region location information field; the spatial region location information field is used to indicate the spatial region location information of the indicated spatial range of the saliency level parameter D c

[0350] In one embodiment, the saliency information data structure B c further includes a point cloud patch information field;

[0351] When the field value of the point cloud patch information field is the first information value, it represents the saliency level parameter D c The indicated spatial range has associated point cloud patches;

[0352] When the field value of the point cloud patch information field is the second information value, it represents the saliency level parameter D c The indicated spatial range does not have associated point cloud patches; the second information value is different from the first information value;

[0353] When the field value of the point cloud patch information field is the first information value, the saliency information data structure B c further includes a point cloud patch quantity field and a point cloud patch identifier field; the point cloud patch quantity field is used to indicate the total quantity of the associated point cloud patches; the point cloud patch identifier field is used to indicate the point cloud patch identifier corresponding to the associated point cloud patches.

[0354] In one embodiment, the saliency information data structure B c further includes a spatial partitioning information field;

[0355] When the field value of the spatial partitioning information field is the third information value, it represents the saliency level parameter D c ​​The indicated spatial range has an associated spatial segmentation;

[0356] When the field value of the point cloud slice information field is the fourth information value, it represents the saliency level parameter D c The indicated spatial range does not have an associated spatial segmentation; the fourth information value is different from the third information value;

[0357] When the field value of the spatial segmentation information field is the third information value, the saliency information data structure B c It further includes a spatial segmentation quantity field and a spatial segmentation identifier field; the spatial segmentation quantity field is used to indicate the total quantity of the associated spatial segmentations; the spatial segmentation identifier field is used to indicate the spatial segmentation identifier corresponding to the associated spatial segmentation.

[0358] In one implementation, the saliency information data box includes a saliency algorithm type field;

[0359] When the field value of the saliency algorithm type field is the first type value, it represents that the saliency information is determined by the saliency detection algorithm;

[0360] When the field value of the saliency algorithm type field is the second type value, it represents that the saliency information is determined by data statistics; the second type value is different from the first type value.

[0361] Please refer to Figure 10 , the media data processing device 1 may further include: a file transmission module 13.

[0362] The file transmission module 13 is used to transmit the transmission signaling for the media file to the client; the transmission instruction carries the saliency information description data; the saliency information description data is used to instruct the client to determine the acquisition order between different media sub-files in the media file when acquiring the media file through the streaming transmission method; the saliency information description data is generated based on the saliency information.

[0363] Among them, the specific implementation manner of the file transmission module 13 can refer to step S204 in the above Figure 7 corresponding embodiment, which will not be elaborated here.

[0364] In one implementation, when the media file includes a saliency information metadata track for indicating the saliency information, the target range includes a time range;

[0365] Among them, the saliency information metadata track includes E sample serial numbers associated with the saliency information; among them, one sample serial number corresponds to one point cloud frame; the point cloud media includes E point cloud frames; E is a positive integer.

[0366] In one implementation, the saliency information metadata track includes the sample serial number F g; g is a positive integer and g is less than or equal to E; the time range includes the sample number F g corresponding point cloud frame;

[0367] The saliency information metadata track includes for the sample number F g saliency level indication field;

[0368] When the field value of the saliency level indication field is the fourth indication value, it means that the saliency level associated with the sample number F g is determined by the reference saliency level parameter; the reference saliency level parameter belongs to the saliency level parameters in the saliency information;

[0369] When the field value of the saliency level indication field is the fifth indication value, it means that the saliency level associated with the sample number F g is determined by the saliency information data structure; the fifth indication value is different from the fourth indication value.

[0370] In one implementation, when the field value of the saliency level indication field is the fourth indication value, it means that the reference saliency level parameter is used to indicate the saliency level of a time range; a time range is the point cloud frame corresponding to the sample number F g corresponding point cloud frame;

[0371] The saliency information metadata track also includes for the sample number F g effective range indication field with a field value of the sixth indication value; the sixth indication value indicates that the reference saliency level parameter is effective within the saliency information metadata track.

[0372] In one implementation, when the field value of the saliency level indication field is the fifth indication value, the saliency information metadata track also includes for the sample number F g effective range indication field;

[0373] When the field value of the effective range indication field is the sixth indication value, it means that the saliency level associated with the sample number F g is effective within the saliency information metadata track;

[0374] When the field value of the effective range indication field is the seventh indication value, it means that the saliency level associated with the sample number F g is effective within the point cloud frame corresponding to the sample number F g ; the seventh indication value is different from the sixth indication value.

[0375] In one implementation, when the field value of the effective range indication field is the seventh indication value, the saliency information metadata track also includes for the sample number F g sample saliency level field; the sample saliency level field is used to indicate the sample number Fg The saliency level parameter of the corresponding point cloud frame.

[0376] In one embodiment, when the field value of the saliency level indication field is the fifth indication value, the saliency information metadata track further includes a data structure quantity field for the sample serial number F g whose field value is T; the data structure quantity field is used to indicate the total quantity of the saliency information data structures; T is a positive integer.

[0377] In one embodiment, the T saliency information data structures include the saliency information data structure U v , where v is a positive integer and v is less than or equal to T;

[0378] The saliency information data structure U v includes a saliency level field whose field value is the saliency level parameter W v , and a target range indication field; the saliency level parameter W v belongs to the saliency level parameters in the saliency information;

[0379] When the field value of the target range indication field is the first indication value, it means that the saliency level parameter W v is used to indicate the saliency level of a spatial region in the point cloud frame corresponding to the sample serial number F g .

[0380] When the field value of the target range indication field is the eighth indication value, it means that the saliency level parameter W v is used to indicate the saliency level of the point cloud frame corresponding to the sample serial number F g ; the eighth indication value is different from the first indication value.

[0381] In one embodiment, when T is a positive integer greater than 1, the saliency information data structure includes a saliency level field and a target range indication field whose field value is the first indication value; the field value of the saliency level field belongs to the saliency level parameters in the saliency information; the first indication value means that the field value of the saliency level field is used to indicate the saliency level of a spatial region in the point cloud frame corresponding to the sample serial number F g .

[0382] In one embodiment, a saliency information data box is included at the sample entry of the saliency information metadata track;

[0383] The saliency information data box includes a saliency algorithm type field; the saliency algorithm type field is used to indicate the determination algorithm type of the saliency information.

[0384] In one embodiment, when the media file includes Z groups of saliency information samples for indicating saliency information, the target range includes a time range;

[0385] Wherein, the total number of mutually different sample serial numbers respectively included in the Z groups of saliency information samples is less than or equal to H; one sample serial number is used to indicate one point cloud frame; the media file includes H point cloud frames; H is a positive integer, and Z is a positive integer and Z is less than H.

[0386] In one embodiment, the Z groups of saliency information samples include the saliency information sample group K m ; m is a positive integer and m is less than or equal to Z; the time range includes the point cloud frames corresponding to the saliency information sample group K m ; the point cloud frames corresponding to the saliency information sample group K m belong to the H point cloud frames;

[0387] The saliency information sample group K m includes an effective range indication field and a data structure quantity field with a field value of I; the data structure quantity field is used to indicate the total quantity of saliency information data structures; I is a positive integer; the I saliency information data structures are used to indicate the saliency levels associated with the saliency information sample group K m ;

[0388] When the field value of the effective range indication field is the sixth indication value, it indicates that the saliency level associated with the saliency information sample group K m is effective within the point cloud track corresponding to the media file;

[0389] When the field value of the effective range indication field is the seventh indication value, it indicates that the saliency level associated with the saliency information sample group K m is effective within the saliency information sample group K m ; the seventh indication value is different from the sixth indication value.

[0390] In one embodiment, when the field value of the effective range indication field is the seventh indication value, the saliency information sample group K m includes a sample saliency level field; the sample saliency level field is used to indicate the saliency level parameter of the point cloud frame corresponding to the saliency information sample group K m ;

[0391] In one embodiment, the I saliency information data structures include the saliency information data structure J n , n is a positive integer, and n is less than or equal to I;

[0392] The saliency information data structure J n includes a field value of the saliency level parameter L nThe saliency level field and the target range indication field; the saliency level parameter L n Belongs to the saliency level parameter in the saliency information;

[0393] When the field value of the target range indication field is the first indication value, it represents the saliency level parameter L n Used to indicate the saliency information sample group K m The saliency level of a spatial region in the point cloud frame corresponding to;

[0394] When the field value of the target range indication field is the eighth indication value, it represents the saliency level parameter L n Used to indicate the saliency information sample group K m The saliency level of the point cloud frame corresponding to; the eighth indication value is different from the first indication value.

[0395] Please refer to again Figure 10 See also, the information encapsulation module 12 is specifically configured to optimize and encode the target range of the point cloud media according to the saliency level parameter in the saliency information to obtain a point cloud bitstream.

[0396] Among them, the specific implementation manner of the information encapsulation module 12 can be referred to the above Figure 7 The steps in the corresponding embodiment of S202 will not be elaborated here.

[0397] Please refer to again Figure 10 See also, the total number of target ranges is at least two, and the at least two target ranges include a first target range and a second target range; the saliency level parameters in the saliency information include a first saliency level parameter corresponding to the first target range and a second saliency level parameter corresponding to the second target range;

[0398] The information encapsulation module 12 may include: a level determination unit 121, an optimization encoding unit 122, and a bitstream generation unit 123.

[0399] The level determination unit 121 is configured to determine a first encoding level of the first target range according to the first saliency level parameter, determine a second encoding level of the second target range according to the second saliency level parameter, and when the first saliency level parameter is greater than the second saliency level parameter, the first encoding level is better than the second encoding level;

[0400] The optimization encoding unit 122 is configured to optimize and encode the first target range through the first encoding level to obtain a first sub-point cloud bitstream, and optimize and encode the second target range through the second encoding level to obtain a second sub-point cloud bitstream;

[0401] The bitstream generation unit 123 is configured to generate a point cloud bitstream according to the first sub-point cloud bitstream and the second sub-point cloud bitstream.

[0402] Among them, for the specific implementation manners of the level determination unit 121, the optimized encoding unit 122, and the bitstream generation unit 123, reference may be made to step S202 in the corresponding embodiment above, which will not be elaborated here. Figure 7 For the corresponding embodiment above, which will not be elaborated here.

[0403] The embodiment of the present application provides a saliency information indication method for immersive media, especially point cloud media. At the file encapsulation level and the signaling transmission level, the embodiment of the present application defines saliency information in different ranges and defines different types of saliency information acquisition algorithms, and associates the saliency information with the spatial information at the signaling level; therefore, the saliency information of different spatial regions and different point cloud frames in the point cloud media can be indicated more flexibly, so as to meet more point cloud application scenarios, enable the server to perform encoding optimization according to the saliency information, and enable the client to perform transmission optimization according to the saliency information.

[0404] Please refer to Figure 11 , Figure 11 is a schematic structural diagram of a media data processing device provided by the embodiment of the present application. Figure Two The media data processing device may be a computer program (including program code) running on a content consumption device. For example, the media data processing device is an application software (such as a video client) in the content consumption device; the device may be used to execute the corresponding steps in the media data processing method provided by the embodiment of the present application. As Figure 11 shown, the media data processing device 2 may include: a file acquisition module 21 and a bitstream decoding module 22.

[0405] The file acquisition module 21 is configured to acquire a media file, unpack the media file, and obtain a point cloud bitstream and the saliency information of the point cloud media; the saliency information includes a saliency level parameter for indicating the target range of the point cloud media; the target range includes a spatial range or a time range;

[0406] The bitstream decoding module 22 is configured to decode the point cloud bitstream to obtain the point cloud media.

[0407] Among them, for the specific implementation manners of the file acquisition module 21 and the bitstream decoding module 22, reference may be made to step S301-step S302 in the corresponding embodiment above, which will not be elaborated here. Figure 9 For the corresponding embodiment above, which will not be elaborated here.

[0408] Please refer to Figure 11 again, the media data processing device 2 may further include: a first determination module 23.

[0409] The first determination module 23 is configured to determine that the target range includes a spatial range when the media file includes a significance information data box for indicating significance information; a significance level parameter in the significance information is used to indicate a spatial range; a spatial range includes a spatial region respectively included in A point cloud frames; the point cloud media includes A point cloud frames; A is a positive integer.

[0410] Among them, for the specific implementation manner of the first determination module 23, reference may be made to step S302 in the corresponding embodiment described above Figure 5 and details are not described herein again.

[0411] Please refer to Figure 11 , the total number of spatial ranges is at least two, and the at least two spatial ranges include a first spatial range and a second spatial range;

[0412] The media data processing device 1 may further include: a first acquisition module 24 and a second determination module 25.

[0413] The first acquisition module 24 is configured to acquire, from the significance information, a significance level parameter O for indicating the first spatial range p , and acquire a significance level parameter O for indicating the second spatial range p+1 ; p is a positive integer, and p is less than the total number of significance level parameters in the significance information;

[0414] The second determination module 25 is configured to, if the significance level parameter O p is greater than the significance level parameter O p+1 , determine that the rendering level corresponding to the first spatial range is better than the rendering level corresponding to the second spatial range;

[0415] The second determination module 25 is further configured to, if the significance level parameter O p is less than the significance level parameter O p+1 , determine that the rendering level corresponding to the second spatial range is better than the rendering level corresponding to the first spatial range.

[0416] Among them, for the specific implementation manners of the first acquisition module 24 and the second determination module 25, reference may be made to step S302 in the corresponding embodiment described above Figure 9 and details are not described herein again.

[0417] Please refer to Figure 11 , the media data processing device 1 may further include: a third determination module 26.

[0418] The third determination module 26 is configured to determine that the target range includes a time range when the media file includes a significance information metadata track for indicating significance information;

[0419] Among them, the saliency information metadata track includes E sample serial numbers associated with the saliency information; among them, one sample serial number corresponds to one point cloud frame; the point cloud media includes E point cloud frames; E is a positive integer.

[0420] Among them, for the specific implementation manner of the third determination module 26, reference can be made to step S302 in the corresponding embodiment above Figure 9 and details are not described herein again.

[0421] Please refer to Figure 11 , the time range includes the first point cloud frame and the second point cloud frame among the E point cloud frames;

[0422] The media data processing apparatus 1 may further include: a second acquisition module 27 and a fourth determination module 28.

[0423] The second acquisition module 27 is configured to acquire, in the saliency information, a saliency level parameter Q for indicating the first point cloud frame r , and acquire a saliency level parameter Q for indicating the second point cloud frame r+1 ; r is a positive integer and r is less than the total number of saliency level parameters in the saliency information;

[0424] The fourth determination module 28 is configured to, if the saliency level parameter Q r is greater than the saliency level parameter Q r+1 , determine that the rendering level corresponding to the first point cloud frame is better than the rendering level corresponding to the second point cloud frame;

[0425] The fourth determination module 28 is further configured to, if the saliency level parameter Q r is less than the saliency level parameter Q r+1 , determine that the rendering level corresponding to the second point cloud frame is better than the rendering level corresponding to the first point cloud frame.

[0426] Among them, for the specific implementation manners of the second acquisition module 27 and the fourth determination module 28, reference can be made to step S302 in the corresponding embodiment above Figure 9 and details are not described herein again.

[0427] Please refer to Figure 11 , the media data processing apparatus 1 may further include: a third acquisition module 29 and a fifth determination module 30.

[0428] The third acquisition module 29 is configured to, if the first point cloud frame includes at least two spatial regions and the saliency levels corresponding to the at least two spatial regions are different, acquire, in the saliency information, a saliency level parameter X corresponding to the first spatial region y , and acquire a saliency level parameter X corresponding to the second spatial region y+1; both the first spatial region and the second spatial region belong to at least two spatial regions; x is a positive integer, and x is less than the total number of saliency level parameters in the saliency information;

[0429] The fifth determination module 30 is configured to, if the saliency level parameter X y is greater than the saliency level parameter X y+1 , determine that the rendering level corresponding to the first spatial region is superior to the rendering level corresponding to the second spatial region;

[0430] The fifth determination module 30 is further configured to, if the saliency level parameter X y is less than the saliency level parameter X y+1 , determine that the rendering level corresponding to the second spatial region is superior to the rendering level corresponding to the first spatial region.

[0431] Among them, the specific implementation manners of the third acquisition module 29 and the fifth determination module 30 can refer to step S302 in the corresponding embodiment above, which will not be elaborated here. Figure 9 The embodiments of the present application can more flexibly indicate the saliency information of different spatial regions and different point cloud frames in the point cloud media, so as to meet more point cloud application scenarios, enable the server to perform encoding optimization according to the saliency information, and enable the client to perform transmission optimization according to the saliency information; during the presentation and rendering process, the client can flexibly allocate computing resources according to the saliency information to optimize the presentation effect of specific regions.

[0432] Please refer to

[0433] which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 12 shown, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to implement connection communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may further be at least one storage device located far from the aforementioned processor 1001. As Figure 12 Figure 12 ​As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0434] In the computer device 1000 as shown in Figure 12 the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an interface for users to input; and the processor 1001 can be used to call the device control application program stored in the memory 1005. It should be understood that the computer device 1000 described in the embodiments of the present application can execute the descriptions of the data processing methods or devices in the previous embodiments, which will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated either.

[0435] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it realizes the descriptions of the data processing methods or devices in the previous embodiments, which will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated either.

[0436] The above computer-readable storage medium may be the data processing device provided in any of the previous embodiments or the internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.

[0437] The embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device can execute the descriptions of the data processing methods or devices in the previous embodiments, which will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated either.

[0438] Further, please refer to Figure 13 , Figure 13It is a schematic structural diagram of a data processing system provided by an embodiment of the present application. The data processing system 3 may include a data processing device 1a and a data processing device 2a. Among them, the data processing device 1a may be the media data processing device 1 in the corresponding embodiment described above. It can be understood that the data processing device 1a may be integrated in the content production device 200A in the corresponding embodiment described above. Therefore, it will not be elaborated here. Among them, the data processing device 2a may be the media data processing device 2 in the corresponding embodiment described above. It can be understood that the data processing device 2a may be integrated in the content consumption device 200B in the corresponding embodiment described above. Therefore, it will not be elaborated here. In addition, the beneficial effects of using the same method will not be described in detail. For the technical details not disclosed in the data processing system embodiment involved in the present application, please refer to the description of the method embodiment of the present application. Figure 10 in the corresponding embodiment described above, the media data processing device 1. It can be understood that the data processing device 1a may be integrated in the content production device 200A in the corresponding embodiment described above. Therefore, it will not be elaborated here. Figure 3 in the corresponding embodiment described above, the content production device 200A. Therefore, it will not be elaborated here. Among them, the data processing device 2a may be the media data processing device 2 in the corresponding embodiment described above. Figure 11 in the corresponding embodiment described above, the media data processing device 2. It can be understood that the data processing device 2a may be integrated in the content consumption device 200B in the corresponding embodiment described above. Therefore, it will not be elaborated here. Figure 3 in the corresponding embodiment described above, the content consumption device 200B. Therefore, it will not be elaborated here. In addition, the beneficial effects of using the same method will not be described in detail. For the technical details not disclosed in the data processing system embodiment involved in the present application, please refer to the description of the method embodiment of the present application.

[0439] In the description of the embodiments of the present application, the terms "first", "second", etc. in the specification, claims and drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the steps or modules listed, but may optionally further include steps or modules not listed, or may optionally further include other steps or units inherent to these processes, methods, devices, products or equipment.

[0440] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0441] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A media data processing method, characterized in that, Including: Determining the saliency information of the point cloud media; The saliency information includes a saliency level parameter for indicating the target range of the point cloud media; the target range includes a spatial range or a temporal range; Encoding the point cloud media to obtain a point cloud bitstream, and encapsulating the point cloud bitstream and the saliency information into a media file; When the media file includes a saliency information metadata track for indicating the saliency information, the target range includes the temporal range; wherein, the saliency information metadata track includes E sample sequence numbers associated with the saliency information; one sample sequence number corresponds to one point cloud frame; the point cloud media includes E point cloud frames; E is a positive integer; The saliency information metadata track includes a sample serial number F g ; g is a positive integer and g is less than or equal to E; the time range includes the point cloud frame corresponding to the sample serial number F g ; The saliency information metadata track includes a saliency level indication field for the sample serial number F g ; When the field value of the salience level indication field is the fourth indication value, it indicates that the salience level associated with the sample serial number F g is determined by the reference salience level parameter; the reference salience level parameter belongs to the salience level parameters in the salience information; When the field value of the salience level indication field is the fifth indication value, it indicates that the salience level associated with the sample serial number F g is determined by the salience information data structure; the fifth indication value is different from the fourth indication value.

2. The method according to claim 1, wherein The media file includes a saliency information data box for indicating the saliency information. When the saliency information data box is included at the sample entry of the point cloud track corresponding to the media file, the target range includes the spatial range; Wherein, one saliency level parameter in the saliency information is used to indicate a spatial range; the one spatial range includes a spatial region respectively included in A point cloud frames; the point cloud media includes the A point cloud frames; A is a positive integer.

3. The method according to claim 2, wherein The saliency information data box includes a data structure quantity field; the data structure quantity field is used to indicate the total quantity of saliency information data structures.

4. The method according to claim 3, wherein The value of the data structure quantity field is S, indicating S saliency information data structures; the S saliency information data structures include the saliency information data structure B c , where both S and c are positive integers, and c is less than or equal to S; The saliency information data structure B c includes a saliency level field with a field value being a saliency level parameter D c and a target range indication field with a field value being a first indication value; the saliency level parameter D c belongs to the saliency level parameter in the saliency information; The first indication value represents the saliency level parameter D c for indicating the saliency level of a spatial range.

5. The method according to claim 4, wherein The saliency information data structure B c further includes a spatial range indication field; When the field value of the spatial range indication field is the second indication value, it indicates the saliency level parameter D c The indicated spatial range is determined by the spatial region identifier; When the field value of the spatial range indication field is the third indication value, it indicates the saliency level parameter D c The indicated spatial range is determined by the spatial region position information; the third indication value is different from the second indication value.

6. The method according to claim 5, characterized in that, When the field value of the spatial range indication field is the second indication value, the saliency information data structure B c further includes a spatial region identifier field; the spatial region identifier field is used to indicate the spatial region identifier of the spatial range indicated by the saliency level parameter D c ​ 7. The method according to claim 5, characterized in that, When the field value of the spatial range indication field is the third indication value, the saliency information data structure B c further includes a spatial region location information field; the spatial region location information field is used to indicate the spatial region location information of the spatial range indicated by the saliency level parameter D c ​ 8. The method according to claim 4, wherein The saliency information data structure B c further includes a point cloud patch information field; When the field value of the point cloud patch information field is the first information value, it indicates that the saliency level parameter D c The indicated spatial range has associated point cloud patches; When the field value of the point cloud patch information field is the second information value, it indicates that the saliency level parameter D c The indicated spatial range does not have associated point cloud patches; the second information value is different from the first information value; Among them, when the field value of the point cloud patch information field is the first information value, the saliency information data structure B c further includes a point cloud patch quantity field and a point cloud patch identifier field; the point cloud patch quantity field is used to indicate the total quantity of the associated point cloud patches; the point cloud patch identifier field is used to indicate the point cloud patch identifier corresponding to the associated point cloud patches.

9. The method according to claim 8, wherein The saliency information data structure B c further includes a spatial partitioning information field; When the field value of the spatial block information field is the third information value, it indicates that the spatial range indicated by the saliency level parameter D c has an associated spatial block; When the field value of the point cloud slice information field is the fourth information value, it indicates that the saliency level parameter D c The indicated spatial range does not have an associated spatial partition; the fourth information value is different from the third information value; Wherein, when the field value of the spatial block information field is the third information value, the saliency information data structure B c further includes a spatial block quantity field and a spatial block identifier field; the spatial block quantity field is used to indicate the total quantity of the associated spatial blocks; the spatial block identifier field is used to indicate the spatial block identifier corresponding to the associated spatial blocks.

10. The method according to claim 2, characterized in that, The saliency information data box includes a saliency algorithm type field; When the field value of the saliency algorithm type field is a first type value, it indicates that the saliency information is determined by a saliency detection algorithm; When the field value of the saliency algorithm type field is a second type value, it indicates that the saliency information is determined by data statistics; the second type value is different from the first type value.

11. The method according to claim 2, wherein The method further includes: Transmitting a transmission signaling for the media file to a client; the transmission signaling carries saliency information description data; the saliency information description data is used to indicate the client to determine the acquisition order between different media sub-files in the media file when obtaining the media file through a streaming transmission method; the saliency information description data is generated based on the saliency information.

12. The method according to claim 1, characterized in that, When the field value of the saliency level indication field is the fourth indication value, it indicates that the reference saliency level parameter is used to indicate the saliency level of a time range; the time range is the point cloud frame corresponding to the sample number F g corresponding thereto; The saliency information metadata track further includes a validity range indication field for the sample number F g with a field value of a sixth indication value; the sixth indication value indicates that the reference saliency level parameter is valid within the saliency information metadata track.

13. The method according to claim 1, characterized in that, When the field value of the salience level indication field is the fifth indication value, the salience information metadata track further includes a validity range indication field for the sample serial number F g ; When the field value of the effective range indication field is the sixth indication value, it indicates that the significance level associated with the sample serial number F g becomes effective within the significance information metadata track; When the field value of the effective range indication field is the seventh indication value, it indicates that the significance level associated with the sample serial number F g is effective within the point cloud frame corresponding to the sample serial number F g ; the seventh indication value is different from the sixth indication value.

14. The method according to claim 13, characterized in that, When the field value of the effective range indication field is the seventh indication value, the saliency information metadata track further includes a sample saliency level field for the sample serial number F g ; the sample saliency level field is used to indicate the saliency level parameter of the point cloud frame corresponding to the sample serial number F g .

15. The method according to claim 1, characterized in that, When the field value of the salience level indication field is the fifth indication value, the salience information metadata track further includes a data structure quantity field for the sample serial number F g whose field value is T; the data structure quantity field is used to indicate the total quantity of the salience information data structures; T is a positive integer.

16. The method according to claim 15, wherein The T saliency information data structures include the saliency information data structure U v , where v is a positive integer and v is less than or equal to T; The saliency information data structure U v includes a saliency level field with a field value of the saliency level parameter W v and a target range indication field; the saliency level parameter W v belongs to the saliency level parameter in the saliency information When the field value of the target range indication field is the first indication value, it indicates the saliency level parameter W v used to indicate the sample sequence number F g the saliency level of a spatial region in the corresponding point cloud frame; When the field value of the target range indication field is the eighth indication value, it indicates the saliency level parameter W v used to indicate the sample serial number F g corresponding to the saliency level of the point cloud frame; the eighth indication value is different from the first indication value.

17. The method according to claim 15, characterized in that When T is a positive integer greater than 1, the saliency information data structure includes a saliency level field and a target range indication field with a field value of a first indication value; the field value of the saliency level field belongs to the saliency level parameter in the saliency information; the first indication value indicates that the field value of the saliency level field is used to indicate the saliency level of a spatial region in the point cloud frame corresponding to the sample sequence number F g corresponding to the saliency level of a spatial region in the point cloud frame corresponding to the sample sequence number F 18. The method according to claim 1, characterized in that The sample entry of the saliency information metadata track includes a saliency information data box; The saliency information data box includes a saliency algorithm type field; the saliency algorithm type field is used to indicate the determination algorithm type of the saliency information.

19. The method according to claim 1, wherein When the media file includes Z saliency information sample groups for indicating the saliency information, the target range includes the temporal range; Wherein, the total quantity of mutually different sample sequence numbers respectively included in the Z saliency information sample groups is less than or equal to H; one sample sequence number is used to indicate one point cloud frame; the media file includes H point cloud frames; H is a positive integer, and Z is a positive integer and Z is less than H.

20. The method according to claim 19, wherein The Z groups of saliency information samples include the saliency information sample group K m ; m is a positive integer and m is less than or equal to Z; the time range includes the saliency information sample group K m corresponding point cloud frame; the saliency information sample group K m corresponding point cloud frame belongs to the H point cloud frames; The saliency information sample group K m includes an effective range indication field and a data structure quantity field with a field value of I; the data structure quantity field is used to indicate the total quantity of saliency information data structures; I is a positive integer; the I saliency information data structures are used to indicate m the saliency levels associated with the saliency information sample group K When the field value of the effective range indication field is the sixth indication value, it indicates that the significance level associated with the significance information sample group K m is effective within the point cloud track corresponding to the media file; When the field value of the effective range indication field is the seventh indication value, it indicates that the significance level associated with the significance information sample group K m is effective within the significance information sample group K m ; the seventh indication value is different from the sixth indication value.

21. The method according to claim 20, characterized in that When the field value of the effective range indication field is the seventh indication value, the saliency information sample group K m further includes a sample saliency level field; the sample saliency level field is used to indicate the saliency level parameter of the point cloud frame corresponding to the saliency information sample group K m to which it corresponds.

22. The method according to claim 20, wherein The I saliency information data structures include the saliency information data structure J n , where n is a positive integer and n is less than or equal to I; The saliency information data structure J n includes a saliency level field with a field value of the saliency level parameter L n and a target range indication field; the saliency level parameter L n belongs to the saliency level parameter in the saliency information; When the field value of the target range indication field is the first indication value, it indicates the saliency level parameter L n For indicating the saliency information sample group K m The saliency level of a spatial region in the corresponding point cloud frame; When the field value of the target range indication field is the eighth indication value, it indicates the saliency level parameter L n used to indicate the saliency information sample group K m corresponding to the saliency level of the point cloud frame; the eighth indication value is different from the first indication value.

23. The method according to claim 1, characterized in that, The encoding the point cloud media to obtain a point cloud bitstream includes: According to the saliency level parameter in the saliency information, performing optimized encoding on the target range of the point cloud media to obtain a point cloud bitstream.

24. The method according to claim 23, wherein The total number of the target ranges is at least two, and the at least two target ranges include a first target range and a second target range; the significance level parameters in the significance information include a first significance level parameter corresponding to the first target range and a second significance level parameter corresponding to the second target range; Optimally encoding the target ranges of the point cloud media according to the significance level parameters in the significance information to obtain a point cloud bitstream, including: Determining a first encoding level of the first target range according to the first significance level parameter, and determining a second encoding level of the second target range according to the second significance level parameter. When the first significance level parameter is greater than the second significance level parameter, the first encoding level is better than the second encoding level; Optimally encoding the first target range through the first encoding level to obtain a first sub-point cloud bitstream, and optimally encoding the second target range through the second encoding level to obtain a second sub-point cloud bitstream; Generating the point cloud bitstream according to the first sub-point cloud bitstream and the second sub-point cloud bitstream.

25. A media data processing method, characterized in that Including: Obtaining a media file, decompressing the media file to obtain a point cloud bitstream and the significance information of the point cloud media; The significance information includes significance level parameters for indicating target ranges of the point cloud media; the target ranges include spatial ranges or temporal ranges; Decoding the point cloud bitstream to obtain the point cloud media; When the media file includes a significance information metadata track for indicating the significance information, determining that the target range includes the temporal range; wherein, the significance information metadata track includes E sample serial numbers associated with the significance information; one sample serial number corresponds to one point cloud frame; the point cloud media includes E point cloud frames; E is a positive integer; the temporal range includes the first point cloud frame among the E point cloud frames; If the first point cloud frame includes at least two spatial regions, and the saliency levels corresponding to the at least two spatial regions are different, then in the saliency information, obtain the saliency level parameter X corresponding to the first spatial region y , obtain the saliency level parameter X corresponding to the second spatial region y+1 ; both the first spatial region and the second spatial region belong to the at least two spatial regions; x is a positive integer, and x is less than the total number of saliency level parameters in the saliency information; If the saliency level parameter X y is greater than the saliency level parameter X y+1 , it is determined that the rendering level corresponding to the first spatial region is better than the rendering level corresponding to the second spatial region; If the saliency level parameter X y is less than the saliency level parameter X y+1 , it is determined that the rendering level corresponding to the second spatial region is better than the rendering level corresponding to the first spatial region.

26. The method according to claim 25, wherein The method further includes: When the media file includes a significance information data box for indicating the significance information, determining that the target range includes the spatial range; one significance level parameter in the significance information is used to indicate one spatial range; the one spatial range includes one spatial region respectively included in A point cloud frames; the point cloud media includes the A point cloud frames; A is a positive integer.

27. The method according to claim 26, wherein The total number of the spatial ranges is at least two, and the at least two spatial ranges include a first spatial range and a second spatial range; The method further includes: In the saliency information, obtain a saliency level parameter O for indicating the first spatial range p , and obtain a saliency level parameter O for indicating the second spatial range p+1 ; p is a positive integer, and p is less than the total number of saliency level parameters in the saliency information; If the saliency level parameter O p is greater than the saliency level parameter O p+1 , it is determined that the rendering level corresponding to the first spatial range is better than the rendering level corresponding to the second spatial range; If the saliency level parameter O p is less than the saliency level parameter O p+1 , then it is determined that the rendering level corresponding to the second spatial range is better than the rendering level corresponding to the first spatial range.

28. The method according to claim 25, characterized in that, The temporal range further includes the second point cloud frame among the E point cloud frames; The method further includes: In the saliency information, obtain a saliency level parameter Q for indicating the first point cloud frame r , and obtain a saliency level parameter Q for indicating the second point cloud frame r+1 ; r is a positive integer and r is less than the total number of saliency level parameters in the saliency information; If the saliency level parameter Q r is greater than the saliency level parameter Q r+1 , it is determined that the rendering level corresponding to the first point cloud frame is better than the rendering level corresponding to the second point cloud frame; If the saliency level parameter Q r is less than the saliency level parameter Q r+1 , it is determined that the rendering level corresponding to the second point cloud frame is better than the rendering level corresponding to the first point cloud frame.

29. A media data processing device, characterized in that, Including: An information determination module for determining the significance information of the point cloud media; The significance information includes significance level parameters for indicating target ranges of the point cloud media; the target ranges include spatial ranges or temporal ranges; An information encapsulation module for encoding the point cloud media to obtain a point cloud bitstream, and encapsulating the point cloud bitstream and the significance information into a media file; When the media file includes a saliency information metadata track for indicating the saliency information, the target range includes the time range; wherein, the saliency information metadata track includes E sample sequence numbers associated with the saliency information; wherein, one sample sequence number corresponds to one point cloud frame; the point cloud media includes E point cloud frames; E is a positive integer; The saliency information metadata track includes a sample serial number F g ; g is a positive integer and g is less than or equal to E; the time range includes the point cloud frame corresponding to the sample serial number F g corresponding thereto. The saliency information metadata track includes a saliency level indication field for the sample number F g ; When the field value of the salience level indication field is the fourth indication value, it indicates that the salience level associated with the sample serial number F g is determined by the reference salience level parameter; the reference salience level parameter belongs to the salience level parameters in the salience information; When the field value of the salience level indication field is the fifth indication value, it indicates that the salience level associated with the sample serial number F g is determined by the salience information data structure; the fifth indication value is different from the fourth indication value.

30. A media data processing device, characterized in that, Comprising: A file acquisition module, configured to acquire a media file, de-encapsulate the media file to obtain a point cloud bitstream and the saliency information of the point cloud media; The saliency information includes a saliency level parameter for indicating the target range of the point cloud media; the target range includes a spatial range or a time range; A bitstream decoding module, configured to decode the point cloud bitstream to obtain the point cloud media; the saliency information is used to determine the rendering effect of the target range when rendering the point cloud media; The file acquisition module is further configured to, when the media file includes a saliency information metadata track for indicating the saliency information, determine that the target range includes the time range; wherein, the saliency information metadata track includes E sample sequence numbers associated with the saliency information; wherein, one sample sequence number corresponds to one point cloud frame; the point cloud media includes E point cloud frames; E is a positive integer; the time range includes the first point cloud frame among the E point cloud frames; The file acquisition module is further configured to, if the first point cloud frame includes at least two spatial regions and the saliency levels corresponding to the at least two spatial regions are different, obtain the saliency level parameter X corresponding to the first spatial region in the saliency information y , and obtain the saliency level parameter X corresponding to the second spatial region y+1 ; both the first spatial region and the second spatial region belong to the at least two spatial regions; x is a positive integer, and x is less than the total number of saliency level parameters in the saliency information; The file acquisition module is further configured to, if the saliency level parameter X y is greater than the saliency level parameter X y+1 , determine that the rendering level corresponding to the first spatial region is better than the rendering level corresponding to the second spatial region; The file acquisition module is further configured to, if the saliency level parameter X y is less than the saliency level parameter X y+1 , determine that the rendering level corresponding to the second spatial region is better than the rendering level corresponding to the first spatial region.

31. A computer device, characterized in that, Comprising: A processor, a memory and a network interface; The processor is connected to the memory and the network interface, wherein, the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program to enable the computer device to execute the method according to any one of claims 1 to 28.

32. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1-28.

Citation Information

Patent Citations

  • Three-dimensional point cloud transmitting and receiving method and device

    CN114422791A