A data processing method, device and equipment of immersive media and a storage medium

CN116643643BActive Publication Date: 2026-09-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210138696.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-15
Publication Date
2026-09-18
Estimated Expiration
2042-02-15

AI Technical Summary

Technical Problem

自由视角视频的呈现需要根据自由视角视频的呈现指示信息来对相机采集的图像数据进行渲染实现;但是实践发现,现有技术中自由视角视频的呈现指示信息难以全面地支持自由视角视频的呈现

Benefits of technology

[0023]In this embodiment, an immersive media file is obtained. This media file contains rendering instructions for free-viewpoint video. The free-viewpoint video is generated based on image data collected by one or more camera groups. The rendering instructions for the free-viewpoint video include instructions for the camera groups. The free-viewpoint video is displayed according to these rendering instructions. Therefore, by grouping multiple cameras used to collect the corresponding image data for the free-viewpoint video and specifying how the free-viewpoint video should be rendered in the rendering instructions, the rendering of the free-viewpoint video can be more comprehensively supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116643643B_ABST
    Figure CN116643643B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a data processing method and device for immersive media, equipment and a storage medium. The method comprises: acquiring a media file of the immersive media, the media file containing presentation indication information of a free-view video, the free-view video being generated based on image data collected by one or more camera groups, the presentation indication information of the free-view video containing indication information of the camera group, and displaying the free-view video according to the presentation indication information of the free-view video. It can be seen that the multiple cameras used to collect image data corresponding to the free-view video are grouped, and the indication information of the camera group is declared in the presentation indication information of the free-view video to indicate how the free-view video is rendered, so that the presentation of the free-view video can be more comprehensively supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to the field of virtual reality (VR) technology, and more particularly to a data processing method for immersive media, a data processing device for immersive media, a computer device, and a computer-readable storage medium. Background Technology

[0002] Free-viewpoint video is an immersive media generated from image data captured by multiple cameras, containing different perspectives and supporting multi-degree-of-freedom interaction for viewers. The presentation of free-viewpoint video requires rendering the camera-captured image data according to the presentation instructions; however, in practice, it has been found that the existing presentation instructions for free-viewpoint video are insufficient to fully support its presentation. Summary of the Invention

[0003] This application provides a data processing method, apparatus, device, and storage medium for immersive media, which can more comprehensively support the presentation of free-viewpoint videos.

[0004] On one hand, embodiments of this application provide a data processing method for immersive media, including:

[0005] Acquire immersive media files containing presentation instructions for free-viewpoint video, which is generated based on image data acquired by one or more camera groups; the presentation instructions for the free-viewpoint video include instructions for the camera groups.

[0006] Display the free-view video according to the presentation instructions.

[0007] On one hand, embodiments of this application provide a data processing method for immersive media, including:

[0008] Acquire image data from one or more camera groups and encode the image data into free-viewpoint video;

[0009] Add presentation instructions to the free-view video according to its application format; the presentation instructions for the free-view video include instructions for the camera group;

[0010] Encapsulate free-viewpoint video and free-viewpoint video presentation instructions into immersive media files.

[0011] On one hand, embodiments of this application provide a data processing apparatus for immersive media, including:

[0012] The acquisition unit is used to acquire media files of immersive media. The media files contain presentation indication information of free-viewpoint video, which is generated based on image data acquired by one or more camera groups. The presentation indication information of the free-viewpoint video includes indication information of the camera groups.

[0013] The processing unit is used to display the free-view video according to the presentation instructions of the free-view video.

[0014] On one hand, embodiments of this application provide a data processing apparatus for immersive media, including:

[0015] The acquisition unit is used to acquire image data collected by one or more camera groups and encode the image data into free-viewpoint video.

[0016] The processing unit is used to add presentation instruction information to the free-view video according to the application format of the free-view video; the presentation instruction information of the free-view video includes the instruction information of the camera group;

[0017] And media files used to encapsulate free-viewpoint video and free-viewpoint video presentation instructions into immersive media.

[0018] Accordingly, this application provides a computer device, the device comprising:

[0019] A processor is used to load and execute computer programs;

[0020] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the data processing method for the immersive media described above.

[0021] Accordingly, this application provides a computer-readable storage medium storing a computer program adapted to be loaded by a processor and executed by the above-described data processing method for immersive media.

[0022] Accordingly, this application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned immersive media data processing method.

[0023] In this embodiment, an immersive media file is obtained. This media file contains rendering instructions for free-viewpoint video. The free-viewpoint video is generated based on image data collected by one or more camera groups. The rendering instructions for the free-viewpoint video include instructions for the camera groups. The free-viewpoint video is displayed according to these rendering instructions. Therefore, by grouping multiple cameras used to collect the corresponding image data for the free-viewpoint video and specifying how the free-viewpoint video should be rendered in the rendering instructions, the rendering of the free-viewpoint video can be more comprehensively supported. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1a A schematic diagram of 6DoF provided in an exemplary embodiment of this application is shown;

[0026] Figure 1b A schematic diagram of 3DoF provided in an exemplary embodiment of this application is shown;

[0027] Figure 1c A schematic diagram of 3DoF+ provided in an exemplary embodiment of this application is shown;

[0028] Figure 1d This illustration shows a schematic diagram of an immersive media process from acquisition to consumption, provided by an exemplary embodiment of this application.

[0029] Figure 1e This illustration shows a schematic diagram of a free-viewpoint video data representation method provided by an exemplary embodiment of this application;

[0030] Figure 2 A flowchart illustrating a data processing method for immersive media provided in an exemplary embodiment of this application is shown;

[0031] Figure 3 A flowchart illustrating another data processing method for immersive media provided by an exemplary embodiment of this application is shown;

[0032] Figure 4 This invention provides a schematic diagram of the structure of a data processing apparatus for immersive media according to an exemplary embodiment of this application.

[0033] Figure 5This invention provides a schematic diagram of the structure of another immersive media data processing apparatus according to an exemplary embodiment of the present application.

[0034] Figure 6 This illustration shows a structural schematic diagram of a content consumption device provided in an exemplary embodiment of this application;

[0035] Figure 7 A schematic diagram of the structure of a content creation device provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0036] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0037] The following describes some technical terms used in the embodiments of this application:

[0038] I. Immersive Media:

[0039] Immersive media refers to media files that provide immersive content, allowing viewers to experience visual, auditory, and other sensory sensations reminiscent of the real world. Based on the degree of freedom viewers have when consuming the media content, immersive media can be categorized as: 6DoF (Degree of Freedom) immersive media, 3DoF immersive media, and 3DoF+ immersive media.

[0040] II. Free-viewpoint video:

[0041] Free-viewpoint video, also known as multi-viewpoint video, is an immersive media generated from image data captured by multiple cameras, containing different perspectives and supporting multi-degree-of-freedom interaction for the viewer. For example, if a free-viewpoint video supports 3DoF+ interaction, then it is a 3DoF+ immersive media. Similarly, if a free-viewpoint video supports 6DoF interaction, then it is a 6DoF immersive media.

[0042] III. Track:

[0043] A track is a collection of media data during the media file encapsulation process. A media file can consist of one or more tracks. For example, a media file can typically contain a video track, an audio track, and a subtitle track.

[0044] IV. Sample:

[0045] A sample is a unit of encapsulation in the media file encapsulation process. A track consists of many samples. For example, a video track can consist of many samples. A sample is usually a video frame.

[0046] 5. DoF (Degree of Freedom, degree of freedom):

[0047] In this application, DoF refers to the degrees of freedom that a viewer has to move and interact with content when watching immersive media (such as free-viewpoint video), which can include 3DoF (three degrees of freedom), 3DoF+, and 6DoF (six degrees of freedom). 3DoF refers to the three degrees of freedom for the viewer's head to rotate around the x, y, and z axes. 3DoF+, in addition to the three degrees of freedom, also grants the viewer limited freedom of movement along the x, y, and z axes. 6DoF, in addition to the three degrees of freedom, also grants the viewer free freedom of movement along the x, y, and z axes.

[0048] 6. ISOBMFF (ISO Based Media File Format): This is a media file encapsulation standard. A typical ISOBMFF file is an MP4 file.

[0049] 7. DASH (Dynamic Adaptive Streaming over HTTP): This is an adaptive bitrate technology that enables high-quality streaming media to be delivered over the Internet through traditional HTTP web servers.

[0050] 8. MPD (Media Presentation Description, in DASH) is used to describe media segment information in a media file.

[0051] 9. Representation: refers to a combination of one or more media components in DASH. For example, a video file of a certain resolution can be regarded as a representation.

[0052] 10. Adaptation Sets: These are collections of one or more video streams in DASH. An Adaptation Set can contain multiple representations.

[0053] This application relates to immersive media data processing technology. Some concepts in the immersive media data processing process will be introduced below. In particular, the following embodiments of this application will use free-viewpoint video as an example for immersive media.

[0054] Figure 1a This illustration shows a schematic diagram of 6DoF provided in an exemplary embodiment of this application. 6DoF is divided into window 6DoF, omnidirectional 6DoF, and 6DoF. Window 6DoF refers to a viewer of immersive media having limited rotational movement along the X and Y axes, and limited translation along the Z axis; for example, the viewer cannot see the scene outside the window frame, and cannot pass through the window. Omnidirectional 6DoF refers to a viewer of immersive media having limited rotational movement along the X, Y, and Z axes; for example, the viewer cannot freely move through the 3D 360-degree VR content within the restricted movement area. 6DoF refers to a viewer of immersive media being able to translate freely along the X, Y, and Z axes; for example, the viewer can move freely within the 3D 360-degree VR content. Similar to 6DoF are 3DoF and 3DoF+ production techniques. Figure 1b This illustration shows a schematic diagram of 3DoF provided in an exemplary embodiment of this application; as shown Figure 1b As shown, 3DoF refers to the viewer of immersive media being fixed at the center point in a three-dimensional space, while the viewer's head rotates along the X, Y, and Z axes to view the images provided by the media content. Figure 1c A schematic diagram of 3DoF+ provided in an exemplary embodiment of this application is shown, such as... Figure 1c As shown, 3DoF+ refers to the ability of immersive media viewers to move their heads within a limited space based on 3DoF to view the images provided by the media content when the virtual scene provided by the immersive media has a certain depth information.

[0055] Figure 1d This illustration shows a schematic diagram of an immersive media process from acquisition to consumption, provided by an exemplary embodiment of this application; as shown... Figure 1d As shown, the process of immersive media from acquisition to consumption includes:

[0056] (1) Video capture: Free-viewpoint video is usually captured by a camera array consisting of multiple cameras from multiple angles to capture the same three-dimensional scene, forming the scene's texture information (color information, etc.) and depth information (spatial distance information, etc.). The content production equipment can construct the immersive media consumed by the viewer on the immersive media side (such as 3DoF immersive media, 3DoF+ immersive media, or 6DoF immersive media, etc.) based on the position information of each virtual viewpoint in the free-viewpoint video and the texture information and depth information from different cameras.

[0057] (2) After obtaining the immersive media, the content production equipment compresses and encodes the immersive media to obtain a free-viewpoint video; for example, the content production equipment can compress and encode the immersive media using AVS3 encoding technology, HEVC encoding technology, etc., to obtain a free-viewpoint video.

[0058] (3) After obtaining the free-viewpoint video, the content production device encapsulates the data stream in the free-viewpoint video. Specifically, the content production device encapsulates the audio and video streams in a file container according to the immersive media file format (such as ISOBMFF (ISOBase Media File Format)) to form an immersive media file resource. This media file resource can be a media file or a media segment forming an immersive media file. The device also records the metadata of the immersive media file resource using Media Presentation Description (MPD) according to the immersive media file format requirements. Here, metadata is a general term for information related to the presentation of immersive media. This metadata may include descriptive information about the media content, descriptive information about the window, and signaling information related to the presentation of the media content, etc.

[0059] (4) The content production device transmits the free-viewpoint video encapsulation file to the content consumption device. This transmission process can be based on various transmission protocols, including but not limited to: DASH (Dynamic Adaptive Streaming over HTTP), HLS (HTTP Live Streaming), SMTP (Smart Media Transport Protocol), TCP (Transmission Control Protocol), etc.

[0060] (5) After obtaining the free-viewpoint video encapsulation file provided by the content production device, the content consumption device decapsulates the free-viewpoint video encapsulation file. The file decapsulation process on the content consumption device side is the reverse of the file encapsulation process on the content production device side. The content consumption device decapsulates the media file resources according to the file format requirements of immersive media to obtain audio and video streams.

[0061] (6) After the free-viewpoint video's encapsulated file is decapsulated, the content consumption device decodes the free-viewpoint video to obtain immersive media. The decoding process of the video stream by the content consumption device includes the following: ① Decoding the video stream to obtain a planar projection image. ② Reconstructing the projection image based on the media presentation description information to convert it into a 3D image. Here, reconstruction refers to the process of reprojecting the two-dimensional projection image into 3D space.

[0062] (7) The content consumption device presents the corresponding immersive media based on the virtual perspective of the viewer. The content consumption device renders the audio content obtained from audio decoding and the 3D image obtained from video decoding based on metadata related to rendering and viewport in the media presentation description information. Once rendering is complete, the 3D image is played and output. Specifically, if 3DoF and 3DoF+ production techniques are used, the content consumption device mainly renders the 3D image based on the current viewpoint, parallax, and depth information. If 6DoF production techniques are used, the content consumption device mainly renders the 3D image within the viewport based on the current viewpoint. Here, viewpoint refers to the viewer's position in the immersive media; parallax refers to the difference in line of sight between the viewer's two eyes or due to motion; and viewport refers to the viewing area.

[0063] Content production equipment and content consumption equipment can together form an immersive media system. Content production equipment refers to the computer equipment used by the provider of immersive media (e.g., the content creator of immersive media). This computer equipment can be a terminal (such as a PC, a smart mobile device (such as a smartphone)) or a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Content consumption equipment refers to the computer equipment used by the user of immersive media (e.g., the viewer of immersive media). This computer equipment can be a terminal (such as a PC, a smart mobile device (such as a smartphone), VR devices (such as VR headsets, VR glasses), smart home appliances, in-vehicle terminals, aircraft, etc.).

[0064] It is understood that the data processing technology for immersive media involved in this application can be implemented using cloud technology; for example, using a cloud server as a content production device. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to realize the computation, storage, processing, and sharing of data.

[0065] In practical applications, free-viewpoint video data can be expressed in different ways. Figure 1e This illustration shows a schematic diagram of a free-viewpoint video data representation method provided by an exemplary embodiment of this application. For example... Figure 1eAs shown, in this application, the free-viewpoint video data consists of image data acquired by multiple cameras and corresponding free-viewpoint video content description metadata. That is, this application expresses free-viewpoint video data through image data acquired by multiple cameras and corresponding free-viewpoint video content description metadata. Among them, the image data acquired by multiple cameras includes: texture maps acquired by multiple cameras and depth maps corresponding to the texture maps of multiple cameras.

[0066] Under the aforementioned free-viewpoint video data expression method, the process of immersive media involved in this application from acquisition to consumption is as follows:

[0067] (1) For content production equipment, firstly, the texture map and depth map acquired by the multi-camera matrix array are obtained; then the texture map information and the corresponding depth map information acquired by the multi-camera are stitched together; it should be noted that if the free-view video needs to consume background information in the subsequent consumption process, the background information should also be stitched into the image frame; then the planar video compression method is used to encode the stitched multi-camera depth map and texture map information; and the metadata description information of the above process is encapsulated into the video file.

[0068] (2) For content consumption devices, the video file is first decapsulated; then, based on the metadata information, the depth map and texture map information of each camera are decoded from the video file; then, based on the algorithm of the free-view video application, the depth map information and texture map information are combined to synthesize the free-view virtual viewpoint requested by the viewer of the immersive media.

[0069] In practical applications, content creation devices can use data boxes to guide content consumption devices in consuming immersive media files. A data box refers to a data block or object that includes metadata; that is, the data box contains the metadata of the corresponding media content. Immersive media can include multiple data boxes, such as rotation data boxes, overlay information data boxes, media file format data boxes, etc. The presentation indication information for free-viewpoint video can be configured in the media format data box of the immersive media, for example, in the free-viewpoint data box. In streaming, the description information corresponding to the presentation indication information can be configured in the transmission signaling file of the immersive media, for example, in the free-viewpoint camera descriptor. The presentation indication information of immersive media can be configured in the free-viewpoint data box (AvsFreeViewInfoBox), content description information, and the free-viewpoint boundary data box (AvsFreeViewBoundaryBox). According to the encoding standard of immersive media (such as AVS), the syntax of the free-viewpoint data box (AvsFreeViewInfoBox) of immersive media can be seen in Table 1 below:

[0070] Table 1

[0071]

[0072]

[0073]

[0074] The semantics of the syntax shown in Table 1 above are as follows:

[0075] In the AvsFreeViewInfoBox, codec_independency indicates the encoding and decoding independence between the texture maps and depth maps corresponding to each camera within the track. See Table 2 for specific indication methods.

[0076] Table 2

[0077]

[0078]

[0079] `depth_padding_size` indicates the width of the guard band in the depth map. `texture_padding_size` indicates the width of the guard band in the texture map. `camera_count` indicates the number of cameras that acquired image data from free-viewpoint video. `ExtCameraInfoStruct()` indicates the extrinsic parameters of the cameras.

[0080] The `IntCameraInfoStruct()` function specifies the intrinsic parameters of the camera. `camera_resolution_x` indicates the resolution width of the texture and depth maps captured by the camera, and `camera_resolution_y` indicates the resolution height. `depth_downsample_factor` indicates the downsampling factor of the depth map; the actual resolution width and height of the depth map are half the resolution width and height captured by the camera (`depth_downsample_factor`).

[0081] `depth_vetex_x` indicates the x-component of the offset of the top-left vertex of the depth map relative to the origin of the plane frame (e.g., the top-left vertex of the plane frame), and `depth_vetex_y` indicates the y-component of the offset of the top-left vertex of the depth map relative to the origin of the plane frame (e.g., the top-left vertex of the plane frame). Similarly, `texture_vetex_x` indicates the x-component of the offset of the top-left vertex of the texture map relative to the origin of the plane frame (e.g., the top-left vertex of the plane frame), and `texture_vetex_y` indicates the y-component of the offset of the top-left vertex of the texture map relative to the origin of the plane frame (e.g., the top-left vertex of the plane frame).

[0082] In the camera extrinsic information structure (ExtCameraInfoStruct), `camera_pos_present` indicates whether the camera position parameter is represented. A value of 1 indicates the presence of the camera position parameter in the media file; a value of 0 indicates its absence. `camera_ori_present` indicates whether the camera orientation parameter is represented. A value of 1 indicates the presence of the camera orientation parameter in the media file; a value of 0 indicates its absence.

[0083] `camera_pos_x` indicates the x-coordinate of the camera position in the global reference coordinate system, expressed in meters; `camera_pos_y` indicates the y-coordinate of the camera position in the global reference coordinate system, expressed in meters; and `camera_pos_z` indicates the z-coordinate of the camera position in the global reference coordinate system, expressed in meters. The values ​​of `camera_pos_x`, `camera_pos_y`, and `camera_pos_z` are incremented by 2. -16 The unit is meters.

[0084] `cam_quat_x` indicates the x-component of the camera rotation quaternion, `cam_quat_y` indicates the y-component, and `cam_quat_z` indicates the z-component. The values ​​of `cam_quat_x`, `cam_quat_y`, and `cam_quat_z` are floating-point values ​​in the range [-1, 1]. When any component of the rotation information is missing, its default value is 0. The calculation rules for each component are as follows:

[0085] qX=cam_quat_x÷2 30 qY = cam_quat_y ÷ 2 30 qZ=cam_quat_z÷2 30

[0086] The fourth component qW can be derived using the mathematical properties of quaternions:

[0087] qW=Sqrt(1–(qX 2 +qY 2 +qZ 2 ))

[0088] The quaternion (w,x,y,z) represents the angle around the vector (x,y,z):

[0089] 2*cos^{-1}(w)=2*sin^{-1}(sqrt(x^{2}+y^{2}+z^{2})) is rotated.

[0090] In the camera intrinsic information structure (IntCameraInfoStruct), `camera_id` indicates the corresponding camera identifier. `camera_depth_present` indicates whether the camera depth parameter is represented. When the value of `camera_depth_present` is 1, it means that the camera depth parameter exists in the media file; when the value of `camera_depth_present` is 0, it means that the camera depth parameter does not exist in the media file. `camera_type` indicates the camera's projection algorithm type. When the value of `camera_type` is 0, it means that the projection algorithm type is ERP (Equi-Rectangular Projection); when the value of `camera_type` is 1, it means that the projection algorithm type is perspective projection; when the value of `camera_type` is 2, it means that the projection algorithm type is orthographic projection; when the value of `camera_type` is 3, it means that the projection algorithm type is a free-viewpoint pinhole camera model; when the value of `camera_type` is 4, it means that the projection algorithm type is a free-viewpoint fisheye camera model; other values ​​are reserved.

[0091] Among them, erp_horizontal_fov is used to indicate the horizontal longitude range of the viewport area when projected via ERP, in radians. The value range of the erp_horizontal_fov field is (0, 2π); erp_vertical_fov is used to indicate the vertical dimension range of the viewport area when projected via ERP, in radians. The value range of the erp_vertical_fov field is (0, π).

[0092] The `perspective_horizontal_fov` field indicates the horizontal range of the viewport area when projected using perspective, in radians. Its value range is (0, π). The `perspective_aspect_ratio` field indicates the aspect ratio (horizontal / vertical) of the viewport when projected using perspective. Its value is a 32-bit floating-point number, and the parsing process follows the IEEE 754 standard.

[0093] `ortho_horizontal_size` indicates the horizontal size of the viewport when projected orthogonally, in meters. This field is a 32-bit floating-point number, and its resolution follows the IEEE 754 standard. `ortho_aspect_ratio` indicates the aspect ratio (horizontal / vertical) of the viewport when projected orthogonally. This field is also a 32-bit floating-point number, and its resolution follows the IEEE 754 standard.

[0094] camera_focal_length_x is used to indicate the x-component of the camera focal length in a free-view pinhole camera model, and camera_focal_length_y is used to indicate the y-component of the camera focal length in a free-view pinhole camera model.

[0095] camera_principle_point_x is used to indicate the x-component of the camera principal optical axis offset in the image coordinate system of the free-view fisheye camera model, and camera_principle_point_y is used to indicate the y-component of the camera principal optical axis offset in the image coordinate system of the free-view fisheye camera model.

[0096] `camera_near_depth` indicates the near-plane depth (or distance) of the camera's associated frustum, and `camera_far_depth` indicates the far-plane depth (or distance) of the camera's associated frustum; the values ​​of `camera_near_depth` and `camera_far_depth` are incremented by 2. -16 The unit is meters.

[0097] Based on the encoding standards for immersive media (such as AVS), the syntax for the content description information of immersive media can be found in Table 3 below:

[0098] Table 3

[0099]

[0100]

[0101] Table 3 above mentions 3DPoint: x, y, z, which represent the x, z, and y coordinates of a 3D point in the Cartesian coordinate system, respectively; the semantics of the syntax shown are as follows:

[0102] Cuboid RegionStruct: cuboid_dx indicates the size of the cuboid subregion in Cartesian coordinates relative to the anchor point along the x-axis, cuboid_dy indicates the size of the cuboid subregion in Cartesian coordinates relative to the anchor point along the y-axis, and cuboid_dz ​​indicates the size of the cuboid subregion in Cartesian coordinates relative to the anchor point along the z-axis.

[0103] SpheroidStruct: radius_x indicates the radius of the spherical sub-region in the x-axis of the Cartesian coordinate system, radius_y indicates the radius of the spherical sub-region in the y-axis of the Cartesian coordinate system, and radius_z indicates the radius of the spherical sub-region in the z-axis of the Cartesian coordinate system.

[0104] HalfspaceStruct: `normal_x` indicates the planar normal of the hemispherical sub-region in the x-axis of the Cartesian coordinate system; `normal_y` indicates the planar normal of the hemispherical sub-region in the y-axis of the Cartesian coordinate system; `normal_z` indicates the planar normal of the hemispherical sub-region in the z-axis of the Cartesian coordinate system. `distance` indicates the distance from the origin along the normal vector to the plane of the hemispherical structure.

[0105] CylinderStruct: radius_x indicates the radius of the cylindrical sub-region in the x-dimensional coordinate system, radius_y indicates the radius of the cylindrical sub-region in the x-dimensional coordinate system, and height indicates the height of the cylindrical sub-region in the z-dimensional coordinate system.

[0106] The AvsFreeViewBoundaryBox is used to indicate information about the boundaries of a free-view video scene. The syntax of the AvsFreeViewBoundaryBox for immersive media, based on the immersive media coding standard (such as AVS), can be found in Table 4 below:

[0107] Table 4

[0108]

[0109]

[0110] The semantics of the AvsFreeViewBoundaryBox syntax shown in Table 4 above are as follows:

[0111] `boundary_camera_count` indicates the number of cameras constituting the boundary of the free-view video scene. `camera_id` indicates the identifier of the camera constituting the boundary of the free-view video scene; when the position, orientation, and viewport area of ​​the boundary camera have already been declared in the AvsFreeViewInfoBox, the relevant information of the boundary camera can be determined only by the camera identifier. `boundary_space_present` indicates whether there is an additional scene rendering boundary. When the value of `boundary_space_present` is 0, the scene rendering boundary is determined by the parameters of the boundary camera, and the shooting range of the boundary camera constitutes the scene rendering boundary; when the value of `boundary_space_present` is 1, the scene rendering boundary is determined by `boundary_space`, which indicates the range of the scene boundary. If the area corresponding to `boundary_space` is larger than the area formed by the boundary camera, the area outside the shooting range of the boundary camera is rendered by the content consumption device based on the content captured by the boundary camera. `bounding_space_type` indicates the type of scene boundary. The meanings of the values ​​of the `bounding_space_type` field are shown in Table 5 below.

[0112] Table 5

[0113]

[0114]

[0115] The default_origin_point field is used to indicate the origin of the scene. When the default_origin_point field is 1, it means that the scene boundary uses the coordinate origin (0,0,0) as the origin of the scene. When the default_origin_point field is 0, it means that the scene boundary uses the specified point as the origin of the scene.

[0116] The `boundary_exceeded_switch` field is used to indicate the handling method when the viewer's viewing behavior in immersive media exceeds the scene boundary. When the `boundary_exceeded_switch` field is 0, a virtual viewport region based on the origin and oriented at (0,0,0) is rendered for the viewer. When the field is 1, the content area captured by the camera corresponding to `dest_camera_id` is rendered for the viewer. When the `boundary_exceeded_switch` field is 2, a virtual viewport region based on the position and orientation specified by `dest_virtual_camera` is rendered for the viewer.

[0117] Combining Tables 1, 3, and 4 above, it can be seen that although the presentation indication information of free-viewpoint video indicates the parameter information related to the camera shooting of free-viewpoint video and the scene boundary information, there are some problems: ① It only considers the problem of a single scene boundary. For application scenarios where multiple camera groups generate free-viewpoint videos separately, the above presentation indication information is not covered. ② It does not support scenarios where ordinary video and free-viewpoint video are consumed together. In response to the above problems, this application has made corresponding improvements and proposed a new data processing method for immersive media for the specific application of free-viewpoint video. Through the technology proposed in this invention, it is possible to: (1) indicate the information of camera groups that constitute different individual scenes and the corresponding scene boundary information; (2) indicate the association relationship between ordinary video and free-viewpoint video; (3) indicate the timing metadata information of free-viewpoint video; (4) generate corresponding transmission signaling based on the information in (1) to (3) to guide the resource selection of content consumption devices in the free-viewpoint video consumption process.

[0118] The data processing flow for immersive media in this application is as follows:

[0119] For content production equipment, on the one hand, it generates a free-viewpoint video bitstream and encapsulates it into a media file. Depending on the application of the free-viewpoint video, it adds corresponding metadata information to the free-viewpoint video file, which may include at least one of the following: a) If the free-viewpoint video contains multiple independent free-viewpoint scenes, it divides the camera groups according to the different free-viewpoint scenes, indicating the camera information contained in each camera group and the scene boundary information corresponding to each camera group. b) If the free-viewpoint video is presented as an enhancement effect of ordinary video, it indicates the relationship between the ordinary video and the free-viewpoint video. c) If the free-viewpoint video is only presented at a specific time, it indicates the presentation information of the free-viewpoint video through timing metadata.

[0120] On the other hand, if the content production device supports streaming, it slices the free-viewpoint video into media segments suitable for streaming (according to the specifications of the transmission standard) and generates corresponding transmission signaling. The transmission signaling indicates at least one of the following: a) camera group information; b) if the free-viewpoint video is presented as an enhancement of the regular video, the relationship between the regular video and the free-viewpoint video; c) if the free-viewpoint video is presented only at specific times, the timing metadata is also streamed and transmitted to the content consumption device. After generating the corresponding transmission signaling, the content production device transmits the signaling file to the content consumption device.

[0121] For content consumption devices, depending on the application, they can download the complete immersive media file and play it locally; or they can establish streaming transmission with content production devices and adaptively select the appropriate free-viewpoint video stream for consumption based on the transmission signaling.

[0122] To support the above implementation process, this application adds several descriptive fields at the system layer. The following example, using extended ISOBMFF data boxes and DASH signaling, defines relevant fields to support free-view video applications. The Free-View Camera Group Data Box (AvsFreeViewCameraGroupInfoBox) indicates information about the camera group contained in the current track (i.e., the track where the Free-View Camera Group Data Box (AvsFreeViewCameraGroupInfoBox) is located, i.e., the track being parsed). The content captured by all cameras within a camera group constitutes a free-view video scene, and all cameras within a camera group share the same coordinate system. The extensions to the Free-View Camera Group Data Box (AvsFreeViewCameraGroupInfoBox) are shown in Table 6 below:

[0123] Table 6

[0124]

[0125]

[0126] The semantics of the AvsFreeViewCameraGroupInfoBox syntax shown in Table 6 above are as follows:

[0127] The same_group_info_flag field is used to indicate whether all the cameras corresponding to the image data contained in the current track belong to one or more camera groups corresponding to the current track. When the same_group_info_flag field is set to the first set value (e.g., same_group_info_flag=1), it means that all the image data contained in the current track belongs to one or more camera groups corresponding to the current track. When the same_group_info_flag field is set to the second set value (e.g., same_group_info_flag=0), it means that at least one camera among the cameras corresponding to the image data contained in the current track belongs to a camera group other than one or more camera groups corresponding to the current track. For example, suppose free-viewpoint video is encapsulated in at least one track, and the at least one track contains a first track containing image data of the free-viewpoint video. The first track corresponds to N camera groups, and each camera group contains one or more cameras, where N is a positive integer. The free-viewpoint camera group data box is encapsulated in the first track. Then, when the same camera group information flag field in the free-viewpoint camera group data box takes a first preset value, it means that all the image data contained in the first track belongs to the N camera groups corresponding to the first track. When the same camera group information flag field in the free-viewpoint camera group data box takes a second preset value, it means that among the cameras corresponding to the image data contained in the first track, at least one camera belongs to a camera group other than the N camera groups.

[0128] The `camera_group_num` field indicates the number of camera groups corresponding to the current track. The `camera_group_id` field indicates the identifier of the camera group.

[0129] Each camera group identifier corresponds to a complete camera group flag field, `complete_group_flag`. The complete camera group flag field corresponding to the current camera group (i.e., the camera group indicated by the camera group identifier field `camera_group_id`, which is the camera group being parsed) is used to indicate whether the current track contains image data acquired by all cameras in the current camera group. When the complete camera group flag field is set to the first set value (e.g., `complete_group_flag = 1`), it means that the current track contains image data acquired by all cameras in the current camera group. When the complete camera group flag field is set to the second set value (e.g., `complete_group_flag = 0`), it means that the current track contains image data acquired by some cameras in the current camera group.

[0130] Furthermore, when the same camera group information flag field is set to the second set value (same_group_info_flag = 0), the free-view camera group data box also includes a camera number field camera_num and a camera identifier field camera_id; wherein, the camera number field camera_num is used to indicate the number of all cameras constituting the current camera group; the camera identifier field camera_id is used to indicate the identifier of the camera.

[0131] The extensions to the AvsFreeViewBoundaryBox are shown in Table 7 below:

[0132] Table 7

[0133]

[0134]

[0135] The semantics of the AvsFreeViewBoundaryBox syntax shown in Table 7 above are as follows: The camera group flag field `camera_group_flag` indicates the type of display boundary corresponding to the current image data; when `camera_group_flag` is set to the first preset value (e.g., 1), it indicates that the current AvsFreeViewBoundaryBox indicates the boundary information of one or more camera groups; when `camera_group_flag` is set to the second preset value (e.g., 0), it indicates that the current AvsFreeViewBoundaryBox indicates the boundary information of the current free-view video. The camera group number field `camera_group_num` indicates the number of camera groups corresponding to the current track. The camera group identifier field `camera_group_id` indicates the identifier of the camera group. The AvsFreeViewBoundaryInfoStruct indicates scene boundary information.

[0136] In practical applications, there are scenarios involving the joint consumption of free-viewpoint and non-free-viewpoint videos. This means that viewers of immersive media are primarily presented with non-free-viewpoint videos, with free-viewpoint videos only shown in certain segments (such as highlight moments). In such cases, it is necessary to associate the non-free-viewpoint and free-viewpoint videos to indicate their relationship. In one implementation, the association between non-free-viewpoint and free-viewpoint videos can be indicated using an AvsFreeViewAssociationGroupBox. The extension of the AvsFreeViewAssociationGroupBox is shown in Table 8 below.

[0137] Table 8

[0138]

[0139]

[0140] The semantics of the free-view association group data box (AvsFreeViewAssociationGroupBox) syntax shown in Table 8 above are as follows: The free-view video flag field freeview_video_flag, when the value of freeview_video_flag is the first set value (such as 1), indicates that the current video is a free-view video; when the value of freeview_video_flag is the second set value (such as 0), it indicates that the current video is a non-free-view video.

[0141] In another implementation, the association between non-free-viewpoint videos and free-viewpoint videos can also be indicated by the AvsFreeViewAssociationBox. The AvsFreeViewAssociationBox is expanded as shown in Table 9 below:

[0142] Table 9

[0143]

[0144] The semantics of the free-view association box syntax shown in Table 9 above can be referenced from the semantics of the free-view association group box syntax shown in Table 8 above, and will not be repeated here.

[0145] In another implementation, free-viewpoint video and non-free-viewpoint video can also be associated via track indexes. Track indexes of type 'afaf' are used to indicate the indexing relationship between free-viewpoint video and associated non-free-viewpoint video, pointing from the track encapsulating the free-viewpoint video to the indexed track encapsulating the non-free-viewpoint video. That is, there is an indexing relationship between the track identifiers of the encapsulated free-viewpoint video and the non-free-viewpoint video; the indexing relationship means that the track identifier of the encapsulated free-viewpoint video can be used to index the track encapsulating the non-free-viewpoint video; or, the track identifier of the non-encapsulated free-viewpoint video can be used to index the track encapsulating the free-viewpoint video.

[0146] In one implementation, a free-view video is encapsulated across multiple video tracks, which can be linked through a free-view track group. The expansion of the free-view track group data box (AvsFreeViewGroupBox) is shown in Table 10 below:

[0147] Table 10

[0148]

[0149] The free-view track group data boxes (AvsFreeViewGroupBox) shown in Table 10 above are identified by the track group type 'a3fg'. Tracks with the same group ID belong to the same track group among all tracks containing track group type data boxes (TrackGroupTypeBox) of type 'afvg'. The semantics of the free-view track group data box (AvsFreeViewGroupBox) syntax are as follows:

[0150] The `camera_group_flag` field indicates whether multiple camera groups exist within the track group corresponding to the free-view track group data box (i.e., the current track group). When `camera_group_flag` is set to the second preset value (e.g., 0), it means that all tracks in the current track group together constitute a free-view video, and this free-view video does not contain multiple camera groups. When `camera_group_flag` is set to the second preset value (e.g., 1), it means that all tracks in the current track group together constitute a free-view video, and this free-view video contains multiple camera groups. In this case, the camera group information corresponding to the current track is indicated by the free-view camera group data box (AvsFreeViewCameraGroupInfoBox), or by relevant information in the current data box (such as the `camera_group_num` field for the number of camera groups and the `camera_group_id` field for the camera group identifier). The `camera_group_num` field indicates the number of camera groups corresponding to the current track. The `camera_group_id` field indicates the identifier of the camera group.

[0151] In another implementation, the free-viewpoint video contains multiple camera groups (i.e., the free-viewpoint video is generated based on image data acquired by multiple camera groups). The presentation time of the camera groups can be indicated by the camera group metadata track. The camera group metadata track contains camera group sample entities (CameraGroupSampleEntry), and the extensions of the camera group sample entities (CameraGroupSampleEntry) are shown in Table 11 below:

[0152] Table 11

[0153]

[0154] The semantics of the syntax for the CameraGroupSampleEntry shown in Table 11 above are as follows:

[0155] When the camera group update flag field camera_group_update_flag corresponding to the current sample is set to the first set value (e.g., 1), it means that the current sample belongs to the camera group, and the information of the camera group to which the current sample belongs is specified by the camera group identifier field (camera_group_id) corresponding to the current sample; when the camera group update flag field corresponding to the current sample is set to the second set value (e.g., 0), it means that the information of the camera group corresponding to the current sample remains unchanged.

[0156] When the camera group cancellation flag field `camera_group_cancel_flag` corresponding to the current sample is set to the first preset value (e.g., 1), it indicates that the camera group information corresponding to the target sample is cancelled, making the camera group information corresponding to the target sample no longer effective. The current sample refers to the sample being parsed; the target sample refers to any historical sample that has been parsed before the current sample. The camera group update flag field corresponding to the target sample is set to the first preset value (e.g., 1), the camera group identifier field corresponding to the target sample is the same as the camera group identifier field (`camera_group_id`) corresponding to the current sample, and the difference between the parsing time of the target sample and the parsing time of the current sample is the minimum difference between the parsing time of each historical sample and the parsing time of the current sample. For example, suppose the sequence number of the current sample is... 10. The parsed historical samples include samples with sample numbers 1-9 (the smaller the historical sample number, the earlier the parsing time of the historical sample), and the camera group identifier field corresponding to the samples with sample numbers 3, 6, and 8 is the same as the camera group identifier field corresponding to the current sample; among them, the camera group update flag field camera_group_update_flag corresponding to the historical sample with sample number 3 and the historical sample with sample number 6 has a value of 1, and the camera group update flag field camera_group_update_flag corresponding to the historical sample with sample number 8 has a value of 0; then when the camera group cancellation flag field camera_group_cancel_flag corresponding to the current sample is 1, it means that the camera group information corresponding to the historical sample with sample number 6 is no longer effective.

[0157] Specifically, when the camera group update flag field corresponding to the current sample is set to the first preset value, the camera group cancellation flag field corresponding to the current sample is not set to the first preset value; the camera group information corresponding to the current sample remains unchanged means that: if the camera group information corresponding to the target sample is in an active state, then the camera group information corresponding to the current sample is in an active state; if the camera group information corresponding to the target sample is in a canceled state, then the camera group information corresponding to the current sample is in a canceled state.

[0158] The description information corresponding to the free-view data box (AvsFreeViewInfoBox) is stored in the free-view camera descriptor (AvsFreeViewCamInfo) provided in this application embodiment. The free-view camera descriptor (AvsFreeViewCamInfo) is a supplemental property element, and its @schemeIdUri attribute is "urn:avs:ims:2018:av3f". The free-view camera descriptor (AvsFreeViewCamInfo) is included in the transmission signaling file, which is encapsulated in the representation layer of the immersive media's media presentation description file, or in the adaptive set layer; an adaptive set layer can contain one or more representation layers. When the description information exists in the target adaptive set layer, it is used to describe all representation layers in the target adaptive set layer; when the description information exists in the target representation layer, it is used to describe the target representation layer. The attributes contained in the transmission signaling file are shown in Table 12 below:

[0159] Table 12

[0160]

[0161]

[0162] In this context, O represents the corresponding attribute as Optional; CM represents the corresponding attribute as Conditional Mandatory; and M represents the corresponding attribute as Mandatory. The expanded free-view camera descriptor (AvsFreeViewCamInfo) adds: AvsFreeViewCam@camera_group_id, AvsFreeViewCam@boundary_camera_flag, and related descriptions of these elements and attributes.

[0163] The description information corresponding to the AvsFreeViewAssociationGroupBox is stored in the AvsFreeViewAssociation video association descriptor provided in this embodiment. The AvsFreeViewAssociation is a supplemental property element with the @schemeIdUri attribute "urn:avs:ims:2018:asf". The AvsFreeViewAssociation is included in the transmission signaling file, which is encapsulated in the representation level of the immersive media's media presentation description file or in the adaptive set level; an adaptive set level can contain one or more representation levels. When the description information exists in the target adaptive set level, it is used to describe all representation levels in the target adaptive set level; when the description information exists in the target representation level, it is used to describe the target representation level. The attributes contained in the transmission signaling file are shown in Table 13 below:

[0164] Table 13

[0165]

[0166] Here, M represents the corresponding attribute as Mandatory. The expanded free-view video association descriptor (AvsFreeViewAssociation) adds: AvsFreeViewAssociation, AvsFreeViewAssociation@group_id, AvsFreeViewAssociation@freeview_video_flag, and related descriptions of these elements and attributes.

[0167] According to the free-viewpoint boundary data box shown in Table 7 of this application embodiment, combined with the description of the free-viewpoint camera descriptor shown in Table 12, the camera group information of each scene contained in the free-viewpoint video and the boundary information of the scene are indicated; according to the free-viewpoint relationship group data box shown in Table 8 or the free-viewpoint relationship data box shown in Table 9 of this application embodiment, combined with the description of the free-viewpoint video association descriptor shown in Table 13, the association relationship between non-free-viewpoint video and free-viewpoint video is indicated; according to the camera group sample entity shown in Table 11 of this application embodiment, the timing metadata information of the free-viewpoint video is indicated; according to the free-viewpoint camera group data box shown in Table 6 and the free-viewpoint track group data box shown in Table 10 of this application embodiment, corresponding transmission signaling is generated to guide the resource selection of the content consumption device in the free-viewpoint video consumption process.

[0168] Figure 2 A flowchart illustrating a data processing method for immersive media provided in an exemplary embodiment of this application is shown; the method can be executed by a content consumption device in an immersive media system, and the method includes the following steps S201-S202:

[0169] S201. Obtain the media files for immersive media.

[0170] The media file contains rendering instructions for the free-viewpoint video, which specify how the video should be presented. For example, rendering instructions can indicate the boundaries of the free-viewpoint video. The free-viewpoint video is generated based on image data acquired by one or more camera groups. Each camera group contains one or more cameras. In practical applications, one or more cameras corresponding to images used to constitute the same scene in the free-viewpoint video can be grouped into the same camera group, or one or more cameras with the same internal camera parameters can be grouped into the same camera group. The image data includes texture maps and corresponding depth maps. The rendering instructions for the free-viewpoint video include camera group instructions, which indicate the number of cameras in the camera group and the camera parameters for each camera.

[0171] In one implementation, the immersive media file is sliced ​​into multiple media segments. The content consumption device acquires the transmission signaling file of the immersive media, which contains descriptive information corresponding to the presentation indication information of the free-viewpoint video. After acquiring the transmission signaling file of the immersive media, the content consumption device acquires the immersive media file based on the transmission signaling file. Specifically, the content consumption device can determine the media segment corresponding to the image data required by the viewer of the immersive media based on the virtual perspective of the viewer and the descriptive information in the transmission signaling file, and retrieve the determined media segment through streaming.

[0172] S202. Display the free-view video according to the presentation instructions for the free-view video.

[0173] The content consumption device determines the image data required for viewing from the media file based on the virtual perspective of the viewer while watching the immersive media; and decodes and displays the image data according to the instructions of the camera group.

[0174] In the first embodiment, the camera group indication information includes camera group attribute indication information. This information can be used to indicate whether the current track contains image data acquired by all cameras in the current camera group (which can be any camera group corresponding to the current track), whether all cameras corresponding to the image data in the current track belong to one or more camera groups corresponding to the current track, the number of camera groups in the current track, the identifiers of each camera group, the number of cameras in each camera group, and the identifiers of each camera in each camera group. The presentation indication information for the free-viewpoint video is metadata information, which includes a free-viewpoint camera group data box (AvsFreeViewCameraGroupInfoBox); the free-viewpoint camera group data box contains the aforementioned camera group attribute indication information.

[0175] In one embodiment, the free-viewpoint video is encapsulated in at least one track, which includes a first track containing image data of the free-viewpoint video. The first track corresponds to N camera groups, each containing one or more cameras, where N is a positive integer. A free-viewpoint camera group data box is encapsulated in the first track. The free-viewpoint camera group data box contains a same-group information flag field, same_group_info_flag. The same-group information flag field is used to indicate whether all cameras corresponding to the image data contained in the current track belong to one or more camera groups corresponding to the current track. When the same-group information flag field is set to a first preset value (e.g., same_group_info_flag = 1), it indicates that all image data contained in the current track belongs to one or more camera groups corresponding to the current track. When the same-group information flag field is set to a second preset value (e.g., same_group_info_flag = 0), it indicates that at least one camera among the cameras corresponding to the image data contained in the current track belongs to a camera group other than one or more camera groups corresponding to the current track. For example, suppose free-viewpoint video is encapsulated in at least one track, and the at least one track contains a first track containing image data of the free-viewpoint video. The first track corresponds to N camera groups, and each camera group contains one or more cameras, where N is a positive integer. The free-viewpoint camera group data box is encapsulated in the first track. Then, when the same camera group information flag field in the free-viewpoint camera group data box takes a first preset value, it means that all the image data contained in the first track belongs to the N camera groups corresponding to the first track. When the same camera group information flag field in the free-viewpoint camera group data box takes a second preset value, it means that among the cameras corresponding to the image data contained in the first track, at least one camera belongs to a camera group other than the N camera groups.

[0176] The free-view camera group data box also includes a camera group number field (camera_group_num) and a camera group identifier field (camera_group_id). The camera group number field indicates the number of camera groups corresponding to the first track, and the value of the camera group number field is N. The camera group identifier field has one value to indicate the identifier of a camera group out of N camera groups.

[0177] The free-view camera group data box also includes a complete camera group flag field, `complete_group_flag`. Each camera group identifier corresponds to a complete camera group flag field, `complete_group_flag`. The complete camera group flag field is used to indicate whether the current track contains image data acquired by all cameras in the current camera group. When the complete camera group flag field is set to the first preset value (e.g., `complete_group_flag = 1`), it means that the current track contains image data acquired by all cameras in the current camera group. When the complete camera group flag field is set to the second preset value (e.g., `complete_group_flag = 0`), it means that the current track contains image data acquired by some cameras in the current camera group.

[0178] Furthermore, when the same group information flag field is set to the second preset value (same_group_info_flag = 0), the free-view camera group data box also includes a camera number field (camera_num) and a camera identifier field (camera_id). The camera number field indicates the number of cameras in the current camera group, and its value is greater than or equal to 1. For example, if camera group A contains 3 cameras, then the camera number field (camera_num) of camera group A is 3. The camera identifier field indicates the identifier of the camera in the current camera group.

[0179] Furthermore, the free-view video contains one or more free-view scenes; in N camera groups, the image data acquired by all cameras in the same camera group belongs to the same free-view scene, and all cameras in the same camera group have the same coordinate system.

[0180] In the second embodiment, the presentation indication information for free-viewpoint video further includes free-viewpoint scene boundary indication information. The presentation indication information for free-viewpoint video is metadata information, which includes a free-viewpoint boundary data box (AvsFreeViewBoundaryBox); the free-viewpoint boundary data box contains free-viewpoint scene boundary indication information. The content consuming device displays the free-viewpoint video according to the presentation indication information, including: determining the display boundary of the image data according to the free-viewpoint scene boundary indication information during the decoding process of the image data.

[0181] In one embodiment, the free-viewpoint video is encapsulated in at least one track, which includes a second track containing image data of the free-viewpoint video. The second track corresponds to P camera groups, where P is a positive integer. A free-viewpoint boundary data box is encapsulated within the second track. The free-viewpoint boundary data box contains a camera group flag field (camera_group_flag), which indicates the type of display boundary corresponding to the image data in the second track. When the camera group flag field is set to a first preset value (e.g., camera_group_flag = 1), it indicates that the free-viewpoint boundary data box is used to indicate the boundary information corresponding to the P camera groups. When the camera group flag field is set to a second preset value (e.g., camera_group_flag = 0), it indicates that the free-viewpoint boundary data box is used to indicate the boundary information of the free-viewpoint video. Furthermore, the free-viewpoint boundary data box also contains a camera group number field (camera_group_num) and a camera group identifier field (camera_group_id). The camera group number field indicates the number of camera groups contained in the second track, and its value is P. The camera group identifier field indicates the identifier of each camera group in the second track.

[0182] In the third implementation, if the free-viewpoint video is associated with a non-free-viewpoint video that is jointly consumed, then the presentation indication information of the free-viewpoint video also includes indication information of the association relationship between the free-viewpoint video and the non-free-viewpoint video. The presentation indication information of the free-viewpoint video is metadata information, which contains a free-view relationship group data box (AvsFreeViewAssociationGroupBox); the free-view relationship group data box contains indication information of the association relationship between the free-viewpoint video and the non-free-viewpoint video.

[0183] The content consumption device displays a free-viewpoint video according to the presentation instructions of the free-viewpoint video, including: during the display of image data, switching the free-viewpoint video to a non-free-viewpoint video according to the association instructions between the free-viewpoint video and the non-free-viewpoint video; or, switching the non-free-viewpoint video to a free-viewpoint video according to the association instructions between the free-viewpoint video and the non-free-viewpoint video.

[0184] In one embodiment, free-viewpoint video is encapsulated in one or more third tracks, and non-free-viewpoint video is encapsulated in one or more fourth tracks; the third and fourth tracks have the same track group type identifier; the free-viewpoint relationship group data box includes a track group type identifier (such as afag) and a free-viewpoint video flag field freeview_video_flag; when the free-viewpoint video flag field is set to a first preset value (such as freeview_video_flag = 1), it indicates that the free-viewpoint video in the third track needs to be presented; when the free-viewpoint video flag field is set to a second preset value (such as freeview_video_flag = 0), it indicates that the non-free-viewpoint video in the fourth track needs to be presented.

[0185] In another embodiment, free-viewpoint video is encapsulated in one or more third tracks, and non-free-viewpoint video is encapsulated in one or more fourth tracks; the free-viewpoint video and the non-free-viewpoint video belong to the same entity group; the free-viewpoint relationship group data box includes an entity group identifier (entity_id, used to indicate the entity group identifier) ​​and a free-viewpoint video flag field (freeview_video_flag); when the free-viewpoint video flag field is set to a first preset value (e.g., freeview_video_flag = 1), it indicates that the free-viewpoint video of the entity corresponding to the entity group identifier in the third track needs to be presented; when the free-viewpoint video flag field is set to a second preset value (e.g., freeview_video_flag = 0), it indicates that the non-free-viewpoint video of the entity corresponding to the entity group identifier in the fourth track needs to be presented.

[0186] In another embodiment, free-viewpoint video is encapsulated in one or more third tracks, and non-free-viewpoint video is encapsulated in one or more fourth tracks; there is an indexing relationship between the track identifiers of the third tracks and the track identifiers of the fourth tracks; the indexing relationship means that the track identifier of the third track can be used to index the fourth track; or, the track identifier of the fourth track can be used to index the third track; for example, suppose the free-viewpoint video is played in the order of track identifiers 1-5, the track identifier for encapsulating the free-viewpoint video is 2, and the track identifier for encapsulating the non-free-viewpoint video is 3, then during the consumption of free-viewpoint video in track 2 by the content consumption device, the track identifier 2 can be used to index the track with identifier 3 that encapsulates the non-free-viewpoint video.

[0187] In the fourth embodiment, if the free-viewpoint video is encapsulated into multiple tracks, the presentation indication information of the free-viewpoint video also includes association indication information between the multiple tracks. The presentation indication information of the free-viewpoint video is metadata information, which includes a free-viewpoint track group data box (AvsFreeViewGroupBox), and the free-viewpoint track group data box contains association indication information between the multiple tracks.

[0188] In one embodiment, multiple tracks are clustered into M track groups, each track carrying a track group identifier; tracks with the same track group identifier belong to the same track group; M is a positive integer. Each track group corresponds to a free-view track group data box. The free-view video contains one or more free-view scenes, each free-view scene being generated based on image data acquired by one or more camera groups; the free-view track group data box contains a camera group flag field, camera_group_flag, which indicates whether there are multiple camera groups in the track group corresponding to the free-view track group data box.

[0189] When the camera group flag field is set to the first preset value (e.g., camera_group_flag = 1), it indicates that all tracks in the track group corresponding to the free-view track group data box constitute the second free-view scene, which is generated based on image data acquired by multiple camera groups. When the camera group flag field is set to the second preset value (e.g., camera_group_flag = 0), it indicates that all tracks in the track group corresponding to the free-view track group data box constitute the first free-view scene, which is generated based on image data acquired by one camera group. The free-view track group data box also includes a camera group number field (camera_group_num) and a camera group identifier field (camera_group_id). The camera group number field indicates the number of camera groups in the track group corresponding to the free-view track group data box, and its value is greater than or equal to 1. Specifically, when camera_group_flag = 1, camera_group_num > 1; when camera_group_flag = 0, camera_group_num = 1. The camera group identifier field indicates the identifier of each camera group in the track group corresponding to the free-view track group data box.

[0190] In the fifth embodiment, the free-viewpoint video is generated based on image data acquired by multiple camera groups. The camera group indication information also includes presentation time indication information for each camera group. The presentation indication information of the free-viewpoint video is metadata information, which includes a camera group sample entity (CameraGroupSampleEntry), and the camera group sample entity contains the presentation time indication information for each camera group. The content consumption device displays the free-viewpoint video according to the presentation indication information of the free-viewpoint video, including: sequentially decoding and displaying the image data corresponding to the corresponding camera group in the image data according to the presentation time indication information for each camera group.

[0191] In one embodiment, the camera group sample entity includes a camera group update flag field camera_group_update_flag, a camera group identifier field camera_group_id, and a camera group cancellation flag field camera_group_cancel_flag;

[0192] When the camera group update flag field camera_group_update_flag corresponding to the current sample is set to the first set value (e.g., 1), it means that the current sample belongs to the camera group, and the information of the camera group to which the current sample belongs is specified by the camera group identifier field (camera_group_id) corresponding to the current sample; when the camera group update flag field corresponding to the current sample is set to the second set value (e.g., 0), it means that the information of the camera group corresponding to the current sample remains unchanged.

[0193] When the camera group cancellation flag field `camera_group_cancel_flag` corresponding to the current sample is set to the first preset value (e.g., 1), it indicates that the camera group information corresponding to the target sample is cancelled, making the camera group information corresponding to the target sample no longer effective. The current sample refers to the sample being parsed; the target sample refers to any historical sample that has been parsed before the current sample. The camera group update flag field corresponding to the target sample is set to the first preset value (e.g., 1), the camera group identifier field corresponding to the target sample is the same as the camera group identifier field (`camera_group_id`) corresponding to the current sample, and the difference between the parsing time of the target sample and the parsing time of the current sample is the minimum difference between the parsing time of each historical sample and the parsing time of the current sample. For example, suppose the sequence number of the current sample is... 10. The parsed historical samples include samples with sample numbers 1-9 (the smaller the historical sample number, the earlier the parsing time of the historical sample), and the camera group identifier field corresponding to the samples with sample numbers 3, 6, and 8 is the same as the camera group identifier field corresponding to the current sample; among them, the camera group update flag field camera_group_update_flag corresponding to the historical sample with sample number 3 and the historical sample with sample number 6 has a value of 1, and the camera group update flag field camera_group_update_flag corresponding to the historical sample with sample number 8 has a value of 0; then when the camera group cancellation flag field camera_group_cancel_flag corresponding to the current sample is 1, it means that the camera group information corresponding to the historical sample with sample number 6 is no longer effective.

[0194] Specifically, when the camera group update flag field corresponding to the current sample is set to a first preset value, the camera group cancellation flag field corresponding to the current sample is not set to the first preset value; the camera group information corresponding to the current sample remains unchanged means that: if the camera group information corresponding to the target sample is in an active state, then the camera group information corresponding to the current sample is in an active state; if the camera group information corresponding to the target sample is in a canceled state, then the camera group information corresponding to the current sample is in a canceled state. In the seventh embodiment, the immersive media file is sliced ​​into multiple media segments; the content consumption device acquires the transmission signaling file of the immersive media, and the transmission signaling file contains descriptive information corresponding to the presentation indication information of the free-viewpoint video.

[0195] In one embodiment, the free-view video includes one or more free-view scenes; the transmission signaling file includes at least one of the following: a free-view camera descriptor and a free-view association descriptor; wherein, the free-view camera descriptor is used to indicate the parameters of each camera in one or more camera groups (such as projection method, camera position, etc.); the free-view association descriptor is used to indicate the association relationship between the free-view video and the corresponding non-free-view video.

[0196] In another embodiment, the transmission signaling file includes a representation layer and an adaptive set layer, wherein the adaptive set layer includes one or more representation layers; when description information exists in the target adaptive set layer, it is used to describe all representation layers in the target adaptive set layer; when description information exists in the target representation layer, it is used to describe the target representation layer; wherein the description information may include one or more of the free-view camera descriptors and free-view association descriptors.

[0197] This application extends the data box and media presentation description file of immersive media. The data processing method for immersive media includes: acquiring the media file of immersive media, the media file containing presentation indication information for free-viewpoint video, the free-viewpoint video being generated based on image data acquired by one or more camera groups, the presentation indication information of the free-viewpoint video containing the indication information of the camera groups, and displaying the free-viewpoint video according to the presentation indication information of the free-viewpoint video. It can be seen that by grouping multiple cameras corresponding to the free-viewpoint video and instructing how the free-viewpoint video to be rendered by declaring the indication information of the camera groups in the presentation indication information of the free-viewpoint video, the presentation of free-viewpoint video can be more comprehensively supported. Furthermore, the display boundary of the free-viewpoint scene is determined by the indication of the free-viewpoint boundary data box; the display of free-viewpoint video or non-free-viewpoint video is determined by the indication of the viewpoint relationship group data box or the free-viewpoint relationship data box.

[0198] Figure 3 A flowchart illustrating another data processing method for immersive media provided by an exemplary embodiment of this application is shown; the method can be performed by a content creation device in an immersive media system, and the method includes the following steps S301-S303:

[0199] S301. Acquire image data from one or more camera groups and encode the image data into free-viewpoint video.

[0200] For a detailed implementation of step S301, please refer to Figure 1d The implementation methods of medium-length video encoding will not be described in detail here.

[0201] S302. Add presentation instruction information to the free-view video according to its application format.

[0202] The presentation indication information for free-viewpoint video includes camera group indication information. Content production equipment adds presentation indication information to the free-viewpoint video based on its application format. This is the reverse process of content consumption equipment displaying the free-viewpoint video based on its presentation indication information. For details, please refer to [reference needed]. Figure 2 The process by which the content consumption device displays the free-view video based on the presentation instructions for the free-view video will not be described in detail here. Taking the free-view camera group data box as an example, assuming that the free-view video is generated based on image data collected by a camera group, and the first track includes image data collected by all cameras in the camera group, then the content production device will configure the value of the complete_group_flag field of the free-view camera group data box to 1.

[0203] S303. Encapsulate the free-viewpoint video and the free-viewpoint video presentation instruction information into an immersive media file.

[0204] For a detailed implementation of step S303, please refer to Figure 1d The implementation methods for encapsulating medium-length video files will not be described in detail here.

[0205] In one embodiment, the content production device supports streaming transmission. After obtaining the media file of the immersive media, the content production device slices the media file to obtain multiple media segments; and generates a transmission signaling file for the immersive media, the transmission signaling file containing descriptive information corresponding to the presentation indication information of the free-viewpoint video; wherein the descriptive information may include at least one of the following: a free-viewpoint camera descriptor and a free-viewpoint association descriptor.

[0206] The free-view camera descriptor is used to indicate the parameters of each camera in one or more camera groups; the free-view correlation descriptor is used to indicate the correlation between free-view video and the corresponding non-free-view video; see Tables 12 and 13 above for details, which will not be repeated here.

[0207] The following two complete examples illustrate the data processing method for immersive media provided in this application:

[0208] In one embodiment, free-viewpoint video consists of images captured by a single camera array.

[0209] Content production equipment: generates free-viewpoint video bitstreams and encapsulates these bitstreams into media files. Based on the application of the free-viewpoint video, it adds corresponding metadata information (i.e., presentation indication information) to the free-viewpoint video file, including:

[0210] Indicates parameters and related information for the camera being used;

[0211] AvsFreeViewInfoBox:

[0212] {Camera1: ID=1; Pos=(100,0,100); orientation=(0,0,0)};

[0213] {Camera2: ID=2; Pos=(100,100,100); orientation=(0.5,0.5,0)};

[0214] {Camera3: ID=3; Pos=(0,0,100); orientation=(0.5,0.5,-0.5)}.

[0215] Furthermore, if the content production device supports streaming, the media file is sliced ​​into media segments suitable for streaming (according to the specifications of the transmission standard), and corresponding transmission signaling is generated, which indicates the following information:

[0216] Indicates camera parameter information and camera group information;

[0217] Fill in the corresponding fields in the free-view camera descriptor (AvsFreeViewCamInfo) in the DASH signaling according to the values ​​of the data boxes in the metadata information mentioned above.

[0218] After completing the above steps, the content creation device will transmit the signaling file to the content consumption device.

[0219] Content consumption device: In one implementation, the content consumption device downloads the complete file and plays it locally. During playback, it selects the appropriate texture map and depth map corresponding to the camera based on the parameters of each camera in the AvsFreeViewInfoBox and the virtual perspective of the viewer, decodes them, and synthesizes the corresponding image.

[0220] In another implementation, the content consumption device establishes a streaming transmission with the content production device. The content requests the appropriate video streams of texture maps and depth maps corresponding to the appropriate cameras based on the parameters of each camera in the AvsFreeViewCamInfo of the transmission signaling and the virtual perspective of the viewer. After receiving the corresponding video streams, the corresponding images are synthesized.

[0221] In another embodiment, the free-viewpoint video consists of images captured by multiple camera groups.

[0222] Content production equipment: generates free-viewpoint video bitstreams and encapsulates these bitstreams into media files. Based on the application of the free-viewpoint video, it adds relevant metadata information to the free-viewpoint video file, including:

[0223] Indicates parameters and related information for the shooting camera; indicates camera group information;

[0224] Track 1:

[0225] AvsFreeViewInfoBox:

[0226] {Camera1:ID=1;Pos=(100,0,100); orientation=(0,0,0)}

[0227] {Camera2: ID=2; Pos=(100,100,100); orientation=(0.5,0.5,0)}

[0228] AvsFreeViewCameraGroupInfoBox:

[0229] {camera_group_num=1; camera_group_id=1; complete_group_flag=1; same_group_info_flag=1}

[0230] AvsFreeViewGroupBox:

[0231] {camera_group_flag=1}

[0232] Track2:

[0233] AvsFreeViewInfoBox:

[0234] {Camera3: ID=3; Pos=(0,0,100); orientation=(0.5,0.5,-0.5)}

[0235] {Camera4: ID=4; Pos=(0,100,0); orientation=(0.5,0.5,0)}

[0236] AvsFreeViewCameraGroupInfoBox:

[0237] {camera_group_num=1; camera_group_id=2; complete_group_flag=1; same_group_info_flag=1}

[0238] AvsFreeViewGroupBox:

[0239] {camera_group_flag=1}

[0240] Track3 (Metadata Track): Metadata track for the free-view video camera group.

[0241] Furthermore, if the content production device supports streaming, the media file is sliced ​​into media segments suitable for streaming (according to the specifications of the transmission standard), and corresponding transmission signaling is generated, which indicates the following information:

[0242] Fill in the corresponding fields of the free-view camera descriptor (AvsFreeViewCamInfo) in the DASH signaling according to the values ​​of the data boxes in the above metadata information, including the camera group identifier field camera_group_id.

[0243] After completing the above steps, the content creation device will transmit the signaling file to the content consumption device.

[0244] Content Consumption Device: In one implementation, the content consumption device downloads the complete file and plays it locally. During playback, the camera group to be played at a specific time is determined based on the information indicated in the free-view video camera group metadata track. The video data track corresponding to a specific camera group is determined based on the free-view camera group data box (AvsFreeViewCameraGroupInfoBox) and the free-view track group data box (AvsFreeViewGroupBox). Then, based on the parameters of each camera in the free-view data box (AvsFreeViewInfoBox) and the viewer's virtual perspective, the appropriate texture map and depth map corresponding to the camera are selected for decoding and composited into the corresponding image. Specifically, camera_group_num = 1 indicates that the number of camera groups in track 1 is 1; camera_group_id = 1 indicates that the identifier of the camera group in track 1 is 1; complete_group_flag = 1 indicates that track 1 contains image data acquired by all cameras in camera group 1; same_group_info_flag = 1 indicates that the image data contained in track 1 all belong to camera group 1 corresponding to track 1. `camera_group_flag = 1` indicates that the current free-view boundary data box is used to indicate the boundary information corresponding to the current camera group. Track 2 is similar to Track 1, and will not be described further here.

[0245] In another implementation, the content consuming device establishes a streaming transmission, requests data from the metadata track of the free-view video camera group, and determines the camera group to be played at a specific time based on the information therein. The content consuming device selects the appropriate video streams for the texture and depth maps corresponding to the cameras based on the parameters of each camera in the transmission signaling free-view camera descriptor (AvsFreeViewCamInfo) and the virtual viewpoint of the viewer. Upon receiving the corresponding video streams, the device synthesizes the corresponding images.

[0246] This application extends the data box and media rendering description file of immersive media. The data processing method for immersive media includes: acquiring image data collected by one or more camera groups and encoding the image data into free-viewpoint video; adding rendering instruction information to the free-viewpoint video according to its application format; the rendering instruction information of the free-viewpoint video includes the instruction information of the camera groups; and encapsulating the free-viewpoint video and its rendering instruction information into a media file for immersive media. It is evident that by grouping multiple cameras corresponding to the free-viewpoint video and specifying how the free-viewpoint video should be rendered by declaring the instruction information of the camera groups in the rendering instruction information of the free-viewpoint video, the rendering of free-viewpoint video can be more comprehensively supported. Furthermore, the camera group information of each scene contained in the free-viewpoint video and the boundary information of that scene are indicated through the free-viewpoint boundary data box; the relationship between non-free-viewpoint video and free-viewpoint video is indicated through the viewpoint relationship group data box or the free-viewpoint relationship data box; the timing metadata information of the free-viewpoint video is indicated through the camera group sample entity; and the corresponding transmission signaling is generated through the free-viewpoint camera group data box and the free-viewpoint track group data box to guide the resource selection of content consumption devices during the free-viewpoint video consumption process.

[0247] The methods of the embodiments of this application have been described in detail above. In order to facilitate better implementation of the above solutions of the embodiments of this application, the apparatus of the embodiments of this application is provided below.

[0248] Please see Figure 4 , Figure 4 This illustration shows a schematic diagram of the structure of a data processing apparatus for immersive media according to an exemplary embodiment of this application; the data processing apparatus for immersive media can be a computer program (including program code) running on a content consumption device, for example, the data processing apparatus for immersive media can be application software in the content consumption device. Figure 4 As shown, the immersive media data processing device includes an acquisition unit 401 and a processing unit 402.

[0249] Please see Figure 4 The detailed descriptions of each unit are as follows:

[0250] The acquisition unit 401 is used to acquire the media file of the immersive media. The media file contains presentation indication information of the free-view video. The free-view video is generated based on image data acquired by one or more camera groups. The presentation indication information of the free-view video contains the indication information of the camera groups.

[0251] The processing unit 402 is used to display the free-view video according to the presentation instruction information of the free-view video.

[0252] In one implementation, the camera group indication information includes camera group attribute indication information; the free-viewpoint video presentation indication information is metadata information, which includes a free-viewpoint camera group data box; the free-viewpoint camera group data box includes camera group attribute indication information.

[0253] In one implementation, the free-viewpoint video is encapsulated in at least one track, the at least one track includes a first track, the first track contains image data of the free-viewpoint video; the first track corresponds to N camera groups, each camera group contains one or more cameras, N is a positive integer; a free-viewpoint camera group data box is encapsulated in the first track; the free-viewpoint camera group data box also contains the same camera group information flag field.

[0254] When the same camera group information flag field is set to the first set value, it means that all the image data contained in the first track belong to N camera groups; when the same camera group information flag field is set to the second set value, it means that among the cameras corresponding to the image data contained in the first track, at least one camera belongs to a camera group other than N camera groups.

[0255] In one implementation, the free-view camera group data box includes a camera group number field and a camera group identifier field. The camera group number field is used to indicate the number of camera groups corresponding to the first track, and the value of the camera group number field is N. One value of the camera group identifier field is used to indicate the identifier of one of the N camera groups.

[0256] In one embodiment, the free-view camera group data box further includes a complete camera group flag field; when the complete camera group flag field is set to a first preset value, it indicates that the image data acquired by all cameras in the camera group indicated by the camera group flag field is included in the first track; when the complete camera group flag field is set to a second preset value, it indicates that the image data acquired by some cameras in the camera group indicated by the camera group flag field is included in the first track.

[0257] In one implementation, when the same camera group information flag field is set to a second preset value, the free-view camera group data box also includes a camera number field and a camera identification field. The camera number field is used to indicate the number of cameras contained in the camera group indicated by the camera group identification field, and the value of the camera number field is greater than or equal to 1. The camera identification field is used to indicate the identifier of the camera in the camera group indicated by the camera group identification field.

[0258] In one implementation, the free-view video includes one or more free-view scenes; in N camera groups, the image data acquired by all cameras in the same camera group belongs to the same free-view scene, and all cameras in the same camera group have the same coordinate system.

[0259] In one implementation, the camera group indication information also includes camera parameter indication information in the camera group; the metadata information also includes a free-view data box, which contains camera parameters and repeating camera parameter indication fields.

[0260] When the repeat camera parameter indicator field is set to the first setting value, it means that the camera parameters in the free-view data box are effective for all cameras; when the repeat camera parameter indicator field is set to the second setting value, it means that different cameras in the free-view data box have their own camera parameters.

[0261] In one implementation, the presentation indication information for free-view video also includes free-view scene boundary indication information.

[0262] In one implementation, the presentation indication information of the free-view video is metadata information, which includes a free-view boundary data box; the free-view boundary data box includes free-view scene boundary indication information.

[0263] In one implementation, the free-viewpoint video is encapsulated in at least one track, the at least one track contains a second track, the second track contains image data of the free-viewpoint video; the second track corresponds to P camera groups, where P is a positive integer; the free-viewpoint boundary data box is encapsulated in the second track;

[0264] The free-view boundary data box contains a camera group flag field, which indicates the type of display boundary corresponding to the image data in the second track. When the camera group flag field is set to a first preset value, it means that the free-view boundary data box is used to indicate the boundary information corresponding to P camera groups. When the camera group flag field is set to a second preset value, it means that the free-view boundary data box is used to indicate the boundary information of the free-view video.

[0265] The free-view boundary data box also includes a camera group number field and a camera group identifier field. The camera group number field indicates the number of camera groups contained in the second track, and the value of the camera group number field is P. The camera group identifier field indicates the identifier of each camera group in the second track.

[0266] In one implementation, if a free-viewpoint video is associated with a non-free-viewpoint video that is jointly consumed, then the presentation indication information of the free-viewpoint video also includes indication information of the association between the free-viewpoint video and the non-free-viewpoint video.

[0267] In one implementation, the presentation indication information for the free-viewpoint video is metadata information, which includes a free-viewpoint relationship group data box; the free-viewpoint relationship group data box includes association indication information between the free-viewpoint video and the non-free-viewpoint video.

[0268] In one implementation, free-viewpoint video is encapsulated in one or more third tracks, and non-free-viewpoint video is encapsulated in one or more fourth tracks; the third tracks and the fourth tracks have the same track group type identifier.

[0269] The free-viewpoint relationship group data box includes a track group type identifier and a free-viewpoint video flag field. When the free-viewpoint video flag field is set to the first set value, it indicates that the free-viewpoint video in the third track needs to be presented. When the free-viewpoint video flag field is set to the second set value, it indicates that the non-free-viewpoint video in the fourth track needs to be presented.

[0270] In one implementation, the free-viewpoint video is encapsulated in one or more third tracks, and the non-free-viewpoint video is encapsulated in one or more fourth tracks; the free-viewpoint video and the non-free-viewpoint video belong to the same entity group.

[0271] The free-viewpoint relationship group data box includes the entity group identifier and the free-viewpoint video flag field. When the free-viewpoint video flag field is set to the first set value, it means that the free-viewpoint video of the entity corresponding to the entity group identifier in the third track needs to be presented. When the free-viewpoint video flag field is set to the second set value, it means that the non-free-viewpoint video of the entity corresponding to the entity group identifier in the fourth track needs to be presented.

[0272] In one implementation, free-viewpoint video is encapsulated in one or more third tracks, and non-free-viewpoint video is encapsulated in one or more fourth tracks; there is an index relationship between the track identifiers of the third tracks and the track identifiers of the fourth tracks.

[0273] An index relationship means that a third track can be indexed to a fourth track; or, a fourth track can be indexed to a third track.

[0274] In one implementation, if the free-viewpoint video is encapsulated into multiple tracks, the presentation indication information of the free-viewpoint video also includes association indication information between the multiple tracks.

[0275] In one implementation, the presentation indication information of the free-viewpoint video is metadata information, which includes a free-viewpoint track group data box, and the free-viewpoint track group data box contains association indication information between multiple tracks.

[0276] In one implementation, multiple orbits are clustered into M orbital groups, each orbital carrying an orbital group identifier; orbitals with the same orbital group identifier belong to the same orbital group; M is a positive integer.

[0277] In one implementation, each track group corresponds to a free-view track group data box, and the free-view video contains one or more free-view scenes. Each free-view scene is generated based on image data acquired by one or more camera groups. The free-view track group data box contains a camera group flag field, which is used to indicate whether there are multiple camera groups in the track group corresponding to the free-view track group data box.

[0278] When the camera group flag field is set to the first set value, it means that all tracks in the track group corresponding to the free-view track group data box constitute the first free-view scene, which is generated based on image data acquired by one camera group; when the camera group flag field is set to the second set value, it means that all tracks in the track group corresponding to the free-view track group data box constitute the second free-view scene, which is generated based on image data acquired by multiple camera groups.

[0279] The free-view track group data box also includes a camera group number field and a camera group identifier field. The camera group number field indicates the number of camera groups in the track group corresponding to the free-view track group data box, and the value of the camera group number field is greater than or equal to 1. The camera group identifier field indicates the identifier of each camera group in the track group corresponding to the free-view track group data box.

[0280] In one implementation, when the free-viewpoint video is generated based on image data acquired by multiple camera groups, the camera group indication information also includes separate presentation time indication information for each camera group.

[0281] In one implementation, the presentation indication information of the free-viewpoint video is metadata information, which includes camera group sample entities, and the camera group sample entities contain separate presentation time indication information for each camera group.

[0282] In one implementation, the camera group sample entity includes a camera group update flag field, a camera group identifier field, and a camera group cancellation flag field;

[0283] When the camera group update flag field corresponding to the current sample is set to the first preset value, it means that the current sample belongs to the camera group, and the information of the camera group to which the current sample belongs is specified by the camera group identifier field corresponding to the current sample; when the camera group update flag field corresponding to the current sample is set to the second preset value, it means that the information of the camera group corresponding to the current sample remains unchanged.

[0284] When the camera group cancellation flag field corresponding to the current sample is set to the first preset value, it indicates that the camera group information corresponding to the target sample is cancelled, and the camera group information corresponding to the target sample is no longer effective; the current sample refers to the sample being parsed; the target sample refers to any of the historical samples that have been parsed before the current sample. The camera group update flag field corresponding to the target sample is set to the first preset value. The camera group identifier field corresponding to the target sample is the same as the camera group identifier field corresponding to the current sample. Furthermore, the difference between the parsing time of the target sample and the parsing time of the current sample is the minimum value among the differences between the parsing time of each sample in the historical samples and the parsing time of the current sample.

[0285] Specifically, when the camera group update flag field corresponding to the current sample is set to the first preset value, the camera group cancellation flag field corresponding to the current sample is not set to the first preset value; the camera group information corresponding to the current sample remains unchanged means that: if the camera group information corresponding to the target sample is in an active state, then the camera group information corresponding to the current sample is in an active state; if the camera group information corresponding to the target sample is in a canceled state, then the camera group information corresponding to the current sample is in a canceled state.

[0286] In one implementation, the immersive media file is sliced ​​into multiple media segments; the acquisition unit 401 is further configured to:

[0287] Obtain the transmission signaling file of the immersive media, which contains descriptive information corresponding to the presentation instructions for the free-viewpoint video.

[0288] In one implementation, the free-view video includes one or more free-view scenes; the transmission signaling file includes at least one of the following: a free-view camera descriptor and a free-view association descriptor;

[0289] Among them, the free-viewpoint camera descriptor is used to indicate the parameters of each camera in one or more camera groups; the free-viewpoint association descriptor is used to indicate the association relationship between free-viewpoint video and the corresponding non-free-viewpoint video.

[0290] In one implementation, the transmission signaling file includes a representation layer and an adaptive set layer, wherein the adaptive set layer includes one or more representation layers; when description information exists in the target adaptive set layer, it is used to describe all representation layers in the target adaptive set layer; when description information exists in the target representation layer, it is used to describe the target representation layer.

[0291] In one embodiment, the processing unit 402 is configured to acquire the media file of the immersive media, specifically for:

[0292] Based on the virtual perspective of the viewer of immersive media and the description information in the transmission signaling file, determine the media segment corresponding to the image data required for viewing;

[0293] The determined media segments are retrieved using streaming.

[0294] In one embodiment, the processing unit 402 is configured to display the free-viewpoint video according to the presentation instruction information of the free-viewpoint video, specifically configured to:

[0295] Based on the virtual perspective of the viewer while watching immersive media, determine the image data required for viewing from the media file;

[0296] The image data is decoded and displayed according to the instructions from the camera group.

[0297] In one embodiment, the presentation indication information of the free-viewpoint video further includes free-viewpoint scene boundary indication information; the processing unit 402 is configured to display the free-viewpoint video according to the presentation indication information, specifically for:

[0298] During the decoding process of image data, the display boundaries of the image data are determined according to the free-view scene boundary indication information.

[0299] In one implementation, if the free-viewpoint video is associated with a non-free-viewpoint video for joint consumption, the presentation indication information of the free-viewpoint video further includes association indication information between the free-viewpoint video and the non-free-viewpoint video; the processing unit 402 is configured to display the free-viewpoint video according to the presentation indication information of the free-viewpoint video, specifically configured to:

[0300] During the display of image data, the free-viewpoint video is switched to a non-free-viewpoint video according to the association information between free-viewpoint and non-free-viewpoint videos; or,

[0301] Based on the association information between free-view and non-free-view videos, switch the non-free-view video to a free-view video.

[0302] In one embodiment, the image data is acquired by multiple camera groups; the camera group indication information also includes presentation time indication information for each camera group; the processing unit 402 is used to decode and display the image data according to the camera group indication information, specifically for:

[0303] According to the presentation time indication information of each camera group, the image data corresponding to the corresponding camera group in the image data is decoded and displayed sequentially.

[0304] According to one embodiment of this application, Figure 2 The data processing method for immersive media shown can involve some steps that can be derived from... Figure 4 The data processing is performed by individual units within the immersive media data processing device shown. For example, Figure 2 Step S201 shown can be performed by Figure 4 The acquisition unit 401 shown is executed, and step S202 can be performed by... Figure 4 The processing unit 402 shown executes. Figure 4 The various units in the immersive media data processing apparatus shown can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the immersive media data processing apparatus may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0305] According to another embodiment of this application, the following can be executed by running on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 2 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 4 The data processing apparatus for immersive media shown herein, and the data processing method for immersive media for implementing embodiments of this application, are described. A computer program may be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the computer-readable recording medium, and executed therein.

[0306] Based on the same inventive concept, the data processing device for immersive media provided in the embodiments of this application solves the problem in a similar principle and with similar beneficial effects as the data processing method for immersive media in the embodiments of this application. For details, please refer to the implementation principle and beneficial effects of the method. For the sake of brevity, these will not be repeated here.

[0307] Please see Figure 5 , Figure 5 This illustration shows a schematic diagram of another immersive media data processing apparatus provided in an exemplary embodiment of this application; the immersive media data processing apparatus can be a computer program (including program code) running on a content production device, for example, the immersive media data processing apparatus can be application software in the content production device. Figure 5As shown, the immersive media data processing apparatus includes an acquisition unit 501 and a processing unit 502. (See also...) Figure 5 The detailed descriptions of each unit are as follows:

[0308] The acquisition unit 501 is used to acquire image data collected by one or more camera groups and encode the image data into free-viewpoint video.

[0309] The processing unit 502 is used to add presentation instruction information to the free-view video according to the application form of the free-view video; the presentation instruction information of the free-view video includes the instruction information of the camera group; and to encapsulate the free-view video and the presentation instruction information of the free-view video into a media file for immersive media.

[0310] In one embodiment, the processing unit 502 is further configured to:

[0311] Slicing a media file to obtain multiple media segments; and,

[0312] Generate a transmission signaling file for immersive media, which contains descriptive information corresponding to the presentation instructions for free-viewpoint video.

[0313] According to one embodiment of this application, Figure 3 The data processing method for immersive media shown can involve some steps that can be derived from... Figure 5 The data processing is performed by individual units within the immersive media data processing device shown. For example, Figure 3 Step S301 shown can be performed by Figure 5 The acquisition unit 501 shown is executed, and steps S302 and S303 can be performed by... Figure 5 The processing unit 502 shown is executed. Figure 5 The various units in the immersive media data processing apparatus shown can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the immersive media data processing apparatus may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0314] According to another embodiment of this application, the following can be executed by running on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 3 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 5 The data processing apparatus for immersive media shown herein, and the data processing method for immersive media for implementing embodiments of this application, are described. A computer program may be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the computer-readable recording medium, and executed therein.

[0315] Based on the same inventive concept, the data processing device for immersive media provided in the embodiments of this application solves the problem in a similar principle and with similar beneficial effects as the data processing method for immersive media in the embodiments of this application. For details, please refer to the implementation principle and beneficial effects of the method. For the sake of brevity, these will not be repeated here.

[0316] Figure 6 This illustration shows a schematic diagram of the structure of a content consumption device provided in an exemplary embodiment of this application; the content consumption device can refer to a computer device used by a user of immersive media, and the computer device can be a terminal (such as a PC, a smart mobile device (such as a smartphone), a VR device (such as a VR headset, VR glasses, etc.)). Figure 6 As shown, the content consumption device includes a receiver 601, a processor 602, a memory 603, and a display / playback device 604. Wherein:

[0317] Receiver 601 is used to enable decoding and transmission interaction with other devices, specifically for the transmission of immersive media between content creation devices and content consumption devices. That is, the content consumption device receives the relevant media resources for immersive media transmitted by the content creation device through receiver 601.

[0318] Processor 602 (or CPU (Central Processing Unit)) is the processing core of the content production device. Processor 602 is adapted to implement one or more program instructions, specifically to load and execute one or more program instructions to achieve... Figure 2 The flowchart illustrates the data processing method for immersive media.

[0319] Memory 603 is a memory device in the content consumption device used to store programs and media resources. It is understood that memory 603 here can include the built-in storage medium of the content consumption device, or it can include extended storage media supported by the content consumption device. It should be noted that memory 603 can be high-speed RAM, or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one memory located remotely from the aforementioned processor. Memory 603 provides storage space for storing the operating system of the content consumption device. Furthermore, this storage space is also used to store computer programs, which include program instructions adapted to be called and executed by the processor to perform the various steps of the immersive media data processing method. In addition, memory 603 can also be used to store the three-dimensional image of the immersive media formed after processor processing, the audio content corresponding to the three-dimensional image, and information required for rendering the three-dimensional image and audio content.

[0320] Display / playback device 604 is used to output rendered sound and 3D images.

[0321] Please see again Figure 6 The processor 602 may include a parser 621, a decoder 622, a converter 623, and a renderer 624; wherein:

[0322] The parser 621 is used to depackage and encapsulate the rendering media from the content production device. Specifically, it depackages the media file resources according to the file format requirements of immersive media to obtain audio and video streams, and provides the audio and video streams to the decoder 622.

[0323] Decoder 622 decodes the audio stream to obtain audio content, which is then provided to the renderer for audio rendering. Additionally, decoder 622 decodes the video stream to obtain a 2D image. Based on the metadata provided by the media presentation description information, if the metadata indicates that the immersive media has undergone a region encapsulation process, the 2D image refers to an encapsulated image; if the metadata indicates that the immersive media has not undergone a region encapsulation process, the planar image refers to a projected image.

[0324] Converter 623 is used to convert 2D images into 3D images. If the immersive media has undergone a region encapsulation process, converter 623 will first decapsulate the encapsulated image to obtain a projected image. Then, the projected image will be reconstructed to obtain a 3D image. If the rendering media has not undergone a region encapsulation process, converter 623 will directly reconstruct the projected image to obtain a 3D image.

[0325] Renderer 624 is used to render the audio content and 3D images of immersive media. Specifically, it renders the audio content and 3D images based on the metadata related to rendering and viewing in the media presentation description information, and then outputs the rendered content to the display / playback device.

[0326] In one exemplary embodiment, the processor 602 (specifically, the devices included in the processor) executes instructions by calling one or more instructions stored in memory. Figure 2 The illustrated steps of the immersive media data processing method. Specifically, the memory stores one or more first instructions, which are adapted to be loaded by the processor 602 and executed in the following steps:

[0327] Acquire immersive media files containing presentation instructions for free-viewpoint video, which is generated based on image data acquired by one or more camera groups; the presentation instructions for the free-viewpoint video include instructions for the camera groups.

[0328] Display the free-view video according to the presentation instructions.

[0329] In one embodiment, the camera group indication information includes camera group attribute indication information; the free-viewpoint video presentation indication information is metadata information, which includes a free-viewpoint camera group data box; the free-viewpoint camera group data box includes camera group attribute indication information.

[0330] In one implementation, the free-viewpoint video is encapsulated in at least one track, the at least one track includes a first track, the first track contains image data of the free-viewpoint video; the first track corresponds to N camera groups, each camera group contains one or more cameras, N is a positive integer; a free-viewpoint camera group data box is encapsulated in the first track; the free-viewpoint camera group data box also contains the same camera group information flag field;

[0331] When the same camera group information flag field is set to the first set value, it means that all the image data contained in the first track belong to N camera groups; when the same camera group information flag field is set to the second set value, it means that among the cameras corresponding to the image data contained in the first track, at least one camera belongs to a camera group other than N camera groups.

[0332] In one embodiment, the free-view camera group data box includes a camera group number field and a camera group identifier field. The camera group number field is used to indicate the number of camera groups corresponding to the first track, and the value of the camera group number field is N. One value of the camera group identifier field is used to indicate the identifier of one of the N camera groups.

[0333] In one embodiment, the free-view camera group data box further includes a complete camera group flag field; when the complete camera group flag field is set to a first preset value, it indicates that the image data acquired by all cameras in the camera group indicated by the camera group flag field is included in the first track; when the complete camera group flag field is set to a second preset value, it indicates that the image data acquired by some cameras in the camera group indicated by the camera group flag field is included in the first track.

[0334] In one embodiment, when the same camera group information flag field is set to a second preset value, the free-view camera group data box also includes a camera number field and a camera identification field. The camera number field is used to indicate the number of cameras contained in the camera group indicated by the camera group identification field, and the value of the camera number field is greater than or equal to 1. The camera identification field is used to indicate the identifier of the camera in the camera group indicated by the camera group identification field.

[0335] In one implementation, the free-view video includes one or more free-view scenes; in N camera groups, the image data acquired by all cameras in the same camera group belongs to the same free-view scene, and all cameras in the same camera group have the same coordinate system.

[0336] In one embodiment, the camera group indication information also includes camera parameter indication information in the camera group; the metadata information also includes a free-view data box, which contains camera parameters and repeating camera parameter indication fields.

[0337] When the repeat camera parameter indicator field is set to the first setting value, it means that the camera parameters in the free-view data box are effective for all cameras; when the repeat camera parameter indicator field is set to the second setting value, it means that different cameras in the free-view data box have their own camera parameters.

[0338] In one implementation, the presentation indication information for free-view video also includes free-view scene boundary indication information.

[0339] In one implementation, the presentation indication information of the free-viewpoint video is metadata information, which includes a free-viewpoint boundary data box; the free-viewpoint boundary data box includes free-viewpoint scene boundary indication information.

[0340] In one implementation, the free-viewpoint video is encapsulated in at least one track, the at least one track contains a second track, the second track contains image data of the free-viewpoint video; the second track corresponds to P camera groups, where P is a positive integer; the free-viewpoint boundary data box is encapsulated in the second track;

[0341] The free-view boundary data box contains a camera group flag field, which indicates the type of display boundary corresponding to the image data in the second track. When the camera group flag field is set to a first preset value, it means that the free-view boundary data box is used to indicate the boundary information corresponding to P camera groups. When the camera group flag field is set to a second preset value, it means that the free-view boundary data box is used to indicate the boundary information of the free-view video.

[0342] The free-view boundary data box also includes a camera group number field and a camera group identifier field. The camera group number field indicates the number of camera groups contained in the second track, and the value of the camera group number field is P. The camera group identifier field indicates the identifier of each camera group in the second track.

[0343] In one implementation, if a free-viewpoint video is associated with a non-free-viewpoint video that is jointly consumed, then the presentation indication information of the free-viewpoint video also includes indication information of the association relationship between the free-viewpoint video and the non-free-viewpoint video.

[0344] In one embodiment, the presentation indication information of the free-viewpoint video is metadata information, which includes a free-viewpoint relationship group data box; the free-viewpoint relationship group data box includes the association relationship indication information between the free-viewpoint video and the non-free-viewpoint video.

[0345] In one implementation, free-viewpoint video is encapsulated in one or more third tracks, and non-free-viewpoint video is encapsulated in one or more fourth tracks; the third tracks and the fourth tracks have the same track group type identifier.

[0346] The free-viewpoint relationship group data box includes a track group type identifier and a free-viewpoint video flag field. When the free-viewpoint video flag field is set to the first set value, it indicates that the free-viewpoint video in the third track needs to be presented. When the free-viewpoint video flag field is set to the second set value, it indicates that the non-free-viewpoint video in the fourth track needs to be presented.

[0347] In one implementation, the free-viewpoint video is encapsulated in one or more third tracks, and the non-free-viewpoint video is encapsulated in one or more fourth tracks; the free-viewpoint video and the non-free-viewpoint video belong to the same entity group.

[0348] The free-viewpoint relationship group data box includes the entity group identifier and the free-viewpoint video flag field. When the free-viewpoint video flag field is set to the first set value, it means that the free-viewpoint video of the entity corresponding to the entity group identifier in the third track needs to be presented. When the free-viewpoint video flag field is set to the second set value, it means that the non-free-viewpoint video of the entity corresponding to the entity group identifier in the fourth track needs to be presented.

[0349] In one implementation, free-viewpoint video is encapsulated in one or more third tracks, and non-free-viewpoint video is encapsulated in one or more fourth tracks; there is an index relationship between the track identifiers of the third tracks and the track identifiers of the fourth tracks.

[0350] An index relationship means that a third track can be indexed to a fourth track; or, a fourth track can be indexed to a third track.

[0351] In one implementation, if the free-viewpoint video is encapsulated into multiple tracks, the presentation indication information of the free-viewpoint video also includes association indication information between the multiple tracks.

[0352] In one embodiment, the presentation indication information of the free-viewpoint video is metadata information, which includes a free-viewpoint track group data box, and the free-viewpoint track group data box contains association indication information between multiple tracks.

[0353] In one implementation, multiple orbits are clustered into M orbital groups, with each orbital carrying an orbital group identifier; orbitals with the same orbital group identifier belong to the same orbital group; M is a positive integer.

[0354] In one implementation, each track group corresponds to a free-view track group data box. The free-view video contains one or more free-view scenes, and each free-view scene is generated based on image data acquired by one or more camera groups. The free-view track group data box contains a camera group flag field, which is used to indicate whether there are multiple camera groups in the track group corresponding to the free-view track group data box.

[0355] When the camera group flag field is set to the first set value, it means that all tracks in the track group corresponding to the free-view track group data box constitute the first free-view scene, which is generated based on image data acquired by one camera group; when the camera group flag field is set to the second set value, it means that all tracks in the track group corresponding to the free-view track group data box constitute the second free-view scene, which is generated based on image data acquired by multiple camera groups.

[0356] The free-view track group data box also includes a camera group number field and a camera group identifier field. The camera group number field indicates the number of camera groups in the track group corresponding to the free-view track group data box, and the value of the camera group number field is greater than or equal to 1. The camera group identifier field indicates the identifier of each camera group in the track group corresponding to the free-view track group data box.

[0357] In one implementation, when the free-viewpoint video is generated based on image data acquired by multiple camera groups, the camera group indication information also includes separate presentation time indication information for each camera group.

[0358] In one embodiment, the presentation indication information of the free-viewpoint video is metadata information, which includes camera group sample entities, and the camera group sample entities contain separate presentation time indication information for each camera group.

[0359] In one implementation, the camera group sample entity includes a camera group update flag field, a camera group identifier field, and a camera group cancellation flag field;

[0360] When the camera group update flag field corresponding to the current sample is set to the first preset value, it means that the current sample belongs to the camera group, and the information of the camera group to which the current sample belongs is specified by the camera group identifier field corresponding to the current sample; when the camera group update flag field corresponding to the current sample is set to the second preset value, it means that the information of the camera group corresponding to the current sample remains unchanged.

[0361] When the camera group cancellation flag field corresponding to the current sample is set to the first preset value, it indicates that the camera group information corresponding to the target sample is cancelled, and the camera group information corresponding to the target sample is no longer effective; the current sample refers to the sample being parsed; the target sample refers to any of the historical samples that have been parsed before the current sample. The camera group update flag field corresponding to the target sample is set to the first preset value. The camera group identifier field corresponding to the target sample is the same as the camera group identifier field corresponding to the current sample. Furthermore, the difference between the parsing time of the target sample and the parsing time of the current sample is the minimum value among the differences between the parsing time of each sample in the historical samples and the parsing time of the current sample.

[0362] Specifically, when the camera group update flag field corresponding to the current sample is set to the first preset value, the camera group cancellation flag field corresponding to the current sample is not set to the first preset value; the camera group information corresponding to the current sample remains unchanged means that: if the camera group information corresponding to the target sample is in an active state, then the camera group information corresponding to the current sample is in an active state; if the camera group information corresponding to the target sample is in a canceled state, then the camera group information corresponding to the current sample is in a canceled state.

[0363] In one embodiment, the immersive media file is sliced ​​into multiple media segments; the computer program in memory 603 is loaded by processor 602 and also performs the following steps:

[0364] Obtain the transmission signaling file of the immersive media, which contains descriptive information corresponding to the presentation instructions for the free-viewpoint video.

[0365] In one implementation, the free-view video includes one or more free-view scenes; the transmission signaling file includes at least one of the following: a free-view camera descriptor and a free-view association descriptor;

[0366] Among them, the free-viewpoint camera descriptor is used to indicate the parameters of each camera in one or more camera groups; the free-viewpoint association descriptor is used to indicate the association relationship between free-viewpoint video and the corresponding non-free-viewpoint video.

[0367] In one embodiment, the transmission signaling file includes a representation layer and an adaptive set layer, wherein the adaptive set layer includes one or more representation layers; when description information exists in the target adaptive set layer, it is used to describe all representation layers in the target adaptive set layer; when description information exists in the target representation layer, it is used to describe the target representation layer.

[0368] In one embodiment, the processor 602 acquires the media file of the immersive media as follows:

[0369] Based on the virtual perspective of the viewer of immersive media and the description information in the transmission signaling file, determine the media segment corresponding to the image data required for viewing;

[0370] The determined media segments are retrieved using streaming.

[0371] In one embodiment, the processor 602 displays the free-viewpoint video according to the presentation instruction information of the free-viewpoint video. The specific implementation method is as follows:

[0372] Based on the virtual perspective of the viewer while watching immersive media, determine the image data required for viewing from the media file;

[0373] The image data is decoded and displayed according to the instructions from the camera group.

[0374] In one embodiment, the presentation indication information for the free-viewpoint video further includes free-viewpoint scene boundary indication information; the processor 602 displays the free-viewpoint video according to the presentation indication information as follows:

[0375] During the decoding process of image data, the display boundaries of the image data are determined according to the free-view scene boundary indication information.

[0376] In one embodiment, if the free-viewpoint video is associated with a non-free-viewpoint video that is jointly consumed, then the presentation indication information of the free-viewpoint video also includes indication information of the association relationship between the free-viewpoint video and the non-free-viewpoint video; the processor 602 displays the free-viewpoint video according to the presentation indication information of the free-viewpoint video. The specific implementation of this method is as follows:

[0377] During the display of image data, the free-viewpoint video is switched to a non-free-viewpoint video according to the association information between free-viewpoint and non-free-viewpoint videos; or,

[0378] Based on the association information between free-view and non-free-view videos, switch the non-free-view video to a free-view video.

[0379] In one embodiment, the image data is acquired by multiple camera groups; the indication information of the camera groups also includes the presentation time indication information for each camera group; the processor 602 decodes and displays the image data according to the indication information of the camera groups. The specific implementation is as follows:

[0380] According to the presentation time indication information of each camera group, the image data corresponding to the corresponding camera group in the image data is decoded and displayed sequentially.

[0381] Based on the same inventive concept, the principle and beneficial effects of the content consumption device provided in the embodiments of this application are similar to the principle and beneficial effects of the immersive media processing method in the embodiments of this application. For details, please refer to the principle and beneficial effects of the method implementation. For the sake of brevity, they will not be repeated here.

[0382] Figure 7 This illustration shows a schematic diagram of a content creation device provided in an exemplary embodiment of this application; the content creation device may refer to a computer device used by an immersive media provider, which may be a terminal (such as a PC, a smart mobile device (such as a smartphone) or a server. Figure 7 As shown, the content creation device includes a capture device 701, a processor 702, a memory 703, and a transmitter 704. Wherein:

[0383] The capture device 701 is used to acquire raw data (including audio and video content synchronized in time and space) of real-world sound-visual scenes to obtain immersive media. The capture device 701 may include, but is not limited to, audio devices, camera devices, and sensing devices. Audio devices may include audio sensors, microphones, etc. Camera devices may include ordinary cameras, stereo cameras, light field cameras, etc. Sensing devices may include laser devices, radar devices, etc.

[0384] Processor 702 (or CPU (Central Processing Unit)) is the processing core of the content production device. Processor 702 is adapted to implement one or more program instructions, specifically to load and execute one or more program instructions to achieve... Figure 3The flowchart illustrates the data processing method for immersive media.

[0385] Memory 703 is a memory device in the content creation apparatus used to store programs and media resources. It is understood that memory 703 here can include both the built-in storage medium of the content creation apparatus and extended storage media supported by the content creation apparatus. It should be noted that the memory can be high-speed RAM or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one memory located remotely from the aforementioned processor. The memory provides storage space for storing the operating system of the content creation apparatus. Furthermore, this storage space is also used to store computer programs, which include program instructions adapted to be invoked and executed by the processor to perform the various steps of the immersive media data processing method. Additionally, memory 703 can also be used to store immersive media files formed after processing by the processor, which include media file resources and media presentation description information.

[0386] The transmitter 704 is used to enable transmission and interaction between the content creation device and other devices, specifically to facilitate the transmission of immersive media between the content creation device and the content playback device. That is, the content creation device uses the transmitter 704 to transmit relevant media resources for immersive media to the content playback device.

[0387] Please see again Figure 7 The processor 702 may include a converter 721, an encoder 722, and a packager 723; wherein:

[0388] Converter 721 performs a series of conversion processes on captured video content to make it suitable for immersive media video encoding. The conversion processes may include stitching and projection; optionally, they may also include region encapsulation. Converter 721 can convert captured 3D video content into 2D images and provide them to the encoder for video encoding.

[0389] Encoder 722 is used to encode the captured audio content to form an audio stream for immersive media. It is also used to encode the 2D image obtained by converter 721 to obtain a video stream.

[0390] The encapsulator 723 encapsulates audio and video streams into a file container according to the immersive media file format (such as ISOBMFF) to form an immersive media file resource. This media file resource can be a media file or a media segment forming an immersive media file. It also records the metadata of the immersive media file resource using media presentation description information according to the immersive media file format requirements. The encapsulated immersive media file obtained by the encapsulator is stored in memory and provided to the content playback device as needed for immersive media presentation.

[0391] The processor 702 (specifically, the various components within the processor) executes instructions by calling one or more instructions from memory. Figure 4 The illustrated steps of the immersive media data processing method are as follows. Specifically, memory 703 stores one or more first instructions, which are adapted to be loaded by processor 702 and executed in the following steps:

[0392] Acquire image data from one or more camera groups and encode the image data into free-viewpoint video;

[0393] Add presentation instructions to the free-view video according to its application format; the presentation instructions for the free-view video include instructions for the camera group;

[0394] Encapsulate free-viewpoint video and free-viewpoint video presentation instructions into immersive media files.

[0395] In one embodiment, the computer program in memory 703 is loaded by processor 702 and further performs the following steps:

[0396] Slicing a media file to obtain multiple media segments; and,

[0397] Generate a transmission signaling file for immersive media, which contains descriptive information corresponding to the presentation instructions for free-viewpoint video.

[0398] Based on the same inventive concept, the principle and beneficial effects of the content production device provided in the embodiments of this application are similar to the principle and beneficial effects of the immersive media processing method in the embodiments of this application. For details, please refer to the principle and beneficial effects of the method implementation. For the sake of brevity, these will not be repeated here.

[0399] This application also provides a computer-readable storage medium storing one or more instructions adapted for a processor to load and execute the immersive media data processing method of the above-described method embodiments.

[0400] This application also provides a computer program product containing instructions that, when run on a computer, causes the computer to execute the immersive media data processing method described in the above method embodiments.

[0401] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned immersive media data processing method.

[0402] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.

[0403] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0404] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0405] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art will understand that all or part of the processes for implementing the above embodiments and equivalent variations made in accordance with the claims of this application are still within the scope of this application.

Claims

1. A data processing method for immersive media, characterized in that, include: The media file of the immersive media is obtained, the media file containing presentation indication information of free-viewpoint video, the free-viewpoint video being generated based on image data acquired by one or more camera groups; the presentation indication information of the free-viewpoint video includes indication information of the camera groups; the indication information of the camera groups includes attribute indication information of the camera groups; the presentation indication information of the free-viewpoint video is metadata information, the metadata information containing free-viewpoint camera group data boxes; The free-view camera group data box contains attribute indication information of the camera group; Display the free-view video according to the presentation instructions for the free-view video; The free-viewpoint video is encapsulated in at least one track, and the at least one track includes a first track containing image data of the free-viewpoint video; the first track corresponds to N camera groups, each camera group contains one or more cameras, and N is a positive integer; the free-viewpoint camera group data box is encapsulated in the first track; the free-viewpoint camera group data box also contains the same camera group information flag field; When the same camera group information flag field is set to a first preset value, it means that all the image data contained in the first track belong to the N camera groups; when the same camera group information flag field is set to a second preset value, it means that among the cameras corresponding to the image data contained in the first track, at least one camera belongs to a camera group other than the N camera groups.

2. The method as described in claim 1, characterized in that, The free-view camera group data box includes a camera group number field and a camera group identifier field. The camera group number field is used to indicate the number of camera groups corresponding to the first track, and the value of the camera group number field is N. One value of the camera group identifier field is used to indicate the identifier of one of the N camera groups.

3. The method as described in claim 2, characterized in that, The free-view camera group data box also includes a complete camera group identifier field; When the complete camera group flag field is set to a first preset value, it means that the image data acquired by all cameras in the camera group indicated by the camera group identifier field is included in the first track; when the complete camera group flag field is set to a second preset value, it means that the image data acquired by some cameras in the camera group indicated by the camera group identifier field is included in the first track.

4. The method as described in claim 2, characterized in that, When the same camera group information flag field is set to a second set value, the free-view camera group data box also includes a camera number field and a camera identification field. The camera number field is used to indicate the number of cameras contained in the camera group indicated by the camera group identification field, and the value of the camera number field is greater than or equal to 1. The camera identification field is used to indicate the identifier of the camera in the camera group indicated by the camera group identification field.

5. The method as described in claim 1, characterized in that, The free-view video contains one or more free-view scenes; among the N camera groups, the image data acquired by all cameras in the same camera group belongs to the same free-view scene, and all cameras in the same camera group have the same coordinate system.

6. The method as described in claim 5, characterized in that, The presentation indication information of the free-view video also includes the boundary indication information of the free-view scene.

7. The method as described in claim 6, characterized in that, The metadata information also includes a free-view boundary data box; the free-view boundary data box contains boundary indication information of the free-view scene.

8. The method as described in claim 7, characterized in that, The at least one track further includes a second track, which contains image data of the free-view video; the second track corresponds to P camera groups, where P is a positive integer; the free-view boundary data box is encapsulated in the second track; The free-view boundary data box includes a camera group flag field, which is used to indicate the type of display boundary corresponding to the image data in the second track. When the camera group flag field is set to a first preset value, it indicates that the free-view boundary data box is used to indicate the boundary information corresponding to the P camera groups. When the camera group flag field is set to a second preset value, it indicates that the free-view boundary data box is used to indicate the boundary information of the free-view video. The free-view boundary data box also includes a camera group number field and a camera group identifier field. The camera group number field is used to indicate the number of camera groups contained in the second track, and the value of the camera group number field is P. The camera group identifier field is used to indicate the identifier of each camera group in the second track.

9. The method as described in claim 1, characterized in that, If the free-viewpoint video is associated with a non-free-viewpoint video that is jointly consumed, then the presentation indication information of the free-viewpoint video also includes the association relationship indication information between the free-viewpoint video and the non-free-viewpoint video.

10. The method as described in claim 9, characterized in that, The metadata information also includes a free-viewpoint relationship group data box; the free-viewpoint relationship group data box contains association information between the free-viewpoint video and the non-free-viewpoint video.

11. The method as described in claim 10, characterized in that, The free-viewpoint video is encapsulated in one or more third tracks, and the non-free-viewpoint video is encapsulated in one or more fourth tracks; The third track and the fourth track have the same track group type identifier; The free-viewpoint relationship group data box includes the track group type identifier and the free-viewpoint video flag field; when the free-viewpoint video flag field is set to a first preset value, it indicates that the free-viewpoint video in the third track needs to be presented; when the free-viewpoint video flag field is set to a second preset value, it indicates that the non-free-viewpoint video in the fourth track needs to be presented.

12. The method as described in claim 10, characterized in that, The free-viewpoint video is encapsulated in one or more third tracks, and the non-free-viewpoint video is encapsulated in one or more fourth tracks; the free-viewpoint video and the non-free-viewpoint video belong to the same entity group; The free-viewpoint relationship group data box includes the identifier of the entity group and a free-viewpoint video flag field; when the free-viewpoint video flag field is set to a first preset value, it indicates that the free-viewpoint video of the entity corresponding to the identifier of the entity group in the third track needs to be presented; when the free-viewpoint video flag field is set to a second preset value, it indicates that the non-free-viewpoint video of the entity corresponding to the identifier of the entity group in the fourth track needs to be presented.

13. The method as described in claim 10, characterized in that, The free-viewpoint video is encapsulated in one or more third tracks, and the non-free-viewpoint video is encapsulated in one or more fourth tracks; there is an index relationship between the track identifiers of the third tracks and the track identifiers of the fourth tracks; The indexing relationship means that the track identifier of the third track can be used to index the fourth track; or, the track identifier of the fourth track can be used to index the third track.

14. The method as described in claim 1, characterized in that, If the free-viewpoint video is encapsulated into multiple tracks, the presentation indication information of the free-viewpoint video also includes the association indication information between the multiple tracks.

15. The method as described in claim 14, characterized in that, The presentation indication information of the free-viewpoint video is metadata information, which includes a free-viewpoint track group data box. The free-viewpoint track group data box contains association indication information between the multiple tracks.

16. The method as described in claim 15, characterized in that, The multiple orbits are clustered into M orbital groups, each orbital carrying an orbital group identifier; orbitals with the same orbital group identifier belong to the same orbital group; M is a positive integer.

17. The method as described in claim 16, characterized in that, Each track group corresponds to a free-view track group data box. The free-view video contains one or more free-view scenes, and each free-view scene is generated based on image data acquired by one or more camera groups. The free-view track group data box contains a camera group flag field, which is used to indicate whether there are multiple camera groups in the track group corresponding to the free-view track group data box. When the camera group flag field is set to a first set value, it indicates that all tracks in the track group corresponding to the free-view track group data box constitute a first free-view scene, which is generated based on image data acquired by one camera group; when the camera group flag field is set to a second set value, it indicates that all tracks in the track group corresponding to the free-view track group data box constitute a second free-view scene, which is generated based on image data acquired by multiple camera groups. The free-view track group data box also includes a camera group number field and a camera group identifier field. The camera group number field is used to indicate the number of camera groups in the track group corresponding to the free-view track group data box, and the value of the camera group number field is greater than or equal to 1. The camera group identifier field is used to indicate the identifier of each camera group in the track group corresponding to the free-view track group data box.

18. The method as described in claim 1, characterized in that, When the free-viewpoint video is generated based on image data acquired by multiple camera groups, the indication information of the camera groups also includes the presentation time indication information of each camera group.

19. The method as described in claim 18, characterized in that, The metadata information also includes camera group sample entities, which contain presentation time indication information for each camera group.

20. The method as described in claim 19, characterized in that, The camera group sample entity includes a camera group update flag field, a camera group identifier field, and a camera group cancellation flag field; When the camera group update flag field corresponding to the current sample is set to a first preset value, it indicates that the current sample belongs to a camera group, and the information of the camera group to which the current sample belongs is specified by the camera group identifier field corresponding to the current sample; when the camera group update flag field corresponding to the current sample is set to a second preset value, it indicates that the camera group information corresponding to the current sample remains unchanged. When the camera group cancellation flag field corresponding to the current sample is set to a first preset value, it indicates that the camera group information corresponding to the target sample is cancelled, and the camera group information corresponding to the target sample is no longer effective; the current sample refers to the sample being parsed; the target sample refers to any one of the historical samples that have been parsed before the current sample, the camera group update flag field corresponding to the target sample is set to a first preset value, the camera group identifier field corresponding to the target sample is the same as the camera group identifier field corresponding to the current sample, and the difference between the parsing time of the target sample and the parsing time of the current sample is the minimum value among the differences between the parsing time of each sample in the historical samples and the parsing time of the current sample; Specifically, when the camera group update flag field corresponding to the current sample is set to a first preset value, the camera group cancellation flag field corresponding to the current sample is not set to the first preset value; the camera group information corresponding to the current sample remains unchanged means that: if the camera group information corresponding to the target sample is in an active state, then the camera group information corresponding to the current sample is in an active state; if the camera group information corresponding to the target sample is in a canceled state, then the camera group information corresponding to the current sample is in a canceled state.

21. The method as described in claim 1, characterized in that, The immersive media file is sliced ​​into multiple media segments; the method further includes: Obtain the transmission signaling file of the immersive media, wherein the transmission signaling file contains descriptive information corresponding to the presentation indication information of the free-viewpoint video.

22. The method as described in claim 21, characterized in that, The free-view video contains one or more free-view scenes; the transmission signaling file contains at least one of the following: a free-view camera descriptor and a free-view association descriptor; The free-view camera descriptor is used to indicate the parameters of each camera in the one or more camera groups; the free-view association descriptor is used to indicate the association relationship between the free-view video and the corresponding non-free-view video.

23. The method as described in claim 21, characterized in that, The transmission signaling file includes a representation layer and an adaptive set layer, wherein the adaptive set layer includes one or more representation layers; when the description information exists in the target adaptive set layer, it is used to describe all representation layers in the target adaptive set layer; when the description information exists in the target representation layer, it is used to describe the target representation layer.

24. The method as described in claim 21, characterized in that, The media files for obtaining immersive media include: Based on the virtual perspective of the viewer of the immersive media and the description information in the transmission signaling file, the media segment corresponding to the image data required for viewing is determined. The determined media segments are retrieved using streaming.

25. The method according to any one of claims 1-24, characterized in that, Displaying the free-view video according to the presentation instruction information of the free-view video includes: Based on the virtual perspective of the viewer of the immersive media, determine the image data required for viewing from the media file; The image data is decoded and displayed according to the instructions from the camera group.

26. The method as described in claim 25, characterized in that, The presentation indication information for the free-view video also includes boundary indication information for the free-view scene; displaying the free-view video according to the presentation indication information further includes: During the decoding process of the image data, the display boundary of the image data is determined according to the boundary indication information of the free-view scene.

27. The method as described in claim 26, characterized in that, If the free-viewpoint video is associated with a non-free-viewpoint video that is jointly consumed, then the presentation indication information of the free-viewpoint video also includes indication information of the association relationship between the free-viewpoint video and the non-free-viewpoint video; displaying the free-viewpoint video according to the presentation indication information of the free-viewpoint video further includes: During the display of the image data, the free-view video is switched to the non-free-view video according to the association information between the free-view video and the non-free-view video; or... According to the association information between the free-view video and the non-free-view video, the non-free-view video is switched to the free-view video.

28. The method as described in claim 25, characterized in that, The image data was acquired by multiple camera groups; the indication information of the camera groups also includes the presentation time indication information of each camera group. The step of decoding and displaying the image data according to the instructions from the camera group further includes: According to the presentation time indication information of each camera group, the image data corresponding to the corresponding camera group in the image data is decoded and displayed sequentially.

29. A data processing method for immersive media, characterized in that, include: Acquire image data from one or more camera groups and encode the image data into free-viewpoint video; According to the application format of the free-view video, presentation indication information is added to the free-view video; the presentation indication information of the free-view video includes the indication information of the camera group; the indication information of the camera group includes the attribute indication information of the camera group; the presentation indication information of the free-view video is metadata information, and the metadata information includes the free-view camera group data box; The free-view camera group data box contains attribute indication information of the camera group; The free-viewpoint video and the presentation instruction information of the free-viewpoint video are encapsulated into a media file of immersive media; The free-viewpoint video is encapsulated in at least one track, and the at least one track includes a first track containing image data of the free-viewpoint video; the first track corresponds to N camera groups, each camera group contains one or more cameras, and N is a positive integer; the free-viewpoint camera group data box is encapsulated in the first track; the free-viewpoint camera group data box also contains the same camera group information flag field; When the same camera group information flag field is set to a first preset value, it means that all the image data contained in the first track belong to the N camera groups; when the same camera group information flag field is set to a second preset value, it means that among the cameras corresponding to the image data contained in the first track, at least one camera belongs to a camera group other than the N camera groups.

30. The method as described in claim 29, characterized in that, The method further includes: The media file is sliced ​​to obtain multiple media segments; and, Generate a transmission signaling file for the immersive media, the transmission signaling file containing descriptive information corresponding to the presentation indication information of the free-viewpoint video.

31. A data processing device for immersive media, characterized in that, The immersive media data processing device includes: An acquisition unit is used to acquire media files of immersive media, wherein the media files contain presentation indication information of free-viewpoint video, which is generated based on image data acquired by one or more camera groups; the presentation indication information of the free-viewpoint video includes indication information of the camera groups, which includes attribute indication information of the camera groups; the presentation indication information of the free-viewpoint video is metadata information, which includes a free-viewpoint camera group data box; the free-viewpoint camera group data box includes attribute indication information of the camera groups. The processing unit is configured to display the free-viewpoint video according to the presentation instruction information of the free-viewpoint video; The free-viewpoint video is encapsulated in at least one track, and the at least one track includes a first track containing image data of the free-viewpoint video; the first track corresponds to N camera groups, each camera group contains one or more cameras, and N is a positive integer; the free-viewpoint camera group data box is encapsulated in the first track; the free-viewpoint camera group data box also contains the same camera group information flag field; When the same camera group information flag field is set to a first preset value, it means that all the image data contained in the first track belong to the N camera groups; when the same camera group information flag field is set to a second preset value, it means that among the cameras corresponding to the image data contained in the first track, at least one camera belongs to a camera group other than the N camera groups.

32. A data processing device for immersive media, characterized in that, The immersive media data processing device includes: An acquisition unit is used to acquire image data collected by one or more camera groups and encode the image data into free-viewpoint video. A processing unit is configured to add presentation indication information to the free-viewpoint video according to the application format of the free-viewpoint video; the presentation indication information of the free-viewpoint video includes the indication information of the camera group; the indication information of the camera group includes the attribute indication information of the camera group; the presentation indication information of the free-viewpoint video is metadata information, the metadata information includes a free-viewpoint camera group data box; the free-viewpoint camera group data box includes the attribute indication information of the camera group; and to encapsulate the free-viewpoint video and the presentation indication information of the free-viewpoint video into a media file for immersive media; The free-viewpoint video is encapsulated in at least one track, and the at least one track includes a first track containing image data of the free-viewpoint video; the first track corresponds to N camera groups, each camera group contains one or more cameras, and N is a positive integer; the free-viewpoint camera group data box is encapsulated in the first track; the free-viewpoint camera group data box also contains the same camera group information flag field; When the same camera group information flag field is set to a first preset value, it means that all the image data contained in the first track belong to the N camera groups; when the same camera group information flag field is set to a second preset value, it means that among the cameras corresponding to the image data contained in the first track, at least one camera belongs to a camera group other than the N camera groups.

33. A computer device, characterized in that, include: Storage devices and processors; A memory, wherein a computer program is stored; A processor is configured to load the computer program to implement the data processing method for immersive media as described in any one of claims 1-28; or, to load the computer program to implement the data processing method for immersive media as described in claim 29 or 30.

34. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and executed as a data processing method for immersive media as described in any one of claims 1-28; or, loaded and executed as a data processing method for immersive media as described in claim 29 or 30.

35. A computer program product, characterized in that, The computer program product includes a computer program adapted to be loaded by a processor and execute the immersive media data processing method as described in any one of claims 1-28; or, to load and execute the immersive media data processing method as described in claim 29 or 30.

Citation Information

Patent Citations

  • Image data transmission device, image data transmission method, and image data receiving device

    CN108471546A

  • Data processing method of immersion media

    CN113766271A