Volumetric video data processing method, device, equipment and readable storage medium

By constructing a mapping relationship between viewpoint groups and spatial positions in volumetric videos and generating timed metadata information, the server optimizes the decoding process, solving the inefficiency problem caused by the video client's indiscriminate decoding of multiple viewpoint groups, and achieving more efficient video decoding and presentation.

CN115481280BActive Publication Date: 2025-09-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110660865.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-15
Publication Date
2025-09-26
Estimated Expiration
2041-06-15

AI Technical Summary

Technical Problem

During volumetric video processing, in the prior art, a video client indiscriminately decodes the coded streams of multiple viewpoint groups, resulting in excessively long decoding time and reduced decoding presentation efficiency.

Method used

The server constructs a mapping relationship between viewpoint groups and spatial position information, generates timed metadata information, and writes it into the encapsulated data box. It only decodes the target viewpoint group, and the video client displays the target video content based on the timed metadata information.

Benefits of technology

By optimizing the decoding process, the decoding and presentation efficiency of the video client is improved, unnecessary decoding operations are reduced, and the speed and efficiency of video playback are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115481280B_ABST
    Figure CN115481280B_ABST
Patent Text Reader

Abstract

The present application provides a volumetric video data processing method, apparatus, device, and readable storage medium. The method comprises: obtaining the i-th viewpoint group among G viewpoint groups of a volumetric video as a target viewpoint group; constructing a mapping relationship between the target viewpoint group and target spatial position information based on the target video content corresponding to the target viewpoint group; writing timed metadata information generated based on the mapping relationship into an encapsulation data box corresponding to the volumetric video to obtain a first extended data box; encapsulating the encoded video streams associated with the G viewpoint groups based on the first extended data box to obtain a video media file of the volumetric video; and delivering the video media file to a video client, so that when the video client obtains the first extended data box based on the video media file, it displays the target video content on the video client based on the target spatial position information in the first extended data box. The present application can improve the decoding and presentation efficiency of volumetric videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, device, and readable storage medium for processing volumetric video data. Background Art

[0002] Volumetric video refers to visual content captured in three-dimensional space, providing users with a viewing experience with multiple degrees of freedom (e.g., 3DoF+ and 6DoF). Volumetric video here primarily refers to multi-perspective video (also known as multi-viewpoint video) with depth information, captured from multiple angles using multiple camera arrays.

[0003] During video processing of volumetric video (i.e., multi-viewpoint video), the different viewpoints of the volumetric video are typically grouped to obtain multiple viewpoint groups, and the media content corresponding to the multiple viewpoint groups can then be encoded to obtain corresponding media encapsulation files. However, the inventors have discovered in practice that when a media encapsulation file contains a volumetric video containing multiple viewpoint groups, when a video client decodes the encoded stream of the volumetric video, all of the encoded streams of the volumetric video are decoded, resulting in indiscriminate decoding of the media content corresponding to each of the multiple viewpoint groups. This consumes a long video decoding time and reduces the decoding and presentation efficiency of the video client. Summary of the Invention

[0004] The embodiments of the present application provide a method, apparatus, device, and readable storage medium for processing volumetric video data, which can improve the decoding and presentation efficiency of a video client.

[0005] In one aspect, an embodiment of the present application provides a method for processing volumetric video data, the method being executed by a server and comprising:

[0006] Obtain G viewpoint groups of the volumetric video, and use the i-th viewpoint group among the G viewpoint groups as the target viewpoint group; i is a non-negative integer less than G;

[0007] Based on the target video content corresponding to the target viewpoint group, a mapping relationship between the target viewpoint group and target spatial position information for viewing the target video content is established, and timed metadata information corresponding to the target viewpoint group is generated based on the mapping relationship;

[0008] Writing the timed metadata information corresponding to the target viewpoint group into the encapsulated data box corresponding to the volumetric video to obtain a first extended data box corresponding to the encapsulated data box; the first extended data box contains the timed metadata information corresponding to each of the G viewpoint groups;

[0009] Obtaining coded video streams associated with the G viewpoint groups, and encapsulating the coded video streams based on the first extended data box to obtain a video media file of the volumetric video;

[0010] The video media file is sent to the video client so that when the video client obtains the first extended data box based on the video media file, the target video content corresponding to the target viewpoint group is displayed on the video client according to the target spatial position information indicated by the timing metadata information corresponding to the target viewpoint group in the first extended data box.

[0011] In one aspect, an embodiment of the present application provides a volumetric video data processing device, including:

[0012] A viewpoint group acquisition module is used to acquire G viewpoint groups of the volumetric video, and take the i-th viewpoint group among the G viewpoint groups as the target viewpoint group; i is a non-negative integer less than G;

[0013] a mapping relationship construction module, configured to construct a mapping relationship between the target viewpoint group and target spatial position information for viewing the target video content based on the target video content corresponding to the target viewpoint group, and to generate timing metadata information corresponding to the target viewpoint group based on the mapping relationship;

[0014] a timed metadata writing module, configured to write the timed metadata information corresponding to the target viewpoint group into the encapsulated data box corresponding to the volumetric video, thereby obtaining a first extended data box corresponding to the encapsulated data box; the first extended data box contains the timed metadata information corresponding to each of the G viewpoint groups;

[0015] a media file delivery module configured to obtain coded video streams associated with the G viewpoint groups, and encapsulate the coded video streams based on the first extended data box to obtain a video media file of the volumetric video;

[0016] The media file sending module is also used to send the video media file to the video client, so that when the video client obtains the first extended data box based on the video media file, it displays the target video content corresponding to the target viewpoint group on the video client according to the target spatial position information indicated by the timing metadata information corresponding to the target viewpoint group in the first extended data box.

[0017] Among them, the viewpoint group acquisition module includes:

[0018] A viewpoint acquisition unit is configured to acquire V viewpoints of the volumetric video and group the V viewpoints based on viewpoint dependencies between the V viewpoints to obtain G viewpoint groups of the volumetric video. V represents the number of viewpoints in the volumetric video and is a positive integer greater than or equal to 2. The viewpoint dependencies are determined by the content relevance between the video content corresponding to each of the V viewpoints.

[0019] Among them, the viewpoint group acquisition module includes:

[0020] a viewpoint group search unit configured to obtain video association information associated with the volumetric video and search the video association information for a designated viewpoint group associated with a content producer; the designated viewpoint group is associated with a shooting intention of the content producer who shot the volumetric video;

[0021] A notification unit is used to obtain the i-th viewpoint group that is irrelevant to the content producer's shooting intention from the G viewpoint groups as the target viewpoint group if the designated viewpoint group associated with the content producer is not found in the video-related information, and to notify the mapping relationship construction module to execute the steps of constructing a mapping relationship between the target viewpoint group and the target spatial position information for viewing the target video content based on the target viewpoint group, and generating timing metadata information corresponding to the target viewpoint group based on the mapping relationship.

[0022] The device further comprises:

[0023] An invalid field setting module is used to add a recommended viewpoint group identification field in the viewpoint group metadata sample of the encapsulated data box, and set the field value of the recommended viewpoint group identification field to an invalid field value, and in the viewpoint group metadata sample of the first extended data box, use the recommended viewpoint group identification field with the invalid field value as the first field indication information; the first field indication information is used to instruct the video client to obtain the timed metadata information of each viewpoint group in the G viewpoint groups from the first extended data box.

[0024] The viewpoint group acquisition module further includes:

[0025] a recommended viewpoint group determining unit, configured to, if a designated viewpoint group associated with a content producer is found in the video-related information, use the found designated viewpoint group as a recommended viewpoint group;

[0026] The target viewpoint group determining unit is configured to obtain, based on the recommended viewpoint group, an i-th viewpoint group related to the shooting intention of the content producer from the G viewpoint groups as the target viewpoint group.

[0027] The device further comprises:

[0028] a recommendation metadata writing module, configured to determine the identifier of the target view group as the recommendation identifier, and to use metadata information describing the recommendation identifier as the recommendation metadata information of the target view group, and to write the recommendation metadata information into an encapsulated data box corresponding to the volumetric video, thereby obtaining a second extended data box corresponding to the encapsulated data box;

[0029] The encapsulation processing module is configured to obtain encoded video streams associated with the G viewpoint groups, encapsulate the encoded video streams based on the second extended data box to obtain a video media file of the volumetric video, and deliver the video media file to a video client, so that when the video client obtains the second extended data box based on the video media file, the video client displays, on the video client, target video content corresponding to the target viewpoint group indicated by the recommendation identifier based on the recommendation metadata information in the second extended data box.

[0030] Before the recommendation metadata writing module writes the recommendation metadata information into the encapsulated data box corresponding to the volumetric video, the apparatus further includes:

[0031] A valid field setting module is used to add a recommended viewpoint group identification field associated with the recommendation identifier in the viewpoint group metadata sample of the encapsulated data box, and set the field value of the recommended viewpoint group identification field to a valid field value, and in the viewpoint group metadata sample of the second extended data box, use the recommended viewpoint group identification field with a valid field value as the second field indication information; the second field indication information is used to instruct the video client to obtain the recommended metadata information from the second extended data box.

[0032] In one aspect, an embodiment of the present application provides a method for processing volumetric video data, which is executed by a video client and includes:

[0033] Receive the video media file of the volumetric video sent by the server, decapsulate the video media file, and obtain the video encoding stream of the volumetric video and the extended data box corresponding to the video encoding stream; the extended data box includes a recommended viewpoint group identification field;

[0034] If the field value of the recommended view group identification field is an invalid value, obtaining timed metadata information corresponding to each of the G view groups of the volumetric video in the first extended data box of the extended data box;

[0035] Obtaining the i-th viewpoint group from the G viewpoint groups, using the spatial position information of the video client as the spatial position information to be compared, and comparing the spatial position information to be compared with the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group to obtain a comparison result; i is a non-negative integer less than G;

[0036] If the comparison result indicates that the to-be-compared spatial position information is the same as the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group, the i-th viewpoint group is used as the matching viewpoint group, and the spatial position information indicated by the timed metadata information corresponding to the matching viewpoint group is used as the target spatial position information;

[0037] Based on the mapping relationship between the matching viewpoint group and the target spatial position information, the matching video content corresponding to the matching viewpoint group is obtained by decoding the video encoding stream of the volumetric video, and the matching video content corresponding to the matching viewpoint group is displayed on the video client.

[0038] In one aspect, an embodiment of the present application provides a volumetric video data processing device, including:

[0039] A media file receiving module is configured to receive a video media file of a volumetric video sent by a server, decapsulate the video media file, and obtain a video encoding stream of the volumetric video and an extended data box corresponding to the video encoding stream; the extended data box includes a recommended viewpoint group identification field;

[0040] a timed metadata acquisition module, configured to acquire, in a first extended data box of the extended data box, timed metadata information corresponding to each of the G view groups of the volumetric video if the field value of the recommended view group identification field is an invalid value;

[0041] An information comparison module is configured to obtain an i-th viewpoint group from the G viewpoint groups, use the spatial position information of the video client as the spatial position information to be compared, and compare the spatial position information to be compared with the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group to obtain a comparison result; i is a non-negative integer less than G;

[0042] a target information determining module, configured to, if the comparison result indicates that the to-be-compared spatial position information is the same as the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group, determine the i-th viewpoint group as a matching viewpoint group, and determine the spatial position information indicated by the timed metadata information corresponding to the matching viewpoint group as target spatial position information;

[0043] The video decoding module is used to decode the matching video content corresponding to the matching viewpoint group from the video encoding stream of the volumetric video based on the mapping relationship between the matching viewpoint group and the target spatial position information, and display the matching video content corresponding to the matching viewpoint group on the video client.

[0044] The timed metadata acquisition module includes:

[0045] a field acquiring unit, configured to acquire, if the field value of the recommended view group identification field is an invalid value, static view group metadata fields associated with the G view groups in a first extended data box of the extended data box;

[0046] A timed metadata acquisition unit is configured to acquire, if the field value of the static viewpoint group metadata field is a numerical value that remains unchanged and is used to describe the mapping relationship between each viewpoint group in the G viewpoint groups, based on the first field indication information associated with the recommended viewpoint group identification field, timed metadata information corresponding to each viewpoint group in the G viewpoint groups.

[0047] Among them, the static viewpoint group metadata field is deployed in the viewpoint group static metadata box of the first extended data box; if the field value of the static viewpoint group metadata field is a numerical value that remains unchanged for describing the mapping relationship of the viewpoint groups in the G viewpoint groups, then the static viewpoint group metadata recorded in the viewpoint group metadata sample entry of the first extended data box: includes the viewpoint group identifier of each viewpoint group in the G viewpoint groups.

[0048] Among them, if the field value of the static viewpoint group metadata field is a numerical value associated with the dynamic viewpoint group metadata, and the dynamic viewpoint group metadata is used to describe a viewpoint group among the G viewpoint groups whose mapping relationship changes over time, then the dynamic viewpoint group metadata recorded in the viewpoint group element sample corresponding to the viewpoint group metadata sample entry includes: an identifier of a variable viewpoint group whose mapping relationship changes at the sample timestamp, and before the sample timestamp, the mapping relationship between the variable viewpoint group and the video content corresponding to the variable viewpoint group remains unchanged; the variable viewpoint group is a viewpoint group among the G viewpoint groups whose mapping relationship changes over time.

[0049] The view group metadata sample entry and the view group element sample are used to constitute a view group timed metadata track of the volumetric video, and the view group timed metadata track is used to index one or more atlas data tracks associated with the volumetric video.

[0050] Among them, the target spatial position information is determined by the type field of the judgment information associated with the viewpoint group recorded by the server in the viewpoint group metadata sample entry of the first extended data box; the type field of the judgment information is deployed in the viewpoint group static metadata box of the viewpoint group metadata sample entry; if the type field of the judgment type is a first value, the target spatial position information having a mapping relationship with the matching viewpoint group is the three-dimensional spatial area information of the matching video content displayed on the video client; if the type field of the judgment type is a second value, the target spatial position information having a mapping relationship with the matching viewpoint group is the viewing position coordinate information of the user watching the matching video content on the video client; if the type field of the judgment type is a third value, the target spatial position information having a mapping relationship with the target viewpoint group is jointly determined by the three-dimensional spatial area information and the viewing position coordinate information.

[0051] The device further comprises:

[0052] an information changing module, configured to, when the spatial position information of the video client is changed from the target spatial position information to the spatial position update information, use, based on the timed metadata information corresponding to the spatial position update information in the first extended data box, a mapping relationship between the spatial position update information and the changed viewpoint group as an update relationship;

[0053] The update decoding module is used to decode the video encoding stream of the volumetric video based on the update relationship to obtain the video content corresponding to the changed viewpoint group, and display the video content corresponding to the changed viewpoint group on the video client.

[0054] The device further comprises:

[0055] a recommendation identifier acquisition module configured to acquire, if the field value of the recommended viewpoint group identification field is a valid value, a recommendation identifier of the recommended viewpoint group indicated by the recommendation metadata information based on second field indication information associated with the recommended viewpoint group identification field in a second extended data box of the extended data box;

[0056] The recommended content display module is used to decode the video encoding stream of the volumetric video to obtain the recommended video content corresponding to the recommended viewpoint group, and display the recommended video content corresponding to the recommended viewpoint group on the video client.

[0057] On the one hand, an embodiment of the present application provides a computer device, the computer device comprising: a processor and a memory;

[0058] The processor is connected to the memory, wherein the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method in any aspect of the embodiments of the present application.

[0059] On one hand, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is suitable for being loaded and executed by a processor so that a computer device having a processor executes the method in any aspect of the embodiment of the present application.

[0060] In one aspect, embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of any aspect of the embodiments of the present application.

[0061] In an embodiment of the present application, when a server (i.e., an encoder) obtains G viewpoint groups of a volumetric video, it can use the i-th viewpoint group among the G viewpoint groups as a target viewpoint group, where i is a non-negative integer less than G. It should be understood that each of the G viewpoint groups can be used as a target viewpoint group, and the i-th viewpoint group is used as an example for illustration. In this way, the server (i.e., the encoder) can further establish a mapping relationship between the target viewpoint group and the target spatial position information for viewing the target video content, and then write the timed metadata information corresponding to the target viewpoint group generated based on the mapping relationship into the encapsulated data box corresponding to the volumetric video to obtain a first extended data box corresponding to the encapsulated data box. It should be understood that the first extended data box can contain timed metadata information corresponding to each of the G viewpoint groups. In other words, the timed metadata information in the first extended data box can be used to describe the one-to-one mapping relationship between each viewpoint group and the spatial position relationship for viewing the corresponding video content. In this way, after the server (i.e., the encoding end) encapsulates the encoded video streams associated with the G viewpoint groups based on the first extended data box, it can obtain a video media file for transmission to the video client (i.e., the decoding end). In this way, when the application client decapsulates the video media file and obtains the first extended data box with the timed metadata information added, it can quickly determine the target viewpoint group mapped to the target spatial position information based on the timed metadata information written in the first extended data box and the current spatial position information of the video client (e.g., the target spatial position information). Then, the target video content corresponding to the target viewpoint group can be decoded on the video client to play the video content corresponding to the specific viewpoint group, thereby improving the decoding and presentation efficiency of the video client during partial access of the volumetric video. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0063] Figure 1 This is an architectural diagram of a volumetric video system provided by an embodiment of the present application;

[0064] Figure 2 1 is a flow chart of a method for processing volumetric video data provided by an embodiment of the present application;

[0065] Figure 3 1 is a flow chart of a method for processing volumetric video data provided by an embodiment of the present application;

[0066] Figure 4 1 is a flow chart of a method for processing volumetric video data provided by an embodiment of the present application;

[0067] Figure 5 1 is a flow chart of a method for processing volumetric video data provided by an embodiment of the present application;

[0068] Figure 6 1 is a schematic structural diagram of a volumetric video data processing device provided in an embodiment of the present application;

[0069] Figure 7 1 is a schematic structural diagram of a volumetric video data processing device provided in an embodiment of the present application;

[0070] Figure 8 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0071] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0072] See Figure 1 , Figure 1 is an architectural diagram of a volumetric video system provided in an embodiment of the present application; Figure 1 As shown, the volumetric video system includes an encoding device and a decoding device. The encoding device can be a computer device used by the volumetric video provider, which can be a terminal (such as a PC (personal computer), a smart mobile device (such as a smartphone), etc.) or a server. The decoding device can be a computer device used by the volumetric video user, which can be a terminal (such as a PC (personal computer), a smart mobile device (such as a smartphone), or a VR device (such as a VR helmet, VR glasses, etc.)). The volumetric video data processing process includes data processing on the encoding device side and data processing on the decoding device side.

[0073] The data processing process on the encoding device side mainly includes: (1) the process of acquiring the media content of the volumetric video; (2) the process of encoding and encapsulating the volumetric video file. The data processing process on the decoding device side mainly includes: (1) the process of decapsulating and decoding the volumetric video file; (2) the process of rendering the volumetric video. In addition, the transmission process involving the volumetric video between the encoding device and the decoding device can be carried out based on various transmission protocols. The transmission protocols here may include but are not limited to: DASH (Dynamic Adaptive Streaming over HTTP, dynamic adaptive streaming media transmission) protocol, HLS (HTTP Live Streaming, dynamic bit rate adaptive transmission) protocol, SMTP (Smart Media Transport Protocol, intelligent media transmission protocol), TCP (Transmission Control Protocol, transmission control protocol), etc.

[0074] The following will be combined Figure 1 , each process involved in the data processing of volumetric video is introduced in detail.

[0075] 1. Data processing on the encoding device side:

[0076] (1) The process of acquiring and producing the media content of volumetric video.

[0077] 1) The process of acquiring the media content of volumetric video.

[0078] Volumetric video media content is obtained by capturing real-world audio and visual scenes using a capture device. In one implementation, the capture device can refer to a hardware component within the encoding device, such as a microphone, camera, or sensor on a terminal. In another implementation, the capture device can also be a hardware device connected to the encoding device, such as a camera connected to a server, providing volumetric video media content acquisition services to the encoding device. The capture device may include, but is not limited to, audio equipment, video equipment, and sensor equipment. Audio equipment may include audio sensors and microphones. Video equipment may include standard cameras, stereo cameras, light field cameras, and other devices. Sensor equipment may include laser equipment, radar equipment, and other devices. There can be multiple capture devices, deployed at specific locations in real space to simultaneously capture audio and video content from different angles within that space, with the captured audio and video content synchronized in both time and space. In embodiments of the present application, the three-dimensional media content captured by capture devices deployed at specific locations to provide a multi-degree-of-freedom viewing experience may be referred to as volumetric video.

[0079] 2) The production process of volumetric video media content.

[0080] It should be understood that the process of producing the media content of the volumetric video involved in the embodiments of the present application can be understood as the process of producing the content of the volumetric video, and the content production of the volumetric video here is mainly produced by multi-viewpoint video, point cloud data, light field and other forms of content captured by cameras or camera arrays deployed in multiple locations. For example, the encoding device can convert the volumetric video from a three-dimensional representation to a two-dimensional representation. The volumetric video here can contain geometric information, attribute information, placeholder information and atlas data, etc. The volumetric video generally needs to be processed specifically before encoding. For example, point cloud data needs to be cut and mapped before encoding. For example, before encoding, the different viewpoints of the multi-view video generally need to be grouped to distinguish between the main viewpoint and the auxiliary viewpoint in each group.

[0081] Specifically, ① the 3D representation data of the captured input volumetric video (i.e., the aforementioned point cloud data) is projected onto a 2D plane, typically using orthogonal projection, perspective projection, or ERP projection. The projected volumetric video is represented by data from geometric components, placeholder components, and attribute components. The geometric component data provides the position information of each point in the volumetric video in 3D space, the attribute component data provides additional attributes of each point in the volumetric video (such as texture or material information), and the placeholder component data indicates whether the data in other components is associated with the volumetric video.

[0082] ② Process the component data of the 2D representation of the volumetric video to generate tiles. Based on the position of the volumetric video represented in the geometric component data, the 2D plane area where the 2D representation of the volumetric video is located is divided into multiple rectangular areas of different sizes. Each rectangular area is a tile, and the tile contains the necessary information to back-project the rectangular area into 3D space.

[0083] ③ Pack the tiles to generate atlases, placing the tiles in a two-dimensional grid and ensuring that the valid parts of each tile do not overlap. The tiles generated by a volumetric video can be packaged into one or more atlases;

[0084] ④ Generate corresponding geometric data, attribute data, and placeholder data based on the atlas data, and combine the atlas data, geometric data, attribute data, and placeholder data to form the final representation of the volumetric video in a two-dimensional plane.

[0085] It should be noted that, during the content production process of volumetric video, the geometry component is mandatory, the placeholder component is conditionally mandatory, and the attribute component is optional.

[0086] Furthermore, it should be noted that since panoramic video can be captured using a capture device, after being processed by an encoding device and transmitted to a decoding device for corresponding data processing, the user on the decoding device needs to perform certain specific actions (such as head rotation) to view the 360-degree video information. However, performing non-specific actions (such as moving the head) does not result in corresponding video changes, resulting in a poor VR experience. Therefore, additional depth information that matches the panoramic video is required to provide users with greater immersion and a better VR experience. This involves 6DoF (Six Degrees of Freedom) production technology. When users can move relatively freely in a simulated scene, it is called 6DoF. When using 6DoF production technology to produce volumetric video content, the capture device generally uses a light field camera, laser equipment, radar equipment, etc. to capture point cloud data or light field data in space.

[0087] (2) The process of encoding and file encapsulation of volumetric video.

[0088] The captured audio content can be directly audio-encoded to form an audio stream of the volumetric video. The captured video content can be video-encoded to obtain a video stream of the volumetric video. It should be noted here that if 6DoF production technology is used, a specific encoding method (such as a point cloud compression method based on traditional video encoding) needs to be used for encoding during the video encoding process. The audio stream and the video stream are encapsulated in a file container according to the file format of the volumetric video (such as ISOBMFF (ISO Base MediaFile Format, ISO media file format)) to form a media file resource of the volumetric video. The media file resource can be a media file of the volumetric video formed by a media file or a media fragment; and according to the file format requirements of the volumetric video, the media presentation description information (MPD) is used to record the metadata of the media file resource of the volumetric video. The metadata here is a general term for information related to the presentation of the volumetric video. The metadata may include description information of the media content, timing metadata information that describes the mapping relationship between each constructed viewpoint group and the spatial position information of the viewed media content, description information of the view window, and signaling information related to the presentation of the media content, etc. As Figure 1 As shown, the encoding device will store the media presentation description information and media file resources formed after the data processing process. The media presentation description information and media file resources here can be encapsulated into a media video file for sending to the decoding device according to a specific media file format.

[0089] Specifically, such as Figure 1As shown, the collected audio will be encoded into the corresponding audio code stream, the geometric information, attribute information and placeholder information of the volumetric video can adopt the traditional video encoding method, and the atlas data of the volumetric video can adopt the entropy encoding method. Then, the encoded media is encapsulated in a file container according to a certain format (such as ISOBMFF, HNSS) and combined with the metadata and window metadata that describe the attributes of the media content to form a media file or an initialization segment and a media segment according to a specific media file format. In an embodiment of the present application, when the encoding device forms a media file or an initialization segment and a media segment according to a specific media file format, the media file or the initialization segment and the media segment can be collectively referred to as a video media file, and then the obtained video media file can be sent to Figure 1 The decoding device shown.

[0090] 2. Data processing on the decoding device side:

[0091] (3) The process of decapsulating and decoding volumetric video files;

[0092] The decoding device can dynamically and adaptively obtain the volumetric video's media file resources and corresponding media presentation description information from the encoding device based on recommendations from the encoding device or based on user needs on the decoding device. For example, the decoding device can determine the user's orientation and position based on head / eye / body tracking information, and then dynamically request the corresponding media file resources from the encoding device based on the determined orientation and position. The media file resources and media presentation description information are transmitted from the encoding device to the decoding device via a transport mechanism (such as DASH and SMT). The file decapsulation process on the decoding device is the inverse of the file encapsulation process on the encoding device. The decoding device decapsulates the media file resources according to the volumetric video file format (e.g., ISO media file format) to obtain audio and video streams. The decoding process on the decoding device is the inverse of the encoding process on the encoding device. The decoding device performs audio decoding on the audio stream to restore the audio content, and the decoding device performs video decoding on the video stream to restore the video content.

[0093] (4) Volumetric video rendering process.

[0094] The decoding device renders the audio content obtained by audio decoding and the video content obtained by video decoding according to the rendering-related metadata in the media presentation description information corresponding to the media file resource. Once the rendering is completed, the playback output of the image is realized.

[0095] The volumetric video system supports data boxes (Boxes). Data boxes are data blocks or objects containing metadata, that is, data boxes contain metadata for corresponding media content. Volumetric video can include multiple data boxes, such as an ISO Base Media File Format Box (ISOBMFF Box), which contains metadata describing the corresponding information during file encapsulation. For example, it can specifically include constructed timed metadata information or constructed recommended metadata information. The ISO file encapsulation box can include a first extended data box and a second extended data box. The metadata information provided by the first extended data box (e.g., timed metadata information) is used to describe the correspondence between the viewpoint groups (i.e., the aforementioned groups) of the volumetric video and the corresponding media content. The metadata information provided by the second extended data box (e.g., recommended metadata information) is used to describe the correspondence between the recommended viewpoint groups associated with the content creator and the corresponding media content. For ease of understanding, in this embodiment of the application, the first and second extended data boxes may be collectively referred to as extended data boxes.

[0096] In an embodiment of the present application, in order to improve the decoding and presentation efficiency of a decoding terminal, an embodiment of the present application proposes a partial access strategy for volumetric video based on the metadata information provided by the above-mentioned extended data box. For example, based on the partial access strategy for volumetric video, when the encoding device obtains the viewpoint group of the volumetric video, it can obtain different metadata indication information based on the determined shooting intent of the viewpoint group of the volumetric video. The decoding device can adaptively provide different decoding and presentation requirements to the user based on the different indication information of the metadata information provided by the extended data box. It can be understood that the shooting intent here can be used to indicate whether there is a viewpoint group in the viewpoint group of the volumetric video that matches the specified viewpoint group specified by the content producer.

[0097] If it does not exist, it indicates that each view group in the volumetric video is independent of the content creator's filming intent. Based on the video content corresponding to each view group, a mapping relationship can be constructed between each view group in the volumetric video and spatial position information (e.g., 3D spatial area and user position information). These constructed mapping relationships can be used as a basis for selecting different view groups in a decoding device, generating timed metadata information corresponding to each view group. In this way, when the encoding device writes the generated timed metadata information corresponding to each view group into the encapsulation data box, it can obtain a first extension data box corresponding to the encapsulation data box. Based on this first extension data box, the encoded video stream associated with each view group (i.e., the media file resource) can be encapsulated to obtain a video media file (e.g., video media file A) for the volumetric video. Based on this, when the decoding device obtains the video media file, it can decapsulate the video media file to obtain the first extension data box and the encoded video stream associated with each view group. At this time, the decoding device can quickly select the viewpoint group (for example, viewpoint group 1) corresponding to the spatial position information that matches the user's current spatial position information based on the timed metadata information of the volumetric video provided by the first extended data box and the user's current spatial position information (for example, the user's current viewing position information), so as to partially decode and present the media content corresponding to viewpoint group 1 from these encoded video streams.

[0098] Optionally, if such a view group exists, the found view group can be directly used as the recommended view group. In this case, the encoding device can determine the identifier of the recommended view group as the recommended identifier, and then use metadata describing the recommended identifier as the recommended metadata information. This recommended metadata information can be written into the encapsulation data box corresponding to the volumetric video, thereby obtaining the aforementioned second extended data box. In this way, upon obtaining the encoded video streams associated with each view group, the encoding device can directly encapsulate these encoded video streams based on the second extended data box to obtain another video media file (e.g., video media file B). Based on this, when the decoding device obtains the video media file, it can decapsulate the video media file to obtain the aforementioned second extended data box and the encoded video streams associated with each view group. In this case, the decoding device can also directly determine the recommended view group corresponding to the recommended identifier based on the recommended metadata information provided in the second extended data box, and then decode and present the media content corresponding to the recommended view group.

[0099] The specific process of the encoding device obtaining the timed metadata information corresponding to each viewpoint group based on the above partial access strategy, or obtaining the recommended metadata information for representing the shooting intention of the content shooter based on the above partial access strategy, can be found in the following Figure 2-Figure 3 The specific process of partial decoding implemented by the decoding device based on the above partial access strategy can be found in the following Figure 4-Figure 5 Description of the corresponding embodiment.

[0100] Further, see Figure 2 , Figure 2 This is a flow chart of a method for processing volumetric video data provided by an embodiment of the present application. The method can be performed by an encoding device in a volumetric video system, for example, the encoding device can be a server. The method can include the following steps S101-S105:

[0101] Step S101: Obtain G viewpoint groups of the volumetric video, and use the i-th viewpoint group among the G viewpoint groups as the target viewpoint group;

[0102] Wherein, i is a non-negative integer less than G;

[0103] Specifically, upon obtaining V viewpoints of the volumetric video, the server can group the V viewpoints based on the viewpoint dependencies between the V viewpoints to obtain G viewpoint groups of the volumetric video. V represents the number of viewpoints in the volumetric video. Since volumetric video here primarily refers to multi-viewpoint video, the number of viewpoints in the volumetric video (i.e., V) is a positive integer greater than or equal to 2. It should be understood that the viewpoint dependencies here are determined by the content relevance between the video content corresponding to each of the V viewpoints.

[0104] It should be understood that multi-viewpoint video is usually captured by a camera array from multiple angles to form the scene's texture information (color information, etc.) and depth information (spatial distance information, etc.), and then the mapping information from the 2D plane frame to the 3D presentation space is added to form 6DoF media that can be consumed on the user side. It should be understood that each camera in the camera array can correspond to a corresponding camera identifier. In the process of reconstructing a three-dimensional scene, it is necessary to use one or more viewpoints of volumetric video to synthesize and render the target viewpoint based on the user's viewing position, direction, etc. The volumetric video corresponding to the auxiliary viewpoint needs to be synthesized based on the volumetric video data of the basic viewpoint.

[0105] It is understood that embodiments of the present application can provide corresponding viewpoint information of a volumetric video through a viewpoint information structure (e.g., ViewInfoStruct) of the volumetric video. For example, the viewpoint information structure of the volumetric video can be used to define the viewpoint identifier of each of the V viewpoints of the volumetric video, the viewpoint group identifier of the viewpoint group to which each viewpoint belongs, and whether each viewpoint carries a valid basic viewpoint identifier, etc., based on the identifier of the camera parameters.

[0106] For ease of understanding, further reference is made to Table 1, which is used to indicate the syntax of a viewpoint information structure of a volumetric video provided in an embodiment of the present application:

[0107] Table 1

[0108]

[0109] The semantics of the syntax shown in Table 1 above are as follows; view_id indicates the viewpoint identifier of the viewpoint. view_group_id indicates the viewpoint group identifier of the viewpoint group to which the viewpoint belongs. view_descritption provides a textual description of the viewpoint, a UTF-8 string ending with a null value. When the field value of basic_view_flag (i.e., the basic viewpoint identifier field) is 1 (i.e., a valid value), it indicates that the current viewpoint is the basic viewpoint; conversely, when the field value of basic_view_flag (i.e., the basic viewpoint identifier field) is 0 (i.e., an invalid value), it indicates that the current viewpoint is not the basic viewpoint.

[0110] Furthermore, the server can determine the viewpoint dependency relationship between the above-mentioned V viewpoints based on the content relevance between the viewpoint contents corresponding to each viewpoint, and then group the V viewpoints based on the viewpoint dependency relationship between the above-mentioned V viewpoints to obtain G viewpoint groups of the volumetric video.

[0111] For ease of understanding, here, an example is used in which G view groups include view group 1, view group 2, view group 3, view group 4, and view group 5. It is understood that embodiments of the present application can provide view grouping information for volumetric video via a view grouping information structure for volumetric video (e.g., ViewGroupInfoStruc), where one view grouping information structure can be used to describe one or more viewpoints. For further understanding, please refer to Table 2, which shows the syntax of a view grouping information structure for volumetric video provided in embodiments of the present application:

[0112] Table 2

[0113]

[0114]

[0115] The semantics of the syntax shown in Table 2 above are as follows: view_group_id indicates the view group identifier for the view group of the volumetric video. For example, the view group identifier of view group 1 above can be identifier A1. view_group_descritption is a text description of the view group. The text description of the view group is a UTF-8 string terminated by a null value. num_views indicates the number of viewpoints in the view group. For example, the number of viewpoints in view group 1 above should be at least one. For ease of understanding, here we take the example of three viewpoint data in view group 1. These three viewpoints (for example, view 1a, view 1b, and view 1c) include one basic viewpoint and two auxiliary viewpoints. Wherein, view_id indicates the viewpoint identifier of the viewpoint in the viewpoint group. For example, the viewpoint identifiers of the three viewpoints in viewpoint group 1 (e.g., viewpoint 1a, viewpoint 1b, and viewpoint 1c) can be identifiers A11, A12, and A13, that is, the viewpoint identifier of the jth viewpoint in the viewpoint group can be recorded as A1j. For any viewpoint group of the divided volumetric video, the number of viewpoints contained in any divided viewpoint group is a positive integer greater than or equal to 1. Wherein, basic_view_flag is the basic viewpoint identifier field in the viewpoint grouping information structure. In this case, when the field value of the basic viewpoint identifier field is 1, it indicates that the viewpoint is a basic viewpoint. For example, for the first viewpoint in the above-mentioned viewpoint group 1 (for example, the viewpoint indicated when j=0) is viewpoint 1a, if the basic viewpoint identification field corresponding to the viewpoint 1a is 1, it indicates that the viewpoint 1a is the basic viewpoint; conversely, optionally, if the basic viewpoint identification field corresponding to the viewpoint 1a is 0, it indicates that the viewpoint 1a is not the basic viewpoint, for example, the viewpoint 1a is an auxiliary viewpoint.

[0116] Furthermore, it can be understood that when the server obtains G viewpoint groups of the volumetric video, in order to help the user corresponding to the video client select a suitable target viewpoint group, the embodiment of the present application proposes a strategy for partial access to the volumetric video, which is applied to the server, video client, and intermediate nodes.

[0117] Before the server uses the i-th viewpoint group among the G viewpoint groups as the target viewpoint group, it must first search among the G viewpoint groups to see if there is a viewpoint group that matches the designated viewpoint group associated with the content creator's shooting intent. It should be understood that the designated viewpoint group here is related to the shooting intent of the content creator who shot the volumetric video.

[0118] Specifically, when the server obtains the video association information associated with the volumetric video, it can search for the designated viewpoint group associated with the content producer in the video association information. Once the content producer specifies a viewpoint group that fits his or her shooting intention (i.e., the designated viewpoint group) when shooting the volumetric video, the video association information will carry the aforementioned designated viewpoint group. Conversely, if the content producer does not specify a designated viewpoint group that fits his or her shooting intention when shooting the volumetric video, the video management information will not carry the aforementioned designated viewpoint group. For example, when shooting the volumetric video, the content producer can refer to the viewpoint group related to his or her shooting intention as a designated viewpoint group, and use the designated viewpoint group that fits his or her shooting intention and / or the camera parameters used to shoot the volumetric video (i.e., the data information used to characterize the content producer's shooting intention) as the video association information associated with the volumetric video to transmit the video association information to the server.

[0119] Based on this, when the server obtains the video association information associated with the volumetric video, it can search for the specified viewpoint group associated with the content producer in the video association information; if the server does not find the specified viewpoint group associated with the content producer in the video association information, it can be determined that these G viewpoint groups are all viewpoint groups that are irrelevant to the content producer's shooting intention, and then the i-th viewpoint group that is irrelevant to the content producer's shooting intention can be obtained from these G viewpoint groups as the target viewpoint group, and the following S102 can be notified to execute, so as to construct a mapping relationship between the target viewpoint group and the target spatial position information for viewing the target video content according to the target video content corresponding to the target viewpoint group, and then the timed metadata information corresponding to the target viewpoint group can be generated based on the constructed mapping relationship.

[0120] Optionally, if a specified viewpoint group associated with the content producer is found in the video-related information, the server can use the found specified viewpoint group as a recommended viewpoint group, and can obtain the i-th viewpoint group related to the content producer's shooting intention from the G viewpoint groups based on the recommended viewpoint group as the target viewpoint group. At this time, the target viewpoint group is a viewpoint group related to the content producer's shooting intention. The server will skip the following steps S102-S104, directly determine the identifier of the target viewpoint group at this time as the recommended identifier, and use the metadata information used to describe the recommended identifier as the recommended metadata information of the target viewpoint group, write the recommended metadata information into the encapsulation data box corresponding to the volumetric video, and obtain a second extended data box corresponding to the encapsulation data box; further, the server can obtain the encoded video streams associated with the G viewpoint groups, and can encapsulate the encoded video streams based on the second extended data box to obtain a video media file of the volumetric video, and can send the video media file to the video client, so that when the video client obtains the second extended data box based on the video media file, it can quickly display the target video content corresponding to the target viewpoint group indicated by the recommended identifier on the video client based on the recommended metadata information in the second extended data box.

[0121] Step S102: constructing a mapping relationship between the target viewpoint group and target spatial position information for viewing the target video content based on the target video content corresponding to the target viewpoint group, and generating timed metadata information corresponding to the target viewpoint group based on the mapping relationship;

[0122] It should be understood that the server Figure 1 Under the system architecture of the volumetric video described, several descriptive fields are added to the system layer. For example, some file encapsulation level field extensions are added to the system layer to extend the ISOBMEF data box (i.e., the above-mentioned encapsulation data box). For example, the embodiment of the present application adds a view group timing metadata track to the above-mentioned encapsulation data box extension. In other words, the embodiment of the present application extends the encapsulation data box on the basis of the existing encapsulation data box to obtain a corresponding extended data box. Among them, the view group timing metadata track can be added to the extended data box. The view group timing metadata track is determined by the view group metadata sample entry (i.e., simply referred to as the sample entry) and the view group metadata sample (i.e., simply referred to as the sample).

[0123] It can be understood that the view group metadata sample entry and view group element sample here can be used to constitute the view group timed metadata track of the volumetric video, and the view group timed metadata track here can be used to index into one or more atlas data tracks associated with the volumetric video.

[0124] It should be understood that the view group timed metadata track can be used to indicate the correspondence between different view groups and media content in the subsequently packaged video media file. For example, the timed metadata track can be directly associated with the corresponding atlas data track, rather than directly associated with the video component track of the volumetric video.

[0125] It can be understood that the view group timed metadata track can be quickly indexed to related tracks and track groups through the track index string "for example, the 'cdtg' string". For example, the 'cdtg' string can be used to directly index an atlas data track associated with the view group timed metadata track. At this time, a view group timed metadata track can be used to associate an atlas data track.

[0126] Optionally, the view group timed metadata track can also index one or more atlas data tracks through another track index string, "for example, a 'cdsc' string". For example, when a view group timed metadata track is used to associate one or more atlas data tracks, the view group timed metadata track can be used to describe each atlas data track separately, and the 'cdsc' string here can have an index sharing function, that is, all atlas data tracks can be indexed through the view group timed metadata track. It can also be understood that the samples in each view group timed metadata track can be marked as synchronized samples. The samples in each view group timed metadata track refer to data boxes with metadata information.

[0127] Among them, it should be understood that the atlas data track here can be used to describe the mapping relationship between the above-mentioned 2D plane and the 3D plane. Since the atlas data track can also be indexed to the component track through other track index strings, the component track contains specific video information such as texture and color. However, in the embodiment of the present application, in order to save computing resources and improve the decoding and presentation efficiency of subsequent video clients, it is emphasized here that the server can directly associate with the corresponding atlas data track through the viewpoint group timing metadata track, so that the subsequent video client can quickly obtain the mapping relationship between the above-mentioned 2D plane and the 3D plane through the atlas data track directly related to the viewpoint group timing metadata track during the decoding process, and then reconstruct the above-mentioned two-dimensional volumetric video into three-dimensional space.

[0128] The definition of the viewpoint group metadata sample entry is as follows:

[0129] Sample entry type: 'vgme';

[0130] Contained in: Sample Description Box ('stsd');

[0131] Is it mandatory: No;

[0132] Quantity: 0 or 1;

[0133] As can be seen from the definition of the view group metadata sample entry above, the sample entry type of the view group metadata sample entry (i.e., the sample entry) here is "6vpt", and the view group metadata sample entry is defined by ViewGroupMetadataSampleEntry. It can be understood that the view group metadata sample entry here must contain ViewGroupStaticMetaBox (i.e., the view group static metadata box). The view group static metadata box here can be used to describe static sample group metadata. The static sample group metadata in the view group static metadata box can be described by the static view group metadata field. For example, if the field value of the static view group metadata field is 1, it means that all sample group metadata corresponding to the view group metadata in the sample entry remain unchanged, and then the mapping relationship of the view groups in the above G view groups can be described by the static sample group metadata recorded in the sample entry. For another example, if the field value of the static view group metadata field is 0, it means that the view group metadata will change over time. At this time, the viewpoint group metadata that changes with time can be collectively referred to as dynamic viewpoint group metadata, and the viewpoint group metadata that changes with time needs to be recorded in the viewpoint group metadata sample corresponding to the viewpoint group sample entry.

[0134] The metadata information carried by the view group sample entry is static view group metadata, which can be used to describe view groups among the G view groups whose mapping relationships remain unchanged (for example, view group 1, view group 2, and view group 3 mentioned above). The metadata information carried by the view group metadata sample is dynamic view group metadata, which is used to describe view groups among the G view groups whose mapping relationships change over time (for example, view group 4 and view group 5). The timed metadata information written into the view group sample entry and view group metadata sample to describe the mapping relationships of the corresponding view groups is not limited herein.

[0135] Among them, the static viewpoint group metadata field is deployed in the viewpoint group static metadata box of the first extended data box; if the field value of the static viewpoint group metadata field is a numerical value that remains unchanged for describing the mapping relationship of the viewpoint groups in the G viewpoint groups, then the static viewpoint group metadata recorded in the viewpoint group metadata sample entry of the first extended data box: includes the viewpoint group identifier of each viewpoint group in the G viewpoint groups.

[0136] For ease of understanding, further information is provided in Table 3, which shows the syntax of the view group element sample entry provided in an embodiment of the present application:

[0137] Table 3

[0138]

[0139]

[0140] Among them, as shown in Table 3 above, the semantics of the view group element sample entry: ViewPositionStruct (i.e., the viewpoint position structure) indicates the viewing position coordinate information of the user corresponding to the video client in the overall space of the multi-view video. For example, as shown in Table 3 above, ViewPositionStruct (i.e., the viewpoint position structure) specifies the 3D spatial position of the viewpoint and the GPS position of the viewpoint. Among them, view_position (i.e., the viewpoint position): indicates the position of the user's viewing. The structure contains the specific x, y, and z coordinate information of the user in the overall space of the multi-view media. Among them, view_orientation (i.e., the viewpoint orientation): indicates the direction in which the user is viewing. The structure contains the rotation information of the user's head. Among them, position_range_flag (i.e., the position movement identification field): indicates whether the user's movement range information is included. Among them, it can be understood that if the field value of the position movement identification field is a valid value (for example, 1), then the position movement identification field can be used to indicate that the viewpoint position structure (i.e., ViewPositionStruct in Table 3 above) contains the user's movement range information; optionally, if the field value of the position movement identification field is an invalid value (for example, 0), then the position movement identification field can be used to indicate that the viewpoint position structure (i.e., ViewPositionStruct in Table 3 above) does not contain the user's movement range information. Among them, position_range: indicates the user's movement range along the x, y, and z axes with view_position as the coordinate starting point.

[0141] As shown in Table 3 above, the view group metadata sample entry must contain an extended ViewGroupStaticMetaBox (i.e., a view group static metadata box, which is deployed in the first extended data box). The first extended data box contains the following file encapsulation-level extension fields. The type field of the view group-associated determination information can be the extended field view_associated_info_type shown in Table 3 above, and the static view group metadata field can be the extended field static_view_group_meta shown in Table 3 above.

[0142] As shown in Table 3 above, view_associated_info_type: indicates the type of determination information associated with the viewpoint group. In one embodiment, if the value of this field is a first value (e.g., 0), it indicates that the associated information is the spatial area of ​​the media content viewed by the user, that is, the target spatial position information having a mapping relationship with the matching viewpoint group is the three-dimensional spatial area information of the matching video content displayed on the video client. Optionally, in another embodiment, if the value of this field is a second value (e.g., 1), it indicates that the associated information is the user viewing position information, that is, the target spatial position information having a mapping relationship with the matching viewpoint group is the viewing position coordinate information of the user viewing the matching video content on the video client. Optionally, in yet another embodiment, if the value of this field is a third value (e.g., 2), it indicates that the associated information is a combination of the spatial area of ​​the media content viewed by the user and the user viewing position information, that is, the target spatial position information having a mapping relationship with the target viewpoint group is determined by the combination of the three-dimensional spatial area information and the viewing position coordinate information.

[0143] It can be seen that the target spatial position information is determined by the type field of the judgment information associated with the viewpoint group recorded by the server in the viewpoint group metadata sample entry of the first extended data box; the type field of the judgment information is deployed in the viewpoint group static metadata box of the viewpoint group metadata sample entry; if the type field of the judgment type is a first value, the target spatial position information having a mapping relationship with the matching viewpoint group is the three-dimensional spatial area information of the matching video content displayed on the video client; if the type field of the judgment type is a second value, the target spatial position information having a mapping relationship with the matching viewpoint group is the viewing position coordinate information of the user watching the matching video content on the video client; if the type field of the judgment type is a third value, the target spatial position information having a mapping relationship with the target viewpoint group is jointly determined by the three-dimensional spatial area information and the viewing position coordinate information.

[0144] As shown in Table 3 above, static_view_group_meta (i.e., the static view group metadata field) indicates whether the view group metadata corresponding to the view group in the G view groups remains unchanged in all samples corresponding to the sample entry. If the field value of the static view group metadata field is a numerical value used to describe that the mapping relationship of the view groups in the G view groups remains unchanged (i.e., if the value of this field is 1), it means that the view group metadata corresponding to the view group in the G view groups remains unchanged in all samples corresponding to the sample entry. In this way, the static view group metadata recorded in the view group metadata sample entry of the first extended data box may include the view group identifier of each view group in the G view groups.

[0145] As shown in Table 3 above, view_group_num (i.e., the number of all viewpoint groups recorded in the above sample entry, for example, the G above): indicates the number of viewpoint groups. view_group_id (i.e., the viewpoint group identifier of the i-th viewpoint group in the G viewpoint groups): indicates the identifier of the viewpoint group. region_num (i.e., the number of spatial regions corresponding to the i-th viewpoint group): indicates the number of 3D spatial regions. 3DSpatialRegionStruct[j] (i.e., 3D spatial region structure): indicates the spatial region of the media content viewed by the user. ViewPositionStruct (i.e., viewpoint position structure): indicates the user viewing position information.

[0146] Among them, for ease of understanding, further, please refer to Table 4, which is the syntax of the spatial region information structure provided in the embodiment of the present application. Among them, the spatial region information structure includes at least the following two substructures: a 3D spatial region structure (3DSpatialRegionStruct) and a 3D border structure (3DBoundingBoxStruct). Among them, the 3D spatial region structure (3DSpatialRegionStruct) provides the spatial region information of the volumetric video (including the offset of the spatial region x, y, and z axes, the width, height, and depth of the 3D spatial region), and the 3D border structure (3DBoundingBoxStruct) provides the border information of the volumetric video. This means that the syntax of the 3D spatial region structure called by Table 3 above can be specifically referred to the following Table 4:

[0147] Table 4

[0148]

[0149] Among them, as shown in Table 4 above, the semantics of the spatial region information structure are as follows: the 3d_region_id: indicates the identifier of the spatial region; x, y, z in the 3D point structure: respectively indicate the x, z, y coordinate values ​​of the 3D point in the Cartesian coordinate system. cuboid_dx, cuboid_dy, cuboid_dz: indicate the dimensions of the rectangular sub-region in the Cartesian coordinate system relative to the anchor point along the x, y, z axes. anchor: indicates a 3D point in the Cartesian coordinate system that serves as the anchor of the 3D spatial region. bb_dx, bb_dy, bb_dz: indicate the dimensions of the extension of the 3D bounding box of the entire volumetric video relative to the origin (0, 0, 0) along the x, y, z axes in the Cartesian coordinate system. dimensions_included_flag: an identifier indicating whether the spatial region dimensions have been marked. When the value of this field is 1, it indicates that the spatial dimension region indicated by 3DSpatialRegionStruct has been marked.

[0150] The rotation structure in Table 3 above indicates the rotation information required to convert the local coordinate axis to the global coordinate axis. The rotation information is expressed in Euler angles or quaternions. In the case of stereo panoramic video, this rotation structure is applicable to each binocular view. For ease of understanding, please refer to Table 5, which is the syntax of the rotation structure provided in the embodiment of the present application:

[0151] Table 5

[0152]

[0153] As shown in Table 5 above, the semantics of the rotation structure are as follows: 3D_rotation_type indicates the representation type of the rotation information. A value of 0 indicates that the rotation information is given in the form of Euler angles; a value of 1 indicates that the rotation information is given in the form of quaternions. The remaining values ​​are reserved. Among them, rotation_yaw, rotation_pitch, and rotation_roll refer to the yaw angle, pitch angle, and roll angle along the X-axis, Y-axis, and Z-axis, respectively, and are used for the conversion of the local coordinate axis of the unit sphere to the global coordinate axis, with 2 -16 For precision, it is related to the global coordinate axis. The range of rotation_yaw is [-180°*2 16 ,180°*2 16 –1], the range of rotation_pitch is [-90°*2 16 ,90°*2 16 ], the range of rotation_roll is [-180°*2 16 ,180°*2 16 –1]. Among them, rotation_x, rotation_y, rotation_z and rotation_w indicate the values ​​of the quaternion x, y, z and w components respectively, which are used to transform the local coordinate axes of the unit sphere to the global coordinate axes.

[0154] Optionally, if the field value of the static viewpoint group metadata field deployed in the viewpoint group static metadata box of the first extended data box is a numerical value associated with the dynamic viewpoint group metadata (that is, if the value of this field is 0), and the dynamic viewpoint group metadata is used to describe a viewpoint group among the G viewpoint groups whose mapping relationship changes over time, then it means that there is a viewpoint group among the G viewpoint groups whose viewpoint group metadata changes over time. In this case, the dynamic viewpoint group metadata recorded in the viewpoint group element sample corresponding to the viewpoint group metadata sample entry includes: an identifier of a variable viewpoint group whose mapping relationship changes at the sample timestamp, and before the sample timestamp, the mapping relationship between the variable viewpoint group and the video content corresponding to the variable viewpoint group remains unchanged; the variable viewpoint group is a viewpoint group among the G viewpoint groups whose mapping relationship changes over time.

[0155] For example, it is understandable that, considering that a view group metadata sample can correspond to multiple view groups, if the static view group metadata field value of view group 2 recorded in the view group sample entry changes from a valid value (e.g., 1) to an invalid value (e.g., 0) at a certain sample timestamp (e.g., the playback timestamp corresponding to the 10th minute of the volumetric video), the server may update the view group 2 record with the mapping change at the 10th minute into the view group metadata sample. This means that before the 10th minute, the mapping relationship (also referred to as the correspondence relationship) indicated by the timed metadata information of view group 2 is valid and unchanged. That is, before the 10th minute, the timed metadata information of view group 2 is the static view group metadata. However, after the 10th minute, the timed metadata information of view group 2 needs to be further defined in the next sample (i.e., the next view group metadata sample). For example, after the 10th minute, the timed metadata information of view group 2 may be redefined as the dynamic view group metadata described above.

[0156] The view group metadata sample is used to indicate metadata information related to the corresponding view group. The metadata of a view group defined in the previous sample will remain unchanged until the metadata information of the view group is redefined in the next sample. For ease of understanding, please refer to Table 6 below, which shows the syntax of the view group metadata sample provided in the embodiment of the present application:

[0157] Table 6

[0158]

[0159]

[0160] As shown in Table 6 above, the semantics of the view group metadata sample are as follows: view_group_num: Indicates the number of view groups included in all current samples whose mapping relationships, as recorded in the view group metadata sample, change over time, at the current time being the sample timestamp. For example, for the five view groups (e.g., G = 5), the mapping relationships of four view groups (G1 = 4) change over time, while the mapping relationship of one view group (e.g., G2 = 1) remains unchanged for all samples in the sample entry. In this case, the timed metadata for this view group can be recorded in the sample entry shown in Table 3 above, and the timed metadata for the four view groups that have changed remain unchanged until the mapping relationships indicated by the timed metadata information for these view groups are redefined in the next sample. The number of G1 and G2 view groups in G is not limited here. Prior to the sample timestamp, the mapping relationships for the G1 view groups remain valid and unchanged. view_group_id: Indicates the identifier of the i-th view group among the G1 view groups recorded in the view group metadata sample. Based on the field value of the type field of the determination information recorded in the sample entry, the spatial position information associated with the type field of the determination information is redefined in the next sample of the viewpoint metadata sample. For example, if the field value of the type field of the determination information is the first value (i.e., 0), then the target spatial position information mapped to the matching viewpoint group is the three-dimensional spatial region information of the matching video content displayed on the video client. In this case, region_num indicates the number of 3D spatial regions; 3DSpatialRegionStruct[j] indicates the spatial region of the media content viewed by the user. Optionally, if the type field of the determination type is the second value, then the target spatial position information mapped to the matching viewpoint group is the viewing position coordinate information of the user viewing the matching video content on the video client. In this case, ViewPositionStruct indicates the user viewing position information. Optionally, if the type field of the determination type is the third value, then the target spatial position information mapped to the target viewpoint group is determined jointly by the three-dimensional spatial region information and the viewing position coordinate information.

[0161] As shown in Table 6 above, the view group metadata sample also contains a recommended view group identifier field (i.e., recommended_view_group_flag). The recommended_view_group_flag field indicates whether a recommended view group is included. A value of 1 indicates that the sample contains a recommended view group, indirectly reflecting that a specific view group related to the content creator's filming intent exists among the G view groups. A value of 0 indicates that the sample does not contain a recommended view group, indirectly reflecting that a specific view group related to the content creator's filming intent does not exist among the G view groups. The rcmd_view_group_id field (i.e., the recommended view group identifier) ​​indicates the identifier of the view group recommended for presentation. It should be understood that the recommended view group identifier can be used to instruct a video client to quickly determine the recommended view group corresponding to the recommended identifier upon receiving the second extended data box. This allows the client to quickly decode the video content corresponding to the recommended view group from the encoded video stream, further improving decoding and presentation efficiency on the video client.

[0162] Step S103: writing the timed metadata information corresponding to the target viewpoint group into the encapsulated data box corresponding to the volumetric video to obtain a first extended data box corresponding to the encapsulated data box;

[0163] The first extended data box includes timing metadata information corresponding to each of the G viewpoint groups;

[0164] Step S104: Obtain coded video streams associated with the G viewpoint groups, and encapsulate the coded video streams based on the first extended data box to obtain a video media file of the volumetric video.

[0165] Step S105: Send the video media file to the video client.

[0166] It should be understood that when the server sends the video media file to the video client, the video client can display the target video content corresponding to the target viewpoint group on the video client according to the target spatial position information indicated by the timing metadata information corresponding to the target viewpoint group in the first extended data box when obtaining the first extended data box based on the video media file.

[0167] It should be understood that in the embodiment of the present application, when the designated viewpoint group related to the content producer's shooting intention is not found in the video association information, the i-th viewpoint group among the G viewpoint groups can be used as the target viewpoint group to construct a mapping relationship between the target viewpoint group and the target spatial position information. It should be understood that referring to the specific process of constructing the mapping relationship between the target viewpoint group and the target control position information, a mapping relationship can be further constructed between each viewpoint group among the G viewpoint groups and the spatial position information for viewing the video content of the corresponding viewpoint group, and then the metadata information used to describe the corresponding mapping relationship can be collectively referred to as the timed metadata information corresponding to each viewpoint group, so that the above steps S103-S104 can be further executed, and then the above-mentioned first extended data box can intelligently provide the timed metadata information corresponding to each viewpoint group. In this way, when the video client unpacks the first extended data box from the video media file, it can match the current spatial position information of the video client with the control position information indicated by the timed metadata information corresponding to a certain viewpoint group among the aforementioned G viewpoint groups. If the spatial position information that matches the current spatial position information of the video client is the aforementioned target spatial position information, the target video content of the target viewpoint group associated with the target spatial position information can be partially decoded from the encoded video stream associated with the G viewpoint groups, and then the target video content of the target viewpoint group can be displayed in the video client. For example, based on the mapping relationship from 2D plane to 3D plane indicated in the atlas track corresponding to the aforementioned target viewpoint group, the above-mentioned two-dimensional volumetric video can be reconstructed into three-dimensional space as accurately as possible, so as to ultimately render and present the three-dimensional target video content on the video client.

[0168] In an embodiment of the present application, when a server (i.e., an encoder) obtains G viewpoint groups of a volumetric video, it can use the i-th viewpoint group among the G viewpoint groups as a target viewpoint group, where i is a non-negative integer less than G. It should be understood that each of the G viewpoint groups can be used as a target viewpoint group, and the i-th viewpoint group is used as an example for illustration. In this way, the server (i.e., the encoder) can further establish a mapping relationship between the target viewpoint group and the target spatial position information for viewing the target video content, and then write the timed metadata information corresponding to the target viewpoint group generated based on the mapping relationship into the encapsulated data box corresponding to the volumetric video to obtain a first extended data box corresponding to the encapsulated data box. It should be understood that the first extended data box can contain timed metadata information corresponding to each of the G viewpoint groups. In other words, the timed metadata information in the first extended data box can be used to describe the one-to-one mapping relationship between each viewpoint group and the spatial position relationship for viewing the corresponding video content. In this way, after the server (i.e., the encoding end) encapsulates the encoded video streams associated with the G viewpoint groups based on the first extended data box, it can obtain a video media file for transmission to the video client (i.e., the decoding end). In this way, when the application client decapsulates the video media file and obtains the first extended data box with the timed metadata information added, it can quickly determine the target viewpoint group mapped to the target spatial position information based on the timed metadata information written in the first extended data box and the current spatial position information of the video client (e.g., the target spatial position information). Then, the target video content corresponding to the target viewpoint group can be decoded on the video client to play the video content corresponding to the specific viewpoint group, thereby improving the decoding and presentation efficiency of the video client during partial access of the volumetric video.

[0169] Further, see Figure 3 , Figure 3 This is a flow chart of a method for processing volumetric video data provided by an embodiment of the present application. The method can be performed by an encoding device in a volumetric video system, which can be the server described above. The method may include the following steps:

[0170] Step S201, obtaining G viewpoint groups of the volumetric video;

[0171] Specifically, the server may obtain V viewpoints of the volumetric video, and may group the V viewpoints based on viewpoint dependencies between the V viewpoints to obtain G viewpoint groups of the volumetric video;

[0172] Where V represents the number of viewpoints in the volumetric video and is a positive integer greater than or equal to 2. The viewpoint dependency is determined by the content relevance between the video content corresponding to each of the V viewpoints. That is, in embodiments of the present application, viewpoints corresponding to video content with high content relevance (e.g., content relevance greater than or equal to a relevance threshold) may be grouped into the same viewpoint group, while viewpoints corresponding to video content with low content relevance (e.g., content relevance less than a relevance threshold) may be grouped into different viewpoint groups.

[0173] Step S202: obtaining video association information associated with the volumetric video, and searching the video association information for a designated viewpoint group associated with the content producer;

[0174] The designated viewpoint group is related to the shooting intention of the content producer who shot the volumetric video. It should be understood that when the encoding device (i.e., the above-mentioned server) performs step S202, if the designated viewpoint group associated with the content producer is not found in the video-related information, the encoding device (i.e., the above-mentioned server) may further perform the following steps S203-S207, and then write the timed metadata information corresponding to each constructed viewpoint group into the encapsulation data box to obtain a first extended data box for indicating that the encoded video stream is to be encapsulated in the above-mentioned ISOBMFF file encapsulation format. Optionally, if the designated viewpoint group associated with the content producer is found in the video-related information, the encoding device (i.e., the above-mentioned server) may further perform the following steps S208-S211.

[0175] Step S203: If the designated viewpoint group associated with the content producer is not found in the video association information, then the i-th viewpoint group that is irrelevant to the content producer's shooting intention is obtained from the G viewpoint groups as the target viewpoint group;

[0176] Step S204 : constructing a mapping relationship between the target viewpoint group and target spatial position information for viewing the target video content based on the target viewpoint group, and generating timed metadata information corresponding to the target viewpoint group based on the mapping relationship.

[0177] Step S205 , writing the timed metadata information corresponding to the target viewpoint group into the encapsulated data box corresponding to the volumetric video to obtain a first extended data box corresponding to the encapsulated data box;

[0178] Among them, the first extended data box contains timing metadata information corresponding to each view group of the G view groups; it should be understood that the encoding device involved in the embodiment of the present application (i.e., the above-mentioned server) will write the timing metadata information into the encapsulated data box and expand the encapsulated data box. In the process, it will also add a recommended view group identification field to the view group metadata sample of the encapsulated data box, and set the field value of the recommended view group identification field to an invalid field value, and in the view group metadata sample of the first extended data box, the recommended view group identification field with an invalid field value is used as the first field indication information; the first field indication information is used to instruct the video client to obtain the timing metadata information of each view group in the G view groups from the first extended data box.

[0179] Step S206: Obtain coded video streams associated with the G viewpoint groups, and encapsulate the coded video streams based on the first extended data box to obtain a video media file of the volumetric video.

[0180] Step S207: Send the video media file to the video client so that when the video client obtains the first extended data box based on the video media file, the target video content corresponding to the target viewpoint group is displayed on the video client according to the target spatial position information indicated by the timing metadata information corresponding to the target viewpoint group in the first extended data box.

[0181] It should be understood that, for the sake of distinction, the embodiment of the present application may refer to the video media file encapsulated through steps S203 to S207 as a first video media file, and the video media file encapsulated through steps S208 to S211 as a second video media file.

[0182] The specific implementation of steps S203 to S207 can be found in the above Figure 2 The description of the specific process of obtaining the first extended data box and encapsulating the encoded video stream based on the first extended data box in the corresponding embodiment will not be repeated here.

[0183] Optionally, in step S208, if a designated viewpoint group associated with the content producer is found in the video-related information, the found designated viewpoint group is used as a recommended viewpoint group;

[0184] Step S209 : Based on the recommended viewpoint groups, an i-th viewpoint group related to the shooting intention of the content producer is obtained from the G viewpoint groups as a target viewpoint group.

[0185] Step S210: Determine the identifier of the target view group as the recommended identifier, use metadata information describing the recommended identifier as recommended metadata information for the target view group, and write the recommended metadata information into the encapsulated data box corresponding to the volumetric video to obtain a second extended data box corresponding to the encapsulated data box.

[0186] The recommendation identifier indicated by the recommendation metadata information written into the second extended data box can refer to the description of the viewpoint group metadata sample in the embodiment corresponding to Table 6 above, and will not be further described here.

[0187] It should be understood that when the encoding device (i.e., the above-mentioned server) writes the recommended metadata information into the encapsulation data box corresponding to the volumetric video, the encoding device (i.e., the above-mentioned server) may also add a recommended viewpoint group identification field associated with the recommendation identifier in the viewpoint group metadata sample of the encapsulation data box, and set the field value of the recommended viewpoint group identification field to a valid field value, and in the viewpoint group metadata sample of the second extended data box, use the recommended viewpoint group identification field with a valid field value as the second field indication information; the second field indication information is used to instruct the video client to obtain the recommended metadata information from the second extended data box; the video client is the client in the decoding device corresponding to the encoding device.

[0188] Step S211: Obtain coded video streams associated with G viewpoint groups, encapsulate the coded video streams based on the second extended data box to obtain a video media file of the volumetric video, and send the video media file to the video client. When the video client obtains the second extended data box based on the video media file, the video client displays the target video content corresponding to the target viewpoint group indicated by the recommendation identifier based on the recommendation metadata information in the second extended data box.

[0189] In an embodiment of the present application, the server may divide the multiple viewpoints of the volumetric video (i.e., the multiple viewpoints of the multi-view video described above) into different viewpoint groups, and may determine whether there is a viewpoint group in these video groups that matches the designated viewpoint group associated with the photographer's shooting intention. If so, the viewpoint group that matches the designated viewpoint group associated with the photographer's shooting intention (e.g., the i-th viewpoint group among the G viewpoint groups) may be used as a recommended viewpoint group. In this case, the recommended viewpoint group here may be regarded as the target viewpoint group described above, and the identifier of the recommended viewpoint group (i.e., the target viewpoint group) may be determined as a recommended identifier, and metadata information used to describe the recommended identifier may be used as recommended metadata information of the recommended viewpoint group (i.e., the target viewpoint group), so as to write the recommended metadata information into the encapsulated data box to obtain a second extended data box. At this time, the server can encapsulate the encoded video stream based on the second extended data box to obtain a video media file (i.e., the aforementioned second video media file) for transmission to the video client. In this way, upon receiving the video media file (e.g., the second video media file), the video client can directly display the video content corresponding to the recommended viewpoint group indicated by the recommendation identifier on the video client based on the recommendation identifier indicated by the recommendation metadata information in the second extended data box. Optionally, if there is no viewpoint group among these video groups that matches the designated viewpoint group associated with the photographer's shooting intention, the i-th viewpoint group can be used as the target viewpoint group among the G viewpoint groups that are unrelated to the content producer's shooting intention. Then, a mapping relationship between the target viewpoint group and the target control position information for viewing the target video content can be constructed based on the target video content corresponding to the target video group. Then, the timed metadata information corresponding to the target viewpoint group generated based on the mapping relationship can be written into the encapsulation data box corresponding to the volumetric video to obtain the aforementioned first extended data box. It should be understood that when the video client obtains the video media file encapsulated based on the first extended data box (i.e., the above-mentioned first video media file), it can unpack the first extended data box, and then compare the spatial position information of the user's current viewing (for example, the spatial area currently viewed by the user and the user's position, etc.) with the spatial position information indicated by the timed metadata information of each viewpoint group in the first extended data box, so as to intelligently select the viewpoint group that best matches the currently viewed spatial position information (for example, the best-matching viewpoint group can be the above-mentioned target viewpoint group), and decode to obtain the file track for carrying the target viewpoint content corresponding to the target viewpoint group, and then render and present the target viewpoint content corresponding to the target viewpoint group in the video client.

[0190] Further, see Figure 4 , Figure 4This is a flow chart of a method for processing volumetric video data provided by an embodiment of the present application. This method can be performed by a decoding device in a volumetric video system. The encoding device can be a user terminal integrated with the aforementioned video client. The method can include the following steps:

[0191] Step S301: Receive a video media file of a volumetric video sent by a server, decapsulate the video media file, and obtain a video encoding stream of the volumetric video and an extended data box corresponding to the video encoding stream.

[0192] The extended data box contains a recommended viewpoint group identification field; the server here can be the above Figure 2 or Figure 3 The server in the corresponding embodiment. It should be understood that the video media file received by the video client running in the user terminal can be a first video media file encapsulated according to the file encapsulation format indicated by the above-mentioned first extended data box, or a second video media file encapsulated according to the file encapsulation format indicated by the above-mentioned second extended data box. It should be understood that whether the video media file here is the first video media file or the second video media file is determined by the field value of the recommended viewpoint group identification field contained in the extended data box. For example, if the field value of the recommended viewpoint group identification field is an invalid value, the video media file received here is the first video media file, and the following steps S302-S305 can be further executed; for another example, if the field value of the recommended viewpoint group identification field is a valid value, the video media file received here is the second video media file, and the following steps S306-S307 can be further executed.

[0193] Step S302: If the value of the recommended view group identification field is an invalid value, then obtaining timed metadata information corresponding to each of the G view groups of the volumetric video in a first extended data box of the extended data box;

[0194] Specifically, if the field value of the recommended viewpoint group identification field is an invalid value (for example, the field value of the recommended viewpoint group identification field is 0), then in the first extended data box of the extended data box, the static viewpoint group metadata field associated with the G viewpoint groups is obtained; if the field value of the static viewpoint group metadata field is a numerical value that remains unchanged for describing the mapping relationship of each viewpoint group in the G viewpoint groups (for example, the field value of the static viewpoint group metadata field is 1), then based on the first field indication information associated with the recommended viewpoint group identification field, the timing metadata information corresponding to each viewpoint group in the G viewpoint groups is obtained to further execute the following step S303.

[0195] Among them, the static viewpoint group metadata field is deployed in the viewpoint group static metadata box of the first extended data box; if the field value of the static viewpoint group metadata field is a numerical value that remains unchanged for describing the mapping relationship of the viewpoint groups in the G viewpoint groups, then the static viewpoint group metadata recorded in the viewpoint group metadata sample entry of the first extended data box: includes the viewpoint group identifier of each viewpoint group in the G viewpoint groups.

[0196] Among them, if the field value of the static viewpoint group metadata field is a numerical value associated with the dynamic viewpoint group metadata, and the dynamic viewpoint group metadata is used to describe a viewpoint group among the G viewpoint groups whose mapping relationship changes over time, then the dynamic viewpoint group metadata recorded in the viewpoint group element sample corresponding to the viewpoint group metadata sample entry includes: an identifier of the variable viewpoint group whose mapping relationship changes at the sample timestamp, and before the sample timestamp, the mapping relationship between the variable viewpoint group and the video content corresponding to the variable viewpoint group remains unchanged. It should be understood that in this embodiment of the present application, the viewpoint groups among the G viewpoint groups whose mapping relationship changes over time can be collectively referred to as variable viewpoint groups.

[0197] The view group metadata sample entry and the view group element sample are used to form the view group timed metadata track of the volumetric video, and the view group timed metadata track is used to index one or more atlas data tracks associated with the volumetric video. Figure 2 Description of the first extended data box in the corresponding embodiment.

[0198] Step S303: obtaining the i-th viewpoint group among the G viewpoint groups, using the spatial position information of the video client as the spatial position information to be compared, and comparing the spatial position information to be compared with the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group to obtain a comparison result;

[0199] Wherein, i is a non-negative integer less than G;

[0200] For example, for each of the G viewpoint groups (for example, viewpoint group 1, viewpoint group 2, viewpoint group 3, viewpoint group 4 and viewpoint group 5 mentioned above), the spatial position information indicated by the timed metadata information corresponding to each viewpoint group can be compared with the spatial position information to be compared (for example, the spatial area currently viewed by the user, user position information, etc.) to obtain a sub-comparison result between the spatial position information to be compared and the spatial position information associated with each viewpoint group. These sub-comparison results can be collectively referred to as comparison results, and then when the spatial position information to be compared is the same as the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group among the G viewpoint groups, the following step S304 can be executed.

[0201] Step S304: If the comparison result indicates that the to-be-compared spatial position information is identical to the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group, the i-th viewpoint group is used as the matching viewpoint group, and the spatial position information indicated by the timed metadata information corresponding to the matching viewpoint group is used as the target spatial position information.

[0202] Among them, the target spatial position information is determined by the type field of the judgment information associated with the viewpoint group recorded by the server in the viewpoint group metadata sample entry of the first extended data box; the type field of the judgment information is deployed in the viewpoint group static metadata box of the viewpoint group metadata sample entry; if the type field of the judgment type is a first value, the target spatial position information having a mapping relationship with the matching viewpoint group is the three-dimensional spatial area information of the matching video content displayed on the video client; if the type field of the judgment type is a second value, the target spatial position information having a mapping relationship with the matching viewpoint group is the viewing position coordinate information of the user watching the matching video content on the video client; if the type field of the judgment type is a third value, the target spatial position information having a mapping relationship with the target viewpoint group is jointly determined by the three-dimensional spatial area information and the viewing position coordinate information.

[0203] Step S305 : Based on the mapping relationship between the matching viewpoint group and the target spatial position information, the matching video content corresponding to the matching viewpoint group is decoded from the video encoding stream of the volumetric video, and the matching video content corresponding to the matching viewpoint group is displayed on the video client.

[0204] Optionally, in step S306, if the field value of the recommended viewpoint group identification field is a valid value, then in a second extended data box of the extended data box, based on second field indication information associated with the recommended viewpoint group identification field, obtaining a recommendation identifier of the recommended viewpoint group indicated by the recommendation metadata information;

[0205] Step S307 : Decode the video encoding stream of the volumetric video to obtain the recommended video content corresponding to the recommended viewpoint group, and display the recommended video content corresponding to the recommended viewpoint group on the video client.

[0206] Thus, it can be seen that in the embodiment of the present application, different field indication information can be adaptively obtained by different metadata information provided in the extended data box. For example, if the obtained field indication information is the first field indication information, the timed metadata information corresponding to each viewpoint group written in the first extended data box can be used, and then, based on the timed metadata information corresponding to each viewpoint group and the current spatial position information of the video client obtained in real time, the spatial position information that matches the current spatial position information of the video client can be intelligently determined, and the determined spatial position information that matches the current spatial position information of the video client (for example, the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group) can be used as a basis for selecting a specific viewpoint group. The selected specific viewpoint group is the matching viewpoint group. At this time, the video client can decode and present the video content corresponding to the matching viewpoint group, and then partial decoding can be achieved during the process of partial access to the volumetric video. In this way, the overall decoding of all video content of the volumetric video can be avoided from the root, thereby saving computing resources. Optionally, if the acquired field indication information is the second field indication information, the recommended metadata information corresponding to the recommended viewpoint group written in the second extended data box can be directly decoded to obtain the video content corresponding to the recommended viewpoint group according to the recommended viewpoint group indicated by the recommended metadata information, so as to further improve the decoding and presentation efficiency of the video client.

[0207] Further, see Figure 5 , Figure 5 This is a flow chart of a method for processing volumetric video data provided by an embodiment of the present application. This method can be performed by a decoding device in a volumetric video system. The encoding device can be a user terminal integrated with the aforementioned video client. The method can include the following steps:

[0208] Step S401: Receive a video media file of a volumetric video sent by a server, decapsulate the video media file, and obtain a video encoding stream of the volumetric video and an extended data box corresponding to the video encoding stream; the extended data box includes a recommended viewpoint group identification field;

[0209] Step S402: If the value of the recommended view group identification field is an invalid value, then obtaining timed metadata information corresponding to each of the G view groups of the volumetric video in a first extended data box of the extended data box;

[0210] Step S403: obtaining the i-th viewpoint group from the G viewpoint groups, using the spatial position information of the video client as the spatial position information to be compared, and comparing the spatial position information to be compared with the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group to obtain a comparison result;

[0211] Wherein, i is a non-negative integer less than G;

[0212] Step S404: If the comparison result indicates that the to-be-compared spatial position information is identical to the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group, the i-th viewpoint group is used as the matching viewpoint group, and the spatial position information indicated by the timed metadata information corresponding to the matching viewpoint group is used as the target spatial position information.

[0213] Step S405 : Based on the mapping relationship between the matching viewpoint group and the target spatial position information, the matching video content corresponding to the matching viewpoint group is decoded from the video encoding stream of the volumetric video, and the matching video content corresponding to the matching viewpoint group is displayed on the video client.

[0214] Step S406: When the spatial position information of the video client is changed from the target spatial position information to the spatial position update information, a mapping relationship between the spatial position update information and the changed viewpoint group is used as an update relationship based on the timed metadata information corresponding to the spatial position update information in the first extended data box;

[0215] Step S407 : Based on the update relationship, the video content corresponding to the changed viewpoint group is obtained by decoding the video encoding stream of the volumetric video, and the video content corresponding to the changed viewpoint group is displayed on the video client.

[0216] It should be understood that the first extended data box includes a viewpoint position structure for describing user position information. The viewpoint position structure includes a view_position structure for indicating the user's viewing position, a view_orientation structure for indicating the user's viewing direction, and a position_range_flag field for indicating whether the user's range of motion information is included. The view_position structure may include the user's specific x, y, and z coordinate information within the multi-view media space. The view_orientation structure may include the user's head rotation information. The semantics of this head rotation information can be found in the description of the rotation structure for indicating user rotation information shown in Table 5. If the position_range_flag field has a value of 1 (i.e., a valid value), it indicates that the viewpoint position structure deployed in the first extended data box includes user range of motion information. Conversely, if the position_range_flag field has a value of 0 (i.e., an invalid value), it indicates that the viewpoint position structure deployed in the first extended data box does not include user range of motion information.

[0217] Based on this, when a user watches the video content corresponding to viewpoint group 1 at the target spatial position information (for example, spatial position information 1, and the spatial position information indicates that the rotation information of the user at viewing position P1 is rotation information W1), once the user's head rotates, the rotation information will be changed, which will in turn cause a change in the spatial position information. At this time, when the spatial position information of the video client used by the user is changed from the target spatial position information to another spatial position information (for example, spatial position information 2, and the spatial position information indicates that the rotation information of the user at viewing position P2 is rotation information W2), the spatial position information 2 is collectively referred to as spatial position update information, and then based on the timed metadata information corresponding to the spatial position update information in the first extended data box, the mapping relationship between the spatial position update information and the changed viewpoint group (for example, viewpoint group 2) can be used as an update relationship, so that the video content corresponding to the changed viewpoint group (i.e., viewpoint group 2) can be decoded from the video encoding stream of the volumetric video based on the update relationship to display the video content corresponding to the changed viewpoint group on the video client. It can be seen from this that the embodiment of the present application can adaptively switch viewpoint groups according to the user's viewing needs based on the mapping relationship associated with each viewpoint group provided by the first extended data box, so as to improve the user's viewing experience while improving the decoding presentation efficiency during the partial access process of the volumetric video.

[0218] Further, see Figure 6 , Figure 6 : This is a schematic diagram of the structure of a volumetric video data processing device provided in an embodiment of the present application. The volumetric video data processing device can be a computer program (including program code) running in an encoding device. For example, the volumetric video data processing device can be an application software in the encoding device. The volumetric video data processing device can be used to perform Figure 2 or Figure 3 The steps of the method for processing volumetric video data in the corresponding embodiment. For further details, please refer to Figure 6 The volumetric video data processing device 1 may include: a viewpoint group acquisition module 11, a mapping relationship construction module 12, a timed metadata writing module 13, and a media file sending module 14. Furthermore, the volumetric video data processing device 1 may also include:

[0219] A viewpoint group acquisition module 11 is configured to acquire G viewpoint groups of the volumetric video and use the i-th viewpoint group among the G viewpoint groups as a target viewpoint group;

[0220] Wherein, i is a non-negative integer less than G;

[0221] The viewpoint group acquisition module 11 includes: a viewpoint acquisition unit 111, a viewpoint group search unit 112, a notification unit 113, a recommended viewpoint group determination unit 114 and a target viewpoint group determination unit 115;

[0222] A viewpoint acquisition unit 111 is configured to acquire V viewpoints of the volumetric video and group the V viewpoints based on viewpoint dependencies between the V viewpoints to obtain G viewpoint groups of the volumetric video. V represents the number of viewpoints in the volumetric video and is a positive integer greater than or equal to 2. The viewpoint dependencies are determined by the content relevance between the video content corresponding to each of the V viewpoints.

[0223] The viewpoint group search unit 112 is configured to obtain video association information associated with the volumetric video and search the video association information for a designated viewpoint group associated with the content producer; the designated viewpoint group is related to the shooting intention of the content producer who shot the volumetric video;

[0224] The notification unit 113 is used to obtain the i-th viewpoint group that is irrelevant to the content producer's shooting intention from the G viewpoint groups as the target viewpoint group if the designated viewpoint group associated with the content producer is not found in the video-related information, and to notify the mapping relationship construction module 12 to execute the steps of constructing a mapping relationship between the target viewpoint group and the target spatial position information for viewing the target video content based on the target viewpoint group, and generating timing metadata information corresponding to the target viewpoint group based on the mapping relationship.

[0225] Optionally, the recommended viewpoint group determining unit 114 is configured to use the found designated viewpoint group as the recommended viewpoint group if a designated viewpoint group associated with the content producer is found in the video associated information;

[0226] The target viewpoint group determining unit 115 is configured to obtain, based on the recommended viewpoint group, an i-th viewpoint group related to the shooting intention of the content producer from the G viewpoint groups as the target viewpoint group.

[0227] It should be understood that after the volumetric video data processing device 1 performs the above steps through the notification unit 113, it can also perform the following steps through the invalid field setting module 14. The specific implementation of the viewpoint acquisition unit 111, the viewpoint group search unit 112, the notification unit 113, the recommended viewpoint group determination unit 114 and the target viewpoint group determination unit 115 can be found in the above Figure 2 The description of the specific process of obtaining the target viewpoint group in the corresponding embodiment will not be repeated here.

[0228] An invalid field setting module 14 is used to add a recommended viewpoint group identification field in the viewpoint group metadata sample of the encapsulated data box, and set the field value of the recommended viewpoint group identification field to an invalid field value, and in the viewpoint group metadata sample of the first extended data box, use the recommended viewpoint group identification field with the invalid field value as the first field indication information; the first field indication information is used to instruct the video client to obtain the timed metadata information of each viewpoint group in the G viewpoint groups from the first extended data box.

[0229] A mapping relationship building module 12 is used to build a mapping relationship between the target viewpoint group and the target spatial position information for viewing the target video content based on the target video content corresponding to the target viewpoint group, and generate timed metadata information corresponding to the target viewpoint group based on the mapping relationship;

[0230] The timed metadata writing module 13 is configured to write the timed metadata information corresponding to the target viewpoint group into the encapsulated data box corresponding to the volumetric video, thereby obtaining a first extended data box corresponding to the encapsulated data box; the first extended data box contains the timed metadata information corresponding to each of the G viewpoint groups.

[0231] The media file delivery module 14 is configured to obtain the encoded video streams associated with the G viewpoint groups, and encapsulate the encoded video streams based on the first extended data box to obtain a video media file of the volumetric video;

[0232] Furthermore, the media file sending module 14 is also used to send the video media file to the video client, so that when the video client obtains the first extended data box based on the video media file, it displays the target video content corresponding to the target viewpoint group on the video client according to the target spatial position information indicated by the timed metadata information corresponding to the target viewpoint group in the first extended data box.

[0233] Optionally, it should be understood that after the volumetric video data processing device 1 performs the above steps through the notification unit 115 , it may also perform the following steps through the recommendation metadata writing module 16 .

[0234] a recommendation metadata writing module 16 configured to determine the identifier of the target view group as the recommendation identifier, and to use metadata information describing the recommendation identifier as the recommendation metadata information of the target view group, and to write the recommended metadata information into an encapsulated data box corresponding to the volumetric video, thereby obtaining a second extended data box corresponding to the encapsulated data box;

[0235] The encapsulation processing module 17 is configured to obtain coded video streams associated with the G viewpoint groups, encapsulate the coded video streams based on the second extended data box to obtain a video media file of the volumetric video, and deliver the video media file to a video client. When the video client obtains the second extended data box based on the video media file, the video client displays the target video content corresponding to the target viewpoint group indicated by the recommendation identifier based on the recommendation metadata information in the second extended data box.

[0236] Optionally, before the recommended metadata writing module 16 writes the recommended metadata information into the encapsulated data box corresponding to the volumetric video, the data processing device 1 of the volumetric video performs the following steps through the valid field setting module 18:

[0237] The valid field setting module 18 is used to add a recommended viewpoint group identification field associated with the recommendation identifier in the viewpoint group metadata sample of the encapsulated data box, and set the field value of the recommended viewpoint group identification field to a valid field value, and in the viewpoint group metadata sample of the second extended data box, the recommended viewpoint group identification field with a valid field value is used as the second field indication information; the second field indication information is used to instruct the video client to obtain the recommended metadata information from the second extended data box.

[0238] The specific implementation of the viewpoint group acquisition module 11, the mapping relationship construction module 12, the timed metadata writing module 13 and the media file sending module 14 can be found in the above Figure 2 The description of steps S101 to S105 in the corresponding embodiment will not be repeated here. Further, the specific implementation of the invalid field setting module 15, the recommended metadata writing module 16, the encapsulation processing module 17 and the valid field setting module 18 can be found in the above Figure 3 The description of steps S201 to S211 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0239] For further information, see Figure 7 , Figure 7 : This is a schematic diagram of the structure of a volumetric video data processing device provided in an embodiment of the present application. The volumetric video data processing device can be a computer program (including program code) running in a decoding device. For example, the volumetric video data processing device can be an application software in the decoding device. The volumetric video data processing device can be used to perform Figure 4 or Figure 5 The steps of the method for processing volumetric video data in the corresponding embodiment. For further details, please refer to Figure 7The volumetric video data processing device 2 may include: a media file receiving module 21, a timed metadata acquisition module 22, an information comparison module 23, a target information determination module 24, and a video decoding module 25. Optionally, the volumetric video data processing device 2 may further include: an information change module 26, an update decoding module 27, a recommendation identifier acquisition module 28, and a recommended content display module 29.

[0240] The media file receiving module 21 is configured to receive a video media file of a volumetric video sent by a server, decapsulate the video media file, and obtain a video encoding stream of the volumetric video and an extended data box corresponding to the video encoding stream; the extended data box includes a recommended viewpoint group identification field;

[0241] a timed metadata acquisition module 22 configured to acquire, in a first extended data box of the extended data box, timed metadata information corresponding to each of the G view groups of the volumetric video if the field value of the recommended view group identification field is an invalid value;

[0242] The timed metadata acquisition module 22 includes: a field acquisition unit 221 and a timed metadata acquisition unit 222;

[0243] The field acquisition unit 221 is configured to acquire, if the field value of the recommended view group identification field is an invalid value, static view group metadata fields associated with the G view groups in a first extended data box of the extended data box;

[0244] The timed metadata acquisition unit 222 is used to obtain the timed metadata information corresponding to each viewpoint group in the G viewpoint groups based on the first field indication information associated with the recommended viewpoint group identification field if the field value of the static viewpoint group metadata field is a numerical value that remains unchanged for describing the mapping relationship of each viewpoint group in the G viewpoint groups.

[0245] The specific implementation of the field acquisition unit 221 and the timing metadata acquisition unit 222 can be found in the above Figure 4 The description of the specific process of obtaining the timing metadata information through the first extended data box in the corresponding embodiment will not be repeated here.

[0246] Among them, the static viewpoint group metadata field is deployed in the viewpoint group static metadata box of the first extended data box; if the field value of the static viewpoint group metadata field is a numerical value that remains unchanged for describing the mapping relationship of the viewpoint groups in the G viewpoint groups, then the static viewpoint group metadata recorded in the viewpoint group metadata sample entry of the first extended data box: includes the viewpoint group identifier of each viewpoint group in the G viewpoint groups.

[0247] Among them, if the field value of the static viewpoint group metadata field is a numerical value associated with the dynamic viewpoint group metadata, and the dynamic viewpoint group metadata is used to describe a viewpoint group among the G viewpoint groups whose mapping relationship changes over time, then the dynamic viewpoint group metadata recorded in the viewpoint group element sample corresponding to the viewpoint group metadata sample entry includes: an identifier of a variable viewpoint group whose mapping relationship changes at the sample timestamp, and before the sample timestamp, the mapping relationship between the variable viewpoint group and the video content corresponding to the variable viewpoint group remains unchanged; the variable viewpoint group is a viewpoint group among the G viewpoint groups whose mapping relationship changes over time.

[0248] The view group metadata sample entry and the view group element sample are used to constitute a view group timed metadata track of the volumetric video, and the view group timed metadata track is used to index one or more atlas data tracks associated with the volumetric video.

[0249] The information comparison module 23 is configured to obtain an i-th viewpoint group from the G viewpoint groups, use the spatial position information of the video client as the spatial position information to be compared, and compare the spatial position information to be compared with the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group to obtain a comparison result; i is a non-negative integer less than G;

[0250] a target information determining module 24 configured to, if the comparison result indicates that the to-be-compared spatial position information is identical to the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group, determine the i-th viewpoint group as a matching viewpoint group, and determine the spatial position information indicated by the timed metadata information corresponding to the matching viewpoint group as target spatial position information;

[0251] Among them, the target spatial position information is determined by the type field of the judgment information associated with the viewpoint group recorded by the server in the viewpoint group metadata sample entry of the first extended data box; the type field of the judgment information is deployed in the viewpoint group static metadata box of the viewpoint group metadata sample entry; if the type field of the judgment type is a first value, the target spatial position information having a mapping relationship with the matching viewpoint group is the three-dimensional spatial area information of the matching video content displayed on the video client; if the type field of the judgment type is a second value, the target spatial position information having a mapping relationship with the matching viewpoint group is the viewing position coordinate information of the user watching the matching video content on the video client; if the type field of the judgment type is a third value, the target spatial position information having a mapping relationship with the target viewpoint group is jointly determined by the three-dimensional spatial area information and the viewing position coordinate information.

[0252] The video decoding module 25 is used to decode the matching video content corresponding to the matching viewpoint group from the video encoding stream of the volumetric video based on the mapping relationship between the matching viewpoint group and the target spatial position information, and display the matching video content corresponding to the matching viewpoint group on the video client.

[0253] Optionally, an information changing module 26 is configured to, when the spatial position information of the video client is changed from the target spatial position information to the spatial position update information, use, based on the timed metadata information corresponding to the spatial position update information in the first extended data box, a mapping relationship between the spatial position update information and the changed viewpoint group as an update relationship;

[0254] The update decoding module 27 is configured to decode the video encoding stream of the volumetric video based on the update relationship to obtain the video content corresponding to the changed viewpoint group, and display the video content corresponding to the changed viewpoint group on the video client.

[0255] Optionally, the recommendation identifier obtaining module 28 is configured to obtain, in a second extended data box of the extended data box, a recommendation identifier of the recommended view group indicated by the recommendation metadata information based on second field indication information associated with the recommended view group identification field if the field value of the recommended view group identification field is a valid value;

[0256] The recommended content display module 29 is configured to decode the video encoding stream of the volumetric video to obtain the recommended video content corresponding to the recommended viewpoint group, and display the recommended video content corresponding to the recommended viewpoint group on the video client.

[0257] The specific implementation of the media file receiving module 21, the timed metadata obtaining module 22, the information comparing module 23, the target information determining module 24 and the video decoding module 25 can be found in the above Figure 4 The description of steps S301 to S305 in the corresponding embodiment will not be repeated here. Figure 4 The description of steps S306-S307 in the corresponding embodiment will not be repeated here; the specific implementation of the recommendation identifier acquisition module 28 and the recommended content display module 29 can be found in the above Figure 5 The description of steps S401 to S407 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0258] Further, see Figure 8 , Figure 8 This is a schematic diagram of a computer device provided in an embodiment of the present application. Figure 8The computer device 1000 shown may include: at least one processor 1001, such as a CPU, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002. The communication bus 1002 is used to implement connection and communication between these components. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 1005 may optionally also be at least one storage device located away from the aforementioned processor 1001. Figure 8 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a device control application program.

[0259] exist Figure 8 In the computer device 1000 shown, the network interface 1004 is mainly used to provide network communication functions; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to execute the above Figure 2 、 Figure 3 、 Figure 4 or Figure 5 The description of the data processing method of the volumetric video in the corresponding embodiment can also be performed as described above. Figure 6 The description of the volumetric video data processing device 1 in the corresponding embodiment can also execute the aforementioned Figure 7 The description of the volumetric video data processing 2 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here either.

[0260] In addition, it should be noted that the embodiment of the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the aforementioned computer device 1000, and the computer program includes program instructions. When the processor executes the program instructions, the aforementioned computer program can be executed. Figure 2 、 Figure 3 、 Figure 4 or Figure 5 The description of the volumetric video data processing method described in the corresponding embodiment will not be repeated here. Furthermore, the beneficial effects of the same method will not be repeated here. For technical details not disclosed in the computer-readable storage medium embodiments of this application, please refer to the description of the method embodiments of this application.

[0261] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0262] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A method for processing volumetric video data, characterized in that: The method is executed by a server and includes: Obtaining G viewpoint groups of a volumetric video, obtaining video association information associated with the volumetric video, and searching the video association information for a designated viewpoint group associated with a content producer; the designated viewpoint group is related to a shooting intention of the content producer who shot the volumetric video; If the designated viewpoint group associated with the content producer is not found in the video-related information, then obtaining an i-th viewpoint group that is irrelevant to the filming intention of the content producer from the G viewpoint groups as the target viewpoint group; wherein i is a non-negative integer less than G; Based on the target video content corresponding to the target viewpoint group, a mapping relationship between the target viewpoint group and target spatial position information for viewing the target video content is constructed, and timed metadata information corresponding to the target viewpoint group is generated based on the mapping relationship; Writing the timed metadata information corresponding to the target viewpoint group into the encapsulated data box corresponding to the volumetric video to obtain a first extended data box corresponding to the encapsulated data box; the first extended data box contains the timed metadata information corresponding to each of the G viewpoint groups; Obtaining coded video streams associated with the G viewpoint groups, and encapsulating the coded video streams based on the first extended data box to obtain a video media file of the volumetric video; The video media file is sent to a video client so that when the video client obtains the first extended data box based on the video media file, the video client matches the current spatial position information of the video client with the target spatial position information indicated by the timing metadata information corresponding to the target viewpoint group in the first extended data box, and when the current spatial position information of the video client matches the target spatial position information, the target video content corresponding to the target viewpoint group is displayed on the video client.

2. The method according to claim 1, characterized in that The obtaining of G viewpoint groups of the volumetric video includes: V viewpoints of a volumetric video are obtained, and based on viewpoint dependencies between the V viewpoints, the V viewpoints are grouped to obtain G viewpoint groups of the volumetric video; wherein V is used to represent the number of viewpoints of the volumetric video, and V is a positive integer greater than or equal to 2; and the viewpoint dependencies are determined by content relevance between video content corresponding to each of the V viewpoints.

3. The method according to claim 1, characterized in that The method further comprises: A recommended viewpoint group identification field is added to the viewpoint group metadata sample of the encapsulated data box, and the field value of the recommended viewpoint group identification field is set to an invalid field value. In the viewpoint group metadata sample of the first extended data box, the recommended viewpoint group identification field with the invalid field value is used as the first field indication information; the first field indication information is used to instruct the video client to obtain the timing metadata information of each viewpoint group in the G viewpoint groups from the first extended data box.

4. The method according to claim 1, wherein The method further comprises: If a designated viewpoint group associated with the content producer is found in the video-related information, the found designated viewpoint group is used as a recommended viewpoint group; Based on the recommended viewpoint group, an i-th viewpoint group related to the shooting intention of the content producer is obtained from the G viewpoint groups as the target viewpoint group.

5. The method according to claim 4, characterized in that The method further comprises: Determining the identifier of the target view group as a recommended identifier, using metadata information describing the recommended identifier as recommended metadata information of the target view group, and writing the recommended metadata information into an encapsulated data box corresponding to the volumetric video to obtain a second extended data box corresponding to the encapsulated data box; Obtain encoded video streams associated with the G viewpoint groups, encapsulate the encoded video streams based on the second extended data box to obtain a video media file of the volumetric video, and send the video media file to a video client, so that when the video client obtains the second extended data box based on the video media file, the video client displays the target video content corresponding to the target viewpoint group indicated by the recommendation identifier based on the recommendation metadata information in the second extended data box.

6. The method according to claim 5, characterized in that When writing the recommended metadata information into the encapsulated data box corresponding to the volumetric video, the method further includes: A recommended viewpoint group identification field associated with the recommendation identifier is added to the viewpoint group metadata sample of the encapsulated data box, and the field value of the recommended viewpoint group identification field is set to a valid field value. In the viewpoint group metadata sample of the second extended data box, the recommended viewpoint group identification field with the valid field value is used as second field indication information; the second field indication information is used to instruct the video client to obtain the recommended metadata information from the second extended data box.

7. A method for processing volumetric video data, characterized in that: The method is executed by a video client and includes: receiving a video media file of a volumetric video sent by a server, decapsulating the video media file to obtain a video encoding stream of the volumetric video and an extended data box corresponding to the video encoding stream; wherein the extended data box includes a recommended viewpoint group identification field; If the field value of the recommended view group identification field is an invalid value, obtaining, in a first extended data box of the extended data box, timed metadata information corresponding to each of the G view groups of the volumetric video; Obtaining an i-th viewpoint group from the G viewpoint groups, using spatial position information of the video client as spatial position information to be compared, and comparing the spatial position information to be compared with spatial position information indicated by timed metadata information corresponding to the i-th viewpoint group to obtain a comparison result; wherein i is a non-negative integer less than G; If the comparison result indicates that the to-be-compared spatial position information is identical to the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group, the i-th viewpoint group is used as the matching viewpoint group, and the spatial position information indicated by the timed metadata information corresponding to the matching viewpoint group is used as the target spatial position information; Based on the mapping relationship between the matching viewpoint group and the target spatial position information, matching video content corresponding to the matching viewpoint group is obtained by decoding the video encoding stream of the volumetric video, and the matching video content corresponding to the matching viewpoint group is displayed on the video client.

8. The method according to claim 7, characterized in that If the field value of the recommended view group identification field is an invalid value, obtaining, in a first extended data box of the extended data box, timed metadata information corresponding to each of the G view groups of the volumetric video, including: If the field value of the recommended view group identification field is an invalid value, obtaining, in a first extended data box of the extended data box, static view group metadata fields associated with the G view groups; If the field value of the static viewpoint group metadata field is a numerical value used to describe that the mapping relationship of each viewpoint group in the G viewpoint groups remains unchanged, then based on the first field indication information associated with the recommended viewpoint group identification field, the timing metadata information corresponding to each viewpoint group in the G viewpoint groups is obtained.

9. The method according to claim 8, characterized in that The static viewpoint group metadata field is deployed in the viewpoint group static metadata box of the first extended data box; if the field value of the static viewpoint group metadata field is a numerical value that remains unchanged for describing the mapping relationship of the viewpoint groups in the G viewpoint groups, then the static viewpoint group metadata recorded in the viewpoint group metadata sample entry of the first extended data box: includes the viewpoint group identifier of each viewpoint group in the G viewpoint groups.

10. The method according to claim 9, characterized in that If the field value of the static viewpoint group metadata field is a numerical value associated with the dynamic viewpoint group metadata, and the dynamic viewpoint group metadata is used to describe a viewpoint group among the G viewpoint groups whose mapping relationship changes over time, then the dynamic viewpoint group metadata recorded in the viewpoint group element sample corresponding to the viewpoint group metadata sample entry includes: an identifier of a variable viewpoint group whose mapping relationship changes at the sample timestamp, and before the sample timestamp, the mapping relationship between the variable viewpoint group and the video content corresponding to the variable viewpoint group remains unchanged; the variable viewpoint group is a viewpoint group among the G viewpoint groups whose mapping relationship changes over time.

11. The method according to claim 10, characterized in that The view group metadata sample entry and the view group element sample are used to constitute a view group timed metadata track of the volumetric video, and the view group timed metadata track is used to index one or more atlas data tracks associated with the volumetric video.

12. The method according to any one of claims 7 to 11, characterized in that: The target spatial position information is determined by a type field of determination information associated with a view group recorded by the server in a view group metadata sample entry of the first extended data box; the type field of the determination information is disposed in a view group static metadata box of the view group metadata sample entry; If the type field of the determination information is a first value, the target spatial position information having the mapping relationship with the matching viewpoint group is the three-dimensional spatial region information of the matching video content displayed on the video client; If the type field of the determination information is a second value, the target spatial position information having the mapping relationship with the matching viewpoint group is viewing position coordinate information of a user watching the matching video content on the video client; If the type field of the determination information is a third value, the target space position information having the mapping relationship with the matching viewpoint group is determined jointly by the three-dimensional space region information and the viewing position coordinate information.

13. The method according to any one of claims 7 to 11, characterized in that The method further comprises: When the spatial position information of the video client is changed from the target spatial position information to the spatial position update information, based on the timed metadata information corresponding to the spatial position update information in the first extended data box, a mapping relationship between the spatial position update information and the changed viewpoint group is used as an update relationship; Based on the update relationship, video content corresponding to the changed viewpoint group is obtained by decoding the video encoding stream of the volumetric video, and the video content corresponding to the changed viewpoint group is displayed on the video client.

14. The method according to any one of claims 7 to 11, characterized in that The method further comprises: If the field value of the recommended view group identification field is a valid value, then obtaining, in a second extended data box of the extended data box, a recommendation identifier of the recommended view group indicated by the recommendation metadata information based on second field indication information associated with the recommended view group identification field; The recommended video content corresponding to the recommended viewpoint group is obtained by decoding the video encoding stream of the volumetric video, and the recommended video content corresponding to the recommended viewpoint group is displayed on the video client.

15. A volumetric video data processing device, characterized in that: include: a viewpoint group acquisition module configured to acquire G viewpoint groups of a volumetric video, acquire video association information associated with the volumetric video, and search the video association information for a designated viewpoint group associated with a content producer; the designated viewpoint group being related to the filming intent of the content producer who filmed the volumetric video; The viewpoint group acquisition module is further configured to, if no designated viewpoint group associated with the content producer is found in the video-related information, acquire an i-th viewpoint group that is unrelated to the content producer's shooting intention from the G viewpoint groups as a target viewpoint group; wherein i is a non-negative integer less than G; a mapping relationship construction module, configured to construct, based on the target video content corresponding to the target viewpoint group, a mapping relationship between the target viewpoint group and target spatial position information for viewing the target video content, and generate timed metadata information corresponding to the target viewpoint group based on the mapping relationship; a timed metadata writing module, configured to write the timed metadata information corresponding to the target viewpoint group into an encapsulated data box corresponding to the volumetric video, thereby obtaining a first extended data box corresponding to the encapsulated data box; the first extended data box contains the timed metadata information corresponding to each of the G viewpoint groups; a media file delivery module, configured to obtain coded video streams associated with the G viewpoint groups, and encapsulate the coded video streams based on the first extended data box to obtain a video media file of the volumetric video; The media file sending module is also used to send the video media file to the video client, so that when the video client obtains the first extended data box based on the video media file, the video client matches the current spatial position information of the video client with the target spatial position information indicated by the timing metadata information corresponding to the target viewpoint group in the first extended data box, and when the current spatial position information of the video client matches the target spatial position information, the target video content corresponding to the target viewpoint group is displayed on the video client.

16. A volumetric video data processing device, characterized in that: include: a media file receiving module configured to receive a video media file of a volumetric video sent by a server, decapsulate the video media file, and obtain a video encoding stream of the volumetric video and an extended data box corresponding to the video encoding stream; the extended data box includes a recommended viewpoint group identification field; a timed metadata acquisition module, configured to, if the field value of the recommended view group identification field is an invalid value, acquire, from a first extended data box of the extended data box, timed metadata information corresponding to each of the G view groups of the volumetric video; an information comparison module, configured to obtain an i-th viewpoint group from the G viewpoint groups, use spatial position information of the video client as spatial position information to be compared, and compare the spatial position information to be compared with spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group to obtain a comparison result; The i is a non-negative integer less than the G; a target information determining module, configured to, if the comparison result indicates that the to-be-compared spatial position information is the same as the spatial position information indicated by the timed metadata information corresponding to the i-th viewpoint group, determine the i-th viewpoint group as a matching viewpoint group, and determine the spatial position information indicated by the timed metadata information corresponding to the matching viewpoint group as target spatial position information; a video decoding module for decoding, based on a mapping relationship between the matching viewpoint group and the target spatial position information, a video encoding stream of the volumetric video to obtain matching video content corresponding to the matching viewpoint group, and displaying the matching video content corresponding to the matching viewpoint group on the video client.

17. A computer device, characterized in that: include: Processor, memory, network interface; The processor is connected to a memory and a network interface, wherein the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program to execute the method according to any one of claims 1 to 14.

18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the method according to any one of claims 1 to 14 is executed.

Citation Information

Patent Citations

  • Media processing method, device and system and readable storage medium

    CN112148115A

  • Encoding and decoding virtual reality video

    US20180035134A1