File packaging method, device, equipment and storage medium for free-viewpoint video
By adding codec independence indication information to the video track, the problem of inefficient decoding efficiency in media files in single-track packaging mode cannot be determined in the prior art, and flexible processing and efficient decoding of media files are realized.
Patent Information
- Application Number
- CN202110913912.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-10
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2041-08-10
AI Technical Summary
In the prior art, the packaging method of video code stream cannot determine whether the media file in single-track packaging mode can decode part of the viewing angle, resulting in low decoding efficiency of media file.
Add codec independence indication information to the video track, indicating whether the video data of a single viewing angle in the video track depends on video data of other viewing angles when encoding and decoding, allowing the client or server to partially decode or repackage based on this information.
Improves the processing flexibility and decoding efficiency of media files, allowing clients to partially decode the texture and depth maps of specific cameras, and the server can decide whether to re-package monorailed videos.
Smart Images

Figure CN115914672B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of video processing technology, and in particular to a file encapsulation method, apparatus, device, and storage medium for free-viewpoint video. Background Art
[0002] Immersive media refers to media content that can provide consumers with an immersive experience. Immersive media can be divided into 3DoF media, 3DoF+ media, and 6DoF media according to the user's degree of freedom when consuming media content.
[0003] However, with the current video stream encapsulation method, the server or client cannot determine whether the media file of the free-viewpoint video encapsulated in the single-track encapsulation mode can decode the media files corresponding to some perspectives, resulting in low decoding efficiency of the media files. Summary of the Invention
[0004] The present application provides a file encapsulation method, apparatus, device and storage medium for free-viewpoint video, by which a server or client can determine whether a media file corresponding to a partial viewpoint in a media file can be decoded, thereby improving the decoding efficiency of the media file.
[0005] In a first aspect, the present application provides a file encapsulation method for a free-viewpoint video, which is applied to a first device, which can be understood as a video encapsulation device. The method includes:
[0006] Obtaining a code stream of free-viewpoint video data, where the free-viewpoint video data includes video data of N viewpoints, where N is a positive integer;
[0007] Encapsulating the code stream of the free-viewpoint video data into at least one video track to obtain a media file of the free-viewpoint video data, wherein the video track includes codec independence indication information and video code streams of M views, wherein the codec independence indication information is used to indicate whether video data of a single view among the M views corresponding to the video track depends on video data of other views during encoding and decoding, where M is a positive integer less than or equal to N;
[0008] The media file of the free-viewpoint video data is sent to a client or a server.
[0009] In a second aspect, the present application provides a file encapsulation method for a free-viewpoint video, which is applied to a client, which can be understood as a video playback device. The method includes:
[0010] Receive a media file of free-viewpoint video data sent by a first device, where the media file includes at least one video track, the free-viewpoint video data includes video data of N perspectives, where N is a positive integer, the video track includes codec independence indication information and video streams of the M perspectives, the codec independence indication information is used to indicate whether video data of a single perspective among the M perspectives corresponding to the video track depends on video data of other perspectives during encoding and decoding, where M is a positive integer less than or equal to N;
[0011] Decapsulating the media file according to the codec independence indication information to obtain a video stream corresponding to at least one perspective;
[0012] The video code stream corresponding to the at least one viewing angle is decoded to obtain reconstructed video data of the at least one viewing angle.
[0013] In a third aspect, the present application provides a file encapsulation method for a free-viewpoint video, applied to a server, the method comprising:
[0014] Receive a media file of free-viewpoint video data sent by a first device, where the media file includes at least one video track, the free-viewpoint video data includes video data of N perspectives, where N is a positive integer, the video track includes codec independence indication information and video streams of the M perspectives, the codec independence indication information is used to indicate whether video data of a single perspective among the M perspectives corresponding to the video track depends on video data of other perspectives during encoding and decoding, where M is a positive integer less than or equal to N;
[0015] Determine whether to decompose the at least one video track into multiple video tracks according to the codec independence indication information.
[0016] In a fourth aspect, the present application provides a multi-view video data processing apparatus, applied to a first device, the apparatus comprising:
[0017] an acquiring unit, configured to acquire a code stream of free-viewpoint video data, wherein the free-viewpoint video data includes video data of N viewpoints, where N is a positive integer;
[0018] an encapsulation unit, configured to encapsulate the code stream of the free-viewpoint video data into at least one video track to obtain a media file of the free-viewpoint video data, wherein the video track includes codec independence indication information and video code streams of M views, wherein the codec independence indication information is used to indicate whether video data of a single view among the M views corresponding to the video track depends on video data of other views during encoding and decoding, where M is a positive integer less than or equal to N;
[0019] The sending unit is configured to send the media file of the free-viewpoint video data to a client or a server.
[0020] In a fifth aspect, the present application provides a multi-view video data processing device, which is applied to a client, and the device includes:
[0021] a receiving unit, configured to receive a media file of free-viewpoint video data sent by a first device, the media file comprising at least one video track, the free-viewpoint video data comprising video data of N perspectives, where N is a positive integer, the video track comprising codec independence indication information and video streams of the M perspectives, the codec independence indication information being used to indicate whether video data of a single perspective among the M perspectives corresponding to the video track depends on video data of other perspectives during encoding and decoding, where M is a positive integer less than or equal to N;
[0022] a decapsulation unit, configured to decapsulate the media file according to the codec independence indication information to obtain a video stream corresponding to at least one viewing angle;
[0023] A decoding unit is used to decode the video code stream corresponding to the at least one perspective to obtain reconstructed video data of the at least one perspective.
[0024] In a sixth aspect, a computing device is provided, comprising: a processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method of the first aspect and / or the second aspect and / or the third aspect.
[0025] In a seventh aspect, a computer-readable storage medium is provided for storing a computer program, wherein the computer program enables a computer to execute the method of the first aspect and / or the second aspect and / or the third aspect.
[0026] In summary, in this application, by adding codec independence indication information to the video track, the codec independence indication information is used to indicate whether the video data of a single perspective among the M perspectives corresponding to the video track depends on the video data of other perspectives during encoding and decoding. In this way, in the single-track encapsulation mode, the client can determine whether the texture map and depth map of a specific camera can be partially decoded based on the codec independence indication information. In addition, in the single-track encapsulation mode, the server can also determine whether the free-viewpoint video encapsulated in the single track can be re-encapsulated as multiple tracks based on the codec independence indication information, thereby improving the processing flexibility of the media file and improving the decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0028] Figure 1 A schematic diagram of three degrees of freedom is shown schematically;
[0029] Figure 2 A schematic diagram of three degrees of freedom + is shown schematically;
[0030] Figure 3 A schematic diagram of six degrees of freedom is shown schematically;
[0031] Figure 4 An architectural diagram of an immersive media system provided in one embodiment of the present application;
[0032] Figure 5 A flowchart of a file packaging method for a free-viewpoint video provided in an embodiment of the present application;
[0033] Figure 6 An interactive flow chart of a file packaging method for a free-viewpoint video provided in an embodiment of the present application;
[0034] Figure 7 An interactive flow chart of a file packaging method for a free-viewpoint video provided in an embodiment of the present application;
[0035] Figure 8 A schematic diagram of the structure of a file packaging device for free-viewpoint video provided in one embodiment of the present application;
[0036] Figure 9 A schematic diagram of the structure of a file packaging device for free-viewpoint video provided in one embodiment of the present application;
[0037] Figure 10 A schematic diagram of the structure of a file packaging device for free-viewpoint video provided in one embodiment of the present application;
[0038] Figure 11 It is a schematic block diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0040] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0041] The embodiments of the present application relate to data processing technology for immersive media.
[0042] Before introducing the technical solution of this application, the following is an introduction to the relevant knowledge of this application:
[0043] Free-viewpoint video: Immersive media video captured by multiple cameras, including different viewpoints, supporting user interaction with 3DoF+ or 6DoF. Also known as multi-view video.
[0044] Track: A media file is a collection of media data in the process of media file encapsulation. A media file can be composed of multiple tracks. For example, a media file can contain a video track, an audio track, and a subtitle track.
[0045] Sample: A sample is a unit of media file encapsulation. A media track consists of many samples. For example, a sample in a video track is usually a video frame.
[0046] DoF: Degree of Freedom. In a mechanical system, this refers to the number of independent coordinates. In addition to translational degrees of freedom, this also includes rotational and vibrational degrees of freedom. In this embodiment, it refers to the degrees of freedom that allow users to move and interact with content while viewing immersive media.
[0047] 3DoF: Three degrees of freedom, refers to the three degrees of freedom of the user's head rotating around the XYZ axes. Figure 1 The schematic diagram of the three degrees of freedom is shown schematically. Figure 1As shown, at a certain location or point, you can rotate on three axes: you can turn your head, tilt it up and down, and even tilt it. This three-degree-of-freedom experience allows users to immerse themselves in a scene 360 degrees. If it's static, it can be considered a panoramic image. If the panoramic image is dynamic, it's a panoramic video, or VR video. However, VR videos have certain limitations: users can't move around or choose a specific location to view.
[0048] 3DoF+: In addition to the three degrees of freedom (DOF), users also have limited freedom of movement along the X, Y, and Z axes. This can also be called restricted six degrees of freedom, and the corresponding media stream can be called a restricted six-DOF media stream. Figure 2 A schematic diagram of three degrees of freedom + is shown schematically.
[0049] 6DoF: In addition to the three degrees of freedom, users also have the freedom to move freely along the X, Y, and Z axes. The corresponding media stream can be called a six-degree-of-freedom media stream. Figure 3 A schematic diagram of six degrees of freedom is shown. 6DoF media refers to six-degree-of-freedom video, meaning it provides users with a high-degree-of-freedom viewing experience, allowing them to freely move their viewpoints along the X, Y, and Z axes of three-dimensional space, as well as freely rotate their viewpoints around the X, Y, and X axes. 6DoF media is a combination of videos captured from different spatial perspectives by a camera array. To facilitate the expression, storage, compression, and processing of 6DoF media, 6DoF media data is represented as a combination of the following information: texture maps captured by multiple cameras, depth maps corresponding to the multi-camera texture maps, and corresponding 6DoF media content description metadata. This metadata includes multi-camera parameters and information describing the 6DoF media's stitching layout and edge protection. At the encoder, the multi-camera texture map information and corresponding depth map information are spliced together, and the description of the splicing method is written into the metadata according to defined syntax and semantics. The spliced multi-camera depth map and texture map information are encoded using a planar video compression method. After being transmitted to the terminal for decoding, the user's requested 6DoF virtual viewpoint is synthesized, providing the user with a 6DoF media viewing experience.
[0050] Depth map: As a way to express three-dimensional scene information, the grayscale value of each pixel in the depth map can be used to represent the distance between a certain point in the scene and the camera.
[0051] AVS: Audio Video Coding Standard, audio and video coding standard.
[0052] AVS3: The third generation audio and video codec standard launched by the AVS working group
[0053] ISOBMFF: ISO Based Media File Format, a media file format based on the ISO (International Standard Organization) standard. ISOBMFF is a media file encapsulation standard, and the most typical ISOBMFF file is the MP4 (Moving Picture Experts Group 4) file.
[0054] DASH: dynamic adaptive streaming over HTTP, dynamic adaptive streaming over HTTP is an adaptive bitrate streaming technology that enables high-quality streaming media to be delivered over the Internet through traditional HTTP network servers.
[0055] MPD: media presentation description, media presentation description signaling in DASH, used to describe media segment information.
[0056] Representation: In DASH, a combination of one or more media components, such as a video file of a certain resolution, can be considered a Representation.
[0057] Adaptation Sets: In DASH, a collection of one or more video streams. An Adaptation Set can contain multiple Representations.
[0058] HEVC: High Efficiency Video Coding, international video coding standard HEVC / H.265.
[0059] VVC: versatile video coding, international video coding standard VVC / H.266.
[0060] Intra(picture)Prediction: Intra-frame prediction.
[0061] Inter(picture)Prediction: Inter-frame prediction.
[0062] SCC: screen content coding, screen content coding.
[0063] Immersive media refers to media content that provides consumers with an immersive experience. Based on the degree of freedom users have when consuming media content, immersive media can be categorized into 3DoF media, 3DoF+ media, and 6DoF media. Common 6DoF media include multi-view video and point cloud media.
[0064] Free-viewpoint video is usually captured by a camera array from multiple angles of the same three-dimensional scene, forming the scene's texture information (color information, etc.) and depth information (spatial distance information, etc.). Based on the user's location information, the texture information and depth information from different cameras are combined to form 6DoF media for user consumption.
[0065] After the free viewpoint video is captured, it needs to be compressed and encoded. In the existing free viewpoint video technology, the video compression algorithm can be implemented by AVS3 encoding technology, HEVC encoding technology, etc.
[0066] Figure 4 This is an architectural diagram of an immersive media system provided in one embodiment of the present application. Figure 4 As shown, the immersive media system includes an encoding device and a decoding device. The encoding device may refer to a computer device used by the provider of the immersive media, and the computer device may be a terminal (such as a PC (Personal Computer), a smart mobile device (such as a smart phone), etc.) or a server. The decoding device may refer to a computer device used by the user of the immersive media, and the computer device may be a terminal (such as a PC (Personal Computer), a smart mobile device (such as a smart phone), a VR device (such as a VR helmet, VR glasses, etc.)). The data processing process of the immersive media includes the data processing process on the encoding device side and the data processing process on the decoding device side.
[0067] The data processing process on the encoding device side mainly includes:
[0068] (1) The acquisition and production process of immersive media content;
[0069] (2) The process of encoding and file packaging of immersive media. The data processing process on the decoding device side mainly includes:
[0070] (3) The process of decapsulating and decoding immersive media files;
[0071] (4) Rendering process of immersive media.
[0072] In addition, the transmission process of immersive media between the encoding device and the decoding device can be carried out based on various transmission protocols. The transmission protocols here may include but are not limited to: DASH (Dynamic Adaptive Streaming over HTTP, dynamic adaptive streaming media transmission) protocol, HLS (HTTP Live Streaming, dynamic bit rate adaptive transmission) protocol, SMTP (Smart Media Transport Protocol, smart media transmission protocol), TCP (Transmission Control Protocol, transmission control protocol), etc.
[0073] The following will be combined Figure 4 , each process involved in the data processing of immersive media is introduced in detail.
[0074] 1. Data processing process on the encoding device side:
[0075] (1) The process of acquiring and producing media content for immersive media.
[0076] 1) The process of acquiring media content of immersive media.
[0077] The media content of immersive media is obtained by capturing the sound and visual scenes of the real world through capture devices.
[0078] In one implementation, the capture device may refer to a hardware component provided in the encoding device, for example, the capture device may refer to a microphone, camera, sensor, etc. of a terminal. In another implementation, the capture device may also be a hardware device connected to the encoding device, for example, a camera connected to a server.
[0079] The capture device may include, but is not limited to, an audio device, a camera device, and a sensor device. The audio device may include an audio sensor, a microphone, etc. The camera device may include a standard camera, a stereo camera, a light field camera, etc. The sensor device may include a laser device, a radar device, etc.
[0080] Multiple capture devices can be deployed at specific locations in real space to simultaneously capture audio and video content from different angles within that space. The captured audio and video content is synchronized in both time and space. The media content collected by the capture devices is called immersive media raw data.
[0081] 2) The production process of media content for immersive media.
[0082] The captured audio content itself is suitable for audio encoding for immersive media. The captured video content undergoes a series of production processes before it becomes suitable for video encoding for immersive media. The production process includes:
[0083] ① Stitching. Since the captured video content is shot by the capture device at different angles, stitching refers to stitching the video content shot at various angles into a complete video that can reflect the 360-degree visual panorama of the real space. In other words, the stitched video is a panoramic video (or spherical video) represented in three-dimensional space.
[0084] ② Projection. Projection refers to the process of mapping a spliced 3D video onto a 2D image. The resulting 2D image is called a projected image. Projection methods include, but are not limited to, latitude and longitude projection and regular hexahedron projection.
[0085] ③Region encapsulation. The projected image can be encoded directly, or the projected image can be region encapsulated before encoding. In practice, it is found that in the data processing process of immersive media, encoding the two-dimensional projection image after region encapsulation can greatly improve the video coding efficiency of the immersive media. Therefore, the region encapsulation technology is widely used in the video processing process of immersive media. The so-called region encapsulation refers to the process of performing conversion processing on the projection image by region. The region encapsulation process converts the projection image into an encapsulated image. The region encapsulation process specifically includes: dividing the projection image into multiple mapping regions, and then performing conversion processing on the multiple mapping regions to obtain multiple encapsulated regions, and mapping the multiple encapsulated regions to a 2D image to obtain an encapsulated image. Among them, the mapping region refers to the region obtained by division in the projection image before performing region encapsulation; the encapsulation region refers to the region located in the encapsulated image after performing region encapsulation.
[0086] The conversion process may include, but is not limited to, mirroring, rotating, rearranging, upsampling, downsampling, changing the resolution of the region, and moving the region.
[0087] It should be noted that since capture devices can only capture panoramic video, after such video is processed by the encoding device and transmitted to the decoding device for appropriate data processing, users on the decoding device can only view 360-degree video information by performing certain specific actions (such as head rotation). Non-specific actions (such as moving the head) do not produce corresponding video changes, resulting in a poor VR experience. Therefore, it is necessary to provide additional depth information that matches the panoramic video to achieve a better immersion and a more optimal VR experience. This involves 6DoF (six degrees of freedom) production technology. When users can move relatively freely in a simulated scene, it is called 6DoF. When using 6DoF production technology to produce immersive media video content, the capture device generally uses a light field camera, laser equipment, radar equipment, etc. to capture point cloud data or light field data in space. During the above production processes ①-③, some specific processing is also required, such as cutting and mapping the point cloud data and calculating depth information.
[0088] (2) The process of encoding and file packaging of immersive media.
[0089] The captured audio content can be directly audio-encoded to form an audio stream of immersive media. After the above-mentioned production process ①-② or ①-③, the projected image or the encapsulated image is video-encoded to obtain a video stream of immersive media. It should be noted here that if 6DoF production technology is adopted, a specific encoding method (such as point cloud encoding) needs to be used for encoding during the video encoding process. The audio stream and the video stream are encapsulated in a file container according to the file format of immersive media (such as ISOBMFF (ISOBase Media File Format, ISO base media file format)) to form a media file resource of immersive media. The media file resource can be a media file or a media fragment to form a media file of immersive media; and according to the file format requirements of immersive media, the media presentation description information (MPD) is used to record the metadata of the media file resource of the immersive media. The metadata here is a general term for information related to the presentation of immersive media. The metadata may include description information of the media content, description information of the window, and signaling information related to the presentation of media content, etc. As Figure 1 As shown, the encoding device stores the media presentation description information and media file resources formed after the data processing process.
[0090] The immersive media system supports data boxes, which are data blocks or objects containing metadata. Specifically, a data box contains metadata about the corresponding media content. Immersive media can include multiple data boxes, such as a Sphere Region Zooming Box, which contains metadata describing sphere region zooming information; a 2D Region Zooming Box, which contains metadata describing 2D region zooming information; a Region Wise Packing Box, which contains metadata describing the corresponding information during the region packing process, and so on.
[0091] 2. Data processing on the decoding device side:
[0092] (3) The process of decapsulating and decoding immersive media files;
[0093] The decoding device can obtain the media file resources and corresponding media presentation description information of the immersive media from the encoding device through the recommendation of the encoding device or adaptively and dynamically according to the user needs of the decoding device. For example, the decoding device can determine the user's orientation and position based on the user's head / eye / body tracking information, and then dynamically request the encoding device to obtain the corresponding media file resources based on the determined orientation and position. The media file resources and media presentation description information are transmitted from the encoding device to the decoding device through a transmission mechanism (such as DASH, SMT). The file decapsulation process on the decoding device side is the opposite of the file encapsulation process on the encoding device side. The decoding device decapsulates the media file resources according to the file format requirements of the immersive media to obtain audio streams and video streams. The decoding process on the decoding device side is the opposite of the encoding process on the encoding device side. The decoding device performs audio decoding on the audio stream to restore the audio content.
[0094] In addition, the decoding process of the video stream by the decoding device includes the following:
[0095] ① Decoding the video stream to obtain a planar image; based on the metadata provided by the media presentation description information, if the metadata indicates that the immersive media has performed a region encapsulation process, the planar image is a packaged image; if the metadata indicates that the immersive media has not performed a region encapsulation process, the planar image is a projected image;
[0096] ② If the metadata indicates that the immersive media has performed a regional encapsulation process, the decoding device will perform regional decapsulation on the encapsulated image to obtain a projected image. Here, regional decapsulation is the opposite of regional encapsulation. Regional decapsulation refers to the process of performing inverse conversion processing on the encapsulated image according to the region. Regional decapsulation converts the encapsulated image into a projected image. The process of regional decapsulation specifically includes: performing inverse conversion processing on multiple encapsulated regions in the encapsulated image according to the instructions of the metadata to obtain multiple mapping regions, and mapping the multiple mapping regions to a 2D image to obtain a projected image. Inverse conversion processing refers to processing that is opposite to the conversion processing. For example, if the conversion processing refers to a counterclockwise rotation of 90 degrees, then the inverse conversion processing refers to a clockwise rotation of 90 degrees.
[0097] ③ Reconstruct the projected image according to the media presentation description information to convert it into a 3D image. The reconstruction process here refers to the process of re-projecting the two-dimensional projected image into a 3D space.
[0098] (4) Rendering process of immersive media.
[0099] The decoding device renders the audio content obtained by audio decoding and the 3D image obtained by video decoding according to the metadata related to rendering and viewport in the media presentation description information. Once the rendering is completed, the playback output of the 3D image is realized. In particular, if 3DoF and 3DoF+ production technologies are adopted, the decoding device mainly renders the 3D image based on the current viewpoint, parallax, depth information, etc. If 6DoF production technology is adopted, the decoding device mainly renders the 3D image in the viewport based on the current viewpoint. Among them, the viewpoint refers to the user's viewing position, the parallax refers to the difference in line of sight between the user's two eyes or the difference in line of sight caused by movement, and the viewport refers to the viewing area.
[0100] The immersive media system supports data boxes, which are data blocks or objects containing metadata. Specifically, a data box contains metadata about the corresponding media content. Immersive media can include multiple data boxes, such as a Sphere Region Zooming Box, which contains metadata describing sphere region zooming information; a 2D Region Zooming Box, which contains metadata describing 2D region zooming information; and a Region Wise Packing Box, which contains metadata describing the corresponding information during the region packing process.
[0101] In some embodiments, the following file encapsulation mode is proposed for encapsulating free-viewpoint videos:
[0102] 1. Free View Track Group
[0103] If a free-viewpoint video is encapsulated into multiple video tracks, these video tracks should be associated through a free-viewpoint track group. The free-viewpoint track group is defined as follows:
[0104]
[0105]
[0106] A free-view track group is obtained by expanding the track group data box and is identified by the 'a3fg' track group type. All tracks containing the 'afvg' type TrackGroupTypeBox have the same group ID and belong to the same track group. The semantics of the fields in AvsFreeViewGroupBox are as follows:
[0107] camera_count: Indicates the number of cameras from which the free-viewpoint texture information or depth information contained in this track comes.
[0108] camera_id: Indicates the camera identifier corresponding to each camera, which corresponds to the value in AvsFreeViewInfoBox in the current track.
[0109] depth_texture_type: indicates the type of texture information or depth information captured by the corresponding camera contained in this track. The values are shown in Table 1 below.
[0110] Table 1
[0111]
[0112] 2. Free View Information Data Box
[0113]
[0114]
[0115] stitching_layout: Indicates whether the texture map and depth map in the track are stitched and encoded. The values are shown in Table 2:
[0116] Table 2
[0117]
[0118] depth_padding_size: The guard band width of the depth map.
[0119] texture_padding_size: The width of the guard band of the texture image.
[0120] camera_model: indicates the camera model type. The values are shown in Table 3:
[0121] Table 3
[0122]
[0123] camera_count: The number of all cameras capturing video.
[0124] camera_id: The camera identifier corresponding to each view.
[0125] camera_pos_x, camera_pos_y, camera_pos_z: indicate the x, y, and z component values of the camera position, respectively.
[0126] focal_length_x, focal_length_y: indicate the x and y component values of the camera focal length respectively.
[0127] camera_resolution_x, camera_resolution_y: The resolution width and height of the texture map and depth map collected by the camera.
[0128] depth_downsample_factor: The depth map downsampling factor. The actual resolution width and height of the depth map is 1 / 2 of the camera acquisition resolution width and height. depth_downsample_factor .
[0129] depth_vetex_x, depth_vetex_y: The x and y component values of the upper left vertex of the depth map relative to the origin of the planar frame (the upper left vertex of the planar frame).
[0130] texture_vetex_x, texture_vetex_y: The x and y component values of the upper left vertex of the texture image relative to the origin of the plane frame (the upper left vertex of the plane frame).
[0131] para_num: The number of user-defined camera parameters.
[0132] para_type: The type of user-defined camera parameters.
[0133] para_length: The length of user-defined camera parameters in bytes.
[0134] camera_parameter: user-defined parameters.
[0135] As can be seen from the above, although the above embodiment indicates the parameter information related to the free-viewpoint video, it also supports multi-track packaging of the free-viewpoint. However, this solution does not indicate the codec independence of the texture and depth maps corresponding to different cameras, making it impossible for the client to determine whether it can partially decode the texture map and depth map of a specific camera in single-track packaging mode. Similarly, in the absence of the codec independence indication information, the server cannot determine whether the free-viewpoint video packaged in a single track can be re-packaged as a multi-track.
[0136] In order to solve the above technical problems, the present application adds codec independence indication information to the video track. The codec independence indication information is used to indicate whether the video data of a single perspective among the M perspectives corresponding to the video track depends on the video data of other perspectives during encoding and decoding. In this way, in the single-track encapsulation mode, the client can determine whether the texture map and depth map of a specific camera can be partially decoded based on the codec independence indication information. In addition, in the single-track encapsulation mode, the server can also determine whether the free-viewpoint video encapsulated in a single track can be re-encapsulated as multiple tracks based on the codec independence indication information, thereby improving the processing flexibility of media files and improving the decoding efficiency of media files.
[0137] The following describes the technical solutions of the embodiments of the present application in detail through some embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0138] Figure 5 A flowchart of a method for packaging a free-viewpoint video file is provided in an embodiment of the present application. Figure 5 As shown, the method includes the following steps:
[0139] S501: A first device obtains a code stream of free-viewpoint video data.
[0140] The free-viewpoint video data includes video data of N viewpoints, where N is a positive integer.
[0141] The free-viewpoint video data of the embodiment of the present application is video data of N viewpoints captured by N cameras. For example, if N is 6, these 6 cameras with different viewpoints capture video data to obtain video data of 6 viewpoints. The video data of these 6 viewpoints constitute the free-viewpoint video data of the embodiment of the present application.
[0142] In some embodiments, the free-view video data is also referred to as multi-view video data.
[0143] In the embodiment of the present application, the first device obtains the code stream of the free-viewpoint video data in the following ways, but is not limited to:
[0144] Method 1: The first device obtains the code stream of the free-viewpoint video data from other devices.
[0145] For example, the first device obtains the code stream of the free-viewpoint video data from the storage device, and obtains the code stream of the free-viewpoint video data from other encoding devices.
[0146] In a second method, the first device encodes the free-viewpoint video data to obtain a code stream of the free-viewpoint video data. For example, the first device is an encoding device. After obtaining the free-viewpoint video data from a capture device (such as a camera), the first device encodes the free-viewpoint video data to obtain a code stream of the free-viewpoint video data.
[0147] The embodiment of the present application does not limit the specific content of the video data. For example, the video data includes at least one of the collected texture map data and depth map data.
[0148] S502. The first device encapsulates the code stream of the free-viewpoint video data into at least one video track to obtain a media file of the free-viewpoint video data, where the video track includes codec independence indication information and video code streams of M viewpoints. The codec independence indication information is used to indicate whether video data of a single viewpoint among the M viewpoints corresponding to the video track depends on video data of other viewpoints during encoding and decoding, where M is a positive integer less than or equal to N.
[0149] Specifically, the first device encapsulates the code stream of the free-viewpoint video data into at least one video track, and the at least one video track forms a media file of the free-viewpoint video data.
[0150] In a possible implementation, a single-track encapsulation mode is adopted to encapsulate the code stream of the free-viewpoint video data into one video track.
[0151] In one possible implementation, a multi-track encapsulation method is used to encapsulate the free-viewpoint video data stream into multiple video tracks. For example, the video stream corresponding to each of N viewpoints is encapsulated into a single video track, resulting in N video tracks. Alternatively, the video stream corresponding to one or more of the N viewpoints is encapsulated into a single video track, resulting in multiple video tracks, each of which may include a video stream corresponding to at least one viewpoint.
[0152] In order to facilitate the processing of media files by the client or server, this application adds a codec independence indication in the video track, so that the client or server processes the media file according to the codec independence indication.
[0153] In some embodiments, codec independence indication information is added to each formed video track, and the codec independence indication information is used to indicate whether video data of a single perspective among multiple perspectives corresponding to the video track depends on video data of other perspectives during encoding and decoding.
[0154] In some embodiments, when encoding a video, the encoding method of the N perspectives is consistent. Therefore, codec independence indication information can be added to one or more video tracks. The codec independence indication information is used to indicate whether the video data of a single perspective among the N perspectives depends on the video data of other perspectives during encoding and decoding.
[0155] In some embodiments, if the value of the codec independence indication information is a first value, it indicates that the texture map data of a single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding, or that the depth map data of a single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding; or
[0156] If the value of the codec independence indication information is the second value, it indicates that the texture map data of a single view depends on the texture map data of other views during encoding and decoding, and the depth map data of a single view depends on the depth map data of other views during encoding and decoding; or
[0157] If the value of the codec independence indication information is the third value, it indicates that the texture map data and depth map data of a single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of a single perspective depend on each other during encoding and decoding; or
[0158] If the value of the codec independence indication information is the fourth value, it indicates that the texture map data and depth map data of a single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of a single perspective do not depend on each other during encoding and decoding.
[0159] The corresponding relationship between the value of the above-mentioned codec independence indication information and the codec independence indicated by the codec independence indication information is shown in Table 4:
[0160] Table 4
[0161]
[0162] The embodiment of the present application does not limit the specific values of the above-mentioned first numerical value, second numerical value, third numerical value and fourth numerical value, which are determined according to actual needs.
[0163] Optionally, the first value is 0.
[0164] Optionally, the second value is 1.
[0165] Optionally, the third value is 2.
[0166] Optionally, the fourth value is 3.
[0167] In some embodiments, the codec independence indication information may be added to a free view information data box of a video track.
[0168] If the encapsulation standard of the above media file is ISOBMFF, the codec independence indication information is represented by the field codec_independency.
[0169] The free viewpoint information data box of the embodiment of the present application includes the following contents:
[0170]
[0171] In the embodiment of the present application, the unsigned int(8)stitching_layout field in the free view information data box is deleted. Among them, stitching_layout: indicates whether the texture map and depth map in the track are stitched and encoded, as shown in Table 5:
[0172] Table 5
[0173]
[0174] Correspondingly, the embodiment of the present application adds a codec_independency field, which indicates the codec independence between the texture map and depth map corresponding to each camera in the track, as shown in Table 6:
[0175] Table 6
[0176]
[0177]
[0178] depth_padding_size: The guard band width of the depth map.
[0179] texture_padding_size: The width of the guard band of the texture image.
[0180] camera_model: indicates the camera model type, as shown in Table 7.
[0181] Table 7
[0182] Camera model 6DoF video camera model 0 Pinhole Model 1 Fisheye model other reserve
[0183] camera_count: The number of all cameras capturing video.
[0184] camera_id: The camera identifier corresponding to each view.
[0185] camera_pos_x, camera_pos_y, camera_pos_z: indicate the x, y, and z component values of the camera position, respectively.
[0186] focal_length_x, focal_length_y: indicate the x and y component values of the camera focal length respectively.
[0187] camera_resolution_x, camera_resolution_y: The resolution width and height of the texture map and depth map collected by the camera.
[0188] depth_downsample_factor: The depth map downsampling factor. The actual resolution width and height of the depth map is 1 / 2 of the camera acquisition resolution width and height. depth_downsample_factor .
[0189] depth_vetex_x, depth_vetex_y: The x and y component values of the upper left vertex of the depth map relative to the origin of the planar frame (the upper left vertex of the planar frame).
[0190] texture_vetex_x, texture_vetex_y: The x and y component values of the upper left vertex of the texture image relative to the origin of the plane frame (the upper left vertex of the plane frame).
[0191] In some embodiments, if the bitstream of the target-view video data is obtained by the first device from another device, the bitstream of the target-view video data includes codec independence indication information. This codec independence indication information is used to indicate whether the video data of a single view among the N views depends on the video data of other views during encoding and decoding. In this way, the first device can determine whether the video data of each view in the target-view video data depends on the video data of other views during encoding based on the codec independence indication information carried in the bitstream, and then add the codec independence indication information to each generated video track.
[0192] In some embodiments, if the target view video data's code stream carries codec independence indication information, the embodiment of the present application extends the free view video bitstream syntax. Taking 6DoF video as an example, the details are shown in Table 8:
[0193] Table 8
[0194]
[0195]
[0196] As shown in Table 8, the field 6DoF video stitching layout stitching_layout is deleted from the free viewpoint video bitstream syntax of the embodiment of the present application. The stitching_layout is an 8-bit unsigned integer used to identify whether the 6DoF video uses a stitching layout of texture map and depth map. The specific values are shown in Table 9.
[0197] Table 9
[0198]
[0199] Correspondingly, the codec_independency field is added to the free-viewpoint video bitstream syntax shown in Table 8. codec_independency is an 8-bit unsigned integer used to identify the codec independence between the texture map and depth map corresponding to each camera of the 6DoF video. The specific values are shown in Table 10.
[0200] Table 10
[0201]
[0202] The following describes the fields in Table 8 above:
[0203] The marker bit marker_bit is a binary variable. A value of 1 is used to prevent a pseudo start code from appearing in the 6DoF video format extension bit stream.
[0204] Texture_padding_size is an 8-bit unsigned integer that represents the number of pixels to pad the texture image. The value ranges from 0 to 255.
[0205] The depth map padding size depth_padding_size is an 8-bit unsigned integer. It indicates the number of pixels to pad the depth map, and the value ranges from 0 to 255.
[0206] The camera_number parameter is an 8-bit unsigned integer ranging from 1 to 255. It indicates the number of cameras used to capture 6DoF video.
[0207] The camera model, camera_model, is an 8-bit unsigned integer with a value of 1 to 255. It is used to indicate the camera model type. The 6DoF video camera model is shown in Table 11:
[0208] Table 11
[0209]
[0210] The camera capture resolution, camera_resolution_x, is a 32-bit unsigned integer. It represents the resolution in the x direction of the camera capture.
[0211] The camera acquisition resolution camera_resolution_y is a 32-bit unsigned integer, indicating the resolution in the y direction of the camera acquisition.
[0212] 6DoF video resolution video_resolution_x is a 32-bit unsigned integer. It represents the resolution of the 6DoF video in the x direction.
[0213] 6DoF video resolution video_resolution_y, a 32-bit unsigned integer, indicates the resolution of the 6DoF video in the y direction.
[0214] The camera's translation matrix camera_translation_matrix[3] is a 3*32-bit floating-point matrix that represents the camera's translation matrix.
[0215] The camera's rotation matrix camera_rotation_matrix[3][3] is a 9*32-bit floating-point matrix that represents the camera's rotation matrix.
[0216] The camera focal length x component, camera_focal_length_x, is a 32-bit floating-point number. It represents the camera's focal length fx.
[0217] The camera focal length y component, camera_focal_length_y, is a 32-bit floating point number. It represents the camera's focal length fy.
[0218] The x component of the camera optical axis offset, camera_principle_point_x, is a 32-bit floating point number. It represents the offset px of the camera's optical axis in the image coordinate system.
[0219] The camera optical axis offset y component, camera_principle_point_y, is a 32-bit floating point number. It represents the camera optical axis offset py in the image coordinate system.
[0220] The x component of the texture image's top left corner coordinate, texture_top_left_x[camera_number], is a 32-bit unsigned integer array of size camera_number, representing the x coordinate of the top left corner of the texture image of the corresponding camera in the 6DoF video frame.
[0221] The y component of the texture image's top left corner coordinate, texture_top_left_y[camera_number], is a 32-bit unsigned integer array of size camera_number, representing the y coordinate of the top left corner of the texture image of the corresponding camera in the 6DoF video frame.
[0222] The x component of the texture image's lower right corner coordinate, texture_bottom_right_x[camera_number], is a 32-bit unsigned integer array of size camera_number, representing the x coordinate of the lower right corner of the texture image of the corresponding camera in the 6DoF video frame.
[0223] The y component of the texture image's lower right corner coordinate, texture_bottom_right_y[camera_number], is a 32-bit unsigned integer array of size camera_number, representing the y coordinate of the lower right corner of the texture image of the corresponding camera in the 6DoF video frame.
[0224] The x component of the top left corner coordinate of the depth map depth_top_left_x[camera_number] is a 32-bit unsigned integer array with an array size of camera_number, which represents the x coordinate of the top left corner of the depth map of the corresponding camera in the 6DoF video frame.
[0225] The y component of the top left corner coordinate of the depth map depth_top_left_y[camera_number] is a 32-bit unsigned integer array with an array size of camera_number, which represents the y coordinate of the top left corner of the depth map of the corresponding camera in the 6DoF video frame.
[0226] The x-component of the coordinate of the lower right corner of the depth map, depth_bottom_right_x[camera_number], is a 32-bit unsigned integer array of size camera_number, representing the x-coordinate of the lower right corner of the depth map of the corresponding camera in the 6DoF video frame.
[0227] The y component of the coordinate of the bottom right corner of the depth map, depth_bottom_right_y[camera_number], is a 32-bit unsigned integer array of size camera_number, representing the y coordinate of the bottom right corner of the depth map of the corresponding camera in the 6DoF video frame.
[0228] The nearest depth distance depth_range_near used for depth map quantization is a 32-bit floating point number. The minimum depth distance from the optical center used for depth map quantization.
[0229] The maximum depth distance depth_range_far used for depth map quantization is a 32-bit floating point number. The maximum depth distance from the optical center used for depth map quantization.
[0230] The depth map downsampling flag depth_scale_flag is an 8-bit unsigned integer ranging from 1 to 255. It is used to indicate the downsampling type of the depth map.
[0231] The background texture flag background_texture_flag is an 8-bit unsigned integer used to indicate whether to transmit the background texture map of multiple cameras.
[0232] The background depth flag background_depth_flag is an 8-bit unsigned integer used to indicate whether to transmit the background depth map of multiple cameras.
[0233] If background depth is applied, the decoded depth map background frame does not participate in viewpoint synthesis, and subsequent frames participate in virtual viewpoint synthesis. The depth map downsampling flag is shown in Table 12:
[0234] Table 12
[0235]
[0236] From the above, it can be seen that the first device determines the specific value of the codec independence indication information based on whether the video data of a single perspective among the M perspectives corresponding to the video track depends on the video data of other perspectives during encoding and decoding, and adds the specific value of the codec independence indication information to the video track.
[0237] In some embodiments, if the video data corresponding to each of the N perspectives does not depend on the video data corresponding to other perspectives during encoding, and the encapsulation mode of the code stream of the free-perspective video data is a single-track encapsulation mode, the first device adds codec independence indication information to the free-perspective information data box of a video track formed by the single-track encapsulation mode, wherein the value of the codec independence indication information is a third numerical value or a fourth numerical value. In this way, the server or client can determine, based on the value of the codec independence indication information, that the video data corresponding to each of the N perspectives does not depend on the video data corresponding to other perspectives during encoding, and can then request the media files corresponding to some perspectives to be decapsulated, or recapsulate a video track encapsulated in the single-track mode into multiple video tracks.
[0238] S503: Send the media file of the free-viewpoint video data to the client or server.
[0239] According to the above method, the code stream of the free-viewpoint video data is encapsulated into at least one video track to obtain a media file of the free-viewpoint video data, and codec independence indication information is added to the video track. Then, the media file including the codec independence indication information is sent to the client or server, so that the media file is sent to the client or server and the client or server processes the media file according to the codec independence indication information carried in the media file. For example, if the codec independence indication information indicates that the video data corresponding to a single viewpoint does not rely on the video data corresponding to other viewpoints during encoding, the client or server can request the media files corresponding to some viewpoints to be decapsulated, or re-encapsulate a video track encapsulated in a single-track mode into multiple video tracks.
[0240] The file encapsulation method for free-viewpoint video provided in an embodiment of the present application adds codec independence indication information to the video track. The codec independence indication information is used to indicate whether the video data of a single perspective among the M perspectives corresponding to the video track depends on the video data of other perspectives during encoding and decoding. In this way, in single-track encapsulation mode, the client can determine whether the texture map and depth map of a specific camera can be partially decoded based on the codec independence indication information. In addition, in single-track encapsulation mode, the server can also determine whether the free-viewpoint video encapsulated in a single track can be re-encapsulated as a multi-track based on the codec independence indication information, thereby improving the processing flexibility of the media file.
[0241] In some embodiments, if the encoding mode of the free-viewpoint video data is AVS3 encoding mode, the media file of the embodiment of the present application may be encapsulated in the form of subsamples. In this case, the method of the embodiment of the present application further includes:
[0242] S500: Encapsulate at least one of header information required for decoding, texture map information of at least one perspective, and depth map information of at least one perspective in the media file in the form of a subsample.
[0243] The subsample data box includes a subsample data box flag and subsample indication information. The subsample data box flag is used to indicate the division method of the subsample, and the subsample indication information is used to indicate the content included in the subsample.
[0244] The content included in the subsample includes at least one of header information required for decoding, texture map information of at least one perspective, and depth map information of at least one perspective.
[0245] In some embodiments, if the value of the subsample indication information is a fifth value, it indicates that a subsample includes header information required for decoding; or,
[0246] If the value of the subsample indication information is the sixth value, it indicates that one subsample includes texture map information corresponding to N perspectives (or cameras) in the current video frame; or,
[0247] If the value of the subsample indication information is the seventh value, it indicates that one subsample includes depth map information corresponding to N perspectives (or cameras) in the current video frame; or,
[0248] If the value of the subsample indication information is the eighth value, it indicates that one subsample includes texture map information and depth map information corresponding to one viewing angle (or camera) in the current video frame; or
[0249] If the value of the subsample indication information is a ninth value, it indicates that one subsample includes texture map information corresponding to one viewing angle (or camera) in the current video frame; or
[0250] If the value of the subsample indication information is the tenth value, it indicates that one subsample includes depth map information corresponding to one viewing angle (or camera) in the current video frame.
[0251] The current video frame is formed by splicing video frames corresponding to the N perspectives. For example, video frames corresponding to N perspectives captured by N cameras at the same time point are spliced to form the current video frame.
[0252] The above-mentioned texture map information and / or depth map information may be understood as data required for decapsulating the texture map code stream or the depth map code stream.
[0253] Optionally, the above-mentioned texture map information and / or depth map information includes the position offset of the texture map code stream and / or depth map code stream in the media file. For example, the texture map code stream corresponding to each perspective is saved at the tail position of the media file, and the texture map information corresponding to perspective 1 includes the offset of the texture map information corresponding to perspective 1 at the tail position in the media file. According to the offset, the position of the texture map code stream corresponding to perspective 1 in the media file can be obtained.
[0254] The corresponding relationship between the values of the above sub-sample indication information and the content included in the sub-sample is shown in Table 13:
[0255] Table 13
[0256]
[0257] The embodiment of the present application does not limit the specific values of the fifth value, sixth value, seventh value, eighth value, ninth value and tenth value mentioned above, which are determined according to actual needs.
[0258] Optionally, the fifth value is 0.
[0259] Optionally, the sixth value is 1.
[0260] Optionally, the seventh value is 2.
[0261] Optionally, the eighth value is 3.
[0262] Optionally, the ninth value is 4.
[0263] Optionally, the tenth value is 5.
[0264] Optionally, the flags field of the subsample data box takes a preset value, such as 1, indicating that the subsample includes valid content.
[0265] In some embodiments, the subsample indication information is represented by payloadType in the codec_specific_parameters field in the SubSampleInformationBox data box.
[0266] In one example, the value of the codec_specific_parameters field in the SubSampleInformationBox data box is as follows:
[0267]
[0268] The values of the payloadType field are shown in Table 14 below:
[0269] Table 14
[0270]
[0271]
[0272] In some embodiments, if the encoding mode of the free-viewpoint video data is AVS3 encoding mode, the video data corresponding to each of the N viewpoints does not depend on the video data corresponding to other viewpoints during encoding, and the encapsulation mode of the free-viewpoint video data stream is single-track encapsulation mode, then the above S500 includes the following S500-A:
[0273] S500-A: Encapsulate at least one of the texture map information and the depth map information corresponding to each of the N perspectives in the form of a subsample in the media file.
[0274] The data box of each sub-sample formed above includes a sub-sample data box flag and sub-sample indication information.
[0275] The value of the subsample data box flag is a preset value, for example, 1, indicating that the subsample is divided into subsamples based on viewing angles.
[0276] The value of the subsample indication information is determined according to the content included in the subsample, and may specifically include the following examples:
[0277] In example 1, if the subsample includes texture map information and depth map information corresponding to a viewing angle in the current video frame, the value of the subsample indication information corresponding to the subsample is the eighth value.
[0278] Example 2: If the sub-sample includes texture map information corresponding to a viewing angle in the current video frame, the value of the sub-sample indication information corresponding to the sub-sample is the ninth value.
[0279] Example 3: If the subsample includes depth map information corresponding to a viewing angle in the current video frame, the value of the subsample indication information corresponding to the subsample is the tenth value.
[0280] In an embodiment of the present application, if the encoding method of the free-viewpoint video data is the AVS3 encoding mode, the video data corresponding to each of the N viewpoints is not dependent on the video data corresponding to other viewpoints during encoding, and the encapsulation mode of the free-viewpoint video data stream is a single-track encapsulation mode, and the texture map information and depth map information corresponding to each of the N viewpoints are encapsulated in the media file in the form of subsamples. This allows the client to decode partial viewpoints according to its own needs after requesting the complete free-viewpoint video, thereby saving client computing resources.
[0281] Figure 6 This is an interactive flow chart of a file packaging method for a free-viewpoint video provided in an embodiment of the present application, such as Figure 6 As shown, the method includes the following steps:
[0282] S601: A first device obtains a code stream of free-viewpoint video data.
[0283] The free-viewpoint video data includes video data of N viewpoints, where N is a positive integer.
[0284] S602: The first device encapsulates the code stream of the free-viewpoint video data into at least one video track to obtain a media file of the free-viewpoint video data.
[0285] Among them, the video track includes codec independence indication information and video streams of M perspectives. The codec independence indication information is used to indicate whether the video data of a single perspective among the M perspectives corresponding to the video track depends on the video data of other perspectives during encoding and decoding, and M is a positive integer less than or equal to N.
[0286] S603: The first device sends the media file of the free-viewpoint video data to the client.
[0287] The execution process of the above S601 to S603 is consistent with the process of the above S501 to S503. Please refer to the detailed description of the above S501 to S503 and will not be repeated here.
[0288] S604: The client decapsulates the media file according to the codec independence indication information to obtain a video stream corresponding to at least one perspective.
[0289] In the media file of the present application, codec independence indication information is added, and the codec independence indication information is used to indicate whether the video data of a single perspective depends on the video data of other perspectives during encoding and decoding. In this way, after the client receives the media file, it can determine whether the video data of a single perspective corresponding to the media file depends on the video data of other perspectives during encoding and decoding based on the codec independence indication information carried by the media file. If it is determined that the video data of a single perspective corresponding to the media file depends on the video data of other perspectives during encoding and decoding, it means that the media file corresponding to a single perspective cannot be decapsulated. Therefore, the client needs to decapsulate the media files corresponding to all perspectives in the media file, and then obtain the video code streams corresponding to N perspectives.
[0290] If it is determined that the single-perspective video data corresponding to the media file does not depend on the video data of other perspectives during encoding and decoding, it means that the media file corresponding to a single perspective in the media file can be decapsulated. In this way, the client can decapsulate the media files of some perspectives as needed to obtain the video code streams corresponding to some perspectives, thereby saving client computing resources.
[0291] As can be seen from Table 4 above, the codec independence indication information indicates whether the video data of a single perspective depends on the video data of other perspectives during encoding and decoding through different values. In this way, the client can determine whether the video data of a single perspective among the N perspectives corresponding to the media file depends on the video data of other perspectives during encoding and decoding based on the value of the codec independence indication information and Table 4.
[0292] In some embodiments, the above S604 includes the following steps:
[0293] S604-A1: If the value of the codec independence indication information is the third value or the fourth value, and the code stream encapsulation mode is the single-track encapsulation mode, the client obtains the user's viewing angle;
[0294] S604-A1. The client determines a target viewing angle corresponding to the user's viewing angle based on the user's viewing angle and viewing angle information in the media file;
[0295] S604-A1. The client decapsulates the media file corresponding to the target viewing angle to obtain a video stream corresponding to the target viewing angle.
[0296] That is, in an embodiment of the present application, if the value of the codec independence indication information carried in the media file is the third value or the fourth value, as shown in Table 4, the third value is used to indicate that the texture map data or depth map data of a single perspective does not depend on the texture map data or depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of a single perspective are mutually dependent during encoding and decoding, and the fourth value is used to indicate that the texture map data or depth map data of a single perspective does not depend on the texture map data or depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of a single perspective are not mutually dependent during encoding and decoding. It can be seen from this that if the value of the codec independence indication information carried in the media file is the third value or the fourth value, it means that the texture map data or depth map data of each perspective in the media file does not depend on the texture map data or depth map data of other perspectives during encoding and decoding, so that the client can decode the texture map data and / or depth map data of some perspectives separately as needed.
[0297] Specifically, based on the user's viewing perspective and the perspective information in the media file, a target perspective that matches the user's viewing perspective is determined, and the media file corresponding to the target perspective is decapsulated to obtain a video stream corresponding to the target perspective. The embodiment of the present application does not limit the method for decapsulating the media file corresponding to the target perspective to obtain the video stream corresponding to the target perspective. Any existing method can be used. For example, based on the information corresponding to the target perspective in the media file, the position of the video stream corresponding to the target perspective in the media file is determined, and then the video stream corresponding to the target perspective is decapsulated to obtain the video stream corresponding to the target perspective, thereby decoding the video data corresponding to a portion of the perspective, saving client computing resources.
[0298] In some embodiments, if the value of the codec independence indication information is the first value or the second value, it means that the texture map data and / or depth map data of each perspective in the media file depends on the texture map data and / or depth map data of other perspectives during encoding and decoding. At this time, the client needs to decapsulate the media files corresponding to all N perspectives in the media file.
[0299] S605: The client decodes the video stream corresponding to at least one perspective to obtain reconstructed video data of at least one perspective.
[0300] After obtaining the video stream corresponding to at least one viewing angle according to the above method, the client can decode the video stream corresponding to the at least one viewing angle and render the decoded video data.
[0301] The process of decoding the video code stream may refer to the description of the existing technology and will not be described in detail here.
[0302] In some embodiments, if the encoding method is AVS3 video encoding method and the media file includes subsamples, the above S605 includes:
[0303] S605-A1. Obtain the subsample data box flag and subsample indication information included in the subsample data box.
[0304] As can be seen from S500 above, the first device can encapsulate the header information required for decoding, the texture map information of at least one view, and the depth map information of at least one view in the form of subsamples in the media file. Based on this, after receiving the media file, when the client detects that the media file includes a subsample, the client obtains the subsample data box flag and subsample indication information included in the subsample data box.
[0305] The subsample data box flag is used to indicate the division method of the subsample, and the subsample indication information is used to indicate the content included in the subsample.
[0306] The content included in the subsample includes at least one of header information required for decoding, texture map information of at least one perspective, and depth map information of at least one perspective.
[0307] S605-A2: Obtain the content included in the subsample according to the subsample data box flag and the subsample indication information.
[0308] For example, when the value of the subsample data box flag is 1, it indicates that the subsample includes valid content. Then, based on the correspondence between the subsample indication information and the content included in the subsample shown in Table 13, the client queries Table 13 for the content included in the subsample corresponding to the above subsample indication information.
[0309] S605-A3: Decapsulate the media file resource according to the content included in the subsample and the codec independence indication information to obtain a video stream corresponding to at least one perspective.
[0310] Among them, the content included in the subsample is shown in Table 13, and the value and indicated information of the codec independence indication information are shown in Table 4. In this way, the media file resource can be decapsulated based on whether the video data of a single perspective indicated by the codec independence indication information depends on the video data of other perspectives during encoding and decoding, and the content included in the subsample to obtain a video stream corresponding to at least one perspective.
[0311] For example, the codec independence indication information is used to indicate that the video data of a single perspective does not depend on the video data of other perspectives during encoding and decoding, and the content included in the subsample is the header information required for decoding. In this way, the client obtains the video stream corresponding to the target perspective according to the codec independence indication information, and decodes the video stream corresponding to the target perspective according to the header information required for decoding included in the subsample.
[0312] For another example, the codec independence indication information may indicate that the video data for a single view depends on the video data for other views during encoding and decoding, and the subsample includes all texture map information within the current video frame. In this case, after obtaining the video streams corresponding to all views based on the codec independence indication information, the client decodes the video streams corresponding to all views based on the texture map information within the current video frame included in the subsample.
[0313] It should be noted that, in the embodiment of the present application, after Table 4 and Table 13 are given, decapsulation of media file resources can be achieved through existing methods based on the content included in the subsample and the codec independence indication information, and no further examples will be given here.
[0314] In some embodiments, if the media file includes a subsample corresponding to each of the N perspectives, S605-A3 includes:
[0315] S605-A31: If the value of the codec independence indication information is the third value or the fourth value, and the code stream encapsulation mode is the single-track encapsulation mode, the client determines the target viewing angle according to the user's viewing angle and the viewing angle information in the media file;
[0316] S605-A32: The client obtains the subsample data box flag and subsample indication information included in the data box of the target subsample corresponding to the target perspective;
[0317] S605-A33: The client determines the content of the target subsample according to the subsample data box flag and subsample indication information included in the data box of the target subsample;
[0318] S605-A34: The client decapsulates the media file corresponding to the target perspective according to the content included in the target sub-sample to obtain the video stream corresponding to the target perspective.
[0319] Specifically, if the value of the codec independence indication information is the third value or the fourth value, it means that the video data of each perspective in the media file does not rely on the video data of other perspectives during encoding and decoding. In this way, the client determines the target perspective based on the user's viewing perspective and the perspective information in the media file. Then, the client obtains the target subsample corresponding to the target perspective from the subsamples corresponding to each perspective in the N perspectives, and then obtains the subsample data box flag and subsample indication information included in the data box of the target subsample, and determines the content included in the target subsample based on the subsample data box flag and subsample indication information included in the data box of the target subsample. For example, if the subsample data box flag is 1 and the value of the subsample indication information is 3, as shown in Table 13, it can be determined that the content included in the target subsample is the texture map information and depth map information corresponding to a single camera in the current video frame. Next, the client decapsulates the media file corresponding to the target perspective based on the content included in the target sub-sample to obtain the video stream corresponding to the target perspective. For example, the client obtains the position of the video stream corresponding to the target perspective in the media file based on the texture map information and depth map information included in the target sub-sample, and then decapsulates the video stream corresponding to the target perspective in the media file to obtain the video stream corresponding to the target perspective.
[0320] In an embodiment of the present application, the first device adds codec independence indication information to the video track. The codec independence indication information is used to indicate whether the video data of a single perspective among the M perspectives corresponding to the video track depends on the video data of other perspectives during encoding and decoding. In this way, in the single-track encapsulation mode, the client can determine whether the texture map and depth map of a specific camera can be partially decoded based on the codec independence indication information, thereby improving the client's processing flexibility for media files, and when the client decodes part of the media files based on the codec independence indication information, it can save the client's computing resources.
[0321] Figure 7 This is an interactive flow chart of a file packaging method for a free-viewpoint video provided in an embodiment of the present application, such as Figure 7 As shown, the method includes the following steps:
[0322] S701: The first device obtains a code stream of free-viewpoint video data.
[0323] The free-viewpoint video data includes video data of N viewpoints, where N is a positive integer.
[0324] S702: The first device encapsulates the code stream of the free-viewpoint video data into at least one video track to obtain a media file of the free-viewpoint video data.
[0325] Among them, the video track includes codec independence indication information and video streams of M perspectives. The codec independence indication information is used to indicate whether the video data of a single perspective among the M perspectives corresponding to the video track depends on the video data of other perspectives during encoding and decoding, and M is a positive integer less than or equal to N.
[0326] S703: The first device sends the media file of the free-viewpoint video data to the server.
[0327] The execution process of the above S701 to S703 is consistent with the above S501 to S503. Please refer to the detailed description of the above S501 to S503 and will not be repeated here.
[0328] S704: The server determines whether to decompose the at least one video track into multiple video tracks according to the codec independence indication information.
[0329] In some embodiments, if the value of the codec independence indication information is the third value or the fourth value, and the encapsulation mode of the code stream is the single-track encapsulation mode, then a video track formed by the single-track encapsulation mode is encapsulated into N video tracks according to the multi-track encapsulation mode, and each of the N video tracks includes video data corresponding to a single perspective.
[0330] In some embodiments, if the value of the codec independence indication information is the first value or the second value, and the code stream encapsulation mode is the single-track encapsulation mode, the server cannot encapsulate a video track formed by the single-track encapsulation mode into N video tracks.
[0331] In some embodiments, if the encoding method of the free-viewpoint video data is the AVS3 video encoding method, and the media file includes subsamples corresponding to each of N viewpoints, and the subsamples corresponding to each viewpoint include at least one of depth map information and texture map information corresponding to each viewpoint, then the above S704 includes:
[0332] S704-A1. For each of the N viewing angles, obtain a subsample data box flag and subsample indication information included in a data box of a subsample corresponding to the viewing angle, where the subsample data box flag is used to indicate a subsample division method, and the subsample indication information is used to indicate content included in the subsample;
[0333] S704-A2, obtaining the content of the subsample corresponding to each perspective according to the subsample data box flag and subsample indication information corresponding to each perspective;
[0334] S704-A3: Encapsulate a video track formed in the single-track encapsulation mode into N video tracks in a multi-track encapsulation mode according to the content included in the sub-sample corresponding to each viewing angle.
[0335] For example, taking a single view among N views as an example, the data box of the subsample corresponding to the view includes a subsample data box flag with a value of 1, and the subsample indication information has a value of 3. As can be seen from Table 13 above, the subsample indication information with a value of 3 indicates that the subsample includes the texture map information and depth map information corresponding to the view. Then, based on the texture map information and depth map information corresponding to the view, the video stream corresponding to the view is queried in the single video track formed in the single-track encapsulation mode. Based on the same method, the video stream corresponding to each view among the N views can be queried, and each corresponding video stream can be encapsulated into a video track to obtain N video tracks.
[0336] In some embodiments, after the server decomposes the single video track into multiple video tracks according to the codec independence indication information, the method of the embodiment of the present application further includes:
[0337] S705: The server decomposes at least one video track into multiple video tracks according to the codec independence indication information, and generates a first signaling.
[0338] In one example, if N=3, the server encapsulates the single-track video file into three video tracks Track1, Track2, and Track3, as follows:
[0339] Track1: {Camera1: ID=1; Pos=(100,0,100); Focal=(10,20)};
[0340] Track2: {Camera2: ID=2; Pos=(100,100,100); Focal=(10,20)};
[0341] Track3: {Camera3: ID=3; Pos=(0,0,100); Focal=(10,20)}.
[0342] The first signaling includes at least one of an identifier of a camera corresponding to each track in the N video tracks, location information of the camera, and focus information of the camera.
[0343] In an example, the first signaling includes three representations: Representation1, Representation2, and Representation3, as shown below:
[0344] Representation1: {Camera1: ID=1; Pos=(100,0,100); Focal=(10,20)};
[0345] Representation2: {Camera2: ID=2; Pos=(100,100,100); Focal=(10,20)};
[0346] Representation3: {Camera3: ID=3; Pos=(0,0,100); Focal=(10,20)}.
[0347] S706: The server sends the first signaling to the client.
[0348] S707. The client generates a first request message according to the first signaling. For example, the client selects and requests the video file corresponding to the target camera according to the user's location information and the shooting position information of each viewing angle in the first signaling (ie, the camera's location information).
[0349] For example, the target cameras are Camera2 and Camera3.
[0350] S708: The client sends a first request message to the server, where the first request message includes identification information of the target camera.
[0351] S709: The server sends the media file resources corresponding to the target camera to the client according to the first request information.
[0352] In an embodiment of the present application, the server decomposes the video track corresponding to the single-track mode into multiple video tracks when determining that the video track corresponding to the single-track mode can be decomposed into multiple video tracks based on the codec independence indication information, thereby supporting the client to request media resources corresponding to partial perspectives, thereby achieving the purpose of partial transmission and partial decoding.
[0353] It should be understood that Figures 5 to 7 This is only an example of the present application and should not be considered as limiting the present application.
[0354] The preferred embodiments of the present application are described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the technical concept of the present application, a variety of simple modifications can be made to the technical solution of the present application, and these simple modifications all fall within the scope of protection of the present application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. In order to avoid unnecessary repetition, the present application will not further explain various possible combinations. For another example, the various different embodiments of the present application can also be arbitrarily combined, and as long as they do not violate the ideas of the present application, they should also be regarded as the contents disclosed in the present application.
[0355] Combined with the above Figures 5 to 7, describes the method embodiment of the present application in detail, and the following is combined with Figures 8 to 11 , describe in detail the device embodiments of the present application.
[0356] Figure 8 This is a schematic diagram of the structure of a file packaging device for a free-viewpoint video provided in an embodiment of the present application. The device 10 is applied to a first device and includes:
[0357] An acquiring unit 11 is configured to acquire a code stream of free-viewpoint video data, wherein the free-viewpoint video data includes video data of N viewpoints, where N is a positive integer;
[0358] an encapsulation unit 12, configured to encapsulate the code stream of the free-viewpoint video data into at least one video track to obtain a media file of the free-viewpoint video data, wherein the video track includes codec independence indication information and video code streams of M views, wherein the codec independence indication information is used to indicate whether video data of a single view among the M views corresponding to the video track depends on video data of other views during encoding and decoding, where M is a positive integer less than or equal to N;
[0359] The sending unit 13 is configured to send the media file of the free-viewpoint video data to a client or a server.
[0360] Optionally, the video data includes at least one of texture map data and depth map data.
[0361] In some embodiments, if the value of the codec independence indication information is a first value, it indicates that the texture map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding, or the depth map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding; or
[0362] If the value of the codec independence indication information is a second value, it indicates that the texture map data of the single view depends on the texture map data of other views during encoding and decoding, and the depth map data of the single view depends on the depth map data of other views during encoding and decoding; or
[0363] If the value of the codec independence indication information is a third value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective depend on each other during encoding and decoding; or
[0364] If the value of the codec independence indication information is the fourth numerical value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective do not depend on each other during encoding and decoding.
[0365] In some embodiments, the encapsulation unit 12 is configured to encapsulate the code stream of the free-viewpoint video data into a video track using a single-track encapsulation mode.
[0366] In some embodiments, if the video data corresponding to each of the N perspectives does not depend on the video data corresponding to other perspectives during encoding, and the encapsulation mode of the code stream of the free perspective video data is a single-track encapsulation mode, the encapsulation unit 12 is also used to add the codec independence indication information to the free perspective information data box of the one video track, and the value of the codec independence indication information is the third value or the fourth value.
[0367] In some embodiments, if the encoding mode of the free-viewpoint video data is AVS3 encoding mode, the encapsulation unit 12 is further configured to encapsulate at least one of header information required for decoding, texture map information of at least one viewpoint, and depth map information of at least one viewpoint in the form of a subsample in the media file;
[0368] The subsample data box includes a subsample data box flag and subsample indication information. The subsample data box flag is used to indicate the division method of the subsample, and the subsample indication information is used to indicate the content included in the subsample.
[0369] In some embodiments, if the value of the subsample indication information is a fifth value, it indicates that a subsample includes header information required for decoding; or,
[0370] If the value of the subsample indication information is the sixth value, it indicates that one subsample includes texture map information corresponding to N perspectives in the current video frame, and the current video frame is spliced by video frames corresponding to the N perspectives; or
[0371] If the value of the subsample indication information is the seventh value, it indicates that one subsample includes depth map information corresponding to N viewing angles in the current video frame; or
[0372] If the value of the subsample indication information is the eighth value, it indicates that one subsample includes texture map information and depth map information corresponding to one viewing angle in the current video frame; or
[0373] If the value of the subsample indication information is a ninth value, it indicates that one subsample includes texture map information corresponding to one viewing angle in the current video frame; or
[0374] If the value of the subsample indication information is the tenth value, it indicates that one subsample includes depth map information corresponding to one viewing angle in the current video frame.
[0375] In some embodiments, if the encoding method of the free-viewpoint video data is the AVS3 encoding mode, the video data corresponding to each of the N viewpoints does not rely on the video data corresponding to other viewpoints during encoding, and the encapsulation mode of the free-viewpoint video data code stream is a single-track encapsulation mode, the encapsulation unit 12 is specifically used to encapsulate the texture map information and depth map information corresponding to each of the N viewpoints in the form of subsamples in the media file.
[0376] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Figure 8 The device 10 shown can execute the method embodiment corresponding to the first device, and the aforementioned and other operations and / or functions of each module in the device 8 are respectively for implementing the method embodiment corresponding to the first device, and for the sake of brevity, they are not repeated here.
[0377] Figure 9 This is a schematic diagram of the structure of a file encapsulation device for a free-viewpoint video provided in an embodiment of the present application. The device 20 is applied to a client and includes:
[0378] A receiving unit 21 is configured to receive a media file of free-viewpoint video data sent by a first device, the media file including at least one video track, the free-viewpoint video data including video data of N perspectives, where N is a positive integer, the video track including codec independence indication information and video streams of M perspectives, the codec independence indication information being used to indicate whether video data of a single perspective among the M perspectives corresponding to the video track depends on video data of other perspectives during encoding and decoding, where M is a positive integer less than or equal to N;
[0379] a decapsulation unit 22, configured to decapsulate the media file according to the codec independence indication information to obtain a video stream corresponding to at least one perspective;
[0380] The decoding unit 23 is configured to decode the video code stream corresponding to the at least one viewing angle to obtain reconstructed video data of the at least one viewing angle.
[0381] Optionally, the video data includes at least one of texture map data and depth map data.
[0382] In some embodiments, if the value of the codec independence indication information is a first value, it indicates that the texture map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding, or the depth map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding; or
[0383] If the value of the codec independence indication information is a second value, it indicates that the texture map data of the single view depends on the texture map data of other views during encoding and decoding, and the depth map data of the single view depends on the depth map data of other views during encoding and decoding; or
[0384] If the value of the codec independence indication information is a third value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective depend on each other during encoding and decoding; or
[0385] If the value of the codec independence indication information is the fourth numerical value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective do not depend on each other during encoding and decoding.
[0386] In some embodiments, the decapsulation unit 22 is configured to obtain the user's viewing perspective if the value of the codec independence indication information is the third value or the fourth value, and the encapsulation mode of the code stream is a single-track encapsulation mode; determine a target viewing perspective based on the user's viewing perspective and the viewing perspective information in the media file; and decapsulate the media file corresponding to the target viewing perspective to obtain a video code stream corresponding to the target viewing perspective.
[0387] In some embodiments, if the encoding method of the free-viewpoint video data is the AVS3 video encoding method, and the media file includes subsamples, then the decapsulation unit 22 is used to obtain the subsample data box flag and the subsample indication information included in the data box of the subsample according to the codec independence indication information, where the subsample data box flag is used to indicate the division method of the subsample, and the subsample indication information is used to indicate the content included in the subsample, and the content included in the subsample includes at least one of header information required for decoding, texture map information of at least one viewpoint, and depth map information of at least one viewpoint; obtain the content included in the subsample according to the subsample data box flag and the subsample indication information; and decapsulate the media file resource according to the content included in the subsample and the codec independence indication information to obtain a video stream corresponding to at least one viewpoint.
[0388] In some embodiments, if the value of the subsample indication information is a fifth value, it indicates that a subsample includes header information required for decoding; or,
[0389] If the value of the subsample indication information is the sixth value, it indicates that one subsample includes texture map information corresponding to N viewing angles in the current video frame; or
[0390] If the value of the subsample indication information is the seventh value, it indicates that one subsample includes depth map information corresponding to N viewing angles in the current video frame; or
[0391] If the value of the subsample indication information is the eighth value, it indicates that one subsample includes texture map information and depth map information corresponding to one viewing angle in the current video frame; or
[0392] If the value of the subsample indication information is a ninth value, it indicates that one subsample includes texture map information corresponding to one viewing angle in the current video frame; or
[0393] If the value of the subsample indication information is the tenth value, it indicates that one subsample includes depth map information corresponding to one viewing angle in the current video frame.
[0394] In some embodiments, if the media file includes a subsample corresponding to each of the N perspectives, the decapsulation unit 22 is specifically configured to, if the value of the codec independence indication information is the third value or the fourth value, and the encapsulation mode of the code stream is the single-track encapsulation mode, determine a target perspective according to the user's viewing perspective and the perspective information in the media file; obtain a subsample data box flag and the subsample indication information included in a data box of a target subsample corresponding to the target perspective; determine content included in the target subsample according to the subsample data box flag and the subsample indication information included in the data box of the target subsample; and decapsulate the media file corresponding to the target perspective according to the content included in the target subsample to obtain a video stream corresponding to the target perspective.
[0395] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Figure 9 The device 20 shown can execute the method embodiment corresponding to the client, and the aforementioned and other operations and / or functions of each module in the device 20 are respectively for implementing the method embodiment corresponding to the client, and for the sake of brevity, they are not repeated here.
[0396] Figure 10This is a schematic diagram of the structure of a file encapsulation device for a free-viewpoint video provided in an embodiment of the present application. The device 30 is applied to a server and includes:
[0397] A receiving unit 31 is configured to receive a media file of free-viewpoint video data sent by a first device, the media file including at least one video track, the free-viewpoint video data including video data of N perspectives, where N is a positive integer, the video track including codec independence indication information and video streams of M perspectives, the codec independence indication information being used to indicate whether video data of a single perspective among the M perspectives corresponding to the video track depends on video data of other perspectives during encoding and decoding, where M is a positive integer less than or equal to N;
[0398] The decomposition unit 32 is configured to determine whether to decompose the at least one video track into multiple video tracks according to the codec independence indication information.
[0399] Optionally, the video data includes at least one of texture map data and depth map data.
[0400] In some embodiments, if the value of the codec independence indication information is a first value, it indicates that the texture map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding, or the depth map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding; or
[0401] If the value of the codec independence indication information is a second value, it indicates that the texture map data of the single view depends on the texture map data of other views during encoding and decoding, and the depth map data of the single view depends on the depth map data of other views during encoding and decoding; or
[0402] If the value of the codec independence indication information is a third value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective depend on each other during encoding and decoding; or
[0403] If the value of the codec independence indication information is the fourth numerical value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and / or depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective do not depend on each other during encoding and decoding.
[0404] In some embodiments, if the value of the codec independence indication information is the third value or the fourth value, and the encapsulation mode of the code stream is a single-track encapsulation mode, the decomposition unit 32 is specifically used to encapsulate a video track formed by the single-track encapsulation mode into N video tracks according to a multi-track encapsulation mode, and each of the N video tracks includes video data corresponding to a single perspective.
[0405] In some embodiments, if the encoding method is the AVS3 video encoding method, and the media file includes a subsample corresponding to each of the N perspectives, and the subsample corresponding to each perspective includes at least one of depth map information and texture map information corresponding to each perspective, then the decomposition unit 32 is specifically used to obtain, for each of the N perspectives, a subsample data box flag and the subsample indication information included in the data box of the subsample corresponding to the perspective, the subsample data box flag is used to indicate the division method of the subsample, and the subsample indication information is used to indicate the content included in the subsample; according to the subsample data box flag and the subsample indication information corresponding to each perspective, the content included in the subsample corresponding to each perspective is obtained; according to the content included in the subsample corresponding to each perspective, a video track formed by the single-track encapsulation mode is encapsulated into N video tracks according to the multi-track encapsulation mode.
[0406] In some embodiments, the apparatus further includes a generating unit 33 and a sending unit 34:
[0407] A generating unit 33 is configured to generate a first signaling, wherein the first signaling includes at least one of an identifier of a camera corresponding to each track in the N video tracks, position information of the camera, and focus information of the camera;
[0408] A sending unit 34, configured to send the first signaling to a client;
[0409] The receiving unit 31 is further configured to receive first request information determined by the client according to the first signaling, where the first request information includes identification information of a target camera;
[0410] The sending unit 34 is further configured to send the media file corresponding to the target camera to the client according to the first request information.
[0411] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Figure 10 The device 30 shown can execute the method embodiment corresponding to the server, and the aforementioned and other operations and / or functions of each module in the device 30 are respectively for implementing the method embodiment corresponding to the server. For the sake of brevity, they are not repeated here.
[0412] The apparatus of the embodiment of the present application is described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.
[0413] Figure 11 It is a schematic block diagram of a computing device provided in an embodiment of the present application, and the computing device may be the above-mentioned first device, server or client.
[0414] like Figure 11 As shown, the computing device 40 may include:
[0415] The memory 41 and the memory 42 are configured to store computer programs and transfer the program code to the memory 42. In other words, the memory 42 can call and run the computer program from the memory 41 to implement the method in the embodiment of the present application.
[0416] For example, the memory 42 may be used to execute the above method embodiments according to the instructions in the computer program.
[0417] In some embodiments of the present application, the memory 42 may include but is not limited to:
[0418] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.
[0419] In some embodiments of the present application, the memory 41 includes but is not limited to:
[0420] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).
[0421] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 41 and executed by the memory 42 to implement the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the video production device.
[0422] like Figure 11 As shown, the computing device 40 may further include:
[0423] The transceiver 40 and the transceiver 43 may be connected to the memory 42 or the memory 41 .
[0424] The memory 42 can control the transceiver 43 to communicate with other devices. Specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 43 may include a transmitter and a receiver. The transceiver 43 may further include an antenna, and the number of antennas may be one or more.
[0425] It should be understood that the various components in the video production device are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.
[0426] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment. In other words, the present application also provides a computer program product containing instructions, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment.
[0427] When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0428] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0429] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0430] Modules described as separate components may or may not be physically separate, and components displayed as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected based on actual needs to achieve the purpose of the present embodiment. For example, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module.
[0431] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A file packaging method for free viewpoint video, characterized in that: Applied to a first device, comprising: Obtaining a code stream of free-viewpoint video data, where the free-viewpoint video data includes video data of N viewpoints, where N is a positive integer; Encapsulating the code stream of the free-viewpoint video data into at least one video track to obtain a media file of the free-viewpoint video data, wherein the video track includes codec independence indication information and video code streams of M views, the codec independence indication information being used to indicate whether video data of a single view among the M views corresponding to the video track depends on video data of other views during encoding and decoding, the video data including at least one of texture map data and depth map data, where M is a positive integer less than or equal to N; If the encoding mode of the free-viewpoint video data is the AVS3 encoding mode, encapsulating at least one of header information required for decoding, texture map information of at least one viewpoint, and depth map information of at least one viewpoint in the form of a subsample in the media file, wherein a data box of the subsample includes a subsample data box flag and subsample indication information, wherein the subsample data box flag is used to indicate a division mode of the subsample, and the subsample indication information is used to indicate content included in the subsample, and the content included in the subsample includes: at least one of the header information required for decoding, texture map information of at least one viewpoint, and depth map information of at least one viewpoint, and at least one of the texture map information and the depth map information includes data required for decapsulating a texture map code stream or a depth map code stream; The media file of the free-viewpoint video data is sent to a client or a server.
2. The method according to claim 1, characterized in that If the value of the codec independence indication information is a first value, it indicates that the texture map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding, or that the depth map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding; or If the value of the codec independence indication information is a second value, it indicates that the texture map data of the single view depends on the texture map data of other views during encoding and decoding, and the depth map data of the single view depends on the depth map data of other views during encoding and decoding; or If the value of the codec independence indication information is a third value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective depend on each other during encoding and decoding; or If the value of the codec independence indication information is the fourth numerical value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective do not depend on each other during encoding and decoding.
3. The method according to claim 2, characterized in that Encapsulating the code stream of the free-viewpoint video data into at least one video track includes: A single-track encapsulation mode is adopted to encapsulate the code stream of the free-viewing-angle video data into a video track.
4. The method according to claim 3, characterized in that If the video data corresponding to each of the N perspectives does not depend on the video data corresponding to other perspectives during encoding, and the encapsulation mode of the code stream of the free-perspective video data is a single-track encapsulation mode, the method further includes: The codec independence indication information is added to the free view information data box of the one video track, and the value of the codec independence indication information is the third value or the fourth value.
5. The method according to any one of claims 1 to 4, characterized in that If the value of the subsample indication information is the fifth value, it indicates that a subsample includes header information required for decoding; or, If the value of the subsample indication information is the sixth value, it indicates that one subsample includes texture map information of N perspectives in the current video frame, and the current video frame is spliced by video frames corresponding to the N perspectives; or If the value of the subsample indication information is the seventh value, it indicates that one subsample includes depth map information corresponding to N perspectives in the current video frame; or, If the value of the subsample indication information is the eighth value, it indicates that one subsample includes texture map information and depth map information corresponding to one viewing angle in the current video frame; or If the value of the subsample indication information is a ninth value, it indicates that one subsample includes texture map information corresponding to one viewing angle in the current video frame; or, If the value of the subsample indication information is the tenth value, it indicates that one subsample includes depth map information corresponding to one viewing angle in the current video frame.
6. The method according to claim 5, characterized in that If the encoding mode of the free-viewpoint video data is the AVS3 encoding mode, the video data corresponding to each of the N viewpoints does not depend on the video data corresponding to other viewpoints during encoding, and the encapsulation mode of the free-viewpoint video data stream is the single-track encapsulation mode, then encapsulating at least one of the header information required for decoding, the texture map information of at least one viewpoint, and the depth map information of at least one viewpoint in the form of a subsample in the media file includes: At least one of the texture map information and the depth map information corresponding to each of the N perspectives is encapsulated in the media file in the form of a subsample.
7. A file packaging method for free viewpoint video, characterized in that: Applied to the client, including: Receive a media file of free-viewpoint video data sent by a first device, the media file comprising at least one video track, the free-viewpoint video data comprising video data of N views, where N is a positive integer, the video track comprising codec independence indication information and video streams of M views, the codec independence indication information being used to indicate whether video data of a single view among the M views corresponding to the video track depends on video data of other views during encoding and decoding, the video data comprising at least one of texture map data and depth map data, where M is a positive integer less than or equal to N, and if the free-viewpoint video data is encoded in an AVS3 encoding mode, the media file further comprising subsamples, a data box of the subsample comprising a subsample data box flag and subsample indication information, the subsample data box flag being used to indicate a division method of the subsample, and the subsample indication information being used to indicate content included in the subsample, the subsample content comprising at least one of header information required for decoding, texture map information of at least one view, and depth map information of at least one view, at least one of the texture map information and the depth map information comprising data required for decapsulating a texture map stream or a depth map stream; Decapsulating the media file according to the codec independence indication information to obtain a video stream corresponding to at least one perspective; The video code stream corresponding to the at least one viewing angle is decoded to obtain reconstructed video data of the at least one viewing angle.
8. The method according to claim 7, characterized in that If the value of the codec independence indication information is a first value, it indicates that the texture map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding, or that the depth map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding; or If the value of the codec independence indication information is a second value, it indicates that the texture map data of the single view depends on the texture map data of other views during encoding and decoding, and the depth map data of the single view depends on the depth map data of other views during encoding and decoding; or If the value of the codec independence indication information is a third value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective depend on each other during encoding and decoding; or If the value of the codec independence indication information is the fourth numerical value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective do not depend on each other during encoding and decoding.
9. The method according to claim 8, characterized in that Decapsulating the media file resource according to the codec independence indication information to obtain a video stream corresponding to at least one perspective includes: If the value of the codec independence indication information is the third value or the fourth value, and the encapsulation mode of the code stream is the single-track encapsulation mode, obtaining the user's viewing angle; determining a target viewing angle according to the user's viewing angle and viewing angle information in the media file; The media file corresponding to the target viewing angle is decapsulated to obtain a video stream corresponding to the target viewing angle.
10. The method according to claim 8 or 9, characterized in that If the encoding mode of the free-viewpoint video data is the AVS3 video encoding mode, then decapsulating the media file resource according to the codec independence indication information to obtain a video stream corresponding to at least one viewpoint includes: Obtaining a subsample data box flag and the subsample indication information included in the subsample data box; Obtaining content included in the subsample according to the subsample data box flag and the subsample indication information; The media file resource is decapsulated according to the content included in the subsample and the codec independence indication information to obtain a video stream corresponding to at least one perspective.
11. The method according to claim 10, characterized in that If the value of the subsample indication information is the fifth value, it indicates that a subsample includes header information required for decoding; or, If the value of the subsample indication information is the sixth value, it indicates that one subsample includes texture map information of N perspectives in the current video frame, and the current video frame is spliced by video frames corresponding to the N perspectives; or If the value of the subsample indication information is the seventh value, it indicates that one subsample includes depth map information corresponding to N perspectives in the current video frame; or, If the value of the subsample indication information is the eighth value, it indicates that one subsample includes texture map information and depth map information corresponding to one viewing angle in the current video frame; or If the value of the subsample indication information is a ninth value, it indicates that one subsample includes texture map information corresponding to one viewing angle in the current video frame; or, If the value of the subsample indication information is the tenth value, it indicates that one subsample includes depth map information corresponding to one viewing angle in the current video frame.
12. The method according to claim 11, characterized in that If the media file includes a subsample corresponding to each of the N perspectives, decapsulating the media file resource according to the content included in the subsample and the codec independence indication information to obtain a video stream corresponding to at least one perspective includes: If the value of the codec independence indication information is the third value or the fourth value, and the encapsulation mode of the code stream is the single-track encapsulation mode, determining the target viewing angle according to the user's viewing angle and the viewing angle information in the media file; Acquire a subsample data box flag and the subsample indication information included in a data box of a target subsample corresponding to the target perspective; Determining the content of the target subsample according to the subsample data box flag included in the data box of the target subsample and the subsample indication information; According to the content included in the target subsample, the media file corresponding to the target perspective is decapsulated to obtain a video stream corresponding to the target perspective.
13. A file packaging method for free viewpoint video, characterized in that: Applicable to servers, including: Receive a media file of free-viewpoint video data sent by a first device, the media file comprising at least one video track, the free-viewpoint video data comprising video data of N views, where N is a positive integer, the video track comprising codec independence indication information and video streams of M views, the codec independence indication information being used to indicate whether video data of a single view among the M views corresponding to the video track depends on video data of other views during encoding and decoding, the video data comprising at least one of texture map data and depth map data, where M is a positive integer less than or equal to N, and if the free-viewpoint video data is encoded in an AVS3 encoding mode, the media file further comprising subsamples, a data box of the subsample comprising a subsample data box flag and subsample indication information, the subsample data box flag being used to indicate a division method of the subsample, and the subsample indication information being used to indicate content included in the subsample, the subsample content comprising at least one of header information required for decoding, texture map information of at least one view, and depth map information of at least one view, at least one of the texture map information and the depth map information comprising data required for decapsulating a texture map stream or a depth map stream; Determine whether to decompose the at least one video track into multiple video tracks according to the codec independence indication information.
14. The method according to claim 13, characterized in that If the value of the codec independence indication information is a first value, it indicates that the texture map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding, or that the depth map data of the single perspective depends on the texture map data and depth map data of other perspectives during encoding and decoding; or If the value of the codec independence indication information is a second value, it indicates that the texture map data of the single view depends on the texture map data of other views during encoding and decoding, and the depth map data of the single view depends on the depth map data of other views during encoding and decoding; or If the value of the codec independence indication information is a third value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective depend on each other during encoding and decoding; or If the value of the codec independence indication information is the fourth numerical value, it indicates that the texture map data and depth map data of the single perspective do not depend on the texture map data and depth map data of other perspectives during encoding and decoding, and the texture map data and depth map data of the single perspective do not depend on each other during encoding and decoding.
15. The method according to claim 14, characterized in that The determining, according to the codec independence indication information, whether to decompose the at least one video track into multiple video tracks includes: If the value of the codec independence indication information is the third value or the fourth value, and the encapsulation mode of the code stream is the single-track encapsulation mode, then a video track formed by the single-track encapsulation mode is encapsulated into N video tracks according to the multi-track encapsulation mode, and each of the N video tracks includes video data corresponding to a single perspective.
16. The method according to claim 15, characterized in that If the encoding mode of the free-viewpoint video data is the AVS3 video encoding mode, and the media file includes a subsample corresponding to each of the N viewpoints, and the subsample corresponding to each viewpoint includes at least one of depth map information and texture map information corresponding to each viewpoint, encapsulating a video track formed by the single-track encapsulation mode into N video tracks according to the multi-track encapsulation mode includes: For each of the N viewing angles, obtaining a subsample data box flag and the subsample indication information included in a data box of a subsample corresponding to the viewing angle; Obtaining the content of the subsample corresponding to each viewing angle according to the subsample data box flag corresponding to each viewing angle and the subsample indication information; According to the content included in the subsample corresponding to each viewing angle, a video track formed by the single-track encapsulation mode is encapsulated into N video tracks according to the multi-track encapsulation mode.
17. The method according to any one of claims 13 to 16, characterized in that: The method further comprises: Generate a first signaling, the first signaling including at least one of an identifier of a camera corresponding to each track of the N video tracks, location information of the camera, and focus information of the camera; Sending the first signaling to the client; receiving first request information determined by the client according to the first signaling, where the first request information includes identification information of a target camera; The media file corresponding to the target camera is sent to the client according to the first request information.
18. A device for processing multi-view video data, characterized in that: Applied to a first device, the apparatus includes: an acquiring unit, configured to acquire a code stream of free-viewpoint video data, wherein the free-viewpoint video data includes video data of N viewpoints, where N is a positive integer; an encapsulation unit, configured to encapsulate the code stream of the free-viewpoint video data into at least one video track to obtain a media file of the free-viewpoint video data, wherein the video track includes codec independence indication information and video code streams of M views, the codec independence indication information being used to indicate whether video data of a single view among the M views corresponding to the video track depends on video data of other views during encoding and decoding, the video data including at least one of texture map data and depth map data, where M is a positive integer less than or equal to N; The encapsulation unit is further configured to, if the encoding mode of the free-viewpoint video data is the AVS3 encoding mode, encapsulate at least one of header information required for decoding, texture map information of at least one viewpoint, and depth map information of at least one viewpoint in the form of a subsample in the media file, wherein a data box of the subsample includes a subsample data box flag and subsample indication information, the subsample data box flag is used to indicate a division mode of the subsample, and the subsample indication information is used to indicate content included in the subsample, and the content included in the subsample includes: the header information required for decoding, at least one of the texture map information of the at least one viewpoint, and the depth map information of the at least one viewpoint, and at least one of the texture map information and the depth map information includes data required for decapsulating a texture map code stream or a depth map code stream; The sending unit is configured to send the media file of the free-viewpoint video data to a client or a server.
19. A device for processing multi-view video data, characterized in that: Applied to a client, the device includes: a receiving unit, configured to receive a media file of free-viewpoint video data sent by a first device, the media file comprising at least one video track, the free-viewpoint video data comprising video data of N views, where N is a positive integer, the video track comprising codec independence indication information and video streams of M views, the codec independence indication information being used to indicate whether video data of a single view among the M views corresponding to the video track depends on video data of other views during encoding and decoding, the video data comprising at least one of texture map data and depth map data, where M is a positive integer less than or equal to N, and if the free-viewpoint video data is encoded in the AVS3 encoding mode, the media file further comprising subsamples, a data box of the subsample comprising a subsample data box flag and subsample indication information, the subsample data box flag being used to indicate a division method of the subsample, and the subsample indication information being used to indicate content included in the subsample, the subsample content comprising at least one of header information required for decoding, texture map information of at least one view, and depth map information of at least one view, at least one of the texture map information and the depth map information comprising data required for decapsulating the texture map stream or the depth map stream; a decapsulation unit, configured to decapsulate the media file according to the codec independence indication information to obtain a video stream corresponding to at least one viewing angle; A decoding unit is used to decode the video code stream corresponding to the at least one perspective to obtain reconstructed video data of the at least one perspective.
20. A device for processing multi-view video data, characterized in that: Applied to a server, the device includes: a receiving unit, configured to receive a media file of free-viewpoint video data sent by a first device, the media file comprising at least one video track, the free-viewpoint video data comprising video data of N views, where N is a positive integer, the video track comprising codec independence indication information and video streams of M views, the codec independence indication information being used to indicate whether video data of a single view among the M views corresponding to the video track depends on video data of other views during encoding and decoding, the video data comprising at least one of texture map data and depth map data, where M is a positive integer less than or equal to N, and if the free-viewpoint video data is encoded in the AVS3 encoding mode, the media file further comprising subsamples, a data box of the subsample comprising a subsample data box flag and subsample indication information, the subsample data box flag being used to indicate a division method of the subsample, and the subsample indication information being used to indicate content included in the subsample, the subsample content comprising at least one of header information required for decoding, texture map information of at least one view, and depth map information of at least one view, at least one of the texture map information and the depth map information comprising data required for decapsulating the texture map stream or the depth map stream; The decomposition unit is configured to determine whether to decompose the at least one video track into multiple video tracks according to the codec independence indication information.
21. A computing device, characterized in that include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 6 or 7 to 12 or 13 to 17.
22. A computer-readable storage medium, characterized in that Used to store a computer program, the computer program causing a computer to execute the method according to any one of claims 1 to 6 or 7 to 12 or 13 to 17.
Citation Information
Patent Citations
Signaling of spatial resolution of depth views in multiview coding file format
US20200336726A1