File packaging method, device and equipment for free view angle video and storage medium
By adding codec independence indication information to the video track, the problem of not being able to determine whether certain viewpoints can be decoded in single-track encapsulation mode is solved, achieving more efficient media file decoding and processing.
Patent Information
- Application Number
- CN202511286418.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-10
- Publication Date
- 2025-11-21
AI Technical Summary
Existing video stream encapsulation methods cannot determine whether media files corresponding to certain viewpoints can be decoded in single-track encapsulation mode, resulting in low media file decoding efficiency.
Add codec independence information to the video track to indicate whether the video data of a single viewpoint in the video track depends on the video data of other viewpoints during encoding and decoding, so that the client and server can determine whether partial decoding or re-encapsulation is possible in single-track encapsulation mode.
It improves the decoding efficiency and processing flexibility of media files. The client can partially decode the texture map and depth map of a specific camera, and the server can decide whether to repackage the single-track free-view video into a multi-track format.
Smart Images

Figure CN121000899A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention application with the application date of August 10, 2021, the Chinese application number of 202110913912.8, and the invention name of "file packaging method, device and equipment and storage medium of free-view video". TECHNICAL FIELD
[0002] Embodiments of the present application relate to the technical field of video processing, in particular to a file packaging method, device, equipment and storage medium of free-view video. BACKGROUND
[0003] Immersive media refers to media content that can bring an immersive experience to consumers. Immersive media can be divided into 3DoF media, 3DoF+ media and 6DoF media according to the degree of freedom of users when consuming media content.
[0004] However, the current video stream packaging method cannot determine whether the media file of the free-view video packaged by the single-track packaging mode can be decoded by the server or the client. The decoding efficiency of the media file is low. SUMMARY
[0005] The present application provides a file packaging method, device, equipment and storage medium of free-view video, which can determine whether the media file corresponding to part of the view angle in the media file can be decoded by the server or the client, thereby improving the decoding efficiency of the media file.
[0006] In a first aspect, the present application provides a file packaging method of free-view video, applied to a first device, which can be understood as a video packaging device. The method comprises:
[0007] Obtaining a code stream of free-view video data, wherein the free-view video data comprises video data of N views, and N is a positive integer;
[0008] Packaging the code stream of the free-view video data into at least one video track to obtain a media file of the free-view video data, wherein the video track comprises codec independence indication information and video code streams of M views, the codec independence indication information is used to indicate whether the video data of a single view in the M views corresponding to the video track depends on the video data of other views when being coded, and M is a positive integer less than or equal to N;
[0009] Sending the media file of the free-view video data to a client or a server.
[0010] In a second aspect, the present application provides a file packaging method of free-view video, applied to a client, which can be understood as a video playing device. The method comprises:
[0011] receiving a media file of free-viewpoint video data sent by a first device, the media file comprising at least one video track, the free-viewpoint video data comprising video data of N viewpoints, N being a positive integer, the video track comprising codec dependency indication information and video bitstreams of M viewpoints, the codec dependency indication information being used to indicate whether video data of a single viewpoint in the M viewpoints corresponding to the video track depends on video data of other viewpoints when being coded, M being a positive integer less than or equal to N;
[0012] decapsulating the media file according to the codec dependency indication information to obtain video bitstreams corresponding to at least one viewpoint;
[0013] decoding the video bitstreams corresponding to the at least one viewpoint to obtain reconstructed video data of the at least one viewpoint.
[0014] In a third aspect, the present application provides a file encapsulation method of free-viewpoint video, applied to a server, the method comprising:
[0015] receiving a media file of free-viewpoint video data sent by a first device, the media file comprising at least one video track, the free-viewpoint video data comprising video data of N viewpoints, N being a positive integer, the video track comprising codec dependency indication information and video bitstreams of M viewpoints, the codec dependency indication information being used to indicate whether video data of a single viewpoint in the M viewpoints corresponding to the video track depends on video data of other viewpoints when being coded, M being a positive integer less than or equal to N;
[0016] determining whether to decompose the at least one video track into multiple video tracks according to the codec dependency indication information.
[0017] In a fourth aspect, the present application provides a processing apparatus of multi-viewpoint video data, applied to a first device, the apparatus comprising:
[0018] an obtaining unit, configured to obtain a code stream of free-viewpoint video data, the free-viewpoint video data comprising video data of N viewpoints, N being a positive integer;
[0019] an encapsulating unit, configured to encapsulate the code stream of the free-viewpoint video data into at least one video track to obtain a media file of the free-viewpoint video data, the video track comprising codec dependency indication information and video bitstreams of M viewpoints, the codec dependency indication information being used to indicate whether video data of a single viewpoint in the M viewpoints corresponding to the video track depends on video data of other viewpoints when being coded, M being a positive integer less than or equal to N;
[0020] The sending unit is configured to send the media file of the free-view video data to a client or a server.
[0021] In a fifth aspect, the present application provides a processing device for multi-view video data, applied to a client, comprising:
[0022] The receiving unit is configured to receive a media file of free-view video data sent by a first device, wherein the media file comprises at least one video track, the free-view video data comprises video data of N views, N is a positive integer, the video track comprises codec independence indication information and video bitstream of M views, the codec independence indication information is used to indicate whether video data of a single view in the M views corresponding to the video track depends on video data of other views during codec, and M is a positive integer less than or equal to N;
[0023] The decapsulation unit is configured to decapsulate the media file according to the codec independence indication information, to obtain video bitstream corresponding to at least one view.
[0024] The decoding unit is configured to decode the video bitstream corresponding to the at least one view, to obtain reconstructed video data of the at least one view.
[0025] In a sixth aspect, a computing device is provided, comprising a processor and a memory, the memory is used to store a computer program, and the processor is used to invoke and run the computer program stored in the memory, to execute the method of the first aspect and / or the second aspect and / or the third aspect.
[0026] In a seventh aspect, a computer readable storage medium is provided, used to store a computer program, the computer program makes a computer execute the method of the first aspect and / or the second aspect and / or the third aspect.
[0027] In summary, in the present application, by adding codec independence indication information in the video track, the codec independence indication information is used to indicate whether video data of a single view in M views corresponding to the video track depends on video data of other views during codec, so that in the single-track encapsulation mode, the client can determine whether the texture map and the depth map of a specific camera can be partially decoded according to the codec independence indication information. In addition, in the single-track encapsulation mode, the server can also determine whether the single-track encapsulated free-view video can be re-encapsulated according to the multi-track according to the codec independence indication information, thereby improving the processing flexibility of the media file and improving the decoding efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiments description. Obviously, the drawings described in the following embodiments are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0029] Figure 1 The schematic diagram of three degrees of freedom is shown schematically.
[0030] Figure 2 The schematic diagram of three degrees of freedom + is shown schematically.
[0031] Figure 3 The schematic diagram of six degrees of freedom is shown schematically.
[0032] Figure 4 The architecture diagram of an immersive media system provided by an embodiment of the present application is shown.
[0033] Figure 5 The flowchart of a file packaging method of a free-view video provided by an embodiment of the present application is shown.
[0034] Figure 6 The interactive flowchart of a file packaging method of a free-view video provided by an embodiment of the present application is shown.
[0035] Figure 7 The interactive flowchart of a file packaging method of a free-view video provided by an embodiment of the present application is shown.
[0036] Figure 8 The structural schematic diagram of a file packaging device of a free-view video provided by an embodiment of the present application is shown.
[0037] Figure 9 The structural schematic diagram of a file packaging device of a free-view video provided by an embodiment of the present application is shown.
[0038] Figure 10 The structural schematic diagram of a file packaging device of a free-view video provided by an embodiment of the present application is shown.
[0039] Figure 11 The schematic block diagram of a computing device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0040] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0041] It should be noted that the terms "first", "second", and the like in the description and in the claims of the present application and above-described accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular sequential or chronological order. It should be understood that the data thus used can be interchanged, where appropriate, so that the embodiments of the application described herein can be carried out in other than the order shown or described herein. Furthermore, the terms "comprising" and "having", and any variations thereof, are intended to cover a non-exclusive inclusion, for example, a process, method, system, product, or server that comprises a list of steps or units not necessarily limited to those explicitly listed, but can include other steps or units not expressly listed or inherent to such processes, methods, products, or apparatuses.
[0042] Embodiments of the present application relate to data processing technology of immersive media.
[0043] Before introducing the technical solutions of the present application, the following first introduces the related knowledge of the present application:
[0044] Free-viewpoint video: immersive media video captured by multiple cameras, containing different viewpoints, supporting 3DoF+ or 6DoF interaction of users. Also known as multi-view video.
[0045] Track: track, a collection of media data in the media file packaging process, a media file can be composed of multiple tracks, such as a media file can contain a video track, an audio track and a subtitle track.
[0046] Sample: sample, packaging unit in the media file packaging process, a media track is composed of many samples. For example, a sample of a video track is usually a video frame.
[0047] DoF: Degree of Freedom, degree of freedom. In mechanical systems, it refers to the number of independent coordinates, in addition to the degree of freedom of translation, there are rotational and vibrational degrees of freedom. In the embodiments of the present application, it refers to the degree of freedom of motion supported by the user when watching immersive media and producing content interaction.
[0048] 3DoF: three degrees of freedom, refers to three degrees of freedom of user head rotation around XYZ axes. Figure 1 A schematic diagram of three degrees of freedom is schematically shown. For example, Figure 1As shown, it can be rotated in three axes at a certain place or point, that is, it can turn its head, lower its head up and down, and shake its head. Through the experience of three degrees of freedom, the user can be immersed in a live scene by 360 degrees. If it is static, it can be understood as a panoramic picture. If the panoramic picture is dynamic, it is a panoramic video, that is, a VR video. However, the VR video has certain limitations, and the user cannot move and cannot choose any place to watch.
[0049] 3DoF+: that is, on the basis of three degrees of freedom, the user also has the freedom of limited motion along the XYZ axis, which can also be called limited six degrees of freedom, and the corresponding media code stream can be called limited six degrees of freedom media code stream. Figure 2 The schematic diagram of three degrees of freedom+ is schematically shown.
[0050] 6DoF: that is, on the basis of three degrees of freedom, the user also has the freedom of motion along the XYZ axis, and the corresponding media code stream can be called six degrees of freedom media code stream. Figure 3 The schematic diagram of six degrees of freedom is schematically shown. The 6DoF media refers to 6 degrees of freedom video, which means that the video can provide a high degree of freedom viewing experience for the user to freely move the viewpoint in the XYZ axis direction of the three-dimensional space and freely rotate the viewpoint around the XYZ axis. The 6DoF media is a combination of videos of different perspectives in space collected by an array of cameras. In order to facilitate the expression, storage, compression and processing of 6DoF media, the 6DoF media data is expressed as a combination of the following information: texture map collected by multiple cameras, depth map corresponding to the multiple camera texture map, and corresponding 6DoF media content description metadata, which includes parameters of the multiple cameras, and description information of stitching layout and edge protection of the 6DoF media. At the encoding end, the texture map information and the corresponding depth map information of the multiple cameras are stitched, and the description data of the stitching method is written into the metadata according to the defined syntax and semantics. The stitched multiple camera depth map and texture map information are encoded by a planar video compression method, and are transmitted to the terminal for decoding, and then the synthesis of the 6DoF virtual viewpoint requested by the user is performed, thereby providing the user with the viewing experience of the 6DoF media.
[0051] Depth map: as a three-dimensional scene information expression method, the gray value of each pixel point of the depth map can be used to represent the distance of a certain point in the scene from the camera.
[0052] AVS: Audio Video Coding Standard, audio and video coding standard.
[0053] AVS3: the third generation audio and video coding standard promoted by the AVS working group
[0054] ISOBMFF: ISO Based Media File Format, media file format based on ISO (International Standard Organization) standard. ISOBMFF is the encapsulation standard of media file, and the most typical ISOBMFF file is MP4 (Moving Picture Experts Group 4) file.
[0055] DASH: dynamic adaptive streaming over HTTP, dynamic adaptive streaming over HTTP is an adaptive bit rate streaming technology that enables high-quality streaming media to be delivered over the Internet through traditional HTTP web servers.
[0056] MPD: media presentation description, media presentation description signaling in DASH, used to describe media segment information.
[0057] Representation: in DASH, a combination of one or more media components, such as a video file of a certain resolution can be regarded as a Representation.
[0058] Adaptation Sets: in DASH, a collection of one or more video streams, and an Adaptation Sets can contain multiple Representations.
[0059] HEVC: High Efficiency Video Coding, international video coding standard HEVC / H.265.
[0060] VVC: versatile video coding, international video coding standard VVC / H.266.
[0061] Intra(picture)Prediction: Intra prediction.
[0062] Inter(picture)Prediction: Inter prediction.
[0063] SCC: screen content coding, screen content coding.
[0064] Immersive media refers to media content that provides consumers with an immersive experience. Based on the degree of freedom users have when consuming media content, immersive media can be divided into 3DoF media, 3DoF+ media, and 6DoF media. Common 6DoF media include multi-view video and point cloud media.
[0065] Free-viewpoint video is typically captured by an array of cameras from multiple angles of the same 3D scene, forming texture information (color information, etc.) and depth information (spatial distance information, etc.) of the scene. Based on the user's location information, combined with the texture and depth information from different cameras, 6DoF media can be consumed by the user.
[0066] After free-viewpoint video is captured, it needs to be compressed and encoded. In existing free-viewpoint video technologies, video compression algorithms can be used, such as AVS3 encoding technology and HEVC encoding technology.
[0067] Figure 4 This is an architectural diagram of an immersive media system provided in one embodiment of this application. Figure 4 As shown, an immersive media system includes encoding and decoding devices. Encoding devices can refer to the computer equipment used by the immersive media provider, which can be a terminal (such as a PC, a smart mobile device, or a smartphone) or a server. Decoding devices can refer to the computer equipment used by the immersive media user, which can be a terminal (such as a PC, a smart mobile device, or a VR device, such as a VR headset or VR glasses). The data processing of immersive media includes data processing on the encoding device side and data processing on the decoding device side.
[0068] The data processing at the encoding device mainly includes:
[0069] (1) The process of acquiring and producing immersive media content;
[0070] (2) The encoding and file encapsulation process of immersive media. The data processing at the decoding device mainly includes:
[0071] (3) The process of decapsulating and decoding immersive media files;
[0072] (4) The rendering process of immersive media.
[0073] In addition, the transmission process between the encoding device and the decoding device involves the immersive media, which can be based on various transmission protocols, including but not limited to: DASH (Dynamic Adaptive Streaming over HTTP) protocol, HLS (HTTP Live Streaming) protocol, SMTP (Smart Media Transport Protocaol) protocol, TCP (Transmission Control Protocol) and the like.
[0074] The following will be combined with Figure 4 , respectively, the various processes involved in the data processing process of immersive media are described in detail.
[0075] I. Data processing process at the encoding device end:
[0076] (1) The acquisition and production process of the media content of the immersive media.
[0077] 1) The acquisition process of the media content of the immersive media.
[0078] The media content of the immersive media is obtained by capturing the sound-visual scene of the real world with a capturing device.
[0079] In one implementation, the capturing device can be a hardware component in the encoding device, such as a microphone, camera, sensor, etc. in the terminal. In another implementation, the capturing device can also be a hardware device connected to the encoding device, such as a camera connected to the server.
[0080] The capturing device can include but is not limited to: audio devices, camera devices, and sensor devices. Among them, the audio device can include audio sensors, microphones, etc. The camera device can include ordinary cameras, stereo cameras, light field cameras, etc. The sensor device can include laser devices, radar devices, etc.
[0081] The number of capturing devices can be multiple, and these capturing devices are deployed at some specific positions in the real space to capture audio content and video content at different angles in the space at the same time, and the captured audio content and video content are kept synchronized in time and space. The media content collected by the capturing device is called the original data of the immersive media.
[0082] 2) The production process of the media content of the immersive media.
[0083] The captured audio content itself is suitable for being encoded as audio of immersive media. The captured video content is suitable for being encoded as video of immersive media after a series of production processes, including:
[0084] ①stitching. Since the captured video content is captured by the capturing device at different angles, stitching refers to stitching the video content captured at each angle into a complete video that can reflect a 360-degree visual panorama of the real space, i.e., the stitched video is a panorama video (or a spherical video) represented in three-dimensional space.
[0085] ②projection. Projection refers to the process of mapping a three-dimensional video formed by stitching to a two-dimensional (3-Dimension, 2D) image. The 2D image formed by projection is called a projection image. The projection method can include but is not limited to: equirectangular projection, cube map projection, and hexahedron projection.
[0086] ③region packing. The projection image can be directly encoded, or the projection image can be region packed before being encoded. In practice, it is found that region packing the two-dimensional projection image before encoding can greatly improve the video encoding efficiency of immersive media, so the region packing technology is widely used in the video processing process of immersive media. Region packing refers to the process of converting the projection image by region. The region packing process converts the projection image into a packed image. The process of region packing specifically includes: dividing the projection image into a plurality of mapping regions, then converting the plurality of mapping regions to obtain a plurality of packing regions, and mapping the plurality of packing regions to a 2D image to obtain a packed image. The mapping region refers to the region divided in the projection image before region packing. The packing region refers to the region in the packed image after region packing.
[0087] The conversion process can include but is not limited to: mirroring, rotating, rearranging, upsampling, downsampling, changing the resolution of the region, and moving.
[0088] It should be noted that, since the capture device can only capture panoramic video, after the video is processed by the encoding device and transmitted to the decoding device for corresponding data processing, the user at the decoding device side can only watch 360-degree video information by performing some specific actions (such as head rotation), and performing non-specific actions (such as moving the head) cannot obtain the corresponding video changes, the VR experience is not good, and therefore it is necessary to additionally provide depth information matching the panoramic video to enable the user to obtain better immersion and better VR experience, which involves 6DoF (Six Degrees of Freedom, six degrees of freedom) production technology. When the user can move more freely in the simulated scene, it is called 6DoF. When using 6DoF production technology to produce video content of immersive media, the capture device generally selects a light field camera, a laser device, a radar device, etc., to capture point cloud data or light field data in the space, and some specific processing is also required during the execution of the above production processes ①-③, such as cutting and mapping of point cloud data, depth information calculation process, etc.
[0089] (2) Process of encoding and file packaging of immersive media.
[0090] The captured audio content can be directly audio encoded to form an audio bitstream of the immersive media. After the above production processes ①-② or ①-③, the projection image or the packaged image is video encoded to obtain a video bitstream of the immersive media. It should be noted that, if the 6DoF production technology is used, a specific encoding method (such as point cloud encoding) is required during video encoding. The audio bitstream and the video bitstream are packaged in a file container according to the file format of the immersive media (such as ISOBMFF (ISO Base Media File Format, ISO base media file format)) to form a media file resource of the immersive media, which can be a media file or a media segment forming a media file of the immersive media; and the metadata of the media file resource of the immersive media is recorded according to the media presentation description information (MPD) required by the file format of the immersive media. The metadata here is a general term for information related to the presentation of the immersive media, which can include description information of the media content, description information of the viewport, and signaling information related to the presentation of the media content, etc. As shown in the following table, the encoding device stores the media presentation description information and the media file resource formed after the data processing process. Figure 1
[0091] The immersive media system supports data boxes, which refer to data blocks or objects including metadata, i.e., the metadata of corresponding media content is contained in the data boxes. The immersive media can include multiple data boxes, such as a sphere region zooming box including metadata for describing sphere region zooming information, a 2D region zooming box including metadata for describing 2D region zooming information, a region wise packing box including metadata for describing corresponding information in a region packing process, and the like.
[0092] II. Data processing process at the decoding device end:
[0093] (3) File unpacking and decoding process of the immersive media;
[0094] The decoding device can obtain the media file resources and corresponding media presentation description information of the immersive media from the encoding device through recommendation of the encoding device or adaptive and dynamic request according to user demand at the decoding device end. For example, the decoding device can determine the orientation and position of the user according to tracking information of the head / eyes / body of the user, and then dynamically request the encoding device to obtain corresponding media file resources based on the determined orientation and position. The media file resources and the media presentation description information are transmitted from the encoding device to the decoding device through a transmission mechanism (such as DASH or SMT). The file unpacking process at the decoding device end is inverse to the file packing process at the encoding device end. The decoding device unpacks the media file resources according to the file format requirements of the immersive media to obtain audio bitstreams and video bitstreams. The decoding process at the decoding device end is inverse to the encoding process at the encoding device end. The decoding device decodes the audio bitstreams to restore the audio content.
[0095] In addition, the decoding process of the video bitstreams by the decoding device includes the following:
[0096] ① Decoding the video bitstreams to obtain planar images; according to the metadata provided by the media presentation description information, if the metadata indicates that the immersive media has performed a region packing process, the planar images refer to packed images; if the metadata indicates that the immersive media has not performed a region packing process, the planar images refer to projection images;
[0097] If the metadata indicates that the immersive media has performed the region packing process, the decoding device region unpacks the packed image to obtain the projection image. Here, the region unpacking is inverse to the region packing, and the region unpacking refers to a process of performing inverse conversion processing on the packed image according to regions. The region unpacking converts the packed image into the projection image. The process of the region unpacking specifically includes: performing inverse conversion processing on the plurality of packed regions in the packed image according to the indication of the metadata to obtain a plurality of mapping regions, and mapping the plurality of mapping regions to a 2D image to obtain the projection image. The inverse conversion processing refers to processing inverse to the conversion processing. For example, the conversion processing refers to counterclockwise rotation by 90 degrees, and the inverse conversion processing refers to clockwise rotation by 90 degrees.
[0098] 3. Reconstructing the projection image according to the media presentation description information to convert the projection image into a 3D image. Here, the reconstructing refers to a process of re-projecting the two-dimensional projection image into a 3D space.
[0099] (4) Rendering process of the immersive media.
[0100] The decoding device renders the audio content obtained by audio decoding and the 3D image obtained by video decoding according to the metadata related to rendering and the viewport in the media presentation description information. The rendering is completed, and the playback output of the 3D image is realized. In particular, if the production technology of 3DoF and 3DoF+ is adopted, the decoding device mainly renders the 3D image based on the current viewpoint, parallax, depth information, and the like. If the production technology of 6DoF is adopted, the decoding device mainly renders the 3D image in the viewport based on the current viewpoint. The viewpoint refers to the viewing position of the user, the parallax refers to the difference in the line of sight of the user's two eyes or the difference in the line of sight due to movement, and the viewport refers to the viewing area.
[0101] The immersive media system supports a data box, which refers to a data block or object including metadata, that is, the data box includes the metadata of the corresponding media content. The immersive media can include a plurality of data boxes, such as a sphere region zooming box including metadata for describing sphere region zooming information, a 2D region zooming box including metadata for describing 2D region zooming information, a region wise packing box including metadata for describing corresponding information in the region packing process, and the like.
[0102] In some embodiments, for packing of the free-view video, the following file packing mode is proposed:
[0103] 1. Free-view angle track group
[0104] If a free view video is encapsulated as multiple video tracks, these video tracks shall be associated through a free view group box, which is defined as follows:
[0105] aligned(8) class AvsFreeViewGroupBox extends TrackGroupTypeBox('afvg')
[0106] {
[0107] / / track_group_id is inherited from TrackGroupTypeBox;
[0108] unsigned int(8) camera_count;
[0109]
[0110] A free view group box is derived from extending the track group data box with the track group type identifier 'a3fg'. In all tracks containing a TrackGroupTypeBox of type 'afvg', tracks with the same group ID belong to the same track group. The semantics of the fields in AvsFreeViewGroupBox are as follows:
[0111] camera_count: indicates the number of cameras whose texture information or depth information are contained in the track.
[0112] camera_id: indicates the camera identifier corresponding to each camera, which corresponds to the value in the AvsFreeViewInfoBox in the current track.
[0113] depth_texture_type: indicates the type of texture information or depth information corresponding to the camera contained in the track, and the value is shown in Table 1.
[0114] Table 1
[0115]
[0116] 2. Free view information data box
[0117]
[0118]
[0119] stitching_layout: indicates whether the texture map and the depth map in the track are stitched and encoded, and the value is shown in Table 2:
[0120] Table 2
[0121]
[0122]
[0123] depth_padding_size: the width of the guard band of the depth map.
[0124] texture_padding_size: the width of the guard band of the texture map.
[0125] camera_model: indicates the model type of the camera, and the values are shown in Table 3:
[0126] Table 3
[0127]
[0128] camera_count: the number of all cameras that collect the video.
[0129] camera_id: the camera identifier corresponding to each view.
[0130] camera_pos_x, camera_pos_y, camera_pos_z: respectively indicate the x, y, z components of the camera position.
[0131] focal_length_x, focal_length_y: respectively indicate the x, y components of the camera focal length.
[0132] camera_resolution_x, camera_resolution_y: the resolution width and height of the texture map and the depth map collected by the camera.
[0133] depth_downsample_factor: the down-sampling factor of the depth map, and the actual resolution width and height of the depth map is 1 / 2 of the camera collected resolution width and height. depth_downsample_factor .
[0134] depth_vetex_x, depth_vetex_y: the x, y component values of the offset of the top-left vertex of the depth map relative to the origin of the plane frame (the top-left vertex of the plane frame).
[0135] texture_vetex_x, texture_vetex_y: the x, y component values of the offset of the top-left vertex of the texture map relative to the origin of the plane frame (the top-left vertex of the plane frame).
[0136] para_num: number of user-defined camera parameters.
[0137] para_type: type of user-defined camera parameters.
[0138] para_length: length of user-defined camera parameters, in bytes.
[0139] camera_parameter: user-defined parameters.
[0140] As can be seen from the above, although the above embodiments indicate the parameter information related to the free-view video, and support multi-track packaging of the free view, the scheme does not indicate the coding independence of the texture and depth map corresponding to different cameras, so that in the single-track packaging mode, the client cannot determine whether the texture map and the depth map of a specific camera can be partially decoded. Similarly, in the case of missing the coding independence indication information, the server cannot determine whether the single-track packaged free-view video can be re-packaged according to the multi-track.
[0141] To solve the above technical problems, the present application adds coding independence indication information in the video track, which is used to indicate whether the video data of a single view in the M views corresponding to the video track depends on the video data of other views when coding. In this way, in the single-track packaging mode, the client can determine whether the texture map and the depth map of a specific camera can be partially decoded according to the coding independence indication information. In addition, in the single-track packaging mode, the server can also determine whether the single-track packaged free-view video can be re-packaged according to the multi-track according to the coding independence indication information, thereby improving the processing flexibility of the media file and improving the decoding efficiency of the media file.
[0142] The technical solutions of the embodiments of the present application will be described in detail in some embodiments. The following several embodiments can be combined with each other, and the same or similar concepts or processes can not be described in some embodiments.
[0143] Figure 5 A flowchart of a file packaging method of a free-view video provided by an embodiment of the present application is shown in FIG. 1, which includes the following steps: Figure 5
[0144] S501, a first device acquires a code stream of free-view video data.
[0145] The free-view video data includes video data of N views, and N is a positive integer.
[0146] The free-view video data of the embodiments of the present application is N-view video data collected by N cameras, for example, N is 6, and the video data of 6 different views is collected by 6 cameras to obtain 6-view video data, and the 6-view video data forms the free-view video data of the embodiments of the present application.
[0147] In some embodiments, the free-view video view is also referred to as multi-view video data.
[0148] In the embodiments of the present application, the manner in which the first device obtains the code stream of the free-view video data includes but is not limited to the following several manners:
[0149] Manner one, the first device obtains the code stream of the free-view video data from other devices.
[0150] For example, the first device obtains the code stream of the free-view video data from a storage device, or obtains the code stream of the free-view video data from other encoding devices.
[0151] Manner two, the first device encodes the free-view video data to obtain the code stream of the free-view video data. For example, the first device is an encoding device, and after the first device obtains the free-view video data from a collection device (for example, a camera), the first device encodes the free-view video data to obtain the code stream of the free-view video data.
[0152] The embodiments of the present application do not limit the specific content of the video data, for example, the video data includes at least one of collected texture map data and depth map data.
[0153] S502, the first device encapsulates the code stream of the free-view video data into at least one video track to obtain a media file of the free-view video data, and the video track includes codec independence indication information and video code streams of M views, and the codec independence indication information is used to indicate whether the video data of a single view in the M views corresponding to the video track depends on the video data of other views when being encoded and decoded, and M is a positive integer less than or equal to N.
[0154] Specifically, the first device encapsulates the code stream of the free-view video data into at least one video track, and the at least one video track forms the media file of the free-view video data.
[0155] In a possible implementation manner, a single-track encapsulation mode is adopted to encapsulate the code stream of the free-view video data into one video track.
[0156] In a possible implementation, a sample multi-track packaging manner is used to package the code stream of the free-view video data into multiple video tracks. For example, the video code stream corresponding to each of the N views is packaged into a video track, thereby obtaining N video tracks. Alternatively, the video code stream corresponding to one or more of the N views is packaged into a video track, thereby obtaining multiple video tracks, and each video track can include the video code stream corresponding to at least one view.
[0157] In order to facilitate the processing of the media file by the client or the server, a codec independence indication is added in the video track, so that the client or the server processes the media file according to the codec independence indication.
[0158] In some embodiments, the codec independence indication information is added in each formed video track, and the codec independence indication information is used to indicate whether the video data of a single view in multiple views corresponding to the video track depends on the video data of other views when being coded.
[0159] In some embodiments, when the N views are coded in a consistent manner, the codec independence indication information can be added in one or more video tracks, and the codec independence indication information is used to indicate whether the video data of a single view in the N views depends on the video data of other views when being coded.
[0160] In some embodiments, if the value of the codec independence indication information is a first value, it indicates that the texture map data of the single view depends on the texture map data and the depth map data of other views when being coded, or the depth map data of the single view depends on the texture map data and the depth map data of other views when being coded; or,
[0161] If the value of the codec independence indication information is a second value, it indicates that the texture map data of the single view depends on the texture map data of other views when being coded, and the depth map data of the single view depends on the depth map data of other views when being coded; or,
[0162] If the value of the codec independence indication information is a third value, it indicates that the texture map data and the depth map data of the single view do not depend on the texture map data and the depth map data of other views when being coded, and the texture map data and the depth map data of the single view depend on each other when being coded; or,
[0163] If the value of the codec independence indication information is a fourth value, it indicates that the texture map data and the depth map data of the single view do not depend on the texture map data and the depth map data of other views when being coded, and the texture map data and the depth map data of the single view do not depend on each other when being coded.
[0164] The correspondence between the value of the codec dependency indication information and the codec dependency indicated by the codec dependency indication information is shown in Table 4.
[0165] Table 4
[0166]
[0167]
[0168] The application does not limit the specific values of the first value, the second value, the third value and the fourth value, and the specific values are determined according to actual needs.
[0169] Optionally, the first value is 0.
[0170] Optionally, the second value is 1.
[0171] Optionally, the third value is 2.
[0172] Optionally, the fourth value is 3.
[0173] In some embodiments, the codec dependency indication information can be added in the free-view information data box of the video track.
[0174] If the packaging standard of the media file is ISOBMFF, the codec dependency indication information is represented by the field codec_independency.
[0175] The free-view information data box of the application includes the following contents:
[0176]
[0177] In the application, the unsigned int (8) stitching_layout field in the free-view information data box is deleted. The stitching_layout indicates whether the texture map and the depth map in the track are stitched and encoded, and the details are shown in Table 5.
[0178] Table 5
[0179]
[0180] Correspondingly, the application adds the codec_independency field, which indicates the codec dependency between the texture map and the depth map corresponding to each camera in the track, and the details are shown in Table 6.
[0181] Table 6
[0182]
[0183]
[0184] depth_padding_size: the width of the guard band of the depth map.
[0185] texture_padding_size: the width of the guard band of the texture map.
[0186] camera_model: the model type of the camera, such as Table 7.
[0187] Table 7
[0188] Camera model 6DoF video camera model 0 Pinhole model 1 Fisheye model Other Preserve
[0189] camera_count: the number of all cameras that capture the video.
[0190] camera_id: the camera identifier corresponding to each view.
[0191] camera_pos_x, camera_pos_y, camera_pos_z: the x, y, z component values indicating the camera position, respectively.
[0192] focal_length_x, focal_length_y: the x, y component values indicating the camera focal length, respectively.
[0193] camera_resolution_x, camera_resolution_y: the width and height of the resolution of the texture map and the depth map captured by the camera.
[0194] depth_downsample_factor: the down-sampling factor of the depth map, the actual resolution width and height of the depth map is 1 / 2 of the camera captured resolution width and height. depth_downsample_factor .
[0195] depth_vetex_x, depth_vetex_y: the x, y component values of the offset of the top-left vertex of the depth map relative to the origin of the plane frame (the top-left vertex of the plane frame).
[0196] texture_vetex_x, texture_vetex_y: the x, y component values of the offset of the top-left vertex of the texture map relative to the origin of the plane frame (the top-left vertex of the plane frame).
[0197] In some embodiments, if the bitstream of the target view video data is obtained from other devices, the codec independence indication information is included in the bitstream of the target view video data, and the codec independence indication information is used to indicate whether the video data of a single view in the N views depends on the video data of other views when being coded. In this way, the first device can obtain whether the video data of each view in the target view video data depends on the video data of other views when being coded according to the codec independence indication information carried in the bitstream, and then add the codec independence indication information in each video track generated.
[0198] In some embodiments, if the codec independence indication information is carried in the bitstream of the target view video data, the free view video bitstream syntax is extended in the embodiments of the present application, and the free view video bitstream syntax is taken as an example of 6DoF video, and the details are shown in Table 8.
[0199] Table 8
[0200]
[0201]
[0202] As shown in Table 8, the field of 6DoF video stitching layout stitching_layout is deleted in the free view video bitstream syntax of the embodiments of the present application, the stitching_layout is an 8-bit unsigned integer, and is used to identify whether the 6DoF video adopts the stitching layout of the texture map and the depth map, and the specific value is shown in Table 9.
[0203] Table 9
[0204]
[0205]
[0206] Correspondingly, the codec_independency field is added in the free view video bitstream syntax shown in Table 8, the codec_independency is an 8-bit unsigned integer, and is used to identify the codec independence between the texture map and the depth map corresponding to each camera of the 6DoF video, and the specific value is shown in Table 10.
[0207] Table 10
[0208]
[0209] The fields in Table 8 are introduced as follows:
[0210] The marker bit marker_bit is a binary variable. The value is 1, which is used to avoid the occurrence of a pseudo start code in the 6DoF video format extension bitstream.
[0211] texture_padding_size, is an 8-bit unsigned integer. It represents the number of padding pixels of the texture map, and its value ranges from 0 to 255.
[0212] depth_padding_size, is an 8-bit unsigned integer. It represents the number of padding pixels of the depth map, and its value ranges from 0 to 255.
[0213] camera_number, is an 8-bit unsigned integer, and its value ranges from 1 to 255. It is used to represent the number of cameras for 6DoF video acquisition.
[0214] camera_model, is an 8-bit unsigned integer, and its value ranges from 1 to 255. It is used to represent the model type of the camera. The 6DoF video camera model is shown in Table 11:
[0215] Table 11
[0216]
[0217] camera_resolution_x, is a 32-bit unsigned integer. It represents the resolution of the camera acquisition in the x direction.
[0218] camera_resolution_y, is a 32-bit unsigned integer. It represents the resolution of the camera acquisition in the y direction.
[0219] video_resolution_x, is a 32-bit unsigned integer. It represents the resolution of the 6DoF video in the x direction.
[0220] video_resolution_y, is a 32-bit unsigned integer. It represents the resolution of the 6DoF video in the y direction.
[0221] camera_translation_matrix[3], is a 3*32-bit floating-point matrix, representing the translation matrix of the camera.
[0222] camera_rotation_matrix[3][3], is a 9*32-bit floating-point matrix, representing the rotation matrix of the camera.
[0223] camera_focal_length_x, is a 32-bit floating-point number. It represents the focal length fx of the camera.
[0224] camera_focal_length_y, which is a 32-bit floating-point number, represents the focal length fy of the camera.
[0225] camera_principle_point_x, which is a 32-bit floating-point number, represents the offset px of the camera in the optical axis in the image coordinate system.
[0226] camera_principle_point_y, which is a 32-bit floating-point number, represents the offset py of the camera in the optical axis in the image coordinate system.
[0227] texture_top_left_x[camera_number], which is an array of 32-bit unsigned integers with a size of camera_number, represents the x-coordinate of the top-left corner of the texture image of the camera with the corresponding serial number in the 6DoF video frame
[0228] texture_top_left_y[camera_number], which is an array of 32-bit unsigned integers with a size of camera_number, represents the y-coordinate of the top-left corner of the texture image of the camera with the corresponding serial number in the 6DoF video frame
[0229] texture_bottom_right_x[camera_number], which is an array of 32-bit unsigned integers with a size of camera_number, represents the x-coordinate of the bottom-right corner of the texture image of the camera with the corresponding serial number in the 6DoF video frame
[0230] texture_bottom_right_y[camera_number], which is an array of 32-bit unsigned integers with a size of camera_number, represents the y-coordinate of the bottom-right corner of the texture image of the camera with the corresponding serial number in the 6DoF video frame
[0231] depth_top_left_x[camera_number], which is an array of 32-bit unsigned integers with a size of camera_number, represents the x-coordinate of the top-left corner of the depth image of the camera with the corresponding serial number in the 6DoF video frame
[0232] A depth map top-left coordinate y component depth_top_left_y[camera_number], which is a 32-bit unsigned integer array with a size of camera_number, represents the y coordinate of the top-left corner of the depth map of the camera with the corresponding serial number in the 6DoF video frame
[0233] A depth map bottom-right coordinate x component depth_bottom_right_x[camera_number], which is a 32-bit unsigned integer array with a size of camera_number, represents the x coordinate of the bottom-right corner of the depth map of the corresponding serial number camera in the 6DoF video frame
[0234] A depth map bottom-right coordinate y component depth_bottom_right_y[camera_number], which is a 32-bit unsigned integer array with a size of camera_number, represents the y coordinate of the bottom-right corner of the depth map of the corresponding serial number camera in the 6DoF video frame
[0235] A nearest depth distance for depth map quantization depth_range_near, which is a 32-bit floating point number, is the minimum depth distance from the optical center for depth map quantization.
[0236] A farthest depth distance for depth map quantization depth_range_far, which is a 32-bit floating point number, is the maximum depth distance from the optical center for depth map quantization.
[0237] A depth map downsampling flag depth_scale_flag, which is an 8-bit unsigned integer, takes values from 1 to 255. It is used to represent the downsampling type of the depth map.
[0238] A background texture flag background_texture_flag, which is an 8-bit unsigned integer, is used to indicate whether to transmit the background texture map of the multi-camera.
[0239] A background depth flag background_depth_flag, which is an 8-bit unsigned integer, is used to indicate whether to transmit the background depth map of the multi-camera.
[0240] If the background depth is applied, the decoded depth map background frame does not participate in the view synthesis, and the subsequent frame participates in the virtual view synthesis. The depth map downsampling flag is shown in Table 12:
[0241] Table 12
[0242]
[0243] According to the above, the first device determines the specific value of the coding dependency indication information according to whether the video data of a single view in the M views corresponding to the video track depends on the video data of other views when being coded, and adds the specific value of the coding dependency indication information in the video track.
[0244] In some embodiments, if the video data corresponding to each of the N views does not depend on the video data corresponding to other views when being coded, and the encapsulation mode of the code stream of the free-view video data is the single-track encapsulation mode, the first device adds the coding dependency indication information in the free-view information data box of a video track formed by the single-track encapsulation mode, where the value of the coding dependency indication information is the third value or the fourth value. In this way, the server or the client can determine that the video data corresponding to each of the N views does not depend on the video data corresponding to other views when being coded according to the value of the coding dependency indication information, and can request to partially encapsulate the media file corresponding to part of the views, or re-encapsulate the single-track encapsulated video track into multiple video tracks.
[0245] S503, sending the media file of the free-view video data to the client or the server.
[0246] According to the above method, the code stream of the free-view video data is encapsulated into at least one video track to obtain a media file of the free-view video data, and the coding dependency indication information is added in the video track. Then, the media file including the coding dependency indication information is sent to the client or the server, so that the client or the server processes the media file according to the coding dependency indication information carried in the media file. For example, if the coding dependency indication information indicates that the video data corresponding to a single view does not depend on the video data corresponding to other views when being coded, the client or the server can request to partially encapsulate the media file corresponding to part of the views, or re-encapsulate the single-track encapsulated video track into multiple video tracks.
[0247] The file encapsulation method of the free-view video provided in the embodiments of the present application adds the coding dependency indication information in the video track, where the coding dependency indication information is used to indicate whether the video data of a single view in the M views corresponding to the video track depends on the video data of other views when being coded. In this way, in the single-track encapsulation mode, the client can determine whether the texture map and the depth map of a specific camera can be partially decoded according to the coding dependency indication information. In addition, in the single-track encapsulation mode, the server can also determine whether the single-track encapsulated free-view video can be re-encapsulated according to the coding dependency indication information, thereby improving the processing flexibility of the media file.
[0248] In some embodiments, if the encoding mode of the free-view video data is the AVS3 encoding mode, the media file of the embodiments of the present application can be encapsulated in the form of a sub-sample, at this time, the method of the embodiments of the present application further comprises:
[0249] S500, encapsulate at least one of the header information required for decoding, the texture map information of at least one view, and the depth map information of at least one view in the form of a sub-sample in the media file.
[0250] Among them, the data box of the sub-sample includes a sub-sample data box flag and sub-sample indication information, the sub-sample data box flag is used to indicate the division mode of the sub-sample, and the sub-sample indication information is used to indicate the content included in the sub-sample.
[0251] Among them, the content included in the sub-sample includes at least one of the header information required for decoding, the texture map information of at least one view, and the depth map information of at least one view.
[0252] In some embodiments, if the value of the sub-sample indication information is the fifth numerical value, it indicates that a sub-sample includes the header information required for decoding; or,
[0253] If the value of the sub-sample indication information is the sixth numerical value, it indicates that a sub-sample includes the texture map information corresponding to N views (or cameras) in the current video frame; or,
[0254] If the value of the sub-sample indication information is the seventh numerical value, it indicates that a sub-sample includes the depth map information corresponding to N views (or cameras) in the current video frame; or,
[0255] If the value of the sub-sample indication information is the eighth numerical value, it indicates that a sub-sample includes the texture map information and the depth map information corresponding to one view (or camera) in the current video frame; or,
[0256] If the value of the sub-sample indication information is the ninth numerical value, it indicates that a sub-sample includes the texture map information corresponding to one view (or camera) in the current video frame; or,
[0257] If the value of the sub-sample indication information is the tenth numerical value, it indicates that a sub-sample includes the depth map information corresponding to one view (or camera) in the current video frame.
[0258] The above current video frame is spliced from the video frames corresponding to the N views, for example, the video frames corresponding to N views collected by N cameras at the same time point are spliced to form the current video frame.
[0259] The texture information and / or the depth information can be understood as data required for unpacking the texture stream or the depth stream.
[0260] Optionally, the texture information and / or the depth information includes a position offset of the texture stream and / or the depth stream in the media file. For example, the texture stream corresponding to the view angle 1 is stored at the tail position of the media file, the texture information corresponding to the view angle 1 includes an offset of the texture information corresponding to the view angle 1 at the tail position of the media file, and the position of the texture stream corresponding to the view angle 1 in the media file can be obtained according to the offset.
[0261] The correspondence between the value of the sub-sample indication information and the content included in the sub-sample is shown in Table 13.
[0262] Table 13
[0263]
[0264] The application does not limit the specific values of the fifth value, the sixth value, the seventh value, the eighth value, the ninth value and the tenth value, and the specific values are determined according to actual needs.
[0265] Optionally, the fifth value is 0.
[0266] Optionally, the sixth value is 1.
[0267] Optionally, the seventh value is 2.
[0268] Optionally, the eighth value is 3.
[0269] Optionally, the ninth value is 4.
[0270] Optionally, the tenth value is 5.
[0271] Optionally, the value of the flags field of the sub-sample data box is a preset value, for example, 1, indicating that the sub-sample includes valid content.
[0272] In some embodiments, the sub-sample indication information is indicated by payloadType in the codec_specific_parameters field of the SubSampleInformationBox data box.
[0273] In an example, the value of the codec_specific_parameters field in the SubSampleInformationBox data box is as follows:
[0274]
[0275] The value of the payloadType field is shown in Table 14.
[0276] Table 14
[0277]
[0278]
[0279] In some embodiments, if the encoding mode of the free-view video data is an AVS3 encoding mode, the video data corresponding to each of the N views is encoded independently of the video data corresponding to other views, and the encapsulation mode of the free-view video data stream is a single-track encapsulation mode, the above S500 includes the following S500-A:
[0280] S500-A, encapsulate at least one of the texture map information and the depth map information corresponding to each of the N views in the form of a sub-sample in the media file.
[0281] The sub-sample data box flag and the sub-sample indication information are included in the data box of each sub-sample formed above.
[0282] The value of the sub-sample data box flag is a preset value, for example, 1, indicating that the sub-sample is divided in units of views.
[0283] The value of the sub-sample indication information is determined according to the content included in the sub-sample, which can include the following examples.
[0284] Example one, if the sub-sample includes the texture map information and the depth map information corresponding to one view in the current video frame, the value of the sub-sample indication information corresponding to the sub-sample is the eighth value.
[0285] Example two, if the sub-sample includes the texture map information corresponding to one view in the current video frame, the value of the sub-sample indication information corresponding to the sub-sample is the ninth value.
[0286] Example three, if the sub-sample includes the depth map information corresponding to one view in the current video frame, the value of the sub-sample indication information corresponding to the sub-sample is the tenth value.
[0287] In an embodiment of the present application, if the encoding mode of the free-view video data is an AVS3 encoding mode, the video data corresponding to each view in the N views is independent of the video data corresponding to other views during encoding, and the encapsulation mode of the free-view video data code stream is a single-track encapsulation mode, the texture map information and the depth map information corresponding to each view in the N views are encapsulated in the media file in the form of sub-samples. At this time, after requesting the complete free-view video, the client can decode part of the views according to its own needs, so as to save the computing resources of the client.
[0288] Figure 6 An interactive flowchart of a file encapsulation method of a free-view video provided in an embodiment of the present application is shown in FIG. 1, and the method comprises the following steps: Figure 6
[0289] S601, a first device acquires a code stream of free-view video data.
[0290] The free-view video data comprises video data of N views, and N is a positive integer.
[0291] S602, the first device encapsulates the code stream of the free-view video data into at least one video track to obtain a media file of the free-view video data.
[0292] The video track comprises codec independence indication information and video code streams of M views, the codec independence indication information is used to indicate whether the video data of a single view in the M views corresponding to the video track is dependent on the video data of other views during encoding, and M is a positive integer less than or equal to N.
[0293] S603, the first device sends the media file of the free-view video data to a client.
[0294] The execution processes of S601 to S603 are the same as those of S501 to S503, and refer to the specific description of S501 to S503, which will not be described here.
[0295] S604, the client decapsulates the media file according to the codec independence indication information to obtain video code streams corresponding to at least one view.
[0296] The codec independence indication information is added in the media file of the present application, and the codec independence indication information is used to indicate whether the single-view video data depends on the video data of other views during coding. Thus, after receiving the media file, the client can determine whether the single-view video data corresponding to the media file depends on the video data of other views during coding according to the codec independence indication information carried in the media file. If it is determined that the single-view video data corresponding to the media file depends on the video data of other views during coding, it is indicated that the media file corresponding to the single view cannot be unpacked. Therefore, the client needs to unpack the media files corresponding to all views in the media file, and then obtain the video bitstream corresponding to N views.
[0297] If it is determined that the single-view video data corresponding to the media file does not depend on the video data of other views during coding, it is indicated that the media file corresponding to the single view can be unpacked. Thus, the client can unpack the media files of part of the views according to the needs, obtain the video bitstream corresponding to part of the views, and then save the computing resources of the client.
[0298] As shown in Table 4, the codec independence indication information indicates whether the single-view video data depends on the video data of other views during coding through different values. Thus, the client can determine whether the single-view video data corresponding to N views in the media file depends on the video data of other views during coding according to the value of the codec independence indication information and Table 4.
[0299] In some embodiments, S604 includes the following steps:
[0300] S604-A1, if the value of the codec independence indication information is the third value or the fourth value, and the packaging mode of the bitstream is the single-track packaging mode, the client obtains the viewing view of the user.
[0301] S604-A1, the client determines the target view corresponding to the viewing view of the user according to the viewing view of the user and the view information in the media file.
[0302] S604-A1, the client unpacks the media file corresponding to the target view to obtain the video bitstream corresponding to the target view.
[0303] That is, in the embodiments of the present application, if the value of the codec independence indication information carried in the media file is the third value or the fourth value, as shown in Table 4, the third value is used to indicate that the texture map data or the depth map data of a single view angle is independent of the texture map data or the depth map data of other view angles in the codec, and the texture map data and the depth map data of the single view angle are dependent on each other in the codec, and the fourth value is used to indicate that the texture map data or the depth map data of a single view angle is independent of the texture map data or the depth map data of other view angles in the codec, and the texture map data and the depth map data of the single view angle are independent of each other in the codec. Therefore, if the value of the codec independence indication information carried in the media file is the third value or the fourth value, it indicates that the texture map data or the depth map data of each view angle in the media file is independent of the texture map data or the depth map data of other view angles in the codec, so that the client can decode the texture map data and / or the depth map data of part of the view angles as needed.
[0304] Specifically, according to the viewing angle of the user and the view angle information in the media file, a target view angle conforming to the viewing angle of the user is determined, and the media file corresponding to the target view angle is unpacked to obtain a video code stream corresponding to the target view angle. The manner in which the embodiments of the present application unpack the media file corresponding to the target view angle to obtain the video code stream corresponding to the target view angle is not limited, and any existing manner can be used, for example, the position of the video code stream corresponding to the target view angle in the media file is determined according to the information corresponding to the target view angle in the media file, and then the video code stream corresponding to the target view angle is unpacked to obtain the video code stream corresponding to the target view angle, so as to realize the decoding of the video data corresponding to part of the view angles and save the computing resources of the client.
[0305] In some embodiments, if the value of the codec independence indication information is the first value or the second value, it indicates that the texture map data and / or the depth map data of each view angle in the media file is dependent on the texture map data and / or the depth map data of other view angles in the codec, and at this time, the client needs to unpack the media file corresponding to all the view angles in the N view angles in the media file.
[0306] S605, the client decodes the video code stream corresponding to at least one view angle to obtain reconstructed video data of the at least one view angle.
[0307] After obtaining the video code stream corresponding to at least one view angle according to the above manner, the client can decode the video code stream corresponding to the at least one view angle and render the decoded video data.
[0308] The process of decoding the video code stream can refer to the description of the prior art, which will not be described here.
[0309] In some embodiments, if the coding mode is the AVS3 video coding mode and the media file includes a subsample, S605 includes:
[0310] S605-A1, obtaining the subsample data box flag and the subsample indication information included in the data box of the subsample.
[0311] As can be seen from S500, the first device can encapsulate the header information required for decoding, the texture map information of at least one view, and the depth map information of at least one view in the media file in the form of a subsample. Based on this, after the client receives the media file, when it is detected that the media file includes a subsample, the client obtains the subsample data box flag and the subsample indication information included in the data box of the subsample.
[0312] The subsample data box flag is used to indicate the division mode of the subsample, and the subsample indication information is used to indicate the content included in the subsample.
[0313] The content included in the subsample includes at least one of the header information required for decoding, the texture map information of at least one view, and the depth map information of at least one view.
[0314] S605-A2, obtaining the content included in the subsample according to the subsample data box flag and the subsample indication information.
[0315] For example, when the value of the subsample data box flag is 1, it indicates that the subsample includes valid content. Then, the client queries the content included in the subsample corresponding to the subsample indication information from Table 13 based on the correspondence between the subsample indication information and the content included in the subsample shown in Table 13.
[0316] S605-A3, decapsulating the media file resource according to the content included in the subsample and the codec independence indication information to obtain the video bitstream corresponding to at least one view.
[0317] The content included in the subsample is shown in Table 13, and the value of the codec independence indication information and the information indicated thereby are shown in Table 4. In this way, the media file resource can be decapsulated according to whether the video data of a single view indicated by the codec independence indication information depends on the video data of other views when being coded, and the content included in the subsample, to obtain the video bitstream corresponding to at least one view.
[0318] For example, the coding dependency indication information is used to indicate that the video data of the single view is coded independently of the video data of other views, and the content included in the sub-sample is the header information required for decoding. Thus, according to the coding dependency indication information, the client obtains the video code stream corresponding to the target view, and decodes the video code stream corresponding to the target view according to the header information required for decoding included in the sub-sample.
[0319] For another example, the coding dependency indication information is used to indicate that the video data of the single view is coded dependent on the video data of other views, and the content included in the sub-sample is the texture map information of all frames in the frame. Thus, according to the coding dependency indication information, the client obtains the video code stream corresponding to all views, and decodes the video code stream corresponding to all views according to the texture map information of all frames in the frame included in the sub-sample.
[0320] It should be noted that, after Table 4 and Table 13 are given in the embodiments of the present application, the unpacking of the media file resource according to the content included in the sub-sample and the coding dependency indication information can be implemented by the existing manner, and thus will not be illustrated one by one here.
[0321] In some embodiments, if the media file includes the sub-sample corresponding to each of the N views, S605-A3 includes:
[0322] S605-A31, if the coding dependency indication information has the third value or the fourth value, and the encapsulation mode of the code stream is the single-track encapsulation mode, the client determines the target view according to the viewing angle of the user and the view information in the media file.
[0323] S605-A32, the client obtains the sub-sample data box flag and the sub-sample indication information included in the data box of the target sub-sample corresponding to the target view;
[0324] S605-A33, the client determines the content included in the target sub-sample according to the sub-sample data box flag and the sub-sample indication information included in the data box of the target sub-sample.
[0325] S605-A34, the client unpacks the media file corresponding to the target view according to the content included in the target sub-sample, and obtains the video code stream corresponding to the target view.
[0326] Specifically, if the coding dependency indication information has the third value or the fourth value, it indicates that the video data of each view in the media file is not dependent on the video data of other views during coding. Thus, the client determines a target view according to a viewing view of a user and the view information in the media file. Then, the client obtains a target sub-sample corresponding to the target view from the sub-sample corresponding to each view of the N views, and further obtains the sub-sample data box flag and the sub-sample indication information included in the data box of the target sub-sample, and determines the content included in the target sub-sample according to the sub-sample data box flag and the sub-sample indication information included in the data box of the target sub-sample. For example, if the sub-sample data box flag is 1 and the sub-sample indication information has the value 3, as shown in Table 13, it can be determined that the content included in the target sub-sample is the texture map information and the depth map information corresponding to a single camera in the current video frame. Then, the client performs de-encapsulation on the media file corresponding to the target view according to the content included in the target sub-sample, to obtain the video bitstream corresponding to the target view, for example, the client obtains the position of the video bitstream corresponding to the target view in the media file according to the texture map information and the depth map information included in the target sub-sample, and then performs de-encapsulation on the video bitstream corresponding to the target view in the media file, to obtain the video bitstream corresponding to the target view.
[0327] In the embodiment of the present application, the first device adds the coding dependency indication information in the video track, which is used to indicate whether the video data of a single view in the M views corresponding to the video track is dependent on the video data of other views during coding. Thus, in the single-track encapsulation mode, the client can determine whether the texture map and the depth map of a specific camera can be partially decoded according to the coding dependency indication information, thereby improving the processing flexibility of the client for the media file, and the client can save the computing resources when decoding part of the media file according to the coding dependency indication information.
[0328] Figure 7 An interactive flowchart of a file encapsulation method of a free-view video provided by the embodiment of the present application is shown in FIG. 1, which includes the following steps: Figure 7
[0329] S701, a first device obtains a bitstream of free-view video data.
[0330] The free-view video data includes video data of N views, and N is a positive integer.
[0331] S702, the first device encapsulates the bitstream of the free-view video data into at least one video track, to obtain a media file of the free-view video data.
[0332] The video track includes codec independence indication information and video code streams of M views, the codec independence indication information is used to indicate whether video data of a single view in the M views corresponding to the video track depends on video data of other views when being coded, and M is a positive integer less than or equal to N.
[0333] S703, the first device sends the media file of the free-view video data to the server.
[0334] The execution processes of S701 to S703 are consistent with those of S501 to S503, and details are described with reference to S501 to S503, which will not be repeated here.
[0335] S704, the server determines whether to decompose the at least one video track into multiple video tracks according to the codec independence indication information.
[0336] In some embodiments, if the value of the codec independence indication information is the third value or the fourth value, and the packaging mode of the code stream is the single-track packaging mode, then the one video track formed by the single-track packaging mode is packaged into N video tracks in the multi-track packaging mode, and each of the N video tracks includes video data corresponding to a single view.
[0337] In some embodiments, if the value of the codec independence indication information is the first value or the second value, and the packaging mode of the code stream is the single-track packaging mode, then the server cannot package the one video track formed by the single-track packaging mode into N video tracks.
[0338] In some embodiments, if the encoding mode of the free-view video data is the AVS3 video encoding mode, and the media file includes a sub-sample corresponding to each of the N views, and each sub-sample corresponding to each view includes at least one of depth map information and texture map information corresponding to each view, then S704 includes:
[0339] S704-A1, for each of the N views, obtaining sub-sample data box flags and sub-sample indication information included in a data box of a sub-sample corresponding to the view, the sub-sample data box flags being used to indicate a division mode of the sub-sample, and the sub-sample indication information being used to indicate content included in the sub-sample;
[0340] S704-A2, obtaining the content included in the sub-sample corresponding to each view according to the sub-sample data box flags and the sub-sample indication information corresponding to each view;
[0341] S704-A3, packaging the one video track formed by the single-track packaging mode into N video tracks in the multi-track packaging mode according to the content included in the sub-sample corresponding to each view.
[0342] For example, taking a single view angle in N view angles as an example, the value of the sub-sample data box flag in the data box of the sub-sample corresponding to the view angle is 1, and the value of the sub-sample indication information is 3. As shown in Table 13, the value of the sub-sample indication information is 3, indicating that the sub-sample includes texture map information and depth map information corresponding to the view angle. Then, according to the texture map information and the depth map information corresponding to the view angle in the single-track encapsulation mode, the video code stream corresponding to the view angle is queried in a single video track. Based on the same manner, the video code stream corresponding to each view angle in N view angles can be queried, and each corresponding video code stream is encapsulated into a video track to obtain N video tracks.
[0343] In some embodiments, after the server decomposes the single video track into multiple video tracks according to the codec independence indication information, the method of the embodiments of the present application further includes:
[0344] S705, after the server decomposes the at least one video track into multiple video tracks according to the codec independence indication information, the server generates first signaling.
[0345] In an example, N = 3, and the server encapsulates the single-track video file into three video tracks Track1, Track2 and Track3, as follows:
[0346] Track1: {Camera1: ID = 1; Pos = (100, 0, 100); Focal = (10, 20)};
[0347] Track2: {Camera2: ID = 2; Pos = (100, 100, 100); Focal = (10, 20)};
[0348] Track3: {Camera3: ID = 3; Pos = (0, 0, 100); Focal = (10, 20)}.
[0349] The first signaling includes at least one of the identification of the camera corresponding to each track in the N video tracks, the position information of the camera, and the focal point information of the camera.
[0350] In an example, the first signaling includes three representations Representation1, Representation2 and Representation3, as shown below:
[0351] Representation1: {Camera1: ID = 1; Pos = (100, 0, 100); Focal = (10, 20)};
[0352] Representation2: {Camera2: ID=2; Pos=(100,100,100); Focal=(10,20)};
[0353] Representation3: {Camera3: ID=3; Pos=(0,0,100); Focal=(10,20)}.
[0354] S706, the server sends the first signaling to the client.
[0355] S707. The client generates first request information based on the first signaling. For example, the client selects the video file corresponding to the target camera and requests it based on the user's location information and the shooting location information of each perspective in the first signaling (i.e., the camera's location information).
[0356] For example, the target cameras are Camera2 and Camera3.
[0357] S708. The client sends a first request message to the server, which includes the identification information of the target camera.
[0358] S709. Based on the first request information, the server sends the media file resources corresponding to the target camera to the client.
[0359] In this embodiment of the application, when the server determines that the video track corresponding to the single-track mode can be decomposed into multiple video tracks based on the encoding / decoding independence indication information, it decomposes the video track corresponding to the single-track mode into multiple video tracks, thereby supporting the client to request media resources corresponding to some viewpoints, achieving the purpose of partial transmission and partial decoding.
[0360] It should be understood that Figure 5 to Figure 7 This is merely an example of what is being done and should not be construed as limiting the scope of this application.
[0361] The preferred embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this application, various simple modifications can be made to the technical solutions of this application, and these simple modifications all fall within the protection scope of this application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this application will not describe the various possible combinations separately. Furthermore, various different embodiments of this application can also be arbitrarily combined, as long as they do not violate the spirit of this application, they should also be considered as the content disclosed in this application.
[0362] The above text combined Figure 5 to Figure 7The method embodiments of the present application are described in detail below, and the device embodiments of the present application are described in detail below Figure 8 to Figure 11 The device embodiments of the present application are described in detail below.
[0363] Figure 8 A structural diagram of a file packaging device for free-view video is provided for an embodiment of the present application. The device 10 is applied to a first device. The device 10 includes:
[0364] An acquisition unit 11 is configured to acquire a code stream of free-view video data. The free-view video data includes video data of N views, and N is a positive integer.
[0365] A packaging unit 12 is configured to package the code stream of the free-view video data into at least one video track to obtain a media file of the free-view video data. The video track includes codec independence indication information and video code streams of M views. The codec independence indication information is used to indicate whether video data of a single view in the M views corresponding to the video track depends on video data of other views when being coded. M is a positive integer less than or equal to N.
[0366] A sending unit 13 is configured to send the media file of the free-view video data to a client or a server.
[0367] Optionally, the video data includes at least one of texture map data and depth map data.
[0368] In some embodiments, if a value of the codec independence indication information is a first numerical value, it indicates that texture map data of the single view depends on texture map data and depth map data of other views when being coded, or depth map data of the single view depends on texture map data and depth map data of other views when being coded; or,
[0369] If the value of the codec independence indication information is a second numerical value, it indicates that texture map data of the single view depends on texture map data of other views when being coded, and depth map data of the single view depends on depth map data of other views when being coded; or,
[0370] If the value of the codec independence indication information is a third numerical value, it indicates that texture map data and depth map data of the single view do not depend on texture map data and depth map data of other views when being coded, and texture map data and depth map data of the single view depend on each other when being coded; or,
[0371] If the value of the codec dependency indication information is a fourth numerical value, it indicates that the texture map data and the depth map data of the single view angle are independent of the texture map data and the depth map data of other view angles in coding, and the texture map data and the depth map data of the single view angle are independent of each other in coding.
[0372] In some embodiments, the encapsulation unit 12 encapsulates the code stream of the free view angle video data into one video track in a single track encapsulation mode.
[0373] In some embodiments, if the video data corresponding to each of the N view angles is independent of the video data corresponding to other view angles in coding, and the encapsulation mode of the code stream of the free view angle video data is a single track encapsulation mode, the encapsulation unit 12 further adds the codec dependency indication information in the free view angle information data box of the one video track, and the value of the codec dependency indication information is the third numerical value or the fourth numerical value.
[0374] In some embodiments, if the encoding mode of the free view angle video data is an AVS3 encoding mode, the encapsulation unit 12 further encapsulates at least one of the header information required for decoding, the texture map information of at least one view angle, and the depth map information of at least one view angle in the form of a sub-sample in the media file.
[0375] The data box of the sub-sample includes a sub-sample data box flag and sub-sample indication information, the sub-sample data box flag is used to indicate the division mode of the sub-sample, and the sub-sample indication information is used to indicate the content included in the sub-sample.
[0376] In some embodiments, if the value of the sub-sample indication information is a fifth numerical value, it indicates that one sub-sample includes the header information required for decoding; or,
[0377] If the value of the sub-sample indication information is a sixth numerical value, it indicates that one sub-sample includes the texture map information corresponding to the N view angles in the current video frame, and the current video frame is spliced from the video frames corresponding to the N view angles; or,
[0378] If the value of the sub-sample indication information is a seventh numerical value, it indicates that one sub-sample includes the depth map information corresponding to the N view angles in the current video frame; or,
[0379] If the value of the sub-sample indication information is an eighth numerical value, it indicates that one sub-sample includes the texture map information and the depth map information corresponding to one view angle in the current video frame; or,
[0380] If the value of the sub-sample indication information is the ninth value, it is indicated that one sub-sample includes the texture map information corresponding to one view angle in the current video frame; or
[0381] If the value of the sub-sample indication information is the tenth value, it is indicated that one sub-sample includes the depth map information corresponding to one view angle in the current video frame.
[0382] In some embodiments, if the encoding mode of the free-view video data is an AVS3 encoding mode, the video data corresponding to each view angle in the N view angles is independent of the video data corresponding to other view angles during encoding, and the encapsulation mode of the free-view video data code stream is a single-track encapsulation mode, the encapsulation unit 12 is specifically configured to encapsulate the texture map information and the depth map information corresponding to each view angle in the N view angles in the form of a sub-sample in the media file.
[0383] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, the foregoing and other operations and / or functions of each module in the device 8 are described below. Figure 8 The device 10 shown can perform the method embodiments corresponding to the first device, and the foregoing and other operations and / or functions of each module in the device 8 are respectively to realize the method embodiments corresponding to the first device, and are not described here for brevity.
[0384] Figure 9 The device 20 provided by an embodiment of the present application is a file encapsulation device of free-view video, and the device 20 is applied to a client. The device 20 includes:
[0385] The receiving unit 21 is configured to receive a media file of free-view video data sent by a first device. The media file includes at least one video track. The free-view video data includes video data of N view angles. The N is a positive integer. The video track includes codec independence indication information and video code streams of M view angles. The codec independence indication information is used to indicate whether the video data of a single view angle in the M view angles corresponding to the video track is dependent on the video data of other view angles during codec. The M is a positive integer less than or equal to the N.
[0386] The decapsulation unit 22 is configured to decapsulate the media file according to the codec independence indication information, to obtain video code streams corresponding to at least one view angle.
[0387] The decoding unit 23 is configured to decode the video code streams corresponding to the at least one view angle, to obtain reconstructed video data of the at least one view angle.
[0388] Optionally, the video data includes at least one of texture map data and depth map data.
[0389] In some embodiments, if the value of the coding dependency indication information is a first numerical value, it indicates that the texture map data of the single view angle is dependent on the texture map data and the depth map data of other view angles during coding, or the depth map data of the single view angle is dependent on the texture map data and the depth map data of other view angles during coding; or,
[0390] If the value of the coding dependency indication information is a second numerical value, it indicates that the texture map data of the single view angle is dependent on the texture map data of other view angles during coding, and the depth map data of the single view angle is dependent on the depth map data of other view angles during coding; or,
[0391] If the value of the coding dependency indication information is a third numerical value, it indicates that the texture map data and the depth map data of the single view angle are not dependent on the texture map data and the depth map data of other view angles during coding, and the texture map data and the depth map data of the single view angle are dependent on each other during coding; or,
[0392] If the value of the coding dependency indication information is a fourth numerical value, it indicates that the texture map data and the depth map data of the single view angle are not dependent on the texture map data and the depth map data of other view angles during coding, and the texture map data and the depth map data of the single view angle are not dependent on each other during coding.
[0393] In some embodiments, the decapsulation unit 22 is configured to, if the value of the coding dependency indication information is the third numerical value or the fourth numerical value, and the encapsulation mode of the code stream is a single-track encapsulation mode, obtain a viewing view angle of a user; determine a target view angle according to the viewing view angle of the user and the view angle information in the media file; and decapsulate the media file corresponding to the target view angle to obtain a video code stream corresponding to the target view angle.
[0394] In some embodiments, if the coding manner of the free-view video data is an AVS3 video coding manner, and the media file includes a sub-sample, the decapsulation unit 22 is configured to obtain a sub-sample data box flag and the sub-sample indication information included in a data box of the sub-sample, the sub-sample data box flag is used to indicate a division manner of the sub-sample, and the sub-sample indication information is used to indicate content included in the sub-sample, the content included in the sub-sample includes at least one of decoding required header information, texture map information of at least one view angle, and depth map information of at least one view angle; obtain the content included in the sub-sample according to the sub-sample data box flag and the sub-sample indication information; and decapsulate the media file resource according to the content included in the sub-sample and the coding dependency indication information to obtain a video code stream corresponding to at least one view angle.
[0395] In some embodiments, if the value of the sub-sample indication information is a fifth value, it indicates that one sub-sample includes the header information required for decoding; or,
[0396] If the value of the sub-sample indication information is a sixth value, it indicates that one sub-sample includes the texture map information corresponding to N views in the current video frame; or,
[0397] If the value of the sub-sample indication information is a seventh value, it indicates that one sub-sample includes the depth map information corresponding to N views in the current video frame; or,
[0398] If the value of the sub-sample indication information is an eighth value, it indicates that one sub-sample includes the texture map information and the depth map information corresponding to one view in the current video frame; or,
[0399] If the value of the sub-sample indication information is a ninth value, it indicates that one sub-sample includes the texture map information corresponding to one view in the current video frame; or,
[0400] If the value of the sub-sample indication information is a tenth value, it indicates that one sub-sample includes the depth map information corresponding to one view in the current video frame.
[0401] In some embodiments, if the media file includes a sub-sample corresponding to each of the N views, the encapsulating unit 22 is specifically configured to, if the value of the codec independence indication information is the third value or the fourth value and the encapsulation mode of the code stream is a single-track encapsulation mode, determine a target view according to the viewing angle of the user and the view information in the media file, obtain the sub-sample data box flag and the sub-sample indication information included in the data box of the target sub-sample corresponding to the target view, determine the content included in the target sub-sample according to the sub-sample data box flag and the sub-sample indication information included in the data box of the target sub-sample, and perform de-encapsulation on the media file corresponding to the target view according to the content included in the target sub-sample, to obtain a video code stream corresponding to the target view.
[0402] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, no longer described here. Specifically, Figure 9 The device 20 shown can perform the corresponding method embodiments of the client, and the foregoing and other operations and / or functions of each module in the device 20 are respectively for realizing the corresponding method embodiments of the client. For the sake of brevity, no longer described here.
[0403] Figure 10The structural diagram of a file packaging device for free-view video is provided for an embodiment of the present application. The device 30 is applied to a server. The device 30 comprises:
[0404] A receiving unit 31 is configured to receive a media file of free-view video data sent by a first device. The media file comprises at least one video track. The free-view video data comprises video data of N views. The N is a positive integer. The video track comprises codec independence indication information and video code streams of M views. The codec independence indication information is used to indicate whether video data of a single view in the M views corresponding to the video track depends on video data of other views when being coded. The M is a positive integer less than or equal to the N.
[0405] A decomposing unit 32 is configured to determine whether to decompose the at least one video track into multiple video tracks according to the codec independence indication information.
[0406] Optionally, the video data comprises at least one of texture map data and depth map data.
[0407] In some embodiments, if the value of the codec independence indication information is a first numerical value, it indicates that the texture map data of the single view depends on the texture map data and the depth map data of other views when being coded, or the depth map data of the single view depends on the texture map data and the depth map data of other views when being coded; or,
[0408] If the value of the codec independence indication information is a second numerical value, it indicates that the texture map data of the single view depends on the texture map data of other views when being coded, and the depth map data of the single view depends on the depth map data of other views when being coded; or,
[0409] If the value of the codec independence indication information is a third numerical value, it indicates that the texture map data and the depth map data of the single view do not depend on the texture map data and the depth map data of other views when being coded, and the texture map data and the depth map data of the single view depend on each other when being coded; or,
[0410] If the value of the codec independence indication information is a fourth numerical value, it indicates that the texture map data and the depth map data of the single view do not depend on the texture map data and / or the depth map data of other views when being coded, and the texture map data and the depth map data of the single view do not depend on each other when being coded.
[0411] In some embodiments, if the value of the coding independence indication information is the third value or the fourth value, and the encapsulation mode of the code stream is a single-track encapsulation mode, the splitting unit 32 is specifically configured to encapsulate one video track formed by the single-track encapsulation mode into N video tracks according to a multi-track encapsulation mode, and each of the N video tracks includes video data corresponding to a single view.
[0412] In some embodiments, if the coding mode is an AVS3 video coding mode, and the media file includes a subsample corresponding to each of the N views, and each subsample corresponding to each view includes at least one of depth map information and texture map information corresponding to each view, the splitting unit 32 is specifically configured to, for each of the N views, obtain a subsample data box flag included in a data box of the subsample corresponding to the view and the subsample indication information, the subsample data box flag is used to indicate a division mode of the subsample, and the subsample indication information is used to indicate content included in the subsample; obtain the content included in the subsample corresponding to each view according to the subsample data box flag of the subsample corresponding to each view and the subsample indication information; and encapsulate one video track formed by the single-track encapsulation mode into N video tracks according to a multi-track encapsulation mode according to the content included in the subsample corresponding to each view.
[0413] In some embodiments, the apparatus further includes a generating unit 33 and a sending unit 34:
[0414] The generating unit 33 is configured to generate first signaling, the first signaling including at least one of an identifier of a camera corresponding to each of the N video tracks, position information of the camera, and focal point information of the camera.
[0415] The sending unit 34 is configured to send the first signaling to a client.
[0416] The receiving unit 31 is further configured to receive first request information determined by the client according to the first signaling, the first request information including identifier information of a target camera.
[0417] The sending unit 34 is further configured to send a media file corresponding to the target camera to the client according to the first request information.
[0418] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, details are not described here. Specifically, Figure 10 The apparatus 30 shown can perform the method embodiments corresponding to the server, and the foregoing and other operations and / or functions of each module in the apparatus 30 are respectively to realize the method embodiments corresponding to the server, and for the sake of brevity, details are not described here.
[0419] The apparatuses of the embodiments of the present application are described above from the perspective of functional modules in combination with the drawings. It should be understood that the functional modules can be realized in the form of hardware, or in the form of instructions of software, or in the form of a combination of hardware and software modules. Specifically, the steps of the method embodiments in the embodiments of the present application can be completed by integrated logic circuits of hardware in a processor and / or instructions of software. The steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware code processing performed by a processor, or be executed by a combination of hardware and software modules in a code processing processor. Alternatively, the software module can be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, and the like. The storage medium is located in a storage, and a processor reads information in the storage, and combines hardware to complete the steps in the above method embodiments.
[0420] Figure 11 is a schematic block diagram of a computing device provided by the embodiments of the present application, which can be the first device, a server or a client described above.
[0421] As shown in Figure 11 , the computing device 40 can include:
[0422] a memory 41 for storing a computer program, and a memory 42 for transmitting the program code. In other words, the memory 42 can call and run the computer program from the memory 41 to implement the method in the embodiments of the present application.
[0423] For example, the memory 42 can be used to execute the above method embodiments according to the instructions in the computer program.
[0424] In some embodiments of the present application, the memory 42 can include but is not limited to:
[0425] a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and the like.
[0426] In some embodiments of the present application, the memory 41 includes but is not limited to:
[0427] The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0428] In some embodiments of the present application, the computer program can be divided into one or more modules, which are stored in the memory 41 and executed by the memory 42 to complete the method provided by the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the video production device.
[0429] As shown in Figure 11 The computing device 40 can further include:
[0430] A transceiver 43, which can be connected to the memory 42 or the memory 41.
[0431] The memory 42 can control the transceiver 43 to communicate with other devices, specifically, can send information or data to other devices, or receive information or data sent by other devices. The transceiver 43 can include a transmitter and a receiver. The transceiver 43 can further include an antenna, and the number of antennas can be one or more.
[0432] It should be understood that the various components within the video production device are connected by a bus system, which includes, in addition to a data bus, a power supply bus, a control bus, and a status signal bus.
[0433] The application also provides a computer storage medium, which stores a computer program, and the computer program enables a computer to execute the method of the above method embodiments when the computer program is executed by the computer. Alternatively, the application embodiments also provide a computer program product containing instructions, and the instructions enable the computer to execute the method of the above method embodiments when the instructions are executed by the computer.
[0434] When implemented by using software, the computer program product can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the flow or function according to the embodiments of the application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital video disc (DVD)), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0435] Those skilled in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the application.
[0436] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the division of the above-described device embodiment is merely an example, and there can be other division manners. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the modules shown or discussed can be indirect coupling or communication connection through some interface, device or module, and can be electrical, mechanical or in other forms.
[0437] The modules described as separated components can or can not be physically separated, and the components displayed as modules can or can not be physical modules, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments. For example, the functional modules in the embodiments of the present application can be integrated into a processing module, or each module can be physically present separately, or two or more modules can be integrated into one module.
[0438] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for encapsulating free-viewpoint video files, characterized in that, Applied to the first device, including: Acquire the bitstream of free-view video data, wherein the free-view video data includes video data from N perspectives, where N is a positive integer; The bitstream of the free-viewpoint video data is encapsulated into at least one video track to obtain a media file of the free-viewpoint video data. The video track includes codec independence information and video bitstreams of M views. The codec independence information is used to indicate whether the video data of a single viewpoint in the M views corresponding to the video track depends on the video data of other views during encoding and decoding. M is a positive integer less than or equal to N. The media file containing the free-viewpoint video data is sent to the client or server.
2. The method according to claim 1, characterized in that, The video data includes at least one of texture map data and depth map data.
3. The method according to claim 2, characterized in that, If the value of the encoding / decoding independence indication information is a first value, it indicates that the texture map data of the single viewpoint depends on the texture map data and depth map data of other viewpoints during encoding / decoding, or that the depth map data of the single viewpoint depends on the texture map data and depth map data of other viewpoints during encoding / decoding; or, If the value of the encoding / decoding independence indication information is a second value, it indicates that the texture map data of the single viewpoint depends on the texture map data of other viewpoints during encoding / decoding, and the depth map data of the single viewpoint depends on the depth map data of other viewpoints during encoding / decoding; or, If the value of the encoding / decoding independence indication information is a third value, it indicates that the texture map data and depth map data of the single viewpoint do not depend on the texture map data and depth map data of other views during encoding / decoding, and that the texture map data and depth map data of the single viewpoint are interdependent during encoding / decoding; or, If the value of the encoding / decoding independence indication information is the fourth value, it indicates that the texture map data and depth map data of the single viewpoint do not depend on the texture map data and depth map data of other views during encoding / decoding, and the texture map data and depth map data of the single viewpoint do not depend on each other during encoding / decoding.
4. The method according to claim 3, characterized in that, The step of encapsulating the bitstream of the free-viewpoint video data into at least one video track includes: A single-track encapsulation mode is used to encapsulate the bitstream of the free-viewpoint video data into a single video track.
5. The method according to claim 4, characterized in that, If the video data corresponding to each of the N perspectives does not depend on the video data corresponding to other perspectives during encoding, and the bitstream of the free-viewpoint video data is encapsulated in a single-track encapsulation mode, then the method further includes: The codec independence indication information is added to the free-view information data box of the video track, and the value of the codec independence indication information is either the third value or the fourth value.
6. The method according to any one of claims 2-5, characterized in that, If the free-viewpoint video data is encoded in AVS3 encoding mode, then the method further includes: At least one of the header information required for decoding, texture map information of at least one view, and depth map information of at least one view is encapsulated in the media file as a subsample; The data box of the subsample includes a subsample data box identifier and subsample indication information. The subsample data box identifier is used to indicate the division method of the subsample, and the subsample indication information is used to indicate the content included in the subsample.
7. The method according to claim 6, characterized in that, If the value of the subsample indication information is the fifth value, then it indicates that a subsample includes the header information required for decoding; or, If the value of the subsample indication information is the sixth value, then it indicates that a subsample includes texture map information from N viewpoints within the current video frame, and the current video frame is composed of video frames corresponding to the N viewpoints; or, If the value of the subsample indication information is the seventh value, then it indicates that a subsample includes the depth map information corresponding to N viewpoints in the current video frame; or, If the value of the subsample indication information is the eighth value, then it indicates that a subsample includes the texture map information and depth map information corresponding to a viewpoint within the current video frame; or, If the value of the subsample indication information is the ninth value, then it indicates that a subsample includes the texture map information corresponding to a viewpoint within the current video frame; or, If the value of the subsample indication information is the tenth value, then it indicates that a subsample includes the depth map information corresponding to a viewpoint within the current video frame.
8. The method according to claim 7, characterized in that, If the encoding method of the free-viewpoint video data is AVS3 encoding mode, and the video data corresponding to each of the N views does not depend on the video data corresponding to other views during encoding, and the encapsulation mode of the free-viewpoint video data stream is single-track encapsulation mode, then encapsulating at least one of the header information required for decoding, texture map information of at least one view, and depth map information of at least one view in the media file in the form of subsamples includes: At least one of the texture map information and depth map information corresponding to each of the N viewpoints is encapsulated in the media file as a subsample.
9. A method for encapsulating free-viewpoint video files, characterized in that, Applied to the client side, including: A media file receiving free-viewpoint video data sent by a first device, the media file including at least one video track, the free-viewpoint video data including video data from N perspectives, where N is a positive integer, the video track including codec independence indication information and video bitstreams from M perspectives, the codec independence indication information being used to indicate whether the video data from a single perspective among the M perspectives corresponding to the video track depends on the video data from other perspectives during encoding and decoding, where M is a positive integer less than or equal to N; Based on the encoding / decoding independence indication information, the media file is decapsulated to obtain a video stream corresponding to at least one viewpoint; Decode the video stream corresponding to the at least one viewpoint to obtain the reconstructed video data of the at least one viewpoint.
10. The method according to claim 9, characterized in that, If the value of the encoding / decoding independence indication information is a first value, it indicates that the texture map data of the single viewpoint depends on the texture map data and depth map data of other viewpoints during encoding / decoding, or that the depth map data of the single viewpoint depends on the texture map data and depth map data of other viewpoints during encoding / decoding; or, If the value of the encoding / decoding independence indication information is a second value, it indicates that the texture map data of the single viewpoint depends on the texture map data of other viewpoints during encoding / decoding, and the depth map data of the single viewpoint depends on the depth map data of other viewpoints during encoding / decoding; or, If the value of the encoding / decoding independence indication information is a third value, it indicates that the texture map data and depth map data of the single viewpoint do not depend on the texture map data and depth map data of other views during encoding / decoding, and that the texture map data and depth map data of the single viewpoint are interdependent during encoding / decoding; or, If the value of the encoding / decoding independence indication information is the fourth value, it indicates that the texture map data and depth map data of the single viewpoint do not depend on the texture map data and depth map data of other views during encoding / decoding, and the texture map data and depth map data of the single viewpoint do not depend on each other during encoding / decoding.
11. The method according to claim 10, characterized in that, Based on the encoding / decoding independence indication information, the media file resources are decapsulated to obtain a video stream corresponding to at least one viewpoint, including: If the value of the encoding / decoding independence indication information is the third value or the fourth value, and the encapsulation mode of the bitstream is a single-track encapsulation mode, then the user's viewing perspective is obtained. The target viewing angle is determined based on the user's viewing angle and the viewing angle information in the media file; The media file corresponding to the target viewpoint is decapsulated to obtain the video stream corresponding to the target viewpoint.
12. The method according to claim 11, characterized in that, If the free-viewpoint video data is encoded using AVS3 video encoding, and the media file also includes sub-samples, then the step of decapsulating the media file resources according to the encoding / decoding independence indication information to obtain a video stream corresponding to at least one viewpoint includes: Obtain the subsample data box identifier and the subsample indication information included in the data box of the subsample. The subsample data box identifier is used to indicate the division method of the subsample. The subsample indication information is used to indicate the content included in the subsample. The content included in the subsample includes at least one of the following: header information required for decoding, texture map information of at least one view, and depth map information of at least one view. The contents included in the subsample are obtained based on the subsample data box identifier and the subsample indication information; Based on the content included in the subsample and the encoding / decoding independence indication information, the media file resource is decapsulated to obtain a video stream corresponding to at least one viewpoint.
13. The method according to claim 12, characterized in that, If the media file includes sub-samples corresponding to each of the N viewpoints, then the step of decapsulating the media file resource according to the content included in the sub-samples and the codec independence indication information to obtain a video stream corresponding to at least one viewpoint includes: If the value of the encoding / decoding independence indication information is the third value or the fourth value, and the encapsulation mode of the bitstream is single-track encapsulation mode, then the target viewing angle is determined based on the user's viewing angle and the viewing angle information in the media file. Obtain the subsample data box identifier and the subsample indication information included in the data box of the target subsample corresponding to the target viewpoint; The contents of the target subsample are determined based on the subsample data box identifier and the subsample indication information included in the data box of the target subsample; Based on the content included in the target subsample, the media file corresponding to the target viewpoint is decapsulated to obtain the video stream corresponding to the target viewpoint.
14. A method for encapsulating free-viewpoint video files, characterized in that, Applied to servers, including: A media file receiving free-viewpoint video data sent by a first device, the media file including at least one video track, the free-viewpoint video data including video data from N perspectives, where N is a positive integer, the video track including codec independence indication information and video bitstreams from M perspectives, the codec independence indication information being used to indicate whether the video data from a single perspective among the M perspectives corresponding to the video track depends on the video data from other perspectives during encoding and decoding, where M is a positive integer less than or equal to N; Based on the encoding / decoding independence indication information, determine whether to decompose the at least one video track into multiple video tracks.
15. A processing apparatus for multi-view video data, characterized in that, Applied to a first device, the device includes: An acquisition unit is used to acquire the bitstream of free-view video data, wherein the free-view video data includes video data from N perspectives, where N is a positive integer. An encapsulation unit is used to encapsulate the bitstream of the free-viewpoint video data into at least one video track to obtain a media file of the free-viewpoint video data. The video track includes codec independence indication information and video bitstreams of M views. The codec independence indication information is used to indicate whether the video data of a single viewpoint in the M views corresponding to the video track depends on the video data of other views during encoding and decoding. M is a positive integer less than or equal to N. The sending unit is used to send the media file of the free-viewpoint video data to the client or server.
16. A processing apparatus for multi-view video data, characterized in that, Applied to a client, the device includes: A receiving unit is configured to receive a media file of free-viewpoint video data sent by a first device. The media file includes at least one video track. The free-viewpoint video data includes video data from N perspectives, where N is a positive integer. The video track includes codec independence indication information and video streams from M perspectives. The codec independence indication information is used to indicate whether the video data from a single perspective among the M perspectives corresponding to the video track depends on the video data from other perspectives during encoding and decoding. The video data includes at least one of texture map data and depth map data, where M is a positive integer less than or equal to N. The decapsulation unit is used to decapsulate the media file according to the encoding / decoding independence indication information to obtain a video stream corresponding to at least one viewpoint; The decoding unit is used to decode the video stream corresponding to the at least one viewpoint to obtain the reconstructed video data of the at least one viewpoint.
17. A processing apparatus for multi-view video data, characterized in that, Applied to a server, the device includes: A receiving unit is configured to receive a media file of free-viewpoint video data sent by a first device. The media file includes at least one video track. The free-viewpoint video data includes video data from N perspectives, where N is a positive integer. The video track includes codec independence indication information and video streams from M perspectives. The codec independence indication information is used to indicate whether the video data from a single perspective among the M perspectives corresponding to the video track depends on the video data from other perspectives during encoding and decoding, where M is a positive integer less than or equal to N. The decomposition unit is used to determine, based on the encoding / decoding independence indication information, whether to decompose the at least one video track into multiple video tracks.
18. A computing device, characterized in that, include: A processor and a memory, the memory for storing a computer program, the processor for calling and running the computer program stored in the memory to perform the method of any one of claims 1 to 8 or 9 to 13 or 14.
19. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method as described in any one of claims 1 to 8 or 9 to 13 or 14.
20. A computer program product, characterized in that, Includes computer program instructions that cause a computer to perform the method as described in any one of claims 1 to 8 or 9 to 13 or 14.