Data processing method and storage medium

The data processing method improves media data efficiency and flexibility by encapsulating texture and depth maps into media tracks, enabling efficient and immersive viewpoint switching in immersive media environments.

JP7862547B2Active Publication Date: 2026-05-19ZTE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ZTE CORP
Filing Date
2023-03-01
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently processing and transmitting media data for flexible viewpoint switching and scene display, particularly in immersive media environments, leading to limitations in user experience and immersion.

Method used

A data processing method that involves acquiring free-view video data, encapsulating texture and depth maps into media tracks, and processing this data based on target view information, utilizing the ISO Base Media File Format for efficient storage and transmission.

Benefits of technology

Enhances data processing efficiency and flexibility, allowing for seamless and immersive viewpoint switching in real-time, reducing the need for multiple data collection equipment and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007862547000001
    Figure 0007862547000001
  • Figure 0007862547000002
    Figure 0007862547000002
  • Figure 0007862547000003
    Figure 0007862547000003
Patent Text Reader

Abstract

Embodiments of the present application provide a data processing method, apparatus, device, computer-readable storage medium, and computer program product, the data processing method, the device, the computer-readable storage medium, and the computer program product comprising: acquiring free view video data including at least one texture map and / or depth map; encapsulating the at least one texture map and / or depth map into at least one media track; and processing the free view video data based on target view information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application is filed based on a Chinese patent application with application number 202210249027.9 and filing date of March 14, 2022, and claims the priority of the Chinese patent application. All the contents of the Chinese patent application are incorporated herein by reference into this application.

[0002] The embodiments of this application relate to the technical field of computers, particularly data processing methods, devices, equipment, storage media, and program products.

Background Art

[0003] With the development of computer technology, users hope to obtain a stronger sense of presence in scenes such as video playback and virtual games by freely switching viewpoints.

[0004] In related technologies, since the efficiency of media data processing and transmission is low, it is difficult for users to perform flexible viewpoint switching and scene display. How to improve data processing efficiency and the user experience has now become an urgent task for consideration and solution.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The embodiments of this application provide a data processing method, device, equipment, computer-readable storage medium, and computer program product aimed at improving data processing efficiency.

Means for Solving the Problems

[0006] In a first aspect, the embodiments of this application include: a step of acquiring free-view video data including at least one collected texture map and / or depth map; a step of encapsulating the at least one texture map and / or depth map into at least one media track; The present invention provides a data processing method that includes the step of processing the free view video data based on target view information.

[0007] In the second aspect, the embodiment of the present application is An acquisition module configured to acquire free-view video data including at least one texture map and / or depth map, and target view information, An encapsulation module configured to encapsulate the at least one texture map and / or depth map in at least one media track, The present invention provides an image display device that includes a processing module configured to process the free-view video data based on the target view information.

[0008] In the third aspect, the embodiments of the present application are as follows: The present invention provides a computer-readable storage medium that stores computer-executable instructions for performing the data processing method described in the first embodiment.

[0009] In the fourth aspect, the embodiments of the present application are as follows: A computer program product that includes a computer program or computer instructions, The computer program or computer instruction is stored in a computer-readable storage medium, and the processor of the computer device reads the computer program or computer instruction from the computer-readable storage medium and executes the computer program or computer instruction, thereby causing the computer device to execute the data processing method described in the first embodiment. [Brief explanation of the drawing]

[0010] [Figure 1] This is a schematic diagram illustrating how a user views live video via a display device. [Figure 2]This is a schematic diagram illustrating how a user views video through another display device. [Figure 3] This is a schematic diagram of the system architecture for an application scenario of a data processing method according to one embodiment of the present invention. [Figure 4] This is a flowchart of a data processing method according to one embodiment of the present invention. [Figure 5] This is a flowchart of a data processing method according to one embodiment of the present invention. [Figure 6] This is a flowchart of a data processing method according to one embodiment of the present invention. [Figure 7] This is a diagram illustrating the file structure of single-track encapsulated free-view video data according to one embodiment of the present invention. [Figure 8] This is a diagram illustrating the structure of a media file for free-view video based on an extension of the single-track encapsulation mode according to one embodiment of the present invention. [Figure 9] This is a flowchart illustrating the processing of free-view video data based on single-track encapsulation according to one embodiment of the present invention. [Figure 10] This is a diagram illustrating the file structure of multi-track encapsulated free-view video data according to one embodiment of the present invention. [Figure 11] This is a diagram illustrating the structure of a media file for FreeView video based on the multi-track encapsulation mode extension according to one embodiment of the present invention. [Figure 12] This is a flowchart illustrating the processing of free-view video data based on multi-track encapsulation according to one embodiment of the present invention. [Figure 13] This is a schematic diagram of media track grouping according to one embodiment of the present invention. [Figure 14] This is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention. [Figure 15] This is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention. [Figure 16] This is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention. [Figure 17]It is a structural schematic diagram of a data processing device according to an embodiment of the present application. [Figure 18] It is a structural schematic diagram of a data processing device according to an embodiment of the present application. [Figure 19] It is a structural schematic diagram of a data processing device according to an embodiment of the present application. **Embodiments for Carrying Out the Invention**

[0011] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be described in more detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described in this specification are only used to explain the present application and are not used to limit the present application.

[0012] Note that the functional module division is shown in the schematic diagram of the device, and the logical order is shown in the flowchart. However, in some cases, it may be different from the module division of the device, or the steps shown or described may be executed in an order different from that shown in the flowchart. Terms such as "first", "second", etc. in the specification, claims, and the aforementioned drawings are not used to explain a specific order or priority, but are used to distinguish similar objects.

[0013] In the description of the embodiments of the present application, unless otherwise clearly defined, terms such as "provide", "attach", "connect", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in the embodiments of the present application in combination with the specific content of the technical solution. In the embodiments of the present application, words such as "further", "exemplarily", or "optionally" are used to indicate examples, illustrations, or exemplifications, and should not be construed as being more preferred or advantageous than other embodiments or design solutions. The use of words such as "further", "exemplarily", or "optionally" is intended to present related concepts in a specific manner.

[0014] The embodiments of this application can be applied to various devices related to the display of images and videos, such as mobile phones, tablets, computers, laptops, wearable devices, in-vehicle devices, liquid crystal displays, cathode ray tube monitors, holographic displays, or terminal equipment such as projectors. They can also be applied to various devices that process image and video data, such as server equipment such as mobile phones, tablets, computers, laptops, wearable devices, and in-vehicle equipment. The embodiments of this application are not limited to these.

[0015] Immersive media uses technologies such as video and audio to provide users with highly realistic virtual environments in terms of sight, sound, and other aspects, creating a sense of being present in that location. Related technologies allow users to view 360-degree video by freely rotating their heads after wearing a head-mounted display device, enabling a 3-degrees-of-freedom (3DOF) immersive experience. Furthermore, to provide an even better visual experience, related technologies can support enhanced 3-degrees-of-freedom (3DOF+) and 6-degrees-of-freedom (6DOF) video, meaning that the user's head and body can move within the space either limited or freely, enabling free switching of views and achieving a more realistic sense of immersion.

[0016] Taking 6DOF video as an example, there are many implementation methods for acquiring and playing back 6DOF video, and freeview video is one of them. Freeview video is typically a collection of videos from different views acquired from a multi-camera matrix array shot facing the same 3D scene. While viewing freeview video, users can freely switch views, viewing the corresponding video images in either a real-world or a synthesized virtual viewpoint. Freeview video data includes multiview texture maps and depth maps, allowing the user to directly view the real-world video acquired by the cameras when switching to a real-world viewpoint, and to synthesize a virtual viewpoint texture map corresponding to the virtual viewpoint in real time when switching to a virtual viewpoint. The position and number of cameras in the freeview video data acquisition process, and the selection, transmission, and processing of video data corresponding to the cameras in the real-time virtual viewpoint image synthesis process, each have different degrees of impact on the time consumption and quality of data synthesis, which directly affects the user experience.

[0017] Freeview video is a new virtual reality (VR) video technology that generally uses multiple cameras to capture the surroundings of a target scene and uses virtual view synthesis technology to obtain images of a virtual view. By using freeview video technology, users can view the target scene from any view they choose, providing a superior viewing experience compared to panoramic video. However, currently, it is not possible to provide users with freeview video of a target scene while the scene is being live-streamed. Therefore, how to allow users to view some of the best videos in a live video stream from their desired view is a technical challenge that those skilled in the art should urgently address.

[0018] To further explain this proposed technology, we will use Figures 1 and 2 as examples to describe related technologies.

[0019] Figure 1 is a schematic diagram of a user viewing live video via a display device. As shown in Figure 1, user 200 views live content via the first display device 100. However, user 200 can only see the live scene content from the view currently being played by the display device, and it is difficult for them to adjust the viewing view according to their preferences.

[0020] Figure 2 is a schematic diagram of a user viewing video via another display device. As shown in Figure 2, user 200 views video content via a second display device 300. User 200 can switch viewing views by rotating their head via the second display device 300, and also by using the remote control 310 of the second display device 300. This method provides a three-dimensional image and allows the user to adjust the viewing angle as desired, but it has limitations in image selection, low transmission and processing efficiency, and is difficult to fully satisfy the application scenes of 3DOF+ and 6DOF. At the same time, it is difficult to apply this technology when displaying or compositing images in real time.

[0021] Embodiments of the present invention provide a data processing method, apparatus, device, computer-readable storage medium, and computer program product that improve data processing efficiency, make data encapsulation and transmission more flexible, and further improve the processing efficiency of free-view video by encapsulating texture maps and depth maps in a media track.

[0022] The embodiments of this application will be further described below with reference to the drawings.

[0023] Figure 3 is a schematic diagram of the system architecture of an application scenario of a data processing method according to one embodiment of the present invention. As shown in Figure 3, this system architecture includes a data processing device 400 and a data acquisition device.

[0024] In this embodiment, two data acquisition devices, a first image acquisition device 510 and a second image acquisition device 520, are used as examples. The first image acquisition device 510 collects image data of the first region A, and the second image acquisition device 520 collects image data of the second region B.

[0025] In this embodiment, the free-view video data is collected by the first image acquisition device 510 and the second image acquisition device 520.

[0026] The data processing device 400 acquires free-view video data.

[0027] In one embodiment, the free-view video data includes a first texture map collected by a first image acquisition device 510 and a second texture map collected by a second image acquisition device 520.

[0028] In one embodiment, the free-view video data includes a first depth map collected by a first image acquisition device 510 and a second depth map collected by a second image acquisition device 520.

[0029] In another embodiment, the free-view video data includes a first texture map and a first depth map from the first image acquisition device 510, and a second texture map and a second depth map from the second image acquisition device 520.

[0030] Texture maps and depth maps may be collected by a data collection device, or they may be obtained by other methods such as computer generation.

[0031] The first view 610 is a view of the first region A, and the second view 620 is a view of the second region B.

[0032] In one embodiment, when viewing a first region A, the data processing device 400 acquires a first texture map collected by a first image acquisition device 510, encapsulates and processes it using an encapsulation module and a processing module to obtain the current free-view video data. When viewing a second region B, the data processing device 400 acquires a second texture map collected by a second image acquisition device 520, encapsulates and processes it using an encapsulation module and a processing module to obtain the current free-view video data. Using the technical proposal of this embodiment, the viewing experience can be improved by switching to the desired view quickly and freely in real time.

[0033] In another embodiment, if it is desirable to view the third region C in a third view 630 (i.e., a virtual viewpoint), the data processing device 400 first determines the first region A and / or the second region B adjacent to the third region C, that is, the first view 610 or the second view 620 adjacent to the third view 630, then synthesizes the first depth map collected by the acquisition device 510 corresponding to the first view 610 and / or the second depth map collected by the acquisition device 520 corresponding to the second view 620 to obtain a third depth map corresponding to the third view 630, and finally combines the first texture map collected by the acquisition device 510 corresponding to the first view 610 and / or the second texture map collected by the acquisition device 520 corresponding to the second view 620 with the third depth map obtained by the synthesis process to synthesize a third texture map corresponding to the third view 630 (i.e., a virtual viewpoint) and obtain the current free view video data. Using the technology described in this embodiment, not only is the amount of data collection equipment used reduced, but the view can also be freely switched in real time, resulting in a more realistic and immersive experience.

[0034] In another embodiment, texture maps and depth maps can also be obtained by other means. For example, depth maps can be obtained by texture map calculations, and simulation data for both texture maps and depth maps can be obtained by computer synthesis.

[0035] Figure 4 is a flowchart of a data processing method according to one embodiment of the present invention. As shown in Figure 4, this data processing method may be used for both terminals and servers. In the embodiment of Figure 4, this data processing method may include, but is not limited to, steps S1000, S2000, and S3000.

[0036] Step S1000: Obtain free view video data.

[0037] In one embodiment, the free-view video data includes a texture map.

[0038] In one embodiment, the freeview video data includes a depth map.

[0039] In one embodiment, the free-view video data includes a texture map and a depth map.

[0040] Texture maps and depth maps may be collected by a data collection device, or they may be obtained by other methods such as computer generation. The following explanation will use data collection by a data collection device as an example.

[0041] Freeview video is a collection of videos from different views collected from a multi-camera matrix array shot toward the same three-dimensional scene. Freeview video data is typically collected by multiple acquisition devices, with devices positioned at different angles and locations to collect images of the scene or subject to be filmed. The acquisition devices may be any devices capable of image acquisition, such as cameras. There may be one or more acquisition devices, and the number of acquisition devices is not limited in the embodiments of this application.

[0042] Freeview video can support the compositing of video images in freeview mode, and its video data typically includes a texture map and a depth map. The texture map describes data about the surface properties of the subject. By rendering the texture map, a realistic video image can be obtained. The depth map, also called a distance image, directly reflects the geometric shape of the visible surface of an object. While a depth map is similar to a grayscale image, each pixel value in the depth map represents the actual distance from the object to the camera. In some application scenarios, such as when the device's memory and computing power are limited, or when the degree of freedom in view switching is limited, the freeview video data may include only either a texture map or a depth map.

[0043] In one embodiment, if there is no acquisition device corresponding to the view after the view switch, virtual view synthesis is performed. In the virtual view synthesis reconstruction process, after determining multiple cameras near the view after the view switch, the depth maps of the cameras are processed, and based on the processed depth maps, the texture maps of the cameras are combined to synthesize a texture map corresponding to the virtual view. Therefore, in the application scene corresponding to this embodiment, the free view video includes texture map data and depth map data.

[0044] In one embodiment, the degree of freedom for viewpoint switching is limited; that is, when the camera switches between corresponding real viewpoints without compositing a virtual viewpoint, or when the acquisition device does not have depth map acquisition capability, the free view video in the application scene corresponding to this embodiment includes only texture map data and does not include depth map data. Alternatively, if there is an acquisition device corresponding to the switched view, a texture map corresponding to the acquisition device is used without compositing a virtual viewpoint using a depth map.

[0045] Step S2000: Encapsulate at least one texture map and / or depth map into at least one media track.

[0046] To support the encapsulated storage, transmission, and processing of freeview video data in media systems, it is necessary to define a rational format structure for freeview video data within media files.

[0047] In the embodiments of this invention, one embodiment involves storing free-view video data in a file based on the International Organization for Standardization Base Media File Format (ISOBMFF). In application scenes with restricted schemes, i.e., when it is necessary to synthesize virtual viewpoints, the ISO Base Media File Format, including information boxes, track reference boxes, and track group boxes, can refer to the MPEG-4 Part 12 ISO Base Media File Format developed by the Moving Picture Experts Group (MPEG) of ISO / IEC JTC1 / SC29 / WG11.

[0048] Based on the ISO Basic Media File Format, all data is encapsulated in boxes. The ISO Basic Media File Format consists of several boxes, each with its own type and length, which can be considered a single data object. Boxes that can contain other boxes are called container boxes. Taking an MP4 file as an example, an MP4 file has one box of type "ftyp," which is the file format identifier and contains some information about the file. An MP4 file also has one box of type "MOOV" (Movie Box), which is a container box, and its subboxes contain media metadata information, which is information to describe the content of the media. The media data of an MP4 file is contained in boxes of type "mdat" (Media Data Box), which are also container boxes, and there may be multiple mdat boxes. If all the media data references other files, there may be no mdat boxes. The structure of the media data is described by metadata, and some general or additional non-timing metadata can be described using boxes of type "meta" (Meta box), which are also container boxes.

[0049] The timing metadata track is an ISOBMFF mechanism for establishing timing metadata associated with a specific sample. Timing metadata is generally "descriptive" and has less coupling with media data.

[0050] In typical application scenarios, free-view video data includes texture map data and depth map data, and virtual viewpoints are often composited based on views that can be freely switched during processing. Therefore, free-view video media tracks should be represented in the media file as restricted video. In the embodiments of this invention, the free-view video media track is described using the SchemeTypeBox of the RestrictedSchemeInfoBox defined in the ISO / IEC 14496-12 standard, where the scheme_type of the SchemeTypeBox is set to "as3f".

[0051] Freeview video tracks are defined in ISO / IEC 19946-12 using VisualSampleEntry. A freeview video media track is a texture map media track if it contains only texture maps. A depth map media track is of type "auxv" using the handler type of the HandlerBox in MediaBox, which means that depth map information is included in the depth map track.

[0052] Furthermore, if the free view video contains only texture map data, the virtual viewpoint compositing process does not occur, and the free view video media track is not represented in the media file as a restricted video; in other words, the free view video media track is not described using the SchemeTypeBox mentioned above.

[0053] In one embodiment, the texture map and depth map collected by the acquisition device are encapsulated in a single media track, i.e., a single media track contains both the texture map and the depth map.

[0054] In one embodiment, the texture maps collected by the acquisition device are encapsulated in a single media track, i.e., a single media track contains the texture maps.

[0055] In one embodiment, the depth maps collected by the acquisition device are encapsulated in a single media track, i.e., one media track contains the depth maps.

[0056] In the above embodiment, the method of encapsulating video data using a single media track may be collectively referred to as single-track encapsulation. Single-track encapsulation is suitable for free-view video data, typically when there are not many acquisition devices and the number of pixels in the texture map of a single frame is not large. In this case, when a terminal requests the acquisition of the texture map and / or depth map of the corresponding acquisition device, the server transmits the texture maps and / or depth maps of all cameras in the media file to the terminal, and the terminal can pre-download the texture maps and / or depth maps of all acquisition devices. When deencapsulating and decoding, the terminal selects the texture maps and / or depth maps of the acquisition devices related to the view after the view has been freely switched, i.e., the acquisition devices for which virtual view synthesis is possible. In this way, the terminal can directly retrieve the texture maps and / or depth maps corresponding to the required acquisition devices locally and decode the texture maps and depth maps corresponding to all or the required related acquisition devices.

[0057] In one embodiment, the texture maps and depth maps collected by the acquisition device are encapsulated in multiple media tracks, i.e., one or more media tracks contain the texture maps and one or more other media tracks contain the depth maps.

[0058] In one embodiment, the texture maps collected by the acquisition device are encapsulated in multiple media tracks, i.e., each media track contains a texture map.

[0059] In one embodiment, the depth map collected by the acquisition device is encapsulated in multiple media tracks, i.e., each media track contains the depth map.

[0060] In the above embodiment, the method of encapsulating video data using multiple media tracks may be collectively referred to as multi-track encapsulation. Unlike single-track encapsulation, multi-track encapsulation is typically applied when there are many acquisition devices, a large number of pixels in a single frame's texture map, and a large amount of data. By encapsulating the texture map and depth map collected by the acquisition devices into different media tracks, texture map data and depth map data from one or more acquisition devices are encapsulated in each media track, depending on different encoding schemes, access requests, and transmission capabilities. The media track that encapsulates texture map data is the texture map media track, and the media track that encapsulates depth map data is the depth map media track.

[0061] The entity that encapsulates the texture map and / or depth map into one or more media tracks may be any processing device with encapsulation capabilities.

[0062] Step S3000: Process the free view video data based on the target view information.

[0063] After the user switches views, the terminal can acquire target view information, which usually includes information such as viewing position and direction, reflecting the free view desired by the user or the playback view desired by the director. Based on this target view information, the terminal can determine whether virtual viewpoint synthesis is necessary and the original video data from the acquisition device required for synthesis.

[0064] In one embodiment, the target view information includes an identifier for the collection device, parameter information for the collection device, position information for the target view, rotation information for the target view, and an identifier for the target view.

[0065] In one embodiment, since the terminal has high data processing capabilities, the terminal itself processes the free-view video data based on the target view information and displays the processed free-view video data.

[0066] Figure 5 is a flowchart of a data processing method according to one embodiment of the present invention. As shown in Figure 5, the data processing process of the terminal may include, but is not limited to, steps S1100, S3110, and S3120.

[0067] Step S1100: The terminal requests the server to retrieve the media file.

[0068] Step S3110: The terminal decrypts / decapsulates the acquired media file and synthesizes a virtual viewpoint texture map based on the free view video data, corresponding metadata information, and target view information within the media file.

[0069] Step S3120: The terminal displays the synthesized virtual viewpoint texture map.

[0070] In one embodiment, since the terminal does not have strong data processing capabilities, the server may process the free-view video data based on target view information, such as by compositing a virtual viewpoint texture map, and send the processed video data to the terminal for display.

[0071] Figure 6 is a flowchart of a data processing method according to one embodiment of the present invention. As shown in Figure 5, the data processing process of the terminal and server will be specifically described and may include, but is not limited to, steps S4000, S3210, and S3220.

[0072] Step S4000: The terminal sends target view information to the server.

[0073] Step S3110: The server retrieves the stored media files based on the target view information, extracts the free view video data and corresponding metadata information, and synthesizes the virtual view texture map.

[0074] Step S3120: The terminal receives the free-viewpoint video data processed by the server and displays the composite virtual view texture map.

[0075] In the embodiments corresponding to Figures 5 and 6, the data processing method according to the present invention is used regardless of whether the terminal or the server performs the synthesis of the virtual view texture map, and the processed texture map and / or depth map are encapsulated in at least one media track.

[0076] Figure 7 shows the file structure of single-track encapsulated FreeView video data. As shown in Figure 7, a media file based on the ISO Basic Media File Format has one media track, which contains a MOOV box and a madt box. The MOOV box contains a FreeViewExtInfoBox subbox and a TrackBox subbox, which encapsulate the media's metadata information, and the madt box encapsulates the FreeView video data in the form of a sample.

[0077] In another example, the MOOV box contains a TrackBox subbox, which in turn contains a FreeViewExtInfoBox subbox.

[0078] Single-track encapsulated freeview video data includes texture map and depth map data collected by multiple original acquisition devices. All of this texture map and depth map data is stored in a single media track; that is, each frame image in the media track contains the texture map and / or depth map of all original acquisition devices at a given point in time. The texture map and depth map corresponding to each original acquisition device are codecized individually, meaning there is no codec dependency on the texture map and depth map data between acquisition devices. When a terminal acquires single-track encapsulated freeview video, all the texture map and depth map data from the original acquisition devices are transmitted.

[0079] The terminal, after acquiring a media file containing freeview video data, selects and decodes the texture maps and depth maps of several original acquisition devices in the view vicinity necessary to synthesize a virtual viewpoint for each frame image. Each sample in a single-track encapsulated media track contains the texture maps and / or depth maps collected by all original acquisition devices at a given point in time, and by dividing the sample into multiple subsamples, each subsample contains the texture map and / or depth map corresponding to one acquisition device. More specifically, the information of the subsamples is described by the SubSampleInformationBox box of SampleTableBox or TrackFragmentBox, and the use of the subsamples must be carried out according to the flags field in the subsample information box (the flags field can identify the three data types stored in the subsample, namely texture maps and / or depth maps), and in particular, the specific configuration structure of each sample corresponding to the freeview video data is described by codec_specific_parameters, and two specific examples are shown below.

[0080] Example 1: The flags indicate the type of information encapsulated in the subsample, and the specific types are as follows: The flags value is 0, and it is based on subsamples of the texture map and depth map corresponding to the original collector, with each subsample containing either a texture map or depth map corresponding to one original collector. The flags value is 1, and it is based on a subsample of the original collector, with one subsample containing a texture map and depth map corresponding to one original collector. if (flags==0){ unsigned int(8) camera_id; unsigned int(1) payload_type; unsigned int(1) All_data; / / Optional bit(21) reserved = 0; } else if (flags==1){ unsigned int(8) camera_id; bit(21) reserved = 0; }

[0081] Example 2: The flags indicate the type of information encapsulated in the subsample, and the specific types are as follows: The flags value is 0, and it is based on a subsample of the data type (original map, depth map), with one subsample containing a texture map or depth map corresponding to one original collection device. The flags value is 1, and it is based on a subsample of the original collector, with one subsample containing the texture map and depth map corresponding to one original collector. if (flags==0){ unsigned int(16) camera_id; bit(16) reserved = 0;} else if (flags==1){ unsigned int(8) payload_type; bit(24) reserved = 0; } Here, camera_id indicates the acquisition device (camera) identifier corresponding to the texture map or depth map in the current subsample. `payload_type` indicates the data type included in the subsample; a value of 0 indicates that it includes a texture map, a value of 1 indicates that it includes a depth map, and a value of 2 indicates that it includes parameter information such as SEI information. All_data indicates that the subsamples contain texture maps or depth maps corresponding to all original acquisition devices.

[0082] The virtual viewpoint synthesis process preferably involves processing the depth map of a selected original acquisition device. To further simplify the encapsulation structure and decoding process, each sample in the media track includes a texture map subsample containing the texture maps of multiple original acquisition devices, a depth map subsample containing the depth maps of multiple original acquisition devices, and a depth map subsample containing the depth maps of all original acquisition devices. The specific subsample structure may be described using the codec_specific_parameters defined in this embodiment.

[0083] In the definition shown in Example 1, the subsample structure information is described using one SubSampleInformationBox, and if flags is 0, camera_id is set to 0, meaning that the subsample contains the texture maps or depth maps of all original acquisition devices. In the definition shown in Example 2, the subsample structure information may be described using two SubSampleInformationBoxes, where the flags value in one of the SubSampleInformationBoxes is 0, camera_id is set to 0, indicating that the depth maps or texture maps of all original acquisition devices are contained within one subsample, and the flags value in the other SubSampleInformationBox is 1, where a flags value of 1 indicates that the subsample contains texture maps or depth maps.

[0084] SubSampleInformationBox can selectively partially decode the texture map and depth map of the original camera necessary to synthesize a virtual viewpoint, based on the size of each subsample in the sample, the corresponding original acquisition device, and the type of image data (texture map or depth map). This improves data processing efficiency and expands the application scenes of free view video.

[0085] In one embodiment, the user's view is simply switched between real viewpoints corresponding to the original acquisition device during a free-switching process, in which case there is no process of compositing virtual viewpoints. In this case, the single-track encapsulated free-view video data is typically the texture map data of the original acquisition device, and there are no depth maps in the subsamples, i.e., the payload_type value is always 0.

[0086] In some embodiments, when the amount of freeview video data is very large, i.e., when there are many original acquisition devices, a large number of pixels in the texture map being acquired, or when grouping relationships exist between the original acquisition devices, the freeview video data may be encapsulated using an extended single-track encapsulation method.

[0087] Figure 8 is a diagram illustrating the structure of a media file for FreeView video based on the single-track encapsulation mode extension. As shown in Figure 8, a media file based on the ISO basic media file format has multiple media tracks, each containing a MOOV box and a madt box. The MOOV box contains a FreeViewExtInfoBox subbox and a SubSampleInfoBox subbox, encapsulating the media's metadata information, while the madt box encapsulates the FreeView video data in the form of a sample.

[0088] In the extended single-track encapsulated media file structure, texture maps and depth maps have a one-to-one correspondence, each sample in a track contains texture maps and depth maps corresponding to multiple original acquisition devices at a given point in time, and information for subsamples is described by a box in SubSampleInformationBox within SampleTableBox or TrackFragmentBox, described using the codec_specific_parameters syntax element defined in the above embodiment, and each subsample contains texture maps and / or depth maps for one or more original acquisition devices.

[0089] In one embodiment, when the texture map and depth map encapsulated in the media track use a video encoding scheme such as AVS2, AVC, or HEVC, and there is an extension or supplement to the encoding information of the free view video, such extension / supplement encoding information may be stored directly in the media track sample, or specific extension / supplement encoding information or other parameter information may be described by defining a free view extension box contained in a MediaInformationBox or VisualSampleEntry (as in ISO / IEC 14496-12), two specific examples of which are shown below.

[0090] Example 1: The media track for single-track encapsulated FreeView video uses a sample entry, VisualSampleEntry, which contains two boxes: a FreeView ExtInfoBox and a Payload InfoBox. The FreeViewExtInfoBox describes the original acquisition device information corresponding to the FreeView video data in the media track, and by associating the view with the acquisition device, camera information can be treated as view information. The specific syntactic semantics are as follows: Box Type: 'fvei' Container: MediaInformationBox or VisualSampleEntry Mandweather: No Quantity: Zero or one. aligned(8) class FreeViewInfoBox extends FullBox(' fvin', 0, 0){ unsigned int(1) texture_in_track; unsigned int(1) depth_in_track; unsigned int(16) num_cameras; for (i=0; i <num_cameras; i++) { unsigned int(16) camera_id; unsigned int(1) int_camera_flag; unsigned int(1) ext_camera_flag; if(int_camera _flag) IntCameraInfoStruct(); if(Ext_camera _flag) ExtCameraInfoStruct();}} The PayloadInfoBox contains the TextureInfostruct() and DepthInfostruct() texture map information corresponding to the original camera / view of the free view video in the media track. The specific syntactic semantics are as follows: aligned(8) class PayloadInfoBox extends FullBox(' plin', 0, 0){ unsigned int(1) texture_info_flag; unsigned int(1) depth_info_flag; unsigned int(16) num_cameras; for (i=0; i <num_cameras; i++) { unsigned int(16) camera_id; if(texture_info_flag) TextureInfostruct(); if(depth_info_flag) DepthInfostruct();}}aligned(8) TextureInfostruct(){ unsigned int(8) texture_padding_size; unsigned int(16) texture_top_left_x; unsigned int(16) texture_top_left_y; unsigned int(16) texture_bottom_right_x; unsigned int(16) texture_bottom_right_y;} aligned(8)DepthInfostruct(){ unsigned int(8) depth_padding_size; unsigned int(16) depth_top_left_x; unsigned int(16) depth_top_left_y; unsigned int(16) depth_bottom_right_x; unsigned int(16) depth_bottom_right_y; unsigned int(8) depth_range_near; unsigned int(8) depth_range_far; unsigned int(8) depth_scale_type;}

[0091] Example 2: The media track for single-track encapsulated FreeView video uses a sample entry, VisualSampleEntry, which includes a FreeViewExtInfoBox. The FreeViewExtInfoBox describes the original acquisition device information corresponding to the FreeView video data in the media track, as well as the texture map and depth map information corresponding to the original acquisition device. aligned(8) class FreeViewExtInfoBox extends FullBox(' fvei', 0, 0){ unsigned int(1) texture_in_track; unsigned int(1) depth_in_track; unsigned int(16) num_cameras; for (i=0; i<num_cameras; i++) { unsigned int(16) camera_id; unsigned int(1) int_camera_flag; unsigned int(1) ext_camera_flag; unsigned int(1) decoding_info_flag; if(int_camera _flag) IntCameraInfoStruct(); if(Ext_camera _flag) ExtCameraInfoStruct(); unsigned int(1) texture_info_flag; unsigned int(1) depth_info_flag; if(texture_info_flag){ unsigned int(8) texture_padding_size; unsigned int(16) texture_top_left_x[i]; unsigned int(16) texture_top_left_y[i]; unsigned int(16) texture_bottom_right_x[i]; unsigned int(16) texture_bottom_right_y[i]; } if(depth_info_flag){ unsigned int(8) depth_padding_size; unsigned int(16) depth_top_left_x[i]; unsigned int(16) depth_top_left_y[i]; unsigned int(16) depth_bottom_right_x[i]; unsigned int(16) depth_bottom_right_y[i]; unsigned int(8) depth_range_near; unsigned int(8) depth_range_far; unsigned int(8) depth_scale_type;}}} Here, `texture_in_track` indicates whether the media track contains a texture map; a value of 1 means a texture map is included, and a value of 0 means a texture map is not included. `depth_in_track` indicates whether the media track contains a depth map; a value of 1 means a depth map is included, and a value of 0 means a depth map is not included. num_cameras indicates the number of original cameras in the media track corresponding to the depth map or texture map. The `camera_id` indicates the identifier of the original camera. The `int_camera_flag` parameter indicates whether the original camera's intrinsic parameters are displayed. A value of 1 displays the intrinsic parameters, while a value of 0 hides them. The syntax for the intrinsic parameters refers to the `IntCameraInfoStruct()` syntax semantics of AVS3-P1. The `ext_camera_flag` flag indicates whether the external parameters of the original camera are displayed. A value of 1 means the external parameters are displayed, and a value of 0 means they are not. The syntax for the external parameters refers to the ExtCameraInfoStruct() syntax semantics of AVS3-P1. The original camera information can be described by the original camera's internal and external parameters. Parameter information for the same original camera is described in only one box. The texture_info_flag indicates whether or not descriptive metadata information about the texture map is included. A value of 1 means that metadata information is included, and a value of 0 means that metadata information is not included. The `depth_info_flag` flag indicates whether or not descriptive metadata information about the depth map is included. A value of 1 means that metadata information is included, and a value of 0 means that metadata information is not included. `texture_padding_size` indicates the number of pixels in the extended edge of the texture map. texture_top_left_x represents the x-coordinate of the top-left corner of the texture map corresponding to the original camera within the video frame plane. texture_top_left_y indicates the y-coordinate of the top-left corner of the texture map corresponding to the original camera within the video frame plane. texture_bottom_right_x indicates the x-coordinate of the bottom right corner of the texture map corresponding to the original camera within the video frame plane. texture_bottom_right_y indicates the y-coordinate of the bottom right corner of the texture map corresponding to the original camera within the video frame plane. `depth_padding_size` indicates the number of pixels to extend the depth map. depth_top_left_x indicates the x-coordinate of the top-left corner of the depth map corresponding to the original camera within the video frame plane. depth_top_left_y indicates the y-coordinate of the top-left corner of the depth map corresponding to the original camera within the video frame plane. depth_bottom_right_x indicates the x-coordinate of the bottom right corner of the depth map corresponding to the original camera within the video frame plane. depth_bottom_right_y indicates the y-coordinate of the bottom right corner of the depth map corresponding to the original camera within the video frame plane. `depth_range_near` indicates the minimum depth distance from the optical center, as quantified by the depth map. `depth_range_far` indicates the maximum depth distance from the optical center, as quantified by the depth map. `depth_scale_type` indicates the type of downsampling used to represent the depth map. This sample entry, VisualSampleEntry, includes a FreeViewExtInfoBox and a PayloadInfoBox.

[0092] In one embodiment, the FreeViewExtInfoBox contains track information, which may include identifiers for the tracks where texture maps exist and / or identifiers for the tracks where depth maps exist.

[0093] In one embodiment, the FreeViewExtInfoBox contains device information for the collection device, which may include the number of collection devices, the identifier of the collection device, and parameter information for the collection device. Information regarding the installation method and attributes of the collection device may also be encapsulated in the FreeViewExtInfoBox as device information.

[0094] In one embodiment, the PayloadInfoBox contains attribute information for the texture map and / or attribute information for the depth map. The attribute information for the texture map may include a texture map information identifier, texture map image information, etc., and the attribute information for the depth map may include a depth map information identifier, depth map image information, etc. Figure 9 is a flowchart of the processing of free-view video data based on single-track encapsulation, including at least steps S1300 and S3300.

[0095] Step S1300: The terminal retrieves the texture maps and depth maps collected by all original collection devices.

[0096] Step S3300: The terminal partially decodes the texture maps and depth maps of several selected original acquisition devices based on the target view information and synthesizes a virtual view texture map.

[0097] Figure 10 is a diagram illustrating the file structure of multi-track encapsulated free-view video data. As shown in Figure 10, a media file based on the ISO Basic Media File Format has multiple media tracks. Each track can encapsulate a texture map (i.e., a texture map media track) or depth map (i.e., a depth map media track) collected by a single original acquisition device. The multi-track encapsulation structure allows for flexible acquisition of any one collected texture map or depth map.

[0098] The depth map may be represented using a grayscale map, which has high encoding efficiency and low decoding complexity, and multiple depth maps from multiple original acquisition devices may be encapsulated in one or more depth map media tracks. Similarly, the texture maps from multiple original cameras may also be encapsulated in one or more texture map media tracks.

[0099] In one embodiment, the number of original acquisition devices corresponding to the texture map media track and the depth map media track may be the same or different, depending on differences in the image size, encoding method, etc., of the texture map or depth map of the original acquisition devices. Specifically, if the number of original acquisition devices corresponding to the texture map media track and the depth map media track are different, the texture maps of m original acquisition devices are encapsulated in i texture map media tracks, and the corresponding depth maps are encapsulated in j depth map media tracks, where i is not equal to j, i is less than or equal to m, and j is less than m. In another embodiment, the depth maps of multiple original acquisition devices may be stored in one media track, and the corresponding texture maps may be stored in multiple media tracks.

[0100] In one embodiment, the original acquisition devices are grouped according to camera attributes such as the positional attributes of the camera placement, the intrinsic parameter attributes of the camera, and the encoding dependencies of the texture map or depth map corresponding to the camera, and the texture maps or depth maps of the original acquisition devices in the same group are encapsulated in a single media track. Figure 11 is a diagram of the media file structure of a FreeView video based on the multi-track encapsulation mode extension. As shown in Figure 11, it is one file structure of multi-track encapsulated FreeView video data based on the same type of multigraph.

[0101] Multitrack encapsulation includes a MOOV box and a madt box. The MOOV box contains a FreeViewExtInfoBox subbox, which encapsulates the media's metadata information, while the madt box encapsulates the FreeView video data in the form of a sample.

[0102] Multitrack encapsulated freeview video data includes texture map and depth map data collected by multiple original acquisition devices, with the texture map and depth map data stored in different media tracks. Furthermore, texture maps are stored in the texture map media track, and depth maps are stored in the depth map media track. That is, each frame image in the texture map media track contains the texture maps of all original acquisition devices at a given point in time, and each frame image in the depth map media track contains the depth maps of all original acquisition devices at a given point in time. The texture maps and depth maps corresponding to each original acquisition device are associated and codecized, meaning that a dependency / mapping relationship exists between the codecs of the texture map and depth map data of a single acquisition device. When a terminal acquires multitrack encapsulated freeview video, it transmits the texture map and depth map data collected by the corresponding acquisition device based on the target view information, which is the user's current view.

[0103] The terminal acquires a media file of free-view video data corresponding to the target view information, and then, for each frame image, selects and decodes the texture maps and depth maps of several original acquisition devices in the view neighborhood necessary to synthesize a virtual viewpoint. Each sample in the multi-track encapsulated texture map media track contains a texture map collected by the original acquisition device at a given point in time, and each sample in the multi-track encapsulated depth map media track contains a depth map collected by the original acquisition device at a given point in time. Furthermore, by dividing the sample into multiple subsamples, each subsample may contain a texture map or depth map corresponding to one acquisition device. More specifically, information about subsamples is described by the SubSampleInformationBox box of the SampleTableBox or TrackFragmentBox, and the use of subsamples must be carried out according to the flags field in the subsample information box (the flags field can identify the data type stored in the subsample), and if the value of the flags field is 0, the texture map and depth map are in the media track, and in particular, the specific configuration structure of each sample corresponding to the free view video data is described by codec_specific_parameters.

[0104] The texture map and depth map are encapsulated in different media tracks; that is, the texture map is encapsulated in the texture map media track, and the depth map is encapsulated in the depth map media track. The process for compositing the virtual view is as follows: First, after the view is switched, several original collections in the vicinity of the view are identified, the originally collected depth maps are processed, and then a texture map corresponding to the virtual view is composited based on the originally collected texture maps and the processed depth maps. An association / reference / mapping relationship is established between the texture map media track and the depth map media track that encapsulate the texture map and depth map of the same original acquisition device. Three specific examples are shown below. Example 1:

[0105] A TrackReferenceBox (reference trackbox), as defined in ISO / IEC 14494-12, can be used for reference relationships between media tracks and describes the auxiliary media (such as a depth map) of the associated / referenced / mapped media track when the defined reference type `reference_type` is "auxl" or "vdep". In this embodiment, the method defined in ISO / IEC 14494-12 is used directly to include a TrackReferenceBox in the depth map media track and associate / reference / map it to the texture map media track corresponding to the same original acquisition device, where `reference_type` is "auxl" or "vdep". Typically, "auxl" is used for reference or association relationships with decoding dependency. Example 2:

[0106] A TrackReferenceBox, as defined in ISO / IEC 14494-12, can be used for reference relationships between tracks, describing the association / reference / mapping of a texture map media track to a depth map media track by defining a new reference type, reference_type. In this example, reference_type is set to "tdrf", and a texture map media track in a TrackReferenceBox containing this reference type references or is associated with the depth map media track indicated by the TrackReferenceBox. Example 3:

[0107] Grouping texture map media tracks and depth map media tracks using a track grouping method is particularly suitable when the original acquisition devices corresponding to the depth map media tracks and the acquisition devices corresponding to the texture map media tracks do not perfectly correspond. For example, depth maps from multiple original acquisition devices may be stored in one media track, while the corresponding texture maps may be stored in multiple media tracks. Track groups are formed by grouping texture map media tracks and depth map media tracks corresponding to the same original acquirer into a single group. Track groups corresponding to texture map media tracks and depth map media tracks of the same original acquirer are indicated by AttrAndDepGroupBox, which extends the TrackGroupTypeBox (track type box) defined in ISO / IEC 14494-12, with track_group_type set to "tadg", and the specific syntactic structure is as follows: aligned(8) class AttrAndDepGroupBox extends TrackGroupTypeBox(' tadg ') { unsigned int(8) num_camera; unsigned int(1) unequal_flag; for(int i=0; i< num_camera; i++) unsigned int(8) camera_id; } Here, num_camera indicates the number of original camera units corresponding to the texture maps within the media track group. The `unequal_flag` flag indicates whether the number of original collectors corresponding to depth map media tracks in a media track group is the same as the number of original collectors corresponding to texture map media tracks. A value of 1 indicates that the number of corresponding original collectors is different, and typically the number of corresponding original collectors in depth map media tracks is greater than the number of corresponding original collectors in texture map media tracks. A value of 0 indicates that the number of corresponding original collectors is the same. The camera_id in the media track group indicates the identifier of the same original acquisition device corresponding to the texture map and depth map.

[0108] In one embodiment, if unequal_flag is 1, a single depth-mapped media track may have multiple AttrAndDepGroupBoxes of the same type, but with different track_group_ids. In one embodiment, when freeview video is encoded in the AVC MVD format, a ViewIdentifierBox of type "vwid" as defined in ISO / IEC 14494-15 can be used to describe the viewpoints corresponding to the original acquisition devices included in each media track. This box may contain information such as the viewpoint identifier and other reference viewpoints corresponding to the view, and this box can also be used to describe whether the media track includes a texture map, a depth map, or both.

[0109] In one embodiment, if the depth map and texture map are on different media tracks, the depth map media track and the texture map media track are associated / referenced / mapped to the texture map media track corresponding to the same original acquisition device by a TrackReferenceBox (this box is included in the depth map media track) box whose reference_type is "deps".

[0110] This sample entry, VisualSampleEntry, includes a free view information box (AvsFreeViewInfoBox) and a payload information box (PayloadInfoBox).

[0111] In one embodiment, the FreeViewInfoBox AvsFreeViewInfoBox contains track information, which may include identifiers for tracks with texture maps or identifiers for tracks with depth maps.

[0112] In one embodiment, the FreeViewInfoBox contains device information for the collection device, which may include the number of collection devices, the identifier of the collection device, and the parameter information of the collection device. Information regarding the installation method and attributes of the collection device may also be encapsulated in the FreeViewInfoBox as device information.

[0113] In one embodiment, the PayloadInfoBox contains attribute information for the texture map and / or attribute information for the depth map. The attribute information for the texture map may include a texture map information identifier, texture map image information, etc., and the attribute information for the depth map may include a depth map information identifier, depth map image information, etc.

[0114] Multitrack encapsulation includes at least a texture map media track and a depth map media track. Each texture map media track and depth map media track's sample entry, VisualSampleEntry, contains a free view information box, AvsFreeViewInfoBox, and a payload information box, PayloadInfoBox, respectively. Figure 12 is a flowchart of the processing of free-view video data based on multi-track encapsulation, including at least steps S1400 and S3400.

[0115] Step S1400: The terminal retrieves the texture map and depth map collected by the original collection device near the view, based on the target view information.

[0116] Step S3400: The terminal retrieves one or more texture map media tracks and depth map media tracks containing the texture maps and depth maps of several required original collectors, extracts the texture map and depth map data of several required original collectors from the media tracks, and composites the texture maps of the virtual viewpoint.

[0117] In single-track encapsulation, the terminal needs to acquire free-view video data collected by all acquisition devices. In contrast, in a multi-track encapsulation structure, the terminal selectively acquires texture maps and / or depth maps corresponding to the target view information in texture map media tracks and depth map media tracks, respectively, based on the target view information, thereby improving data transmission and processing efficiency.

[0118] In some application scenarios, when viewing a free-view video, users can switch views autonomously according to their needs, or they can switch views passively, such as through director recommendations. In this case, it is necessary to define dynamic view timing metadata to describe dynamically changing view information and the corresponding information from the original data collection device. In other words, it is necessary to associate views with video playback time. Therefore, during the playback of a free-view video, it is necessary to synthesize the content of the corresponding view that needs to be presented based on the recommended view at a certain time and the information from the original data collection device in its vicinity. The timing metadata is encapsulated in a timing metadata track that is correlated with the media track where the free-view video data corresponding to the target view resides.

[0119] Specifically, the sample entry type for the timing metadata track of a dynamic view is "dyfv", and the syntax is defined as follows: DynamicFreeViewSampleEntry extends MetaDataSampleEntry ('dyfv'){ string freeview_description; unsigned int(8) play_camera_id; } Here, `freeview_description` is a string that ends with a whitespace character and provides text description information for the free view. `play_camera_id` indicates the identifier of the original camera to which the view corresponds during the initial playback, and usually does not require virtual view compositing. Each sample within the dynamic view timing metadata track describes the dynamic state of a particular view using the free view sample format. aligned(8) FreeviewSample(){ unsigned int(1) dynamic_switch_flag; if(dynamic_switch_flag) unsigned int(1) virtual_view_flag; if(virtual_view_flag){ unsigned int(8) original_camera_num; for(int i=0; i< original_camera_num; i++){ unsigned int(8) original_camera_id[i]; unsigned int(1) int_camera _flag; unsigned int(1) ext_camera _flag; if(int_camera_flag) IntCameraInfoStruct(); if(ext_camera_flag) ExtCameraInfoStruct(); FreeViewStruct(); } else unsigned int(8) play_camera_id;}} aligned(8) class FreeViewStruct(){ signed int(32) freeview_x; signed int(32) freeview_y; signed int(32) freeview_z; unsigned int(1) rotation_flag; bit(7) reserved = 0; if(rotation_flag) { signed int(32) freeview_yaw; signed int(32) freeview_pitch; signed int(32) freeview_roll;}} Here, dynamic_switch_flag indicates whether the view in the sample will switch or not. A value of 1 indicates that the view will switch, while a value of 0 indicates that the view will not switch and the view from the previous sample will be used. The `virtual_camera_flag` flag indicates whether or not the virtual viewpoint needs to be composited. A value of 1 indicates that compositing is necessary, while a value of 0 indicates that compositing is not necessary and the view should be switched directly to the original camera. `original_camera_num` indicates the number of original cameras used to synthesize the virtual viewpoint. `original_camera_id` indicates the identifier of the original camera from which the virtual viewpoint is synthesized. `play_camera_id` indicates the identifier of the original camera that is played back directly without the need to composite a virtual viewpoint. FreeViewStruct() shows the parameter structure of a free view. Freeview_x, Freeview_y, and Freeview_z represent the x, y, and z axis coordinates of the free view relative to a common reference coordinate system, respectively. The `grotation_flag` flag indicates whether or not to represent rotation in the free view; a value of 1 means it is represented, and a value of 0 means it is not. freeview_yaw, freeview_pitch, and freeview_roll each indicate the rotation angle of the free view relative to a common reference coordinate system.

[0120] If the original camera parameter information exists in MediaInformationBox or SampleEntry, it is not necessary to repeatedly describe it in the sample format, and the values ​​of int_camera_flag and ext_camera_flag are 0.

[0121] The aforementioned dynamic view metadata track can be associated with or reference to one or more media tracks, including the depth map and / or texture map of the original acquisition device, by referencing a Track Reference Box of type "cdsc".

[0122] By adopting the director-recommended view switching method described above, content corresponding to the recommended view can be pre-composited based on the recommended view and the original camera information in its vicinity. This reduces the data processing requirements on devices such as terminals, further improves video processing efficiency, enhances the user's viewing experience, and expands the application scenes for free-view video.

[0123] In some application scenarios, attributes such as the placement of the original acquisition device, the frequency of image acquisition by the original acquisition device, and the compression method of the images acquired by the original acquisition device differ. As a result, differences often occur during media processing when users freely switch views. For example, texture maps and depth maps acquired by original acquisition devices in the same area can be used to synthesize a virtual viewpoint, but texture maps and depth maps acquired by original acquisition devices in non-identical areas, such as two rooms, cannot be used to synthesize a virtual viewpoint together. Therefore, in this embodiment, we propose grouping the texture map media tracks and depth map media tracks of the original acquisition devices according to the attributes of the original acquisition device, and grouping texture map media tracks and depth map media tracks of original acquisition devices with the same attribute (e.g., position attribute) into a single track group. Figure 13 is a schematic diagram of media track grouping, showing an example of a track group of texture map media tracks and depth map media tracks grouped based on the attributes of the original camera.

[0124] As shown in the diagram, Tracks 1-4 are texture map media tracks, Tracks 5-7 are depth map media tracks, Tracks 3, 4, and 7 meet the pre-set track grouping conditions, and Tracks 1, 2, 5, and 6 also meet the pre-set track grouping conditions, so Tracks 3, 4, and 7 are considered one media track group, and Tracks 1, 2, 5, and 6 are considered one media track group. When a terminal requests free view video data, the data from one media track group is transmitted uniformly, so the terminal can accurately acquire image data corresponding to the target view and accurately process the data group.

[0125] In one embodiment, the track groups of texture map media tracks and depth map media tracks based on the attributes of the original acquisition device are represented by a CamAttrGroupBox, which extends the TrackGroupTypeBox defined in ISO / IEC 14494-12, with track_group_type set to "caag", and the specific syntactic structure is as follows: aligned(8) class CamAttrGroupBox extends TrackGroupTypeBox(' caag ',0,0) { group_description; unsigned int(8) num_camera; for(int i=0; i< num_camera; i++) unsigned int(8) camera_id; } Here, group_description is a whitespace-ending string that provides text description information grouped based on the attributes of the original collector. num_camera indicates the number of original camera units corresponding to the texture map within the media track group. The camera_id indicates the identifier of the corresponding original collection device within the track group. In one embodiment, the track groups of texture map media tracks and depth map media tracks based on the attributes of the original acquisition device are described by extending EntityToGroupBox as defined in ISO / IEC 14494-12. The grouping type grouping_type is set to "caeg", and the syntactic structure is as follows: aligned(8) class CamAttrEntityGroupBox extends EntityToGroupBox(' caeg ',0,0) { group_description; unsigned int(8) num_camera; for(int i=0; i< num_camera; i++) unsigned int(8) camera_id; }

[0126] In one embodiment, grouping based on elements such as the original camera attributes may be grouping of only depth map media tracks, or grouping of only texture map media tracks, or grouping of both texture map and corresponding depth map media tracks. In another embodiment, a similar effect can be achieved even when only depth map media tracks or texture map media tracks are grouped, due to the association between the depth map media tracks and the texture map media tracks.

[0127] Figure 14 is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention. As shown in Figure 14, the data processing device according to the embodiment of the present invention is applied to a terminal and can perform the data processing method according to the embodiment of the present invention, and this terminal has a functional module and technical effects corresponding to the performance of the method. This device may be implemented as software, hardware, or a combination of software and hardware, and includes an acquisition module 700, an encapsulation module 800, and a processing module 900.

[0128] The acquisition module 700 is configured to acquire free-view video data.

[0129] In one embodiment, the acquisition module 700 is also configured to acquire target view information.

[0130] The encapsulation module 800 is configured to encapsulate the texture map into a media track. In one embodiment, the encapsulation module 800 is configured to encapsulate the texture maps collected by the acquisition device into a single media track. In one embodiment, the encapsulation module 800 is configured to encapsulate the texture maps collected by the acquisition device into multiple media tracks. In one embodiment, the encapsulation module 800 is configured to encapsulate the depth maps collected by the acquisition device into a single media track. In one embodiment, the encapsulation module 800 is configured to encapsulate the depth map collected by the acquisition device into multiple media tracks. In one embodiment, the encapsulation module 800 is configured to encapsulate the texture map and depth map collected by the acquisition device into a single media track. In one embodiment, the encapsulation module 800 is configured to encapsulate the texture map and depth map collected by the acquisition device into multiple media tracks.

[0131] The processing module 900 is configured to process free-view video data based on target view information.

[0132] In one embodiment, the processing module 900 is configured to decode / decapsulate the requested free view video data, obtain the corresponding free video view data based on the target view information, and perform processing such as extracting a texture map corresponding to the real viewpoint or a texture map corresponding to the synthesized virtual viewpoint. Furthermore, it decodes / decapsulates the requested free view video media file or processed media data, extracts the corresponding free view video data based on the requested media file and the viewing position and viewing direction after the user switches views, and performs media processing.

[0133] Figure 15 is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention. As shown in Figure 15, the data processing device further includes a display module 1000 configured to display an image in accordance with the free-view video data processed by the processing module. That is, depending on the processing result by the processing module 900, a texture map of the real viewpoint corresponding to the current view is displayed to the user, or a texture map of the virtual viewpoint corresponding to the current view is displayed to the user.

[0134] Figure 16 is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention. As shown in Figure 16, the data processing device further includes a transmission module 1100 configured to send data requests to a server.

[0135] In one embodiment, the transmission module 1100 is also configured to transmit target view information to the server.

[0136] The data processing device may also include a storage module 1200 instead of an encapsulation module 800. Figure ** is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention. As shown in Figure **, it includes an acquisition module 700, a processing module 900, and a storage module 1200. The acquisition module 700 and the processing module 900 operate in the same manner as in the previously described embodiment, so their description is omitted here. The storage module 1200 is configured to store media files of encapsulated free-view video data.

[0137] Figure 17 is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention. As shown in Figure 17, the data processing device according to the embodiment of the present invention is applied to a media server and can execute the data processing method according to the embodiment of the present invention, and the server has functional modules and technical effects corresponding to the execution of the method. This device may be implemented as software, hardware, or a combination of software and hardware, and includes a transmission module 1300, an arithmetic module 1400, and a storage module 1500.

[0138] The transmission module 1300 is configured to receive a request message from a terminal, transmit target view information according to the request information, and process free view video data. Specifically, it transmits media files stored by the storage module 1500 or media data processed by the arithmetic module 1400. The above reception or transmission can be achieved by a wireless network provided by a communications vendor, a locally configured wireless local network, or a wired method.

[0139] The arithmetic module 1400 is configured to retrieve the necessary media files from the storage module 1500 according to the request message, extract the corresponding free-view video data from the media files, and perform arithmetic processing. The storage module 1500 is configured to store media files of encapsulated free-view video data.

[0140] Figure 18 is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention. As shown in Figure 18, this device includes a memory 1600, a processor 1700, and a communication device 1800. The number of memory 1600 and processor 1700 may be one or more, and Figure 16 illustrates one memory 1600 and one processor 1700. The memory 1600 and processor 1700 of the device may be connected by a bus or by other means, and Figure 16 illustrates a connection via a bus.

[0141] The memory 1600 may be used as a computer-readable storage medium to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the data processing method according to any embodiment of the present application. The processor 1700 implements the above-described data processing method by executing the software programs, instructions, and modules stored in the memory 1600.

[0142] Memory 1600 may primarily include an operating system, a program storage area capable of storing application programs required for at least one function, and a data storage area. Memory 1600 may also include high-speed random-access memory, and further may include non-volatile memory such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state memory device. In some examples, memory 1600 may further include memory located remotely from the processor 1700, and these remotely located memories may be connected to equipment via a network. Examples of networks include, but are not limited to, the Internet, a corporate intranet, a local area network, a mobile communication network, and combinations thereof.

[0143] The communication device 1800 is configured to transmit and receive information in accordance with the control of the processor 1700. In one embodiment, the communication device 1800 includes a receiver 1810 and a transmitter 1820. The receiver 1810 is a combination of modules or devices that receive data within the electronic device. The transmitter 1820 is a combination of modules or devices that transmit data within the electronic device.

[0144] Figure 19 is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention. As shown in Figure 19, in one embodiment, the data processing device may further include an input device 1400 and an output device 1500.

[0145] The input device 1400 can be used to receive input numerical or character information and generate key signal inputs related to user settings and function control of the device. The output device 1500 may include a display device such as a display screen. One embodiment of the present invention also provides a computer-readable storage medium that stores computer-executable instructions for performing a data processing method according to any embodiment of the present invention.

[0146] Another embodiment of the present invention provides a computer program product comprising a computer program or computer instruction stored on a computer-readable storage medium, wherein the processor of a computer device reads the computer program or computer instruction from the computer-readable storage medium, and the processor executes the computer program or computer instruction, thereby causing the computer device to execute the data processing method according to any embodiment of the present invention.

[0147] The system architectures and application scenarios described in the embodiments of this application are intended to more clearly illustrate the technical proposals of the embodiments and do not limit the technical proposals of the embodiments. Those skilled in the art will understand that, as system architectures evolve and new application scenarios emerge, the technical proposals of the embodiments of this application can be similarly applied to similar technical problems.

[0148] Those skilled in the art will understand that all or part of the steps in the methods disclosed above, or the functional modules / units in a system or device, may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0149] In hardware embodiments, the distinctions between functional modules / units described above do not necessarily correspond to distinctions between physical components. For example, one physical component may have multiple functions, or multiple physical components may work together to perform a single function or step. Some or all of the physical components may be implemented as software executed by a processor such as a central processing unit, digital signal processing unit, or microprocessor, or as hardware, or as an integrated circuit such as an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-temporary media) and communication media (or temporary media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information (computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage devices, magnetic boxes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other media that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media may include any information distribution media, typically containing computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms.

[0150] As used herein, terms such as “component,” “module,” and “system” are used to mean computer-related entities, hardware, firmware, hardware-software combinations, software, or running software. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, or a computer. As illustrated, applications and computers running on computing equipment may both be components. One or more components may reside within a process or execution thread, and components may be located on a single computer or distributed across two or more computers. Furthermore, these components may run from various computer-readable media that store various data structures. Components may communicate via local or remote processes according to signals, for example, having one or more data packets (e.g., data from two components interacting with other components across a local system, a distributed system, or a network, e.g., data from the Internet interacting with other systems via signals).

[0151] Some embodiments of this application are described with reference to the accompanying drawings, but this does not limit the scope of the rights of this application. Any amendments, equivalent substitutions, and improvements made by a person skilled in the art without departing from the scope and essence of this application shall all be within the scope of the rights of this application.

Claims

1. A data processing method performed by a data processing device, A step of obtaining free view video data that includes at least one texture map and / or depth map, The steps include encapsulating at least one texture map and / or depth map collected by at least one acquisition device into a first media track, A method comprising the step of processing the free view video data based on target view information.

2. The sample of the first media track includes at least one subsample, The method according to claim 1, wherein the subsample includes the texture map and / or depth map.

3. The subsample includes the texture map and / or depth map, If the subsample includes texture maps and depth maps collected by multiple acquisition devices, The method according to claim 2, wherein the subsample includes at least one of a texture map and / or depth map collected by one of the acquisition devices.

4. The sample entry for the first media track includes a first box and / or a second box. The first box contains device information of the collection device, The method according to claim 1, wherein the second box includes attribute information of the texture map and / or depth map.

5. The aforementioned first box further contains track information, The method according to claim 4, wherein the track information includes at least one of the identifiers of tracks in which a texture map exists and the identifier of a track in which a depth map exists.

6. The device information of the collection device includes at least one of the following: the number of collection devices, the identifier of the collection device, and the parameter information of the collection device. The attribute information of the aforementioned texture map includes at least one of a texture map information identifier and texture map image information. The method according to claim 4, wherein the attribute information of the depth map includes at least one of a depth map information identifier and depth map image information.

7. This further includes the step of obtaining target view information, The above step of obtaining free view video data is, The method according to claim 1, further comprising the step of obtaining free view video data corresponding to a target view based on target view information.

8. The step of obtaining free view video data corresponding to the target view based on the aforementioned target view information is: The method according to claim 7, further comprising the step of obtaining a texture map and / or depth map corresponding to the target view based on the target view information.

9. The step of acquiring target view information includes the step of acquiring target view information according to a timing metadata track, wherein the target view changes dynamically according to the instructions of the timing metadata track. The method according to claim 7, wherein the timing metadata track is associated with the first media track containing the free view video data corresponding to the target view.

10. The target view information includes at least one of the following: an identifier for the collection device, parameter information for the collection device, position information for the target view, rotation information for the target view, and an identifier for the target view. The step of processing the free view video data based on the target view information is: Based on the free-view video data and the target view information, a texture map and depth map collected by the acquisition device corresponding to the target view are selected, and the texture map corresponding to the target view is synthesized. Or, The method according to claim 1, further comprising the step of selecting a texture map collected by an acquisition device corresponding to the target view based on the free view video data and the target view information.

11. A data processing method performed by a data processing device, A step of obtaining free view video data that includes at least one texture map and / or depth map, The steps include encapsulating the texture maps collected by at least one acquisition device into at least one texture map media track, and encapsulating the depth maps collected by at least one acquisition device into at least one depth map media track, A method comprising the step of processing the free view video data based on target view information.

12. The sample entry for the aforementioned texture map media track includes a first texture map box, The sample entry for the depth map media track includes a first depth map box, The first texture map box includes device information of the acquisition device corresponding to the texture map media track or attribute information of the texture map. The first depth map box includes device information of the acquisition device corresponding to the depth map media track or attribute information of the depth map, The aforementioned texture map box further includes an identifier for the track in which the texture map exists. The method according to claim 11, wherein the first depth map box further includes an identifier for the track in which the depth map exists.

13. The sample entry for the aforementioned texture map media track further includes a second texture map box, The sample entry in the aforementioned depth map media track further includes a second depth map box, The second texture map box includes device information of the acquisition device corresponding to the texture map media track or attribute information of the texture map. The second depth map box includes device information of the acquisition device corresponding to the depth map media track or attribute information of the depth map, The method according to claim 12, wherein the first texture map box and the second texture map box contain different information, and the first depth map box and the second depth map box contain different information.

14. The third box is used to associate the texture map media track with the depth map media track, The fourth box is used to associate the depth map media track with the texture map media track, The method according to claim 11, further comprising at least one of the steps of associating the depth map media track and the texture map media track as a track group by a fifth box.

15. The steps of encapsulating at least two of the texture map media tracks as a first media track group based on pre-set information, A step of encapsulating at least two of the depth map media tracks as a second media track group based on pre-configured information, The method further includes at least one of the following steps: encapsulating at least one of the texture map media track and at least one of the depth map media track as a third media track group based on pre-configured information; The method according to claim 14, wherein the pre-set information includes at least one of the following: the location of the collection device, the identifier of the collection device, the image collection frequency, and the encoding scheme.

16. This further includes the step of obtaining target view information, The above step of obtaining free view video data is, The method according to claim 11, further comprising the step of obtaining free view video data corresponding to a target view based on target view information.

17. The step of obtaining free view video data corresponding to the target view based on the aforementioned target view information is: The method according to claim 16, further comprising the step of obtaining a texture map and / or depth map corresponding to the target view based on the target view information.

18. The step of acquiring target view information includes the step of acquiring target view information according to a timing metadata track, wherein the target view changes dynamically according to the instructions of the timing metadata track. The method according to claim 16, wherein the timing metadata track is associated with a media track containing the free view video data corresponding to the target view.

19. The target view information includes at least one of the following: an identifier for the collection device, parameter information for the collection device, position information for the target view, rotation information for the target view, and an identifier for the target view. The step of processing the free view video data based on the target view information is: Based on the free-view video data and the target view information, a texture map and depth map collected by the acquisition device corresponding to the target view are selected, and the texture map corresponding to the target view is synthesized. Or, The method according to claim 11, further comprising the step of selecting a texture map collected by an acquisition device corresponding to the target view based on the free view video data and the target view information.

20. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, realize the data processing method described in any one of claims 1 to 19.