Immersive Video Encoding and Decoding

Metadata-based differentiation of original and inpainted data in immersive video encoding addresses reprojection inaccuracies and resource demands, enhancing view synthesis efficiency and reducing bitrate.

JP7803337B2Active Publication Date: 2026-01-21KONINKLIJKE PHILIPS NV

Patent Information

Application Number
JP2023518747
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-30
Filing Date
2021-09-23
Publication Date
2026-01-21
Estimated Expiration
2041-09-23

AI Technical Summary

Technical Problem

Inpainting during post-processing in immersive video encoding leads to inaccurate reprojection and increased resource requirements due to inpainted depth information quality issues and conflicts in texture atlas packing, resulting in higher bitrates and resource demands.

Method used

Implement metadata in immersive video encoding and decoding to distinguish between original and inpainted patch data units, using fields to indicate the presence and quality of inpainted data, allowing for controlled rendering and reduced bitrate through LoD downscaling and separate mesh representation.

Benefits of technology

Enhances view synthesis by accurately handling inpainted data, reducing bitrate and resource requirements, and optimizing rendering processes for immersive video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007803337000001
    Figure 0007803337000001
  • Figure 0007803337000002
    Figure 0007803337000002
  • Figure 0007803337000003
    Figure 0007803337000003
Patent Text Reader

Abstract

[0003] Concepts for encoding and decoding multi-view data for immersive video are disclosed. In the encoding method, metadata is generated that includes a field indicating whether a patch data unit of the multi-view data includes inpainted data to represent missing data. The generated metadata provides a means for distinguishing patch data units that include original texture and depth data from patch data units that include inpainted data (e.g., inpainted texture and depth data). Providing such information within the metadata of the immersive video can address issues related to blending and pruned view reconstruction. Also provided are an encoder and decoder for multi-view data for immersive video, and a corresponding bitstream including the metadata.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to immersive video, and more particularly to a method and apparatus for encoding and decoding multi-view data for immersive video. [Background technology]

[0002] Immersive video, also known as six degrees of freedom (6DoF) video, is a video of a three-dimensional (3D) scene that allows a view of the scene to be reconstructed for a viewpoint with varying position and orientation. It represents an evolution of three degrees of freedom (3DoF) video, which allows a view to be reconstructed for a viewpoint with any orientation, but only at a fixed point in space. In 3DoF, the degrees of freedom are angles: pitch, roll, and yaw. 3DoF video supports head rotation; in other words, a user consuming video content can look in any direction within a scene, but cannot move to a different location within the scene. 6DoF video supports head rotation and also supports the selection of a location within the scene from which the scene is observed.

[0003] To generate 6DoF video, multiple cameras are required to record a scene. Each camera generates image data (often called texture data in this context) and corresponding depth data. For each pixel, the depth data represents the depth at which the corresponding image pixel data is observed. Each of the multiple cameras provides a respective view of the scene.

[0004] A problem with generating a target view is that only image data available for the view from the source cameras can be synthesized. Some image regions of the target view may not be available from the transmitted video stream (e.g., because they were not visible from any of the source cameras). To address this problem, it is typical to fill or "paint" these image regions using color data available from other background regions. Such "inpainting" is performed as a post-processing step (e.g., in a decoder) after the view synthesis step. This is a complex operation, especially if the size of the regions of missing data is large. Summary of the Invention [Problem to be solved by the invention]

[0005] An alternative to inpainting during post-processing is to do the inpainting during data encoding (e.g., in the encoder) and then pack the resulting texture atlas with regular patches. However, this has the following associated drawbacks: (i) Inpainted image regions require texture and depth information. Depth information is required for the necessary reprojection. In addition to texture information, the inpainted depth information is also likely to be of lower quality than the original depth information. As a result, reprojection of regions of inpainted data is less accurate. (ii) During reconstruction of a pruned (redundant) source view from the encoded data, a problem arises when the texture atlas is packed with additional inpainted image regions: both inpainted patches and patches with original image data may be mapped to the same location in the reconstructed view, causing conflicts. (iii) Packing additional inpainted textures into the video stream increases the bitrate, which in turn increases the required (active) frame size, i.e., pixel rate, of the textures and depth atlases, which increases the resource requirements on the client device (which typically has limited resources). [Means for solving the problem]

[0006] The invention is defined by the claims.

[0007] According to an example according to an aspect of the present invention, there is provided a method for encoding multi-view data for immersive video as claimed in claim 1.

[0008] The proposed concept aims to provide schemes, solutions, concepts, designs, methods, and systems related to encoding multi-view data for immersive video. Specifically, embodiments aim to provide a concept for distinguishing patch data units that retain original texture and depth information from patch data units that retain inpainted data. Thus, blending and pruned view reconstruction issues can be addressed. Specifically, embodiments propose using metadata of immersive video to provide a strategy for indicating whether a patch data unit of multi-view data contains inpainted data to represent missing data. In this way, existing features of immersive video can be leveraged to indicate the presence of inpainted data in the multi-view data.

[0009] For example, according to a proposed embodiment, metadata for an immersive video can be generated to include a field (i.e., a syntax element, metadata field, metadata element, or input element populated with data) that indicates whether a patch data unit contains inpainted data.

[0010] This field may contain at least two sets of allowed values. A first value in this set may indicate that the patch data unit of the multi-view data contains original image data captured from at least one viewpoint, and a second value in this set may indicate that the patch data unit of the multi-view data contains inpainted data. For example, this field may be a binary flag or a Boolean indicator and may therefore consist of a simple bit (representing a Boolean value of "0" / "low" or "1" / "high"). This field may have the form of a syntax element in the bitstream. Alternatively, this field may be derived from other fields. For example, a first other field may represent the total number of views present in the bitstream, and a second other field may indicate the total number of non-inpainted views. If the view index exceeds the total number of non-inpainted views, the (derived) field is "1," otherwise it is "0," or vice versa. Therefore, such an implementation would require minimal or minor modifications to conventional immersive video metadata.

[0011] However, in some embodiments, the set of allowed values ​​may include more than two allowed values. For example, the value of the field may indicate the Level of Detail (LoD) of the patch data unit. One value of this field may indicate that the patch data unit comprises original / captured data of the highest quality (and therefore highest priority for use, i.e., lossless). Another value of this field may indicate that the patch data unit includes data synthesized from captured data (i.e., somewhat lower fidelity, but still good quality). Yet another value of this field may indicate that the patch data unit includes inpainted data of the lowest quality (and therefore lowest priority for use, i.e., inpaint loss). In this way, the field can provide further information about the inpainted data (such as LoD details of the inpainted data). Thus, some embodiments may use a field with more than two allowed values. Thus, this field may comprise multiple bits (e.g., one or more bytes).

[0012] Multi-view data may be encoded, and this field may be associated with a frame of the encoded multi-view data and may contain a description (or definition) of one or more patch data units of this frame that have inpainted data.

[0013] In some embodiments, this field includes an identifier or address of a stored value. Such stored values ​​may include, for example, rendering parameter values. That is, this field may include information that allows one or more values ​​to be searched for or "looked up." For example, different rendering parameter sets may be predefined, each stored using its own unique identifier (e.g., address). The identifier / address included in the field for a patch data unit may then be used to identify and retrieve the parameter set (i.e., set of parameter values) for use with that patch data unit. That is, the field associated with a patch data unit may include an identifier or address for finding additional information related to the patch data unit.

[0014] Some embodiments may further include determining whether a patch data unit of the multi-view data includes original image data captured from at least one viewpoint or includes inpainted data to represent missing image data, and, based on the result of this determination, defining a field value to indicate whether the patch data unit includes original image data or inpainted data. That is, some embodiments may include a process of analyzing the patch data unit to determine whether it includes inpainted data and then setting the value of the field according to the analysis result. Such a process may be performed, for example, when information regarding inpainted data in the multi-view data is not provided by other means (e.g., via user input or from a separate data analysis process).

[0015] According to some embodiments, the field value may include a view parameter. Determining whether a patch data unit of the multi-view data includes original image data captured from at least one viewpoint or inpainted data to represent missing image data may include determining that the patch data unit of the multi-view data includes inpainted data in response to identifying that the patch data unit has a reference to an inpainted view. In such embodiments, the field may be part of the view parameter, and a patch may be identified as an inpainted patch when it references an inpainted view. This may be particularly beneficial for implementations that create a composite background view that is inpainted into the patch data unit.

[0016] Furthermore, embodiments may also include defining a Level of Detail (LoD) value representing a data subsampling factor to be applied to the patch data unit based on the result of the determination. By employing LoD functionality, embodiments may support downscaling of inpainted patch data units.

[0017] The multi-view data may be video data having multiple source views, each of which has texture and depth values. In other words, the method for encoding multi-view data as described above can be applied to a method for encoding immersive video.

[0018] According to another aspect of the present invention, there is provided a method for decoding multi-view data for immersive video as set forth in claim 8. The proposed concept therefore aims to provide schemes, solutions, concepts, designs, methods, and systems related to decoding multi-view data for immersive video. Specifically, embodiments aim to provide a concept for decoding a bitstream including multi-view data and associated metadata encoded according to the proposed embodiment. In such a concept, rendering parameters of a patch data unit are set based on a field indicating that the patch data unit of the multi-view data contains inpainted data. In this way, the proposed field of metadata associated with the multi-view data can be utilized to control view synthesis for the patch data unit, such as rendering priority, rendering order, or blending weights.

[0019] By way of example, in one embodiment, this field may include an identifier for a rendering parameter value. Setting the rendering parameters for a patch data unit may include identifying a rendering parameter value based on the identifier and setting the rendering parameters to the identified rendering parameter value. In this manner, proposed embodiments may be configured to "look up" one or more rendering parameters using the field. For example, multiple rendering parameter sets may be predefined, each having a respective unique identifier, and a parameter set may be selected for use with a patch data unit according to its identifier being included in the field for that patch data unit.

[0020] In some embodiments, the rendering parameters include a rendering priority. Setting the rendering parameters for the patch data unit may include setting the rendering priority of the patch data unit to a first priority value in response to a field indicating that the patch data unit of the multi-view data includes inpainted data, and setting the rendering priority of the patch data unit to a second, different priority value in response to a field indicating that the patch data unit of the multi-view data includes original image data captured from at least one viewpoint. Thus, the importance, or “weight,” of rendering a patch data unit may be controlled according to whether the field associated with the patch data unit indicates that it includes inpainted data. This may allow the order of rendering or view synthesis to be controlled according to preferences or requirements related to the inpainted data.

[0021] Also disclosed is a computer program comprising computer code that, when executed on a processing system, causes the processing system to perform the methods summarized above. The computer program can be stored in a computer-readable storage medium, which may be a non-transitory storage medium.

[0022] There is also provided an encoder for encoding multi-view data for immersive video, as claimed in claim 14.

[0023] Further, there is provided a decoder for decoding multi-view data for immersive video as claimed in claim 16.

[0024] According to yet another aspect, there is provided a bitstream comprising multi-view data for immersive video and associated metadata as claimed in claim 17.

[0025] The bitstream can be encoded and decoded using the methods summarized above, which can be embodied on a computer-readable medium or as a signal modulated onto an electromagnetic carrier wave.

[0026] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]

[0027] For a better understanding of the present invention and to show more clearly how the same may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings in which: [Figure 1] 1 is a flowchart of a method for encoding multi-view data for immersive video according to a first embodiment of the present invention; [Figure 2] 2 is a block diagram of an encoder according to one embodiment configured to perform the method shown in FIG. 1; [Figure 3] 6 is a flowchart illustrating a method for decoding multi-view data for immersive video according to a second embodiment of the present invention. [Figure 4] 4 is a block diagram of a decoder according to an embodiment configured to perform the method shown in FIG. 3; DETAILED DESCRIPTION OF THE INVENTION

[0028] The present invention will now be described with reference to the drawings.

[0029] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the apparatus, systems and methods, are intended for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects and advantages of the apparatus, systems and methods of the present invention will become better understood from the following description, the appended claims and the accompanying drawings. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0030] Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprise" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.

[0031] It should be understood that the drawings are schematic only and are not drawn to scale, and that the same reference numerals will be used throughout the drawings to denote the same or similar parts.

[0032] Implementations according to the present disclosure relate to various techniques, methods, schemes, and / or solutions related to encoding and decoding multi-view data for immersive video. According to the proposed concepts, several solutions may be implemented separately or together. That is, although these possible solutions may be described separately below, two or more of these possible solutions may be implemented in one or another combination.

[0033] MPEG Immersive Video (MIV) has three data streams: texture data, depth data (also called geometry or range data), and metadata. The content is encoded using a standard compression codec (e.g., HEVC), and the metadata includes camera parameters and patch data.

[0034] The term "patch" or "patch data unit" refers to a (rectangular) region (i.e., a patch) within an encoded multiview frame (atlas) of an immersive video. Thus, pixels in a patch refer to a portion in a certain source view and are transformed and projected equivalently. A patch data unit may correspond to a frustum slice or an entire projection plane. That is, a patch is not necessarily limited to a region smaller in size than the entire frame (i.e., a subregion of the frame), but may include the entire frame.

[0035] On the source side, the multiview data corresponds to the entire (i.e., captured) view. In immersive video, the coded multiview frames are usually called atlases and consist of one or more texture and depth (geometry) images.

[0036] Also, references to "rendering priority" should be interpreted as referring to importance or relative weighting, not order. Thus, assigning a high rendering priority to a patch data unit may, but does not necessarily, result in that patch data unit moving toward the front of the rendering queue. Rather, a higher rendering priority may affect the rendering order but may not ultimately change or alter the rendering order due to the relative importance or weighting of other factors in the patch data unit. That is, priority does not necessarily imply a temporal order. The rendering order may depend on the implementation, and different rendering orders for the inpainted data and the original data are possible.

[0037] According to the proposed concept, a method for encoding and decoding multi-view data for immersive video is disclosed. In the proposed encoding method, metadata is generated that includes a field indicating whether a patch data unit of the multi-view data includes inpainted data to represent missing data. The generated metadata provides a way to distinguish between patch data units that include original texture and depth data and patch data units that include inpainted data (e.g., inpainted texture and depth data). Providing such information within the metadata of the immersive video can address issues related to blending and pruned view reconstruction (as part of target view synthesis).

[0038] By providing metadata that includes a field indicating whether a patch data unit of multi-view data includes inpainted data, embodiments can provide a means for indicating the location of the inpainted data within the immersive video. This can also allow patch data units with inpainted data to employ a reduced level of detail LoD, thereby reducing the required bitrate and pixel rate.

[0039] Thus, according to the proposed concept, metadata for an immersive video may be enhanced to indicate the presence, location, and extent of inpainted data within the multi-view data of the immersive video. The proposed encoding method may output (extended) metadata indicating the inpainted data in one or more patches. This (extended) metadata may be used by a corresponding decoding method to render or synthesize a view. Encoders and decoders for multi-view data, as well as corresponding bitstreams with such (extended) metadata, are also provided.

[0040] Fig. 1 shows an encoding method according to a first embodiment of the present invention, and Fig. 2 is a schematic block diagram of an encoder for carrying out the method of Fig. 1.

[0041] The encoder 200 includes an input interface 210 , an analyzer 220 , a metadata encoder 230 , and an output section 240 .

[0042] In step 110, the input interface 210 receives multi-view data including patch data units. In this embodiment, the multi-view data is immersive video data including multiple source views. Each source view includes texture and depth values. The encoding of the texture and depth values ​​is outside the scope of the present invention and will not be described further here. The input interface 210 is coupled to an analyzer 220.

[0043] In step 120, analyzer 220 determines whether a patch data unit of multi-view data contains original image data captured from at least one viewpoint or whether it contains inpainted data to represent missing image data.

[0044] In step 125, the analyzer defines a field value to indicate whether the patch data unit contains original image data or inpainted data based on the determination.

[0045] The task of the analyzer is therefore to identify whether a patch data unit contains original image data or inpainted data and to indicate the results of such analysis. The analyzer 220 provides the results of the analysis to the metadata encoder 230.

[0046] In step 130, the metadata encoder 230 generates metadata 140 including a field indicating whether a patch data unit of the multi-view data includes inpainted data to represent missing data. In this example, the field comprises a binary flag with two permissible values ​​(e.g., a single bit with permissible values ​​of "0" (logical low) and "1" (logical high)). The first value "0" indicates that the patch data unit of the multi-view data includes original image data captured from at least one viewpoint. The second value "1" indicates that the patch data unit of the multi-view data includes inpainted data.

[0047] Thus, the task of the metadata encoder 230 is to generate (extended) metadata that includes a binary flag indicating whether a patch data unit of the multi-view data includes inpainted data to represent missing data. This (extended) metadata includes information defining the patch data units that include inpainted data. Although not in this embodiment, a field of the metadata may be configured to indicate / include further information about the inpainted data of the patch data unit, such as, for example, the LoD of the inpainted data. However, this may not be necessary in some embodiments. For example, the LoD of the inpainted data may be predetermined and / or standardized.

[0048] The output unit 240 generates and outputs the generated (extended) metadata. It may output the metadata as part of a bitstream containing multiview data (i.e., texture and depth data streams) or separately from the bitstream.

[0049] Figure 3 is a flow chart illustrating a method for decoding encoded multi-view data for immersive video according to a second embodiment of the present invention, and Figure 4 is a schematic block diagram of a decoder for performing the method of Figure 3.

[0050] The decoder 400 comprises an input interface 410, a metadata decoder 420, and an output section 430. Optionally, it may also include a renderer 440.

[0051] In step 310, the input interface 410 receives a bitstream containing texture and depth data 305. The input interface 410 also receives metadata 140 describing the bitstream. The metadata may be embedded in the bitstream or separate. The metadata 140 in this example is created according to the method of FIG. 1 above. Thus, the metadata includes a field indicating whether a patch data unit of the multi-view data contains inpainted data to represent missing data. Note that the metadata input to the decoder 400 is typically a version of the metadata output by the encoder 300, which may have subsequently undergone compression (and possibly error-prone communication over a transmission channel).

[0052] In step 320, the metadata decoder 420 decodes the metadata. This includes setting rendering parameters for the patch data units of the multi-view data based on an associated field indicating whether the patch data units contain inpainted data. In this example, the rendering parameter is a rendering priority. In response to the field indicating that the patch data unit contains inpainted data, the rendering priority for the patch data unit is set to a first priority value (e.g., low). In response to the field indicating that the patch data unit contains original image data captured from at least one viewpoint, the rendering priority for the patch data unit is set to a second, higher priority value (e.g., high).

[0053] The metadata decoder 420 provides the rendering parameters to the output unit 430. The output unit 430 outputs the rendering parameters (step 330).

[0054] If the decoder 400 includes an optional renderer 440, the data decoder 420 may provide the decoded rendering parameters to the renderer 440, which reconstructs one or more views according to the rendering parameters. In this case, the renderer 440 may provide the reconstructed views to the output unit 430, which may output the reconstructed views (e.g., to a frame buffer).

[0055] There are various ways in which metadata fields can be defined and used, some of which will now be described in more detail.

[0056] Variation A In some embodiments, a field of the metadata includes a binary flag (e.g., a single bit) that indicates whether a patch data unit of multi-view data contains original image data captured from at least one viewpoint or contains inpainted data to represent missing data.

[0057] In the encoder, the flag is set (i.e., asserted, set to logic high, set to value "1", etc.) when the patch data unit contains original content, and the flag is not set (i.e., negated, set to logic low, set to value "0", etc.) when the patch data unit contains inpainted content.

[0058] At the decoder: When textures of unflagged patches are blended, the blend weight is set to a low value, so when other texture data (flagged) is mapped to the same output location, it gets a substantially higher blending priority, resulting in more optimal quality.

[0059] When a decoder uses "pruned view reconstruction" before actual view synthesis: The reconstruction process is done by selectively allowing only flagged patches, effectively ignoring inpainted data (i.e., treating inpainted data as a low priority). Then, during actual view synthesis, patches that retain inpainted content (i.e., those without the flag set) are used only for areas of missing data.

[0060] Variation B In an alternative embodiment, the metadata is extended so that, for each atlas frame, an "inpaint patch region" (e.g., a rectangle) is specified that is dedicated to patches containing inpainted data. Such a region can be initially specified using user parameters (e.g., as a percentage of the available atlas frame size) or can be automatically determined to balance the available space (determined by the maximum pixel rate) for the original data versus the inpainted data. In this way, a field of the metadata is associated with a frame of encoded multiview data and contains a description (i.e., definition) of one or more patch data units of the frame that contain inpainted data.

[0061] The encoder considers an "inpaint patch region" into which patch data units containing inpainted content are placed, and other patches (containing the original content) are left outside of that region.

[0062] At the decoder, the same behavior applies as described in the previous embodiment: the video encoder may be instructed to encode this region using larger quantization values ​​(i.e., lower quality) for the texture and / or depth video components.

[0063] For MIV implementations where multiple atlas components are packed into a video frame, the patch data units may be part of separate atlases, and the atlases may be packed into the video frame, i.e., one or more portions of the video frame may be reserved for the video data associated with these inpainted patch data units.

[0064] Note that Variant A requires the least amount of changes to the current MIV (draft) standard, as it only adds flags associated with patch data units. It also allows all patch data units to be packed more efficiently (compared to Variant B). Using quality values ​​(e.g., bytes) instead of quality flags (e.g., bits) may have the added advantage that quality may be further optimized.

[0065] Variant B does not require metadata syntax per patch data unit and therefore requires a lower metadata bitrate. Furthermore, patches holding inpainted content can be compactly packed together, which may enable the creation of a separate mesh of triangles to be used for a dedicated inpaint rendering stage (e.g., first creating a backdrop with inpainted data, then compositing using regular patch data).

[0066] As mentioned above in the Background, inpainting missing data in the encoder increases the bit rate and pixel rate. We now describe extensions and / or modifications to the proposed embodiment that aim to limit this increase.

[0067] Downscaling patches with inpainted content It is proposed that painted content be packed into patches using a smaller scale (i.e., a reduced LoD) to reduce bit rates and pixel rates. In particular, it is proposed that some embodiments may be configured to specify an LoD per patch data unit, such that patch data units with in-painted content can use a lower LoD to reduce bit rates and pixel rates.

[0068] The transmission standard used may support a syntax / semantic whereby per-patch data unit LoD specification is enabled by default for inpainted patch data units and disabled by default for regular patches (i.e., patches consisting of original data). Default LoD parameter values ​​may be specified for bitstreams containing inpainted patches.

[0069] A typical implementation may be configured to subsample inpainted data by a factor of 2 and not subsample regular patches, however embodiments may still be configured to override the default LoD parameter values ​​on a per-patch basis (e.g., to use a lower LoD for low-texture portions of a scene).

[0070] Use of low-resolution meshes to represent the background A specific mesh with a minimal / sparse set of vertices can be used to represent the missing background content. The vertices can be accompanied by color data (or references to color data in an existing texture). Such an approach offers the advantage that relatively large background areas can be represented with only a small number of vertices.

[0071] Such a low-resolution mesh may be constructed at the encoder side from the depth map of the source view, however this is not necessarily the case and a textured graphics model may also be used as the background mesh, i.e. a combination of artificial (graphics) data and real camera data may be used.

[0072] The low-resolution mesh with associated texture does not need to be represented in the same projection space as the source view. For example, when the source view has a perspective projection with a given field of view (FoV), a low-resolution background mesh can be defined for a perspective projection with a larger FoV to avoid uncovering at the viewport boundaries. It can also be useful to choose a spherical projection of the background mesh.

[0073] A low-resolution background mesh may require associated metadata to be defined / generated. Accordingly, some embodiments may include generating metadata including a field for defining and / or describing the associated low-resolution mesh. For example, in its simplest form, this field may include a binary flag indicating the presence of a background mesh. Alternatively, this field may be in a form that allows further information to be indicated, such as the position and / or standard projection parameters of depth and texture data. In the absence of such additional information (e.g., rendering parameters), default parameters may be used.

[0074] In the exemplary embodiments described above, the field is described as containing a binary flag or Boolean indicator. However, it should be understood that the proposed field for indicating whether a patch data unit of multi-view data contains inpainted data may be configured to provide additional information beyond a simple binary indication. For example, in some embodiments, the field may contain one or more bytes to indicate a larger range of possible values. The possible values ​​may also include identifiers or addresses of stored values, thus enabling the information to be searched or "looked up."

[0075] For example, multiple rendering parameter sets may be predefined and each stored with its own unique identifier (e.g., address). The identifier included in a field for a patch data unit may then be used to select and retrieve a parameter set for use with the patch data unit. That is, the field associated with the patch data unit may include an identifier or address to identify additional information about the patch data unit.

[0076] Of course, the proposed metadata fields can also be used to provide other information about the inpainted patch data unit. Such features may include (but are not limited to) data quality, rendering preferences, one or more identifiers, etc. Such information may be used in its entirety or piecewise in combination with other information or rendering parameters.

[0077] Embodiments of the present invention rely on the use of metadata to describe patch data units. Because the metadata is important to the decoding process, it is beneficial if the metadata is encoded with additional error detection or error correction codes. Suitable codes are known in the field of communications theory.

[0078] The encoding and decoding methods of Figures 1 and 3 and the encoders and decoders of Figures 2 and 4 may be implemented in hardware or software, or a mixture of both (e.g., as firmware executing on a hardware device). To the extent that an embodiment is implemented partially or entirely in software, the functional steps illustrated in the process flowcharts may be performed by appropriately programmed physical computing devices, such as one or more central processing units (CPUs) or graphics processing units (GPUs). Each process, and its individual component steps illustrated in the flowcharts, may be performed by the same or different computing devices. According to an embodiment, a computer-readable storage medium stores a computer program including computer program code configured to cause one or more physical computing devices to perform an encoding or decoding method as described above when the program is executed on the one or more physical computing devices.

[0079] The storage medium may include volatile and non-volatile computer memory such as RAM, PROM, EPROM, and EEPROM, optical disks (such as CDs, DVDs, and BDs), and magnetic storage media (such as hard disks and tapes). Various storage media may be installed in a mobile computing device or may be transportable such that one or more programs stored on the storage medium are read by a processor.

[0080] Metadata according to an embodiment may be stored on a storage medium. A bitstream according to an embodiment may be stored on the same storage medium or a different storage medium. The metadata may be embedded in the bitstream, but this is not required. Similarly, the metadata and / or the bitstream (with the metadata in the bitstream or separate from it) may be transmitted as a signal modulated onto an electromagnetic carrier wave. The signal may be defined according to a standard for digital communication. The carrier wave may be an optical carrier wave, a radio frequency wave, a millimeter wave, or a short-range communication wave. It may be wired or wireless.

[0081] To the extent that an embodiment is implemented partially or entirely in hardware, the blocks shown in the block diagrams of Figures 2 and 4 may be separate physical components, logical subdivisions of a single physical component, or all integrated into one physical component. The functionality of a block shown in the figures may be split among multiple components in implementation, or the functionality of multiple blocks shown in the figures may be combined into a single component in implementation. Hardware components suitable for use in embodiments of the present invention include, but are not limited to, conventional microprocessors, application-specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs). One or more blocks may be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuitry to perform other functions.

[0082] Variations to the disclosed embodiments can be understood and implemented by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used to advantage. Where a computer program is described above, the computer program can be stored or distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but it may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems. When the term "adapted for" is used in the claims or the description, it has the same meaning as the term "configured to." Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. 1. A method for encoding multi-view data for immersive video, comprising: generating metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; 10. A method according to claim 9, wherein the field has a set of at least two allowed values, a first value in the set indicating that the patch data unit of the multi-view data has original image data captured from at least one viewpoint, a second value in the set indicating that the patch data unit of the multi-view data has inpainted data, and the value of the field indicates a level of detail for the patch data unit.

2. A method for encoding multi-view data for immersive video, comprising: generating metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; The method wherein the field comprises an identifier or address of a stored value.

3. The method of claim 2 , wherein the stored values ​​include rendering parameter values.

4. A method for encoding multi-view data for immersive video, the method comprising: generating metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data, the step of generating metadata comprising: determining whether the patch data unit of the multi-view data comprises original image data captured from at least one viewpoint or comprises inpainted data representing missing image data; determining a value for the field to indicate whether the patch data unit contains original image data or inpainted data based on the result of the determination; A method comprising:

5. the value of the field comprises a view parameter; determining whether a patch data unit of multi-view data has original image data captured from at least one viewpoint or has inpainted data representing missing image data, determining that the patch data unit of multiview data comprises inpainted data in response to identifying that the patch data unit has a reference to an inpainted view; 5. The method of claim 4 dependent on claim 2.

6. 1. A method for decoding multi-view data for immersive video, comprising: receiving a bitstream having multi-view data and associated metadata, the metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; decoding the patch data unit of the multi-view data, the decoding comprising setting rendering parameters of the patch data unit based on the field indicating that the patch data unit of the multi-view data has inpainted data; 10. A method according to claim 9, wherein the field has a set of at least two allowed values, a first value in the set indicating that the patch data unit of the multi-view data has original image data captured from at least one viewpoint, a second value in the set indicating that the patch data unit of the multi-view data has inpainted data, and the value of the field indicates a level of detail for the patch data unit.

7. A method for decoding multi-view data for immersive video, comprising: receiving a bitstream having multi-view data and associated metadata, the metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; decoding the patch data unit of the multi-view data, the decoding comprising setting rendering parameters of the patch data unit based on the field indicating that the patch data unit of the multi-view data has inpainted data; the field having an identifier or address of a stored value; setting the rendering parameters of the patch data unit, Identifying the stored value based on the identifier or address; Setting the rendering parameters based on the stored values.

8. A method for decoding multi-view data for immersive video, comprising: receiving a bitstream having multi-view data and associated metadata, the metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; decoding the patch data unit of the multi-view data, the decoding comprising setting rendering parameters of the patch data unit based on the field indicating that the patch data unit of the multi-view data has inpainted data; the rendering parameters include a rendering priority; setting the rendering parameters of the patch data unit, setting a rendering priority of the patch data unit to a first priority value in response to the field indicating that the patch data unit of the multi-view data has inpainted data; A method for setting the rendering priority of the patch data unit of the multi-view data to a second different priority value depending on the field indicating that the patch data unit has original image data captured from at least one viewpoint.

9. the field is associated with a frame of the multi-view data and contains a description of one or more patch data units of the frame having inpainted data; decoding the patch data unit of the multi-view data, analyzing the description to determine if the patch data unit has inpainted data; setting rendering parameters for the patch data unit based on the results of the analysis; 9. The method according to any one of claims 6 to 8.

10. A method for decoding multi-view data for immersive video, comprising: receiving a bitstream having multi-view data and associated metadata, the metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; decoding the patch data unit of the multi-view data, the decoding comprising setting rendering parameters of the patch data unit based on the field indicating that the patch data unit of the multi-view data has inpainted data; the field is associated with a frame of the multi-view data and contains a description of one or more patch data units of the frame having inpainted data; decoding the patch data unit of the multi-view data, analyzing the description to determine if the patch data unit has inpainted data; setting rendering parameters for the patch data unit based on results of the analysis; The method, wherein the value of the field is a view parameter, and analyzing the description includes determining whether the description has a reference to an inpaint view.

11. A computer program when executed by a processing system causes the processing system to carry out the method of any one of claims 1 to 10.

12. 1. An encoder for encoding multi-view data for immersive video, comprising: a metadata encoder configured to generate metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; an encoder, wherein the field has a set of at least two allowed values, a first value in the set indicating that the patch data unit of the multi-view data has original image data captured from at least one viewpoint, a second value in the set indicating that the patch data unit of the multi-view data has inpainted data, and the value of the field indicates a level of detail for the patch data unit.

13. An encoder for encoding multi-view data for immersive video, the encoder having a metadata encoder configured to generate metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data, the field having an identifier or address of a stored value.

14. The encoder of claim 13, wherein the stored values ​​include rendering parameter values.

15. An encoder for encoding multi-view data for immersive video, comprising: a metadata encoder configured to generate metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data, the metadata encoder comprising: determining whether the patch data unit of the multi-view data comprises original image data captured from at least one viewpoint or comprises inpainted data representing missing image data; an encoder that generates the metadata by determining a value for the field to indicate whether the patch data unit has original image data or inpainted data based on the result of the determination.

16. The value of the field has a view parameter, Determining whether a patch data unit of the multi-view data has original image data captured from at least one viewpoint or has inpainted data representing missing image data includes:

16. An encoder according to claim 15 when dependent on claim 13, wherein the encoder is configured to determine that the patch data unit of multi-view data comprises inpainted data in response to identifying that the patch data unit has a reference to an inpainted view.

17. 1. A decoder for decoding multi-view data for immersive video, comprising: an input interface configured to receive a bitstream having multi-view data and associated metadata, the metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; and a data decoder configured to decode the patch data unit of the multi-view data, the data decoder setting rendering parameters of the patch data unit based on the field indicating that the patch data unit of the multi-view data has inpainted data; a decoder, wherein the field has a set of at least two allowed values, a first value in the set indicating that the patch data unit of the multi-view data has original image data captured from at least one viewpoint, a second value in the set indicating that the patch data unit of the multi-view data has inpainted data, and the value of the field indicates a level of detail for the patch data unit.

18. A decoder for decoding multi-view data for immersive video, comprising: an input interface configured to receive a bitstream having multi-view data and associated metadata, the metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; and a data decoder configured to decode the patch data unit of the multi-view data, the data decoder setting rendering parameters of the patch data unit based on the field indicating that the patch data unit of the multi-view data has inpainted data; the field having an identifier or address of a stored value; setting the rendering parameters of the patch data unit, Identifying the stored value based on the identifier or address; A decoder that sets the rendering parameters based on the stored values.

19. A decoder for decoding multi-view data for immersive video, comprising: an input interface configured to receive a bitstream having multi-view data and associated metadata, the metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; and a data decoder configured to decode the patch data unit of the multi-view data, the data decoder setting rendering parameters of the patch data unit based on the field indicating that the patch data unit of the multi-view data has inpainted data; the rendering parameters include a rendering priority; setting the rendering parameters of the patch data unit, setting a rendering priority of the patch data unit to a first priority value in response to the field indicating that the patch data unit of the multi-view data has inpainted data; A decoder that sets the rendering priority of the patch data unit of the multi-view data to a second different priority value in response to the field indicating that the patch data unit has original image data captured from at least one viewpoint.

20. The field is associated with a frame of the multi-view data and has a description of one or more patch data units of the frame having inpainted data; decoding the patch data units of the multi-view data, analyzing the description to determine if the patch data unit has inpainted data; A decoder according to any one of claims 17 to 19, configured to set rendering parameters of the patch data unit based on the results of the analysis.

21. A decoder for decoding multi-view data for immersive video, comprising: an input interface configured to receive a bitstream having multi-view data and associated metadata, the metadata having a field indicating whether a patch data unit of the multi-view data has inpainted data representing missing data; and a data decoder configured to decode the patch data unit of the multi-view data, the data decoder setting rendering parameters of the patch data unit based on the field indicating that the patch data unit of the multi-view data has inpainted data; the field is associated with a frame of the multi-view data and contains a description of one or more patch data units of the frame having inpainted data; decoding the patch data units of the multi-view data, analyzing the description to determine if the patch data unit has inpainted data; setting rendering parameters for the patch data unit based on results of the analysis; A decoder, wherein the value of the field is a view parameter, and analyzing the description includes determining whether the description has a reference to an inpaint view.

Citation Information

Patent Citations

  • 3D video format

    JP2012518367A

  • Point cloud compression using hybrid transforms

    US20190122393A1

Cited By

  • Encoding and decoding of immersive video

    JP2026032008A