Method and device for encoding and decoding a multi-view video sequence

The method optimizes patch arrangement and transformation in atlases to enhance immersive video quality and reduce computational complexity in devices with limited resources, addressing inefficiencies in existing immersive video encoding and decoding methods.

JP7708786B2Active Publication Date: 2025-07-15オランジュ
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022564436
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-04-22
Filing Date
2021-03-29
Publication Date
2025-07-15
Estimated Expiration
2041-03-29

AI Technical Summary

Technical Problem

Existing immersive video encoding and decoding methods fail to provide sufficient visual quality and immersion due to inefficient processing and transmission of multi-view video data, particularly in devices with limited resources.

Method used

A method and device for encoding and decoding multi-view video that optimizes patch arrangement and transformation in atlases by applying transformations such as oversampling, subsampling, and pixel value corrections to reduce computational complexity and improve compression efficiency.

Benefits of technology

Enhances immersive video quality by optimizing pixel occupancy and reducing computational complexity, allowing devices with limited resources to process multi-view videos more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708786000001
    Figure 0007708786000001
  • Figure 0007708786000002
    Figure 0007708786000002
  • Figure 0007708786000003
    Figure 0007708786000003
Patent Text Reader

Abstract

The coded data stream includes coded data representing an atlas, the atlas corresponding to an image including patches, the patches corresponding to sets of pixels extracted from components of views of a multi-view video, the views not being coded in the coded data stream. The method includes decoding the atlas including decoding the patches from the coded data stream, determining for the decoded patches whether and which transformations must be applied to the decoded patches, the transformations belonging to a group of transformations including oversampling of the patches or modifying pixel values ​​of the patches, and applying the determined transformations to the decoded patches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to immersive videos representing scenes captured by one or more cameras. More particularly, the present invention relates to the encoding and decoding of such videos.

Background Art

[0002] In the context of immersive videos, i.e., when the viewer has the feeling of being immersed in the scene, the scene is generally captured by a set of cameras as shown in FIG. 1. These cameras can be of type 2D (cameras C1, C2, C3, C4 in FIG. 1) or of type 360, i.e., those that capture the entire scene 360 degrees around the camera (camera C5 in FIG. 1).

[0003] All of these captured views have conventionally been encoded and then decoded by the viewer's terminal. However, it is not sufficient to display only the captured views to provide sufficient immersive quality and thus the visual quality of the scene displayed to the viewer and a good immersion into that scene.

[0004] To improve the immersion into the scene, one or more views, commonly called intermediate views, are usually calculated from the decoded views.

[0005] These intermediate views can be calculated by a view synthesis algorithm.

[0006] Generally, for example, in the MIV system (Metadata for Immersive Video) which is currently being standardized, not all of the original views, i.e., all those captured by the cameras, are necessarily sent to the decoder. A selection, also called "pruning", of the data that can be used to synthesize intermediate viewpoints is made from at least a part of the original views.

[0007] FIG. 2 shows an example of an encoding-decoding system that synthesizes an intermediate view on the decoder side using such data selection of multi-view video.

[0008] According to this method, one or more basic views (T in FIG. 2 b , D b ) are encoded by a 2D encoder, such as an HEVC encoder, or by a multi-view encoder.

[0009] The remaining views (T s , D s ) are processed to extract a specific zone from each of these views. The extracted zones, hereinafter also referred to as patches, are collected in an image called an atlas. The atlas is encoded, for example, by a conventional 2D video encoder, such as an HEVC encoder. On the decoder side, the atlas is decoded, whereby the decoded patches are supplied to a view synthesis algorithm to generate an intermediate view from the basic views and the decoded patches. Overall, the patches enable the same zone to be transmitted from different viewpoints. Specifically, the patches enable occlusion, i.e., parts of the scene that are invisible from a given view of the scene, to be transmitted.

[0010] The MIV system (MPEG-I Part 12) generates an atlas formed by a set of patches in its reference implementation (TMIV, representing "Test Model for Immersive Video").

[0011] Figure 3 shows an example of extracting patches (Patch 2, Patch 5, Patch 8, Patch 3, Patch 7) from views (V0, V1, V2) and creating related atlases, for example, two atlases A0 and A1. These atlases A0 and A1 each include texture images T0, T1 and corresponding depth maps D0, D1. Atlas A0 has texture T0 and depth D0, and atlas A1 has texture T1 and depth D1.

[0012] As described in Figure 2, the patches are collected within the image and encoded by a conventional 2D video encoder. To avoid the extra cost of signal transmission and encoding of the extracted patches, it is necessary to perform an optimal arrangement of the patches within the atlas. Furthermore, if there is a large amount of information to be processed by the decoder to reconstruct the views of the multi-view video, it is necessary not only to reduce the cost of compressing such patches, but also to reduce the number of pixels that the decoder has to process. In fact, in many application fields, the devices for playing such videos have more limited resources than the devices for encoding such videos.

Prior Art Documents

Non-Patent Documents

[0013]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0014] Therefore, it is necessary to improve the prior art.

Means for Solving the Problems

[0015] The present invention improves the prior art. For this purpose, the present invention relates to a method for decoding a coded data stream representing multi-view video, wherein the coded data stream includes coded data representing at least one atlas, the at least one atlas corresponds to an image including at least one patch, the at least one patch corresponds to a set of pixels extracted from at least one component of a view of the multi-view video, and the view is not coded in the coded data stream. This decoding method includes - decoding the at least one atlas, including decoding the at least one patch, from the coded data stream; - for the at least one decoded patch, determining whether a transformation must be applied to the at least one decoded patch and which transformation must be applied, the transformation belonging to a group of transformations including at least one oversampling of the patch or a correction of the pixel values of the patch; - applying the determined transformation to the decoded patch and includes.

[0016] In this regard, the present invention relates to a method for coding a data stream representing multi-view video, the coding method including - extracting at least one patch corresponding to a set of pixels of the component from at least one component of a view of the multi-view video that is not coded in the data stream; - for the at least one extracted patch, determining whether a transformation must be applied to the at least one patch and which transformation must be applied, the transformation belonging to a group of transformations including at least one subsampling of the patch or a correction of the pixel values of the patch; - Applying the determined transformation to the at least one patch; - Encoding at least one atlas in the data stream, wherein the at least one atlas corresponds to an image that at least includes the at least one patch; also relates to a method including the above.

[0017] Thus, according to the present invention, it becomes possible to identify which patches of the decoded atlas must be transformed during reconstruction. Such a transformation corresponds to the inverse transformation of the transformation applied during the encoding of the atlas.

[0018] The present invention can also apply a transformation having different transformations for each patch, or different parameters, to the patches of the atlas.

[0019] Thus, the arrangement of the patches in the atlas is optimized for compression. In fact, the transformation used for the patches of the atlas can optimize the pixel occupancy rate of the atlas by using transformations such as rotation, subsampling, encoding, etc. to arrange the patches in the atlas image.

[0020] On the other hand, the transformation can be optimized by modifying the cost of compressing the patches, in particular by reducing the pixel values of these patches, for example by reducing the dynamic range of the pixels, or by using subsampling which results in fewer pixels to be encoded, or by using the optimal arrangement of the patches in the atlas image that allows the fewest possible pixels to be obtained. Reducing the pixel occupancy rate of the atlas reduces the ratio of pixels to be processed by the decoder, and thus reduces the computational complexity of decoding.

[0021] According to a particular embodiment of the present invention, whether the conversion must be applied to the at least one decoded patch is determined from at least one syntax element decoded from the coded data stream for the at least one patch. According to this particular embodiment of the present invention, syntax elements are explicitly coded in the data stream to indicate whether the conversion must be applied to the decoded patch and which conversion must be applied.

[0022] According to another particular embodiment of the present invention, the at least one decoded syntax element includes at least one indicator indicating whether the conversion must be applied to the at least one patch, and if the indicator indicates that the conversion must be applied to the at least one patch, the at least one syntax element optionally includes at least one parameter of the conversion. According to this particular embodiment of the present invention, the conversion to be applied to the patch is coded in the form of an indicator indicating whether the conversion must be applied to the patch and, in the case where it must be applied, probably one or more parameters of the conversion to be applied. For example, a binary indicator can indicate whether the conversion must be applied to the patch, and if it must be applied, the code indicates which conversion is used and probably one or more parameters of the conversion, such as magnification, a correction function for the pixel dynamic range, a rotation angle, etc.

[0023] In other embodiments, the parameters of the conversion can be set by default in the encoder.

[0024] According to another particular embodiment of the present invention, the at least one parameter of the conversion to be applied to the patch has a value that is predictively coded with respect to a predicted value. Thus, this particular embodiment of the present invention can save the signaling cost of the parameters of the conversion.

[0025] According to another specific embodiment of the present invention, the predicted value is encoded within the header of the view, or within the header of the components of the atlas, or within the header of the atlas.

[0026] According to another specific embodiment of the present invention, the predicted value is - patches processed previously according to the processing order of the patches of the atlas, - patches processed previously, extracted from the same component as the component to which at least one patch of the views of the multi-view video belongs, - patches selected from a set of candidate patches using an index encoded within the data stream, - patches selected from a set of candidate patches using a criterion corresponding to the values of the parameters of the transformation applied to the patches belonging to a group including

[0027] According to another specific embodiment of the present invention, the determination as to whether a transformation must be applied to the at least one decoded patch for the at least one decoded patch is carried out when a syntax element decoded from the header of the data stream indicates the activation of the application of the transformation to the patches encoded within the data stream, and the syntax element is encoded within the header of the view, or within the header of the components of the view, or within the header of the atlas. According to this specific embodiment of the present invention, high-level syntax elements are encoded within the data stream in order to signal the use of the transformation to be applied to the patches of the multi-view video. Thus, the additional cost incurred by encoding the parameters of the transformation at the patch level is avoided when these transformations are not used. In addition, this specific embodiment of the present invention can limit the computational complexity of the decoding when these transformations are not used.

[0028] According to another specific embodiment of the present invention, if the characteristics of the at least one decoded patch meet the criteria, it is determined that a transformation must be applied to the decoded patch. According to this specific embodiment of the present invention, an indication indicating the use of the transformation to be applied to the patch is not explicitly encoded in the data stream. Such an indication is inferred from the characteristics of the decoded patch. In this specific embodiment of the present invention, patch transformation can be used without the need for the additional cost of encoding to signal the use of the transformation.

[0029] According to another specific embodiment of the present invention, the characteristic corresponds to a ratio R = H / W, where H corresponds to the height of the at least one decoded patch and W corresponds to the width of the at least one decoded patch, and when the ratio is included within a determined interval, the transformation to be applied to the at least one patch corresponds to vertical oversampling at a predetermined magnification. Thus, according to this specific implementation mode of the present invention, it is possible to mix within the same atlas "long" patches for which subsampling is not of interest to be performed on them and "long" patches for which subsampling is performed on them without the need to signal that subsampling is to be performed on them.

[0030] According to another specific implementation mode of the present invention, the characteristic corresponds to an energy E calculated from the pixel values of the at least one decoded patch, and when the energy E is lower than a threshold value, the transformation to be applied to the at least one patch corresponds to the multiplication of the pixel values by a determined magnification.

[0031] According to another specific embodiment of the present invention, when several transformations must be applied to the same patch, the order in which the transformations must be applied is predefined. In this specific embodiment of the present invention, signaling to indicate the order in which the transformations are applied is not required. This order is defined in the encoder and decoder and does not change for all patches to which these transformations are applied.

[0032] The present invention relates to a device for decoding a coded data stream representing multi-view video, wherein the coded data stream includes coded data representing at least one atlas, the at least one atlas corresponds to an image including at least one patch, the at least one patch corresponds to a set of pixels extracted from at least one component of a view of the multi-view video, the view is not coded within the coded data stream, and the decoding device - decoding the at least one atlas, including decoding the at least one patch, from the coded data stream, - for the at least one decoded patch, determining whether a transformation must be applied to the at least one decoded patch and which transformation must be applied, the transformation belonging to a group including at least one oversampling of the patch or correction of pixel values of the patch, - applying the determined transformation to the decoded patch and also relates to a device comprising a processor and a memory configured to perform the above.

[0033] According to a particular embodiment of the present invention, such a device is included within a terminal.

[0034] The present invention also relates to a device for coding a data stream representing multi-view video, - extracting at least one patch corresponding to a set of pixels of the component from at least one component of a view of the multi-view video that is not coded within the data stream, - Determining, for the at least one extracted patch, whether the at least one patch must be subject to a transformation and which transformation must be applied, the transformation belonging to a group of transformations including at least one subsampling of the patch or a correction of the pixel values of the patch - Applying the determined transformation to the at least one patch - Encoding at least one atlas in the data stream, the at least one atlas corresponding to an image including at least the at least one patch Also relates to a device comprising a processor and a memory configured to perform the above

[0035] According to a particular embodiment of the invention, such a device is included within a terminal

[0036] The encoding method or decoding method according to the invention can each be implemented in various ways, in particular in wired form or in software form. According to a particular embodiment of the invention, the encoding method or decoding method is implemented by a computer program. The invention also relates to a computer program comprising instructions for implementing an encoding method or a decoding method according to any one of the specific embodiments described above when the program is executed by a processor. Such a program can use any programming language. Such a program can be downloaded from a communication network and / or recorded on a computer-readable medium

[0037] This program can use any programming language and can take the form of source code, object code, or an intermediate code between source code and object code in a partially compiled form or any other desired form

[0038] The present invention also relates to a computer-readable storage medium or data medium containing the instructions of the computer program described above. The above-described recording medium can be any entity or device capable of storing a program. For example, the medium can include a ROM, a storage means such as a CD-ROM or a super-small electronic circuit ROM, a USB flash drive, or a magnetic recording means such as a hard drive. On the other hand, the recording medium can correspond to a transmissible medium such as an electrical signal or an optical signal that can be transmitted through an electrical cable or an optical cable, wirelessly, or by other means. The program according to the present invention can be downloaded particularly on an Internet type network.

[0039] Alternatively, the recording medium can correspond to an integrated circuit in which the program is embedded, and this circuit is adapted to execute the method or to be used in the execution of the method.

[0040] Other features and advantages of the present invention will become more clearly understood from the following description of a specific embodiment provided as a simple, illustrative, non-limiting example, and the accompanying drawings.

Brief Description of the Drawings

[0041]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

DETAILED DESCRIPTION OF THE INVENTION

[0042] FIG. 4 shows steps of a multi-view video encoding method in at least one encoded data stream according to a specific embodiment of the present invention.

[0043] According to the present invention, a multi-view video is encoded according to the encoding method shown with respect to FIG. 2, where one or more base views are encoded in the data stream, and sub-images or patches including texture and depth data are also encoded in the data stream. These patches are derived from additional views that are not necessarily fully encoded in the data stream. These patches and one or more base views enable the decoder to synthesize other views of the scene, hereinafter also referred to as virtual views, or synthesized views, or even intermediate views. These synthesized views are not encoded in the data stream. The steps of such an encoding method according to a specific embodiment of the present invention will be described below.

[0044] For example, here, it is considered that the scene is captured by a set of cameras C1, C2,..., C N as shown in FIG. 1. Each camera generates a view including at least one so-called texture component that varies over time. In other words, the texture component of a view is a sequence of 2D images corresponding to the images captured by the camera placed at the viewpoint of that view. Each view also includes a depth component called a depth map, which is determined for each image in the view.

[0045] A depth map can be generated by known methods, such as by estimating depth using texture or by capturing volumetric data from a scene using light detection and ranging (Lidar) technology.

[0046] Hereinafter, the term "view" is used to denote a sequence of texture images and depth maps representing a scene captured from a certain viewpoint. By misuse of language, the term "view" may also mean the texture image and depth map of a view at a given time.

[0047] When views of a multi-view video are captured, the encoder then proceeds with the steps described below, for example, according to the coding method defined in Basel Salahieh, Bart Kroon, Joel Jung, Marek Domanski, Test Model 4 for Immersive Video, ISO / IEC JTC 1 / SC 29 / WG 11 N19002, Brussels, BE - January 2020.

[0048] In step E40, one or more base views are selected from the captured views of the multi-view video.

[0049] The base views are selected from the set of captured views of the multi-view video by known methods. For example, spatial subsampling can be performed to select one view from two views. In another example, the content of the views can be used to determine which views should be retained as base views. In yet another example, camera parameters (position, orientation, focus) can be used to determine the views that must be selected as base views. At the end of step E40, a certain number of views to be base views are selected.

[0050] The remaining views not selected as base views are called "additional views".

[0051] In step E41, apply the pruning method to the additional views to identify one or more patches to be sent to the decoder for each additional view. In this step, the patches to be sent are determined by extracting the zones necessary for intermediate view synthesis from the additional view images. For example, such zones correspond to occlusion zones that are invisible within the base view, or zones that are visible within the base view and have undergone a change in illumination or are of lower quality. The extracted zones can be of any size and shape. Implement clustering of pixels connected to their neighbours to create one or more rectangular patches from the extracted zones of the same view that are easier to encode and place.

[0052] In step E42, for each patch, the encoder determines one or more transformations to be applied to the patch when the patch is placed within the atlas.

[0053] It will be recalled that the patches can be patches having a texture component and / or a depth component.

[0054] The patches are placed within the atlas so as to minimize the encoding cost of the atlas and / or to reduce the number of pixels to be processed by the decoder. To achieve this, the patches can undergo transformations including · Subsampling by a factor of Nv in the vertical dimension · Subsampling by a factor of Nh in the horizontal dimension · Subsampling by a factor of Ne in each dimension · Modification of the pixel values contained within the patch · Rotation of the patch by an angle of i*90°, where i = 0, 1, 2, or 3 The encoder then examines each patch once and determines one or more transformations to be applied to the patch.

[0055] ​

[0056] In one variant form, the list of conversions to be tested on the patch may also include the "identity" conversion, in other words the zero transformation (no transformation).

[0057] The selection of a conversion from among the possible conversions can be done by evaluating a rate-distortion criterion which is calculated using, for the reconstructed signal, the rate required to encode the transformed patch and the distortion calculated between the original patch and the transformed patch that has been encoded and then reconstructed. The selection can also be done based on an assessment of the quality of additional views synthesized using the patch being processed.

[0058] For each conversion, one or more parameters can be tested.

[0059] For example, in the case of subsampling, various magnification factors Nv, Nh, and Ne can be tested. In a preferred embodiment, the magnification factors Nv, Nh, and Ne are equal to 2. In other embodiments, other values such as 4, 8, or 16 are possible.

[0060] A conversion corresponding to a change in pixel values is also called a "mapping". Such a mapping conversion can consist, for example, of dividing all the pixel values of the patch by a given value Dv. For example, Dv is equal to 2. However, other values such as 4, 8, or 16 are possible.

[0061] In another example, the mapping is a parameterized function f PIt may be to convert the x value of a pixel to a new y value using (x)=y. Such a function may be, for example, a linear function per part, and each part is parameterized by its starting abscissa x1, as well as the parameters a and b of the linear function y = ax + b. In that case, the parameter P of the conversion is a triple list (x1, a, b) for each linear part of the mapping.

[0062] In another example, the mapping can also be a look-up table (LUT) which is a table associating the value y with the input x.

[0063] In the case of a rotation transformation, the rotation transformation can be a 180° vertical rotation, also known as a vertical flip. It is also possible to test angular values defined by other rotation parameter values, for example, i * 90°, where i = 0, 1, 2, or 3.

[0064] In determining the transformation associated with a patch, the number of atlases available for the encoding of multi-view video can be taken into account to globally optimize the rate / distortion cost of the encoding of the atlas or the quality of the intermediate view synthesis, and the arrangement of the patches within the atlas can also be simulated.

[0065] At the end of step E42, a list of transformed patches becomes available. Each patch is associated with the transformation determined for that patch and the related parameters.

[0066] During step E43, these patches are arranged within one or more atlases. The number of atlases is determined according to parameters defined as input to the encoder, such as, for example, the size (length and height) of the atlas, or the maximum number M of pixels for texture and depth for all atlases per given time or image. This maximum number M corresponds to the number of pixels that should be processed at once by the decoder for multi-view video.

[0067] In the specific embodiment described herein, each base view is considered to constitute a patch that includes a texture component and a depth component at a given time of that base view and is encoded within the atlas. In this specific embodiment, there are as many atlases as there are base views, and only as many atlases as are necessary to transfer all the patches extracted from the additional views.

[0068] Depending on the size of the atlas provided as input, the atlas can consist of base views and patches, or, if the size of the views is larger than the size of the atlas, the base views can be split and represented on several atlases.

[0069] According to the specific embodiment described herein, the patches of the atlas can in that case correspond to the entire image of the base view, or to a part of the base view, or to zones extracted from additional views.

[0070] The texture pixels of the patch are arranged within the texture component of the atlas, and the depth pixels of the patch are arranged within the depth component of the atlas.

[0071] The atlas can include only one texture component or depth component, or can include both a texture component and a depth component. In other examples, the atlas can include other types of components that contain information useful for intermediate view synthesis. For example, other types of components can include information such as a reflectance index for indicating how transmissive the corresponding zone is, or reliability information about the depth value at that location.

[0072] During step E43, the encoder scans all patches in the patch list. For each patch, the encoder determines in which atlas this patch is encoded. This list includes both transformed and non-transformed patches. Non-transformed patches include patches that have received a zero transformation or an identity transformation, zones extracted from additional views, or patches that include the image of the base view. Here, when a patch must be transformed, consider that the patch has already been transformed.

[0073] An atlas is a set of patches spatially rearranged within an image that is intended to be encoded. The purpose of this arrangement is to maximize the use of the space within the atlas image to be encoded. In fact, one of the purposes of video encoding is to minimize the number of pixels to be decoded before a view can be synthesized. For this purpose, the patches are arranged within the atlas such that the number of patches within the atlas is maximized. Such a method is described in Basel Salahieh, Bart Kroon, Joel Jung, Marek Domanski, Test Model 4 for Immersive Video, ISO / IEC JTC 1 / SC 29 / WG 11 N19002, Brussels, BE - January 2020.

[0074] Upon receiving step E43, a list of patches for each atlas is generated. Note that this arrangement also determines the number of atlases to be encoded for a given time.

[0075] During step E44, the atlases are encoded into the data stream. In this step, each atlas containing a texture component and / or a depth component in the form of a 2D image is encoded using a conventional video encoder such as HEVC, VVC, MV-HEVC, 3D-HEVC. As explained above, the base view is considered here as a patch. Thus, the encoding of the atlases involves the encoding of the base view.

[0076] During step E45, the information associated with each atlas is encoded within the data stream. This information is typically encoded by an entropy encoder.

[0077] For each atlas, the list of patches includes the following items for each patch in the list. · The location of the patch within the atlas, in the form of 2D coordinates, e.g., the position of the upper left corner of the rectangle representing the patch. · The location of the patch within the original view of the patch, i.e., the position of the patch within the image of the view from which the patch was extracted, in the form of 2D coordinates, e.g., the position of the upper left corner of the rectangle representing the patch within the image. · The dimensions (length and height) of the patch. · The identifier of the original view of the patch. · Information regarding the transformation applied to the patch.

[0078] In step E45, for at least some patches of the atlas, information regarding the transformation to be applied to the patch during decoding is encoded within the data stream. The transformation to be applied to the patch during decoding corresponds to the inverse transformation determined above that was applied to the patch when placing the patches of the atlas.

[0079] In a particular embodiment of the present invention, information indicating the transformation to be applied is transmitted for each patch.

[0080] In the particular embodiment described herein, it is considered that what is indicated is the transformation to be applied during decoding, rather than the transformation applied during encoding (which corresponds to the inverse transformation of decoding). For example, when subsampling is applied during encoding, oversampling is applied during decoding. In other particular embodiments of the present invention, it will be clearly understood that the transmitted information regarding the transformation to be applied may correspond to the information indicating the transformation applied during encoding, and in that case, the decoder infers the transformation to be applied from this information.

[0081] For example, the information indicating the transformation to be applied can be an index indicating the transformation to be applied within a list of possible transformations. Such a list can further include an identity transformation. Therefore, in the case where the zero transformation is applied to the patch, the index indicating the identity transformation can be encoded.

[0082] In another embodiment, a binary indicator can be encoded to indicate whether the patch is transformed. If the binary indicator indicates that the patch has been transformed, an index indicating which transformation from the list of possible transformations should be applied is encoded.

[0083] In one embodiment where only one transformation to be applied is possible, only a binary indicator can be encoded to indicate whether the patch is transformed.

[0084] The list of possible transformations can be known to the decoder and thus need not be transmitted within the data stream. In other embodiments, the list of possible transformations can be encoded within the data stream, for example, within the header of a view or within the header of a multi-view video.

[0085] The parameters associated with the transformation to be applied can also be defined by default and can be known to the decoder. In another specific embodiment of the present invention, the parameters associated with the transformation applied to the patch are encoded within the data stream for each patch.

[0086] When the transformation corresponds to oversampling in one or both dimensions (equivalent to the same subsampling during encoding), the parameters associated with the transformation can correspond to the values of the interpolation to be applied for all dimensions or the values of the interpolation to be applied for each dimension.

[0087] When the transformation corresponds to the modification of the pixel values of the patch to be coded by mapping using parameters, the parameters of this transformation correspond to the characteristics of the mapping to be applied, i.e., the parameters of a piecewise linear linear function, a look-up table (LUT), etc. In particular, the possible LUTs can be made known to the decoder.

[0088] When the transformation corresponds to a rotation, the parameter corresponds to the rotation angle selected from among the possible rotations.

[0089] The parameters associated with the transformation can be coded as they are or by prediction relative to a predicted value.

[0090] In one embodiment according to a variant form, in order to predict the value of the parameter, a predicted value is determined and it can be coded in the header of the view containing the current patch, or in the header of the component, or in the header of the view image, or even in the header of the atlas, in the data stream.

[0091] Thus, for a given atlas, the value P of the parameter is predicted by the value Ppred coded at the atlas level. In that case, the difference between Ppred and P is coded for each patch of the atlas.

[0092] In another embodiment, in order to predict the value of the parameter, the predicted value Ppred may correspond to the value of the parameter used for the previously processed patch. For example, the previously processed patch can also be the previous patch in the patch processing order, or the previous patch belonging to the same view as the current patch.

[0093] The predicted value of the parameter can also be obtained by a mechanism similar to the "merge" mode of the HEVC encoder. For each patch, a list of candidate patches is determined and an index indicating one of these candidate patches is coded for that patch.

[0094] In another embodiment, it is not necessary to transmit an index because a patch can be identified from the list of candidate patches using a criterion. Thus, for example, a patch that maximizes the degree of similarity to the current patch can be selected, or even a patch whose dimensions are closest to those of the current patch can be selected.

[0095] In other alternative embodiments, when the use of a transformation is enabled, information indicating whether a patch must undergo the transformation can be decomposed into a part indicating the use of the transformation (e.g., a binary indicator) and a part indicating the parameters of the transformation. This signaling mechanism can be used independently for each possible transformation for a patch.

[0096] In a particular embodiment of the present invention, a binary indicator can be coded at the level of the header of the atlas, or of the view, or of the component, in order to activate the use of a determined transformation on a patch of the atlas, or on a patch of the view, or on a patch of the component. In that case, the application of the determined transformation to the patch depends on the value of this binary indicator.

[0097] For example, two binary indicators I A and I B associated with the activation of transformation A and the activation of transformation B respectively are coded within the header of the atlas. The value of binary indicator I A indicates that the use of transformation A is possible, while the value of binary indicator I B indicates that the use of transformation B is not possible. In this example, for each patch, the binary indicator indicates whether transformation A is applied to that patch and perhaps the associated parameters. In this example, for each patch, it is not necessary to code a binary indicator to indicate whether transformation B is applied to that patch.

[0098] This particular embodiment, which activates the use of the transformation at the patch level or at a higher level, can, in particular, save the cost of signaling when no patch uses this transformation.

[0099] When this binary activation indicator is coded at the view or component level, its value applies to all patches belonging to that view or component, regardless of which atlas the patch is coded in. Thus, an atlas may contain patches to which a particular transformation can be applied according to the indicator coded for that patch, and patches to which the same transformation cannot be applied. In the case of this latter type of patch, the indicator for this transformation is not coded within the patch information.

[0100] In another particular embodiment of the invention, information indicating the transformation is not coded at the patch level. The transformation is inferred from the characteristics of the patch in the decoder. The transformation is then applied to the patch as soon as a particular criterion is met. This particular mode will be described in more detail below with respect to the decoding process.

[0101] FIG. 5 shows the steps of a method for decoding a coded data stream representing a multi-view video according to a particular embodiment of the invention. For example, the coded data stream is generated by the coding method described with respect to FIG. 4.

[0102] During step E50, the atlas information is decoded. This information is generally decoded by an appropriate entropy decoder.

[0103] The information includes a list of patches and the following elements for each patch. - The location of the patch in the atlas, in the form of coordinates, - The location of the patch in the original view of the patch, in the form of coordinates, - The dimensions of the patch, - Identifier of the original view of the patch, - Information indicating whether the conversion must be applied to the patch.

[0104] Similar to the coding method, this information can be an index indicating the conversion from a list of possible conversions, or an indicator indicating whether the conversion must be applied to the patch for each possible conversion.

[0105] In the case of a conversion corresponding to the same oversampling in both dimensions, the information can be a binary indicator indicating the use of the conversion, or the value of the interpolation to be applied to all dimensions.

[0106] In the case of a conversion corresponding to separate oversampling in two dimensions, the information can correspond to a binary indicator indicating the use of the conversion, or for each dimension, correspond to the value of the interpolation to be applied.

[0107] In the case of a conversion corresponding to the correction of the pixels of the patch to be decoded by mapping using parameters, the information can include an information item indicating the use of the mapping and perhaps information representing the characteristics of the mapping to be applied (parameters of a partially linear linear function, look-up table, etc.).

[0108] In the case of a conversion corresponding to rotation, the parameter indicates which rotation was selected from the possible rotations.

[0109] The transmitted information that can identify the conversion to be applied to the patch is decoded in a manner suitable for the applied coding. Therefore, the information can be decoded as it is (direct decoding) or prediction-decoded in a manner similar to the encoder.

[0110] According to a particular embodiment of the present invention, when the use of the conversion is activated, the information for identifying the conversion to be applied to the patch can include a part indicating the use of the conversion (binary indicator) and a part indicating the parameters of the conversion.

[0111] Similar to the case of the coding method, according to a particular embodiment of the present invention, the decoding of an item of information for identifying the conversion to be applied to a given patch can depend on the activation binary indicator encoded in the header of the atlas, in the header of the view, or in the header of the component to which the patch belongs.

[0112] According to another particular embodiment of the present invention, the information for identifying the conversion to be applied to the patch is not encoded together with the patch information, but is obtained from the characteristics of the decoded patch.

[0113] For example, in one embodiment, the energy of the decoded pixels in the patch is measured by calculating the root mean square error of the patch. If this energy is below a given threshold, for example, a root mean square error of less than 100, the pixel values of the patch are converted by multiplying all the values of the patch by a specified magnification Dv. For example, Dv = 2. Other thresholds and other patch value correction magnifications are possible.

[0114] According to another variant, if the ratio of the decoded dimensions of the patch, where H is the height of the patch and W is the length of the patch, is within a given range, for example, 0.75 < H / W < 1.5, the patch is interpolated in the vertical dimension at a given magnification, for example, magnification 2. The patch dimensions considered here are the patch dimensions decoded from the atlas information in which the patch is encoded. These are the dimensions of the patch before the conversion for the decoder (and thus after the conversion for the encoder). When it is determined that the H / W ratio is within the determined range, the patch is oversampled and its dimensions are recalculated as a result.

[0115] With this variant, it becomes possible to mix, within the same atlas, "long" patches for which subsampling is not of interest and to which it is applied, without signaling the criteria that would allow them to be interpolated in the decoder, with "long" patches for which subsampling is applied without signaling. Other thresholds, for example more restrictive values such as 0.9 < H / W < 1.1, can be used.

[0116] During step E51, the components of the atlas are decoded. Each atlas containing 2D texture components and / or 2D depth components is decoded using a conventional video decoder such as AVC or HEVC, VVC, MV-HEVC, 3D-HEVC, etc.

[0117] During step E52, the transformation identified in step E50 is applied to the texture components and / or depth components of each patch within the atlas of each patch, depending on whether the transformation is applied to the texture component, the depth component, or both components, so that the decoded patches are reconstructed.

[0118] For additional views, this step consists of modifying each patch individually by applying the transformation identified for this patch. This can be done in several ways, for example, by modifying the pixels of that patch within the atlas containing the patch, by copying the modified patch into a buffer memory zone, or by copying the transformed patch into its associated view.

[0119] Depending on the information decoded earlier, one of the following transformations can be applied to each patch to be reconstructed. · Subsampling by a factor of Nv in the vertical dimension, · Subsampling by a factor of Nh in the horizontal dimension, · Subsampling by a factor of Ne in each dimension, · Modification of the pixel values contained within the patch ·Rotation of the patch.

[0120] The modification of the pixel values is similar to encoding and decoding. Note that the transmitted mapping parameters can be the parameters of the encoder mapping (in which case the decoder must apply the inverse function of that mapping) or the parameters of the decoder mapping (in which case the encoder must apply the inverse function of that mapping).

[0121] According to a particular embodiment of the present invention, it is possible to apply several transformations to the patch in the encoder. These transformations are signaled either within the information encoded for that patch in the stream or, otherwise, are estimated from the characteristics of the decoded patch. For example, the encoder may be subsampled by a factor of two in each dimension of the patch, followed by mapping of the pixel values of the patch and then rotation may follow.

[0122] According to this particular embodiment of the present invention, the order of the transformations to be applied is predefined and known to the encoder and the decoder. For example, in the encoder the order is as follows: rotation, then subsampling, then mapping.

[0123] When reconstructing the patch in the decoder, when several transformations have to be applied to the patch, the reverse order is applied to the patch (mapping, oversampling, then rotation). Thus, both the decoder and the encoder know in which order to apply the transformations in order to produce the same result.

[0124] At the end of step E52, a set of reconstructed patches becomes available.

[0125] During step E53, at least one intermediate view is synthesized using at least one base view and at least one previously reconstructed patch. The selected virtual view synthesis algorithm is applied to the decoded multi-view video sent to the decoder and the reconstructed data. As previously explained, this algorithm generates a view from the perspective between cameras using the pixels of the base view component and the patch view component.

[0126] For example, the synthesis algorithm uses at least two textures and two depth maps from the base view and / or additional views to generate an intermediate view. Synthesizers are known, and synthesizers belong to, for example, the DIBR category (Depth Image Based Rendering). For example, the algorithms frequently used by standard organizations are as follows. - In VSRS, which represents the View Synthesis Reference Software initiated by Nagoya University and enhanced by MPEG, the forward projection of the depth map is applied using the homography between the reference view and the intermediate view, followed by a filling step to remove the artifacts of forward warping. - In RVS, which represents the Reference View Synthesizer initiated by the University of Brussels and improved by Philips, it starts by projecting the reference view using the calculated disparity. The reference is partitioned into triangles and distorted. Then, the deformed views of each reference are blended, and basic inpainting is applied to fill the disocclusions. - In the VVS, which represents the Versatile View Synthesizer developed by Orange, the references are sorted, a transformation of certain depth map information is applied, and then these depths are conditionally merged. Next, inverse warping of the texture is applied, followed by the merging of various textures and depths. Finally, spatio-temporal restoration is applied and then spatial filtering of the intermediate image is performed.

[0127] FIG. 6 shows an example of a data stream according to a particular embodiment of the present invention, and in particular shows atlas information used to identify one or more transformations to be encoded in the stream and applied to patches of the atlas. For example, this data stream is generated by an encoding method according to any one of the particular embodiments described with respect to FIG. 4 and is suitable for being decoded by a decoding method according to any one of the particular embodiments described with respect to FIG. 5.

[0128] According to this particular embodiment of the present invention, such a stream includes, among other things, the following. - An Act indicator encoded in the header of the atlas to indicate whether a given transformation is activated Trf indicator, - A predicted value Ppred that serves as a predicted value for the transformation parameter value, - The number Np of encoded patches in the atlas, - For each patch of the atlas, patch information, in particular a Trf indicator indicating whether the transformation is used for the patch, - When the Trf indicator indicates the use of a transformation for the patch, the transformation parameter Par, which takes the form of a residual obtained therefor, for example when the predicted value Ppred is encoded.

[0129] As described with respect to the encoding method and the decoding method described above, further particular embodiments of the present invention are possible in terms of the transformation-related information encoded for the patches.

[0130] Figure 7 shows a simplified structure of an encoding device COD adapted to implement an encoding method according to any one of the specific embodiments of the present invention.

[0131] According to a specific embodiment of the present invention, the steps of the encoding method are implemented by computer program instructions. For this purpose, the encoding device COD has a standard architecture of a computer, and in particular includes a memory MEM and, for example, a processor PROC, and a processing device UT driven by a computer program PG stored in the memory MEM. The computer program PG includes instructions for implementing the steps of the encoding method described above when the program is executed by the processor PROC.

[0132] At initialization, the code instructions of the computer program PG are loaded, for example, into a RAM memory (not shown) and then executed by the processor PROC. Specifically, the processor PROC of the processing device UT implements the steps of the encoding method described above according to the instructions of the computer program PG.

[0133] Figure 8 shows a simplified structure of a decoding device DEC adapted to implement a decoding method according to any one of the specific embodiments of the present invention.

[0134] According to a specific embodiment of the present invention, the decoding device DEC has a standard architecture of a computer, and in particular includes a memory MEM0 and, for example, a processor PROC0, and a processing device UT0 driven by a computer program PG0 stored in the memory MEM0. The computer program PG0 includes instructions for implementing the steps of the decoding method described above when the program is executed by the processor PROC0.

[0135] Upon initialization, the code instructions of the computer program PG0 are loaded into, for example, a RAM memory (not shown) and then executed by the processor PROC0. Specifically, the processor PROC0 of the processing device UT0 implements the steps of the above-described decoding method according to the instructions of the computer program PG0.

Description of Signs

[0136] A0 Atlas A1 Atlas COD Encoding Device C1 Camera C2 Camera C3 Camera C4 Camera C5 Camera DEC Decoding Device D b Basic View D s Remaining Views D0 Depth Map, Depth D1 Depth Map, Depth MEM Memory MEM0 Memory PG Computer Program PG0 Computer Program PROC Processor PROC0 Processor T b Basic View T s Remaining Views T0 Texture Image, Texture T1 Texture Image, Texture UT Processing Device UT0 Processing Device V0 View V1 View V2 View

Claims

1. A method for decoding a coded data stream representing multi-view video, wherein the coded data stream includes coded data representing at least one atlas, the at least one atlas corresponds to an image including at least one patch, the at least one patch corresponds to a set of pixels extracted from at least one component of a view of the multi-view video, the view is not coded within the coded data, and the method for decoding includes - decoding the at least one atlas, including decoding the at least one patch, from the coded data stream; - determining, for the at least one decoded patch, whether a transformation must be applied to the at least one decoded patch and which transformation must be applied, the transformation belonging to a group of transformations including at least one oversampling of the patch or a correction of pixel values of the patch; - applying the determined transformation to the decoded patch A method comprising the above steps.

2. The method according to claim 1, wherein whether a transformation must be applied to the at least one decoded patch is determined from at least one syntax element decoded from the coded data stream for the at least one patch.

3. The method according to claim 2, wherein the at least one decoded syntax element includes at least one indicator indicating whether a transformation must be applied to the at least one patch, and when the indicator indicates that a transformation must be applied to the at least one patch, the at least one syntax element optionally includes at least one parameter of the transformation.

4. The method according to claim 3, wherein the at least one parameter of the transformation to be applied to the patch has a value that is prediction-coded with respect to a predicted value.

5. The method according to claim 4, wherein the predicted value is coded within a view header, or within a header of a component of the atlas, or within a header of the atlas.

6. The predicted value is - Patches processed previously according to the processing order of the patches of the atlas, - Patches processed previously, extracted from the same component as the component to which the at least one patch belongs, of the views of the multi-view video, - Patches selected from a set of candidate patches using an index encoded in the data stream, - Patches selected from a set of candidate patches using selection criteria The method according to claim 4, corresponding to the values of the parameters of the transformation, applied to patches belonging to a group comprising **Claim 7** The method according to claim 1, wherein the determination as to whether a transformation must be applied to the at least one decoded patch is carried out when a syntax element decoded from the header of the data stream indicates the activation of the application of the transformation to the patch encoded in the data stream, the syntax element being encoded in the header of a view, or in the header of a component of a view, or in the header of the atlas. **Claim 8** The method according to claim 1, wherein it is determined that a transformation must be applied to the decoded patch if the characteristics of the at least one decoded patch meet a criterion. **Claim 9** The method according to claim 8, wherein the characteristic corresponds to a ratio R = H / W, H corresponding to the height of the at least one decoded patch, W corresponding to the width of the at least one decoded patch, and when the ratio is included within a determined interval, the transformation to be applied to the at least one patch corresponds to vertical oversampling at a predetermined magnification. **Claim 10** The method according to claim 8, wherein the characteristic corresponds to an energy E calculated from the pixel values of the at least one decoded patch, and when the energy E is lower than a threshold value, the transformation to be applied to the at least one patch corresponds to the multiplication of the values of the pixels by a determined magnification. **Claim 11** A method for encoding a data stream representing a multi-view video, the method for encoding comprising: - Extracting at least one patch corresponding to a set of pixels of the component from at least one component of a view of the multi-view video that is not encoded in the data stream; - For each of the at least one extracted patch, determining whether the at least one patch must be subjected to a transformation and which transformation must be applied, the transformation belonging to a group of transformations including at least one subsampling of the patch or a correction of the pixel values of the patch; - Applying the determined transformation to the at least one patch; - Encoding at least one atlas in the data stream, the at least one atlas corresponding to an image including at least the at least one patch; A method, comprising.

12. The method according to claim 1 or 11, wherein when several transformations must be applied to the same patch, the order in which the transformations must be applied is predefined.

13. A device for decoding a coded data stream representing a multi-view video, the coded data stream including coded data representing at least one atlas, the at least one atlas corresponding to an image including at least one patch, the at least one patch corresponding to a set of pixels extracted from at least one component of a view of the multi-view video, the view not being coded in the coded data stream, the device for decoding comprising - Decoding the at least one atlas from the coded data stream, including decoding the at least one patch; - For each of the at least one decoded patch, determining whether the at least one decoded patch must be subjected to a transformation and which transformation must be applied, the transformation belonging to a group including at least one oversampling of the patch or a correction of the pixel values of the patch; - Applying the determined transformation to the decoded patch A device comprising a processor and a memory configured to perform the above.

14. A device for encoding a data stream representing a multi-view video, - Extracting at least one patch corresponding to a set of pixels of the component from at least one component of the views of the multi-view video that is not encoded in the data stream; - Determining, for the at least one extracted patch, whether at least one transformation must be applied to the at least one patch and which transformation must be applied, wherein the transformation belongs to a group of transformations including at least one subsampling of the patch or a correction of the pixel values of the patch; - Applying the determined transformation to the at least one patch; - Encoding at least one atlas in the data stream, wherein the at least one atlas corresponds to an image including at least the at least one patch; A device comprising a processor and a memory configured to perform the above.

15. A computer program comprising instructions for implementing the decoding method according to any one of claims 1 to 10 and 12 and / or instructions for implementing the encoding method according to claim 11 or 12 when the computer program is executed by a processor.

Citation Information

Patent Citations

  • Efficient patch rotation in point cloud coding

    JP2022517118A

  • Immersive Video Coding Technology for 3DoF+ / MIV and V-PCC

    JP2022532302A