Encoding and Decoding of Omnidirectional Images

The method addresses inefficiencies in encoding 360° videos by offering two encoding techniques, one for conventional high-quality reconstruction and another for processed data, reducing transmission costs and enhancing intermediate view synthesis.

JP7693065B2Active Publication Date: 2025-06-16オランジュ
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024108903
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-10-01
Filing Date
2024-07-05
Publication Date
2025-06-16
Estimated Expiration
2039-09-25

AI Technical Summary

Technical Problem

Existing video coders for 360° and omnidirectional videos are inefficient in compression and lack inter-image prediction, leading to high data encoding and transmission costs.

Method used

A method that selects between two encoding techniques for each view image: a conventional encoding method for high-quality reconstruction and an innovative encoding method that processes and encodes image data, reducing transmission costs by omitting unnecessary data.

Benefits of technology

Significantly reduces signal transmission costs by processing and encoding only necessary data, while maintaining the quality of reconstructed view images for efficient synthesis of intermediate views.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693065000001
    Figure 0007693065000001
  • Figure 0007693065000002
    Figure 0007693065000002
  • Figure 0007693065000003
    Figure 0007693065000003
Patent Text Reader

Abstract

To provide a method for coding views forming part of a plurality of views.SOLUTION: A process for coding an image of a view from among a plurality of views, includes the following steps of: selecting a first or a second coding method to code image data from the image (C1); generating a data signal containing information (flag_proc) indicating whether the selected method is the first or the second coding method (C10, C12a; C10, C12b; C10, C12c), and, if it is the first coding method, coding original image data so as to provide coded original data (C11a), and, if it is the second coding method, coding processed image data from the image obtained by image processing of the original image data so as to provide coded processed data (C11b); and coding information describing the image processing which has been applied (C11b).SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of omnidirectional video, and more specifically to videos such as 360° and 180°. More specifically, the present invention relates to the encoding and decoding of views such as 360° and 180° captured for generating such videos, and the synthesis of non-captured intermediate viewpoints.

[0002] The present invention can be applied, in particular (but not exclusively), to video encoding implemented in current AVC and HEVC video coders and their extensions (such as MVC, 3D-AVC, MV-HEVC, 3D-HEVC, etc.) and corresponding video decoding.

Background Art

[0003] In order to generate omnidirectional video (such as 360° video, etc.), it is common to use a 360° camera. Such a 360° camera is formed by a number of 2D (two-dimensional) cameras installed on a spherical platform. Each 2D camera captures a specific angle of a 3D (three-dimensional) scene, and a set of views captured by the cameras enables the generation of a video representing a 3D scene with a 360°×180° field of view. It is also possible to capture a 3D scene with a 360°×180° field of view using a single 360° camera. Naturally, such a field of view can be smaller (for example, 270°×135°).

[0004] Then, with such 360° video, the user can view the scene as if the user were located at the center, or look around the entire surroundings of the user (over 360°), thereby providing a new way of viewing videos. Such videos are generally played back in a virtual reality headset, also known as an HMD, which corresponds to a "head-mounted device". However, those videos can also be displayed on a 2D screen equipped with appropriate user interaction means. The number of 2D cameras for capturing a 360° scene varies depending on the platform used.

[0005] To generate a 360° image, the various views captured by various 2D cameras are joined end-to-end and arranged, taking into account the overlap between the views, to create a panoramic 2D image. This step is also known as "stitching". For example, equirectangular projection (ERP) is one possible projection for obtaining such a panoramic image. According to this projection, the views captured by each of the 2D cameras are projected onto a spherical surface. Other types of projections are also possible, such as cube mapping type projections (projections onto the faces of a cube). The views projected onto the surface are then projected onto a 2D plane to obtain a 2D panoramic image, which includes all the views of the scene captured at a given time.

[0006] To increase the sense of immersion, a number of 360° cameras of the above type can be used simultaneously to capture a scene, and these cameras can be positioned in the scene in any way. The 360° cameras can be either real cameras (i.e., physical objects) or virtual cameras. In the case of virtual cameras, the views are obtained by view generation software. Specifically, such virtual cameras enable the generation of views representing viewpoints of a 3D scene that were not captured by real cameras.

[0007] Then, the 360° view images obtained using a single 360° camera or the 360° view images obtained using a number of 360° cameras (real cameras and virtual cameras) are, for example, - conventional 2D video coders (e.g., coders compliant with the HEVC (abbreviation for "High Efficiency Video Coding") standard) - conventional 3D video coders (e.g., coders compliant with the MV-HEVC and 3D-HEVC standards) used for encoding. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION

[0008] Such coders, needless to say, encode a large number of 360° view images and the specific geometry of the 360° representation of 3D scenes using such 360° views. Considering that a very large amount of data images are encoded in one 360° view, they are not sufficiently efficient from the perspective of compression. Moreover, since the views captured by the 2D cameras of 360° cameras are diverse, the aforementioned coders are not suitable for encoding different images of 360° views because almost no or no inter-image prediction is used by these coders. In particular, there is little similar content that can be predicted between the two views captured by two 2D cameras respectively. Therefore, all the images of 360° views are compressed in the same way. Specifically, in these coders, for the image of the current 360° view to be encoded, an analysis for determining whether it makes sense to encode all the data of this image or only a part of the data of this image is not performed as part of the synthesis of the non-captured intermediate view image that will use this image of the view to be decoded next after encoding.

[0009] One object of the present invention is to correct the drawbacks of the aforementioned prior art.

Means for Solving the Problems

[0010] For that purpose, one aspect of the present invention is a method for encoding images of view-forming portions of a number of views, implemented by an encoding device, wherein the number of views simultaneously represents a 3D scene from different viewing angles or positions, - Selecting a first encoding method or a second encoding method for encoding the image of the view, - Generating a data signal including information indicating whether the selected one is the first encoding method or the second encoding method, - When the first encoding method is selected, encoding the original data of the image of the view, wherein the first encoding method provides the encoded original data, - When the second encoding method is selected, - Encoding the processed data of the view image, where these data are obtained by image processing applied to the original data of the view image, and the encoding provides processed and encoded data; - Encoding the description information of the applied image processing; The method includes: - The generated data signal is - When the first encoding method is selected, the encoded original data of the view image, - When the second encoding method is selected, the processed and encoded data of the view image and the encoded description information of the image processing. It further relates to the method.

[0011] According to the present invention, for each image of each view scheduled for encoding, from among a large number of images from the current view scheduled for encoding of the aforementioned type (the above images represent a very large amount of data scheduled for encoding and thus for signal transmission), two encoding techniques are provided, namely: - The first encoding technique, by which the images of one or more views are encoded in a conventional manner (e.g., HEVC, MVC-HEVC, 3D-HEVC), and reconstructed images forming views of very good quality are obtained respectively; - The second innovative encoding technique, by which the processed data of the images of one or more other views are encoded, and the processed image data is obtained on the decoding side. Thus, the processed image data does not match the original data of these images, but the benefit is obtained that the signal transmission cost of the processed and encoded data of these images is significantly reduced. It is possible to combine them.

[0012] As a result, for each image of each of the other views, what is obtained on the decoding side (processed data encoded according to the second encoding method) is the corresponding processed data of the view image and the explanatory information of the image processing applied to the original data of the view image on the encoding side. Next, such processed data can be processed using the corresponding image processing explanatory information to form the view image, and with that view, at least one of the view images reconstructed according to the conventional first decoding method, it is possible to synthesize the image of the non-captured intermediate view in a particularly efficient and effective manner.

[0013] The present invention is also a method for decoding a data signal representing images of view forming portions of a number of views, which are implemented by a decoding device, wherein the number of views simultaneously represent a 3D scene from different viewing angles or positions. - Reading an item of information indicating whether to decode the view image according to a first decoding method or according to a second decoding method based on the data signal; - In the case of the first decoding method: - Reading the encoded data associated with the view image in the data signal; - Reconstructing the view image based on the read encoded data, wherein the reconstructed view image includes the original data of the view image; - In the case of the second decoding method: - Reading the encoded data associated with the view image in the data signal; - Reconstructing the view image based on the read encoded data, wherein the reconstructed view image of the view includes the processed data of the view image associated with the explanatory information of the image processing used to obtain the processed data; and relates to a method including the above.

[0014] According to a specific embodiment, - The processed data of the view image is the data of the view image that was not deleted after the application of the cropping of the view image. - The description information of the image processing is information about the location of one or more cropped areas in the view image.

[0015] By such cropping processing applied to the view image, encoding of a part of the original data can be avoided, and since the data belonging to one or more cropped areas is neither encoded nor transmitted to the decoder, there is an advantage that the transmission rate of the encoded data associated with the view image is significantly reduced. The reduction in rate depends on the size of one or more areas to be cropped. Therefore, the view image that is reconstructed after decoding and then potentially processed with the processed data using the corresponding image processing description information does not include all of its original data, or at least is different from the original image of the view. However, obtaining such an image of the view cropped in this way does not harm the effect of synthesizing the intermediate image that will use such an image of the cropped view at the time of reconstruction. In fact, such synthesis using one or more images reconstructed using a conventional decoder (e.g., HEVC, MVC-HEVC, 3D-HEVC), the view image, and the images reconstructed in the conventional manner can recover the original area of the intermediate view.

[0016] According to another specific embodiment, - The processed data of the view image is the data of at least one area of the view image sampled according to a predetermined sampling factor and in at least one predetermined direction. - The description information of the image processing includes at least one item of information about the location of at least one sampled area in the view image.

[0017] Such processing also aids in the uniform degradation of the image of the view for the purpose of optimizing the reduction in the rate of data resulting from the subsequent encoding following the application of sampling in this case. Even if the subsequent reconstruction of such an image of the view sampled in this way provides a reconstructed image of a view that is degraded from the original image of the view / different from the original image (the original data having been encoded following sampling), the effect of the synthesis of the intermediate image that would use such an image of the sampled and reconstructed view is not impaired. In fact, such a synthesis using one or more images reconstructed using a conventional decoder (e.g., HEVC, MVC-HEVC, 3D-HEVC), in these one or more images reconstructed in a conventional manner, the original region corresponding to the filtered region of the image of the view can be recovered.

[0018] According to another specific embodiment, - The processed data of the image of the view is the data of at least one region of the image of the view where filtering has been performed. - The explanatory information of the image processing includes at least one item of information about the location of at least one filtered region in the image of the view.

[0019] Such processing aids in the deletion of the data of the image of the view that is considered unnecessary for encoding for the purpose of optimizing the reduction in the rate of the encoded data, and the encoded data is advantageously formed only of the filtered data of the image.

[0020] Subsequent reconstruction of such an image of the filtered view in this way, even if it provides a reconstructed image of a view that has degraded from the original image of the view / differs from the original image, (the original data having been encoded following filtering), does not harm the effect of the synthesis of the intermediate image that will use such an image of the filtered and reconstructed view. Indeed, such a synthesis using one or more images reconstructed using a conventional decoder (e.g., HEVC, MVC-HEVC, 3D-HEVC), the filtered regions of the image of the view and the images reconstructed in a conventional manner, can recover the original regions of the intermediate view.

[0021] According to another specific embodiment, - The processed data of the image of the view is the pixels of the image of the view corresponding to occlusions detected using the image of another view among a number of views. - The explanatory information of the image processing includes indicators of the pixels of the image of the view seen in the image of another view.

[0022] Similar to the previous embodiments, such processing helps to delete the data of the image of the view that is considered unnecessary for encoding, for the purpose of optimizing the reduction of the rate of the encoded data, which is advantageously formed only by the pixels of the image of the view, the lack of which has been detected in another image of the current view among the said number of views.

[0023] Even if the subsequent reconstruction of such an image of the view provides a reconstructed image of the view that has degraded from the original image of the view / differs from the original image (only the occlusion regions are encoded), the effect of the synthesis of the intermediate image to be used with such a reconstructed image of the view is not impaired. In fact, such synthesis using one or more images reconstructed using a conventional decoder (e.g., HEVC, MVC-HEVC, 3D-HEVC), the image of the current view, and the image reconstructed in the conventional manner can recover the original regions of the intermediate view.

[0024] According to another specific embodiment, - The processed data of the encoded / decoded view image is - Based on the original data of the view image, - Based on the original data of at least one other view image encoded / decoded using a first encoding / decoding method, - Potentially, based on the original data of at least one other view image encoded / decoded using a second encoding / decoding method where the processed data is Pixels that are calculated. - The description information of the above image processing is - Indicators of the pixels of the calculated view image, - Information about the location of the original data used to calculate the pixels of the view image in at least one other view image encoded / decoded using a first encoding / decoding method, - Potentially, information about the location of the original data used to calculate the pixels of the view image in at least one other view image where the processed data is encoded / decoded is included.

[0025] According to another specific embodiment, the processed data of the image of the first view and the processed data of the image of at least one second view are combined into a single image.

[0026] According to the above embodiment, the processed data of the view image obtained according to the second decoding method includes the processed data of the first view image and the processed data of at least one second view image.

[0027] According to a specific embodiment, - The processed and encoded / decoded data of the view image is data of the image type. - The encoded / decoded description information of the image processing is data of the image type and / or text type.

[0028] Furthermore, the present invention is an encoding device for encoding images of view forming parts of a number of views, where the number of views simultaneously represents a 3D scene from different viewing angles or positions. At the current time, - Selecting a first encoding method or a second encoding method for encoding the view image; - Generating a data signal including information indicating whether the selected one is the first encoding method or the second encoding method; - When the first encoding method is selected, encoding the original data of the view image, where the first encoding method provides the encoded original data; - When the second encoding method is selected, - Encoding the processed data of the view image, where the processed data is obtained by image processing applied to the original data of the view image, and the encoding provides the processed and encoded data; - Encoding the description information of the applied image processing; An encoding device including a processor configured to perform the above; - The generated data signal is - When the first encoding method is selected, the encoded original data of the view image; - When the second encoding method is selected, it also relates to an encoding device that further includes the processed and encoded data of the view image and the encoded description information of the image processing. Specifically, such an encoding device can implement the aforementioned encoding method.

[0029] Specifically, such an encoding device can implement the aforementioned encoding method.

[0030] Furthermore, the present invention relates to a device for decoding a data signal representing an image of a view forming portion of a plurality of views, where the plurality of views simultaneously represent a 3D scene from different viewing angles or positions. At the current time, - Read an item of information in the data signal indicating whether to decode the view image according to the first decoding method or the second decoding method. - In the case of the first decoding method, - Read the encoded data associated with the view image in the data signal. - Reconstruct the view image based on the read encoded data, where the reconstructed view image includes the original data of the view image. - In the case of the second decoding method, - Read the encoded data associated with the view image in the data signal. - Reconstruct the view image based on the read encoded data, where the reconstructed view image of the view includes the processed data of the view image associated with the description information of the image processing used to obtain the processed data. It also relates to a decoding device including a processor configured to perform the above.

[0031] Specifically, such a decoding device can implement the aforementioned decoding method.

[0032] Furthermore, the present invention also relates to a data signal including data encoded according to the aforementioned encoding method.

[0033] Furthermore, the present invention also relates to the above program including instructions for implementing the decoding method or encoding method according to the present invention according to any one of the specific embodiments described above when the computer program is executed by a processor.

[0034] This program can use any programming language and can be in the form of source code, object code, or intermediate code between source code and object code, such as a partially compiled form or any other desired form.

[0035] Furthermore, the present invention also targets a computer-readable recording medium or information medium including computer program instructions such as those mentioned above.

[0036] The recording medium can be any entity or device capable of storing a program. For example, the medium can include storage means such as a ROM (e.g., CD-ROM or ultra-small electronic circuit ROM) or magnetic recording means (e.g., USB key or hard disk).

[0037] Moreover, the recording medium can also be a transmissible medium such as an electrical or optical signal that can be transmitted via an electrical or optical cable by wireless or other means. The program according to the present invention can specifically be downloaded from an Internet-type network.

[0038] Alternatively, the recording medium can be an integrated circuit in which the program is incorporated, and the circuit is adapted to execute or be used when executing the aforementioned encoding or decoding method.

[0039] Other features and advantages will become more clearly apparent by reading through some preferred embodiments provided as purely illustrative and non - limiting examples, which are described below with reference to the accompanying drawings.

Brief Description of the Drawings

[0040]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 3A

Figure 3B

Figure 4A

Figure 4B

Figure 4C

Figure 4D

Figure 4E

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12A

Figure 12B

Figure 13

Figure 14

Figure 15A

Figure 15B

Figure 15C

Figure 15D

Figure 16

Figure 17

Figure 18

[0041] The present invention mainly proposes a scheme for encoding a number of current images of a number of views, where the number of views represents a 3D scene at a current time by a predetermined position or a predetermined viewing angle, and two encoding techniques, namely, - A first encoding technique, by which at least one current image of a view is encoded using a conventional encoding mode (e.g., HEVC, MV-HEVC, 3D-HEVC, etc.) - A second innovative encoding technique, by which the processed data of at least one current image of a view (obtained from the application of specific image processing to the original data of this image) is encoded using the aforementioned type of conventional encoding mode and / or any other suitable encoding mode, thereby significantly reducing the signaling cost of the encoded data of this image obtained from the processing performed before the encoding step are available.

[0042] Accordingly, the present invention provides two decoding techniques, namely, - A first decoding technique, by which at least one current image of the encoded view is reconstructed using a conventional decoding mode (e.g., HEVC, MV-HEVC, 3D-HEVC, etc.) corresponding to the conventional encoding mode used during encoding and signaled to the decoder, thereby obtaining at least one reconstructed image of the view with very good quality - A second innovative decoding technique (whereby, using a decoding mode corresponding to the encoding mode signaled to the decoder (i.e., the conventional encoding mode and / or other suitable encoding mode), the encoded processed data of at least one image of the view is decoded, whereby processed image data and explanatory information of the image processing that is the source of the obtained processed data are obtained. Thus, the processed data obtained on the decoding side for this image is different from the image data decoded according to the first decoding technique and does not match its original data) Propose a decoding scheme that allows them to be combined.

[0043] The image of the view subsequently reconstructed based on such processed and decoded image data and image processing explanatory information will be different from the original image of the view (i.e., the one before being encoded following the processing of its original data). However, such a reconstructed image of the view, together with the images of other views reconstructed according to the first conventional decoding technique, constitutes an image of the view that enables the synthesis of intermediate view images in a particularly efficient and effective way.

[0044] 6. Implementation of an exemplary encoding scheme Hereinafter, for example, a method for encoding 360°, 180° or other omnidirectional video that can use any type of multi-view video coder conforming to the 3D-HEVC or MV-HEVC standard or the like will be described.

[0045] Referring to FIG. 1, such an encoding method is applied to the current image of the view that forms part of a number of views V1,..., V N wherein the number of views each represent a 3D scene by a number of viewing angles or a number of positions / orientations.

[0046] According to a common example, in the case where three omnidirectional cameras are used to generate a video (e.g., a 360° video), - The first omnidirectional camera can be placed at the center of the 3D scene, for example, with a 360° × 180° field of view, - The second omnidirectional camera can be placed on the left side of the 3D scene, for example, with a 360° × 180° field of view, - The third omnidirectional camera can be placed on the right side of the 3D scene, for example, with a 360° × 180° field of view.

[0047] According to another, more unconventional example, in the case where three omnidirectional cameras are used to generate an α° video (0° < α ≤ 360°), - The first omnidirectional camera can be placed at the center of the 3D scene, for example, with a 360° × 180° field of view, - The second omnidirectional camera can be placed on the left side of the 3D scene, for example, with a 270° × 135° field of view, - The third omnidirectional camera can be placed on the right side of the 3D scene, for example, with a 180° × 90° field of view.

[0048] Of course, other configurations are also possible.

[0049] At least two of the above-mentioned multiple views may or may not represent the 3D scene from the same field of view.

[0050] The encoding method according to the present invention, at the current time, - The image IV1 of view V1, - The image IV2 of view V2, -…, - The image IV k of view V k , -…, - The image IV N of view V N is to be encoded.

[0051] The image of the target view can similarly be a texture image or a depth image. The image of the target view (e.g., image IV k) includes original data (d1 k …, dQ k ) of quantity Q (Q≧1), such as Q pixels for example.

[0052] Next, the encoding method includes the following for at least one image IV k of the view V k to be encoded.

[0053] In C1, for the image IV k , the first encoding method MC1 or the second encoding method MC2 is selected.

[0054] When the first encoding method MC1 is selected, in C10, in order to indicate that the encoding method MC1 is selected, for example, a bit is set to 0 and information called flag_proc is encoded.

[0055] In C11a, for example, using a conventional coder such as one conforming to standards such as HEVC, MV-HEVC, 3D-HEVC, etc., the Q original data (pixels) d1 k of the image IV k , …, dQ k are encoded. When the encoding C11a is completed, the encoded image IVC k of the view V k is obtained. Next, the encoded image IVC k includes Q encoded original data dc1 k , dc2 k , …, dcQ k .

[0056] In C12a, a data signal F1 k is generated. As shown in FIG. 2A, the data signal F1 k includes information flag_proc = 0 related to the selection of the first encoding method MC1 and the encoded original data dc1 k , dc2 k , …, dcQ k .

[0057] When the second encoding method MC2 is selected, in C10, in order to indicate that the encoding method MC2 is selected, for example, a bit is set to 1 and information called flag_proc is encoded.

[0058] In C11b, data DT k resulting from the processing of the image IV k performed before the encoding step is applied with the encoding method MC2.

[0059] Such data DT k is - data of an image type (pixels) corresponding to all or part of the original data of the image IV k processed using a specific image process before the encoding step (various detailed examples thereof will be further described in the description), - description information of the image process applied to the image IV k before the encoding step C11b (such description information is, for example, text and / or an image type) and includes.

[0060] When the encoding C11b is completed, processed and encoded data DTC k is obtained. Those processed and encoded data DTC k represent processed and encoded image IVTC k and.

[0061] Therefore, the processed data DT k does not match the original data of the image IV k and.

[0062] For example, these processed data DT k correspond to an image with a resolution higher or lower than that of the image IV k before processing. Therefore, the processed image IV k is, for example, larger because it is obtained based on an image of another view, or conversely, the image IV kIt can be smaller because it is obtained by deleting one or more original pixels from

[0063] According to another example, these processed data DT k are the representation (such as YUV, RGB, etc.) format of the processed image IV k which is different from the original format of the image IV before processing k and corresponds to an image with a different number of bits (such as 16 bits, 10 bits, 8 bits, etc.) for representing pixels

[0064] According to yet another example, these processed data DT k correspond to a color or texture component that is degraded with respect to the original texture or color component of the image IV before processing k

[0065] According to still another example, these processed data DT k correspond to a specific representation of the original content of the image IV before processing (for example, the representation of the filtered original content of the image IV k k ).

[0066] In the case where the processed data DT k is just image data (i.e., in the form of a grid of pixels, for example), the encoding method MC2 can be implemented by a coder similar to the coder that implements the first encoding method MC1. It can be a lossy or lossless coder. In the case where the processed data DT k is different from image data (such as text type data, etc.) or includes both image data and data of a type other than image data, the encoding method MC2 is - In particular, for encoding text type data, by a lossless coder - In particular, for encoding image data, by a lossy or lossless coder (such a coder may or may not be the same as the coder that implements the first encoding method MC1)​​ It can be implemented.

[0067] In C12b, the data signal F2 k is generated. As shown in FIG. 2B, the data signal F2 k is, in the case where all of these data are image data, the information flag_proc = 1 related to the selection of the second encoding method MC2, and the processed and encoded data DTC k and.

[0068] Alternatively, in C12c, in the case where the processed and encoded data DTC k includes both image data and text-type data, two signals F3 k , F'3 k are generated.

[0069] As shown in FIG. 2C, - The data signal F'3 includes the information flag_proc = 1 related to the selection of the second encoding method MC2, and the processed and encoded data DTC k of the image type, - The data signal F'3 k includes the processed and encoded data DTC k of the text type.

[0070] The encoding method just described above can then be implemented for some, or only for some, of each of the N views of the images IV1, IV2,..., IV N scheduled for encoding, or limited to the image IV k (e.g., k = 1).

[0071] According to two exemplary embodiments shown in FIGS. 3A and 3B, for example, among the N images IV1,..., IV N scheduled for encoding, the following is assumed. - The n first images IV1,..., IV n are encoded using the first encoding technique MC1. The n first views V1~Vn is called the master view because, when reconstructed, the images of the n master views contain all of their original data and are suitable for use with one or more of the N - n other views to synthesize the images of any view required by the user. - N - n other images IV n+1 , …, IV N are processed before being encoded using the second encoding method MC2. These processed N - n other images belong to what is called additional views.

[0072] If n = 0, all the images IV1, …, IV of all views N are processed. If n = N - 1, the image of a single view out of N (e.g., the image of the first view) is processed.

[0073] N - n other images IV n+1 , …, IV N When the processing of the N - n other images IV, …, IV is completed, M - n processed data are obtained. If M = N, there is processed data for the number of views to be processed. If M < N, at least one of the N - n views is deleted during processing. In this case, when the encoding C11b by the second encoding method MC2 is completed, the processed data related to the image of the deleted view are encoded using a lossless coder and contain only the image and the information indicating the absence of this view. Then, for this image deleted by the processing, no pixel is encoded using the second encoding method MC2 using a coder of types such as HEVC, MVC - HEVC, 3D - HEVC.

[0074] In the example of Figure 3A, in C11a, using a conventional coder of the HEVC type, the n first images IV1, …, IV n are encoded independently of each other. When the encoding C11a is completed, n encoded images IVC1, …, IVC n are obtained respectively. In C12a, n data signals F11, …, F1n are each generated. In C13a, these n data signals F11, …, F1 n are concatenated to generate a data signal F1.

[0075] As represented by the dashed arrows, images IV1, …, IV n are the N - n other images IV n+1 , …, IV N that can be used during the processing of.

[0076] Still referring to FIG. 3A, in C11b, using a conventional HEVC - type coder, M - n processed data DT n+1 , …, DT M are encoded independently of each other (if all of these M - n processed data are of image type), or, for processed data of image type, they are encoded independently of each other using a conventional HEVC - type coder, and for processed data of text type, they are encoded by a lossless coder (if those data are of both image type and text type). When the encoding C11b is completed, N - n processed and encoded data DTC n+1 , …, DTC N are each obtained as the N - n encoded data associated with each of the N - n images IV n+1 , …, IV N In C12b, in the case where all of the M - n processed data are of image type, N - n data signals F2 n+1 , …, DTC N each containing the N - n processed and encoded data DTC n+1 , …, F2 N are each generated. In C12c, in the case where the M - n processed data are of both image type and text type, - N - n data signals F3 each containing the N - n processed and encoded data of image type n+1 , …, F3 N are generated, - N - n data signals F’3 each containing text - type, processed, and encoded data of N - n n+1 …, F’3 N are generated.

[0077] In C13c, data signals F3 n+1 , …F3 N and F’3 n+1 , …, F’3 N are concatenated to generate data signal F3.

[0078] In C14, signals F1 and F2 are concatenated, or signals F1 and F3 are concatenated to provide a data signal F that can be decoded by a decoding method. The decoding method will be further described in the description.

[0079] In the example of FIG. 3B, in C11a, using a conventional coder of the MV - HEVC or 3D - HEVC type, n first images IV1, …, IV n are encoded simultaneously. When the encoding C11a is completed, n encoded images IVC1, …, IVC n are obtained respectively. In C12a, a single signal F1 is generated, and the single signal F1 contains the encoded original data associated with each of these n encoded images.

[0080] As represented by the dashed - arrow, the first images IV1, …, IV n can be used during the processing of N - n other images IV n+1 , …, IV N .

[0081] Still referring to FIG. 3B, in C11b, using a conventional coder of the MV - HEVC or 3D - HEVC type, M - n processed data DT of the image type n+1 , …, DT MAre they encoded simultaneously (when all of these M - n processed data are of the image type), or for the processed data of the image type, they are encoded simultaneously using a conventional coder of the MV - HEVC or 3D - HEVC type, and for the processed data of the text type, they are encoded by a lossless coder (when those data are of both the image type and the text type). When the encoding C11b is completed, N - n processed and encoded data DTC n+1 , …, DTC N are respectively obtained as N - n processed and encoded data associated with each of the N - n processed images IV n+1 , …, IV N . In C12b, in the case where all of the M - n processed data are of the image type, a single signal F2 including the N - n processed and encoded data DTC n+1 , …, DTC N is generated. In C12c, in the case where the M - n processed data are of both the image type and the text type, - A signal F3 including the processed and encoded data of the image type is generated, - A signal F’3 including the processed and encoded data of the text type is generated.

[0082] In C14, the signals F1 and F2 are concatenated, or the signals F1, F3 and F’3 are concatenated, and a data signal F that can be decoded by a decoding method is provided. The decoding method will be further described in the explanation.

[0083] Of course, other combinations of encoding methods are also possible.

[0084] According to one possible variant of FIG. 3A, the coder implementing the encoding method MC1 is a coder of the HEVC type, and the coder implementing the encoding method MC2 is a coder of the MV - HEVC or 3D - HEVC type or can include a coder of the MV - HEVC or 3D - HEVC type and a lossy coder.

[0085] According to one possible deformation of FIG. 3B, the coder implementing the encoding method MC1 is a coder of the MV-HEVC or 3D-HEVC type, and the coder implementing the encoding method MC2 can be a coder of the HEVC type or can include a coder of the HEVC type and a lossless coder.

[0086] Here, with reference to FIGS. 4A to 4E, a first embodiment of the processing applied to the original data of the image IV before the encoding step C11b (FIG. 1) according to the second encoding method MC2 will be described. k A description of a first embodiment of the processing applied to the original data of the image IV is provided.

[0087] In the examples shown in these figures, the processing applied to the original data of the image IV k is horizontal or vertical or simultaneous cropping in both directions of one or more regions of this image.

[0088] In the example of FIG. 4A, the left boundary B1 and the right boundary B2 of the image IV k are cropped. The cropping means deleting the pixels of the rectangular region of the image IV k formed by each of the boundaries B1 and B2.

[0089] In the example of FIG. 4B, the upper boundary B3 and the lower boundary B4 of the image IV k are cropped. The cropping means deleting the pixels of the rectangular region of the image IV k formed by each of the boundaries B3 and B4.

[0090] In the example of FIG. 4C, the cropping is applied vertically to the rectangular region Z1 located in the image IV k In the example of FIG. 4D, the cropping is applied horizontally to the rectangular region Z2 located in the image IV

[0091] In the example of FIG. 4D, the cropping is applied horizontally to the rectangular region Z2 located in the image IV k In the example of FIG. 4D, the cropping is applied horizontally to the rectangular region Z2 located in the image IV

[0092] In the example of FIG. 4E, cropping is applied in both the horizontal and vertical directions to the region Z3 located in the image IV k Next, the processed data DT

[0093] scheduled for encoding k is - the pixels of the remaining region Z k of the image IV that was not deleted after cropping (FIGS. 4A, 4B, 4E) R or the pixels of the remaining region Z1 k of the image IV that was not deleted after cropping R and Z2 R (FIGS. 4C, 4D), and - information explaining the applied cropping .

[0094] For example, in the cases of FIGS. 4A, 4B, and 4E, the information explaining the applied cropping is of text type and - the coordinates of the pixel located at the upper leftmost corner of the remaining region Z k of the image IV R and - the coordinates of the pixel located at the lower rightmost corner of the remaining region Z k of the image IV R . .

[0095] For example, in the cases of FIGS. 4C and 4D, the information explaining the applied cropping is - the coordinates of the pixel located at the upper leftmost corner of the remaining region Z1 k of the image IV R and - the coordinates of the pixel located at the lower rightmost corner of the remaining region Z1 k of the image IV R and - the coordinates of the pixel located at the upper leftmost corner of the remaining region Z2 k of the image IV R and - the coordinates of the pixel located at the lower rightmost corner of the remaining region Z2 k of the image IV R . .

[0096] Next, in C11b (FIG. 1), the original data (pixels) of the rectangular region defined by the coordinates of the pixel located at the top left of each of the remaining regions Z R (Z1 R and Z2 R ), and the coordinates of the pixel located at the bottom right of the remaining region Z R (Z1 R and Z2 R ) are encoded. The description information of the applied cropping is encoded for that part by a lossless coder in C11b (FIG. 1).

[0097] As a variant, the information describing the applied cropping includes the number of rows and / or columns of pixels to be deleted, and the positions of these rows and / or columns in the image IV k .

[0098] According to one embodiment, the amount of data deleted by cropping is fixed. For example, the amount can be determined to systematically delete X rows and / or Y columns from the image of the view in question. In this case, the description information includes only the information about the cropping (or not for each view).

[0099] According to another embodiment, the amount of data deleted by cropping can be changed between the image IV of the view V k and the image of another available view. k

[0100] Also, the amount of data deleted by cropping can depend, for example, on the position of the 3D scene of the camera that captured the image IV k . Thus, for example, if the image of another view among the above N views is captured by a camera having a position / orientation in a 3D scene different from the position / orientation of the camera that captured the image IV k , then, for example, the image IV k ​A data deletion amount different from the data deletion amount deleted for

[0101] The application of cropping may further depend on the time when Image IV k is encoded. For example, at the current time, while deciding to apply cropping to Image IV k at times before and after the current time, it can be decided not to apply such cropping or any processing related thereto to Image IV k .

[0102] Finally, cropping can be applied to the images of one or more views at the current time. The cropped area of Image IV k of View V k may or may not be the same as the cropped area of the image of another view scheduled for encoding at the current time.

[0103] Here, with reference to FIGS. 5A to 5D, a second embodiment of the processing applied to the original data of Image IV k before the encoding step C11b (FIG. 1) by the second encoding method MC2 will be described.

[0104] In the examples shown in these figures, the processing applied to the original data of Image IV k is the downsampling in the horizontal or vertical direction of one or more regions of this image.

[0105] In the example of FIG. 5A, the downsampling is applied vertically to Region Z4 of Image IV k .

[0106] In the example of FIG. 5B, the downsampling is applied horizontally to Region Z5 of Image IV k .

[0107] In the example of FIG. 5C, the downsampling is applied to the whole of Image IV k .

[0108] In the example of FIG. 5D, downsampling is applied both horizontally and vertically to the region Z6 of the image IV k in the image IV.

[0109] Next, the processed data DT scheduled for encoding k is - the downsampled image data (pixels), and - information explaining the applied downsampling and includes information explaining the applied downsampling, for example - the downsampling factor used, - the downsampling direction used, - in the cases of FIGS. 5A, 5B, and 5D, the locations of the filtered regions Z4, Z5, Z6 of the image IV k in the image IV, or - in the case of FIG. 5C, the coordinates of the pixel located at the top left of the image IV k in the image IV, and the coordinates of the pixel located at the bottom right of the image IV k in the image IV (thus defining the complete area of this image) and so on.

[0110] Next, in C11b (FIG. 1), the downsampled image data (pixels) is encoded by a coder of a type such as HEVC, 3D-HEVC, MV-HEVC, etc. The explanatory information of the applied downsampling is encoded by a lossless coder in C11b (FIG. 1) for that part.

[0111] The value of the downsampling factor can be fixed or can depend on the position of the 3D scene of the camera that captured the image IV k in the image IV. Thus, for example, if an image of another view among the above N views is captured by a camera having a position / orientation in a 3D scene different from the position / orientation of the camera that captured the image IV k in the image IV, for example, a different downsampling factor is used.

[0112] The application of downsampling may further depend on the time at which the image of view V k is encoded. For example, at the current time, it is determined to apply downsampling to image IV k , while at times before and after the current time, it can be determined not to apply such downsampling or any processing in that regard to the image of view V k .

[0113] Finally, downsampling can be applied to the images of one or more views at the current time. The downsampled area of image IV k may or may not be the same as the downsampled area of the image of another view scheduled for encoding at the current time.

[0114] Here, referring to FIG. 6, a third embodiment of the processing applied to the original data of image IV k before the encoding step C11b (FIG. 1) by the second encoding method MC2 is provided.

[0115] In the example shown in FIG. 6, the processing applied to the original data of image IV k is the detection of contours by filtering this image. For example, two contours ED1 and ED2 exist in image IV k . In a method well known per se, such filtering includes, for example, - applying a contour detection filter to contours ED1 and ED2, and - applying an expansion of contours ED1 and ED2 to increase the area surrounding each of contours ED1 and ED2 (such as the area represented by the hatching in FIG. 6), and - deleting all of the original data that does not form a part of the hatched area (and thus is considered unnecessary for encoding) from image IV k and .

[0116] Next, the processed data DT scheduled for encoding k in such filtering cases - the original pixels included in the shaded area, and - information explaining the applied filtering (e.g., values predefined for each pixel not included in the shaded area, e.g., represented by the predefined value YUV = 000) and is included.

[0117] Next, in C11b (FIG. 1), the image data (pixels) corresponding to the shaded area is encoded by a coder of the type such as HEVC, 3D-HEVC, MV-HEVC. The information explaining the applied filtering is encoded for that part by a lossless coder in C11b (FIG. 1).

[0118] Image IV k The filtering just described in connection with Image IV can be applied to one or more images (regions of these one or more images that may vary depending on the view image) of other views among the above N views.

[0119] Here, referring to FIG. 7, an explanation is provided for a fourth embodiment of the processing applied to the original data of Image IV before the encoding step C11b (FIG. 1) by the second encoding method MC2. k For the example shown in FIG. 7, the processing applied to the original data of Image IV

[0120] in the example shown in FIG. 7, the processing applied to the original data of Image IV k is the detection of occlusion of at least one region Z p of Image IV p using at least one image IV k of another view V OC among the N views (1 ≦ p ≦ N).

[0121] In a method well known per se, such occlusion detection is based on Image IV, for example, using disparity estimation. p Based on Image IVk area Z OC is searched for. Next, the occlusion area Z OC is enlarged, for example, using a mathematical morphology algorithm. The thus enlarged area Z OC is represented by hatching in FIG. 7. The image IV OC that does not form the shaded area (and thus is considered not to require coding) k has all its original data deleted.

[0122] Next, the processed data DT k to be coded, in such an occlusion detection case, - the original pixels included in the shaded area, and - information explaining the applied occlusion detection (such as a value predefined for each pixel not included in the shaded area, represented, for example, by the predefined value YUV = 000) is included.

[0123] Next, in C11b (FIG. 1), the image data (pixels) corresponding to the shaded area are coded by a coder of the type such as HEVC, 3D - HEVC, MV - HEVC. The information explaining the applied occlusion detection is coded for that part by a lossless coder in C11b (FIG. 1).

[0124] Here, with reference to FIG. 8, an explanation is provided for a fifth embodiment of the processing applied to the original data of the image IV k before the coding step C11b (FIG. 1) by the second coding method MC2.

[0125] In the example shown in FIG. 8, the processing applied to the original data of the image IV k is - based on the original pixels of the image IV k and - One or more images IV of other views encoded using the first encoding method MC1 in C11a (Figure 1) j Based on the original pixels of (1 ≤ j ≤ n), - Potentially, an image IV of at least one other view in which the processed pixels in C11b (Figure 1) are encoded using the second encoding method MC2 l Based on the original pixels of (n + 1 ≤ l ≤ N), is to calculate pixels.

[0126] Next, the processed data DT scheduled for encoding k in such a case of calculation,[[]] - Indicators of the pixels of the calculated image IV of the view k - Information about the location of the original pixels used to calculate the pixels of the image IV in the image IV - Image IV j in the image IV k - Potentially, information about the location of the original pixels used to calculate the pixels of the image IV in the image IV - Image IV l in the image IV k including information about the location of the original pixels used to calculate the pixels of the image IV is included.

[0127] The aforementioned calculation is, for example, to subtract the original pixels of the image IV of the view from the original pixels of the image IV j and potentially the original pixels of the image IV k l l of the original pixels of the image IV.

[0128] Here, referring to Figure 9, an explanation of the sixth embodiment of the processing applied to the original data of the image IV before the encoding step C11b (Figure 1) by the second encoding method MC2 is provided. k

[0129] In the example shown in Figure 9, the processing applied to the original data of the image IV k is - Processing the original pixels of the image IV k to obtain the processed data DT’k provides, - processes the original pixels of the image IVs of views Vs (1 ≦ s ≦ N) to provide processed data DTs, - the processed data DT’ k of the image IV k and the processed data DTs of the image IVs are combined into a single image IV one (the obtained original processed data DT k is encoded using the second encoding method MC2 in C11b (Figure 1)), is.

[0130] Figure 10 shows a simple structure of an encoding device COD designed to implement an encoding method according to any one of the specific embodiments of the present invention.

[0131] According to a specific embodiment of the present invention, the actions performed by the encoding method are implemented by computer program instructions. For that purpose, the encoding device COD has a conventional architecture of a computer and specifically includes a memory MEM_C and a processing unit UT_C. The processing unit UT_C is equipped with, for example, a processor PROC_C and is driven by a computer program PG_C stored in the memory MEM_C. The computer program PG_C includes instructions for performing actions of an encoding method such as those described above when the program is executed by the processor PROC_C.

[0132] In initialization, the code instructions of the computer program PG_C are loaded, for example, into a RAM memory (not shown) before being executed by the processor PROC_C. The processor PROC_C of the processing unit UT_C specifically performs the actions of the encoding method described above according to the instructions of the computer program PG_C.

[0133] 7. Implementation of an exemplary decoding scheme In the following, a method for decoding 360°, 180° or other omnidirectional images will be described, which can use any type of multi-view video decoder conforming to, for example, the 3D-HEVC or MV-HEVC standard or the like.

[0134] Referring to FIG. 11, such a decoding method is applied to a data signal representing the current image of the views that form part of the above-mentioned multiple views V1, …, V N The present invention's decoding method is applied to a data signal representing the current image of the views that form part of the above-mentioned multiple views V1, …, V

[0135] The decoding method according to the present invention is - a data signal representing the encoded data associated with the image IV1 of view V1, - a data signal representing the encoded data associated with the image IV2 of view V2, - …, - a data signal representing the encoded data associated with the image IV k of view V k and - …, - a data signal representing the encoded data associated with the image IV N of view V N and decodes the data signals.

[0136] The image of the view to be reconstructed using the above decoding method can similarly be a texture image or a depth image.

[0137] The decoding method includes, for the data signal F1 k representing at least one image IV k of the view V to be reconstructed, k F2 k or F3 k and F'3 k the following.

[0138] In D1, as shown in FIGS. 2A, 2B, and 2C respectively, for the data signal F1 k F2 k or F3 k and F'3 kIn this case, for the image IV k information named flag_proc indicating whether it is encoded using the first encoding method MC1 or the second encoding method MC2 is read.

[0139] For the signal F1 k in this case, the information named flag_proc is 0.

[0140] In D11a, for the data signal F1 k the encoded image IVC k and the encoded data dc1 k associated with it, dc2 k ..., dcQ k are read.

[0141] In D12a, based on the encoded data dc1 k read in D11a, dc2 k ..., dcQ k the decoding method MD1 corresponding to the encoding method MC1 applied to the encoding in C11a of FIG. 1 is used to reconstruct the image IVD k For this purpose, the image IV k is reconstructed using a conventional decoder, such as one conforming to standards such as HEVC, MVC-HEVC, 3D-HEVC, etc.

[0142] When the decoding in D12a is completed, the image IVD k reconstructed in this way contains the original data d1 k of the image IV k encoded in C11a of FIG. 1, d2 k ..., dQ k .

[0143] Since the image IVD k is identical to the original image IV k it is suitable for use, for example, in constructing a master image in the context of synthesizing intermediate views.

[0144] In D1, for the signal F2k or F3 k In the case of, the determined information of flag_proc is 1.

[0145] Signal F2 k In the case of, in D11b, the data signal F2 k In, the processed and encoded image IVTC as obtained in C11b of FIG. 1 k and the processed and encoded data DTC associated therewith k is read.

[0146] These read processed and encoded data DTC k are only of the image type data.

[0147] In D12b, based on the encoded data DTC read in D11b k using the decoding method MD2 corresponding to the encoding method MC2 applied to the encoding in C11b of FIG. 1, the processed image IVTD k is reconstructed. For that purpose, the encoded data DTC k is decoded using a conventional decoder, such as one conforming to standards such as HEVC, MVC-HEVC, 3D-HEVC, etc.

[0148] When the decoding D12b is completed, the processed image IVTD thus reconstructed k corresponds to the decoded data DTC k and includes the processed data DT of those pre-encoding images IV in C11b of FIG. 1 k k including.

[0149] The processed and reconstructed image IVTD k includes the image data (pixels) corresponding to all or part of the original data of the image IV processed using specific image processing, and various detailed examples thereof are described with reference to FIGS. 4 to 9. k

[0150] ​​ Signal F3 k In the case of, in D11c, the processed and encoded image IVTC k and the processed and encoded data DTC associated therewith k are read.

[0151] For that purpose, - In signal F3 k the processed and encoded data DTC of the image type k are read, - In signal F’3 k the processed and encoded data DTC different from the image data, such as, for example, data of the text type, or data including both the image data and data of a type other than the image data k are read.

[0152] The processed and encoded data DTC k is decoded in D12b using the decoding method MD2, and the decoding method MD2 - In particular, for decoding the image data, by a lossy or lossless decoder (such a decoder may be the same as or different from the decoder implementing the first decoding method MD1), - In particular, for decoding the text type data, by a lossless decoder can be implemented.

[0153] When the decoding D12b is completed, - The processed image IVTD thus reconstructed corresponding to the decoded data DTC of the image type k , k - The text type processed data corresponding to the description information of the processing applied to the image IV k before the encoding C11b (FIG. 1) is obtained.

[0154] Such processed and reconstructed image IVTD by the second decoding method MD2 k ​is the image IV before being encoded by C11b following the process k does not include all of the original data of. However, such a reconstructed image of the view by the second decoding method MD2 can be used in addition to the image of the master view reconstructed using the first decoding method MD1, for example, in the context of synthesizing intermediate images, in order to obtain a synthesized image of a view of excellent quality.

[0155] Just now, the decoding methods described above are then, at the current time, for the encoded images IVC1, IVC2,..., IVC N scheduled for reconstruction that are available, for only some of them, or, for the image IV k (for example, k = 1) can be implemented limitedly.

[0156] According to two exemplary embodiments shown in FIGS. 12A and 12B, for example, among the N encoded images IVC1,..., IVC N scheduled for reconstruction, it is assumed as follows. - The n first encoded images IVC1,..., IVC n are reconstructed using the first decoding technique MD1 in order to obtain each image of the n master views respectively. - The N - n other encoded images IVC n+1 ,..., IVC N are reconstructed using the second decoding method MD2 in order to obtain each image of the N - n additional views respectively.

[0157] When n = 0, the images IVC1,..., IVC N of the 1 to N views are reconstructed using the second decoding method MD2. When n = N - 1, an image of a single view is reconstructed using the second decoding method MD2.

[0158] In the example of FIG. 12A, at D100, the data signal F as generated at C14 in FIG. 3A is - Two data signals (data signal F1 as generated at C13a in FIG. 3A, data signal F2 as generated at C13b in FIG. 3A), or, - Three data signals (data signal F1 as generated at C13a in FIG. 3A, data signals F3 and F’3 as generated at C13c in FIG. 3A) are separated.

[0159] In the case of signals F1 and F2, at D110a, the data signal F1 is then separated into n data signals F11, …, F1 n each representing one of the n encoded images IVC1, …, IVC n of the view.

[0160] At D11a, for each of the n data signals F11, …, F1 n the encoded original data dc11, …, dcQ1, …, dc1 n …, dcQ n …, dcQ

[0161] At D12a, based on those respective encoded original data read at D11a, using a conventional HEVC type decoder, the images IVD1, …, IVD n are reconstructed independently of each other.

[0162] Still referring to FIG. 12A, at D110b, the data signal F2 is then separated into N - n data signals F2 n+1 …, DTC N each representing one of the N - n processed and encoded data DTC n+1 …, F2 N …, F2

[0163] At D11b, for each of the N - n data signals F2 n+1 …, F2 N the N - n images IV n+1 …, IV NN - n processed and encoded data DTC corresponding to each of them n+1 , …, DTC N are each read.

[0164] In D12b, based on the N - n processed and encoded data DTC n+1 , …, DTC N read in D11b, using a conventional HEVC - type decoder, the processed images are reconstructed independently of each other. Then, the processed and reconstructed images IVTD n+1 , …, IVTD N are obtained.

[0165] In the case of signals F1, F3, and F’3, in D110a, the data signal F1 is then separated into n encoded image IVC1, …, IVC n represented by n data signals F11, …, F1 n respectively.

[0166] In D11a, for each of the n data signals F11, …, F1 n , the encoded original data dc11, …, dcQ1, …, dc1 n , …, dcQ n associated with each of these n encoded images are each read.

[0167] In D12a, based on their respective encoded original data read in D11a, using a conventional HEVC - type decoder, images IVD1, …, IVD n are reconstructed independently of each other.

[0168] In D110c, - The data signal F3 is then separated into N - n processed and encoded data DTC n+1 , …, DTC N represented by N - n data signals F3 n+1 , …, F3 N respectively, - The data signal F’3 is then separated into N-n processed and encoded data DTC of text or another type, DTC n+1 , …, DTC N represented by N-n data signals F’3 n+1 , …, F’3 N respectively.

[0169] In D11c, - For each of the N-n data signals F3 n+1 , …, F3 N , N-n processed and encoded data DTC of the image type corresponding to each of the N-n images IV n+1 , …, IV N to be reconstructed are each read, n+1 , …, DTC N and - For each of the N-n data signals F’3 n+1 , …, F’3 N , N-n processed and encoded data DTC of text or another type corresponding to the description information of the processing for each of the N-n images IV n+1 , …, IV N to be reconstructed are each read. n+1 , …, DTC N

[0170] In D12b, based on the N-n processed and encoded data DTC n+1 , …, DTC N read in D11b, N-n processed images are each independently reconstructed using a conventional HEVC type decoder. Then, processed and reconstructed images IVTD n+1 , …, IVTD N are obtained.

[0171] Also in D12b, based on the N-n processed and encoded data DTC of text or other types n+1 , …, DTC NBased on this, using a decoder corresponding to the lossless coder used in encoding, the pre-encoded images IV in C11b (FIG. 3A) n+1 , …, IV N There is also reconstructed explanatory information on the processing applied to each of them.

[0172] In the example of FIG. 12B, at D100, the data signal F generated as in C14 of FIG. 3B is - separated into two data signals (data signal F1 generated as in C12a of FIG. 3B, data signal F2 generated as in C12b of FIG. 3B), or - separated into three data signals (data signal F1 generated as in C12a of FIG. 3B, data signals F3 and F'3 generated as in C12c of FIG. 3B) is separated.

[0173] In the case of signals F1 and F2, at D11a, in the data signal F1, the encoded original data dc11, …, dcQ1, …, dc1 1,i , …, IVC n,i respectively associated with each of the encoded views of the image IVC n , …, dcQ n are read.

[0174] At D12a, based on the respective encoded original data read at D11a, using a conventional decoder of the MV-HEVC or 3D-HEVC type, simultaneously, the images IVD 1,i , …, IVD n,i are reconstructed.

[0175] At D11b in FIG. 12B, in the data signal F2, the N-n views of the image IV n+1 , …, IV N each associated with the N-n processed and encoded data DTC n+1 , …, DTC N are read.

[0176] In D12b, based on the N - n processed and encoded data DTC n+1 , …, DTC N read in D11b, using a conventional decoder of the MV - HEVC or 3D - HEVC type, simultaneously, N - n processed images are each reconstructed. Then, the processed and reconstructed images IVTD n+1 , …, IVTD N are obtained.

[0177] In the case of signals F1, F3, and F’3, in D11a, in the data signal F1, the encoded original data dc11, …, dcQ1, …, dc1 n respectively associated with each of the n encoded view images IVC1, …, IVC n , …, dcQ n are read.

[0178] In D12a, based on those respective encoded original data read in D11a, using a conventional decoder of the MV - HEVC or 3D - HEVC type, simultaneously, images IVD1, …, IVD n are reconstructed.

[0179] In D11c of FIG. 12B, in the data signal F3, the N - n image types of processed and encoded data DTC n+1 , …, IV N respectively associated with each of the N - n images IV n+1 , …, DTC N are read.

[0180] In D12b, based on the N - n processed and encoded data DTC n+1 , …, DTC N read in D11c, using a conventional decoder of the MV - HEVC or 3D - HEVC type, simultaneously, the processed images are each reconstructed. Then, the view IVTD n+1 , …, IVTD NA processed and reconstructed image is obtained.

[0181] Also, in D12b of FIG. 12B, the text or other type of N-n processed and encoded data DTC read in D11c n+1 , …, DTC N Based on these, the pre-encoded images IV in C11b (FIG. 3B) that are decoded using a decoder corresponding to the lossless coder used in the encoding n+1 , …, IV N There is also reconstructed description information of the processing applied to each of them.

[0182] Naturally, other combinations of decoding methods are also possible.

[0183] According to one possible variant of FIG. 12A, it is also possible that the decoder implementing the decoding method MD1 is a HEVC type decoder and the decoder implementing the decoding method MD2 is an MV-HEVC or 3D-HEVC type decoder.

[0184] According to one possible variant of FIG. 12B, it is also possible that the decoder implementing the decoding method MD1 is an MV-HEVC or 3D-HEVC type decoder and the decoder implementing the decoding method MD2 is a HEVC type coder.

[0185] FIG. 13 shows a simple structure of a decoding device DEC designed to implement a decoding method according to any one of the specific embodiments of the present invention.

[0186] According to a particular embodiment of the invention, the actions performed by the decoding method are implemented by computer program instructions. To that end, the decoding device DEC has a conventional architecture of a computer and in particular comprises a memory MEM_D and a processing unit UT_D, for example equipped with a processor PROC_D and driven by a computer program PG_D stored in the memory MEM_D. The computer program PG_D comprises instructions for carrying out the actions of the decoding method, such as those described above, when the program is executed by the processor PROC_D.

[0187] At initialization, the code instructions of the computer program PG_D are loaded, for example, into a RAM memory (not shown) before being executed by the processor PROC_D, which in particular performs the actions of the decoding method described above according to the instructions of the computer program PG_D.

[0188] According to one embodiment, the decoding device DEC is for example comprised in a terminal.

[0189] 8. Exemplary Applications of the Invention to Image Processing As already explained above, N reconstructed images IVD1, ..., IVD n and IVTD n+1 , …, I.V.T.D. N can be used to synthesize the intermediate views required by the user.

[0190] As shown in FIG. 14, in the case where a user needs to synthesize images of arbitrary views, the views IVD1, ..., IVD2, which are considered as master views, are n The n initial reconstructed images are sent to an image synthesis module in S1.

[0191] Nn processed and reconstructed images of views IVTD n+1 , …, I.V.T.D. NAs additional images of the view, in order to be able to be used in the synthesis of images, in S2, it may be necessary to process using those images and the decoded image processing description information respectively associated therewith.

[0192] When process S2 is completed, N - n reconstructed images IVD of the view n+1 , …, IVD N are obtained.

[0193] Next, the N - n reconstructed images IVD of the view n+1 , …, IVD N are sent to the image synthesis module in S3.

[0194] In S4, at least one of the n first reconstructed view images IVD1, …, IVD n and potentially at least one of the N - n images IVD of the N - n reconstructed views n+1 , …, IVD N are used to synthesize the view image.

[0195] Next, when the synthesis S4 is completed, the image IV of the synthesized view SY is obtained.

[0196] It should also be noted that the n reconstructed images IVD1, …, IVD n can also perform process S2. Such a process S2 can be seen to be necessary in cases where the user UT requires an image of a view whose represented viewing angle does not match one or more of the viewing angles of the n reconstructed images IVD1, …, IVD n . For example, while the user UT requests an image of a view representing a 120×90 viewing field, each of the n reconstructed images IVD1, …, IVD n may represent a 360×180 viewing angle. Such a possibility of processing for the reconstructed images IVD1, …, IVD n is represented by the dashed arrow in FIG. 14. In addition, in relation to the types of processes described with reference to FIGS. 8 and 9, the reconstructed images IVD1, …, IVDn is for use during the processing of N - n other images IV n+1 , …, IV N and can be used during the processing thereof.

[0197] Here, with reference to FIGS. 15A to 15C, a first embodiment of the processing applied to the data of the processed and reconstructed image IVTD k will be described. Such processing obtains the initial resolution of the corresponding view image IV k sampled before encoding in C11b of FIG. 1.

[0198] In the example of FIG. 15A, the processing applied before encoding is assumed to be vertical downsampling of the region Z4 of the image IV k as shown in FIG. 5A.

[0199] The processing applied to the processed and reconstructed image IVTD k uses the information explaining the applied downsampling to return to the initial resolution of the image IV k , specifically, - The downsampling factor used (from which the corresponding upsampling factor can be determined), - The downsampling direction used (from which the corresponding upsampling direction can be determined), - The location of the downsampled region Z4 of the image IV k and the like to apply upsampling corresponding to the downsampling applied in FIG. 5A to the region Z4.

[0200] In the example of FIG. 15B, the processing applied before encoding is assumed to be horizontal downsampling of the region Z5 of the image IV k as shown in FIG. 5B.

[0201] The processing applied to the processed and reconstructed image IVTD k uses the information explaining the applied downsampling to return to the initial resolution of the image IVk To return to the initial resolution of - The downsampling factor used (from which the corresponding upsampling factor can be determined), - The downsampling direction used (from which the corresponding upsampling direction can be determined), - Image IV k The location of the downsampled region Z5 of etc. are used to apply upsampling corresponding to the downsampling applied in Figure 5B to region Z5.

[0202] In the example of Figure 15C, the processing applied before encoding is assumed to be the overall downsampling of Image IV k as shown in Figure 5C.

[0203] The processed and reconstructed Image IVTD k The processing applied to k To return to the initial resolution of Image IV - The downsampling factor used (from which the corresponding upsampling factor can be determined), - The downsampling direction used (from which the corresponding upsampling direction can be determined) etc. are used to apply upsampling corresponding to the downsampling applied in Figure 5C to all of the image data of Image IVTD k

[0204] In the example of Figure 15D, the processing applied before encoding is assumed to be the downsampling in both the horizontal and vertical directions of region Z6 of Image IV k as shown in Figure 5D.

[0205] The processed and reconstructed Image IVTD​k The process applied to k restore the initial resolution of Image IV, uses the information explaining the applied downsampling, specifically, - The downsampling factor used (by which the corresponding upsampling factor can be determined), - The downsampling direction used (by which the corresponding upsampling direction can be determined), - Image IV k The location of the downsampled area Z6 of etc. to apply upsampling corresponding to the downsampling applied in Figure 5D to area Z6.

[0206] Here, referring to Figure 16, a second embodiment of the process applied to the data of the processed and reconstructed Image IVTD k is described. Such a process restores one or more contours of Image IV k of the view that has been filtered before encoding of Image IV k is.

[0207] In the example of Figure 16, the process applied before encoding is assumed to be the filtering of the contours ED1 and ED2 of Image IV k as shown in Figure 6.

[0208] Next, the process applied to the processed and reconstructed view of Image IVTD k uses the information explaining the applied filtering, specifically, predefined values (specifically, predefined value YUV = 000) for each pixel that was not filtered, etc., to restore the contours ED1 and ED2 of Image IV k is.

[0209] Here, referring to Figure 17, the processed and reconstructed Image IVTD kA third embodiment of the process applied to the data will be described. Such a process reconstructs the pixels of the image IV of the view calculated before encoding according to the embodiment of the process of FIG. 8 k of the pixels.

[0210] Next, the process applied to the processed and reconstructed image IVTD k is - Using the indicators read in the data signal, such as the indicators of the pixels of the calculated view image IV k to recover the pixels of the above-mentioned view image IV k calculated in the encoding, - Using the information about the location of the pixels in the image IV k used to calculate the pixels of the image IV j to recover the pixels of at least one other view image IV j (1 ≤ j ≤ n) reconstructed using the first decoding method MD1, - Potentially, the processed pixels are the pixels of at least one other view image IV l (n + 1 ≤ l ≤ N) decoded using the second decoding method MD2 to recover the pixels.

[0211] Next, the decoding of the processed data DT k is - Based on the pixels of at least one other view image IV j (1 ≤ j ≤ n), - Potentially, based on the pixels of at least one other view image IV l (n + 1 ≤ l ≤ N), to calculate the pixels of the image IV k .

[0212] The aforementioned calculations are, for example, to combine the pixels of the image IV k with the pixels of the reconstructed image IV j and potentially with the pixels of the reconstructed image IV l .

[0213] Here, with reference to FIG. 18, a fourth embodiment of the processing applied to the data of the processed and reconstructed image IVTD k will be described. Such processing reconstructs the pixels of the view image IV k calculated before encoding according to the embodiment of the processing in FIG. 9.

[0214] The processing is first applied to the reconstructed image IVD one by the second decoding method MD2. Then, the processing is based on the image IVD one to - the processed data DT' k of the image IV k decoded according to the second decoding method MD2, - based on the processed data DTs of the images IVDs of the decoded views Vs (1 ≦ s ≦ N), the pixels of the image IV k are reconstructed.

[0215] 9. Exemplary Specific Applications of the Present Invention According to the first example, it is considered that six images of a view with a resolution of 4096×2048 pixels are each captured by six 360° type cameras. A depth estimation method is applied to provide six corresponding 360° depth maps.

[0216] The image IV0 of view V0 is encoded using the first encoding method MC1 in the conventional manner, and the other five images IV1, IV2, IV3, IV4, IV5 of views V1, V2, V3, V4, V5 are cropped before encoding. The process applied to each of the images IV1, IV2, IV3, IV4, IV5 is to remove columns of constants (e.g., 200) on the right and left sides of each of these images. The number of columns to be removed is selected such that the viewing angle is reduced from 360° to 120°. Similarly, the process applied to each of the images IV1, IV2, IV3, IV4, IV5 is to delete rows of constants (e.g., 100) from the top and bottom of each of these images respectively. The number of rows to be deleted is selected such that the viewing angle is reduced from 180° to 120°.

[0217] The item of information named flag_proc is set to 0 in relation to the image IV0, and the item of information named flag_proc is set to 1 in relation to the images IV1, IV2, IV3, IV4, IV5.

[0218] The image IV0 of the view is encoded using an HEVC coder. A single data signal F10 is generated, and the said data signal includes the encoded original data of the image IV0 and the information that flag_proc = 0.

[0219] The data of the remaining area after cropping of each of the images IV1, IV2, IV3, IV4, IV5 is encoded using an HEVC coder. Five data signals F21, F22, F23, F24, F25 are generated, and the data signals respectively include the encoded data of the remaining area after cropping of each of the images IV1, IV2, IV3, IV4, IV5 associated with the information that flag_proc = 1. The data signals F10, F21, F22, F23, F24, F25 are concatenated and then transmitted to the decoder.

[0220] The five data signals F21, F22, F23, F24, F25 may further include the coordinates of the cropped area in the following manner. IV1, IV2, IV3, IV4, IV5: flag_proc = 1, upper left point (h, v) = (0 + 200, 0 + 100), lower right point (h, v) = (4096 - 200, 2048 - 100), where "h" is the horizontal direction and "v" is the vertical direction.

[0221] On the decoder side, information called flag_proc is read.

[0222] When flag_proc = 0, the view image IV0 is reconstructed using the HEVC decoder.

[0223] When flag_proc = 1, the images IV1, IV2, IV3, IV4, IV5 corresponding to the processed and encoded data are reconstructed using the HEVC decoder. Since the reconstruction of the data of these images deleted by cropping is impossible, no processing is applied to the reconstructed images IV1, IV2, IV3, IV4, IV5. However, the synthesis algorithm uses the six reconstructed images IV0, IV1, IV2, IV3, IV4, IV5 to generate an image of any view required by the user.

[0224] In the case where the five data signals F21, F22, F23, F24, F25 further include the coordinates of the cropped area, these coordinates are used by the synthesis algorithm to generate an image of the view required by the user.

[0225] According to the second example, ten images IV0, …, IV9 having a resolution of 4096×2048 for each pixel generated by a computer to simulate ten 360° type cameras are considered. It is determined not to process images IV0 and IV9. The texture components of images IV1 to IV8 are downsampled by a factor of 2 in the horizontal direction and a factor of 2 in the vertical direction for that part, and the corresponding depth components are downsampled by a factor of 4 in the horizontal direction and a factor of 4 in the vertical direction for that part. As a result, the resolution of the texture components of images IV1 to IV8 becomes 2048×1024, and the resolution of the depth components of images IV1 to IV8 becomes 1024×512.

[0226] Processing data such as image data related to images IV1 to IV8 includes eight downsampled texture components with a resolution of 2048×1024 of images IV1 to IV8 and eight downsampled depth components with a resolution of 1025×512 of images IV1 to IV8.

[0227] In addition, the aforementioned processing data includes text type data indicating the downsampling factor for each of images IV1 to IV8 of views 1 to 8. Those data are - Texture of IV1 to IV8: se_h = 2, se_v = 2 (“se” is downsampling, “h” is the horizontal direction, “v” is the vertical direction) - Depth of IV1 to IV8: se_h = 4, ss_v = 4 described as follows.

[0228] An item of information named flag_proc is set to 0 in relation to images IV0 and IV9, and an item of information named flag_proc is set to 1 in relation to images IV1 to IV8.

[0229] Images IV0 and IV9 are simultaneously encoded using an MV-HEVC type coder, thereby generating a single data signal F1 including information flag_proc = 0 and the encoded original data of images IV0 and IV9.

[0230] Also, processing data such as image data related to Images IV1 to IV8, including eight downsampled texture components with a resolution of 2048×1024 and eight downsampled depth components with a resolution of 1025×512 of Images IV1 to IV8, is also simultaneously encoded using an MV-HEVC type coder, whereby a single data signal F2 including information of flag_proc = 1 associated with the downsampled and encoded texture and depth data is generated. Text type data indicating the downsampling factor for each of Images IV1 to IV8 of Views 1 to 8 is losslessly encoded for that portion. Data signals F1 and F2 are concatenated and then transmitted to the decoder.

[0231] On the decoder side, information of flag_proc is read.

[0232] When flag_proc = 0, Images IV0 and IV9 are simultaneously reconstructed using an MV-HEVC decoder. Then, reconstructed Images IVD0 and IVD9 at their initial resolutions are obtained.

[0233] When flag_proc = 1, Images IV1 to IV8 corresponding to their respective downsampled and encoded texture and depth data are simultaneously reconstructed using an MV-HEVC decoder. Then, downsampled and reconstructed Images IVDT1 to IVDT8 are obtained. Text type data corresponding to each of the eight Images IV1 to IV8 is also decoded, providing the downsampling factor used for each of Images IV1 to IV8.

[0234] Next, the downsampled and reconstructed images IVDT1 to IVDT8 are processed using their corresponding downsampling factors. When the processing is completed, the reconstructed images IVD1 to IVD8 are obtained, and their eight respective texture components are at their initial resolution of 4096×2048, and their eight respective depth components are at their initial resolution of 4096×2048.

[0235] The compositing algorithm uses the ten view images thus constructed at their initial resolutions to generate the images of the views required by the user.

[0236] According to a third example, three images IV0 to IV2 each having a pixel resolution of 4096×2048 generated by a computer to simulate four 360°-type cameras are considered. Then, three texture components and three corresponding depth components are obtained. It is determined not to process image IV0 and to extract occlusion maps for each of images IV1 and IV2. For that purpose, a disparity estimation is performed between image IV1 and image IV0, and an occlusion mask of image IV1 (i.e., the pixels of image IV1 not seen in image IV0) is generated. Also, a disparity estimation is performed between image IV2 and image IV0, and an occlusion mask of image IV2 is generated.

[0237] The processing data such as the image data related to images IV1 and IV2 includes two texture components of the occlusion masks of images IV1 and IV2 and two depth components of the occlusion masks of images IV1 and IV2.

[0238] An item of information named flag_proc is set to 0 in relation to image IV0, and an item of information named flag_proc is set to 1 in relation to images IV1 and IV2.

[0239] Image IV0 is encoded using the HEVC coder. A single data signal F10 is generated, and the data signal includes the encoded original data of Image IV0 and the information flag_proc = 0.

[0240] The image data (texture and depth) of the occlusion mask for each of Images IV1 and IV2 is encoded using the HEVC coder. Two data signals F21 and F22 are generated, and the data signals respectively include the encoded image data of the occlusion masks for Images IV1 and IV2 associated with the information flag_proc = 1. The data signals F10, F21, and F22 are concatenated and then transmitted to the decoder.

[0241] On the decoder side, the information flag_proc is read.

[0242] If flag_proc = 0, Image IV0 is reconstructed using the HEVC decoder.

[0243] If flag_proc = 1, Images IV1 and IV2 corresponding to the encoded image data (texture and depth) of the occlusion masks for each of Images IV1 and IV2 are reconstructed using the HEVC decoder. Since it is impossible to reconstruct the data of these images that were deleted upon completion of occlusion detection, no processing is applied to the reconstructed Images IV1 and IV2. However, the synthesis algorithm can use the reconstructed Images IV0, IV1, and IV2 to generate the image of the view required by the user.

[0244] According to a fourth example, it is considered that two images IV0 and IV1 with a resolution of 4096×2048 pixels are respectively captured by two 360° type cameras. The image IV0 of the first view is encoded using the first encoding method MC1 in the conventional manner, and the image IV1 of the second view is processed before encoding by the second encoding method MC2. Such processing is - Extract the contour of the image IV1 using a filter (e.g., Sobel filter, etc.); - Apply dilation of the contour using, for example, a mathematical morphology operator to increase the area around the contour; and include;

[0245] The processing data such as the image data related to the image IV1 includes pixels within the area around the contour and pixels set to 0 corresponding to each of the pixels located outside the area around the contour.

[0246] In addition, text type data is generated in the form of marking information (e.g., YUV = 000) indicating that the pixels located outside the area around the contour are set to 0. The pixels set to 0 are neither encoded nor transmitted to the decoder.

[0247] The image IV0 is encoded using an HEVC coder, thereby generating a data signal F1 including information of flag_proc = 0 and the encoded original data of the image IV0.

[0248] The image data of the area around the contour of the image IV1 is encoded using an HEVC coder, and the marking information is encoded using a lossless coder. Then, a data signal F2 is generated, and the above signal includes information of flag_proc = 1, the encoded pixels of the area around the contour of the image IV1, and an item of the encoded marking information.

[0249] On the decoder side, the information of flag_proc is read.

[0250] When flag_proc = 0, the image IV0 is reconstructed at its original resolution using an HEVC decoder.

[0251] When flag_proc = 1, the image IV1 corresponding to the image data of the area around the contour of the image IV1 is reconstructed by the HEVC decoder using the marking information that enables the restoration of the pixels surrounding the area with the value set to 0.

[0252] The synthesis algorithm can use the two reconstructed images IV0 and IV1 to generate the images of the views required by the user.

[0253] According to the fifth example, it is considered that four images IV0 to IV3 with a resolution of 4096×2048 pixels are respectively captured by four 360° type cameras. The image IV0 is encoded using the first encoding method MC1 in the conventional manner, and the images IV1 to IV3 are processed before encoding by the second encoding method MC2. Such processing is the filtering of the images IV1 to IV3, during which the region of interest ROI is calculated. The region of interest includes, for example, one or more regions of each of the images IV1 to IV3 that are considered to be the most relevant because they contain a lot of details.

[0254] Such filtering is performed, for example, according to one of the following two methods. - Calculate the saliency map of each of the images IV1 to IV3 by filtering. - Filter the depth map of each of the images IV1 to IV3. The depth map is characterized by whether each texture pixel is a near-distance depth value or a far-distance depth value in the 3D scene. A threshold is defined, and each pixel of the images IV1, IV2, and IV3 located below this threshold is associated with an object in the scene close to the camera. Then, all pixels located below this threshold are considered to be the region of interest.

[0255] The processed data such as the image data related to the images IV1 to IV3 includes the pixels within their respective regions of interest, and the pixels corresponding to the pixels located outside these regions of interest are set to 0.

[0256] In addition, for example, text type data is generated in the form of marking information indicating that pixels located outside the region of interest are set to 0. The pixels set to 0 are neither encoded nor transmitted as signals to the decoder.

[0257] Image IV0 is encoded using the HEVC coder, thereby generating a data signal F10 including information of flag_proc = 0 and the encoded original data of image IV0.

[0258] The image data of the regions of interest of each of the images IV1, IV2, and IV3 is encoded using the HEVC coder, and the marking information is encoded using a lossless coder. Three data signals F21, F22, and F23 are generated, and these signals respectively include the encoded image data of the regions of interest of each of the images IV1, IV2, and IV3 associated with the information of flag_proc = 1 and the corresponding items of the encoded marking information. The data signals F10, F21, F22, and F23 are concatenated and then transmitted to the decoder.

[0259] On the decoder side, the information of flag_proc is read.

[0260] When flag_proc = 0, image IV0 is reconstructed at its original resolution using the HEVC decoder.

[0261] When flag_proc = 1, each of the images IV1 to IV3 corresponding to the image data of their respective regions of interest is reconstructed using the HEVC decoder with the marking information that enables the restoration of the pixels surrounding the region with the value set to 0.

[0262] The synthesis algorithm can directly use the four reconstructed images IV0, IV1, IV2, and IV3 to generate the image of the view required by the user.

[0263] Needless to say, the embodiments described above are provided purely as a completely non-limiting illustration, and those skilled in the art can easily make many changes without departing from the scope of the present invention.

Claims

1. 1. A method for encoding an image of a view-forming portion of a number of views, the number of views simultaneously representing a 3D scene from different viewing angles or positions, the method being implemented by an encoding device, - selecting (C1) a first encoding method or a second encoding method for encoding the image of the view, the first encoding method encoding original data of the image of the view and the second encoding method encoding data resulting from processing of the image of the view; generating a data signal including information (flag_proc) indicating whether the first or the second encoding method has been selected (C10, C12a; C10, C12b; C10, C12c); - if said first encoding method is selected, encoding (C11a) said original data of said image of said view, said first encoding method providing encoded original data; if said second encoding method is selected, - encoding (C11b) processed data of the image of the view, the data corresponding to at least one remaining region of the image of the view, the remaining region being obtained by a cropping applied to the original data of the image of the view, the second encoding method providing at least one encoded remaining region; - encoding (C11b) description information of said cropping, said description information being information about the position of said remaining area in said image of said view; Including, said generated data signal being if the first encoding method is selected, the encoded original data of the image of the view; if the second encoding method is selected, the encoded description of the encoded remaining area and of the cropping of the image of the view; The method further comprising:

2. the remaining region of the image of the view corresponds to pixels of the image of the view that have not been deleted after applying the cropping to the image of the view, The method of claim 1.

3. the explanatory information of the cropping comprises some coordinates of a pixel located at the top left of the remaining area of ​​the image of the view and some coordinates of another pixel located at the bottom right of the remaining area of ​​the image of the view, The method of claim 1.

4. the explanatory information of the cropping comprises the number of rows and / or columns of pixels that have been removed in the image of the view, and the positions of the rows and / or columns in the image of the view; The method of claim 1.

5. The cropping is configured to remove a certain amount of pixels of the current image of the view. The method of claim 1.

6. 1. A method for decoding a data signal representing an image of a view-forming portion of a number of views, the number of views simultaneously representing a 3D scene from different viewing angles or positions, the method being implemented by a decoding device, the method comprising: reading (D1) in said data signal an item of information (flag_proc) indicating whether said image of said view should be decoded according to a first or a second decoding method, said first decoding method decoding original data encoded for said image of said view and said second decoding method decoding data resulting from processing of said image of said view that has been encoded; if it is the first decoding method, - reading (D11a) in said data signal the encoded data associated with said image of said view; reconstructing (D12a) an image of said view based on said encoded data, said image of said view comprising said original data of said image of said view; in the case of said second decoding method, - reading (D11b; D11c) of said data signal, - coded data associated with said image of said view, said coded data corresponding to at least one remaining area of ​​said image of said view that has been coded, said remaining area being obtained by a cropping applied to the original data of the current image of said view; - reading (D11b; D11c) explanatory information of said cropping, said explanatory information being information about the position of said remaining area in the image of said view, - reconstructing (D12b) said image of said view based on said coded residual region and on said cropping description information; The method includes:

7. the remaining region of the image of the view corresponds to pixels of the image of the view that have not been deleted after applying the cropping to the image of the view, The method according to claim 6.

8. the explanatory information of the cropping includes a number of coordinates of a top left pixel in the remaining area of ​​the image of the view, and a number of coordinates of another bottom right pixel in the remaining area of ​​the image of the view; The method according to claim 6.

9. The method of claim 6 , wherein the explanatory information of the cropping includes a number of rows and / or columns of pixels that have been removed in the image of the view, and the positions of the rows and / or columns in the image of the view.

10. 1. A device for encoding an image of a view-forming portion of a multiple number of views, said multiple views simultaneously representing a 3D scene from different viewing angles or positions, - selecting a first encoding method or a second encoding method for encoding data of the image of the view, the first encoding method encoding original data of the image of the view and the second encoding method encoding data resulting from processing of the image of the view; generating a data signal including information indicating whether said first encoding method or said second encoding method has been selected; - if the first encoding method is selected, encoding the original data of the image of the view, the first encoding method providing encoded original data; if said second encoding method is selected, - encoding processed data of the image of the view, the data corresponding to at least one remaining region of the image of the view, the remaining region being obtained by a cropping applied to the original data of the image of the view, the second encoding method providing at least one encoded remaining region; - encoding explanatory information of the cropping, the explanatory information being information about the position of the remaining area in the image of the view; a processor configured to implement said generated data signal being if the first encoding method is selected, the encoded original data of the image of the view, if the second encoding method is selected, further comprising the encoded description of the cropping and the encoded remaining area of ​​the image of the view, device.

11. 1. A device for decoding a data signal representing an image of a view-forming portion of a multiple number of views, said multiple views simultaneously representing a 3D scene from different viewing angles or positions, reading, in said data signal, an item of information (flag_proc) indicating whether said image of said view should be decoded according to a first or a second decoding method, said first decoding method decoding original data encoded of said image of said view and said second decoding method decoding data resulting from processing of said image of said view that has been encoded; if it is the first decoding method, - reading, in said data signal, encoded data associated with said image of said view; reconstructing an image of said view based on the encoded data read, said reconstructed image of said view including said original data of said image of said view; in the second decoding method, in the data signal: - coded data related to said image of said view, said coded data corresponding to at least one remaining area of ​​said image of said view that has been coded, said remaining area being obtained by a cropping applied to the original data of the current image of said view; - reading explanatory information of the cropping, the explanatory information being information about the position of the remaining area in the image of the view; - reconstructing the image of the view based on the coded residual region and on the cropping description information; a processor configured to implement device.

12. A computer program comprising instructions comprising program code instructions for carrying out the steps of the encoding method according to any one of claims 1 to 5, when said computer program is executed on a computer.

13. A computer readable recording medium readable by a computer and comprising instructions for the computer program of claim 12.

14. A computer program comprising instructions, which when executed on a computer include program code instructions for carrying out the steps of the decoding method according to any one of claims 6 to 9.

15. A computer-readable recording medium that can be read by a computer and contains instructions for the computer program according to claim 14.

Citation Information

Patent Citations

  • Stereoscopic video encoding device, stereoscopic video decoding device, stereoscopic video encoding method, stereoscopic video decoding method, stereoscopic video encoding program, and stereoscopic video decoding program

    JP2014132721A

  • Encoding device and program, decoding device and program, and distribution system

    JP2017152905A

  • Three-dimensional video with asymmetric spatial resolution

    WO2013022540A1

  • Image processing device and image processing method

    WO2018150933A1