Encoding and decoding of omnidirectional video
By employing two encoding methods to encode the 360° video view, encoding the original data and the processed data respectively, a data signal for cropping description information is generated, solving the problem of low encoding efficiency in existing technologies and realizing a highly efficient encoding and decoding process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ORANGE SA
- Filing Date
- 2019-09-25
- Publication Date
- 2026-04-21
AI Technical Summary
Existing 360° video coding technologies are not efficient enough in terms of compression and coding efficiency, especially for 3D scenes with multiple 360° views. Conventional encoders cannot effectively utilize the predictive information between images, resulting in large data volume and low coding efficiency.
Two encoding methods are employed: the first encoding method encodes the original data of the view, and the second encoding method encodes the processed data of the view, generating a data signal containing clipping description information to optimize the transmission rate of the encoded data.
By combining two encoding techniques, the transmission rate of encoded data is significantly reduced, while ensuring that the synthesis effect of intermediate views is not compromised, thus achieving an efficient encoding and decoding process.
Smart Images

Figure CN118317113B_ABST
Abstract
Description
[0001] This patent application is a divisional application of the following invention patent application:
[0002] Application Number: 201980064753.X
[0003] Application date: September 25, 2019
[0004] Invention Title: Encoding and Decoding of Omnidirectional Video Technical Field
[0005] This invention generally relates to the field of omnidirectional video, such as specifically 360°, 180°, etc. More specifically, this invention relates to the encoding and decoding of 360°, 180°, etc. views captured to generate such videos, and to the synthesis of uncaptured intermediate viewpoints.
[0006] This invention can be specifically, but not exclusively, applied to video encoding implemented in current AVC and HEVC video encoders and their extensions (MVC, 3D-AVC, MV-HEVC, 3D-HEVC, etc.), and to corresponding video decoding. Background Technology
[0007] To generate omnidirectional video, such as 360° video, the common practice is to use a 360° camera. This 360° camera consists of multiple 2D cameras mounted on a spherical platform. Each 2D camera captures a specific angle of the 3D scene, and the set of views captured by the cameras allows for the generation of video representing a 3D scene with a 360° × 180° field of view. A single 360° camera can also be used to capture a 3D scene with a 360° × 180° field of view. This field of view can, of course, be smaller, such as 270° × 135°.
[0008] Subsequently, such 360° videos allow users to view the scene as if they were at the center of it, looking around in a 360° range, thus providing a new way to watch videos. These videos are typically reproduced on virtual reality headsets (also known as “head-mounted devices” or HMDs). However, they can also be displayed on a 2D screen equipped with suitable user interaction devices. The number of 2D cameras used to capture 360° scenes varies depending on the platform used.
[0009] To generate 360° video, divergent views captured by various 2D cameras are placed end-to-end, taking into account overlap between views, to create a panoramic 2D image. This step is also known as "stitching." For example, isometric projection (ERP) is one possible projection for obtaining such a panoramic image. According to this projection, the view captured by each 2D camera is projected onto a spherical surface. Other types of projection are also possible, such as cube mapping (projection onto a cube face). The views projected onto the surface are then projected onto a 2D plane to obtain a 2D panoramic image that includes all views of the scene captured at a given time.
[0010] To enhance immersion, multiple 360° cameras of the aforementioned type can be used simultaneously to capture the scene, with these cameras positioned arbitrarily within the scene. The 360° cameras can be actual cameras (i.e., physical objects) or virtual cameras (in which case the view is obtained through view generation software). Specifically, such virtual cameras enable the generation of views representing perspectives of a 3D scene not captured by actual cameras.
[0011] Subsequently, images of a 360° view obtained using a single 360° camera or images of a 360° view obtained using multiple 360° cameras (real and virtual) are encoded using devices such as:
[0012] - Conventional 2D video encoders, such as encoders conforming to the HEVC (short for "High-Efficiency Video Coding") standard,
[0013] - Conventional 3D video encoders, such as encoders compliant with MV-HEVC and 3D-HEVC standards.
[0014] Given the sheer volume of data required to encode a single 360° view image, let alone multiple 360° view images, and considering the specific geometry of the 360° representation of a 3D scene using such 360° views, this encoder is insufficiently efficient for compression. Furthermore, since the views captured by the 2D cameras of a 360° camera are divergent, the aforementioned encoders are unsuitable for encoding different images of a 360° view, as they make almost no use of inter-image prediction. Specifically, there is little predictable similarity between two views captured separately by two 2D cameras. Therefore, all images of a 360° view are compressed in the same way. Specifically, for the current 360° view image to be encoded, no analysis is performed in these encoders to determine whether it makes sense to encode all or only some of the data of this image as part of a synthesis of uncaptured intermediate view images that will be used to encode this image of the subsequently decoded view. Summary of the Invention
[0015] One of the objectives of this invention is to correct the shortcomings of the prior art.
[0016] Therefore, one aspect of the present invention relates to a method, implemented by an encoding device, for encoding an image of a view forming part of a plurality of views, the plurality of views simultaneously representing a 3D scene from different viewpoints or positions, the method comprising the following:
[0017] - Select either a first encoding method or a second encoding method to encode the image of the view.
[0018] - Generate a data signal containing information indicating whether the first encoding method or the second encoding method has been selected.
[0019] - If the first encoding method is selected, the raw data of the image of the view is encoded, and the first encoding method provides the encoded raw data.
[0020] -If the second encoding method is selected:
[0021] The second encoding method encodes the processed data of the image of the view, the processed data corresponding to at least one remaining region of the image of the view, the remaining region being obtained by applying cropping to the original data of the image of the view, and providing at least one encoded remaining region.
[0022] • Encode the descriptive information of the cropping, which is information about the location of the remaining region in the image of the view.
[0023] -The generated data signal further includes:
[0024] • If the first encoding method has been selected, then the original data of the encoded image of the view,
[0025] • If the second encoding method has been selected, then the remaining area of the image of the view and the description information of the encoded cropped area.
[0026] With the aid of this invention, in multiple images of a current view to be encoded of the type described above, where the images represent a very large amount of data to be encoded and therefore transmitted, two encoding techniques may be combined for each image of each view to be encoded:
[0027] - A first encoding technique, according to which images of one or more views are encoded in a conventional manner (e.g., HEVC, MVC-HEVC, 3D-HEVC) to obtain reconstructed images that form views of very high quality.
[0028] - A second innovative coding technique, according to which processed data of images of one or more other views are encoded so that processed image data not corresponding to the original data of these images is obtained during decoding, but with the benefit of significantly reducing the signal transmission cost of the encoded processed data of these images.
[0029] Subsequently, for each image of each other view (whose processed data has been encoded according to the second encoding method), what is obtained during decoding is the corresponding processed data of the view's image, along with image processing descriptive information applied to the original data of the view's image during encoding. This processed data can then be processed using the corresponding image processing descriptive information to form an image of the view, which will be used in conjunction with at least one of the images of the view reconstructed according to the first method of conventional decoding, enabling the synthesis of images of uncaptured intermediate views in a particularly efficient and effective manner.
[0030] The present invention also relates to a method for decoding data signals of images representing a portion of a plurality of views, implemented by a decoding device, of a 3D scene simultaneously from different viewpoints or positions, the method comprising the following:
[0031] - In this data signal, information indicating whether the image of the view will be decoded according to a first decoding method or a second decoding method is read.
[0032] -If it is the first decoding method:
[0033] Then, the encoded data associated with the image of the view is read from the data signal.
[0034] • An image of the view is reconstructed based on the encoded data, and the reconstructed image of the view contains the original data of the original image of the view.
[0035] -If it is the second decoding method:
[0036] Then read from this data signal:
[0037] - Encoded data associated with the image of the view, the encoded data corresponding to at least one remaining region of the image of the view that has already been encoded, the remaining region being obtained by applying cropping to the original data of the current image of the view.
[0038] - The cropping description information, the encoded description information being information about the location of the remaining region within the image of the view.
[0039] • The image of the view is reconstructed based on the encoded remaining region and the description information of the cropping.
[0040] This cropping process applied to the image of the view allows for the avoidance of encoding a portion of its original data, offering the benefit of significantly reducing the transmission rate of encoded data associated with the image of the view, since data belonging to one or more cropped regions is neither encoded nor transmitted to the decoder. The rate reduction will depend on the size of the one or more cropped regions. Therefore, the image of the view to be reconstructed after decoding, and subsequently processing its processed data using appropriate image processing description information, will not contain all of its original data or will at least be different from the original image of the view. However, obtaining such an image of the view cropped in this way does not impair the effectiveness of the synthesis of intermediate images, which will be used once reconstructed. Indeed, using this synthesis of one or more images reconstructed using conventional decoders (e.g., HEVC, MVC-HEVC, 3D-HEVC), the original regions can be retrieved in the intermediate view from the image of the view and the conventionally reconstructed image.
[0041] According to another specific embodiment:
[0042] - These processed data of the image of the view are data of at least one region of the image of the view, which has been sampled according to a given sampling factor and in at least one given direction.
[0043] - The descriptive information of the image processing includes at least one piece of information about the location of the at least one sampling region in the image of the view.
[0044] This processing facilitates uniform degradation of the image of the view, again aiming to optimize the data rate reduction caused by the applied sampling, followed by encoding. Subsequent reconstruction of this image of the view sampled in this manner does not compromise the effectiveness of the synthesis of intermediate images (which will use this image of the reconstructed sampled view), even if the reconstruction provides a reconstructed image of a view that is downgraded / different from the original image of the view, whose original data has been sampled and subsequently encoded. Indeed, using this synthesis of one or more images reconstructed using conventional decoders (e.g., HEVC, MVC-HEVC, 3D-HEVC), the original regions corresponding to the filtered regions of the image of the view can be retrieved in these one or more conventionally reconstructed images.
[0045] According to another specific embodiment:
[0046] - These processed data of the image in the view are data from at least one region of the image in the view that has undergone filtering.
[0047] The descriptive information of the image processing includes at least one piece of information regarding the location of the at least one filtered region in the image of the view.
[0048] This process facilitates the removal of image data from the view that is deemed unnecessary to encode, which, in order to optimize the reduction in the rate of encoded data, is advantageously formed solely from filtered image data.
[0049] Subsequent reconstruction of such an image of a view filtered in this manner does not impair the effectiveness of the synthesis of intermediate images (which will use such an image of the reconstructed filtered view), even if the reconstruction provides a reconstructed image of a view that is downgraded / different from the original image of the view, whose original data has been filtered and subsequently encoded. Indeed, using such synthesis of one or more images reconstructed using conventional decoders (e.g., HEVC, MVC-HEVC, 3D-HEVC), the original region can be retrieved in the intermediate view through the filtered regions of the image of the view and the conventionally reconstructed image.
[0050] According to another specific embodiment:
[0051] - These processed data for the image of this view are the pixels of the image of this view, which correspond to the occlusion detected using the image of another view among multiple views.
[0052] - The descriptive information of the image processing includes indicators of these pixels of the image found in another view of the image.
[0053] Similar to the foregoing embodiments, this process facilitates the removal of image data of the view that is deemed unnecessary to encode. In order to optimize the reduction in the rate of data encoding, this encoded data is advantageously formed solely from the pixels of the image of the view, the absence of which has been detected in another image of the current view among the plurality of views.
[0054] Subsequent reconstruction of this image of the view does not compromise the effectiveness of the synthesis of intermediate images (which will use this reconstructed image of the view), even if the reconstruction provides a reconstructed image of a view that is downgraded / different from the original image of the view, encoding only its occluded regions. Indeed, using this synthesis of one or more images reconstructed using conventional decoders (e.g., HEVC, MVC-HEVC, 3D-HEVC), the original regions can be retrieved in intermediate views from the image of the current view and the conventionally reconstructed image.
[0055] According to another specific embodiment:
[0056] - The processed data of the image of the encoded / decoded view is calculated from pixels in the following way:
[0057] -Based on this raw data of the image in this view,
[0058] - Based on the original data of at least one other view image encoded / decoded using the first encoding / decoding method,
[0059] - and may encode / decode the processed data of the image of at least one other view using the second encoding / decoding method, based on the raw data of the image of at least one other view.
[0060] -The description information of the image processing includes:
[0061] - An indicator of the calculated pixels of the view image.
[0062] -Information regarding the location of the raw data of pixels used to calculate the view image in at least one other view image, which has been encoded / decoded using the first encoding / decoding method.
[0063] -And possibly, information about the location of the raw data of pixels used to calculate the view image in the image of at least one other view has been encoded / decoded from the processed data of that at least one other view.
[0064] According to another specific embodiment, processed data of an image from a first view and processed data of an image from at least one second view are combined into a single image.
[0065] Accordingly, relative to the above embodiments, the processed data of the view image obtained according to the second decoding method includes the processed data of the first view image and the processed data of at least one second view image.
[0066] According to a specific embodiment:
[0067] -The encoded / decoded data of the image in the view is image-type data.
[0068] - The encoding / decoding description information for image processing is data of image type and / or text type.
[0069] The present invention also relates to an apparatus for encoding an image of a view forming part of a plurality of views, the plurality of views simultaneously representing a 3D scene from different viewpoints or positions, the encoding apparatus including a processor configured to perform the following operations at the present time:
[0070] - Select either a first encoding method or a second encoding method to encode the data of the image in the view.
[0071] - Generate a data signal containing information indicating whether the first encoding method or the second encoding method has been selected.
[0072] - If the first encoding method is selected, the raw data of the image of the view is encoded, and the first encoding method provides the encoded raw data.
[0073] -If the second encoding method is selected:
[0074] The second encoding method encodes the processed data of the image of the view, the processed data corresponding to at least one remaining region of the image of the view, the remaining region being obtained by applying cropping to the original data of the image of the view, and providing at least one encoded remaining region.
[0075] • Encode the descriptive information of the cropping, which is information about the location of the remaining region in the image of the view.
[0076] -The generated data signal further includes:
[0077] • If the first encoding method has been selected, then the original data of the encoded image of the view,
[0078] • If the second encoding method has been selected, then the remaining area of the image of the view and the description information of the encoded cropped area.
[0079] This encoding device is specifically capable of implementing the aforementioned encoding method.
[0080] The present invention also relates to an apparatus for decoding data signals of images representing a portion of a plurality of views, which simultaneously represent a 3D scene from different viewpoints or positions, the decoding apparatus including a processor configured to perform the following operations at the present time:
[0081] - In this data signal, information indicating whether the image of the view will be decoded according to a first decoding method or a second decoding method is read.
[0082] -If it is the first decoding method:
[0083] Then, the encoded data associated with the image of the view is read from the data signal.
[0084] • Reconstruct an image of the view based on the read encoded data, the reconstructed image of the view containing the original data of the image of the view.
[0085] -If it is the second decoding method:
[0086] Then read from this data signal:
[0087] - Encoded data associated with the image of the view, the encoded data corresponding to at least one remaining region of the image of the view that has already been encoded, the remaining region being obtained by applying cropping to the original data of the current image of the view.
[0088] - The cropping description information, the encoded description information being information about the location of the remaining region within the image of the view.
[0089] • The image of the view is reconstructed based on the encoded remaining region and the description information of the cropping.
[0090] This decoding device is specifically capable of implementing the aforementioned decoding method.
[0091] The present invention also relates to a data signal comprising data encoded according to the above-described encoding method.
[0092] The present invention also relates to a computer program comprising instructions for implementing the decoding method or encoding method according to any one of the above specific embodiments when the program is executed by a processor.
[0093] This program can be used in any programming language and can be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form or in any other desired form.
[0094] The present invention also relates to a computer-readable recording medium or information medium comprising computer program instructions as mentioned above.
[0095] The recording medium can be any entity or device capable of storing programs. For example, the medium can include storage devices such as ROMs, such as CD ROMs or microelectronic circuit ROMs, or other magnetic recording devices such as USB keys or hard drives.
[0096] Furthermore, the recording medium can be a transmittable medium (such as an electrical or optical signal) that can be transmitted via cable or optical fiber, radio, or other means. Specifically, the program according to the invention can be downloaded from an internet-type network.
[0097] Alternatively, the storage medium may be an integrated circuit into which the program is incorporated, the circuit being adapted to perform or to perform the aforementioned encoding or decoding methods. Attached Figure Description
[0098] Other features and advantages will become more apparent from reading the following preferred embodiments, given only by way of illustrative and non-limiting examples and described below with reference to the accompanying drawings, in which:
[0099] - Figure 1 The main actions performed by the encoding method according to an embodiment of the present invention are shown.
[0100] - Figure 2A It demonstrates the ability to implement Figure 1 The first type of data signal generated after the encoding method,
[0101] - Figure 2B It demonstrates the ability to implement Figure 1 The second type of data signal generated after the encoding method,
[0102] - Figure 2C It demonstrates the ability to implement Figure 1 The third type of data signal generated after the encoding method,
[0103] - Figure 3A A first embodiment of a method for encoding all images of a view available at the current time is shown.
[0104] - Figure 3B A second embodiment of a method for encoding all images of a view available at the current time is shown.
[0105] - Figures 4A to 4E Examples of processing applied to a view image according to the first embodiment are shown respectively.
[0106] - Figures 5A to 5D Each illustrates an example of processing applied to a view image according to the second embodiment.
[0107] - Figure 6 An example of processing applied to a view image according to a third embodiment is shown.
[0108] - Figure 7 An example of processing applied to a view image according to the fourth embodiment is shown.
[0109] - Figure 8 An example of processing applied to a view image according to the fifth embodiment is shown.
[0110] - Figure 9An example of processing applied to a view image according to the sixth embodiment is shown.
[0111] - Figure 10 Implementation shown Figure 1 The encoding device of the encoding method,
[0112] - Figure 11 The main actions performed by the decoding method according to an embodiment of the present invention are shown.
[0113] - Figure 12A A first embodiment of a method for decoding all images of a view available at the current time is shown.
[0114] - Figure 12B A second embodiment of a method for decoding all images of a view available at the current time is shown.
[0115] - Figure 13 Implementation shown Figure 11 Decoding devices for decoding methods,
[0116] - Figure 14 An embodiment of a composite view image is shown, in which an embodiment is used according to Figure 11 The view image reconstructed by the decoding method
[0117] - Figures 15A to 15D Examples of processing applied to the view image after reconstruction according to the first embodiment are shown respectively.
[0118] - Figure 16 An example of the processing applied to the view image after reconstruction, according to the second embodiment, is shown.
[0119] - Figure 17 An example of processing applied to the view image after reconstruction, according to the third embodiment, is shown.
[0120] - Figure 18 An example of the processing applied to the view image after reconstruction, according to the fourth embodiment, is shown. Detailed Implementation
[0121] This invention primarily proposes a scheme for encoding current images corresponding to multiple views, which represent a 3D scene at a given location or viewpoint at the current time, wherein two encoding techniques are available:
[0122] - A first encoding technique, according to which at least one current image of the view is encoded using a conventional encoding mode, such as HEVC, MV-HEVC, or 3D-HEVC.
[0123] - A second innovative coding technique, according to which processing data of at least one current image of a view is encoded using conventional coding patterns of the type described above and / or any other suitable coding pattern, in order to significantly reduce the signal transmission cost of the encoded data of the image (due to processing performed prior to the encoding step), which is obtained by applying processing to the original data of the image using specific image processing.
[0124] Accordingly, this invention proposes a decoding scheme that allows for the combination of two decoding techniques:
[0125] - A first decoding technique, according to which at least one current image of the encoded view is reconstructed using a conventional decoding mode, such as HEVC, MV-HEVC, 3D-HEVC, and corresponding to a conventional encoding mode used in encoding and transmitted to the decoder, in order to obtain at least one reconstructed image of the view of very high quality.
[0126] - A second innovative decoding technique, according to which the encoded processing data of at least one image of the view is decoded using a decoding mode corresponding to the encoding mode transmitted to the decoder, namely a conventional encoding mode and / or another suitable encoding mode, to obtain processed image data and descriptive information of the image processing, the obtained processed data originating from the image processing. Therefore, unlike the image data decoded according to the first decoding technique, the processed data obtained for the image during decoding does not correspond to its original data.
[0127] The image of the view reconstructed based on the processed image data and image processing description information obtained from this decoding is different from the original image of the view; that is, its original data is subsequently encoded before processing. However, this reconstructed image of the view constitutes the image of the view, which, when used in conjunction with images of other views reconstructed according to the first conventional decoding technique, will enable the synthesis of images of intermediate views in a particularly efficient and effective manner.
[0128] 6. Exemplary Encoding Scheme Implementation
[0129] The following describes a method for encoding 360°, 180°, or other omnidirectional video, which can use any type of multi-view video encoder, such as those conforming to 3D-HEVC or MV-HEVC standards.
[0130] refer to Figure 1 This encoding method is applied to forming multiple views V1...V N A portion of the current image of the view, which represents a 3D scene based on multiple viewpoints or multiple positions / orientations.
[0131] Based on a common example, in the case of generating video (e.g., 360° video) using three omnidirectional cameras:
[0132] - For example, the first omnidirectional camera can be placed in the center of the 3D scene with a field of view of 360° × 180°.
[0133] - For example, a second omnidirectional camera can be placed on the left side of the 3D scene, with a field of view of 360° × 180°.
[0134] - For example, a third omnidirectional camera can be placed on the right side of the 3D scene with a field of view of 360°×180°.
[0135] According to another, more atypical example, in the case of using three omnidirectional cameras to generate α° video (where 0° < α ≤ 360°):
[0136] - For example, the first omnidirectional camera can be placed in the center of the 3D scene with a field of view of 360° × 180°.
[0137] - For example, a second omnidirectional camera can be placed on the left side of the 3D scene, with a field of view of 270° × 135°.
[0138] - For example, a third omnidirectional camera can be placed on the right side of the 3D scene with a field of view of 180°×90°.
[0139] Of course, other configurations are also possible.
[0140] At least two of the multiple views can represent the 3D scene from the same or different perspectives.
[0141] The encoding method according to the present invention includes encoding the following at the current time:
[0142] -Image IV1 of view V1,
[0143] -Image IV2 of view V2,
[0144] -……,
[0145] -View V k Image IV k ,
[0146] -……,
[0147] -View V N Image IV N ,
[0148] The image of the view in question can be either a texture image or a depth image. For example, image IV represents the image of the view in question. kThe original data (d1) contains a quantity Q (Q≥1). k ……dQ k For example, Q pixels.
[0149] For view V k At least one image IV k The encoding method then includes the following to be encoded:
[0150] In C1, select the image IV. k The first encoding method is MC1 or the second encoding method is MC2.
[0151] If the first encoding method MC1 is selected, then in C10, for example, the information flag_proc is encoded on bits set to 0 to indicate that encoding method MC1 has been selected.
[0152] In C11a, conventional encoders, such as those conforming to HEVC, MV-HEVC, and 3D-HEVC standards, are used to process image IVs. k Q raw data (pixels) d1 k ……dQ k Perform encoding. After completing encoding C11a, obtain view V. k Encoded image IVC k Subsequently, the encoded image IVC k raw data dc1 containing Q codes k dc2 k ..., dcQ k .
[0153] In C12a, data signal F1 is generated. k .like Figure 2A As shown, data signal F1 k It contains information related to the selection of the first encoding method MC1, flag_proc=0, and the original encoded data dc1. k dc2 k ……dcQ k .
[0154] If the second encoding method MC2 is selected, then in C10, for example, the information flag_proc is encoded on the bit set to 1 to indicate that the encoding method MC2 has been selected.
[0155] In C11b, the encoding method MC2 is applied to the image IV before the encoding step. k The data obtained from the processing of DT k .
[0156] This type of data DT k include:
[0157] - Corresponding to image IV k The raw data of an image type, including all or some of the original data (pixels), has been processed using specific image processing techniques prior to the encoding step. Various detailed examples of this data will be further described in the description.
[0158] - Applied to image IV prior to encoding step C11b k The image processing description information, which may be text and / or image type.
[0159] After completing the encoding of C11b, the encoded processed data DTC is obtained. k They represent the encoded, processed images (IVTC). k .
[0160] Therefore, the processed data DT k Not corresponding to image IV k The original data.
[0161] For example, these processed data DT k This corresponds to an image whose resolution is higher or lower than that of the image IV before processing. k The resolution. Therefore, the processed image IV k It can be, for example, larger because it is obtained based on images from other views, or conversely, smaller because it is due to the image IV. k It is obtained by deleting one or more original pixels.
[0162] According to another example, these processed data DT k Corresponding to an image, the processed image IV k The representation format (YUV, RGB, etc.) and the image IV before processing k The original format differs from the number of bits used to represent pixels (16-bit, 10-bit, 8-bit, etc.).
[0163] According to yet another example, these processed data DT k Corresponding to the image IV relative to the image before processing k The original texture or color component is downgraded to a lower color or texture component.
[0164] According to yet another example, these processed data DT k Corresponding to the image IV before processing k A specific representation of the original content, such as an image IV. k The original content is represented after filtering.
[0165] Post-processed data DTk In the case of image data only, that is, in the form of a pixel grid, encoding method MC2 can be implemented by an encoder similar to the encoder implementing the first encoding method MC1. It can be a lossy or lossless encoder. After processing the data DT... k Unlike image data, data such as text, or data including both image data and other types of data, can be encoded using the MC2 encoding method as follows:
[0166] - Use a lossless encoder specifically for encoding text-type data.
[0167] - Image data is encoded specifically using a lossy or lossless encoder, which may be the same as or different from the encoder that implements the first encoding method MC1.
[0168] In C12b, data signal F2 is generated. k .like Figure 2B As shown, data signal F2 k Includes information related to the selection of the second encoding method MC2, flag_proc=1, and the encoded processed data DTC. k This is assuming that all of these data are image data.
[0169] As an alternative, in C12c, the data DTC is processed after encoding. k In the case of both image data and text data, two signals F3 are generated. k and F'3 k .
[0170] like Figure 2C As shown:
[0171] - Data signal F'3 contains information related to the selection of the second encoding method MC2, flag_proc=1, and the processed data DTC of the image type encoding. k ,
[0172] -Data signal F'3 k DTC (Data Transmission Controlled Transcript) contains encoded text data. k .
[0173] Then, for each image IV1, IV2...IV of the N available views to be encoded, N Implement the encoding method just described above, only for some of these images, or perhaps only for image IVs. k For example, k=1.
[0174] according to Figure 3A and Figure 3BThe two exemplary embodiments shown, for example, assume that among the N images IV1……IV to be encoded N :
[0175] - n first images IV1……IV are encoded using the first encoding technique MC1 n : The n first views V1 to V n are called the main views because once reconstructed, the images of the n main views will contain all their original data and are suitable for use with one or more of the N - n other views to synthesize an image of any view desired by the user.
[0176] - Before encoding the N - n other images IV n+1 ……IV N using the second encoding method MC2, these images are processed: These N - n other images to be processed belong to the so-called additional views.
[0177] If n = 0, then all images IV1……IV of all views are processed N . If n = N - 1, then the image of a single view among the N views is processed, for example, the image of the first view.
[0178] After completing the processing of the N - n other images IV n+1 ……IV N , M - n processed data are obtained. If M = N, there are as many processed data as there are views to be processed. If M < N, at least one of the N - n views has been deleted during processing. In this case, after completing the encoding C11b according to the second encoding method MC2, the processed data related to the image of the deleted view is encoded using a lossless encoder, which only includes information indicating the absence of the image and the view. Subsequently, for the image deleted by processing, using the second encoding method MC2, an encoder of the type such as HEVC, MVC - HEVC, 3D - HEVC, etc., will not encode the pixels. Figure 1
[0179] In the Figure 3A example, in C11a, the n first images IV1……IV are encoded independently of each other using a conventional HEVC - type encoder. After completing the encoding C11a, n encoded images IVC1……IVC n are obtained respectively. In C12a, n data signals F11……F1 n are generated respectively. In C13a, these n data signals F11……F1 n are cascaded to generate the data signal F1. n
[0180] As shown by the dashed arrow, it is possible to process Nn other images IV. n+1 ...IV N Images IV1...IV were used during the period n .
[0181] Still referencing Figure 3A In C11b, if all Mn processed data are image types, then a regular encoder of HEVC type is used to process the Mn processed data DT. n+1 ……DT M The data is encoded independently of each other. Alternatively, if the data is both image and text, the processed image data is encoded independently using a regular HEVC encoder, while the processed text data is encoded using a lossless encoder. After encoding C11b, Nn encoded processed data DTCs are obtained. n+1…… DTC N , respectively, as IV of Nn images n+1 ...IV N The associated Nn encoded data. In C12b, when Mn processed data are all image types, DTCs containing Nn encoded data are generated respectively. n+1 ...DTC N Nn data signals F2 n+1 ...F2 N In C12c, when Mn processed data are both image and text types:
[0182] - Generate Nn data signals F3, each containing Nn encoded processed data of the image type. n+1 ...F3 N ,as well as
[0183] - Generate Nn data signals F'3, each containing Nn encoded text data. n+1 ...F'3 N .
[0184] In C13c, the data signal F3 n+1 ...F3 N and F'3 n+1 ...F'3 N Cascaded to generate data signal F3.
[0185] In C14, cascading signals F1 and F2, or cascading signals F1 and F3, provides a data signal F that can be decoded by a decoding method, which will be further described in the specification.
[0186] exist Figure 3B In the example, in C11a, a conventional encoder of type MV-HEVC or 3D-HEVC is used to simultaneously process n first images IV1...IV1. n Encoding is performed. After encoding C11a, n encoded images IVC1...IVC are obtained respectively. n In C12a, a single signal F1 is generated, which contains the raw data of the encoding associated with each of the n encoded images.
[0187] As shown by the dashed arrow, it is possible to process Nn other images IV. n+1 ...IV N The process uses the first image IV1...IV n .
[0188] Still referencing Figure 3B In C11b, if these Mn processed data are all image types, then a regular encoder of type MV-HEVC or 3D-HEVC is used to process the Mn processed data DT of the image type. n+1 ……DT M Simultaneous encoding is performed, or if the data is both image and text, the processed image data is encoded simultaneously using a conventional encoder of type MV-HEVC or 3D-HEVC, while the processed text data is encoded using a lossless encoder. After completing the encoding of C11b, Nn encoded processed data DTCs are obtained respectively. n+1…… DTC N , respectively, as IV of the Nn images that have been processed. n+1 ...IV N The associated Nn encoded processed data. In C12b, when Mn processed data are all image types, Nn encoded processed data DTCs containing image types are generated. n+1 ...DTC N A single signal F2. In C12c, in the case where Mn processed data are both image and text types:
[0189] - Generate signal F3, which contains processed data encoded in the image type.
[0190] - Generate signal F'3, which contains processed data encoded in text type.
[0191] In C14, cascading signals F1 and F2, or cascading signals F1, F3, and F'3, provides a data signal F that can be decoded by a decoding method, which will be further described in the specification.
[0192] Of course, other combinations of encoding methods are possible.
[0193] according to Figure 3A One possible variation is that the encoder implementing encoding method MC1 can be an HEVC type encoder, and the encoder implementing encoding method MC2 can be an MV-HEVC or 3D-HEVC type encoder, or include an MV-HEVC or 3D-HEVC type encoder and a lossy encoder.
[0194] according to Figure 3B One possible variation is that the encoder implementing encoding method MC1 can be an MV-HEVC or 3D-HEVC type encoder, and the encoder implementing encoding method MC2 can be an HEVC type encoder, or include an HEVC type encoder and a lossy encoder.
[0195] Now refer to Figures 4A to 4E The encoding steps C11b according to the second encoding method MC2 are given. Figure 1 Previously applied to image IV k The description of the first embodiment of processing the raw data.
[0196] In the examples shown in these figures, the image IV is applied. k The processing of the raw data involves cropping one or more regions of the image in the horizontal or vertical direction, or simultaneously in both directions.
[0197] exist Figure 4A In the example, image IV k The left-hand boundary B1 and right-hand boundary B2 are cropped, which means that the image IV formed by each of boundaries B1 and B2 is deleted. k The pixels of the rectangular area.
[0198] exist Figure 4B In the example, image IV k The top boundary B3 and bottom boundary B4 are cropped, which means that the image IV formed by each of boundaries B3 and B4 is deleted. k The pixels of the rectangular area.
[0199] exist Figure 4C In the example, cropping is applied vertically to the image IV. k The rectangular region Z1 in the middle.
[0200] exist Figure 4D In the example, cropping is applied horizontally to the image IV. k The rectangular region Z2 in the middle.
[0201] exist Figure 4E In the example, cropping is applied in both the horizontal and vertical directions to the image IV. k The rectangular region Z3 in the text.
[0202] The processed data DT to be encoded k include:
[0203] -Image IV k The remaining region Z R The pixels, these pixels in the cropping ( Figure 4A , Figure 4B , Figure 4E (This was not deleted afterward, or image IV) k The remaining region Z1 R and Z2 R ( Figure 4C , Figure 4D (Pixels that were not deleted after cropping)
[0204] - Describe the information about the clipping applied.
[0205] In example Figure 4A , Figure 4B , Figure 4E In this case, the information describing the applied clipping is a text type and includes:
[0206] -Located in image IV k The remaining region Z R The coordinates of the top and furthest pixels in the image are those of the leftmost pixel.
[0207] -Located in image IV k The remaining region Z R The coordinates of the bottom and farthest pixels in the image are those of the rightmost pixel.
[0208] In example Figure 4C and Figure 4D In this case, the information describing the applied clipping includes:
[0209] -Located in image IV k The remaining region Z1 R The coordinates of the top and furthest pixels in the image are those of the leftmost pixel.
[0210] -Located in image IV k The remaining region Z1 R The coordinates of the bottom and furthest pixels in the image are those of the rightmost pixel.
[0211] -Located in image IVk The remaining region Z2 R The coordinates of the top and furthest pixels in the image are those of the leftmost pixel.
[0212] -Located in image IV k The remaining region Z2 R The coordinates of the bottom and farthest pixels in the image are those of the rightmost pixel.
[0213] Subsequently in C11b ( Figure 1 In this context, encoders of types such as HEVC, 3D-HEVC, and MV-HEVC are used to process data located in the remaining region Z. R (respectively Z1) R and Z2 R The coordinates of the top, farthest left pixel in the region, and the coordinates of the pixels located in the remaining Z region. R (respectively Z1) R and Z2 R The original data (pixels) of a rectangular region defined by the coordinates of the bottom, furthest, and rightmost pixels in C11b are encoded. Figure 1 In this process, a portion of the descriptive information of the applied clipping is encoded by a lossless encoder.
[0214] As a variation, the information describing the applied cropping includes the number of rows and / or columns of pixels to be removed, and the number of these rows and / or columns in the image IV. k The position in the middle.
[0215] According to one embodiment, the amount of data to be deleted by cropping is fixed. For example, it may be decided to systematically delete × rows and / or Y columns from the image of the view in question. In this case, the description information only contains information about whether cropping is performed for each view.
[0216] According to another embodiment, the amount of data to be deleted by cropping is in view V. k Image IV k The image is variable between the view and another available view.
[0217] For example, the amount of data to be removed by cropping may also depend on the captured image IV. k The position of the camera in the 3D scene. Therefore, for example, if the image of another view among the N views has been captured by an image IV in the 3D scene... k If the camera is positioned / oriented differently, then capturing the image will, for example, use a different camera than the one used for the image IV. k The amount of data to be deleted: A certain amount of data needs to be deleted.
[0218] The application of cropping can further depend on the image IV. kThe encoding time. At the current time, for example, it can be determined that the image IV... k Apply cropping, and decide whether to exclude the image IV from time before or after the current time. k Apply this cutting or any processing in this regard.
[0219] Finally, at the current time, cropping can be applied to images in one or more views. View V k Image IV k The cropped area in the image can be the same as or different from the cropped area of another view that is to be encoded at the current time.
[0220] Now refer to Figures 5A to 5D The encoding steps C11b according to the second encoding method MC2 are given. Figure 1 Previously applied to image IV k The second embodiment of the processing of raw data is described.
[0221] In the examples shown in these figures, the image IV is applied. k The processing of the raw data involves downsampling one or more regions of the image in the horizontal or vertical direction.
[0222] exist Figure 5A In the example, downsampling is applied to image IV in the vertical direction. k Region Z4.
[0223] exist Figure 5B In the example, downsampling is applied to the image IV in the horizontal direction. k Region Z5.
[0224] exist Figure 5C In the example, downsampling is applied to the entire image IV. k .
[0225] exist Figure 5D In the example, downsampling is applied to image IV in both the horizontal and vertical directions. k Region Z6.
[0226] The processed data DT to be encoded k include:
[0227] - Downsampled image data (pixels),
[0228] - Describe the information applied to the downsampling, such as:
[0229] - The downsampling factor used
[0230] - The downsampling direction used
[0231] -exist Figure 5A , Figure 5B , Figure 5D In the case of filtering regions Z4, Z5, and Z6 in image IV k The position in the middle, or,
[0232] -exist Figure 5C In this case, the entire region defining the image is located in image IV. k The coordinates of the top, farthest left pixel, and the pixel located in image IV. k The coordinates of the bottom and farthest pixels in the image are those of the rightmost pixel.
[0233] Subsequently, in C11b ( Figure 1 In C11b, downsampled image data (pixels) is encoded using encoders of types such as HEVC, 3D-HEVC, and MV-HEVC. Figure 1 In this process, a portion of the description information of the applied downsampled data is encoded by a lossless encoder.
[0234] The value of the downsampling factor can be fixed or depend on the captured image IV. k The position of the camera in the 3D scene. Therefore, for example, if the image of another view among the N views has been captured by an image IV in the 3D scene... k If the camera is captured at a different position / orientation, then a different sampling factor will be used, for example.
[0235] The application of downsampling can further depend on the view V k The time for encoding the image. At the current time, for example, it can be decided to encode the image IV. k Applying downsampling, and considering times before or after the current time, can determine whether to exclude view V. k The image is subjected to this downsampling or any processing in that regard.
[0236] Finally, at the current time, downsampling can be applied to the image of one or more views. Image IV k The downsampled region in the image can be the same as or different from the downsampled region of another view to be encoded at the current time.
[0237] Now refer to Figure 6 The encoding steps C11b according to the second encoding method MC2 are given. Figure 1 Previously applied to image IV k The third embodiment of processing the raw data is described.
[0238] exist Figure 6 In the example shown, it is applied to image IV kThe processing of the raw data involves detecting contours by filtering the image. In image IV... k For example, there are two contours, ED1 and ED2. In a manner known per se, this filtering includes, for example, the following:
[0239] - Apply the contour detection filter to contours ED1 and ED2.
[0240] - Apply extensions to contours ED1 and ED2 to expand the area surrounding each contour ED1 and ED2, such an area... Figure 6 The middle part is indicated by a shading line.
[0241] - Extract all raw data from image IV k The portions of the original data that do not form the shaded area are deleted and are therefore considered not to require encoding.
[0242] Subsequently, in the case of such filtering, the processed data DT to be encoded... k include:
[0243] - The original pixels contained in the shaded area,
[0244] - Information describing the applied filter, such as predefined values for each pixel not included in the shadow area, for example, represented by the predefined value YUV=000.
[0245] Subsequently, in C11b ( Figure 1 In C11b, image data (pixels) corresponding to the shadow areas are encoded using encoders of types such as HEVC, 3D-HEVC, and MV-HEVC. Figure 1 In this process, a portion of the descriptive information of the applied filter is encoded by a lossless encoder.
[0246] Just now about image IV k The described filtering can be applied to one or more images of other views in the N views, and the transition from one image of a view to another image of a view can be a region of these one or more different images.
[0247] Now refer to Figure 7 The encoding steps C11b according to the second encoding method MC2 are given. Figure 1 Previously applied to image IV k The fourth embodiment of processing the raw data is described.
[0248] exist Figure 7 In the example shown, it is applied to image IV k The processing of the raw data involves using another view V out of N views. p At least one image IV p(1≤p≤N) to detect image IV k At least one region Z OC The obstruction.
[0249] In a manner known per se, this occlusion detection involves using, for example, disparity estimation based on image IVs. p Finding Image IV k Region Z OC Then, for example, mathematical morphology algorithms are used to expand the occluded region Z. OC Therefore, the extended region Z OC exist Figure 7 The image is represented by a shaded line. Image IV... k All original data is deleted; this original data does not form the shaded area Z. OC This part is considered to require no encoding.
[0250] Subsequently, under this occlusion removal condition, the processed data DT to be encoded... k include:
[0251] - The original pixels contained in the shaded area,
[0252] - Describes information about the applied occlusion removal, such as predefined values for each pixel not included in the shadow area, for example, represented by the predefined value YUV=000.
[0253] Subsequently, in C11b ( Figure 1 In C11b, image data (pixels) corresponding to the shadow areas are encoded using encoders of types such as HEVC, 3D-HEVC, and MV-HEVC. Figure 1 In the process, a portion of the descriptive information for the applied occlusion removal is encoded by a lossless encoder.
[0254] Now refer to Figure 8 The encoding steps C11b according to the second encoding method MC2 are given. Figure 1 Previously applied to image IV k The fifth embodiment of processing the raw data is described.
[0255] exist Figure 8 In the example shown, it is applied to image IV k The processing of raw data includes calculating pixels:
[0256] -Based on image IV k The original pixels,
[0257] -Based on C11a( Figure 1 Image IVs of one or more other views encoded using the first encoding method MC1. jThe original pixels of (1≤j≤n)
[0258] - and may be based on an image IV from at least one other view. l The original pixels (n+1≤l≤N) in C11b( Figure 1 The second encoding method MC2 is used to encode the processed pixels of the image.
[0259] Subsequently, under this computational condition, the processed data DT to be encoded... k include:
[0260] - The calculated image IV of the view k The pixel indicator,
[0261] -Regarding the image IV already used for calculation k The original pixels of the pixels in image IV j Information about the location in the middle,
[0262] - and possibly regarding the image IV that has been used to calculate the image IV k The original pixels of the pixels in image IV j Information about the location in the middle,
[0263] The aforementioned calculations include, for example, converting the image IV of the view. k The original pixels and possible image IV l The original pixels from image IV j Subtract the original pixels.
[0264] Now refer to Figure 9 The encoding steps C11b according to the second encoding method MC2 are given. Figure 1 Previously applied to image IV k The sixth embodiment of processing the raw data is described.
[0265] exist Figure 9 In the example shown, it is applied to image IV k The processing of raw data includes:
[0266] -Image Processing IV k The original pixels provide the processed data DT' k ,
[0267] -Process view V S Image IV of (1≤s≤N) s The original pixels provide the processed data DT s ,
[0268] -Transfer image IV k The processed data DT' k and Image IVs DT processed data s Merge into a single image IV one In C11b ( Figure 1 The raw processed data DT obtained from this single image is obtained using the second encoding method MC2. k Encode it.
[0269] Figure 10 A simplified structure of an encoding device COD is shown, which is designed to implement an encoding method according to any of the specific embodiments of the present invention.
[0270] According to a specific embodiment of the present invention, the actions performed by this encoding method are implemented by computer program instructions. For this purpose, the encoding device COD has a conventional computer architecture and specifically includes a memory MEM_C, a processing unit UT_C equipped with, for example, a processor PROC_C, and driven by a computer program PG_C stored in the memory MEM_C. The computer program PG includes instructions for implementing actions such as the encoding method described above when the program is executed by the processor PROC_C.
[0271] During initialization, the code instructions of the computer program PG_C are loaded into RAM memory (not shown) before being executed by the processor PROC_C. The processor PROC_C of the processing unit UT_C specifically implements the operations of the encoding method described above according to the instructions of the computer program PG_C.
[0272] 7. Exemplary Decoding Scheme Implementation
[0273] The following describes a method for decoding 360°, 180°, or other omnidirectional video. This method can use any type of multi-view video decoder, such as those conforming to the 3D-HEVC or MV-HEVC standards.
[0274] refer to Figure 11 This decoding method is applied to represent the formation of the multiple views V1...V N A portion of the view's current image data signal.
[0275] The decoding method according to the present invention includes decoding the following:
[0276] - A data signal representing the encoded data associated with image IV1 of view V1.
[0277] - A data signal representing the encoded data associated with image IV2 of view V2.
[0278] -……,
[0279] -Representative and View Vk Image IV k The associated encoded data signal,
[0280] -……,
[0281] -Representative and View V N Image IV N The associated encoded data signal.
[0282] The image of the view in question that is to be reconstructed using the aforementioned decoding method can also be a texture image or a depth image.
[0283] The decoding method includes the following: for the view V representing the one to be reconstructed k At least one image IV k Data signal F1 k F2 k Or F3 k and F'3 k :
[0284] In D1, as shown in... Figure 2A , Figure 2B and Figure 2C The data signal F1 shown in the figure k F2 k Or F3 k and F'3 k Read the information flag_proc, which indicates whether the first encoding method MC1 or the second encoding method MC2 was used to encode the image IV. k Encode it.
[0285] If it is signal F1 k If so, the information flag_proc will be 0.
[0286] In D11a, in data signal F1 k Image reading and encoding in IVC k Associated encoded data dc1 k dc2 k ……dcQ k .
[0287] In D12a, the corresponding method is used. Figure 1 The decoding method MD1, which is applied to the encoding method MC1 in C11a, is based on the encoded data dc1 read in D11a. k dc2 k ……dcQ k To reconstruct images IVD k Therefore, conventional decoders conforming to standards such as HEVC, MVC-HEVC, and 3D-HEVC are used to reconstruct the image IV.k .
[0288] After completing the decoding of D12a, the image IVD is reconstructed in this way. k Included Figure 1 Image IV encoded in C11a k The original data d1 k d2 k ……dQ k .
[0289] Due to image IVD k With the original image IV k To maintain consistency, it constitutes the main image that is suitable for use, for example, in the context of compositing intermediate views.
[0290] In D1, if it is signal F2 k Or F3 k If so, the determined information flag_proc is 1.
[0291] If it is signal F2 k Then in D11b, in data signal F2 k Reading and such Figure 1 The encoded processed image IVTC obtained from C11b k Associated encoded processed data DTC k .
[0292] Read these encoded processed data DTC k It is only image-type data.
[0293] In D12b, the corresponding method is used. Figure 1 The decoding method MD2, which is applied to the encoding method MC2 in C11b, is based on the encoded data DTC read from D11b. k To reconstruct the processed image IVTD k Therefore, conventional decoders conforming to standards such as HEVC, MVC-HEVC, and 3D-HEVC are used to process the encoded data using DTC. k Decode it.
[0294] After completing the decoding of D12b, the corresponding decoded data DTC k The reconstructed image after IVTD k Included Figure 1 Image IV before encoding in C11b k DT processed data k .
[0295] Reconstructed processed image IVTD kIncludes an IV corresponding to an image that has been processed using specific image processing techniques. k All or some of the original image data (pixels), as shown in Figures 4 to 5. Figure 9 Various detailed examples are described.
[0296] If it is signal F3 k Then, the processed image IVTC, which is read and encoded in D11c, is... k Associated encoded processed data DTC k .
[0297] to this end:
[0298] -In signal F3 k DTC (Digital Transmission Control Center) reads the encoded data of the image type. k ,
[0299] -In signal F′3 k Read the encoded processed data DTC k This data differs from image data; for example, it could be text-type data, or data that includes both image data and other types of data.
[0300] The encoded DTC data is processed in D12b using the MD2 decoding method. k Decoding can be performed as follows:
[0301] - Use a lossy or lossless decoder specifically for decoding image data. This decoder may be the same as or different from the decoder that implements the first decoding method MD1.
[0302] - Use a lossless decoder to specifically decode text-type data.
[0303] After decoding D12b, the following was obtained:
[0304] - Data DTC corresponding to the decoded image type k The reconstructed image after IVTD k ,
[0305] - Corresponds to encoding C11b( Figure 1 Previously applied to image IV k The processed data is text-based and describes the information being processed.
[0306] The reconstructed image IVTD based on the second decoding method MD2 k The image IV is not included before processing and subsequent encoding in C11b. kAll the original data. However, for example in the context of synthesizing intermediate images, in addition to the image of the main view reconstructed using the first decoding method MD1, such a reconstructed image of the view according to the second decoding method MD2 can be used to obtain a synthetic image of the view with high quality.
[0307] Then, at the current time, the available encoded images to be reconstructed, IVC1, IVC2...IVC, can be used. N Each implementation of the decoding method just described above is only for some of these encoded images, or may be limited to image IVs. k For example, k=1.
[0308] according to Figure 12A and Figure 12B The two exemplary embodiments shown, for example, assume that there are N coded images IVC1...IVC to be reconstructed. N middle:
[0309] - Reconstruct n first-coded images IVC1...IVC using the first decoding technique MD1. n In order to obtain the image of each of the n main views,
[0310] - Reconstruct Nn other encoded images IVC using the second decoding method MD2. n+1 ...IVC N In order to obtain the image of each of the Nn other views respectively.
[0311] If n = 0, then the second decoding method MD2 is used to reconstruct the images of the view from 1 to N: IVC1...IVC N If n = N-1, then the image of a single view is reconstructed using the second decoding method, MD2.
[0312] exist Figure 12A In the example, in D100, Figure 3A The data signal F generated in C14 is divided into:
[0313] - Two data signals: in Figure 3A The data signal F1 generated in C13a and in Figure 3A The data signal F2 generated in C13b is
[0314] -Or three data signals: in Figure 3A The data signal F1 generated in C13a and in Figure 3A The data signals F3 and F'3 are generated in C13c.
[0315] In the case of signals F1 and F2, in D110a, data signal F1 is further divided into signals representing views IVC1...IVC respectively. n n data signals F11……F1 of n encoded images n .
[0316] In D11a, there are n data signals F11……F1 n In each of the n encoded images, the original encoded data dc11……dcQ1……dc1 is determined respectively. n ……dcQ n .
[0317] In D12a, a regular HEVC-type decoder is used, based on the images IVD1...IVD read in D11a. n The original data, each encoded separately, are used to reconstruct these images independently.
[0318] Still referencing Figure 12A In D110b, the data signal F2 is further divided into processed data DTCs representing Nn codes. n+1 ...DTC N Nn data signals F2 n+1 ...F2 N .
[0319] In D11b, there are Nn data signals F2 n+1 ...F2 N In each of them, read Nn encoded processed data DTCs respectively. n+1 ...DTC N These data correspond to the Nn images IV to be reconstructed. n+1 ...IV N Each of them.
[0320] In D12b, a regular decoder of HEVC type is used, based on the processed data DTC of Nn encoded data read in D11b. n+1 ...DTC N The processed images are reconstructed independently of each other. The reconstructed processed image IVTD is then obtained. n+1 ...IVTD N .
[0321] In the case of signals F1, F3, and F'3, in D110a, the data signal F1 is further divided into n coded images IVC1...IVC, each representing a different coded image. n n data signals F11……F1 n .
[0322] In D11a, there are n data signals F11……F1 n In each of the n encoded images, read the original encoded data dc11……dcQ1……dc1. n ……dcQ n .
[0323] In D12a, a regular HEVC-type decoder is used, based on the images IVD1...IVD read in D11a. n The original data, each encoded separately, are used to reconstruct these images independently.
[0324] In D110c:
[0325] - The data signal F3 is then divided into Nn coded DTCs, each representing an image type. n+1 ...DTC N Nn data signals F3 n+1 ...F3 N .
[0326] - The data signal F'3 is then divided into Nn encoded DTCs, each representing either text or another type of data. n+1 ...DTC N Nn data signals F'3 n+1 ...F'3 N .
[0327] In D11c:
[0328] -In Nn data signals F3 n+1 ...F3 N In each of them, read Nn encoded processed data DTCs of the image type respectively. n+1 ...DTC N These data correspond to the Nn images IV to be reconstructed. n+1 ...IV N Each of them,
[0329] -In Nn data signals F'3 n+1 ...F'3 N In each of them, read Nn encoded processed data of text or another type, DTC. n+1 ...DTC N These data correspond to the Nn images IVs to be reconstructed. n+1 ...IV N Each of these contains descriptive information about its processing.
[0330] In D12b, a regular decoder of HEVC type is used, based on the processed data DTC of Nn encoded data read in D11b. n+1 ...DTC N Each of the Nn processed images is reconstructed independently. Then, the reconstructed processed image IVTD is obtained. n+1 ...IVTD N .
[0331] Also in D12b, there exists an application for image IV. n+1 ...IV N The reconstruction description information of each of the processes in C11b ( Figure 3A Before encoding these images, the processed data DTC (Data Transmission Controlled Transcription) is read in D11c using Nn encoded text or other types of data, based on a decoder corresponding to the lossless encoder used in the encoding. n+1 ...DTC N .
[0332] exist Figure 12B In the example, in D100, Figure 3B The data signal F generated in C14 is divided into:
[0333] - Two data signals: in Figure 3B The data signal F1 generated in C12a and in Figure 3B The data signal F2 generated in C12b is
[0334] -Or three data signals: in Figure 3B The data signal F1 generated in C12a and in Figure 3B The data signals F3 and F'3 are generated in C12c.
[0335] In the case of signals F1 and F2, in D11a, in data signal F1, the encoded raw data dc11……dcQ1……dc1 is read. n ……dcQ n These data are respectively compared with the images of n encoded views (IVC). 1,i ...IVC n,i Each of them is associated with something else.
[0336] In D12a, a conventional decoder of MV-HEVC or 3D-HEVC type is used, based on the image IVD read in D11a. 1,i IVD n,i The original data, each encoded separately, are used to reconstruct these images simultaneously.
[0337] In D11b, Figure 12BIn the data signal F2, Nn encoded processed data DTCs are read. n+1 ...DTC N These data are respectively compared with the images IV of the Nn views to be reconstructed. n+1 ...IV N Each of them is associated with something else.
[0338] In D12b, a conventional decoder of MV-HEVC or 3D-HEVC type is used, based on the processed data DTC of Nn encoded data read in D11b. n+1 ...DTC N Simultaneously, Nn processed images are reconstructed. Then, the reconstructed processed images IVTD are obtained. n+1 ...IVTD N .
[0339] In the case of signals F1, F3, and F'3, in D11a, in the data signal F1, read the encoded raw data dc11……dcQ1……dc1 n ……dcQ n These data are respectively associated with images IVC1...IVC of n encoded views. n Each of them is associated with something else.
[0340] In D12a, a conventional decoder of MV-HEVC or 3D-HEVC type is used, based on the image IVD1...IVD read in D11a. n The original data, each encoded separately, are used to reconstruct these images simultaneously.
[0341] In D11c, Figure 12B In the data signal F3, Nn encoded processed data DTCs are read. n+1 ...DTC N These data are respectively compared with the Nn images IV to be reconstructed. n+1 ...IV N Each of them is associated with something else.
[0342] In D12b, a conventional decoder of MV-HEVC or 3D-HEVC type is used, based on the processed data DTC of Nn encoded data read in D11c. n+1 ...DTC N Simultaneously, the processed images are reconstructed. Then, the reconstructed processed image IVTD of the view is obtained. n+1 ...IVTD N .
[0343] Also in D12b, Figure 12B In, there exists an application to image IV. n+1...IV N The reconstruction description information of each of the processes in C11b ( Figure 3B Before encoding these images, the processed data DTC (Data Transmission Controlled Transcript) is based on Nn encoded text or other types of data read in D11c and decoded using a decoder corresponding to the lossless encoder used in the encoding. n+1 ...DTC N .
[0344] Of course, other combinations of decoding methods are possible.
[0345] according to Figure 12A One possible variation is that the decoder implementing decoding method MD1 can be an HEVC type decoder, and the decoder implementing decoding method MD2 can be an MV-HEVC or 3D-HEVC type decoder.
[0346] according to Figure 12B One possible variation is that the decoder implementing decoding method MD1 can be an MV-HEVC or 3D-HEVC type decoder, and the decoder implementing decoding method MD2 can be an HEVC type encoder.
[0347] Figure 13 A simplified structure of a decoding device DEC is shown, which is intended to implement a decoding method according to any of the specific embodiments of the present invention.
[0348] According to a specific embodiment of the present invention, the actions performed by the decoding method are implemented by computer program instructions. For this purpose, the decoding device DEC has a conventional computer architecture and specifically includes a memory MEM_D, a processing unit UT_D equipped with, for example, a processor PROC_D, and driven by a computer program PG_D stored in the memory MEM_D. The computer program PG_D includes instructions for implementing actions such as the decoding method described above when the program is executed by the processor PROC_D.
[0349] During initialization, the code instructions of the computer program PG_D are loaded into RAM memory (not shown) before being executed by the processor PROC_D. The processor PROC_D of the processing unit UT_D specifically implements the decoding method described above according to the instructions of the computer program PG_D.
[0350] According to one embodiment, the decoding device DEC is included, for example, in the terminal.
[0351] 8. Exemplary applications of the present invention in image processing
[0352] As explained above, N reconstructed images IVD1...IVDn and IVTD n+1 ...IVTD N It can be used to synthesize images that provide the intermediate view required by the user.
[0353] like Figure 14 As shown, when the user needs to synthesize images of arbitrary views, in S1, n first reconstructed images IVD1...IVD of the view considered as the main view are used. n Transmitted to the image compositing module.
[0354] Nn reconstructed processed images of the view IVTD n+1 ...IVTD N In order to use additional images as views in the synthesized image, it may be necessary to process them in S2 using decoded image processing description information associated with these images respectively.
[0355] After completing process S2, Nn reconstructed images (IVD) of the view are obtained. n+1 IVD N .
[0356] Then, in S3, the Nn reconstructed images are IVD n+1 IVD N Transmitted to the image compositing module.
[0357] In S4, n images of the first reconstructed view, IVD1...IVD, are used. n At least one of the Nn images of the Nn possible reconstructed views IVD n+1 IVD N At least one of them is used to synthesize the image of the view.
[0358] Subsequently, upon completion of synthesis S4, the synthesis view IV is obtained. SY The image.
[0359] It should be noted that n reconstructed images IVD1...IVD n It can also go through processing S2. The viewpoint that the user's UT needs to represent does not correspond to n reconstructed images IVD1...IVD n In the case of images representing one or more viewpoints, this processing S2 may prove necessary. The user UT may, for example, request images representing a 120×90 field of view, with n reconstructed images IVD1…IVD… n Each represents a 360×180 perspective. Figure 14 The dashed arrows in the diagram indicate the processing of reconstructed images IVD1...IVD n This possibility. Additionally, regarding references... Figure 8 and Figure 9The described processing type can be applied to Nn other images IV. n+1 ...IV N The reconstructed image IVD1...IVD is used during processing. n .
[0360] Now refer to Figures 15A to 15C Description of the processed image IVTD applied to reconstruction k The first embodiment of data processing. This processing includes obtaining an image IV of the corresponding view. k At its initial resolution, the image is in Figure 1 It was sampled before being encoded in C11b.
[0361] exist Figure 15A In the example, it is assumed that the processing applied before encoding is as follows: Figure 5A The image IV is shown in the vertical direction. k Downsampling of region Z4.
[0362] IVTD applied to the reconstructed image k The processing includes corresponding to Figure 5A The downsampling applied in the middle will be upsampled to region Z4 so that information describing the applied downsampling can be used to return to image IV. k The initial resolution, for example:
[0363] - The downsampling factor used, which allows the corresponding upsampling factor to be determined.
[0364] - The downsampling direction used, which allows the corresponding upsampling direction to be determined.
[0365] - The downsampled region Z4 in image IV k The position in the middle.
[0366] exist Figure 15B In the example, it is assumed that the processing applied before encoding is as follows: Figure 5B The image IV shown is positioned horizontally. k Downsampling of region Z5.
[0367] IVTD applied to the reconstructed image k The processing includes corresponding to Figure 5B The downsampling applied in the middle will be upsampled to region Z5 so that information describing the applied downsampling can be used to return to image IV. k The initial resolution, for example:
[0368] - The downsampling factor used, which allows the corresponding upsampling factor to be determined.
[0369] - The downsampling direction used, which allows the corresponding upsampling direction to be determined.
[0370] - The downsampled region Z5 in image IV k The position in the middle.
[0371] exist Figure 15C In the example, it is assumed that the processing applied before encoding is as follows: Figure 5C The image shown is for the entire IV. k Downsampling.
[0372] IVTD applied to the reconstructed image k The processing includes corresponding to Figure 5C The downsampling applied in the middle will be upsampling applied to the image IVTD. k All image data, so that information describing the downsampling applied can be used to reconstruct the image IV. k The initial resolution, for example:
[0373] - The downsampling factor used, which allows the corresponding upsampling factor to be determined.
[0374] - The downsampling direction used, which allows the corresponding upsampling direction to be determined.
[0375] exist Figure 15D In the example, it is assumed that the processing applied before encoding is as follows: Figure 5D The image IV is shown in both the horizontal and vertical directions. k Downsampling of region Z6.
[0376] IVTD applied to the reconstructed image k The processing includes corresponding to Figure 5D The downsampling applied in the middle will be upsampled to region Z6 so that information describing the applied downsampling can be used to return to image IV. k The initial resolution, for example:
[0377] - The downsampling factor used, which allows the corresponding upsampling factor to be determined.
[0378] - The downsampling direction used, which allows the corresponding upsampling direction to be determined.
[0379] - The downsampled region Z6 in image IV k The position in the middle.
[0380] refer to Figure 16 The processed image IVTD applied to the reconstruction will now be described. k A second embodiment of data processing. This processing includes restoring the image IV of the view.k One or more contours of the image, which has been filtered before encoding.
[0381] exist Figure 16 In the example, it is assumed that the processing applied before encoding is as follows: Figure 6 The image IV shown k Filtering of the contours ED1 and ED2.
[0382] Subsequently, the image IVTD of the reconstructed post-processed view is applied. k The processing includes using information describing the applied filters to recover the image IV. k The outlines ED1 and ED2, this information such as the predefined value for each unfiltered pixel, specifically the predefined value YUV=000.
[0383] refer to Figure 17 The processed image IVTD applied to the reconstruction will now be described. k The third embodiment of data processing. This processing includes reconstructing the image IV of the view. k The pixels, pre-coded according to Figure 8 The processing implementation calculates these pixels.
[0384] Subsequently, the processed image IVTD was applied to the reconstruction. k The processing includes:
[0385] -Use the already calculated view image IV k The pixel indicator is used to retrieve the view image IV that has been calculated in the encoding. k The pixels, these indicators are read from the data signal.
[0386] -Use information already used to calculate image IV k The pixels of the image IV j Information about the position in (1≤j≤n) is used to retrieve the image IV of at least one other view that has been reconstructed using the first decoding method MD1. j pixels,
[0387] - and may retrieve the image IV of at least one other view. l The pixels (n+1≤l≤N) have been decoded using the second decoding method MD2 after processing the image.
[0388] DT of processed data k Decoding then includes calculating the image IV. k pixels:
[0389] - Image IV based on at least one other view j Pixels of (1≤j≤n)
[0390] - and may be based on an image IV from at least one other view. l (n+1≤l≤N) pixels.
[0391] The aforementioned calculations include, for example, converting the image IV k Pixels and reconstructed image IV j The pixels and possible reconstruction of image IV l The pixels are combined.
[0392] refer to Figure 18 The processed image IVTD applied to the reconstruction will now be described. k The fourth embodiment of data processing. This processing includes reconstructing the image IV of the view. k The pixels, pre-coded according to Figure 9 The processing implementation calculates these pixels.
[0393] First, the processing is applied to the reconstructed image IVD according to the second decoding method MD2. one It subsequently included image-based IVD. one Reconstructed Image IV k pixels:
[0394] -Based on the image IV that has been decoded according to the second decoding method MD2 k The processed data DT' k ,
[0395] -Based on the decoded view V S Image IVD of (1≤s≤N) s DT processed data s .
[0396] 9. Exemplary Specific Applications of the Invention
[0397] Based on the first example, consider capturing six images of a view with a resolution of 4096 × 2048 pixels each using six cameras of type 360°. Apply a depth estimation method to provide six corresponding 360° depth maps.
[0398] Image IV0 of view V0 is encoded using the first encoding method MC1 as usual, while the five other images IV1, IV2, IV3, IV4, and IV5 of views V1, V2, V3, V4, and V5 are cropped before encoding. The processing applied to each of images IV1, IV2, IV3, IV4, and IV5 involves removing a fixed number of columns, e.g., 200, from the right and left sides of each of these images. The number of columns to be removed is selected such that the viewing angle is reduced from 360° to 120°. Similarly, the processing applied to each of images IV1, IV2, IV3, IV4, and IV5 involves deleting a fixed number of rows, e.g., 100, from the top and bottom of each of these images, respectively. The number of rows to be deleted is selected such that the viewing angle is reduced from 180° to 120°.
[0399] The information flag_proc is set to 0 in association with image IV0, and the information flag_proc is set to 1 in association with images IV1, IV2, IV3, IV4, and IV5.
[0400] The image IV0 of the view is encoded using an HEVC encoder. A single data signal F10 is generated, which contains the raw data of the encoded image IV0, with the information flag_proc = 0.
[0401] The HEVC encoder encodes the data of the remaining regions after cropping each of images IV1, IV2, IV3, IV4, and IV5. Five data signals F21, F22, F23, F24, and F25 are generated in association with the information flag_proc=1. These five data signals contain the encoded data of the remaining regions after cropping each of images IV1, IV2, IV3, IV4, and IV5, respectively. The data signals F10, F21, F22, F23, F24, and F25 are concatenated and then transmitted to the decoder.
[0402] The five data signals F21, F22, F23, F24, and F25 can additionally include the coordinates of the clipping region in the following way:
[0403] IV1, IV2, IV3, IV4, IV5: flag_proc = 1, point top_left(h,v) = (0+200, 0+100), point bot_right(h,v) = (4096-200, 2048-100), where "h" represents horizontal and "v" represents vertical.
[0404] At the decoder, the information flag_proc is read.
[0405] If flag_proc=0, the view's image IV0 is reconstructed using the HEVC decoder.
[0406] If flag_proc = 1, the HEVC decoder is used to reconstruct images IV1, IV2, IV3, IV4, and IV5 corresponding to the encoded processed data. No processing is performed on the reconstructed images IV1, IV2, IV3, IV4, and IV5, as it is impossible to reconstruct the data of these images that have been removed by cropping. However, the synthesis algorithm uses the six reconstructed images IV0, IV1, IV2, IV3, IV4, and IV5 to generate an image with any view desired by the user.
[0407] In addition to the coordinates of the cropped area, the five data signals F21, F22, F23, F24, and F25 are used by the synthesis algorithm to generate an image of the view desired by the user.
[0408] According to the second example, consider 10 computer-generated images IV0…IV9, each with a resolution of 4096×2048 pixels, to simulate 10 cameras of a 360° type. It is decided that images IV0 and IV9 will not be processed. The texture components of images IV1 to IV8 themselves undergo a 2x downsampling in the horizontal direction and a 2x downsampling in the vertical direction, and the corresponding depth components themselves undergo a 4x downsampling in the horizontal direction and a 4x downsampling in the vertical direction. Therefore, the resolution of the texture components of images IV1 to IV8 becomes 2048×1024, and the resolution of the depth components of images IV1 to IV8 becomes 1024×512.
[0409] Processing data, such as image data related to images IV1 to IV8, including eight downsampled texture components of images IV1 to IV8 with a resolution of 2048×1024 and eight downsampled depth components of images IV1 to IV8 with a resolution of 1025×512.
[0410] In addition, the aforementioned processed data includes text-type data, which indicates the view Figures 1 to 8 The downsampling factors for each image, IV1 through IV8, are defined as follows:
[0411] -IV1 to IV8 textures: se_h = 2, se_v = 2 ("se" represents downsampling, "h" represents horizontal, and "v" represents vertical).
[0412] -IV1 to IV8 depth: se_h = 4, ss_v = 4.
[0413] The information flag_proc is set to 0 in association with images IV0 and IV9, and to 1 in association with images IV1 through IV8.
[0414] An MV-HEVC type encoder is used to encode images IV0 and IV9 simultaneously. This encoder generates a single data signal F1 containing the information flag_proc=0 and the raw data of the encoded images IV0 and IV9.
[0415] Associated with the encoded downsampled texture and depth data, an encoder of the MV-HEVC type also simultaneously encodes processing data (such as image data related to images IV1 through IV8). This processing data includes eight downsampled texture components of images IV1 through IV8 at a resolution of 2048 × 1024 and eight downsampled depth components of images IV1 through IV8 at a resolution of 1025 × 512. This encoder generates a single data signal F2 containing the information flag_proc = 1. Figures 1 to 8 The text-type data for each image, with downsampling factors IV1 through IV8, is itself losslessly encoded. The data signals F1 and F2 are concatenated and then transmitted to the decoder.
[0416] At the decoder, the information flag_proc is read.
[0417] If flag_proc=0, the MV-HEVC decoder is used to reconstruct images IV0 and IV9 simultaneously. Reconstructed images IVD0 and IVD9 at their initial resolution are then obtained.
[0418] If flag_proc=1, then the MV-HEVC decoder is used to simultaneously reconstruct images IV1 through IV8, which correspond to their respective encoded downsampled texture and depth data. The reconstructed downsampled images IVDT1 through IVDT8 are then obtained. The text-type data corresponding to each of the eight images IV1 through IV8 is also decoded, providing the downsampling factor already used for each image IV1 through IV8.
[0419] Subsequently, these images are processed using the corresponding downsampling factors of the reconstructed downsampled images IVDT1 to IVDT8. After processing, reconstructed images IVD1 to IVD8 are obtained, with each of the eight texture components of these reconstructed images at its initial resolution of 4096×2048, and each of the eight depth components at its initial resolution of 4096×2048.
[0420] The synthesis algorithm uses images of 10 views reconstructed at their initial resolution to generate an image of the view desired by the user.
[0421] According to the third example, consider three computer-generated images IV0 to IV2, each with a resolution of 4096 × 2048 pixels, to simulate four cameras of a 360° type. Three texture components and three corresponding depth components are then obtained. It is decided not to process image IV0 and to extract occlusion maps for images IV1 and IV2 separately. To do this, disparity estimation is performed between image IV1 and image IV0 to generate an occlusion mask for image IV1, that is, pixels of image IV1 that are not found in image IV0. Disparity estimation is also performed between image IV2 and image IV0 to generate an occlusion mask for image IV2.
[0422] Process data, such as image data related to images IV1 and IV2, including two texture components of the occlusion mask of images IV1 and IV2 and two depth components of the occlusion mask of images IV1 and IV2.
[0423] The information flag_proc is set to 0 in association with image IV0, and to 1 in association with images IV1 and IV2.
[0424] Image IV0 is encoded using an HEVC encoder. A single data signal F10 is generated, which contains the raw data of the encoded image IV0, with the information flag_proc = 0.
[0425] An HEVC encoder is used to encode the image data (texture and depth) of the occlusion mask for each of images IV1 and IV2. Two data signals, F21 and F22, are generated in association with the information flag_proc=1. These two data signals contain the encoded image data of the occlusion mask for each of images IV1 and IV2, respectively. The data signals F10, F21, and F22 are concatenated and then transmitted to the decoder.
[0426] At the decoder, the information flag_proc is read.
[0427] If flag_proc=0, the image IV0 is reconstructed using the HEVC decoder.
[0428] If flag_proc = 1, the encoded image data (texture and depth) corresponding to the occlusion mask of each of images IV1 and IV2 are used to reconstruct images IV1 and IV2 using the HEVC decoder. No processing is performed on the reconstructed images IV1 and IV2, as it is impossible to reconstruct the data of these images that were removed after occlusion removal was completed. However, the synthesis algorithm can use the reconstructed images IV0, IV1, and IV2 to generate an image representing the view desired by the user.
[0429] According to the fourth example, consider two images, IV0 and IV1, with a resolution of 4096×2048 pixels, captured by two cameras of the 360° type, respectively. Image IV0 of the first view is encoded using the first encoding method MC1 as usual, while image IV1 of the second view is processed before being encoded according to the second encoding method MC2. This processing includes the following:
[0430] - Use filters (such as the Sobel filter) to extract the contours of image IV1.
[0431] - For example, mathematical morphology operators can be used to apply expansion to a contour in order to increase the area around the contour.
[0432] Processing data, such as image data related to image IV1, includes pixels in the region surrounding the contour and pixels set to 0 corresponding to pixels located outside the region surrounding the contour.
[0433] Alternatively, text-type data can be generated, for example, in the form of marker information (e.g., YUV=000), which indicates that pixels located outside the region surrounding the outline are set to 0. Pixels set to 0 are neither encoded nor transmitted to the decoder.
[0434] The image IV0 is encoded using an HEVC encoder, which generates a data signal F1 containing the information flag_proc=0 and the raw data encoded in the image IV0.
[0435] The image data of the region surrounding the outline of image IV1 is encoded using an HEVC encoder, while the marker information is encoded using a lossless encoder. A data signal F2 is then generated, which includes the information flag_proc=1, the encoded pixels of the region surrounding the outline of image IV1, and the encoded marker information.
[0436] At the decoder, the information flag_proc is read.
[0437] If flag_proc=0, the image is reconstructed using the HEVC decoder at the original resolution of image IV0.
[0438] If flag_proc=1, then the image data corresponding to the region around the outline of image IV1 is reconstructed by the HEVC decoder using marker information that allows the values of pixels around the region to be set to 0 to be recovered.
[0439] The synthesis algorithm can use two reconstructed images, IV0 and IV1, to generate an image of the view desired by the user.
[0440] According to the fifth example, consider four images IV0 to IV3, each with a resolution of 4096 × 2048 pixels, captured by four cameras of a 360° type. Image IV0 is encoded conventionally using the first encoding method MC1, while images IV1 to IV3 are processed before being encoded according to the second encoding method MC2. This processing involves filtering images IV1 to IV3, during which a Region of Interest (ROI) is calculated. The ROI contains one or more regions in each image IV1 to IV3 that are considered most relevant, for example, because they contain a great deal of detail.
[0441] For example, this filtering can be performed using one of the following two methods:
[0442] - The saliency maps IV1 to IV3 of each image are calculated by filtering.
[0443] - Filter the depth maps of each image IV1 through IV3: For each texture pixel, the depth map is characterized by near or far depth values in the 3D scene. Define a threshold such that each pixel of images IV1, IV2, and IV3 below this threshold is associated with an object in the scene near the camera. Subsequently, all pixels below this threshold are considered regions of interest.
[0444] Processing data, such as image data related to images IV1 to IV3, includes pixels within their respective regions of interest and pixels set to 0 corresponding to pixels located outside these regions of interest.
[0445] Alternatively, text-type data can be generated, for example, in the form of marker information indicating that pixels located outside the region of interest are set to 0. Pixels set to 0 are neither encoded nor transmitted to the decoder.
[0446] The image IV0 is encoded using an HEVC encoder, which generates a data signal F10 containing the information flag_proc=0 and the raw data encoded in the image IV0.
[0447] The HEVC encoder encodes the region of interest (ROI) image data for each of images IV1, IV2, and IV3, while a lossless encoder encodes the marker information. Three data signals F21, F22, and F23, along with corresponding encoded marker information, are generated in association with the information flag_proc=1. These data signals contain the encoded image data for the ROI of each of images IV1, IV2, and IV3. The data signals F10, F21, F22, and F23 are concatenated and then transmitted to the decoder.
[0448] At the decoder, the information flag_proc is read.
[0449] If flag_proc=0, the image is reconstructed using the HEVC decoder at the original resolution of image IV0.
[0450] If flag_proc=1, the image data corresponding to the respective regions of interest for each image IV1 to IV3 are reconstructed using the HEVC decoder with marker information that allows the values of pixels around the region to be set to 0 to be recovered.
[0451] The synthesis algorithm can directly use the four reconstructed images IV0, IV1, IV2, and IV3 to generate an image of the view required by the user.
[0452] It goes without saying that the embodiments described above have been given only by way of completely non-limiting instruction, and that many modifications can be readily made by those skilled in the art without otherwise departing from the scope of the invention.
Claims
1. A method, implemented by an encoding device, for encoding an image of a view forming part of a plurality of views, the plurality of views simultaneously representing a 3D scene from different viewpoints or positions, the method comprising: - Select (C1) either the first encoding method or the second encoding method to encode the image of the view. - Generate (C10, C12a; C10, C12b; C10, C12c) a data signal containing information indicating whether the first encoding method or the second encoding method has been selected. - If the first encoding method is selected, the raw data of the image of the view is encoded (C11a), wherein the first encoding method provides the encoded raw data, The generated data signal further includes: the encoded raw data of the image of the view. -If the second encoding method is selected: • The processed data of the image of the view is encoded (C11b), the processed data being data of at least one remaining region of the image of the view, the remaining region being one or more portions of the image retained after cropping the original data of the image of the view, the second encoding method providing at least one encoded remaining region. • Encode (C11b) the descriptive information of the cropping, which is information about the location of the remaining region in the image of the view. The generated data signal further includes: the encoded remaining region of the image of the view and the encoded descriptive information of the cropping.
2. The method of claim 1, wherein: - The remaining region of the image of the view corresponds to the pixels of the image of the view that have not been deleted after the cropping of the image of the view.
3. The method of claim 1, wherein the cropping description information includes some coordinates of the top leftmost pixel located in the remaining area of the image of the view, and some coordinates of another pixel located in the bottom rightmost area of the remaining area of the image of the view.
4. The method of claim 1, wherein the cropped descriptive information includes rows and / or columns of pixels deleted from the image of the view, and the positions of the rows and / or columns in the image of the view.
5. The method of claim 1, wherein the cropping is configured to delete a fixed number of pixels of the image of the view.
6. A method, implemented by a decoding device, for decoding data signals representing images that form part of a plurality of views, the plurality of views simultaneously representing a 3D scene from different viewpoints or positions, the method comprising: - In this data signal, the read (D1) indicates whether the image of the view will be decoded according to the first decoding method or the second decoding method. -If it is the first decoding method: • Then, read (D11a) the encoded data associated with the image of the view from the data signal. • Reconstruct (D12a) an image of the view based on the encoded data, wherein the reconstructed image of the view contains the original data of the image of the view. -If it is the second decoding method: • Then read (D11b; D11c) from this data signal: - Encoded data associated with the image of the view, the encoded data being encoded data of at least one remaining region of the image of the view, the remaining region being one or more portions of the image retained after cropping the original data of the image of the view. - The cropping description information, the encoded description information being information about the location of the remaining region within the image of the view. • Based on the encoded remaining region and the description information of the cropping, reconstruct (D12b) the image of the view.
7. The method of claim 6, wherein the remaining region corresponds to pixels of the image of the view that have not been deleted after the cropping of the image of the view.
8. The method of claim 6, wherein the cropping description information includes some coordinates of the top leftmost pixel located in the remaining area of the image of the view, and some coordinates of another pixel located in the bottom rightmost area of the remaining area of the image of the view.
9. The method of claim 6, wherein the cropped descriptive information includes rows and / or columns of pixels deleted from the image of the view, and the positions of the rows and / or columns in the image of the view.
10. An apparatus for encoding an image of a view forming part of a plurality of views, the plurality of views simultaneously representing a 3D scene from different viewpoints or positions, the apparatus comprising a processor configured to perform the following operations: - Select either a first encoding method or a second encoding method to encode the data of the image in the view. - Generate a data signal containing information indicating whether the first encoding method or the second encoding method has been selected. - If the first encoding method is selected, the raw data of the image of the view is encoded, wherein the first encoding method provides the encoded raw data, The generated data signal further includes: the encoded raw data of the image of the view. -If the second encoding method is selected: • The processed data of the image of the view is encoded. The processed data is data of at least one remaining region of the image of the view. The remaining region is one or more portions of the image retained after cropping the original data of the image of the view. The second encoding method provides at least one encoded remaining region. • Encode the descriptive information of the cropping, which is information about the location of the remaining region in the image of the view. The generated data signal further includes: the encoded remaining region of the image of the view and the encoded descriptive information of the cropping.
11. An apparatus for decoding data signals of images representing a portion of a plurality of views, the plurality of views simultaneously representing a 3D scene from different viewpoints or positions, the apparatus comprising a processor configured to perform the following operations: - In this data signal, information indicating whether the image of the view will be decoded according to a first decoding method or a second decoding method is read. -If it is the first decoding method: • Then, the encoded data associated with the image of the view is read from the data signal. • Reconstruct an image of the view based on the read encoded data, the reconstructed image of the view containing the original data of the original image of the view. -If it is the second decoding method: • Then read from this data signal: - Encoded data associated with the image of the view, the encoded data being encoded data of at least one remaining region of the image of the view, the remaining region being one or more portions of the image retained after cropping the original data of the image of the view. - The cropping description information, the encoded description information being information about the location of the remaining region within the image of the view. • The image of the view is reconstructed based on the encoded remaining region and the description information of the cropping.
12. A storage medium that is computer-readable and includes a computer program that, when executed on a computer, performs the steps of the method as claimed in any one of claims 1 to 9.
Citation Information
Patent Citations
3D video encoding and decoding methods and apparatus
US9485494B1
An apparatus, a method and a computer program for video coding and decoding
WO2018002425A2