Method and apparatus for decoding multi-view video, and method, computer program and apparatus for image synthesis
By providing standardized metadata to the image processing module during decoding, the method improves the quality and efficiency of virtual view synthesis in virtual reality applications, addressing the challenges of suboptimal image quality and complexity in existing technologies.
Patent Information
- Application Number
- JP2024175158
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-05-03
- Filing Date
- 2024-10-04
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2039-04-16
AI Technical Summary
Existing virtual reality applications face challenges in synthesizing high-quality virtual views due to the lack of standardized methods for decoding multi-view video data, leading to suboptimal image quality and increased computational complexity.
A method and apparatus for decoding multi-view video that provides metadata in a standardized format to an image processing module, allowing it to control the virtual view synthesis process more efficiently, reducing computational complexity and improving image quality.
The proposed solution enhances the quality of synthesized virtual views and simplifies the decoding process, enabling smoother transitions and reducing the number of cameras required for capturing scenes.
Smart Images

Figure 0007745061000002 
Figure 0007745061000003 
Figure 0007745061000004
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to the field of 3D image processing, and more particularly to the simulation of multi-view images. This involves decoding the sequence and synthesizing intermediate views. [Background technology]
[0002] In the field of virtual reality, free navigation allows the viewer to View the scene from any viewpoint, where that viewpoint corresponds to the viewpoint captured by the camera. whether it corresponds to a viewpoint not captured by the camera or Such views that are not captured by a camera are called virtual views. This view is also called the intermediate view, because it is the view captured by the camera. Because it is between the captured views and must be composited for restoration. is.
[0003] Free navigation within a scene involves all the movements of the user viewing the multi-view video. This ensures that the image is properly managed and that the viewer does not perceive it as being suboptimal. It is necessary to avoid the discomfort that can result.
[0004] Generally, the user's movements are recorded by a virtual device, such as a head-mounted device (HMD). This is properly taken into account by the reality headset.
[0005] However, regardless of the user's movement (rotation or translation), Providing the pixels is still a problem. In practice, the calculation of the image to be displayed is , some captures are used to display additional images of the virtual (i.e., synthesized) view. Such a virtual view requires the use of a virtual camera captured by the camera. and the decoded captured view and the related It is required to calculate them from the associated depths.
[0006] Therefore, codecs that provide free navigation functionality require several views and It is effective for encoding the image and associated depth, while also providing optimal resolution for the virtual view. rendering must be possible, i.e., the compositing algorithm must be able to be used for display. and requires.
[0007] Multiview video encoder designed to encode multiview sequences standard MV-HEVC or 3D-HEVC (Series H: Audiovisual l and multimedia systems - Infrastructure of audio visual services - Coding of m oving video, High Efficiency Video Coding, Recommendation ITU-T H.265, Internati National Telecommunication Union, December 2016) is known.
[0008] The MV-HEVC encoder applies very basic inter-view prediction, while 3D-H The EVC encoder includes some additional tools to handle more than just temporal redundancy. It also exploits inter-view redundancy. In addition, 3D-HEVC efficiently encodes depth maps. These two codecs, and especially 3D-HEV, have specific tools for C is a 2D standard such as the HEVC standard when encoding multiple views with associated depth. Effectively reduces bitrate compared to traditional video codecs for handling video sequences do.
[0009] In a virtual reality situation, it is captured by a camera and encoded into a data stream. After decoding the currently captured view, a virtual view is synthesized based on, for example, the user's movements. It is possible.
[0010] For example, the VSRS tool (Wegner, St ankiewicz, Tanimoto, Domanski, Enhanced view synthesis reference software (VSRS) for free-viewpoint television, ISO / IEC JTC1 / SC29 / WG11 m31520, October 2013, Gen eva, Switzerland) is known.
[0011] Figure 1 shows how a decoder DEC (e.g., 3D-HEVC) decodes a data stream STR. The conventional free navigation system generates decoded views (VD1, VD2) by In this case, such views are synthesized by a view synthesizer SYNT used by H (e.g., VSRS) to generate the synthesized view VS(1+2). Then the decoded and synthesized views are restored according to the user's movements. Displayed by the source device DISP.
[0012] A conventional decoder DEC is shown in Figure 2. Typically, such a decoder Perform analysis of the data stream STR (E20) to obtain the relevant data to be decoded. Then, a decoding process (E21) is applied to generate the virtual view. Decoded views (VD1, VD2) that can be used later by SYNTH Reconstruct.
[0013] Therefore, the process of decoding the views from the data stream and the process of synthesizing the virtual views are In particular, the synthesis process is a difficult process that does not involve a decoder. The decoder simply extracts the decoded video signal reconstructed from the data stream. The queue is made available to the synthesis module.
[0014] The technical problem facing virtual reality applications is that encoders and decoders, in particular In the case of free navigation, advance knowledge of the final viewpoint required by the user Multi-view video encoders and decoders also have the following features: , and has no knowledge of the compositing process that will ultimately be used to synthesize the virtual view. The synthesis method used to synthesize the virtual view is different from that used in multiview video decoders. is not currently standardized and therefore is not used by virtual reality applications. The synthetic method remains a unique tool.
[0015] Therefore, the quality of the synthesized virtual view is limited by the quality of the virtual view used by such applications. Generally, such quality depends on the synthesis tools and algorithms used. The complexity of the synthesis tools used and the resources of the equipment implementing these synthesis tools will determine the do.
[0016] Virtual reality applications, and more particularly virtual reality with free navigation The application must be real-time. The virtual view synthesis module: Generally, decoding is difficult, especially if the number of views captured and decoded is insufficient. and the reconstructed captured view is of high visual quality, but of medium quality. provides a virtual view of Summary of the Invention [Problem to be solved by the invention]
[0017] The present invention improves upon the state of the art. [Means for solving the problem]
[0018] The present invention relates to a method for decoding a data stream representing a multi-view video, which is implemented by a decoding device. A method for decoding a data stream, comprising: extracting a syntax element from at least one portion of a data stream; The syntax elements are then retrieved and the number of views of the video is calculated from the retrieved syntax elements. Advantageously, this also relates to a decoding method, which includes reconstructing one image at a time. The decoding method involves extracting a small amount of metadata in a predetermined format from at least one syntax element. At least one item of metadata must be acquired and at least one item of metadata must be transferred to the image processing module. and providing the module with the information.
[0019] Therefore, such a decoding method may be implemented externally to an image processing module, e.g., a decoder. The synthesis module represents the data of the video stream and uses it by the image processing module. This allows the image processing module to provide metadata that can be used to The processing performed within the virtual view synthesis module is less complex. In this case, the data used by the synthesis algorithm and available to the decoder Furthermore, the present invention provides a method in which the image processing module: It has access to data that it cannot compute on its own and uses that data to control its behavior. For example, in the case of the virtual view synthesis module, the decoder , we can provide an occlusion map to the synthesis module, and such occlusion The motion is determined by the synthesis module solely from the reconstructed image of the video view. It is difficult to do so.
[0020] Therefore, the processing performed in the image processing module can be improved. ,because the computational complexity of obtaining data that is available at the decoder level is reduced. This reduces the image quality and therefore requires more complex and therefore more powerful image processing algorithms. This allows the system to be more easily implemented in the image processing module.
[0021] In the case of the virtual view synthesis module, the quality of the virtual view is thus improved. This improves multiview video by providing smoother transitions between views. It also improves the user's free navigation in the scene. This also reduces the number of cameras needed to capture a scene.
[0022] By providing metadata in a predetermined format, decoders and image processing modules For example, metadata can be indexed and It is provided in the form of a standardized table, so that the image processing module can easily For each box, you know what metadata is stored in this index. do.
[0023] It is known to use metadata in video data communication, e.g., H.264 / AV The SEI (Supplemental Enhancement Information) message introduced in the C standard The message is data about optional processing operations performed at the decoder level. The SEI message is transmitted to the decoder via the video data bitstream. However, such SEI message data is created at the encoder level. and is used only by the decoder, and is optionally decoded and reconstructed Improves the quality of the views.
[0024] According to a particular embodiment of the invention, obtaining at least one item of metadata extracts at least one part of the syntax element from at least one item of the above metadata The method further includes calculating the eye.
[0025] Such particular embodiments of the present invention may, for example, be implemented using a decoder to reconstruct the views. information not used by the system, e.g., confidence values calculated for depth information or in another form. information used by the decoder, e.g., rather than the granularity used when reconstructing an image. It allows the computation of new metadata corresponding to coarse-grained motion information. do.
[0026] According to another particular embodiment of the invention, at least one item of said metadata is are not used to reconstruct at least one image.
[0027] According to another particular embodiment of the invention, at least one item of said metadata comprises: The following, namely: - camera parameters, - decoded and scaled motion vectors, -Segmentation of the reconstructed image, - the reference image used by the image block of the reconstructed view, - coding mode of the image of the reconstructed view, - the quantization parameter value of the image of the reconstructed view, - prediction residual values of the image of the reconstructed view, - a map representing the motion within the image of the reconstructed view, - a map representing the presence of occlusions in the image of the reconstructed view, - a map representing confidence values associated with the depth map, corresponds to one item of information contained within a group containing
[0028] According to another particular embodiment of the invention, the predetermined format is at least one of the metadata It corresponds to an indexed table, where entries are stored in association with an index.
[0029] According to another particular embodiment of the invention, at least one item of said metadata is It is obtained based on the level of granularity specified in the code device.
[0030] According to this particular embodiment of the present invention, the metadata generated from the syntax elements is For example, motion information can be obtained at different granularity levels. the granularity used in the coder (i.e., as used by the decoder), or (For example, provide one motion vector per block of size 64x64. It is possible to provide motion vectors with a coarser granularity (by
[0031] According to another particular embodiment of the present invention, the decoding method comprises: This request indicates at least one item of metadata required by the processing module. According to this particular embodiment of the present invention, For example, the image processing module indicates to the decoder the information that the image processing module needs. Therefore, the decoder makes only the necessary metadata available to the image processing module. This limits the complexity and memory resource usage in the decoder. It will be what is created.
[0032] According to another particular embodiment of the invention, the request includes a predetermined list of available metadata. It contains at least one index pointing to an item of required metadata in
[0033] The present invention also relates to decoding according to any one of the specific embodiments defined above. The present invention relates to a decoding device configured to carry out the method. It can include different features related to the decoding device according to the invention. The features and advantages of the decoding device are the same as those of the decoding method and are further detailed below. do not have.
[0034] According to a particular embodiment of the invention, such a decoding device is located in a terminal or in a server. Included.
[0035] The present invention provides a method for decoding a video signal from at least one image of a view decoded by a decoding device, The present invention also relates to a method of image synthesis, which comprises generating at least one image of a virtual view. According to the disclosure, such an image processing method involves processing at least one item of metadata in a predetermined format. and reading at least one item of said metadata by a decoding device. Therefore, at least one symbol obtained from a data stream representing a multi-view video is the at least one image is obtained from the tax element, and the at least one image has at least one of the metadata It is generated using one retrieved item.
[0036] Therefore, the image synthesis method utilizes the metadata available to the decoder to Such metadata is used by the image processor to generate virtual view images of the audio video. This can correspond to data to which the user does not have access or that can be recalculated. However, the calculations become very complicated.
[0037] The virtual view here refers to the sequence of images captured by the camera of the scene capture system. It means a new perspective view of a scene that has not been captured before.
[0038] According to a particular embodiment of the present invention, the image synthesis method includes: The method further includes sending a request indicating at least one item of metadata necessary for the
[0039] The present invention also relates to an image processing method according to any one of the specific embodiments defined above. The present invention relates to an image processing device configured to perform the method. It can include different features related to the image processing method according to the invention. The features and advantages of the processing device are the same as those of the image processing method and are further detailed below. do not have.
[0040] According to a particular embodiment of the present invention, such an image processing device is installed in a terminal or a server. Included.
[0041] The present invention also provides a method for extracting multiview video from a data stream representing the multiview video. An image processing system for displaying a decoding image according to any one of the above embodiments. and an image processing device according to any one of the above embodiments. Regarding the system.
[0042] The decoding method, respectively the image processing method according to the invention, can be implemented in various ways, in particular by It can be implemented linearly or in software. The decoding method and each image processing method are implemented by a computer program. The present invention also provides a method for implementing the above-described specific embodiments when executed by a processor. A computer program containing instructions for implementing any one of the decoding or image processing methods Such programs may use any programming language. The program can be downloaded from a communication network and / or installed on a computer. It can be recorded on a readable medium.
[0043] This program can be written in any programming language, and can be object code, intermediate code between source code and object code, e.g. The file may be in a partially compiled format, or in any other desired format.
[0044] The present invention relates to a computer-readable storage medium containing instructions for the computer program described above. The above-mentioned recording medium can store a program. The medium may be any entity or device. For example, the medium may be a storage means, e.g. ROM, such as CD-ROM or microelectronic circuit ROM, USB flash drive On the other hand, the recording medium may include a hard drive, a hard disk, or a magnetic recording means. , which can be transmitted over electrical or optical cables, by radio or other means. The program according to the present invention can correspond to a transmittable medium such as an electrical signal or an optical signal. The RAM can be downloaded, especially over an Internet-type network.
[0045] Alternatively, the recording medium may correspond to an integrated circuit in which the program is embedded; The circuitry is adapted to perform or be used in performing the method.
[0046] Further features and advantages of the present invention will become apparent from the following description, which is given by way of example only and not by way of limitation, with reference to the accompanying drawings, in which: This will become more apparent upon reading the following description of specific embodiments, given by way of example. [Brief explanation of the drawings]
[0047] [Figure 1] 1 is a diagrammatic illustration of a system for free navigation in multi-view video according to the prior art; FIG. [Figure 2] 1 shows diagrammatically a decoder of a data stream representing multi-view video according to the prior art; FIG. [Figure 3] FIG. 1 illustrates a schematic diagram of a system for free navigation in multi-view video, in accordance with certain embodiments of the present invention. [Figure 4]FIG. 3 illustrates steps of a method for decoding a data stream representing multi-view video according to a particular embodiment of the invention. [Figure 5] FIG. 2 shows diagrammatically a decoder of a data stream representing multi-view video according to a particular embodiment of the invention; [Figure 6] 3A-3D illustrate steps of an image processing method according to a particular embodiment of the present invention; [Figure 7] 5 illustrates steps of a decoding method and an image processing method according to another particular embodiment of the invention; [Figure 8] 1 shows diagrammatically an apparatus adapted to implement a decoding method according to a particular embodiment of the invention; [Figure 9] 1 shows diagrammatically an apparatus adapted to carry out an image processing method according to a particular embodiment of the invention; [Figure 10] FIG. 1 is a diagram illustrating the arrangement of views in a multi-view capture system. DETAILED DESCRIPTION OF THE INVENTION
[0048] The present invention modifies the decoding process of a data stream representing multi-view video. This allows image processing based on the views reconstructed by the decoding process. For example, the image processing process is simplified for the process of synthesizing virtual views. For this purpose, the decoder only receives the images of the reconstructed views from the data stream. Instead, it also provides the metadata associated with such images, which are then , can be used for the synthesis of virtual views. Advantageously, such metadata is formatted, i.e., to facilitate interoperability between decoders and synthesizers. Therefore, to synthesize the virtual view, Any synthesizer configured to read the metadata in the image can be used.
[0049] FIG. 3 illustrates a method for free navigation in multiview video according to a particular embodiment of the present invention. The system in Figure 3 is similar to that described in relation to Figure 1. It operates in the same way as the system described above, but the decoder DEC outputs a reconstructed The difference is that it provides metadata MD1 and MD2 in addition to the images of views VD1 and VD2. Such metadata MD1, MD2 are provided at the input to the synthesizer. ,Then, the synthesizer generates a virtual view VS( 1+2) and the decoder DEC and the synthesizer SYNTH generate the The decoder DEC and the synthesizer SYNTH form an image processing system. It can be contained within a single device, or it can be contained within two separate devices that can communicate with each other. can.
[0050] For example, but not limited to and non-exhaustively, such metadata may correspond to: This can be done. - the camera parameters of the view reconstructed by the decoder, - decoding and scaling of the image reconstructed by the decoder; Kutlu, -Segmentation of the reconstructed image, - indication of the reference image used by the block of the reconstructed image, - coding mode of the reconstructed image, the quantization parameter value of the reconstructed image, -Prediction residual values of the reconstructed image.
[0051] Such information can be provided for use by the decoder. Alternatively, such information may be provided by a decoder, e.g., to a decoder using a granularity parameter. The grain size can be processed to provide a finer or coarser grain size than the grain size.
[0052] Metadata can also be computed and shared by decoders, for example: do. A map representing the global motion within one image or group of images of the reconstructed view For example, such a map can be created by thresholding the motion vectors of an image or group of images. It can be a binary map obtained by processing. - A map representing the presence of occlusions in the image of the reconstructed view. Such a map indicates the level of information contained in the prediction residual for each pixel in the case of inter-view prediction. It can be a binary map obtained by considering the occlusion The information of possible locations of the motion is derived from the disparity vector or edge map of the image. It is possible. A map representing confidence values associated with the depth map. For example, such a map could be: By comparing the texture coding mode with the corresponding depth coding mode It can be calculated by the decoder.
[0053] Some of the output metadata may be data relating to a single view. If , this output metadata is specific to that view. Other metadata is 2 The metadata can be obtained from more than one view. In this case, the metadata is Difference or correlation (difference in camera parameters, occlusion map, decode mode) etc.)
[0054] FIG. 4 illustrates a data stream representing a multi-view video according to a particular embodiment of the present invention. 1 shows steps of a method for decoding a
[0055] The data stream STR is applied to the input of the decoder DEC, e.g. as a bit stream. The data stream STR is provided in the form of a A conventional video encoder adapted to encode multi-view video is or single-view video encoding applied separately to each view of a multi-view video. It contains multi-view video data encoded by the decoder.
[0056] In step E20, the decoder DEC determines whether the decoded syntax elements are Decode at least one portion of the data stream to be decoded. 20 is, for example, the current view of the view to be reconstructed, e.g., the view viewed by the user. Parsing the data stream to extract the syntax elements needed to reconstruct the image Such a syntax supports decoding the bitstream's entropy. The data element may contain, for example, the coding mode of the blocks of the current picture, inter-picture prediction or inter-view prediction. In the case of estimation, it corresponds to a motion vector, a quantization coefficient of a prediction residual, etc.
[0057] Conventionally, during step E21, the current image of the view (VD1, VD2) to be reconstructed is An image consists of decoded syntax elements and possibly their views or other previous Such a reconstruction of the current image is the coding mode used at the encoder level to encode the image of It is carried out according to forecasting techniques.
[0058] The images of the reconstructed views are provided at the input of the image processing module SYNTH. do.
[0059] In step E23, at least one item of metadata is Such items of metadata are obtained from coded syntax elements. Such a predetermined format may be, for example, the format in which the data is transmitted or It corresponds to a specific syntax configured to be stored in memory. If the decoder is a compliant decoder, the metadata syntax is: For example, the specific standard or the standard associated with the specific decoding standard. It can be assumed that the
[0060] According to a particular embodiment of the invention, the predetermined format is defined by at least one item of metadata: corresponds to an indexed table in which the According to a particular embodiment, each metadata type is associated with an index. An example of such a table is shown in Table 1 below.
[0061] [Table 1]
[0062] Each item of metadata is associated with its index and is based on the metadata type. It is stored in a suitable format.
[0063] For example, the camera parameters of a view each correspond to, for example, the position of the camera in the scene. The coordinates of the point in the corresponding 3D coordinate system and the three angles in the 3D coordinate system a triplet of data including orientation information defined by the values of and stored.
[0064] According to another example, the motion vectors are given by , are stored in the form of a table containing the corresponding motion vector values.
[0065] The metadata table shown below is a non-limiting example only. The metadata may be stored in other predefined For example, if only one metadata type is possible, It is not necessary to associate an index with that metadata type.
[0066] According to a particular embodiment of the invention, in step E22, at least one of the metadata An item is at least one of the decoded syntax elements before the acquisition step E23. is also calculated from one part.
[0067] Thus, according to such a particular embodiment of the present invention, the current It is not used to reconstruct the image, but synthesizes a virtual view from the current reconstructed image. It is possible to obtain metadata, e.g., an occlusion map, that can be used to It becomes possible.
[0068] According to such a particular embodiment of the present invention, the grain size used to reconstruct the current image is It is also possible to obtain metadata with a different granularity from the motion For example, if the block size is 64x64 pixels on the whole image, the vector is 64x The reconstructed motion vectors of all sub-blocks of the current image contained within the block of 64 are From the motion vector, it can be calculated more coarsely, e.g., for each 64x64 block. The motion vector is the minimum or maximum of the motion vectors of the sub-blocks. The value may be calculated by selecting the mean or median, or any other function.
[0069] In step E24, the metadata MD1, MD2 obtained in step E23 are 2 is an image processing module SYNTH external to the decoder DEC, e.g., a virtual view synthesis module The decoder's external module is responsible for decoding the data stream. This behavior is necessary for the decoder to display the reconstructed view. means a module that is not
[0070] For example, the metadata may be stored in a memory accessible to the image processing module. According to this example, the metadata is stored in a memory where the decoder and the image processing module are integrated in the same device. If the image data is to be transmitted to the image processing module via a connection link such as a data transmission bus, or a cable if the decoder and image processing module are integrated in separate devices. The image data is transmitted to the image processing module via a wired or wireless connection.
[0071] FIG. 5 illustrates a data stream representing a multi-view video according to a particular embodiment of the present invention. 1 shows a schematic of a decoder for
[0072] Conventionally, decoding of the views reconstructed from the data stream STR is as follows: The decoding of the reconstructed views is performed image by image and block by block for each image. For each block to be reconstructed, the elements corresponding to that block are Decoded from the data stream STR by the entropy decoding module D, Decoded syntax elements SE (texture encoding mode, motion vector tors, disparity vector, depth encoding mode, reference image index, ...) and quantity A set of coefficients coeff is provided.
[0073] The quantization coefficients coeff are passed through the inverse quantization module (Q -1 ) and then the inverse transformation model Joules (T -1 ) to obtain the prediction residual value res of the block. rec is provided. The coded syntax elements (SE) are sent to a prediction module (P) to Reconstructed image I ref (A portion of the current image or a previously reconstructed view) The prediction block pre is also calculated using a reference image (or a reference image of another previously reconstructed view). Then, the current block is calculated by dividing the prediction pred by the prediction residual re of the block. s rec is reconstructed by adding to (B rec ). Then the reconstructed block B rec ) can be used later to reconstruct the current image or another image or another view. The data is stored in the memory MEM so that the data can be read.
[0074] According to the invention, at the output of the entropy decoding module, the decoding of the blocks The encoded syntax element SE and optional quantization coefficients are and selecting at least a portion of the quantization coefficients SE and the optional quantization coefficients, and dividing them into predetermined a module FORM configured to store the reconstructed image in the form , or metadata MD about a group of images is provided.
[0075] The selection of the decoded syntax element SE to be formatted is determined by, for example, the decoder. It can be fixed as specified in the standard that describes the operation of The choice of different types can be fixedly defined, e.g. via a decoder profile. The decoder parameterization is performed by the corresponding syntax of the format module FORM. This can be configured to select the x element. The decoder exchanges with the image processing module to which it provides metadata. In this case, the image processing module can send the received image to the decoder. The decoder module FORM explicitly indicates the type of metadata it wishes to receive. , selects only the requested decoded syntax elements.
[0076] Providing metadata at a different level of granularity than that used by the decoder If possible, such a level of granularity should be specified in the standard describing the decoder's behavior. It can be defined in the image processing model or fixedly via a decoder profile. When the module communicates with the decoder to obtain metadata, the image processing module , the decoder specifies the granularity level at which this image processing module wants to receive parts of the metadata. can be shown explicitly in
[0077] According to a particular embodiment of the invention, at the output of the entropy decoding module: The decoded syntax element SE and optional quantization coefficients are stored in the syntax element S a module CALC configured to calculate the metadata from E and / or the quantized coefficients As mentioned above, the computed metadata describes the decoder's behavior. may be explicitly defined in the standard or according to a different profile or otherwise. Alternatively, it can be determined from the exchange with the image processing module in question.
[0078] According to a particular embodiment of the invention, the module FORM is particularly adapted to process the views to be reconstructed. Select the camera parameters.
[0079] To synthesize a new viewpoint, the synthesis module synthesizes each pixel of the original (reconstructed) view. A model must be created that describes how the pixels are projected onto the virtual view. Most compositors, e.g., compositors based on DIBR (Depth Image Based Rendering) techniques, The device uses depth information to project the pixels of the reconstructed view into 3D space. Then, the corresponding points in 3D space are projected onto the camera plane from the new viewpoint.
[0080] Such a projection of an image point in 3D space is calculated using the following formula: M=K.RT.M' where M is the coordinate matrix of the points in 3D space and K is is the matrix of intrinsic parameters of the virtual camera, and RT is the A row of the camera's extrinsic parameters (camera position and orientation in 3D space) column, and M' is the pixel matrix of the current image.
[0081] If the camera parameters are not sent to the synthesis module, the synthesis module will and their camera parameters must be calculated at the expense of accuracy, and the calculation is It cannot be done in real time or must be acquired by an external sensor. Therefore, by providing these parameters through the decoder, the synthesis module This makes it possible to limit the complexity of the rules.
[0082] According to another particular embodiment of the invention, the module FORM is in particular adapted to reproduce the current image. Select the syntax elements for the reference images that will be used to construct it.
[0083] To generate the virtual view, a synthesis module uses various previously reconstructed available images. If there is a possibility to select a reference image from among the images of the views, the synthesis module Which reference view was used when coding the view used for For example, Figure 10 shows a multi-camera system with 16 cameras. The view capture system's view arrangement is shown. The arrows between each frame indicate The view decoding order is shown. The synthesis module is between view V6 and view V10. 10. The virtual view VV Conventionally, when a virtual view needs to be generated, the synthesis module constructs the best virtual view. To achieve this, the availability of each view must be checked.
[0084] According to certain embodiments described herein, for a view, If the synthesis module has metadata indicating the reference view used to reconstruct the image, The virtual viewer uses the virtual image to determine which images to use to generate the virtual view. Only the closest available view to the point (view V6 in the case of Figure 10) can be selected. For example, if a block in view V6 uses an image in view V7 as a reference image, The composition module is used by view V6 and therefore needs to be available. It may also be decided to use V7. By avoiding the need to check the availability of each view in a synthesis module, Reduce noise.
[0085] According to another particular embodiment of the invention, the module CALC is particularly Select a syntax element related to the motion vector to generate a group.
[0086] In areas with little motion, virtual view synthesis generally suffers from inaccuracies in the depth map. These incoherences are due to the virtual viewpoint. This is very hindering for visualization of the
[0087] In this particular embodiment, the decoder module CALC is the motion vectors, i.e., the inverse prediction of the motion vectors and the motion vectors The module CALC selects the motion vector after scaling of the motion map. , typically by using the reconstructed motion vectors of each block to generate a binary map. In a binary map, each element takes the value 0 or 1, and the area is Binary maps indicate whether a region has motion or not. For example, erosion, expansion, opening, closing This can be improved by using (closing)).
[0088] The motion binary map is then scaled to the desired granularity (pixel-level map, block Level map or sub-block level map, or for a specific block size within an image The motion is formatted according to the map defined for the view, and the It can indicate whether or not
[0089] A synthesis module that receives such a motion map then computes, for example, Apply different compositing processes depending on whether or not is marked as having motion For example, the time incoherence can be reduced. To solve the problem, the traditional synthesis process is disabled in fixed (motionless) regions. and simply inherit the pixel values from the previous image.
[0090] Of course, the synthesis module may be implemented using other means, e.g., as an encoder. By estimating the motion, a motion map can be generated. However, such an operation increases the complexity of the synthesis algorithm and the resulting model. This significantly impacts the accuracy of the decoder's output because the encoder The purpose is to estimate motion from uncoded images that are no longer available. .
[0091] In the example shown in FIG. 10 and in the embodiment described above, the closest available view is Effectiveness can be achieved not only by using the reference view but also by averaging the reference views in the vicinity of the virtual viewpoint. For example, we can calculate the reference views V6, V7, V10, and V 11 can be averaged by the decoder module CALC, resulting in The resulting average view can be sent to a synthesis module.
[0092] In another variant, the decoder module CALC calculates the occlusion map. where the occlusion map can be calculated for each pixel or block of the image as Indicates whether the region corresponds to an occlusion region. For example, the module CALC information about the reference image(s) used by the decoder to reconstruct the region By using For example, in the case of Figure 10, most of the blocks in the image of view V6 are predicted using temporal prediction. and some blocks in the image of view V6 are inter-view predicted, e.g., When using inter-view prediction for 2, these blocks correspond to occlusion regions. It is highly likely that this will occur.
[0093] A synthesis module that receives such an occlusion map then calculates whether the region is occluded. It is possible to decide to apply different compositing processes depending on whether the area is marked as a region or not. This can be done.
[0094] According to another particular embodiment of the invention, the module CALC is in particular Select the coding mode associated with the texture of the image and the depth map of the image. do.
[0095] According to the prior art, compositing algorithms mainly use depth maps. The map typically shows errors that create artifacts in the synthesized virtual view. By comparing the encoding modes between the texture and the depth map, the decoder , a confidence measure associated with the depth map, e.g., whether depth and texture are correlated (value A binary map can be derived that indicates whether the correlation is significant (value 1) or not (value 0).
[0096] For example, the confidence value can be derived from the encoding mode. The code mode and depth encoding mode are different, for example, one is intra mode (intr a mode and the other is inter mode, this means that the texture and This means that the depth is not correlated with the image, so the confidence value is low, e.g., 0. .
[0097] The confidence values can also be arranged according to the motion vectors. If the textures have different motion vectors, this indicates that the texture and depth are uncorrelated. Therefore, the confidence value is low, e.g., 0.
[0098] The confidence values can also be arranged according to the reference images used by the texture and depth. If the reference images are different, this means that the texture and depth are uncorrelated. Therefore, the confidence value is low, e.g., 0.
[0099] The synthesis module that receives such a confidence map then determines whether the region has a low confidence value. You can decide to apply different compositing processes depending on whether the For example, for such a region, another reference that provides a better confidence value for the region may be used. The views can be used to synthesize corresponding regions.
[0100] Figure 6 illustrates steps of an image processing method according to a particular embodiment of the invention. Such processing can be decoded and reconstructed, for example, by the decoding method described in connection with FIG. This is performed, for example, by a virtual view synthesis module from the generated views.
[0101] At step E60, at least one item of metadata (MD1, MD2) is Read by the synthesis module. Metadata read by the synthesis module corresponds to a syntax element decoded from a stream representing multiview video. , which are associated with one or more views, which are the sequences of the decoded syntax elements. It can also accommodate information calculated during the method of decoding the stream. The data is stored or transmitted to the synthesis module in a predetermined format, so that it can be read in a suitable manner. Any composite module that has a read module can read it.
[0102] In step E61, the synthesis module receives at its input, for example, the The view reconstructed by the multiview video decoder according to the described decoding method The synthesis module receives at least one image from each of the received images (VD1, VD2). Using the received views VD1 and VD2 and the read metadata MD1 and MD2, Generate at least one image from the imaginary viewpoint VS(1+2). In particular, metadata MD1 , MD2 are used by the synthesis module to determine the The compositing algorithm to be used to generate the virtual view images is determined or The view is determined.
[0103] FIG. 7 shows steps of a decoding method and an image processing method according to another particular embodiment of the present invention. This shows:
[0104] Generally, a multiview video decoder uses the synthesis equipment used to generate the virtual viewpoints. In other words, the decoder has no knowledge of the type of synthesis algorithm used. I don't know if it's possible to do this or what metadata types are useful to a decoder.
[0105] According to the particular embodiment described herein, the decoder and synthesis module , are adapted to be able to exchange in both directions. For example, the synthesis module ,decodes the list of metadata that the synthesis module needs to achieve better synthesis. Before or after a request from the synthesis module, the decoder can Informs the decoder module of the metadata that the decoder can send to the synthesis module. Advantageously, the list of metadata that the decoders can share can be are standardized, meaning that all decoders conforming to the decoding standard will be able to read the metadata on the list. Therefore, for a given decoding standard, it is necessary to share the data. Thus, the synthesis module knows what metadata is available. The list of data can also be adapted according to the profile of the decoder standard. For example: For profiles aimed at decoders requiring low computational complexity, the list of metadata is: It contains only the decoded syntax elements of the stream, while at the same time it has a higher computational complexity. For profiles for decoders that can handle it, the list of metadata is Decoded data from the stream, such as a segmentation map, occlusion map, or confidence map. It can also contain metadata that is derived by calculation from syntax elements.
[0106] At step E70, the synthesis module instructs the decoder to generate images from a virtual viewpoint. A request is sent indicating at least one item of metadata required to contains an index or list of indexes, each corresponding to the required metadata. nothing.
[0107] Such requests are made according to a predetermined format, i.e., the synthesis module and the decoder must It is transmitted according to a predetermined syntax so that it can be understood by anyone. For example, Such a syntax could be: nb For an integer i in the range 0 to nb-1, list[i] where the syntax element nb is the number of metadata required by the synthesis module. , which indicates the number of indices to be read by the decoder, and list[i ] indicates the index of each of the required metadata.
[0108] By way of example, taking the metadata example given by Table 1 above, the composite module The rule is to specify nb=2 and the camera parameters and occlusion map respectively in the request. The corresponding indices 0 and 9 can be shown.
[0109] According to one variant, the synthesis module calculates the index of the item of required metadata. associated with, for example, the "grlevel" syntax element associated with a metadata item. The level of granularity can also be indicated by specifying a predetermined value for the occlusion. In the case of a saturation map, the compositing module computes the occlusion map at the pixel level. If you want a "level" element with a value of 1 associated with index 9, or a coarser level, For example, you want an occlusion map for a block of size 8x8 at a low level. could indicate a value of 2 for the "level" element associated with index 9.
[0110] In step E71, the decoder retrieves the corresponding metadata. For this purpose, According to the example described above in connection with FIG. 4 or FIG. 5, the decoder It finds the decoded syntax elements needed to generate the replay, such as the occlusion map. Calculate the metadata that is not used by the decoder for construction. Then, The metadata is stored in a file according to a given format so that the synthesis module can read it. It will be formatted.
[0111] In step E72, the decoder transmits the metadata to the synthesis module, after which ,The synthesis module can use the metadata in its synthesis algorithms. .
[0112] FIG. 8 shows a block diagram of a digital video decoder for implementing the decoding method according to the above-described specific embodiment of the present invention. 1 shows a diagram of the adapted device DEC.
[0113] Such a decoding device comprises a memory MEM and, for example, a processor PROC. A processing unit UT controlled by a computer program PG stored in a memory The computer program PG is configured to include the following: instructions that, when executed by the ROC, implement the steps of the decoding method described above. include.
[0114] According to a particular embodiment of the invention, the decoding device DEC is, inter alia, characterized in that: receiving a data stream representing multi-view video over a communications network It has a communication interface COM0 that enables
[0115] According to another particular embodiment of the invention, the decoding device DEC is adapted to decode the composite model. The metadata is transmitted to an image processing device such as a module and reconstructed from the data stream. It has a communication interface COM1 that allows transmitting the images of the generated view. .
[0116] At initialization, the code instructions of the computer program PG are transmitted to, for example, the processor PRO It is loaded into memory before being executed by C. In particular, the processor of the processing unit UT PROC is a program according to the instructions of the computer program PG in relation to FIGS. 4, 5 and 7. The memory MEM carries out the steps of the described decoding method, among others, storing the data of a predetermined format and adapted to store the metadata obtained during the decoding method.
[0117] According to a particular embodiment of the invention, the above-described decoding device is suitable for use in a television receiver, a mobile mobile phones (e.g., smartphones), set-top boxes, virtual reality headsets, etc. Included in the device.
[0118] FIG. 9 illustrates a block diagram of a computer system for implementing an image processing method according to the above-described specific embodiment of the present invention. 1 shows diagrammatically the adapted device SYNTH.
[0119] Such a device comprises a memory MEM9 and, for example, a processor PROC9, A processing unit UT controlled by a computer program PG9 stored in EM9 The computer program PG9 is configured by the program When executed by PROC9, it performs the steps of the image processing method as described above. This includes an order to
[0120] According to a particular embodiment of the invention, the device SYNTH is a device as described above. It receives metadata sent from a decoding device such as DEC, and also receives the metadata sent from the device DEC. receiving images of the reconstructed views from a data stream representing a multi-view video; It is equipped with a communication interface COM9 that enables
[0121] At initialization, the code instructions of the computer program PG9 are written to the processor PR It is loaded into memory before being executed by OC9. In particular, the process of the processing unit UT9 Processor PROC9 executes the steps associated with FIGS. 6 and 7 according to the instructions of computer program PG9. The image processing method steps described above are carried out.
[0122] According to a particular embodiment of the invention, the device SYNTH is Output interface AFF9 that allows sending images to a device, e.g. a screen For example, such an image may be a combination of the images of the reconstructed views and the images received from the device DEC. The image from the virtual viewpoint is generated by the device SYNTH using the transmitted metadata. We can respond to this.
[0123] According to a particular embodiment of the invention, the device SYNTH is a synthesis module. Modules include television sets, mobile phones (e.g., smartphones), set-top boxes, devices such as smartphones, tablets, and virtual reality headsets.
[0124] The principles of the present invention are described in the case of a multi-view video decoding system, where multiple views are decoded from the same stream (bitstream) and metadata is obtained for each view. The principles equally apply when multi-view video is encoded using multiple streams (bitstreams), with one view encoded per stream. In this case, each view decoder provides metadata associated with the view it decodes. In order to maintain the disclosure of the present application as originally filed, the contents of claims 1 to 14 as originally filed are added below. (Claim 1) 1. A method for decoding a data stream representing multi-view video, performed by a decoding device, comprising: Obtaining syntax elements from at least one portion of said data stream (E20); - reconstructing (E21) at least one image of a view of said video from said obtained syntax elements; and the decoding method comprises: Obtaining at least one item of metadata in a predetermined format from at least one syntax element (E23); providing (E24) at least one item of said metadata to an image synthesis module; 20. The decoding method according to claim 19, further comprising: (Claim 2) The method of decoding of claim 1 , wherein obtaining at least one item of metadata further comprises computing the at least one item of metadata from at least one portion of the syntax element. (Claim 3) 3. A method of decoding according to claim 1 or 2, wherein at least one item of said metadata is not used to reconstruct said at least one image. (Claim 4) At least one item of said metadata is one of the following: camera parameters, the decoded and scaled motion vectors, Segmentation of the reconstructed image; a reference image used by a block of an image of said reconstructed view; a coding mode of the image of the reconstructed view; a quantization parameter value of the image of the reconstructed view; a prediction residual value of the image of the reconstructed view; a map representing motion within the image of said reconstructed view; a map representing the presence of occlusions in the image of the reconstructed view; a map representing confidence values associated with the depth map; 4. The decoding method according to claim 1, wherein the decoding method corresponds to an item of information contained in a group including: (Claim 5) 5. The decoding method according to claim 1, wherein the predetermined format corresponds to an indexed table in which at least one item of metadata is stored in association with an index. (Claim 6) 6. The decoding method according to claim 1, wherein at least one item of the metadata is acquired based on a granularity level specified in the decoding device. (Claim 7) 7. The decoding method of claim 1, further comprising receiving, by said decoding device, a request from said image synthesis module indicating at least one item of metadata required by said image synthesis module. (Claim 8) 8. The method of claim 7, wherein the request includes at least one index indicating the required metadata item from a predetermined list of available metadata. (Claim 9) 1. An apparatus for decoding a data stream representing a multi-view video, comprising: The device comprises: obtaining syntax elements from at least one portion of the data stream; reconstructing at least one image of a view of the video from the retrieved syntax elements; It is configured as follows (UT, MEM, COM1), The decoding device obtaining at least one item of metadata in a predetermined format from at least one syntax element; providing at least one item of said metadata to an image synthesis module; The decoding device is further configured to: (Claim 10) 1. A method of image synthesis, comprising generating at least one image of a virtual view from at least one image of a view decoded by a decoding device, the method comprising: - reading (E60) at least one item of metadata in a predetermined format, said at least one item of metadata being obtained by said decoding device from at least one syntax element obtained from a data stream representing multi-view video; generating said at least one image (E61) including using at least one retrieved item of said metadata; An image synthesis method comprising: (Claim 11) 11. The image synthesis method of claim 10, further comprising sending to said decoding device a request indicating at least one item of metadata required to generate said image. (Claim 12) an image synthesis device configured to generate at least one image of a virtual view from at least one image of a view decoded by a decoding device, the image synthesis device comprising: the image synthesis device is configured to read at least one item of metadata in a predetermined format (UT9, MEM9, COM9), the at least one item of metadata being obtained by the decoding device from at least one syntax element obtained from a data stream representing multi-view video; and said at least one retrieved item of metadata is used when said at least one image is generated; An image synthesis device comprising: (Claim 13) 1. An image processing system for displaying multi-view video from a data stream representing said multi-view video, comprising: a decoding device according to claim 9; an image synthesis device according to claim 12; An image processing system comprising: (Claim 14) A computer program comprising instructions for implementing the decoding method according to any one of claims 1 to 8 or the image synthesis method according to claim 10 or 11 when executed by a processor.
Claims
1. 1. A method for decoding a data stream representing multi-view video, performed by a decoding device, comprising: Obtaining (E20) syntax elements from at least one portion of said data stream; - reconstructing (E21) at least one image of a view of the multiview video from the obtained syntax elements; and the decoding method comprises: - obtaining (E23) from at least one syntax element at least one item of metadata in a predetermined format, said at least one item of metadata being associated with said at least one reconstructed image and corresponding to a quantization parameter value of said at least one reconstructed image; providing (E24) said at least one item of metadata to an image synthesis module, said image synthesis module being configured to synthesize from said at least one reconstructed image and said at least one item of metadata at least one virtual view different from a view of said multiview video; 2. A decoding method, comprising:
2. The method of decoding of claim 1 , wherein obtaining at least one item of metadata further comprises computing the at least one item of metadata from at least one portion of the syntax element.
3. 3. A method of decoding according to claim 1 or 2, wherein at least one item of said metadata is not used to reconstruct said at least one image.
4. 4. A decoding method according to claim 1 or 3, wherein the predetermined format corresponds to an indexed table, in which at least one item of metadata is stored in association with an index.
5. 5. A decoding method according to claim 1 or 4, wherein the at least one item of metadata is obtained based on a specified level of granularity in the decoding device.
6. 6. A method of decoding as claimed in claim 1 or 5, further comprising receiving, by said decoding device, a request from said image synthesis module indicating at least one item of metadata required by said image synthesis module.
7. 7. A method of decoding according to claim 6, wherein the request includes at least one index indicating the required metadata item from a predetermined list of available metadata.
8. 1. An apparatus for decoding a data stream representing a multi-view video, comprising: The device comprises: obtaining syntax elements from at least one portion of the data stream; reconstructing at least one image of a view of the multiview video from the retrieved syntax elements; (UT, MEM, COM1), The decoding device obtaining at least one item of metadata in a predetermined format from at least one syntax element, said at least one item of metadata being associated with said at least one reconstructed image and corresponding to a quantization parameter value of said at least one reconstructed image; providing the at least one item of metadata to an image synthesis module, the image synthesis module synthesizing at least one virtual view different from a view of the multi-view video from the at least one reconstructed image and the at least one item of metadata; The decoding device is further configured to:
9. 1. A method of image synthesis, comprising generating, from at least one image of a view of a multi-view video decoded by a decoding device, at least one image of a virtual view that is different from a view of the multi-view video, the image synthesis method comprising: - reading (E60) at least one item of metadata in a predetermined format, wherein said at least one item of metadata is obtained by said decoding device from at least one syntax element obtained from a data stream representing a multi-view video, said at least one item of metadata corresponding to a quantization parameter value of said at least one reconstructed image; generating said at least one image (E61) including using at least one retrieved item of said metadata; An image synthesis method comprising:
10. 10. The method of claim 9, further comprising sending to said decoding device a request indicating at least one item of metadata required to generate said image.
11. 1. An image synthesis device configured to generate, from at least one image of a view of a multi-view video decoded by a decoding device, at least one image of a virtual view that is different from a view of the multi-view video, the image synthesis device comprising: the image synthesis device reads at least one item of metadata in a predetermined format, the at least one item of metadata being obtained by the decoding device from at least one syntax element obtained from a data stream representing multi-view video, the at least one item of metadata corresponding to a quantization parameter value of the at least one reconstructed image; generating the at least one image including using the at least one retrieved item of metadata. An image synthesis device configured as follows (UT9, MEM9, COM9).
12. 1. An image processing system for displaying multi-view video from a data stream representing said multi-view video, comprising: a decoding device according to claim 8; The image synthesis device according to claim 11; An image processing system comprising:
13. A computer program comprising instructions which, when executed by a processor, perform the decoding method according to any one of claims 1 to 7 or the image synthesis method according to claim 9 or 10.
Citation Information
Patent Citations
Stereoscopic video encoding device, stereoscopic video decoding device, stereoscopic video encoding method, stereoscopic video decoding method, stereoscopic video encoding program, and stereoscopic video decoding program
JP2014132721A
Camera and / or depth parameter signaling
JP2014528190A
Controller, control method, and program
JP2017212592A
JPP7371090B
JPP7569434B