Image processing device and image encoding and decoding method
The method and device address inefficiencies in compressing and transmitting large-scale light field images by selecting reference frames and using multi-scale interpolation, achieving reduced data transmission and high-quality image reconstruction.
Patent Information
- Application Number
- PCT/KR2025/000399
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-23
- Filing Date
- 2025-01-08
- Publication Date
- 2025-08-28
AI Technical Summary
Existing methods are inefficient for compressing, transmitting, and restoring large-scale light field images due to their high data requirements and the need for optimizing synthesis quality for each scene.
A method and device that select reference frames and encode only these frames along with additional information, using pixel warping-based multi-scale frame interpolation to generate composite frames, reducing data transmission and enabling high-quality image reconstruction.
Efficiently compresses and transmits large-capacity light field images by reducing data volume while maintaining high image quality through optimized synthesis, allowing for accurate reconstruction of composite frames.
Smart Images

Figure KR2025000399_28082025_PF_FP_ABST
Abstract
Description
Image processing device and method for encoding and decoding images
[0001] The present disclosure relates to a method for compressing, transmitting and restoring large-capacity light field images.
[0002] A light field is a field used to express the intensity and direction of light reflected from a subject in three-dimensional space. Light fields, a new approach to three-dimensional image processing, rely on large amounts of data. Therefore, a method is needed to efficiently compress, transmit, and restore large-scale light field images.
[0003] According to one aspect of the present disclosure, a method for decoding and outputting light field images may be provided. The method may include obtaining an encoded first frame and an encoded second frame, which represent reference frames captured from different viewpoints, and an encoded synthesis parameter and an encoded occlusion map for generating a composite frame. The method may include decoding the encoded first frame, the encoded second frame, the encoded synthesis parameter, and the encoded occlusion map to reconstruct the first frame, the second frame, the disparity vector, the synthesis weight, and the occlusion map. The method may include generating a composite frame based on the first frame, the second frame, the disparity vector, the synthesis weight, and the occlusion map. The method may include outputting light field images including the first frame, the second frame, and the composite frame.
[0004] According to one aspect of the present disclosure, an image processing device may be provided. The image processing device may include a memory storing one or more instructions; and one or more processors executing the one or more instructions stored in the memory. The one or more processors may, by executing the one or more instructions, obtain an encoded first frame and an encoded second frame representing reference frames captured from different viewpoints, and an encoded synthesis parameter and an encoded occlusion map for generating a synthesis frame. The one or more processors may, by executing the one or more instructions, decode the encoded first frame, the encoded second frame, the encoded synthesis parameter, and the encoded occlusion map to reconstruct the first frame, the second frame, the disparity vector, the synthesis weight, and the occlusion map. The one or more processors can generate a composite frame based on the first frame, the second frame, the disparity vector, the composite weight, and the occlusion map by executing the one or more instructions. The one or more processors can output light field images including the first frame, the second frame, and the composite frame by executing the one or more instructions.
[0005] According to one aspect of the present disclosure, a method for encoding light field images may be provided. The method may include a step of selecting reference frames including a first frame and a second frame from among original light field images including a plurality of viewpoints. The method may include a step of generating a disparity vector, a synthesis weight, and an occlusion map for inferring frames between the first frame and the second frame. The method may include a step of encoding the first frame, the second frame, the disparity vector, the synthesis weight, and the occlusion map.
[0006] FIG. 1 is a drawing for explaining a light field image generated by a decoding device according to one embodiment of the present disclosure.
[0007] FIG. 2 is a drawing for explaining an operation of an image processing device according to one embodiment of the present disclosure to acquire light field images.
[0008] FIG. 3 is a diagram for explaining the encoding and decoding process of light field images according to one embodiment of the present disclosure.
[0009] FIG. 4 is a diagram for explaining a compression transmission method of light field images according to one embodiment of the present disclosure.
[0010] FIG. 5 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate a composite frame.
[0011] FIG. 6 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate a composite frame.
[0012] FIG. 7 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate a composite frame.
[0013] FIG. 8 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to optimize a synthesis weight and a viewpoint interval of a reference frame.
[0014] FIG. 9 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to optimize a synthesis weight and a viewpoint interval of a reference frame.
[0015] FIG. 10 is a diagram for explaining the encoding and decoding operations of an occlusion map used in the encoding and decoding method of the present disclosure.
[0016] FIG. 11 is a drawing for explaining a key occlusion map according to one embodiment of the present disclosure.
[0017] FIG. 12 is a diagram for explaining an operation of generating occlusion maps based on a key occlusion map according to one embodiment of the present disclosure.
[0018] FIG. 13 is a drawing for explaining an occlusion map according to one embodiment of the present disclosure.
[0019] FIG. 14 is a block diagram illustrating a configuration of an image processing device according to one embodiment of the present disclosure.
[0020] FIG. 15 is a block diagram illustrating a configuration of an image processing device according to one embodiment of the present disclosure.
[0021] The terms used in this specification will be briefly explained, followed by a detailed description of the present disclosure. The terms used in this disclosure have been selected from widely used and common terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, in which case their meanings will be described in detail in the relevant description. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.
[0022] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein. Furthermore, terms containing ordinal numbers, such as "first" or "second," used herein may be used to describe various components, but such components should not be limited by such terms. Such terms are used solely to distinguish one component from another.
[0023] When a part of the specification is said to "include" a component, unless otherwise specifically stated, this does not exclude other components but rather implies the inclusion of other components. Furthermore, terms such as "part" and "module" used in the specification refer to a unit that processes at least one function or operation, which may be implemented in hardware, software, or a combination of hardware and software.
[0024] Below, with reference to the attached drawings, embodiments of the present disclosure are described in detail so that those skilled in the art can easily practice the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, portions irrelevant to the description have been omitted for clarity of explanation, and similar reference numerals have been used throughout the specification to designate similar parts.
[0025] The present disclosure will be described below with reference to the attached drawings.
[0026] FIG. 1 is a drawing for explaining a light field image generated by a decoding device according to one embodiment of the present disclosure.
[0027] Referring to FIG. 1, restored light field images (100) generated by a decoding device are illustrated. The restored light field images (100) may be composed of one or more reference frames (110) and one or more composite frames (120).
[0028] In one embodiment, the restored light field images (100) may be generated by receiving encoded data based on original light field images from an encoding device and restoring the same. The original light field images may be images corresponding to multiple viewpoints. For example, the original light field images composed of N images may include a total of N viewpoints, with each image being acquired from a different viewpoint. The original light field images may be images acquired by multiple cameras at different viewpoints. Alternatively, the original light field images may be images acquired by a single camera at different viewpoints while moving.
[0029] In one embodiment, the original light field images are encoded in an encoding device. In this case, not all of the original light field images are encoded, but only some images are selected, encoded, and transmitted to the decoding device. Furthermore, the remaining images, other than the selected images, are encoded as additional information rather than encoding the images themselves and transmitted to the decoding device.
[0030] The reference frame (110) may refer to images selected from among the original light field images. The reference frame (110) is encoded in an encoding device and transmitted to a decoding device. The decoding device can restore the reference frame (110) by decoding the encoded reference frame (110). As a result, examples of reference frames (110) restored in the decoding device, namely, a first reference frame (112), a second reference frame (114), and a third reference frame (116), are illustrated in FIG. 1.
[0031] Meanwhile, the reference frame (110) may be referred to by various expressions representing the same or similar concepts. The reference frame (110) may be referred to by expressions such as a reference image, a reference view, a key frame, an anchor frame, an anchor view, etc., and is not limited to the examples described above.
[0032] The composite frame (120) may refer to images generated through a synthesis process in a decoding device, which are images corresponding to the remaining frames other than the encoded and transmitted reference frames (110) among the original light field images. The decoding device may generate composite frames (120) having viewpoints between the reference frames (110) by referring to the reference frames (110). For example, the decoding device may generate a first composite frame (122), ..., second composite frame (124), which are composite frames (120) having viewpoints between the first reference frame (112) and the second reference frame (114), based on a first reference frame (112) and a second reference frame (114). Additionally, for example, the decoding device can generate third composite frames (126), which are composite frames (120) having points of view between the second reference frame (114) and the third reference frame (116), based on the second reference frame (114) and the third reference frame (116).
[0033] Meanwhile, the synthetic frame (120) may be referred to by various expressions representing the same / similar concepts. The synthetic frame (120) may be referred to by expressions such as synthetic frame, synthetic view, predicted frame, predicted image, predicted view, inferred frame, inferred image, inferred view, etc., and is not limited to the examples described above.
[0034] In one embodiment, when the decoding device generates composite frames (120), additional information (pixel movement information, composite weight information, etc.) transmitted from the encoding device may be utilized.
[0035] An encoding / decoding method according to one embodiment relates to a method for efficiently compressing, transmitting, and restoring large-capacity original light field images. Large-capacity original light field images have different viewpoints, but there are overlapping portions between the images. Therefore, compressing and transmitting all images is a relatively inefficient method.
[0036] The encoding device can compress and transmit only the reference frames (110), which are some images selected using the encoding method of the present disclosure, and compress and transmit additional information corresponding to each image for the remaining images. Since the encoded additional information is a relatively small amount of data compared to the encoded frame, the amount of data transmitted from the encoding device to the decoding device can be reduced. In addition, the encoding device can generate additional information optimized for each scene through an optimization process. By compressing and transmitting the additional information, the encoding device can enable the decoding device to generate high-quality composite frames (120).
[0037] A decoding device can restore data received from an encoding device using the decoding method of the present disclosure. Furthermore, the decoding device can generate a composite frame (120) using the pixel warping-based multi-scale frame interpolation method of the present disclosure. In this case, the composite frame (120) can be generated based on a restored reference frame (110) and restored additional information.
[0038] Meanwhile, in the present disclosure, the encoding device and decoding device may also be referred to as "image processing devices." For example, separate image processing devices may each perform encoding and decoding, or a single image processing device may perform both encoding and decoding.
[0039] The encoding and decoding methods and the operations of the encoding and decoding devices of the present disclosure will be described in more detail through the drawings and descriptions thereof described below.
[0040] FIG. 2 is a drawing for explaining an operation of an image processing device according to one embodiment of the present disclosure to acquire light field images.
[0041] In explaining the operations of the method for decoding light field images of the present disclosure in FIG. 2, the decoding device, which is the subject performing the decoding operation, will be referred to as an “image processing device.”
[0042] In one embodiment, an image processing device can decode and output light field images. Below, operations for decoding light field images are described. Light field images can be composed of images having multiple different viewpoints.
[0043] In operation S210, the image processing device can obtain an encoded first frame, an encoded second frame, an encoded synthesis parameter, and an encoded occlusion map.
[0044] In one embodiment, the image processing device can obtain encoded reference frames. The encoded reference frames may be received from the encoding device or encoded in the image processing device.
[0045] The image processing device can adaptively select reference frames, and the number of selected reference frames may be two or more. The following description will illustrate an example of two reference frames. Multiple different viewpoints may be included between the reference frames.
[0046] The encoded first frame and the encoded second frame represent reference frames captured from different viewpoints. For example, the first frame may be an image of a first viewpoint obtained by capturing an object with a camera positioned at a first location, and the second frame may be an image of a second viewpoint obtained by capturing the object with a camera positioned at a second location. A plurality of different viewpoints may be further included between the first viewpoint and the second viewpoint. The image processing device may generate composite frames corresponding to different viewpoints between the first viewpoint and the second viewpoint.
[0047] The encoded synthesis parameters represent parameters for generating a synthesis frame having a viewpoint between the first frame and the second frame. The encoded synthesis parameters may include pixel movement information and synthesis weight information.
[0048] In one embodiment, the encoded first frame, the encoded second frame, and the encoded synthesis parameters may be obtained using various types of video coding techniques. For example, HEVC (High Efficiency Video Coding) may be used as the encoding method. However, the present invention is not limited thereto, and other methods such as H.264, VVC (Versatile Video Coding), or AV1 (AOMedia Video) may also be used.
[0049] The encoded occlusion map represents an occluded area caused by an object between a first viewpoint corresponding to the first frame and a second viewpoint corresponding to the second frame. The encoded occlusion map may include confidence weight information for processing the occluded area. The encoded occlusion map may include information related to the occluded area (e.g., information on the number of occluded areas, information on the maximum length of the occluded area).
[0050] In one embodiment, the occlusion map can be encoded and decoded using an occlusion map encoder / decoder.
[0051] In operation S220, the image processing device can decode the encoded first frame, the encoded second frame, the encoded synthesis parameters, and the encoded occlusion map to restore the first frame, the second frame, the disparity vector, the synthesis weights, and the occlusion map.
[0052] In one embodiment, the encoded first frame and the encoded second frame can be decoded and reconstructed into the first frame and the second frame. The encoded synthesis parameters can be decoded and reconstructed into a disparity vector and a synthesis weight. The disparity vector represents pixel movement information from the reference frame, and the synthesis weight represents a weight used for image combination. The encoded occlusion map can be decoded and reconstructed into an occlusion map. The decoding method can be performed by applying a decoding method corresponding to the encoding method.
[0053] In operation S230, the image processing device can generate a composite frame based on the first frame, the second frame, the disparity vector, the composite weight, and the occlusion map.
[0054] In one embodiment, the image processing device can generate warped images based on the restored reference frame and the disparity vector. For example, if the image processing device wants to generate a composite frame corresponding to point n between a first point and a second point, the image processing device can generate a first warped image in which pixels are shifted from the first point, which is a reference point, to point n using the first frame and the disparity vector, and can generate a second warped image in which pixels are shifted from the second point, which is a reference point, to point n using the second frame and the disparity vector.
[0055] An image processing device can combine warped images to create a composite frame.
[0056] In one embodiment, the image processing device can generate warped images of multiple scales. For example, the image processing device can generate warped images of a first scale and warped images of a second scale. The image processing device can independently combine the warped images of each scale to generate intermediate images. Independently combining means that when an intermediate image of a particular scale is generated, warped images of other scales are not involved. For example, the image processing device can combine warped images of a first scale to generate a first intermediate image, and combine warped images of a second scale to generate a second intermediate image. The image processing device can generate a composite frame by combining the intermediate images. For example, the first intermediate image and the second intermediate image can be combined to generate a composite frame.
[0057] In operation S240, the image processing device can output light field images including a first frame, a second frame, and one or more composite frames.
[0058] In one embodiment, the image processing device can transmit light field images to an external device (e.g., a user device, etc.). The image processing device can transmit the light field images so that the light field images can be displayed on the external device.
[0059] In one embodiment, if the image processing device includes a display, the image processing device can display the restored light field images on a screen of the image processing device.
[0060] FIG. 3 is a diagram for explaining the encoding and decoding process of light field images according to one embodiment of the present disclosure.
[0061] Referring to FIG. 3, the encoding device (302) compresses and transmits large-capacity light field images, and the decoding device (304) can restore encoded data received from the encoding device (302) back into light field images. The decoded light field images can be displayed to the user.
[0062] In one embodiment, the encoding device (302) can acquire a large amount of light field images, compress the light field images, and transmit them. The light field images may be, for example, N viewpoint images (e.g., 60 views) of 4K resolution, but are not limited thereto. When transmitting the light field images, the encoding device (302) may not compress and transmit all images, but may compress and transmit only reference frames. In addition, additional information for generating the remaining frames other than the reference frames may be generated, and only the additional information may be transmitted to the decoding device (304). Accordingly, the decoding device (304) can generate composite frames corresponding to the remaining frames.
[0063] The encoding device (302) may include various modules for performing the functions of reference viewpoint selection (310), disparity estimation (320), encoding (330), and occlusion map encoding (340). The operations of each module are described below. Each module may be implemented as a program, command, or code for implementing specific functions, and may be executed by the processor of the encoding device (302).
[0064] The reference viewpoint selection (310) operation selects reference frames from among light field images. For example, if the light field images are images corresponding to N viewpoints, reference viewpoints corresponding to M viewpoints (N>M) from among the N viewpoints can be selected.
[0065] In one embodiment, reference frames may be selected based on a preset view stride. For example, frames at every view stride V may be selected as reference frames. In one embodiment, the preset view stride may be adaptively adjusted. The encoding device (302) may evaluate the synthesis quality of the synthesized frame and increase or decrease the view stride V for selecting reference frames based on the synthesis quality.
[0066] The disparity estimation (320) operation produces a disparity vector and a synthesis weight. The disparity vector is used to generate a warped image by shifting pixel information based on a reference frame, and the synthesis weight can be used to generate an intermediate image by combining warped images or to generate a composite frame by combining intermediate images.
[0067] In one embodiment, the disparity vector and synthesis weights may be optimized for each scene. Instead of directly applying a pre-trained model, the encoding device (302) may perform scene-specific optimization to find the optimal disparity vector and synthesis weight for each input (light field image). Through scene-specific optimization, the encoding device (302) may produce the optimal disparity vector and weights suitable for the input data.
[0068] The encoding (340) operation can encode a reference frame, a disparity vector, and a synthesis weight. For example, HEVC (High Efficiency Video Coding) can be used as the encoding method. However, the present invention is not limited thereto, and other methods such as H.264, VVC (Versatile Video Coding), or AV1 (AOMedia Video) can also be used.
[0069] The occlusion map encoding (330) operation encodes an occlusion map. The occlusion map may include reliability weights that enable processing of occlusion areas without pixel information when generating a composite frame. The occlusion map may be generated based on original data (input light field images). For example, information about occlusion areas may be generated by shifting pixels of frames at each viewpoint to pixel locations at other viewpoints and then re-shifting them to their original locations. Based on the occlusion map, reliability weights used when generating a composite frame may be generated. The reliability weight may mean that the pixels used for synthesis when using the composite frame can be trusted. In other words, it may mean that the pixels corresponding to the location of the reliability weight can be used when generating the composite frame.
[0070] The encoding device (302) can transmit the encoded reference frame, the encoded disparity vector, the encoded synthesis weights, and the encoded occlusion map to the decoding device (304). The encoded disparity vector and the encoded synthesis weights may be referred to as encoded synthesis parameters.
[0071] In one embodiment, the decoding device (304) can reconstruct light field images by generating reference frames and synthesized frames. The decoding device (304) may include various modules for performing the functions of decoding (350), occlusion map decoding (360), and adaptive view synthesis (370). The operations of each module are described below. Each module may be implemented as a program, instruction, or code for implementing specific functions, and may be executed by the processor of the decoding device (304).
[0072] The decoding (350) operation decodes the data received from the encoding device (302) to restore the reference frame, disparity vector, and synthesis weights. The decoding method of the decoding (350) operation can be performed by applying a decoding method corresponding to the encoding method. For example, if the encoding method is HEVC, the decoding method can also be HEVC.
[0073] The occlusion map decoding (360) operation restores the encoded occlusion map received from the encoding device (302).
[0074] The adaptive view synthesis (370) operation synthesizes frames corresponding to viewpoints between reference frames. There may be one or more synthesized frames generated.
[0075] The decoding device (304) can generate multi-scale warped images based on reference frames and disparity vectors. The decoding device (304) can combine multi-scale warped images to generate multiple multi-scale intermediate images. In this case, a synthesis weight can be used in the process of generating the intermediate images. The decoding device (304) can combine multiple intermediate images to generate a composite frame. In this case, a synthesis weight can be used in the process of generating the composite frame.
[0076] FIG. 4 is a diagram for explaining a compression transmission method of light field images according to one embodiment of the present disclosure.
[0077] Referring to FIG. 4, the encoding device can obtain input light field images (410). For example, the input light field images (410) can be images corresponding to 60 viewpoints at 4K resolution.
[0078] In one embodiment, the encoding device can select reference frames from among the input light field images (410). The encoding device can encode the reference frames to generate an encoded reference frame (420).
[0079] In one embodiment, the encoding device may generate meta information (430) corresponding to the remaining frames other than the reference frames. The meta information may include information for representing viewpoints between the reference frames. The meta information (430) may include, for example, a disparity vector indicating pixel movement information. A warped image representing an image moved by n viewpoints from the reference frame may be obtained by moving pixels of the reference frame based on the disparity vector. The meta information (430) may include, for example, a synthesis weight. The synthesis weight may include a first synthesis weight used to synthesize the warped image to obtain an intermediate image and a second synthesis weight used to synthesize the intermediate image to obtain a synthesized frame. The meta information (430) may include, for example, an occlusion map. The occlusion map may include a reliability weight representing a reliability of pixels used when generating the synthesized frame. The encoding device may encode the meta information (430).
[0080] Since the encoding device transmits the encoded reference frames (420) and encoded meta information (430) to the decoding device, a smaller amount of data can be transmitted to the decoding device than if all of the input light field images (410) were encoded and transmitted. In other words, the amount of data transmitted can be reduced.
[0081] In one embodiment, the decoding device can obtain a restored reference frame (440) by decoding the encoded reference frame (420). In addition, the decoding device can obtain restored meta information (450) by decoding the encoded meta information (430). The restored meta information (450) may include, but is not limited to, a disparity vector, a synthesis weight, and an occlusion map, for example.
[0082] The decoding device can generate composite frames through a view synthesis process.
[0083] The decoding device can generate a warped image representing an image shifted by n points in time from the reference frame based on the restored reference frame (440) and the disparity vector. There may be multiple warped images. In addition, the warped images may be multi-scale images having different scales (e.g., a first scale, a second scale, ..., a kth scale).
[0084] The decoding device can synthesize multiple warped images to generate intermediate images. For example, a first intermediate image can be generated using warped images of a first scale, and a second intermediate image can be generated using warped images of a second scale. A synthesis weight can be applied to the warped images to generate the intermediate images. In addition, a confidence weight included in the occlusion map can be applied to generate the intermediate images.
[0085] A decoding device can generate a composite frame by synthesizing a plurality of intermediate images. The intermediate images may be images of multiple scales. The decoding device can adjust the scale of the intermediate images to match the scale of the final composite frame and synthesize the intermediate images to generate the composite frame. For example, a first intermediate image of a first scale and a second intermediate image of a second scale may be adjusted to be unified to a predetermined scale and synthesized to generate a composite frame. A synthesis weight may be applied to the intermediate images to generate the composite frame.
[0086] In one embodiment, the decoding device can output output light field images (460). The output light field images (460) can include restored reference frames (440) and composite frames, which are images corresponding to points in time between the reference frames (440). In other words, even if not all images of the input light field images (410) are encoded and transmitted by the encoding device, since the decoding device can generate composite frames, all images of the output light field images (460) can correspond to all images of the original input light field images (410). In addition, since not all of the input light field images (410) are encoded and transmitted, the amount of computation required for decoding can be reduced compared to decoding all of the input light field images (410).
[0087] FIG. 5 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate a composite frame.
[0088] In describing FIG. 5, the image processing device may be a decoding device. In addition, when the first reference frame (510) is the 0th image at the left viewpoint and the second reference frame (520) is the Nth image at the right viewpoint, the nth image between the first reference frame (510) and the second reference frame (520) is synthesized as an example.
[0089] In one embodiment, the image processing device can generate a composite frame using a pixel-based multi-scale synthesis method that combines pixel values into independent flows of multiple scales. The pixel-based multi-scale synthesis method refers to a method of warping pixels to create and synthesize images of multiple scales. The image processing device can generate the composite frame using multi-scale disparity vectors and multi-scale synthesis weights corresponding to each of the multiple scales. The composite frame can be an image having an intermediate viewpoint of reference frames.
[0090] An image processing device can generate warped images using restored reference frames and disparity vectors. The warped images may be multi-scale warped images. In this case, multi-scale warped images may be generated using multi-scale disparity vectors. For example, the number of multi-scales Ns may be 3.
[0091] An image processing device can generate warped images based on a first reference frame (510) and multi-scale first disparity vectors (512, 514, 516). For example, the image processing device can generate first warped images by warping a viewpoint of the first reference frame (510) from 0 to n using the first disparity vectors (512, 514, 516) for moving a viewpoint from 0 to n. The first disparity vectors (512, 514, 516) for generating multi-scale warped images based on the first reference frame (510) can have different scales. For example, the first disparity vector A (512) may have a scale of 0 (e.g., 1x), the first disparity vector B (514) may have a scale of 1 (e.g., 0.5x), and the first disparity vector C (516) may have a scale of 2 (e.g., 0.25x).
[0092] In the same manner, the image processing device can generate warped images based on the second reference frame (520) and the second disparity vectors (522, 524, 526) of multiple scales. For example, the image processing device can generate second warped images by warping the viewpoint of the second reference frame (520) from N to n using the second disparity vectors (522, 524, 526) for moving the viewpoint from N to n. The second disparity vectors (522, 524, 526) for generating the multi-scale warped images can have different scales. For example, the second disparity vector A (522) may have a scale of 0 (e.g., 1x), the second disparity vector B (524) may have a scale of 1 (e.g., 0.5x), and the second disparity vector C (526) may have a scale of 2 (e.g., 0.25x).
[0093] An image processing device can generate multi-scale intermediate images based on multi-scale disparity vectors. For example, the image processing device can generate a first intermediate image (532) by combining a first warped image A and a second warped image A. The first intermediate image (532) may be an image with a scale of 0 (e.g., 1x). In addition, the image processing device can generate a second intermediate image (534) having a size of a scale of 1 (e.g., 0.5x) by combining the first warped image B and the second warped image B, and can generate a third intermediate image (536) having a size of a scale of 2 (e.g., 0.25x) by combining the first warped image C and the second warped image C.
[0094] The image processing device can generate a composite frame (540) based on intermediate images of multiple scales. For example, the image processing device can generate the composite frame (540) by combining a first intermediate image (532), a second intermediate image (534), and a third intermediate image (536). In the example of FIG. 5, the generated composite frame (540) can be an image having an n-th viewpoint between 0 and N. The image processing device can restore the scale of the intermediate images to generate the composite frame (540). For example, since the second intermediate image (534) is an image with a scale of 0.5 times, the scale can be adjusted by a factor of 2 to match the scale of the first intermediate image (532) with the original scale. In addition, since the third intermediate image (536) is an image with a scale of 0.25 times, the scale can be adjusted by a factor of 4 to match the scale of the first intermediate image (532) with the original scale.
[0095] In one embodiment, the image processing device may utilize the synthesis weights when generating the synthesis frame (540).
[0096] The image processing device may apply a first synthesis weight when generating an intermediate image. For example, when the number of multi-scales Ns=3, the first synthesis weight may be applied to the disparity vectors corresponding to each scale s=0, 1, and 2. This can be expressed mathematically as follows.
[0097]
[0098] In the above equation, is the middle image, is the first warped image, is the first weight applied to the first warped image, is the second warped image, represents a second weight applied to the second warped image. The first weight may be different for each scale.
[0099] The image processing device may apply a second synthesis weight when generating a composite frame (540). For example, when the number of multi-scales Ns=3, the second synthesis weight may be applied to the intermediate images corresponding to each scale s=0, 1, and 2. This can be expressed mathematically as follows.
[0100]
[0101] In the above equation, is a composite frame (540), is the middle image, represents the second weight applied to the intermediate image, represents the restoration of the scale that was adjusted (2x, 4x). The second weight can be different for each scale.
[0102] In one embodiment, the image processing device may utilize confidence weights included in the occlusion map to generate an intermediate image. The application of confidence weights is described in the description of FIG. 7.
[0103] Meanwhile, in explaining FIG. 5, when the first reference frame (510) is an image corresponding to time point 0 and the second reference frame (520) is an image corresponding to time point N, an example in which a composite frame (540) corresponding to time point n, which is an intermediate time point between 0 and N, is generated has been described, but is not limited thereto. For example, composite frames corresponding to each of time points 1, 2, ..., n, ..., N-1 can be generated through pixel-based multi-scale synthesis in the same manner.
[0104] Meanwhile, the number of multi-scales and the multi-scale ratios described above are examples for convenience of explanation, and thus the number of multi-scales and the multi-scale ratios are not limited to the examples described above. The scale weights representing the scale ratios can be determined by optimization operations during the encoding process. The optimization operations are described in the descriptions of FIGS. 8 and 9.
[0105] FIG. 6 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate a composite frame.
[0106] In describing Fig. 6, the image processing device may be a decoding device. In addition, when the first reference frame is the 0th image from the left viewpoint and the second reference frame is the Nth image from the right viewpoint, the nth image between the first reference frame and the second reference frame is synthesized as an example.
[0107] In one embodiment, the image processing device may generate a plurality of intermediate images to generate a composite frame. The image processing device may generate a plurality of intermediate images at different scales and apply a composite weight to generate the composite frame.
[0108] For example, the image processing device may generate a first warped image A corresponding to the n viewpoint based on a 1x scale image (612) and a disparity vector of a first reference frame, and generate a second warped image A corresponding to the n viewpoint based on a 1x scale image (622) and a disparity vector of a second reference frame, and then combine the first warped image A and the second warped image A to generate a first intermediate image (632). The first intermediate image (632) may be an image corresponding to the n viewpoint and having a 1x scale. In the same manner, the image processing device can generate a second intermediate image (634) corresponding to the n time point and having a 0.5-fold scale based on the 0.5-fold scale images (614, 624) of the reference images, and can generate a third intermediate image (636) corresponding to the n time point and having a 0.25-fold scale based on the 0.25-fold scale images (616, 626) of the reference images.
[0109] Meanwhile, in the above example, the scale of the reference frame is first adjusted and then a warped image is generated as an example, but a warped image may be first generated based on the reference frame and then the scale of the warped image may be adjusted.
[0110] The image processing device can generate a composite frame by synthesizing intermediate images. For example, the image processing device can generate a composite frame through a weighted sum of a first intermediate image (632), a second intermediate image (634), and a third intermediate image (636).
[0111] An image processing device according to one embodiment can combine pixel values into independent flows of multiple scales. That is, the image processing device can synthesize a new image by referencing images of various scales, thereby improving the expressiveness of the image by reflecting nonlinear changes between viewpoints of the image.
[0112] Referring to Fig. 6, there is no reflected light in the images at point 0, and reflected light exists in the images at point N. For example, in Fig. 6, three light reflection areas are illustrated in each of the images at point N. The image processing device can utilize images of various scales for image synthesis to reflect the nonlinear change in reflected light existing in the images at point n.
[0113] To be more specific, when the first intermediate image (632), the second intermediate image (634), and the third intermediate image (632) are generated, the images are processed at different scales, so the areas referenced for image synthesis may be different. Referring to FIG. 6, the slightly different areas referenced when generating each intermediate image due to the different scales are represented by arrows. Therefore, in the example of FIG. 6, each of three different light reflection areas is referenced. However, each arrow is merely an exemplary illustration for the convenience of explanation, and does not limit the reference to the area pointed by the actual arrow. When the image processing device uses multi-scale images for image synthesis, a synthesized frame with higher expressiveness can be obtained than when synthesizing a single-scale image.
[0114] In one embodiment, the image processing device may utilize scale weights representing multi-scale magnification. The scale weights may be determined by optimization operations during the encoding process. The optimization operations are further described with reference to FIGS. 8 and 9 .
[0115] Meanwhile, the image processing device can improve the quality of the combined images while being robust to errors by performing weighted combinations using optimized first synthesis weights when combining warped images to generate an intermediate image, and by performing weighted combinations using optimized second synthesis weights when combining intermediate images to generate a composite frame. The optimization process for the synthesis weights can be performed during the encoding process. This will be further described with reference to FIGS. 8 and 9.
[0116] FIG. 7 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate a composite frame.
[0117] In describing Fig. 7, the image processing device may be a decoding device. For convenience of explanation, Fig. 7 illustrates and describes only one scale among multi-scale pixel synthesis.
[0118] An image processing device can obtain encoded data from a transmission system (e.g., an encoding device). For example, the image processing device can obtain reference frame 0, reference frame N, disparity vectors and synthesis weights including pixel movement information from 0 to N, and disparity vectors and synthesis weights including pixel movement information from N to 0.
[0119] In one embodiment, the image processing device can generate a warped image based on a reference frame and a disparity vector. The warped image may also be referred to as a predicted image.
[0120] For example, a warped image 0, which is an image at point n predicted from reference frame 0, can be generated. The warped image 0 is generated by disparating pixels of reference frame 0 from n (0) based on a disparity vector containing pixel movement information from 0 to N. <n<N) 에 대응하도록 이동시킴으로써 획득될 수 있다. 같은 방식으로, 레퍼런스 프레임 N 및, N으로부터 0까지의 픽셀 이동 정보를 포함하는 디스패리티 벡터에 기초하여, 레퍼런스 프레임 N으로부터 예측된 n 시점의 이미지인 워핑된 이미지 N이 생성될 수 있다.
[0121] In one embodiment, the image processing device can generate a weight map. The weight map can include at least one of synthesis weight information and reliability weight information. The weight map can be of the same scale as the warped image. For example, when warped image 0 and warped image N are weight-combined to generate an intermediate image, weight map 0 (710) can be applied to warped image 0, and weight map N (720) can be applied to warped image N.
[0122] Since the synthesis weights have been described above in the descriptions of FIGS. 5 and 6, a repetitive description will be omitted for convenience of explanation, and only the occlusion map and confidence weights will be described below. The confidence weights are generated based on the occlusion map and may represent weights indicating that the pixels used for synthesis are reliable for processing occluded areas without pixel information.
[0123] An image processing device can obtain an occlusion map. The occlusion map may be aligned with a reference frame. For example, occlusion map information (712) aligned with reference frame 0 and occlusion map information (722) aligned with reference frame N may be obtained. The occlusion map information (712) aligned with reference frame 0 may include information about an occlusion area when a viewpoint moves from 0 to N, and the occlusion map information (722) aligned with reference frame N may include information about an occlusion area when a viewpoint moves from N to 0.
[0124] The reliability weight included in the weight map 0 (710) can be used to ensure that only reliable pixels are used when moving the reference frame 0 to point n using a disparity vector containing pixel movement information from 0 to N. For example, if the reliability weight is greater than or equal to a preset value, the pixel is a reliable pixel and thus can be processed as a valid pixel when creating a warped image. Alternatively, if the reliability weight is less than the preset value, the pixel is an unreliable pixel and thus can be unused or processed as a pixel with a relatively low weight when creating a warped image.
[0125] Similarly, the reliability weight included in the weight map N (720) can be used to ensure that only reliable pixels are used when moving the reference frame N to point n using a disparity vector including pixel movement information from N to 0, or that pixels with low reliability are given a relatively low weight.
[0126] In other words, the image processing device can generate a synthetic frame based on reference frames, disparity vectors, synthesis weights, and occlusion maps.
[0127] A more specific example of the operations performed when a composite frame is created at point n between 0 and N is as follows.
[0128] - Example 1
[0129] 1) A warped image 0 is generated based on a disparity vector containing pixel movement information from reference frame 0, 0 to N, and a reliability weight included in a weight map 0 (710).
[0130] 2) A warped image N is generated based on a reference frame N, a disparity vector containing pixel movement information from N to 0, and a reliability weight included in a weight map N (720).
[0131] 3) An intermediate image is generated based on the warped image 0, the synthetic weights included in the weight map 0 (710), the warped image N, and the synthetic weights included in the weight map N (710).
[0132] 4) A composite frame is generated by weighting and combining multiple intermediate images generated by performing 1) to 3) above as independent flows of multiple scales using a composite weight.
[0133] - Example 2
[0134] 1) A warped image 0 is generated based on a disparity vector containing pixel movement information from reference frame 0, 0 to N.
[0135] 2) A warped image N is generated based on a reference frame N and a disparity vector containing pixel movement information from N to 0.
[0136] 3) An intermediate image is generated based on the warped image 0, the synthesis weights and confidence weights included in the weight map 0 (710), the warped image N, the synthesis weights and confidence weights included in the weight map N (710).
[0137] 4) A composite frame is generated by weighting and combining multiple intermediate images generated by performing 1) to 3) above as independent flows of multiple scales using a composite weight.
[0138] The operations of the first or second example described above are examples of creating a composite frame at n points in time, and the image processing device can create composite frames corresponding to multiple points in time within the reference frames. For example, composite frames corresponding to each of the points in time between 0 and N, i.e., points in time 1, 2, ..., n, ..., N-1, can be created through pixel-based multi-scale synthesis in the same manner.
[0139] FIG. 8 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to optimize a synthesis weight and a viewpoint interval of a reference frame.
[0140] The optimization operations described below can be performed by an encoding device. In describing the optimization operations of the present disclosure in FIG. 8, the encoding device, which is the subject performing the optimization operations, will be referred to as an "image processing device."
[0141] In one embodiment, the image processing device can obtain original light field images (or videos). The original light field images can be images corresponding to N viewpoints.
[0142] In operation S810, the image processing device can select reference frames from among the original light field images.
[0143] The image processing device can select reference frames based on a preset time interval V. For example, among the original light field images corresponding to N time points, M (M <N)의 프레임들이 레퍼런스 프레임들로 선택될 수 있다.
[0144] In operation S820, the image processing device can predict warped images.
[0145] An image processing device can predict warped images corresponding to viewpoints between reference frames. The image processing device can predict disparity vectors for generating the warped images.
[0146] In operation S830, the image processing device can generate synthetic frames through pixel-based multi-scale synthesis.
[0147] An image processing device can generate an intermediate image by combining warped images. Furthermore, the intermediate images can be combined to generate a composite frame. The image processing device can apply scale weights indicating scale ratios when generating multi-scale warped images. The image processing device can apply composite weights when generating intermediate images and composite frames.
[0148] In operation S840, the image processing device can measure loss between original and synthesized frames and perform optimization.
[0149] The image processing device can obtain an original frame corresponding to the synthesized frame and measure the loss by comparing the synthesized frame with the original frame. For example, if the synthesized frame corresponds to the nth time point, the image processing device can obtain the original frame corresponding to the nth time point among the original light field images.
[0150] The image processing device can optimize parameters used for synthesis based on the loss between the original and the synthesis. The image processing device can update at least one of a predicted disparity vector, a scale weight indicating a multi-scale scale factor, a synthesis weight used when synthesizing an intermediate image based on warped images, and a synthesis weight used when synthesizing a synthesized frame based on the intermediate images.
[0151] The image processing device can repeatedly perform operations S830 to S840 until the loss value is minimized. The image processing device can generate multi-scale warped images based on the updated disparity vector and scale weights.
[0152] An image processing device can generate multi-scale intermediate images based on multi-scale warped images and synthesis weights. The image processing device can generate a synthesis frame based on the multi-scale intermediate images and synthesis weights. In this case, the synthesis weights (first synthesis weights) for generating a plurality of intermediate images and the synthesis weights (second synthesis weights) for generating the synthesis frame may be different.
[0153] The image processing device can improve the quality of the synthesized frame by comparing the generated synthesized frame with the original frame, re-measuring the loss, and updating the weights.
[0154] In one embodiment, the image processing device can evaluate the synthesis quality of the synthesized frame. Based on the synthesis quality, the image processing device can adjust a preset time interval V for selecting reference frames.
[0155] For example, if the quality of the synthesized frame is lower than the target value even after the optimization process is performed, it may be because the viewpoint interval between the reference frames is too far, making it difficult to estimate the disparity, thereby reducing the quality of the warped image. The image processing device may reduce the preset viewpoint interval V and repeat operations S810 to S840. If the quality of the synthesized frame is higher than the target value after the optimization process is performed, the image processing device may encode the reference frames and encode information about the optimized warped images (e.g., a disparity vector, a scale weight, a synthesis weight, etc.).
[0156] In operation S850, a pixel-based multi-scale synthesis operation can be performed. Operation S850 can be performed in a decoding device. The decoding device can be the same device as the image processing device that performed operations S810 to S840, or a different device. For example, operation S850 can be performed in an image processing device that performed operations S810 to S840, or operations S810 to S840 can be performed in a first image processing device and operation S850 can be performed in a second image processing device. As a result of performing operation S850, light field images having N viewpoints can be restored. The restored light field images can be composed of restored reference frames and synthesized frames.
[0157] FIG. 9 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to optimize a synthesis weight and a viewpoint interval of a reference frame.
[0158] Figure 9 further illustrates the optimization operations of Figure 8. The optimization operations may be performed by an encoding device. In describing Figure 9, the encoding device, which is the entity performing the optimization operations, will be referred to as an "image processing device."
[0159] In one embodiment, the disparity vector, weights, and viewpoint spacing of the reference frame can be optimized for each scene. That is, in the present disclosure, optimized values are used for each input data, rather than trained or predetermined values.
[0160] For example, the image processing device can optimize parameters for restoring the first light field images. The image processing device can select reference images and predict disparity vectors for the first light field images. The image processing device can restore the first light field images by generating composite frames through pixel-based multi-scale synthesis using the reference frame, the disparity vector, and the initial weights (scale weights, synthesis weights). The image processing device can measure loss by comparing the restored first light field images with the original first light field images, and update at least one of the predicted disparity vector, the scale weight indicating the scale ratio of the multi-scale, the synthesis weight used when synthesizing an intermediate image based on the warped images, and the synthesis weight used when synthesizing a composite frame based on the intermediate images. Since this has been described above with reference to FIG. 8, a repeated description thereof will be omitted.
[0161] In one embodiment, the optimization of the image processing device is scene-by-scene optimization. Accordingly, optimized parameters can be determined for each light field image of a different scene. In the example described above, the optimized parameters are parameters optimized for the first light field images. Accordingly, the encoding device can encode the optimized parameters and reference frames for the first light field images, and the decoding device can decode them to restore the first light field images. In addition, when the image processing device encodes second light field images, which are light field images of another scene, the image processing device can perform optimization operations for the second light field images to determine the optimized parameters.
[0162] FIG. 10 is a diagram for explaining the encoding and decoding operations of an occlusion map used in the encoding and decoding method of the present disclosure.
[0163] In operation S1010, the encoding device may generate an occlusion map. The encoding device may generate an occlusion map indicating an occluded area at viewpoints between the reference frames by moving pixels by a preset viewpoint interval based on reference frames. For example, occlusion maps may be generated for all viewpoints n between reference frame 0 and reference frame N. The occlusion map may include confidence weight information.
[0164] In operation S1020, the encoding device can extract key occlusion map information for all time points n.
[0165] In one embodiment, the occlusion area gradually increases or decreases depending on the viewpoint distance. Therefore, occlusion areas corresponding to multiple viewpoints can be processed with a single piece of information. The key occlusion map may mean a single map that collects occlusion maps corresponding to each viewpoint. In other words, since the key occlusion map is generated based on the occlusion maps corresponding to each viewpoint, the key occlusion map information may be information available for all viewpoints n between the viewpoints of the reference frames.
[0166] The key occlusion map information may include information related to the reliability weights available for all intermediate viewpoints n between reference frames. Specifically, the key occlusion map information may include information indicating at which intermediate viewpoint n a pixel of the warped image is a usable pixel. For example, the key occlusion map information may include the number of occlusion regions and the maximum length of the occlusion regions.
[0167] In operation S1030, the encoding device can transmit key occlusion map information to the decoding device. The encoding device can encode the key occlusion map and transmit it to the decoding device. Since the occlusion map and the key occlusion map are mutually convertible, the encoded key occlusion map may also be referred to as an encoded occlusion map.
[0168] In operation S1040, the decoding device can generate a key occlusion map from the decoded key occlusion map information. The decoding device can obtain the key occlusion map by decoding the encoded key occlusion map information. The key occlusion map can be data available for all viewpoints n between viewpoints of reference frames.
[0169] In operation S1050, the decoding device can generate an occlusion map indicating reliability weights used in synthesizing n frames. Since the key occlusion map is data available for all viewpoints n between viewpoints of reference frames, the decoding device can generate occlusion maps indicating reliability weights corresponding to intermediate viewpoints n between reference frames using the key occlusion map.
[0170] For example, when the viewpoints of reference frame 0 and reference frame N are 0 and N, respectively, and an intermediate viewpoint is n, the decoding device can generate occlusion maps corresponding to warped images having intermediate viewpoints between 0 and N with respect to reference frame 0. Specifically, the decoding device can generate occlusion maps indicating reliability weights used in synthesizing each of the synthesized frames 1, ..., N-1.
[0171] FIG. 11 is a drawing for explaining a key occlusion map according to one embodiment of the present disclosure.
[0172] A key occlusion map can be defined as a single map that represents occlusion maps corresponding to each viewpoint. In other words, since a key occlusion map is generated based on occlusion maps corresponding to each viewpoint, key occlusion map information can be available for all viewpoints n between the viewpoints of reference frames.
[0173] A key occlusion map may include information related to reliability weights available for all intermediate viewpoints n between reference frames. Specifically, the information of the key occlusion map may include information related to occlusion regions, which indicates which pixels of the warped image can be used at which intermediate viewpoint n. For example, the information of the key occlusion map may include information on the number of occlusion regions (1110) and information on the maximum length of the occlusion regions (1120).
[0174] The number of occluded areas information (1110) may refer to the number of occluded areas existing along the horizontal axis of the image. For example, if the number of occluded areas information (1110) is 6, this may mean that the number of occluded areas existing in the corresponding row of the image is 6.
[0175] The maximum length information (1120) of the occluded area may refer to the maximum length of the occluded area existing along the horizontal axis of the image. The maximum length information (1120) of the occluded area may include information indicating the starting point and maximum length of the occluded area.
[0176] The occlusion map may be aligned with a first reference frame or a second reference frame. For example, occlusion maps corresponding to viewpoints between the first reference frame and the second reference frame may be generated based on the first reference frame or the second reference frame. Here, the maximum length may refer to the length of the occlusion area in the occlusion map corresponding to the viewpoint furthest from the first reference frame (in the case of an occlusion map aligned with the first reference frame). Alternatively, it may refer to the length of the occlusion area in the occlusion map corresponding to the viewpoint furthest from the second reference frame (in the case of an occlusion map aligned with the second reference frame). Since the alignment of the occlusion map with respect to the reference frame has been mentioned in the description of FIG. 7, a repeated description thereof will be omitted.
[0177] In one embodiment, the length of the occlusion region may increase linearly. For example, if the value of the maximum length information (1120) of the occlusion region is (115, 4), the x value of the start coordinate of the occlusion region may be 115, and the maximum length may be 4. In this case, the length of the occlusion region of the occlusion map at the farthest viewpoint distance from the reference frame may be the maximum length, 4, and the length of the occlusion region of the occlusion map corresponding to the intermediate viewpoints may increase linearly until it becomes 4.
[0178] In one embodiment, the length of the occlusion region may increase nonlinearly. For example, the length of the occlusion region may change according to a quadratic function. In this case, the length information included in the maximum length information (1120) of the occlusion region may include information for interpreting the length information. For example, the information for interpreting the length information may include a quadratic coefficient used in the nonlinear increase / decrease calculation. Similarly, the length information included in the maximum length information (1120) of the occlusion region may include one or more coefficients included in an n-th order function for interpreting the length information.
[0179] FIG. 12 is a diagram for explaining an operation of generating occlusion maps based on a key occlusion map according to one embodiment of the present disclosure.
[0180] In describing Fig. 12, it is described as an example that light field images are composed of frames 1 to 9, reference frames are frame numbers 1 and 9, and occlusion maps aligned to reference frame 1 are generated. However, this is not limited thereto, and occlusion maps aligned to reference frame 9 may also be generated in a method identical / similar to that described below.
[0181] In one embodiment, the decoding device can decode an encoded occlusion map (or an encoded key occlusion map). The decoding device can generate an occlusion map based on the key occlusion map.
[0182] For example, the decoding device can generate an occlusion map (1210) corresponding to frame 3 based on the key occlusion map. Specifically, if the value of the maximum length information of the occlusion area of the key occlusion map is (115, 4), this means that the x value of the start coordinate of the occlusion area is 115, and the maximum length of the occlusion area is 4. In this case, since frame 3 is not the farthest frame from frame 1, the length of the occlusion area is not determined as the maximum length of 4, but may be recalculated as a value representing the occlusion area at an intermediate point in time. For example, as illustrated in FIG. 12, the length of the occlusion area corresponding to frame 3 may be calculated as 1. In this case, the x coordinates of the pixels corresponding to the occlusion area in the corresponding row of the image become 115 and 116. Since the occlusion map represents a reliability weight indicating the degree of reliability of a pixel, it may mean that the pixels at x-coordinates 115 and 116 corresponding to the occlusion area are unreliable pixels. In other words, the decoding device generates a synthetic frame corresponding to frame 3, and a warped image and an intermediate image are generated in the process of generating the synthetic frame. Here, when the decoding device generates a warped image, which is intermediate data obtained in the process of generating the synthetic frame corresponding to frame 3, or generates an intermediate image, the pixels corresponding to the occlusion area may not be used or may be used after a relatively low weight is applied, depending on the reliability weight of the occlusion map (1210) corresponding to frame 3.
[0183] Also, for example, the decoding device can generate an occlusion map (1220) corresponding to frame 9 based on the key occlusion map. Specifically, if the value of the maximum length information of the occlusion area of the key occlusion map is (115, 4), this means that the x value of the start coordinate of the occlusion area is 115 and the maximum length is 4. In this case, since frame 9 is the frame farthest from frame 1, the length of the occlusion area can be determined as the maximum length, 4. When the decoding device generates a synthetic frame corresponding to frame 9, by referring to the occlusion map (1220) corresponding to frame 9, pixels corresponding to the occlusion area can be unused or used after a relatively low weight is applied. The result of generating the occlusion map will be further described with reference to FIG. 13.
[0184] FIG. 13 is a drawing for explaining an occlusion map according to one embodiment of the present disclosure.
[0185] In explaining Fig. 13, the light field images are explained as being composed of frame 1 (1310) to frame 9 (1320) as an example. In addition, the reference frame is explained as being selected as an example in which frame 1 (1310) and frame 9 (1320) are reference frames.
[0186] In one embodiment, occlusion maps aligned to frame 1 (1310) may be obtained by a decoding device based on frame 1 (1310). For example, as illustrated in FIG. 13, an occlusion map (1312) corresponding to frame 3 and an occlusion map (1322) corresponding to frame 9 may be generated.
[0187] The decoding device can generate a composite frame 3 corresponding to frame 3. When generating the composite frame, the decoding device can generate multi-scale intermediate images based on multi-scale disparity vectors, and generate the composite frame based on the multi-scale intermediate images.
[0188] When the decoding device generates composite frame 3, warped images generated based on frame 1 (1310), which is a reference frame, and warped images generated based on frame 9 (1320), which is a reference frame, are combined at different scales to generate a multi-scale image. As shown in FIG. 13, the occlusion map (1312) corresponding to frame 3 represents the reliability weight at the intermediate viewpoint 3 when the viewpoint moves from frame 1 (1310) to frame 9 (1320). Similarly, the occlusion map (1322) corresponding to frame 9 represents the reliability weight at the final viewpoint 9 when the viewpoint moves from frame 1 (1310) to frame 9 (1320).
[0189] Referring to FIG. 13, it can be seen that as the viewpoint moves from frame 1 (1310) to frame 9 (1320) and the viewpoint distance increases, the length of the occlusion area increases. For example, the occlusion area (1314) of the occlusion map (1312) corresponding to frame 3 has a relatively shorter length than the occlusion area (1324) of the occlusion map (1322) corresponding to frame 9.
[0190] That is, as the viewpoint distance from the reference frame increases, the size of the area without pixel information due to the occlusion area increases, so the number of unusable pixels increases as the viewpoint distance increases. When generating a synthetic frame, the decoding device can utilize the reliability weight of the occlusion map to ensure that pixels with low reliability are not used or are given a low weight.
[0191] FIG. 14 is a block diagram illustrating a configuration of an image processing device according to one embodiment of the present disclosure.
[0192] In one embodiment, the image processing device (2000) may correspond to the encoding device and / or the decoding device of the present disclosure. The image processing device (2000) may perform the encoding operation or the decoding operation of the present disclosure. For example, the image processing device (2000) may perform the encoding operations of the present disclosure as an encoding device. Alternatively, the image processing device (2000) may perform the decoding operations of the present disclosure as a decoding device. The encoding device and the decoding device may be implemented as the same device or as different devices. For example, the image processing device (2000) may perform the encoding / decoding operations as an encoding / decoding device. Alternatively, a first image processing device may perform the encoding operations as an encoding device, and a second image processing device may perform the decoding operations as a decoding device.
[0193] In one embodiment, the image processing device (2000) may include a memory (2100) and a processor (2200).
[0194] The memory (2100) may store instructions, data structures, and program codes that can be read by the processor (2200). Operations performed by the processor (2200) may be implemented by executing instructions or codes of a program stored in the memory (2100).
[0195] The memory (2200) may include a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), and may include a non-volatile memory including at least one of a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, and an optical disk, and a volatile memory such as a RAM (Random Access Memory) or an SRAM (Static Random Access Memory).
[0196] The memory (2100) may store one or more instructions and / or programs that cause the image processing device (2000) to operate to process an image. For example, the memory (2100) may store instructions and / or programs for implementing functions of reference viewpoint selection, encoding, disparity estimation, and occlusion map encoding of an encoding device. Additionally, for example, the memory (2100) may store instructions and / or programs for implementing functions of decoding, occlusion map decoding, and adaptive viewpoint synthesis of a decoding device.
[0197] The processor (2200) is a circuit device that can control the overall operations of the image processing device (2000). For example, the processor (2200) can control the overall operations of the image processing device (2000) for performing encoding, decoding, and image synthesis by executing one or more instructions of a program stored in the memory (2100). There may be one or more processors (2200).
[0198] The processor (2200) may be configured with at least one of, but is not limited to, a central processing unit, a microprocessor, a graphic processing unit, an application specific integrated circuits (ASICs), a digital signal processor (DSPs), a digital signal processing device (DSPDs), a programmable logic device (PLDs), a field programmable gate array (FPGAs), an application processor, a neural processing unit, or an artificial intelligence processor designed with a hardware structure specialized for processing an artificial intelligence model.
[0199] The processor (2200) can write data to the memory (2100) or read data stored in the memory (2100). For example, the processor (2200) can load one or more instructions and / or programs stored in the non-volatile memory into the volatile memory, and process data and program code. Descriptions related to operations of reference viewpoint selection, encoding, disparity estimation, and occlusion map encoding of the encoding device executed by the processor (2200) and operations of decoding, occlusion map decoding, and adaptive viewpoint synthesis of the decoding device have already been described in the description of the previous drawings, and therefore, a repeated description will be omitted.
[0200] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an AI-specific processor). Here, an AI-specific processor, which is an example of the second processor, may perform operations for training / inference of an AI model. However, the embodiments of the present disclosure are not limited thereto.
[0201] One or more processors according to the present disclosure may be implemented as a single-core processor or as a multi-core processor.
[0202] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core or may be performed by a plurality of cores included in one or more processors.
[0203] FIG. 15 is a block diagram illustrating a configuration of an image processing device according to one embodiment of the present disclosure.
[0204] In one embodiment, the image processing device (2000) may include a memory (2100), a processor (2200), a communication interface (2300), and a display. The memory (2100) and the processor (2200) of FIG. 15 may correspond to the memory (2100) and the processor (2200) of FIG. 14. Therefore, for the sake of brevity, a repeated description is omitted.
[0205] The communication interface (2300) can perform data communication with other electronic devices under the control of the processor (2200). For example, when the image processing device (2000) operates as an encoding device, the image processing device (2000) can transmit encoded data to a decoding device. Alternatively, when the image processing device (2000) operates as a decoding device, the image processing device (2000) can receive encoded data from the encoding device.
[0206] The communication interface (2300) may include a communication circuit that can perform data communication between the image device (2000) and another electronic device using at least one of data communication methods including, for example, wired LAN, wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliances (WiGig), and RF communication.
[0207] The display (2400) can output image signals to the screen of the image processing device (2000) under the control of the processor (2200). For example, the display (2400) can output original light field images to the screen. For example, the display (2400) can output restored light field images including composite frames to the screen.
[0208] The present disclosure relates to a method for efficiently compressing, transmitting, and restoring large-capacity light field images. The technical challenges addressed by the present disclosure are not limited to those described above, and other technical challenges not mentioned herein will be readily apparent to those skilled in the art, based on the description herein.
[0209] According to one aspect of the present disclosure, a method for decoding and outputting light field images may be provided.
[0210] The method may include obtaining an encoded first frame and an encoded second frame, which represent reference frames captured from different viewpoints, and an encoded synthesis parameter and an encoded occlusion map for generating a synthesis frame.
[0211] The method may include decoding the encoded first frame, the encoded second frame, the encoded synthesis parameters, and the encoded occlusion map to restore the first frame, the second frame, the disparity vector, the synthesis weights, and the occlusion map.
[0212] The method may include a step of generating a synthetic frame based on the first frame, the second frame, the disparity vector, the synthetic weight, and the occlusion map.
[0213] The method may include a step of outputting light field images including the first frame, the second frame, and the composite frame.
[0214] The above encoded synthesis parameters may include pixel movement information and synthesis weight information.
[0215] The encoded occlusion map may represent an occlusion area caused by an object between a first viewpoint corresponding to the first frame and a second viewpoint corresponding to the second frame, and may include reliability weight information for processing the occlusion area and information related to the occlusion area.
[0216] The decoding step may include a step of restoring multi-scale disparity vectors and multi-scale synthesis weights based on a plurality of encoded synthesis parameters.
[0217] The step of generating the above synthetic frame may include the step of generating a plurality of intermediate images of multiple scales based on the disparity vectors of the multiple scales.
[0218] The step of generating the composite frame may include a step of generating the composite frame by combining the plurality of intermediate images.
[0219] The step of generating the plurality of intermediate images may include the step of generating a first intermediate image by applying a first synthesis weight to the first frame and a first disparity vector corresponding to the first frame, and the second frame and a second disparity vector corresponding to the second frame.
[0220] The step of generating the above-described composite frame may include a step of generating the above-described composite frame by applying a second composite weight to the first intermediate image and the second intermediate image.
[0221] The step of generating the above composite frame may be generating a plurality of composite frames corresponding to a plurality of time points between a first time point corresponding to the first frame and a second time point corresponding to the second frame.
[0222] According to one aspect of the present disclosure, an image processing device may be provided. The image processing device may include a memory storing one or more instructions; and one or more processors executing the one or more instructions stored in the memory.
[0223] The one or more processors can obtain an encoded first frame and an encoded second frame, which represent reference frames captured from different points of view, and an encoded synthesis parameter and an encoded occlusion map for generating a synthesis frame, by executing the one or more instructions.
[0224] The one or more processors can decode the encoded first frame, the encoded second frame, the encoded synthesis parameters, and the encoded occlusion map to restore the first frame, the second frame, the disparity vector, the synthesis weights, and the occlusion map by executing the one or more instructions.
[0225] The one or more processors can generate a composite frame based on the first frame, the second frame, the disparity vector, the composite weight, and the occlusion map by executing the one or more instructions.
[0226] The one or more processors can output light field images including the first frame, the second frame, and the composite frame by executing the one or more instructions.
[0227] The above encoded synthesis parameters may include pixel movement information and synthesis weight information.
[0228] The encoded occlusion map may represent an occlusion area caused by an object between a first viewpoint corresponding to the first frame and a second viewpoint corresponding to the second frame, and may include reliability weight information for processing the occlusion area and information related to the occlusion area.
[0229] The one or more processors can restore multi-scale disparity vectors and multi-scale synthesis weights based on a plurality of encoded synthesis parameters by executing the one or more instructions.
[0230] The one or more processors can generate a plurality of multi-scale intermediate images based on the multi-scale disparity vectors by executing the one or more instructions.
[0231] The one or more processors can generate the composite frame by combining the plurality of intermediate images by executing the one or more instructions.
[0232] The one or more processors can generate a first intermediate image by applying a first synthesis weight to the first frame and a first disparity vector corresponding to the first frame, and to the second frame and a second disparity vector corresponding to the second frame, by executing the one or more instructions.
[0233] The one or more processors can generate the composite frame by applying a second composite weight to the first intermediate image and the second intermediate image by executing the one or more instructions.
[0234] The one or more processors can generate a plurality of composite frames corresponding to a plurality of time points between a first time point corresponding to the first frame and a second time point corresponding to the second frame by executing the one or more instructions.
[0235] According to one aspect of the present disclosure, a method of encoding light field images may be provided.
[0236] The method may include a step of selecting reference frames including a first frame and a second frame from among original light field images including a plurality of viewpoints.
[0237] The method may include generating a disparity vector, a synthesis weight, and an occlusion map for inferring frames between the first frame and the second frame.
[0238] The method may include encoding the first frame, the second frame, the disparity vector, the synthesis weight, and the occlusion map.
[0239] The step of selecting the above reference frames may include a step of selecting a plurality of reference frames including the first frame and the second frame based on a preset time interval.
[0240] The method may include a step of generating a synthetic frame based on the first frame, the second frame, the disparity vector, the synthetic weight, and the occlusion map.
[0241] The method may include a step of comparing the synthesized frame with an original frame and measuring a difference.
[0242] The method may include a step of updating the synthetic weight based on the difference.
[0243] The step of generating the above composite frame may include the step of generating multi-scale warped images based on the first frame, the second frame, and the disparity vectors.
[0244] The step of generating the above composite frame may include the step of generating a plurality of intermediate images of multiple scales based on the warped images of multiple scales.
[0245] The step of generating the composite frame may include a step of generating the composite frame by combining the plurality of intermediate images.
[0246] The above synthesis weights may include a first synthesis weight for generating the plurality of intermediate images and a second synthesis weight for generating the synthesis frame.
[0247] The step of updating the above synthetic weight may include the step of updating at least one of the first synthetic weight and the second synthetic weight.
[0248] The method may include a step of evaluating the synthesis quality of the synthesized frame.
[0249] The method may include a step of adjusting the preset time interval for selecting the reference frames based on the synthesis quality.
[0250] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include computer storage media and communication media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.
[0251] Additionally, a computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0252] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0253] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.
[0254] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.
Claims
1. In a method for decoding and outputting light field images, A step of obtaining an encoded first frame and an encoded second frame, which represent reference frames captured from different points of view, and an encoded synthesis parameter and an encoded occlusion map for generating a synthesis frame; A step of decoding the encoded first frame, the encoded second frame, the encoded synthesis parameters and the encoded occlusion map to restore the first frame, the second frame, the disparity vector, the synthesis weights and the occlusion map; A step of generating a composite frame based on the first frame, the second frame, the disparity vector, the composite weight, and the occlusion map; and A method comprising the step of outputting light field images including the first frame, the second frame, and the composite frame.
2. In paragraph 1, A method wherein the encoded synthesis parameters include pixel movement information and synthesis weight information.
3. In paragraph 2, A method wherein the encoded occlusion map represents an occlusion area caused by an object occurring between a first viewpoint corresponding to the first frame and a second viewpoint corresponding to the second frame, and includes reliability weight information for processing the occlusion area and information related to the occlusion area.
4. In paragraph 1, The above decoding step is, A method comprising the step of restoring multi-scale disparity vectors and multi-scale synthesis weights based on a plurality of encoded synthesis parameters.
5. In paragraph 4, The step of generating the above composite frame is: A step of generating multiple intermediate images of multiple scales based on the above multi-scale disparity vectors; and A method comprising the step of combining the plurality of intermediate images to generate the composite frame.
6. In paragraph 5, The step of generating the above multiple intermediate images is: A step of generating a first intermediate image by applying a first synthesis weight to the first frame and a first disparity vector corresponding to the first frame, and the second frame and a second disparity vector corresponding to the second frame, The step of generating the above composite frame is: A method comprising the step of generating the composite frame by applying a second composite weight to the first intermediate image and the second intermediate image.
7. In paragraph 1, The step of generating the above composite frame is: A method for generating a plurality of composite frames corresponding to a plurality of time points between a first time point corresponding to the first frame and a second time point corresponding to the second frame.
8. In the image processing device, Memory that stores one or more instructions; and comprising one or more processors that execute one or more instructions stored in the memory; The one or more processors, by executing the one or more instructions, Obtaining an encoded first frame and an encoded second frame, which represent reference frames captured from different viewpoints, and an encoded synthesis parameter and an encoded occlusion map for generating a synthesis frame, Decoding the encoded first frame, the encoded second frame, the encoded synthesis parameters, and the encoded occlusion map to restore the first frame, the second frame, the disparity vector, the synthesis weights, and the occlusion map, Generating a composite frame based on the first frame, the second frame, the disparity vector, the composite weight, and the occlusion map, An image processing device that outputs light field images including the first frame, the second frame, and the composite frame.
9. In paragraph 8, An image processing device, wherein the encoded synthesis parameters include pixel movement information and synthesis weight information.
10. In paragraph 9, An image processing device, wherein the encoded occlusion map represents an occlusion area caused by an object occurring between a first viewpoint corresponding to the first frame and a second viewpoint corresponding to the second frame, and includes reliability weight information for processing the occlusion area and information related to the occlusion area.
11. In paragraph 8, The one or more processors, by executing the one or more instructions, An image processing device that restores multi-scale disparity vectors and multi-scale synthesis weights based on a plurality of encoded synthesis parameters.
12. In paragraph 11, The one or more processors, by executing the one or more instructions, Based on the above multi-scale disparity vectors, multiple intermediate images of multi-scales are generated, An image processing device that combines the plurality of intermediate images to generate the composite frame.
13. In paragraph 12, The one or more processors, by executing the one or more instructions, Generating a first intermediate image by applying a first synthesis weight to the first frame and a first disparity vector corresponding to the first frame, and to the second frame and a second disparity vector corresponding to the second frame, An image processing device that generates the composite frame by applying a second composite weight to the first intermediate image and the second intermediate image.
14. In paragraph 11, The one or more processors, by executing the one or more instructions, An image processing device that generates a plurality of composite frames corresponding to a plurality of time points between a first time point corresponding to the first frame and a second time point corresponding to the second frame.
15. In a method for encoding light field images, A step of selecting reference frames including a first frame and a second frame from among original light field images including multiple viewpoints; A step of generating a disparity vector, a synthesis weight, and an occlusion map for inferring frames between the first frame and the second frame; and A method comprising the steps of encoding the first frame, the second frame, the disparity vector, the synthesis weight, and the occlusion map.
Citation Information
Patent Citations
Systems and methods for encoding and decoding light field image files
KR102002165B1
Face Recognition System and Method for Activating Attendance Menu
KR102325251B1
Methods for full parallax compressed light field 3D imaging systems
KR102450688B1
Neural blending for novel view synthesis
KR102612529B1
KR20230106714A