Image processing device and image encoding and decoding method

The method efficiently compresses and restores large-scale light field images by encoding reference frames and residual images, reducing data transmission through selective encoding and inference-based restoration processes.

WO2025192865A1PCT designated stage Publication Date: 2025-09-18SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/000917
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-14
Filing Date
2025-01-15
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing methods are inefficient for compressing, transmitting, and restoring large-scale light field images due to their high data requirements and overlapping image portions.

Method used

A method and device that selectively encode and transmit reference frames, flow maps, and residual images, allowing for the generation of composite frames through inference and restoration processes without requiring additional metadata, using techniques like HEVC, VVC, or AV1 for encoding, and generating occlusion and index maps to apply residual images to composite frames.

Benefits of technology

This approach reduces data transmission by encoding only reference frames and residual images, enabling efficient compression and restoration of large-capacity light field images with reduced data volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025000917_18092025_PF_FP_ABST
    Figure KR2025000917_18092025_PF_FP_ABST
Patent Text Reader

Abstract

A method of processing light field images is provided. The method may comprise the steps of: obtaining a residual image, a flow map for generating a synthetic frame, and a first frame and a second frame representing reference frames photographed from different viewpoints; generating a synthetic frame on the basis of the first frame, the second frame, and the flow map; generating an occlusion map corresponding to the synthetic frame; generating an index map corresponding to the residual image on the basis of the occlusion map; applying the residual image to the synthetic frame on the basis of the index map; and outputting light field images including the first frame, the second frame, and the synthetic frame.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device and method for encoding and decoding images

[0001] The present disclosure relates to a method for compressing, transmitting and restoring large-capacity light field images.

[0002] A light field is a field used to express the intensity and direction of light reflected from a subject in three-dimensional space. Light fields, a new approach to three-dimensional image processing, rely on large amounts of data. Therefore, a method is needed to efficiently compress, transmit, and restore large-scale light field images.

[0003] According to one aspect of the present disclosure, a method for processing light field images may be provided. The method may include obtaining a first frame representing reference frames captured from different viewpoints, a second frame, and a flow map and a residual image for generating a composite frame. The method may include generating a composite frame based on the first frame, the second frame, and the flow map. The method may include generating an occlusion map corresponding to the composite frame. The method may include generating an index map corresponding to the residual image based on the occlusion map. The method may include applying the residual image to the composite frame based on the index map. The method may include outputting light field images including the first frame, the second frame, and the composite frame.

[0004] According to one aspect of the present disclosure, an image processing device may be provided. The image processing device may include a memory storing one or more instructions; and one or more processors executing the one or more instructions stored in the memory. The one or more processors may, by executing the one or more instructions, obtain a first frame representing reference frames captured from different viewpoints, a second frame, and a flow map and a residual image for generating a composite frame. The one or more processors may, by executing the one or more instructions, generate a composite frame based on the first frame, the second frame, and the flow map. The one or more processors may, by executing the one or more instructions, generate an occlusion map corresponding to the composite frame. The one or more processors may, by executing the one or more instructions, generate an index map corresponding to the residual image based on the occlusion map. The one or more processors may apply the residual image to the composite frame based on the index map by executing the one or more instructions. The one or more processors may output light field images including the first frame, the second frame, and the composite frame by executing the one or more instructions.

[0005] According to one aspect of the present disclosure, a method for compressing light field images may be provided. The method may include selecting reference frames including a first frame and a second frame from among original light field images including a plurality of viewpoints. The method may include generating a flow map for inferring frames between the first frame and the second frame. The method may include generating a synthesized frame based on the first frame, the second frame, and the flow map. The method may include generating an occlusion map corresponding to the synthesized frame. The method may include generating an index map corresponding to a residual image based on the occlusion map. The method may include generating a residual image including a residual to be applied to the synthesized frame based on the index map. The method may include encoding the first frame, the second frame, the flow map, and the residual image.

[0006] FIG. 1 is a drawing for explaining light field images generated by an image processing device according to one embodiment of the present disclosure.

[0007] FIG. 2 is a drawing for explaining an operation of an image processing device according to one embodiment of the present disclosure to acquire light field images.

[0008] FIG. 3 is a diagram for explaining the encoding and decoding process of light field images according to one embodiment of the present disclosure.

[0009] FIG. 4 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate light field images.

[0010] FIG. 5 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate an occlusion map corresponding to a reference frame.

[0011] FIG. 6 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate an occlusion map corresponding to a synthesized frame.

[0012] FIG. 7 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate an index map.

[0013] FIG. 8 is a drawing for explaining a mask and index generated by an image processing device according to one embodiment of the present disclosure.

[0014] FIG. 9 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to estimate the location of a restoration area.

[0015] FIG. 10 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate an index map.

[0016] FIG. 11 is a drawing for generally explaining a process of restoring light field images by an image processing device according to one embodiment of the present disclosure.

[0017] FIG. 12 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate an index map.

[0018] FIG. 13 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to restore an image using an index map.

[0019] FIG. 14 is a block diagram illustrating a configuration of an image processing device according to one embodiment of the present disclosure.

[0020] FIG. 15 is a block diagram illustrating a configuration of an image processing device according to one embodiment of the present disclosure.

[0021] The terms used in this specification will be briefly explained, followed by a detailed description of the present disclosure. The terms used in this disclosure have been selected from widely used and common terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, in which case their meanings will be described in detail in the relevant description. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.

[0022] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein. Furthermore, terms containing ordinal numbers, such as "first" or "second," used herein may be used to describe various components, but such components should not be limited by such terms. Such terms are used solely to distinguish one component from another.

[0023] When a part of the specification is said to "include" a component, unless otherwise specifically stated, this does not exclude other components but rather implies the inclusion of other components. Furthermore, terms such as "part" and "module" used in the specification refer to a unit that processes at least one function or operation, which may be implemented in hardware, software, or a combination of hardware and software.

[0024] Below, with reference to the attached drawings, embodiments of the present disclosure are described in detail so that those skilled in the art can easily practice the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, portions irrelevant to the description have been omitted for clarity of explanation, and similar reference numerals have been used throughout the specification to designate similar parts.

[0025] The present disclosure will be described below with reference to the attached drawings.

[0026] FIG. 1 is a drawing for explaining light field images generated by an image processing device according to one embodiment of the present disclosure.

[0027] Restored light field images (100) generated by a decoding device are illustrated. The restored light field images (100) may be composed of one or more reference frames (110) and one or more composite frames (120).

[0028] In one embodiment, the restored light field images (100) may be generated by receiving encoded data based on original light field images from an encoding device and restoring the same. The original light field images may be images corresponding to multiple viewpoints. For example, the original light field images composed of N images may include a total of N viewpoints, with each image being acquired from a different viewpoint. The original light field images may be images acquired by multiple cameras at different viewpoints. Alternatively, the original light field images may be images acquired by a single camera at different viewpoints while moving.

[0029] In one embodiment, the original light field images are encoded in an encoding device. In this case, not all of the original light field images are encoded, but only some of the images are selected, encoded, and transmitted to the decoding device. In addition to the selected images, the remaining images, including the flow map and residual images, are encoded and transmitted to the decoding device.

[0030] The reference frame (110) may refer to reference images to be encoded among the original light field images. The reference frame (110) is encoded in an encoding device and transmitted to a decoding device. The decoding device can restore the reference frame (110) through decoding the encoded reference frame (110). As a result, examples of reference frames (110) restored in the decoding device, namely, a first reference frame (112), a second reference frame (114), and a third reference frame (116), are illustrated in FIG. 1.

[0031] Meanwhile, the reference frame (110) may be referred to by various expressions representing the same or similar concepts. The reference frame (110) may be referred to by expressions such as a reference image, a reference view, a key frame, an anchor frame, an anchor view, etc., and is not limited to the examples described above.

[0032] A composite frame (120) may refer to images corresponding to the remaining frames other than the encoded and transmitted reference frames (110) among the original light field images, and may refer to images generated through a composite process in a decoding device.

[0033] The decoding device can generate composite frames (120) having viewpoints between the reference frames (110) by referring to the reference frames (110). For example, the decoding device can generate, based on a first reference frame (112) and a second reference frame (114), a first composite frame (122), ..., a second composite frame (124), which are composite frames (120) having viewpoints between the first reference frame (112) and the second reference frame (114). Also, for example, the decoding device can generate, based on a second reference frame (114) and a third reference frame (116), third composite frames (126), which are composite frames (120) having viewpoints between the second reference frame (114) and the third reference frame (116).

[0034] Meanwhile, the synthetic frame (120) may be referred to by various expressions representing the same / similar concepts. The synthetic frame (120) may be referred to by expressions such as synthetic frame, synthetic view, predicted frame, predicted image, predicted view, inferred frame, inferred image, inferred view, restored frame, restored image, restored view, etc., and is not limited to the examples described above.

[0035] In one embodiment, when the decoding device generates composite frames (120), metadata for applying the residual image may not be received from the encoding device. The decoding device may directly perform operations for performing image restoration by applying the residual image to the composite frame.

[0036] An encoding / decoding method according to one embodiment relates to a method for efficiently compressing, transmitting, and restoring large-capacity original light field images. Large-capacity original light field images have different viewpoints, but there are overlapping portions between the images. Therefore, compressing and transmitting all images is a relatively inefficient method.

[0037] An encoding device can compress and transmit only some of the images, which are reference frames (110), selected using the encoding method of the present disclosure. In addition, the remaining images other than the reference frames can be compressively transmitted, such as a flow map and a residual image, so that they can be generated through inference and restoration operations in a decoding device. Since the flow map and the residual image are relatively small-capacity data compared to the original image, compressing and transmitting the flow map and the residual image can reduce the amount of data transmitted from the encoding device to the decoding device compared to compressing and transmitting the original image.

[0038] The decoding device can obtain restored light field images (100) by restoring data received from the encoding device using the decoding method of the present disclosure. The decoding device can generate an index map for mapping the composite frame (120) and the residual image by performing the restoration area estimation method of the present disclosure. In this case, residual compensation can be applied to the composite frame (120) based on the index map.

[0039] Meanwhile, in the present disclosure, the encoding device and decoding device may also be referred to as "image processing devices." For example, separate image processing devices may each perform encoding and decoding, or a single image processing device may perform both encoding and decoding.

[0040] The encoding and decoding methods and the operations of the encoding and decoding devices of the present disclosure will be described in more detail through the drawings and descriptions thereof described below.

[0041] FIG. 2 is a drawing for explaining an operation of an image processing device according to one embodiment of the present disclosure to acquire light field images.

[0042] In one embodiment, the image processing device may be a decoding device that performs operations of a method of decoding light field images.

[0043] In operation S210, the image processing device may obtain a first frame, a second frame, a flow map, and a residual image. The first frame and the second frame may refer to reference frames captured from different viewpoints. For example, the first frame may be an image of a first viewpoint obtained by capturing an object by a camera positioned at a first location, and the second frame may be an image of a second viewpoint obtained by capturing the object by a camera positioned at a second location. A plurality of different viewpoints may be further included between the first viewpoint and the second viewpoint. The number of reference frames may be two or more, but the present disclosure will be described as an example in which there are two reference frames (the first frame and the second frame). The residual image may be an aggregate of areas in which there are differences between original light field images of a plurality of viewpoints and synthesized frames of a plurality of viewpoints, in units of pixels.

[0044] In one embodiment, the first frame, the second frame, the flow map, and the residual image may be encoded. The image processing device may obtain the encoded first frame, the encoded second frame, and the encoded flow map and the encoded residual image. The encoding operation may be performed by the image processing device, or may be performed by another encoding device (e.g., the second image processing device) other than the image processing device and received by the image processing device. Various types of video coding technologies may be applied for the encoding. For example, HEVC (High Efficiency Video Coding) may be used as the encoding method. However, the present invention is not limited thereto, and other methods such as H.264, VVC (Versatile Video Coding), or AV1 (AOMedia Video) may also be used.

[0045] The image processing device can decode the encoded first frame, the encoded second frame, the encoded flow map, and the encoded residual image to obtain the first frame, the second frame, the flow map, and the residual image.

[0046] In operation S220, the image processing device can generate a composite frame based on the first frame, the second frame, and the flow map.

[0047] In one embodiment, the composite frame may be an image corresponding to a point in time between reference frames. For example, if the first frame, which is a reference frame, is an image of the first point in time, and the second frame, which is a reference frame, is an image of the second point in time, the composite frame may be an image corresponding to point n between the first and second points in time.

[0048] When the image processing device wants to generate a composite frame corresponding to time n between a first time point and a second time point, the image processing device can generate a first warped image in which pixels are moved from the first time point, which is a reference point, to time n using the first frame and the first flow map corresponding to the first frame. In addition, the image processing device can generate a second warped image in which pixels are moved from the second time point to time n using the second frame and the second flow map. The image processing device can generate a composite frame corresponding to time n using at least one of the first warped image and the second warped image. Meanwhile, since time n included between the first time point and the second time point may be plural (n = 1, 2, 3, ...), the composite frames corresponding to any time point n may also be plural.

[0049] In operation S230, the image processing device can generate an occlusion map corresponding to the synthesized frame.

[0050] An occlusion map contains information about error areas that occur due to viewpoint differences when a composite frame is generated between a first viewpoint corresponding to the first frame and a second viewpoint corresponding to the second frame. For example, the occlusion map may contain information about an occluded area caused by an object that occurs due to a viewpoint shift between the first and second viewpoints, but is not limited thereto. For example, the occlusion map may also contain information about an area where pixel information is uncertain as a result of moving pixels through a flow map for viewpoint shift.

[0051] In one embodiment, the image processing device can generate a base occlusion map, which is an occlusion map corresponding to a reference frame. For example, the image processing device can generate a first base occlusion map corresponding to a first frame and a second base occlusion map corresponding to a second frame. The operation of generating the base occlusion map is further described in the description of FIG. 5 .

[0052] An image processing device can generate an occlusion map corresponding to a synthesized frame. For example, the image processing device can generate an occlusion map corresponding to a synthesized frame of point n, which is a point between a first point and a second point. The occlusion map corresponding to the synthesized frame can be generated based on at least one of a first base occlusion map and a second base occlusion map. The operation of generating the occlusion map is further described in the description of FIG. 6.

[0053] In operation S240, the image processing device can generate an index map corresponding to the residual image based on the occlusion map.

[0054] In one embodiment, the image processing device can estimate a restoration area of ​​a synthetic frame based on an occlusion map. The restoration area of ​​a synthetic frame refers to an area in the synthetic frame generated in operation S220 where pixel information is missing or inaccurate, and refers to an area to be restored using a residual image.

[0055] An image processing device can estimate a restoration area using an occlusion map and generate an index map for matching the restoration area with a residual image. The index map can include at least one of viewpoint information and pixel position information. By directly estimating the restoration area and generating the index map based on the occlusion map, the image processing device can apply the residual image to a composite frame without receiving metadata indicating the restoration area from an encoding device. The operation of generating the index map is further described in the descriptions of FIGS. 7, 10, and 12.

[0056] In operation S250, a residual image can be applied to a composite frame based on an index map.

[0057] In one embodiment, the residual image may be an aggregate of pixel-wise areas of differences between original light field images of multiple viewpoints and synthesized frames of multiple viewpoints. Furthermore, since the index map includes information regarding where a pixel of the residual image is to be applied to which synthesized frame, the image processing device can apply the pixels of the residual image to the synthesized frame by referring to the index map to restore the image.

[0058] Therefore, in the image restoration task, each pixel in the residual image collected by referring to the index map can be distributed to the pixel positions of the original synthesized frame. For example, the first pixel of the residual image can be restored to a specific position (x1, y1) of the first synthesized frame, and the second pixel of the residual image can be distributed to a specific position (x2, y2) of the second synthesized frame. The first synthesized frame and the second synthesized frame may correspond to different points in time.

[0059] In operation S260, the image processing device may output light field images including a first frame, a second frame, and a composite frame. The composite frame may be restored in operation S250.

[0060] In one embodiment, the image processing device can transmit light field images to an external device (e.g., a user device, etc.). The image processing device can transmit the light field images so that the light field images can be displayed on the external device.

[0061] In one embodiment, if the image processing device includes a display, the image processing device can display the restored light field images on a screen of the image processing device.

[0062] FIG. 3 is a diagram for explaining the encoding and decoding process of light field images according to one embodiment of the present disclosure.

[0063] In one embodiment, an encoding device compresses and transmits light field images, and a decoding device can decompress encoded data received from the encoding device back into light field images. The decoded light field images can be displayed to a user.

[0064] When transmitting light field images, the encoding device can compress and transmit only reference frames, rather than compressing and transmitting all images. Furthermore, data for generating the remaining frames, excluding the reference frames, can be generated and transmitted to the decoding device. This allows the decoding device to generate composite frames corresponding to the remaining frames.

[0065] The encoding device can calculate difference images (330) between original images (310) and synthetic images (320). The original images (310) may refer to originals of the remaining frames excluding reference frames among the light field images. The synthetic images (320) may refer to inferred frames excluding reference frames among the light field images. The synthetic images (320) may be inferred using the reference frame and a flow map. The difference images (330) may represent errors between the inferred synthetic images (320) and the original images (310). The original images (310), the synthetic images (320), and the difference images (330) may be images each corresponding to a plurality of viewpoints.

[0066] The encoding device can generate an index map (340) including information on an area where an error exists between the original and the composite based on the difference images (330). The index map (340) can include at least one of viewpoint information indicating at which viewpoint the error is between the original and the composite and pixel position information indicating at which pixel position the error is between the original and the composite. After generating the index map (340), the encoding device can generate a residual image (350) that collects, on a pixel basis, areas where there are differences between the original images (310) of a plurality of viewpoints and the composite images (320) of a plurality of viewpoints based on the index map (340).

[0067] The encoding device can encode data and transmit it to the decoding device. For example, the encoding device can encode reference frames. For example, the encoding device can encode flow maps. For example, the encoding device can encode residual images (350). Meanwhile, the encoding device may not encode original images (310), composite images (320), and metadata for image restoration (e.g., index map (340), etc.). In this case, since the encoding device does not encode and transmit all light field images, it can transmit a smaller amount of data to the decoding device than if it encodes and transmits all light field images. In other words, the amount of data transmitted can be reduced.

[0068] The decoding device can decode encoded data and restore light field images. For example, the decoding device can decode encoded reference frames, encoded flow maps, and encoded residual images (350). The decoding device can use the reference frames and the synthesized flow maps to generate synthetic images (322) generated by the decoding device.

[0069] In one embodiment, the decoding device can directly generate metadata for image restoration since it restores light field images without metadata. For example, the decoding device can generate an index map (342). Based on the index map (342) generated by the decoding device, the decoding device can apply the residual image (350) to the synthetic images (322) to obtain restored images (360). The restored images (360) can be images corresponding to the original images (310) generated by the decoding device through restoration area inference and residual compensation.

[0070] FIG. 4 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate light field images.

[0071] In one embodiment, the image processing device may include various modules for performing functions of image decompression and synthetic frame generation (410), restoration area estimation (420), index generation (430), and residual compensation (450). Each module may be implemented as a program, instruction, or code for implementing specific functions, and may be executed by a processor of the image processing device.

[0072] An image processing device can obtain encoded data (400). The encoded data can include an encoded reference frame, an encoded flow map, and an encoded residual image.

[0073] The image decompression and composite frame generation (410) operation decodes encoded data (400) and generates a composite frame. The result (412) of the image decompression and composite frame generation (410) operation may include a reference frame, a flow map, a residual image, and a composite frame.

[0074] An image processing device can decode encoded data (400) to obtain a reference frame, a flow map, and a residual image. There may be multiple reference frames. For example, the reference frames may include a first frame and a second frame representing images at different points in time.

[0075] An image processing device can use reference frames and a flow map to generate a composite frame representing a time point between the reference frames. There may be multiple composite frames. For example, if a first frame corresponding to a first time point and a second frame corresponding to a second time point are reference frames, the composite frames may include multiple images corresponding to time points between the first time point and the second time point.

[0076] The restoration area estimation (420) operation estimates an area to be restored within a synthesized frame. The result of the restoration area estimation (420) operation may include a restoration area mask (422).

[0077] An image processing device can generate occlusion maps corresponding to the synthesized frames. The image processing device can generate restoration area masks (422) using the occlusion maps. The occlusion map and / or restoration area mask (422) can include restoration area information indicating errors in the synthesized frames, and enable pixels in the corresponding restoration area to be restored by the residual image.

[0078] The index generation (430) operation generates an index for mapping the estimated restoration area and the residual image. As a result of the index generation (430) operation, an index map (432) may be generated. The index map (432) is also used in the encoding process of light field images, but is not encoded as metadata and therefore is not included in the encoded data (400) acquired by the image processing device. The image processing device can directly infer the index map (432) in the restoration process of light field images. The index map (432) may include at least viewpoint information and pixel position information. The viewpoint information may be information indicating which viewpoint of a synthesized frame a pixel of a residual image should be applied to among synthesized frames. The pixel position information may be information indicating which position of a pixel a pixel of a residual image should be applied to within a synthesized frame.

[0079] The residual compensation (440) operation compensates for residuals in a synthesized frame using a residual image. The image processing device can apply pixels of the residual image to restored areas of synthesized frames corresponding to multiple viewpoints using viewpoint information and / or pixel position information included in the index map (432).

[0080] The image processing device can obtain restored light field images through the aforementioned operations. The restored light field images may include a reference frame and a synthesized frame.

[0081] FIG. 5 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate an occlusion map corresponding to a reference frame.

[0082] In explaining Fig. 5, it is explained as an example that the light field images are composed of frames 1 to 9, and that frame numbers 1 and 9 are reference frames (first frame, second frame). In other words, it is explained as an example that the light field images are composed of images of nine different viewpoints corresponding to each frame.

[0083] In one embodiment, an image processing device may use a reference frame and a flow map to generate a base occlusion map, which is an occlusion map corresponding to the reference frame. The occlusion map corresponding to the reference frame may include information indicating an error between an inferred image and the original reference image when inferring the original reference frame based on another reference frame.

[0084] For example, referring to FIG. 5, when the first frame (510) is warped using the first frame (510), which is a reference frame, and the flow map (flow 1) including pixel movement information from the first frame to the second frame (520), the second frame (520) can be inferred. In addition, when the second frame (520) is warped using the second frame (520), which is a reference frame, and the flow map (flow 9) including pixel movement information from the second frame to the first frame (510), the first frame (510) can be inferred.

[0085] An image processing device can reverse warp an image using an inverse flow map, which is an inverse transformation of a flow map. For example, a reverse-warped first frame (512) may be obtained by inferring the first frame (510) using a second frame (520) and an inverse flow map (flow 1'). Also, for example, a reverse-warped second frame (522) may be obtained by inferring the second frame (520) using the first frame (510) and an inverse flow map (flow 9').

[0086] An image processing device can generate an occlusion map corresponding to a reference frame using an original reference frame and an inferred reference frame. For example, the image processing device can calculate a difference between the original reference frame and the inferred reference frame. Specifically, the image processing device can obtain a first occlusion map (514) by calculating a difference between a first frame (510) and a dewarped first frame (512), and can obtain a second occlusion map (524) by calculating a difference between a second frame (520) and a dewarped second frame (522). The first occlusion map (514) corresponds to the first frame (510) and may be referred to as a first base occlusion map. The second occlusion map (524) corresponds to the second frame (520) and may be referred to as a second base occlusion map.

[0087] Meanwhile, occlusion maps can include occlusion intensity, which indicates the degree of error. A larger occlusion intensity indicates a greater difference between the original and the inference, while a smaller occlusion intensity indicates a smaller difference between the original and the inference.

[0088] FIG. 6 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate an occlusion map corresponding to a synthesized frame.

[0089] In explaining FIG. 6, the example of FIG. 5, in which the light field images are composed of frames 1 to 9, is continued.

[0090] In one embodiment, the image processing device can generate an occlusion map corresponding to the synthesized frame. The image processing device can generate an occlusion map corresponding to intermediate viewpoints using the base occlusion map and the flow map.

[0091] For example, referring to FIG. 6, an occlusion map (630) corresponding to a viewpoint T between the first base occlusion map (610) and the second base occlusion map (620) can be generated. In the example of FIG. 6, since the light field images are composed of nine images, the number of viewpoints corresponding between the first base occlusion map (610) and the second base occlusion map (620) can be seven, and accordingly, the number of occlusion maps (630) corresponding to viewpoint T can be seven.

[0092] An image processing device can generate an occlusion map corresponding to an intermediate viewpoint T using base occlusion maps and a flow map. For example, the image processing device can generate a first warped occlusion map using a first base occlusion map (610) and a flow map for moving the viewpoint of the first base occlusion map (610) to viewpoint N. In addition, the image processing device can generate a second warped occlusion map using a second base occlusion map (620) and a flow map for moving the viewpoint of the second base occlusion map (620) to viewpoint N. The image processing device can generate an occlusion map corresponding to viewpoint N by synthesizing the first warped occlusion map and the second warped occlusion map. In this case, synthesis weights having different values ​​may be applied to the first warped occlusion map and the second warped occlusion map.

[0093] In one embodiment, the image processing device may generate masks representing a reconstructed area of ​​the synthesized frame based on occlusion maps corresponding to the synthesized frame. For example, the image processing device may generate a mask (640) representing a reconstructed area of ​​the synthesized frame at time point T based on an occlusion map (630) corresponding to time point T.

[0094] The image processing device can generate a mask based on an occlusion intensity of an occlusion map.

[0095] For example, the image processing device may determine the top n pixels with high occlusion intensities as pixels corresponding to the restoration area based on a preset threshold value n. The image processing device may generate a restoration area mask by designating the values ​​of the pixels corresponding to the restoration area as 1 and the values ​​of the pixels corresponding to the non-restoration area as 0.

[0096] For example, the image processing device may determine pixels having an occlusion intensity value greater than or equal to m as restoration areas based on a preset threshold value m. The image processing device may generate a restoration area mask by designating the values ​​of pixels corresponding to the restoration area as 1 and the values ​​of pixels corresponding to the non-restoration area as 0.

[0097] In one embodiment, the threshold applied by the image processing device may be adaptively changed. Information regarding the adaptively changed threshold may be received from the encoding device or may be changed according to the data transmission environment (e.g., transmission speed) of the image processing device.

[0098] When the threshold value changes, the number of pixels determined as the restoration area changes. For example, when the threshold value increases, the number of pixels determined as the restoration area decreases, and when the threshold value decreases, the number of pixels determined as the restoration area increases. From an encoding perspective, as the threshold value adaptively changes, the amount of restoration area changes, and the amount of residual data for restoring the restoration area changes. In other words, the amount of residual data transmitted can be adjusted. In other words, the encoding device can adaptively change the threshold value according to the data transmission environment of the encoding device to adjust the amount of residual data transmitted.

[0099] In terms of decoding, the threshold value used during encoding can be received from the encoding device, or the threshold value can be adaptively changed according to the data transmission environment of the image processing device. Accordingly, the image processing device can adjust the size of the restoration area to correspond to the amount of encoded residual data.

[0100] FIG. 7 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate an index map.

[0101] In one embodiment, the image processing device can generate a mask (710) representing a restoration area. Before describing the operation of the image processing device generating an index map, the operation of the image processing device generating the mask (710) will be described in more detail.

[0102] An image processing device can identify the occlusion intensity of an occlusion map. The image processing device can sort pixels of the occlusion map based on the occlusion intensity. Pixels with high occlusion intensity may be referred to as "restored pixels" as they represent restoration areas that need to be restored.

[0103] The image processing device can identify the index (730) of the restored pixel. The index of the restored pixel may correspond to the sorted pixels of the occlusion map. For example, referring to FIG. 7, the first value of the index (730) of the restored pixel may indicate the index of the pixel with the highest occlusion intensity in the occlusion map.

[0104] The index value may include viewpoint information and pixel position information. That is, viewpoint information and pixel position information may be calculated based on the index value.

[0105] For example, the view information view can be the quotient of the index value divided by the image resolution (H*W). This can be expressed as a formula as follows.

[0106] view = [Index / (H*W)]

[0107] For example, the vertical position information h of a pixel can be the quotient of the remainder after dividing the index value by the image resolution (H*W) and then dividing it by the horizontal resolution W. This can be expressed as a formula as follows.

[0108] h=[Index % (H*W) / W]

[0109] For example, the horizontal position information w of a pixel can be the remainder obtained by dividing the index value by the image resolution (H*W), which is then divided by the horizontal resolution W. This can be expressed as a formula as follows.

[0110] h=Index % (H*W) % W

[0111] Referring again to Figure 7, when the resolution of the occlusion map is 3840*2160 as an example, the first value of the index (730) of the restored pixel is '1289', which means that the pixel with the highest occlusion intensity among the occlusion maps is the 1289th pixel in width and the 0th pixel in height (0,1289,0) at the 0th point in time.

[0112] The image processing device can select restored pixels by applying an adaptively changing threshold value after sorting pixels of an occlusion map according to occlusion intensity and generating an index (730) of restored pixels corresponding to the pixels of the sorted occlusion map. For example, the top n pixels among the pixels of the sorted occlusion map can be determined as restored pixels, or the pixels of the sorted occlusion map having an occlusion intensity value of m or greater can be determined as restored pixels.

[0113] The image processing device can generate a mask (710) based on the determined restored pixels.

[0114] In one embodiment, the image processing device can process the mask (710) to generate a filtered mask (720). The image processing device can generate the filtered mask (720) by applying a morphological operation. The morphological operation process can include applying a Gaussian filter to remove noise from the mask (710). By applying the morphological operation to the mask (710) to generate the filtered mask (720), the compression efficiency can be increased during the encoding process, and the visibility of the restoration can be increased during the decoding process.

[0115] The image processing device can generate an index map (740) for mapping the restoration area and the residual image using the mask (710) and / or the filtered mask (720).

[0116] FIG. 8 is a drawing for explaining a mask and index generated by an image processing device according to one embodiment of the present disclosure.

[0117] In explaining Fig. 8, for convenience of explanation, an image resolution of 26*14 is used as an example.

[0118] Referring to FIG. 8, image 1 (810) illustrates a mask and index corresponding to the first synthesized frame (time point N=0). The shaded area within image 1 (810) represents the restored area of ​​the mask. The value within each pixel represents an index value. The index value may include time point information and pixel position information, which are information indicating which pixel of which time point the restored area is.

[0119] For example, in image 1 (810), each pixel has an index value that starts from 0 and sequentially increases. For example, an index value of 0 may mean the (0,0) pixel of the synthesized frame corresponding to the intermediate time point N=0 (the time point of the first synthesized frame). For example, an index value of 26 may mean the (0,1) pixel of the synthesized frame corresponding to the intermediate time point N=0. Similarly, an index value of 363 may mean the (25,13) pixel of the synthesized frame corresponding to the intermediate time point N=0 (the last pixel of the first synthesized frame). Since the method of calculating the time point information and the pixel position information from the index values ​​has been described above in the description of FIG. 7, a repeated description is omitted for brevity.

[0120] Image 2 (820) illustrates a mask and index corresponding to the second composite frame (time point N=1). Since image 2 (820) corresponds to a composite frame at a different time point than image 1 (810), the position of the mask indicating the restoration area may be different from that of image 1 (810).

[0121] In image 2 (820), each pixel has a value that sequentially increases starting from 364. For example, an index value of 364 may mean the (0,0) pixel of the synthesized frame corresponding to the intermediate time point N=1 (the time point of the second synthesized frame). For example, an index value of 390 may mean the (0,1) pixel of the synthesized frame corresponding to the intermediate time point N=1. Similarly, an index value of 727 may mean the (25,13) pixel of the synthesized frame corresponding to the intermediate time point N=1 (the last pixel of the second synthesized frame). Since the method of calculating the time point information and the pixel position information from the index value has been described above in the description of FIG. 7, a repeated description is omitted for brevity.

[0122] Meanwhile, although only two images are shown in FIG. 8 for illustration purposes, masks and indices can be generated for all intermediate time points (N=0,1,2, ...) between the time points of the reference frames.

[0123] FIG. 9 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to estimate the location of a restoration area.

[0124] In one embodiment, the image processing device can estimate the location of the restoration area based on a mask (910). The mask (910) may be the mask (710) or the filtered mask (720) illustrated in FIG. 7.

[0125] The image processing device can scan the mask (910) in block units having a preset size. The image processing device can identify a valid restoration area by scanning the mask (910) in block units having a preset size. For example, if the ratio of the restoration area within a block unit is greater than or equal to a preset value (e.g., n%), the image processing device can identify the block as a valid block and determine it as a restoration area. The determined restoration area corresponds to residual pixels of the block unit stored in the residual image (920).

[0126] The image processing device can identify a restoration area on a pixel-by-pixel basis by repeatedly scanning the mask (910) while gradually reducing the block size. The initial block size and the gradually reduced block size may be preset.

[0127] For example, the initial size of a block unit may be 64*64. The image processing device may scan the mask (910) in a size of 64*64 and identify the block as a valid block if the ratio of the restoration area within the block is n1% or more. The image processing device may store valid block information and change the pixel values ​​that have been stored within the mask (910) from 1 to 0 to prevent duplicate scanning. In this case, a pixel value of 1 may represent a restoration area, and a pixel value of 0 may represent a non-restoration area.

[0128] After the first scan of the mask (910) is performed, the image processing device can reduce the block size and perform a scan on the mask (910) again. For example, the image processing device can scan the mask (910) in units of blocks of 32*32 size and identify the block as a valid block if the ratio of the restoration area within the block is n2% or more. The image processing device can estimate the restoration area within the mask (910) on a pixel basis by repeating the scan on the mask (910) while reducing the block size to 16*16, 8*8, 4*4, 2*2, and 1*1. The estimated restoration area may correspond to the residual blocks stored in the residual image (920).

[0129] Meanwhile, since there are multiple synthetic frames corresponding to multiple time points, there may be multiple masks (910) for estimating the restoration area of ​​the synthetic frames. In this case, a block-by-block iterative scan may be performed on the multiple masks (910). For example, the image processing device may estimate the positions of the restoration area of ​​the synthetic frames by scanning the multiple masks (910) with an initial block size. After the scan for the multiple masks (910) with the initial block size is performed, the image processing device may repeatedly perform a scan for the multiple masks (910) while reducing the block size, thereby estimating the restoration area on a pixel-by-pixel basis.

[0130] Meanwhile, the image processing device can generate an index map for mapping the synthesized frame and the residual image (920) based on the locations of the estimated restoration area. This is further described with reference to FIG. 10.

[0131] FIG. 10 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate an index map.

[0132] Referring to FIG. 10, an image (1010) illustrates a mask and an index. An image processing device can generate an index map (1020) by arranging index values ​​indicating the location of a restoration area within the mask on a block basis.

[0133] The size of the index map (1020) may correspond to the size of the residual image. That is, the pixels of the index map (1020) correspond one-to-one to the pixels of the residual image. In addition, the index values, which are pixel values ​​of the index map (1020), may include viewpoint information and pixel position information. Therefore, based on the index value information of the index map (1020), the pixels of the residual image can be applied to the composite frame.

[0134] For example, the viewpoint information and pixel position information of 'pixel a', which is a pixel at a predetermined position in the index map (1020), may indicate the nth viewpoint and pixel position (x, y). In this case, the residual pixel at the position corresponding to 'pixel a' in the residual image may be applied to the pixel position (x, y) of the synthesized frame at the nth viewpoint, thereby compensating for the residual.

[0135] Referring again to FIG. 10, the image processing device can generate an index map (1020) by storing index values ​​in block units. In this case, only valid blocks determined to be restoration areas within the mask can be included in the index map (1020).

[0136] For example, the shaded area in the image (1010) represents the restored area of ​​the mask. The image processing device can scan the mask in block units of a preset size (for example, 2*2 size).

[0137] As a result of the block-by-block scan, a block whose restoration area ratio within the block is greater than or equal to a preset value can be identified as a valid block. Accordingly, a block including index values ​​[70, 71, 96, 97] can be identified as the first valid block. The image processing device (1020) can store an index indicating a restoration area identified as a valid block in the index map (1020).

[0138] And, as a result of the continuous block-by-block scan, a block including index values ​​[72, 73, 98, 99] may be identified as a second valid block. The image processing device (1020) may store an index indicating a restoration area identified as a valid block in the index map (1020). In this case, the valid block may include a non-restoration area (mask value: 0) in addition to the restoration area (mask value: 1). Since the non-restoration area is an area that does not need to be restored in the synthesized frame, the image processing device may set the index value corresponding to the non-restoration area to a dummy value (e.g., 0). As a result, the index values ​​of [72, 0, 98, 0] may be stored in the index map (1020). Meanwhile, since the residual image stores residuals corresponding to the restoration area, the value corresponding to the non-restoration area in the residual image may also be 0. In other words, the value of a pixel in the residual image corresponding to a pixel corresponding to a dummy index of the index map (1020) may be 0.

[0139] In addition, if block-by-block scanning is continuously performed, a block containing index values ​​[124, 125, 150, 151] represents a restoration area only for pixels corresponding to index value 124. In this case, the block may not be identified as a valid block because the restoration area value is less than a preset value. Later, the block may be identified as a valid block in a 1*1 pixel-by-pixel scanning process when the block size is reduced.

[0140] Additionally, if block-by-block scanning is continuously performed, a block containing index values ​​[266, 267, 292, 293] may be identified as the third valid block. The image processing device may store only the indices corresponding to the restoration area within the valid block in the index map (1020). As a result, the index values ​​of [0, 267, 292, 293] may be stored in the index map (1020).

[0141] The image processing device can identify which pixel of the residual image is which pixel in the synthesized frame at which point in time based on pixel information (index value) of the index map (1020) corresponding to the pixel in the residual image, and apply residual compensation to the synthesized frame.

[0142] FIG. 11 is a drawing for generally explaining a process of restoring light field images by an image processing device according to one embodiment of the present disclosure.

[0143] An image processing device can obtain compressed data of a large amount of light field video from an encoding device. The encoded data can include an encoded reference frame, an encoded flow map, and an encoded residual image.

[0144] In operation S1110, the image processing device can perform an image decoding task. The image processing device can decode the encoded data, the encoded reference frame, the encoded flow map, and the encoded residual image, thereby obtaining the reference frame, the flow map, and the residual image.

[0145] In operation S1115, the image processing device can perform a synthesis operation. The image processing device can use the reference frame and the flow map to generate synthesis frames corresponding to viewpoints between reference frames having different viewpoints. As a result of the synthesis operation, synthesis frames corresponding to N viewpoints can be generated.

[0146] In operation S1120, the image processing device can calculate an occlusion map. The image processing device can generate a base occlusion map corresponding to the reference frame using the reference frame and the flow map. The image processing device can generate an occlusion map corresponding to the synthesized frame using the base occlusion map and the flow map.

[0147] In operation S1130, the image processing device can perform an adaptive thresholding operation. The image processing device can apply a threshold value that adaptively changes to the occlusion intensity of the occlusion map to generate a mask representing a restoration area.

[0148] In operation S1140, the image processing device may generate a filtered mask by applying a morphological operation to the mask. Once the filtered mask is obtained, the image processing device may perform a restoration area location estimation operation to identify pixels corresponding to the restoration area within the mask. The image processing device may identify index values ​​including viewpoint information and pixel location information of pixels estimated to be in the restoration area.

[0149] In operation S1150, the image processing device can rearrange index values ​​in units of blocks to generate an index map. The image processing device can scan a mask in units of blocks to identify valid blocks estimated to be restoration areas within the mask. The image processing device can store the index values ​​of valid blocks in units of blocks to generate an index map.

[0150] In operation S1160, the image processing device can perform a restoration operation. The image processing device can apply the residual image to the composite frames by referring to the index map. The composite frames may be generated in operation S1115. As a result, the image processing device can obtain restored light field images. The restored light field images may include a decoded reference frame and a composite frame in which the residual is compensated.

[0151] Meanwhile, since encoding and decoding are processes in opposite directions, light field images can be processed in an encoding device in a manner similar to operations S1110 to S1160. The encoding device may be an image processing device of the present disclosure. Alternatively, a first image processing device may perform encoding, and a second image processing device may receive encoded data from the first image processing device and perform decoding.

[0152] Below, the operations performed by the image processing device to perform encoding are described. During the encoding process, the image processing device may perform the aforementioned operations of the image processing device to find a restoration area and generate a residual image.

[0153] The image processing device can select reference images from among the light field images. The reference images can be selected based on preset criteria (e.g., a predefined frame interval, a predefined viewpoint interval, etc.).

[0154] An image processing device can generate a flow map including pixel movement information. Various known algorithms for generating flow maps can be used to generate the flow map.

[0155] An image processing device can generate a composite frame corresponding to viewpoints between the reference frames using reference frames and a flow map. There may be multiple composite frames.

[0156] An image processing device can encode reference frames and a flow map. Furthermore, the image processing device can decode the encoded reference frames and flow map again. That is, the image processing device can generate a composite frame using the reference frames and flow map, which have undergone encoding and decoding operations in the encoding step. This is to reflect noise in advance when generating a composite frame in the subsequent decoding step.

[0157] An image processing device can generate a base occlusion map corresponding to the reference frames using the reference frames and the flow map. Furthermore, the image processing device can generate an occlusion map corresponding to the synthesized frame using the base occlusion map and the flow map. In the example described above, the base occlusion map may be generated using the reference frames and the flow map that have undergone encoding and decoding operations.

[0158] An image processing device can perform adaptive thresholding. The image processing device can apply a threshold value that adaptively changes based on the occlusion intensity of an occlusion map to generate a mask representing a restoration area.

[0159] An image processing device can generate a filtered mask by applying a morphological operation to the mask. Once the filtered mask is obtained, the image processing device can perform a restoration area location estimation operation to identify pixels corresponding to the restoration area within the mask. The image processing device can identify index values ​​including viewpoint information and pixel location information of pixels estimated to be in the restoration area.

[0160] An image processing device can generate an index map by rearranging index values ​​block by block. The image processing device can scan a mask block by block to identify valid blocks within the mask that are estimated to be restoration areas. The image processing device can generate an index map by storing the index values ​​of valid blocks block by block.

[0161] An image processing device can generate a residual image. The image processing device can generate a residual image including a residual to be applied to synthesized frames by referring to an index map. The residual image may have the same size as the index map. The residual image may be an aggregate of areas where there are differences between original light field images of multiple viewpoints and synthesized frames of multiple viewpoints, on a pixel basis. Meanwhile, the synthesized frame used to generate the residual image may be generated using reference frames and a flow map that have undergone encoding and decoding operations in the encoding step. In other words, the image processing device may generate the residual image based on a workpiece that has undergone an encoding-decoding step in advance in the encoding step in order to consider noise that may occur in the subsequent decoding step. Accordingly, the quality of the restored image generated by the image restoration operation S1160 that applies the residual and the synthetic frame generated by the synthetic frame generation operation S1115 in the subsequent decoding step can be improved.

[0162] FIG. 12 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to generate an index map.

[0163] In one embodiment, the image processing device can generate an index map using the occlusion map.

[0164] An image processing device can generate base occlusion maps corresponding to reference frames and generate occlusion maps (1210) corresponding to synthesized frames based on the base occlusion maps. The generation operation of the occlusion maps (1210) has been described in the description of the previous drawings, and thus a repeated description is omitted. The occlusion maps (1210) may correspond to each of a plurality of different viewpoints.

[0165] The image processing device may generate an index map (1220) based on the occlusion intensity of each pixel of the occlusion maps. For example, the image processing device may generate the index map (1220) using the arg max function. Specifically, the image processing device may obtain an array of pixels located at the same position in the occlusion maps. In this case, the arg max function may return viewpoint information of the occlusion map having the largest occlusion intensity value within the array of pixels located at the same position in the occlusion maps. When the image processing device applies arg max to the occlusion maps, an index value representing viewpoint information of the occlusion map, such as 1, 2, 3, ..., may be extracted for each pixel of the index map (1220). The image processing device may store the index value representing the returned viewpoint information in the index map (1220).

[0166] The generated index map (1220) can be used to generate a residual image that stores the residual between the original and synthesized frames during the encoding process, and can be used to compensate the residual of the residual image to the synthesized frame during the decoding process.

[0167] Meanwhile, the image processing device can maintain the pixel coordinates of the residual image representing the residual between the original and the composite by using an index map utilizing arg max. In other words, the image processing device can improve compression efficiency by maintaining locality when compressing data by using an index map utilizing arg max.

[0168] FIG. 13 is a diagram for explaining an operation of an image processing device according to one embodiment of the present disclosure to restore an image using an index map.

[0169] In explaining Fig. 13, it is explained as an example that the composite frames are composite frames corresponding to four viewpoints. Referring to Fig. 13, a residual image (1310), an index map (1320), and restored images (1330) are illustrated.

[0170] In one embodiment, the image processing device can generate restored images by applying the residual image (1310) to the composite frame with reference to the index map (1320).

[0171] In this case, the index map (1320) includes viewpoint information. For example, if the index value of (h, w) in the index map (1320) is n (n=1, 2, 3, 4), the pixel located at (h, w) in the residual image is applied to (h, w) of the synthesized frame. Specifically, for example, if the index value of the pixel (128, 128) in the index map (1320) is 3, the pixel (128, 128) of the residual image can be used to restore the synthesized frame corresponding to viewpoint 3.

[0172] FIG. 14 is a block diagram illustrating a configuration of an image processing device according to one embodiment of the present disclosure.

[0173] In one embodiment, the image processing device (2000) may correspond to the encoding device and / or the decoding device of the present disclosure. The image processing device (2000) may perform the encoding operation or the decoding operation of the present disclosure. For example, the image processing device (2000) may perform the encoding operations of the present disclosure as an encoding device. Alternatively, the image processing device (2000) may perform the decoding operations of the present disclosure as a decoding device. The encoding device and the decoding device may be implemented as the same device or as different devices. For example, the image processing device (2000) may perform the encoding / decoding operations as an encoding / decoding device. Alternatively, a first image processing device may perform the encoding operations as an encoding device, and a second image processing device may perform the decoding operations as a decoding device.

[0174] In one embodiment, the image processing device (2000) may include a memory (2100) and a processor (2200).

[0175] The memory (2100) may store instructions, data structures, and program codes that can be read by the processor (2200). Operations performed by the processor (2200) may be implemented by executing instructions or codes of a program stored in the memory (2100).

[0176] The memory (2200) may include a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), and may include a non-volatile memory including at least one of a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, and an optical disk, and a volatile memory such as a RAM (Random Access Memory) or an SRAM (Static Random Access Memory).

[0177] The memory (2100) may store one or more instructions and / or programs that cause the image processing device (2000) to operate to process an image. For example, the memory (2100) may store instructions and / or programs for implementing functions of encoding, flow map generation, synthetic frame generation, occlusion map generation, restoration area estimation, index map generation, and residual image generation of an encoding device. In addition, for example, the memory (2100) may store instructions and / or programs for implementing functions of decoding, synthetic frame generation, occlusion map generation, restoration area estimation, index map generation, and residual compensation of a decoding device.

[0178] The processor (2200) is a circuit device that can control the overall operations of the image processing device (2000). For example, the processor (2200) can control the overall operations of the image processing device (2000) to perform encoding, decoding, image synthesis, image restoration, etc. by executing one or more instructions of a program stored in the memory (2100). There may be one or more processors (2200).

[0179] The processor (2200) may be configured with at least one of, but is not limited to, a central processing unit, a microprocessor, a graphic processing unit, an application specific integrated circuits (ASICs), a digital signal processor (DSPs), a digital signal processing device (DSPDs), a programmable logic device (PLDs), a field programmable gate array (FPGAs), an application processor, a neural processing unit, or an artificial intelligence processor designed with a hardware structure specialized for processing an artificial intelligence model.

[0180] The processor (2200) can write data to the memory (2100) or read data stored in the memory (2100). For example, the processor (2200) can load one or more instructions and / or programs stored in a non-volatile memory into a volatile memory and process data and program code. The operations of encoding, flow map generation, synthetic frame generation, occlusion map generation, restoration area estimation, index map generation, and residual image generation of the encoding device executed by the processor (2200) and the operations of decoding, synthetic frame generation, occlusion map generation, restoration area estimation, index map generation, and residual compensation of the decoding device have already been described in the description of the previous drawings, and therefore, a repeated description will be omitted.

[0181] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an AI-specific processor). Here, an AI-specific processor, which is an example of the second processor, may perform operations for training / inference of an AI model. However, the embodiments of the present disclosure are not limited thereto.

[0182] One or more processors according to the present disclosure may be implemented as a single-core processor or as a multi-core processor.

[0183] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core or may be performed by a plurality of cores included in one or more processors.

[0184] FIG. 15 is a block diagram illustrating a configuration of an image processing device according to one embodiment of the present disclosure.

[0185] In one embodiment, the image processing device (2000) may include a memory (2100), a processor (2200), a communication interface (2300), and a display. The memory (2100) and the processor (2200) of FIG. 15 may correspond to the memory (2100) and the processor (2200) of FIG. 14. Therefore, for the sake of brevity, a repeated description is omitted.

[0186] The communication interface (2300) can perform data communication with other electronic devices under the control of the processor (2200). For example, when the image processing device (2000) operates as an encoding device, the image processing device (2000) can transmit encoded data to a decoding device. Alternatively, when the image processing device (2000) operates as a decoding device, the image processing device (2000) can receive encoded data from the encoding device.

[0187] The communication interface (2300) may include a communication circuit that can perform data communication between the image device (2000) and another electronic device using at least one of data communication methods including, for example, wired LAN, wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliances (WiGig), and RF communication.

[0188] The display (2400) can output image signals to the screen of the image processing device (2000) under the control of the processor (2200). For example, the display (2400) can output original light field images to the screen. For example, the display (2400) can output restored light field images including composite frames to the screen.

[0189] The present disclosure relates to a method for efficiently compressing, transmitting, and restoring large-capacity light field images. The technical challenges addressed by the present disclosure are not limited to those described above, and other technical challenges not mentioned herein will be readily apparent to those skilled in the art, based on the description herein.

[0190] According to one aspect of the present disclosure, a method for processing light field images may be provided.

[0191] The method may include obtaining a first frame representing reference frames captured from different points in time, a second frame, and a flow map and residual image for generating a composite frame.

[0192] The method may include a step of generating a composite frame based on the first frame, the second frame, and the flow map.

[0193] The method may include a step of generating an occlusion map corresponding to the synthesized frame.

[0194] The method may include a step of generating an index map corresponding to the residual image based on the occlusion map.

[0195] The method may include a step of applying the residual image to the synthesized frame based on the index map.

[0196] The method may include a step of outputting light field images including the first frame, the second frame, and the composite frame.

[0197] The method may include obtaining an encoded first frame, an encoded second frame, an encoded flow map, and an encoded residual image.

[0198] The step of obtaining the first frame, the second frame, the flow map, and the residual image may include the step of decoding the encoded first frame, the encoded second frame, the encoded flow map, and the encoded residual image.

[0199] The above residual image may be an aggregate of areas where there are differences between original light field images of multiple viewpoints and synthesized frames of multiple viewpoints, in pixel units.

[0200] The pixels of the above index map may correspond to pixels of the above residual image.

[0201] Each pixel of the above index map may include viewpoint information and pixel location information.

[0202] The step of generating the occlusion map may include the step of generating a base occlusion map corresponding to the first frame and the second frame.

[0203] The step of generating the occlusion map may include a step of warping the base occlusion map to generate the occlusion map corresponding to the synthesized frame.

[0204] The method may include a step of generating a mask representing a restoration area of ​​the synthesized frame based on the occlusion map.

[0205] The method may include a step of estimating a location of a restoration area of ​​the synthetic frame based on the mask.

[0206] The step of generating the mask may be to generate the mask by applying an adaptively changing threshold value to the occlusion map.

[0207] The step of estimating the location of the above restoration area may be to estimate the location of the restoration area by scanning the mask in block units having a preset size.

[0208] The step of generating the above index map may be to generate the index map corresponding to the size of the residual image by arranging index values ​​indicating the location of the restoration area according to the block unit.

[0209] The step of generating the above occlusion map may be generating a plurality of occlusion maps corresponding to a plurality of synthesized frames.

[0210] The step of generating the index map may be to generate an index map including viewpoint information indicating which pixel of the residual image is applied to which of the plurality of synthesized frames, based on pixel values ​​existing at the same location in the plurality of occlusion maps.

[0211] According to one aspect of the present disclosure, an image processing device may be provided.

[0212] The image processing device may include a memory storing one or more instructions; and one or more processors executing the one or more instructions stored in the memory.

[0213] The one or more processors can obtain a first frame representing reference frames captured from different points in time, a second frame, and a flow map and residual image for generating a composite frame by executing the one or more instructions.

[0214] The one or more processors can generate a composite frame based on the first frame, the second frame, and the flow map by executing the one or more instructions.

[0215] The one or more processors can generate an occlusion map corresponding to the synthesized frame by executing the one or more instructions.

[0216] The one or more processors can generate an index map corresponding to the residual image based on the occlusion map by executing the one or more instructions.

[0217] The one or more processors can apply the residual image to the composite frame based on the index map by executing the one or more instructions.

[0218] The one or more processors can output light field images including the first frame, the second frame, and the composite frame by executing the one or more instructions.

[0219] The one or more processors can obtain an encoded first frame, an encoded second frame, an encoded flow map, and an encoded residual image by executing the one or more instructions.

[0220] The one or more processors can decode the encoded first frame, the encoded second frame, the encoded flow map, and the encoded residual image by executing the one or more instructions.

[0221] The above residual image may be an aggregate of areas where there are differences between original light field images of multiple viewpoints and composite frames of multiple viewpoints, in pixel units.

[0222] The pixels of the above index map may correspond to pixels of the above residual image.

[0223] Each pixel of the above index map may include viewpoint information and pixel location information.

[0224] The one or more processors can generate a base occlusion map corresponding to the first frame and the second frame by executing the one or more instructions.

[0225] The one or more processors can generate the occlusion map corresponding to the synthesized frame by warping the base occlusion map by executing the one or more instructions.

[0226] The one or more processors can generate a mask representing a restoration area of ​​the synthesized frame based on the occlusion map by executing the one or more instructions.

[0227] The one or more processors can estimate the location of the restoration area of ​​the synthetic frame based on the mask by executing the one or more instructions.

[0228] The one or more processors can generate the mask by applying an adaptively changing threshold value to the occlusion map by executing the one or more instructions.

[0229] The one or more processors can estimate the location of the restoration area by scanning the mask in units of blocks having a preset size by executing the one or more instructions.

[0230] The one or more processors can generate the index map corresponding to the size of the residual image by arranging index values ​​indicating the location of the restoration area in units of blocks by executing the one or more instructions.

[0231] The one or more processors can generate a plurality of occlusion maps corresponding to the plurality of synthesized frames by executing the one or more instructions.

[0232] The one or more processors can generate an index map including viewpoint information indicating which pixel of the residual image is applied to which of the plurality of synthesized frames based on pixel values ​​existing at the same location in the plurality of occlusion maps by executing the one or more instructions.

[0233] According to one aspect of the present disclosure, a method for compressing light field images may be provided.

[0234] The method may include a step of selecting reference frames including a first frame and a second frame from among original light field images including a plurality of viewpoints.

[0235] The method may include a step of generating a flow map for inferring frames between the first frame and the second frame.

[0236] The method may include a step of generating a composite frame based on the first frame, the second frame, and the flow map.

[0237] The method may include a step of generating an occlusion map corresponding to the synthesized frame.

[0238] The method may include a step of generating an index map corresponding to a residual image based on the occlusion map.

[0239] The method may include a step of generating a residual image including a residual to be applied to a synthesized frame based on the index map.

[0240] The method may include encoding the first frame, the second frame, the flow map, and the residual image.

[0241] The method may include a step of transmitting the encoded first frame, the encoded second frame, the encoded flow map, and the encoded residual image to a decoding device.

[0242] The above residual image may be an aggregate of areas where there are differences between original light field images of multiple viewpoints and synthesized frames of multiple viewpoints, in pixel units.

[0243] The step of generating the occlusion map may include the step of generating a base occlusion map corresponding to the first frame and the second frame.

[0244] The step of generating the occlusion map may include a step of warping the base occlusion map to generate the occlusion map corresponding to the synthesized frame.

[0245] The method may include a step of generating a mask representing a restoration area of ​​the synthesized frame based on the occlusion map.

[0246] The method may include a step of estimating a location of a restoration area of ​​the synthetic frame based on the mask.

[0247] The step of generating the mask may be to generate the mask by applying an adaptively changing threshold value to the occlusion map.

[0248] The step of estimating the location of the above restoration area may be to estimate the location of the restoration area by scanning the mask in block units having a preset size.

[0249] The step of generating the above index map may be to generate the index map corresponding to the size of the residual image by arranging index values ​​indicating the location of the restoration area according to the block unit.

[0250] The step of generating the above occlusion map may be generating a plurality of occlusion maps corresponding to a plurality of synthesized frames.

[0251] The step of generating the index map may be to generate an index map including viewpoint information indicating which pixel of the residual image is applied to which of the plurality of synthesized frames, based on pixel values ​​existing at the same location in the plurality of occlusion maps.

[0252] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include computer storage media and communication media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.

[0253] Additionally, a computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0254] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0255] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.

[0256] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.

Claims

1. In the method of processing light field images, A step of acquiring a first frame representing reference frames captured from different points of view, a second frame, and a flow map and residual image for generating a composite frame; A step of generating a composite frame based on the first frame, the second frame, and the flow map; A step of generating an occlusion map corresponding to the above synthesized frame; A step of generating an index map corresponding to the residual image based on the occlusion map; A step of applying the residual image to the synthesized frame based on the index map; and A method comprising the step of outputting light field images including the first frame, the second frame, and the composite frame.

2. In paragraph 1, The above method, Further comprising the steps of obtaining an encoded first frame and an encoded second frame, and an encoded flow map and an encoded residual image, The step of obtaining the first frame, the second frame, the flow map, and the residual image comprises: A method comprising the steps of decoding the encoded first frame, the encoded second frame, the encoded flow map, and the encoded residual image.

3. In paragraph 1, The pixels of the above index map correspond to the pixels of the above residual image, A method wherein each pixel of the above index map includes viewpoint information and pixel location information.

4. In paragraph 1, The steps for generating the above occlusion map are: A step of generating a base occlusion map corresponding to the first frame and the second frame; and A method comprising the step of warping the base occlusion map to generate the occlusion map corresponding to the synthesized frame.

5. In paragraph 4, The above method, A step of generating a mask representing a restoration area of ​​a synthetic frame based on the above occlusion map; and A method further comprising a step of estimating the location of the restoration area of ​​the synthetic frame based on the mask.

6. In paragraph 5, The steps for creating the above mask are: A method for generating the mask by applying an adaptively changing threshold value to the occlusion map.

7. In paragraph 5, The step of estimating the location of the above restoration area is: A method for estimating the location of the restoration area by scanning the mask in block units having a preset size.

8. In paragraph 7, The steps for creating the above index map are: A method for generating an index map corresponding to the size of the residual image by arranging index values ​​indicating the location of the restoration area according to the block unit.

9. In the image processing device, Memory that stores one or more instructions; and comprising one or more processors that execute one or more instructions stored in the memory; The one or more processors, by executing the one or more instructions, Acquire a first frame representing reference frames captured from different points of view, a second frame, and a flow map and residual image for generating a composite frame, Generating a composite frame based on the first frame, the second frame, and the flow map, Generate an occlusion map corresponding to the above composite frame, Based on the above occlusion map, an index map corresponding to the residual image is generated, Applying the residual image to the synthesized frame based on the index map, An image processing device that outputs light field images including the first frame, the second frame, and the composite frame.

10. In paragraph 9, The one or more processors, by executing the one or more instructions, Obtaining an encoded first frame and an encoded second frame, an encoded flow map, and an encoded residual image, An image processing device that decodes the encoded first frame, the encoded second frame, the encoded flow map, and the encoded residual image.

11. In paragraph 9, The pixels of the above index map correspond to the pixels of the above residual image, An image processing device, wherein each pixel of the above index map includes viewpoint information and pixel position information.

12. In paragraph 9, The one or more processors, by executing the one or more instructions, Generate a base occlusion map corresponding to the first frame and the second frame, An image processing device that warps the base occlusion map to generate the occlusion map corresponding to the synthesized frame.

13. In paragraph 12, The one or more processors, by executing the one or more instructions, Based on the above occlusion map, a mask representing the restoration area of ​​the synthesized frame is generated, An image processing device that estimates the location of a restoration area of ​​the synthetic frame based on the mask.

14. In paragraph 13, The one or more processors, by executing the one or more instructions, An image processing device that estimates the location of the restoration area by scanning the mask in block units having a preset size.

15. In paragraph 14, The one or more processors, by executing the one or more instructions, An image processing device that generates an index map corresponding to the size of the residual image by arranging index values ​​indicating the location of the restoration area according to the block unit.

Citation Information

Patent Citations

  • Light field image super-resolution reconstruction method based on frequency domain analysis and deep learning

    CN113139898A

  • Light field rendering method based on scene layering

    CN116503536A

  • Tridimensional rendering with adjustable disparity direction

    KR1020170075656A

  • Multi-view scene segmentation and propagation

    KR1020180132946A

  • Method and Apparatus for Synthesis of Face Light Field On-Device

    KR102294806B1