Image processing apparatus, image processing method, and program
The image processing device enhances foreground contour accuracy by using a series of units to correct and refine contours, reducing processing load and maintaining image quality in virtual viewpoint image generation.
Patent Information
- Application Number
- JP2024021731
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-16
- Publication Date
- 2025-08-28
AI Technical Summary
Existing technologies struggle to achieve high accuracy in extracting foreground contours from reduced images when generating virtual viewpoint images, leading to degraded image quality.
An image processing device that includes an acquisition unit, reduction unit, estimation unit, enlargement unit, extraction unit, removal unit, contour correction unit, image correction unit, and mask generation unit, which work together to enhance the accuracy of foreground region extraction by correcting contours using edge images from the input image.
The device reduces processing load and achieves highly accurate foreground region extraction, minimizing image degradation in virtual viewpoint images.
Smart Images

Figure 2025125660000001_ABST
Abstract
Description
[Technical Field]
[0001] SUMMARY This disclosure relates to image processing techniques for identifying image regions that correspond to foreground. [Background technology]
[0002] There is a technology for generating an image (hereinafter referred to as a "virtual viewpoint image") corresponding to a view from an arbitrary virtual viewpoint using multiple images (hereinafter referred to as "captured images") obtained by capturing images using multiple imaging devices. For example, when live broadcasting captured images of a baseball or basketball game in real time, it is required to provide highlight scenes created using virtual viewpoint images in real time. In order to generate high-quality virtual viewpoint images in real time, it is necessary to achieve both high image quality and high speed in image processing. Patent Document 1 discloses a technology for achieving high image quality by performing image processing with a high processing load on a low-resolution image obtained by reducing an input image, thereby speeding up processing, and using pixels of the input image to calculate complementary pixels when restoring the resolution of the image after image processing. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-272878 Summary of the Invention [Problem to be solved by the invention]
[0004] In the process of extracting an image area corresponding to the foreground from a captured image (hereinafter referred to as the "foreground area"), which is one of the processes in the process of generating a virtual viewpoint image, it is necessary to identify the contour of the foreground with high accuracy. However, the technology disclosed in Patent Document 1 has a problem in that when the foreground area extracted from a reduced image is restored to the original resolution, it is not possible to obtain a high-accuracy foreground area. [Means for solving the problem]
[0005] The image processing device according to the present disclosure includes an acquisition means for acquiring data of an input image, a reduction means for acquiring a reduced image by reducing the input image, an estimation means for acquiring a foreground estimated image by estimating a foreground region in the reduced image, an enlargement means for acquiring an enlarged estimated image having the same resolution as the input image by enlarging the foreground estimated image, an extraction means for acquiring a first edge image by extracting contours from the input image, a limitation means for acquiring a second edge image by limiting the contours included in the first edge image to those necessary for correcting the enlarged estimated image, an image correction means for acquiring a corrected image by correcting the enlarged estimated image using the second edge image, and a generation means for generating, from the corrected image, a mask image indicating a foreground region corresponding to an image of an object that is in the foreground in the input image. [Effects of the Invention]
[0006] According to the present disclosure, it is possible to obtain a highly accurate foreground region while reducing the amount of calculation required for processing to extract the foreground region from an image. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a block diagram showing an example of an image processing system according to a first embodiment. [Figure 2] 1 is a block diagram showing an example of a hardware configuration of an image processing device according to a first embodiment. [Figure 3] 4 is a flowchart showing an example of a processing flow of the image processing device according to the first embodiment. [Figure 4] 1A and 1B are diagrams showing examples of images acquired or generated by the image processing apparatus according to the first embodiment. [Figure 5] 5A to 5C are diagrams for explaining an example of correction processing in a contour correction unit according to the first embodiment. [Figure 6] 10A and 10B are diagrams for explaining another example of correction processing in the contour corrector according to the first embodiment. [Figure 7] 10 is a flowchart showing an example of a processing flow of an image processing device according to a second embodiment. [Figure 8] FIG. 10 is a diagram showing an example of an image acquired or generated by the image processing apparatus according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. However, the means for solving the problems according to the present disclosure are not limited to the following embodiments. Furthermore, not all of the combinations of features described in the following embodiments are necessarily essential to the means for solving the problems according to the present disclosure. Note that the same components are given the same reference numerals, and redundant explanations will be omitted or simplified.
[0009] [Embodiment 1] <Image Processing System Overview> FIG. 1 is a block diagram showing an example of an image processing system 1 according to a first embodiment. The image processing system 1 includes a plurality of image capture devices 10, an image processing device 100, and an image generation device 11. The image processing device 100 and the image generation device 11 do not need to be built into the same housing, but may be configured as separate devices connected via a signal path that allows communication between them. Furthermore, the image processing device 100 does not need to be built into the same housing as the image capture devices 10, but may be configured as separate devices connected to the image capture devices 10 via a signal path that allows communication between them. FIG. 1 also shows an example of the functional configuration of the image processing device 100 according to the first embodiment. The image processing device 100 includes, as its functional configuration, an acquisition unit 101, a reduction unit 102, an estimation unit 103, an enlargement unit 104, an extraction unit 105, a removal unit 106, a contour correction unit 107, an image correction unit 108, a mask generation unit 109, and an expansion / contraction unit 110. The processing of each unit included in the image processing device 100 as its functional configuration will be described later.
[0010] The image processing system 1 generates an image (virtual viewpoint image) corresponding to the view from an arbitrary virtual viewpoint using an image (captured image) obtained by capturing an image using the imaging device 10. The virtual viewpoint image generated by the image processing system 1 is also called a free viewpoint image, and outputs a signal related to the virtual viewpoint image corresponding to a virtual viewpoint arbitrarily designated by a user. For example, the image processing system 1 generates a virtual viewpoint image corresponding to a virtual viewpoint selected by a user from a plurality of preset virtual viewpoint candidates, and outputs a signal related to the virtual viewpoint image. The designation of the virtual viewpoint is not limited to selection by the user, but may be performed automatically by AI (artificial intelligence) based on the results of image analysis, etc. Furthermore, the virtual viewpoint image may be a moving image or a still image.
[0011] The virtual viewpoint image is generated, for example, by the following method. First, an imaging area including a target object (hereinafter simply referred to as "object") is imaged from multiple directions by multiple imaging devices 10, and data of the captured images obtained by the imaging is output. Here, the imaging area is, for example, an area surrounded by a stadium field and an arbitrary height. The imaging area may be an area corresponding to a three-dimensional space from which the three-dimensional shape of the object is estimated. Furthermore, the three-dimensional space may be the entire imaging area or a part of the imaging area. Furthermore, the imaging area may be a concert hall, an imaging studio, or the like. The multiple imaging devices 10 are installed at different positions and in different directions (orientations) to surround the imaging area, and capture images synchronously with each other. Note that the multiple imaging devices 10 do not need to be installed around the entire circumference of the imaging area; they may be installed in only a partial direction relative to the imaging area depending on installation location restrictions, etc. Furthermore, the number of imaging devices 10 is not limited to a predetermined number. For example, if the imaging area is a rugby stadium, tens to hundreds of imaging devices 10 may be installed around the stadium.
[0012] The multiple imaging devices 10 may also include imaging devices 10 with different angles of view, such as telephoto cameras and wide-angle cameras. Here, a telephoto camera is an imaging device 10 with an optical system (lens) having a long focal length, and a wide-angle camera is an imaging device 10 with an optical system (lens) having a short focal length. For example, a virtual viewpoint image with higher resolution can be generated by capturing high-resolution images of athletes using a telephoto camera. Furthermore, when the imaging target is a ball game, the ball moves over a wide range. Therefore, in such a case, the number of imaging devices 10 can be reduced by capturing a wide range of the imaging target using a wide-angle camera. Furthermore, capturing images using a combination of wide-angle cameras and telephoto cameras can improve the flexibility of the installation location of the imaging devices 10. The imaging devices 10 are synchronized with a common time, and captured images are assigned information that identifies the time at which the image was captured. For example, when the imaging device 10 captures a moving image, each frame in the moving image is assigned information that identifies the time at which the frame was captured.
[0013] Next, the image processing device 100 extracts an area (hereinafter referred to as a "foreground area") corresponding to an image of a target object such as a person or a ball from each of the multiple captured images obtained by the imaging device 10. The image processing device 100 acquires a foreground image corresponding to the extracted foreground area and a background image corresponding to an area other than the foreground area (hereinafter referred to as a "background area"). The data of the foreground image and the background image includes texture information such as color information.
[0014] Finally, image generating device 11 generates data representing the three-dimensional shape of the object (hereinafter referred to as the "foreground model") and texture data for coloring the foreground model based on the foreground image. Image generating device 11 also generates texture data for coloring data representing the three-dimensional shape of the background of a stadium or the like (hereinafter referred to as the "background model") based on the background image. Image generating device 11 maps the corresponding texture data to the generated foreground model and background model, and generates a virtual viewpoint image by performing rendering according to the specified position of the virtual viewpoint and the line of sight at the virtual viewpoint.
[0015] A foreground image is an image in which a region of a foreground object (foreground region) is extracted from a captured image obtained by imaging by the imaging device 10. A foreground object refers to, for example, a dynamic object that moves when captured over time from the same direction, specifically, a dynamic object whose position or shape may change, i.e., a moving body. For example, if the subject of imaging is a sport, foreground objects include natural people such as players or umpires who are present on the field where the ball game is being played, and if the subject of imaging is a ball game, foreground objects include not only natural people but also the ball, etc. In addition, in a concert or entertainment show, natural people such as singers, musicians, performers, or presenters are included in the foreground objects.
[0016] A background image is an image that represents at least an area different from the area corresponding to the image of a foreground object, i.e., an area that serves as the background. Specifically, a background image is an image in which the image of the object in the foreground area has been removed from the captured image. Furthermore, the background refers to an imaged object, such as a structure or floor, that remains stationary or nearly stationary when images are captured in time series from the same direction. Examples of imaged background objects include a stage for a concert, a stadium where an event such as a sport is held, and structures and fields such as goals used in ball games. However, a background area is an area in an image that does not include at least the foreground area. Note that the imaged object may include, in addition to the foreground object and the background object, other objects that are different from these.
[0017] <Hardware configuration of image processing device> FIG. 2 is a block diagram showing an example of the hardware configuration of the image processing device 100 according to the first embodiment. The image processing device 100 is configured by a computer such as a personal computer, and includes a CPU 201, a ROM 202, a RAM 203, an auxiliary storage device 204, a display unit 205, an operation unit 206, a communication I / F 207, and a bus 208. The CPU 201 controls the entire image processing device 100 using computer programs and data stored in the ROM 202 or the RAM 203, thereby realizing the various units shown in FIG. 1 that are included in the functional configuration of the image processing device 100. Note that the image processing device 100 may include one or more dedicated processing hardware units different from the CPU 201, and at least a portion of the processing by the CPU 201 may be executed by the dedicated processing hardware units. Examples of the processing hardware include an FPGA (field programmable gate array) and a DSP (digital signal processor).
[0018] The ROM 202 is a storage medium for storing programs and the like that do not require modification. The RAM 203 is a storage medium for temporarily storing programs and data supplied from the auxiliary storage device 204 and data and the like supplied from the outside via the communication I / F 207. The RAM 203 is also used as a work area for the CPU 201. The auxiliary storage device 204 is configured with a hard disk drive or the like and is a storage device for storing various data, such as image data and audio data. The display unit 205 is configured with a liquid crystal display, LEDs, or the like and displays, for example, a GUI (Graphical User Interface) for the user to operate the image processing device 100. The operation unit 206 is configured with a keyboard, mouse, joystick, touch panel, or the like and inputs various instructions to the CPU 201 in response to, for example, user operations. The CPU 201 also operates as a display control unit that controls the display unit 205 and an operation control unit that controls the operation unit 206.
[0019] The communication I / F 207 is used for communication with devices external to the image processing device 100. For example, when the image processing device 100 is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 207, and when the image processing device 100 has a function for wireless communication with external devices, the communication I / F 207 is equipped with an antenna. The bus 208 is a transmission path that connects each unit that the image processing device 100 has as a hardware configuration so that they can communicate with each other. In this embodiment, the display unit 205 and the operation unit 206 are described as being present inside the image processing device 100, but at least one of the display unit 205 and the operation unit 206 may be present as a separate device outside the image processing device 100.
[0020] <Functional configuration of image processing device> 1, the processing of each functional configuration of the image processing device 100 will be described. The acquisition unit 101 acquires data of captured images obtained by imaging using each imaging device 10 as input image data. The input image data acquired by the acquisition unit 101 is transmitted to the reduction unit 102, extraction unit 105, and image generation device 11.
[0021] The reduction unit 102 receives input image data transmitted from the acquisition unit 101 and reduces the input image by an arbitrary magnification to generate a reduced image. In this embodiment, the reduced image generation process is described as being performed, for example, by thinning out pixels in the input image. By generating a reduced image by thinning out the pixels, the circuit scale of the processing hardware or the amount of calculation in the CPU 201 in the input image reduction process can be reduced. In this way, by reducing the input image to lower the resolution and then performing subsequent image processing, the amount of calculation in the image processing can be reduced, thereby realizing faster image processing. Note that the reduction process in the reduction unit 102 is not limited to the above-described thinning out process. For the reduction process, a method of thinning out pixels after filter processing or a known method such as a bicubic algorithm may also be applied. The reduced image data generated by the reduction unit 102 is transmitted to the estimation unit 103.
[0022] Estimation unit 103 receives data of the reduced image transmitted from reduction unit 102, estimates an area (foreground area) corresponding to the object in the reduced image, and generates a foreground estimated image showing the foreground area. Specifically, for example, estimation unit 103 inputs data of the reduced image into a trained model obtained as a result of learning such as machine learning, and acquires data of the foreground estimated image output by the trained model as an estimation result, thereby generating the foreground estimated image. For example, the trained model calculates, for each pixel of the input reduced image, an estimated value expressed using grayscale, which indicates the likelihood that the pixel contains an image of a foreground object. In the following, as an example, a description will be given assuming that a larger estimated value indicates a higher probability that the pixel contains an image of a foreground object.
[0023] In this embodiment, an estimated value is calculated using a trained model that has been trained so that the foreground region in the foreground estimation image output from the trained model covers the outer periphery of the actual foreground region, including several pixels. This is because, in the processing performed by the image correction unit 108 at the subsequent stage, an edge image generated from the input image is subtracted from a foreground region that is slightly larger than the actual foreground region, thereby faithfully reproducing the contour of the foreground estimation region. Details of the foreground estimation image will be described later using FIG. 4. Data of the foreground estimation image generated by the estimation unit 103 is transmitted to the enlargement unit 104.
[0024] Enlargement unit 104 receives data of the foreground estimated image output from estimation unit 103 and generates an enlarged estimated image by enlarging the foreground estimated image by a factor corresponding to the reciprocal of the reduction factor used in the reduction process of the input image by reduction unit 102. In this embodiment, pixel interpolation in the enlargement process of the foreground estimated image is described as being performed by repeatedly setting, in the enlarged estimated image, pixel values of the foreground estimated image corresponding to an area in the enlarged estimated image. By repeatedly setting pixel values in this manner, it is possible to reduce the circuit scale of the processing hardware or the amount of calculation in CPU 201 in the enlargement process of the foreground estimated image. The resolution of the enlarged estimated image generated by enlargement unit 104 is equal to the resolution of the input image.
[0025] The enlargement process in the enlargement unit 104 is not limited to the process of repeatedly setting pixel values as described above, and known methods such as a method of interpolating by filter processing, a method of interpolating by bicubic method, or super-resolution technology may be applied to the enlargement process. Details of the enlarged estimated image will be described later with reference to FIG. 4. Data of the enlarged estimated image generated by the enlargement unit 104 is transmitted to the removal unit 106, the contour correction unit 107, and the image correction unit 108.
[0026] The extraction unit 105 receives input image data transmitted from the acquisition unit 101 and extracts pixels corresponding to the contour of an object from the input image, thereby generating an edge image representing the contour. In this embodiment, the contour extraction process is described as being performed using a Sobel filter. Note that the contour extraction process is not limited to the method using the Sobel filter described above, and known methods using other algorithms capable of extracting the contour of an object in an image, such as a Laplacian filter or a Roberts filter, may also be applied to the contour extraction process. Hereinafter, the edge image generated by the extraction unit 105 will be referred to as a first edge image. Details of the first edge image will be described later with reference to FIG. 4. Data of the first edge image generated by the extraction unit 105 is transmitted to the removal unit 106.
[0027] The removal unit 106 receives data of the enlarged estimated image generated by the enlargement unit 104 and data of the first edge image generated by the extraction unit 105, and generates a second edge image by using the enlarged estimated image to remove unnecessary contours in the first edge image. The second edge image generated by the removal unit 106 is used to correct the enlarged estimated image in subsequent processing. In this embodiment, the removal unit 106 is described as removing contours in an area corresponding to the foreground area of the enlarged estimated image from the first edge image. Details of the second edge image will be described later using FIG. 4. The data of the second edge image generated by the removal unit 106 is transmitted to the contour correction unit 107.
[0028] The contour correction unit 107 receives data of the enlarged estimated image generated by the enlargement unit 104 and data of the second edge image generated by the removal unit 106, and generates a third edge image by emphasizing the contours included in the second edge image using the enlarged estimated image. Details of the third edge image will be described later with reference to Figs. 4 to 6. The data of the third edge image generated by the contour correction unit 107 is transmitted to the image correction unit 108.
[0029] The image correction unit 108 receives data of the enlarged estimated image generated by the enlargement unit 104 and data of the third edge image generated by the contour correction unit 107, and corrects the contour portion of the foreground region in the enlarged estimated image using the third edge image. Hereinafter, the image obtained by this correction process will be referred to as the corrected image. Specifically, the image correction unit 108 generates the corrected image by subtracting the value of each pixel in the enlarged estimated image from the value of the corresponding pixel in the third edge image. Details of the corrected image will be described later with reference to FIG. 4. The data of the corrected image generated by the image correction unit 108 is transmitted to the mask generation unit 109.
[0030] Mask generation unit 109 receives data of the corrected image generated by image correction unit 108 and binarizes the corrected image using a predetermined threshold value to generate a foreground mask image (hereinafter referred to as a "first foreground mask image") that indicates the foreground region in the input image. In the first foreground mask image according to this embodiment, pixels with a pixel value of "1" represent the foreground region, and pixels with a pixel value of "0" represent the background region, but the pixel values representing the foreground region and the background region may be reversed. Details of the corrected image will be described later using FIG. 4. Data of the first foreground mask image generated by mask generation unit 109 is transmitted to expansion / contraction unit 110.
[0031] Dilation / contraction unit 110 receives data of the first foreground mask image generated by mask generation unit 109 and performs erosion processing and dilation processing on the first foreground mask image to remove foreground dust generated by the correction processing by image correction unit 108. Hereinafter, the image obtained by this removal processing will be referred to as a second foreground mask image. Details of the second foreground mask image and the foreground dust generated by the correction processing by image correction unit 108 will be described later using FIG. 4. The second foreground mask image generated by expansion / contraction unit 110 is transmitted to image generation device 11 via communication I / F 207.
[0032] <Correction process for enlarged estimated image> In the image processing device 100 according to this embodiment, in order to reduce the processing load of the estimation unit 103, the reduction unit 102 reduces the input image to be processed, the estimation unit 103 performs foreground estimation processing on the reduced image, and the enlargement unit 104 enlarges the estimated foreground image. Here, if missing pixels are generated by interpolation processing when enlarging the estimated foreground image, the contour of the foreground region in the enlarged estimated foreground image (enlarged estimated image) does not necessarily faithfully reproduce the contour of the foreground region in the input image before reduction. If jagged edges occur in the contour of the foreground region in the enlarged estimated image, jagged edges also occur in the contour of the foreground region in the foreground mask image (first foreground mask image) generated by the mask generation unit 109, resulting in degradation of the image quality of the virtual viewpoint image. Therefore, the image processing device 100 according to this embodiment is configured to generate an edge image from the input image and correct the enlarged estimated image using the generated edge image. The correction processing of the enlarged estimated image in the image processing device 100 according to this embodiment will be described below with reference to FIGS. 3 to 6.
[0033] FIG. 3 is a flowchart showing an example of a processing flow of the image processing device 100 according to the first embodiment. Note that the processing of each step shown in the flowchart shown in FIG. 3 and the flowcharts described below is realized by the CPU 201 executing a computer program stored in a memory such as the ROM 202 or the auxiliary storage device 204. Also, the letter "S" added to the beginning of a reference symbol indicates a step (process). FIG. 4 is a diagram showing an example of an image acquired or generated by the image processing device 100 according to the first embodiment. The operation of the image processing device 100 according to this embodiment will be described below in accordance with the flowchart shown in FIG. 3, using the image shown in FIG. 4.
[0034] The image processing device 100 starts the processing of this flowchart when the operation unit 206 receives an operation to start this image processing from the user. First, in S301, the acquisition unit 101 acquires captured image data as input image data. Fig. 4(a) is a diagram showing an example of an input image 400. The input image 400 includes an image 403 corresponding to a captured object in the background and an image 401 of an object in the foreground, and the image 401 includes an image 402 corresponding to the pattern of clothing worn by the object.
[0035] Next, in S302, reduction unit 102 generates a reduced image by reducing input image 400 acquired in S301 by an arbitrary magnification. Next, in S303, estimation unit 103 generates a foreground estimated image by estimating, from the reduced image generated in S302, an area that covers an area corresponding to the image of the foreground object, including several pixels outside that area, as a foreground area. Hereinafter, the number of pixels in the foreground estimated image generated in S303 is assumed to be the same as the number of pixels in the reduced image generated in S302, but the number of pixels in the foreground estimated image is not limited to this.
[0036] Next, in S304, enlargement unit 104 generates an enlarged estimated image by enlarging the foreground estimated image by a magnification corresponding to the reciprocal of the reduction magnification in reduction unit 102. Fig. 4(b) is a diagram showing an example of enlarged estimated image 410. In enlarged estimated image 410, region 411 is an enlarged region of the foreground region estimated in S303, and is an area whose size has expanded by several pixels compared to image 401 of the foreground object shown in Fig. 4(a).
[0037] Next, in S305, the extraction unit 105 generates an edge image (first edge image) by extracting the contour of the image 401 of the foreground object from the input image 400 acquired in S301. FIG. 4(c) is a diagram showing an example of the first edge image 420. In the first edge image 420, the contour 421 is the contour of the image 401 of the foreground object, and the contour 423 is the contour of the image 403 corresponding to the captured object in the background. Similarly, the contour 422 is the contour of the image 402 corresponding to the pattern of the clothing worn by the object.
[0038] Next, in S306, the removal unit 106 generates a second edge image by removing a contour 422 existing inside a contour 421 of the image 401 of the foreground object from the first edge image 420 generated in S305. This is because, when subtracting the edge image from the enlarged estimated image 410 in a subsequent processing step, if other contours such as the contour 422 remain inside the contour 421, a hole or gap will appear in the foreground region of the enlarged estimated image after subtraction. FIG. 4(e) is a diagram showing an example of a second edge image 440. In the second edge image 440, the contour 422 in the first edge image has been removed.
[0039] Specifically, for example, the removal unit 106 removes the contour 422 by subtracting an image obtained by contracting the enlarged estimated image 410 from the first edge image 420 generated in S305. Fig. 4(d) is a diagram showing an example of an image 430 obtained by contracting the enlarged estimated image 410. The image 430 is an image obtained by contracting the area 411 in the enlarged estimated image 410 to an area 431 that fits inside the image 401 of the foreground object shown in Fig. 4(a). The removal unit 106 generates a second edge image 440 shown in Fig. 4(e) by subtracting each pixel value of the image shown in Fig. 4(d) from the value of the corresponding pixel in the first edge image 420 shown in Fig. 4(c).
[0040] Next, in S307, the contour correction unit 107 uses the enlarged estimated image 410 generated in S304 to correct the contours 421 and 423 included in the second edge image 440 generated in S306 so as to emphasize the contours from the inside to the outside of the contours. This correction process generates a third edge image. The foreground estimated image estimated and generated by the estimation unit 103 has a characteristic in that pixel values in the image of the foreground object decrease with a gradation toward the background along the contour of the image. The contour correction unit 107 utilizes this characteristic to correct the second edge image 440 by increasing the pixel values outside the contours 421 and 423 in the second edge image 440, which have a width of several pixels. This is because by more strongly emphasizing the outside of the contours 421 and 423, foreground dust that occurs during correction of the enlarged estimated image in a subsequent processing step can be minimized.
[0041] Fig. 4(f) is a diagram showing an example of a third edge image 450. The third edge image 450 is generated by multiplying the value of each pixel outside the contours 421 and 423 in the second edge image 440 shown in Fig. 4(e) by a gain corresponding to the pixel value of the enlarged estimated image 410 shown in Fig. 4(b). Contour 451 is the contour corresponding to contour 421, and contour 453 is the contour corresponding to contour 423. The above-mentioned gain adjustment method and correction processing of the second edge image 440 will be described with reference to Figs. 5 and 6.
[0042] FIG. 5 is a diagram illustrating an example of the correction process performed by the contour correction unit 107 according to the first embodiment and a method for adjusting a gain used in the correction process. Specifically, FIG. 5 illustrates an example of a gain adjustment method for when the gradient of pixel values near the boundary between the foreground and background regions of the enlarged estimated image is steep. FIG. 5(a) illustrates an example of an enlarged estimated image 500. A boundary 501 is the boundary between the foreground and background regions in the enlarged estimated image 500, and is a boundary with a gradation and a width of approximately one or two pixels. FIG. 5(b) illustrates an example of a second edge image 510 corresponding to the enlarged estimated image 500 shown in FIG. 5(a). In the second edge image 510, a contour 511 corresponds to the contour of the foreground region in the enlarged estimated image 500, and similarly, a contour 512 corresponds to the contour of the background region.
[0043] 5(c) is a graph showing an example of the relationship between pixel values and gains in the enlarged estimated image 500. When the gradient of pixel values near the boundary is steep, the center in the width direction of the wide line representing the contour 511 of the foreground region shown in FIG. 5(b) is located inside the boundary 501 shown in FIG. 5(a), i.e., on the foreground region side. The pixel values inside the boundary 501 in the enlarged estimated image 500 are likely to be included in the foreground region as estimated by the estimation unit 103, and are therefore often the maximum value (255 when pixel values are expressed in 8 bits). Therefore, the contour correction unit 107 adjusts the gain so that the gain increases as the pixel value decreases, using as a reference a gain obtained by multiplying the value of the pixel corresponding to the center in the width direction of the wide line representing the contour 511 by the gain corresponding to the maximum estimated value.
[0044] Fig. 5(d) is a diagram showing an example of a third edge image 530 corresponding to the second edge image 510 shown in Fig. 5(b). In the third edge image 530, a contour 531 corresponds to the contour 511 shown in Fig. 5(b), and has been corrected so that the contour is emphasized from the center to the outside of a line having a width indicating the contour. Therefore, the contour 531 is slightly expanded on the outside compared to the contour 511. The series of correction processes described above by the contour corrector 107 is performed, for example, by calculation using the following equation (1). H(x,y)=G(F(x,y))×E(x,y)...Equation (1)
[0045] Here, x and y represent the coordinates of corresponding pixels in the second edge image, the third edge image, and the enlarged estimated image. H(x, y) is a pixel value of the third edge image, F(x, y) is a pixel value of the enlarged estimated image, E(x, y) is a pixel value of the second edge image, and G() is a function that takes the pixel value of the enlarged estimated image as an independent variable and returns the gain corresponding to the pixel value as a dependent variable. Figure 5(e) will be described later.
[0046] FIG. 6 is a diagram illustrating another example of the correction process performed by the contour correction unit 107 according to the first embodiment and a gain adjustment method used in the correction process. Specifically, FIG. 6 illustrates an example of a gain adjustment method for when the gradient of pixel values near the boundary between the foreground and background regions of the enlarged estimated image is gentle. FIG. 6(a) illustrates an example of an enlarged estimated image 600. A boundary 601 is the boundary between the foreground and background regions in the enlarged estimated image 600, and is a boundary with a gradation and a width of approximately 5 or 6 pixels. FIG. 6(b) illustrates an example of a second edge image 610 corresponding to the enlarged estimated image 600 shown in FIG. 6(a). In the second edge image 610, a contour 611 corresponds to the contour of the foreground region in the enlarged estimated image 600, and a contour 612 corresponds to the contour existing inside the foreground region.
[0047] FIG. 6(c) is a graph showing an example of the relationship between pixel values and gains in the enlarged estimated image 600. When the gradient of pixel values near the boundary is gentle, the center of the width of the line with a certain width representing the contour 611 of the foreground region shown in FIG. 6(b) is located within the gradation representing the boundary 601 shown in FIG. 6(a). A pixel value 621 represents the value of a pixel located at the center of the line representing the contour 611. The contour correction unit 107 uses the gain corresponding to the pixel value 621 as a reference for the gain to be multiplied by the value of the pixel located at the center of the line representing the contour 611, and adjusts the gain so that the smaller the pixel value, the higher the gain, and the lower the gain, so that the larger the pixel value, the lower the gain. With this adjustment, by reducing the gain to be multiplied by the value of the pixel corresponding to the contour located inside the foreground region, a contour 612 located inside the region surrounded by the contour 611 in the second edge image 610 can be removed. Therefore, in this case, the contour removal process in the removal unit 106 can be omitted, and the contour correction unit 107 only needs to perform the above-described correction process on the first edge image generated by the extraction unit 105.
[0048] FIG. 6(d) is a diagram illustrating an example of a third edge image 630 corresponding to the second edge image 610 shown in FIG. 6(b). In the third edge image 630, a region 631 represents a region in which a contour 612 existing inside the region surrounded by the contour 611 in the second edge image 610 shown in FIG. 6(b) has been removed by gain adjustment by the contour correction unit 107. Note that in this embodiment, the degree of gradient (steepness or gentleness) of pixel values near the boundary between the foreground region and the background region in the enlarged estimated image is determined in advance by the estimation unit 103. Specifically, for example, the degree of gradient may be determined based on the accuracy or performance of the foreground region estimation by the trained model, or may be determined by intentionally learning the degree of gradient when training the learning model. FIG. 6(e) will be described later.
[0049] After S307, in S308, the image corrector 108 subtracts the value of each pixel in the third edge image generated in S307 from the value of the corresponding pixel in the enlarged estimated image generated in S304. In this way, the image corrector 108 corrects the enlarged estimated image so as to reproduce the contour of the foreground region in the input image, thereby generating a corrected image that is the enlarged estimated image after correction.
[0050] FIG. 4(g) is a diagram showing an example of a corrected image 460 generated by the image correction unit 108. The corrected image 460 is generated by subtracting the third edge image 450 shown in FIG. 4(f) from the enlarged estimated image 410 shown in FIG. 4(b). Region 461 represents a foreground region that reproduces the contours of the foreground region in the input image. Region 462 represents an expanded portion of the foreground resulting from the subtraction process described above; hereinafter, this expanded portion will be referred to as "foreground dust."
[0051] FIGS. 5(e) and 6(e) are diagrams illustrating examples of corrected images 540 and 640 obtained by performing simulations using actual captured images. Specifically, the corrected image 540 shown in FIG. 5(e) was obtained by subtracting the third edge image 530 shown in FIG. 5(d) from the enlarged estimated image 500 shown in FIG. 5(a). Region 541 is a region of foreground dust resulting from the subtraction process. The corrected image 640 shown in FIG. 6(e) was obtained by subtracting the third edge image 630 shown in FIG. 6(d) from the enlarged estimated image 600 shown in FIG. 6(a). Region 641 is a region of foreground dust resulting from the subtraction process. The contour correction process in S307 makes it possible to suppress, to a certain extent, foreground dust that appears in the corrected image after the correction process of the enlarged estimated image in S308. Remaining foreground dust, whose pixel values have decreased or whose shapes have become smaller due to the foreground dust suppression effect of the correction process in S307, is removed in subsequent processing steps.
[0052] After S308, in S309, mask generation unit 109 binarizes the corrected image generated in S308 using a predetermined threshold value, thereby generating a foreground mask image (first foreground mask image). Fig. 4(h) is a diagram showing an example of first foreground mask image 470. In first foreground mask image 470, area 471 outside the black line surrounding foreground area 472 represents foreground dust remaining after binarization processing by mask generation unit 109.
[0053] Next, in S310, the expansion / contraction unit 110 performs an erosion process and an expansion process on the first foreground mask image generated in S309, thereby generating a final foreground mask image (second foreground mask image). FIG. 4(i) is a diagram showing an example of a second foreground mask image 480. A white area 481 in the second foreground mask image 480 is the foreground area. The expansion / contraction unit 110 first performs an erosion process on the first foreground mask image 470 shown in FIG. 4(h) to remove foreground dust indicated by area 471, and then performs an expansion process to restore the size of the mask of the foreground area to its original size. After S310, the image processing device 100 ends the processing of the flowchart shown in FIG. 3.
[0054] In this embodiment, image processing device 100 is configured to generate a foreground estimated image from a reduced image obtained by reducing an input image, and then enlarge the generated foreground estimated image by correcting the contours of the enlarged foreground estimated image (enlarged estimated image) using contours extracted from the input image. Image processing device 100 configured in this manner can reduce the amount of calculation required for the process of extracting a foreground region from an input image, obtain a highly accurate foreground region, and generate a foreground mask whose contour portion is faithful to the input image. As a result, degradation of image quality in the virtual viewpoint image can be suppressed.
[0055] [Embodiment 2] In the first embodiment, a method for correcting and improving the contour of the foreground area in an enlarged estimated image was described, by subtracting an edge image (third edge image) generated based on an input image from the enlarged estimated image in which a foreground area slightly larger than the actual foreground area is estimated. In the second embodiment, a method for specifying a foreground area with high accuracy corresponding to the actual foreground area by adding an edge image of the input image to the estimated enlarged estimated image will be described, which is a method different from the improvement method in the first embodiment.
[0056] <Functional configuration of image processing device> The functional configuration of an image processing device 100 according to the second embodiment will be described with reference to Fig. 1. The functional configuration of the image processing device 100 according to the second embodiment is the same as the functional configuration of the image processing device 100 according to the first embodiment, but processing in some of the functional configuration differs from that of the first embodiment. Therefore, in the following, a description of the same matters as those in the first embodiment will be omitted, and only the differences will be described.
[0057] The estimation unit 103 estimates an area (foreground area) corresponding to the object in the reduced image and generates a foreground estimated image showing the foreground area. Specifically, for example, the estimation unit 103 inputs data of the reduced image into a trained model obtained as a result of learning such as machine learning, and acquires data of the foreground estimated image output as an estimation result by the trained model, thereby generating the foreground estimated image. The trained model described above according to this embodiment is assumed to have been trained to estimate a foreground area of the foreground estimated image that exactly overlaps with the foreground area in the input reduced image.
[0058] The removal unit 106 generates a second edge image by using the enlarged estimated image to remove unnecessary contours from the contours included in the first edge image generated by the extraction unit 105. In this embodiment, contours present in an area corresponding to an area outside the area (foreground area) corresponding to the object in the enlarged estimated image are removed from the contours included in the first edge image. The second edge image generated by the removal unit 106 is used in correction processing of the foreground area in the enlarged estimated image in subsequent processing. Details of the second edge image will be described later using FIG. 8.
[0059] The contour correction unit 107 corrects the contours by enhancing the contours included in the second edge image, thereby generating a third edge image. In this embodiment, when correcting the contours included in the second edge image, the contours are corrected by multiplying each pixel value of the second edge image by an externally set gain without using pixel values of the enlarged estimated image. The image correction unit 108 adds the value of each pixel in the enlarged estimated image to the value of the corresponding pixel in the third edge image, thereby generating an image (corrected image) in which the contour portion of the foreground region in the enlarged estimated image has been corrected. Details of the corrected image will be described later using FIG. 8.
[0060] Dilation / contraction unit 110 generates a second foreground mask image by performing an erosion process on the first foreground mask image. Lines having a width corresponding to the contour of the foreground region, calculated by applying a Sobel filter, exist inside and outside the foreground region with respect to the boundary between the foreground region and the background region. Due to the influence of these contour lines existing outside, the foreground region of the enlarged estimated image (corrected image) after correction generated by image correction unit 108 expands by approximately one pixel relative to the original foreground region. Therefore, in mask generation unit 109, a foreground mask image (first foreground mask image) generated from this expanded corrected image also expands similarly to the corrected image according to the first embodiment. Therefore, dilation / contraction unit 110 performs an erosion process on the first foreground mask image to remove the expanded portion, thereby restoring the foreground region of the foreground mask image to the same size as the original foreground region.
[0061] <Correction process for enlarged estimated image> The correction processing of the enlarged estimated image according to the second embodiment will be described with reference to FIGS. 7 and 8. FIG. 7 is a flowchart showing an example of the processing flow of the image processing device 100 according to the second embodiment. Note that in FIG. 7, the same processing steps as those of the image processing device 100 according to the first embodiment are denoted by the same reference numerals as those shown in FIG. 3, and description thereof will be omitted. FIG. 8 is a diagram showing an example of an image acquired or generated by the image processing device 100 according to the second embodiment. The operation of the image processing device 100 according to this embodiment will be described below in accordance with the flowchart shown in FIG. 7, using the image shown in FIG. 8. First, the image processing device 100 executes the processing of S301 and S302. After S302, in S703, the estimation unit 103 generates a foreground estimated image by estimating a foreground region from the reduced image generated in S302 so that the foreground region exactly overlaps with the image of the foreground object. After S703, the image processing device 100 executes the processing of S304 and S305.
[0062] After S305, in S706, the removal unit 106 generates a second edge image by removing contours that exist outside the image of the object that is the foreground from the first edge image generated in S305. This is because if contours remain in the area outside the image of the object, i.e., the area that is the background, when the edge image is added to the enlarged estimated image in a later processing step, the contours that correspond to the background will be added as a foreground area in the enlarged estimated image after addition.
[0063] FIG. 8(a) shows an example of an input image 800. The input image 800 is similar to the input image 400 shown in FIG. 4(a). In the input image 800, the images 801 to 803 are similar to the images 401 to 403 shown in FIG. 4(a), respectively. FIG. 8(b) shows an example of an enlarged estimated image 810, and FIG. 8(c) shows an example of a first edge image 820. In the enlarged estimated image 810, an area 811 is an area obtained by enlarging the foreground area estimated in S703, and is an area that exactly overlaps with the image 401 of the foreground object shown in FIG. 4(a). In the first edge image 820, the contours 821 to 823 are similar to the contours 421 to 423 shown in FIG. 4(c), respectively. The removal unit 106 removes contours that exist in areas of the first edge image 820 that correspond to the background area of the enlarged estimated image 810 by referring to the pixel values of the enlarged estimated image 810. Specifically, for example, the removal unit 106 removes the contour using the following equation (2). TIFF2025125660000002.tif13150
[0064] Here, x and y represent pixel coordinates. H(x, y) represents the pixel value of the edge image (second edge image) after contour removal, E(x, y) represents the pixel value of the edge image (first edge image) before contour removal, and F(x, y) represents the pixel value of the enlarged estimated image. α represents a pixel value threshold for distinguishing between background and foreground regions in the enlarged estimated image. If the value of a pixel in the enlarged estimated image is equal to or less than the threshold α, the removal unit 106 determines that the pixel in the second edge image corresponding to the pixel is included in the background region and replaces the value of the pixel in the first edge image with 0. This generates a second edge image in which the contours present in the region of the first edge image corresponding to the background in the enlarged estimated image have been removed. FIG. 8D shows an example of a second edge image 830. In the second edge image 830, the contour 823 in the first edge image 820 shown in FIG. 8C has been removed.
[0065] Next, in S707, the contour correction unit 107 corrects the second edge image 830 by multiplying each pixel value of the second edge image 830 generated in S706 by an externally set gain. This correction process generates a third edge image in which the contours included in the second edge image 830 are emphasized. Next, in S708, the image correction unit 108 adds the value of each pixel in the third edge image generated in S707 to the value of the corresponding pixel in the enlarged estimated image generated in S304. In this way, the image correction unit 108 corrects jaggies that have occurred in the contour portion of the foreground region in the enlarged estimated image by the enlargement process so as to reproduce the contours of the image of the object in the input image, thereby generating a corrected image that is an enlarged estimated image after correction.
[0066] Fig. 8(e) shows an example of a corrected image 840 generated by the image correcting unit 108. The corrected image 840 is obtained by adding an image (third edge image) obtained after the second edge image shown in Fig. 8(d) has been corrected by the contour correcting unit 107 to the enlarged estimated image 810 shown in Fig. 8(b). In the corrected image 840, contours 841 and 842 are corrected so that contours 821 and 822 shown in Fig. 8(d) are emphasized, in that order.
[0067] After S708, the image processing device 100 executes the process of S309. That is, after S708, in S309, the mask generation unit 109 generates a first mask image by binarizing the corrected image 840 generated in S708. After S309, in S710, the expansion / contraction unit 110 performs an erosion process on the first foreground mask image generated in S309 to generate a final foreground mask image (second foreground mask image). FIG. 8(f) shows an example of the second foreground mask image 850. A region 851 shown in white in the second foreground mask image 850 is the foreground region. After S710, the image processing device 100 ends the process of the flowchart shown in FIG. 7.
[0068] In this embodiment, image processing device 100 is configured to generate a foreground estimated image from a reduced image obtained by reducing an input image, and then enlarge the generated foreground estimated image by correcting the contours of the enlarged foreground estimated image (enlarged estimated image) using contours extracted from the input image. Image processing device 100 configured in this manner can reduce the amount of calculation required for the process of extracting a foreground region from an input image, obtain a highly accurate foreground region, and generate a foreground mask whose contour portion is faithful to the input image. As a result, degradation of image quality in the virtual viewpoint image can be suppressed.
[0069] [Other embodiments] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0070] It should be noted that within the scope of the present disclosure, the embodiments may be freely combined, any component of each embodiment may be modified, or any component of each embodiment may be omitted.
[0071] [Configuration of the present disclosure] The present disclosure includes the following configurations, methods, and programs.
[0072] <Configuration 1> an acquisition means for acquiring input image data; a reduction means for reducing the input image to obtain a reduced image; an estimation means for estimating a foreground region in the reduced image to obtain a foreground estimation image; an enlargement means for enlarging the estimated foreground image to obtain an enlarged estimated image having the same resolution as the input image; an extraction means for extracting a contour from the input image to obtain a first edge image; a limiting means for obtaining a second edge image by limiting the contours included in the first edge image to contours necessary for correcting the enlarged estimated image; an image correcting means for correcting the enlarged estimated image using the second edge image to obtain a corrected image; a generation means for generating, from the corrected image, a mask image indicating a foreground region corresponding to an image of an object that is in the foreground in the input image; 1. An image processing device comprising:
[0073] <Configuration 2> a contour correction means for correcting the second edge image using pixel values of the enlarged estimated image to obtain a third edge image; the image correction means corrects the enlarged estimated image by using the third edge image instead of the second edge image to obtain the corrected image; 2. The image processing device according to configuration 1,
[0074] <Configuration 3> the limiting means acquires the second edge image by removing, from the first edge image, a contour that exists in a region corresponding to a foreground region in the enlarged estimated image; 3. The image processing device according to configuration 2,
[0075] <Configuration 4> the image correction means obtains the corrected image by subtracting the third edge image from the enlarged estimated image; 4. The image processing device according to configuration 2 or 3, characterized in that:
[0076] <Configuration 5> Further comprising an expansion / contraction means for acquiring a corrected mask image by performing an expansion / contraction process on the pixels of the mask image. 5. The image processing device according to any one of configurations 1 to 4, characterized in that:
[0077] <Configuration 6> the expansion / contraction means performs a contraction process on the pixels of the mask image and then performs a contraction process on the pixels of the mask image; 6. The image processing device according to configuration 5,
[0078] <Configuration 7> the estimation means estimates, as the foreground region, an area that is slightly larger than an image of the foreground object in the reduced image; 7. The image processing device according to any one of configurations 1 to 6,
[0079] <Configuration 8> the estimation means estimates, as a foreground region, a region corresponding to an image of an object in the foreground of the reduced image; 2. The image processing device according to configuration 1,
[0080] <Configuration 9> a contour correction means for correcting the second edge image using pixel values of the enlarged estimated image to obtain a third edge image; the image correction means corrects the enlarged estimated image by using the third edge image instead of the second edge image to obtain the corrected image; 9. The image processing device according to configuration 8,
[0081] <Configuration 10> the limiting means acquires the second edge image by removing, from the first edge image, a contour that exists in an area corresponding to an area outside the foreground area in the enlarged estimated image; 10. The image processing device according to configuration 9,
[0082] <Configuration 11> the image correction means obtains the corrected image by subtracting the third edge image from the enlarged estimated image; 11. The image processing device according to configuration 9 or 10,
[0083] <Configuration 12> further comprising a contraction means for performing a contraction process on the pixels of the mask image to obtain a corrected mask image; 12. The image processing device according to any one of configurations 8 to 11,
[0084] <Method> an acquisition step of acquiring input image data; a reduction step of reducing the input image to obtain a reduced image; an estimation step of estimating a foreground region in the reduced image to obtain a foreground estimation image; an enlargement step of enlarging the foreground estimated image to obtain an enlarged estimated image with the same resolution as the input image; an extraction step of extracting a contour from the input image to obtain a first edge image; a limiting step of acquiring a second edge image by limiting the contours included in the first edge image to contours necessary for correcting the enlarged estimated image; an image correction step of correcting the enlarged estimated image using the second edge image to obtain a corrected image; a generation step of generating, from the corrected image, a mask image indicating a foreground region corresponding to an image of an object that is in the foreground in the input image; An image processing method comprising:
[0085] <Program> 13. A program for causing a computer to function as the image processing device according to any one of configurations 1 to 12. [Explanation of symbols]
[0086] 100: Image processing device 101: Acquisition Department 102: Reduction section 103: Estimation part 104: Enlarged section 105: Extraction part 106:Removal section 108: Image correction unit 109: Mask generation unit
Claims
1. an acquisition means for acquiring input image data; a reduction means for reducing the input image to obtain a reduced image; an estimation means for estimating a foreground region in the reduced image to obtain a foreground estimation image; an enlargement means for enlarging the estimated foreground image to obtain an enlarged estimated image having the same resolution as the input image; an extraction means for extracting a contour from the input image to obtain a first edge image; a limiting means for obtaining a second edge image by limiting the contours included in the first edge image to contours necessary for correcting the enlarged estimated image; an image correcting means for correcting the enlarged estimated image using the second edge image to obtain a corrected image; a generation means for generating, from the corrected image, a mask image indicating a foreground region corresponding to an image of an object that is in the foreground in the input image; 1. An image processing device comprising:
2. a contour correction unit for correcting the second edge image using pixel values of the enlarged estimated image to obtain a third edge image; the image correcting means corrects the enlarged estimated image using the third edge image instead of the second edge image to obtain the corrected image; 2. The image processing device according to claim 1, wherein:
3. the limiting means acquires the second edge image by removing, from the first edge image, a contour that exists in a region corresponding to a foreground region in the enlarged estimated image; 3. The image processing device according to claim 2, wherein:
4. the image correction means obtains the corrected image by subtracting the third edge image from the enlarged estimated image; 3. The image processing device according to claim 2, wherein:
5. Further comprising an expansion / contraction means for acquiring a corrected mask image by performing an expansion / contraction process on the pixels of the mask image.
2. The image processing device according to claim 1, wherein:
6. the expansion / contraction means performs a contraction process on the pixels of the mask image and then performs a contraction process on the pixels of the mask image; 6. The image processing device according to claim 5,
7. the estimation means estimates, as the foreground region, an area that is slightly larger than an image of the foreground object in the reduced image; 2. The image processing device according to claim 1, wherein:
8. the estimation means estimates, as a foreground region, a region corresponding to an image of an object in the foreground of the reduced image; 2. The image processing device according to claim 1, wherein:
9. a contour correction unit for correcting the second edge image using pixel values of the enlarged estimated image to obtain a third edge image; the image correcting means corrects the enlarged estimated image using the third edge image instead of the second edge image to obtain the corrected image; The image processing device according to claim 8 ,
10. the limiting means acquires the second edge image by removing, from the first edge image, a contour that exists in an area corresponding to an area outside the foreground area in the enlarged estimated image; The image processing device according to claim 9 ,
11. the image correction means obtains the corrected image by subtracting the third edge image from the enlarged estimated image; The image processing device according to claim 9 ,
12. further comprising a contraction means for performing a contraction process on the pixels of the mask image to obtain a corrected mask image; The image processing device according to claim 8 ,
13. an acquisition step of acquiring input image data; a reduction step of reducing the input image to obtain a reduced image; an estimation step of estimating a foreground region in the reduced image to obtain a foreground estimation image; an enlargement step of enlarging the foreground estimated image to obtain an enlarged estimated image with the same resolution as the input image; an extraction step of extracting a contour from the input image to obtain a first edge image; a limiting step of acquiring a second edge image by limiting the contours included in the first edge image to contours necessary for correcting the enlarged estimated image; an image correction step of correcting the enlarged estimated image using the second edge image to obtain a corrected image; a generation step of generating, from the corrected image, a mask image indicating a foreground region corresponding to an image of an object that is in the foreground in the input image; An image processing method comprising:
14. A program for causing a computer to function as the image processing device according to any one of claims 1 to 12.
Citation Information
Patent Citations
Image processing program and image processing device
JP2007272878A