Image processing device, image processing method and program
The image processing device enhances foreground extraction by adjusting background hues, addressing the challenge of similar foreground and background colors, thereby simplifying the process and improving accuracy.
Patent Information
- Application Number
- JP2020197430
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-11-27
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2040-11-27
AI Technical Summary
Existing image processing techniques for extracting foreground regions face challenges when the foreground and background colors are similar, requiring additional equipment like infrared light detection, which complicates the imaging process.
An image processing device that adjusts pixel values in a captured image by shifting the hue of background colors to enhance the difference between foreground and background, allowing accurate extraction using background subtraction.
Enables simple and appropriate extraction of foreground regions, reducing the need for additional equipment and improving accuracy in foreground detection.
Smart Images

Figure 0007746006000005 
Figure 0007746006000006 
Figure 0007746006000007
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to an image processing technology for extracting a foreground region from a captured image. [Background technology]
[0002] Techniques for extracting foreground regions (image regions corresponding to objects of interest, such as people) from captured images are used for various purposes, and there are a variety of techniques for doing so. A typical technique is background subtraction. Background subtraction compares an input captured image with its background image (an image that does not include the object of interest) acquired by some method, and extracts pixels as foreground regions where the difference in pixel values between corresponding pixels is equal to or greater than a predetermined threshold. This background subtraction technique has the drawback that, when the colors of the foreground object and the background are similar, the difference in pixel values becomes small, making it difficult to accurately extract the foreground region. In this regard, Patent Document 1 discloses a technique for reliably extracting the foreground region even when the foreground and background are the same or similar in color, by using a camera capable of detecting invisible infrared light and a lighting fixture that emits infrared light in addition to a visible light camera. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-021397 Summary of the Invention [Problem to be solved by the invention]
[0004] The technology of Patent Document 1 requires a separate device for detecting and irradiating invisible light such as infrared light in addition to the visible light camera, which increases the effort required for imaging and requires large-scale equipment.
[0005] An object of the present disclosure is to easily and appropriately extract a foreground region from a captured image. [Means for solving the problem]
[0006] The image processing device according to the present disclosure includes an acquisition unit for acquiring a captured image including a foreground object and a background image not including the object, and an acquisition unit for acquiring a captured image including a foreground object and a background image not including the object, the Pixels in a particular image region based on minimum and maximum component values according to a given color space For example, the color indicated by the pixel value is The range between the minimum value and the maximum value The image processing device is characterized by having an adjustment means for adjusting pixel values so that the pixel values are outside the range of the object, and a generation means for generating an image showing the area of the object based on the difference between the adjusted background image and the captured image. [Effects of the Invention]
[0007] According to the technology of the present disclosure, it is possible to extract a foreground region simply and appropriately. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 2A is a diagram showing a schematic configuration of an image processing system, and FIG. 2B is a diagram showing a hardware configuration of an image processing device. [Figure 2] FIG. 1 is a block diagram showing the main functional configuration of an image processing system. [Figure 3] FIG. 3 is a diagram showing the internal configuration of a background generation unit according to the first embodiment. [Figure 4] 5 is a flowchart showing the flow of a series of processes for generating a foreground silhouette image from an input image according to the first embodiment. [Figure 5] FIG. 1A is a diagram showing an example of an input image, and FIG. 1B is a diagram showing an example of a background image. [Figure 6] FIG. 1A is a diagram showing an example of a foreground silhouette image obtained by a conventional method, and FIG. 1B is a diagram showing an example of a foreground silhouette image obtained by the method of the first embodiment. [Figure 7] 10A and 10B are diagrams showing examples of histograms of the difference between an input image and a background image. [Figure 8] FIG. 10 is a diagram showing the internal configuration of a background generation unit according to the second embodiment. [Figure 9] 10A and 10B are diagrams illustrating the setting of a correction area. [Figure 10] FIG. 4 is a diagram illustrating the setting of correction values. [Figure 11] 10 is a flowchart showing the flow of a series of processes for generating a foreground silhouette image from an input image according to the second embodiment. [Figure 12] FIG. 10 is a diagram showing an example of a background image in which only a specific image region within the background image has been corrected. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, the present invention will be described in detail based on preferred embodiments thereof with reference to the drawings. Note that the configurations shown in the following embodiments are merely examples and are not limited to the configurations shown in the drawings.
[0010] [Embodiment 1] In this embodiment, an application example will be described in which a foreground region is extracted from a captured image by a background subtraction method, and a foreground silhouette image required for generating a virtual viewpoint image is generated.
[0011] First, a brief overview of virtual viewpoint images will be provided. There is a technology that generates a virtual viewpoint image from any virtual viewpoint using multiple viewpoint images captured from multiple viewpoints. For example, using virtual viewpoint images, it is possible to view highlight scenes of a soccer or basketball game from various angles, which gives the user a more realistic feeling than with normal images.
[0012] When generating a virtual viewpoint image, a foreground portion representing the shape of an object (subject) is separated from a background portion, modeled, and then rendered. Modeling the foreground requires information about the shape (silhouette) of the object as viewed from multiple image capture devices and information about the foreground texture (e.g., R, G, and B color information for each pixel in the foreground portion). The process of separating the foreground portion from the background portion is called "foreground-background separation processing." This foreground-background separation processing typically uses a "background subtraction method," which calculates the difference between a captured image containing the foreground and its background image, and defines a region consisting of a collection of pixels whose difference value is determined to be equal to or greater than a predetermined threshold as the foreground region. In this embodiment, the accuracy of extracting the foreground region is improved by correcting the background image used in the background subtraction method so that the difference from the captured image containing the foreground increases.
[0013] In this embodiment, the extraction of a foreground region for generating a foreground silhouette image in a system for generating a virtual viewpoint image will be described as an example, but the application of the foreground extraction method disclosed in this embodiment is not limited to the generation of a foreground silhouette image. For example, the foreground extraction method of this embodiment is also effective for detecting moving objects for use in danger prediction, etc., in surveillance imaging devices installed in various facilities, or in surveillance imaging devices installed in remote locations or outdoors.
[0014] <System configuration> FIG. 1(a) is a diagram illustrating the schematic configuration of an image processing system 100 according to this embodiment. A sports event, such as soccer, is taking place in a stadium 101, and a foreground person 102 is present within the stadium 101. The foreground object may be a specific person, such as a player, a coach, or a referee, or an object with a predetermined image pattern, such as a ball or a goal. The foreground object may be either a moving or a stationary object. Multiple camera image processing devices 103 are arranged around the stadium 101, enabling synchronized image capture of a soccer game or other event taking place in the stadium 101 from multiple viewpoints. Each of the multiple camera image processing devices 103 has an image capturing function and an image processing function. The camera image processing devices 103 are connected to each other via a ring-shaped network, for example, using a network cable 104, and are configured to sequentially transmit image data to adjacent camera image processing devices 103 via the network. In other words, each camera image processing device 103 is configured to transmit the received image data and image data obtained by capturing and processing the image data itself to the adjacent camera image processing device 103. The image data processed in each camera image processing device 103 is finally sent to the integrated image processing device 105. The integrated image processing device 105 uses the received image data to generate a virtual viewpoint image. Note that the system configuration shown in Fig. 1(a) is just an example, and is not limited to a ring-type network connection, and other connection topologies such as a star type may also be used.
[0015] <Hardware configuration> FIG. 1(b) is a block diagram showing the basic hardware configuration common to the camera image processing device 103 and the integrated image processing device 105. The image processing device 103 / 105 includes a CPU 11, a ROM 12, a RAM 13, an auxiliary storage device 14, a display unit 15, an operation unit 16, a communication I / F 17, and a bus 18. The CPU 11 controls the entire device using computer programs and data stored in the ROM 12 and RAM 13, thereby realizing each function of the image processing device 103 / 105. Note that one or more dedicated hardware components separate from the CPU 11 may be included, and at least some of the processing performed by the CPU 11 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (application-specific integrated circuit), an FPGA (field-programmable gate array), and a DSP (digital signal processor). The ROM 12 stores programs that do not require modification. The RAM 13 temporarily stores programs and data supplied from the auxiliary storage device 14 and data supplied from an external device via the communication I / F 17. The auxiliary storage device 14 is composed of, for example, a hard disk drive or the like, and stores various data such as image data and audio data. The display unit 15 is composed of, for example, an LCD display or LEDs, and displays a GUI (Graphical User Interface) for a user to operate the image processing device 105. The operation unit 16 is composed of, for example, a keyboard, mouse, joystick, touch panel, etc., and inputs various instructions to the CPU 11 in response to user operations. The CPU 11 operates as a display control unit that controls the display unit 15 and an operation control unit that controls the operation unit 16. The communication I / F 17 is used for communication with devices external to the image processing device 103 / 105. For example, when connected to an external device via a wired connection, a communication cable is connected to the communication I / F 17. Furthermore, when the communication I / F 17 has the function of wirelessly communicating with an external device, it is equipped with an antenna. The bus 18 connects various components within the device to transmit information. In this embodiment, the display unit 15 and the operation unit 16 are described as being present inside the device, but at least one of the display unit 15 and the operation unit 16 may be present as a separate device outside the device.Furthermore, the display unit 15 and the operation unit 16 are not essential components of the camera image processing device 103, and the system may be configured to be remotely operable from, for example, an external controller (not shown).
[0016] <Functional configuration> 2 is a block diagram showing the main functional configuration of the image processing system 100 of this embodiment. The camera image processing device 103 comprises an imaging control unit 201, an image pre-processing unit 202, a background generation unit 203, a foreground silhouette generation unit 204, and a foreground texture generation unit 205. The integrated image processing device 105 comprises a shape model generation unit 211 and a virtual viewpoint image generation unit 212. Below, each unit constituting the camera image processing device 103 and the integrated image processing device 105 will be explained in order.
[0017] First, the internal configuration of the camera image processing device 103 will be described. The imaging control unit 201 controls an optical system (not shown) to capture moving images at a predetermined frame rate (e.g., 60 fps). The image sensor built into the imaging control unit 201 has an image sensor that detects the visible light range. Alternatively, the image sensor may allocate one pixel of a color filter to the near-infrared range and simultaneously acquire a near-infrared image in addition to a normal color image (RGB image). In this case, the foreground / background separation process (described below) can be performed using four channels: three RGB channels and one IR (near-infrared) channel, enabling foreground extraction with higher accuracy. The captured image obtained by the imaging control unit 201 is input to the image pre-processing unit 202. The image pre-processing unit 202 performs pre-processing, such as development, distortion correction, and vibration correction, on the input captured image. The pre-processed captured image is input frame by frame to the background generation unit 203, foreground silhouette generation unit 204, and foreground texture generation unit 205, respectively. The pre-processed captured image input to each of these units frame by frame will be referred to as the "input image" below. The background generation unit 203 generates a background image based on the input image from the image pre-processing unit 202. At this time, a predetermined correction process is performed on the generated background image, the details of which will be described later. The foreground silhouette generation unit 204 performs foreground / background separation processing using a background subtraction method, using the input image from the image pre-processing unit 202 and the background image from the background generation unit 203, to generate a foreground silhouette image. The foreground silhouette image is also called a "foreground mask." Data on the generated foreground silhouette image is input to the foreground texture generation unit 205. The foreground texture generation unit 205 extracts color information of the portion corresponding to the foreground silhouette from the input image from the image pre-processing unit 202 to generate a foreground texture. Image data 210 (hereinafter referred to as "camera image data"), which includes the foreground silhouette image and foreground texture obtained as described above, is then sequentially transmitted from all camera image processing devices 103 and collected by the integrated image processing device 105.
[0018] Next, the internal configuration of the integrated image processing device 105 will be described. The shape model generation unit 211 generates shape data (hereinafter referred to as a "shape model") representing the three-dimensional shape of an object based on the foreground silhouette image included in the received camera image data 210 corresponding to each camera image processing unit 103. The virtual viewpoint image generation unit 212 applies a foreground texture to the shape model according to information on the virtual viewpoint position and orientation, colors it, and synthesizes it with the background image to generate a virtual viewpoint image representing the view from the virtual viewpoint.
[0019] The above-mentioned functional units are implemented inside an ASIC or FPGA. Alternatively, the CPU 11 may read a program stored in a ROM 12 or the like into a RAM 13 and execute the program, thereby causing the CPU 11 to function as each unit shown in Fig. 2. In other words, the camera image processing device 103 and the integrated image processing device 105 may realize each functional module shown in Fig. 2 as a software module.
[0020] <Details of the background generation section> Next, the processing performed by the background generation unit 203 will be described in detail. Fig. 3 is a functional block diagram showing the internal configuration of the background generation unit 203 according to this embodiment. As shown in Fig. 3, the background generation unit 203 has a background region extraction unit 301, a correction processing unit 302, and a correction parameter setting unit 303. Each functional unit will be described below.
[0021] The background region extraction unit 301 extracts a background region from an input image to generate a base background image. One method for extracting a background region is a method (inter-frame difference method) that extracts a still region as a background region from multiple input images arranged in chronological order. Specifically, the difference between the image information of the nth input frame and the (n-1)th frame immediately preceding it is calculated, and if the difference exceeds a certain value, it is determined to be a moving region, while other regions are determined to be still regions, thereby extracting the background region. This method considers a pixel region where the difference between the two frames is below a threshold as a background region, since the two adjacent frames are captured with a fixed angle of view and only differ slightly in time. The image showing the extracted background region is input to the correction processing unit 302 as a base background image to be subjected to the correction processing described below.
[0022] The correction processor 302 performs correction processing on the background image generated by the background region extraction unit 301 to enable more accurate extraction of the foreground region. In this embodiment, the correction processing involves changing the hue of the colors contained in the background image. Generally, hue is given in the range of 0 degrees to 360 degrees, with 0 degrees and 360 degrees representing red, 60 degrees representing yellow, 120 degrees representing green, 180 degrees representing cyan, 240 degrees representing blue, and 300 degrees representing magenta. In this embodiment, for each color contained in the input background image, a process is performed to shift (offset) the hue angle by a predetermined angle (e.g., 10 degrees) in the plus or minus direction. Each pixel in the input image has an RGB value, and each pixel in the background image generated from the input image also has an RGB value. Therefore, in order to correct the hue of the colors contained in the background image, the correction processor 302 first converts the color space of the background image from RGB to HSV. Specifically, the hue H of each pixel is calculated using the following equations (1) to (3). When min(R,G,B)=B, H=(GR) / S×60+60 Formula (1) When min(R,G,B)=R, H=(BG) / S×60+180 Formula (2) When min(R,G,B)=G, H=(RB) / S×60+300 Formula (3) In the above formulas (1) to (3), S represents saturation, which is calculated by the following formula (4). S=max(R,G,B)-min(R,G,B) Equation (4) Then, for each pixel that makes up the background image, the value of the specified angle is added to (or subtracted from) the hue angle, which is the component value that represents the hue, to calculate a new hue H. Once the hue of each pixel has been corrected in this way, the pixel values are then returned to the RGB color space by performing an inverse calculation (inverse conversion) on the conversion formula. This results in a new background image in which the hues of the colors contained in the background image have been changed.
[0023] The correction parameter setting unit 303 sets an offset value as a correction parameter for the correction process performed by the correction processing unit 302 to apply the offset. In this embodiment, in which the hue of a color included in a background image is changed, the correction parameter setting unit 303 sets the predetermined angle, which defines the amount of hue change, as an offset value and provides it to the correction processing unit 302 as a correction parameter. Assume that an offset value of "+10 degrees" is set, for example. In this case, for yellow-green pixels in the background image, the hue angle is shifted by 10 degrees in the positive direction from the "90 degrees" hue angle, resulting in an increase in blue. If the offset value is "-10 degrees," the hue angle is shifted by 10 degrees in the negative direction from the "90 degrees" hue angle, resulting in an increase in yellow. The offset value is set based on user input; specifically, the user specifies it via a user interface screen (not shown) taking into consideration the color of the main object, etc. Note that if the offset value is too large, there is a greater possibility that a portion that should be background will be erroneously extracted as the foreground region in the foreground region extraction process described below.
[0024] In this embodiment, the specific color in the base background image is corrected by the above-described processing, and a new background image is generated.
[0025] (Processing of background generation unit and foreground silhouette generation unit) Fig. 4 is a flowchart showing the flow of a series of processes for generating a foreground silhouette image from an input image in the camera image processing device 103 of this embodiment. A detailed explanation will be given below with reference to the flowchart in Fig. 4. Note that the symbol "S" indicates a step.
[0026] In S401, the correction processing unit 302 of the background generation unit 203 acquires the correction parameters (offset values in this embodiment) set by the correction parameter setting unit 303. At this time, the correction processing unit 302 acquires the correction parameters (offset values in this embodiment) set by the correction parameter setting unit 303 by reading information on the offset values set and saved in advance from the auxiliary storage device 14 or the like. Alternatively, a UI screen (not shown) for setting correction parameters may be displayed on the display unit 15, and the correction parameter setting unit 303 may first set the value of the predetermined angle input via the UI screen, and then acquire information on the set offset values.
[0027] In S402, the background generation unit 203 and the foreground silhouette generation unit 204 acquire an image (input image) of a frame of interest from the video that has undergone image preprocessing.
[0028] In S403, the background region extraction unit 301 in the background generation unit 203 generates a background image from the input image of the frame of interest using the method described above. Data on the generated background image is input to the correction processing unit 302.
[0029] In S404, the correction processing unit 302 in the background generation unit 201 performs a correction process on the input background image based on the correction parameters acquired in S401. As described above, in this embodiment, the correction process is a process of changing the hues of the colors included in the background image by a preset offset value. Data of the background image obtained by this correction process, in which the hues of the colors of each pixel have been changed (hereinafter referred to as the "corrected background image"), is input to the foreground silhouette generation unit 204.
[0030] In S405, the foreground silhouette generation unit 204 uses the corrected background image generated in S404 to extract the foreground region in the input image of the frame of interest acquired in S401 by background subtraction, thereby generating a foreground silhouette image. The foreground silhouette image is a binary image in which the extracted foreground region is represented by "1" and the other background regions by "0". In extracting the foreground region, the difference (diff) between the input image of the frame of interest and the corrected background image is first calculated. Here, the difference (diff) is expressed by the following equation (1).
[0031]
number
[0032] In the above formula (1), (R in , G in , B in ) represents the pixel value in the input image, and (R bg , G bg , B bg ) represents the pixel value in the corrected background image. R , K. G , K. B represents the weight of the difference between the R, G, and B components.
[0033] Once the difference (diff) between the input image of the frame of interest and the corrected background image is calculated, a binarization process is then performed using a predetermined threshold value TH. This results in a foreground silhouette image in which the foreground region is represented as white (1) and the background region as black (0). The foreground silhouette image may be the same size (same resolution) as the input image, or it may be a partial image obtained by cutting out only the circumscribing rectangle of the extracted foreground region from the input image. The data of the generated foreground silhouette image is input to the foreground texture generation unit 205 and is also sent to the shape model generation unit 211 as part of the camera image data.
[0034] In S406, it is determined whether processing has been completed for all frames constituting the moving image to be processed. If processing has not been completed for all frames, the process returns to S402 to obtain an input image for the next frame of interest and continue processing. On the other hand, if processing has been completed for all frames, the process ends.
[0035] The above is the flow of processing in this embodiment up to generating a foreground silhouette image from an input image. Note that, in this embodiment, the background generation unit 203 and the foreground silhouette generation unit 204 have been described as sequentially acquiring input images of frames of interest from the image pre-processing unit 202 on a frame-by-frame basis, but this is not limiting. For example, the background generation unit 203 and the foreground silhouette generation unit 204 may each acquire data for all frames of the moving image to be processed, and each may perform processing on a frame-by-frame basis in synchronization.
[0036] Here, the foreground silhouette image obtained by the method of this embodiment will be compared with the conventional technique, and the differences and advantages will be explained.
[0037] FIG. 5(a) shows an input image of a frame of interest containing a person object 501, and FIG. 5(b) shows a background image of the frame of interest. In this case, the color of a star-shaped mark 502 on the clothing worn by the person 501 in the foreground is similar to the color of the floor 503 in the background. FIGS. 6(a) and 6(b) show foreground silhouette images obtained based on the input image of FIG. 5(a) and the background image of FIG. 5(b), with FIG. 6(a) corresponding to a conventional method and FIG. 6(b) corresponding to the method of this embodiment. In the foreground silhouette image shown in FIG. 6(a), a silhouette portion 601 of the person 501 is represented by white pixels representing the foreground, and the remaining portion 602 is represented by black pixels representing the background. Meanwhile, a portion 603 corresponding to the star-shaped mark 502 is also represented by black pixels. This is because the color difference (diff) between the star-shaped mark 502, which is the pattern on the clothing, and the floor 503, which is the background, is similar, so the difference (diff) does not exceed the threshold value TH for binarization processing, and the star-shaped mark 502 is determined to be a background region. Fig. 7(a) is a diagram illustrating the binarization process performed when generating the foreground silhouette image of Fig. 6(a) according to the conventional method, and is a histogram with the x-coordinate of the A-A' cross section in Fig. 5(a) on the horizontal axis and the difference diff value on the vertical axis. It can be seen that the difference diff value for the part corresponding to star mark 502 does not exceed the threshold TH.
[0038] In contrast, in the foreground silhouette image shown in FIG. 6(b) according to this embodiment, the entire silhouette 611 of the person 501, including the portion corresponding to the star mark 502, is represented by white pixels, indicating the foreground region, and the remaining portion 612 is represented by black pixels. FIG. 7(b) is a histogram illustrating the binarization process used to generate the foreground silhouette image shown in FIG. 6(b) according to this embodiment. Unlike the histogram of the conventional method shown in FIG. 7(a), the difference diff value of the portion corresponding to the star mark 502 also exceeds the threshold TH. Thus, in the method of this embodiment, the correction process to the background image increases the difference diff between the input image and the background image to a level that exceeds the threshold TH, allowing the portion of the star mark 502, which has a similar color to the background, to be extracted as the foreground region.
[0039] <Modification> In this embodiment, a background image is generated for each frame, but it is not necessary to generate it for each frame. For example, in an image capture scene where the background does not change due to sunlight, such as an indoor sports game, a fixed background image may be used. A fixed background image can be obtained, for example, by capturing an image in a state where no foreground objects are present (for example, before the start of the game).
[0040] Furthermore, in the correction process of this embodiment, the color space of the background image is converted from RGB to HSV and an offset is applied to the hue H, but the content of the correction process is not limited to this. For example, an offset may be applied to each component value (or one or two component values of RGB) while the background image remains in the RGB color space. Alternatively, the background image may be converted to a color space other than HSV, such as YUV, and an offset may be applied to the luminance Y.
[0041] As described above, according to this embodiment, by performing a correction process on a background image, a large difference (diff) can be obtained even if the color of a foreground object in an input image is similar to the color of the background. As a result, it becomes less likely that a part of the foreground region will be erroneously determined to be background in the binarization process for foreground / background separation, and it becomes possible to extract the foreground region appropriately.
[0042] [Embodiment 2] In the method of the first embodiment, the entire background image is uniformly corrected with a predetermined correction value. However, when the method corrects the entire background image with a fixed correction value, the correction may increase the difference value from the input image, resulting in pixels that actually constitute the background being erroneously extracted as a foreground region. Furthermore, for example, in cases where so-called chromakey imaging is performed in a dedicated studio, the color of a green or blue background may be reflected in part of an object. In such cases, it is difficult to prevent the foreground region in the input image where the reflection occurs from being erroneously extracted as a background region using a method that uniformly corrects the entire background image with a fixed correction value. Therefore, as the second embodiment, a mode in which the region to be corrected (the correction region) and the correction value are adaptively determined and correction is performed only on the necessary region in the background image will be described. Note that the description of the basic system configuration and other aspects common to the first embodiment will be omitted or simplified, and the following description will focus on the correction process for the background image, which is the difference between the first embodiment and the second embodiment.
[0043] <Details of the background generation section> Fig. 8 is a functional block diagram showing the internal configuration of the background generation unit 203 according to this embodiment. As shown in Fig. 8, the components of the background generation unit 203 are basically the same as those in the first embodiment, and include a background region extraction unit 301, a correction processing unit 302', and a correction parameter setting unit 303'. A major difference from the first embodiment is that information for specifying a correction region, which is necessary for limited correction of the background image in the correction processing unit 302', is set as a correction parameter. Below, each functional unit will be described, focusing on the differences from the first embodiment.
[0044] The background region extraction unit 301 extracts a background region from an input image and generates a base background image, as in the first embodiment. However, in this embodiment, the generated background image is input to the correction parameter setting unit 303′ in addition to the correction processing unit 302.
[0045] The correction processing unit 302' performs correction processing on a partial image region within the background image generated by the background region extraction unit 301, based on the correction parameters set by the correction parameter setting unit 303'. Note that in this embodiment, an example will be described in which the background image is not subjected to color space conversion, but is instead subjected to correction processing while remaining in the RGB color space.
[0046] The correction parameter setting unit 303' adaptively determines a correction area and a correction value within the background image, and provides them as correction parameters to the correction processing unit 302'. Here, the setting of the correction area and the setting of the correction value will be explained separately.
[0047] <<Correction area settings>> The correction area is set using, for example, spatial information. Here, spatial information refers to, for example, a mask image or ROI (Region of Interest). In the case of a mask image, the correction area in the background image is represented by a region of white (or black) pixels. In the case of an ROI, the correction area in the background image is represented by elements such as the coordinate position (x, y) and the width (w) and height (h) of the target image area. Here, a method for determining the correction area when representing the correction area using a mask image will be explained using the input image shown in FIG. 5(a) and the background image shown in FIG. 5(b) as an example. The area requiring correction processing in the background image shown in FIG. 5(b) is an area that has a similar hue to the foreground (here, the person object 501) and is therefore at risk of being erroneously extracted. Therefore, first, the difference (diff) between the input image shown in FIG. 5(a) and the background image shown in FIG. 5(b) is calculated using the aforementioned equation (1). FIG. 9(a) is a histogram of the calculated difference (diff), where the vertical axis represents the number of pixels and the horizontal axis represents the difference value. The histogram in Fig. 9(a) shows two thresholds (TH low and T.H. high These two thresholds are set by the user via a UI screen (not shown) or the like, focusing on the difference between the foreground and background, for example, when a person in the foreground is wearing clothes of a similar color to the background. Based on the two thresholds set by the user, the difference value is changed from 0 to the threshold TH lowThe range 901 from TH to TH is an area with a high probability of being the background. low From TH high The range 902 from the point A to the point B is an area where it is unclear whether it is the foreground or the background. high The range 903 exceeding the threshold TH is the area with a high probability of being the foreground. low and T.H. high 9(b) shows the three types of image regions separated by the above equation, represented by three types of pixels: white, gray, and black. In FIG. 9(b), the white pixel region 911 corresponds to the foreground region 903, the two gray pixel regions 912, consisting of the edge of the person object 501 and the star mark 502, correspond to the ambiguous region 902, and the black pixel region 913 corresponds to the background region 901. Finally, the two gray pixel regions 912 obtained by threshold processing using the above two thresholds are determined as correction regions, and a mask image (hereinafter referred to as a "correction region mask") is generated in which the determined correction regions are indicated by white pixels and the remaining regions are indicated by black pixels. Using this correction region mask enables limited correction processing to be performed on only image regions of the background image whose color is similar to that of the foreground object.
[0048] <<Setting the correction value>> The correction value is calculated by dividing the above two thresholds TH low and T.H. high In this embodiment, the correction process is performed in the RGB color space, and the offset values of each RGB component (R n ,G n ,B n ) is calculated using, for example, the following formula (2) or formula (3) and the weights W of each of RGB.
[0049]
number
[0050]
number
[0051] In this case, the weight W is, for example, equally distributed for each component (R n :G n :B n ) = (1:1:1). Now, if the above two thresholds are TH low =50,TH high = 30 and the above formula (3) is applied. In this case, the range of the image area that is likely to be the foreground is the range of the difference diff value from "40" to "50", and by the above formula (3),
[0052]
number
[0053] The value of is "10". And, since the weight W is equal for each component, the offset value (R n ,G n ,B n )=(3,3,3).
[0054] Here, we will confirm the meaning of determining the offset value by the above formula (2) or formula (3). Fig. 10 is a histogram of the difference diff between the input image of a certain frame f and its background image, and like the histogram in Fig. 9(a), the vertical axis represents the number of pixels and the horizontal axis represents the difference value. As mentioned above, the difference diff is determined by two thresholds (TH high and T.H. low The image area corresponding to the range 1000 between the two points 1000 and 1000 is an image area (intermediate area) where it is unclear whether it is foreground or background. This intermediate area includes both image areas that should be judged as background and image areas that should be judged as foreground. Of these, the image area that should be judged as foreground is the area in the range 1000 that is smaller than the threshold TH high Therefore, for example, by using the above formula (2), an offset value is determined that enables the image area corresponding to the upper half range 1001 of the above sandwiched range 1000 to be extracted as the foreground.
[0055] (Processing of background generation unit and foreground silhouette generation unit) Fig. 11 is a flowchart showing the flow of a series of processes for generating a foreground silhouette image from an input image in the camera image processing device 103 of this embodiment. A detailed explanation will be given below with reference to the flowchart in Fig. 11. Note that the symbol "S" indicates a step.
[0056] In S1101, the correction parameter setting unit 303′ of the background generation unit 203 sets two preset thresholds (TH low and T.H. high ) is used to determine the offset value as a correction parameter using the above-mentioned method. low and T.H. high ) may be read from the auxiliary storage device 14 or the like. Information on the determined offset value is held in the RAM 13. 4 in the first embodiment, the background generation unit 203 and the foreground silhouette generation unit 204 acquire an input image of the target frame to be processed. In this embodiment, the input image data acquired by the background generation unit 203 is sent to the background region extraction unit 301 as well as to the correction parameter setting unit 303′.
[0057] In S1103, the background region extraction unit 301 in the background generation unit 203 generates a background image using the input image of the frame of interest. Data of the generated background image is input to the correction processing unit 302′.
[0058] In S1104, the correction parameter setting unit 303' generates the above-mentioned correction area mask using the background image generated in S1103 and the input image of the frame of interest.
[0059] In S1105, the correction processing unit 302' determines a pixel of interest in the background image generated in S1103, and determines whether the position of the pixel of interest is within the mask area (white pixel area) indicated by the correction area mask generated in S1104. If the result of the determination shows that the position of the pixel of interest is within the mask area, the process proceeds to S1106. On the other hand, if the position of the pixel of interest is outside the mask area, the process skips S1106 and proceeds to S1107.
[0060] In S1106, the correction processing unit 302' corrects the pixel value of the pixel of interest in the background image using the offset value determined in S1101. In this embodiment, in which the correction processing is performed in the RGB color space, the offset value (R n ,G n ,B n For example, if the offset value is (R n ,G n ,B n )=(3,3,3) and the pixel value of the pixel of interest is (R,G,B)=(100,100,50). In this case, the pixel value of the pixel of interest in the corrected background image will be (R,G,B)=(103,103,53). Note that since it is only necessary to apply an offset, subtraction processing may be performed instead of addition processing.
[0061] In S1107, the correction processing unit 302' determines whether the determination process of S1105 has been completed for all pixels in the background image generated in S1103. If there are any unprocessed pixels, the process returns to S1105 to determine the next pixel of interest and continue processing. On the other hand, if the process has been completed for all pixels in the background image, the process proceeds to S1108.
[0062] In S1108, similar to S405 in the flow of FIG. 4 in the first embodiment, the foreground silhouette generation unit 204 extracts the foreground region in the input image of the frame of interest using the background image corrected by the above-mentioned correction process, and generates a foreground silhouette image. Figure 12 shows a background image corresponding to the above-mentioned FIG. 9(b) in which correction process has been performed on only a specific correction region. Comparing the background image before correction shown in FIG. 5(b) reveals that a region 1201 corresponding to the gray region 912 in FIG. 9(b) (i.e., the edge of the person object 501 and the portion of the star-shaped mark 502) has changed due to the correction process.
[0063] In S1109, similar to S406 in the flow of Fig. 4 in the first embodiment, it is determined whether or not processing has been completed for all frames constituting the moving image to be processed. If processing has not been completed for all frames, the process returns to S1102 to determine the next frame of interest and continue processing. On the other hand, if processing has been completed for all frames, the process ends.
[0064] The above is the flow of processing in this embodiment for generating a foreground silhouette image from an input image. Through the above processing, correction processing is performed only on pixel values of pixels that belong to a specific image region of the background image.
[0065] <Modification> In the above explanation, spatial information is used to identify the correction area, but color space information may also be used. For example, the minimum and maximum values of the hue H in the HSV color space may be specified, or the minimum and maximum values of each RGB component value in the RGB color space may be specified. This makes it possible to identify an image area made up of pixels with a specified range of hue or RGB values as the correction area. For example, when capturing images of a rugby or soccer match, this method is effective in cases where the color of the players' uniforms is similar to the color of the grass in the background. When using the HSV color space, the lower limit of the hue H may be set to H min = 100 degrees, upper limit H max= 140 degrees. By using color space information in this way, it is possible to set as the correction area only the image area of the background image where it is unclear whether it is foreground or background.
[0066] As a correction method, instead of applying an offset, each pixel included in the image area of the specified range may be filled with a specific color, or may be matched to the color of the pixels in the surrounding area.
[0067] As described above, in this embodiment, the foreground extraction process is performed based on a background image in which only image areas that satisfy certain conditions have been corrected, which makes it possible to more effectively prevent erroneous extractions, such as erroneously extracting an actual background area as a foreground area, or conversely, erroneously treating an actual foreground area as background.
[0068] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]
[0069] 103 Camera image processing device 203 Background generation section 302 Correction processing unit 303 Correction parameter setting section
Claims
1. an acquisition means for acquiring a captured image including a foreground object and a background image not including the object; an adjustment means for adjusting pixel values of pixels in a specific image area included in the background image and based on minimum and maximum values of component values according to a predetermined color space, so that the color indicated by the pixel value falls outside the range between the minimum and maximum values; a generating means for generating an image showing a region of the object based on a difference between the adjusted background image and the captured image; 1. An image processing device comprising:
2. The adjustment is a process of multiplying component values according to the predetermined color space by an offset.
2. The image processing device according to claim 1, wherein:
3. The adjusting means is converting a first color space in the background image into the predetermined color space different from the first color space; performing a process of multiplying the offset by the component value according to the predetermined color space; 3. The image processing device according to claim 2.
4. the first color space is RGB; the predetermined color space is HSV, the adjustment means performs processing of adding or subtracting an offset value to a component value representing a hue among component values of pixels constituting the background image whose color space has been converted to HSV.
4. The image processing device according to claim 3.
5. the first color space is RGB; the predetermined color space is YUV, The adjustment means adjusts the component values of the pixels that make up the background image, the color space of which has been converted to YUV. Adding or subtracting an offset value to or from the component value representing luminance.
4. The image processing device according to claim 3.
6. An image processing device as described in any one of claims 1 to 5, characterized in that the minimum value and the maximum value are specified by a user.
7. 7. The image processing apparatus according to claim 1, wherein the adjustment means performs the adjustment based on a mask image that expresses the specific image region and other image regions in binary.
8. 2. The image processing apparatus according to claim 1, wherein the component values according to the predetermined color space are component values representing hues in an HSV color space or RGB values in an RGB color space.
9. The adjustment is a process of replacing the pixel values of the pixels belonging to the specific image region with pixel values representing a different color.
2. The image processing device according to claim 1, wherein:
10. 2. The image processing device according to claim 1, wherein the adjustment is a process of adjusting the pixel values of the pixels belonging to the specific image region to the pixel values of pixels belonging to a peripheral region of the specific image region.
11. 11. The image processing apparatus according to claim 1, wherein the object is a target for generating three-dimensional shape data.
12. an acquisition step of acquiring a captured image including a foreground object and a background image not including the object; an adjustment step of adjusting pixel values of pixels in a specific image area included in the background image based on minimum and maximum values of component values according to a predetermined color space, so that the color indicated by the pixel value falls outside the range between the minimum and maximum values; a generating step of generating an image showing a region of the object based on a difference between the adjusted background image and the captured image; An image processing method comprising:
13. an acquisition means for acquiring a captured image including a foreground object and a background image not including the object; an adjustment means for adjusting pixel values of pixels in a specific image area included in the background image and based on minimum and maximum values of component values according to a predetermined color space, so that the color indicated by the pixel value falls outside the range between the minimum and maximum values; a generating means for generating an image showing a region of the object based on a difference between the adjusted background image and the captured image; An image processing system comprising:
14. A program for causing a computer to function as the image processing device according to any one of claims 1 to 11.
Citation Information
Patent Citations
Device and method for extracting silhouette, and device and method for generating three-dimensional shape data
JP2007017364A
Image processing apparatus, image processing method, program, and storage medium
JP2013257843A
Object detection device
JP2016194778A
Image processing device, image processing method, and image processing program
JP2020021397A
Video image correction device and program
JP2020135653A