Program, model generation method, image processing method, image processing device, and image processing system
A learning model-based program accurately removes backgrounds from images by classifying pixels, addressing limitations of traditional methods and enabling flexible foreground extraction and composition.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DAI NIPPON PRINTING CO LTD
- Filing Date
- 2022-02-14
- Publication Date
- 2026-04-14
AI Technical Summary
Existing background removal techniques, such as chroma key processing and background difference methods, are limited by the need for a green screen and struggle with accurately extracting foreground areas when subjects have similar colors to the background or when shooting conditions vary.
A program that uses a learning model to process captured images and background images, training the model to output a background-removed image by classifying each pixel as background or foreground, enabling accurate background removal without a green screen.
Accurately removes the background from captured images, allowing for precise extraction of foreground areas even when subjects have similar colors or shooting conditions change, and supports flexible composition with arbitrary backgrounds.
Smart Images

Figure 0007844910000001 
Figure 0007844910000002 
Figure 0007844910000003
Abstract
Description
Technical Field
[0001] This application relates to a program, a model generation method, an image processing method, an image processing apparatus, and an image processing system.
Background Art
[0002] Image composition is performed in which an area that becomes the foreground is extracted from an image captured by a camera and synthesized with another image. As a technique for removing the background area from an image, for example, in Patent Document 1, chroma key processing for extracting the captured area of a subject from a captured image by capturing the subject against a green background is disclosed. Also, background removal processing by a background difference method for extracting a subject that does not exist in a previously acquired image by calculating the difference between the previously acquired image and a newly acquired image is also performed.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When performing background removal by chroma key processing, it is necessary to install a green background, so background removal by chroma key processing cannot be performed in places where a green background cannot be installed. Also, in background removal by chroma key processing, there is a problem that when the subject that becomes the foreground has the same color as the green background, the area of the same color is removed as the background. In background removal by the background difference method, even an area that has the same color as the background is removed as the background even if it is a foreground area. Also, in background removal by the background difference method, since areas with different brightness (luminance) are also extracted as difference areas, it is difficult to accurately extract the foreground area when shooting conditions such as brightness are different.
[0005] This disclosure is made in view of the circumstances described herein, and its purpose is to provide a program, etc., that can accurately remove the background area from captured images without using a green screen or the like. [Means for solving the problem]
[0006] A program according to one aspect of the present invention acquires a background image by taking a picture of the background, acquires a captured image including the background and the subject, and when the background image taken of the background and the captured image including the background and the subject are input to a learning model that has been trained to output a background-removed image in which the background region has been removed from the captured image, the program inputs the acquired background image and the captured image to a learning model that has been trained to output a background-removed image in which the background region of the input background image has been removed from the input captured image, causing the computer to execute the process of acquiring a background-removed image in which the background region of the input background image has been removed from the input captured image. [Effects of the Invention]
[0007] In one aspect of the present invention, the background region can be accurately removed from a captured image without using a green screen or the like. [Brief explanation of the drawing]
[0008] [Figure 1] This is an explanatory diagram showing an example configuration of an image processing system. [Figure 2] This is an explanatory diagram of image processing using an image processing device. [Figure 3] This is a block diagram showing an example of the configuration of an image processing device. [Figure 4] This is an explanatory diagram showing an example of a learning model configuration. [Figure 5] This is an explanatory diagram showing an example of an image set for training data. [Figure 6] This is a flowchart illustrating an example of the process for generating a learning model. [Figure 7] This flowchart shows an example of the procedure for providing a composite image. [Figure 8] This flowchart shows an example of the procedure for providing a composite image. [Figure 9] It is an explanatory diagram of background removal processing. [Figure 10] It is an explanatory diagram showing an example of a screen. [Figure 11] It is an explanatory diagram showing an example of a screen. [Figure 12] It is an explanatory diagram showing an example of a screen. [Figure 13] It is a flowchart showing an example of the procedure for providing a composite image in Embodiment 2. [Figure 14] It is an explanatory diagram showing an example of a screen. [Figure 15] It is an explanatory diagram of background removal processing in Embodiment 2. [Figure 16] It is a flowchart showing an example of the procedure for providing a composite image in Embodiment 3. [Figure 17] It is an explanatory diagram showing an example of a screen. [Figure 18] It is a flowchart showing another example of the procedure for providing a composite image in Embodiment 3. [Figure 19] It is a flowchart showing an example of the procedure for providing a composite image in Embodiment 4. [Figure 20] It is an explanatory diagram showing an example of a screen. [Figure 21] It is an explanatory diagram showing a configuration example of an image processing system in Embodiment 5.
Mode for Carrying Out the Invention
[0009] Hereinafter, the program, model generation method, image processing method, image processing apparatus, and image processing system of the present disclosure will be described in detail based on the drawings showing their embodiments.
[0010] (Embodiment 1) FIG. 1 is an explanatory diagram showing a configuration example of an image processing system. In the present embodiment, an image processing system that extracts a subject shooting area from a shooting image in which a subject to be a foreground is included in the background using a background image obtained by previously shooting the background, and synthesizes the extracted subject shooting area into another image will be described. The image processing system of the present embodiment includes an image processing apparatus 10 and a camera 20 (shooting apparatus), and the image processing apparatus 10 and the camera 20 are communicatively connected via a network N. The network N may be the Internet or a LAN (Local Area Network) constructed within the facility where the image processing system is provided. Further, the image processing apparatus 10 and the camera 20 may be configured to directly transmit and receive information by wired communication or wireless communication via a cable.
[0011] The camera 20 is an imaging device having an imaging unit including a lens and an imaging element, etc., for example, a communication unit for connecting to the network N, and a shooting button for instructing shooting timing. The camera 20 performs shooting by the imaging unit according to a shooting instruction received via the shooting button to acquire image data, and performs a process of transmitting the acquired image data from the communication unit (transmission unit) to the image processing apparatus 10. The camera 20 is configured to perform a shooting process for acquiring one piece of image data (still image) corresponding to one shooting instruction and a shooting process for acquiring, for example, 30 or 15 pieces of image data (moving image) per second, and the shooting button has a button for instructing execution of still image shooting and a button for instructing start and end of moving image shooting. Note that the camera 20 of the present embodiment may be configured to perform only the shooting process of still images. Further, the camera 20 may be a smartphone, a tablet terminal, etc. having a camera, or may be built in the image processing apparatus 10.
[0012] The image processing device 10 is an information processing device capable of various information processing and information transmission and reception, such as a personal computer or a server computer. The image processing device 10 acquires image data captured by the camera 20 and performs predetermined image processing on the acquired image data. Figure 2 is an explanatory diagram of image processing by the image processing device 10. In the image processing system of this embodiment, the camera 20 takes pictures with its shooting position fixed, for example, using a tripod. However, it is sufficient that the background in each image taken at different times is the same scenery. For example, if the scenery captured as the background is the same when taking pictures, by using a zoom (magnification / reduction) function, the shooting position of the camera 20 does not need to be fixed.
[0013] Figure 2A is a captured image taken by a camera 20 with a fixed shooting position, and this image becomes the background image, capturing the background. Figure 2B is a captured image taken with a subject within the shooting range of the camera 20, and the background in this captured image is the same scenery as the background in the background image. The image processing device 10 of this embodiment uses the learning model 12M (see Figure 4) to generate a mask image, as shown in Figure 2C, from the background image shown in Figure 2A and the captured image shown in Figure 2B, in which the background area is represented in black and the foreground area (the area of the subject) is represented in white. Furthermore, the image processing device 10 of this embodiment generates a subject image, as shown in Figure 2D, from the original captured image shown in Figure 2B, based on the original captured image shown in Figure 2B and the mask image shown in Figure 2C, in which only the foreground area (the area of the subject) is extracted. By performing the above processing, the image processing device 10 of this embodiment realizes background removal processing, which removes the background area from the captured image (Figure 2B).
[0014] Figure 3 is a block diagram showing an example configuration of the image processing device 10. The image processing device 10 includes a control unit 11, a storage unit 12, a communication unit 13, an input unit 14, a display unit 15, etc., and these units are connected via a bus. The control unit 11 includes one or more processors such as a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), a GPU (Graphics Processing Unit), or an AI chip (AI semiconductor). The control unit 11 executes the information processing and control processing that the image processing device 10 should perform by appropriately executing the program 12P stored in the storage unit 12.
[0015] The memory unit 12 includes RAM (Random Access Memory), flash memory, hard disk, SSD (Solid State Drive), etc. The memory unit 12 stores the program 12P (program product) executed by the control unit 11 and various data. The memory unit 12 also temporarily stores data generated when the control unit 11 executes the program 12P. The program 12P and various data may be written to the memory unit 12 during the manufacturing stage of the image processing device 10, or the control unit 11 may download them from other devices via the communication unit 13 and store them in the memory unit 12. The memory unit 12 also stores, for example, a trained model 12M that has been trained on training data by machine learning. The trained model 12M is a trained model that has been trained to output a mask image as shown in Figure 2C when a background image as shown in Figure 2A and a captured image as shown in Figure 2B are input. The trained model 12M is intended to be used as a program module that constitutes artificial intelligence software. The learning model 12M performs a predetermined operation on the input value and outputs the operation result. The memory unit 12 stores data such as the coefficients and thresholds of the function that defines this operation as the learning model 12M.
[0016] The memory unit 12 also stores the training DB 12a, the background image for background removal 12b, and the background DB 12c for synthesis. The training DB 12a is a database in which training data used for the learning process of the learning model 12M is stored. The background image for background removal 12b is an image used by the image processing device 10 when performing background removal processing, and is a single background image taken by a camera 20 with a fixed shooting position. The background DB 12c for synthesis is a database in which background images for synthesis (synthesis images) are stored to synthesize subject images generated by removing the background region. The background images for synthesis are of multiple types and may be still images or moving images. Some or all of the learning model 12M, the training DB 12a, the background image for background removal 12b, and the background DB 12c may be stored in other storage devices connected to the image processing device 10, or in other storage devices that the image processing device 10 can communicate with.
[0017] The communication unit 13 is a communication module for performing wired or wireless communication processing, and transmits and receives information with other devices via the network N. The input unit 14 receives user input and sends control signals corresponding to the operation to the control unit 11. The display unit 15 is a liquid crystal display or an organic EL display, etc., and displays various information according to instructions from the control unit 11. Part of the input unit 14 and the display unit 15 may be an integrated touch panel, or the touch panel may be externally attached to the image processing device 10.
[0018] In this embodiment, the image processing device 10 may be a multicomputer consisting of multiple computers, a virtual machine virtually constructed by software, or a cloud server. Furthermore, the image processing device 10 does not necessarily have an input unit 14 and a display unit 15, and may be configured to accept operations through a connected computer, or to output the information to be displayed to an external display device. The image processing device 10 may also include a reading unit that reads a non-temporary computer-readable portable storage medium 10a, and may use the reading unit to read the program 12P from the portable storage medium 10a and store it in the storage unit 12. Note that the program 12P may be executed on a single computer, or on multiple computers interconnected via a network N.
[0019] In the image processing apparatus 10 of this embodiment, the control unit 11 reads and executes a program 12P stored in the storage unit 12, and based on a background image previously captured by the camera 20 (background image for background removal 12b) and a newly captured image by the camera 20, it performs a process to remove the background area from the captured image, extract the foreground area (subject area), and generate a subject image. The control unit 11 also performs a process to combine the subject image generated from the captured image with a background image for composition. Therefore, the image processing system of this embodiment can provide the user with a composite image in which the subject image generated from the captured image captured by the camera 20 is combined with an arbitrary background image for composition. In the image processing apparatus 10 of this embodiment, the control unit 11 uses a learning model 12M when extracting the subject area based on the background image and the captured image.
[0020] Figure 4 is an explanatory diagram showing an example configuration of the learning model 12M. The learning model 12M is a model that recognizes predetermined objects contained in an input image. The learning model 12M is a model that can classify objects in an image on a pixel-by-pixel basis, for example, by semantic segmentation. The learning model 12M of this embodiment takes one captured image and one background image as input, and performs calculations to recognize the background region and foreground region (subject not included in the background image) contained in the captured image based on the input captured image and background image, and outputs the recognized result. Specifically, the learning model 12M classifies each pixel of the input captured image into a background region and a foreground region (subject region), and outputs a classified captured image (hereinafter referred to as a labeled image) in which each pixel is associated with a label for each region. The learning model 12M can be composed of, for example, U-Net, FCN (Fully Convolutional Network), SegNet, etc.
[0021] The 12M learning model has an input layer, an intermediate layer, and an output layer. The input layer receives the captured image to be processed and the background image. The intermediate layer has a convolutional layer, a pooling layer, and a deconvolutional layer. The convolutional layer extracts image features from the pixel information of the image input via the input layer and generates a feature map, and the pooling layer compresses the generated feature map. Multiple convolutional and pooling layers are repeated, and the feature maps generated by multiple convolutional and pooling layers are output to the deconvolutional layer. The deconvolutional layer scales (maps) the feature maps generated by the convolutional and pooling layers to the original image size. The deconvolutional layer also identifies which objects are located where in the image on a pixel-by-pixel basis based on the features extracted by the convolutional layer, and generates a label image indicating which object each pixel corresponds to. Figure 4 shows an example where the captured image shown in Figure 2B and the background image shown in Figure 2A are input to the learning model 12M. The label image output by the learning model 12M is an image in which each pixel of the captured image is classified into the background region and the foreground region, and a pixel value corresponding to each region is assigned. In the label image shown in Figure 4, pixels classified as background region are shown in black, and pixels classified as foreground region are shown in white. By outputting such a label image, the learning model 12M is configured to output a background-removed image in which the background region has been removed from the captured image.
[0022] With the configuration described above, when a captured image and a background image are input, the learning model 12M outputs a label image that classifies each pixel in the captured image into areas that are visible in the background image (background area) and areas that are not visible in the background image (foreground area). The image processing device 10 uses the label image output from the learning model 12M as a mask image with the background area masked, and extracts the unmasked foreground area from the captured image to generate a foreground image (subject image), which is a background-removed image with the background area removed. Therefore, the image processing device 10 generates a foreground image from the captured image in which the foreground area not visible in the background image is extracted.
[0023] The learning model 12M can be generated by preparing training data that includes a training background image and a captured image, and a label image in which data indicating the object to be classified (in this case, the background region and the foreground region) is labeled for each pixel in the captured image, and then using this training data to train an untrained learning model. In the training label image, the captured image is labeled with the coordinate range corresponding to the region of each object and the type of each object. In this embodiment, as shown in Figure 4, a background-removed image obtained by removing the background region from the captured image can be used as the training label image.
[0024] Figure 5 is an explanatory diagram showing an example of an image set for training data. The training data used for training the learning model 12M includes a background image common to all training data, a captured image in which a subject in the background is photographed with the subject in the background image as the background, and a label image in which the captured image is masked by the region in the background image. Such background image, captured image, and label image are associated and stored in the training DB 12a as a single training data.
[0025] The learning model 12M is trained to output a label image included in the training data when the background image and captured image included in the training data are input. Specifically, the learning model 12M performs calculations in the hidden layer based on the input background image and captured image, and obtains detection results for each object in the captured image (in this case, the background region and foreground region). More specifically, the learning model 12M obtains a label image as output, in which a value indicating the type of object classified is labeled for each pixel in the captured image. The learning model 12M then compares the obtained detection results (label image) with the coordinate range and object type of the correct object region indicated by the training data, and optimizes parameters such as the weights (connection coefficients) between neurons so that the two approximate each other. The method of parameter optimization is not particularly limited, but methods such as the steepest descent method and backpropagation can be used. As a result, a learning model 12M is obtained that outputs a label image indicating the background region and foreground region in the captured image when the background image and captured image are input.
[0026] The image processing device 10 prepares such a learning model 12M in advance and uses it for background removal processing to remove the background region from the captured image taken by the camera 20 and extract the foreground region. The learning model 12M only needs to be able to identify the position and shape of the background region or foreground region in the captured image. The learning model 12M may be trained by another learning device. The trained learning model 12M generated by training on another learning device is downloaded from the learning device to the image processing device 10, for example, via the network N or the portable storage medium 10a, and stored in the storage unit 12.
[0027] The following describes the process of generating a learning model 12M by learning from the training data described above. Figure 6 is a flowchart showing an example of the procedure for generating the learning model 12M. The following process is performed by the control unit 11 of the image processing device 10 according to the program 12P stored in the storage unit 12, but it may also be performed by other learning devices. In the following process, it is assumed that the training data using the image set described above is stored in the training DB 12a.
[0028] The control unit 11 of the image processing device 10 acquires one training data from the training DB 12a (S11). Specifically, the control unit 11 reads one image set from the training DB 12a as shown in Figure 5. Then, the control unit 11 performs training processing on the learning model 12M using the acquired training data (S12). Here, the control unit 11 inputs the background image and captured image included in the training data into the learning model 12M and acquires the label image output from the learning model 12M based on the input of the background image and captured image. The learning model 12M performs calculations based on the input background image and captured image to calculate the label image to be output from the output layer. The control unit 11 compares each pixel value of the label image output from the learning model 12M with each pixel value of the correct label image included in the training data and trains the learning model 12M so that the two approximate each other. In the learning process, the control unit 11 optimizes parameters such as the weights (connection coefficients) between neurons in the hidden layer using a backpropagation method, which sequentially updates the learning model 12M from the output layer to the input layer.
[0029] The control unit 11 determines whether there is any unprocessed training data stored in the training DB 12a that has not undergone learning processing (S13). If it determines that there is unprocessed training data (S13: YES), the control unit 11 returns to the process of step S11 and performs the processing of steps S11 to S12 on the training data that has not undergone learning processing. If it determines that there is no unprocessed training data (S13: NO), the control unit 11 terminates the series of processes.
[0030] The learning process described above generates a learning model 12M that, when a captured image and a background image are input, outputs a mask image (label image) in which the background region of the captured image is masked. Furthermore, the learning model 12M can be further optimized by repeatedly performing the learning process using the training data described above. In addition, an already trained learning model 12M can be retrained by performing the above process, in which case a learning model 12M with higher discrimination accuracy can be generated.
[0031] The following describes the process in this embodiment of the image processing system for generating a foreground image by performing background removal processing on an image captured using the camera 20, and then combining the generated foreground image with an arbitrary background image for composition and providing it. Figures 7 and 8 are flowcharts showing an example of the procedure for providing a composite image, Figure 9 is an explanatory diagram of the background removal process, and Figures 10 to 12 are explanatory diagrams showing example screens. The following processing is performed by the control unit 11 of the image processing device 10 according to the program 12P stored in the storage unit 12.
[0032] When providing a composite image using the image processing system of this embodiment, the camera 20 is used to photograph the foreground subjects with each subject in a pre-photographed background image (background image 12b for background removal) as the background, and to acquire a captured image. The camera 20 determines whether or not it has received a shooting instruction via the shooting button (S21), and if it determines that it has not received one (S21: NO), it waits. If it determines that it has received a shooting instruction (S21: YES), the camera 20 performs shooting with the imaging unit and acquires a captured image (S22). For example, the camera 20 acquires a captured image as shown in Figure 9A. The camera 20 transmits the acquired captured image to the image processing device 10 (S23).
[0033] The control unit 11 (image acquisition unit) of the image processing device 10 acquires the captured image (a captured image including the background and foreground subjects) transmitted from the camera 20 (S24). The control unit 11 (background image acquisition unit) reads from the storage unit 12 the background image 12b for background removal, which has been previously captured by the camera 20 and stored in the storage unit 12 (S25). Thus, the control unit 11 acquires the captured image as shown in Figure 9A and the background image (background image for background removal 12b) as shown in Figure 9B. Then, the control unit 11 performs a background removal process on the captured image based on the acquired captured image and background image for background removal 12b (S26) and generates a foreground image (background removed image) from which the background region has been removed from the captured image (S27). Specifically, the control unit 11 (background removed image acquisition unit) inputs the captured image and background image for background removal 12b to the learning model 12M and acquires the label image (background removed image) output from the learning model 12M. Here, the control unit 11 acquires a label image as shown in Figure 9C. Then, the control unit 11 (extraction unit) uses the acquired label image as a mask image and extracts the region corresponding to the white area in the mask image from the captured image, thereby extracting the foreground region from the captured image and generating a foreground image.
[0034] The control unit 11 displays the generated foreground image on the display unit 15 (S28). For example, the control unit 11 displays a screen as shown in Figure 10A on the display unit 15 and presents the generated foreground image to the user. The screen shown in Figure 10A is configured to accept the selection of any subject included in the foreground image via the displayed foreground image. The user makes a selection of any subject on the screen shown in Figure 10A by selecting a subject via the input unit 14 and operating the OK button. The control unit 11 determines whether or not the selection of any subject has been accepted (S29). In the example shown in Figure 10A, there are three subjects in the foreground image, and the state in which the selection of the central subject has been accepted is shown. Note that multiple subjects may be selected. Furthermore, when displaying the foreground image, the control unit 11 may be configured to detect each subject in the foreground image by, for example, performing object detection processing on the foreground image, and present each detected subject with a bounding box. In this case, the system can be configured so that the user can select any subject by selecting one of the bounding boxes.
[0035] If the control unit 11 determines that it has not received a selection for a subject (S29: NO), it waits until it receives one. If the control unit 11 determines that it has received a selection for a subject (S29: YES), it extracts the area of the selected subject from the foreground image and generates a subject image (S30). Here, as shown in Figure 9D, the control unit 11 generates a subject image that includes only the selected subject.
[0036] Next, the control unit 11 displays a list of composite background images (S31) to accept the selection of a composite background image to be composited with the generated subject image. Here, the control unit 11 reads a composite background image stored in the composite background DB 12c and displays a screen as shown in Figure 10B on the display unit 15, presenting the user with a list of composite background images. The screen shown in Figure 10B displays still images and thumbnail images of videos (for example, the first image) prepared as composite background images, and is configured to accept the selection of any of the composite background images. The screen shown in Figure 10B also displays the total playback time and an indicator showing the playback position relative to the total playback time for the video composite background image, and the playback position of the video displayed on the screen can be changed by moving the playback position via the indicator. The user selects any of the composite background images on the screen shown in Figure 10B by selecting one of the composite background images via the input unit 14 and operating the OK button. The control unit 11 determines whether or not it has accepted a selection for any of the background images to be composited (S32). If it determines that it has not accepted a selection (S32: NO), it waits until it accepts a selection.
[0037] When the control unit 11 determines that it has received a selection for one of the background images for compositing (S32: YES), it displays a setting screen on the display unit 15 for receiving settings for image processing to be performed on the subject image when compositing the subject image generated in step S30 onto the selected background image for compositing (S33). For example, as shown in Figure 11A, the control unit 11 displays a setting screen that displays the background image for compositing and the subject image, and has input fields for inputting the processing details of the image processing to be performed on the subject image. The image processing to be performed on the subject image includes scaling processing to enlarge or reduce the subject image, and rotation processing to rotate the subject image. Therefore, the setting screen has input fields for scaling ratios for the width and height of the subject image, and input fields for rotation angles. Each input field may be configured to allow input of any numerical value, and a pull-down menu may be provided for selecting any one of several options. Note that the setting screen shown in Figure 11A may also be configured to allow input of the enlargement or reduction ratio for the subject image by, for example, pinch-out and pinch-in operations. Furthermore, the settings screen is configured to allow the user to specify the compositing position (compositing position within the compositing image) where the subject image is composited onto the compositing background image by dragging, as shown in Figure 11B. Note that the image processing that can be performed on the subject image is not limited to scaling and rotation.
[0038] On the screens shown in Figures 11A and 11B, the user inputs the scaling factor for the width and height of the subject image, the rotation angle, and the position for compositing the subject image relative to the background image via the input unit 14, and then operates the composite execution button to instruct the execution of the subject image compositing process onto the background image. The control unit 11 determines whether or not it has received input for the image processing on the setting screens shown in Figures 11A and 11B (S34). If it determines that it has received input (S34: YES), it displays the input processing details (specifically, scaling factor and rotation angle) in each input field. The control unit 11 then executes the image processing of the input processing details on the displayed subject image (S35). Here, if a scaling factor is input, the control unit 11 executes scaling processing (enlargement or reduction processing) at the input scaling factor on the displayed subject image, and if a rotation angle is input, it executes rotation processing at the input rotation angle on the displayed subject image. The control unit 11 then updates the displayed subject image to the subject image after image processing. The screen in Figure 11B displays the subject image after it has been scaled down compared to the subject image in Figure 11A. Image processing on the subject image is not always necessary, and if the control unit 11 determines that it has not received input for the image processing content (S34: NO), it skips the process in step S35.
[0039] Next, the control unit 11 determines whether it has received input for the composition position of the subject image relative to the background image on the setting screen shown in Figures 11A and B (S36). If it determines that it has received input (S36: YES), it moves the subject image to the specified composition position, as shown in Figure 11B (S37). Then, the control unit 11 determines whether the composition execution button has been operated (S38). If it determines that it has not been operated (S38: NO), it returns to the process in step S34 and continues to accept input for the image processing content and composition position. The control unit 11 also returns to the process in step S34 if it determines that it has not received input for the composition position (S36: NO).
[0040] If the control unit 11 determines that the composite execution button has been operated (S38: YES), it performs a composite process to composite the selected subject image onto the selected background image for composite (S39). Specifically, the control unit 11 generates a composite image by compositing the subject image processed in step S35 onto the background image for composite selected in step S32 at the compositing position specified in steps S36 to S37. The control unit 11 outputs the generated composite image (S40) and terminates the series of processes. For example, the control unit 11 displays the generated composite image to the user by displaying a screen like the one shown in Figure 12 on the display unit 15.
[0041] The screen shown in Figure 12 displays the composite image and has a "Redo" button to instruct the generation of the composite image to be redone, a "Print" button to instruct the printing of the composite image, and a "Send" button to instruct the transmission of the composite image. If the "Redo" button is operated on the screen shown in Figure 12, the control unit 11 of the image processing device 10 returns to the process of step S33 in Figure 8, for example, and executes the image processing and compositing process again. The control unit 11 may return to the process of step S31 in Figure 7 if it is to redo the selection of the background image for compositing, or to the process of step S28 in Figure 7 if it is to redo the selection of the subject. If the "Print" button is operated on the screen shown in Figure 12, the control unit 11 of the image processing device 10 sends the composite image to a printer via the network N or directly connected to print the composite image. If the "Send" button is operated, the control unit 11 transmits the composite image to a designated terminal via the network N or by short-range wireless communication, etc. The control unit 11 may print the generated composite image using a printer or transmit it to a predetermined terminal without displaying the screen shown in Figure 12.
[0042] Through the processing described above, the image processing system of this embodiment can achieve background removal processing that accurately extracts the area of the foreground subject from the captured image taken by the camera 20 without using a green screen or the like. Therefore, even in facilities where a green screen cannot be installed, or in places where installing a green screen would spoil the scenery, it is possible to generate a subject image (foreground image) with only the subject area extracted by highly accurate background removal processing based on a captured image of a subject taken with any background. Furthermore, scaling and rotation processing are possible for the subject image, and any background image for compositing can be selected, so the subject image can be freely processed. Moreover, when multiple subjects are captured, the subject to be composited can be selected, so even if, for example, an unintended subject is captured, background removal processing can be performed to leave only the subject selected by the user, and a subject image containing only the desired subject can be generated.
[0043] In this embodiment, a mask image used to remove the background region from a captured image is generated using a learning model 12M. The learning model 12M automatically extracts image features from the input captured image and background image and outputs a mask image. By training it with a large amount of training data, it becomes possible to generate a mask image that accurately classifies the background region and foreground region in the captured image. Furthermore, since the learning model 12M is used to classify the background region and foreground region (subject) in the captured image, even subjects with similar colors to the background can be appropriately classified as foreground regions, enabling appropriate background removal processing. Moreover, if, for example, the shooting environment changes, such as if the brightness of the room where the camera 20 is installed changes, the learning model 12M can be fine-tuned to maintain its processing accuracy. Specifically, by retraining the learning model 12M using a background image (background image 12b for background removal) and a captured image taken in the changed shooting environment, along with a label image (mask image) generated from that captured image, it becomes possible to generate a highly accurate mask image even for captured images taken in the changed shooting environment.
[0044] In this embodiment, the learning process of the learning model 12M using training data, the background removal process using the learning model 12M, or the process of compositing the subject image onto the background image for compositing are not limited to being performed locally by the image processing device 10. For example, a server that performs the learning process of the learning model 12M may be provided. In this case, the image processing device 10 is configured to send training data to the server, and the server generates the learning model 12M through learning processing using the training data and sends it to the image processing device 10. In this case as well, the image processing device 10 can perform background removal processing using the learning model 12M generated by the server. Alternatively, a server that performs background removal processing using the learning model 12M may be provided. In this case, the image processing device 10 is configured to send the target captured image and background image to the server, and a mask image generated by the server using the learning model 12M is sent to the image processing device 10. In this case as well, the image processing device 10 can perform background removal processing to remove the background region from the captured image using the mask image generated by the server. Furthermore, a server that performs the process of compositing the subject image onto the background image for compositing may be provided. In this case, the image processing device 10 is configured to send the subject image to be composited and the background image for composite to the server, and the composite image, in which the subject image is composited with the background image for composite, is sent back to the image processing device 10. In this case as well, the image processing device 10 can provide the composite image generated by the server to the user by printing or sending it. Even with the configuration described above, the same processing as in this embodiment is possible and the same effects can be obtained.
[0045] In this embodiment, a subject image is generated from a captured image and combined with a background image for synthesis to produce a composite image, but the configuration is not limited to this. For example, a configuration in which a subject image is generated by performing a background removal process on a captured image and then provided to the user may also be used.
[0046] (Embodiment 2) In Embodiment 1 described above, the system was configured to accept the selection of a subject to be composited via the obtained subject image (foreground image) after background removal processing was performed on the captured image. In Embodiment 2, an image processing system is described in which the selection of a subject to be composited via the captured image is accepted before background removal processing is performed. The image processing system of this embodiment is implemented using the same apparatus as the image processing system of Embodiment 1 shown in Figures 1 and 3, so a description of the configuration will be omitted.
[0047] Figure 13 is a flowchart showing an example of the composite image provisioning process procedure of Embodiment 2, Figure 14 is an explanatory diagram showing an example screen, and Figure 15 is an explanatory diagram of the background removal process of Embodiment 2. The process shown in Figure 13 is the same as the process shown in Figures 7 and 8, but with steps S51 to S54 added between steps S25 and S26, and step S55 added instead of steps S27 to S29. The same steps as in Figures 7 and 8 are omitted from the explanation. Note that steps S32 to S40 in Figures 7 and 8 are not shown in Figure 13.
[0048] In the image processing system of this embodiment, the camera 20 performs the same processing as in steps S21 to S23 in Figure 7, and the control unit 11 of the image processing device 10 performs the same processing as in steps S24 to S25 in Figure 7. As a result, in this embodiment as well, the image processing device 10 acquires the captured image taken by the camera 20 and reads the background image 12b for background removal from the storage unit 12.
[0049] In the image processing apparatus 10 of this embodiment, after processing in step S25, the control unit 11 displays the captured image acquired from the camera 20 on the display unit 15 (S51). For example, the control unit 11 displays a screen as shown in Figure 14 and presents the captured image taken by the camera 20 to the user. The screen shown in Figure 14 is configured to accept the selection of any subject included in the captured image via the displayed captured image. The user makes a selection of any subject on the screen shown in Figure 14 by selecting a subject via the input unit 14 and operating the OK button. The control unit 11 determines whether or not a selection of any subject has been accepted via the captured image (S52). In the example shown in Figure 14, it shows a state in which a selection of the central subject among three subjects in the captured image has been accepted. Multiple subjects may be selected here as well. Furthermore, when displaying the captured image, the control unit 11 may be configured to detect each subject in the captured image by, for example, performing object detection processing on the captured image, and to present each detected subject with a bounding box.
[0050] If the control unit 11 determines that it has not received a selection for a subject (S52: NO), it waits until it receives one. If the control unit 11 determines that it has received a selection for a subject (S52: YES), it extracts the area of the selected subject from the captured image (S53). Here, the control unit 11 extracts a rectangular area (subject area) containing the selected subject, as shown by the rectangle in Figure 14. If multiple subjects are selected, the control unit 11 may extract a single subject area containing all the selected subjects, or it may extract a subject area for each subject. If a subject area is extracted for each subject, the control unit 11 executes the processes in steps S54, S26, and S55 for each subject area.
[0051] Next, the control unit 11 extracts the region corresponding to the subject region extracted in step S53 (background region for background removal) from the background image 12b for background removal read in step S25 (S54). Here, the control unit 11 extracts the region corresponding to the subject region from the background image 12b for background removal, as shown by the rectangle in Figure 15A. Specifically, when the control unit 11 extracts the subject region from the captured image, it obtains the coordinate values of the upper left and lower right pixels of the subject region, for example, using a coordinate system where the upper left of the captured image is the origin, the right direction is the X-axis, and the downward direction is the Y-axis. As a result, the control unit 11 can extract the region corresponding to the subject region from the background image 12b for background removal based on the obtained coordinate values.
[0052] Then, the control unit 11 performs background removal processing on the subject area based on the subject area extracted from the captured image in step S53 and the area corresponding to the subject area (background area for background removal) extracted from the background removal background image 12b in step S54 (S26). Here, as shown in Figure 15B, the control unit 11 inputs the subject area and the background area for background removal into the learning model 12M and obtains a label image output from the learning model 12M. The label image shown in Figure 15B is a background-removed image obtained by removing subjects that appear in the background area for background removal from the subject area input into the learning model 12M. In this label image, not only the selected subject but also unselected subjects are classified as foreground areas (white areas). Therefore, the control unit 11 changes the areas of subjects other than the selected subject (unselected subjects) into background areas in the label image obtained using the learning model 12M (S55). In the label image shown in Figure 15B, the areas of the subject on the left and the subject on the right are changed to the background area (in this case, the black area), thereby allowing the control unit 11 to obtain the label image shown in Figure 15C.
[0053] The control unit 11 uses the label image generated in step S55 as a mask image to extract the area corresponding to the white area in the mask image from the captured image or the subject area extracted in step S53, thereby generating a subject image that includes only the area of the selected subject (S30). Here too, the control unit 11 can generate a subject image as shown in Figure 9D. After that, the control unit 11 executes the processing from step S31 onwards. As a result, in this embodiment as well, it is possible to select any background image for compositing, and it is possible to specify the image processing content to be performed on the subject image and the compositing position when compositing the subject image onto the background image for compositing. Therefore, it is possible to generate a composite image in which a subject image that has undergone arbitrary scaling and rotation processing is composited at any position on the background image for compositing.
[0054] In this embodiment, the same effects as in Embodiment 1 described above can be obtained. Furthermore, in this embodiment, the selection of a subject to be composited onto an arbitrary background image can be received via the captured image. In this embodiment as well, the modifications described as appropriate in Embodiment 1 described above can be applied.
[0055] (Embodiment 3) In the embodiments 1 and 2 described above, the process of selecting the subject to be composited from the subjects in the foreground image or captured image was performed manually by the user. Embodiment 3 describes an image processing system that automatically identifies the subject to be composited by performing object detection processing on the foreground image or captured image. In this embodiment, the image processing system automatically detects subjects in the foreground image or captured image, and if only one subject is detected, that subject is selected as the subject to be composited. If multiple subjects are detected, the user selects the subject to be composited. The image processing system of this embodiment is implemented using the same apparatus as the image processing system of Embodiment 1 shown in Figures 1 and 3, so a description of the configuration will be omitted.
[0056] Figure 16 is a flowchart showing an example of the composite image provisioning process procedure of Embodiment 3, and Figure 17 is an explanatory diagram showing an example screen. The process shown in Figure 16 is the same as the process shown in Figures 7 and 8, but with steps S61 to S62 added between steps S27 and S28, and step S63 added between steps S28 and S29. The same steps as in Figures 7 and 8 are omitted from the explanation. Note that in Figure 16, steps S21 to S25 and steps S32 to S40 in Figures 7 and 8 are not shown.
[0057] In the image processing system of this embodiment, the camera 20 performs the same processing as in steps S21 to S23 in Figure 7, and the control unit 11 of the image processing device 10 performs the same processing as in steps S24 to S27 in Figure 7. As a result, in this embodiment as well, the image processing device 10 performs background removal processing using the learning model 12M on the captured image taken by the camera 20, and can generate a foreground image (background-removed image) from the captured image by removing the subject that appears in the background removal background image 12b.
[0058] In the image processing apparatus 10 of this embodiment, after the processing in step S27, the control unit 11 performs object detection processing on the generated foreground image (S61). The object detection processing can be performed, for example, using a trained model that has been trained to determine whether the subject in the image is one of the previously trained subjects (objects or people) when an image is input. Therefore, by inputting the foreground image into such a trained model, the control unit 11 can detect the subject in the foreground image based on the output information from the trained model.
[0059] The control unit 11 determines, based on the object detection process, whether or not there is only one subject in the foreground image (S62). If it determines that there is more than one subject (S62: NO), i.e., if multiple subjects are detected, the control unit 11 proceeds to step S28 and displays the foreground image generated in step S27 on the display unit 15 (S28). The control unit 11 also displays bounding boxes surrounding each subject detected by the object detection process in the displayed foreground image (S63). For example, the control unit 11 displays a screen on the display unit 15 as shown in Figure 17, presenting the foreground image and the subjects in the foreground image to the user. This allows the user to easily understand which subjects are selectable, and to easily select any subject by selecting a bounding box. In the foreground image in Figure 17, three subjects are shown by bounding boxes, with solid bounding boxes indicating selected subjects and dashed bounding boxes indicating unselected subjects.
[0060] Subsequently, the control unit 11 executes the processing from step S29 onward. This allows the system to accept the selection of the subject to be synthesized via the screen shown in Figure 17 and generate a subject image by extracting the region of the selected subject.
[0061] If the control unit 11 determines that there is only one subject in the foreground image (S62: YES), it proceeds to the process of step S30, extracting the region of the single subject detected by the object detection process from the foreground image and generating a subject image (S30). Here, a subject image similar to that shown in Figure 9D is generated. After that, the control unit 11 executes the processes from step S31 onwards. In this way, even in this embodiment, it is possible to generate a composite image by combining a subject image that has undergone arbitrary scaling and rotation processing with an arbitrary background image for compositing.
[0062] In the process described above, the configuration may be limited to people as the subject to be synthesized. In this case, the control unit 11 may determine in step S62, based on the results of the object detection process, whether the subject in the foreground image is one person or multiple people. In such a configuration, even if objects other than people are captured in the foreground image, for foreground images containing one person, the detected single person can be identified as the subject to be synthesized.
[0063] In this embodiment, the same effects as in Embodiments 1 and 2 described above can be obtained. Furthermore, in this embodiment, if only one subject is visible in the foreground image generated by background removal processing on the captured image, that subject can be identified (selected) as the subject to be composited. Therefore, it becomes unnecessary for the user to select the subject to be composited, and the processing can be simplified. Also, if multiple subjects are visible in the foreground image, the multiple subjects can be presented to the user, and the user can select the subject to be composited, thereby allowing any subject to be used as the subject to be composited. In this embodiment as well, the modifications described as appropriate in Embodiments 1 and 2 described above can be applied.
[0064] The configuration of this embodiment is applicable to the image processing systems of Embodiments 1 and 2 described above, and the same effects can be obtained even when applied to the image processing systems of Embodiments 1 and 2. Figure 18 is a flowchart showing another example of the composite image provisioning process of Embodiment 3. The process shown in Figure 18 is the process when the configuration of this embodiment is applied to the image processing system of Embodiment 2. The process shown in Figure 18 is the process shown in Figure 13 with steps S61 to S62 added between steps S25 and S51, and step S63 added between steps S51 and S52. The same steps as in Figures 13 and 16 will not be explained. Note that in Figure 18, steps S21 to S24 and steps S31 to S40 in Figure 13 are not shown.
[0065] In the process shown in Figure 18, the control unit 11 of the image processing device 10 executes steps S61 to S62 after the processing in step S25. In step S61, the control unit 11 performs object detection processing on the captured image acquired from the camera 20. If the control unit 11 determines, as a result of the object detection processing, that there is more than one subject in the captured image (S62: NO), it proceeds to step S51. After the processing in step S51, the control unit 11 executes step S63. This makes it possible to display bounding boxes surrounding each subject in the captured image shown on the screen in Figure 14. Subsequently, the control unit 11 executes steps S52 and onward. If the control unit 11 determines, again, that there is only one subject in the captured image (S62: YES), it proceeds to step S53.
[0066] Through the process described above, in the image processing system of Embodiment 2, when only one subject is present in the captured image, that subject can be identified (selected) as the subject to be composited, eliminating the need for the user to select the subject to be composited. Furthermore, when multiple subjects are present in the captured image, the user can select the subjects to be composited, and any subject can be used as the subject to be composited.
[0067] (Embodiment 4) In embodiments 1 to 3 described above, the subject image to be composited onto the background image was a still image. Embodiment 4 describes an image processing system that composites a moving subject image onto a background image. In the image processing system of this embodiment, a video is captured using the camera 20, and background removal processing is performed on each still image included in the video to generate multiple foreground images (subject images) from which the background area has been removed. The subject image of the video is generated by stitching together these multiple subject images. The image processing system of this embodiment is implemented using the same apparatus as the image processing system of Embodiment 1 shown in Figures 1 and 3, so a description of the configuration will be omitted.
[0068] Figure 19 is a flowchart showing an example of the composite image provisioning process procedure of Embodiment 4, and Figure 20 is an explanatory diagram showing an example screen. The process shown in Figure 19 is the same as the process shown in Figures 7 and 8, but with steps S71 to S73 added instead of steps S22 to S24, step S74 added between steps S25 and S26, step S75 added after step S27, steps S76 to S77 added instead of step S28, and step S78 added instead of step S30. The same steps as in Figures 7 and 8 are omitted from the explanation. Note that steps S32 to S40 in Figure 8 are not shown in Figure 19.
[0069] In the image processing system of this embodiment, when shooting a video using the camera 20, the camera instructs the camera to start and stop shooting the video by operating the shooting button. The camera 20 determines whether or not it has received a video shooting instruction via the shooting button (S21), and if it determines that it has received a video shooting instruction (S21:YES), it takes pictures with the imaging unit, for example, 30 or 15 times per second, to acquire a video composed of multiple captured images (still images) (S71). Here, too, a video of the foreground subject is shot with each subject in the pre-shot background image 12b for background removal as the background. The camera 20 transmits the acquired video to the image processing device 10 (S72).
[0070] The control unit 11 acquires the video (captured image including multiple still images) transmitted from the camera 20 (S73). After processing in step S25, the control unit 11 extracts one still image from the video acquired in step S73 (S74). Then, based on the one still image (captured image) and the background removal background image 12b, the control unit 11 performs background removal processing on the one still image (S26) to generate a foreground image (background-removed image) from which the background region has been removed (S27). Specifically, the control unit 11 inputs the one still image and the background removal background image 12b into the learning model 12M and acquires a label image output from the learning model 12M. Here again, the control unit 11 acquires a label image as shown in Figure 9C, and uses the acquired label image as a mask image to extract the foreground region from the still image to be processed and generate a foreground image.
[0071] The control unit 11 determines whether the background removal process has been completed for all still images included in the video acquired in step S73 (S75). If it determines that the process has not been completed (S75: NO), it returns to the process in step S74. The control unit 11 then extracts one unprocessed still image from the video acquired in step S73 (S74), performs the processes in steps S26 to S27 on the extracted still image, and extracts the foreground region from each still image to generate a foreground image.
[0072] If the control unit 11 determines that the background removal process has been completed for all still images (S75: YES), it generates a foreground video by stitching together the foreground images generated from each still image included in the video (S76). The control unit 11 then displays the generated foreground video on the display unit 15 (S77). For example, the control unit 11 presents the foreground video to the user by displaying a screen as shown in Figure 20A. The screen shown in Figure 20A has the same configuration as the screen shown in Figure 10A, and further displays an indicator 15a that shows the playback position relative to the total playback time of the foreground video being displayed. By moving the playback position using the indicator 15a, it is possible to change the playback position of the foreground video displayed on the screen. The screen shown in Figure 20A is also configured to accept selection of any subject via the foreground video being displayed.
[0073] If the control unit 11 determines that it has received a selection for any subject via the foreground video being displayed (S29: YES), it extracts the region of the selected subject from the foreground video and generates a subject image (S78). Specifically, the control unit 11 extracts the region of the selected subject from each foreground image (still image) included in the foreground video to generate a subject image, and then generates a subject video by stitching together the generated subject images.
[0074] Subsequently, the control unit 11 executes the processing from step S31 onward. In this embodiment, in step S33, the control unit 11 displays a settings screen as shown in Figure 20B. The settings screen shown in Figure 20B has the same configuration as the screens shown in Figures 11A and 11B, and further displays an indicator 15b that shows the total playback time and the playback position relative to the total playback time for the subject video currently being displayed, and an indicator 15c that shows the total playback time and the playback position relative to the total playback time for the background image to be composited currently being displayed. Note that in the settings screen shown in Figure 20B, the state when a video is selected as the background image to be composited is shown, so the indicator 15c is displayed for the background image to be composited, but if a still image is selected as the background image to be composited, the indicator 15c is not displayed. Note that even if the subject image to be composited is a video, the background image to be composited may be a still image or a video.
[0075] Furthermore, the settings screen shown in Figure 20B may have a configuration that includes input fields for the total playback time, the number of times to repeat playback, and the playback speed for the subject video to be composited. In this case, not only scaling and rotation processing can be performed on the subject video, but the total playback time, number of playbacks, and playback speed can also be specified, and arbitrary editing processing can be performed on the subject video. Therefore, in step S35, the control unit 11 performs video editing processing based on the input content, in addition to scaling and rotation processing, on the subject video displayed on the settings screen. Moreover, the settings screen shown in Figure 20B may have a configuration that allows specifying an arbitrary playback position of the background image for compositing via the indicator 15c, and specifying the compositing position of the subject video by dragging the image (still image) at the specified playback position, as shown in Figure 11B. In this case, it becomes possible to composite the subject image from an arbitrary playback position of the background image for compositing, which is a video, to an arbitrary compositing position. Therefore, in step S39, the control unit 11 can generate a composite video by compositing the subject video, which has undergone image processing and video editing processing in step S35, from an arbitrary playback position of the composite background image to the compositing position specified for each image (still image). Through the above processing, a composite video is generated in which the subject video is composited with the video of an arbitrary composite background image.
[0076] In this embodiment, the same effects as in embodiments 1 to 3 described above can be obtained. Furthermore, in this embodiment, by performing background removal processing on each still image included in a video of the subject, a video of the subject with the background area removed can be generated. Therefore, since not only still images but also videos of the subject can be used as the composite target, it becomes possible to generate composite images with a higher degree of freedom. In this embodiment as well, the modifications described as appropriate in embodiments 1 to 3 described above can be applied.
[0077] The configuration of this embodiment is applicable to the image processing systems of Embodiments 1 to 3 described above, and the same effects can be obtained even when applied to the image processing systems of Embodiments 1 to 3. Specifically, when the configuration of this embodiment is applied to the image processing system of Embodiment 2, the process of selecting the subject to be composited can be performed before the background removal process is executed for each still image included in the video in which the subject was filmed. Furthermore, when the configuration of this embodiment is applied to the image processing system of Embodiment 3, if only one subject is shown in the video in which the subject was filmed, that subject can be identified as the subject to be composited.
[0078] (Embodiment 5) The system configurations for installing the image processing systems of Embodiments 1 to 4 described above in various facilities will now be explained. Figure 21 is an explanatory diagram showing an example configuration of the image processing system of Embodiment 5. The image processing system of this embodiment is installed and used in places such as amusement parks, theme parks, aquariums, zoos, botanical gardens, exhibitions, and tourist destinations. In addition to the image processing device 10 and camera 20, the image processing system of this embodiment has a touch panel 30 and a printer 40, and each device 10, 20, 30, and 40 may communicate via a network N, or they may communicate directly via wired or wireless communication.
[0079] In this embodiment of the image processing system, the various screens that were displayed on the display unit 15 of the image processing device 10 in embodiments 1 to 4 described above are displayed on the touch panel 30, and the various information that was input via the input unit 14 is input via the touch panel 30. Even with this configuration, the same processing as in the image processing systems of embodiments 1 to 4 described above is possible, and the same effects can be obtained.
[0080] The embodiments disclosed herein should be considered in all respects to be illustrative and not restrictive. The scope of the invention is indicated by the claims, not in the sense described above, and all modifications are intended to be in the sense and scope equivalent to the claims. [Explanation of Symbols]
[0081] 10 Image Processing Device 11 Control Unit 12 Storage section 13 Communications Department 14 Input section 15 Display 20 cameras 12M Learning Model
Claims
1. Obtain a background image by taking a picture of the background, A photographic image including the background and subject is obtained. Select at least one of the aforementioned subjects, From the aforementioned captured image, extract the subject area containing the selected subject. From the aforementioned background image, the background region corresponding to the subject region is extracted. A trained model, which is trained to output a background-removed image from which the background region has been removed when given a background image taken of the background and a captured image including the background and the subject, is given an image of the background region extracted from the background image and an image of the subject region extracted from the captured image, and a background-removed image is obtained from the input subject region image, in which the background region in the input background region image has been removed. For the acquired background-removed image, the areas of subjects other than the selected subject are changed to background areas. A program that instructs a computer to perform a process.
2. The background-removed image is an image obtained by extracting from the captured image the areas that are not visible in the background image. The program according to claim 1.
3. The subject in the captured image is detected, Based on the detected subjects, identify the subject to select. The program according to claim 1 or 2, which causes the computer to perform the processing.
4. Based on the image of the subject area and the background removal image, the shooting area of the subject is extracted from the image of the subject area. A program according to any one of claims 1 to 3 that causes the computer to perform a process.
5. Select one of the multiple images for compositing, The selected image for synthesis is composited with the subject's shooting area, which is extracted from the image of the subject's area. The program according to claim 4, which causes the computer to perform the processing.
6. Multiple images are acquired as described above, For each of the multiple captured images, the image of the background region extracted from the background image and the image of the subject region extracted from the captured image are input to the learning model to obtain a background-removed image from the image of the subject region in which the background region of the background region image has been removed. For each of the aforementioned multiple captured images, the acquired background-removed image is modified by changing the areas of subjects other than the selected subject into background areas. For each of the aforementioned multiple captured images, the captured area of the subject is extracted from the image of the subject area based on the image of the subject area and the background removal image. The shooting regions of the subject, extracted from each of the aforementioned multiple captured images, are combined as a video into a composite image. The program according to claim 4 or 5, which causes the computer to perform the processing.
7. The system accepts input for image processing to be performed on the subject's shooting area when compositing the subject's shooting area into the composite image, and input for the compositing position for the composite image. The image processing performed on the subject's shooting area is then applied, and the processed subject's shooting area is composited into the composite image at the designated composite position. The program according to claim 5 or 6, which causes the computer to perform the processing.
8. Obtain a background image by taking a picture of the background, A photographic image including the background and subject is obtained. Select at least one of the aforementioned subjects, From the aforementioned captured image, extract the subject area containing the selected subject. From the aforementioned background image, the background region corresponding to the subject region is extracted. A trained model, which is trained to output a background-removed image from which the background region has been removed when given a background image taken of the background and a captured image including the background and the subject, is given an image of the background region extracted from the background image and an image of the subject region extracted from the captured image, and a background-removed image is obtained from the input subject region image, in which the background region in the input background region image has been removed. For the acquired background-removed image, the areas of subjects other than the selected subject are changed to background areas. An image processing method in which a computer performs the processing.
9. In an image processing apparatus having a control unit, The control unit, Obtain a background image by taking a picture of the background, A photographic image including the background and subject is obtained. Select at least one of the aforementioned subjects, From the aforementioned captured image, extract the subject area containing the selected subject. From the aforementioned background image, the background region corresponding to the subject region is extracted. A trained model, which is trained to output a background-removed image from which the background region has been removed when given a background image taken of the background and a captured image including the background and the subject, is given an image of the background region extracted from the background image and an image of the subject region extracted from the captured image, and a background-removed image is obtained from the input subject region image, in which the background region in the input background region image has been removed. For the acquired background-removed image, the areas of subjects other than the selected subject are changed to background areas. Image processing device.
10. An image processing system including a camera and an image processing device, The aforementioned imaging device is The system includes a transmission unit that sends a background image, which is a photograph of the background, and a captured image, which includes the background and the subject, to the image processing device. The aforementioned image processing device is The background image is acquired from the aforementioned imaging device. The captured image is acquired from the aforementioned imaging device. Select at least one of the aforementioned subjects, From the aforementioned captured image, extract the subject area containing the selected subject. From the aforementioned background image, the background region corresponding to the subject region is extracted. A trained model, which is trained to output a background-removed image from which the background region has been removed when given a background image taken of the background and a captured image including the background and the subject, is given an image of the background region extracted from the background image and an image of the subject region extracted from the captured image, and a background-removed image is obtained from the input subject region image, in which the background region in the input background region image has been removed. For the acquired background-removed image, the areas of subjects other than the selected subject are changed to background areas. Based on the image of the subject region extracted from the aforementioned captured image and the background-removed image in which the regions of subjects other than the selected subject have been changed to background regions, the captured region of the subject is extracted from the image of the subject region. The extracted shooting area of the subject is combined into the composite image. It includes a control unit that performs processing. Image processing system.
Citation Information
Patent Citations
Image processor, image processing method, and program
JP2009177358A
Image processing apparatus
JP2013143069A
Information processing method and information processing apparatus
JP2020014051A
Photographing apparatus, method for making photographic image, and program
JP2020127165A
Image processing apparatus, image processing method, and program
JP2021182320A