Image processing apparatus, image processing system, image processing method, learning apparatus, learning method, and program

The image processing apparatus employs two learned models to accurately extract foreground regions from images, addressing the challenges of overfitting and computational load, especially when specific subjects are present.

JP7682694B2Active Publication Date: 2025-05-26CANON KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021089252
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-27
Publication Date
2025-05-26
Estimated Expiration
2041-05-27

AI Technical Summary

Technical Problem

Existing foreground extraction methods face challenges in accurately extracting foreground regions from images, especially when a specific subject with repeating patterns is present, leading to decreased extraction accuracy due to overfitting and increased computational load.

Method used

The proposed solution involves an image processing apparatus that uses two learned models to extract foreground regions. The first model processes regions without specific subjects, while the second model, trained on images with specific subjects, provides higher extraction accuracy when specific subjects are present. This approach sets distinct regions in the image for processing, thereby managing computational load effectively.

Benefits of technology

This method enables accurate foreground extraction regardless of the presence of specific subjects, while maintaining a controlled computational load, thus improving extraction accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682694000001
    Figure 0007682694000001
  • Figure 0007682694000002
    Figure 0007682694000002
  • Figure 0007682694000003
    Figure 0007682694000003
Patent Text Reader

Abstract

To extract a foreground region from a captured image at high accuracy while suppressing an increase in computation load in extracting the foreground region irrespective of whether or not a specific object shows up in the captured image.SOLUTION: A first region not including an image region where a specific object shows up, and a second region including an image region where the specific object shows up are set to a captured image. A first foreground region indicating a foreground region included in the first region extracted by a first learned model based on the captured image and the first region, and a second foreground region indicating a foreground region included in the second region extracted by a second learned model based on the captured image and the second region are obtained. Here, extraction accuracy of the second foreground region extracted by the second learned model based on the captured image and on the second region is higher than extraction accuracy of the second foreground region extracted by the first learned model based on the captured image and on the second region.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image processing technique for extracting a foreground region in an image.

Background Art

[0002] There is a technique for generating an image (hereinafter referred to as a "virtual viewpoint image") seen from a virtual viewpoint specified by a user using an image captured by a plurality of cameras (hereinafter referred to as a "captured image"). According to the virtual viewpoint image, for example, it is possible to view highlight scenes of a game such as soccer or basketball from various angles.

[0003] When generating a virtual viewpoint image, foreground extraction is performed to extract an image region corresponding to an object from the captured image as a foreground region. As one of the foreground extraction methods, a method of extracting from a captured image as a foreground region using a learned model obtained by machine learning is known. Non-Patent Document 1 discloses a foreground extraction method that combines a convolutional neural network (hereinafter referred to as "CNN (Convolutional neural networks)"), which is a learned model obtained by machine learning, and a deconvolution network.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the foreground extraction method disclosed in Non-Patent Document 1, when a part of the image area corresponding to a predetermined moving object such as a person to be extracted as the foreground area overlaps with a part of the image area corresponding to a specific subject, the extraction accuracy of the foreground area may decrease. Here, the specific subject is a subject having a repeating pattern such as a goal net. In the conventional method, in order to improve the extraction accuracy of the foreground area in such a captured image, when generating a learned model by machine learning, many captured images in which a specific subject such as a soccer goal appears are used as learning data and the learning model is made to learn. However, in such learning, due to overfitting, the extraction accuracy of the foreground area in a captured image in which a specific subject does not appear may decrease. Also, whether or not a specific subject appears, it is also conceivable to make the learning model learn a learning model with a complicated model structure in order to improve the extraction accuracy of the foreground area in the captured image. However, in the learned model obtained by such a learning model with a complicated model structure, the computational load when extracting the foreground area becomes heavy.

[0006] An object of the present disclosure is to enable accurate extraction of a foreground area from a captured image while suppressing an increase in the computational load when extracting the foreground area, whether or not a specific subject appears in the captured image.

Means for Solving the Problems

[0007] The image processing apparatus according to the present disclosure includes an image acquisition unit that acquires a captured image, a setting unit that sets a first region that does not include an image region in which a specific subject appears in the captured image and a second region that includes the image region in which the specific subject appears in the captured image, a foreground acquisition unit that acquires a first foreground region indicating a foreground region included in the first region, which is extracted by a first learned model based on the captured image and the first region, and a second foreground region indicating a foreground region included in the second region, which is extracted by a second learned model based on the captured image and the second region. The extraction accuracy of the second foreground region extracted by the second learned model based on the captured image and the second region is higher than the extraction accuracy of the second foreground region extracted by the first learned model based on the captured image and the second region.

Advantages of the Invention

[0008] According to the present disclosure, even when a specific subject appears or does not appear in the captured image, it is possible to accurately extract the foreground region from the captured image while suppressing an increase in the computational load when extracting the foreground region.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Embodiments for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Note that the configurations shown in the following embodiments are merely examples, and the scope of the present disclosure is not limited only to those configurations.

[0011] [First Embodiment] Hereinafter, with reference to FIGS. 1 to 8, the image processing apparatus 100 according to the first embodiment will be described. FIG. 1 is a schematic diagram showing an example of the configuration of an image processing system 1 to which the image processing apparatus 100 according to the first embodiment is applied. The image processing system 1 includes a plurality of imaging devices 10, a plurality of image processing devices 100, and an image computing server 11.

[0012] The imaging device 10 is composed of a digital video camera, a digital still camera, or the like, and outputs an image obtained by imaging an imaging target (hereinafter referred to as an "imaging image") as imaging image information. The plurality of imaging devices 10 are arranged so as to surround a field 12 that is the imaging target. The imaging image information output from each of the plurality of imaging devices 10 is input to the image processing device 100 corresponding to the imaging device 10. Hereinafter, as shown in FIG. 1, the imaging device 10 and the image processing device 100 correspond one-to-one, and one image processing device 100 will be described as acquiring the imaging image information output from a predetermined one imaging device 10 corresponding to the image processing device 100. Note that the configuration shown in FIG. 1 is an example, and the image processing device 100 may acquire the imaging image information output from a plurality of imaging devices 10.

[0013] Each of the plurality of image processing devices 100 acquires the imaging image information output from the corresponding imaging device 10. The image processing device 100 extracts a foreground region of a subject such as a person in the imaging image indicated by the imaging image information (hereinafter referred to as an "object"), and generates information indicating the foreground region (hereinafter referred to as "foreground information"). The image processing device 100 outputs the imaging image information and the foreground information corresponding to the imaging image information to the image computing server 11. Details of the image processing device 100 will be described later.

[0014] The image computing server 11 acquires the captured image information and foreground information output by each of the plurality of image processing apparatuses 100. Based on the foreground information output by each of the plurality of image processing apparatuses 100, the image computing server 11 generates three-dimensional shape data of an object. Specifically, for example, the image computing server 11 generates three-dimensional shape data of an object using the technology of Visual Hull. Since the method of generating three-dimensional shape data of an object using the technology of Visual Hull is well-known, the description thereof is omitted. Based on the generated three-dimensional shape data and the acquired captured image information, the image computing server 11 generates an image (hereinafter referred to as a "virtual viewpoint image") viewed from a viewpoint specified by a user or the like, and outputs information indicating the generated virtual viewpoint image. Since the method of generating a virtual viewpoint image based on three-dimensional shape data and a captured image is well-known, the description thereof is omitted.

[0015] Note that, as an example, the image processing system 1 shown in FIG. 1 has a star configuration in which the image processing apparatuses 100 to which a plurality of imaging apparatuses 10 are connected are each connected to the image computing server 11. The image processing system 1 is not limited to the star configuration. For example, the image processing system 1 may have a configuration in which the image processing apparatuses 100 are connected to each other in a daisy chain, and one or more of the plurality of image processing apparatuses 100 are connected to the image computing server 11. Also, in FIG. 1, ten imaging apparatuses 10 and ten image processing apparatuses 100 are shown, but the number of imaging apparatuses 10 and image processing apparatuses 100 may be nine or less or eleven or more, and the number is not limited. Further, the image processing apparatus 100 may be provided as an image processing unit inside the imaging apparatus 10.

[0016] FIG. 2 is a functional block diagram showing an example of the configuration of the image processing apparatus 100 according to the first embodiment. The image processing apparatus 100 includes an image acquisition unit 110, a setting unit 120, a candidate extraction unit 130, a generation unit 140, an inference unit 150, a foreground acquisition unit 160, and an output unit 170. In the first embodiment, the image processing apparatus 100 is described as including the generation unit 140 and the candidate extraction unit 130, but the generation unit 140 and the candidate extraction unit 130 are not essential components in the image processing apparatus 100. Hereinafter, with reference to FIGS. 2 and 3, the processing of each component included in the image processing apparatus 100 will be described. Also, the image processing apparatus 100 is described as being connected one-to-one to the corresponding imaging apparatus 10, but the image processing apparatus 100 may be connected one-to-many to a plurality of corresponding imaging apparatuses 10. In this case, for example, the image processing apparatus 100 includes a plurality of setting units 120, candidate extraction units 130, generation units 140, inference units 150, and foreground acquisition units 160 corresponding to the plurality of corresponding imaging apparatuses 10.

[0017] The image acquisition unit 110 acquires a captured image. Specifically, the image acquisition unit 110 acquires a captured image by acquiring the captured image information output by the imaging apparatus 10. The image acquisition unit 110 may also acquire a captured image by reading out and acquiring the captured image information from a storage device (not shown in FIG. 1) in which the captured image information is stored in advance. FIG. 3(a) is a diagram showing an example of the captured image 310 acquired by the image acquisition unit 110 according to the first embodiment. Specifically, FIG. 3(a) shows, as an example, a captured image 310 obtained by an imaging apparatus 10 capturing a soccer game. The captured image 310 includes players 311, 312, 313, a field 301 to be a background area, and a goal mouth 303 including a goal net 302 (hereinafter, the goal net 302 and the goal mouth 303 are collectively referred to as "goal 304"). The image acquisition unit 110 outputs the acquired captured image to the candidate extraction unit 130, the generation unit 140, and the output unit 170. Hereinafter, the image area to be extracted as the foreground area from the captured image 310 is described as being the image area corresponding to each of the players 311, 312, 313.

[0018] The candidate extraction unit 130 extracts candidate regions (hereinafter simply referred to as "candidate regions") that are candidates for the foreground region from the captured image, and outputs information indicating the extracted candidate regions to the generation unit 140. Specifically, for example, the candidate extraction unit 130 extracts candidate regions by detecting moving objects appearing in the captured image from the captured image by the background difference method. For example, the candidate extraction unit 130 extracts a rectangular range that includes the outermost periphery of the region of the detected moving object in the captured image as the candidate region. The method by which the candidate extraction unit 130 extracts candidate regions is not limited to the background difference method, and the shape of the candidate regions extracted by the candidate extraction unit 130 is not limited to a rectangle as long as it includes the outermost periphery of the region of the moving object in the captured image. FIG. 3(b) is a diagram showing an example of candidate regions 321, 322, and 323 extracted by the candidate extraction unit 130 from the captured image 310. The candidate regions 321, 322, and 323 are candidate regions corresponding to the players 311, 312, and 313. Although the candidate regions 321, 322, and 323 shown as an example in FIG. 3(b) are all rectangular images, the candidate regions output by the candidate extraction unit 130 only need to be information indicating at least the range of the foreground candidates, and do not necessarily have to be in the shape of a rectangle or in the form of an image.

[0019] The setting unit 120 sets a first region and a second region for the captured image. Here, the second region is, for example, a region including an image region in which a specific subject (hereinafter referred to as "specific subject") in the captured image is captured, and the first region is, for example, a region not including an image region in which a specific subject in the captured image is captured. The specific subject is, for example, a subject having a repeating pattern such as a goal net 302, or a subject having a color scheme similar to the uniforms worn by players 311, 312, and 313. The subject having a color scheme similar to the uniform is, for example, a field logo arranged on the field 301 or a signboard arranged around the field 301. The first region and the second region can be determined in advance by the user, for example, for each imaging device 10 or for each angle of view preset for each imaging device 10. In this case, for example, the setting unit 120 acquires the first region and the second region by reading information indicating the first region and the second region prepared in advance from a storage device (not shown in FIGS. 1 and 2) provided inside or outside the image processing apparatus 100, and sets the first region and the second region.

[0020] The setting unit 120 may identify an object shown in the captured image by analyzing the captured image, determine whether a specific subject is shown in the captured image, and if so, extract the region in the captured image where the specific subject is shown, and set a first region and a second region. For example, the setting unit 120 may set, as the second region, a region including the region in the captured image where the specific subject is shown, and set, as the first region, a region in the captured image where the specific subject is not shown. When a specific subject is shown in the captured image, the setting unit 120 may set the entire captured image as the second region, and when a specific subject is not shown in the captured image, the setting unit 120 may set the entire captured image as the first region. Further, when a specific subject is shown in the captured image, the setting unit 120 may set the first region and the second region as follows. For example, when the ratio of the area of the image region corresponding to the specific subject to the area of the entire captured image is greater than a predetermined threshold, the setting unit 120 sets the entire captured image as the second region, and when the ratio is less than or equal to the threshold, the setting unit 120 sets the entire captured image as the first region. Since the method of identifying an object shown in an image and the method of extracting the region where the object is shown are well-known, the description thereof is omitted.

[0021] Hereinafter, it is assumed that the specific subject in the captured image 310 shown in FIG. 3(a) is the goal 304. FIG. 3(c) is a diagram showing an example of the first region 331 and the second region 332 set by the setting unit 120 for the captured image 310. In FIG. 3(c), as the second region 332, a rectangular region including the image region where the goal 304, which is the specific subject, is shown is indicated by a one-dot chain line. Although the second region 332 shown in FIG. 3(c) is a rectangular region, the second region set by the setting unit 120 only needs to be a region including at least the image region where the goal 304 in the captured image 310 is shown, and does not necessarily have to be a rectangular region. In FIG. 3(c), the first region 331 is the entire region other than the second region 332 in the entire region of the captured image 310. Further, the first region 331 set by the setting unit 120 does not necessarily have to be the entire region other than the second region 332 in the entire region of the captured image 310, and may be a part of the region other than the second region 332 in the entire region of the captured image 310.

[0022] The inference unit 150 consists of a learned model corresponding to the learning result by machine learning. Using the learned model, it infers and extracts a foreground region, which is an image region corresponding to an object, from an image (hereinafter referred to as "input image") input as an explanatory variable. Further, the inference unit 150 outputs information (foreground information) indicating the extracted foreground region as an inference result. Specifically, the inference unit 150 consists of a first learned model (hereinafter referred to as "first learned model 151") and a second learned model (hereinafter referred to as "second learned model 152"). The inference unit 150 extracts the foreground region in the captured image 310 using the first or second learned model 151, 152, and outputs the foreground information to the foreground acquisition unit 160 as an inference result. Details of the first and second learned models 151, 152 will be described later. Note that the learning method for training the first and second learned models 151, 152 is not limited to machine learning, and may be, for example, a learning method in which a natural person sets, adjusts, or modifies parameters, etc. Hereinafter, the first and second learned models 151, 152 will be described as learned models corresponding to the learning result by machine learning.

[0023] The generation unit 140 generates input images for inputting into each of the first and second learned models 151 and 152. Specifically, the generation unit 140 generates a first image for inputting as an input image into the first learned model 151 by cutting out at least a part of the first region 331 from the captured image 310. Also, the generation unit 140 generates a second image for inputting as an input image into the second learned model 152 by cutting out at least a part of the second region 332 from the captured image 310. Further, the generation unit 140 outputs the generated first image and second image to the inference unit 150. More specifically, for example, first, the generation unit 140 acquires the captured image 310 output by the image acquisition unit 110 and the first region 331 and the second region 332 set by the setting unit 120. The generation unit 140 generates a first image by cutting out the region corresponding to the first region 331 from the captured image 310, and outputs the generated first image to the inference unit 150. Also, the generation unit 140 generates a second image by cutting out the region corresponding to the second region 332 from the captured image 310, and outputs the generated second image to the inference unit 150. The inference unit 150 inputs the first image as an explanatory variable into the first learned model, and inputs the second image as an explanatory variable into the second learned model. The first learned model 151 infers and extracts the foreground region from the first image input as an explanatory variable, and the second learned model 152 infers and extracts the foreground region from the second image input as an explanatory variable. Each of the first and second learned models 151 and 152 outputs information (foreground information) indicating the extracted foreground region as an inference result to the foreground acquisition unit 160.

[0024] Note that the generation unit 140 is not an essential component in the image processing apparatus 100. When the image processing apparatus 100 does not include the generation unit 140, the inference unit 150 extracts a foreground region in the captured image based on the captured image 310 and the first and second regions, and outputs foreground information. Specifically, in this case, the first learned model 151 extracts a foreground region in the captured image included in the first region based on the captured image and the first region, and outputs information indicating the extracted foreground region (foreground information) to the foreground acquisition unit 160. More specifically, for example, the first learned model 151 cuts out a first image from the captured image based on the captured image and the first region, and extracts a foreground region in the first image based on the first image. Similarly, the second learned model 152 extracts a foreground region in the captured image included in the second region based on the captured image and the second region, and outputs information indicating the extracted foreground region (foreground information) to the foreground acquisition unit 160. More specifically, for example, the second learned model 151 cuts out a second image from the captured image based on the captured image and the second region, and extracts a foreground region in the second image based on the second image.

[0025] When the image processing apparatus 100 includes the candidate extraction unit 130, for example, the generation unit 140 cuts out an image corresponding to each of one or more candidate regions (for example, candidate regions 322 and 323) including at least a part of the first region 331. Thereby, the generation unit 140 generates one or more first images. The generation unit 140 sequentially outputs each of the generated one or more first images to the inference unit 150, and the inference unit 150 sequentially inputs each first image received from the generation unit 140 to the first learned model 151. Further, the generation unit 140 cuts out an image corresponding to each of one or more candidate regions (for example, candidate region 321) including at least a part of the second region 332. Thereby, the generation unit 140 generates one or more second images. The generation unit 140 sequentially outputs each of the generated one or more second images to the inference unit 150. The inference unit 150 sequentially inputs each second image received from the generation unit 140 to the second learned model 152. By inputting such first images and second images generated by the generation unit 140 to the first or second learned models 151 and 152, the processing load of foreground region extraction in the first and second learned models 151 and 152 can be reduced.

[0026] When the image processing apparatus 100 does not include the candidate extraction unit 130, the generation unit 140 may generate the first image and the second image as follows. For example, first, the generation unit 140 generates one or more first images by dividing the first image corresponding to the first region 331 in the captured image 310 into a plurality of images such as N×M (N and M are integers of 1 or more, and at least one of them is an integer of 2 or more). The generation unit 140 sequentially outputs the generated one or more first images to the inference unit 150, so that each first image is input to the first learned model 151. Similarly, the generation unit 14 generates one or more second images by dividing the second image corresponding to the second region 332 in the captured image 310 into a plurality of images such as N×M. The generation unit 140 sequentially outputs the generated one or more second images to the inference unit 150, so that each second image is input to the second learned model 152. In this way, for each of the divided images, the first and second learned models 151 and 152 may be made to infer the foreground region. By configuring in this way, the processing load of foreground region extraction in the first or second learned models 151 and 152 can be reduced compared to the case where the entire first image or second image is input to the corresponding first or second learned models 151 and 152.

[0027] In addition, when the image processing apparatus 100 includes the candidate extraction unit 130 and does not include the generation unit 140, the inference unit 150 extracts the foreground region in the captured image based on the captured image 310, the first region and the second region, and the candidate region, and outputs foreground information. Specifically, in this case, the first learned model 151 extracts the foreground region in the captured image included in the first region based on the captured image, the first region, and the candidate region, and outputs information (foreground information) indicating the extracted foreground region to the foreground acquisition unit 160. More specifically, for example, the first learned model 151 cuts out one or more first images from the captured image based on the captured image, the first region, and the candidate region, and extracts the foreground region in each first image based on the first image. Similarly, the second learned model 152 extracts the foreground region in the captured image included in the second region based on the captured image, the second region, and the candidate region, and outputs information (foreground information) indicating the extracted foreground region to the foreground acquisition unit 160. More specifically, for example, the second learned model 151 cuts out one or more second images from the captured image based on the captured image, the second region, and the candidate region, and extracts the foreground region in each of the second images based on the second image.

[0028] FIG. 3(d) is a diagram showing an example of the input image 342 output from the generation unit 140 to the inference unit 150. The input image 342 shown as an example in FIG. 3(d) is an image obtained by the generation unit 140 cutting out an image region corresponding to the candidate region 322 from the captured image 310. Since the candidate region 322 shown as an example in FIG. 3(d) includes at least a part of the first region 331, the input image 342 is a first image, and the inference unit 150 inputs the input image 342 as an explanatory variable to the first learned model 151.

[0029] The foreground acquisition unit 160 acquires the foreground information output by the inference unit 150 as the inference result, that is, the foreground information output by the first and second learned models 151 and 152 as the inference result, and acquires the foreground area in the captured image 310 based on the foreground information. Specifically, the foreground acquisition unit 160 acquires the foreground information output by the first learned model 151 as the inference result, and acquires the foreground area in the first area 331. In addition, the foreground acquisition unit 160 acquires the foreground information output by the second learned model 152 as the inference result, and acquires the foreground area in the second area 332. FIG. 3(e) is a diagram showing an example of the foreground area 352 indicated by the foreground information acquired by the foreground acquisition unit 160 from the inference unit 150. Specifically, the foreground area 352 shown as an example in FIG. 3(e) corresponds to the foreground information output by the first learned model 151 when the inference unit 150 inputs the input image 342 as an explanatory variable. More specifically, for example, the foreground acquisition unit 160 acquires the foreground information as an image in which the foreground area corresponding to the object (player), which is the area corresponding to the foreground area 352 shown as an example in FIG. 3(e), is binarized with white pixels and the other areas (background areas) are binarized with black pixels.

[0030] The foreground acquisition unit 160 may acquire the foreground information output by the first learned model 151 and the second learned model 152 respectively, and generate a mask image in which a plurality of foreground areas corresponding to the foreground information are arranged in one image. FIG. 3(f) is a diagram showing an example of the mask image 360 generated by the foreground acquisition unit 160. The mask image 360 shown as an example in FIG. 3(f) corresponds to the entire image area of the captured image 310, and is an image in which the foreground area is binarized with white pixels and the background area is binarized with black pixels. In FIG. 3(f), foreground areas 361, 362, and 363 are shown as foreground areas corresponding to players 311, 312, and 313.

[0031] The output unit 170 outputs information indicating the foreground area acquired by the foreground acquisition unit 160. Specifically, the output unit 170 integrates the information indicating the foreground area acquired by the foreground acquisition unit 160 and the information indicating the captured image 310 output by the image acquisition unit 110, and outputs the integrated information to the image computing server 11.

[0032] Referring to FIG. 4, the first and second learned models 151 and 152 will be described. FIG. 4(a) is a configuration diagram showing an example of the configuration of the first learned model 151 according to the first embodiment, and FIG. 4(b) is a configuration diagram showing an example of the configuration of the second learned model 152 according to the first embodiment. As shown as an example in FIG. 4(a), the first learned model 151 is constituted by, for example, a neural network 410 having an input layer 411, an intermediate layer 412 having one or more layers, and an output layer 413. Further, each of the input layer 411, the layers in the intermediate layer 412, and the output layer 413 has one or more neurons indicated by circles in FIG. 4(a). Similarly, as shown as an example in FIG. 4(b), the second learned model 152 is constituted by, for example, a neural network having an input layer 421, an intermediate layer 422 having one or more layers, and an output layer 423. Further, each of the input layer 421, the layers in the intermediate layer 422, and the output layer 423 has one or more neurons indicated by circles in FIG. 4(b). The second learned model 152 may be constituted by the same neural network 410 as the first learned model 151 having the input layer 411, the intermediate layer 412, and the output layer 413 shown as an example in FIG. 4(a).

[0033] Specifically, for example, the first and second learned models 151 and 152 are constituted by a convolutional neural network (hereinafter referred to as "CNN (Convolutional neural networks)"), which is one of neural networks. The first and second learned models 151 and 152 may be constituted by a combination of the CNN disclosed in Non-Patent Document 1 and a deconvolution network. Further, the CNN is merely an example, and the first and second learned models 151 and 152 may be constituted by neural networks other than the CNN. Further, the first and second learned models 151 and 152 are not limited to those constituted by neural networks as long as they are learned models corresponding to the learning results by learning.

[0034] Both the first and second learned models 151 and 152 perform inference on the foreground region for the input image input as an explanatory variable and output foreground information as the inference result. Both the first and second learned models 151 and 152 perform inference on the foreground region, but they are different from each other. Hereinafter, a method for generating the first and second learned models 151 and 152 will be described. For example, the first and second learned models 151 and 152 are generated by a learning device.

[0035] With reference to FIG. 5, the configuration of the learning device 500 that generates the first and second learned models 151 and 152 will be described. FIG. 5 is a functional block diagram showing an example of the configuration of the learning device 500 according to the first embodiment. The learning device 500 includes an image group acquisition unit 510, a model acquisition unit 520, a learning unit 530, and a model output unit 540.

[0036] The image group acquisition unit 510 acquires an image group composed of a plurality of images as learning data. Specifically, the image group acquisition unit 510 acquires a first image group (hereinafter referred to as the "first image group") as the first learning data (hereinafter referred to as the "first learning data"). Further, the image group acquisition unit 510 acquires a second image group (hereinafter referred to as the "second image group") as the second learning data (hereinafter referred to as the "second learning data"). For example, the image group acquisition unit 510 acquires the image group by reading an image corresponding to the first or second image group from a storage device provided inside or outside the image processing device 100 or the learning device 500. Here, the first image group and the second image group are different image groups from each other. Specifically, for example, the second image group acquired as the second learning data includes more images in which a specific subject is captured than the first image group acquired as the first learning data. Also, for example, the second image group has a ratio of the number of images in which a specific subject is captured that is greater than a predetermined ratio. Here, the ratio of the number of images in which a specific subject is captured is, for example, the ratio of the total number of images in which a specific subject is captured to the total number of images included in the second image group.

[0037] The model acquisition unit 520 acquires a learning model before learning or during learning. Specifically, the model acquisition unit 520 acquires a first learning model (hereinafter referred to as the "first learning model"), which becomes the first learned model 151 corresponding to the learning result through learning by machine learning. Also, the model acquisition unit 520 acquires a second learning model (hereinafter referred to as the "second learning model"), which becomes the second learned model 152 corresponding to the learning result through learning by machine learning.

[0038] The learning unit 530 causes the learning model to learn by machine learning using the image group as learning data, and generates a learned model capable of inferring the foreground region in the input image input as an explanatory variable. Specifically, the learning unit 530 generates the first learned model 151 by causing the first learning model to perform machine learning using the first image group as the first learning data. Also, the learning unit 530 generates the second learned model 152 by causing the second learning model to perform machine learning using the second image group as the second learning data. Such a second learned model 152 can perform inference of the foreground region in the input image with high accuracy when an input image in which a specific subject is captured is input as an explanatory variable. That is, the extraction accuracy of the foreground region extracted by the second learned model based on the input image in which a specific subject is captured is higher than the extraction accuracy of the foreground region extracted by the first learned model based on the same input image.

[0039] For example, the learning unit 530 causes the first learning model and the second learning model to learn by machine learning with supervised learning, and generates the first and second learned models 151 and 152. When the learning unit 530 performs supervised learning, the teacher data corresponding to each image of the first image group and the teacher data corresponding to each image of the second image group are acquired by, for example, the image group acquisition unit 510. Since the machine learning method with supervised learning is well-known, the description thereof is omitted. Note that the machine learning method in the learning unit 530 is not limited to machine learning with supervised learning, and may be machine learning with unsupervised learning such as reinforcement learning.

[0040] The first learning model and the second learning model may be different from each other or may be similar to each other. When the first learning model and the second learning model are similar to each other, the first learned model 151 and the second learned model 152 have the same configuration, specifically, for example, the configuration shown as an example in FIG. 4(a).

[0041] On the other hand, when the first learning model and the second learning model are different from each other, the first learning model and the second learning model have differences as follows, for example. For example, when the second learning model is subjected to machine learning in the second learning model using an image in which a specific subject is captured as learning data, the value of the loss function is set to be smaller than when the same machine learning is performed on the first learning model. The second learned model 152 generated by subjecting such a second learning model to machine learning can perform inference on the foreground region in the input image with high accuracy when the input image in which the specific subject is captured is input as an explanatory variable, as compared with the first learned model 151.

[0042] Also, for example, the second learning model may have a larger number of intermediate layers or a larger number of neurons in the intermediate layer than the first learning model. As a result, the second learned model 152 has a larger number of intermediate layers or a larger number of neurons in the intermediate layer than the first learned model 151. Such a second learned model 152 can perform inference on the foreground region in the input image with high accuracy when the input image in which the specific subject is captured is input as an explanatory variable, as compared with the first learned model 151.

[0043] The model output unit 540 outputs the first and second learned models 151 and 152. More specifically, for example, the model output unit 540 outputs the first and second learned models 151 and 152 to a storage device provided inside or outside the image processing apparatus 100 or the learning apparatus 500. The model output unit 540 causes the output first and second learned models 151 and 152 to be stored in the storage device. The inference unit 150 included in the image processing apparatus 100 reads the first and second learned models 151 and 152 pre-stored in the storage device, and performs inference using the first and second learned models 151 and 152.

[0044] The processing of each unit included in the image processing apparatus 100 is performed by hardware such as an ASIC (application specific integrated circuit) built in the image processing apparatus 100. The processing may be performed by hardware such as an FPGA (field programmable gate array). Further, the processing may be performed by software using a CPU (Central Processor Unit) or a GPU (Graphic Processor Unit), and a memory.

[0045] Referring to FIG. 6, the hardware configurations of the image processing apparatus 100 and the learning apparatus 500 when each unit included in the image processing apparatus 100 and the learning apparatus 500 operates as software will be described. FIG. 6(a) is a block diagram showing an example of the hardware configuration of the image processing apparatus 100 according to the first embodiment. FIG. 6(b) is a block diagram showing an example of the hardware configuration of the learning apparatus 500 according to the first embodiment. Each of the image processing apparatus 100 and the learning apparatus 500 is configured by, for example, a computer. The computer has a GPU 610, a CPU 611, a ROM 612, a RAM 613, an auxiliary storage device 614, a communication I / F 617, and a system bus 618 as shown as an example in FIG. 6(a) or FIG. 6(b).

[0046] The CPU 611 controls the computer using programs or data stored in the ROM 612, the RAM 613, or the like. Thereby, the CPU 611 functions as each part included in the image processing apparatus 100 shown in FIG. 2 or the learning apparatus 500 shown in FIG. 5. Similarly to the CPU 611, the GPU 610 controls the computer using programs or data stored in the ROM 612, the RAM 613, or the like. The GPU 610 can perform efficient calculations by processing more data in parallel. When performing machine learning repeatedly a plurality of times using a learning model such as deep learning, it is effective to perform the calculation processing in the machine learning by the GPU 610. Also, when performing calculation processing using a learned model corresponding to the learning result by deep learning, it is effective to perform the calculation processing by the GPU 610.

[0047] Therefore, the processing in the inference unit 150 included in the image processing apparatus 100 and the processing in the learning unit 530 included in the learning apparatus 500 are performed by, for example, the CPU 611 and the GPU 610. Specifically, for example, the inference by the first and second learned models 151 and 152 in the inference unit 150 included in the image processing apparatus 100 is performed by the CPU 611 and the GPU 610 cooperating to perform calculations. Note that the processing of each part other than the inference unit 150 included in the image processing apparatus 100 and the processing of each part other than the learning unit 530 included in the learning apparatus 500 may be performed by only the CPU 611 or the GPU 610. Also, the image processing apparatus 100 and the learning apparatus 500 have one or more dedicated hardware different from the CPU 611 or the GPU 610, and at least a part of the processing by the CPU 611 or the GPU 610 may be executed by the dedicated hardware. Examples of the dedicated hardware include ASIC, FPGA, and DSP (digital signal processor).

[0048] The ROM 612 stores programs and the like that do not require modification. The RAM 613 temporarily stores programs or data supplied from the auxiliary storage device 614, or data and the like supplied from the outside via the communication I / F 617. The auxiliary storage device 614 is composed of, for example, a hard disk drive or the like, and stores various data such as image data or audio data. The communication I / F 617 is used for communication between the image processing apparatus 100 and an external device such as the image computing server 11, and for communication between the learning apparatus 500 and an external device. For example, when the image processing apparatus 100 or the learning apparatus 500 is wired-connected to an external device, a communication cable is connected to the communication I / F 617. When the image processing apparatus 100 or the learning apparatus 500 has a function of wireless communication with an external device, the communication I / F 617 is provided with an antenna. The system bus 618 connects each part of the image processing apparatus 100 and also connects each part of the learning apparatus 500 to transmit information.

[0049] Referring to FIG. 7, the operation of the learning apparatus 500 will be described. FIG. 7 is a flowchart showing an example of a processing flow in the learning apparatus 500 according to the first embodiment. Hereinafter, as an example, with reference to FIG. 7, the operation of the learning apparatus 500 for generating the first learned model 151 will be described. The operation when the learning apparatus 500 generates the second learned model 152 is the same as the operation of the learning apparatus 500 for generating the first learned model 151. Specifically, since it is the same as the operation in which the first is replaced with the second, the description thereof will be omitted. In the following description, the symbol "S" means step.

[0050] First, in S701, the model acquisition unit 520 acquires the first learning model before or during learning. Next, in S702, the image group acquisition unit 510 acquires the first image group. Next, in S703, the learning unit 530 selects input images to be input to the first learning model from a plurality of images in the first image group and inputs them to the first learning model. Next, in S704, the learning unit 530 compares the foreground information output by the first learning model with the teacher data corresponding to the input images, and changes the parameters of the first learning model to cause the first learning model to be learned by machine learning. Next, in S705, the learning unit 530 determines whether a condition for ending machine learning (hereinafter referred to as the "ending condition") is satisfied. Here, the ending condition is when all images in the first image group are selected, when a predetermined period has elapsed since the start of machine learning, or when a predetermined number of images have been selected and input since the start of machine learning, etc.

[0051] If the learning unit 530 determines that the ending condition is not satisfied, the learning device 500 returns to the process of S703 and repeatedly executes the processes from S703 to S705 until the learning unit 530 determines that the ending condition is satisfied. If the learning unit 530 determines that the ending condition is satisfied, in S710, the learning unit 530 outputs the first learning model as the first learned model 151 to the model output unit 540. After S710, in S711, the model output unit 540 outputs information indicating the first learned model 151 to an external device such as the image processing device 100. After S711, the learning device 500 ends the process of the flowchart shown in FIG. 7. Note that the order of the processes in S701 and S702 is arbitrary.

[0052] Referring to FIG. 8, the operation of the image processing apparatus 100 will be described. FIG. 8 is a flowchart showing an example of a processing flow in the image processing apparatus 100 according to the first embodiment. First, in S801, the setting unit 120 sets the first region 331 and the second region 332. Next, in S802, the inference unit 150 acquires information indicating the first learned model 151 and information indicating the second learned model 152. After S802, the image processing apparatus 100 repeatedly executes the processing from S810 to S822 until a predetermined end condition is satisfied. Here, the end condition is, for example, when the power of the image processing apparatus 100 is turned off, when imaging image information is not acquired from the imaging apparatus 10 for a predetermined period, or when the angle of view of the imaging apparatus 10 is changed. Note that the order of the processes in S801 and S802 is arbitrary. After S802, in S810, the image processing apparatus 100 determines whether the end condition is satisfied.

[0053] If it is determined that the end condition is satisfied, the image processing apparatus 100 ends the processing of the flowchart shown in FIG. 8. If it is determined that the end condition is not satisfied, in S811, the image acquisition unit 110 acquires the captured image 310. After S811, in S812, the candidate extraction unit 130 extracts one or more candidate regions 321, 322, 323. For example, the candidate extraction unit 130 assigns an index number that can be distinguished from each other to each candidate region 321, 322, 323. After S812, in S813, the generation unit 140 selects one of the one or more extracted candidate regions 321, 322, 323. For example, the generation unit 130 selects the candidate region corresponding to the smallest index number among the candidate regions that have not been selected so far. The image processing apparatus 100 repeatedly executes the processing from S813 to S821 until it is determined in S821 described later that all the candidate regions 321, 322, 323 have been selected. After S813, in S814, the generation unit 140 determines whether at least a part of the selected candidate region includes at least a part of the second region 332.

[0054] When it is determined that at least a part of the candidate area includes at least a part of the second area 332, in S815, the generation unit 140 cuts out an image corresponding to the candidate area from the captured image 310 as a second image and outputs it to the inference unit 150. Further, the inference unit 150 inputs the image as an explanatory variable to the second learned model 152. After S815, in S816, the second learned model 152 in the inference unit 150 infers a foreground area from the image and outputs foreground information as an inference result. After S816, in S817, the foreground acquisition unit 160 acquires the foreground information output from the inference unit 150 and acquires the foreground area of the image. When it is determined that at least a part of the candidate area does not include at least a part of the second area 332, in S818, the generation unit 140 cuts out an image corresponding to the candidate area from the captured image 310 as a first area and outputs it to the inference unit 150. Further, the inference unit 150 inputs the image as an explanatory variable to the first learned model 151. After S818, in S819, the first learned model 151 in the inference unit 150 infers a foreground area from the image and outputs foreground information as an inference result. After S819, in S820, the foreground acquisition unit 160 acquires the foreground information output from the inference unit 150 and acquires the foreground area of the image. After S817 or S820, in S821, the generation unit 140 determines whether all the candidate areas 321, 322, 323 have been selected.

[0055] When it is determined that not all of the candidate areas 321, 322, 323 have been selected, the image processing apparatus 100 returns to the process of S813 and repeatedly executes the processes from S813 to S821. When it is determined that all of the candidate areas 321, 322, 323 have been selected, in S822, the output unit 170 outputs information indicating the foreground area acquired in S817 or S820. After S822, the image processing apparatus 100 returns to the process of S810 and repeatedly executes the processes from S810 to S822.

[0056] As described above, when at least a part of the candidate region is included in at least a part of the second region 332, the image processing apparatus 100 inputs the image corresponding to the candidate region as an explanatory variable to the second learned model 152 to perform inference on the foreground region. Thereby, it is possible to extract the foreground region with high accuracy. Also, as described above, when the candidate region is included in the first region 331, the image corresponding to the candidate region is input as an explanatory variable to the first learned model 151 to perform inference on the foreground region, whereby it is possible to extract the foreground region with high accuracy. As a result, according to the image processing apparatus 100, whether or not a specific subject is captured in the captured image 310, it is possible to accurately extract the foreground region from the captured image 310 while suppressing an increase in the calculation load when extracting the foreground region.

[0057] So far, the form in which the setting unit 120 sets two regions, the first region 331 and the second region 332, has been described. However, the setting unit 120 may set three or more regions. Specifically, for example, the second region 332 is a region in which a first specific subject is captured, the third region is a region in which a second specific subject different from the first specific subject is captured, and the first region 331 is a region in which neither the first nor the second specific subject is captured. In this case, the inference unit 150 includes a first learned model 151 to which a first image corresponding to the first region 331 is input, a second learned model 152 to which a second image corresponding to the second region 332 is input, and a third learned model to which an image corresponding to the third region is input.

[0058] As described above, the first region 331 and the second region 332 can be determined in advance for each imaging device 10 or for each preset angle of view for each imaging device 10. That is, when the imaging device 10 and the image processing device 100 are installed around the field 301, it is determined whether or not a specific subject appears in the captured image 310 obtained by the imaging device 10. That is, it is also determined whether or not the image processing device 100 corresponding to the imaging device 10 performs inference of the foreground region by the second learned model 152. The second learned model 152 has a higher processing load when inferring the foreground region compared to the first learned model 151. Therefore, there is a concern that the image processing device 100 that performs inference of the foreground region by the second learned model 152 generates more heat due to an increase in power consumption compared to the image processing device 100 that does not perform such inference. Therefore, for the image processing device 100 for which it is predetermined to perform inference of the foreground region by the second learned model 152, it is preferable to cool the image processing device 100 by blowing air with an air-cooling device such as a fan. By cooling the image processing device 100 with a cooling device, the CPU 610, GPU 611, etc. included in the image processing device 100 are cooled, and the efficiency of arithmetic processing is improved. As a result, the power consumption of the image processing device 100 can be reduced and the cost can be reduced. The cooling device may cool the image processing device 100 by air cooling or by a method other than air cooling.

[0059] Incidentally, although the form in which the image processing apparatus 100 includes the inference unit 150 has been described so far, the inference unit 150, that is, the functional block that infers the foreground region using the first and second learned models 151 and 152, does not have to be included in the image processing apparatus 100. When the image processing apparatus 100 does not include the inference unit 150, each of the image processing system 1 and the image processing apparatus 100 can be configured as follows, for example. The image processing system 1 includes an external device such as a cloud server (not shown in FIG. 1) that has a functional block corresponding to the inference unit 150. The generation unit 140 outputs the first image and the second image to the functional block that the external device has via the communication I / F 617. The functional block that the external device has inputs the first image as an explanatory variable to the first learned model 151 and inputs the second image as an explanatory variable to the second learned model 152. The foreground acquisition unit 160 acquires the foreground information output by the first and second learned models 151 and 152 in the functional block that the external device has via the communication I / F 617. With such a configuration, since the arithmetic processing load in the image processing apparatus 100 is reduced, it is possible to employ hardware such as the CPU 610 or GPU 611 included in the image processing apparatus 100 that has low arithmetic processing ability, for example, an inexpensive one.

[0060] [Second Embodiment] With reference to FIGS. 9 to 11, the image processing apparatus 100 according to the second embodiment will be described. The image processing apparatus 100 according to the second embodiment includes an image acquisition unit 110, a setting unit 120, a candidate extraction unit 130, a generation unit 140, an inference unit 150, a foreground acquisition unit 160, and an output unit 170, similarly to the image processing apparatus 100 according to the first embodiment. Each functional block of the generation unit 140, the inference unit 150, the foreground acquisition unit 160, and the output unit 170 included in the image processing apparatus 100 according to the second embodiment has a function different from that of each functional block included in the image processing apparatus 100 according to the first embodiment.

[0061] Hereinafter, to distinguish between the image processing apparatus 100 according to the first embodiment and the image processing apparatus 100 according to the second embodiment, the image processing apparatus 100 according to the second embodiment will be described as the image processing apparatus 100a. Similarly, hereinafter, the generation unit 140, the inference unit 150, the foreground acquisition unit 160, and the output unit 170 included in the image processing apparatus 100a will be described as the generation unit 140a, the inference unit 150a, the foreground acquisition unit 160a, and the output unit 170a. Further, the image processing system 1 according to the second embodiment includes a plurality of imaging devices 10, a plurality of image processing apparatuses 100a, and an image computing server 11, similar to the image processing system 1 according to the first embodiment. Since the image acquisition unit 110, the setting unit 120, and the candidate extraction unit 130 included in the image processing apparatus 100a are the same as those included in the image processing apparatus 100 according to the first embodiment, the description thereof will be omitted.

[0062] The inference unit 150a is composed of the first and second learned models 151 and 152. Using the first and second learned models 151 and 152, it infers a foreground region, which is an image region corresponding to an object, from the input input image, and outputs foreground information as an inference result. However, the first and second learned models 151 and 152 according to the second embodiment output, as an inference result, information indicating the reliability of the inference of the foreground region indicated by the foreground information (hereinafter referred to as "reliability information") in addition to the foreground information.

[0063] The inference unit 150a inputs the candidate region output by the candidate extraction unit 130 as an input image to the inference unit 150, that is, the first or second learned model 151, 152. Specifically, for a candidate region that includes at least a part of the first region and at least a part of the second region among the candidate regions, the inference unit 150a inputs an image corresponding to the candidate region as an explanatory variable to the first learned model and the second learned model 152.

[0064] Note that, for a candidate region that includes at least a part of the first region and does not include at least a part of the second region among the candidate regions, for example, the inference unit 150a inputs an image corresponding to the candidate region as an explanatory variable to the first learned model 151. The inference unit 150a may input an image corresponding to the candidate region as an explanatory variable to the first and second learned models 151 and 152. Further, for a candidate region that does not include at least a part of the first region and includes at least a part of the second region among the candidate regions, for example, the inference unit 150a inputs an image corresponding to the candidate region as an explanatory variable to the second learned model 152. The inference unit 150a may input an image corresponding to the candidate region as an explanatory variable to the first and second learned models 151 and 152.

[0065] Referring to FIG. 9, a candidate region that includes at least a part of the first region 331 and includes at least a part of the second region will be described. FIG. 9 is a diagram showing an example of candidate regions 321, 322, 924 extracted from the captured image 910, and a first region 331 and a second region 332 set for the captured image 910. In the captured image 910 shown in FIG. 9, players 311, 312, 914 and a goal 304 to be a background region are shown. The candidate regions 321, 322, 924 are candidate regions corresponding to the players 311, 322, 914. A candidate region that includes at least a part of the first region 331 and includes at least a part of the second region 332 is, for example, a candidate region such as the candidate region 924. Further, a candidate region that includes at least a part of the first region 331 and does not include at least a part of the second region 332 is, for example, a candidate region such as the candidate region 322. Further, a candidate region that does not include at least a part of the first region 331 and includes at least a part of the second region 332 is, for example, a candidate region such as the candidate region 321.

[0066] When an input image corresponding to a candidate region such as candidate region 924 is input, in at least one of the first and second learned models 151, 152, accurate inference may not be performed. The accuracy of inference can be evaluated by the reliability of the inference. Therefore, by inputting an image corresponding to a candidate region such as candidate region 924 into the first and second learned models 151, 152 and comparing the reliabilities indicated by the reliability information output by each, a foreground region with higher inference accuracy can be selected. Details of the reliability will be described later.

[0067] The foreground acquisition unit 160a acquires information (foreground information) indicating the foreground region in the captured image 910 that is output by the inference unit 150a as an inference result, that is, output by the first and second learned models 151, 152 as an inference result. When the input image output by the generation unit 140a is input to either the first or second learned model 151, 152, the foreground acquisition unit 160a acquires the foreground information output by the first or second learned model 151, 152 to which the input image is input. Further, the foreground acquisition unit 160a acquires the foreground region in the input image based on the acquired foreground information.

[0068] On the other hand, when the input image output by the generation unit 140a is input to the first or second learned model 151, 152, the foreground acquisition unit 160a acquires the foreground information and reliability information output by the first and second learned models 151, 152 as inference results. In this case, the foreground acquisition unit 160a selects either one of the foreground information output by the first or second learned model 151, 152 based on the reliability indicated by the reliability information output by the first and second learned models 151, 152. Specifically, in this case, the foreground acquisition unit 160a selects the foreground information with a high reliability among the foreground information output by the first and second learned models 151, 152, and acquires the foreground region in the input image based on the selected foreground information. The output unit 170a outputs information indicating the foreground region acquired by the foreground acquisition unit 160a.

[0069] Referring to FIG. 10, the operation of the image processing apparatus 100a will be described. FIG. 10 is a flowchart showing an example of a processing flow in the image processing apparatus 100 (image processing apparatus 100a) according to the second embodiment. Note that the flowchart shown in FIG. 10 is obtained by adding the processes from S1014 to S1017 to the flowchart shown in FIG. 8. In the flowchart shown in FIG. 10, the same processes as those in the flowchart shown in FIG. 8 are denoted by the same reference numerals and the description thereof is omitted.

[0070] First, the image processing apparatus 100a executes the processes from S801 to S814. If it is determined at S814 that at least a part of the candidate region does not include at least a part of the second region 332, the image processing apparatus 100a executes the processes from S818 to S820. If it is determined at S814 that at least a part of the candidate region includes at least a part of the second region 332, at S1014, the generation unit 140a determines whether at least a part of the selected candidate region includes at least a part of the first region 331. If it is determined at S1014 that at least a part of the candidate region selected by the generation unit 140a does not include at least a part of the first region 331, the image processing apparatus 100a executes the processes from S815 to S817. If it is determined at S1014 that at least a part of the candidate region selected by the generation unit 140a includes at least a part of the first region 331, the image processing apparatus 100a executes the processes from S1015 to S1017.

[0071] Specifically, in S1015, the generation unit 140a cuts out the image corresponding to the candidate region from the captured image 910 and outputs it to the inference unit 150a. Further, the inference unit 150a inputs the image as an explanatory variable to the first and second learned models 151 and 152. After S1015, in S1016, each of the first and second learned models 151 and 152 in the inference unit 150a infers the foreground region from the image and outputs the foreground information and the confidence information as the inference result. After S1016, in S1017, the foreground acquisition unit 160a acquires the two pieces of foreground information and the confidence information output from the inference unit 150a, and selects one of the two pieces of foreground information based on the confidence level indicated by each confidence information. Further, the foreground acquisition unit 160a acquires the foreground region of the image based on the selected foreground information. After S817, S820, or S1017, in S821, the generation unit 140a determines whether all candidate regions have been selected.

[0072] Referring to FIG. 11, an example of a method for calculating the confidence level in the first and second learned models 151 and 152 will be described. FIG. 11 is a diagram showing an example of a mask image 1101 indicating the foreground region inferred in the first or second learned model 151 or 152, and an example of an evaluation value group 1102 corresponding to a partial pixel group of the mask image 1101. Each evaluation value in the evaluation value group 1102 shown in FIG. 11 is a numerical value obtained by normalizing the evaluation of whether the corresponding pixel belongs to the foreground region in 100 levels from 0 to 99.

[0073] When inferring the foreground region in the input image input as an explanatory variable, the first and second learned models 151 and 152 determine whether each pixel in the image corresponds to the foreground region by calculating an evaluation value corresponding to each pixel. For example, when the calculated evaluation value is 80 or more, the pixel corresponding to the evaluation value is regarded as corresponding to the foreground region, and the pixel is set as a white pixel in the mask image 1101. Similarly, when the evaluation value is less than 80, the pixel corresponding to the evaluation value is regarded as corresponding to the background region, and the pixel is set as a black pixel in the mask image 1101. The first and second learned models 151 and 152 estimate the foreground region in the input image by determining whether all the pixels in the input image correspond to the foreground region or the background region, and output the mask image 1101 as foreground information. Note that the evaluation value is not limited to being normalized to 100 levels, and may be at a level of 99 or less or 101 or more. Also, the threshold value of the evaluation value for the foreground region or the background region is not limited to 80, and may be less than 80 or greater than 80.

[0074] The first and second learned models 151 and 152 calculate, for example, the average value of the evaluation values corresponding to all the pixels regarded as the foreground region as a numerical value indicating the reliability. A higher numerical value indicating the reliability indicates higher reliability of the extraction of the extracted foreground region. Note that using the average value of the evaluation values corresponding to all the pixels regarded as the foreground region as the reliability is an example, and the reliability may be a statistical value such as the median or the mode of the evaluation values corresponding to all the pixels regarded as the foreground region.

[0075] As described above, when the candidate region straddles the first region 331 and the second region 332, the foreground region is estimated by both the first learned model 151 and the second learned model 152, and based on the reliability of the estimation, one of the foreground regions is selected. Thereby, even when the candidate region straddles the first region 331 and the second region 332, the image processing apparatus 100a can accurately acquire the foreground region.

[0076] [Other Embodiments] The present disclosure can also be implemented by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. It can also be implemented by a circuit (for example, ASIC) that realizes one or more functions.

[0077] Note that within the scope of the present disclosure, any combination of the embodiments, any modification of any component of each embodiment, or any omission of any component in each embodiment is possible.

Description of Reference Numerals

[0078] 100 Image processing apparatus 110 Image acquisition unit 120 Setting unit 130 Candidate extraction unit 140 Generation unit 150 Inference unit 160 Foreground acquisition unit 170 Output unit 10 Imaging device 11 Image computing server

Claims

1. an image acquisition means for acquiring a captured image; a setting means for setting a first area that does not include an image area in which a specific subject appears in the captured image and a second area that includes the image area in which the specific subject appears in the captured image; a foreground acquisition means for acquiring a first foreground area indicating a foreground area included in the first area, which is extracted by a first learned model based on the captured image and the first area, and a second foreground area indicating a foreground area included in the second area, which is extracted by a second learned model based on the captured image and the second area; having the extraction accuracy of the second foreground area extracted by the second learned model based on the captured image and the second area is higher than the extraction accuracy of the second foreground area extracted by the first learned model based on the captured image and the second area An image processing apparatus characterized by the above.

2. a candidate extraction means for extracting a candidate area that is a candidate for the foreground area from the captured image; further comprising the first learned model extracts the first foreground area based on the candidate area in addition to the captured image and the first area, the second learned model extracts the second foreground area based on the candidate area in addition to the captured image and the second area The image processing apparatus according to claim 1, characterized by the above.

3. a generation means for generating a first image for input to the first learned model based on the captured image and the first area, and generating a second image for input to the second learned model based on the captured image and the second area; further comprising the first learned model extracts the first foreground area based on the first image based on the captured image and the first area, the second learned model extracts the second foreground area based on the second image based on the captured image and the second area The image processing apparatus according to claim 1, characterized by the above.

4. Generating means for generating a first image for inputting to the first pre-trained model based on the captured image, the first region, and the candidate region, and generating a second image for inputting to the second pre-trained model based on the captured image, the second region, and the candidate region, wherein the generating means generates one or more of the first images for inputting to the first pre-trained model by cutting out regions corresponding to each of one or more of the candidate regions including at least a part of the first region from the captured image, and generates one or more of the second images for inputting to the second pre-trained model by cutting out regions corresponding to each of one or more of the candidate regions including at least a part of the second region from the captured image, further comprising the first pre-trained model extracts the first foreground region based on the first image based on the captured image, the first region, and the candidate region, the second pre-trained model extracts the second foreground region based on the second image based on the captured image, the second region, and the candidate region The image processing apparatus according to claim 2, characterized in that

5. In addition to extracting the foreground region in the image corresponding to the candidate region, the first pre-trained model and the second pre-trained model acquire information indicating the reliability of the foreground region, For the candidate region that includes at least a part of the first region and at least a part of the second region among the candidate regions, the image corresponding to the candidate region is input to the first pre-trained model and the second pre-trained model, The foreground acquisition means acquires, in addition to the foreground region in the image corresponding to the candidate region extracted by the first pre-trained model and the second pre-trained model, information indicating the reliability of the foreground region, and when the image corresponding to the same candidate region is input to the first pre-trained model and the second pre-trained model, based on the reliability, selects one of the foreground regions extracted by the first pre-trained model or the second pre-trained model, and acquires the selected foreground region as the foreground region of the image corresponding to the candidate region The image processing apparatus according to claim 4, characterized in that

6. Inference means comprising the first pre-trained model and the second pre-trained model, further comprising The generation means outputs the generated first image and the second image to the inference means. The inference means inputs the first image to the first pre-trained model and inputs the second image to the second pre-trained model. The foreground acquisition means acquires the first foreground region extracted by the first pre-trained model of the inference means and the second foreground region extracted by the second pre-trained model. The image processing apparatus according to any one of claims 3 to 5, characterized by the above.

7. The first pre-trained model corresponds to a learning result obtained by learning using an image group composed of a plurality of learning images as first learning data. The second pre-trained model corresponds to a learning result obtained by learning using an image group composed of a plurality of learning images as second learning data. The image processing apparatus according to any one of claims 1 to 6, characterized in that the first learning data and the second learning data are different image groups from each other.

8. The image group that is the second learning data includes more learning images in which the specific subject is captured than the image group that is the first learning data. The image processing apparatus according to claim 7, characterized by the above.

9. In the image group that is the second learning data, the ratio of the number of learning images in which the specific subject is captured is larger than a predetermined ratio. The image processing apparatus according to claim 7 or 8, characterized by the above.

10. When the second pre-trained model is trained using the learning image in which the specific subject is captured as learning data in the second learning model corresponding to the second pre-trained model, compared with the case where the learning image in which the specific subject is captured is used as learning data in the first learning model corresponding to the first pre-trained model, the second pre-trained model is a pre-trained model corresponding to a learning result obtained by training so that the value of the loss function becomes smaller. The image processing apparatus according to any one of claims 7 to 9, characterized by the above.

11. Both the first pre-trained model and the second pre-trained model are pre-trained models configured by a neural network. The image processing apparatus according to any one of claims 1 to 10, characterized by the above.

12. The second pre-trained model has more intermediate layers or more neurons in the intermediate layer than the first pre-trained model. The image processing apparatus according to claim 11, characterized by the above.

13. When the specific subject appears in the captured image, the setting means sets the entire area of the captured image as the second region, and when the specific subject does not appear in the captured image, the setting means sets the entire area of the captured image as the first region. The image processing apparatus according to any one of claims 1 to 12, characterized in that.

14. When the specific subject appears in the captured image and the ratio of the area of the image region corresponding to the specific subject to the area of the entire captured image is greater than a predetermined value, the setting means sets the entire area of the captured image as the second region. When the specific subject does not appear in the captured image, or when the specific subject appears in the captured image and the ratio of the area of the image region corresponding to the specific subject to the area of the entire captured image is less than or equal to the predetermined value, the setting means sets the entire area of the captured image as the first region. The image processing apparatus according to any one of claims 1 to 12, characterized in that.

15. A plurality of the image processing apparatuses according to claim 6, One or more cooling devices for cooling the image processing apparatus, An image processing system comprising: The cooling device cools the image processing apparatus from which the second foreground region is extracted by the second learned model of the inference means among the plurality of image processing apparatuses. An image processing system characterized in that.

16. A program for causing a computer to operate as each means of the image processing apparatus according to any one of claims 1 to 14.

17. An image acquisition step in which an image processing apparatus acquires a captured image, A setting step in which the image processing apparatus sets a first region that does not include an image region in which a specific subject appears in the captured image and a second region that includes an image region in which the specific subject appears in the captured image, A foreground acquisition step in which the image processing apparatus acquires a first foreground region indicating a foreground region included in the first region, which is extracted by a first learned model based on the captured image and the first region, and a second foreground region indicating a foreground region included in the second region, which is extracted by a second learned model based on the captured image and the second region, Including, The extraction accuracy of the second foreground region extracted by the second learned model based on the captured image and the second region is higher than the extraction accuracy of the second foreground region extracted by the first learned model based on the captured image and the second region. An image processing method characterized by

18. An image group acquisition means for acquiring an image group composed of a plurality of learning images, A model acquisition means for acquiring a learning model, A learning means for generating a learned model that extracts a foreground region in an input image by training the learning model using the image group as learning data, A model output means for outputting the learned model, having The learning means generates a first learned model using the first image group acquired by the image group acquisition means as the learning data, and generates a second learned model using the second image group acquired by the image group acquisition means, which includes more learning images in which a specific subject is captured than the first image group, as the learning data A learning device characterized by

19. A program for causing a computer to operate as each means of the learning device according to Claim 18.

20. An image processing apparatus includes an image group acquisition step of acquiring an image group composed of a plurality of learning images, a model acquisition step in which the image processing apparatus acquires a learning model, a learning step in which the image processing apparatus trains the learning model using the image group as learning data to generate a learned model that extracts a foreground region in an input image, a model output step in which the image processing apparatus outputs the learned model, including In the learning step, a first learned model is generated using the first image group acquired in the image group acquisition step as the learning data, and a second learned model is generated using the second image group acquired in the image group acquisition step, which includes more learning images in which a specific subject is captured than the first image group, as the learning data A learning method characterized by

Citation Information

Patent Citations

  • Image processing device, image processing method, and program

    JP2020129276A

  • Image processing device, image processing method, and program

    JP2021056960A

  • Deep salient content neural networks for efficient digital object segmentation

    US20190130229A1