Image processing system and image processing method
The image processing system addresses the challenges of undetected objects and false detections by employing a dual approach of learned models and background subtraction methods, enabling high-precision three-dimensional shape data generation while reducing costs.
Patent Information
- Application Number
- JP2023107741
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2043-06-30
AI Technical Summary
Existing techniques for generating silhouette images to create three-dimensional shape data face challenges such as undetected objects and false detections, particularly in regions with stationary objects, similar background colors, and areas with shadows or virtual images, leading to incomplete or inaccurate three-dimensional shape data.
The proposed image processing system employs a dual approach by using a learned model for one imaging device to generate silhouette images without a background image, and another imaging device uses the background subtraction method to generate silhouette images based on a background image, allowing for high-precision three-dimensional shape data generation while reducing the mounting cost of image processing apparatuses.
This system effectively generates high-precision three-dimensional shape data while minimizing the cost of image processing apparatuses by selectively using learned models and background subtraction methods based on the imaging device's viewing angle and specific image processing requirements.
Smart Images

Figure 0007696956000001 
Figure 0007696956000002 
Figure 0007696956000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for generating a silhouette image indicating a foreground region in an image.
Background Art
[0002] There is a technique for generating three-dimensional shape data indicating the three-dimensional shape of an object (hereinafter simply referred to as an "object") using data of a plurality of images (hereinafter referred to as "imaging images") obtained by synchronized imaging by an imaging device. The three-dimensional shape data is generated by a method such as a volume intersection method using a silhouette image generated by extracting a region corresponding to the object (hereinafter referred to as an "object region") from each imaging image. As a method for generating a silhouette image from an imaging image, for example, there are a background difference method and a machine learning method. In the background difference method, in the viewing angle imaged by a certain imaging device, an imaging image obtained by imaging during a period when no object exists is used as a background image, and a difference is taken between this background image and an imaging image obtained by imaging during a period when the object exists. Further, based on this difference, the object region and the background region in the imaging image are separated to generate a silhouette image indicating the object region (hereinafter referred to as a "silhouette image of the object"). In the machine learning method, first, a sufficient number of learning data pairs are prepared, where each pair consists of an imaging image obtained by imaging during a period when an object exists and data indicating the object region in the imaging image, which is teacher data. Subsequently, a learned model, which is the result of training a learning model using the learning data, is generated. Further, using this learned model, the object region and the background region in the imaging image are separated to generate a silhouette image of the object.
[0003] Depending on the characteristics included in the angle of view of the imaging device, a suitable separation method for separating the object region and the background region in the captured image is different. When using the background subtraction method as the separation method, in the following regions, the object region may not be detected from the captured image (hereinafter referred to as "undetected object"). For example, a region where an object with little movement exists, a region where a fixed background with a color similar to that of the object exists, and a region where there is movement and a background with a color similar to that of the object exists in part of it. Also, in the following regions, a region other than the object region may be erroneously detected as the foreground region from the captured image (hereinafter referred to as "false detection of object"). For example, a region where a shadow of an object is generated, and a region where a virtual image is generated due to reflection of the image of the object on a shiny floor or a wet field surface, etc.
[0004] These undetected or falsely detected objects can cause missing or remaining parts of the three-dimensional shape data. Therefore, for regions where undetected or false detection may occur in the separation process by the background subtraction method, it is useful to perform the separation process by the machine learning method. On the other hand, generally, in the separation process by the machine learning method, the computational cost becomes higher than that in the separation process by the background subtraction method. Also, in the separation process by the machine learning method, the accuracy of the boundary of the object region in the silhouette image deteriorates more than that in the separation process by the background subtraction method because the shape of the object is inferred from the statistical information of the surrounding pixels. Patent Document 1 discloses a technique for sharing a region where separation processing is performed by the background subtraction method and a region where separation processing is performed by the machine learning method within the angle of view of a certain imaging device.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, in the technique disclosed in Patent Document 1, the image processing apparatus corresponding to one imaging device needs to be provided with a configuration for performing both the separation process by the background difference method and the separation process by the machine learning method, and the mounting cost for one image processing apparatus increases.
Means for Solving the Problems
[0007] The image processing system according to the present disclosure inputs the data of the image obtained by imaging an object by the first imaging means as the data of the first input image into a learned model, thereby generating the data of the first silhouette image indicating the region where the object exists in the first input image. A first image processing apparatus, and the data of the image obtained by imaging the object by a second imaging means different from the first imaging means is used as the data of the second input image, and the second input image and the second imaging means are used. A second image processing apparatus that generates data of a second silhouette image indicating a region where the object exists in the second input image based on a difference from a background image that is an image obtained by imaging by the second imaging means in a state where the object does not exist in the region to be imaged, The first image processing device does not have a configuration for generating a silhouette image indicating a region where the object exists based on a difference between the first input image and a background image corresponding to the first input image. The first image processing apparatus generates the first silhouette image without inputting a background image into a learned model, The second image processing device does not have a configuration for generating a silhouette image indicating a region where the object exists using a learned model. The second image processing apparatus generates the second silhouette image without using a learned model.
Advantages of the Invention
[0008] According to the present disclosure, it is possible to generate a silhouette image capable of generating high-precision three-dimensional shape data while suppressing the mounting cost on the image processing apparatus.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0010] Hereinafter, embodiments for carrying out the present disclosure will be described with reference to the drawings. The following embodiments do not limit the present disclosure, and not all combinations of the features described in the present embodiments are essential for the solution means of the present disclosure. The same components will be described with the same reference numerals. Also, for terms that differ only in the alphabet attached after the number in the reference numeral, they are assumed to be devices having the same function and indicating different devices from each other. For example, the first imaging unit 110A and the first imaging unit 110B shown in FIG. 1 indicate that they are different devices having the same function. Note that having the same function means having at least a specific function such as an imaging function with each other. For example, some of the functions and performances of the first imaging unit 110A and the first imaging unit 110B may be different from each other.
[0011] <Embodiment 1> In this embodiment, in the separation process of the region (foreground region) corresponding to the object in the image and the background region by the background subtraction method, a case will be described where a region where the foreground region may not be detected exists within the imaging angle of view of the imaging device.
[0012] [Configuration of the System] FIG. 1 is a block diagram showing an example of the functional configuration of an image processing system 100 according to Embodiment 1. The image processing system 100 generates three-dimensional shape data (hereinafter referred to as "three-dimensional shape data of the object") 160 indicating the three-dimensional shape of the object. The first imaging unit 110 images the imaging target, acquires the data of the captured image obtained by the imaging, and outputs the acquired data of the captured image to the first image processing unit 120. The first image processing unit 120 receives the data of the captured image output by the first imaging unit 110, separates the region (object region) corresponding to the object to be generated with the three-dimensional shape data in the captured image and the background region, and generates a silhouette image of the object. The data of the silhouette image generated by the first image processing unit 120 is transmitted to the third image processing unit 150.
[0013] The first image processing unit 120 includes a first input unit 121, a first separation unit 122, and a transmission unit 124. The first input unit 121 receives the data of the captured image output by the first imaging unit 110 and stores the received data of the captured image in a storage device provided inside the first image processing unit 120. The first separation unit 122 separates the foreground region, which is the object region included in the captured image, and the background region, and generates a silhouette image of the object. Specifically, the first separation unit 122 performs separation processing by a machine learning method to... To separate the object... ToSeparate the area to generate a silhouette image. More specifically, the first separation unit 122 holds a learned model 123 inside. The learned model 123 is a model learned to output data of an object's silhouette image with the data of a captured image including an object area as input. The learned model 123 separates the foreground area and the background area in the captured image to generate a silhouette image, for example, by performing a Semantic Segmentation task of classifying each pixel of the input image.
[0014] As implementation methods of the learned model 123, there are various methods using CNN (Convolutional Neural Network), and there are a plurality of network structures using an Encoder layer, a Decorder layer, and a Skip structure. For example, the learned model 123 is realized by using a structure called SegNet or Unet. The data of the generated silhouette image and the data of the texture image holding the color information of the object area in the silhouette image are output to the transmission unit 124. The transmission unit 124 receives the data of the silhouette image and the texture image output from the first separation unit 122 and transmits them to the third image processing unit 150. The transmission is performed via a communication I / F (interface) such as a LAN (Local Area Network) or a WAN (Wide Area Network).
[0015] The second imaging unit 130 captures an imaging target, acquires data of a captured image obtained by the imaging, and outputs the acquired data of the captured image to the second image processing unit 140. The second image processing unit 140 receives the data of the captured image output by the second imaging unit 130, separates the area (object area) corresponding to the object that is the generation target of the three-dimensional shape data in the captured image and the background area, and generates a silhouette image of the object. The data of the silhouette image generated by the second imaging unit 130 is transmitted to the third image processing unit 150.
[0016] The second imaging unit 130 includes a second input unit 141, a second separation unit 142, and a transmission unit 144. The second input unit 141 stores the data of the received captured image in a storage device provided inside the second image processing unit 140. The second separation unit 142 separates the object region and the background region included in the captured image to generate a silhouette image of the object. Specifically, the second separation unit 142 performs separation processing by the background difference method to separate the object region and the background region included in the captured image and generate a silhouette image of the object. More specifically, for example, the second separation unit 142 calculates the difference between the background image 143, which is a captured image obtained in a state where no object to be captured exists at the viewing angle of the second imaging unit 130, and the captured image obtained in a state where an object exists. Further, the second separation unit 142 generates a silhouette image of the object by separating the object region and the background region in the captured image based on the calculated difference.
[0017] As a method for obtaining the background image 143 used in the separation processing by the background difference method, there is a method of using a captured image at the moment when no object exists as the background image. Also, a method of generating a background image using a plurality of captured images obtained over a certain period may be used. Specifically, for example, the change in pixel values in each captured image is observed in units of pixels or small regions, and when the change in pixel values is within a certain amount, the average value or the latest pixel value of the pixels within a certain period is used to generate the background image. The data of the silhouette image generated by the second separation unit 142 and the data of the texture image holding the color information of the object region in the silhouette image are output to the transmission unit 144. The transmission unit 144 has the same function as the transmission unit 124, receives the data of the silhouette image and the texture image output from the second separation unit 142, and transmits these data to the third image processing unit 150.
[0018] The third image processing unit 150 generates three-dimensional shape data 160 of an object using the data of the silhouette image and the texture image received from the first image processing unit 120 and the second image processing unit 140. The third image processing unit 150 includes a receiving unit 151 and a shape generation unit 152. The receiving unit 151 receives the data of the silhouette image and the texture image transmitted from the first image processing unit 120 and the second image processing unit 140, and stores these data in a storage device provided inside the third image processing unit 150.
[0019] The shape generation unit 152 performs a shape estimation process and a coloring process for generating three-dimensional shape data 160 using the data of the silhouette image and the texture image received by the receiving unit 151. Specifically, the shape generation unit 152 first executes a shape estimation process, and then executes a coloring process to generate three-dimensional shape data 160. As the shape estimation process, for example, the volume intersection method can be used. For example, in the volume intersection method, first, rectangular parallelepipeds of unit volume called voxels are spread over the generation target space of the three-dimensional shape data. Hereinafter, the set of voxels spread over the generation target space of the three-dimensional shape data is called a voxel group. Subsequently, each silhouette image generated by performing a separation process on the captured image of the first or second imaging unit 110, 130 is projected onto the voxel group in the imaging range of the first or second imaging unit 110, 130. Subsequently, using all the silhouette images, voxels that are not included in the region where the object region in each silhouette image is projected are removed from this voxel group, thereby generating uncolored three-dimensional shape data. Subsequently, as the coloring process, the data of the texture image corresponding to each silhouette image is used as color information, and the three-dimensional shape data is colored by pasting a texture on each voxel of the uncolored three-dimensional shape data generated by the shape estimation process. By performing the above processes, the shape generation unit 152 generates three-dimensional shape data 160.
[0020] Here, cases where the shape estimation process fails will be described. One case is, for example, a state where the silhouette image of an object does not include the region corresponding to the object (object region), that is, a case where the object is not detected. Another case is, for example, a case where a region other than the region corresponding to the object is included as a foreground region in the silhouette image of the object, that is, a case where the object is misdetected.
[0021] When the object is not detected, the voxels corresponding to the object will be erroneously deleted from the voxel group. Therefore, the object should not be undetected in the captured images of all the first and second imaging units 110 and 130. When the object is misdetected, three-dimensional shape data of a false shape different from the shape of the object will be generated only when the image regions corresponding to the same voxels are misdetected in the captured images of all the first and second imaging units 110 and 130. Therefore, in the case of misdetection of the object, it is sufficient if the misdetection can be suppressed in the captured image of at least one of the first or second imaging units 110 and 130 that images the imaging target space corresponding to the space where the voxels constituting the false shape exist.
[0022] Note that the three-dimensional shape data 160 generated by the above-described process becomes a three-dimensional point cloud that is a set of colored voxels, but the form of the three-dimensional shape data is not limited to this. For example, the three-dimensional shape data 160 may be three-dimensional shape data in the form of a three-dimensional polygon mesh generated from the colored three-dimensional point cloud. In this case, for example, the shape generation unit 152 may include a process of generating data of a three-dimensional polygon mesh from the generated three-dimensional point cloud as the three-dimensional shape data.
[0023] FIG. 2 is a block diagram showing an example of the hardware configuration of the first image processing unit 120, the second image processing unit 140, and the third image processing unit 150. The hardware configurations of the first image processing unit 120, the second image processing unit 140, and the third image processing unit 150 (hereinafter collectively referred to as the "image processing apparatus 200") are the same as each other. The image processing apparatus 200 includes a CPU 201, a ROM 202, a RAM 203, an auxiliary storage device 204, a display unit 205, an operation unit 206, a communication I / F 207, and a bus 208.
[0024] The CPU 201 controls the entire image processing apparatus 200 using computer programs and data stored in the ROM 202 or the RAM 203, and realizes each unit that the image processing apparatus 200 has as a functional configuration. Note that the image processing apparatus 200 may have one or more dedicated hardware different from the CPU 201, and at least a part of the processing by the CPU 201 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), and a DSP (digital signal processor). The ROM 202 stores programs and the like that do not require modification. The RAM 203 is used as a work space for the CPU 201, and temporarily stores programs or data supplied from the auxiliary storage device 204 and data etc. supplied from the outside via the communication I / F 207. The auxiliary storage device 204 is constituted by a large-capacity storage device such as a hard disk drive, and stores various data etc. The RAM 203 and the auxiliary storage device 204 hold the input data, intermediate data during processing, and output data of the image processing apparatus 200.
[0025] In the RAM 203 or the auxiliary storage device 204 of the first image processing unit 120, data of the captured image received from the first imaging unit 110, the learned model 123 used in the first separation unit 122, and intermediate data in the separation process using the learned model 123 are retained. Also, in the RAM 203 or the auxiliary storage device 204 of the first image processing unit 120, data of the silhouette image and the texture image generated by the first separation unit 122 are retained. In the RAM 203 or the auxiliary storage device 204 of the second image processing unit 140, data of the captured image received from the second imaging unit 130, data of the background image 143 used by the second separation unit 142, and intermediate data for generating the background image 143 are retained. Also, in the RAM 203 or the auxiliary storage device 204 of the second image processing unit 140, data of the silhouette image and the texture image generated by the second separation unit 142 are retained. In the RAM 203 or the auxiliary storage device 204 of the third image processing unit 150, data of the silhouette image and the texture image received from the first image processing unit 120 and the second image processing unit 140 are retained. Also, in the RAM 203 or the auxiliary storage device 204 of the third image processing unit 150, intermediate data in the shape estimation process in the shape generation unit 152, and the uncolored three-dimensional shape data after the shape estimation process are retained.
[0026] The display unit 205 is composed of a liquid crystal display or an LED (light-emitting diode), etc., and a GUI (Graphical User Interface) etc. for the user to operate the image processing apparatus 200 is displayed. The operation unit 206 is composed of a keyboard, a mouse, a joystick, or a touch panel, etc., and receives an operation by the user and inputs various instructions to the CPU 201. The CPU 201 operates as a display control unit that controls the display unit 205 and an operation control unit that controls the operation unit 206. The communication I / F 207 is used for communication between the image processing apparatus 200 and an external device.
[0027] The first image processing unit 120 has, as a communication I / F 207, an I / F for receiving data of a captured image from the first imaging unit 110, and an I / F for transmitting data of a silhouette image and a texture image to the third image processing unit 150. The second image processing unit 140 has, as a communication I / F 207, an I / F for receiving data of a captured image from the second imaging unit 130, and an I / F for transmitting data of a silhouette image and a texture image to the third image processing unit 150. The third image processing unit 150 has, as a communication I / F 207, an I / F for receiving data transmitted from the first image processing unit 120 and the second image processing unit 140. The communication I / F 207 is realized by a wired network I / F such as Ethernet, a wireless network I / F such as a wireless LAN, or an SDI (Serial Digital Interface) I / F for transmitting and receiving video signals. The bus 208 connects each part that the image processing apparatus 200 has as a hardware configuration and transmits information. In the present embodiment, it is assumed that the display unit 205 and the operation unit 206 exist inside the image processing apparatus 200, but at least one of the display unit 205 and the operation unit 206 may exist outside the image processing apparatus 200 as another device.
[0028] [Application Example of the System] FIG. 3 is a diagram showing an application example 300 of the image processing system 100 according to Embodiment 1. The application example 300 shown in FIG. 3 shows an example in which the image processing system 100 is applied to an imaging target of a soccer stadium. In the imaging target space 301, there are a playing field 305 and an object 306 that is a generation target of three-dimensional shape data. In the case of soccer, the object 306 includes objects 306A to 306C that are players and an object 306D that is a soccer ball. Among these, the objects 306A and 306C are field players, and the object 306B is a goalkeeper. The object 306 may include a referee (not shown in FIG. 3). In the vicinity of the playing field 305, there are a bench 307 where players wait and a sign 308 that is painted on the floor of the soccer stadium or laid as a sign-shaped three-dimensional object. Also, around the playing field 305, there is a lighting device 309 that is lit when the illuminance is insufficient at night or in rainy weather.
[0029] The sign-shaped sign 308 also includes a digital sign that is displayed on a display and whose display content changes over time. The lighting device 309 may be a light source that is locally arranged so that a dark shadow of the object 306 is generated in a specific direction and irradiates strong light. Also, the lighting device 309 may be a light source that is arranged at regular intervals around the playing field 305 so that a faint shadow of the object 306 is evenly generated around it. Also, on the floor surface of the playing field 305, when the field surface is wet due to rain or the like, a virtual image of the object 306 may be generated by specular reflection. Note that the processing of the image processing system 100 when a shadow or virtual image of the object 306 occurs will be described in Embodiment 2.
[0030] Around the imaging target space 301, imaging devices 302A to 302P, which are the first imaging unit 110 or the second imaging unit 130, are arranged. Each imaging device 302 images at least a part of the imaging target space 301. FIG. 3 shows an aspect in which the imaging devices 302 are arranged so as to surround the periphery of the imaging target space 301 on a plan view looking down on the imaging target space 301. However, the imaging devices 302 are also arranged with variations in the elevation angle of the imaging devices 302 with respect to the imaging target space 301. With such an arrangement of the imaging devices 302, it becomes possible to image the object 306 from various angles, and it becomes possible to generate high-quality three-dimensional shape data.
[0031] Also, when a shadow or virtual image of the object 306 described in Embodiment 2 occurs, each imaging device 302 is determined to be the first imaging unit 110 or the second imaging unit 130 according to the elevation angle of the imaging device 302. The imaging device 302 outputs the data of the captured image to an imaging device control box 303, which is the first image processing unit 120 or the second image processing unit 140. The imaging device control box 303 receives the data of the captured image, and when there is a region corresponding to the object 306 (hereinafter referred to as the "region of the object 306") in the captured image, it generates the data of the silhouette image and the texture image of the object 306. Further, the imaging device control box 303 transmits the generated data of the silhouette image and the texture image to an image processing server 304, which is the third image processing unit 150. The image processing server 304 generates three-dimensional shape data using the data of the silhouette image and the texture image received from each imaging device control box 303.
[0032] Referring to FIGS. 4 and 5, a method for determining whether to use the imaging device 302 and the imaging device control box 303 as the first imaging unit 110 and the first image processing unit 120, or as the second imaging unit 130 and the second image processing unit 140 will be described. FIG. 4 is a flowchart showing an example of the flow of the determination process of the first or second imaging units 110, 130 and the first or second image processing units 120, 140 according to Embodiment 1. The determination of the first or second imaging units 110, 130 and the first or second image processing units 120, 140 is made before imaging using the image processing system 100 based on the viewing angle of each imaging device 302 when the imaging device 302 is installed. In the following description, the leading letter "S" of the reference numeral means "step". In S401, the user sets a specific area. In Embodiment 1, the specific area is an area where the object 306 may not be detected even though the area of the object 306 exists in the separation process of the area of the object 306 by the background difference method performed by the second image processing unit 140.
[0033] Referring to FIG. 5, the specific area will be described. FIG. 5 is a diagram showing an example of the viewing angle of the imaging device 302 according to Embodiment 1. FIG. 5(a) shows an example of the viewing angle 510 of the imaging device 302 that captures the objects 306E and 306G, which are field players, and the object 306F, which is a ball, existing in the competition field 305. In the viewing angle 510, since all the objects 306 move at a certain speed or more, the same object 306 does not continue to exist at a certain pixel for a certain period or more. Therefore, a background image without the object 306 can be generated. Here, it is assumed that there is a difference of a certain value or more between the color value of the object 306 and the pixel value of the background image in the viewing angle 510. In this case, the background image can be created, and the difference between the pixel value of the area of the object 306 in the captured image captured at the viewing angle 510 and the pixel value of the corresponding area in the background image is also a certain value or more. Therefore, the viewing angle 510 of the imaging device 302 is not considered to include the specific area.
[0034] FIG. 5(b) shows the viewing angle 520 of the imaging device 302 that captures the objects 306I and 306J, which are field players, the object 306H, which is a ball, and the object 306K, which is a goalkeeper, present in the competition field 305. Also, in the viewing angle 520 of this imaging device 302, there is also a sign 308E painted on the floor surface of the soccer stadium. FIG. 5(c) shows the viewing angle 521 that is the same as the viewing angle 520 of the imaging device 302 shown in FIG. 5(b), and shows the viewing angle 521 after a certain period of time has elapsed from the state of the viewing angle 520. In the viewing angle 521, there are the object 306 and the sign 308 that existed in the state of the viewing angle 520. Among these, for the object 306 that has moved, in FIG. 5(c), an apostrophe (’) is added to the end of its symbol. For example, the object 306I is the same field player as the object 306I’, indicating that the field player is moving.
[0035] The specific region in the viewing angle 520 (521) will be described. First, the object 306K, which is a goalkeeper, has hardly changed its position even after a certain period of time has elapsed. In the region in the captured image corresponding to the object 306 that is substantially stationary in this way, in the method for generating a background image that determines pixels with no change in pixel value over a certain period of time as the background region, it is determined to be a region that constitutes the background image. That is, the region of the object 306K enters the background image. As a result, when performing the separation process of the object region by the background difference method, in the region of the object 306K, no difference occurs between the captured image and the background image, and the object 306K is not detected. Therefore, the silhouette image of the object 306K cannot be generated. For a region where there may be an object 306 with little movement, that is, a substantially stationary object 306, like in the case of the object 306K which is a goalkeeper, the user sets it as a specific region.
[0036] On the one hand, since the signage 308 painted on the floor of the soccer stadium does not move, it is incorporated into the background image. Here, consider the case where the difference value between the pixel value of the area corresponding to the signage 308 in the captured image and the pixel value of the uniform of the player, which is the object 306, is smaller than a predetermined threshold value. That is, consider the case where the color of the signage 308 and the color of the uniform are similar colors. In this case, since the difference value between the pixel values of the player's uniform and the signage 308 does not exceed the predetermined threshold value, the non-detection of the object 306, which is the player, occurs, and the silhouette image indicating the area corresponding to the player is not generated. Thus, for the background area where a color similar to the color of the object 306, which is the generation target of the three-dimensional shape data, may exist in a part of the background image, the user sets it as a specific area.
[0037] FIG. 5(d) shows the viewing angle 530 of the imaging device 302 that captures the objects 306M and 306N, which are field players, the object 306L, which is a ball, and the bench 307C where substitute players wait, existing in the playing field 305. Here, it is assumed that neither the bench 307C nor the substitute players existing therein are the generation targets of the three-dimensional shape data. Since the substitute players in the bench 307C are natural persons, the substitute players have movements associated with the passage of time. When generating the background image at the viewing angle 530 of such an imaging device 302, a part of the area corresponding to the substitute players in the captured image does not become pixels that do not change for a certain period of time or more, and thus is not incorporated into the background image. When performing the separation process of the object area by the background difference method on the captured image using such a background image, a part of the area corresponding to the substitute players is separated as the foreground area, and the silhouette image indicating the foreground area is generated.
[0038] On the other hand, in the captured image, a part of the area corresponding to the stationary reserve players does not change for a certain period of time or more, so it is captured as part of the background image. In this case, there is no difference greater than or equal to a certain value between the pixel values of the area corresponding to the body or uniform of the reserve player captured in the background image and the pixel values of the area corresponding to the body or uniform of the field player, which is the object 306, in the captured image. Therefore, the object area cannot be correctly separated, undetected parts occur in a part of the object 306, and a silhouette image with a part of the area corresponding to the object 306 missing is generated. Thus, a stable background image cannot be generated, and for a background area where a color similar to the color of the object 306 may exist in a part of the background image, the user sets it as a specific area.
[0039] FIG. 5(e) shows the viewing angle 540 of the imaging device 302 that captures the objects 306P and 306Q, which are field players, and the object 306O, which is a ball, existing in the competition field 305. In addition, in the viewing angle 540 of this imaging device 302, there is a signage 308F, which is a billboard-shaped digital signage installed outside the competition field 305. FIG. 5(f) shows the viewing angle 541, which is the same as the viewing angle 540 of the imaging device 302 shown in FIG. 5(e), after a certain period of time has elapsed. In the viewing angle 541, there are the object 306 and the signage 308 that existed in the state of the viewing angle 540 of the imaging device 302. Among these, for the moving object 306 and signage 308, in FIG. 5(f), an apostrophe (') is added to the end of their symbols. For example, since the display content of the digital signage 308F is changing, it is denoted as 308F'.
[0040] Hereinafter, a specific area in the imaging angle 540 (541) will be described. Since the display content of the digital signage 308 changes, the pixels in the area corresponding to the signage 308 in the captured image will not be pixels with unchanged pixel values over a certain period of time, and this area will not be incorporated into the background image. Therefore, if the object area separation process by the background difference method is performed on the captured image in the state of imaging angle 540 and the captured image in the state of imaging angle 541, the area of the signage 308F will be erroneously separated as the foreground area. As a result, a silhouette image of the signage 308F, which is not a generation target of the three-dimensional shape data, will be generated.
[0041] On the other hand, if the pixel values of the area corresponding to the signage 308 in the captured image do not change over a certain period of time, this area will be incorporated into the background image. In this case, if the color used for the display content of the digital signage when it is incorporated into the background image is similar to the color of the object 306, the object 306 will not be detected. As a result, the silhouette image of the object 306 will not be generated. Such an area where a silhouette image of an object that is not a generation target of the three-dimensional shape data can be generated, and an area where the background image cannot be stably generated and the object 306 may not be detected, the user sets as a specific area.
[0042] To summarize the above, the specific area according to Embodiment 1 includes an area where an object with little movement may exist, and an area where a fixed background of a color similar to the color of the object may exist. Also, the specific area according to Embodiment 1 includes an area where there is movement and a background area of a color similar to the color of the object 306 may exist in a part of the background image. Note that the specific area is not limited to the above, and an area including a factor causing non-detection of the object during the object area separation process by the background difference method may be set as the specific area. By setting such a specific area, the user can apply the present disclosure even in cases other than the above-described cases described with reference to FIG. 5.
[0043] Return to the description of the flowchart shown in FIG. 4. After S401, the user determines, for all imaging devices 302, whether it is the first imaging unit 110 or the second imaging unit 130, and whether the imaging image data output by the imaging device 302 is to be processed by the first image processing unit 120 or the second image processing unit 140. Specifically, the user performs the steps from S403 to S405 for each imaging device 302. Specifically, for example, in the case of the application example 300 shown in FIG. 3, the user performs the steps from S403 to S405 for all imaging devices 302 from the imaging device 302A to the imaging device 302P. Note that S402 indicates the start of the loop of the steps from S403 to S405.
[0044] In S402, the user selects an arbitrary imaging device 302 from one or more imaging devices 302 that have not been selected so far. Next, in S403, the user determines whether the imaging angle of view of the imaging device 302 selected in S402 includes the specific area. If it is determined in S403 that the specific area is included, then in S404, the user determines that the imaging device 302 selected in S402 is the first imaging unit 110, and the imaging device control box 303 corresponding to the imaging device 302 is the first image processing unit 120. If it is determined in S403 that the specific area is not included, then in S405, the user determines that the imaging device 302 selected in S402 is the second imaging unit 130, and the imaging device control box 303 corresponding to the imaging device 302 is the second image processing unit 140.
[0045] S406 indicates the end of the loop of the steps from S403 to S405. At S406, the user determines whether all the imaging devices 302 were selected at S402. If, at S406, it is determined that not all the imaging devices 302 have been selected, that is, there are imaging devices 302 that have not been selected, the user returns to S402 and selects any of the imaging devices 302 that have not been selected so far. Thereafter, the steps from S402 to S406 are repeatedly performed until it is determined at S406 that all the imaging devices 302 have been selected. If it is determined at S406 that all the imaging devices 302 have been selected, the user ends the determination process shown in the flowchart of FIG. 4.
[0046] With reference to the viewing angles 510, 520, 530, and 540 shown in FIG. 5, a specific example of the method for determining the imaging device 302 and the imaging device control box 303 will be described. Since the viewing angle 510 of the imaging device 302 does not include a specific area, the imaging device 302 is the second imaging unit 130, and the imaging device control box 303 connected to the imaging device 302 is determined to be the second image processing unit 140. On the other hand, since the viewing angles 520, 530, and 540 of the imaging device 302 all include a specific area, the imaging device 302 is the first imaging unit 110, and the imaging device control box 303 connected to the imaging device 302 is determined to be the first image processing unit 120.
[0047] For example, first, the user installs all the imaging devices 302 at desired positions and adjusts the viewing angles of the respective imaging devices 302 to the desired viewing angles. Subsequently, based on the determination of the determination process shown in the flowchart of FIG. 4, the user installs the first image processing unit 120, for example, in the vicinity of the imaging device 302 determined to be the first imaging unit 110, and connects the first image processing unit 120 to the first imaging unit 110 and the third image processing unit 150. Similarly, the second image processing unit 140 is installed, for example, in the vicinity of the imaging device 302 determined to be the second imaging unit 130, and the second image processing unit 140 is connected to the second imaging unit 130 and the third image processing unit 150.
[0048] Referring to FIG. 6, the operations of the first image processing unit 120, the second image processing unit 140, and the third image processing unit 150 will be described. FIG. 6(a) is a flowchart showing an example of the processing flow of the first image processing unit 120 according to Embodiment 1. First, at S601, the first input unit 121 receives the data of the captured image output by the first imaging unit 110. Next, at S602, the first separation unit 122 inputs the data of the captured image received at S601 into the learned model 123. Next, at S603, the first separation unit 122 acquires the data of the silhouette image generated by the learned model 123. Next, at S604, the transmission unit 124 outputs the data of the silhouette image acquired at S603 to the third image processing unit 150. After S604, the first image processing unit 120 ends the processing of the flowchart shown in FIG. 6(a), and repeats the processing of the flowchart shown in FIG. 6(a) each time the first imaging unit 110 outputs the data of a new captured image.
[0049] FIG. 6(b) is a flowchart showing an example of the processing flow of the second image processing unit 140 according to Embodiment 1. First, at S611, the second input unit 141 receives the data of the captured image output by the second imaging unit 130. Next, at S612, the second separation unit 142 generates a silhouette image by performing separation processing by the background difference method on the data of the captured image received at S601 using the data of the background image. Next, at S613, the transmission unit 144 outputs the data of the silhouette image generated at S612 to the third image processing unit 150. After S613, the second image processing unit 140 ends the processing of the flowchart shown in FIG. 6(b), and repeats the processing of the flowchart shown in FIG. 6(b) each time the second imaging unit 130 outputs the data of a new captured image.
[0050] FIG. 6(c) is a flowchart showing an example of the processing flow of the third image processing unit 150 according to Embodiment 1. First, in S621, the receiving unit 151 receives data of the silhouette image and the texture image transmitted from the first image processing unit 120 and the second image processing unit 140. Here, the first imaging unit 110 and the second imaging unit 130 perform imaging synchronized with each other, for example, and the receiving unit 151 receives data of the silhouette image and the texture image based on the data of the captured image obtained by imaging at the same time. Note that the same time here is not limited to exactly the same time, but includes approximately the same time.
[0051] After S621, in S622, the shape generation unit 152 executes a shape estimation process of the three-dimensional shape data using the data of the silhouette image received in S621. Next, in S623, the shape generation unit 152 performs a coloring process on the three-dimensional shape data generated in S622 using the data of the texture image to generate three-dimensional shape data 160. After S623, the third image processing unit 150 ends the processing of the flowchart shown in FIG. 6(c). Thereafter, the third image processing unit 150 repeatedly executes the processing of the flowchart shown in FIG. 6(c) every time the first image processing unit 120 and the second image processing unit 140 output data of a new silhouette image and texture image.
[0052] As described above, the image processing system 100 is configured such that an area where undetected objects may occur in the background subtraction method is set as a specific area. Further, with respect to the data of the captured image output from the imaging device 302 whose viewing angle includes the specific area, the image processing system 100 is configured to input it to at least an image processing device that separates object areas using the learned model 123. Furthermore, with respect to the data of the captured image output from the imaging device 302 whose viewing angle does not include the specific area, the image processing system 100 is configured to input it to an image processing device that separates object areas by the background subtraction method. Moreover, the image processing system 100 is configured to determine an image processing device that appropriately separates object areas according to the viewing angle of the imaging device 302. According to the image processing system 100 configured as described above, it is not necessary for a single image processing device to have a configuration that performs both separation processing by the background subtraction method and separation processing by the machine learning method. Therefore, it is possible to generate a silhouette image that can generate high-precision three-dimensional shape data while suppressing the implementation cost of the image processing device.
[0053] <Embodiment 2> In this embodiment, in the separation process of the object area by the background subtraction method, a case will be described in which an area where there is a possibility of erroneously detecting an area outside the area of the object 306 as a foreground area exists within the viewing angle of the imaging device 302. Areas where there is a possibility of erroneous detection include areas where the shadow of the object 306 may occur, and areas where the image of the object 306 may be reflected on the surface of the competition field 305 to generate a virtual image. Since the functional configuration of the image processing system according to Embodiment 2 is the same as the functional configuration of the image processing system 100 shown in FIG. 1, the differences from Embodiment 1 will be described below.
[0054] [System Application Example] As an application example, similar to Embodiment 1, it will be described as being the same as Application Example 300 of the image processing system shown in FIG. 3. With reference to FIGS. 7 and 8, a method for determining whether to use the imaging device 302 and the imaging device control box 303 as the first imaging unit 110 and the first image processing unit 120, or as the second imaging unit 130 and the second image processing unit 140 will be described. FIG. 7 is a flowchart showing an example of the flow of the determination process of the first or second imaging unit 110, 130 and the first or second image processing unit 120, 140 according to Embodiment 2.
[0055] The determination process shown in the flowchart of FIG. 7 is performed before the processing of the image processing system is started, based on the angle of view determined at the time of installation of the imaging device 302 and the weather or lighting conditions when using the image processing system. It is assumed that the changes in the weather or lighting conditions and the changes in the shadow or virtual image of the object 306 that may occur in the angle of view of each imaging device 302 accordingly have been investigated in advance. Note that only whether to use a certain imaging device 302 as the first imaging unit 110 or the second imaging unit 130 is described in the flowchart of FIG. 7. However, for the imaging device control box 303, it is also determined as the first image processing unit 120 or the second image processing unit 140 in accordance with the determination of the first imaging unit 110 or the second imaging unit 130. Specifically, when it is determined that a certain imaging device 302 is the first imaging unit 110, the imaging device control box 303 corresponding to the imaging device 302 is determined to be the first image processing unit 120. On the other hand, when it is determined that a certain imaging device 302 is the second imaging unit 130, the imaging device control box 303 corresponding to the imaging device 302 is determined to be the second image processing unit 140.
[0056] First, in S701, the user sets a specific area. In Embodiment 2, the specific area is an area where a shadow or virtual image of the object 306 can occur. Next, in S702, the user determines whether a dark shadow is generated in a specific direction. If it is determined in S702 that a dark shadow is generated in a specific direction, then in S703, the user determines a part of one or more imaging devices 302 capable of imaging the area where the dark shadow can occur as the first imaging unit 110. Since each imaging device 302 covers the field without omission while having an overlapping area between the viewing angles, there are a plurality of imaging devices 302 whose viewing angles include the area where a dark shadow can occur. The user determines one or more of them as the first imaging unit 110.
[0057] FIG. 8 is a diagram for explaining an example of the shadow and virtual image of the object 306 according to Embodiment 2. FIG. 8(a) shows an example of a shadow 811 generated at the viewing angle 810 of the imaging device 302. In the example shown in FIG. 8(a), due to the strong light from a specific direction by the lighting device 309C, a shadow 811 of the object 306R is generated as a dark shadow on the surface of the competition field 305. FIG. 8(b) shows an example of an aerial image 812 obtained by virtually looking down on FIG. 8(a) from directly above. As shown in FIG. 8(b), when a dark shadow 811 is generated, the shadow 811 is generated in a specific direction with respect to the object 306. FIGS. 8(a) and (b) show an example where the shadow 811 is generated in one direction, but depending on the number and arrangement of the lighting devices, the shadow of the object 306 may be generated in two or more directions. However, when the shadow of the object 306 is generated in two or more directions, the subsequent shadow will not be a dark shadow. This is because the surface of the competition field 305 in the direction where there is a lighting device that irradiates strong light as viewed from the object 306 is illuminated by the light from that lighting device, and the shadow generated in the area in that direction will not be a dark shadow.
[0058] When it is presumed that a dark shadow 811 as shown in the bird's-eye view image 812 of FIG. 8(b) occurs, the user designates as the first imaging unit 110 a part of one or more imaging devices 302 existing in the direction in which the dark shadow 811 occurs as viewed from the object 306R. Since it suffices to be able to delete the voxels corresponding to the region where the dark shadow 811 occurs, if there is an imaging device 302 that captures the entire dark shadow 811 within its field of view, only the imaging device 302 may be designated as the first imaging unit 110 as part of the above. In actual imaging, it is rare that the conditions are met such that only one imaging device 302 can delete all the voxels corresponding to the region where a dark shadow may occur. In such a case, for the entire competition field 305, the number of imaging devices 302 sufficient to perform deletion of the voxels corresponding to the region where a dark shadow may occur is determined as the first imaging unit 110.
[0059] After S703, or when it is determined in S702 that no dark shadow occurs in a specific direction, in S704, the user determines whether a shadow of the object 306 is evenly generated around the object 306. When it is determined in S704 that a shadow is evenly generated around the object 306, the user performs the process of S705. Specifically, in S705, the user designates as the first imaging unit 110 a predetermined number of imaging devices 302 out of all the imaging devices 302 that can image at least a part of the region where a shadow can be evenly generated around the object 306. This is to suppress the remaining voxels due to the shadow evenly generated around the object 306.
[0060] FIG. 8(c) shows an example of a shadow 821 that evenly occurs around the object 306 at the angle of view 820 of the imaging device 302. In the example shown in FIG. 8(c), due to the light irradiated by a large number of lighting devices 309 existing around the competition field 305, a shadow 821 is generated around the object 306S. FIG. 8(d) shows an example of an aerial image 822 obtained by virtually looking down on FIG. 8(c) from directly above. As shown in FIG. 8(d), when a shadow 821 can evenly occur around the object 306, the user determines a predetermined number of imaging devices 302 among all the imaging devices 302 in which at least a part of the area where the shadow 821 can occur is included in the angle of view as the first imaging unit 110.
[0061] The predetermined number is determined, for example, based on how much the optical axes of the respective imaging devices 302 are set at an angular difference with respect to the object 306 or the imaging target space. The user determines a combination of a predetermined number of imaging devices 302 that can delete the voxels corresponding to the area of the shadow 821 generated around the object 306S, and determines the imaging devices 302 included in the determined combination as the first imaging unit 110. Specifically, for example, the user determines two or three imaging devices 302 selected such that the angular difference of the optical axes of the imaging devices 302 is 120 to 180 degrees among a plurality of imaging devices 302 that image a part of the imaging target space as the first imaging unit 110. According to the two or three imaging devices 302 selected in this way, the shadow 821 can be imaged without being blocked by the object 306S.
[0062] After S705, or when it is determined in S704 that no shadow evenly occurs around the object 306, in S706, the user determines whether a virtual image occurs around the object 306. When it is determined in S706 that a virtual image occurs around the object 306, in S707, the user determines a part of one or more imaging devices 302 that can image the area where the virtual image occurs as the first imaging unit 110.
[0063] FIG. 8(e) shows an example of a virtual image 831 of an object 306 generated at the viewing angle 830 of the imaging device 302. In the example shown in FIG. 8(e), the virtual image 831 is generated by the image of the object 306T being reflected on the surface of the competition field 305. FIG. 8(f) is a cross-sectional image 832 obtained by virtually viewing FIG. 8(e) from directly side-on. The virtual image of the object 306 is generated on the straight line where the plane determined by the straight line connecting the imaging device 302 and the object 306 and the vector orthogonal to the surface of the competition field 305 intersects the surface of the competition field 305. Therefore, the location where the virtual image is generated is different for each imaging device 302. For example, in the cross-sectional image 832, the virtual image generated by the image of the object 306T being reflected on the field surface 833 appears as the virtual image 834 in the imaging device 302Q and appears as the virtual image 835 in the imaging device 302R. Therefore, the false three-dimensional shape data due to the virtual image does not occur in the region corresponding to the region above the field surface 8 33.
[0064] However, when the generation target space of the three-dimensional shape data is set to the region corresponding to the region below the field surface 8 33, false three-dimensional shape data due to the virtual image occurs in the following space in the generation target space of the three-dimensional shape data. Specifically, in this case, false three-dimensional shape data due to the virtual image occurs in the space corresponding to the region that is symmetric to the region of the object 306 in the generation target space of the three-dimensional shape data with respect to the field surface 8 33. For example, in the cross-sectional image 832, the three-dimensional shape data 836 is generated as false three-dimensional shape data due to the virtual image. Therefore, the user determines a sufficient number of imaging devices 302 as the first imaging unit 110 in order to cut out the voxels of the false three-dimensional shape data generated in the space corresponding to the region below the field surface 8 33.
[0065] Unlike the shadow of the object 306, the virtual image of the object 306 is not blocked by the object 306. Therefore, basically, the user may determine any one of the plurality of imaging devices 302 that image the area where the virtual image can be generated as the first imaging unit 110 of the imaging device 302. However, due to the concentration of a plurality of objects 306, the virtual image of a certain object 306 may be hidden by another object 306. Considering such a case, the user may determine two or more of the plurality of imaging devices 302 that image the area where the virtual image can be generated as the first imaging unit 110.
[0066] After S707, or when it is determined in S706 that no virtual image is generated around the object 306, in S708, the user determines whether a shadow is generated in the area surrounded by the plurality of objects 306. If it is determined in S708 that a shadow is generated, in S709, the user determines a part of one or more imaging devices 302 whose elevation angle of the optical axis is equal to or greater than a predetermined angle as the first imaging unit 110. By determining the imaging device 302 whose elevation angle of the optical axis is equal to or greater than a predetermined angle as the first imaging unit 110, it is possible to suppress omission of the three-dimensional shape data due to the shadow generated in the area surrounded by the plurality of objects 306.
[0067] FIG. 8(g) shows an example of a shadow 841 generated in a region surrounded by an object 306 at the viewing angle 840 of the imaging device 302. In the example shown in FIG. 8(g), a shadow 841 is generated in a region surrounded by objects 306U, 306V, 306W, and 306X by the light irradiated by a large number of lighting devices 309 existing around the competition field 305. FIG. 8(h) shows an example of a cross-sectional image 842 obtained by virtually viewing FIG. 8(g) from the side. Voxels corresponding to the region where the shadow 841 as shown in FIGS. 8(g) and 8(h) exists can be removed, for example, using a silhouette image corresponding to the captured image of the imaging device 302 as follows. For example, as shown in the cross-sectional image 842, the imaging device 302 has an elevation angle of the optical axis larger than the elevation angle 843 of the tangent line drawn from the field surface at the center of the plurality of dense objects 306 to the upper part of the object 306U.
[0068] For example, the user sets such an elevation angle as an angle threshold based on the object 306 to be imaged or the characteristics of the competition, etc., and determines a part of one or more imaging devices 302 whose elevation angle of the optical axis is equal to or greater than this angle threshold as the first imaging unit 110. Specifically, the user selects two or three imaging devices 302 among the plurality of imaging devices 302 that satisfy the above-described conditions for imaging a part of the imaging target space, where the angular difference between the optical axes of the imaging devices 302 is 120 degrees to 180 degrees, in the same manner as in the case of the shadow 821 or the virtual image 831. Further, the user determines the selected two or three imaging devices 302 as the first imaging unit 110.
[0069] After S709, if it is determined in S708 that no shadow is generated, in S710, the user determines whether a shadow is generated below the object 306. If it is determined in S710 that a shadow is generated below the object 306, in S711, the user determines a part of one or more imaging devices 302 whose elevation angle is equal to or less than a predetermined angle as the first imaging unit 110. By determining a part of one or more imaging devices 302 whose elevation angle is equal to or less than a predetermined angle as the first imaging unit 110, it is possible to suppress the remaining uncut three-dimensional shape data due to the shadow generated below the object 306.
[0070] Figure 8(i) shows an example of a shadow 847 generated at the lower part of an object 846. Note that Figure 8(i) shows a cross-sectional image 844 when the object 846 is viewed from the horizontal direction. Here, instead of the object 306 used in the previous description, the shadow 847 generated by an object 846 such as a table will be used for explanation. Voxels corresponding to the region where the shadow 847 is generated at the lower part of the object 846 such as a table can be removed using a silhouette image corresponding to the captured image of the imaging device 302 as follows. For example, as shown in the cross-sectional image 844, the imaging device 302 has an elevation angle of the optical axis smaller than the elevation angle 848 of the tangent line drawn from the field plane at the center of the object 846 to the end of the upper surface of the object 846.
[0071] For example, based on the shape of the object 846 to be imaged, the user sets such an elevation angle as an angle threshold, and determines a part of one or more imaging devices 302 whose elevation angle of the optical axis is equal to or less than this angle threshold as the first imaging unit 110. Similar to the cases of the shadow 821, the virtual image 831, and the shadow 841, the user selects two or three imaging devices 302 among the plurality of imaging devices 302 that satisfy the above-described conditions for imaging a part of the imaging target space, and the angle difference between the optical axes of the imaging devices 302 is 120 degrees to 180 degrees. Further, the user determines the selected two or three imaging devices 302 as the first imaging unit 110.
[0072] After S711, or if it is determined in S710 that no shadow is generated at the lower part of the object 306, in S712, the user determines the remaining imaging devices 302 as the second imaging unit 130. Specifically, the user determines all of the imaging devices 302 that have not been determined as the first imaging unit 110 in the previous steps as the second imaging unit 130.
[0073] As described above, the image processing system 100 is configured such that, in the background subtraction method, an area where false detection of an object area may occur is set as a specific area. Here, the specific area according to Embodiment 2 is an area where a shadow or a virtual image of the object 306 may occur. Further, with respect to the data of the captured image output from at least a part of one or more imaging devices 302 whose viewing angle includes the specific area, the image processing system 100 is configured to input it to an image processing device that separates the object area by the learned model 123. Furthermore, with respect to the data of the captured image output from the imaging device 302 whose viewing angle does not include the specific area, the image processing system 100 is configured to input it to an image processing device that separates the object area by the background subtraction method. According to the image processing system 100 configured as described above, it is not necessary to perform both separation processes of the background subtraction method and the machine learning method in one image processing device, and a silhouette image capable of generating high-precision three-dimensional shape data can be generated.
[0074] <Other Embodiments> In Embodiment 1 and Embodiment 2, as an example, the case where the image processing system 100 is applied to a soccer game has been described, but the application destination of the image processing system 100 is not limited to this. An example of applying the image processing system 100 to other imaging targets is shown in FIG. 9. FIG. 9 shows an example of factors when non-detection or false detection of an object occurs in the separation process of the object area by the background subtraction method for each imaging target. An area including these factors is set as a specific area by the same setting method as in Embodiment 1 and Embodiment 2. Further, based on the specific area and the viewing angle of the imaging device 302, the imaging device 302 is determined as the first imaging unit 110 or the second imaging unit 130. Thereby, a silhouette image capable of generating high-precision three-dimensional shape data without missing or remaining cut portions with respect to the shape of the object 306 can be generated.
[0075] In addition, in Embodiment 1, as an example, in the process of separating the object region by the background difference method, the case where the object is not detected was described. Further, in Embodiment 2, as an example, in the process of separating the object region by the background difference method, the case where the object is erroneously detected was described. In each embodiment, examples of cases where the object is not detected and erroneously detected are shown separately, but the image processing system 100 can also be applied when these factors coexist. In this case, first, by the determination method shown in Embodiment 1, the imaging device 302 in which the factor causing non-detection exists is determined as the first imaging unit 110. Subsequently, by the determination method shown in Embodiment 2, a part of the imaging devices 302 in which the factor causing erroneous detection exists among the remaining imaging devices 302 is determined as the first imaging unit 110, and further, for the remaining imaging devices 302, they are determined as the second imaging unit 130. By determining the first imaging unit 110 and the second imaging unit 130 in this way, it is possible to generate a silhouette image capable of generating high-precision three-dimensional shape data without missing or remaining cut portions with respect to the shape of the object 306.
[0076] Note that the present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to each device of the image processing system 100 via a network or a storage medium, and a process in which one or more processors in the device read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0077] In addition, within the scope of the present disclosure, any combination of the embodiments, modification of any component of each embodiment, or omission of any component in each embodiment is possible.
[0078] <Configuration of the Present Disclosure> The present disclosure includes the following configurations and methods.
[0079] [Configuration 1] By inputting the data of the image obtained by imaging an object with a first imaging means that images an area including at least a part of a specific area as the data of a first input image into a learned model, a first image processing apparatus that generates data of a first silhouette image indicating an area where the object exists in the first input image; By using the data of the image obtained by imaging the object with a second imaging means different from the first imaging means as the data of a second input image, calculating the difference between the second input image and a background image which is an image obtained by imaging with the second imaging means in a state where the object does not exist in the area imaged by the second imaging means, a second image processing apparatus that generates data of a second silhouette image indicating an area where the object exists in the second input image; An image processing system, characterized by comprising the above.
[0080] [Configuration 2] The second imaging means images an area that does not include the specific area. The image processing system according to Configuration 1, characterized by the above.
[0081] [Configuration 3] The specific area is an area where the object can exist in a substantially stationary state. The image processing system according to Configuration 1 or 2, characterized by the above.
[0082] [Configuration 4] The specific area is an area where the difference between the color of the object and the color of the background of the area imaged by the imaging device which is the first imaging means or the second imaging means is smaller than a predetermined standard. The image processing system according to any one of Configurations 1 to 3, characterized by the above.
[0083] [Configuration 5] The specific area is an area where a shadow of the object or a virtual image caused by reflection of an image of the object can occur in a part of the area imaged by the imaging device which is the first imaging means or the second imaging means. The image processing system according to any one of Configurations 1 to 4, characterized by
[0084] [Configuration 6] Setting, as the first imaging means, a part of one or more imaging devices that at least include a part of the specific area within the imaging angle The image processing system according to any one of Configurations 1 to 5, characterized by
[0085] [Configuration 7] Setting, as the first imaging means, an imaging device among two or more imaging devices that at least include a part of the specific area within the imaging angle, where the angle formed by the optical axis vectors between the imaging devices is equal to or greater than a predetermined angle The image processing system according to any one of Configurations 1 to 6, characterized by
[0086] [Configuration 8] Setting, as the first imaging means, an imaging device among one or more imaging devices that at least include a part of the specific area within the imaging angle, where the angle formed by the optical axis vector of the imaging device and the field plane of the imaging target of the imaging device is equal to or greater than a predetermined angle or equal to or less than a predetermined angle The image processing system according to any one of Configurations 1 to 7, characterized by
[0087] [Configuration 9] The first image processing device A first acquisition means for acquiring, as the data of the first input image, the data of the image obtained by imaging the object by the first imaging means A first generation means for generating the data of the first silhouette image by inputting the acquired data of the first input image into the learned model A first output means for outputting the generated first silhouette image And having The second image processing device A second acquisition means for acquiring, as the data of the second input image, the data of the image obtained by imaging the object by the second imaging means Second generation means for generating data of the second silhouette image by calculating the difference between the obtained second input image and the background image; Second output means for outputting the generated second silhouette image; characterized in that it has; The image processing system according to any one of Configurations 1 to 8.
[0088] [Configuration 10] A third image processing device that generates three-dimensional shape data indicating the shape of the object using the data of the first silhouette image generated by the first image processing device and the data of the second silhouette image generated by the second image processing device; further characterized in that it has; The image processing system according to any one of Configurations 1 to 9.
[0089] [Configuration 11] The third image processing device includes: Silhouette acquisition means for acquiring the data of the first silhouette image output by the first image processing device and the data of the second silhouette image output by the second image processing device; Shape generation means for generating the three-dimensional shape data using the acquired data of the first silhouette image and the data of the second silhouette image; characterized in that it has; The image processing system according to Configuration 10.
[0090] [Method] A first imaging step in a first image processing device, including a first acquisition step of acquiring, as data of a first input image, data of an image obtained by imaging an object by first imaging means that images a region including at least a part of a specific region, a first step of generating data of a first silhouette image indicating a region where the object exists in the first input image by inputting the data of the first input image into a learned model, and a first output step of outputting the data of the first silhouette image; A second image processing step in the second image processing apparatus, including: a second acquisition step of acquiring, as data of a second input image, data of an image obtained by imaging the object by a second imaging unit different from the first imaging unit; a second generation step of generating data of a second silhouette image indicating a region where the object exists in the second input image by calculating a difference between the second input image and a background image which is an image obtained by imaging by the second imaging unit in a state where the object does not exist in a region imaged by the second imaging unit; and a second output step of outputting the data of the second silhouette image; An image processing method characterized by including the above.
Explanation of Signs
[0091] 100 Image processing system 110 First imaging unit 120 First image processing unit 121 First input unit 122 First separation unit 123 Learned model 124 Transmission unit 130 Second imaging unit 140 Second image processing unit 141 Second input unit 142 Second separation unit 143 Background image 144 Transmission unit
Claims
1. A first image processing apparatus that generates a first silhouette image indicating a region where the object exists in the first input image by inputting a first input image obtained by imaging the object with a first imaging means into a learned model; A second image processing apparatus that generates a second silhouette image indicating a region where the object exists in the second input image based on a difference between a second input image obtained by imaging the object with a second imaging means different from the first imaging means and a background image that is an image obtained by imaging with the second imaging means in a state where the object does not exist in the region imaged by the second imaging means; having The first image processing apparatus does not have a configuration for generating a silhouette image indicating a region where the object exists based on a difference between the first input image and a background image corresponding to the first input image. The first image processing apparatus generates the first silhouette image without inputting a background image into the learned model. The second image processing apparatus does not have a configuration for generating a silhouette image indicating a region where the object exists using a learned model. The second image processing apparatus generates the second silhouette image without using a learned model. An image processing system characterized by the above.
2. The first imaging means images a region including at least a part of a specific region. The second imaging means images a region that does not include the specific region. The image processing system according to claim 1, characterized by the above.
3. The specific region is a region where the object can exist in a substantially stationary state. The image processing system according to claim 2, characterized by the above.
4. The specific area is an area where the difference between the color of the object and the color of the background of the area imaged by the imaging device that is the first imaging means or the second imaging means is smaller than a predetermined standard. The image processing system according to claim 2, characterized in that.
5. The specific area is an area where a virtual image due to the reflection of the image of the object can occur in a part of the shadow of the object or the area imaged by the imaging device that is the first imaging means or the second imaging means. The image processing system according to claim 2, characterized in that.
6. The first image processing device A first acquisition means for acquiring the first input image; A first generation means for generating the first silhouette image; A first output means for outputting the generated first silhouette image; and has The second image processing device A second acquisition means for acquiring the second input image; A second generation means for generating the second silhouette image; A second output means for outputting the generated second silhouette image; and has The image processing system according to claim 1, characterized in that.
7. A third image processing device that generates three-dimensional shape data indicating the shape of the object using the first silhouette image generated by the first image processing device and the second silhouette image generated by the second image processing device; further having The image processing system according to claim 1, characterized in that.
8. The third image processing device A silhouette acquisition means for acquiring the first silhouette image output by the first image processing device and the second silhouette image output by the second image processing device; Shape generation means for generating the three-dimensional shape data using the obtained first silhouette image and the second silhouette image, having, An image processing system according to claim 7, characterized in that.
9. A step in a first image processing apparatus, including: a first acquisition step of acquiring a first input image obtained by imaging an object by first imaging means; a first generation step of generating a first silhouette image indicating a region where the object exists in the first input image by inputting the first input image into a learned model; and a first output step of outputting the first silhouette image, which is a first image processing step, A step in a second image processing apparatus, including: a second acquisition step of acquiring a second input image obtained by imaging the object by second imaging means different from the first imaging means; a second generation step of generating a second silhouette image indicating a region where the object exists in the second input image based on a difference between the second input image and a background image, which is an image obtained by imaging by the second imaging means in a state where the object does not exist in the region imaged by the second imaging means; and a second output step of outputting the second silhouette image, which is a second image processing step, including, The first image processing step does not include a step of generating a silhouette image indicating a region where the object exists based on a difference between the first input image and a background image corresponding to the first input image, In the first image processing step, the first silhouette image is generated without inputting a background image into a learned model, The second image processing step does not include a step of generating a silhouette image indicating a region where the object exists using a learned model, In the second image processing step, the second silhouette image is generated without using a learned model An image processing method characterized by that.
Citation Information
Patent Citations
Image processing device, image processing method, and program
JP2021056960A