Image processing system and image processing method

By combining background difference method and machine learning method in the image processing system, silhouette images are generated, and the problem of object undetected and error detection is solved, and high-accuracy 3D shape data generation is achieved, while reducing the cost of image processing equipment.

JP2025074247AActive Publication Date: 2025-05-13CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025032990
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-13
Estimated Expiration
2043-06-30

Smart Images

  • Figure 2025074247000001_ABST
    Figure 2025074247000001_ABST
Patent Text Reader

Abstract

To generate silhouette images with which highly accurate three-dimensional shape data can be generated while suppressing implementation cost of image processing apparatuses.SOLUTION: An image processing system 100 includes: a first image processing apparatus configured to generate data of a first silhouette image representing a region in which an object exists in a first input image by inputting data of an image, as data of the first input image, into a trained model, the image being obtained by first imaging means imaging the object, the first imaging means imaging a region including at least part of a specific region; and a second image processing apparatus configured to output data of a second silhouette image representing a region in which the object exists in a second input image by calculating a difference between the second input image and a background image captured by second imaging means in a state in which the object does not exist, using data of an image, as data of the second input image, which is obtained by the second imaging means that images the object, the second imaging means being different from the first imaging means.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to techniques for generating silhouette images that show foreground regions in an image. [Background technology]

[0002] There is a technique for generating three-dimensional shape data showing the three-dimensional shape of an object of interest (hereinafter simply referred to as an object) using data of a plurality of images (hereinafter referred to as "captured images") obtained by synchronous imaging using an imaging device. The three-dimensional shape data is generated by a method such as a visual volume intersection method using a silhouette image generated by extracting an area corresponding to the object from each captured image (hereinafter referred to as an "object area"). Methods for generating a silhouette image from a captured image include, for example, a background subtraction method or a machine learning method. In the background subtraction method, a captured image obtained by imaging during a period in which an object does not exist in an angle of view captured by a certain imaging device is used as a background image, and a difference is taken between this background image and a captured image obtained by imaging during a period in which an object exists. Furthermore, an object area and a background area in the captured image are separated based on this difference to generate a silhouette image showing the object area (hereinafter referred to as a "silhouette image of an object"). In the machine learning method, first, a sufficient number of learning data are prepared in which a captured image obtained by imaging during a period in which an object exists is paired with data showing the object area in the captured image, which is training data. Next, a trained model is generated as a result of training a training model using the training data, and a silhouette image of the object is generated by separating the object region and the background region in the captured image using the trained model.

[0003] A suitable separation method for separating an object region and a background region in a captured image varies depending on the characteristics included in the angle of view of the imaging device. When a background subtraction method is used as a separation method, an object region may not be detected from a captured image in the following regions (hereinafter referred to as "non-detection of an object"). For example, a region where an object with little movement exists, a region where a stationary background of a color similar to that of the object exists, and a region where there is movement and a background of a color similar to that of the object exists in part. In addition, in the following regions, a region other than the object region may be erroneously detected as a foreground region in the captured image (hereinafter referred to as "erroneous detection of an object"). For example, a region where a shadow of an object occurs, and a region where a virtual image occurs due to the image of an object being reflected on a glossy floor or a wet field surface.

[0004] The non-detection or erroneous detection of these objects causes chipping or remaining scraping of three-dimensional shape data. Therefore, for areas where non-detection or erroneous detection may occur in separation processing by the background subtraction method, it is useful to perform separation processing by the machine learning method. On the other hand, separation processing by the machine learning method generally has a higher calculation cost than separation processing by the background subtraction method. In addition, since separation processing by the machine learning method infers the shape of an object from statistical information of surrounding pixels, the accuracy of the boundary of the object area in the silhouette image is degraded compared to separation processing by the background subtraction method. Patent Document 1 discloses a technology for sharing a region where separation processing by the background subtraction method and a region where separation processing by the machine learning method are performed within the angle of view of a certain imaging device. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent Publication No. 2021-56960 Summary of the Invention [Problem to be solved by the invention]

[0006] However, in the technology disclosed in Patent Document 1, an image processing device corresponding to one imaging device needs to be equipped with configurations for performing both separation processing using the background subtraction method and separation processing using the machine learning method, which increases the implementation cost for one image processing device. [Means for solving the problem]

[0007] The image processing system according to the present disclosure includes a first image processing device that generates first silhouette image data indicating the area in which the object exists in the first input image by inputting image data obtained by capturing an image of an object using a first imaging means that captures an area including at least a part of a specific area as first input image data into a trained model, and a second image processing device that generates second silhouette image data indicating the area in which the object exists in the second input image by using image data obtained by capturing an image of the object using a second imaging means different from the first imaging means as second input image data and calculating a difference between the second input image and a background image, which is an image obtained by capturing an image using the second imaging means in a state in which the object is not present in the area captured by the second imaging means. Effect of the Invention

[0008] According to the present disclosure, it is possible to generate a silhouette image that enables generation of highly accurate three-dimensional shape data while suppressing implementation costs in an image processing device. [Brief description of the drawings]

[0009] [Figure 1] FIG. 2 is a block diagram showing an example of a functional configuration of the image processing system. [Diagram 2] 2 is a block diagram showing an example of a hardware configuration of a first image processing unit, a second image processing unit, and a third image processing unit. FIG. [Diagram 3] FIG. 1 is a diagram illustrating an application example of an image processing system. [Figure 4]13 is a flowchart showing an example of a flow of a determination step of the first or second imaging section, and the first or second image processing section. [Diagram 5] FIG. 2 is a diagram illustrating an example of the angle of view of an imaging device. [Figure 6] 6 is a flowchart showing an example of a processing flow of a first image processing unit, a second image processing unit, and a third image processing unit. [Figure 7] 13 is a flowchart showing an example of a flow of a determination step of the first or second imaging section, and the first or second image processing section. [Figure 8] 1A and 1B are diagrams for explaining an example of a shadow and a virtual image of an object. [Figure 9] 11A and 11B are diagrams illustrating examples of application of the image processing system to other imaging targets. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, with reference to the drawings, an embodiment for carrying out the present disclosure will be described. The following embodiments do not limit the present disclosure, and not all of the combinations of features described in the present embodiment are necessarily essential to the solution of the present disclosure. The same configurations will be described with the same reference numerals. In addition, terms with different alphabets added after the reference numerals indicate devices having the same functions but different from each other. For example, the first imaging unit 110A and the first imaging unit 110B shown in FIG. 1 indicate that they are different devices having the same functions. Note that having the same function refers to having at least a specific function such as an imaging function, and for example, some of the functions and performances of the first imaging unit 110A and the first imaging unit 110B may be different from each other.

[0011] <Embodiment 1> In this embodiment, a case will be described in which, in a process of separating an area corresponding to an object in an image (foreground area) from the background area using the background difference method, an area in which it is possible that the foreground area cannot be detected exists within the field of view of an imaging device.

[0012] [System configuration] FIG. 1 is a block diagram showing an example of a functional configuration of an image processing system 100 according to the first embodiment. The image processing system 100 generates three-dimensional shape data (hereinafter referred to as "three-dimensional shape data of an object") 160 indicating a three-dimensional shape of an object. The first imaging unit 110 captures an image of an object, acquires captured image data obtained by the image capture, and outputs the acquired captured image data to the first image processing unit 120. The first image processing unit 120 receives the captured image data output by the first imaging unit 110, separates a region (object region) corresponding to an object for which three-dimensional shape data is to be generated from a background region in the captured image, and generates a silhouette image of the object. The silhouette image data generated by the first image processing unit 120 is transmitted to the third image processing unit 150.

[0013] The first image processing unit 120 includes a first input unit 121, a first separation unit 122, and a transmission unit 124. The first input unit 121 receives data of the captured image output by the first imaging unit 110, and stores the received data of the captured image in a storage device provided inside the first image processing unit 120. The first separation unit 122 separates a foreground region and a background region, which are object regions, included in the captured image to generate a silhouette image of the object. Specifically, the first separation unit 122 separates the object region and generates a silhouette image by performing a separation process using a machine learning method. More specifically, the first separation unit 122 holds a trained model 123 therein. The trained model 123 is a model that has been trained to receive data of a captured image including an object region as input and output data of a silhouette image of an object. The trained model 123 separates a foreground region and a background region in the captured image to generate a silhouette image, for example, by executing a Semantic Segmentation task that performs class identification for each pixel of an input image.

[0014] There are various methods using a CNN (Convolutional Neural Network) to realize the trained model 123, and there are a plurality of network structures using an Encoder layer, a Decoder layer, and a Skip structure. For example, the trained model 123 is realized by using a structure called SegNet or Unet. Data of the generated silhouette image and data of a texture image holding color information of an object region in the silhouette image are output to a transmission unit 124. The transmission unit 124 receives the data of the silhouette image and the texture image output from the first separation unit 122, and transmits them to the third image processing unit 150. The transmission is performed via a communication I / F (interface) such as a LAN (Local Area Network) or a WAN (Wide Area Network).

[0015] The second imaging section 130 captures an image of an object to be imaged, acquires captured image data obtained by the image capture, and outputs the acquired captured image data to the second image processing section 140. The second image processing section 140 receives the captured image data output by the second imaging section 130, separates an area (object area) corresponding to an object for which three-dimensional shape data is to be generated in the captured image from a background area, and generates a silhouette image of the object. The silhouette image data generated by the second imaging section 130 is transmitted to the third image processing section 150.

[0016] The second imaging section 130 includes a second input section 141, a second separation section 142, and a transmission section 144. The second input section 141 stores the received captured image data in a storage device provided inside the second image processing section 140. The second separation section 142 separates an object region and a background region included in the captured image to generate a silhouette image of the object. Specifically, the second separation section 142 separates an object region and a background region included in the captured image by performing a separation process using a background difference method to generate a silhouette image of the object. More specifically, for example, the second separation section 142 calculates a difference between a background image 143, which is a captured image acquired in a state where the object to be imaged does not exist in the angle of view of the second imaging section 130, and a captured image acquired in a state where the object exists. Furthermore, the second separation section 142 generates a silhouette image of the object by separating the object region and the background region in the captured image based on the calculated difference.

[0017] A method of acquiring the background image 143 used in the separation process using the background difference method may be a method of using a captured image at a moment when no object exists as the background image. Alternatively, a method of generating a background image using a plurality of captured images captured over a certain period of time may be used. Specifically, for example, a change in pixel value in each captured image is observed in pixel or small region units, and when the change in pixel value is within a certain amount, a background image is generated using an average value of pixels within a certain period or the latest pixel value. The data of the silhouette image generated by the second separation unit 142 and the data of the texture image holding the color information of the object region in the silhouette image are output to the transmission unit 144. The transmission unit 144 has the same function as the transmission unit 124, and receives the data of the silhouette image and the texture image output from the second separation unit 142 and transmits these data to the third image processing unit 150.

[0018] The third image processing unit 150 generates three-dimensional shape data 160 of the object using the silhouette image and texture image data received from the first image processing unit 120 and the second image processing unit 140. The third image processing unit 150 has a receiving unit 151 and a shape generating unit 152. The receiving unit 151 receives the silhouette image and texture image data transmitted from the first image processing unit 120 and the second image processing unit 140, and stores these data in a storage device provided inside the third image processing unit 150.

[0019] The shape generating unit 152 performs a shape estimation process and a coloring process for generating the three-dimensional shape data 160 using the silhouette image and texture image data received by the receiving unit 151. Specifically, the shape generating unit 152 first performs a shape estimation process, and then performs a coloring process to generate the three-dimensional shape data 160. For example, the shape estimation process can be performed using a volume intersection method. For example, in the volume intersection method, first, a rectangular parallelepiped having a unit volume called a voxel is laid out in a target space for generating the three-dimensional shape data. Hereinafter, a set of voxels laid out in a target space for generating the three-dimensional shape data is referred to as a voxel group. Next, each silhouette image generated by performing a separation process on the captured image of the first or second imaging unit 110, 130 is projected onto a voxel group in the imaging range of the first or second imaging unit 110, 130. Next, using all silhouette images, voxels not included in the area where the object area in each silhouette image is projected are scraped off from this voxel group to generate uncolored three-dimensional shape data. Next, in the coloring process, texture image data corresponding to each silhouette image is used as color information, and texture is applied to each voxel of the uncolored three-dimensional shape data generated by the shape estimation process to color the three-dimensional shape data. By performing the above processes, shape generation unit 152 generates three-dimensional shape data 160.

[0020] Here, we will explain cases where the shape estimation process fails. One is, for example, when a silhouette image of an object does not include a region corresponding to the object (object region), that is, when the object is not detected. The other is, for example, when a silhouette image of an object includes a region other than the region corresponding to the object as a foreground region, that is, when the object is erroneously detected.

[0021] When an object is not detected, the voxels corresponding to the object are mistakenly scraped off from the voxel group. Therefore, the object should not be not detected in all the images captured by the first and second imaging units 110 and 130. When an object is not detected, three-dimensional shape data of a false shape different from the shape of the object is generated only when an image area corresponding to the same voxel is similarly falsely detected in all the images captured by the first and second imaging units 110 and 130. Therefore, in the case of false detection of an object, it is sufficient to suppress false detection in at least one of the images captured by the first or second imaging unit 110 or 130 that captures the imaging target space corresponding to the space in which the voxels constituting the false shape exist.

[0022] The three-dimensional shape data 160 generated by the above-mentioned process is a three-dimensional point cloud that is a collection of colored voxels, but the form of the three-dimensional shape data is not limited to this. For example, the three-dimensional shape data 160 may be three-dimensional shape data based on a three-dimensional polygon mesh generated from a colored three-dimensional point cloud. In this case, for example, the shape generating unit 152 may include a process of generating three-dimensional polygon mesh data from the three-dimensional point cloud as the generated three-dimensional shape data.

[0023] 2 is a block diagram showing an example of a hardware configuration of the first image processing unit 120, the second image processing unit 140, and the third image processing unit 150. The hardware configurations of the first image processing unit 120, the second image processing unit 140, and the third image processing unit 150 (hereinafter collectively referred to as "image processing device 200") are the same as each other. The image processing device 200 has a CPU 201, a ROM 202, a RAM 203, an auxiliary storage device 204, a display unit 205, an operation unit 206, a communication I / F 207, and a bus 208.

[0024] The CPU 201 uses computer programs and data stored in the ROM 202 or the RAM 203 to control the entire image processing device 200, and realizes each part that the image processing device 200 has as a functional configuration. The image processing device 200 may have one or more dedicated hardware different from the CPU 201, and at least a part of the processing by the CPU 201 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor). The ROM 202 stores programs that do not require modification. The RAM 203 is used as a workspace for the CPU 201, and temporarily stores programs or data supplied from the auxiliary storage device 204, and data supplied from the outside via the communication I / F 207. The auxiliary storage device 204 is composed of a large-capacity storage device such as a hard disk drive, and stores various data. The RAM 203 and the auxiliary storage device 204 hold input data, intermediate data during processing, and output data of the image processing device 200.

[0025] The RAM 203 or the auxiliary storage device 204 of the first image processing unit 120 holds data of the captured image received from the first imaging unit 110, the trained model 123 used in the first separation unit 122, and intermediate data in the separation process using the trained model 123. The RAM 203 or the auxiliary storage device 204 of the first image processing unit 120 holds data of the silhouette image and the texture image generated by the first separation unit 122. The RAM 203 or the auxiliary storage device 204 of the second image processing unit 140 holds data of the captured image received from the second imaging unit 130, data of the background image 143 used by the second separation unit 142, and intermediate data for generating the background image 143. The RAM 203 or the auxiliary storage device 204 of the second image processing unit 140 holds data of the silhouette image and the texture image generated by the second separation unit 142. The RAM 203 or the auxiliary storage device 204 of the third image processing unit 150 holds the silhouette image and texture image data received from the first image processing unit 120 and the second image processing unit 140. The RAM 203 or the auxiliary storage device 204 of the third image processing unit 150 also holds intermediate data in the shape estimation process in the shape generation unit 152, and uncolored three-dimensional shape data after the shape estimation process.

[0026] The display unit 205 is configured with a liquid crystal display, an LED (light-emitting diode), or the like, and displays a GUI (Graphical User Interface) or the like for the user to operate the image processing device 200. The operation unit 206 is configured with a keyboard, a mouse, a joystick, a touch panel, or the like, and receives operations by the user to input various instructions to the CPU 201. The CPU 201 operates as a display control unit that controls the display unit 205, and an operation control unit that controls the operation unit 206. The communication I / F 207 is used for communication between the image processing device 200 and an external device.

[0027] The first image processing unit 120 has, as a communication I / F 207, an I / F for receiving captured image data from the first imaging unit 110, and an I / F for transmitting silhouette image and texture image data to the third image processing unit 150. The second image processing unit 140 has, as a communication I / F 207, an I / F for receiving captured image data from the second imaging unit 130, and an I / F for transmitting silhouette image and texture image data to the third image processing unit 150. The third image processing unit 150 has, as a communication I / F 207, an I / F for receiving data transmitted from the first image processing unit 120 and the second image processing unit 140. The communication I / F 207 is realized by a wired network I / F such as Ethernet, a wireless network I / F such as wireless LAN, or an SDI (Serial Digital Interface) I / F for transmitting and receiving video signals. The bus 208 connects each unit that the image processing device 200 has as a hardware configuration to transmit information. In this embodiment, the display unit 205 and the operation unit 206 are assumed to exist inside the image processing device 200, but at least one of the display unit 205 and the operation unit 206 may exist outside the image processing device 200 as a separate device.

[0028] [System application examples] FIG. 3 is a diagram showing an application example 300 of the image processing system 100 according to the first embodiment. The application example 300 shown in FIG. 3 shows an example in which the image processing system 100 is applied to a soccer stadium as an image capture target. In an image capture target space 301, there are a competition field 305 and an object 306 for which three-dimensional shape data is to be generated. In the case of soccer, the object 306 includes objects 306A to 306C which are players, and an object 306D which is a soccer ball. Of these, the objects 306A and 306C are field players, and the object 306B is a goalkeeper. The object 306 may include a referee not shown in FIG. 3. In addition, in the vicinity of the competition field 305, there are a bench 307 where the players wait, and a signage 308 which is painted on the floor of the soccer stadium or laid down as a sign-like three-dimensional object. In addition, there are lighting devices 309 around the competition field 305 which are turned on when the illuminance is insufficient at night or in rainy weather.

[0029] The signboard-like signage 308 includes digital signage that is displayed on a display and whose display content changes over time. The lighting device 309 may be a light source that is locally arranged and emits strong light so that a deep shadow of the object 306 is generated in a specific direction. The lighting device 309 may also be a light source that is arranged at regular intervals around the competition field 305 so that a light shadow of the object 306 is evenly generated around it. In addition, when the field surface is wet due to rain or the like, a virtual image of the object 306 may be generated on the floor surface of the competition field 305 due to specular reflection. Note that the processing of the image processing system 100 when a shadow or virtual image of the object 306 is generated will be described in the second embodiment.

[0030] Imaging devices 302A to 302P, which are the first imaging section 110 or the second imaging section 130, are arranged around the imaging target space 301. Each imaging device 302 captures at least a part of the imaging target space 301. FIG. 3 shows an aspect in which the imaging devices 302 are arranged so as to surround the imaging target space 301 on a plan view overlooking the imaging target space 301, but the imaging devices 302 are arranged with variations in the elevation angle of the imaging devices 302 with respect to the imaging target space 301. Such an arrangement of the imaging devices 302 makes it possible to capture images of the object 306 from various angles, and to generate high-quality three-dimensional shape data.

[0031] In addition, when a shadow or a virtual image of the object 306 described in the second embodiment occurs, each imaging device 302 is determined to be the first imaging unit 110 or the second imaging unit 130 according to the elevation angle of the imaging device 302. The imaging device 302 outputs captured image data to the imaging device control box 303 which is the first image processing unit 120 or the second image processing unit 140. The imaging device control box 303 receives the captured image data, and generates silhouette image and texture image data of the object 306 when an area corresponding to the object 306 (hereinafter referred to as the "area of ​​the object 306") exists in the captured image. Furthermore, the imaging device control box 303 transmits the generated silhouette image and texture image data to the image processing server 304 which is the third image processing unit 150. The image processing server 304 generates three-dimensional shape data using the silhouette image and texture image data received from each imaging device control box 303.

[0032] A method of determining whether the imaging device 302 and the imaging device control box 303 are the first imaging section 110 and the first image processing section 120 or the second imaging section 130 and the second image processing section 140 will be described with reference to Figs. 4 and 5. Fig. 4 is a flowchart showing an example of a flow of a process of determining the first or second imaging section 110, 130 and the first or second image processing section 120, 140 according to the first embodiment. The determination of the first or second imaging section 110, 130 and the first or second image processing section 120, 140 is performed before imaging is performed using the image processing system 100, based on the angle of view of each imaging device 302 determined when the imaging device 302 is installed. In the following description, the letter "S" at the beginning of the code means a step (process). In S401, the user sets a specific area. In embodiment 1, a specific region is a region in which there is a possibility that object 306 will not be detected even though the region of object 306 exists, in the process of separating the region of object 306 using the background subtraction method performed by the second image processing unit 140.

[0033] The specific region will be described with reference to FIG. 5. FIG. 5 is a diagram showing an example of the angle of view of the imaging device 302 according to the first embodiment. FIG. 5(a) shows an example of the angle of view 510 of the imaging device 302 capturing the objects 306E and 306G, which are field players, and the object 306F, which is a ball, present on the competition field 305. In the angle of view 510, all the objects 306 move at a speed equal to or higher than a certain value, so that the same object 306 does not continue to exist in a pixel for a certain period of time or more. Therefore, a background image in which the object 306 does not exist can be generated. Here, it is assumed that there is a difference of a certain value or more between the color value of the object 306 and the pixel value of the background image in the angle of view 510. In this case, it is possible to create a background image, and the difference between the pixel value of the region of the object 306 in the captured image captured at the angle of view 510 and the pixel value of the region corresponding to the object 306 in the background image is also a certain value or more. Therefore, the angle of view 510 of the imaging device 302 does not include the specific region.

[0034] FIG. 5(b) shows a field of view 520 of the imaging device 302 capturing objects 306I and 306J, which are field players, an object 306H, which is a ball, and an object 306K, which is a goalkeeper, which are present on the competition field 305. In addition, the field of view 520 of the imaging device 302 also includes a signage 308E painted on the floor of the soccer stadium. FIG. 5(c) shows a field of view 521 that is the same as the field of view 520 of the imaging device 302 shown in FIG. 5(b), and shows the field of view 521 after a certain time has passed since the state of the field of view 520. In the field of view 521, the object 306 and the signage 308 that existed in the state of the field of view 520 exist. Among them, the object 306 that has moved is shown in FIG. 5(c) with a "'" at the end of its reference number. For example, the object 306I is the same field player as the object 306I', and it is shown that the field player is moving.

[0035] The specific region in the angle of view 520 (521) will be described. First, the position of the object 306K, which is a goalkeeper, has hardly changed even after a certain time has passed. In a background image generation method in which pixels whose pixel values ​​do not change over a certain time are determined to be background regions, the region in the captured image corresponding to the object 306, which is substantially still, is determined to be a region constituting the background image. That is, the region of the object 306K is included in the background image. As a result, when a separation process of the object region is performed by the background difference method, no difference occurs between the captured image and the background image in the region of the object 306K, and the object 306K is not detected. Therefore, a silhouette image of the object 306K is not generated. As in the case of the object 306K, which is a goalkeeper, the user sets a region where the object 306, which is less moving, i.e., where the object 306 may be substantially still, as a specific region.

[0036] On the other hand, the signage 308 painted on the floor of the soccer stadium is not moving, and is therefore included in the background image. Here, a case will be considered in which the difference between the pixel values ​​of the area corresponding to the signage 308 in the captured image and the pixel values ​​of the uniform of the player, which is the object 306, is smaller than a predetermined threshold. That is, a case will be considered in which the color of the signage 308 and the color of the uniform are similar. In this case, the difference between the pixel values ​​of the uniform of the player and the signage 308 is not equal to or greater than the predetermined threshold, so that the object 306, which is the player, is not detected, and a silhouette image showing the area corresponding to the player is not generated. In this way, the user sets, as a specific area, a background area in which a color similar to the color of the object 306 for which three-dimensional shape data is to be generated may exist in a part of the background image.

[0037] FIG. 5(d) shows the angle of view 530 of the imaging device 302 capturing the objects 306M and 306N, which are field players, the object 306L, which is a ball, and the bench 307C, where the reserve players wait, which are present on the competition field 305. Here, the bench 307C and the reserve players present therein are not targets for generating 3D shape data. Since the reserve players on the bench 307C are natural persons, the reserve players move over time. When a background image is generated in such an angle of view 530 of the imaging device 302, a part of the area corresponding to the reserve players in the captured image is not included in the background image because it is not a pixel that remains unchanged for a certain period of time. When such a background image is used to perform a process of separating object areas using a background difference method on the captured image, a part of the area corresponding to the reserve players is separated as a foreground area, and a silhouette image showing the foreground area is generated.

[0038] On the other hand, a part of the region in the captured image corresponding to a stationary substitute player has pixels that do not change for a certain period of time or more, and is therefore incorporated into the background image. In this case, there is no difference of a certain amount between the pixel values ​​of the region corresponding to the body or uniform of the substitute player incorporated into the background image and the pixel values ​​of the region in the captured image corresponding to the body or uniform of the field player, which is the object 306. Therefore, separation of the object region cannot be performed correctly, and part of the object 306 is not detected, resulting in the generation of a silhouette image in which part of the region corresponding to the object 306 is missing. In this way, the user sets the background region in which a stable background image cannot be generated and a color similar to the color of the object 306 may exist in part of the background image as a specific region.

[0039] FIG. 5(e) shows a field angle 540 of the imaging device 302 capturing objects 306P and 306Q, which are field players, and object 306O, which is a ball, present on the competition field 305. Also, in the field angle 540 of the imaging device 302, there is a signage 308F, which is a billboard-like digital signage installed outside the competition field 305. FIG. 5(f) shows a field angle 541, which is the same as the field angle 540 of the imaging device 302 shown in FIG. 5(e), after a certain time has elapsed. In the field angle 541, there are an object 306 and a signage 308 that existed in the state of the field angle 540 of the imaging device 302. Among them, the object 306 and the signage 308 that have moved are indicated by adding "'" to the end of their reference numerals in FIG. 5(f). For example, signage 308F, which is digital signage, is referred to as signage 308F' because the display content has changed.

[0040] The specific region in the angle of view 540 (541) will be described below. Since the display content of the digital signage, which is the signage 308, changes, the pixels in the region corresponding to the signage 308 in the captured image do not become pixels whose pixel values ​​do not change for a certain period of time or more, and the region is not incorporated into the background image. Therefore, when the captured image in the state of the angle of view 540 and the captured image in the state of the angle of view 541 are subjected to object region separation processing by the background difference method, the region of the signage 308F is also erroneously separated as being a foreground region. As a result, a silhouette image of the signage 308F, which is not the target for generating three-dimensional shape data, is generated.

[0041] On the other hand, when the pixel values ​​of a region in the captured image corresponding to the signage 308 do not change for a certain period of time or more, the region is incorporated into the background image. In this case, when the color used in the display content of the digital signage when incorporated into the background image is similar to the color of the object 306, the object 306 goes undetected. As a result, a silhouette image of the object 306 is not generated. The user sets, as the specific region, such a region where a silhouette image of an object not targeted for three-dimensional shape data generation may be generated, and a region where a stable background image cannot be generated and the object 306 goes undetected.

[0042] In summary, the specific region according to the first embodiment includes a region where an object with little movement may exist and a region where a fixed background with a color similar to that of the object may exist. The specific region according to the first embodiment also includes a region where there is movement and a background region with a color similar to that of the object 306 may exist in a part of the background image. The specific region is not limited to the above, and a region including a factor that causes an object not to be detected during the object region separation process using the background difference method may be set as the specific region. By setting such a specific region, the user can apply the present disclosure to cases other than the above case described with reference to FIG. 5.

[0043] Returning to the description of the flowchart shown in FIG. 4, after S401, the user determines whether all the imaging devices 302 are the first imaging unit 110 or the second imaging unit 130, and whether the captured image data output by the imaging devices 302 is to be processed by the first image processing unit 120 or the second image processing unit 140. Specifically, the user performs steps from S403 to S405 for each imaging device 302. Specifically, for example, in the case of the application example 300 shown in FIG. 3, the user performs steps from S403 to S405 for all the imaging devices 302 from the imaging device 302A to the imaging device 302P. Note that S402 indicates the start of a loop of steps from S403 to S405.

[0044] In S402, the user selects an arbitrary imaging device 302 from among one or more imaging devices 302 that have not been selected so far. Next, in S403, the user determines whether or not the angle of view captured by the imaging device 302 selected in S402 includes a specific area. If it is determined in S403 that the angle of view includes a specific area, in S404, the user determines that the imaging device 302 selected in S402 is the first imaging unit 110 and that the imaging device control box 303 corresponding to the imaging device 302 is the first image processing unit 120. If it is determined in S403 that the angle of view does not include a specific area, in S405, the user determines that the imaging device 302 selected in S402 is the second imaging unit 130 and that the imaging device control box 303 corresponding to the imaging device 302 is the second image processing unit 140.

[0045] S406 indicates the end of the loop of steps from S403 to S405. In S406, the user judges whether or not all of the imaging devices 302 have been selected in S402. If it is judged in S406 that all of the imaging devices 302 have not been selected, that is, that there are imaging devices 302 that have not been selected, the user returns to S402 and selects any imaging device 302 that has not been selected so far. Thereafter, the steps from S402 to S406 are repeatedly performed until it is judged in S406 that all of the imaging devices 302 have been selected. If it is judged in S406 that all of the imaging devices 302 have been selected, the user ends the determination step shown in the flowchart of FIG. 4.

[0046] 5, a specific example of a method for determining the imaging device 302 and the imaging device control box 303 will be described. Since the angle of view 510 of the imaging device 302 does not include a specific region, the imaging device 302 is determined to be the second imaging section 130, and the imaging device control box 303 connected to the imaging device 302 is determined to be the second image processing section 140. On the other hand, since the angles of view 520, 530, and 540 of the imaging device 302 all include a specific region, the imaging device 302 is determined to be the first imaging section 110, and the imaging device control box 303 connected to the imaging device 302 is determined to be the first image processing section 120.

[0047] For example, the user first installs all the imaging devices 302 at desired positions and adjusts the angle of view of each imaging device 302 to the desired angle of view. Next, the user installs the first image processing unit 120, for example, near the imaging device 302 determined to be the first imaging unit 110 based on the determination in the determination step shown in the flowchart of Fig. 4, and connects the first image processing unit 120 to the first imaging unit 110 and the third image processing unit 150. Similarly, the user installs the second image processing unit 140, for example, near the imaging device 302 determined to be the second imaging unit 130, and connects the second image processing unit 140 to the second imaging unit 130 and the third image processing unit 150.

[0048] The operations of the first image processing unit 120, the second image processing unit 140, and the third image processing unit 150 will be described with reference to FIG. 6. FIG. 6(a) is a flowchart showing an example of a processing flow of the first image processing unit 120 according to the first embodiment. First, in S601, the first input unit 121 receives data of the captured image output by the first imaging unit 110. Next, in S602, the first separation unit 122 inputs the captured image data received in S601 to the trained model 123. Next, in S603, the first separation unit 122 acquires data of the silhouette image generated by the trained model 123. Next, in S604, the transmission unit 124 outputs the data of the silhouette image acquired in S603 to the third image processing unit 150. After S604, the first image processing unit 120 ends the processing of the flowchart shown in FIG. 6(a), and repeatedly executes the processing of the flowchart shown in FIG. 6(a) every time the first imaging unit 110 outputs new captured image data.

[0049] FIG. 6B is a flowchart showing an example of a processing flow of the second image processing unit 140 according to the first embodiment. First, in S611, the second input unit 141 receives captured image data output by the second imaging unit 130. Next, in S612, the second separation unit 142 generates a silhouette image by performing separation processing by a background difference method on the captured image data received in S601 using the background image data. Next, in S613, the transmission unit 144 outputs the silhouette image data generated in S612 to the third image processing unit 150. After S613, the second image processing unit 140 ends the processing of the flowchart shown in FIG. 6B, and repeatedly executes the processing of the flowchart shown in FIG. 6B every time the second imaging unit 130 outputs new captured image data.

[0050] 6(c) is a flowchart showing an example of a processing flow of the third image processing unit 150 according to the first embodiment. First, in S621, the receiving unit 151 receives silhouette image and texture image data transmitted from the first image processing unit 120 and the second image processing unit 140. Here, the first imaging unit 110 and the second imaging unit 130 perform imaging synchronously with each other, for example, and the receiving unit 151 receives silhouette image and texture image data based on captured image data obtained by imaging at the same time. Note that the same time here is not limited to the exact same time, but includes approximately the same time.

[0051] After S621, in S622, the shape generation unit 152 executes a shape estimation process for the three-dimensional shape data using the silhouette image data received in S621. Next, in S623, the shape generation unit 152 executes a coloring process for the three-dimensional shape data generated in S622 using the texture image data to generate three-dimensional shape data 160. After S623, the third image processing unit 150 ends the process of the flowchart shown in FIG. 6(c). Thereafter, the third image processing unit 150 repeatedly executes the process of the flowchart shown in FIG. 6(c) every time the first image processing unit 120 and the second image processing unit 140 output new silhouette image and texture image data.

[0052] As described above, the image processing system 100 is configured to set an area where an object may not be detected in the background subtraction method as a specific area. The image processing system 100 is configured to input at least captured image data output from the imaging device 302 whose angle of view includes the specific area to an image processing device that separates the object area using the learned model 123. The image processing system 100 is configured to input captured image data output from the imaging device 302 whose angle of view does not include the specific area to an image processing device that separates the object area using the background subtraction method. The image processing system 100 is also configured to determine an image processing device that appropriately separates the object area according to the angle of view of the imaging device 302. According to the image processing system 100 configured as described above, it is not necessary to provide a configuration that performs both separation processing using the background subtraction method and separation processing using the machine learning method in one image processing device. Therefore, it is possible to generate a silhouette image that can generate high-precision three-dimensional shape data while suppressing the implementation cost of the image processing device.

[0053] <Embodiment 2> In this embodiment, a case will be described in which an area other than the area of ​​the object 306 may be erroneously detected as a foreground area in the object area separation process using the background difference method, is present within the angle of view of the imaging device 302. Areas that may be erroneously detected include an area where a shadow of the object 306 may be generated, and an area where a virtual image may be generated due to the image of the object 306 being reflected on the surface of the competition field 305. The functional configuration of the image processing system according to the second embodiment is the same as that of the image processing system 100 shown in FIG. 1, so the differences from the first embodiment will be described below.

[0054] [System application examples] As in the first embodiment, the application example will be described as being similar to application example 300 of the image processing system shown in Fig. 3. A method for determining whether the imaging device 302 and the imaging device control box 303 are the first imaging section 110 and the first image processing section 120, or the second imaging section 130 and the second image processing section 140 will be described with reference to Figs. 7 and 8. Fig. 7 is a flowchart showing an example of the flow of a process for determining the first or second imaging section 110, 130, and the first or second image processing section 120, 140 according to the second embodiment.

[0055] The determination process shown in the flowchart of FIG. 7 is performed before the processing of the image processing system is started, based on the angle of view determined when the imaging device 302 is installed and the weather or lighting conditions when the image processing system is used. It is assumed that changes in the weather or lighting conditions and changes in the shadow or virtual image of the object 306 that may occur in the angle of view of each imaging device 302 due to the changes have been investigated in advance. Note that the flowchart of FIG. 7 only describes whether a certain imaging device 302 is to be the first imaging unit 110 or the second imaging unit 130. However, it is assumed that the imaging device control box 303 is also to be determined to be the first image processing unit 120 or the second image processing unit 140 in accordance with the determination of the first imaging unit 110 or the second imaging unit 130. Specifically, when a certain imaging device 302 is determined to be the first imaging unit 110, the imaging device control box 303 corresponding to the imaging device 302 is determined to be the first image processing unit 120. On the other hand, when a certain imaging device 302 is determined to be the second imaging section 130 , the imaging device control box 303 corresponding to that imaging device 302 is determined to be the second image processing section 140 .

[0056] First, in S701, the user sets a specific area. In the second embodiment, the specific area is an area where a shadow or a virtual image of the object 306 may occur. Next, in S702, the user determines whether a dark shadow occurs in a specific direction. If it is determined in S702 that a dark shadow occurs in a specific direction, in S703, the user determines a part of one or more imaging devices 302 capable of capturing an area where a dark shadow may occur as the first imaging unit 110. Since each imaging device 302 captures the entire field while having an overlapping area between the angles of view, there are multiple imaging devices 302 whose angles of view include an area where a dark shadow may occur. The user determines one or more of them as the first imaging unit 110.

[0057] FIG. 8 is a diagram for explaining an example of a shadow and a virtual image of an object 306 according to the second embodiment. FIG. 8(a) shows an example of a shadow 811 occurring in an angle of view 810 of the imaging device 302. In the example shown in FIG. 8(a), a shadow 811 of an object 306R occurs as a dark shadow on the surface of the competition field 305 due to strong light from a specific direction by the lighting device 309C. FIG. 8(b) shows an example of an overhead image 812 in which FIG. 8(a) is virtually viewed from directly above. When a dark shadow 811 occurs as shown in FIG. 8(b), the shadow 811 occurs in a specific direction relative to the object 306. In FIGS. 8(a) and (b), an example is shown in which the shadow 811 occurs in one direction, but the shadow of the object 306 may occur in two or more directions depending on the number and arrangement of the lighting devices. However, when the shadow of the object 306 occurs in two or more directions, the shadow from this will not be a dark shadow. This is because the surface of the playing field 305 in the direction of the lighting device that emits strong light as viewed from the object 306 is illuminated by the light from that lighting device, and the shadow that occurs in the area in that direction is not a deep shadow.

[0058] When it is predicted that a dark shadow 811 as shown in the overhead image 812 of FIG. 8(b) will occur, the user determines a part of one or more imaging devices 302 present in the direction in which the dark shadow 811 occurs as viewed from the object 306R as the first imaging unit 110. Since it is only necessary to delete voxels corresponding to the area in which the dark shadow 811 occurs, if there is an imaging device 302 that captures the entire dark shadow 811 in the angle of view, only that imaging device 302 may be determined as the first imaging unit 110 as the part described above. In actual imaging, conditions are rarely met such that only one imaging device 302 can completely remove voxels corresponding to areas in which a dark shadow may occur. In such a case, a sufficient number of imaging devices 302 are determined as the first imaging unit 110 to delete voxels corresponding to areas in which a dark shadow may occur for the entire competition field 305.

[0059] After S703, or if it is determined in S702 that no dark shadow occurs in a particular direction, in S704, the user determines whether or not the shadow of the object 306 occurs evenly around the object 306. If it is determined in S704 that a shadow occurs evenly around the object 306, the user performs the step of S705. Specifically, in S705, the user determines a predetermined number of imaging devices 302 out of all imaging devices 302 that can capture at least a part of an area in which a shadow may occur evenly around the object 306, as the first imaging unit 110. This is to suppress the uncut voxels caused by the shadow that occurs evenly around the object 306.

[0060] FIG. 8(c) shows an example of a shadow 821 that occurs evenly around the object 306 in the angle of view 820 of the imaging device 302. In the example shown in FIG. 8(c), a shadow 821 occurs around the object 306S due to light emitted from a large number of lighting devices 309 around the competition field 305. FIG. 8(d) shows an example of an overhead image 822 that is a virtual overhead view of FIG. 8(c). As shown in FIG. 8(d), when the shadow 821 may occur evenly around the object 306, the user determines a predetermined number of imaging devices 302 as the first imaging unit 110 from among all imaging devices 302 in which at least a part of the area in which the shadow 821 may occur is included in the angle of view.

[0061] The predetermined number is determined, for example, by the angular difference between the optical axes of the imaging devices 302 and the object 306 or the imaging target space. The user determines a combination of a predetermined number of imaging devices 302 that can remove voxels corresponding to the region of the shadow 821 generated around the object 306S, and determines the imaging devices 302 included in the determined combination as the first imaging unit 110. Specifically, for example, the user determines, as the first imaging unit 110, two or three imaging devices 302 selected from among a plurality of imaging devices 302 that capture a part of the imaging target space such that the angular difference between the optical axes of the imaging devices 302 is 120 to 180 degrees. With the two or three imaging devices 302 selected in this way, it is possible to capture the shadow 821 without being blocked by the object 306S.

[0062] After S705, or if it is determined in S704 that a shadow does not occur evenly around the object 306, the user determines in S706 whether or not a virtual image occurs around the object 306. If it is determined in S706 that a virtual image occurs around the object 306, the user determines in S707 as the first imaging unit 110 a part of the one or more imaging devices 302 that can capture an image of the area where the virtual image occurs.

[0063] FIG. 8(e) shows an example of a virtual image 831 of an object 306 generated in a field of view 830 of the imaging device 302. In the example shown in FIG. 8(e), a virtual image 831 is generated by the image of an object 306T being reflected by the surface of the competition field 305. FIG. 8(f) is a cross-sectional image 832 virtually seen from the side of FIG. 8(e). The virtual image of the object 306 is generated on a line where the plane determined by the line connecting the imaging device 302 and the object 306 and the vector perpendicular to the surface of the competition field 305 intersects with the surface of the competition field 305. Therefore, the location where the virtual image is generated differs for each imaging device 302. For example, in the cross-sectional image 832, a virtual image generated by the image of the object 306T being reflected by the surface of the competition field 833 appears as a virtual image 834 in the imaging device 302Q, and as a virtual image 835 in the imaging device 302R. Therefore, false three-dimensional shape data due to a virtual image does not occur in an area corresponding to an area above the field surface 733 in the space for generating three-dimensional shape data.

[0064] However, when the target space for generating the three-dimensional shape data is set up to a region corresponding to a region below the field surface 733, false three-dimensional shape data due to a virtual image is generated in the following space in the target space for generating the three-dimensional shape data. Specifically, in this case, false three-dimensional shape data due to a virtual image is generated in a space corresponding to a region that is plane-symmetrical with respect to the field surface 733 of the region where the object 306 exists in the target space for generating the three-dimensional shape data. For example, in the cross-sectional image 832, three-dimensional shape data 836 is generated as false three-dimensional shape data due to a virtual image. Therefore, the user determines, as the first imaging unit 110, a sufficient number of imaging devices 302 to cut off the voxels of the false three-dimensional shape data generated in the space corresponding to the region below the field surface 733.

[0065] Unlike the shadow of the object 306, the virtual image of the object 306 is not blocked by the object 306. Therefore, basically, the user may determine any one of the multiple imaging devices 302 capturing an image of an area where a virtual image may occur as the first imaging unit 110 of the imaging device 302. However, if multiple objects 306 are crowded together, the virtual image of one object 306 may be hidden by another object 306. Taking such a case into consideration, the user may determine two or more imaging devices 302 of the multiple imaging devices 302 capturing an image of an area where a virtual image may occur as the first imaging unit 110.

[0066] After S707, or if it is determined in S706 that no virtual image will occur around the object 306, the user determines in S708 whether or not a shadow will occur in the area surrounded by the multiple objects 306. If it is determined in S708 that a shadow will occur, the user determines in S709 as the first imaging unit 110 a part of the one or more imaging devices 302 whose optical axis elevation angle is equal to or greater than a predetermined angle. By determining as the first imaging unit 110 an imaging device 302 whose optical axis elevation angle is equal to or greater than a predetermined angle, it is possible to suppress the remaining shavings of the 3D shape data caused by a shadow occurring in the area surrounded by the multiple objects 306.

[0067] FIG. 8(g) shows an example of a shadow 841 that occurs in an area surrounded by the object 306 in the angle of view 840 of the imaging device 302. In the example shown in FIG. 8(g), a shadow 841 occurs in an area surrounded by the objects 306U, 306V, 306W, and 306X due to light emitted from a large number of lighting devices 309 around the competition field 305. FIG. 8(h) shows an example of a cross-sectional image 842 in which FIG. 8(g) is virtually viewed from the side. The voxels corresponding to the area in which the shadow 841 exists as shown in FIG. 8(g) and FIG. 8(h) can be removed using, for example, a silhouette image corresponding to the image captured by the imaging device 302 as follows. For example, as shown in the cross-sectional image 842, the imaging device 302 has an elevation angle of the optical axis larger than an elevation angle 843 of a tangent drawn from the field surface at the center of the densely packed multiple objects 306 to the top of the object 306U.

[0068] For example, the user sets such an elevation angle as an angle threshold based on the characteristics of the object 306 or the sport to be imaged, and determines some of the one or more imaging devices 302 whose optical axis elevation angles are equal to or greater than this angle threshold as the first imaging unit 110. Specifically, the user selects two or three imaging devices 302 whose optical axis angle difference between the imaging devices 302 is 120 degrees to 180 degrees from among the multiple imaging devices 302 that image a part of the imaging target space and satisfy the above-mentioned conditions, as in the case of the shadow 821 or the virtual image 831. Furthermore, the user determines the selected two or three imaging devices 302 as the first imaging unit 110.

[0069] After S709, if it is determined in S708 that a shadow will not be generated, then in S710 the user determines whether or not a shadow will be generated below the object 306. If it is determined in S710 that a shadow will be generated below the object 306, then in S711 the user determines a part of the one or more imaging devices 302 having an elevation angle equal to or less than a predetermined angle as the first imaging unit 110. By determining a part of the one or more imaging devices 302 having an elevation angle equal to or less than a predetermined angle as the first imaging unit 110, it is possible to suppress the remaining shavings of the 3D shape data due to the shadow generated below the object 306.

[0070] FIG. 8(i) shows an example of a shadow 847 that occurs under an object 846. FIG. 8(i) shows a cross-sectional image 844 when the object 846 is viewed from the horizontal direction. Here, a shadow 847 that occurs due to a table-like object 846 will be used for explanation, instead of the object 306 used in the explanation so far. The voxels corresponding to the region where the shadow 847 occurs under the table-like object 846 can be removed using a silhouette image corresponding to the image captured by the imaging device 302 as shown below. For example, as shown in the cross-sectional image 844, the imaging device 302 has an elevation angle of the optical axis that is smaller than an elevation angle 848 of a tangent drawn from the field surface at the center of the object 846 to the end of the upper surface of the object 846.

[0071] For example, the user sets such an elevation angle as an angle threshold based on the shape of the object 846 to be imaged, and determines some of the one or more imaging devices 302 whose optical axis elevation angles are equal to or less than this angle threshold as the first imaging unit 110. As in the cases of the shadow 821, the virtual image 831, and the shadow 841, the user selects two or three imaging devices 302 whose optical axis angle difference between the imaging devices 302 is 120 degrees to 180 degrees from among the multiple imaging devices 302 that image a part of the imaged space and satisfy the above-mentioned conditions. Furthermore, the user determines the selected two or three imaging devices 302 as the first imaging unit 110.

[0072] After S711, or when it is determined in S710 that no shadow will be cast below the object 306, in S712, the user determines the remaining imaging devices 302 as the second imaging unit 130. Specifically, the user determines, as the second imaging unit 130, all of the imaging devices 302 that have not been determined as the first imaging unit 110 in the previous steps.

[0073] As described above, the image processing system 100 is configured to set an area where a false detection of an object area may occur in the background subtraction method as a specific area. Here, the specific area according to the second embodiment is an area where a shadow or a virtual image of the object 306 may occur. The image processing system 100 is configured so that captured image data output from at least a part of one or more imaging devices 302 whose angle of view includes the specific area is input to an image processing device that separates the object area using the trained model 123. Furthermore, the image processing system 100 is configured so that captured image data output from an imaging device 302 whose angle of view does not include the specific area is input to an image processing device that separates the object area using the background subtraction method. According to the image processing system 100 configured as described above, it is not necessary to perform separation processing by both the background subtraction method and the machine learning method in one image processing device, and a silhouette image capable of generating highly accurate three-dimensional shape data can be generated.

[0074] <Other embodiments> In the first and second embodiments, the image processing system 100 is applied to a soccer game as an example, but the application of the image processing system 100 is not limited to this. FIG. 9 shows an example of the application of the image processing system 100 to another imaging target. FIG. 9 shows an example of factors that may cause an object to be undetected or erroneously detected in the object region separation process by the background difference method for each imaging target. A region including these factors is set as a specific region by the same setting method as in the first and second embodiments. Furthermore, based on the specific region and the angle of view of the imaging device 302, the imaging device 302 is determined as the first imaging unit 110 or the second imaging unit 130. This makes it possible to generate a silhouette image that can generate high-precision three-dimensional shape data without chipping or remaining parts for the shape of the object 306.

[0075] In the first embodiment, as an example, a case where there is a factor that causes an object to be undetected in the object region separation process using the background difference method has been described. In the second embodiment, as an example, a case where there is a factor that causes an object to be erroneously detected in the object region separation process using the background difference method has been described. In each embodiment, an example where there is a factor that causes an object to be undetected and erroneously detected has been described separately, but the image processing system 100 can also be applied to a case where these factors coexist. In this case, first, the imaging device 302 in which there is a factor that causes an object to be undetected is determined as the first imaging unit 110 by the determination method shown in the first embodiment. Next, a part of the imaging devices 302 in which there is a factor that causes erroneous detection among the remaining imaging devices 302 is determined as the first imaging unit 110 by the determination method shown in the second embodiment, and further, the remaining imaging devices 302 are determined as the second imaging unit 130. By determining the first imaging unit 110 and the second imaging unit 130 in this manner, a silhouette image that can generate high-precision three-dimensional shape data without chipping or remaining parts in the shape of the object 306 can be generated.

[0076] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiment to each device of the image processing system 100 via a network or a storage medium, and having one or more processors in the device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0077] Furthermore, within the scope of the present disclosure, the embodiments may be freely combined, any component of each embodiment may be modified, or any component of each embodiment may be omitted.

[0078] <Configuration of the present disclosure> The present disclosure includes the following configurations and methods.

[0079] [Configuration 1] A first image processing device that generates data of a first silhouette image indicating an area in which the object exists in the first input image by inputting data of an image obtained by imaging an object using a first imaging means that images an area including at least a part of a specific area as data of a first input image into a trained model; a second image processing device that uses image data obtained by capturing an image of the object by a second imaging means different from the first imaging means as data of a second input image, and calculates a difference between the second input image and a background image, which is an image obtained by capturing an image by the second imaging means in a state in which the object does not exist in an area captured by the second imaging means, thereby generating data of a second silhouette image that indicates an area in which the object exists in the second input image; 1. An image processing system comprising:

[0080] [Configuration 2] The second imaging means images an area not including the specific area; 2. The image processing system according to claim 1,

[0081] [Configuration 3] the specific area is an area in which the object may exist in a substantially stationary state; 3. The image processing system according to configuration 1 or 2, characterized in that

[0082] [Configuration 4] The specific area is an area in which a difference between a color of the object and a color of a background of an area captured by an imaging device that is the first imaging means or the second imaging means is smaller than a predetermined standard; 4. The image processing system according to any one of configurations 1 to 3,

[0083] [Configuration 5] The specific area is an area where a shadow of the object or a virtual image due to a reflection of an image of the object may occur in a part of an area captured by an imaging device that is the first imaging means or the second imaging means; 5. The image processing system according to any one of configurations 1 to 4,

[0084] [Configuration 6] setting a part of one or more imaging devices, the part including at least a part of the specific region in an angle of view, as the first imaging means; 6. The image processing system according to any one of configurations 1 to 5,

[0085] [Configuration 7] Among two or more imaging devices whose angle of view includes at least a part of the specific region, an imaging device whose optical axis vectors between the imaging devices form an angle equal to or larger than a predetermined angle is set as the first imaging means; 7. The image processing system according to any one of configurations 1 to 6,

[0086] [Configuration 8] Among one or more imaging devices whose angle of view includes at least a part of the specific region, an imaging device in which an angle between an optical axis vector of the imaging device and a field plane of an imaging target of the imaging device is equal to or larger than a predetermined angle or equal to or smaller than a predetermined angle is set as the first imaging means; 8. The image processing system according to any one of configurations 1 to 7,

[0087] [Configuration 9] The first image processing device is a first acquisition means for acquiring image data obtained by imaging the object by the first imaging means as data of the first input image; A first generation means for generating data of the first silhouette image by inputting data of the acquired first input image into the trained model; a first output means for outputting the generated first silhouette image; having The second image processing device is a second acquisition means for acquiring image data obtained by imaging the object by the second imaging means as data of the second input image; a second generation means for generating data of the second silhouette image by calculating a difference between the acquired second input image and the background image; a second output means for outputting the generated second silhouette image; having 9. The image processing system according to any one of configurations 1 to 8,

[0088] [Configuration 10] a third image processing device that generates three-dimensional shape data indicating a shape of the object, using data of the first silhouette image generated by the first image processing device and data of the second silhouette image generated by the second image processing device;

[0033] 10. The image processing system according to any one of configurations 1 to 9,

[0089] [Configuration 11] The third image processing device is a silhouette acquisition means for acquiring data of the first silhouette image output by the first image processing device and data of the second silhouette image output by the second image processing device; a shape generating means for generating the three-dimensional shape data using the acquired data of the first silhouette image and the acquired data of the second silhouette image; having 11. The image processing system according to claim 10,

[0090] [method] A first image processing step in a first image processing device, the first image processing step including: a first acquisition step of acquiring image data obtained by imaging an object by a first imaging means that images an area including at least a part of a specific area as data of a first input image; a first step of generating data of a first silhouette image indicating an area in the first input image where the object exists by inputting the data of the first input image to a trained model; and a first output step of outputting the data of the first silhouette image; a second image processing step in a second image processing device, the second image processing step including: a second acquisition step of acquiring image data obtained by imaging the object by a second imaging means different from the first imaging means as second input image data; a second generation step of generating data of a second silhouette image showing an area in which the object exists in the second input image by calculating a difference between the second input image and a background image which is an image obtained by imaging by the second imaging means in a state in which the object does not exist in an area imaged by the second imaging means; and a second output step of outputting the data of the second silhouette image; 13. An image processing method comprising: [Explanation of symbols]

[0091] 100 Image Processing System 110 First imaging unit 120 First image processing unit 121 First input section 122 1st separation section 123 trained models 124 Transmission Unit 130 Second imaging unit 140 Second image processing unit 141 Second input section 142 Second separation section 143 Background Images 144 Transmission Unit

Claims

1. a first image processing device that generates data of a first silhouette image indicating an area in which the object exists in the first input image by inputting data of an image obtained by imaging an object using a first imaging means that images an area including at least a part of a specific area as data of a first input image into a trained model; a second image processing device that uses image data obtained by capturing an image of the object by a second imaging means different from the first imaging means as data of a second input image, and calculates a difference between the second input image and a background image, which is an image obtained by capturing an image by the second imaging means in a state in which the object is not present in an area captured by the second imaging means, thereby generating data of a second silhouette image that indicates an area in the second input image where the object exists; 1. An image processing system comprising:

2. The second imaging means images an area not including the specific area; 2. The image processing system according to claim 1,

3. the specific area is an area in which the object may exist in a substantially stationary state; 2. The image processing system according to claim 1,

4. the specific area is an area in which a difference between a color of the object and a color of a background of an area captured by an imaging device that is the first imaging means or the second imaging means is smaller than a predetermined standard; 2. The image processing system according to claim 1,

5. The specific area is an area where a shadow of the object or a virtual image due to a reflection of an image of the object may occur in a part of an area captured by an imaging device that is the first imaging means or the second imaging means; 2. The image processing system according to claim 1,

6. setting a part of one or more imaging devices having an angle of view including at least a part of the specific region as the first imaging means; 2. The image processing system according to claim 1,

7. Among two or more image capturing devices whose angle of view includes at least a part of the specific region, an image capturing device whose optical axis vectors between the image capturing devices form an angle equal to or larger than a predetermined angle is set as the first image capturing means; 2. The image processing system according to claim 1,

8. Among one or more imaging devices whose angle of view includes at least a part of the specific region, an imaging device in which an angle between an optical axis vector of the imaging device and a field plane of an imaging target of the imaging device is equal to or larger than a predetermined angle or equal to or smaller than a predetermined angle is set as the first imaging means; 2. The image processing system according to claim 1,

9. The first image processing device is a first acquisition means for acquiring image data obtained by imaging the object by the first imaging means as data of the first input image; A first generation means for generating data of the first silhouette image by inputting data of the acquired first input image into the trained model; a first output means for outputting the generated first silhouette image; having The second image processing device is a second acquisition means for acquiring image data obtained by imaging the object by the second imaging means as data of the second input image; a second generation means for generating data of the second silhouette image by calculating a difference between the acquired second input image and the background image; a second output means for outputting the generated second silhouette image; having 2. The image processing system according to claim 1,

10. a third image processing device that generates three-dimensional shape data indicating a shape of the object, using data of the first silhouette image generated by the first image processing device and data of the second silhouette image generated by the second image processing device; [0033] 2. The image processing system according to claim 1,

11. The third image processing device is a silhouette acquisition means for acquiring data of the first silhouette image output by the first image processing device and data of the second silhouette image output by the second image processing device; a shape generating means for generating the three-dimensional shape data using the acquired data of the first silhouette image and the data of the second silhouette image; having 11. The image processing system according to claim 10,

12. A first image processing step in a first image processing device, the first image processing step including: a first acquisition step of acquiring image data obtained by imaging an object by a first imaging means that images an area including at least a part of a specific area as data of a first input image; a first step of generating data of a first silhouette image indicating an area in the first input image where the object exists by inputting the data of the first input image to a trained model; and a first output step of outputting the data of the first silhouette image; a second image processing step in a second image processing device, the second image processing step including: a second acquisition step of acquiring, as second input image data, image data obtained by imaging the object by a second imaging means different from the first imaging means; a second generation step of generating second silhouette image data showing an area in which the object exists in the second input image by calculating a difference between the second input image and a background image, which is an image obtained by imaging by the second imaging means in a state in which the object does not exist in an area imaged by the second imaging means; and a second output step of outputting the second silhouette image data; 13. An image processing method comprising:

Citation Information

Patent Citations

  • Image processing device, image processing method, and program

    JP2021056960A