Information processing device, program, and information processing method

The information processing apparatus adjusts thresholds based on shielding elements to accurately estimate the three-dimensional shape of a subject using multiple cameras, overcoming occlusions and background interference.

WO2025141776A1PCT designated stage expired Publication Date: 2025-07-03MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/046966
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Conventional methods struggle to accurately estimate the three-dimensional shape of a subject when there is a shielding element between the subject and the camera, especially when the shielding element is not clearly shown in the image or indistinguishable from the background.

Method used

An information processing apparatus that generates threshold information to determine which three-dimensional objects to remove from a space, adjusting thresholds based on the presence of shielding elements, and uses a combination of cameras and image processing to estimate the subject's shape by projecting voxels onto discrimination images.

Benefits of technology

Enables reliable estimation of the three-dimensional shape of a subject even with shielding elements by adjusting thresholds and using multiple cameras to account for occlusions, ensuring accurate shape reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023046966_03072025_PF_FP_ABST
    Figure JP2023046966_03072025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device (130) is provided with: a threshold value map generation unit (135) that generates, for each of a plurality of voxels, a threshold value map that indicates a threshold value for determining whether or not to remove the voxel; a silhouette extraction unit (136) that generates, from a plurality of images obtained from a plurality of cameras, a plurality of silhouette images for distinguishing a silhouette of a subject; and an estimation unit (140) that, by counting the number of cameras corresponding to the silhouette images in which a position obtained by projecting one voxel selected from the plurality of voxels onto each of the plurality of silhouette images is included in the silhouette, and comparing the number of cameras with the threshold value, repeatedly performs processing for determining whether or not to remove the one voxel, and estimates the shape of the subject using the remaining voxels after voxels have been removed from the space. When there is a shielding element that causes shielding between a voxel and the plurality of cameras, the threshold value map generation unit (135) makes the threshold value for the voxel smaller than when no shielding element is present.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, program, and information processing method

[0001] The present disclosure relates to an information processing device, a program, and an information processing method.

[0002] A technique has been known in the past for estimating the three-dimensional shape of a subject by using images of the subject captured by cameras installed at multiple locations and removing voxels that are not part of the subject from voxels arranged in three-dimensional space. However, if an occluding element exists between the subject and the camera, it may not be possible to accurately estimate the three-dimensional shape of the subject.

[0003] The information processing device described in Patent Document 1 is capable of estimating the three-dimensional shape of a subject by integrating, based on images captured by multiple cameras, a mask for removing voxels in areas other than the subject from voxels arranged in three-dimensional space, and a mask for removing voxels in areas other than structures that act as occluding elements.

[0004] Japanese Patent Application Laid-Open No. 2019-197523

[0005] However, conventional techniques are unable to accurately estimate the three-dimensional shape of a subject in situations where it is difficult to grasp the shape of an occluding element from an image, such as when the occluding element is not clearly visible in the image or when the occluding element blends into the background in the image.

[0006] Therefore, one or more aspects of the present disclosure aim to enable reliable estimation of the three-dimensional shape of a subject when an occluding element is present between a camera and the subject.

[0007] An information processing device according to one aspect of the present disclosure includes a threshold information generating unit that generates threshold information indicating a threshold for determining whether or not to remove each of a plurality of three-dimensional objects virtually arranged so as to fill a predetermined space including the subject at a location where the subject is imaged; a distinguishing image generating unit that generates a plurality of distinguishing images that are images for distinguishing a first region that is a region of the subject and a second region that is a region other than the subject from a plurality of images obtained from a plurality of imaging devices that image the subject at different positions at the location; and a distinguishing image generating unit that generates a distinguishing image in which a position of one three-dimensional object selected from the plurality of three-dimensional objects is projected onto each of the plurality of distinguishing images and is included in the first region. and an estimation unit that counts the number of imaging devices that captured an image captured by the object and compares the count with the threshold value to determine whether or not to remove the one three-dimensional object from the space, repeating this process until all of the plurality of three-dimensional objects are selected as the one three-dimensional object, thereby identifying one or more three-dimensional objects to be removed from the space, and estimates the shape of the subject from one or more three-dimensional objects that remain after removing the one or more three-dimensional objects from the space, wherein the threshold information generation unit, when a blocking element is present that blocks a target three-dimensional object that is one of the plurality of three-dimensional objects, and at least one of the plurality of imaging devices, sets the threshold value for the target three-dimensional object to a value smaller than when the blocking element is not present.

[0008] A program according to one aspect of the present disclosure includes a computer including: a threshold information generating unit that generates threshold information indicating a threshold for determining whether or not to remove a plurality of three-dimensional objects virtually arranged so as to fill a predetermined space including the subject at a location where the subject is imaged; a distinguishing image generating unit that generates a plurality of distinguishing images that distinguish a first region that is a region of the subject from a second region that is a region other than the subject, from a plurality of images obtained from a plurality of imaging devices that image the subject at different positions at the location; and a distinguishing image generating unit that generates a distinguishing image in which a position of one three-dimensional object selected from the plurality of three-dimensional objects projected onto each of the plurality of distinguishing images is included in the first region. The method counts the number of imaging devices that captured the image used, and compares this number with the threshold value to determine whether or not to remove the one three-dimensional object from the space. This process is repeated until all of the plurality of three-dimensional objects are selected as the one three-dimensional object, thereby identifying one or more three-dimensional objects to be removed from the space, and functions as an estimation unit that estimates the shape of the subject using one or more three-dimensional objects that remain after removing the one or more three-dimensional objects from the space. The threshold information generation unit is characterized in that, when there is a blocking element that blocks a target three-dimensional object that is one of the plurality of three-dimensional objects, and at least one of the plurality of imaging devices, sets the threshold value for the target three-dimensional object to be smaller than when there is no blocking element.

[0009] An information processing method according to one aspect of the present disclosure includes generating threshold information indicating a threshold for determining whether to remove each of a plurality of three-dimensional objects virtually arranged to fill a predetermined space including the subject at a location where the subject is imaged, generating a plurality of distinguishing images that are images for distinguishing a first region that is a region of the subject from a second region that is a region other than the subject from a plurality of images obtained from a plurality of imaging devices that image the subject at different positions at the location, and calculating an image used to generate the distinguishing images in which a position of one three-dimensional object selected from the plurality of three-dimensional objects projected onto each of the plurality of distinguishing images is included in the first region. The information processing method involves counting the number of imaging devices that have taken an image, comparing this number with the threshold value, and repeating the process of determining whether or not to remove the one three-dimensional object from the space until all of the multiple three-dimensional objects have been selected as the one three-dimensional object, thereby identifying one or more three-dimensional objects to remove from the space, and estimating the shape of the subject from one or more three-dimensional objects that remain after removing the one or more three-dimensional objects from the space, and is characterized in that when there is an occluding element that blocks the gap between a target three-dimensional object that is one of the multiple three-dimensional objects and at least one of the multiple imaging devices, the threshold value for the target three-dimensional object is made smaller than when the occluding element is not present.

[0010] According to one or more aspects of the present disclosure, it is possible to reliably estimate the three-dimensional shape of a subject when an occluding element is present between a camera and the subject.

[0011] FIG. 1 is a block diagram schematically showing the configuration of an information processing system according to Embodiments 1 to 3. FIG. 2 is a block diagram schematically showing the configuration of an information processing device according to Embodiment 1. FIG. 3 is a schematic diagram for explaining a method for generating a threshold map. (A) and (B) are schematic diagrams showing an example of a case where it is determined whether or not to remove a voxel using a uniform threshold. (A) and (B) are schematic diagrams showing an example of a case where it is determined whether or not to remove a voxel by varying the threshold. FIG. 4 is a schematic diagram showing an overview of processing in a coordinate transformation unit. FIG. 5 is a block diagram schematically showing the configuration of a PC. FIG. 6 is a flowchart showing the operation of the information processing device according to Embodiment 1. FIG. 7 is a flowchart showing the operation of the information processing device according to Embodiment 2. FIG. 8 is a schematic diagram for explaining a location where an image of a subject is captured. FIG. 9 is a block diagram schematically showing the configuration of an information processing device according to Embodiment 3. FIG. 10 is a flowchart showing the operation of the information processing device according to Embodiment 3.

[0012] 1 is a block diagram showing a schematic configuration of an information processing system 100 according to embodiment 1. The information processing system 100 includes a plurality of cameras 110 and an information processing device 130. The plurality of cameras 110 and the information processing device 130 are connected to a network 101.

[0013] The camera 110 is an imaging device that captures images including a subject whose three-dimensional shape is to be estimated. The cameras 110 capture images at different positions within a location where the subject is to be captured. The images captured by the camera 110 are sent to the information processing device 130 via the network 101.

[0014] 2 is a block diagram schematically showing the configuration of information processing device 130 according to embodiment 1. Information processing device 130 includes receiving unit 131, camera parameter acquisition unit 132, camera parameter storage unit 133, structure data storage unit 134, threshold map generation unit 135, silhouette extraction unit 136, coordinate conversion unit 137, object model generation unit 138, and output unit 139. Here, coordinate conversion unit 137 and object model generation unit 138 form estimation unit 140.

[0015] The receiving unit 131 receives various data from the network 101. For example, the receiving unit 131 receives image data from the camera 110. The received image data is provided to the camera parameter acquisition unit 132 during calibration, and to the silhouette extraction unit 136 at other times. The receiving unit 131 also receives structure data, which will be described later. The received structure data is stored in the structure data storage unit 134.

[0016] The camera parameter acquisition unit 132 calculates camera parameters indicating the position and orientation of the camera 110 from image data captured by the camera 110 , and stores the calculated camera parameters in the camera parameter storage unit 133 .

[0017] Here, the camera parameters are composed of extrinsic parameters and intrinsic parameters. The extrinsic parameters are composed of a rotation matrix and a translation matrix, and indicate the position and orientation of the camera 110. The intrinsic parameters include information such as the focal length and optical center of the camera 110, and indicate the angle of view of the camera 110, the size of the imaging sensor, etc.

[0018] The process of calculating the camera parameters is called calibration, and the camera parameters can be determined by using the correspondence between points in a three-dimensional world coordinate system obtained using multiple image data captured by camera 110 of a specific pattern such as a checkerboard, and corresponding points on a two-dimensional system.

[0019] The camera parameter storage unit 133 stores the camera parameters acquired by the camera parameter acquisition unit 132 .

[0020] The structure data storage unit 134 stores structure data. The structure data is transmitted from, for example, another device (not shown) capable of using CAD (Computer-Aided Design) and received by the receiving unit 131. Specifically, the structure data is data indicating the position and shape of static elements, which are elements such as stationary objects that are placed at a location where an image of a subject is captured. Here, the static elements include objects that move at their locations, such as robots, as long as their locations do not change.

[0021] The threshold map generation unit 135 generates a threshold map as threshold information indicating a threshold for determining whether to remove a voxel when estimating the three-dimensional shape of the subject in the subject model generation unit 138. For example, the threshold map generation unit 135 is a threshold information generation unit that generates threshold information indicating a threshold for determining whether to remove a voxel for each of multiple voxels, which are three-dimensional objects virtually arranged to fill a predetermined space including the subject at a location where the subject is imaged. Here, when an occluding element exists between a target voxel, which is one of the multiple voxels filling the space including the subject, and at least one of the multiple cameras 110, the threshold map generation unit 135 sets a lower threshold for the target voxel than when the occluding element does not exist. For example, the threshold map generation unit 135 may set the threshold for the target voxel to the number of cameras 110 that are not occluded by an occluding element between the target voxel and the target voxel, among the multiple cameras 110 that image the subject. The threshold map generation unit 135 determines whether a static element, which is a stationary element placed at a location where an image of a subject is captured, is an occluding element by referring to structural data indicating the position and shape of the static element.

[0022] 3 is a schematic diagram for explaining a method for generating a threshold map. As shown in FIG. 3, the following explanation will be given taking an example in which an object 150 is imaged by four cameras 110A to 110D.

[0023] 3, a shielding element 151 is disposed between camera 110A and subject 150. Therefore, when camera 110A captures an image including subject 150, a portion of subject 150 is shielded by shielding element 151, and the entire subject 150 is not included in the image. However, no object that acts as a shielding element exists between the images captured by cameras 110B to 110D and subject 150, and the images captured by cameras 110B to 110D include the entire subject 150.

[0024] 4A and 4B are schematic diagrams showing an example of a case where it is determined whether or not to remove a voxel using a uniform threshold value without performing processing in the threshold map generating unit 135.

[0025] As shown in FIG. 4A, in order to estimate the three-dimensional shape of the subject 150, a large number of cubic voxels 152 shown in FIG. 4A are placed at the position where the subject 150 is located.

[0026] 4A indicates the number of cameras 110 that captured an image in which the voxel is included in the silhouette of the subject 150. For example, a voxel 152 indicated with the number "4" indicates that the voxel is present in the silhouette of the subject 150 in an image captured by four cameras 110.

[0027] In the example shown in FIG. 4(A), if the threshold for whether or not to remove a voxel is uniformly set to "4", as shown in FIG. 4(B), the voxel that should be left as the subject 150 is actually removed because it is occluded by the occluding element 151.

[0028] To avoid the situation shown in Figures 4(A) and (B), the threshold map generation unit 135 sets a smaller threshold for voxels 152 where an occluding element 151 exists between the camera 110 and the voxel 152 than for voxels 152 where an occluding element 151 does not exist.

[0029] The threshold map generation unit 135 sets the threshold for each voxel 152 to the number of cameras 110 that have no occluding elements 151 between them, as shown in the numbers in the voxels 152 in Figure 5 (A), and can accurately estimate the three-dimensional shape of the subject 150 without removing voxels 152 that have occluding elements 151 between them and the camera 110A, as shown in Figure 5 (B).

[0030] 2 , as described above, the threshold map generation unit 135 determines, for each voxel, a threshold for determining whether or not to remove that voxel. Here, the threshold map generation unit 135 sets a smaller threshold for a voxel that has an occluding element between it and the camera 110 than for a voxel that has no occluding element between it and the camera 110. Specifically, for a voxel that has no occluding element between it and the camera 110, the threshold is set to the number of cameras whose angle of view is the subject 150, and for a voxel that has an occluding element between it and the camera 110, the threshold is set to a value obtained by subtracting the number of cameras whose angle of view is the subject 150 and the number of cameras whose angle of view is the subject 150.

[0031] The silhouette extraction unit 136 is a distinguishing image generation unit that generates a plurality of silhouette images, which are distinguishing images that distinguish a first area as a silhouette, which is the area of ​​the subject, from a second area, which is the area other than the subject, from each of a plurality of images represented by a plurality of image data obtained from a plurality of cameras 110.

[0032] For example, the silhouette extraction unit 136 extracts a silhouette, which is a region of a subject, from the image data for each camera 110. Specifically, the silhouette extraction unit 136 may store background image data indicating a background image in advance and extract a silhouette by subtracting the background image from the background image. Alternatively, the silhouette extraction unit 136 may divide an object region by segmentation and extract a region corresponding to the subject as a silhouette from the divided object region. The silhouette image data indicating the extracted silhouette is provided to the coordinate conversion unit 137. The silhouette image data may be binary data with the same resolution as the image indicated by the image data from the camera 110, with the silhouette portion being "1" and the other portions being "0".

[0033] The estimation unit 140 counts the number of cameras 110 that captured images used to generate silhouette images in which the position of a voxel selected from the plurality of voxels projected onto each of the plurality of silhouette images is included in the silhouette, and compares the count with the threshold value of the voxel indicated in the threshold map to determine whether to remove the voxel from the space filled with the plurality of voxels. This process is repeated until all of the plurality of voxels filling the space are selected, thereby identifying one or more three-dimensional objects to remove from the space. The estimation unit 140 then estimates the shape of the subject from one or more three-dimensional objects remaining after removing the identified one or more three-dimensional objects from the space. The estimation unit 140 is composed of a coordinate conversion unit 137 and a subject model generation unit 138.

[0034] The coordinate transformation unit 137 transforms the positions of the voxels into the camera coordinate system based on the camera parameters stored in the camera parameter storage unit 133. In other words, the coordinate transformation unit 137 projects the three-dimensional positions of the voxels onto the positions of the two-dimensional image.

[0035] 6 is a schematic diagram showing an outline of the processing in the coordinate conversion unit 137. The coordinate conversion unit 137 converts the position of a voxel (X w , Y w , Z w) is converted into a position (u, v) in image coordinates, which is the camera coordinate system, using an external parameter indicating rotation, an external parameter indicating translation, and an internal parameter. The camera coordinate system here is the coordinate system of the camera 110. By such a conversion, the position of the voxel can be converted into the position of the h-louette image.

[0036] The object model generating unit 138 generates an object model that indicates the three-dimensional shape of the object. Here, the object model generating unit 138 generates the object model that indicates the three-dimensional shape of the object using, for example, a Space Cavity Method.

[0037] Specifically, subject model generation unit 138 determines, for each voxel, whether the position of the voxel in the camera coordinate system is included in a silhouette generated based on image data from the corresponding camera 110. This allows subject model generation unit 138 to count the number of cameras 110 included in the silhouette for each voxel. Then, by referring to the threshold map generated by threshold map generation unit 135, subject model generation unit 138 can remove voxels that are less than the threshold set for each voxel, and estimate the three-dimensional shape of the subject from the remaining voxels.

[0038] Note that the object model generating unit 138 may generate, as the object model, data indicating the three-dimensional shape of the object, or may generate data obtained by rendering the three-dimensional shape of the object.

[0039] The output unit 139 outputs the subject model. As a method of output, data of the subject model may be transmitted to an external device, the data of the subject model may be recorded on a portable recording medium, or an image rendered as the subject model may be displayed.

[0040] The information processing device 130 described above can be realized by, for example, a computer such as the PC 10 shown in Fig. 7. The PC 10 includes a storage 11 such as a hard disk drive (HDD) and a solid state drive (SSD), a memory 12, a processor 13 such as a central processing unit (CPU), a communication interface (I / F) 14 such as a network interface card (NIC), an input interface 15 such as a keyboard and a mouse, and a display 16. Although not shown, the PC 10 may also include a connection I / F for connecting a portable memory such as a universal serial bus (USB).

[0041] For example, camera parameter storage unit 133 and structure data storage unit 134 can be realized by storage 11 or memory 12. Receiving unit 131 can be realized by communication I / F 14. Camera parameter acquisition unit 132, threshold map generation unit 135, silhouette extraction unit 136, coordinate conversion unit 137, and object model generation unit 138 can be realized by processor 13 reading a program stored in storage 11 into memory 12 and executing the program. Output unit 139 can be realized by communication I / F 14, display 16, or a connection I / F (not shown), depending on how the object model is output.

[0042] The above program may be downloaded to the storage 11 from a recording medium (not shown) via a reader / writer (not shown) or from the network 101 via the communication I / F 14, and then loaded onto the memory 12 and executed by the processor 13. Alternatively, the program may be directly loaded onto the memory 12 from a recording medium via the reader / writer or from the network 101 via the communication I / F 14, and then executed by the processor 13. In other words, the program may be provided by a program product such as a recording medium.

[0043] 8 is a flowchart showing the operation of the information processing device 130 in Embodiment 1. It is assumed here that the camera parameters have already been stored in the camera parameter storage unit 133, and the structure data have already been stored in the structure data storage unit 134.

[0044] First, threshold map generation unit 135 references the structure data stored in structure data storage unit 134 and generates, for each voxel, a threshold map indicating a threshold for determining whether or not to remove that voxel (S10). Here, threshold map generation unit 135 sets, for each voxel, the threshold value as the number of cameras 110 that have that voxel within their angle of view and that have no occluding elements between them. The generated threshold map is provided to subject model generation unit 138.

[0045] The silhouette extraction unit 136 receives the image data from the receiving unit 131, extracts a silhouette, which is the area of ​​the subject, from the image data for each camera 110, and generates silhouette image data that shows a silhouette image for distinguishing the extracted silhouette from other areas (S11). The generated silhouette image data is provided to the coordinate conversion unit 137.

[0046] Next, the coordinate transformation unit 137 sequentially selects one voxel as a target voxel from the multiple voxels arranged in the space containing the subject, and repeatedly executes the processes of steps S12 to S17 until all voxels have been processed as target voxels (S12).

[0047] Coordinate conversion unit 137 converts the position of the target voxel into the camera coordinate system of each camera 110 based on the camera parameters stored in camera parameter storage unit 133 (S13). Coordinate conversion unit 137 then provides silhouette image data and correspondence information indicating the correspondence between the three-dimensional position of the target voxel (in other words, the position in the world coordinate system) and the two-dimensional position of the target voxel (in other words, the position in the camera coordinate system) to subject model generation unit 138.

[0048] The subject model generation unit 138 counts the number of cameras whose silhouettes contain the target voxel by determining, for each camera, whether the two-dimensional position of the target voxel is contained within the silhouette shown in the silhouette image (S14).

[0049] Then, the subject model generation unit 138 compares the number counted in step S14 with the threshold value of the target voxel indicated in the threshold map to determine whether the number is less than the threshold value (S15). If the number is less than the threshold value (YES in S15), the process proceeds to step S16, and if the number is equal to or greater than the threshold value (NO in S15), the process proceeds to step S17.

[0050] In step S16, the object model generation unit 138 removes the target voxel from the space, and the process then proceeds to step S17.

[0051] In step S17, the subject model generation unit 138 determines whether all voxels have been selected as target voxels and whether or not to remove the target voxels, and if there are voxels that have not yet been selected as target voxels, the process returns to S12. On the other hand, if all voxels have been selected as target voxels and whether or not to remove the target voxels has been determined, the process proceeds to step S18.

[0052] In step S18, the subject model generation unit 138 estimates the three-dimensional shape of the subject using the voxels that remain in the space without being removed, and provides a subject model representing the three-dimensional shape to the output unit 139, which then outputs the subject model.

[0053] As described above, according to the first embodiment, a threshold map indicating the threshold for whether or not to remove voxels can be generated from structural data indicating the position of an object in space. Therefore, even if an occluding element is not clearly captured in the image data, the three-dimensional shape of the subject can be reliably estimated, excluding the influence of such an occluding element.

[0054] 1, an information processing system 200 according to the second embodiment includes a plurality of cameras 110 and an information processing device 230. The plurality of cameras 110 of the information processing system 200 according to the second embodiment are similar to the plurality of cameras 110 of the information processing system 100 according to the first embodiment.

[0055] In the second embodiment, a case where a dynamic object is included in the occluding element will be described.

[0056] 9 is a block diagram schematically showing the configuration of information processing device 230 in embodiment 2. Information processing device 230 includes receiving unit 231, camera parameter acquisition unit 132, camera parameter storage unit 133, structure data storage unit 134, threshold map generation unit 235, silhouette extraction unit 136, coordinate conversion unit 137, object model generation unit 238, output unit 139, dynamic model storage unit 241, and dynamic position acquisition unit 242. In embodiment 2, coordinate conversion unit 137 and object model generation unit 238 form an estimation unit 240.

[0057] The camera parameter acquisition unit 132, camera parameter storage unit 133, structure data storage unit 134, silhouette extraction unit 136, coordinate conversion unit 137 and output unit 139 of the information processing device 230 in embodiment 2 are similar to the camera parameter acquisition unit 132, camera parameter storage unit 133, structure data storage unit 134, silhouette extraction unit 136, coordinate conversion unit 137 and output unit 139 of the information processing device 130 in embodiment 1.

[0058] The receiving unit 231 receives various data from the network 101. For example, the receiving unit 231 receives image data from the camera 110. The received image data is provided to the camera parameter acquisition unit 132 during calibration, and to the silhouette extraction unit 136 at other times. The receiving unit 231 also receives structure data. The received structure data is stored in the structure data storage unit 134.

[0059] The receiving unit 231 receives a dynamic model, which will be described later. The received dynamic model is stored in the dynamic model storage unit 241.

[0060] The dynamic model storage unit 241 stores a dynamic model, which is data indicating the shape of a dynamic element, which is an element that moves at a location where the subject is imaged by the camera 110.

[0061] The dynamic position acquisition unit 242 acquires dynamic position information indicating the positions of moving dynamic elements in a location where the cameras 110 capture images of the subject. Here, the dynamic position acquisition unit 242 acquires dynamic position information at an image capture time, which is the time when the multiple cameras 110 capture images of the subject. Note that the image capture time can be specified, for example, by metadata included in the image data. The dynamic position information is provided to the threshold map generation unit 235.

[0062] A known method may be used to acquire the dynamic position information. For example, the position of a dynamic element such as an automatic guided vehicle (AGV) may be identified by reading an AR (Argumented Reality) marker attached to the dynamic element from image data captured by the camera 110. Alternatively, the dynamic element may identify its own position using a global positioning system (GPS) or the like and transmit position information indicating the identified position to the information processing device 230. Furthermore, the dynamic position acquisition unit 242 may store a schedule for the movement of the dynamic element, and the position of the dynamic element at the time of image capture may be identified based on the schedule.

[0063] The threshold map generation unit 235 generates a threshold map as threshold information indicating thresholds for determining whether or not to remove voxels when estimating the three-dimensional shape of the subject in the subject model generation unit 238. In the second embodiment, the threshold map generation unit 235 can determine whether or not a dynamic element is an occluding element by referring to the dynamic model and dynamic position information.

[0064] For example, the threshold map generating unit 235 generates a threshold map by assuming that, at the time when the subject is imaged, the dynamic element having the shape stored in the dynamic model storing unit 241 is present at the position indicated by the dynamic position information from the dynamic position acquiring unit 242. The method of generating the threshold map itself is the same as in the first embodiment.

[0065] The object model generating unit 238 generates an object model that indicates the three-dimensional shape of the object. Here, the object model generating unit 238 also generates the object model that indicates the three-dimensional shape of the object using, for example, a Space Cavity Method.

[0066] The method for generating the subject model itself is the same as in embodiment 1, but the subject model generation unit 238 generates the subject model using a threshold map generated according to the position where the dynamic element exists at the imaging time, which is the time when the image data corresponding to the silhouette image is captured.

[0067] The information processing device 230 described above can also be realized by a computer such as the PC 10 shown in FIG.

[0068] For example, the dynamic model storage unit 241 can also be realized by the storage 11 or the memory 12. The dynamic position acquisition unit 242 can also be realized by the processor 13 reading a program stored in the storage 11 into the memory 12 and executing the program.

[0069] 10 is a flowchart showing the operation of the information processing device 230 in embodiment 2. It is also assumed here that the camera parameters have already been stored in the camera parameter storage unit 133 and the structure data has already been stored in the structure data storage unit 134.

[0070] First, the dynamic position acquisition unit 242 acquires dynamic position information indicating the position where the dynamic element exists at the time when the subject is imaged by the camera 110 (S20).

[0071] Next, the threshold map generation unit 235 refers to the structure data stored in the structure data storage unit 134 and the dynamic model stored in the dynamic model storage unit 241, and generates a threshold map indicating a threshold for determining whether or not to remove each voxel, assuming that the dynamic element is present at the position indicated by the dynamic position information (S21). The generated threshold map is provided to the subject model generation unit 238.

[0072] The silhouette extraction unit 136 receives the image data from the receiving unit 131, extracts a silhouette, which is the area of ​​the subject, from the image data for each camera 110, and generates silhouette image data (S22). The image data here represents the image captured at the capture time. The generated silhouette image data is provided to the coordinate conversion unit 137.

[0073] Next, the coordinate transformation unit 137 sequentially selects one voxel as a target voxel from the multiple voxels arranged in the space containing the subject, and repeatedly executes the processes of steps S23 to S28 until all voxels have been processed as target voxels (S23).

[0074] Coordinate conversion unit 137 converts the position of the target voxel into the camera coordinate system of each camera 110 based on the camera parameters stored in camera parameter storage unit 133 (S24). Then, coordinate conversion unit 137 provides silhouette image data and correspondence information indicating the correspondence between the three-dimensional position of the target voxel and the two-dimensional position of the target voxel to subject model generation unit 238.

[0075] The subject model generation unit 238 counts the number of cameras whose silhouettes contain the target voxel by determining for each camera whether the two-dimensional position of the target voxel is contained within the silhouette shown in the silhouette image (S25).

[0076] Then, the subject model generation unit 238 compares the number counted in step S25 with the threshold value of the target voxel indicated in the threshold map corresponding to the imaging time to determine whether the number is less than the threshold value (S26). If the number is less than the threshold value (YES in S26), the process proceeds to step S27, and if the number is equal to or greater than the threshold value (NO in S26), the process proceeds to step S28.

[0077] In step S27, the object model generation unit 238 removes the target voxel from the space, and the process then proceeds to step S28.

[0078] In step S28, the subject model generation unit 238 determines whether all voxels have been selected as target voxels and whether or not the target voxels have been removed, and if there are any voxels that have not yet been selected as target voxels, the process returns to S23. On the other hand, if all voxels have been selected as target voxels and whether or not the target voxels have been removed has been determined, the process proceeds to step S29.

[0079] In step S29, the subject model generation unit 238 estimates the three-dimensional shape of the subject using the voxels that remain in the space without being removed, and provides a subject model representing that three-dimensional shape to the output unit 139, which then outputs the subject model.

[0080] As described above, according to the second embodiment, a dynamic occluding element having a shape indicated by a dynamic model can be placed at a position indicated by dynamic position information indicating the position of the dynamic occluding element in space, and a threshold map indicating a threshold for whether or not to remove a voxel can be generated. Therefore, even if a moving occluding element is included, the influence of such an occluding element can be removed and the three-dimensional shape of the subject can be reliably estimated.

[0081] 1, an information processing system 300 according to the third embodiment includes a plurality of cameras 110 and an information processing device 330. The plurality of cameras 110 of the information processing system 300 according to the third embodiment are similar to the plurality of cameras 110 of the information processing system 100 according to the first embodiment.

[0082] In the third embodiment, as shown in Fig. 11, a case will be described in which an image captured by a plurality of cameras 110 includes a subject 350A whose shape is to be estimated and a subject 350B whose shape is not to be estimated. For example, a user of information processing system 300 may select subject 350B whose shape is not to be estimated as a subject that the user wishes to erase from the image. Here, subject 350B whose shape is not to be estimated is also referred to as a non-target subject.

[0083] 12 is a block diagram schematically showing the configuration of information processing device 330 in embodiment 3. Information processing device 330 includes receiving unit 331, camera parameter acquisition unit 132, camera parameter storage unit 133, structure data storage unit 134, threshold map generation unit 335, silhouette extraction unit 136, coordinate conversion unit 137, object model generation unit 138, output unit 339, eliminated region model storage unit 343, non-target position acquisition unit 344, and model integration unit 345. In embodiment 3, coordinate conversion unit 137 and object model generation unit 138 form estimation unit 140.

[0084] Camera parameter acquisition unit 132, camera parameter storage unit 133, structure data storage unit 134, silhouette extraction unit 136, coordinate conversion unit 137, and object model generation unit 138 of information processing device 330 in embodiment 3 are similar to camera parameter acquisition unit 132, camera parameter storage unit 133, structure data storage unit 134, silhouette extraction unit 136, coordinate conversion unit 137, and object model generation unit 138 of information processing device 130 in embodiment 1. However, in embodiment 3, object model generation unit 138 provides an object model to model integration unit 345.

[0085] The receiving unit 331 receives various data from the network 101. For example, the receiving unit 331 receives image data from the camera 110. The received image data is provided to the camera parameter acquisition unit 132 during calibration, and to the silhouette extraction unit 136 at other times. The receiving unit 331 also receives structure data. The received structure data is stored in the structure data storage unit 134.

[0086] The receiving unit 331 receives an erasure area model that indicates the shape of a three-dimensional area from which a non-target subject is to be erased. The received erasure area model is stored in the erasure area model storage unit 343. Here, the erasure area model may be data that indicates the three-dimensional shape of the non-target subject, or may be a predetermined three-dimensional shape that can accommodate the non-target subject therein, such as a rectangular parallelepiped or a cylinder. The erasure area model storage unit 343 stores the erasure area model.

[0087] The non-target position acquisition unit 344 acquires non-target position information indicating the positions of non-target subjects, which are subjects to be erased, in the space captured by the camera 110. The non-target position information is provided to the threshold map generation unit 335 and the model integration unit 345.

[0088] A known method may be used to acquire the non-target position information. For example, the position of a non-target subject may be identified by reading an AR marker attached to the non-target subject from image data captured by the camera 110. Alternatively, the non-target subject may identify its own position using a GPS or the like and transmit position information indicating the identified position to the information processing device 330.

[0089] The threshold map generation unit 335 generates a threshold map as threshold information indicating thresholds used to determine whether to remove voxels when estimating the three-dimensional shape of the subject in the subject model generation unit 138. In the third embodiment, when a non-target subject is included in the location where the subject is imaged by the cameras 110, the threshold map generation unit 335 sets the thresholds of voxels in a three-dimensional region including at least the non-target subject to be greater than the number of the multiple cameras 110. Here, the threshold map generation unit 335 may specify the three-dimensional region by referring to the eliminated region model and the non-target position information. Specifically, the three-dimensional region may be specified by placing the three-dimensional shape indicated by the eliminated region model at the position indicated by the non-target position information.

[0090] For example, the method for generating the threshold map itself is the same as in embodiment 1. However, in embodiment 3, the threshold map generation unit 335 generates a threshold map in the same way as in embodiment 1, but for voxels at positions indicated by non-target position information and included in the three-dimensional shape indicated by the eliminated area model, the threshold is set to be larger than the maximum value of the thresholds for the other voxels. In other words, the threshold for voxels in the three-dimensional area to be eliminated as non-target subjects is set to a value larger than the total number of cameras 110.

[0091] In the third embodiment, the threshold value of voxels in the three-dimensional region to be eliminated as non-target subjects is greater than the total number of cameras 110, and therefore, even if object model generation unit 138 counts the number of cameras 110 whose silhouettes contain the projection positions of voxels, the number will never exceed the threshold value. For this reason, voxels in the three-dimensional region to be eliminated as non-target subjects are always removed, and the three-dimensional shapes of the non-target subjects are not generated as object models.

[0092] The model integration unit 345 generates an integrated model by adding a predetermined replacement model to the subject model from the subject model generation unit 138 at the position indicated by the non-target position information from the non-target position acquisition unit 344, and provides the integrated model to the output unit 339.

[0093] Here, the replacement model can be, for example, CAD data of the non-target subject, which can be a more precise model than the model identified from the image data. Furthermore, by using a replacement model that is different from the non-target subject, even if the image data contains a non-target subject that the user does not want to disclose, the non-target subject can be replaced with another object, person, etc. The user of the information processing device 330 can set such a replacement model in the model integration unit 345 in advance.

[0094] The output unit 339 outputs the integrated model. As a method of output, data of the subject model and the replacement model may be transmitted to an external device, data of the subject model and the replacement model may be recorded on a portable recording medium, or images rendered as the subject model and the replacement model may be displayed.

[0095] The information processing device 330 described above can also be realized by a computer such as the PC 10 shown in FIG.

[0096] For example, the erasure area model storage unit 343 can also be realized by the storage 11 or the memory 12. In addition, the non-target position acquisition unit 344 and the model integration unit 345 can also be realized by the processor 13 reading a program stored in the storage 11 into the memory 12 and executing the program.

[0097] 13 is a flowchart showing the operation of the information processing device 330 in embodiment 3. It is also assumed here that the camera parameters have already been stored in the camera parameter storage unit 133 and the structure data has already been stored in the structure data storage unit 134.

[0098] First, the non-target position acquisition unit 344 acquires non-target position information indicating the position where the non-target subject is present (S30).

[0099] Next, the threshold map generation unit 335 refers to the structure data stored in the structure data storage unit 134 to identify, for each voxel, a threshold value for determining whether or not to remove that voxel, and also refers to the erasure area model stored in the erasure area model storage unit 343 to place a three-dimensional shape indicated by the erasure area model at the position indicated by the non-target position information, and generates a threshold map by making the threshold of the voxels included in the placed three-dimensional shape larger than the total number of cameras 110 (S31). The generated threshold map is provided to the subject model generation unit 138.

[0100] The silhouette extraction unit 136 receives the image data from the receiving unit 131, and extracts a silhouette, which is the area of ​​the subject, from the image data for each camera 110 to generate silhouette image data (S32). The generated silhouette image data is provided to the coordinate conversion unit 137.

[0101] Next, the coordinate transformation unit 137 sequentially selects one voxel as a target voxel from the multiple voxels arranged in the space containing the subject, and repeatedly executes the processes of steps S33 to S38 until all voxels have been processed as target voxels (S33).

[0102] Coordinate conversion unit 137 converts the position of the target voxel into the camera coordinate system of each camera 110 based on the camera parameters stored in camera parameter storage unit 133 (S34). Coordinate conversion unit 137 then provides silhouette image data and correspondence information indicating the correspondence between the three-dimensional position of the target voxel and the two-dimensional position of the target voxel to subject model generation unit 138.

[0103] The subject model generation unit 138 counts the number of cameras whose silhouettes contain the target voxel by determining, for each camera, whether the two-dimensional position of the target voxel is contained within the silhouette shown in the silhouette image (S35).

[0104] Then, the subject model generation unit 138 compares the number counted in step S35 with the threshold value of the target voxel indicated in the threshold map corresponding to the imaging time to determine whether the number is less than the threshold value (S36). If the number is less than the threshold value (YES in S36), the process proceeds to step S37, and if the number is equal to or greater than the threshold value (NO in S36), the process proceeds to step S38.

[0105] In step S37, the object model generation unit 138 removes the target voxel from the space, and the process then proceeds to step S38.

[0106] In step S38, the subject model generation unit 138 determines whether all voxels have been selected as target voxels and whether or not to remove the target voxels, and if there are any voxels that have not yet been selected as target voxels, the process returns to S33. On the other hand, if all voxels have been selected as target voxels and whether or not to remove the target voxels has been determined, the process proceeds to step S39.

[0107] In step S39, object model generation unit 238 estimates the three-dimensional shape of the object from the voxels that remain in the space without being removed, and provides an object model indicating the three-dimensional shape to model integration unit 345. Model integration unit 345 generates an integrated model by adding a predetermined replacement model to the object model from object model generation unit 138 at the position indicated by the non-target position information from non-target position acquisition unit 344, and provides the integrated model to output unit 339.

[0108] Next, the output unit 139 outputs the integrated model (S40).

[0109] As described above, according to the third embodiment, when there is an object to be erased in a space, by assigning a large threshold value to voxels included in the three-dimensional shape represented by the erasure area model at the position indicated by the non-target position information indicating the position of the object, even if such an object is included, it is possible to remove such an object and reliably estimate the three-dimensional shape of the required object.

[0110] In the third embodiment described above, a replacement model is added to the subject model, but the third embodiment is not limited to this example, and only the subject model may be output from output unit 339. In such a case, model integration unit 345 is not necessary.

[0111] 100, 200, 300 Information processing system, 110 Camera, 130, 230, 330 Information processing device, 131, 231, 331 Receiving unit, 132 Camera parameter acquisition unit, 133 Camera parameter storage unit, 134 Structure data storage unit, 135, 235, 335 Threshold map generation unit, 136 Silhouette extraction unit, 137 Coordinate conversion unit, 138, 238 Subject model generation unit, 139, 339 Output unit, 241 Dynamic model storage unit, 242 Dynamic position acquisition unit, 343 Erasure area model storage unit, 344 Non-target position acquisition unit, 345 Model integration unit.

Claims

1. A threshold information generation unit that generates threshold information indicating a threshold for determining whether to remove each of a plurality of three-dimensional objects virtually arranged so as to fill a predetermined space including the subject at a location where the subject is imaged; A discrimination image generation unit that generates a plurality of discrimination images that are images for discriminating a first region that is a region of the subject and a second region that is a region other than the subject from each of a plurality of images obtained from a plurality of imaging devices that image the subject at different positions at the location; Count the number of imaging devices that have imaged the image used to generate a discrimination image in which the position where one three-dimensional object selected from the plurality of three-dimensional objects is projected onto each of the plurality of discrimination images is included in the first region, and compare the number with the threshold value. By repeating the process of determining whether to remove the one three-dimensional object from the space until all of the plurality of three-dimensional objects are selected as the one three-dimensional object, one or more three-dimensional objects to be removed from the space are specified, and the shape of the subject is estimated by one or more remaining three-dimensional objects after removing the one or more three-dimensional objects from the space. The threshold information generation unit reduces the threshold value of the target three-dimensional object when there is a shielding element that shields between the target three-dimensional object, which is one of the plurality of three-dimensional objects, and at least one of the plurality of imaging devices, compared to the case where the shielding element does not exist. An information processing apparatus characterized by the above.

2. The threshold information generation unit determines whether the static element becomes the shielding element by referring to structure data indicating the position and shape of the static element, which is an element arranged at the location and is stationary. The information processing apparatus according to claim 1, characterized by the above.

3. The threshold information generation unit determines whether the dynamic element becomes the shielding element by referring to a dynamic model indicating the shape of the dynamic element, which is an element that moves at the location, and dynamic position information indicating the position of the dynamic element when imaging the subject. The information processing apparatus according to claim 1 or 2, characterized by the above.

4. The threshold information generation unit increases the threshold of the three-dimensional object in a three-dimensional region including at least the non-target object to be greater than the number of the plurality of imaging devices when there is a non-target object that is different from the subject and does not have its shape estimated at the location. The information processing apparatus according to any one of claims 1 to 3, characterized in that.

5. The threshold information generation unit specifies the three-dimensional region by referring to an erasure region model indicating the shape of the three-dimensional region and non-target position information indicating the position of the non-target object. The information processing apparatus according to claim 4, characterized in that.

6. The threshold information generation unit sets the threshold of the target three-dimensional object to the number of imaging devices among the plurality of imaging devices that are not shielded by the shielding element from the target three-dimensional object. The information processing apparatus according to any one of claims 1 to 5, characterized in that.

7. A computer, at a location where a subject is imaged, for each of a plurality of three-dimensional objects virtually arranged to fill a predetermined space including the subject, a threshold information generation unit that generates threshold information indicating a threshold for determining whether to remove the object; a discrimination image generation unit that generates a plurality of discrimination images, which are images for discriminating a first region that is a region of the subject and a second region that is a region other than the subject, from each of a plurality of images obtained from a plurality of imaging devices that image the subject at different positions at the location; and a process of counting the number of imaging devices that have imaged an image used to generate a discrimination image in which a position where one three-dimensional object selected from the plurality of three-dimensional objects is projected onto each of the plurality of discrimination images is included in the first region, and comparing the number with the threshold to determine whether to remove the one three-dimensional object from the space, and repeating the process until all of the plurality of three-dimensional objects are selected as the one three-dimensional object, thereby identifying one or more three-dimensional objects to be removed from the space, and estimating the shape of the subject by the remaining one or more three-dimensional objects after removing the one or more three-dimensional objects from the space, the threshold information generation unit reducing the threshold of the target three-dimensional object when there is a shielding element that shields between the target three-dimensional object, which is one of the plurality of three-dimensional objects, and at least one of the plurality of imaging devices, to be smaller than when there is no shielding element. A program characterized by the above.

8. At a location where a subject is imaged, threshold information indicating a threshold for determining whether to remove each of a plurality of three-dimensional objects virtually arranged so as to fill a predetermined space including the subject is generated. A plurality of discrimination images, which are images for distinguishing a first region that is a region of the subject and a second region that is a region other than the subject, are generated from each of a plurality of images obtained from a plurality of imaging devices that image the subject at different positions at the location. The number of imaging devices that have imaged an image used to generate a discrimination image in which a position obtained by projecting one three-dimensional object selected from the plurality of three-dimensional objects onto each of the plurality of discrimination images is included in the first region is counted. By comparing the number with the threshold, a process of determining whether to remove the one three-dimensional object from the space is repeatedly performed until all of the plurality of three-dimensional objects are selected as the one three-dimensional object, thereby identifying one or more three-dimensional objects to be removed from the space. By removing the one or more three-dimensional objects from the space, the remaining one or more three-dimensional objects are used to estimate the shape of the subject. An information processing method, characterized in that when there is a shielding element that shields between a target three-dimensional object, which is one of the plurality of three-dimensional objects, and at least one of the plurality of imaging devices, the threshold of the target three-dimensional object is made smaller than when the shielding element does not exist.

Citation Information

Patent Citations

  • Object Movement System

    JP2022548009A

  • Generation device, generation method and program for three-dimensional model

    WO2019116942A1